The first instinct when an OCR model gets most of the job done but fumbles a few labels is to reach for a bigger hammer. A CRF, a BiLSTM, maybe a GNN if you really want to impress someone. It is a natural response, and it is almost certainly the wrong one for the problem described here. The user has already diagnosed the core issue: the model sees a centered line of text and decides it is body copy because its geometric features do not match the patterns it learned from other documents. That is not a sequence labeling failure. It is a context problem. The deeper issue is that they are conflating two very different tasks. The first is classifying a block as title or text, which is a visual and semantic judgment. The second is reconstructing the hierarchy of those blocks, which is a structural puzzle. A CRF is a reasonable tool for the second task if you have clean labels to work with. But the user has already admitted that the labels are noisy. Feeding noisy labels into a probabilistic graphical model does not fix the noise. It just bakes it into the structure. The related work on noisy text in RAG pipelines shows a similar pattern: when the input layer is unreliable, every downstream step inherits that unreliability. Noisy Text in RAG: Typos, OCR, and the Gap Classical Spell-Check Leaves makes that point well. Garbage in, garbage out, even when the garbage is only occasional. What would we actually tell this person? Start with the numbering. Legal documents of this type are ruthlessly hierarchical. ANNEX, TITLE, A., 1., (a), (b), (c). That pattern is more reliable than any visual feature. The fact that a title is centered or left-aligned is a rendering choice. The numbering is a legal fact. A simple rule that first checks the numbering pattern, then uses geometry to break ties, would catch most of the errors the OCR model makes. That is not overcomplicating. That is matching the tool to the domain. The GNN idea is a distraction. Graph neural networks are good at learning from relational data when you have a clear graph structure. But building that graph from noisy bounding boxes is a project in itself, and the payoff here is unclear. The user wants to process legal documents, not publish a paper on layout analysis. The more interesting question is whether they can build something that generalizes beyond this one document type. The answer is probably not with a CRF or a GNN. Those models will learn the visual patterns of legal PDFs, and then someone will feed them a memo or a contract with a different layout, and the accuracy will collapse. That is the trap. The related work on GNNs and tabular leakage is a warning here: models that look sophisticated often just memorize the surface features of the training data. [Your GNN is probably just an overcomplicated MLP (Tabular Leakage). We built SynthFin-AML to enforce strict causal boundaries. \[P\]](/post/your-gnn-is-probably-just-an-overcomplicated-mlp-tabular-lea-cmthwjk5m0wstmi9zyf9pu7kq) is a useful cautionary tale. The model is not learning a general notion of "title." It is learning that certain pixel positions and font sizes correlate with a label. That works until it does not. The practical move is to stop treating this as a machine learning problem and start treating it as an information extraction problem. Use the OCR for what it is good at, which is turning pixels into text. Then apply a deterministic layer that understands the structure of numbered legal documents. That layer can be a few hundred lines of Python. It will be faster, cheaper, and easier to debug than any sequence model. The user asked if they are overcomplicating it. Yes. The fact that TITLE I was mislabeled as text is not a reason to build a new model. It is a reason to write a rule that says "if the text matches the pattern for a top-level section marker and it is centered, it is a title." The challenge is not the classification. It is deciding when to trust the model and when to trust the structure. The answer is to trust the structure more and the model less.
OCR
Simplify legal document extraction with smarter structural modeling.
A single mislabeled `TITLE I` is a small crack, but in legal documents, that crack runs deep.
5 min readMachine Learning
I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before committing to it.
What I've done so far: I render each PDF page to an image and run it through Baidu's DeepSeek-OCR model. It returns each detected block with a bounding box [x0, y0, x1, y1], a label (title, text, list, table, header, footer, etc.), and the recognized text. The OCR quality itself is genuinely good as the text comes out clean.