generative AI automation

SGOCR: Teaching AI to See Text Where It Lives

Introducing SGOCR: a pioneering open-source dataset pipeline designed for spatially-grounded Optical Character Recognition (OCR) and Visual Question Answering (VQA) tuples.

3 min readMachine Learning

There is a quiet confidence in what the developer behind SGOCR has built, and it deserves attention. Rather than chasing another model that tries to reason about text in images, they stepped back and asked a simpler question: what if we taught the model where the text actually lives first? That grounding-first instinct is the right one, and it's one that too many dataset builders skip in favor of more glamorous tasks. The result is an open source pipeline that treats spatial grounding as the foundation, not an afterthought, and that distinction matters for anyone training vision-language models on real-world documents.

The practical takeaway here is about efficiency and control. By using a smaller teacher model like Gemini 2.5 Flash for verification, the developer found that high-quality, spatially grounded annotations did the heavy lifting. That is a useful lesson for teams working with limited compute or budget: you do not always need the largest model if your data is precise and your verification loop is tight. The shift from three OCR models and multiple grounding systems down to a leaner stack also reflects a maturity that many projects never reach. It is easy to assume more models mean more accuracy, but this work shows that a well-designed pipeline with clear roles for each component can outperform a cluttered one.

What stands out most is the human-in-the-loop approach. Building a review frontend to store accept, reject, and maybe marks, then bootstrapping those judgments into a quality score, is a practical way to keep human oversight meaningful without slowing automation. That is not just a nice feature; it is a reminder that good datasets are built with intention, not just scraped and labeled at scale. The agentic optimization loop, inspired by Karpathy's work but adapted for broader sweeps, also shows a willingness to question the default playbook. That kind of iterative, observation-driven development is what moves the field forward, even if the changes are incremental.

For anyone working on similar VLMs, the message is clear: start with grounding, keep your pipeline simple, and let human feedback guide your quality metrics. SGOCR is not a flashy demo or a benchmark-topping claim. It is a practical tool that fills a real gap, and that is exactly the kind of work worth exploring. If you have been frustrated by datasets that ask models to reason about text without first teaching them to see it, this is a project worth your time.

From Machine Learning

I've been independently researching & developing small-but-powerful vision-language models (VLMs) and noticed a gap in visual datasets - none were teaching my model to simply ground text in imagery, but trying to get it to reason about the text or about the scene itself. This lead me down a 2 week side-side-project to create SGOCR, an open source dataset pipeline for generating spatially-grounded, OCR-focused VQA tuples with tons of rich metadata to support diverse VLM training strategies.

Read the original at Machine Learning