Explore how open-source DharmaOCR makes document AI more accessible and actionable.

We are excited to announce the open-sourcing of DharmaOCR on Hugging Face, making our specialized SLM models and datasets publicly available for exploration.

3 min readMachine Learning

We are watching the old wall between powerful document AI and everyday users collapse, and DharmaOCR is holding the wrecking bar. The decision to open-source both models and datasets on Hugging Face, paired with a paper that lays out every experiment, is a direct challenge to the prevailing logic that advanced OCR must be locked behind APIs or expensive enterprise contracts. This is not a marketing gesture. It is a substantive contribution to the public toolkit, and we think it changes the practical calculus for anyone who has felt boxed in by proprietary document processing.

The numbers tell a compelling story, but what matters more is how they were achieved. Dharma-AI took small language models, 3 billion and 7 billion parameters, and refined them with supervised fine-tuning and direct preference optimization. They did not chase scale for its own sake. The result: a 7B model scoring 0.925 against GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6. That is not just competitive; it is a demonstration that specialization beats brute force when the task is well defined. For users, this means you do not need a data center or a massive cloud budget to process complex documents with high accuracy. The model that outperformed the big-name APIs fits on hardware you likely already have.

The innovation that stands out is the negative sampling strategy. By using the model's own degenerate outputs as rejected examples during DPO training, Dharma-AI cut the failure rate by 87.6 percent. That is not a tweak around the edges. It is a direct method for teaching the model what not to do, and it has practical consequences. Fewer failures mean less time manually correcting outputs, less frustration when processing batches of invoices or contracts, and more trust that the system will handle edge cases without crashing or hallucinating. This is the kind of user-facing improvement that technical teams talk about internally but rarely deliver so cleanly.

On the cost side, the AWQ quantization results are worth attention. A 22 percent drop in per-page inference cost with negligible performance loss removes one more reason to stay with proprietary services. For a team processing thousands of pages a month, that savings compounds quickly. Open-source models already eliminate per-seat licensing; now they are closing the operational cost gap as well. The practical message is clear: you can experiment with DharmaOCR today, run it on your own data, evaluate it against your own documents, and decide if the trade-offs work for your workflow. The paper, the models, and the datasets are all public. Go look at the rejection strategy, run your own benchmarks, and see whether a 7B parameter specialist outperforms the generalist giant on the documents you actually handle.

From Machine Learning

Hey everyone, we just open-sourced DharmaOCR on Hugging Face. Models and datasets are all public, free to use and experiment with.

We also published the paper documenting all the experimentation behind it, for those who want to dig into the methodology.

Read the original at Machine Learning