Baidu

Transcribing Long Documents with a More Intelligent OCR Approach

A month after Baidu launched Unlimited-OCR, the model's focus on long-document transcription stands out.

4 min readAnalytics Vidhya
Transcribing Long Documents with a More Intelligent OCR Approach

Baidu's new Unlimited-OCR model is a quiet answer to a very loud problem. For anyone who has tried to transcribe a fifty-page research report or a dense legal brief with a vision-language model, the pain is familiar: the system starts strong, then slows, stumbles, or simply forgets the top of the page by the time it reaches the bottom. The culprit is the Key-Value cache, a memory buffer that grows with every token the model processes. As the document lengthens, the cache expands, inference time balloons, and accuracy takes a hit. Baidu's approach tackles this bottleneck head-on, designing a system that keeps long-document transcription both fast and stable. That is not a small feat, and it is worth pausing to consider why.

What makes this interesting is not just the engineering win, but what it signals about the direction of the field. We have spent a lot of time talking about how AI agents learn by editing context rather than model weights, a concept explored in our piece on Agentic Context Learning. The same logic applies here: the model is not being retrained to handle longer documents; it is being given a more efficient way to manage the context it already has. That is a meaningful distinction. It suggests that progress in this space is not only about bigger models or more data, but about smarter memory management. And when you look at how LLMs navigate token space, as we discussed in Exploring Paragraph Structure: How LLMs Navigate Token Space, the idea of a cache that does not spiral out of control becomes even more central. The paragraph is a structural unit that helps the model organize information; the KV cache is what keeps that organization viable at scale.

For the practical user, the takeaway is straightforward: the era of "just feed it more text" is ending. If you have been holding back on adopting AI-powered transcription because the results on long documents were mediocre, Baidu's work is a signal that the bottleneck is being addressed. But do not mistake progress for perfection. Unlimited-OCR is a step forward, not a finish line. The real test will be how it handles edge cases: handwritten notes, poor scans, tables that refuse to align. We would tell a reader to watch for third-party evaluations that stress-test the model on messy, real-world documents, not just clean PDFs. The underlying architecture is promising, but the proof is in the pages people actually need to transcribe.

The deeper lesson here is about the economics of inference. A model that can handle long documents without a runaway cache is not just more convenient; it is more affordable, because it requires less compute per page. That is the kind of improvement that quietly expands who can use the technology. We have written before about the practical side of Unlocking LLM Training: A Practical Guide to Distributed Algorithms, and the same spirit applies here: the best innovations are the ones that lower the barrier to entry. Baidu has done that for long-document OCR. The question now is whether others will follow suit, and how quickly. Watch for the next model to claim similar efficiency, and hold them to the same standard: show us the cache, and show us the speed. That is the detail worth tracking.

From Analytics Vidhya

About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page documents with high accuracy while delivering fast and stable inference. Unlike conventional vision-language OCR systems, Unlimited-OCR addresses a major bottleneck in long-document transcription: the rapidly growing Key-Value (KV) cache, […]

The post How Baidu Unlimited-OCR Works: Solving Long-Document Transcription appeared first on Analytics Vidhya.

Read the original at Analytics Vidhya