Scaling OCR to 1M PDFs? This pipeline hits 1200 images per second.

Introducing TurboOCR, an innovative solution designed to tackle the challenges of processing vast amounts of PDFs efficiently.

3 min readMachine Learning

The gap between what OCR models can understand and what they can process is widening, and the pipeline's performance makes that tension impossible to ignore. The author took on nearly a million PDFs and found that PaddleOCR, the strongest non-VLM open source option, managed only 15 images per second on an RTX 5090. That is not a minor performance gap; it is the difference between a batch job that finishes overnight and one that stretches into weeks. The decision to rebuild the pipeline in C++ and CUDA, swap FP32 for FP16 TensorRT, fuse kernels, and pool multi-stream inference is the kind of pragmatic engineering that actually moves the needle for people with real workloads.

What stands out is the reported jump to 1,200 images per second on sparse documents. That is not an incremental improvement; it is two orders of magnitude over the baseline. For anyone running real-time RAG pipelines or bulk processing large collections, that throughput changes what is possible. You stop designing around the OCR bottleneck and start treating document understanding as a fast, cheap layer in your stack. The toggleable layout detection, adding only about 20 percent to inference time, is a thoughtful touch because it lets you pay for structure only when you need it. That kind of control matters when you are processing millions of pages and every millisecond has a cost.

The pipeline is also honest about the trade-offs, and that honesty is refreshing. Complex table extraction and structured output like invoice-to-JSON still require VLM-based approaches, which crawl at 2 images per second. So the tool is not a universal replacement; it is a targeted solution for a specific class of problems. That distinction is important because it tells you when to reach for this pipeline and when to stick with a slower, more capable model. The stated roadmap, bringing structured extraction, markdown output, and table parsing to the faster pipeline while preserving speed, is the right direction, but it is also where the hard work begins.

For teams sitting on large PDF archives, the practical takeaway is straightforward: the cost of OCR is no longer a given. You can process a million pages in under fifteen minutes on a single GPU, assuming the sparse-document average holds, and that opens the door to indexing, searching, and feeding downstream systems in ways that were previously impractical. The author has shown that speed and accuracy are not always opposing forces; sometimes you just need to stop running Python and let the hardware do what it does best. That is a lesson worth applying beyond OCR.

From Machine Learning

I had about 940,000 PDFs to process. Running VLMs over a million pages is slow and expensive, and that gap is only getting worse as OCR moves toward transformer and VLM-based approaches. They’re great for complex understanding, but throughput and cost can become a bottleneck at scale.

PaddleOCR (the non VL version), in my opinion the best non-VLM open source OCR, only handled ~15 img/s on my RTX 5090, which was still too slow. PaddleOCR-VL was crawling at 2 img/s with vLLM.

Read the original at Machine Learning