Ten years of manual Photoshop work sounds like the opposite of a dataset. It sounds like the thing you do *before* you get to the actual machine learning. But the team behind Ibteda Digital Library, a private community archive in Pakistan, recognized something else in those 575,729 finished pages: a decade of human decisions, quietly recorded in the geometry of every crop. By registering those finished pages back to their raw photos, they recovered a supervision signal that no amount of manual annotation could have matched. This is a powerful inversion, and it deserves attention. It also forces a reckoning with how we think about scale. The community has been quick to celebrate bigger backbones and higher resolutions, but this project suggests that the missing information is often not in the pixels at all. It is in the operator. That is a humbling thought, and a useful one. It echoes the kind of lesson we have seen elsewhere, like in [The evaluation resolution has been shown to have a significant impact on the identification of the "learning rule" that exhibits the most brain-like characteristics at V1. [R]](/post/the-evaluation-resolution-has-been-shown-to-have-a-significa-cmt76swa30nyzmi9zuweibfoh), where the choice of metric, not the model, determines what you learn. The negative results here are the real story. Scaling from 378 to 572 training books did nothing for held-out performance. Neither did ResNet-50, higher resolution inputs, or a spatial head. The model fit the training data better, but it did not generalize. The error analysis explains why: the failures were near-constant offsets per volume, reflecting a preferred margin inset that simply is not present in the pixels of a new book. In other words, the model learned to imitate the average crop, but not the *preference* behind it. That is a different kind of problem than the one most computer vision pipelines are built to solve. It is not about detecting structure; it is about inferring taste. And the workaround was strikingly simple. Ten operator-corrected crops per book, combined element-wise as a median residual, lifted pass@80 from 0.71 to 0.83 on held-out volumes. Ten labels beat every scaling lever they tried. That is not a small result. It is a direct challenge to the assumption that more data, or more compute, is the default answer. It also connects to the kind of practical, hands-on problem-solving seen in Jigsaw Jeeves: Building a Puzzle Assistant using Computer Vision, where the solution is not a bigger model but a smarter composition of tools. The retouching pipeline is just as instructive. Instead of asking a neural net to generate pixels, they kept it to detection only. A U-Net proposes a support mask; classical OpenCV reconstructs the paper; everything outside the mask is byte-identical to the original. And they used a stricter label set, with REMOVE, KEEP, and IGNORE states, where any erased Urdu diacritic vetoed deployment regardless of IoU. That stricter cut improved mark IoU from 0.56 to 0.60 and drove diacritic false positives to zero. The lesson is not that classical methods are better. It is that the *definition of success* matters more than the model. If you care about zero alteration outside a declared region, then a diffusion model that might hallucinate a new stroke is not an improvement, it is a liability. For archival work, that is the only honest standard. The authors ask whether anyone has modeled document boundaries that depend on an invisible human preference, and whether there is prior work on per-instance residual calibration. The honest answer is that we should all be paying attention to this line of inquiry. As for constrained inpainting, the question is not whether it is possible, but whether you would trust it. We would not, yet. The practical takeaway here is simple: before you scale up, look at your residuals. The information you need might be hiding in the gap between what the model does and what the operator would have done. That is the gap worth closing. And it might not require a bigger backbone, just a better question, like the kind raised in [Bad but typical NeurIPS experience? [D]](/post/bad-but-typical-neurips-experience-d-cmsdje3b601k9mi9zwjnhncka) about whether we are optimizing for the right thing in the first place. Watch for the few-shot inset inference.
book digitization
Ten years of manual crop data unlock automated book digitization
A decade of manual Photoshop work taught one archivist what no model could: crop boundaries are a human preference, not a pixel pattern.
5 min readMachine Learning
Author here. Ibteda Digital Library is a private community archive in Pakistan — for ten years we digitized rare Urdu books (lithographs, dictionaries, periodicals) on a DIY camera rig, finishing every page by hand in Photoshop. When we wound down daily operations, I realized those 575,729 finished pages across 1,765 books recorded a decade of crop decisions, so I registered them back to their raw photos (SIFT + MAGSAC with conservative acceptance gates) and used the recovered geometry as supervision.