Every researcher has been where this reader is: staring at a table of statistics that refuses to line up with reality, wondering if the problem is your process or the paper itself. The gap between what a preprint promises and what raw data delivers is one of the quiet frustrations of reproducible science. Our reader tried every reasonable interpretation of the preprocessing steps, and still landed an order of magnitude off for one dataset. That is not a rounding error. That is a signal that something structural is missing, whether in the write-up, the code, or the data release itself.
The instinct to document the mismatch and move forward with a larger-but-valid version is a good one, but it deserves a sharper edge. We would tell this reader to treat the discrepancy as a finding, not a failure. If the paper's stated pipeline cannot be reproduced from public artifacts, that is worth stating plainly in any write-up or code repository. Sampling down to match the reported size might feel like a compromise, but it risks creating a second dataset that no one else can verify. The better path is to keep the version you can defend, publish your preprocessing code, and let the numbers speak for themselves. This connects to a broader theme we have explored in Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where the gap between benchmark conditions and real-world deployment often comes down to exactly these kinds of undocumented data decisions.
On the question of escalation, patience has a limit, but it should be measured in months, not days. The reader followed up once, which is reasonable, but authors are often buried under review requests and teaching loads. A second, polite follow-up that asks for a specific artifact, like the exact preprocessing script or a list of excluded files, is more likely to get a response than a general nudge. If that fails, contacting the journal's editor is not an act of aggression; it is a standard part of the scientific record. Journals have a vested interest in reproducibility, and many will issue a correction or at least a note on the paper's page. That said, we would caution against assuming malice. Sometimes the raw data is simply too messy, and the authors made judgment calls that did not make it into the methods section.
The deeper issue here is one we have touched on in Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning: the difference between a formula that works in theory and a pipeline that works in practice. Just as the Forrester function is a clean benchmark that hides real-world noise, a paper's Table 1 is often a cleaned-up version of a messier reality. The reader's experience is a reminder that reproducibility is not just about sharing code; it is about documenting the decisions that turn raw logs into a usable dataset. And when that documentation is missing, the burden falls on the reader to decide what counts as good enough.
One concrete takeaway we would offer: never let the absence of a response from authors stop you from publishing your own preprocessing notes. If the data is public, your version is a contribution, even if it differs from the paper's. The community benefits more from a transparent mismatch than from silent acceptance of an unreproducible result. As for the journal route, we would say this: escalate when the discrepancy affects the paper's conclusions, not just its statistics. If the order-of-magnitude gap changes the findings, it is worth a formal query. If it only changes the numbers in a table, document it and move on. The field needs more people willing to ask these questions, and fewer papers that treat "data available on request" as a dead end.