[D] How do you get preprocessed dataset of a paper [D]
Our take
The frustration expressed in this recent Reddit post resonates deeply within the AI research community. Reproducibility, a cornerstone of scientific progress, is consistently challenged by the opaque practices surrounding data availability. This particular instance, where a researcher meticulously followed preprocessing steps outlined in a paper only to find significant discrepancies in dataset statistics, highlights a systemic problem. It’s a problem exacerbated by the common, yet often unreliable, “data available on request” clause. The scenario described is far from uncommon; as we discussed in NeurIPS 2026: handling of multiple venue locations seems bad, the logistical and organizational hurdles within large conferences can inadvertently contribute to these data access issues. This underscores a need for more robust and standardized data sharing practices, moving beyond reactive requests and towards proactive, readily accessible repositories.
The proposed solutions – documenting the mismatch and using a smaller, reproducible sample – are pragmatic, albeit imperfect, compromises. Sampling, however, introduces its own set of biases and irreproducibility concerns. The question of when to escalate to the journal is a difficult one, balancing the need for persistence with the potential for strained author relationships. The lack of response from the authors, a sadly familiar occurrence, further complicates matters. It’s a testament to the pressures researchers face, often juggling numerous responsibilities and potentially overlooking data requests amidst those demands. The emergence of powerful foundation models, like the recently released TabPFN-3.5 is released as the next SOTA tabular foundation model emphasizes the importance of reliable and accessible data; these models thrive on high-quality datasets, and their effectiveness is undermined when the underlying data is difficult to verify. The challenges surrounding data reproducibility are not merely academic exercises; they directly impact the credibility and advancement of the field.
This situation calls for a broader cultural shift within the AI research community. Journals could adopt stricter data availability policies, mandating the release of preprocessed datasets alongside accepted papers, or at least providing clear and detailed documentation of preprocessing steps alongside a publicly available source for the raw data. The research community itself could develop standardized tools and protocols for data sharing and verification, simplifying the process for both authors and those seeking to reproduce results. Platforms like those discussed in How much work in progress can a workshop submission be could be adapted to facilitate data sharing and collaboration, fostering a more open and transparent research environment. The current reliance on individual email requests and ad-hoc data sharing is simply unsustainable as the field continues to grow in complexity and scale.
Ultimately, addressing this issue requires a collective commitment to transparency and reproducibility. While acknowledging the inherent challenges in data sharing, the potential benefits – increased trust in research findings, accelerated scientific progress, and the ability to build upon existing work with confidence – far outweigh the costs. The question now is whether the community will proactively embrace these changes, or continue to grapple with the frustrating and time-consuming process of chasing elusive datasets. What mechanisms, beyond journal policy, can incentivize and support authors in prioritizing data accessibility and reproducibility?
Hi all,
I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described.
I've tried all reasonable interpretations of the filtering described in the paper and the closest I can get is still an order of magnitude off for one of the datasets. The paper says "data available on request" — I emailed the authors and followed up once, no reply so far.
For those who've been in this spot:
- Do you just keep the larger-but-valid version you can reproduce and document the mismatch?
- Is it worth sampling to match the reported size or does that just create a different irreproducible dataset?
- When do you escalate to the journal vs just waiting?
How have you successfully gotten preprocessed files from authors? Any etiquette around follow-ups or journal contacts that actually worked?
Thanks for any advice.
[link] [comments]
Read on the original site
Open the publisher's page for the full experience