[D] How do you get preprocessed dataset of a paper [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Hi all,
I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described.
I've tried all reasonable interpretations of the filtering described in the paper and the closest I can get is still an order of magnitude off for one of the datasets. The paper says "data available on request" — I emailed the authors and followed up once, no reply so far.
For those who've been in this spot:
- Do you just keep the larger-but-valid version you can reproduce and document the mismatch?
- Is it worth sampling to match the reported size or does that just create a different irreproducible dataset?
- When do you escalate to the journal vs just waiting?
How have you successfully gotten preprocessed files from authors? Any etiquette around follow-ups or journal contacts that actually worked?
Thanks for any advice.
[link] [comments]
More from r/MachineLearning
-
Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]
Sep 28
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.