r/MachineLearning · · 1 min read

Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

Since OpenAI retired SWE-bench Verified in February (every frontier model tested could reproduce reference fixes for some tasks; underspecified tests rewarded knowing the intended fix), I've been trying to write down precisely what a decontamination report can and can't establish.

The claim: training-side decontamination has a hard floor. A report is a claim by the party whose score depends on it, over a corpus nobody else can inspect, using matching that misses paraphrase and synthetic derivatives. Corpus commitments and PSI make the lab's claims more precise but not checkable by anyone else, because neither proves the model was trained on the declared corpus and nothing else, and proof-of-training proposals so far have been shown spoofable.

The alternative is to make the evaluation side rule out prior exposure by construction: the submission never receives labels, no network at evaluation, the evaluator builds from a named commit and reproduces the score itself, forward-dated test data where the problem allows it.

I've implemented this for small tabular models and the post separates what's built from what's design. It also lists four things reproduction does not prove: benchmark validity, resistance to adaptive overfitting via repeated submissions (no per-solver budget yet; this is the open gap), funder-side leakage, and third-party re-runnability without the data. The record today is an audit receipt, not a portable proof; the post is explicit about what a complete ZK proof of a result would have to bind (model, inputs, scoring, all to one evaluation) and why proof-of-inference alone isn't it.

https://holdoutlabs-ai.github.io/reproduce-it-or-it-doesnt-count/

Interested in where the argument breaks, especially from anyone who's run a persistent private leaderboard.

submitted by /u/NoahPersaud
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning