The field of contamination mitigation has long been trapped in a dual predicament of \"unmeasurable and incurable\": metrics conceal failures through cancellation, while strategies gamble on pre-hoc estimation. This paper replaces both the yardstick and the remedy.</p>\n","updatedAt":"2026-08-10T07:05:22.553Z","author":{"_id":"6583dadd3a84a40185a3c110","avatarUrl":"/avatars/a5c02048ef77204dd3d454f854356a89.svg","fullname":"hou","name":"onnookk","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9106653928756714},"editors":["onnookk"],"editorAvatarUrls":["/avatars/a5c02048ef77204dd3d454f854356a89.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07341","authors":[{"_id":"6a7975f48e9301703eaa5fb1","name":"Ruijie Hou","hidden":false},{"_id":"6a7975f48e9301703eaa5fb2","name":"Yueyang Jiao","hidden":false},{"_id":"6a7975f48e9301703eaa5fb3","name":"Zhao Wang","hidden":false},{"_id":"6a7975f48e9301703eaa5fb4","name":"Yingming Li","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination","submittedOnDailyBy":{"_id":"6583dadd3a84a40185a3c110","avatarUrl":"/avatars/a5c02048ef77204dd3d454f854356a89.svg","isPro":false,"fullname":"hou","user":"onnookk","type":"user","name":"onnookk"},"summary":"Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the G-AP (Gap of Aggregate Performance), is flawed. Discrete correct/incorrect readouts cannot characterize per-question performance, averaging before differencing lets over- and under-suppression cancel out, and uniform per-question weighting invites strategies to push solve probabilities onto the clean model's high-frequency values. We propose SA-PPG (Stratified Aggregate of Per-question Probability Gaps): estimate each question's solve probability by sampling, difference it against the clean model per question, and aggregate within groups defined by the clean model's solve probability. Existing mitigation strategies first estimate where contamination lies and then operate on the estimate, so they are only as correct as the estimate. RailCap instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, accumulating suppression until the response distribution becomes sufficiently dispersed. Across multiple contaminated models and benchmarks, SA-PPG reveals that prior strategies' restoration is substantially overestimated, while RailCap attains the lowest SA-PPG.","upvotes":2,"discussionId":"6a7975f48e9301703eaa5fb5","organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6583dadd3a84a40185a3c110","avatarUrl":"/avatars/a5c02048ef77204dd3d454f854356a89.svg","isPro":false,"fullname":"hou","user":"onnookk","type":"user"},{"_id":"64c76c6717bd8060455b2f35","avatarUrl":"/avatars/ff3a8e4abf2665fcf6da0b67593a7069.svg","isPro":false,"fullname":"Yueyang Jiao","user":"yyjiao","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07341.md","query":{}}">
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
Published on Aug 7
· Submitted by hou on Aug 10 Abstract
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the G-AP (Gap of Aggregate Performance), is flawed. Discrete correct/incorrect readouts cannot characterize per-question performance, averaging before differencing lets over- and under-suppression cancel out, and uniform per-question weighting invites strategies to push solve probabilities onto the clean model's high-frequency values. We propose SA-PPG (Stratified Aggregate of Per-question Probability Gaps): estimate each question's solve probability by sampling, difference it against the clean model per question, and aggregate within groups defined by the clean model's solve probability. Existing mitigation strategies first estimate where contamination lies and then operate on the estimate, so they are only as correct as the estimate. RailCap instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, accumulating suppression until the response distribution becomes sufficiently dispersed. Across multiple contaminated models and benchmarks, SA-PPG reveals that prior strategies' restoration is substantially overestimated, while RailCap attains the lowest SA-PPG.
Community
The field of contamination mitigation has long been trapped in a dual predicament of "unmeasurable and incurable": metrics conceal failures through cancellation, while strategies gamble on pre-hoc estimation. This paper replaces both the yardstick and the remedy.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.07341 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.07341 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.07341 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.