The Discovery Certification Protocol turns AI research claims into testable evidence. Validate the gain. Challenge its recovery. Measure the contribution of feedback.</p>\n","updatedAt":"2026-09-10T02:47:09.995Z","author":{"_id":"6310a812a23f0327bce68778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg","fullname":"Ethan Ning","name":"ethanning","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8664777278900146},"editors":["ethanning"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.09219","authors":[{"_id":"6aa219d1a2aeb74440b1de07","name":"Jingjie Ning","hidden":false},{"_id":"6aa219d1a2aeb74440b1de08","name":"Shanshan Zhong","hidden":false},{"_id":"6aa219d1a2aeb74440b1de09","name":"Xiaochuan Li","hidden":false},{"_id":"6aa219d1a2aeb74440b1de0a","name":"Ji Zeng","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6310a812a23f0327bce68778/dhdqSTT2RVtXQqVf0O9cr.png"],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents","submittedOnDailyBy":{"_id":"6310a812a23f0327bce68778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg","isPro":false,"fullname":"Ethan Ning","user":"ethanning","type":"user","name":"ethanning"},"summary":"AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.","upvotes":13,"discussionId":"6aa219d2a2aeb74440b1de0b","projectPage":"https://cxcscmu.github.io/Discovery-Certification-Protocol/","githubRepo":"https://github.com/cxcscmu/Discovery-Certification-Protocol","githubRepoAddedBy":"user","ai_summary":"The Discovery Certification Protocol validates AI research agent outcomes through executable recovery tests, controlled audits, and deterministic verification with finite-sample recovery bounds.","ai_keywords":["Discovery Certification Protocol","recovery and feedback tests","sealed evaluation","Core veto","finite-sample bound","null calibration","SQLite optimization","virtual catalyst control","LLM-free verifier","frozen evidence"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"697f71ec552ef6b0ee9d2703","avatarUrl":"/avatars/f41428695b08898c6918f6934a12f09a.svg","isPro":false,"fullname":"yiyi","user":"racheleven11","type":"user"},{"_id":"654346031767b6a27a0e775d","avatarUrl":"/avatars/8ed561692b65aa6d88bcc0efeb3cf3d4.svg","isPro":false,"fullname":"Xinheng He","user":"Xinheng","type":"user"},{"_id":"64b103cf372d434077206750","avatarUrl":"/avatars/ba0eb4fc712a8b9b93ceb30d11859ec2.svg","isPro":true,"fullname":"Xiaochuan Li","user":"lixiaochuan2020","type":"user"},{"_id":"6a800797ab9839fbc41e04d8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a800797ab9839fbc41e04d8/kYFufvg8I9235g2-59htN.jpeg","isPro":false,"fullname":"伊藤 悠人","user":"Iinouewinnie","type":"user"},{"_id":"6a8513340c5afd86c7d5a008","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8513340c5afd86c7d5a008/0OoFRjN1Q0DrmWzCzDuEU.jpeg","isPro":false,"fullname":"Justin","user":"justinri","type":"user"},{"_id":"6a8739f03f1512e63494dd5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8739f03f1512e63494dd5e/XEM8vni6LQD6ySWePF6Y0.jpeg","isPro":false,"fullname":"Rahul","user":"mumi-yer30","type":"user"},{"_id":"6a861854962577a22faf5f85","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a861854962577a22faf5f85/MrHHBWjsfkZUKOyQo5ha7.jpeg","isPro":false,"fullname":"Kelvin Aquino","user":"kelvinaquinobury","type":"user"},{"_id":"69dee65c5813014ed721a624","avatarUrl":"/avatars/a0d15bac872d8340e4a62b6eb13340ea.svg","isPro":false,"fullname":"XL","user":"lotrrrr","type":"user"},{"_id":"6310a812a23f0327bce68778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg","isPro":false,"fullname":"Ethan Ning","user":"ethanning","type":"user"},{"_id":"6a8d322b07f45a1964934a89","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8d322b07f45a1964934a89/sfL5IDbzCfD-qFTnqKQd6.jpeg","isPro":false,"fullname":"Javier Ruiz","user":"javi-errem","type":"user"},{"_id":"6a81eecc37acba2154e945a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a81eecc37acba2154e945a2/TCMK7joGU8ZFHkjgBH4aV.jpeg","isPro":false,"fullname":"Manon Roux","user":"manonkroux3","type":"user"},{"_id":"64ba19a63269cb9386db003d","avatarUrl":"/avatars/86001ca10b50b0a1f8794521509348f7.svg","isPro":false,"fullname":"Kangrui Mao","user":"karrymao","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.09219.md","query":{}}">
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Abstract
The Discovery Certification Protocol validates AI research agent outcomes through executable recovery tests, controlled audits, and deterministic verification with finite-sample recovery bounds.
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.
Community
The Discovery Certification Protocol turns AI research claims into testable evidence. Validate the gain. Challenge its recovery. Measure the contribution of feedback.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.09219 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.09219 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.09219 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.