Hugging Face Daily Papers · · 4 min read

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The Discovery Certification Protocol turns AI research claims into testable evidence. Validate the gain. Challenge its recovery. Measure the contribution of feedback.</p>\n","updatedAt":"2026-09-10T02:47:09.995Z","author":{"_id":"6310a812a23f0327bce68778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg","fullname":"Ethan Ning","name":"ethanning","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8664777278900146},"editors":["ethanning"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.09219","authors":[{"_id":"6aa219d1a2aeb74440b1de07","name":"Jingjie Ning","hidden":false},{"_id":"6aa219d1a2aeb74440b1de08","name":"Shanshan Zhong","hidden":false},{"_id":"6aa219d1a2aeb74440b1de09","name":"Xiaochuan Li","hidden":false},{"_id":"6aa219d1a2aeb74440b1de0a","name":"Ji Zeng","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6310a812a23f0327bce68778/dhdqSTT2RVtXQqVf0O9cr.png"],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents","submittedOnDailyBy":{"_id":"6310a812a23f0327bce68778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg","isPro":false,"fullname":"Ethan Ning","user":"ethanning","type":"user","name":"ethanning"},"summary":"AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.","upvotes":13,"discussionId":"6aa219d2a2aeb74440b1de0b","projectPage":"https://cxcscmu.github.io/Discovery-Certification-Protocol/","githubRepo":"https://github.com/cxcscmu/Discovery-Certification-Protocol","githubRepoAddedBy":"user","ai_summary":"The Discovery Certification Protocol validates AI research agent outcomes through executable recovery tests, controlled audits, and deterministic verification with finite-sample recovery bounds.","ai_keywords":["Discovery Certification Protocol","recovery and feedback tests","sealed evaluation","Core veto","finite-sample bound","null calibration","SQLite optimization","virtual catalyst control","LLM-free verifier","frozen evidence"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"697f71ec552ef6b0ee9d2703","avatarUrl":"/avatars/f41428695b08898c6918f6934a12f09a.svg","isPro":false,"fullname":"yiyi","user":"racheleven11","type":"user"},{"_id":"654346031767b6a27a0e775d","avatarUrl":"/avatars/8ed561692b65aa6d88bcc0efeb3cf3d4.svg","isPro":false,"fullname":"Xinheng He","user":"Xinheng","type":"user"},{"_id":"64b103cf372d434077206750","avatarUrl":"/avatars/ba0eb4fc712a8b9b93ceb30d11859ec2.svg","isPro":true,"fullname":"Xiaochuan Li","user":"lixiaochuan2020","type":"user"},{"_id":"6a800797ab9839fbc41e04d8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a800797ab9839fbc41e04d8/kYFufvg8I9235g2-59htN.jpeg","isPro":false,"fullname":"伊藤 悠人","user":"Iinouewinnie","type":"user"},{"_id":"6a8513340c5afd86c7d5a008","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8513340c5afd86c7d5a008/0OoFRjN1Q0DrmWzCzDuEU.jpeg","isPro":false,"fullname":"Justin","user":"justinri","type":"user"},{"_id":"6a8739f03f1512e63494dd5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8739f03f1512e63494dd5e/XEM8vni6LQD6ySWePF6Y0.jpeg","isPro":false,"fullname":"Rahul","user":"mumi-yer30","type":"user"},{"_id":"6a861854962577a22faf5f85","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a861854962577a22faf5f85/MrHHBWjsfkZUKOyQo5ha7.jpeg","isPro":false,"fullname":"Kelvin Aquino","user":"kelvinaquinobury","type":"user"},{"_id":"69dee65c5813014ed721a624","avatarUrl":"/avatars/a0d15bac872d8340e4a62b6eb13340ea.svg","isPro":false,"fullname":"XL","user":"lotrrrr","type":"user"},{"_id":"6310a812a23f0327bce68778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310a812a23f0327bce68778/dp9JherZV-RYiSRKnvxsh.jpeg","isPro":false,"fullname":"Ethan Ning","user":"ethanning","type":"user"},{"_id":"6a8d322b07f45a1964934a89","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8d322b07f45a1964934a89/sfL5IDbzCfD-qFTnqKQd6.jpeg","isPro":false,"fullname":"Javier Ruiz","user":"javi-errem","type":"user"},{"_id":"6a81eecc37acba2154e945a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a81eecc37acba2154e945a2/TCMK7joGU8ZFHkjgBH4aV.jpeg","isPro":false,"fullname":"Manon Roux","user":"manonkroux3","type":"user"},{"_id":"64ba19a63269cb9386db003d","avatarUrl":"/avatars/86001ca10b50b0a1f8794521509348f7.svg","isPro":false,"fullname":"Kangrui Mao","user":"karrymao","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.09219.md","query":{}}">
Papers
arxiv:2609.09219

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

Published on Sep 7
· Submitted by
Ethan Ning
on Sep 10
Authors:
,

Abstract

The Discovery Certification Protocol validates AI research agent outcomes through executable recovery tests, controlled audits, and deterministic verification with finite-sample recovery bounds.

AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.

Community

Paper submitter about 5 hours ago

The Discovery Certification Protocol turns AI research claims into testable evidence. Validate the gain. Challenge its recovery. Measure the contribution of feedback.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.09219
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.09219 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.09219 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.09219 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers