Excited to share IdeaAMBIG!<br>We ask a simple question: when an AI system is given a research idea, can it tell whether the idea is actually specified well enough to implement?<br>We introduce a benchmark of 660 evidence-grounded specification gaps and evaluate 13 LLMs on detecting, localizing, and resolving implementation-critical defects.<br>Our most striking result: LLMs are much better at fixing a gap than finding it. On real-world cases, the best model achieves only 9.6% defect recovery, but 80.6% clarification success once the defect is given. Providing the gold resolution raises downstream codification readiness from 14% to 98%.<br>This points to an important bottleneck for research agents: before asking models to implement an idea, we may first need them to reliably recognize what is missing, ambiguous, or inconsistent.<br>Would love to hear the community’s thoughts on this!</p>\n","updatedAt":"2026-09-11T15:11:02.283Z","author":{"_id":"69227307168e7fa5935fc7a6","avatarUrl":"/avatars/20f959a3314ce098fcd224ae999e2a8f.svg","fullname":"YilingMa","name":"YilingMa","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9084301590919495},"editors":["YilingMa"],"editorAvatarUrls":["/avatars/20f959a3314ce098fcd224ae999e2a8f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.10539","authors":[{"_id":"6aa37d1a47a406da7901e797","name":"Yiling Ma","hidden":false},{"_id":"6aa37d1a47a406da7901e798","name":"Yilun Zhao","hidden":false},{"_id":"6aa37d1a47a406da7901e799","name":"Sihong Wu","hidden":false},{"_id":"6aa37d1a47a406da7901e79a","name":"Manasi Patwardhan","hidden":false},{"_id":"6aa37d1a47a406da7901e79b","name":"Arman Cohan","hidden":false}],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications","submittedOnDailyBy":{"_id":"69227307168e7fa5935fc7a6","avatarUrl":"/avatars/20f959a3314ce098fcd224ae999e2a8f.svg","isPro":false,"fullname":"YilingMa","user":"YilingMa","type":"user","name":"YilingMa"},"summary":"A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.","upvotes":0,"discussionId":"6aa37d1a47a406da7901e79c","githubRepo":"https://github.com/Yiling-Ma/IdeaAMBIG","githubRepoAddedBy":"user","ai_summary":"The study introduces a benchmark to evaluate whether research methods are specified clearly enough for implementation, finding that identifying missing details is the primary challenge for language models.","ai_keywords":["codification readiness","evidence-grounded specifications","IdeaAMBIG","defect localization","clarification action generation","Macro Defect Recovery Rate","Macro Clarification Action Success Rate"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6532df27d690f3012efde84c","name":"yale-nlp","fullname":"Yale NLP Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65204db5b0e0d57453cb1809/9OAeiZ-BrN2g1h1yd6-1W.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"organization":{"_id":"6532df27d690f3012efde84c","name":"yale-nlp","fullname":"Yale NLP Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65204db5b0e0d57453cb1809/9OAeiZ-BrN2g1h1yd6-1W.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.10539.md","query":{}}">
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
Abstract
The study introduces a benchmark to evaluate whether research methods are specified clearly enough for implementation, finding that identifying missing details is the primary challenge for language models.
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.
Community
Excited to share IdeaAMBIG!
We ask a simple question: when an AI system is given a research idea, can it tell whether the idea is actually specified well enough to implement?
We introduce a benchmark of 660 evidence-grounded specification gaps and evaluate 13 LLMs on detecting, localizing, and resolving implementation-critical defects.
Our most striking result: LLMs are much better at fixing a gap than finding it. On real-world cases, the best model achieves only 9.6% defect recovery, but 80.6% clarification success once the defect is given. Providing the gold resolution raises downstream codification readiness from 14% to 98%.
This points to an important bottleneck for research agents: before asking models to implement an idea, we may first need them to reliably recognize what is missing, ambiguous, or inconsistent.
Would love to hear the community’s thoughts on this!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.10539 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.10539 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.10539 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.