Hugging Face Daily Papers · · 4 min read

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

How do agents fail at autonomous research? We stress-tested them on 100 real frontier-research tasks (800 trajectories, 8 harness–model combos). Every failure traced back to one thing — agents can't self-doubt or self-correct. The bottleneck is metacognition, not capability. We release ARFT, a taxonomy of 45 failure modes.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6343d871e64579a11f168cec/S7iUKEyTXEykPfMR_jniK.jpeg\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6343d871e64579a11f168cec/S7iUKEyTXEykPfMR_jniK.jpeg\" alt=\"general_overview_page-0001 (1)\"></a></p>\n","updatedAt":"2026-08-18T08:52:58.951Z","author":{"_id":"6343d871e64579a11f168cec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3PW9RV30CZ0jl3TOP2-h5.png","fullname":"xinmiao yu","name":"chelseyu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7578538060188293},"editors":["chelseyu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3PW9RV30CZ0jl3TOP2-h5.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.14905","authors":[{"_id":"6a83c001675db694db8cd475","name":"Yanlin Fei","hidden":false},{"_id":"6a83c001675db694db8cd476","name":"Nazhou Liu","hidden":false},{"_id":"6a83c001675db694db8cd477","user":{"_id":"6343d871e64579a11f168cec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3PW9RV30CZ0jl3TOP2-h5.png","isPro":false,"fullname":"xinmiao yu","user":"chelseyu","type":"user","name":"chelseyu"},"name":"Xinmiao Yu","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.107Z","hidden":false},{"_id":"6a83c001675db694db8cd478","name":"Shaolong Chen","hidden":false},{"_id":"6a83c001675db694db8cd479","name":"Lei Li","hidden":false},{"_id":"6a83c001675db694db8cd47a","name":"Rahul Thapa","hidden":false},{"_id":"6a83c001675db694db8cd47b","name":"Madalina Ciobanu","hidden":false},{"_id":"6a83c001675db694db8cd47c","name":"Qingqing Mao","hidden":false},{"_id":"6a83c001675db694db8cd47d","name":"Ritankar Das","hidden":false}],"publishedAt":"2026-08-14T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks","submittedOnDailyBy":{"_id":"6343d871e64579a11f168cec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3PW9RV30CZ0jl3TOP2-h5.png","isPro":false,"fullname":"xinmiao yu","user":"chelseyu","type":"user","name":"chelseyu"},"summary":"AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.","upvotes":17,"discussionId":"6a83c001675db694db8cd47e","projectPage":"https://titanresearchlabs.github.io/AutoResearchEval-site","ai_summary":"Autonomous research agents evaluated across the full scientific lifecycle reveal a pervasive lack of metacognitive self-correction, motivating a new benchmark and failure taxonomy.","ai_keywords":["LLMs","agentic scaffolds","AutoResearch","AutoResearchEval","ARFT","metacognitive loop","agent-as-a-judge"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6a59f0714165d80fbafa0221","name":"PrentisAI","fullname":"Prentis AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a59ef21d47661de0b7b2db8/UcUTJ9pqZsT-j8wQKVVQE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a6c200b101ebc51fc395fea","avatarUrl":"/avatars/0d775aaa3f5451c0950b4355c0f16ee8.svg","isPro":false,"fullname":"yu","user":"xinmiao28","type":"user"},{"_id":"666280fcc36ae4c83f69dd2f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/666280fcc36ae4c83f69dd2f/9HwsQwAA2pOu-9NMbx6Hq.jpeg","isPro":false,"fullname":"Joe Liu Nazhou","user":"JoeLiu996","type":"user"},{"_id":"67e1e0e013360658b946e77e","avatarUrl":"/avatars/b003acc36785d6f546c8354b684b4e67.svg","isPro":false,"fullname":"linda","user":"lindafei001","type":"user"},{"_id":"69708b3f2418e43e8e56bb35","avatarUrl":"/avatars/b8ae1177c95c853bdaecdcf77a5680e1.svg","isPro":false,"fullname":"Li Lei","user":"Bill-MK","type":"user"},{"_id":"65b705e9cc412b887626f459","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b705e9cc412b887626f459/wZENkgjoDiuvgIGnv3mKe.jpeg","isPro":false,"fullname":"Mingkun Lei","user":"Leimingkun","type":"user"},{"_id":"647443bf7d131daf633ac21f","avatarUrl":"/avatars/41944f26aea31962a069122febbb69c1.svg","isPro":false,"fullname":"splendon chen","user":"splendon","type":"user"},{"_id":"6343d871e64579a11f168cec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3PW9RV30CZ0jl3TOP2-h5.png","isPro":false,"fullname":"xinmiao yu","user":"chelseyu","type":"user"},{"_id":"69e760de7acd5e283f5c7bea","avatarUrl":"/avatars/033420e527d76f432f069f5d3a1b5d41.svg","isPro":false,"fullname":"bool","user":"boolzhou12","type":"user"},{"_id":"6a83c257e6d5fedfd8d72122","avatarUrl":"/avatars/d52e363e247253b0bc89e6ce11ceefc3.svg","isPro":false,"fullname":"Chenghua Wang","user":"chwang24","type":"user"},{"_id":"661e44f85b996b856943e711","avatarUrl":"/avatars/e8c54832ddccdd45d3860d5e2561b78e.svg","isPro":false,"fullname":"tjcaspar","user":"tjcaspar","type":"user"},{"_id":"660fc9599760d0856d3df1ae","avatarUrl":"/avatars/90378f325adbd0ccf409b746207eddad.svg","isPro":false,"fullname":"Shaolong Chen","user":"chenshaolong","type":"user"},{"_id":"67d65060883fd5f8a91966b2","avatarUrl":"/avatars/19dde03180121af4030629ab44437bf8.svg","isPro":false,"fullname":"CHENG Man Him","user":"BR1ANCHENG","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a59f0714165d80fbafa0221","name":"PrentisAI","fullname":"Prentis AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a59ef21d47661de0b7b2db8/UcUTJ9pqZsT-j8wQKVVQE.png"},"query":{}}">
Papers
arxiv:2608.14905

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

Published on Aug 14
· Submitted by
xinmiao yu
on Aug 18
Authors:
,

Abstract

Autonomous research agents evaluated across the full scientific lifecycle reveal a pervasive lack of metacognitive self-correction, motivating a new benchmark and failure taxonomy.

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.

Community

Paper author Paper submitter about 1 hour ago

How do agents fail at autonomous research? We stress-tested them on 100 real frontier-research tasks (800 trajectories, 8 harness–model combos). Every failure traced back to one thing — agents can't self-doubt or self-correct. The bottleneck is metacognition, not capability. We release ARFT, a taxonomy of 45 failure modes.
general_overview_page-0001 (1)

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.14905 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.14905 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.14905 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers