Hugging Face Daily Papers · · 6 min read

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence.<br>To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts.<br>On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: \\url{<a href=\"https://github.com/zjunlp/AutoSciRub%7D\" rel=\"nofollow\">https://github.com/zjunlp/AutoSciRub}</a>).</p>\n","updatedAt":"2026-09-01T03:53:43.981Z","author":{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","fullname":"Shumin Deng","name":"231sm","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8825733661651611},"editors":["231sm"],"editorAvatarUrls":["/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.31076","authors":[{"_id":"6a964bf9cd6ebc484732ed1e","name":"Xuehai Wang","hidden":false},{"_id":"6a964bf9cd6ebc484732ed1f","name":"Haowei Qin","hidden":false},{"_id":"6a964bf9cd6ebc484732ed20","name":"Tongxin Liu","hidden":false},{"_id":"6a964bf9cd6ebc484732ed21","name":"Junkai Li","hidden":false},{"_id":"6a964bf9cd6ebc484732ed22","name":"Buqiang Xu","hidden":false},{"_id":"6a964bf9cd6ebc484732ed23","name":"Jintian Zhang","hidden":false},{"_id":"6a964bf9cd6ebc484732ed24","name":"Yijun Chen","hidden":false},{"_id":"6a964bf9cd6ebc484732ed25","name":"Zirui Xue","hidden":false},{"_id":"6a964bf9cd6ebc484732ed26","name":"Shumin Deng","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents","submittedOnDailyBy":{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","isPro":false,"fullname":"Shumin Deng","user":"231sm","type":"user","name":"231sm"},"summary":"Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).","upvotes":5,"discussionId":"6a964bfacd6ebc484732ed27","githubRepo":"https://github.com/zjunlp/AutoSciRub","githubRepoAddedBy":"user","ai_summary":"AutoSciRub improves autonomous scientific agents by generating task-specific executable rubrics that guide experiments, verify criteria, and iteratively refine outputs.","ai_keywords":["AutoSciRub","executable rubric","atomic scientific goals","rubric-guided verification","ResearchClawBench","AstaBench"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":14,"organization":{"_id":"620a6fcd8d5e5dfed284bc91","name":"zjunlp","fullname":"ZJUNLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1644851027419-620a61cba53066560e226d30.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","isPro":false,"fullname":"Shumin Deng","user":"231sm","type":"user"},{"_id":"6698c1c3157ceb76c48ff996","avatarUrl":"/avatars/2f1d732c4d9df4f5b554268ee1949dda.svg","isPro":false,"fullname":"徐步强","user":"Xubqpanda","type":"user"},{"_id":"68d8fc00ff474874c83a1c99","avatarUrl":"/avatars/17e3a2f5197274536bf68d949c5416db.svg","isPro":false,"fullname":"huminclu","user":"huminclu","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"620a6fcd8d5e5dfed284bc91","name":"zjunlp","fullname":"ZJUNLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1644851027419-620a61cba53066560e226d30.png"},"query":{}}">
Papers
arxiv:2608.31076

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Published on Aug 31
· Submitted by
Shumin Deng
on Sep 1
Authors:
,

Abstract

AutoSciRub improves autonomous scientific agents by generating task-specific executable rubrics that guide experiments, verify criteria, and iteratively refine outputs.

Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).

Community

Paper submitter about 4 hours ago

Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence.
To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts.
On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: \url{https://github.com/zjunlp/AutoSciRub}).

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.31076 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.31076 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.31076 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers