Hugging Face Daily Papers · · 5 min read

Evo-Bench: Can Language Models Improve Agent Harness?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.</p>\n","updatedAt":"2026-08-11T03:20:54.699Z","author":{"_id":"668fbc76611c65fd76c5ec43","avatarUrl":"/avatars/b8fd35b4ec0188e65deb9adcea9b4f6b.svg","fullname":"Huang Lisheng","name":"hlsheng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8838836550712585},"editors":["hlsheng"],"editorAvatarUrls":["/avatars/b8fd35b4ec0188e65deb9adcea9b4f6b.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09096","authors":[{"_id":"6a7a899e019ce76dc7b3a9ab","user":{"_id":"668fbc76611c65fd76c5ec43","avatarUrl":"/avatars/b8fd35b4ec0188e65deb9adcea9b4f6b.svg","isPro":false,"fullname":"Huang Lisheng","user":"hlsheng","type":"user","name":"hlsheng"},"name":"Lisheng Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.429Z","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9ac","user":{"_id":"66c7e9a596661a3c4ea8e2d0","avatarUrl":"/avatars/af40c11aa79b95497180eb34020751a9.svg","isPro":false,"fullname":"Chen Yang","user":"flust","type":"user","name":"flust"},"name":"Chen Yang","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.422Z","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9ad","name":"Hao Zhou","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9ae","name":"Huatong Song","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9af","name":"Zongchao Chen","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9b0","name":"Ran Le","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9b1","name":"Yang Song","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9b2","name":"Wayne Xin Zhao","hidden":false},{"_id":"6a7a899e019ce76dc7b3a9b3","name":"Tao Zhang","hidden":false}],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Evo-Bench: Can Language Models Improve Agent Harness?","submittedOnDailyBy":{"_id":"668fbc76611c65fd76c5ec43","avatarUrl":"/avatars/b8fd35b4ec0188e65deb9adcea9b4f6b.svg","isPro":false,"fullname":"Huang Lisheng","user":"hlsheng","type":"user","name":"hlsheng"},"summary":"Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.","upvotes":9,"discussionId":"6a7a899e019ce76dc7b3a9b4","projectPage":"https://evobench.org/","githubRepo":"https://github.com/RUCAIBox/Evo-Bench","githubRepoAddedBy":"user","ai_summary":"Evo-Bench evaluates autonomous harness optimization across agent domains using sensitivity-aware task construction and reveals strong but domain-dependent evolution gains.","ai_keywords":["harness evolution","autonomous agents","Evo-Bench","harness-guided construction","auxiliary-task evolution","sensitivity-aware stratified splitting","cross-suite generalization","reasoning structures"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"6704ef33935b1a7c59795566","name":"RUC-AIBOX","fullname":"RUC-AIBOX","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61b8405b516a20acdf3b85ff/Q3_mJHjNqZYfArFl1ZpAL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"668fbc76611c65fd76c5ec43","avatarUrl":"/avatars/b8fd35b4ec0188e65deb9adcea9b4f6b.svg","isPro":false,"fullname":"Huang Lisheng","user":"hlsheng","type":"user"},{"_id":"66c7e9a596661a3c4ea8e2d0","avatarUrl":"/avatars/af40c11aa79b95497180eb34020751a9.svg","isPro":false,"fullname":"Chen Yang","user":"flust","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"65c747f1bbc318a59eceb452","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65c747f1bbc318a59eceb452/W5ERLsLFmwhbt-blcNslJ.jpeg","isPro":false,"fullname":"Shuang Sun","user":"SNHE","type":"user"},{"_id":"64b0a5037a475fba70a7260d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg","isPro":false,"fullname":"Zhen Fang","user":"CostaliyA","type":"user"},{"_id":"63bb1845df5897db7f0431ea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63bb1845df5897db7f0431ea/1v1JP13Gsp7nYpwruijBh.jpeg","isPro":false,"fullname":"Gemini Light","user":"GeminiLight","type":"user"},{"_id":"6757c1bc1866a87cbc3860ed","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6757c1bc1866a87cbc3860ed/XZoIw6Fwj9PWMfwQl_QR7.jpeg","isPro":false,"fullname":"Liang Qiliang","user":"unknowncloudw","type":"user"},{"_id":"651c80a26ba9ab9b9582c273","avatarUrl":"/avatars/e963452eafd21f517d800f2e58e0f918.svg","isPro":false,"fullname":"siyeng feng","user":"siyengfeng","type":"user"},{"_id":"698f3e5bbb8868af0007ee11","avatarUrl":"/avatars/93af16a11c456644bc5dd3a8504a0237.svg","isPro":false,"fullname":"Rauno Ryystö","user":"raunoryys","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6704ef33935b1a7c59795566","name":"RUC-AIBOX","fullname":"RUC-AIBOX","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61b8405b516a20acdf3b85ff/Q3_mJHjNqZYfArFl1ZpAL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09096.md","query":{}}">
Papers
arxiv:2608.09096

Evo-Bench: Can Language Models Improve Agent Harness?

Published on Aug 10
· Submitted by
Huang Lisheng
on Aug 11
Authors:

Abstract

Evo-Bench evaluates autonomous harness optimization across agent domains using sensitivity-aware task construction and reveals strong but domain-dependent evolution gains.

Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.

Community

Paper author Paper submitter about 16 hours ago

Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.09096
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.09096 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.09096 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers