Your agent evaluation can be cheaper!</p>\n","updatedAt":"2026-09-03T03:18:04.648Z","author":{"_id":"645b0c3ec35da9c7afd95421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645b0c3ec35da9c7afd95421/vYBrCDagHsXAo6J2p-uG0.jpeg","fullname":"Yuling","name":"YerbaPage","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":110,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4953441917896271},"editors":["YerbaPage"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/645b0c3ec35da9c7afd95421/vYBrCDagHsXAo6J2p-uG0.jpeg"],"reactions":[{"reaction":"👀","users":["YerbaPage","Silin-Chen"],"count":2},{"reaction":"🤗","users":["YerbaPage","Silin-Chen"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.02783","authors":[{"_id":"6a98e6b5fea818274321fedc","user":{"_id":"645b0c3ec35da9c7afd95421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645b0c3ec35da9c7afd95421/vYBrCDagHsXAo6J2p-uG0.jpeg","isPro":false,"fullname":"Yuling","user":"YerbaPage","type":"user","name":"YerbaPage"},"name":"Yuling Shi","status":"claimed_verified","statusLastChangedAt":"2026-09-03T08:15:38.634Z","hidden":false},{"_id":"6a98e6b5fea818274321fedd","user":{"_id":"6295c00307dbe3912697e983","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6295c00307dbe3912697e983/J6Eyes3Xurv0bvjTbjTjT.jpeg","isPro":false,"fullname":"v587su","user":"zhensuuu","type":"user","name":"zhensuuu"},"name":"Zhensu Sun","status":"claimed_verified","statusLastChangedAt":"2026-09-03T08:15:36.780Z","hidden":false},{"_id":"6a98e6b5fea818274321fede","name":"Junsen Dong","hidden":false},{"_id":"6a98e6b5fea818274321fedf","name":"Chengcheng Wan","hidden":false},{"_id":"6a98e6b5fea818274321fee0","name":"David Lo","hidden":false},{"_id":"6a98e6b5fea818274321fee1","name":"Xiaodong Gu","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction","submittedOnDailyBy":{"_id":"645b0c3ec35da9c7afd95421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645b0c3ec35da9c7afd95421/vYBrCDagHsXAo6J2p-uG0.jpeg","isPro":false,"fullname":"Yuling","user":"YerbaPage","type":"user","name":"YerbaPage"},"summary":"Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. Prior efforts, centered on benchmark distillation, reduce the number of evaluation tasks but leave the cost of executing each retained task untouched. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task. Our key insight is that an agent's final outcome is often evident from its intermediate behavior well before execution completes. We instantiate this idea in EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features, and halts an agent run the moment either classifier crosses a calibrated confidence threshold, adding negligible per-step overhead. Across three benchmarks, SWE-bench Verified, TerminalBench, and Toolathlon, EarlyEval can eliminate 13%-26% of agent steps and up to 44.1% input tokens and 29.4% output tokens at 89%-97% prediction accuracy, while perturbing per-agent resolve rates by only one to two percentage points on average.","upvotes":76,"discussionId":"6a98e6b5fea818274321fee2","githubRepo":"https://github.com/inphotoo/earlyeval","githubRepoAddedBy":"user","ai_summary":"EarlyEval predicts agent outcomes from intermediate behavior to reduce evaluation cost by halting runs early with minimal accuracy loss.","ai_keywords":["LLM agents","early outcome prediction","EarlyEval","LightGBM classifiers","behavioral features","SWE-bench Verified","TerminalBench","Toolathlon"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"645b0c3ec35da9c7afd95421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645b0c3ec35da9c7afd95421/vYBrCDagHsXAo6J2p-uG0.jpeg","isPro":false,"fullname":"Yuling","user":"YerbaPage","type":"user"},{"_id":"6295c00307dbe3912697e983","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6295c00307dbe3912697e983/J6Eyes3Xurv0bvjTbjTjT.jpeg","isPro":false,"fullname":"v587su","user":"zhensuuu","type":"user"},{"_id":"65684c80a9a1a6a50d779f58","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65684c80a9a1a6a50d779f58/it534ZdH5LxRub1M_o3uM.jpeg","isPro":false,"fullname":"Silin Chen","user":"Silin-Chen","type":"user"},{"_id":"65f40d605d44012d77f6251c","avatarUrl":"/avatars/aee2b1c29899111419534427e17aa6f6.svg","isPro":false,"fullname":"Yunbo Lyu","user":"Ra1nbow99","type":"user"},{"_id":"66b9fe571caaa1d77c44e79e","avatarUrl":"/avatars/19dc7c2b8ad9ad3de65fa2b8271be2eb.svg","isPro":false,"fullname":"Bohou Zhang","user":"yiwenz","type":"user"},{"_id":"6408823b92033c15073b59d5","avatarUrl":"/avatars/24633b955436e638a18a186749f42530.svg","isPro":true,"fullname":"iCSawyer","user":"iCSawyer","type":"user"},{"_id":"67062b238587c775e4f14191","avatarUrl":"/avatars/e6cc278cff345048b195e62df7b99647.svg","isPro":false,"fullname":"Lihan","user":"SJTU-furrina","type":"user"},{"_id":"6787e222d156a4ec4158641f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6787e222d156a4ec4158641f/Dxf5d3dllkWJPriO_Sy_Q.jpeg","isPro":false,"fullname":"TangLi","user":"A-SAD","type":"user"},{"_id":"69d87ec6112eaf0b8aa17174","avatarUrl":"/avatars/1a8c456261f6a02847607a1b7409c437.svg","isPro":false,"fullname":"xiaose","user":"mactavishe","type":"user"},{"_id":"69b7745c11a0bbbbef8fd4e0","avatarUrl":"/avatars/d8dbe27a8099a8cee87fe67e80abfe62.svg","isPro":false,"fullname":"LUZQ","user":"ZhuqunL","type":"user"},{"_id":"69414310774f687eba4af1ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69414310774f687eba4af1ff/gH8iTWI24TyFW9GEsXOIK.jpeg","isPro":false,"fullname":"Andrew C","user":"SkkuAC","type":"user"},{"_id":"69d87ca70194463f23e1e994","avatarUrl":"/avatars/8767348ed3e5168798b4849db26ed152.svg","isPro":false,"fullname":"Shu Zhang","user":"Orgicat","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.02783.md","query":{}}">
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Abstract
EarlyEval predicts agent outcomes from intermediate behavior to reduce evaluation cost by halting runs early with minimal accuracy loss.
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. Prior efforts, centered on benchmark distillation, reduce the number of evaluation tasks but leave the cost of executing each retained task untouched. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task. Our key insight is that an agent's final outcome is often evident from its intermediate behavior well before execution completes. We instantiate this idea in EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features, and halts an agent run the moment either classifier crosses a calibrated confidence threshold, adding negligible per-step overhead. Across three benchmarks, SWE-bench Verified, TerminalBench, and Toolathlon, EarlyEval can eliminate 13%-26% of agent steps and up to 44.1% input tokens and 29.4% output tokens at 89%-97% prediction accuracy, while perturbing per-agent resolve rates by only one to two percentage points on average.
Community
Your agent evaluation can be cheaper!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.02783 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.02783 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.02783 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.