Hugging Face Daily Papers · · 4 min read

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

hi!</p>\n","updatedAt":"2026-09-15T03:59:12.101Z","author":{"_id":"637c7503fe115289cfecbe6b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676361945047-637c7503fe115289cfecbe6b.jpeg","fullname":"Wenhao Chai","name":"wchai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":45,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"uz","probability":0.24673785269260406},"editors":["wchai"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1676361945047-637c7503fe115289cfecbe6b.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.15309","authors":[{"_id":"6aa8c2455dd4cb9b4cc028d3","name":"Kaiyuan Liu","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028d4","name":"Qiuyang Mang","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028d5","name":"Bo Peng","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028d6","name":"Wenhao Chai","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028d7","name":"Hanchen Li","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028d8","name":"Shreyas Pimpalgaonkar","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028d9","name":"Luke Zettlemoyer","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028da","name":"Alex Dimakis","hidden":false},{"_id":"6aa8c2455dd4cb9b4cc028db","name":"Alvin Cheung","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis","submittedOnDailyBy":{"_id":"637c7503fe115289cfecbe6b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676361945047-637c7503fe115289cfecbe6b.jpeg","isPro":false,"fullname":"Wenhao Chai","user":"wchai","type":"user","name":"wchai"},"summary":"Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. We apply it to four general-purpose agents on four open-ended benchmarks, with sessions of up to 100M tokens, and to three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference, for which Elo grows linearly with log compute. Against this reference, agents can initially convert tokens into Elo faster than independent sampling, but their marginal gains diminish and eventually fall below the reference. In contrast, the strongest historical human contestants improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, providing evidence of continual learning and substantial headroom after agents slow down. We define the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference. Using this point as the per-session budget, we split 100M tokens across parallel sessions on FrontierCS Polyomino Packing, gaining +264 Elo over one long session and +355 over ten short sessions.","upvotes":6,"discussionId":"6aa8c2455dd4cb9b4cc028dc","ai_summary":"Elo-per-token analysis reveals that LLM agents initially scale faster than independent sampling but eventually slow, while parallel short sessions improve performance over single long runs.","ai_keywords":["Elo-per-token analysis","Bradley-Terry model","open-ended tasks","LLM agents","test-time compute","scaling inflection point","independent sampling","AtCoder Heuristic Contest"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69446b15835f00df604cbc7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/203QH9qkSLhigPnPA2SZW.webp","isPro":false,"fullname":"Qiuyang Mang","user":"qmang","type":"user"},{"_id":"68ed8956c93562a97c50d0a7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Ijc0rJjsnw9azaPzAjOBp.png","isPro":false,"fullname":"Runyuan He","user":"runyuanhe","type":"user"},{"_id":"637c7503fe115289cfecbe6b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676361945047-637c7503fe115289cfecbe6b.jpeg","isPro":false,"fullname":"Wenhao Chai","user":"wchai","type":"user"},{"_id":"642b970ceb31218a5f204a29","avatarUrl":"/avatars/582287f477bbb1a0842787145e375fd3.svg","isPro":false,"fullname":"andy-yang","user":"andy-yang","type":"user"},{"_id":"66ce751a8ec9fda2cf5a9e85","avatarUrl":"/avatars/c17093ca81dad007b3e50bae503955a7.svg","isPro":false,"fullname":"Haocheng Xi","user":"xihc-ucb","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Papers
arxiv:2609.15309

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

Published on Sep 14
· Submitted by
Wenhao Chai
on Sep 15
Authors:
,

Abstract

Elo-per-token analysis reveals that LLM agents initially scale faster than independent sampling but eventually slow, while parallel short sessions improve performance over single long runs.

Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. We apply it to four general-purpose agents on four open-ended benchmarks, with sessions of up to 100M tokens, and to three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference, for which Elo grows linearly with log compute. Against this reference, agents can initially convert tokens into Elo faster than independent sampling, but their marginal gains diminish and eventually fall below the reference. In contrast, the strongest historical human contestants improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, providing evidence of continual learning and substantial headroom after agents slow down. We define the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference. Using this point as the per-session budget, we split 100M tokens across parallel sessions on FrontierCS Polyomino Packing, gaining +264 Elo over one long session and +355 over ten short sessions.

Community

Paper submitter about 4 hours ago
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.15309 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.15309 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.15309 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers