Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning, ACL 2026 Findings</p>\n","updatedAt":"2026-08-14T07:22:16.833Z","author":{"_id":"643407dd4b34368fdb0149e8","avatarUrl":"/avatars/9477b9267d5692a4fe59e30590e9639d.svg","fullname":"Xinyan Guan","name":"xinyan233333","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8141764998435974},"editors":["xinyan233333"],"editorAvatarUrls":["/avatars/9477b9267d5692a4fe59e30590e9639d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.29211","authors":[{"_id":"6a7ec19b42823931a1f177c9","user":{"_id":"643407dd4b34368fdb0149e8","avatarUrl":"/avatars/9477b9267d5692a4fe59e30590e9639d.svg","isPro":false,"fullname":"Xinyan Guan","user":"xinyan233333","type":"user","name":"xinyan233333"},"name":"Xinyan Guan","status":"claimed_verified","statusLastChangedAt":"2026-08-14T08:45:04.805Z","hidden":false},{"_id":"6a7ec19b42823931a1f177ca","name":"Jiali Zeng","hidden":false},{"_id":"6a7ec19b42823931a1f177cb","name":"Chunlei Xin","hidden":false},{"_id":"6a7ec19b42823931a1f177cc","name":"Yaojie Lu","hidden":false},{"_id":"6a7ec19b42823931a1f177cd","name":"Hongyu Lin","hidden":false},{"_id":"6a7ec19b42823931a1f177ce","name":"Xianpei Han","hidden":false},{"_id":"6a7ec19b42823931a1f177cf","name":"Le Sun","hidden":false},{"_id":"6a7ec19b42823931a1f177d0","name":"Fandong Meng","hidden":false}],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning","submittedOnDailyBy":{"_id":"643407dd4b34368fdb0149e8","avatarUrl":"/avatars/9477b9267d5692a4fe59e30590e9639d.svg","isPro":false,"fullname":"Xinyan Guan","user":"xinyan233333","type":"user","name":"xinyan233333"},"summary":"Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this futile reasoning phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce CaRL (Capability-aligned Reinforcement Learning), which aligns model behavior with capability boundaries through reward shaping that incentivizes refusal over futile reasoning and hindsight refusal augmentation that converts failures into refusal supervision. Experiments demonstrate a substantial reduction in futile reasoning while preserving performance across task difficulties, effectively achieving capability-aligned behavior without sacrificing utility. https://github.com/icip-cas/Knowing-When-to-Quit","upvotes":0,"discussionId":"6a7ec19c42823931a1f177d1","githubRepo":"https://github.com/icip-cas/Knowing-When-to-Quit","githubRepoAddedBy":"user","ai_summary":"CaRL uses reinforcement learning with refusal incentives and hindsight augmentation to reduce futile reasoning in large language models while preserving task performance.","ai_keywords":["futile reasoning","capability overreach","miscalibration","specious reasoning","CaRL","capability-aligned reinforcement learning","reward shaping","hindsight refusal augmentation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.29211.md","query":{}}">
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
Abstract
CaRL uses reinforcement learning with refusal incentives and hindsight augmentation to reduce futile reasoning in large language models while preserving task performance.
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this futile reasoning phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce CaRL (Capability-aligned Reinforcement Learning), which aligns model behavior with capability boundaries through reward shaping that incentivizes refusal over futile reasoning and hindsight refusal augmentation that converts failures into refusal supervision. Experiments demonstrate a substantial reduction in futile reasoning while preserving performance across task difficulties, effectively achieving capability-aligned behavior without sacrificing utility. https://github.com/icip-cas/Knowing-When-to-Quit
Community
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning, ACL 2026 Findings
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.29211 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.29211 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.29211 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.