Hugging Face Daily Papers · June 26, 2026 · 5 min read

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Like Read original ↗

Our Code is available at <a href=\"https://github.com/hypasd-art/Tool-RL-Box\" rel=\"nofollow\">https://github.com/hypasd-art/Tool-RL-Box</a>.</p>\n","updatedAt":"2026-06-26T03:29:41.603Z","author":{"_id":"643379416c6ecd58798421b3","avatarUrl":"/avatars/831db7eab2663abc33b176cf386b02f2.svg","fullname":"Zhuoran Jin","name":"jinzhuoran","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7592245936393738},"editors":["jinzhuoran"],"editorAvatarUrls":["/avatars/831db7eab2663abc33b176cf386b02f2.svg"],"reactions":[],"isReport":false}},{"id":"6a3fbb4d2a7f2e5aa7749fb4","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-06-27T12:00:13.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"The analysis of catastrophic collapse in multi-step tool-use RL is a critical find. We've all seen models suddenly 'forget' how to call a tool despite having the capability in the base weights; seeing this attributed to probability spikes in control tokens rather than a loss of logic is a huge distinction. Using supervisory signals to stabilize the structured execution makes a lot of sense for anyone building production agentic systems. It shifts the problem from 'teaching the model to reason' to 'maintaining the integrity of the output format' during RL. This is the kind of engineering-grounded insight that actually helps in deploying reliable agents.","html":"<p>The analysis of catastrophic collapse in multi-step tool-use RL is a critical find. We've all seen models suddenly 'forget' how to call a tool despite having the capability in the base weights; seeing this attributed to probability spikes in control tokens rather than a loss of logic is a huge distinction. Using supervisory signals to stabilize the structured execution makes a lot of sense for anyone building production agentic systems. It shifts the problem from 'teaching the model to reason' to 'maintaining the integrity of the output format' during RL. This is the kind of engineering-grounded insight that actually helps in deploying reliable agents.</p>\n","updatedAt":"2026-06-27T12:00:13.992Z","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9476225972175598},"editors":["O96a"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.26027","authors":[{"_id":"6a3df1623b43e283349ec1c8","name":"Yupu Hao","hidden":false},{"_id":"6a3df1623b43e283349ec1c9","name":"Zhuoran Jin","hidden":false},{"_id":"6a3df1623b43e283349ec1ca","name":"Huanxuan Liao","hidden":false},{"_id":"6a3df1623b43e283349ec1cb","name":"Kang Liu","hidden":false},{"_id":"6a3df1623b43e283349ec1cc","name":"Jun Zhao","hidden":false}],"publishedAt":"2026-06-24T00:00:00.000Z","submittedOnDailyAt":"2026-06-26T00:00:00.000Z","title":"Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It","submittedOnDailyBy":{"_id":"643379416c6ecd58798421b3","avatarUrl":"/avatars/831db7eab2663abc33b176cf386b02f2.svg","isPro":false,"fullname":"Zhuoran Jin","user":"jinzhuoran","type":"user","name":"jinzhuoran"},"summary":"Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often leads to instability or limited gains in tool-use tasks. In our experiments, some models exhibit catastrophic collapse, where performance abruptly drops and tool-invocation structures fail. The analysis reveals that these failures stem from unexpected probability spikes in specific control tokens, disrupting structured execution, yet the underlying tool-use capability remains intact, merely obscured by specific formats. To address this, we systematically investigate a diverse set of supervisory signals, including off-policy supervision, hint-based guidance, erroneous example supervision, and others, applied under both synchronous and interleaved training schemes. We find that interleaving supervised fine-tuning (SFT) with RL substantially improves stability, but exhibits degraded performance under format and content out-of-distribution (OOD) evaluation. We also analyze the impact of learning rates and generalization across settings. These results highlight the importance of understanding RL failures and demonstrate how diverse supervisory signals can guide exploratory learning, enabling robust training of LLMs for complex, multi-step tool-use tasks. Our Code is available at https://github.com/hypasd-art/Tool-RL-Box.","upvotes":15,"discussionId":"6a3df1623b43e283349ec1cd","githubRepo":"https://github.com/hypasd-art/Tool-RL-Box","githubRepoAddedBy":"user","ai_summary":"Research investigates how different supervisory signals and training strategies improve the stability and performance of large language models in tool-use tasks, addressing issues like catastrophic collapse and format sensitivity through interleaved supervised fine-tuning and reinforcement learning.","ai_keywords":["agentic reinforcement learning","tool-use tasks","catastrophic collapse","control tokens","supervised fine-tuning","off-policy supervision","hint-based guidance","erroneous example supervision","interleaved training","synchronous training","learning rates","out-of-distribution evaluation","exploratory learning"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"640a887796aae649741a586f","name":"CASIA","fullname":"Chinese Academic of Science Institute of Automation","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1678411888885-6388984e8a5dbe2f3dc5afee.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"643379416c6ecd58798421b3","avatarUrl":"/avatars/831db7eab2663abc33b176cf386b02f2.svg","isPro":false,"fullname":"Zhuoran Jin","user":"jinzhuoran","type":"user"},{"_id":"64b89dfa6a68a9a715df407e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b89dfa6a68a9a715df407e/FpBAdClhr-oVAv11Bjwjs.jpeg","isPro":false,"fullname":"Jiachun Li","user":"Septzzz","type":"user"},{"_id":"6a30bc0f7246336616384101","avatarUrl":"/avatars/8a15598f1f35a2a16589e121902e0a5f.svg","isPro":false,"fullname":"wlings","user":"wlings","type":"user"},{"_id":"67d15f29bacfd19231a6e178","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/I0CxXy7s-Huy495htZlJC.png","isPro":false,"fullname":"Lu Wang","user":"wanglu666","type":"user"},{"_id":"67c6744ffaf82ef97dbe871d","avatarUrl":"/avatars/fe9f2c8bb3c98af3177a1170a6829913.svg","isPro":false,"fullname":"Longxiang Wang","user":"wlxxx","type":"user"},{"_id":"6307612bfd79b417f1bc3fa3","avatarUrl":"/avatars/e86ed202106c43d5ba65bc3ff1f0c1fd.svg","isPro":false,"fullname":"ricky_33","user":"ricky333","type":"user"},{"_id":"6538d3089d66a6c304ee2f4b","avatarUrl":"/avatars/6e93ea6c937128765e5323fb4a5d1cea.svg","isPro":false,"fullname":"HongbangYuan","user":"HongbangYuan","type":"user"},{"_id":"6555ca4b5891609e4558dd2a","avatarUrl":"/avatars/30395217fc2e6dbed8e261e1fa4883a6.svg","isPro":true,"fullname":"Tianyi Men","user":"MultimodalAgent","type":"user"},{"_id":"650fa426877b574970bc1f0c","avatarUrl":"/avatars/55945605b53f92f715f447aab0b1ee95.svg","isPro":false,"fullname":"zkj","user":"JokerJan","type":"user"},{"_id":"646def60df618b303b419323","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646def60df618b303b419323/JLJGYen4-5M8ivsLsSk0w.jpeg","isPro":false,"fullname":"Lei Wang","user":"demolei","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"69f0bb9a53592156859aab90","avatarUrl":"/avatars/122aeb140c584b7842c50ae693c2a27e.svg","isPro":false,"fullname":"mini09999","user":"mini09999","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"640a887796aae649741a586f","name":"CASIA","fullname":"Chinese Academic of Science Institute of Automation","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1678411888885-6388984e8a5dbe2f3dc5afee.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.26027.md","query":{}}">

Papers

arxiv:2606.26027

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

Published on Jun 24

· Submitted by

Zhuoran Jin on Jun 26

Chinese Academic of Science Institute of Automation

Upvote

Authors:

Abstract

Research investigates how different supervisory signals and training strategies improve the stability and performance of large language models in tool-use tasks, addressing issues like catastrophic collapse and format sensitivity through interleaved supervised fine-tuning and reinforcement learning.

Generated by Qwen/Qwen2.5-Coder-32B-Instruct

Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often leads to instability or limited gains in tool-use tasks. In our experiments, some models exhibit catastrophic collapse, where performance abruptly drops and tool-invocation structures fail. The analysis reveals that these failures stem from unexpected probability spikes in specific control tokens, disrupting structured execution, yet the underlying tool-use capability remains intact, merely obscured by specific formats. To address this, we systematically investigate a diverse set of supervisory signals, including off-policy supervision, hint-based guidance, erroneous example supervision, and others, applied under both synchronous and interleaved training schemes. We find that interleaving supervised fine-tuning (SFT) with RL substantially improves stability, but exhibits degraded performance under format and content out-of-distribution (OOD) evaluation. We also analyze the impact of learning rates and generalization across settings. These results highlight the importance of understanding RL failures and demonstrate how diverse supervisory signals can guide exploratory learning, enabling robust training of LLMs for complex, multi-step tool-use tasks. Our Code is available at https://github.com/hypasd-art/Tool-RL-Box.

View arXiv page View PDF GitHub 2 Add to collection

Community

jinzhuoran

Paper submitter 2 days ago

Our Code is available at https://github.com/hypasd-art/Tool-RL-Box.

O96a

about 21 hours ago

The analysis of catastrophic collapse in multi-step tool-use RL is a critical find. We've all seen models suddenly 'forget' how to call a tool despite having the capability in the base weights; seeing this attributed to probability spikes in control tokens rather than a loss of logic is a huge distinction. Using supervisory signals to stabilize the structured execution makes a lot of sense for anyone building production agentic systems. It shifts the problem from 'teaching the model to reason' to 'maintaining the integrity of the output format' during RL. This is the kind of engineering-grounded insight that actually helps in deploying reliable agents.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Get this paper in your agent:

hf papers read 2606.26027

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.26027 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.26027 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.26027 in a Space README.md to link it from this page.

Collections including this paper 1

Discussion (0)

No comments yet. Sign in and be the first to say something.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

Abstract

Community

Models citing this paper 0

Datasets citing this paper 0

Spaces citing this paper 0

Collections including this paper 1

Discussion (0)

More from Hugging Face Daily Papers