Hugging Face Daily Papers · May 27, 2026 · 5 min read

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Like Read original ↗

We propose AKBE (Agentic Knowledge Boundary Enhancement), an on-policy method that dynamically probes the model's intrinsic knowledge boundary through dual-path (with-tool and no-tool) rollouts during training. We define the knowledge boundary as the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targeted supervisory signals that guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improves task accuracy by +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% higher tool productivity without any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category.</p>\n","updatedAt":"2026-05-27T02:52:36.566Z","author":{"_id":"6462271493f702673bf99c0b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6462271493f702673bf99c0b/PyWyI2uoJGr0kpugGGr0t.jpeg","fullname":"Dingwei Chen","name":"CuSO4-Chen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8898244500160217},"editors":["CuSO4-Chen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6462271493f702673bf99c0b/PyWyI2uoJGr0kpugGGr0t.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2605.26952","authors":[{"_id":"6a165649e9aa3c8e322db45f","name":"Dingwei Chen","hidden":false},{"_id":"6a165649e9aa3c8e322db460","name":"Zefang Zong","hidden":false},{"_id":"6a165649e9aa3c8e322db461","name":"Zhipeng Ma","hidden":false},{"_id":"6a165649e9aa3c8e322db462","name":"Leo Luo","hidden":false},{"_id":"6a165649e9aa3c8e322db463","name":"Yang Li","hidden":false},{"_id":"6a165649e9aa3c8e322db464","name":"Chengming Li","hidden":false},{"_id":"6a165649e9aa3c8e322db465","name":"Peng Chen","hidden":false},{"_id":"6a165649e9aa3c8e322db466","name":"Jie Jiang","hidden":false}],"publishedAt":"2026-05-26T00:00:00.000Z","submittedOnDailyAt":"2026-05-27T00:00:00.000Z","title":"Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement","submittedOnDailyBy":{"_id":"6462271493f702673bf99c0b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6462271493f702673bf99c0b/PyWyI2uoJGr0kpugGGr0t.jpeg","isPro":false,"fullname":"Dingwei Chen","user":"CuSO4-Chen","type":"user","name":"CuSO4-Chen"},"summary":"Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that agentic RL training induces increasing redundant tool calls and blurs the model's intrinsic knowledge boundary, where the model fails to distinguish when tools are needed versus when parametric knowledge suffices. Existing solutions based on reward shaping create coarse-grained optimization targets that tend to incentivize indiscriminate tool-call suppression, leading to reward hacking. In this paper, we propose AKBE (Agentic Knowledge Boundary Enhancement), an on-policy method that dynamically probes the model's intrinsic knowledge boundary through dual-path (with-tool and no-tool) rollouts during training. We define the knowledge boundary as the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targeted supervisory signals that guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improves task accuracy by +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% higher tool productivity without any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category. Our code is available at https://github.com/CuSO4-Chen/AKBE.","upvotes":8,"discussionId":"6a165649e9aa3c8e322db467","githubRepo":"https://github.com/CuSO4-Chen/AKBE","githubRepoAddedBy":"user","ai_summary":"AKBE enhances LLM agent training by dynamically identifying when tools are needed versus when internal knowledge suffices, improving accuracy and reducing unnecessary tool usage through targeted supervisory signals.","ai_keywords":["agentic reinforcement learning","tool-use capabilities","reward shaping","reward hacking","on-policy method","dual-path rollouts","knowledge boundary","supervisory signals","task accuracy","tool productivity"],"githubStars":4,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6462271493f702673bf99c0b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6462271493f702673bf99c0b/PyWyI2uoJGr0kpugGGr0t.jpeg","isPro":false,"fullname":"Dingwei Chen","user":"CuSO4-Chen","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"65745569839aa08899ea5d27","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/4X8waDwiphbfKZySrYlFy.jpeg","isPro":false,"fullname":"Kailin Jiang","user":"kailinjiang","type":"user"},{"_id":"678652f54760fc44ca4adac2","avatarUrl":"/avatars/d196326c54f52fbcf132fa4b21d304ca.svg","isPro":false,"fullname":"Siyu Zhai","user":"AlmightyFish","type":"user"},{"_id":"637c99bbfe115289cfedfb44","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637c99bbfe115289cfedfb44/p4uSY0TKufJfcHpvEb_ZQ.jpeg","isPro":false,"fullname":"ssz","user":"ssz1111","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"66fb33e2cb9763996580fa24","avatarUrl":"/avatars/e54945cfcd5e425fc616214b9b8d98b5.svg","isPro":true,"fullname":"J","user":"jrcrittenden","type":"user"},{"_id":"644a1dbb9c340e5e1e713153","avatarUrl":"/avatars/21cb93ad067a798a39829ef7e67c70b8.svg","isPro":false,"fullname":"JGC","user":"Nothing2Say","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2605/2605.26952.md"}">

Papers

arxiv:2605.26952

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Published on May 26

· Submitted by

Dingwei Chen on May 27

Tencent

Upvote

Authors:

Abstract

AKBE enhances LLM agent training by dynamically identifying when tools are needed versus when internal knowledge suffices, improving accuracy and reducing unnecessary tool usage through targeted supervisory signals.

AI-generated summary

Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that agentic RL training induces increasing redundant tool calls and blurs the model's intrinsic knowledge boundary, where the model fails to distinguish when tools are needed versus when parametric knowledge suffices. Existing solutions based on reward shaping create coarse-grained optimization targets that tend to incentivize indiscriminate tool-call suppression, leading to reward hacking. In this paper, we propose AKBE (Agentic Knowledge Boundary Enhancement), an on-policy method that dynamically probes the model's intrinsic knowledge boundary through dual-path (with-tool and no-tool) rollouts during training. We define the knowledge boundary as the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targeted supervisory signals that guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improves task accuracy by +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% higher tool productivity without any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category. Our code is available at https://github.com/CuSO4-Chen/AKBE.

View arXiv page View PDF GitHub 4 Add to collection

Community

CuSO4-Chen

Paper submitter about 9 hours ago

•

edited about 8 hours ago

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Get this paper in your agent:

hf papers read 2605.26952

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2605.26952 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2605.26952 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2605.26952 in a Space README.md to link it from this page.

Collections including this paper 1

Discussion (0)

No comments yet. Sign in and be the first to say something.

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Abstract

Community

Models citing this paper 0

Datasets citing this paper 0

Spaces citing this paper 0

Collections including this paper 1

Discussion (0)

More from Hugging Face Daily Papers