We propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most.</p>\n","updatedAt":"2026-08-06T03:04:36.922Z","author":{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","fullname":"Zhengxi Lu","name":"LZXzju","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8703550100326538},"editors":["LZXzju"],"editorAvatarUrls":["/avatars/dfd802a24bd63e509728159ebb1769f6.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.00782","authors":[{"_id":"6a73f997c5e410d076869ab6","name":"Zhuowen Han","hidden":false},{"_id":"6a73f997c5e410d076869ab7","name":"Jinwei Xiao","hidden":false},{"_id":"6a73f997c5e410d076869ab8","user":{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","isPro":false,"fullname":"Zhengxi Lu","user":"LZXzju","type":"user","name":"LZXzju"},"name":"Zhengxi Lu","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.521Z","hidden":false},{"_id":"6a73f997c5e410d076869ab9","name":"Renren Jin","hidden":false},{"_id":"6a73f997c5e410d076869aba","name":"Zhiyuan Yao","hidden":false},{"_id":"6a73f997c5e410d076869abb","name":"Yuxin Liu","hidden":false},{"_id":"6a73f997c5e410d076869abc","name":"Hongyan Hao","hidden":false},{"_id":"6a73f997c5e410d076869abd","name":"Yueqing Sun","hidden":false},{"_id":"6a73f997c5e410d076869abe","name":"Yu Yang","hidden":false},{"_id":"6a73f997c5e410d076869abf","name":"Qi GU","hidden":false},{"_id":"6a73f997c5e410d076869ac0","name":"Xunliang Cai","hidden":false},{"_id":"6a73f997c5e410d076869ac1","name":"Deyi Xiong","hidden":false}],"publishedAt":"2026-08-01T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance","submittedOnDailyBy":{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","isPro":false,"fullname":"Zhengxi Lu","user":"LZXzju","type":"user","name":"LZXzju"},"summary":"Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals and loses gradients entirely when all responses within a group receive identical rewards. On-policy distillation (OPD) offers a natural remedy by providing dense, token-level supervision from a teacher model. However, naively combining GRPO with OPD leads to degraded performance, due to three underlying causes: not all samples benefit from distillation; fitting too quickly to the teacher undermines the exploratory capacity of RL; and OPD's advantages are asymmetric, suppressing most tokens. To address these challenges, we propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most. At the sample level, OPD is restricted to negative zero-variance prompts with each sample weighted by the teacher's confidence score. At the token level, distillation targets only tokens with high student entropy or large teacher-student divergence. We further augment training with SFT on correct trajectories generated by the teacher model, injecting positive gradient signals where RL yields none. Experiments demonstrate that RSTG substantially outperforms naive GRPO+OPD by +4.02% on math and +3.05% on code.","upvotes":9,"discussionId":"6a73f997c5e410d076869ac2"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","isPro":false,"fullname":"Zhengxi Lu","user":"LZXzju","type":"user"},{"_id":"61185170a7169fc81917af72","avatarUrl":"/avatars/5a7f22e08f1bd104684cfb2c45414efa.svg","isPro":false,"fullname":"MengWang","user":"poker125","type":"user"},{"_id":"697c8b15a7f796854ef333c4","avatarUrl":"/avatars/94de3a736fac914944f1b57609e3819a.svg","isPro":false,"fullname":"Joel Wang","user":"joelhenwang","type":"user"},{"_id":"6a6aa002977fbfce4bad937b","avatarUrl":"/avatars/04ced1974d585c0bd7507cda07ae61c7.svg","isPro":false,"fullname":"Charles White","user":"zenithLens","type":"user"},{"_id":"6a6c8532da65172f47ef3f9d","avatarUrl":"/avatars/4c207375dd1c09ba60f2cc666d6b01eb.svg","isPro":false,"fullname":"Brian Williams","user":"EmberGlade","type":"user"},{"_id":"64ca39391867f2d13736f040","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ca39391867f2d13736f040/imMCce0U5jTybxKGw19fR.jpeg","isPro":false,"fullname":"Yu Wang","user":"Wloner0809","type":"user"},{"_id":"6a6d3d4d1480b2a9801f8880","avatarUrl":"/avatars/28660f6b9f514f3974c4e2d2d430d86c.svg","isPro":false,"fullname":"Edward Miller","user":"sableRidgeX","type":"user"},{"_id":"6a6dcb13097052be15667902","avatarUrl":"/avatars/be85cdc648b1d02daaa28bafba165ba9.svg","isPro":false,"fullname":"Joseph Williams","user":"robert-2844509","type":"user"},{"_id":"6a1fd9d481eee8267ebfe1f6","avatarUrl":"/avatars/4703ee72cb11285299f7fbc7128524a9.svg","isPro":false,"fullname":"Zhuowen Han","user":"Zhuowen02","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.00782.md","query":{}}">
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals and loses gradients entirely when all responses within a group receive identical rewards. On-policy distillation (OPD) offers a natural remedy by providing dense, token-level supervision from a teacher model. However, naively combining GRPO with OPD leads to degraded performance, due to three underlying causes: not all samples benefit from distillation; fitting too quickly to the teacher undermines the exploratory capacity of RL; and OPD's advantages are asymmetric, suppressing most tokens. To address these challenges, we propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most. At the sample level, OPD is restricted to negative zero-variance prompts with each sample weighted by the teacher's confidence score. At the token level, distillation targets only tokens with high student entropy or large teacher-student divergence. We further augment training with SFT on correct trajectories generated by the teacher model, injecting positive gradient signals where RL yields none. Experiments demonstrate that RSTG substantially outperforms naive GRPO+OPD by +4.02% on math and +3.05% on code.
Community
We propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.00782 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.00782 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.00782 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.