Hi everyone! We're excited to introduce our new work, Stable Advantage Fusion (SAF) — a lightweight advantage-fusion framework for jointly training reinforcement learning (RL) with on-policy distillation (OPD). Instead of naively combining GRPO and OPD with a fixed mixing coefficient, SAF separately regulates the magnitude and timing of the teacher signal, fully exploiting dense token-level guidance while preserving room for continued exploration driven by verifiable rewards.<br>🚀 Key Highlights:<br>✦ Diagnosing why 1+1<2: We find that fixed-weight fusion causes two types of mismatch: a small number of OPD tokens can have advantage values far larger than those from GRPO, dominating the update; meanwhile, sustained full-strength distillation pulls the student too close to the teacher, causing entropy collapse and a surge in response length early in training, while limiting further exploration later on.<br>✦ A four-stage stable fusion framework: SAF applies Top-k sparsification → tanh-bounded compression → KL-triggered warm-up → linear annealing solely to the OPD advantage values, jointly resolving both the token-level magnitude mismatch and the training-phase timing mismatch.<br>✦ Lightweight, modular, and easy to integrate: SAF requires no additional models, auxiliary losses, or extra forward passes. All four stages can be toggled independently, making it a plug-and-play replacement for existing GRPO+OPD pipelines.<br>💡 Additional finding: Staying closer to the teacher does not necessarily lead to better final performance. Training dynamics show that fixed fusion achieves the lowest student–teacher KL divergence, yet plateaus in accuracy earlier. By dynamically adjusting the trust placed in the teacher signal, SAF achieves a better balance between teacher guidance and autonomous exploration driven by verifiable rewards.</p>\n","updatedAt":"2026-08-03T11:56:31.764Z","author":{"_id":"6a7033350d14dae90b507561","avatarUrl":"/avatars/4c3aaa1142d3689e2a35588388c5c89b.svg","fullname":"dingyi","name":"dingyii","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8981351256370544},"editors":["dingyii"],"editorAvatarUrls":["/avatars/4c3aaa1142d3689e2a35588388c5c89b.svg"],"reactions":[{"reaction":"👍","users":["YoshuaYL"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.29209","authors":[{"_id":"6a70339dbbe824e6bcc466ee","user":{"_id":"6a7033350d14dae90b507561","avatarUrl":"/avatars/4c3aaa1142d3689e2a35588388c5c89b.svg","isPro":false,"fullname":"dingyi","user":"dingyii","type":"user","name":"dingyii"},"name":"Yifan Ding","status":"claimed_verified","statusLastChangedAt":"2026-08-03T07:57:10.690Z","hidden":false},{"_id":"6a70339dbbe824e6bcc466ef","name":"Xincheng Wei","hidden":false},{"_id":"6a70339dbbe824e6bcc466f0","name":"Yoshua Y. Li","hidden":false},{"_id":"6a70339dbbe824e6bcc466f1","name":"Ziheng Li","hidden":false},{"_id":"6a70339dbbe824e6bcc466f2","name":"Yuquan Lu","hidden":false},{"_id":"6a70339dbbe824e6bcc466f3","name":"Siyu Zhang","hidden":false},{"_id":"6a70339dbbe824e6bcc466f4","name":"Dongsheng Ma","hidden":false},{"_id":"6a70339dbbe824e6bcc466f5","name":"Rongxiang Weng","hidden":false},{"_id":"6a70339dbbe824e6bcc466f6","name":"Xunliang Cai","hidden":false},{"_id":"6a70339dbbe824e6bcc466f7","name":"Yun Chen","hidden":false}],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-03T00:00:00.000Z","title":"SAF-OPD: Stable Advantage Fusion for On-Policy Distillation","submittedOnDailyBy":{"_id":"6a7033350d14dae90b507561","avatarUrl":"/avatars/4c3aaa1142d3689e2a35588388c5c89b.svg","isPro":false,"fullname":"dingyi","user":"dingyii","type":"user","name":"dingyii"},"summary":"Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers entropy collapse from two miscalibrations: a magnitude mismatch, where token-level OPD advantages can spike far beyond the bounded RLVR advantage and erase its signal, and a temporal mismatch, where sustained full-strength OPD keeps pulling the student toward the teacher and limits exploration needed to surpass it. We propose SAF, a Stable Advantage Fusion framework that resolves both issues via a lightweight, four-stage pipeline applied only to the OPD advantage: a sparsify-then-compress mechanism for magnitude control paired with a warm-up-then-anneal mechanism for temporal control, with each stage independently switchable and adding negligible overhead. Instantiating RLVR with GRPO, we evaluate SAF across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B: SAF avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.51-2.70% across all six model-domain settings while achieving more stable training.","upvotes":17,"discussionId":"6a70339dbbe824e6bcc466f8"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a7033350d14dae90b507561","avatarUrl":"/avatars/4c3aaa1142d3689e2a35588388c5c89b.svg","isPro":false,"fullname":"dingyi","user":"dingyii","type":"user"},{"_id":"636a7459eb076ec3f4030e7d","avatarUrl":"/avatars/832dc709211e3a2ea5e93caea3768122.svg","isPro":false,"fullname":"Ziheng Li","user":"ChillingDream","type":"user"},{"_id":"6742e356f8883755e01c6053","avatarUrl":"/avatars/e8eae3b3dc934ac49cb682d8b9b57362.svg","isPro":false,"fullname":"Yoshua Li","user":"YoshuaYL","type":"user"},{"_id":"66e5a80fc86016ecd987b94a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66e5a80fc86016ecd987b94a/I6K2fzW4jufKzMxZ-X_Ql.png","isPro":false,"fullname":"QLY","user":"QLYYLQ","type":"user"},{"_id":"658255f192b5a9664de86250","avatarUrl":"/avatars/bfe5eb8fc55faec28c5f37576b5ccd86.svg","isPro":false,"fullname":"zhuzhuhao","user":"zhuzhuhao","type":"user"},{"_id":"66915f24f4ac69749d45781f","avatarUrl":"/avatars/80f4da9dad1c38583ccf538c988247e8.svg","isPro":false,"fullname":"ads","user":"sxcasf","type":"user"},{"_id":"6353a8b543577a0f542148c6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1666427070682-6353a8b543577a0f542148c6.png","isPro":false,"fullname":"XrazyMee","user":"xTuS","type":"user"},{"_id":"67f503cbf41790a64ecdddbf","avatarUrl":"/avatars/e4c6dc54c264c74f8f07ba4a04836fe8.svg","isPro":false,"fullname":"Hande Huang","user":"Timiduck","type":"user"},{"_id":"679068116d5aed184a2a3423","avatarUrl":"/avatars/b99f9262ab8bf77e840ddb692c837c76.svg","isPro":false,"fullname":"ForgotDream","user":"ForgotDream","type":"user"},{"_id":"6405445e0ab5e22719fbbbec","avatarUrl":"/avatars/42a89a8ea031e92ca22a284d2f305b89.svg","isPro":false,"fullname":"Wong","user":"Ce1este","type":"user"},{"_id":"65095591cc02352e1e106d0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65095591cc02352e1e106d0a/QgNhKBItxxHY927pBVBai.jpeg","isPro":false,"fullname":"Tom Lu","user":"eigentom","type":"user"},{"_id":"6643666b9d2a3c125cdde735","avatarUrl":"/avatars/c9c495ad4789bbfba3c500c90a0f54d3.svg","isPro":false,"fullname":"福生无量摸鱼天尊","user":"TanL","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.29209.md","query":{}}">
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Published on Jul 31
· Submitted by dingyi on Aug 3 Abstract
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers entropy collapse from two miscalibrations: a magnitude mismatch, where token-level OPD advantages can spike far beyond the bounded RLVR advantage and erase its signal, and a temporal mismatch, where sustained full-strength OPD keeps pulling the student toward the teacher and limits exploration needed to surpass it. We propose SAF, a Stable Advantage Fusion framework that resolves both issues via a lightweight, four-stage pipeline applied only to the OPD advantage: a sparsify-then-compress mechanism for magnitude control paired with a warm-up-then-anneal mechanism for temporal control, with each stage independently switchable and adding negligible overhead. Instantiating RLVR with GRPO, we evaluate SAF across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B: SAF avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.51-2.70% across all six model-domain settings while achieving more stable training.
Community
Hi everyone! We're excited to introduce our new work, Stable Advantage Fusion (SAF) — a lightweight advantage-fusion framework for jointly training reinforcement learning (RL) with on-policy distillation (OPD). Instead of naively combining GRPO and OPD with a fixed mixing coefficient, SAF separately regulates the magnitude and timing of the teacher signal, fully exploiting dense token-level guidance while preserving room for continued exploration driven by verifiable rewards.
🚀 Key Highlights:
✦ Diagnosing why 1+1<2: We find that fixed-weight fusion causes two types of mismatch: a small number of OPD tokens can have advantage values far larger than those from GRPO, dominating the update; meanwhile, sustained full-strength distillation pulls the student too close to the teacher, causing entropy collapse and a surge in response length early in training, while limiting further exploration later on.
✦ A four-stage stable fusion framework: SAF applies Top-k sparsification → tanh-bounded compression → KL-triggered warm-up → linear annealing solely to the OPD advantage values, jointly resolving both the token-level magnitude mismatch and the training-phase timing mismatch.
✦ Lightweight, modular, and easy to integrate: SAF requires no additional models, auxiliary losses, or extra forward passes. All four stages can be toggled independently, making it a plug-and-play replacement for existing GRPO+OPD pipelines.
💡 Additional finding: Staying closer to the teacher does not necessarily lead to better final performance. Training dynamics show that fixed fusion achieves the lowest student–teacher KL divergence, yet plateaus in accuracy earlier. By dynamically adjusting the trust placed in the teacher signal, SAF achieves a better balance between teacher guidance and autonomous exploration driven by verifiable rewards.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.29209 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.29209 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.29209 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.