Hugging Face Daily Papers · · 5 min read

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We introduce β-OPSD, a principled generalization of on-policy self-distillation for reasoning language models.<br>🔍 Our key observation is that vanilla OPSD is exactly the β = 1 case of a broader KL-regularized policy-optimization objective. Its optimal policy geometrically interpolates between a reference policy and a privileged teacher.<br>⚡ Instead of running costly, high-variance reinforcement learning, β-OPSD converts this closed-form solution into an efficient token-level distillation target. We also use return-to-go credit assignment to better align token updates with sequence-level outcomes.<br>📈 Across mathematical reasoning benchmarks and Qwen3 models from 1.7B to 8B, β-OPSD improves average performance and training stability over vanilla OPSD—including a +5.74-point average gain at the 1.7B scale.<br>🌐 Project page: <a href=\"https://umd-huang-lab.github.io/beta-opsd/\" rel=\"nofollow\">https://umd-huang-lab.github.io/beta-opsd/</a></p>\n","updatedAt":"2026-07-31T14:18:22.495Z","author":{"_id":"668f330016a6ba9e78ec66ac","avatarUrl":"/avatars/cac0181818a5d91a866d030758bc67e5.svg","fullname":"Minghui Liu","name":"minghuiliu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8766970634460449},"editors":["minghuiliu"],"editorAvatarUrls":["/avatars/cac0181818a5d91a866d030758bc67e5.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28582","authors":[{"_id":"6a6cadc726c3ed41815c87fd","name":"Jiawei Xu","hidden":false},{"_id":"6a6cadc726c3ed41815c87fe","user":{"_id":"668f330016a6ba9e78ec66ac","avatarUrl":"/avatars/cac0181818a5d91a866d030758bc67e5.svg","isPro":false,"fullname":"Minghui Liu","user":"minghuiliu","type":"user","name":"minghuiliu"},"name":"Minghui Liu","status":"claimed_verified","statusLastChangedAt":"2026-07-31T16:45:27.497Z","hidden":false},{"_id":"6a6cadc726c3ed41815c87ff","name":"Juzheng Zhang","hidden":false},{"_id":"6a6cadc726c3ed41815c8800","name":"Tom Goldstein","hidden":false},{"_id":"6a6cadc726c3ed41815c8801","name":"Furong Huang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/668f330016a6ba9e78ec66ac/odY26JA-7xKwoOUiN4eg6.png"],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation","submittedOnDailyBy":{"_id":"668f330016a6ba9e78ec66ac","avatarUrl":"/avatars/cac0181818a5d91a866d030758bc67e5.svg","isPro":false,"fullname":"Minghui Liu","user":"minghuiliu","type":"user","name":"minghuiliu"},"summary":"On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the β=1 member of a broader policy-optimization family, where β weights the KL penalty anchoring the student to a reference policy. This equivalence turns β from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity to a reference policy against privileged teacher guidance. We introduce β-OPSD and derive its optimal policy as a geometric interpolation between the reference policy and the privileged teacher. Directly optimizing this objective with reinforcement learning, however, would be costly and high-variance. Rather than optimize the RL objective directly, we turn its closed-form solution into a distillation target. Each value of β selects a target along the reference-to-teacher path, which we implement efficiently by mixing their token-level logits. In this way, inexpensive distillation approximates the solution of expensive policy optimization. Return-to-go credit assignment further aligns token updates with the sequence-level objective while retaining the simplicity of OPSD. Experiments on mathematical reasoning benchmarks show that β-OPSD consistently outperforms vanilla OPSD, improving optimization stability and downstream reasoning performance. Our results provide a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.","upvotes":13,"discussionId":"6a6cadc826c3ed41815c8802","projectPage":"https://umd-huang-lab.github.io/beta-opsd/","organization":{"_id":"64cbc5468174e45ae060ec46","name":"furonghuang-lab","fullname":"Furong Huang's Lab at UMD","avatar":"https://www.gravatar.com/avatar/add71ee6bbcef2277b077b42b3cba002?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"668f330016a6ba9e78ec66ac","avatarUrl":"/avatars/cac0181818a5d91a866d030758bc67e5.svg","isPro":false,"fullname":"Minghui Liu","user":"minghuiliu","type":"user"},{"_id":"655524c9710bb1dad208d680","avatarUrl":"/avatars/8eab009a3b6bfeb11d828d09f36d5f3a.svg","isPro":false,"fullname":"Jiawei Xu","user":"JimmyXUJW","type":"user"},{"_id":"653962e75c8e4863e1a2068f","avatarUrl":"/avatars/d4f5f5da141f37d53ca1986ff17b325e.svg","isPro":false,"fullname":"Mengting Ai","user":"famous-blue-raincoat","type":"user"},{"_id":"653429efe983fb23fa31485a","avatarUrl":"/avatars/3fa367efc49b82f9f1b2b909d085fc13.svg","isPro":false,"fullname":"Kaiyu He","user":"Meanstudent1","type":"user"},{"_id":"64cbc3e2a257a3212c00a115","avatarUrl":"/avatars/836e61be4aeda2080ddf2db9f2626cc6.svg","isPro":false,"fullname":"Furong Huang Lab at UMD","user":"furongh-lab","type":"user"},{"_id":"68adc450ec734f30b187f957","avatarUrl":"/avatars/29601811313ac40041eb025fb2c743cf.svg","isPro":true,"fullname":"Yoonkyo Jung","user":"yoonkyojung","type":"user"},{"_id":"66720ab819bebc69b5b93685","avatarUrl":"/avatars/b2f1314d9a26f6f5eaf6cebdb0d28812.svg","isPro":true,"fullname":"Yijun Liang","user":"joliang17","type":"user"},{"_id":"6790acb39551780939ae9d3d","avatarUrl":"/avatars/6bdd572d75b7b6736d25f48da09438cc.svg","isPro":false,"fullname":"Baicheng Chen","user":"Danny-1223","type":"user"},{"_id":"65ffbd04321ac5cd3409df17","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ffbd04321ac5cd3409df17/WFEtwl-oWK85HiZqxp1YW.jpeg","isPro":false,"fullname":"Lingzhi Yuan","user":"YuanXiaopang","type":"user"},{"_id":"64eacb4218d79efd5345aadf","avatarUrl":"/avatars/0345d5e6038268a3e992db37e6d540a0.svg","isPro":false,"fullname":"Mengxuan","user":"Mengxuan03021","type":"user"},{"_id":"638f26bb3783be5e1d04a86b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638f26bb3783be5e1d04a86b/iLDzwTKPAQcZJv7s6ZLcp.jpeg","isPro":false,"fullname":"Sy-Tuyen Ho","user":"hosytuyen","type":"user"},{"_id":"63b86a3daaa8cf17f0963b45","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1673030173042-noauth.png","isPro":false,"fullname":"Andrew","user":"mendeza","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64cbc5468174e45ae060ec46","name":"furonghuang-lab","fullname":"Furong Huang's Lab at UMD","avatar":"https://www.gravatar.com/avatar/add71ee6bbcef2277b077b42b3cba002?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28582.md","query":{}}">
Papers
arxiv:2607.28582

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Published on Jul 30
· Submitted by
Minghui Liu
on Jul 31
Authors:
,

Abstract

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the β=1 member of a broader policy-optimization family, where β weights the KL penalty anchoring the student to a reference policy. This equivalence turns β from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity to a reference policy against privileged teacher guidance. We introduce β-OPSD and derive its optimal policy as a geometric interpolation between the reference policy and the privileged teacher. Directly optimizing this objective with reinforcement learning, however, would be costly and high-variance. Rather than optimize the RL objective directly, we turn its closed-form solution into a distillation target. Each value of β selects a target along the reference-to-teacher path, which we implement efficiently by mixing their token-level logits. In this way, inexpensive distillation approximates the solution of expensive policy optimization. Return-to-go credit assignment further aligns token updates with the sequence-level objective while retaining the simplicity of OPSD. Experiments on mathematical reasoning benchmarks show that β-OPSD consistently outperforms vanilla OPSD, improving optimization stability and downstream reasoning performance. Our results provide a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.

Community

We introduce β-OPSD, a principled generalization of on-policy self-distillation for reasoning language models.
🔍 Our key observation is that vanilla OPSD is exactly the β = 1 case of a broader KL-regularized policy-optimization objective. Its optimal policy geometrically interpolates between a reference policy and a privileged teacher.
⚡ Instead of running costly, high-variance reinforcement learning, β-OPSD converts this closed-form solution into an efficient token-level distillation target. We also use return-to-go credit assignment to better align token updates with sequence-level outcomes.
📈 Across mathematical reasoning benchmarks and Qwen3 models from 1.7B to 8B, β-OPSD improves average performance and training stability over vanilla OPSD—including a +5.74-point average gain at the 1.7B scale.
🌐 Project page: https://umd-huang-lab.github.io/beta-opsd/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.28582
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.28582 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.28582 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.28582 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers