Hugging Face Daily Papers · · 3 min read

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Code: <a href=\"https://github.com/YuYingLi0/FiRe-OPD\" rel=\"nofollow\">https://github.com/YuYingLi0/FiRe-OPD</a></p>\n","updatedAt":"2026-06-04T04:41:48.249Z","author":{"_id":"649d54b314afbb10ce2a9eeb","avatarUrl":"/avatars/15c325d8c2273ff63569f23015e98486.svg","fullname":"Hangjie Yuan","name":"JacobYuan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8621976971626282},"editors":["JacobYuan"],"editorAvatarUrls":["/avatars/15c325d8c2273ff63569f23015e98486.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.02684","authors":[{"_id":"6a202b6615100c5272a841d5","name":"Yuying Li","hidden":false},{"_id":"6a202b6615100c5272a841d6","name":"Leqi Zheng","hidden":false},{"_id":"6a202b6615100c5272a841d7","name":"Yongzi Yu","hidden":false},{"_id":"6a202b6615100c5272a841d8","name":"Wenrui Zhou","hidden":false},{"_id":"6a202b6615100c5272a841d9","name":"Xuchang Zhong","hidden":false},{"_id":"6a202b6615100c5272a841da","name":"Xing Hu","hidden":false},{"_id":"6a202b6615100c5272a841db","name":"Jing Jin","hidden":false},{"_id":"6a202b6615100c5272a841dc","name":"Huangjie Yuan","hidden":false},{"_id":"6a202b6615100c5272a841dd","name":"Tao Feng","hidden":false}],"publishedAt":"2026-06-01T00:00:00.000Z","submittedOnDailyAt":"2026-06-04T00:00:00.000Z","title":"Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation","submittedOnDailyBy":{"_id":"649d54b314afbb10ce2a9eeb","avatarUrl":"/avatars/15c325d8c2273ff63569f23015e98486.svg","isPro":false,"fullname":"Hangjie Yuan","user":"JacobYuan","type":"user","name":"JacobYuan"},"summary":"On-Policy distillation (OPD) in large language models is shifting from full-trace KL supervision toward more selective training paradigms. Recent OPD methods increasingly focus on selecting which trajectories to learn from, which tokens are most informative, and which supervision signals are most reliable. Motivated by this trend, we rethink optimization granularity of OPD and propose \\fireicon\\ FiRe-OPD (Filter, then Reweight), which jointly adjusts supervision signals at both trajectory and token levels. In details, FiRe-OPD first filters trajectories to remove low-quality rollout samples, and then applies soft reweighting within the retained trajectories to emphasize informative tokens. Compared with hard token selection, FiRe-OPD leverages a soft-weighting mechanism to effectively mitigate information loss and enhance optimization stability, thereby achieving finer-grained OPD optimization. We validate the effectiveness of FiRe-OPD across strong-to-weak, single-teacher, and multi-teacher settings, and demonstrate its superiority over recent token-level OPD methods ( (e.g., +6.25 on AIME 2024 in strong-to-weak, +18.81 on Miner in multi-teacher). Our code is available at https://github.com/YuYingLi0/FiRe-OPD.","upvotes":8,"discussionId":"6a202b6615100c5272a841f0","githubRepo":"https://github.com/YuYingLi0/FiRe-OPD","githubRepoAddedBy":"user","ai_summary":"FiRe-OPD improves on-policy distillation in large language models by filtering low-quality trajectories and applying soft reweighting to enhance informative token selection and optimization stability.","ai_keywords":["on-policy distillation","KL supervision","trajectory selection","token selection","supervision signals","soft reweighting","optimization granularity","FiRe-OPD","trajectory filtering","token weighting"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":4,"organization":{"_id":"628735cbc83a2d6ab8d14a66","name":"Tsinghua","fullname":"Tsinghua University","avatar":"https://www.gravatar.com/avatar/6c5c1441e3283e7543342e59277ea219?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"649d54b314afbb10ce2a9eeb","avatarUrl":"/avatars/15c325d8c2273ff63569f23015e98486.svg","isPro":false,"fullname":"Hangjie Yuan","user":"JacobYuan","type":"user"},{"_id":"64f44021220c2e5e96628595","avatarUrl":"/avatars/43bcb389760d1037dde1f81c4398d260.svg","isPro":true,"fullname":"gdwind LQ","user":"gdwind","type":"user"},{"_id":"644d2532d185572dd1e48f90","avatarUrl":"/avatars/5831acebb02d8bc8f80f56b7b11c7c69.svg","isPro":false,"fullname":"Zhu","user":"zzzhu","type":"user"},{"_id":"680b3903f209ad3256527538","avatarUrl":"/avatars/23d715718f23d00cc8c370e2e21c514e.svg","isPro":false,"fullname":"XuzhaoLi","user":"LXZ1OOO","type":"user"},{"_id":"678609789a285d232ee14157","avatarUrl":"/avatars/a6cb2c571d9ef6deb0b1659f754afe7f.svg","isPro":false,"fullname":"Weichu Xie","user":"akarinmoe","type":"user"},{"_id":"6708cca079b2d9e035c4a54e","avatarUrl":"/avatars/ed7daf61627ea910ec2de4582abf5a0d.svg","isPro":false,"fullname":"xichen","user":"xichen-fy","type":"user"},{"_id":"66a246cae22bfd8d72560696","avatarUrl":"/avatars/97cec177a0eb0a8465e3532b888ef1d4.svg","isPro":false,"fullname":"zrchen","user":"zrchen03","type":"user"},{"_id":"69bcc905518f6d1f3d2b86df","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/t-ZEqYEkqPI8-saats0nT.jpeg","isPro":false,"fullname":"Liu Jingyi","user":"gao-wenxuan2","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"628735cbc83a2d6ab8d14a66","name":"Tsinghua","fullname":"Tsinghua University","avatar":"https://www.gravatar.com/avatar/6c5c1441e3283e7543342e59277ea219?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.02684.md"}">
Papers
arxiv:2606.02684

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation

Published on Jun 1
· Submitted by
Hangjie Yuan
on Jun 4
Authors:
,
,
,
,
,
,
,
,

Abstract

FiRe-OPD improves on-policy distillation in large language models by filtering low-quality trajectories and applying soft reweighting to enhance informative token selection and optimization stability.

On-Policy distillation (OPD) in large language models is shifting from full-trace KL supervision toward more selective training paradigms. Recent OPD methods increasingly focus on selecting which trajectories to learn from, which tokens are most informative, and which supervision signals are most reliable. Motivated by this trend, we rethink optimization granularity of OPD and propose \fireicon\ FiRe-OPD (Filter, then Reweight), which jointly adjusts supervision signals at both trajectory and token levels. In details, FiRe-OPD first filters trajectories to remove low-quality rollout samples, and then applies soft reweighting within the retained trajectories to emphasize informative tokens. Compared with hard token selection, FiRe-OPD leverages a soft-weighting mechanism to effectively mitigate information loss and enhance optimization stability, thereby achieving finer-grained OPD optimization. We validate the effectiveness of FiRe-OPD across strong-to-weak, single-teacher, and multi-teacher settings, and demonstrate its superiority over recent token-level OPD methods ( (e.g., +6.25 on AIME 2024 in strong-to-weak, +18.81 on Miner in multi-teacher). Our code is available at https://github.com/YuYingLi0/FiRe-OPD.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.02684
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.02684 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.02684 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.02684 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers