Hugging Face Daily Papers · · 4 min read

Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.</p>\n","updatedAt":"2026-09-03T02:02:05.273Z","author":{"_id":"65037565da2d88e201f63b7a","avatarUrl":"/avatars/d1b6ce17236360e9583b8bb4cb87e506.svg","fullname":"Runpeng Dai","name":"Leo-Dai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8628391623497009},"editors":["Leo-Dai"],"editorAvatarUrls":["/avatars/d1b6ce17236360e9583b8bb4cb87e506.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.29846","authors":[{"_id":"6a98d4fbfea818274321fe55","name":"Run Yang","hidden":false},{"_id":"6a98d4fbfea818274321fe56","user":{"_id":"65037565da2d88e201f63b7a","avatarUrl":"/avatars/d1b6ce17236360e9583b8bb4cb87e506.svg","isPro":false,"fullname":"Runpeng Dai","user":"Leo-Dai","type":"user","name":"Leo-Dai"},"name":"Runpeng Dai","status":"claimed_verified","statusLastChangedAt":"2026-09-03T08:17:05.210Z","hidden":false},{"_id":"6a98d4fbfea818274321fe57","name":"Jie Sun","hidden":false},{"_id":"6a98d4fbfea818274321fe58","name":"Jielei Zhang","hidden":false},{"_id":"6a98d4fbfea818274321fe59","name":"Fan Zhou","hidden":false},{"_id":"6a98d4fbfea818274321fe5a","name":"Hongtu Zhu","hidden":false},{"_id":"6a98d4fbfea818274321fe5b","name":"Peiyi Li","hidden":false},{"_id":"6a98d4fbfea818274321fe5c","name":"Longwen Gao","hidden":false}],"publishedAt":"2026-08-30T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation","submittedOnDailyBy":{"_id":"65037565da2d88e201f63b7a","avatarUrl":"/avatars/d1b6ce17236360e9583b8bb4cb87e506.svg","isPro":false,"fullname":"Runpeng Dai","user":"Leo-Dai","type":"user","name":"Leo-Dai"},"summary":"Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.","upvotes":14,"discussionId":"6a98d4fcfea818274321fe5d","ai_summary":"Influence-Directed Adaptive On-Policy Distillation improves diversity transfer in reasoning model distillation by selectively preserving entropy-expanding updates and replacing entropy-contracting ones with adaptive advantage shrinkage.","ai_keywords":["on-policy distillation","First-Order Local Entropy Influence","entropy contraction","IDA-OPD","divergence-adaptive advantage shrinkage","pass@k","reasoning-oriented distillation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"66b60f9953d81b9a9abde088","name":"Leander001","fullname":"Blibli","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66b5249a6cef22e6ba45a492/DYg1_jfzyxr8fxDFo018c.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65037565da2d88e201f63b7a","avatarUrl":"/avatars/d1b6ce17236360e9583b8bb4cb87e506.svg","isPro":false,"fullname":"Runpeng Dai","user":"Leo-Dai","type":"user"},{"_id":"6a80a8db8bbb1c5aa1d81682","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a80a8db8bbb1c5aa1d81682/z6CVFKBftseCN5KfoSt-4.jpeg","isPro":false,"fullname":"Clara Schroeder","user":"SCHROEDERva","type":"user"},{"_id":"64e6c617ecce34cb442cb208","avatarUrl":"/avatars/ebc61bf6a043314cb2089b1efd5e6a18.svg","isPro":false,"fullname":"JieSun(SII)","user":"Sunshine279","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"69bcd98076fe7a5ea0a81b7f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/naciiDu7-0tHRc4I0CVwi.png","isPro":false,"fullname":"Luo Yuxuan","user":"joseph-moore521","type":"user"},{"_id":"647073267fd7ecdbd0e9f06c","avatarUrl":"/avatars/8d7d7d257d764edeab0580e16a606571.svg","isPro":false,"fullname":"runyang","user":"0v01111","type":"user"},{"_id":"69b391a2c8ca078940878a38","avatarUrl":"/avatars/63c8213a5b9c2a7e8f16619cf99ebd87.svg","isPro":false,"fullname":"Qiang Sun","user":"qsunstats","type":"user"},{"_id":"6a98e7785fb7580424e2350b","avatarUrl":"/avatars/feeba5225a041ec6df488e2b09f0b859.svg","isPro":false,"fullname":"Wenxin E.E","user":"wenxinee","type":"user"},{"_id":"6468ed2ae134d050a58b4c4e","avatarUrl":"/avatars/e54f43f80b5da6ab12ace5d34fef3d1f.svg","isPro":false,"fullname":"mathsfan-spark","user":"ml-llm-12345","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6a5d9c726a7dffea4340db6c","avatarUrl":"/avatars/76823670a4fe8b59a8a8c3f7a8891683.svg","isPro":false,"fullname":"오예은","user":"florineh22","type":"user"},{"_id":"6a9910b24b64afb72baeef59","avatarUrl":"/avatars/88fc854e95ea268fceb257c6afd6345f.svg","isPro":false,"fullname":"22y","user":"22yuu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66b60f9953d81b9a9abde088","name":"Leander001","fullname":"Blibli","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66b5249a6cef22e6ba45a492/DYg1_jfzyxr8fxDFo018c.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.29846.md","query":{}}">
Papers
arxiv:2608.29846

Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation

Published on Aug 30
· Submitted by
Runpeng Dai
on Sep 3
Authors:
,

Abstract

Influence-Directed Adaptive On-Policy Distillation improves diversity transfer in reasoning model distillation by selectively preserving entropy-expanding updates and replacing entropy-contracting ones with adaptive advantage shrinkage.

Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.

Community

Paper author Paper submitter about 7 hours ago

Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.29846
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.29846 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.29846 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.29846 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers