Hugging Face Daily Papers · · 5 min read

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.</p>\n","updatedAt":"2026-08-24T04:54:27.915Z","author":{"_id":"65a0aade5fafc248c2156e95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a0aade5fafc248c2156e95/S9YjJMTuKc-U1cFizqUMA.jpeg","fullname":"DeyangKong","name":"DeyangKong","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9316017627716064},"editors":["DeyangKong"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65a0aade5fafc248c2156e95/S9YjJMTuKc-U1cFizqUMA.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16647","authors":[{"_id":"6a8bce833d26296ea30919ab","name":"Zhaoyi Li","hidden":false},{"_id":"6a8bce833d26296ea30919ac","name":"Deyang Kong","hidden":false},{"_id":"6a8bce833d26296ea30919ad","name":"Yuan Wei","hidden":false},{"_id":"6a8bce833d26296ea30919ae","name":"Evan Yang","hidden":false},{"_id":"6a8bce833d26296ea30919af","name":"Ranran Shen","hidden":false},{"_id":"6a8bce833d26296ea30919b0","name":"Mahardika Krisna Ihsani","hidden":false},{"_id":"6a8bce833d26296ea30919b1","name":"Ming Yang","hidden":false},{"_id":"6a8bce833d26296ea30919b2","name":"Wei Zhang","hidden":false},{"_id":"6a8bce833d26296ea30919b3","name":"Chuan Hao","hidden":false},{"_id":"6a8bce833d26296ea30919b4","name":"Jian Yang","hidden":false},{"_id":"6a8bce833d26296ea30919b5","name":"Ran Tao","hidden":false},{"_id":"6a8bce833d26296ea30919b6","name":"Bryan Dai","hidden":false},{"_id":"6a8bce833d26296ea30919b7","name":"Shikun Zhang","hidden":false},{"_id":"6a8bce833d26296ea30919b8","name":"Wei Ye","hidden":false},{"_id":"6a8bce833d26296ea30919b9","name":"Ying Wei","hidden":false},{"_id":"6a8bce833d26296ea30919ba","name":"Defu Lian","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-24T00:00:00.000Z","title":"Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models","submittedOnDailyBy":{"_id":"65a0aade5fafc248c2156e95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a0aade5fafc248c2156e95/S9YjJMTuKc-U1cFizqUMA.jpeg","isPro":false,"fullname":"DeyangKong","user":"DeyangKong","type":"user","name":"DeyangKong"},"summary":"On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.","upvotes":9,"discussionId":"6a8bce833d26296ea30919bb","ai_summary":"On-policy distillation transfers reasoning behaviors rather than specific answers, with generalization strongly tied to teacher-student origin alignment and multi-teacher combinations causing capability trade-offs.","ai_keywords":["on-policy distillation","reasoning behavior","cross-domain transfer","multi-teacher","same-origin","cross-origin","distribution shift"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6937e1578cd2568ba4c9d3f9","name":"IQuestLab","fullname":"IQuest","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68c80cc3a1ca9a73c17d29f7/pjOEigZuk_UQ4dqyS482V.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65a0aade5fafc248c2156e95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a0aade5fafc248c2156e95/S9YjJMTuKc-U1cFizqUMA.jpeg","isPro":false,"fullname":"DeyangKong","user":"DeyangKong","type":"user"},{"_id":"64212c24b286e8c464f95f13","avatarUrl":"/avatars/30d2e6aea3812859e85eaac6d99bd6c5.svg","isPro":false,"fullname":"Li Zhaoyi","user":"joeylee2333","type":"user"},{"_id":"6915af0d5d35d03a50a22b91","avatarUrl":"/avatars/ac111a1e33c21b3924f3fedbd1029195.svg","isPro":false,"fullname":"Qing Yi","user":"qyi-iquest","type":"user"},{"_id":"6499a36cb1365e4af3519420","avatarUrl":"/avatars/4db5c21bbe5eaab9eec74f0421ea82ed.svg","isPro":false,"fullname":"duolaxianren","user":"duolaxianren","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"67c035c5e9ef20cf3c1ebbef","avatarUrl":"/avatars/1e2471a1795006e9600b3d32be3818cf.svg","isPro":false,"fullname":"Louis Wong","user":"louiswng","type":"user"},{"_id":"6a5c5d55f1a1f0c587b22359","avatarUrl":"/avatars/be08bb87649f7f32f56a9dd63b9a0391.svg","isPro":false,"fullname":"qingfeng","user":"qf-iquest","type":"user"},{"_id":"63edb098679c2cc40abc6c2e","avatarUrl":"/avatars/288c7229937c2c3f29fda6d17c7df2eb.svg","isPro":false,"fullname":"Xiangyu","user":"xixy","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6937e1578cd2568ba4c9d3f9","name":"IQuestLab","fullname":"IQuest","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68c80cc3a1ca9a73c17d29f7/pjOEigZuk_UQ4dqyS482V.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16647.md","query":{}}">
Papers
arxiv:2608.16647

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

Published on Aug 17
· Submitted by
DeyangKong
on Aug 24
#3 Paper of the day
Authors:
,

Abstract

On-policy distillation transfers reasoning behaviors rather than specific answers, with generalization strongly tied to teacher-student origin alignment and multi-teacher combinations causing capability trade-offs.

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

Community

Paper submitter about 3 hours ago

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16647
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16647 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16647 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16647 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers