Hugging Face Daily Papers · · 4 min read

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We investigate the data side and the training dynamics of on-policy distillation (OPD), and try to answer how far the training set of OPD can be reduced, and find that a single training example already induces most of the states a full dataset visits.</p>\n<p>Check the repo: <a href=\"https://github.com/Thinking-Space/One-Shot-OPD\" rel=\"nofollow\">https://github.com/Thinking-Space/One-Shot-OPD</a></p>\n","updatedAt":"2026-09-04T02:32:45.310Z","author":{"_id":"64c5e944979493279b700cb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vjFuPWw8Vl7b7gXB19Sk-.jpeg","fullname":"Bingxiang He","name":"hbx","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9203958511352539},"editors":["hbx"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vjFuPWw8Vl7b7gXB19Sk-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04172","authors":[{"_id":"6a9a2c968f7c3b755723948c","user":{"_id":"6711f02ee6f26ad913f2e030","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6711f02ee6f26ad913f2e030/THMWxLyZDiXzkj1UN85H6.jpeg","isPro":false,"fullname":"Zixuan Fu","user":"ZixuanFu","type":"user","name":"ZixuanFu"},"name":"Zixuan Fu","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.265Z","hidden":false},{"_id":"6a9a2c968f7c3b755723948d","name":"Bingxiang He","hidden":false},{"_id":"6a9a2c968f7c3b755723948e","name":"Yuxin Zuo","hidden":false},{"_id":"6a9a2c968f7c3b755723948f","user":{"_id":"6860d8e658095f461a850ff3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6860d8e658095f461a850ff3/BpOMOqf5B0s9e_8hVHy6W.jpeg","isPro":false,"fullname":"Haohuan Huang","user":"hhh675597","type":"user","name":"hhh675597"},"name":"Haohuan Huang","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.271Z","hidden":false},{"_id":"6a9a2c968f7c3b7557239490","name":"Jinqian Zhang","hidden":false},{"_id":"6a9a2c968f7c3b7557239491","name":"Ruhang Xiao","hidden":false},{"_id":"6a9a2c968f7c3b7557239492","name":"Cheng Qian","hidden":false},{"_id":"6a9a2c968f7c3b7557239493","name":"Qinyu Luo","hidden":false},{"_id":"6a9a2c968f7c3b7557239494","name":"Huan-ang Gao","hidden":false},{"_id":"6a9a2c968f7c3b7557239495","name":"Yudong Wang","hidden":false},{"_id":"6a9a2c968f7c3b7557239496","name":"Zhiyuan Liu","hidden":false},{"_id":"6a9a2c968f7c3b7557239497","name":"Ning Ding","hidden":false},{"_id":"6a9a2c968f7c3b7557239498","name":"Chaojun Xiao","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"Rethinking On-Policy Distillation of Large Language Models II: One Training Example","submittedOnDailyBy":{"_id":"64c5e944979493279b700cb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vjFuPWw8Vl7b7gXB19Sk-.jpeg","isPro":false,"fullname":"Bingxiang He","user":"hbx","type":"user","name":"hbx"},"summary":"On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure state coverage, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \\(71.5\\%\\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \\(98.9\\%\\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.","upvotes":47,"discussionId":"6a9a2c968f7c3b7557239499","ai_summary":"On-policy distillation improves over hundreds of steps from a single query by rapidly covering teacher states, yet student alignment remains slow, indicating the method is algorithm-starved rather than data-starved.","ai_keywords":["on-policy distillation","token-level supervision","state coverage","multi-teacher OPD","alignment rate","step efficiency"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6a9973c4c66b46d52427b4ec","name":"Thinking-Space","fullname":"Thinking Space","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/622474f38dc6b0b64f5e903d/korkveUGvDBf5V1iJtXSq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64c5e944979493279b700cb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vjFuPWw8Vl7b7gXB19Sk-.jpeg","isPro":false,"fullname":"Bingxiang He","user":"hbx","type":"user"},{"_id":"622474f38dc6b0b64f5e903d","avatarUrl":"/avatars/d6b60a014277a8ec7d564163c5f644aa.svg","isPro":false,"fullname":"Yuxin Zuo","user":"yuxinzuo","type":"user"},{"_id":"6711f02ee6f26ad913f2e030","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6711f02ee6f26ad913f2e030/THMWxLyZDiXzkj1UN85H6.jpeg","isPro":false,"fullname":"Zixuan Fu","user":"ZixuanFu","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6468ed2ae134d050a58b4c4e","avatarUrl":"/avatars/e54f43f80b5da6ab12ace5d34fef3d1f.svg","isPro":false,"fullname":"mathsfan-spark","user":"ml-llm-12345","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"65697feb9fb2d79a79e14e0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65697feb9fb2d79a79e14e0a/wVGaBjn8pQIJneZWSFIwS.jpeg","isPro":false,"fullname":"haodi lei","user":"bingyang-lei","type":"user"},{"_id":"663f07d029be04778ba97871","avatarUrl":"/avatars/fb7c9d4a2c537d918a3267e7cbc03f04.svg","isPro":false,"fullname":"Xingtai Lv","user":"XingtaiHF","type":"user"},{"_id":"66f15f9125991d4ac9d87e4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f15f9125991d4ac9d87e4d/nbBws6rC6-CcciYz6w0lt.png","isPro":false,"fullname":"陈英豪","user":"cyh2004","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6a6910e1611ac85b10b249d1","avatarUrl":"/avatars/fafa03d615903010d5a85bf25babf746.svg","isPro":false,"fullname":"quietlane","user":"quietlane","type":"user"},{"_id":"69804cdd555beac9be536071","avatarUrl":"/avatars/380c3021671962416fab240c179972da.svg","isPro":false,"fullname":"mingzhang","user":"Mingzh21","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a9973c4c66b46d52427b4ec","name":"Thinking-Space","fullname":"Thinking Space","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/622474f38dc6b0b64f5e903d/korkveUGvDBf5V1iJtXSq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04172.md","query":{}}">
Papers
arxiv:2609.04172

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Published on Sep 3
· Submitted by
Bingxiang He
on Sep 4
Authors:

Abstract

On-policy distillation improves over hundreds of steps from a single query by rapidly covering teacher states, yet student alignment remains slow, indicating the method is algorithm-starved rather than data-starved.

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure state coverage, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.

Community

Paper submitter about 6 hours ago

We investigate the data side and the training dynamics of on-policy distillation (OPD), and try to answer how far the training set of OPD can be reduced, and find that a single training example already induces most of the states a full dataset visits.

Check the repo: https://github.com/Thinking-Space/One-Shot-OPD

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.04172
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.04172 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.04172 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers