Hugging Face Daily Papers · · 3 min read

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Github Repo: <a href=\"https://github.com/tuandattt/dOPSD\" rel=\"nofollow\">https://github.com/tuandattt/dOPSD</a></p>\n","updatedAt":"2026-07-07T05:30:42.805Z","author":{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","fullname":"LIQIIIII","name":"LIQIIIII","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7811819911003113},"editors":["LIQIIIII"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04428","authors":[{"_id":"6a4c8e8425849b193a834304","name":"Phuong Tuan Dat","hidden":false},{"_id":"6a4c8e8425849b193a834305","name":"Qi Li","hidden":false},{"_id":"6a4c8e8425849b193a834306","name":"Xinchao Wang","hidden":false}],"publishedAt":"2026-07-05T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"dOPSD: On-Policy Self-Distillation for Diffusion Language Models","submittedOnDailyBy":{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","isPro":false,"fullname":"LIQIIIII","user":"LIQIIIII","type":"user","name":"LIQIIIII"},"summary":"Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-of-domain code generation, outperforming supervised and on-policy baselines.","upvotes":14,"discussionId":"6a4c8e8425849b193a834307","githubRepo":"https://github.com/tuandattt/dOPSD","githubRepoAddedBy":"user","ai_summary":"Diffusion large language models face challenges in reasoning enhancement through post-training, but a novel on-policy self-distillation method using internal denoising trajectories improves mathematical reasoning and code generation performance.","ai_keywords":["diffusion large language models","autoregressive models","supervised fine-tuning","exposure bias","reinforcement learning","on-policy self-distillation","teacher-student model","privileged information","denoising trajectory","in-domain math reasoning","out-of-domain code generation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":7},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","isPro":false,"fullname":"LIQIIIII","user":"LIQIIIII","type":"user"},{"_id":"677fbbf5f2e19477cb809830","avatarUrl":"/avatars/51af04f28038870f3ec418cc4909ecd0.svg","isPro":false,"fullname":"Tianbo Pan","user":"pan7386","type":"user"},{"_id":"6627cccfded9b7936d5d1d21","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6627cccfded9b7936d5d1d21/LKGr7EP7AirjmkbFZSn4o.jpeg","isPro":true,"fullname":"Guangnian Wan","user":"bigglesworthnotcat","type":"user"},{"_id":"6a0c4875512e8cf10c427be1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a0c4875512e8cf10c427be1/eMm3o7aXACXkRf2XQ_oEy.jpeg","isPro":false,"fullname":"Siao Tang","user":"ttu1818","type":"user"},{"_id":"6517ec0d593b3af3120c8cc6","avatarUrl":"/avatars/c2c3d69776f08cb24bd1524be0bb000e.svg","isPro":false,"fullname":"Nguyen Trong Huy","user":"Huyisbeee","type":"user"},{"_id":"6860fe55a1ab4d5c885c3edf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/HNSwVGoTd33mgbyjxA-fc.jpeg","isPro":false,"fullname":"QIN ZHIBIN","user":"tuantuan0321","type":"user"},{"_id":"67a4a26d5e65aa63c6d30e68","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67a4a26d5e65aa63c6d30e68/GtodlJGw-_IL2DTXQTucz.jpeg","isPro":false,"fullname":"Sicheng Feng","user":"FSCCS","type":"user"},{"_id":"668e740f1173ab43d9d9ed5e","avatarUrl":"/avatars/caa9b47c2a5f6d6d679759b8b234a0ab.svg","isPro":false,"fullname":"Zeqing Wang","user":"INV-WZQ","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"69fa9c2f8023820e9f89048b","avatarUrl":"/avatars/b93a22ac521faf63064d264695563ce3.svg","isPro":false,"fullname":"Elijah McMahon","user":"hello-world-i-am-human","type":"user"},{"_id":"63139adb289cf15634c8b18e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63139adb289cf15634c8b18e/Mb3Ev-ZrlvIN3P_hoervL.jpeg","isPro":true,"fullname":"Rand Arete","user":"sixwing","type":"user"},{"_id":"69a2c48af29b3c3292e5d754","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/zc6q7JS2ZgTsl3BH-vFIi.png","isPro":false,"fullname":"Siyu Luo","user":"luke-smith","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04428.md","query":{}}">
Papers
arxiv:2607.04428

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

Published on Jul 5
· Submitted by
LIQIIIII
on Jul 7
Authors:
,
,

Abstract

Diffusion large language models face challenges in reasoning enhancement through post-training, but a novel on-policy self-distillation method using internal denoising trajectories improves mathematical reasoning and code generation performance.

Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-of-domain code generation, outperforming supervised and on-policy baselines.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.04428
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.04428 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.04428 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.04428 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers