Hugging Face Daily Papers · · 4 min read

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

SA-OPD addresses a previously overlooked failure mode in on-policy distillation: teacher signals can appear confident, informative, and learnable while being driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates, producing large gradients with little task-improving direction. By comparing token-level teacher–student divergence under the original input and a residual no-prompt context, SA-OPD provides a lightweight proxy for input-groundedness—without requiring external verification labels or auxiliary judges. It then filters only tokens that combine low input-groundedness with high absolute distillation divergence, targeting the high-impact portion of dense OPD supervision most likely to induce harmful updates. Across LLM and VLM distillation settings, SA-OPD consistently improves over Vanilla OPD and strong selective OPD baselines on mathematical reasoning, visual understanding, and visual reasoning benchmarks, offering a practical direction for more reliable, input-grounded distillation signal selection in on-policy distillation.</p>\n","updatedAt":"2026-08-06T02:30:31.099Z","author":{"_id":"637b6b0056db0404b7c74e3e","avatarUrl":"/avatars/2e010f03c1682fecec2165ead3f3f85a.svg","fullname":"Yinuo Jiang","name":"jenoj","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8377023339271545},"editors":["jenoj"],"editorAvatarUrls":["/avatars/2e010f03c1682fecec2165ead3f3f85a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.03632","authors":[{"_id":"6a72f3115e0a61a8ccd03ba6","user":{"_id":"637b6b0056db0404b7c74e3e","avatarUrl":"/avatars/2e010f03c1682fecec2165ead3f3f85a.svg","isPro":false,"fullname":"Yinuo Jiang","user":"jenoj","type":"user","name":"jenoj"},"name":"Yinuo Jiang","status":"claimed_verified","statusLastChangedAt":"2026-08-05T16:45:04.560Z","hidden":false},{"_id":"6a72f3115e0a61a8ccd03ba7","name":"Yongjie Ye","hidden":false},{"_id":"6a72f3115e0a61a8ccd03ba8","name":"Zhou Tao","hidden":false},{"_id":"6a72f3115e0a61a8ccd03ba9","name":"Xiang Zhuang","hidden":false},{"_id":"6a72f3115e0a61a8ccd03baa","name":"Qiang Zhang","hidden":false},{"_id":"6a72f3115e0a61a8ccd03bab","name":"Huajun Chen","hidden":false},{"_id":"6a72f3115e0a61a8ccd03bac","name":"Tiankai Li","hidden":false}],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation","submittedOnDailyBy":{"_id":"637b6b0056db0404b7c74e3e","avatarUrl":"/avatars/2e010f03c1682fecec2165ead3f3f85a.svg","isPro":false,"fullname":"Yinuo Jiang","user":"jenoj","type":"user","name":"jenoj"},"summary":"On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.","upvotes":18,"discussionId":"6a72f3115e0a61a8ccd03bad","githubRepo":"https://github.com/jjjyinuo/SA-OPD","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64d201b1c2bd235422fb1d14","avatarUrl":"/avatars/e50581aa66391cedae94e116e759b9ec.svg","isPro":false,"fullname":"wang","user":"stormthunder","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"},{"_id":"623ad8954e235e6a6961f063","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648023697580-noauth.jpeg","isPro":false,"fullname":"Xuecheng Wu","user":"Conna","type":"user"},{"_id":"637b6b0056db0404b7c74e3e","avatarUrl":"/avatars/2e010f03c1682fecec2165ead3f3f85a.svg","isPro":false,"fullname":"Yinuo Jiang","user":"jenoj","type":"user"},{"_id":"66095a7e362a1d713a8a8c85","avatarUrl":"/avatars/8ef3850285573798fb8916f9dd1f1641.svg","isPro":false,"fullname":"zhou tao","user":"gulitaozhou","type":"user"},{"_id":"65a892fc162efc9aef8f896d","avatarUrl":"/avatars/0c0a84fe75fd7d514b9ccd093fccc663.svg","isPro":false,"fullname":"ZHANG","user":"YIWEN123","type":"user"},{"_id":"64fea99a01aedd0e8615a8e9","avatarUrl":"/avatars/bb14ed43b3fc22c24ec177b7456b7e99.svg","isPro":false,"fullname":"Xiang C.H.","user":"LockOnN","type":"user"},{"_id":"64328f1a6c2a26ae66cfdde6","avatarUrl":"/avatars/9732f63dbdbba2f10950cc2c8242fb00.svg","isPro":false,"fullname":"Jenny","user":"Smilingk","type":"user"},{"_id":"66e93dec6065253cb09bbf7c","avatarUrl":"/avatars/80acc6f5107cca926f4ae419c4dd2655.svg","isPro":false,"fullname":"myR_001","user":"myR-001","type":"user"},{"_id":"655d6461d246e013226d431f","avatarUrl":"/avatars/cc3e31c387719bf00469e6b27faa8b84.svg","isPro":false,"fullname":"Xiang Zhuang","user":"XiangZH","type":"user"},{"_id":"6782116fa8027d56f0225f4c","avatarUrl":"/avatars/b87ae0c5123fdde6ce1d954edd4bc52a.svg","isPro":false,"fullname":"Shaobo Ju","user":"Luminousllsa","type":"user"},{"_id":"668e4c1034b1d6e49fddaf49","avatarUrl":"/avatars/77aa8592be32bb94a83f5ec149b698ac.svg","isPro":false,"fullname":"wang","user":"siqi04","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.03632.md","query":{}}">
Papers
arxiv:2608.03632

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

Published on Aug 4
· Submitted by
Yinuo Jiang
on Aug 6
Authors:

Abstract

On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.

Community

Paper author Paper submitter about 8 hours ago

SA-OPD addresses a previously overlooked failure mode in on-policy distillation: teacher signals can appear confident, informative, and learnable while being driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates, producing large gradients with little task-improving direction. By comparing token-level teacher–student divergence under the original input and a residual no-prompt context, SA-OPD provides a lightweight proxy for input-groundedness—without requiring external verification labels or auxiliary judges. It then filters only tokens that combine low input-groundedness with high absolute distillation divergence, targeting the high-impact portion of dense OPD supervision most likely to induce harmful updates. Across LLM and VLM distillation settings, SA-OPD consistently improves over Vanilla OPD and strong selective OPD baselines on mathematical reasoning, visual understanding, and visual reasoning benchmarks, offering a practical direction for more reliable, input-grounded distillation signal selection in on-policy distillation.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.03632
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.03632 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.03632 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.03632 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers