Hugging Face Daily Papers · · 4 min read

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Standard OPD may under-cover plausible alternatives, while EOPD relies solely on teacher entropy and uncalibrated targets. SPOT sparsely probes informative positions and uses verifier-scored continuations to construct outcome-calibrated targets. Across three student scales, SPOT improves macro Avg@8/Pass@8 over OPD by 0.47–1.48/4.55–5.28 points and over EOPD by 0.29–0.68/2.49–3.19 points.</p>\n","updatedAt":"2026-08-11T02:22:33.127Z","author":{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","fullname":"Zhongxiang Dai","name":"dzxagent","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8731234669685364},"editors":["dzxagent"],"editorAvatarUrls":["/avatars/3367920f104ca45305944f6bc1d6b8df.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04419","authors":[{"_id":"6a782d898e9301703eaa5c6b","name":"Zikun Qu","hidden":false},{"_id":"6a782d898e9301703eaa5c6c","name":"Min Zhang","hidden":false},{"_id":"6a782d898e9301703eaa5c6d","name":"Mingze Kong","hidden":false},{"_id":"6a782d898e9301703eaa5c6e","name":"Zhiwei Shang","hidden":false},{"_id":"6a782d898e9301703eaa5c6f","name":"Yikun Ban","hidden":false},{"_id":"6a782d898e9301703eaa5c70","name":"Shuang Qiu","hidden":false},{"_id":"6a782d898e9301703eaa5c71","user":{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user","name":"dzxagent"},"name":"Zhongxiang Dai","status":"claimed_verified","statusLastChangedAt":"2026-08-10T08:45:04.281Z","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation","submittedOnDailyBy":{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user","name":"dzxagent"},"summary":"On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-k candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.","upvotes":16,"discussionId":"6a782d898e9301703eaa5c72","githubRepo":"https://github.com/QuZikun/SPOT","githubRepoAddedBy":"user","ai_summary":"SPOT improves on-policy distillation by selectively probing uncertain positions and calibrating targets to downstream outcomes, boosting reasoning quality and coverage.","ai_keywords":["on-policy distillation","reverse-KL","teacher entropy","sparse probing","outcome-calibrated targets","acquisition-exploration-exploitation","verifier-scored continuations","KL-regularized target"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6223644d0129f2097d69a407","name":"CUHKSZ","fullname":"Chinese University of Hong Kong, Shenzhen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1646486592158-6108ae87823007eaf0c7bd1e.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d53b638a15934c10f2c16d","avatarUrl":"/avatars/a8431e4554d68708562f45bc6a8544d2.svg","isPro":false,"fullname":"fafa","user":"kunkun0919","type":"user"},{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user"},{"_id":"67c54394bb902949d6548c5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/bpfbtewGpK5CdP6DZmiPb.png","isPro":false,"fullname":"Zhiwei Shang","user":"javenshang","type":"user"},{"_id":"67d654e797767f4925a48fee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/0X51HpzNzAdcSJnd9k9Oz.png","isPro":false,"fullname":"Xinyi Yuan","user":"xinyyyiii","type":"user"},{"_id":"69bac4dde4325f7946a34143","avatarUrl":"/avatars/20bd4d0afdbafea13c8ab82acadc00fe.svg","isPro":false,"fullname":"s","user":"wsghc","type":"user"},{"_id":"66bf1156eb4f43ee8a525f2f","avatarUrl":"/avatars/3d738d7d1c535a2ef8d402f2b580e9df.svg","isPro":false,"fullname":"Tywh202","user":"Tywh202","type":"user"},{"_id":"65df436f487081eb2c160e57","avatarUrl":"/avatars/dda2d6fc83c988791f1710ee36d1c066.svg","isPro":false,"fullname":"Mingze Kong","user":"EmbPhy","type":"user"},{"_id":"6920035cd48b817fb1297da3","avatarUrl":"/avatars/109a16cb17427cec1248e70ff09ac239.svg","isPro":false,"fullname":"Mingrong Gong","user":"Gmr5233","type":"user"},{"_id":"67c0acd7ec3563f6b93d3f51","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/S3eKPZswOffvGVY6U5HIN.png","isPro":false,"fullname":"zifan wang","user":"aCodeDog","type":"user"},{"_id":"6a7962e9306352b010760546","avatarUrl":"/avatars/2abb83142b69c7d51a299ccc71cccd7a.svg","isPro":false,"fullname":"LAIqp","user":"qpLAI","type":"user"},{"_id":"68673435069b1aa9224b88f1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Fxq0yk1NLb0oAwxodlcuD.png","isPro":false,"fullname":"Junfeng Liao","user":"Junfeng01","type":"user"},{"_id":"68345345f4bbf856e2d708e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68345345f4bbf856e2d708e2/L5H2HNCuWje3ti2tNbC5p.jpeg","isPro":false,"fullname":"Yikun Ban","user":"Yikunb","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6223644d0129f2097d69a407","name":"CUHKSZ","fullname":"Chinese University of Hong Kong, Shenzhen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1646486592158-6108ae87823007eaf0c7bd1e.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04419.md","query":{}}">
Papers
arxiv:2608.04419

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

Published on Aug 5
· Submitted by
Zhongxiang Dai
on Aug 11
Authors:
,

Abstract

SPOT improves on-policy distillation by selectively probing uncertain positions and calibrating targets to downstream outcomes, boosting reasoning quality and coverage.

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-k candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.

Community

Paper author Paper submitter about 17 hours ago

Standard OPD may under-cover plausible alternatives, while EOPD relies solely on teacher entropy and uncalibrated targets. SPOT sparsely probes informative positions and uses verifier-scored continuations to construct outcome-calibrated targets. Across three student scales, SPOT improves macro Avg@8/Pass@8 over OPD by 0.47–1.48/4.55–5.28 points and over EOPD by 0.29–0.68/2.49–3.19 points.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04419
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.04419 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.04419 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.04419 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers