Standard OPD may under-cover plausible alternatives, while EOPD relies solely on teacher entropy and uncalibrated targets. SPOT sparsely probes informative positions and uses verifier-scored continuations to construct outcome-calibrated targets. Across three student scales, SPOT improves macro Avg@8/Pass@8 over OPD by 0.47–1.48/4.55–5.28 points and over EOPD by 0.29–0.68/2.49–3.19 points.</p>\n","updatedAt":"2026-08-11T02:22:33.127Z","author":{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","fullname":"Zhongxiang Dai","name":"dzxagent","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8731234669685364},"editors":["dzxagent"],"editorAvatarUrls":["/avatars/3367920f104ca45305944f6bc1d6b8df.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04419","authors":[{"_id":"6a782d898e9301703eaa5c6b","name":"Zikun Qu","hidden":false},{"_id":"6a782d898e9301703eaa5c6c","name":"Min Zhang","hidden":false},{"_id":"6a782d898e9301703eaa5c6d","name":"Mingze Kong","hidden":false},{"_id":"6a782d898e9301703eaa5c6e","name":"Zhiwei Shang","hidden":false},{"_id":"6a782d898e9301703eaa5c6f","name":"Yikun Ban","hidden":false},{"_id":"6a782d898e9301703eaa5c70","name":"Shuang Qiu","hidden":false},{"_id":"6a782d898e9301703eaa5c71","user":{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user","name":"dzxagent"},"name":"Zhongxiang Dai","status":"claimed_verified","statusLastChangedAt":"2026-08-10T08:45:04.281Z","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation","submittedOnDailyBy":{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user","name":"dzxagent"},"summary":"On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-k candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.","upvotes":16,"discussionId":"6a782d898e9301703eaa5c72","githubRepo":"https://github.com/QuZikun/SPOT","githubRepoAddedBy":"user","ai_summary":"SPOT improves on-policy distillation by selectively probing uncertain positions and calibrating targets to downstream outcomes, boosting reasoning quality and coverage.","ai_keywords":["on-policy distillation","reverse-KL","teacher entropy","sparse probing","outcome-calibrated targets","acquisition-exploration-exploitation","verifier-scored continuations","KL-regularized target"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6223644d0129f2097d69a407","name":"CUHKSZ","fullname":"Chinese University of Hong Kong, Shenzhen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1646486592158-6108ae87823007eaf0c7bd1e.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d53b638a15934c10f2c16d","avatarUrl":"/avatars/a8431e4554d68708562f45bc6a8544d2.svg","isPro":false,"fullname":"fafa","user":"kunkun0919","type":"user"},{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user"},{"_id":"67c54394bb902949d6548c5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/bpfbtewGpK5CdP6DZmiPb.png","isPro":false,"fullname":"Zhiwei Shang","user":"javenshang","type":"user"},{"_id":"67d654e797767f4925a48fee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/0X51HpzNzAdcSJnd9k9Oz.png","isPro":false,"fullname":"Xinyi Yuan","user":"xinyyyiii","type":"user"},{"_id":"69bac4dde4325f7946a34143","avatarUrl":"/avatars/20bd4d0afdbafea13c8ab82acadc00fe.svg","isPro":false,"fullname":"s","user":"wsghc","type":"user"},{"_id":"66bf1156eb4f43ee8a525f2f","avatarUrl":"/avatars/3d738d7d1c535a2ef8d402f2b580e9df.svg","isPro":false,"fullname":"Tywh202","user":"Tywh202","type":"user"},{"_id":"65df436f487081eb2c160e57","avatarUrl":"/avatars/dda2d6fc83c988791f1710ee36d1c066.svg","isPro":false,"fullname":"Mingze Kong","user":"EmbPhy","type":"user"},{"_id":"6920035cd48b817fb1297da3","avatarUrl":"/avatars/109a16cb17427cec1248e70ff09ac239.svg","isPro":false,"fullname":"Mingrong Gong","user":"Gmr5233","type":"user"},{"_id":"67c0acd7ec3563f6b93d3f51","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/S3eKPZswOffvGVY6U5HIN.png","isPro":false,"fullname":"zifan wang","user":"aCodeDog","type":"user"},{"_id":"6a7962e9306352b010760546","avatarUrl":"/avatars/2abb83142b69c7d51a299ccc71cccd7a.svg","isPro":false,"fullname":"LAIqp","user":"qpLAI","type":"user"},{"_id":"68673435069b1aa9224b88f1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Fxq0yk1NLb0oAwxodlcuD.png","isPro":false,"fullname":"Junfeng Liao","user":"Junfeng01","type":"user"},{"_id":"68345345f4bbf856e2d708e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68345345f4bbf856e2d708e2/L5H2HNCuWje3ti2tNbC5p.jpeg","isPro":false,"fullname":"Yikun Ban","user":"Yikunb","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6223644d0129f2097d69a407","name":"CUHKSZ","fullname":"Chinese University of Hong Kong, Shenzhen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1646486592158-6108ae87823007eaf0c7bd1e.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04419.md","query":{}}">
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
Abstract
SPOT improves on-policy distillation by selectively probing uncertain positions and calibrating targets to downstream outcomes, boosting reasoning quality and coverage.
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-k candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.
Community
Standard OPD may under-cover plausible alternatives, while EOPD relies solely on teacher entropy and uncalibrated targets. SPOT sparsely probes informative positions and uses verifier-scored continuations to construct outcome-calibrated targets. Across three student scales, SPOT improves macro Avg@8/Pass@8 over OPD by 0.47–1.48/4.55–5.28 points and over EOPD by 0.29–0.68/2.49–3.19 points.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.04419 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.04419 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.04419 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.