Hugging Face Daily Papers · · 2 min read

On-Policy Delta Distillation for Multilingual Math Reasoning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Extending On-Policy Delta Distillation (OPD²) to multilingual reasoning</p>\n","updatedAt":"2026-08-07T02:01:44.293Z","author":{"_id":"64b9feed96676e40d0fa89a7","avatarUrl":"/avatars/2154a6ceb87677ad2c9d9620de5b18ec.svg","fullname":"Byeongho Heo","name":"bhheo","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6512726545333862},"editors":["bhheo"],"editorAvatarUrls":["/avatars/2154a6ceb87677ad2c9d9620de5b18ec.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05802","authors":[{"_id":"6a753c19e1228e04b323816b","user":{"_id":"64b9feed96676e40d0fa89a7","avatarUrl":"/avatars/2154a6ceb87677ad2c9d9620de5b18ec.svg","isPro":false,"fullname":"Byeongho Heo","user":"bhheo","type":"user","name":"bhheo"},"name":"Byeongho Heo","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.497Z","hidden":false},{"_id":"6a753c19e1228e04b323816c","name":"Jaehui Hwang","hidden":false},{"_id":"6a753c19e1228e04b323816d","name":"Sangdoo Yun","hidden":false},{"_id":"6a753c19e1228e04b323816e","name":"Dongyoon Han","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"On-Policy Delta Distillation for Multilingual Math Reasoning","submittedOnDailyBy":{"_id":"64b9feed96676e40d0fa89a7","avatarUrl":"/avatars/2154a6ceb87677ad2c9d9620de5b18ec.svg","isPro":false,"fullname":"Byeongho Heo","user":"bhheo","type":"user","name":"bhheo"},"summary":"On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.","upvotes":20,"discussionId":"6a753c1ae1228e04b323816f","organization":{"_id":"64ffe603efd273eec7768bde","name":"naver-ai","fullname":"NAVER AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ff1755b75685dd7a46e146/Zj2bxgq31oSqwVrw16IE_.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b9feed96676e40d0fa89a7","avatarUrl":"/avatars/2154a6ceb87677ad2c9d9620de5b18ec.svg","isPro":false,"fullname":"Byeongho Heo","user":"bhheo","type":"user"},{"_id":"6a6a91ad4e0203edaf1f7ae6","avatarUrl":"/avatars/764cfb5bc47dbfca6d720260a64cfe16.svg","isPro":false,"fullname":"Patricia Rodriguez","user":"patricia-rodriguez","type":"user"},{"_id":"6a6a936a4a49113ba435f3d1","avatarUrl":"/avatars/b2523cc97cb8e1532377c6eae4ff1ca8.svg","isPro":false,"fullname":"Patricia Thompson","user":"Patricia-Thompson","type":"user"},{"_id":"6a6a9c3dbbca071c7189c742","avatarUrl":"/avatars/1420f2ad85cb529e3ad9ce5becef48d2.svg","isPro":false,"fullname":"Richard Anderson","user":"Quiet-Eli","type":"user"},{"_id":"6a6aa3723eb4da20611a1610","avatarUrl":"/avatars/5c290d2a94ef9b1a60b49f06f0c97fc9.svg","isPro":false,"fullname":"Timothy Harris","user":"Lunar-BloomW","type":"user"},{"_id":"6a6c7c751964ab520526857c","avatarUrl":"/avatars/0e4c180ab58ef319875181a57bf82239.svg","isPro":false,"fullname":"William Wilson","user":"SableWilliam","type":"user"},{"_id":"6a6c842cac918664e64e490d","avatarUrl":"/avatars/b574b8e8ef7f19cb0d5fc52c69f0bd86.svg","isPro":false,"fullname":"Jessica Davis","user":"Atlas-Bloom","type":"user"},{"_id":"6a6c869c8330caf6ccf1be14","avatarUrl":"/avatars/1100359e86699888d742e0ab17746953.svg","isPro":false,"fullname":"Matthew Miller","user":"MatthewMiller","type":"user"},{"_id":"6a6c8d382fbbe664ac698737","avatarUrl":"/avatars/c9842309a88cb9944b66bf5b06814980.svg","isPro":false,"fullname":"George Gonzalez","user":"zenithbloom","type":"user"},{"_id":"6a6aa1a9719ef934225fa075","avatarUrl":"/avatars/c7528de1b34e4f12e112800214751369.svg","isPro":false,"fullname":"Matthew Thomas","user":"orbitMatthew","type":"user"},{"_id":"6a6da7e008e6705013bc9e63","avatarUrl":"/avatars/c3d8bc73df06931ba549417b8af48e6b.svg","isPro":false,"fullname":"Linda Gonzalez","user":"NimbusLens","type":"user"},{"_id":"6a6dc6d49caa2ff7db786e2a","avatarUrl":"/avatars/efdc260d9bb75fcf331f574b1d1efbe6.svg","isPro":false,"fullname":"Jessica Gonzalez","user":"Azure-Rune","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64ffe603efd273eec7768bde","name":"naver-ai","fullname":"NAVER AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ff1755b75685dd7a46e146/Zj2bxgq31oSqwVrw16IE_.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05802.md","query":{}}">
Papers
arxiv:2608.05802

On-Policy Delta Distillation for Multilingual Math Reasoning

Published on Aug 6
· Submitted by
Byeongho Heo
on Aug 7
Authors:

Abstract

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

Community

Paper author Paper submitter about 15 hours ago

Extending On-Policy Delta Distillation (OPD²) to multilingual reasoning

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.05802
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.05802 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.05802 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.05802 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers