Hugging Face Daily Papers · · 6 min read

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

\n\n### Contributions\n\nOur key contributions are as follows:\n\n- **Weak-to-Strong Generalization.** We study how post-training gains from weaker models can be transferred to stronger students in two practical scenarios: successive model transfer and multi-domain consolidation. We identify the central challenge as exploiting these gains without making either the weak policy or its policy shift a separate optimization target.\n\n- **On-Policy Reverse Distillation.** We introduce OPRD, which evaluates a weak teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. Because OPRD only rescales verifier-supported updates, it accelerates the student's own optimization while preserving its stationary points, allowing the student to move beyond the teacher.\n\n- **Empirical Evaluation and Analysis.** Across successive-model and multi-teacher settings, OPRD reaches the final performance of competing methods substantially earlier and ultimately outperforms both RL and distillation baselines. We further confirm that these gains extend to conventional strong-to-weak distillation. We also compare against recent weak-to-strong methods, analyze the design and dynamics of teacher guidance, examine practical challenges and mitigations, and study student reasoning and response style under teacher guidance.","html":"<img src=\"https://cdn-uploads.huggingface.co/production/uploads/6a749bdd331b223584a91a34/qWxMSrWcQFYHTFiXdSwmr.png\" width=\"740\">\n\n<h3 class=\"relative group flex items-baseline\">\n\t<a id=\"contributions\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#contributions\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tContributions\n\t</span>\n</h3>\n<p>Our key contributions are as follows:</p>\n<ul>\n<li><p><strong>Weak-to-Strong Generalization.</strong> We study how post-training gains from weaker models can be transferred to stronger students in two practical scenarios: successive model transfer and multi-domain consolidation. We identify the central challenge as exploiting these gains without making either the weak policy or its policy shift a separate optimization target.</p>\n</li>\n<li><p><strong>On-Policy Reverse Distillation.</strong> We introduce OPRD, which evaluates a weak teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. Because OPRD only rescales verifier-supported updates, it accelerates the student's own optimization while preserving its stationary points, allowing the student to move beyond the teacher.</p>\n</li>\n<li><p><strong>Empirical Evaluation and Analysis.</strong> Across successive-model and multi-teacher settings, OPRD reaches the final performance of competing methods substantially earlier and ultimately outperforms both RL and distillation baselines. We further confirm that these gains extend to conventional strong-to-weak distillation. We also compare against recent weak-to-strong methods, analyze the design and dynamics of teacher guidance, examine practical challenges and mitigations, and study student reasoning and response style under teacher guidance.</p>\n</li>\n</ul>\n","updatedAt":"2026-09-09T04:14:26.700Z","author":{"_id":"6a749bdd331b223584a91a34","avatarUrl":"/avatars/94f7549ad4b2c46497b24b87e5fd5155.svg","fullname":"Sangmin Bae","name":"bsmn0223","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.880849301815033},"editors":["bsmn0223"],"editorAvatarUrls":["/avatars/94f7549ad4b2c46497b24b87e5fd5155.svg"],"reactions":[{"reaction":"👍","users":["HwanChang0106"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08798","authors":[{"_id":"6aa0da83d0174964227bed81","name":"Youngrok Park","hidden":false},{"_id":"6aa0da83d0174964227bed82","name":"Sangmin Bae","hidden":false},{"_id":"6aa0da83d0174964227bed83","name":"Hojung Jung","hidden":false},{"_id":"6aa0da83d0174964227bed84","user":{"_id":"64b7628af902508f0d7ae112","avatarUrl":"/avatars/83c155254486e80c1dfd14676fdf9215.svg","isPro":false,"fullname":"Jongwoo Ko","user":"jongwooko","type":"user","name":"jongwooko"},"name":"Jongwoo Ko","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.682Z","hidden":false},{"_id":"6aa0da83d0174964227bed85","name":"Yunseon Choi","hidden":false},{"_id":"6aa0da83d0174964227bed86","name":"Young Jin Kim","hidden":false},{"_id":"6aa0da83d0174964227bed87","name":"Pashmina Cameron","hidden":false},{"_id":"6aa0da83d0174964227bed88","name":"Aaron Courville","hidden":false},{"_id":"6aa0da83d0174964227bed89","name":"Se-Young Yun","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation","submittedOnDailyBy":{"_id":"6a749bdd331b223584a91a34","avatarUrl":"/avatars/94f7549ad4b2c46497b24b87e5fd5155.svg","isPro":true,"fullname":"Sangmin Bae","user":"bsmn0223","type":"user","name":"bsmn0223"},"summary":"Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.","upvotes":74,"discussionId":"6aa0da83d0174964227bed8a","githubRepo":"https://github.com/raymin0223/on_policy_reverse_distillation","githubRepoAddedBy":"user","ai_summary":"On-Policy Reverse Distillation enables stronger models to exceed weak supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, accelerating optimization without imposing capacity limits.","ai_keywords":["weak-to-strong generalization","On-Policy Reverse Distillation","policy gradient","verifier-driven policy optimization","multi-teacher distillation","RL distillation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"671532b6cf2b9210e588a38f","avatarUrl":"/avatars/40e3a445adcd6c30e1878e9db7ed588b.svg","isPro":true,"fullname":"Youngrok Park","user":"yr-park","type":"user"},{"_id":"6a749bdd331b223584a91a34","avatarUrl":"/avatars/94f7549ad4b2c46497b24b87e5fd5155.svg","isPro":true,"fullname":"Sangmin Bae","user":"bsmn0223","type":"user"},{"_id":"64049a20ad54665351d7d8e2","avatarUrl":"/avatars/db72824521826bb9f3f4d849be4d36df.svg","isPro":false,"fullname":"Sangmin Hwang","user":"BBang3","type":"user"},{"_id":"68e49b401b1f4cba97b91dfd","avatarUrl":"/avatars/994106c561c315867bc8310326ff051a.svg","isPro":false,"fullname":"jungwoo Lee","user":"Jungwoo1031","type":"user"},{"_id":"67cfc505182d970d408a3c77","avatarUrl":"/avatars/fbea63c4e0b3e33259b1205a60342ee2.svg","isPro":false,"fullname":"Junghyun Lee","user":"nick-jhlee","type":"user"},{"_id":"67d823d86d6baf84fba04392","avatarUrl":"/avatars/8dfe7800aa3f9d9166963a1f5da5657a.svg","isPro":false,"fullname":"segyu lee","user":"segyulee","type":"user"},{"_id":"6786410e825069b89ac10a7b","avatarUrl":"/avatars/6746c862a22d62610ec1464f3f3c3cc3.svg","isPro":false,"fullname":"Nam Cao","user":"mrClumsy1207","type":"user"},{"_id":"677e21e781018fcad9872808","avatarUrl":"/avatars/f8dded5e043b93313c71642a7d48654f.svg","isPro":false,"fullname":"Jaehyun Kwak","user":"Jackwaky","type":"user"},{"_id":"61b672df32e9f98551babc32","avatarUrl":"/avatars/bc698c7416801f5ee3101d8a82230472.svg","isPro":false,"fullname":"Changhae Lee","user":"chlee973","type":"user"},{"_id":"63f0c2ac9cf89c9ed1bdd25c","avatarUrl":"/avatars/856b2cb482250fb83c6fe793e29dfd74.svg","isPro":false,"fullname":"Sungnyun Kim","user":"sungnyun","type":"user"},{"_id":"65f836339e3737dc3040f3be","avatarUrl":"/avatars/70b479f58338b71192a05331dfa1bb15.svg","isPro":true,"fullname":"Hojung Jung","user":"cossmoss","type":"user"},{"_id":"6556bb7f66423b57b2dfe749","avatarUrl":"/avatars/c61267ad7d2192380eb30b93d8681421.svg","isPro":false,"fullname":"JihwanOh","user":"ericoh929","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08798.md","query":{}}">
Papers
arxiv:2609.08798

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Published on Sep 8
· Submitted by
Sangmin Bae
on Sep 9
Authors:
,

Abstract

On-Policy Reverse Distillation enables stronger models to exceed weak supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, accelerating optimization without imposing capacity limits.

Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.

Community

Paper submitter about 10 hours ago

Contributions

Our key contributions are as follows:

  • Weak-to-Strong Generalization. We study how post-training gains from weaker models can be transferred to stronger students in two practical scenarios: successive model transfer and multi-domain consolidation. We identify the central challenge as exploiting these gains without making either the weak policy or its policy shift a separate optimization target.

  • On-Policy Reverse Distillation. We introduce OPRD, which evaluates a weak teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. Because OPRD only rescales verifier-supported updates, it accelerates the student's own optimization while preserving its stationary points, allowing the student to move beyond the teacher.

  • Empirical Evaluation and Analysis. Across successive-model and multi-teacher settings, OPRD reaches the final performance of competing methods substantially earlier and ultimately outperforms both RL and distillation baselines. We further confirm that these gains extend to conventional strong-to-weak distillation. We also compare against recent weak-to-strong methods, analyze the design and dynamics of teacher guidance, examine practical challenges and mitigations, and study student reasoning and response style under teacher guidance.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.08798
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.08798 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.08798 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.08798 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers