Hugging Face Daily Papers · · 5 min read

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.</p>\n","updatedAt":"2026-07-14T06:46:08.739Z","author":{"_id":"64673506b990713c503079bc","avatarUrl":"/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg","fullname":"Fu","name":"fudaocheng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"editors":["fudaocheng"],"editorAvatarUrls":["/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg"],"reactions":[{"reaction":"👍","users":["yuyang-cloud"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.11505","authors":[{"_id":"6a55d8b0a9d74d6e65bbd807","name":"Daocheng Fu","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd808","name":"Rong Wu","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd809","name":"Yu Yang","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80a","name":"Xuemeng Yang","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80b","name":"Jianbiao Mei","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80c","name":"Licheng Wen","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80d","name":"Pinlong Cai","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80e","name":"Yong Liu","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80f","name":"Botian Shi","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd810","name":"Yu Qiao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64673506b990713c503079bc/29rDaJ_2_kZvgKS-jjw8z.png"],"publishedAt":"2026-07-13T12:56:21.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals","submittedOnDailyBy":{"_id":"64673506b990713c503079bc","avatarUrl":"/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg","isPro":false,"fullname":"Fu","user":"fudaocheng","type":"user","name":"fudaocheng"},"summary":"Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.","upvotes":7,"discussionId":"6a55d8b1a9d74d6e65bbd811","organization":{"_id":"68e8716fa897e565aaec87ca","name":"KnowledgeXLab","fullname":"KnowledgeXLab@Shanghai AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a7c43ae940d769194055df/oPtAVxSeVl8XHm5mo0Nq8.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64673506b990713c503079bc","avatarUrl":"/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg","isPro":false,"fullname":"Fu","user":"fudaocheng","type":"user"},{"_id":"656882da5025f8e01b424eda","avatarUrl":"/avatars/7fda244634c90b4647b49f83fdb25a9b.svg","isPro":false,"fullname":"yxm","user":"jokester-yxm","type":"user"},{"_id":"66fa9fa56b620c614cdb1b57","avatarUrl":"/avatars/907c101ba4528de05d2ce4bb1ad82005.svg","isPro":false,"fullname":"Rong Wu","user":"Edaizi","type":"user"},{"_id":"66b4c849d3ffe590456020f4","avatarUrl":"/avatars/9eeac7a494ac1070721f75423d1e8482.svg","isPro":false,"fullname":"HR","user":"ZHHHR","type":"user"},{"_id":"67e41575baa288f7f82af73e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67e41575baa288f7f82af73e/0Wx_4SmB6XsjUghQZTRLj.jpeg","isPro":false,"fullname":"YANG YU","user":"yuyang-cloud","type":"user"},{"_id":"64a7c43ae940d769194055df","avatarUrl":"/avatars/441ccadd62e039fb8cb112f138ed917d.svg","isPro":false,"fullname":"Licheng Wen","user":"Wayne-lc","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68e8716fa897e565aaec87ca","name":"KnowledgeXLab","fullname":"KnowledgeXLab@Shanghai AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a7c43ae940d769194055df/oPtAVxSeVl8XHm5mo0Nq8.png"},"query":{}}">
Papers
arxiv:2607.11505

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

Published on Jul 13
· Submitted by
Fu
on Jul 14
Authors:
,

Abstract

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.

Community

Paper submitter about 13 hours ago

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.11505 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.11505 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.11505 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers