Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.</p>\n","updatedAt":"2026-07-14T06:46:08.739Z","author":{"_id":"64673506b990713c503079bc","avatarUrl":"/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg","fullname":"Fu","name":"fudaocheng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"editors":["fudaocheng"],"editorAvatarUrls":["/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg"],"reactions":[{"reaction":"👍","users":["yuyang-cloud"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.11505","authors":[{"_id":"6a55d8b0a9d74d6e65bbd807","name":"Daocheng Fu","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd808","name":"Rong Wu","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd809","name":"Yu Yang","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80a","name":"Xuemeng Yang","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80b","name":"Jianbiao Mei","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80c","name":"Licheng Wen","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80d","name":"Pinlong Cai","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80e","name":"Yong Liu","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd80f","name":"Botian Shi","hidden":false},{"_id":"6a55d8b0a9d74d6e65bbd810","name":"Yu Qiao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64673506b990713c503079bc/29rDaJ_2_kZvgKS-jjw8z.png"],"publishedAt":"2026-07-13T12:56:21.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals","submittedOnDailyBy":{"_id":"64673506b990713c503079bc","avatarUrl":"/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg","isPro":false,"fullname":"Fu","user":"fudaocheng","type":"user","name":"fudaocheng"},"summary":"Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.","upvotes":7,"discussionId":"6a55d8b1a9d74d6e65bbd811","organization":{"_id":"68e8716fa897e565aaec87ca","name":"KnowledgeXLab","fullname":"KnowledgeXLab@Shanghai AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a7c43ae940d769194055df/oPtAVxSeVl8XHm5mo0Nq8.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64673506b990713c503079bc","avatarUrl":"/avatars/17396c835f1eeae00ec2972b6a24a2dd.svg","isPro":false,"fullname":"Fu","user":"fudaocheng","type":"user"},{"_id":"656882da5025f8e01b424eda","avatarUrl":"/avatars/7fda244634c90b4647b49f83fdb25a9b.svg","isPro":false,"fullname":"yxm","user":"jokester-yxm","type":"user"},{"_id":"66fa9fa56b620c614cdb1b57","avatarUrl":"/avatars/907c101ba4528de05d2ce4bb1ad82005.svg","isPro":false,"fullname":"Rong Wu","user":"Edaizi","type":"user"},{"_id":"66b4c849d3ffe590456020f4","avatarUrl":"/avatars/9eeac7a494ac1070721f75423d1e8482.svg","isPro":false,"fullname":"HR","user":"ZHHHR","type":"user"},{"_id":"67e41575baa288f7f82af73e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67e41575baa288f7f82af73e/0Wx_4SmB6XsjUghQZTRLj.jpeg","isPro":false,"fullname":"YANG YU","user":"yuyang-cloud","type":"user"},{"_id":"64a7c43ae940d769194055df","avatarUrl":"/avatars/441ccadd62e039fb8cb112f138ed917d.svg","isPro":false,"fullname":"Licheng Wen","user":"Wayne-lc","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68e8716fa897e565aaec87ca","name":"KnowledgeXLab","fullname":"KnowledgeXLab@Shanghai AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a7c43ae940d769194055df/oPtAVxSeVl8XHm5mo0Nq8.png"},"query":{}}">
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Published on Jul 13
· Submitted by Fu on Jul 14 Abstract
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.
Community
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.11505 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.11505 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.11505 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.