Hugging Face Daily Papers · · 3 min read

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Blog Page: <a href=\"https://jyyang26.github.io/stable_async_analysis\" rel=\"nofollow\">https://jyyang26.github.io/stable_async_analysis</a><br>arXiv: <a href=\"https://arxiv.org/abs/2607.18722\" rel=\"nofollow\">https://arxiv.org/abs/2607.18722</a><br>GitHub: <a href=\"https://github.com/jyyang26/SAT\" rel=\"nofollow\">https://github.com/jyyang26/SAT</a></p>\n","updatedAt":"2026-07-22T03:18:22.793Z","author":{"_id":"642447e873f7a0d40b30d677","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642447e873f7a0d40b30d677/hbUnfIhKZCeSSPv9wGAzk.png","fullname":"LZX","name":"zli12321","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.43244683742523193},"editors":["zli12321"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642447e873f7a0d40b30d677/hbUnfIhKZCeSSPv9wGAzk.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18722","authors":[{"_id":"6a602f0e7e7f152167e470b3","user":{"_id":"66926e96f658d6ee95ef6d9d","avatarUrl":"/avatars/cb523e7ebaa5e80bb9cb33715ec1e6e1.svg","isPro":false,"fullname":"Junyao Yang","user":"TberiusJunyao","type":"user","name":"TberiusJunyao"},"name":"Junyao Yang","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.299Z","hidden":false},{"_id":"6a602f0e7e7f152167e470b4","user":{"_id":"64beb6b6140491ca9f803ebf","avatarUrl":"/avatars/0daa2e813a13668b8b708cd8c12763d9.svg","isPro":false,"fullname":"Yucheng SHi","user":"YuchengShi","type":"user","name":"YuchengShi"},"name":"Yucheng Shi","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.306Z","hidden":false},{"_id":"6a602f0e7e7f152167e470b5","user":{"_id":"642447e873f7a0d40b30d677","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642447e873f7a0d40b30d677/hbUnfIhKZCeSSPv9wGAzk.png","isPro":false,"fullname":"LZX","user":"zli12321","type":"user","name":"zli12321"},"name":"Zongxia Li","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:40:38.735Z","hidden":false},{"_id":"6a602f0e7e7f152167e470b6","name":"Zhongzhi Li","hidden":false},{"_id":"6a602f0e7e7f152167e470b7","name":"Ruhan Wang","hidden":false},{"_id":"6a602f0e7e7f152167e470b8","user":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","isPro":false,"fullname":"Xiangxin Zhou","user":"zhouxiangxin","type":"user","name":"zhouxiangxin"},"name":"Xiangxin Zhou","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.314Z","hidden":false},{"_id":"6a602f0e7e7f152167e470b9","name":"Kishan Panaganti","hidden":false},{"_id":"6a602f0e7e7f152167e470ba","name":"Haitao Mi","hidden":false},{"_id":"6a602f0e7e7f152167e470bb","name":"Leowei Liang","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning","submittedOnDailyBy":{"_id":"642447e873f7a0d40b30d677","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642447e873f7a0d40b30d677/hbUnfIhKZCeSSPv9wGAzk.png","isPro":false,"fullname":"LZX","user":"zli12321","type":"user","name":"zli12321"},"summary":"Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a result, high-staleness updates remain weakly controlled in the asynchronous regime where stale rollouts matter most.\n We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies high-mismatch tails within each batch via staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. This preserves baseline behavior on ordinary tokens while enforcing more conservative updates on newly intercepted outward bands. We prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness.\n We evaluate SAT in a decoupled asynchronous RL setup built on Qwen3-30B-A3B-Base, using SGLang as the inference engine and Megatron for training. In this setting, SAT-GSPO w/ R3 achieves the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. Adaptive clipping and routing replay act as complementary stabilizers targeting mismatch tails and routing inconsistency, respectively. Overall, aligning clip intervals with staleness heterogeneity effectively stabilizes asynchronous RL.","upvotes":24,"discussionId":"6a602f0f7e7f152167e470bc","projectPage":"https://jyyang26.github.io/stable_async_analysis/","githubRepo":"https://github.com/jyyang26/SAT","githubRepoAddedBy":"user","githubStars":9,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66926e96f658d6ee95ef6d9d","avatarUrl":"/avatars/cb523e7ebaa5e80bb9cb33715ec1e6e1.svg","isPro":false,"fullname":"Junyao Yang","user":"TberiusJunyao","type":"user"},{"_id":"64beb6b6140491ca9f803ebf","avatarUrl":"/avatars/0daa2e813a13668b8b708cd8c12763d9.svg","isPro":false,"fullname":"Yucheng SHi","user":"YuchengShi","type":"user"},{"_id":"6723e878dce5fb21b5a96bde","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/zXUmFNtamuwwjbz6FEEc2.png","isPro":false,"fullname":"Ruhan Wang","user":"ruhwang","type":"user"},{"_id":"642447e873f7a0d40b30d677","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642447e873f7a0d40b30d677/hbUnfIhKZCeSSPv9wGAzk.png","isPro":false,"fullname":"LZX","user":"zli12321","type":"user"},{"_id":"654a238a3b78e73b439afb7c","avatarUrl":"/avatars/ee582bd50ca389bece1baff29f08cc78.svg","isPro":false,"fullname":"pangpangxuan","user":"pangxuan","type":"user"},{"_id":"68ee1311c53422e9855b20e9","avatarUrl":"/avatars/38008cf7b90c4ed2bed4d08a09819c33.svg","isPro":false,"fullname":"Nan Lu","user":"neoluxx","type":"user"},{"_id":"67679b5cfeac1e9f62571cf9","avatarUrl":"/avatars/4b0a0348dd0bf871aa40f8ff37703efa.svg","isPro":false,"fullname":"Zhuowen Liang","user":"SetonLiang2","type":"user"},{"_id":"640f0409c025ddf618963046","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/640f0409c025ddf618963046/ZBsLKfS6S5tBJYRPNhOTs.png","isPro":false,"fullname":"BorisGuo","user":"BorisGuo","type":"user"},{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","isPro":false,"fullname":"Xiangxin Zhou","user":"zhouxiangxin","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6672f7c82376de4b7ab9fbd5","avatarUrl":"/avatars/60fe64dbe4c814fd8b1df7ce2bebc951.svg","isPro":true,"fullname":"Guo","user":"Rongjin03","type":"user"},{"_id":"6214e4ee1e35c843d42d1f88","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6214e4ee1e35c843d42d1f88/fj-9wuIdPhvogh3BrcXTB.jpeg","isPro":false,"fullname":"Longxu Dou","user":"dreamerdeo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"},"query":{}}">
Papers
arxiv:2607.18722

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Published on Jul 21
· Submitted by
LZX
on Jul 22

Abstract

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a result, high-staleness updates remain weakly controlled in the asynchronous regime where stale rollouts matter most. We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies high-mismatch tails within each batch via staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. This preserves baseline behavior on ordinary tokens while enforcing more conservative updates on newly intercepted outward bands. We prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness. We evaluate SAT in a decoupled asynchronous RL setup built on Qwen3-30B-A3B-Base, using SGLang as the inference engine and Megatron for training. In this setting, SAT-GSPO w/ R3 achieves the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. Adaptive clipping and routing replay act as complementary stabilizers targeting mismatch tails and routing inconsistency, respectively. Overall, aligning clip intervals with staleness heterogeneity effectively stabilizes asynchronous RL.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.18722 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.18722 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.18722 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers