Group-relative policy optimization methods for reinforcement learning with verifiable rewards (RLVR) typically use a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation of this design: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. In particular, rollouts with low group success, few correct solutions within a rollout group, tend to exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by uniform clipping.</p>\n<p>To address this issue, we propose \\textit{Group Adaptive Clipping Policy Optimization (GAPO)}, a simple plug-in modification to group-relative policy optimization that adapts the IS clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. Importantly, GAPO requires no reward shaping or objective modification, preserving the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama base models, GAPO consistently improves both pass@1 and pass@k over fixed clipping and advantage-shaping baselines on mathematical reasoning benchmarks.</p>\n","updatedAt":"2026-09-07T01:58:23.269Z","author":{"_id":"66aa6b432795c28eb3a0b66b","avatarUrl":"/avatars/dff15a8279c2365e14a36de588f7e2e3.svg","fullname":"Sheng Jia","name":"shengjia-toronto","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8843257427215576},"editors":["shengjia-toronto"],"editorAvatarUrls":["/avatars/dff15a8279c2365e14a36de588f7e2e3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00444","authors":[{"_id":"6a9e1a2ede5ea82090db653a","name":"Sheng Jia","hidden":false},{"_id":"6a9e1a2ede5ea82090db653b","name":"Xiao Wang","hidden":false},{"_id":"6a9e1a2ede5ea82090db653c","name":"Shiva Prasad Kasiviswanathan","hidden":false},{"_id":"6a9e1a2ede5ea82090db653d","name":"Rein Houthooft","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-07T00:00:00.000Z","title":"Group Adaptive Clipping Policy Optimization","submittedOnDailyBy":{"_id":"66aa6b432795c28eb3a0b66b","avatarUrl":"/avatars/dff15a8279c2365e14a36de588f7e2e3.svg","isPro":true,"fullname":"Sheng Jia","user":"shengjia-toronto","type":"user","name":"shengjia-toronto"},"summary":"Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. Rollouts with low group success exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by fixed clipping.\n To address this, we propose Group Adaptive Clipping Policy Optimization (GAPO), a plug-in modification to GRPO methods that adapts the clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. GAPO requires no reward shaping and preserves the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama models, GAPO consistently improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math reasoning and coding benchmarks where the pass rates by the base model are relatively low.","upvotes":3,"discussionId":"6a9e1a2ede5ea82090db653e","githubRepo":"https://github.com/Sheng-J/GAPO","githubRepoAddedBy":"user","ai_summary":"GAPO adaptively adjusts importance-sampling clipping thresholds based on rollout advantage to preserve stronger gradient signals from low-success groups in reinforcement learning with verifiable rewards.","ai_keywords":["Group Adaptive Clipping Policy Optimization","GRPO","importance-sampling ratio clipping","rollout advantage","reverse-KL trust-region","PPO surrogate","reinforcement learning with verifiable rewards"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"5ffdfbadbba2ae614d771970","name":"amazon","fullname":"Amazon","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f19ed428ae41c20c470792/8y7msN6A6W82LdQhQd85a.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66aa6b432795c28eb3a0b66b","avatarUrl":"/avatars/dff15a8279c2365e14a36de588f7e2e3.svg","isPro":true,"fullname":"Sheng Jia","user":"shengjia-toronto","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"68b2a4157f881fc640ba7d80","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lMTgr3pe7pOHtMe7bVF7F.png","isPro":false,"fullname":"khtsly","user":"khtsly","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"5ffdfbadbba2ae614d771970","name":"amazon","fullname":"Amazon","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f19ed428ae41c20c470792/8y7msN6A6W82LdQhQd85a.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00444.md","query":{}}">
Group Adaptive Clipping Policy Optimization
Abstract
GAPO adaptively adjusts importance-sampling clipping thresholds based on rollout advantage to preserve stronger gradient signals from low-success groups in reinforcement learning with verifiable rewards.
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. Rollouts with low group success exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by fixed clipping.
To address this, we propose Group Adaptive Clipping Policy Optimization (GAPO), a plug-in modification to GRPO methods that adapts the clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. GAPO requires no reward shaping and preserves the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama models, GAPO consistently improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math reasoning and coding benchmarks where the pass rates by the base model are relatively low.
Community
Group-relative policy optimization methods for reinforcement learning with verifiable rewards (RLVR) typically use a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation of this design: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. In particular, rollouts with low group success, few correct solutions within a rollout group, tend to exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by uniform clipping.
To address this issue, we propose \textit{Group Adaptive Clipping Policy Optimization (GAPO)}, a simple plug-in modification to group-relative policy optimization that adapts the IS clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. Importantly, GAPO requires no reward shaping or objective modification, preserving the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama base models, GAPO consistently improves both pass@1 and pass@k over fixed clipping and advantage-shaping baselines on mathematical reasoning benchmarks.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.00444 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.00444 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.00444 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.