Hugging Face Daily Papers · June 10, 2026 · 4 min read

Rethinking the Divergence Regularization in LLM RL

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Like Read original ↗

Divergence Regularized Policy Optimization (DRPO) is our attempt to make mask-based trust regions smoother without losing what makes them effective.\nOur key insight is that trust-region geometry matters: DPPO-style trust-region geometry works well for LLM token updates, but its hard mask can be brittle.\nDRPO keeps this geometry and replaces the hard cutoff with a smooth, advantage-weighted regularizer, turning discarded gradients into continuous corrective signals for more stable RL training.\n","updatedAt":"2026-06-10T04:26:50.534Z","author":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","fullname":"Xiangxin Zhou","name":"zhouxiangxin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":3,"identifiedLanguage":{"language":"en","probability":0.46447476744651794},"editors":["zhouxiangxin"],"editorAvatarUrls":["/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg"],"reactions":[{"reaction":"❤️","users":["sumailmao"],"count":1}],"isReport":false}},{"id":"6a28e870d09225daee42d324","author":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","fullname":"Xiangxin Zhou","name":"zhouxiangxin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false},"createdAt":"2026-06-10T04:30:40.000Z","type":"comment","data":{"edited":true,"hidden":false,"latest":{"raw":"\n![rollout_prob_hist_cdf](https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/RRto_5MmsxeVokP96Bzj6.png)","html":"<a href=\"https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/RRto_5MmsxeVokP96Bzj6.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/RRto_5MmsxeVokP96Bzj6.png\" alt=\"rollout_prob_hist_cdf\"></a>\n","updatedAt":"2026-06-10T04:30:51.120Z","author":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","fullname":"Xiangxin Zhou","name":"zhouxiangxin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.4681173264980316},"editors":["zhouxiangxin"],"editorAvatarUrls":["/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg"],"reactions":[],"isReport":false}},{"id":"6a28e883c4b98b9925426337","author":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","fullname":"Xiangxin Zhou","name":"zhouxiangxin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false},"createdAt":"2026-06-10T04:30:59.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\n![main_accuracy](https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/88Fk1bLDEWv5F0k4bhXab.png)\n","html":"<a href=\"https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/88Fk1bLDEWv5F0k4bhXab.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/88Fk1bLDEWv5F0k4bhXab.png\" alt=\"main_accuracy\"></a>\n","updatedAt":"2026-06-10T04:30:59.950Z","author":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","fullname":"Xiangxin Zhou","name":"zhouxiangxin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.44497916102409363},"editors":["zhouxiangxin"],"editorAvatarUrls":["/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.09821","authors":[{"_id":"6a28d9c8e7d78ea7587e54a7","name":"Jiarui Yao","hidden":false},{"_id":"6a28d9c8e7d78ea7587e54a8","name":"Xiangxin Zhou","hidden":false},{"_id":"6a28d9c8e7d78ea7587e54a9","name":"Penghui Qi","hidden":false},{"_id":"6a28d9c8e7d78ea7587e54aa","name":"Wee Sun Lee","hidden":false},{"_id":"6a28d9c8e7d78ea7587e54ab","name":"Liefeng Bo","hidden":false},{"_id":"6a28d9c8e7d78ea7587e54ac","name":"Tianyu Pang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66129c7b50350afe76757262/pC9zoJlLEJJnoE8dzbU0e.png"],"publishedAt":"2026-06-08T00:00:00.000Z","submittedOnDailyAt":"2026-06-10T00:00:00.000Z","title":"Rethinking the Divergence Regularization in LLM RL","submittedOnDailyBy":{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","isPro":false,"fullname":"Xiangxin Zhou","user":"zhouxiangxin","type":"user","name":"zhouxiangxin"},"summary":"Reinforcement learning (RL) has become a key component of post-training large language models (LLMs). In practice, LLM RL is often off-policy because of training-inference mismatch and policy staleness, making trust-region control essential for stable optimization. Mainstream methods such as PPO and GRPO approximate this control with a ratio-clipping mechanism, but the importance ratio can be a poor proxy for distributional shift in long-tailed vocabularies. Recent work such as DPPO addresses this mismatch by replacing ratio-based clipping with a divergence-based mask, yielding a trust region defined by the sampled token's absolute probability shift. However, DPPO still relies on a hard mask: once a token crosses the trust-region boundary in a harmful direction, its gradient is discarded rather than corrected. To address this, we propose Divergence Regularized Policy Optimization (DRPO), which replaces the hard mask with a smooth advantage-weighted quadratic regularizer on policy shift. DRPO preserves the same trust-region geometry as DPPO while inducing bounded, continuous gradient weights that attenuate diverging updates and provide corrective signals beyond the boundary. Experiments across model scales, architectures, and precision settings show that DRPO improves the stability and efficiency of LLM RL training.","upvotes":26,"discussionId":"6a28d9c9e7d78ea7587e54ad","githubRepo":"https://github.com/Tencent-Hunyuan/UniRL","githubRepoAddedBy":"user","ai_summary":"DRPO improves LLM reinforcement learning stability by replacing hard masks with smooth regularization that provides continuous gradient corrections beyond trust-region boundaries.","ai_keywords":["reinforcement learning","large language models","off-policy","trust-region control","PPO","GRPO","ratio-clipping","importance ratio","DPPO","divergence-based mask","policy shift","advantage-weighted quadratic regularizer"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":408,"organization":{"_id":"6a24e31c749f04abbbb5105d","name":"Tencent-Hunyuan-Multimodal-RL","fullname":"Tencent-Hunyuan-Multimodal-RL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66129c7b50350afe76757262/AMk8iMwAaEjA0M5Cdq0NE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66129c7b50350afe76757262","avatarUrl":"/avatars/a2f4fac076b9d658a0d904ed54960f6f.svg","isPro":false,"fullname":"Xiangxin Zhou","user":"zhouxiangxin","type":"user"},{"_id":"63885f1d0bebb233d8ad6e5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1669881620925-noauth.jpeg","isPro":false,"fullname":"Penghui Qi","user":"QPHutu","type":"user"},{"_id":"63d91b6d255ef6add20e1b38","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1675921369867-63d91b6d255ef6add20e1b38.jpeg","isPro":false,"fullname":"Tianyu Pang","user":"P2333","type":"user"},{"_id":"66f8689725464a7989b75845","avatarUrl":"/avatars/43a61a528c5779103eaf5687ba44ee14.svg","isPro":false,"fullname":"Jiarui Yao","user":"FlippyDora","type":"user"},{"_id":"649369b34f0e40ee1a0ed5ba","avatarUrl":"/avatars/50d0e77883579d5002906c8d29c26ec5.svg","isPro":false,"fullname":"Maxwell Yao","user":"MaxwellJryao","type":"user"},{"_id":"644fe6a9e1d7a97f3b66e906","avatarUrl":"/avatars/ad1a45f0b1c8a4d03ba87f2a3ce5a8f8.svg","isPro":false,"fullname":"Yuanming-Li","user":"Lymann","type":"user"},{"_id":"65706b2b670035a6071c5f53","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/xpQVkP9Se4SISq-6VXE8G.png","isPro":false,"fullname":"Lazy Beaver","user":"Jayce-Ping","type":"user"},{"_id":"642e7a12ccdcf5da7f9657a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642e7a12ccdcf5da7f9657a0/w8jW5BagTuTp6EvC6KEyR.png","isPro":true,"fullname":"Jiaqi Tang","user":"Jiaqi-hkust","type":"user"},{"_id":"6486b09e8315b19342f0bf5e","avatarUrl":"/avatars/bc5f22f231c884146d373fe1042d81bd.svg","isPro":false,"fullname":"Xiangyan Liu","user":"xyliu6","type":"user"},{"_id":"6541fa406be058da06580347","avatarUrl":"/avatars/fc217146d5ec611bd7f4fb355a5939b3.svg","isPro":false,"fullname":"Wu","user":"Hai-Tao","type":"user"},{"_id":"6710815a07325c4b0ad7b6d4","avatarUrl":"/avatars/a6012a4ee9bdea862cea28ada5e506a0.svg","isPro":true,"fullname":"He Guangxin","user":"gxhe","type":"user"},{"_id":"64b76c8453d91a364aae131f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b76c8453d91a364aae131f/4fHfPJ8QT8zVssBAnDPQ1.png","isPro":false,"fullname":"Lvfang Tao","user":"MeowFET","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a24e31c749f04abbbb5105d","name":"Tencent-Hunyuan-Multimodal-RL","fullname":"Tencent-Hunyuan-Multimodal-RL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66129c7b50350afe76757262/AMk8iMwAaEjA0M5Cdq0NE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.09821.md"}">

Papers

arxiv:2606.09821

Rethinking the Divergence Regularization in LLM RL

Published on Jun 8

· Submitted by

Xiangxin Zhou on Jun 10

Tencent-Hunyuan-Multimodal-RL

Upvote

Authors:

Abstract

DRPO improves LLM reinforcement learning stability by replacing hard masks with smooth regularization that provides continuous gradient corrections beyond trust-region boundaries.

Generated by Qwen/Qwen2.5-Coder-32B-Instruct

Reinforcement learning (RL) has become a key component of post-training large language models (LLMs). In practice, LLM RL is often off-policy because of training-inference mismatch and policy staleness, making trust-region control essential for stable optimization. Mainstream methods such as PPO and GRPO approximate this control with a ratio-clipping mechanism, but the importance ratio can be a poor proxy for distributional shift in long-tailed vocabularies. Recent work such as DPPO addresses this mismatch by replacing ratio-based clipping with a divergence-based mask, yielding a trust region defined by the sampled token's absolute probability shift. However, DPPO still relies on a hard mask: once a token crosses the trust-region boundary in a harmful direction, its gradient is discarded rather than corrected. To address this, we propose Divergence Regularized Policy Optimization (DRPO), which replaces the hard mask with a smooth advantage-weighted quadratic regularizer on policy shift. DRPO preserves the same trust-region geometry as DPPO while inducing bounded, continuous gradient weights that attenuate diverging updates and provide corrective signals beyond the boundary. Experiments across model scales, architectures, and precision settings show that DRPO improves the stability and efficiency of LLM RL training.

View arXiv page View PDF GitHub 408 Add to collection

Community

zhouxiangxin

Paper submitter about 13 hours ago

•

edited about 13 hours ago

Divergence Regularized Policy Optimization (DRPO) is our attempt to make mask-based trust regions smoother without losing what makes them effective.

Our key insight is that trust-region geometry matters: DPPO-style trust-region geometry works well for LLM token updates, but its hard mask can be brittle.

DRPO keeps this geometry and replaces the hard cutoff with a smooth, advantage-weighted regularizer, turning discarded gradients into continuous corrective signals for more stable RL training.