Page: <a href=\"https://dodojordi.github.io/SP3O/\" rel=\"nofollow\">https://dodojordi.github.io/SP3O/</a><br>Repo: <a href=\"https://github.com/Dodojordi/SP3O\" rel=\"nofollow\">https://github.com/Dodojordi/SP3O</a></p>\n","updatedAt":"2026-09-17T06:34:52.657Z","author":{"_id":"644915c5e87a77e872e61350","avatarUrl":"/avatars/46ba7bdf04ad4c1b0ad79155010dc684.svg","fullname":"Luo","name":"ramiroluo","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4368331730365753},"editors":["ramiroluo"],"editorAvatarUrls":["/avatars/46ba7bdf04ad4c1b0ad79155010dc684.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.18708","authors":[{"_id":"6aab88ba1d9cc4dec79625a4","user":{"_id":"6785e04ce63a669874f166e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6785e04ce63a669874f166e9/xSUEAmRm-KadDNOYiOZew.jpeg","isPro":false,"fullname":"dodojorid","user":"yizhuoli","type":"user","name":"yizhuoli"},"name":"Yizhuo Li","status":"claimed_verified","statusLastChangedAt":"2026-09-17T09:09:42.043Z","hidden":false},{"_id":"6aab88ba1d9cc4dec79625a5","name":"Jianhao Yan","hidden":false},{"_id":"6aab88ba1d9cc4dec79625a6","name":"Yun Luo","hidden":false},{"_id":"6aab88ba1d9cc4dec79625a7","name":"Zhi Wang","hidden":false},{"_id":"6aab88ba1d9cc4dec79625a8","name":"Futing Wang","hidden":false},{"_id":"6aab88ba1d9cc4dec79625a9","name":"Rong-Xi Tan","hidden":false},{"_id":"6aab88ba1d9cc4dec79625aa","name":"Kanghui Tian","hidden":false},{"_id":"6aab88ba1d9cc4dec79625ab","name":"Ganqu Cui","hidden":false},{"_id":"6aab88ba1d9cc4dec79625ac","name":"Ning Ding","hidden":false},{"_id":"6aab88ba1d9cc4dec79625ad","name":"Peilin Zhao","hidden":false},{"_id":"6aab88ba1d9cc4dec79625ae","name":"Yafu Li","hidden":false},{"_id":"6aab88ba1d9cc4dec79625af","name":"Yu Cheng","hidden":false}],"publishedAt":"2026-09-16T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening","submittedOnDailyBy":{"_id":"644915c5e87a77e872e61350","avatarUrl":"/avatars/46ba7bdf04ad4c1b0ad79155010dc684.svg","isPro":false,"fullname":"Luo","user":"ramiroluo","type":"user","name":"ramiroluo"},"summary":"In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empirical analyses relate Value Flattening to an implicit variance penalty in the critic loss and redundant updates from temporally correlated states with similar gradients. Motivated by these findings, we introduce SParse Proximal Policy Optimization (SP^3O), which applies the value loss to only a few well-separated states in each response to mitigate both effects. Experiments on Qwen3-Base show that SP^3O with only three states supervised per response can mitigate Value Flattening and consistently improve the learned policy across model sizes and evaluation suites. Together, our results identify Value Flattening as an important yet overlooked failure mode of critic learning in standard PPO and show that a simple sparse supervision strategy can mitigate it.","upvotes":55,"discussionId":"6aab88ba1d9cc4dec79625b0","projectPage":"https://dodojordi.github.io/SP3O/","githubRepo":"https://github.com/Dodojordi/SP3O","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"6a4fb75a1c66dbf208e7ddb6","name":"Shanghai-AI-Laboratory","fullname":"Shanghai AI Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65cd955637be1841d0b75397/Rao_Kq6NMtTVfSqLUIR4k.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"644915c5e87a77e872e61350","avatarUrl":"/avatars/46ba7bdf04ad4c1b0ad79155010dc684.svg","isPro":false,"fullname":"Luo","user":"ramiroluo","type":"user"},{"_id":"6785e04ce63a669874f166e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6785e04ce63a669874f166e9/xSUEAmRm-KadDNOYiOZew.jpeg","isPro":false,"fullname":"dodojorid","user":"yizhuoli","type":"user"},{"_id":"6086838b19137b3a6ba760e7","avatarUrl":"/avatars/d63eea3e39b22c6e65b82c28192696f1.svg","isPro":false,"fullname":"Jianhao Yan","user":"Elliott","type":"user"},{"_id":"65141bfb5f99d14097bf72a7","avatarUrl":"/avatars/8497810edf56a3928dc9233c36fc74d7.svg","isPro":false,"fullname":"Han Cui","user":"hancui","type":"user"},{"_id":"6aab8f291f88dc502db53fb0","avatarUrl":"/avatars/8815d9b29cce76b1a8666e1ae32bdb1c.svg","isPro":false,"fullname":"WZY","user":"Wangzysiuuuu","type":"user"},{"_id":"6a69ec931b577a27e51d8379","avatarUrl":"/avatars/d2226010b481746cf7609a641aea8b77.svg","isPro":false,"fullname":"David Anderson","user":"david-anderson-research","type":"user"},{"_id":"6a6a9a25a5b9c4c08baf22cb","avatarUrl":"/avatars/d3dbbe26cea0e3549d6051687e82fe4b.svg","isPro":false,"fullname":"Daniel Brown","user":"cobalttrail","type":"user"},{"_id":"6a6c7a87101ebc51fc3d89c0","avatarUrl":"/avatars/0248ad3542097101c81488b27faf7a63.svg","isPro":false,"fullname":"Mary Perez","user":"MaryPerez","type":"user"},{"_id":"6a6c83e57dcdd359f72cd2b7","avatarUrl":"/avatars/a1c338f5fca9f2121c663c41758e10ed.svg","isPro":false,"fullname":"Jessica Brown","user":"Indigo-Jessica","type":"user"},{"_id":"6a6d43cae1c088f52420afb3","avatarUrl":"/avatars/90870bb92e6c143ba173694c33629187.svg","isPro":false,"fullname":"Edward Williams","user":"Orbit-Noah","type":"user"},{"_id":"6a6da8267ad403b19cd12072","avatarUrl":"/avatars/f2ce8e6ace279a88cc4082f30a9a2b8e.svg","isPro":false,"fullname":"Joshua Clark","user":"AtlasDawn","type":"user"},{"_id":"6a6dca4ec1e23c5ff6a36aba","avatarUrl":"/avatars/71ed0a47e7c665ce9820aaa414cd2ff9.svg","isPro":false,"fullname":"Jennifer Lopez","user":"cobaltRemy","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6a4fb75a1c66dbf208e7ddb6","name":"Shanghai-AI-Laboratory","fullname":"Shanghai AI Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65cd955637be1841d0b75397/Rao_Kq6NMtTVfSqLUIR4k.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.18708.md","query":{}}">
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
Abstract
In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empirical analyses relate Value Flattening to an implicit variance penalty in the critic loss and redundant updates from temporally correlated states with similar gradients. Motivated by these findings, we introduce SParse Proximal Policy Optimization (SP^3O), which applies the value loss to only a few well-separated states in each response to mitigate both effects. Experiments on Qwen3-Base show that SP^3O with only three states supervised per response can mitigate Value Flattening and consistently improve the learned policy across model sizes and evaluation suites. Together, our results identify Value Flattening as an important yet overlooked failure mode of critic learning in standard PPO and show that a simple sparse supervision strategy can mitigate it.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.18708 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.18708 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.18708 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.