A paper focusing on LLM RL and reward shaping</p>\n","updatedAt":"2026-09-03T02:13:20.882Z","author":{"_id":"638d601b5e14c2f38678fb3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg","fullname":"韩沛煊","name":"HakHan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.912456750869751},"editors":["HakHan"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.02817","authors":[{"_id":"6a98d1dffea818274321fdf8","name":"Peixuan Han","hidden":false},{"_id":"6a98d1dffea818274321fdf9","name":"Runhui Wang","hidden":false},{"_id":"6a98d1dffea818274321fdfa","name":"Ketan Ramaneti","hidden":false},{"_id":"6a98d1dffea818274321fdfb","name":"Jie Hao","hidden":false},{"_id":"6a98d1dffea818274321fdfc","name":"Gerald Friedland","hidden":false},{"_id":"6a98d1dffea818274321fdfd","name":"Chris Kong","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"Cliff: Learning Process Rewards from the First Mistake","submittedOnDailyBy":{"_id":"638d601b5e14c2f38678fb3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg","isPro":false,"fullname":"韩沛煊","user":"HakHan","type":"user","name":"HakHan"},"summary":"Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.","upvotes":15,"discussionId":"6a98d1dffea818274321fdfe","projectPage":"https://x.com/peixuanhakhan/status/2095334137909359011","ai_summary":"Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.","ai_keywords":["reinforcement learning with verifiable rewards","process reward modeling","on-policy distillation","reward shaping","token-level advantages","GRPO"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63624ffd2a84d82a8c8d3f60","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63624ffd2a84d82a8c8d3f60/NhDFHhlCWJeMevXN7aQUX.png","isPro":false,"fullname":"Chumeng Liang","user":"chumengl","type":"user"},{"_id":"66f8689725464a7989b75845","avatarUrl":"/avatars/43a61a528c5779103eaf5687ba44ee14.svg","isPro":false,"fullname":"Jiarui Yao","user":"FlippyDora","type":"user"},{"_id":"6a8cbea247cc488ec586dbf0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8cbea247cc488ec586dbf0/nhbFguivVtsnO8AjW8m7m.jpeg","isPro":false,"fullname":"Z. Zhao","user":"Xidianstatistics","type":"user"},{"_id":"638d601b5e14c2f38678fb3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg","isPro":false,"fullname":"韩沛煊","user":"HakHan","type":"user"},{"_id":"662a7759efa616e734ab493d","avatarUrl":"/avatars/8b79c6ec01f13d6b82414a6ad1b2d588.svg","isPro":false,"fullname":"Ke Yang","user":"EmpathYang","type":"user"},{"_id":"655601f1ae085c2ba7a22b95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/4UmxFrc_TEiXcnm3RewZM.jpeg","isPro":false,"fullname":"Xiaoji Zheng","user":"Student-Xiaoji","type":"user"},{"_id":"6a8c87c7ae19072fca4cd0e4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8c87c7ae19072fca4cd0e4/kx3LaDZXVSYnk8bfaDEm1.jpeg","isPro":false,"fullname":"欧阳晨风","user":"wwilliamsdaniel","type":"user"},{"_id":"6449dbd8df4e6cb7eaef943e","avatarUrl":"/avatars/41a549a7b1cfe1d59ea16b3cbd2168cc.svg","isPro":false,"fullname":"ChengQ","user":"0Cheng0","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6270ff726417aed8a7340c8b","avatarUrl":"/avatars/3f14913c55cc4fc78678ac43fb603e80.svg","isPro":false,"fullname":"Xiusi Chen","user":"XtremSup","type":"user"},{"_id":"699daec187f4416b4e205d08","avatarUrl":"/avatars/1b76013615423c9798f083fa37283357.svg","isPro":false,"fullname":"Carter Sanchez","user":"cartersanchez8","type":"user"},{"_id":"6a8097824030eff1e049bab6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8097824030eff1e049bab6/yQmOwpZnklbuDLYtAb3pn.jpeg","isPro":false,"fullname":"徐军","user":"zhenghuinim","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.02817.md","query":{}}">
Cliff: Learning Process Rewards from the First Mistake
Published on Sep 2
· Submitted by 韩沛煊 on Sep 3 Abstract
Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.
Community
A paper focusing on LLM RL and reward shaping
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.02817 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.02817 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.02817 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.