Hugging Face Daily Papers · · 4 min read

Cliff: Learning Process Rewards from the First Mistake

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A paper focusing on LLM RL and reward shaping</p>\n","updatedAt":"2026-09-03T02:13:20.882Z","author":{"_id":"638d601b5e14c2f38678fb3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg","fullname":"韩沛煊","name":"HakHan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.912456750869751},"editors":["HakHan"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.02817","authors":[{"_id":"6a98d1dffea818274321fdf8","name":"Peixuan Han","hidden":false},{"_id":"6a98d1dffea818274321fdf9","name":"Runhui Wang","hidden":false},{"_id":"6a98d1dffea818274321fdfa","name":"Ketan Ramaneti","hidden":false},{"_id":"6a98d1dffea818274321fdfb","name":"Jie Hao","hidden":false},{"_id":"6a98d1dffea818274321fdfc","name":"Gerald Friedland","hidden":false},{"_id":"6a98d1dffea818274321fdfd","name":"Chris Kong","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"Cliff: Learning Process Rewards from the First Mistake","submittedOnDailyBy":{"_id":"638d601b5e14c2f38678fb3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg","isPro":false,"fullname":"韩沛煊","user":"HakHan","type":"user","name":"HakHan"},"summary":"Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.","upvotes":15,"discussionId":"6a98d1dffea818274321fdfe","projectPage":"https://x.com/peixuanhakhan/status/2095334137909359011","ai_summary":"Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.","ai_keywords":["reinforcement learning with verifiable rewards","process reward modeling","on-policy distillation","reward shaping","token-level advantages","GRPO"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63624ffd2a84d82a8c8d3f60","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63624ffd2a84d82a8c8d3f60/NhDFHhlCWJeMevXN7aQUX.png","isPro":false,"fullname":"Chumeng Liang","user":"chumengl","type":"user"},{"_id":"66f8689725464a7989b75845","avatarUrl":"/avatars/43a61a528c5779103eaf5687ba44ee14.svg","isPro":false,"fullname":"Jiarui Yao","user":"FlippyDora","type":"user"},{"_id":"6a8cbea247cc488ec586dbf0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8cbea247cc488ec586dbf0/nhbFguivVtsnO8AjW8m7m.jpeg","isPro":false,"fullname":"Z. Zhao","user":"Xidianstatistics","type":"user"},{"_id":"638d601b5e14c2f38678fb3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638d601b5e14c2f38678fb3a/Elu-TTd97lGy7YL7eKRuZ.jpeg","isPro":false,"fullname":"韩沛煊","user":"HakHan","type":"user"},{"_id":"662a7759efa616e734ab493d","avatarUrl":"/avatars/8b79c6ec01f13d6b82414a6ad1b2d588.svg","isPro":false,"fullname":"Ke Yang","user":"EmpathYang","type":"user"},{"_id":"655601f1ae085c2ba7a22b95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/4UmxFrc_TEiXcnm3RewZM.jpeg","isPro":false,"fullname":"Xiaoji Zheng","user":"Student-Xiaoji","type":"user"},{"_id":"6a8c87c7ae19072fca4cd0e4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8c87c7ae19072fca4cd0e4/kx3LaDZXVSYnk8bfaDEm1.jpeg","isPro":false,"fullname":"欧阳晨风","user":"wwilliamsdaniel","type":"user"},{"_id":"6449dbd8df4e6cb7eaef943e","avatarUrl":"/avatars/41a549a7b1cfe1d59ea16b3cbd2168cc.svg","isPro":false,"fullname":"ChengQ","user":"0Cheng0","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6270ff726417aed8a7340c8b","avatarUrl":"/avatars/3f14913c55cc4fc78678ac43fb603e80.svg","isPro":false,"fullname":"Xiusi Chen","user":"XtremSup","type":"user"},{"_id":"699daec187f4416b4e205d08","avatarUrl":"/avatars/1b76013615423c9798f083fa37283357.svg","isPro":false,"fullname":"Carter Sanchez","user":"cartersanchez8","type":"user"},{"_id":"6a8097824030eff1e049bab6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8097824030eff1e049bab6/yQmOwpZnklbuDLYtAb3pn.jpeg","isPro":false,"fullname":"徐军","user":"zhenghuinim","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.02817.md","query":{}}">
Papers
arxiv:2609.02817

Cliff: Learning Process Rewards from the First Mistake

Published on Sep 2
· Submitted by
韩沛煊
on Sep 3
Authors:
,

Abstract

Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

Community

Paper submitter about 7 hours ago

A paper focusing on LLM RL and reward shaping

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.02817
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.02817 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.02817 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.02817 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers