Hugging Face Daily Papers · · 3 min read

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We propose Contrastive Policy Optimization (CPO), a framework that uses the disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. We provide a theoretical explanation for why this disagreement reliably indicates token-level correctness. We further show that on-policy self-distillation can be viewed as a special instantiation of CPO.</p>\n","updatedAt":"2026-07-20T08:47:43.921Z","author":{"_id":"64118689756b9e455c7eac62","avatarUrl":"/avatars/cdb3da22593facf545a0bafbf548b07e.svg","fullname":"Xu Weiwen","name":"xww033","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9103520512580872},"editors":["xww033"],"editorAvatarUrls":["/avatars/cdb3da22593facf545a0bafbf548b07e.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.14614","authors":[{"_id":"6a5dde696a69ce099f4d6f54","name":"Weiwen Xu","hidden":false},{"_id":"6a5dde696a69ce099f4d6f55","name":"Jia Liu","hidden":false},{"_id":"6a5dde696a69ce099f4d6f56","name":"Hou Pong Chan","hidden":false},{"_id":"6a5dde696a69ce099f4d6f57","name":"Long Li","hidden":false},{"_id":"6a5dde696a69ce099f4d6f58","name":"Deng Cai","hidden":false},{"_id":"6a5dde696a69ce099f4d6f59","name":"Min Chen","hidden":false},{"_id":"6a5dde696a69ce099f4d6f5a","name":"Hao Zhang","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-20T00:00:00.000Z","title":"Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization","submittedOnDailyBy":{"_id":"64118689756b9e455c7eac62","avatarUrl":"/avatars/cdb3da22593facf545a0bafbf548b07e.svg","isPro":false,"fullname":"Xu Weiwen","user":"xww033","type":"user","name":"xww033"},"summary":"Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.","upvotes":5,"discussionId":"6a5dde6a6a69ce099f4d6f5b"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64118689756b9e455c7eac62","avatarUrl":"/avatars/cdb3da22593facf545a0bafbf548b07e.svg","isPro":false,"fullname":"Xu Weiwen","user":"xww033","type":"user"},{"_id":"65f2ad65cd968972d1b03aea","avatarUrl":"/avatars/74a8c05d06f9683cebc15c5d595f5458.svg","isPro":false,"fullname":"Jia Liu","user":"jialiu0330hust","type":"user"},{"_id":"64b7cd74ff6d81ae297feded","avatarUrl":"/avatars/880fbc96cc093f5e901ce84f32a1d21d.svg","isPro":false,"fullname":"ZHANG HAO","user":"26hzhang","type":"user"},{"_id":"6a15df7df60916c42e28a4b8","avatarUrl":"/avatars/a49d3e8b654560812c43ed246f6ca6da.svg","isPro":false,"fullname":"子涵 郑","user":"grace-martinez3","type":"user"},{"_id":"6a1477d5922424601198e4ea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/zJn1ZF0Os0pDoOErZHFa7.png","isPro":false,"fullname":"장예은","user":"EvelynYoung2026","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.14614.md","query":{}}">
Papers
arxiv:2607.14614

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Published on Jul 16
· Submitted by
Xu Weiwen
on Jul 20
Authors:
,

Abstract

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.

Community

Paper submitter about 2 hours ago

We propose Contrastive Policy Optimization (CPO), a framework that uses the disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. We provide a theoretical explanation for why this disagreement reliably indicates token-level correctness. We further show that on-policy self-distillation can be viewed as a special instantiation of CPO.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.14614
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.14614 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.14614 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.14614 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers