Hugging Face Daily Papers · · 4 min read

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

CoRT addresses a simple but important mismatch in rubric-guided RL: rich criterion-level feedback is ultimately collapsed into one response-level advantage and broadcast uniformly to every generated token. By replaying the same response with and without the rubric criteria, CoRT turns policy-internal log-probability changes into normalized token-level credit weights—without training an auxiliary token scorer, changing the reward, or requiring additional generation and verifier calls. Across multiple models, reward settings, and instruction-following benchmarks, CoRT improves matched response-level GRPO by 4.4 points on average while remaining competitive with learned token-relevance methods. A lightweight and practical direction for fine-grained credit assignment in LLM reinforcement learning.</p>\n","updatedAt":"2026-07-30T03:33:17.587Z","author":{"_id":"636d2965a756aeb5c4af05e5","avatarUrl":"/avatars/439650942cb1298aaec98460753dbb9d.svg","fullname":"quinn","name":"jwhe","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8963597416877747},"editors":["jwhe"],"editorAvatarUrls":["/avatars/439650942cb1298aaec98460753dbb9d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25659","authors":[{"_id":"6a6ac4c14463a8a84bdc3fd4","name":"Bo-Wen Zhang","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd5","name":"Junwei He","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd6","name":"Wen Wang","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd7","name":"Song-Lin Lv","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd8","name":"Wentao Ma","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd9","name":"Rongyi Lin","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fda","name":"Shuhan Zhong","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fdb","name":"Lan-Zhe Guo","hidden":false}],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization","submittedOnDailyBy":{"_id":"636d2965a756aeb5c4af05e5","avatarUrl":"/avatars/439650942cb1298aaec98460753dbb9d.svg","isPro":false,"fullname":"quinn","user":"jwhe","type":"user","name":"jwhe"},"summary":"Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.","upvotes":32,"discussionId":"6a6ac4c24463a8a84bdc3fdc","organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"636d2965a756aeb5c4af05e5","avatarUrl":"/avatars/439650942cb1298aaec98460753dbb9d.svg","isPro":false,"fullname":"quinn","user":"jwhe","type":"user"},{"_id":"66894d95c1c9eaffe4a82b11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66894d95c1c9eaffe4a82b11/BqrZvU_L0ulB_U45hmEy1.jpeg","isPro":false,"fullname":"Bowen Zhang","user":"Cbphcr","type":"user"},{"_id":"663835feccadfaaeac43ef48","avatarUrl":"/avatars/c317e30d3f8451d0f59d1a49330c4bd9.svg","isPro":false,"fullname":"MQliu","user":"EstarHash","type":"user"},{"_id":"69a3beb716005ca14c16649b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/T947I0eMbx6lZdLJWiTtn.png","isPro":false,"fullname":"Арсений Волков","user":"mason-flores4","type":"user"},{"_id":"6629d26acecbf3a7f825d716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6629d26acecbf3a7f825d716/imWEAQK1IRNPeIUWJQf5d.jpeg","isPro":false,"fullname":"Ferry Li","user":"Ferry30","type":"user"},{"_id":"6604c1c68726fd47390cef63","avatarUrl":"/avatars/889c3d2f9208a6d5f90b1150778621f6.svg","isPro":false,"fullname":"yangjunqi23","user":"yangjunqi23","type":"user"},{"_id":"6a147f4f695d577a5249a9c8","avatarUrl":"/avatars/1194f0e6d9c809dd7767c413d64cd889.svg","isPro":false,"fullname":"Emily Brown","user":"emily-brown2025","type":"user"},{"_id":"6a146c2131430965c63f5ecb","avatarUrl":"/avatars/28e5680cbea1978e4cc6a005614e66a6.svg","isPro":false,"fullname":"Lin Wenxuan","user":"linwenxuan6","type":"user"},{"_id":"6a15a30886efa551cffd4509","avatarUrl":"/avatars/55bb0dd90b93a63002e4ea83a5ade0da.svg","isPro":false,"fullname":"Lin Haoran","user":"liha5y","type":"user"},{"_id":"6a14686491d4af20d4030a76","avatarUrl":"/avatars/6c36c9a3d564174094e7ddbfc5260ae7.svg","isPro":false,"fullname":"Huang Siyu","user":"huangsiy3","type":"user"},{"_id":"6a1467d8af3fe6cd43bdbad6","avatarUrl":"/avatars/ce26446e5a8452573162d27fee66a458.svg","isPro":false,"fullname":"Zhou Wenxuan","user":"zwenxuan","type":"user"},{"_id":"6a14682131430965c63f1bc3","avatarUrl":"/avatars/06ff73863b78bc69f313c8d2015ff241.svg","isPro":false,"fullname":"Hu Linxi","user":"hulinxi","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25659.md","query":{}}">
Papers
arxiv:2607.25659

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Published on Jul 28
· Submitted by
quinn
on Jul 30
Authors:
,

Abstract

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.

Community

Paper submitter about 4 hours ago

CoRT addresses a simple but important mismatch in rubric-guided RL: rich criterion-level feedback is ultimately collapsed into one response-level advantage and broadcast uniformly to every generated token. By replaying the same response with and without the rubric criteria, CoRT turns policy-internal log-probability changes into normalized token-level credit weights—without training an auxiliary token scorer, changing the reward, or requiring additional generation and verifier calls. Across multiple models, reward settings, and instruction-following benchmarks, CoRT improves matched response-level GRPO by 4.4 points on average while remaining competitive with learned token-relevance methods. A lightweight and practical direction for fine-grained credit assignment in LLM reinforcement learning.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.25659
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.25659 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.25659 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.25659 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers