CoRT addresses a simple but important mismatch in rubric-guided RL: rich criterion-level feedback is ultimately collapsed into one response-level advantage and broadcast uniformly to every generated token. By replaying the same response with and without the rubric criteria, CoRT turns policy-internal log-probability changes into normalized token-level credit weights—without training an auxiliary token scorer, changing the reward, or requiring additional generation and verifier calls. Across multiple models, reward settings, and instruction-following benchmarks, CoRT improves matched response-level GRPO by 4.4 points on average while remaining competitive with learned token-relevance methods. A lightweight and practical direction for fine-grained credit assignment in LLM reinforcement learning.</p>\n","updatedAt":"2026-07-30T03:33:17.587Z","author":{"_id":"636d2965a756aeb5c4af05e5","avatarUrl":"/avatars/439650942cb1298aaec98460753dbb9d.svg","fullname":"quinn","name":"jwhe","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8963597416877747},"editors":["jwhe"],"editorAvatarUrls":["/avatars/439650942cb1298aaec98460753dbb9d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25659","authors":[{"_id":"6a6ac4c14463a8a84bdc3fd4","name":"Bo-Wen Zhang","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd5","name":"Junwei He","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd6","name":"Wen Wang","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd7","name":"Song-Lin Lv","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd8","name":"Wentao Ma","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fd9","name":"Rongyi Lin","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fda","name":"Shuhan Zhong","hidden":false},{"_id":"6a6ac4c14463a8a84bdc3fdb","name":"Lan-Zhe Guo","hidden":false}],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization","submittedOnDailyBy":{"_id":"636d2965a756aeb5c4af05e5","avatarUrl":"/avatars/439650942cb1298aaec98460753dbb9d.svg","isPro":false,"fullname":"quinn","user":"jwhe","type":"user","name":"jwhe"},"summary":"Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.","upvotes":32,"discussionId":"6a6ac4c24463a8a84bdc3fdc","organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"636d2965a756aeb5c4af05e5","avatarUrl":"/avatars/439650942cb1298aaec98460753dbb9d.svg","isPro":false,"fullname":"quinn","user":"jwhe","type":"user"},{"_id":"66894d95c1c9eaffe4a82b11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66894d95c1c9eaffe4a82b11/BqrZvU_L0ulB_U45hmEy1.jpeg","isPro":false,"fullname":"Bowen Zhang","user":"Cbphcr","type":"user"},{"_id":"663835feccadfaaeac43ef48","avatarUrl":"/avatars/c317e30d3f8451d0f59d1a49330c4bd9.svg","isPro":false,"fullname":"MQliu","user":"EstarHash","type":"user"},{"_id":"69a3beb716005ca14c16649b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/T947I0eMbx6lZdLJWiTtn.png","isPro":false,"fullname":"Арсений Волков","user":"mason-flores4","type":"user"},{"_id":"6629d26acecbf3a7f825d716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6629d26acecbf3a7f825d716/imWEAQK1IRNPeIUWJQf5d.jpeg","isPro":false,"fullname":"Ferry Li","user":"Ferry30","type":"user"},{"_id":"6604c1c68726fd47390cef63","avatarUrl":"/avatars/889c3d2f9208a6d5f90b1150778621f6.svg","isPro":false,"fullname":"yangjunqi23","user":"yangjunqi23","type":"user"},{"_id":"6a147f4f695d577a5249a9c8","avatarUrl":"/avatars/1194f0e6d9c809dd7767c413d64cd889.svg","isPro":false,"fullname":"Emily Brown","user":"emily-brown2025","type":"user"},{"_id":"6a146c2131430965c63f5ecb","avatarUrl":"/avatars/28e5680cbea1978e4cc6a005614e66a6.svg","isPro":false,"fullname":"Lin Wenxuan","user":"linwenxuan6","type":"user"},{"_id":"6a15a30886efa551cffd4509","avatarUrl":"/avatars/55bb0dd90b93a63002e4ea83a5ade0da.svg","isPro":false,"fullname":"Lin Haoran","user":"liha5y","type":"user"},{"_id":"6a14686491d4af20d4030a76","avatarUrl":"/avatars/6c36c9a3d564174094e7ddbfc5260ae7.svg","isPro":false,"fullname":"Huang Siyu","user":"huangsiy3","type":"user"},{"_id":"6a1467d8af3fe6cd43bdbad6","avatarUrl":"/avatars/ce26446e5a8452573162d27fee66a458.svg","isPro":false,"fullname":"Zhou Wenxuan","user":"zwenxuan","type":"user"},{"_id":"6a14682131430965c63f1bc3","avatarUrl":"/avatars/06ff73863b78bc69f313c8d2015ff241.svg","isPro":false,"fullname":"Hu Linxi","user":"hulinxi","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25659.md","query":{}}">
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Published on Jul 28
· Submitted by quinn on Jul 30 Abstract
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.
Community
CoRT addresses a simple but important mismatch in rubric-guided RL: rich criterion-level feedback is ultimately collapsed into one response-level advantage and broadcast uniformly to every generated token. By replaying the same response with and without the rubric criteria, CoRT turns policy-internal log-probability changes into normalized token-level credit weights—without training an auxiliary token scorer, changing the reward, or requiring additional generation and verifier calls. Across multiple models, reward settings, and instruction-following benchmarks, CoRT improves matched response-level GRPO by 4.4 points on average while remaining competitive with learned token-relevance methods. A lightweight and practical direction for fine-grained credit assignment in LLM reinforcement learning.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.25659 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.25659 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.25659 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.