We introduce <strong>CriPO</strong> (<strong>Criterion-Distilled Policy Optimization</strong>), a simple on-policy framework that improves rubric-based reinforcement learning for open-ended LLM post-training. We identify two overlooked failure modes—<strong>Unexplored Criteria</strong> and <strong>Suppressed Criteria</strong>—and address them with localized self-distillation and token-level advantage correction, without introducing train–inference mismatch. Across medicine and science benchmarks, CriPO consistently outperforms existing rubric-based RL methods while reaching the same performance with ~2× fewer optimization steps.</p>\n","updatedAt":"2026-08-03T03:06:34.735Z","author":{"_id":"65dc040952eca001fd0bb142","avatarUrl":"/avatars/cd8e54ceef7c9e4a3bb4b0900c47a8b6.svg","fullname":"Mingxuan Xia","name":"MingxuanXia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8479852080345154},"editors":["MingxuanXia"],"editorAvatarUrls":["/avatars/cd8e54ceef7c9e4a3bb4b0900c47a8b6.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18082","authors":[{"_id":"6a7003b0bbe824e6bcc465ff","name":"Mingxuan Xia","hidden":false},{"_id":"6a7003b0bbe824e6bcc46600","name":"Yuhang Yang","hidden":false},{"_id":"6a7003b0bbe824e6bcc46601","name":"Chao Ye","hidden":false},{"_id":"6a7003b0bbe824e6bcc46602","name":"Shuai Zhu","hidden":false},{"_id":"6a7003b0bbe824e6bcc46603","name":"Shenzhi Yang","hidden":false},{"_id":"6a7003b0bbe824e6bcc46604","name":"Guangcheng Zhu","hidden":false},{"_id":"6a7003b0bbe824e6bcc46605","name":"Yuhang Zhang","hidden":false},{"_id":"6a7003b0bbe824e6bcc46606","name":"Cheng Peng","hidden":false},{"_id":"6a7003b0bbe824e6bcc46607","name":"Haobo Wang","hidden":false},{"_id":"6a7003b0bbe824e6bcc46608","name":"Siqing Wang","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-08-03T00:00:00.000Z","title":"Enhancing Rubric-based RL via Self-Distillation","submittedOnDailyBy":{"_id":"65dc040952eca001fd0bb142","avatarUrl":"/avatars/cd8e54ceef7c9e4a3bb4b0900c47a8b6.svg","isPro":false,"fullname":"Mingxuan Xia","user":"MingxuanXia","type":"user","name":"MingxuanXia"},"summary":"Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these exploration-focused approaches overlook a fundamentally different failure mode that we term Suppressed Criteria (SC) -- criteria that are satisfied by some rollouts yet whose learning signals are lost during optimization because scalar reward aggregation assigns them non-positive aggregate advantages. Our analysis reveals that SC are remarkably prevalent: over 57% of samples exhibit this failure mode throughout training, with an average of 1.8 SC per sample. To simultaneously address both UC and SC without introducing training-inference mismatch, we propose Criterion-Distilled Policy Optimization (CriPO), which enhances rubric-based RL via on-policy self-distillation. For UC, CriPO constructs a criterion-injection self-teacher and computes a localized forward-KL loss to inject missing behaviors into the policy. For SC, CriPO employs a counterfactual self-teacher to locate criterion-relevant tokens in negative-advantage rollouts and flips their token-level advantages to positive values, preserving useful patterns that would otherwise be suppressed. Experiments on medicine and science benchmarks demonstrate that CriPO consistently outperforms rubric-based RL, achieving stronger final performance with approximately 2times fewer optimization steps.","upvotes":13,"discussionId":"6a7003b1bbe824e6bcc46609"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65dc040952eca001fd0bb142","avatarUrl":"/avatars/cd8e54ceef7c9e4a3bb4b0900c47a8b6.svg","isPro":false,"fullname":"Mingxuan Xia","user":"MingxuanXia","type":"user"},{"_id":"6a6a82d4ce8b4ee20326608f","avatarUrl":"/avatars/a01c7ba7e8e706939874b22c4804665d.svg","isPro":false,"fullname":"George Martin","user":"eugene-4477412","type":"user"},{"_id":"6a6aa0a8b172d8c070b6b478","avatarUrl":"/avatars/a3e77445bdf99a9d00175ac942c99e81.svg","isPro":false,"fullname":"Edward Wilson","user":"timothy-2108992","type":"user"},{"_id":"6a6aa3fbb58832f7d0fd7e32","avatarUrl":"/avatars/e2b99552157ef3356352ec0af387396e.svg","isPro":false,"fullname":"Sarah Smith","user":"stephen-0908817","type":"user"},{"_id":"6a6c7c392cdbd2d07b4e26fa","avatarUrl":"/avatars/4097ee3de713cb0bd315c851e81ba635.svg","isPro":false,"fullname":"Michael Garcia","user":"bradley-4448477","type":"user"},{"_id":"6a6c8374ed2d5f6d8078cdd1","avatarUrl":"/avatars/81df2448657fcb21486b8fdd739ed5c6.svg","isPro":false,"fullname":"Karen Anderson","user":"mary-6767104","type":"user"},{"_id":"6a6c8c8d8248f83bcaa6bb5c","avatarUrl":"/avatars/afa9c97281da39900902888dbc037c7a.svg","isPro":false,"fullname":"Steven Sanchez","user":"gary-8653498","type":"user"},{"_id":"6a6aa1a9719ef934225fa075","avatarUrl":"/avatars/c7528de1b34e4f12e112800214751369.svg","isPro":false,"fullname":"Matthew Thomas","user":"justin-7911330","type":"user"},{"_id":"6a6c9d27ac4f2e43d4f7a737","avatarUrl":"/avatars/97782f8f69a2442067c878822d355aea.svg","isPro":false,"fullname":"Brian Jones","user":"austin-5261133","type":"user"},{"_id":"6a6dc954cc72fa11589a9b29","avatarUrl":"/avatars/31b8a3bea5f6bbd4834cd5460a295790.svg","isPro":false,"fullname":"Karen Wilson","user":"mary-0339432","type":"user"},{"_id":"6a6de8afcd50c6f8f59c8f11","avatarUrl":"/avatars/f8fd05ff9b2a6bc3f31955e17d292bdb.svg","isPro":false,"fullname":"Sarah Miller","user":"michael-0506344","type":"user"},{"_id":"6a6dec339caa2ff7db7a720b","avatarUrl":"/avatars/8d103b3fb8475daf9ce93af90cdda06b.svg","isPro":false,"fullname":"Paul Gonzalez","user":"david-7490294","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18082.md","query":{}}">
Enhancing Rubric-based RL via Self-Distillation
Abstract
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these exploration-focused approaches overlook a fundamentally different failure mode that we term Suppressed Criteria (SC) -- criteria that are satisfied by some rollouts yet whose learning signals are lost during optimization because scalar reward aggregation assigns them non-positive aggregate advantages. Our analysis reveals that SC are remarkably prevalent: over 57% of samples exhibit this failure mode throughout training, with an average of 1.8 SC per sample. To simultaneously address both UC and SC without introducing training-inference mismatch, we propose Criterion-Distilled Policy Optimization (CriPO), which enhances rubric-based RL via on-policy self-distillation. For UC, CriPO constructs a criterion-injection self-teacher and computes a localized forward-KL loss to inject missing behaviors into the policy. For SC, CriPO employs a counterfactual self-teacher to locate criterion-relevant tokens in negative-advantage rollouts and flips their token-level advantages to positive values, preserving useful patterns that would otherwise be suppressed. Experiments on medicine and science benchmarks demonstrate that CriPO consistently outperforms rubric-based RL, achieving stronger final performance with approximately 2times fewer optimization steps.
Community
We introduce CriPO (Criterion-Distilled Policy Optimization), a simple on-policy framework that improves rubric-based reinforcement learning for open-ended LLM post-training. We identify two overlooked failure modes—Unexplored Criteria and Suppressed Criteria—and address them with localized self-distillation and token-level advantage correction, without introducing train–inference mismatch. Across medicine and science benchmarks, CriPO consistently outperforms existing rubric-based RL methods while reaching the same performance with ~2× fewer optimization steps.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.18082 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.18082 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.18082 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.