Hugging Face Daily Papers · · 3 min read

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

[COLM 2026] From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement</p>\n","updatedAt":"2026-08-03T01:36:35.307Z","author":{"_id":"66128e8b7e0e7a64652dbbdf","avatarUrl":"/avatars/b90f06e74c52286fd421c4d9bbbc2c9f.svg","fullname":"Wang","name":"Qinsi1","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8370137214660645},"editors":["Qinsi1"],"editorAvatarUrls":["/avatars/b90f06e74c52286fd421c4d9bbbc2c9f.svg"],"reactions":[{"reaction":"👍","users":["Moeus"],"count":1}],"isReport":false}},{"id":"6a70432dd614a264ff80c5e0","author":{"_id":"63cc1a06336d9d56170db39c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674320327203-noauth.png","fullname":"Nitish Pandey","name":"nitishpandey04","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-08-03T07:28:45.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"mind = blown","html":"<p>mind = blown</p>\n","updatedAt":"2026-08-03T07:28:45.760Z","author":{"_id":"63cc1a06336d9d56170db39c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674320327203-noauth.png","fullname":"Nitish Pandey","name":"nitishpandey04","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4562975764274597},"editors":["nitishpandey04"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674320327203-noauth.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.23802","authors":[{"_id":"6a6fb250bbe824e6bcc4653c","name":"Qinsi Wang","hidden":false},{"_id":"6a6fb250bbe824e6bcc4653d","name":"Jing Shi","hidden":false},{"_id":"6a6fb250bbe824e6bcc4653e","name":"Huazheng Wang","hidden":false},{"_id":"6a6fb250bbe824e6bcc4653f","user":{"_id":"66274e02348a5304435dc9cc","avatarUrl":"/avatars/bda87559cd497c310597c2fc8430b31f.svg","isPro":false,"fullname":"Kun Wan","user":"timecuriosity","type":"user","name":"timecuriosity"},"name":"Kun Wan","status":"claimed_verified","statusLastChangedAt":"2026-08-03T07:54:50.257Z","hidden":false},{"_id":"6a6fb250bbe824e6bcc46540","name":"Yiran Wu","hidden":false},{"_id":"6a6fb250bbe824e6bcc46541","user":{"_id":"635e3a76106f984574c36409","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667120725800-635e3a76106f984574c36409.png","isPro":false,"fullname":"Bo Liu","user":"Benjamin-eecs","type":"user","name":"Benjamin-eecs"},"name":"Bo Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-03T00:45:04.570Z","hidden":false},{"_id":"6a6fb250bbe824e6bcc46542","name":"Qingyun Wu","hidden":false},{"_id":"6a6fb250bbe824e6bcc46543","name":"Hai Helen Li","hidden":false},{"_id":"6a6fb250bbe824e6bcc46544","name":"Yiran Chen","hidden":false},{"_id":"6a6fb250bbe824e6bcc46545","name":"Handong Zhao","hidden":false},{"_id":"6a6fb250bbe824e6bcc46546","name":"Wentian Zhao","hidden":false}],"publishedAt":"2026-07-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-03T00:00:00.000Z","title":"From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement","submittedOnDailyBy":{"_id":"66128e8b7e0e7a64652dbbdf","avatarUrl":"/avatars/b90f06e74c52286fd421c4d9bbbc2c9f.svg","isPro":false,"fullname":"Wang","user":"Qinsi1","type":"user","name":"Qinsi1"},"summary":"Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verified. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs.Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a multi-agent self-play environment inspired by Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at https://github.com/wangqinsi1/SpyRL.","upvotes":52,"discussionId":"6a6fb251bbe824e6bcc46547","projectPage":"https://github.com/wangqinsi1/RLSVR/tree/SpyRL","githubRepo":"https://github.com/wangqinsi1/RLSVR","githubRepoAddedBy":"user"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66128e8b7e0e7a64652dbbdf","avatarUrl":"/avatars/b90f06e74c52286fd421c4d9bbbc2c9f.svg","isPro":false,"fullname":"Wang","user":"Qinsi1","type":"user"},{"_id":"68aa1467759a11a9c65a9841","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68aa1467759a11a9c65a9841/9Ex1R4gfZwlRN_U3loRNl.jpeg","isPro":false,"fullname":"Jinghan Ke","user":"JinghanKe","type":"user"},{"_id":"68effdf018af859e89e3ca92","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/m5dgHAtaI8AsfRrJXva98.png","isPro":false,"fullname":"Jack Wang","user":"Jack231duke","type":"user"},{"_id":"64b5198c25882acb62fb77ef","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b5198c25882acb62fb77ef/HX9pfMEPQlfjvSAgSLplY.png","isPro":false,"fullname":"Yueqian Lin","user":"linyueqian","type":"user"},{"_id":"6798fd6d37bfd1568c58606e","avatarUrl":"/avatars/e92e6ea5a44d258d84d2e9b8976b4156.svg","isPro":false,"fullname":"uu","user":"JayZc","type":"user"},{"_id":"67c71505ec03b94956e112f3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67c71505ec03b94956e112f3/D_07LgN7mG3BO0IryhZ_X.jpeg","isPro":false,"fullname":"WeiLiang","user":"Bright45","type":"user"},{"_id":"67ccf98da5588863846e615f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ccf98da5588863846e615f/WmsnCEWd4Q2wQ-uTZ_Yaw.jpeg","isPro":false,"fullname":"ZhengHao","user":"ZhengHao-L","type":"user"},{"_id":"67b19f81615a3737b5772b3b","avatarUrl":"/avatars/2ca4c53715e82a7e17c4b631eaf34042.svg","isPro":false,"fullname":"Fenzhif","user":"FANCERTA","type":"user"},{"_id":"67b188864b65dd4cee3d9f9b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b188864b65dd4cee3d9f9b/gHEwPbAFjZITDqodAS_Qw.png","isPro":false,"fullname":"TLL","user":"ilHuGGing","type":"user"},{"_id":"66443629b23fe8d3f7f2d0c7","avatarUrl":"/avatars/98ff088036aa382f33a05c232604c565.svg","isPro":false,"fullname":"Wentian Zhao","user":"zwt123home123","type":"user"},{"_id":"67c0858901cef6d4b982f5a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67c0858901cef6d4b982f5a8/kFdHRMGAlT8AiMNq2iGk9.jpeg","isPro":false,"fullname":"XUTianyuu","user":"XuTianyuu","type":"user"},{"_id":"698c87259bc3b459ef48df03","avatarUrl":"/avatars/358ef5b37e7dcdb7936c8ebecaea4a35.svg","isPro":false,"fullname":"Zhuxinlin","user":"xinliliu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.23802.md","query":{}}">
Papers
arxiv:2607.23802

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Published on Jul 26
· Submitted by
Wang
on Aug 3
#1 Paper of the day
Authors:
,

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verified. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs.Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a multi-agent self-play environment inspired by Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at https://github.com/wangqinsi1/SpyRL.

Community

Paper submitter about 6 hours ago

[COLM 2026] From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.23802
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.23802 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.23802 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.23802 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers