Can an AI system improve not only its answers, but also the criteria it uses to judge those answers? We introduce DecoEvo, a score-decoupled co-evolution framework that jointly evolves a solver and a rubric generator in text space. By extracting structured feedback from solution audits and analyzing discrepancies across multiple rollouts, DecoEvo continually refines both problem-solving strategies and evaluation standards. Our results show that this co-evolution process provides stronger learning signals and consistently improves reasoning performance across diverse tasks.</p>\n","updatedAt":"2026-07-30T03:02:51.537Z","author":{"_id":"66a083c16580efecfb69a41f","avatarUrl":"/avatars/51b14942dc86e1d39d82fed9b3b1d744.svg","fullname":"chenjiangwang","name":"jwchen2001","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9248611330986023},"editors":["jwchen2001"],"editorAvatarUrls":["/avatars/51b14942dc86e1d39d82fed9b3b1d744.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25675","authors":[{"_id":"6a6969f29d3a1231d492b861","user":{"_id":"66a083c16580efecfb69a41f","avatarUrl":"/avatars/51b14942dc86e1d39d82fed9b3b1d744.svg","isPro":false,"fullname":"chenjiangwang","user":"jwchen2001","type":"user","name":"jwchen2001"},"name":"Jiangwang Chen","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.472Z","hidden":false},{"_id":"6a6969f29d3a1231d492b862","user":{"_id":"6579c94a1e0e383c6f9954c9","avatarUrl":"/avatars/37b3e953c5a0d592e0b776513cca242f.svg","isPro":false,"fullname":"Zixin Song","user":"prophesier","type":"user","name":"prophesier"},"name":"Zixin Song","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.487Z","hidden":false},{"_id":"6a6969f29d3a1231d492b863","user":{"_id":"69045a2e9c3d523381fcf489","avatarUrl":"/avatars/8b1982e043a7ab28006b707dfff6c5dc.svg","isPro":false,"fullname":"JunlinLiu","user":"AaronLiu0702","type":"user","name":"AaronLiu0702"},"name":"Junlin Liu","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.480Z","hidden":false},{"_id":"6a6969f29d3a1231d492b864","name":"Shuaiyu Zhou","hidden":false},{"_id":"6a6969f29d3a1231d492b865","name":"Haiyan Wu","hidden":false},{"_id":"6a6969f29d3a1231d492b866","name":"Haihan Shi","hidden":false},{"_id":"6a6969f29d3a1231d492b867","name":"Chenxi Zhou","hidden":false},{"_id":"6a6969f29d3a1231d492b868","name":"Hanqing Li","hidden":false},{"_id":"6a6969f29d3a1231d492b869","name":"Xiao Yang","hidden":false},{"_id":"6a6969f29d3a1231d492b86a","name":"Da Zhu","hidden":false},{"_id":"6a6969f29d3a1231d492b86b","name":"Guanjun Jiang","hidden":false},{"_id":"6a6969f29d3a1231d492b86c","name":"Hai Wan","hidden":false},{"_id":"6a6969f29d3a1231d492b86d","name":"Xibin Zhao","hidden":false}],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space","submittedOnDailyBy":{"_id":"66a083c16580efecfb69a41f","avatarUrl":"/avatars/51b14942dc86e1d39d82fed9b3b1d744.svg","isPro":false,"fullname":"chenjiangwang","user":"jwchen2001","type":"user","name":"jwchen2001"},"summary":"Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\\% relative gains over SkillOpt in the five-benchmark average.","upvotes":45,"discussionId":"6a6969f29d3a1231d492b86e","organization":{"_id":"6a6841e7107886ba1a151b03","name":"QwenBusinessUnit","fullname":"Qwen Business Unit","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f79b323fe089b75e9e0c04/MlefZsdry-JuhKzAwxjQl.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6579c94a1e0e383c6f9954c9","avatarUrl":"/avatars/37b3e953c5a0d592e0b776513cca242f.svg","isPro":false,"fullname":"Zixin Song","user":"prophesier","type":"user"},{"_id":"66a5e1c42756d40a3fa9f0da","avatarUrl":"/avatars/ada30e3cd4efffcaf73d29f5d19d6e4f.svg","isPro":false,"fullname":"shuaiyv zhou","user":"monster290","type":"user"},{"_id":"69a825f0dfa5dd290305789f","avatarUrl":"/avatars/19c506fc8f058697c9fb69c37287d3b9.svg","isPro":false,"fullname":"james","user":"james011112","type":"user"},{"_id":"69a7f0beaa631530a75c5cee","avatarUrl":"/avatars/4baaa57608ea82378e301598578da572.svg","isPro":false,"fullname":"HanqingLi","user":"HanqingLi","type":"user"},{"_id":"66a083c16580efecfb69a41f","avatarUrl":"/avatars/51b14942dc86e1d39d82fed9b3b1d744.svg","isPro":false,"fullname":"chenjiangwang","user":"jwchen2001","type":"user"},{"_id":"65afd2ed7e5d5a4ecc2c6254","avatarUrl":"/avatars/ef27a06fd75373b099bfc8be1d80a057.svg","isPro":false,"fullname":"Bo Li","user":"EchoO0","type":"user"},{"_id":"67d27fb95785e2093c553182","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/P0JWb5xIVV4Az-1mvXNcA.png","isPro":false,"fullname":"ZianHuang","user":"ZianHuang","type":"user"},{"_id":"69d4c8fa2d99e481567ddcdf","avatarUrl":"/avatars/9f20051ffda505465d300eb4e03e710e.svg","isPro":false,"fullname":"He","user":"houxinxinxinxin","type":"user"},{"_id":"67f60a9c8f60757d94f9e325","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/8GaHEzcS4IksrPpD9qrNy.png","isPro":false,"fullname":"孙桂宇","user":"starShooting","type":"user"},{"_id":"6915acbb907428dbc99519fe","avatarUrl":"/avatars/860a461ffa45b8ec8c969fb2428f9fee.svg","isPro":false,"fullname":"aaa","user":"yhrbb02","type":"user"},{"_id":"6656a923de16d232e0e1c5fe","avatarUrl":"/avatars/6f5168817994b078f0c28d86832a3d4e.svg","isPro":false,"fullname":"Jiazheng Kang","user":"Jiazhengg","type":"user"},{"_id":"650856b118f9c580cde97d07","avatarUrl":"/avatars/b60adcb7ad64bdf964739ec3f9e0eea7.svg","isPro":false,"fullname":"Shi HaiHan","user":"shh924","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6a6841e7107886ba1a151b03","name":"QwenBusinessUnit","fullname":"Qwen Business Unit","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f79b323fe089b75e9e0c04/MlefZsdry-JuhKzAwxjQl.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25675.md","query":{}}">
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Abstract
Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.
Community
Can an AI system improve not only its answers, but also the criteria it uses to judge those answers? We introduce DecoEvo, a score-decoupled co-evolution framework that jointly evolves a solver and a rubric generator in text space. By extracting structured feedback from solution audits and analyzing discrepancies across multiple rollouts, DecoEvo continually refines both problem-solving strategies and evaluation standards. Our results show that this co-evolution process provides stronger learning signals and consistently improves reasoning performance across diverse tasks.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.25675 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.25675 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.25675 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.