How much does RLVR performance depend on the data policy? DataFlex-RL provides a controlled evaluation platform for studying rollout selection, reweighting, and domain-mixture strategies under a common GRPO recipe. Across 13 configurations, 12 matched seeds, and 12 math, logic, and science benchmarks, uniform sampling improves average accuracy by 7.76 points, while no alternative policy shows a statistically reliable improvement. The study also shows that benchmark composition can substantially change conclusions, highlighting the need for balanced and reproducible evaluation in RLVR.</p>\n","updatedAt":"2026-09-14T12:18:40.451Z","author":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","fullname":"Hao Liang","name":"lhpku20010120","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.90956711769104},"editors":["lhpku20010120"],"editorAvatarUrls":["/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06107","authors":[{"_id":"6aa0cf6ed0174964227bec9e","user":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":true,"fullname":"Hao Liang","user":"lhpku20010120","type":"user","name":"lhpku20010120"},"name":"Hao Liang","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.584Z","hidden":false},{"_id":"6aa0cf6ed0174964227bec9f","name":"Mingrui Chen","hidden":false},{"_id":"6aa0cf6ed0174964227beca0","name":"Hengyi Feng","hidden":false},{"_id":"6aa0cf6ed0174964227beca1","name":"Meiyi Qiang","hidden":false},{"_id":"6aa0cf6ed0174964227beca2","name":"Wentao Zhang","hidden":false}],"publishedAt":"2026-09-05T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"DataFlex-RL: An Evaluation Platform for RLVR Data Policies","submittedOnDailyBy":{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":true,"fullname":"Hao Liang","user":"lhpku20010120","type":"user","name":"lhpku20010120"},"summary":"Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.","upvotes":84,"discussionId":"6aa0cf6ed0174964227beca3","ai_summary":"DataFlex-RL evaluates reinforcement learning data policies and finds that uniform sampling matches or exceeds adaptive rollout selection, reweighting, and domain mixing across math, logic, and science benchmarks.","ai_keywords":["RLVR","GRPO","rollout-selection","reweighting","adaptive mixtures","domain-balanced average accuracy","Qwen2.5-7B-Base","Llama-3.1-8B-Base","GPQA-Diamond"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6751a4fedf636b0140a9b873","avatarUrl":"/avatars/d75f7f6cfbfb4d646e0e557d1cfacdce.svg","isPro":true,"fullname":"Hao Liang","user":"lhpku20010120","type":"user"},{"_id":"6217599529500f41901123f8","avatarUrl":"/avatars/8a0fe54e53fe6527c70a78598a0cd941.svg","isPro":false,"fullname":"Hao Liang","user":"lhbit20010120","type":"user"},{"_id":"66ac9567c97d2f0c88c3ac72","avatarUrl":"/avatars/14df8b5eed4ea756c93f61999c75e44f.svg","isPro":false,"fullname":"PKU_Baichuan","user":"PKU-Baichuan","type":"user"},{"_id":"67ca931e163cf5cf898e49b3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ca931e163cf5cf898e49b3/Fh7RsNtPSM2iBSY4dJZOW.jpeg","isPro":false,"fullname":"scuuy","user":"scuuy666","type":"user"},{"_id":"65b7098af327f1f4e315294d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b7098af327f1f4e315294d/38JtHebJ_XYWciMA15vJ3.jpeg","isPro":false,"fullname":"Runming He","user":"blackBOX25I47","type":"user"},{"_id":"68400c7b50cb0ac62e5fd9f2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68400c7b50cb0ac62e5fd9f2/UqFfQFbFsCxjLIcwIwdFx.png","isPro":false,"fullname":"Qihan Lin","user":"tunaaa126","type":"user"},{"_id":"69b38a5b3fc81ea209b001df","avatarUrl":"/avatars/28497c47bb579f4b702b19c603bd29e8.svg","isPro":false,"fullname":"Jincen Luo","user":"ArolaLuo","type":"user"},{"_id":"668cfdacf278aa900c8e00be","avatarUrl":"/avatars/0c03d223736448480dd434d54ada249e.svg","isPro":false,"fullname":"Hengyi Feng","user":"Heinz217","type":"user"},{"_id":"6618a60721d5003025004c96","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6618a60721d5003025004c96/_9DZG3lKbIt5KOj4WxsIJ.jpeg","isPro":false,"fullname":"Meiyi Qiang","user":"MeiyiQiang","type":"user"},{"_id":"664ac1f5f604081903dedac3","avatarUrl":"/avatars/e08f99828644f251dbfb6c1650b8bb95.svg","isPro":false,"fullname":"xubinrui","user":"xu220811","type":"user"},{"_id":"6880e7df575954f53124c68c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/NG-VcF3FvLUxmVkAUTu9w.png","isPro":false,"fullname":"Qifeng Xia","user":"PiarPP","type":"user"},{"_id":"6a64b707f530e420d1721c1f","avatarUrl":"/avatars/652cfb530c649fe807c59b8f449cc20f.svg","isPro":false,"fullname":"NJQ","user":"MYNJQ","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06107.md","query":{}}">
DataFlex-RL: An Evaluation Platform for RLVR Data Policies
Abstract
DataFlex-RL evaluates reinforcement learning data policies and finds that uniform sampling matches or exceeds adaptive rollout selection, reweighting, and domain mixing across math, logic, and science benchmarks.
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
Community
How much does RLVR performance depend on the data policy? DataFlex-RL provides a controlled evaluation platform for studying rollout selection, reweighting, and domain-mixture strategies under a common GRPO recipe. Across 13 configurations, 12 matched seeds, and 12 math, logic, and science benchmarks, uniform sampling improves average accuracy by 7.76 points, while no alternative policy shows a statistically reliable improvement. The study also shows that benchmark composition can substantially change conclusions, highlighting the need for balanced and reproducible evaluation in RLVR.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.06107 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.06107 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.06107 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.