Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.</p>\n","updatedAt":"2026-09-10T02:10:08.602Z","author":{"_id":"66b1dee3bb6fcf48f1bee343","avatarUrl":"/avatars/c56049647323313bf31b81dcd19c1348.svg","fullname":"Youngjun Yu","name":"colin31472","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9055525660514832},"editors":["colin31472"],"editorAvatarUrls":["/avatars/c56049647323313bf31b81dcd19c1348.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08650","authors":[{"_id":"6aa14579ff4bf7311191aaf9","user":{"_id":"66b1dee3bb6fcf48f1bee343","avatarUrl":"/avatars/c56049647323313bf31b81dcd19c1348.svg","isPro":false,"fullname":"Youngjun Yu","user":"colin31472","type":"user","name":"colin31472"},"name":"Youngjun Yu","status":"claimed_verified","statusLastChangedAt":"2026-09-09T16:45:04.623Z","hidden":false},{"_id":"6aa14579ff4bf7311191aafa","name":"Sanghwan Jang","hidden":false},{"_id":"6aa14579ff4bf7311191aafb","name":"Hwanjo Yu","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR","submittedOnDailyBy":{"_id":"66b1dee3bb6fcf48f1bee343","avatarUrl":"/avatars/c56049647323313bf31b81dcd19c1348.svg","isPro":false,"fullname":"Youngjun Yu","user":"colin31472","type":"user","name":"colin31472"},"summary":"Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.","upvotes":0,"discussionId":"6aa1457aff4bf7311191aafc","githubRepo":"https://github.com/colin31472/DATPO","githubRepoAddedBy":"user","ai_summary":"DATPO improves reasoning coverage in large models by using difficulty-adaptive tree-structured rollouts with sentence-entropy-guided branching and diversity-aware optimization.","ai_keywords":["RLVR","pass@k","difficulty-adaptive rollout","tree-based rollout","sentence-entropy-guided forking","DATPO","sibling-diversity advantage","test-time scaling"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08650.md","query":{}}">
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Abstract
DATPO improves reasoning coverage in large models by using difficulty-adaptive tree-structured rollouts with sentence-entropy-guided branching and diversity-aware optimization.
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.
Community
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.08650 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.08650 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.08650 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.