Hugging Face Daily Papers · · 6 min read

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

RLVR-induced solution collapse is an access failure, not an execution failure: models lose reasoning diversity at initial computational entrance while retaining latent downstream capability. Reasoning breadth is lost at the door, not inside the room.</p>\n","updatedAt":"2026-09-04T19:23:46.965Z","author":{"_id":"633f536250d83f5065d28f6d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/633f536250d83f5065d28f6d/eZCdJJpU6OJFYUWYoUAVn.jpeg","fullname":"Ruizhe Li","name":"rzdiversity","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.919245719909668},"editors":["rzdiversity"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/633f536250d83f5065d28f6d/eZCdJJpU6OJFYUWYoUAVn.jpeg"],"reactions":[],"isReport":false}},{"id":"6a9b6f366f26088671a8fc08","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-05T01:24:06.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [Verifier-Induced Support Reshaping in On-Policy Optimization](https://huggingface.co/papers/2608.00220) (2026)\n* [Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information](https://huggingface.co/papers/2607.19313) (2026)\n* [Boosting LLM Exploration via Weak-Model Guidance in RLVR](https://huggingface.co/papers/2608.27420) (2026)\n* [Contrastive Branch Policy Optimization](https://huggingface.co/papers/2608.24300) (2026)\n* [Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies](https://huggingface.co/papers/2608.12679) (2026)\n* [SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation](https://huggingface.co/papers/2608.04419) (2026)\n* [From Base Rollouts to RL Reasoning: A Budgeted Search Perspective](https://huggingface.co/papers/2609.01274) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2608.00220\">Verifier-Induced Support Reshaping in On-Policy Optimization</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.19313\">Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.27420\">Boosting LLM Exploration via Weak-Model Guidance in RLVR</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.24300\">Contrastive Branch Policy Optimization</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.12679\">Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.04419\">SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2609.01274\">From Base Rollouts to RL Reasoning: A Budgeted Search Perspective</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-05T01:24:06.959Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7407609224319458},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.29188","authors":[{"_id":"6a9aaa1b8f7c3b75572397ab","name":"Qiancheng Zhou","hidden":false},{"_id":"6a9aaa1b8f7c3b75572397ac","user":{"_id":"633f536250d83f5065d28f6d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/633f536250d83f5065d28f6d/eZCdJJpU6OJFYUWYoUAVn.jpeg","isPro":false,"fullname":"Ruizhe Li","user":"rzdiversity","type":"user","name":"rzdiversity"},"name":"Ruizhe Li","status":"claimed_verified","statusLastChangedAt":"2026-09-05T00:45:04.103Z","hidden":false}],"publishedAt":"2026-08-29T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space","submittedOnDailyBy":{"_id":"633f536250d83f5065d28f6d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/633f536250d83f5065d28f6d/eZCdJJpU6OJFYUWYoUAVn.jpeg","isPro":false,"fullname":"Ruizhe Li","user":"rzdiversity","type":"user","name":"rzdiversity"},"summary":"Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator, across PPO on Qwen2.5-3B and GRPO on Qwen2.5-3B-Instruct. Across both training setups, solution coverage falls by up to 67%, halving even on problems solved across all checkpoints. We show that this contraction is heavily concentrated at the entrance: per-token likelihood shifts are 11x--16x larger prior to the first arithmetic operation than during downstream reasoning. Supplying only an unselected entrance prefix restores completion rates in low-access families by over an order of magnitude (0.018 -> 0.212 under PPO), demonstrating that alternative solutions remain executable but are no longer initiated. Guided by this localization, we find that while surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1. Finally, we show that early-step entropy collapse recurs across six math benchmarks with 7B and 14B models, but is not an inevitable byproduct of reasoning optimization: an SFT baseline preserves more than double the coverage, and staged SFT--DPO--RLVR pipelines retain early-step entropy. In summary, reasoning breadth is lost at the door, not inside the room. Code: https://github.com/ershiyidian/early-branch-locking.","upvotes":9,"discussionId":"6a9aaa1b8f7c3b75572397ad","githubRepo":"https://github.com/ershiyidian/early-branch-locking","githubRepoAddedBy":"user","ai_summary":"Reinforcement learning with verifiable rewards narrows reasoning diversity primarily at the initial solution step rather than during execution, and targeted interventions can restore coverage without sacrificing accuracy.","ai_keywords":["RLVR","PPO","GRPO","Countdown task","solution coverage","entrance families","per-token likelihood","parameter interpolation","entropy collapse","SFT-DPO-RLVR"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"66a10360f1fe01fc346b0772","name":"UniversityofBirmingham","fullname":"University of Birmingham","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66a10305d79afdd54d920b3c/XerDR4BUpF85rJ_s2Lzc4.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"633f536250d83f5065d28f6d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/633f536250d83f5065d28f6d/eZCdJJpU6OJFYUWYoUAVn.jpeg","isPro":false,"fullname":"Ruizhe Li","user":"rzdiversity","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a6aa05b3550efadfe66fdbf","avatarUrl":"/avatars/4232e88a186fb67e814d28ed96574c02.svg","isPro":false,"fullname":"Joseph Jackson","user":"driftwisp","type":"user"},{"_id":"6a6c809402e1b4f71ce1b8c1","avatarUrl":"/avatars/c37fe722d7a515512b8d19ae39db9b27.svg","isPro":false,"fullname":"Edward Thomas","user":"PrismStack39","type":"user"},{"_id":"6a6da8c0f372a51769692d07","avatarUrl":"/avatars/5b90452cb5968daa2d4822e2e9fc677a.svg","isPro":false,"fullname":"Robert Martinez","user":"IndigoPulse","type":"user"},{"_id":"6a6decee9caa2ff7db7a7c19","avatarUrl":"/avatars/761577ab980666cc634f42048cc00b82.svg","isPro":false,"fullname":"Joshua Clark","user":"ZenithVale","type":"user"},{"_id":"6a9ab7969a727a5f5bc59ae0","avatarUrl":"/avatars/a4b0ff987e74607bf28c72f4a80ec586.svg","isPro":false,"fullname":"Robert Walters","user":"leemary","type":"user"},{"_id":"6a9b54e1eed21fdbb9ecb120","avatarUrl":"/avatars/3c8330aa1100a9b771df3a25f2143312.svg","isPro":false,"fullname":"송영환","user":"Pixel-Remy","type":"user"},{"_id":"6819bd8214b764ef8da980da","avatarUrl":"/avatars/c92ba50ae1f43b4395d67a297e2766ce.svg","isPro":false,"fullname":"Zhou Qc","user":"esyd","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66a10360f1fe01fc346b0772","name":"UniversityofBirmingham","fullname":"University of Birmingham","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66a10305d79afdd54d920b3c/XerDR4BUpF85rJ_s2Lzc4.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.29188.md","query":{}}">
Papers
arxiv:2608.29188

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

Published on Aug 29
· Submitted by
Ruizhe Li
on Sep 4
Authors:
,

Abstract

Reinforcement learning with verifiable rewards narrows reasoning diversity primarily at the initial solution step rather than during execution, and targeted interventions can restore coverage without sacrificing accuracy.

Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator, across PPO on Qwen2.5-3B and GRPO on Qwen2.5-3B-Instruct. Across both training setups, solution coverage falls by up to 67%, halving even on problems solved across all checkpoints. We show that this contraction is heavily concentrated at the entrance: per-token likelihood shifts are 11x--16x larger prior to the first arithmetic operation than during downstream reasoning. Supplying only an unselected entrance prefix restores completion rates in low-access families by over an order of magnitude (0.018 -> 0.212 under PPO), demonstrating that alternative solutions remain executable but are no longer initiated. Guided by this localization, we find that while surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1. Finally, we show that early-step entropy collapse recurs across six math benchmarks with 7B and 14B models, but is not an inevitable byproduct of reasoning optimization: an SFT baseline preserves more than double the coverage, and staged SFT--DPO--RLVR pipelines retain early-step entropy. In summary, reasoning breadth is lost at the door, not inside the room. Code: https://github.com/ershiyidian/early-branch-locking.

Community

Paper author Paper submitter about 23 hours ago

RLVR-induced solution collapse is an access failure, not an execution failure: models lose reasoning diversity at initial computational entrance while retaining latent downstream capability. Reasoning breadth is lost at the door, not inside the room.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.29188
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.29188 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.29188 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.29188 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers