Parameter Exploration for RLVR via Variational Learning</p>\n","updatedAt":"2026-08-13T10:31:21.763Z","author":{"_id":"65e5c02e4664a2c9e97e34be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65e5c02e4664a2c9e97e34be/uH6Vrf_NTXxzrmXDzaqC6.jpeg","fullname":"Vatsal Venkatkrishna","name":"wetsoledrysoul","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.573104202747345},"editors":["wetsoledrysoul"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65e5c02e4664a2c9e97e34be/uH6Vrf_NTXxzrmXDzaqC6.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09805","authors":[{"_id":"6a7d9c3c42823931a1f173ba","name":"Vatsal Venkatkrishna","hidden":false},{"_id":"6a7d9c3c42823931a1f173bb","name":"Nico Daheim","hidden":false},{"_id":"6a7d9c3c42823931a1f173bc","name":"Iryna Gurevych","hidden":false}],"publishedAt":"2026-08-10T16:28:07.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"Parameter Exploration for RLVR via Variational Learning","submittedOnDailyBy":{"_id":"65e5c02e4664a2c9e97e34be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65e5c02e4664a2c9e97e34be/uH6Vrf_NTXxzrmXDzaqC6.jpeg","isPro":true,"fullname":"Vatsal Venkatkrishna","user":"wetsoledrysoul","type":"user","name":"wetsoledrysoul"},"summary":"Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.","upvotes":1,"discussionId":"6a7d9c3c42823931a1f173bd","githubRepo":"https://github.com/insait-institute/C3PO","githubRepoAddedBy":"user","ai_summary":"Parameter-space exploration via perturbed policy sampling improves LLM reinforcement learning by diversifying rollouts and reducing training failures compared to action-space methods.","ai_keywords":["parameter-space exploration","Perturbed Parameter Policy Optimization (3PO)","GRPO","rollout grouping","reward estimation","zero-advantage groups","action-space exploration","temperature scaling","posterior sampling","LLM reinforcement learning"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"69690d291ab46be99d235ea1","name":"BayesRL","fullname":"BayesRL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65e5c02e4664a2c9e97e34be/g46lfnQQ9hRCXJZ1SfmCO.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65e5c02e4664a2c9e97e34be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65e5c02e4664a2c9e97e34be/uH6Vrf_NTXxzrmXDzaqC6.jpeg","isPro":true,"fullname":"Vatsal Venkatkrishna","user":"wetsoledrysoul","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69690d291ab46be99d235ea1","name":"BayesRL","fullname":"BayesRL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65e5c02e4664a2c9e97e34be/g46lfnQQ9hRCXJZ1SfmCO.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09805.md","query":{}}">
Parameter Exploration for RLVR via Variational Learning
Abstract
Parameter-space exploration via perturbed policy sampling improves LLM reinforcement learning by diversifying rollouts and reducing training failures compared to action-space methods.
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.
Community
Parameter Exploration for RLVR via Variational Learning
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.09805 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.09805 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.