In this paper, we introduce exploration-guided prompt scaffolding for multimodal RL post-training. Our Exploration Potential Score (EPS) identifies low-utility prompts from existing rollouts, guiding a teacher model to rewrite them into more informative training inputs rather than providing answers to imitate. Integrated with GRPO, our approach achieves up to 9.7% relative improvement in-domain, with gains of 11.5% on MathVision and 11.1% on MMMU-Pro.</p>\n<p>🌐 Project page: <a href=\"https://mqleet.github.io/EPS-ProjectPage/\" rel=\"nofollow\">https://mqleet.github.io/EPS-ProjectPage/</a></p>\n","updatedAt":"2026-09-15T02:40:09.170Z","author":{"_id":"6448b2f53e7b3c11be684348","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448b2f53e7b3c11be684348/QvlUQG3pWf8ZyEVBV6F7w.jpeg","fullname":"Qianli Ma","name":"Mqleet","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8431652188301086},"editors":["Mqleet"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6448b2f53e7b3c11be684348/QvlUQG3pWf8ZyEVBV6F7w.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.15051","authors":[{"_id":"6aa8a1cc5dd4cb9b4cc0248e","name":"Yuanhao Yue","hidden":false},{"_id":"6aa8a1cc5dd4cb9b4cc0248f","name":"Qianli Ma","hidden":false},{"_id":"6aa8a1cc5dd4cb9b4cc02490","name":"Chengyu Wang","hidden":false},{"_id":"6aa8a1cc5dd4cb9b4cc02491","name":"Haoting Wang","hidden":false},{"_id":"6aa8a1cc5dd4cb9b4cc02492","name":"Lei Shen","hidden":false},{"_id":"6aa8a1cc5dd4cb9b4cc02493","name":"Jun Huang","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training","submittedOnDailyBy":{"_id":"6448b2f53e7b3c11be684348","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448b2f53e7b3c11be684348/QvlUQG3pWf8ZyEVBV6F7w.jpeg","isPro":true,"fullname":"Qianli Ma","user":"Mqleet","type":"user","name":"Mqleet"},"summary":"Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard training. We propose an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs). Central to our approach is the Exploration Potential Score (EPS), a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory, computable directly from on-policy rollout statistics without additional overhead. Rather than discarding low-utility prompts, we use a teacher model to generate scaffolded rewrites that preserve the original task intent while making subsequent training more informative, reframing teacher supervision as training-data refinement rather than output imitation. Integrated with GRPO on Geo3K and MMK12, our method consistently outperforms the baseline on both in-domain and out-of-distribution benchmarks, achieving up to 9.7\\% relative improvement in-domain and gains of 11.5\\% on MathVision and 11.1\\% on MMMU-Pro.","upvotes":4,"discussionId":"6aa8a1cc5dd4cb9b4cc02494","projectPage":"https://mqleet.github.io/EPS-ProjectPage/","ai_summary":"The framework dynamically adjusts training prompts via exploration potential scoring and scaffolded rewrites to improve reinforcement learning for multimodal language models.","ai_keywords":["online reinforcement learning","exploration-guided prompt scaffolding","Exploration Potential Score","KL-regularized policy improvement","multimodal large language models","GRPO","scaffolded rewrites"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"625e37fb943e346492b8feed","name":"alibaba-pai","fullname":"Alibaba-PAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1650342890988-623c6253389748c9f72ca287.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6448b2f53e7b3c11be684348","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448b2f53e7b3c11be684348/QvlUQG3pWf8ZyEVBV6F7w.jpeg","isPro":true,"fullname":"Qianli Ma","user":"Mqleet","type":"user"},{"_id":"62aba5ebab9ed4f63c36b1e2","avatarUrl":"/avatars/c1f53a6c3ca558f82d102dd8995a530f.svg","isPro":false,"fullname":"Yue","user":"Bohr","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"6aa8b1683910e2035c1c266b","avatarUrl":"/avatars/ad8f044a37d83c49e696713373a712d5.svg","isPro":false,"fullname":"Haoting Wang","user":"HaotingWang66","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"625e37fb943e346492b8feed","name":"alibaba-pai","fullname":"Alibaba-PAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1650342890988-623c6253389748c9f72ca287.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.15051.md","query":{}}">
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
Abstract
The framework dynamically adjusts training prompts via exploration potential scoring and scaffolded rewrites to improve reinforcement learning for multimodal language models.
Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard training. We propose an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs). Central to our approach is the Exploration Potential Score (EPS), a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory, computable directly from on-policy rollout statistics without additional overhead. Rather than discarding low-utility prompts, we use a teacher model to generate scaffolded rewrites that preserve the original task intent while making subsequent training more informative, reframing teacher supervision as training-data refinement rather than output imitation. Integrated with GRPO on Geo3K and MMK12, our method consistently outperforms the baseline on both in-domain and out-of-distribution benchmarks, achieving up to 9.7\% relative improvement in-domain and gains of 11.5\% on MathVision and 11.1\% on MMMU-Pro.
Community
In this paper, we introduce exploration-guided prompt scaffolding for multimodal RL post-training. Our Exploration Potential Score (EPS) identifies low-utility prompts from existing rollouts, guiding a teacher model to rewrite them into more informative training inputs rather than providing answers to imitate. Integrated with GRPO, our approach achieves up to 9.7% relative improvement in-domain, with gains of 11.5% on MathVision and 11.1% on MMMU-Pro.
🌐 Project page: https://mqleet.github.io/EPS-ProjectPage/
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.15051 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.15051 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.15051 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.