We introduce <strong>SpectraReward</strong>, a <strong>training-free reward function</strong> for image-generation reinforcement learning.</p>\n<ul>\n<li><strong>Prompt recovery as reward:</strong> We measure how well an MLLM can recover the original prompt from a generated image.</li>\n<li><strong>No additional training:</strong> We require <strong>no preference labels, reward-model fine-tuning, or decomposed verification questions</strong>.</li>\n<li><strong>Closed-loop self-improvement:</strong> With <strong>Self-SpectraReward</strong>, a unified multimodal model uses its own understanding branch to reward its generation branch.</li>\n<li><strong>Broad evaluation:</strong> We study <strong>2 diffusion models, 3 RL algorithms, 9 MLLMs ranging from 4B to 235B, and 5 out-of-distribution benchmarks</strong>.</li>\n<li><strong>Key finding:</strong> <strong>Larger reward models are not always better. Reward-policy alignment matters more.</strong></li>\n</ul>\n<p>Project Page: <a href=\"https://huangrh99.github.io/SpectraReward/\" rel=\"nofollow\">https://huangrh99.github.io/SpectraReward/</a></p>\n","updatedAt":"2026-07-15T02:35:43.345Z","author":{"_id":"630f0542cc8ed75decb03b68","avatarUrl":"/avatars/f76c3603b2700591f33a8a931f7ca664.svg","fullname":"huangrh9","name":"huangrh9","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7887012362480164},"editors":["huangrh9"],"editorAvatarUrls":["/avatars/f76c3603b2700591f33a8a931f7ca664.svg"],"reactions":[],"isReport":false}},{"id":"6a56f32d99e0dbd7f9d4bfb4","author":{"_id":"665eab8bf8cb81b0a4915cdc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S2sg92CriAvhSJ4U1b4F2.png","fullname":"Trump","name":"willaihappy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-07-15T02:40:45.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Great Work!","html":"<p>Great Work!</p>\n","updatedAt":"2026-07-15T02:40:45.003Z","author":{"_id":"665eab8bf8cb81b0a4915cdc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S2sg92CriAvhSJ4U1b4F2.png","fullname":"Trump","name":"willaihappy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5589598417282104},"editors":["willaihappy"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S2sg92CriAvhSJ4U1b4F2.png"],"reactions":[{"reaction":"🚀","users":["willaihappy","chandan-ku"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.11886","authors":[{"_id":"6a56ef3ae548eb96f98cca64","name":"Runhui Huang","hidden":false},{"_id":"6a56ef3ae548eb96f98cca65","name":"Qihui Zhang","hidden":false},{"_id":"6a56ef3ae548eb96f98cca66","name":"Zhe Liu","hidden":false},{"_id":"6a56ef3ae548eb96f98cca67","name":"Yu Gao","hidden":false},{"_id":"6a56ef3ae548eb96f98cca68","name":"Jie Wu","hidden":false},{"_id":"6a56ef3ae548eb96f98cca69","name":"Hengshuang Zhao","hidden":false}],"publishedAt":"2026-07-13T00:00:00.000Z","submittedOnDailyAt":"2026-07-15T00:00:00.000Z","title":"Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation","submittedOnDailyBy":{"_id":"630f0542cc8ed75decb03b68","avatarUrl":"/avatars/f76c3603b2700591f33a8a931f7ca664.svg","isPro":false,"fullname":"huangrh9","user":"huangrh9","type":"user","name":"huangrh9"},"summary":"In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: https://huangrh99.github.io/SpectraReward/","upvotes":29,"discussionId":"6a56ef3ae548eb96f98cca6a","projectPage":"https://huangrh99.github.io/SpectraReward/","organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d39e72df34ad328c20f0c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/qQBezE-XbMyow_gbtBvXA.png","isPro":false,"fullname":"Ruizhe Zhong","user":"Yukino271828","type":"user"},{"_id":"630f0542cc8ed75decb03b68","avatarUrl":"/avatars/f76c3603b2700591f33a8a931f7ca664.svg","isPro":false,"fullname":"huangrh9","user":"huangrh9","type":"user"},{"_id":"665eab8bf8cb81b0a4915cdc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S2sg92CriAvhSJ4U1b4F2.png","isPro":false,"fullname":"Trump","user":"willaihappy","type":"user"},{"_id":"641d6c27c5f150c3af0bf879","avatarUrl":"/avatars/39808793703406b4fe997e351b6a6bfb.svg","isPro":false,"fullname":"Ray Yang","user":"rayruiyang","type":"user"},{"_id":"6507f434dacc94cd6c2a8700","avatarUrl":"/avatars/71bf2dff7119f9fd2c01ddf925229239.svg","isPro":false,"fullname":"Minbin Huang","user":"centaurus-alpha","type":"user"},{"_id":"656b2c240bbc114fe6e29e05","avatarUrl":"/avatars/ad28bc90a3820265501b95bfa5bb9418.svg","isPro":false,"fullname":"Zhou","user":"Zanwei","type":"user"},{"_id":"66a0c72a813431cd7aa0fdf6","avatarUrl":"/avatars/52a97401826f124030aadef094dc725f.svg","isPro":true,"fullname":"Kun Xiang","user":"Kun-Xiang","type":"user"},{"_id":"63aaf2a2a4bdd629b7eb2b5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63aaf2a2a4bdd629b7eb2b5b/WOa3nAUNy5D3MsFUV9B8Z.jpeg","isPro":false,"fullname":"Junyi Li","user":"ProvenceStar","type":"user"},{"_id":"64dd8a1e22f93c32882eb2fe","avatarUrl":"/avatars/a5e84a5c5a14058d91c342f8c89a36a2.svg","isPro":false,"fullname":"yanxin long","user":"yestinl","type":"user"},{"_id":"64e9c855233101ed99ca2315","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/ozwIxW0ofJMR6kj623p1T.png","isPro":false,"fullname":"YANG, Zhenya","user":"ANIYA673","type":"user"},{"_id":"675acaa3e21ed19ca522973e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/t0PqwfaFbVQ9R-cr_Lzzs.png","isPro":false,"fullname":"Fangyu","user":"rslinfy","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.11886.md","query":{}}">
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
Abstract
In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: https://huangrh99.github.io/SpectraReward/
Community
We introduce SpectraReward, a training-free reward function for image-generation reinforcement learning.
- Prompt recovery as reward: We measure how well an MLLM can recover the original prompt from a generated image.
- No additional training: We require no preference labels, reward-model fine-tuning, or decomposed verification questions.
- Closed-loop self-improvement: With Self-SpectraReward, a unified multimodal model uses its own understanding branch to reward its generation branch.
- Broad evaluation: We study 2 diffusion models, 3 RL algorithms, 9 MLLMs ranging from 4B to 235B, and 5 out-of-distribution benchmarks.
- Key finding: Larger reward models are not always better. Reward-policy alignment matters more.
Project Page: https://huangrh99.github.io/SpectraReward/
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.11886 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.11886 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.11886 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.