Hugging Face Daily Papers · · 4 min read

WorldReward: Reward Modeling for Camera-Conditioned World Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

WorldReward: Reward Modeling for Camera-Conditioned World Models</p>\n","updatedAt":"2026-09-04T03:09:57.219Z","author":{"_id":"654c6845bac6e6e49895a5b5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KXQaAxulqr8jNBSpEaYM4.png","fullname":"SII-Yibin Wang","name":"CodeGoat24","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":50,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5754196643829346},"editors":["CodeGoat24"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KXQaAxulqr8jNBSpEaYM4.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03952","authors":[{"_id":"6a9a35218f7c3b75572394ae","name":"Yibin Wang","hidden":false},{"_id":"6a9a35218f7c3b75572394af","name":"Zehan Wang","hidden":false},{"_id":"6a9a35218f7c3b75572394b0","name":"Junshu Tang","hidden":false},{"_id":"6a9a35218f7c3b75572394b1","name":"Zhimin Li","hidden":false},{"_id":"6a9a35218f7c3b75572394b2","name":"Yujie Zhou","hidden":false},{"_id":"6a9a35218f7c3b75572394b3","name":"Jiazi Bu","hidden":false},{"_id":"6a9a35218f7c3b75572394b4","name":"Pengyang Ling","hidden":false},{"_id":"6a9a35218f7c3b75572394b5","name":"Feng Han","hidden":false},{"_id":"6a9a35218f7c3b75572394b6","name":"Zhixiong Zhang","hidden":false},{"_id":"6a9a35218f7c3b75572394b7","name":"Long Xing","hidden":false},{"_id":"6a9a35218f7c3b75572394b8","name":"Shengyuan Ding","hidden":false},{"_id":"6a9a35218f7c3b75572394b9","name":"Ziang Li","hidden":false},{"_id":"6a9a35218f7c3b75572394ba","name":"Cheng Jin","hidden":false},{"_id":"6a9a35218f7c3b75572394bb","user":{"_id":"63859cf3b2906edaf83af9f0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63859cf3b2906edaf83af9f0/kajwuVzd4pDucSPlwghxo.png","isPro":true,"fullname":"Yuhang Zang","user":"yuhangzang","type":"user","name":"yuhangzang"},"name":"Yuhang Zang","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.275Z","hidden":false},{"_id":"6a9a35218f7c3b75572394bc","name":"Jiaqi Wang","hidden":false},{"_id":"6a9a35218f7c3b75572394bd","name":"Tianyu Pang","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"WorldReward: Reward Modeling for Camera-Conditioned World Models","submittedOnDailyBy":{"_id":"654c6845bac6e6e49895a5b5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KXQaAxulqr8jNBSpEaYM4.png","isPro":false,"fullname":"SII-Yibin Wang","user":"CodeGoat24","type":"user","name":"CodeGoat24"},"summary":"Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.","upvotes":7,"discussionId":"6a9a35218f7c3b75572394be","projectPage":"https://codegoat24.github.io/WorldReward/","ai_summary":"WorldReward is a vision-language reward model that evaluates camera-conditioned world models by aligning video chunks with actions and aggregating preferences for both execution consistency and visual quality.","ai_keywords":["vision-language model","camera-conditioned world models","pairwise preference reward model","action-aligned chunks","structured visual evidence","reasoning-augmented preference dataset","tool-based agent auditing","WorldReward-Bench","RL post-training"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"654c6845bac6e6e49895a5b5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KXQaAxulqr8jNBSpEaYM4.png","isPro":false,"fullname":"SII-Yibin Wang","user":"CodeGoat24","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"661240c953de8ec06a9be648","avatarUrl":"/avatars/4fdded6b35b76b6ffb93bf7f6a3e4f97.svg","isPro":false,"fullname":"Feng Han(SII)","user":"maplebb","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"},{"_id":"6425761a175bd295228311a0","avatarUrl":"/avatars/dcd0d267445563d0616d5a31b5d754b7.svg","isPro":false,"fullname":"zehan wang","user":"sleetwang6","type":"user"},{"_id":"63859cf3b2906edaf83af9f0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63859cf3b2906edaf83af9f0/kajwuVzd4pDucSPlwghxo.png","isPro":true,"fullname":"Yuhang Zang","user":"yuhangzang","type":"user"},{"_id":"646cd947da8e99940b6e55cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646cd947da8e99940b6e55cf/9c0P0WppFqNW9pdo8LgOS.jpeg","isPro":false,"fullname":"Shengyuan Ding","user":"ChrisDing1105","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03952.md","query":{}}">
Papers
arxiv:2609.03952

WorldReward: Reward Modeling for Camera-Conditioned World Models

Published on Sep 3
· Submitted by
SII-Yibin Wang
on Sep 4
Authors:
,

Abstract

WorldReward is a vision-language reward model that evaluates camera-conditioned world models by aligning video chunks with actions and aggregating preferences for both execution consistency and visual quality.

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.

Community

Paper submitter about 6 hours ago

WorldReward: Reward Modeling for Camera-Conditioned World Models

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03952
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.03952 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers