Hugging Face Daily Papers · · 3 min read

Wonder: Video World Model Done Better

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<video src=\"https://cdn-uploads.huggingface.co/production/uploads/68a34b763d0ab4811c881195/M6Gt3ueffojoixIcI3xvF.webm\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-29T04:52:41.472Z","author":{"_id":"68a34b763d0ab4811c881195","avatarUrl":"/avatars/6d2e2957dcd54887881fdf028a2c21b3.svg","fullname":"alq","name":"somnvz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.278155654668808},"editors":["somnvz"],"editorAvatarUrls":["/avatars/6d2e2957dcd54887881fdf028a2c21b3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.26037","authors":[{"_id":"6a69618f9d3a1231d492b7b6","name":"Jiacong Xu","hidden":false},{"_id":"6a69618f9d3a1231d492b7b7","name":"Hanwen Jiang","hidden":false},{"_id":"6a69618f9d3a1231d492b7b8","name":"Zhixin Shu","hidden":false},{"_id":"6a69618f9d3a1231d492b7b9","name":"Kalyan Sunkavalli","hidden":false},{"_id":"6a69618f9d3a1231d492b7ba","name":"Vishal M. Patel","hidden":false},{"_id":"6a69618f9d3a1231d492b7bb","name":"Yiqun Mei","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/J4rEn3_UurIFAFXZfsiHd.mp4"],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-29T00:00:00.000Z","title":"Wonder: Video World Model Done Better","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.","upvotes":7,"discussionId":"6a69618f9d3a1231d492b7bc","projectPage":"https://wonder-world-model.github.io/","organization":{"_id":"61e5d14f77496de0a6d95c6b","name":"adobe","fullname":"Adobe","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1645217431826-61e35e517ac6b6d06cfa8081.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ca8e060609f1def7e6548a","avatarUrl":"/avatars/1da7947840cb87d5f77c0af9ee11f9c2.svg","isPro":true,"fullname":"Yi Jung","user":"YJ-142150","type":"user"},{"_id":"67f5e63688b2c5303ab5be7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f5e63688b2c5303ab5be7a/QSH8-QZH6l3KradXqNxJT.png","isPro":false,"fullname":"Chengxuan Qian","user":"Raymond-Qiancx","type":"user"},{"_id":"647f01d3838ac3601fc6c17b","avatarUrl":"/avatars/9605767bdf8a896b9c99f974797715c5.svg","isPro":false,"fullname":"Wangjianxiong","user":"wjx138819","type":"user"},{"_id":"65ce020f8b4adee8718cbcff","avatarUrl":"/avatars/b82f2286fd4f335dbe4bf5fe86d39410.svg","isPro":false,"fullname":"yusong hu","user":"songsongh","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"6848933f729cfa426d465951","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/IjRV2AW77lTJ3CZtyixfe.png","isPro":false,"fullname":"Maya","user":"Yamanjuinc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61e5d14f77496de0a6d95c6b","name":"adobe","fullname":"Adobe","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1645217431826-61e35e517ac6b6d06cfa8081.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.26037.md","query":{}}">
Papers
arxiv:2607.26037

Wonder: Video World Model Done Better

Published on Jul 28
· Submitted by
taesiri
on Jul 29
Authors:
,

Abstract

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.26037
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.26037 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.26037 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.26037 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers