Hugging Face Daily Papers · · 3 min read

EchoWM: Open and Enterable Omnimodal World Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

An omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music, and speech evolve together.</p>\n","updatedAt":"2026-08-25T07:14:24.146Z","author":{"_id":"6411c801e872ae3fb1e2c96e","avatarUrl":"/avatars/f8898dc13d700e545eedbbfab1c18353.svg","fullname":"Franklin","name":"Franklinzhang","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8951196670532227},"editors":["Franklinzhang"],"editorAvatarUrls":["/avatars/f8898dc13d700e545eedbbfab1c18353.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23189","authors":[{"_id":"6a8d10005add2537c32e9747","name":"Songchun Zhang","hidden":false},{"_id":"6a8d10005add2537c32e9748","name":"Yaowei Li","hidden":false},{"_id":"6a8d10005add2537c32e9749","name":"Junhao Zhuang","hidden":false},{"_id":"6a8d10005add2537c32e974a","user":{"_id":"66608add236f958513d21d2e","avatarUrl":"/avatars/53eca0891c98cbb93be899885160a983.svg","isPro":false,"fullname":"Weiyang Jin","user":"Wayne-King","type":"user","name":"Wayne-King"},"name":"Weiyang Jin","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:45.197Z","hidden":false},{"_id":"6a8d10005add2537c32e974b","name":"Haoyu Wang","hidden":false},{"_id":"6a8d10005add2537c32e974c","name":"Xin Lu","hidden":false},{"_id":"6a8d10005add2537c32e974d","name":"Yilang Sun","hidden":false},{"_id":"6a8d10005add2537c32e974e","name":"Shiyi Zhang","hidden":false},{"_id":"6a8d10005add2537c32e974f","name":"Haoran Li","hidden":false},{"_id":"6a8d10005add2537c32e9750","name":"Xiaoxiao Ma","hidden":false},{"_id":"6a8d10005add2537c32e9751","name":"Yuming Li","hidden":false},{"_id":"6a8d10005add2537c32e9752","name":"Yijun Liu","hidden":false},{"_id":"6a8d10005add2537c32e9753","name":"Yaofeng Su","hidden":false},{"_id":"6a8d10005add2537c32e9754","name":"Yanwen Ma","hidden":false},{"_id":"6a8d10005add2537c32e9755","name":"Haoyu Wu","hidden":false},{"_id":"6a8d10005add2537c32e9756","name":"Zihan Su","hidden":false},{"_id":"6a8d10005add2537c32e9757","name":"Yue Ma","hidden":false},{"_id":"6a8d10005add2537c32e9758","name":"Lvmin Zhang","hidden":false},{"_id":"6a8d10005add2537c32e9759","name":"Haoyang Huang","hidden":false},{"_id":"6a8d10005add2537c32e975a","name":"Zeyue Xue","hidden":false},{"_id":"6a8d10005add2537c32e975b","name":"Anyi Rao","hidden":false},{"_id":"6a8d10005add2537c32e975c","name":"Nan Duan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6411c801e872ae3fb1e2c96e/tFKX_WWKtU2PBi-ojObJW.mp4"],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-25T00:00:00.000Z","title":"EchoWM: Open and Enterable Omnimodal World Models","submittedOnDailyBy":{"_id":"6411c801e872ae3fb1e2c96e","avatarUrl":"/avatars/f8898dc13d700e545eedbbfab1c18353.svg","isPro":true,"fullname":"Franklin","user":"Franklinzhang","type":"user","name":"Franklinzhang"},"summary":"We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \\model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.","upvotes":48,"discussionId":"6a8d10005add2537c32e975d","projectPage":"https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/wm/","githubRepo":"https://github.com/jd-opensource/JoyAI-Echo","githubRepoAddedBy":"user","ai_summary":"EchoWM is an omnimodal world model that generates synchronized high-resolution video, sound, music, and speech while following continuous 6-DoF navigation trajectories across first- and third-person views.","ai_keywords":["omnimodal world model","enterable generative media","camera intent","6-DoF trajectory","dataset-level calibration","audio-visual generation","progressive training","autoregressive post-training","long-horizon generation","world-model benchmarks"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1882},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user"},{"_id":"6362801380c1a705a6ea54ac","avatarUrl":"/avatars/041ad5abf9be42e336938f51ebb8746c.svg","isPro":false,"fullname":"Yaowei Li","user":"Yw22","type":"user"},{"_id":"64292edd69bb3e94e99de995","avatarUrl":"/avatars/717c1e364f983c8e5ff91f3abd0fa41c.svg","isPro":false,"fullname":"yuanbo","user":"sumyyyyy","type":"user"},{"_id":"644a16eded295eb43e6815e1","avatarUrl":"/avatars/0ee0f8c649247c1c618111899f0a577d.svg","isPro":false,"fullname":"Jiahao Shao","user":"jhshao","type":"user"},{"_id":"6590f7880c993129053a2344","avatarUrl":"/avatars/5cf5c5185d10feeaecd5e8add7c4330d.svg","isPro":false,"fullname":"Haoyu wu","user":"Haoyuwu","type":"user"},{"_id":"68da9ca61e61717a96d9894a","avatarUrl":"/avatars/f8f33d5639539969daff359cdf5b6469.svg","isPro":true,"fullname":"Yihao Quan","user":"0x33B","type":"user"},{"_id":"66aa39349238d9c3a1c7f9dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aa39349238d9c3a1c7f9dc/mj6r7uxEYXM502x296UMf.jpeg","isPro":false,"fullname":"Xin Jin","user":"Xin1118","type":"user"},{"_id":"6411c801e872ae3fb1e2c96e","avatarUrl":"/avatars/f8898dc13d700e545eedbbfab1c18353.svg","isPro":true,"fullname":"Franklin","user":"Franklinzhang","type":"user"},{"_id":"64b1303bf460afaefcf922c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/0zhdXXravx0CKJLTSVDwU.jpeg","isPro":false,"fullname":"Huan-ang Gao","user":"c7w","type":"user"},{"_id":"65781534e390cfd40998d7af","avatarUrl":"/avatars/85d3be7f74b9d959ba9b3ccc04398536.svg","isPro":false,"fullname":"Hongyang Wei","user":"nonwhy","type":"user"},{"_id":"658ad59e304552ba0c034d35","avatarUrl":"/avatars/ab0c9978b774e68b4d63eef8cca4417c.svg","isPro":false,"fullname":"Lu Dai","user":"stellaludai","type":"user"},{"_id":"64560c21babbbbd3486df362","avatarUrl":"/avatars/69a37786be8d5ec964d25a565afa3103.svg","isPro":false,"fullname":"Xu Huang","user":"Savoia","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23189.md","query":{}}">
Papers
arxiv:2608.23189

EchoWM: Open and Enterable Omnimodal World Models

Published on Aug 24
· Submitted by
Franklin
on Aug 25
#3 Paper of the day
Authors:
,

Abstract

EchoWM is an omnimodal world model that generates synchronized high-resolution video, sound, music, and speech while following continuous 6-DoF navigation trajectories across first- and third-person views.

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.

Community

An omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music, and speech evolve together.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23189
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.23189 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.23189 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23189 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers