Hugging Face Daily Papers · · 3 min read

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video</p>\n","updatedAt":"2026-09-15T06:59:28.349Z","author":{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","fullname":"Chuanhao","name":"ChuanhaoLi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6883113980293274},"editors":["ChuanhaoLi"],"editorAvatarUrls":["/avatars/86fd4c6d7d33435de49b35659bf65265.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14462","authors":[{"_id":"6aa8eb7e5dd4cb9b4cc02a6e","name":"Jiaming Tan","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a6f","name":"Mingliang Zhai","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a70","name":"Zhen Li","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a71","name":"Yuwei Wu","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a72","name":"Chuanhao Li","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a73","name":"Kaipeng Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/674ea59a8f2e7614a6c72f26/ELktmTq9YLja36YHmCR5X.mp4"],"publishedAt":"2026-09-13T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video","submittedOnDailyBy":{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","isPro":false,"fullname":"Chuanhao","user":"ChuanhaoLi","type":"user","name":"ChuanhaoLi"},"summary":"Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.","upvotes":14,"discussionId":"6aa8eb7e5dd4cb9b4cc02a74","projectPage":"https://alaya-lab.github.io/AlayaVista/","githubRepo":"https://github.com/AlayaLab/AlayaVista","githubRepoAddedBy":"user","ai_summary":"AlayaVista decouples panoramic scene evolution from perspective video synthesis to enable efficient, high-fidelity interactive world modeling, supported by the MUGEN dataset.","ai_keywords":["camera-controllable streaming video world model","panoramic latent state","latent viewport renderer","perspective refiner","chunk-autoregressive generation","panorama expansion model","MUGEN","panoramic video dataset"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":15,"organization":{"_id":"689f08c50df4fcf7fddc0b08","name":"AlayaLab","fullname":"Alaya Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63342778d92c5842ae728aef/dNCvNz9MMshksG2xspIbM.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","isPro":false,"fullname":"Chuanhao","user":"ChuanhaoLi","type":"user"},{"_id":"6369fcae64aad59d4d45c848","avatarUrl":"/avatars/eafb0ef55d222b3c52852cf021408859.svg","isPro":false,"fullname":"zhai mingliang","user":"zmling","type":"user"},{"_id":"659fa7002c4538d113c7296a","avatarUrl":"/avatars/eff46ce65ce83e09b8223bad335cd0fc.svg","isPro":false,"fullname":"Ziqi Cai","user":"GhostCai","type":"user"},{"_id":"66040e5acfb9b90f95a31d76","avatarUrl":"/avatars/bd5b8a4c38aac0531126fd036debb5ed.svg","isPro":false,"fullname":"Zheng-Hui Huang","user":"Brian9999","type":"user"},{"_id":"6511689bcac39a7d888fa6ce","avatarUrl":"/avatars/1d8ed283b8474b0ea211f349d06926d8.svg","isPro":false,"fullname":"Ruicong Liu","user":"MickeyLLG","type":"user"},{"_id":"6307a98795b2ab342fec0cf7","avatarUrl":"/avatars/85b261bcdda4717a6e40491f6c7b7a89.svg","isPro":false,"fullname":"Zhixiang Wang","user":"wangzx1994","type":"user"},{"_id":"66e7de0b3340bbe522ba57e8","avatarUrl":"/avatars/b296607096ee3f43ee274a3aa68bb521.svg","isPro":false,"fullname":"joseph_lin","user":"JosephLin1999","type":"user"},{"_id":"6672fe26c33b5004b69a1d6a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Ff8cOS6Y0TPUSihx_hOMe.png","isPro":false,"fullname":"YouZhe","user":"YouZhe","type":"user"},{"_id":"6428591f53b748123d4c95bf","avatarUrl":"/avatars/ad9800c1588c30a877fc74fe1e54620b.svg","isPro":false,"fullname":"solytia","user":"solytia","type":"user"},{"_id":"6352593e507b679c3c5bf5dc","avatarUrl":"/avatars/eab2eced53cc20cddd5ea3b89ff6d14c.svg","isPro":false,"fullname":"Chenchen Jing","user":"zxs1996","type":"user"},{"_id":"68323f961e5e5c17eb1f0de4","avatarUrl":"/avatars/e0b56c721c2aec1daf52b67f05093a2c.svg","isPro":false,"fullname":"sue","user":"jzf0634","type":"user"},{"_id":"659cb2671d398a2381625b2f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/659cb2671d398a2381625b2f/-y_odTNSgvlcABJ-9GeFf.jpeg","isPro":false,"fullname":"SII-YuanyangYin","user":"SII-YuanyangYin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"689f08c50df4fcf7fddc0b08","name":"AlayaLab","fullname":"Alaya Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63342778d92c5842ae728aef/dNCvNz9MMshksG2xspIbM.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14462.md","query":{}}">
Papers
arxiv:2609.14462

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

Published on Sep 13
· Submitted by
Chuanhao
on Sep 15
Authors:
,

Abstract

AlayaVista decouples panoramic scene evolution from perspective video synthesis to enable efficient, high-fidelity interactive world modeling, supported by the MUGEN dataset.

Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.

Community

Paper submitter about 1 hour ago

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.14462
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.14462 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.14462 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.14462 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers