AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video</p>\n","updatedAt":"2026-09-15T06:59:28.349Z","author":{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","fullname":"Chuanhao","name":"ChuanhaoLi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6883113980293274},"editors":["ChuanhaoLi"],"editorAvatarUrls":["/avatars/86fd4c6d7d33435de49b35659bf65265.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14462","authors":[{"_id":"6aa8eb7e5dd4cb9b4cc02a6e","name":"Jiaming Tan","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a6f","name":"Mingliang Zhai","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a70","name":"Zhen Li","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a71","name":"Yuwei Wu","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a72","name":"Chuanhao Li","hidden":false},{"_id":"6aa8eb7e5dd4cb9b4cc02a73","name":"Kaipeng Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/674ea59a8f2e7614a6c72f26/ELktmTq9YLja36YHmCR5X.mp4"],"publishedAt":"2026-09-13T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video","submittedOnDailyBy":{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","isPro":false,"fullname":"Chuanhao","user":"ChuanhaoLi","type":"user","name":"ChuanhaoLi"},"summary":"Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.","upvotes":14,"discussionId":"6aa8eb7e5dd4cb9b4cc02a74","projectPage":"https://alaya-lab.github.io/AlayaVista/","githubRepo":"https://github.com/AlayaLab/AlayaVista","githubRepoAddedBy":"user","ai_summary":"AlayaVista decouples panoramic scene evolution from perspective video synthesis to enable efficient, high-fidelity interactive world modeling, supported by the MUGEN dataset.","ai_keywords":["camera-controllable streaming video world model","panoramic latent state","latent viewport renderer","perspective refiner","chunk-autoregressive generation","panorama expansion model","MUGEN","panoramic video dataset"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":15,"organization":{"_id":"689f08c50df4fcf7fddc0b08","name":"AlayaLab","fullname":"Alaya Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63342778d92c5842ae728aef/dNCvNz9MMshksG2xspIbM.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","isPro":false,"fullname":"Chuanhao","user":"ChuanhaoLi","type":"user"},{"_id":"6369fcae64aad59d4d45c848","avatarUrl":"/avatars/eafb0ef55d222b3c52852cf021408859.svg","isPro":false,"fullname":"zhai mingliang","user":"zmling","type":"user"},{"_id":"659fa7002c4538d113c7296a","avatarUrl":"/avatars/eff46ce65ce83e09b8223bad335cd0fc.svg","isPro":false,"fullname":"Ziqi Cai","user":"GhostCai","type":"user"},{"_id":"66040e5acfb9b90f95a31d76","avatarUrl":"/avatars/bd5b8a4c38aac0531126fd036debb5ed.svg","isPro":false,"fullname":"Zheng-Hui Huang","user":"Brian9999","type":"user"},{"_id":"6511689bcac39a7d888fa6ce","avatarUrl":"/avatars/1d8ed283b8474b0ea211f349d06926d8.svg","isPro":false,"fullname":"Ruicong Liu","user":"MickeyLLG","type":"user"},{"_id":"6307a98795b2ab342fec0cf7","avatarUrl":"/avatars/85b261bcdda4717a6e40491f6c7b7a89.svg","isPro":false,"fullname":"Zhixiang Wang","user":"wangzx1994","type":"user"},{"_id":"66e7de0b3340bbe522ba57e8","avatarUrl":"/avatars/b296607096ee3f43ee274a3aa68bb521.svg","isPro":false,"fullname":"joseph_lin","user":"JosephLin1999","type":"user"},{"_id":"6672fe26c33b5004b69a1d6a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Ff8cOS6Y0TPUSihx_hOMe.png","isPro":false,"fullname":"YouZhe","user":"YouZhe","type":"user"},{"_id":"6428591f53b748123d4c95bf","avatarUrl":"/avatars/ad9800c1588c30a877fc74fe1e54620b.svg","isPro":false,"fullname":"solytia","user":"solytia","type":"user"},{"_id":"6352593e507b679c3c5bf5dc","avatarUrl":"/avatars/eab2eced53cc20cddd5ea3b89ff6d14c.svg","isPro":false,"fullname":"Chenchen Jing","user":"zxs1996","type":"user"},{"_id":"68323f961e5e5c17eb1f0de4","avatarUrl":"/avatars/e0b56c721c2aec1daf52b67f05093a2c.svg","isPro":false,"fullname":"sue","user":"jzf0634","type":"user"},{"_id":"659cb2671d398a2381625b2f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/659cb2671d398a2381625b2f/-y_odTNSgvlcABJ-9GeFf.jpeg","isPro":false,"fullname":"SII-YuanyangYin","user":"SII-YuanyangYin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"689f08c50df4fcf7fddc0b08","name":"AlayaLab","fullname":"Alaya Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63342778d92c5842ae728aef/dNCvNz9MMshksG2xspIbM.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14462.md","query":{}}">
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Abstract
AlayaVista decouples panoramic scene evolution from perspective video synthesis to enable efficient, high-fidelity interactive world modeling, supported by the MUGEN dataset.
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.
Community
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.14462 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.14462 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.14462 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.