Great method</p>\n","updatedAt":"2026-08-12T06:52:23.383Z","author":{"_id":"655b32bf7b098f9cb54571c5","avatarUrl":"/avatars/21c103e4db8d95a74935b2399a684cb4.svg","fullname":"jasca","name":"An-Jack","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7547126412391663},"editors":["An-Jack"],"editorAvatarUrls":["/avatars/21c103e4db8d95a74935b2399a684cb4.svg"],"reactions":[],"isReport":false}},{"id":"6a7c2aceccc1cd992a793150","author":{"_id":"682d3a381760873904707d3a","avatarUrl":"/avatars/c3cde49d1cf913d605641cf0f26d3dd1.svg","fullname":"Shu","name":"gdace829","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-12T08:11:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"good","html":"<p>good</p>\n","updatedAt":"2026-08-12T08:11:58.659Z","author":{"_id":"682d3a381760873904707d3a","avatarUrl":"/avatars/c3cde49d1cf913d605641cf0f26d3dd1.svg","fullname":"Shu","name":"gdace829","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8952815532684326},"editors":["gdace829"],"editorAvatarUrls":["/avatars/c3cde49d1cf913d605641cf0f26d3dd1.svg"],"reactions":[],"isReport":false}},{"id":"6a7c4571cf1ef357b0d48604","author":{"_id":"67864a46d02b5aadd99e41ea","avatarUrl":"/avatars/1774002e38a3f6e5f8ae0432e76a7066.svg","fullname":"Ruihang Xu","name":"ruihangxu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-08-12T10:05:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"nice and efficient method!","html":"<p>nice and efficient method!</p>\n","updatedAt":"2026-08-12T10:05:37.918Z","author":{"_id":"67864a46d02b5aadd99e41ea","avatarUrl":"/avatars/1774002e38a3f6e5f8ae0432e76a7066.svg","fullname":"Ruihang Xu","name":"ruihangxu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8343386054039001},"editors":["ruihangxu"],"editorAvatarUrls":["/avatars/1774002e38a3f6e5f8ae0432e76a7066.svg"],"reactions":[{"reaction":"🤗","users":["ruihangxu","HKX"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.10744","authors":[{"_id":"6a7bdd071653ef87c6af1bdd","name":"Zihao Liu","hidden":false},{"_id":"6a7bdd071653ef87c6af1bde","name":"Xiaolong Shen","hidden":false},{"_id":"6a7bdd071653ef87c6af1bdf","name":"Zhenglin Zhou","hidden":false},{"_id":"6a7bdd071653ef87c6af1be0","name":"Ruijie Quan","hidden":false},{"_id":"6a7bdd071653ef87c6af1be1","name":"Yi Yang","hidden":false}],"publishedAt":"2026-08-11T00:00:00.000Z","submittedOnDailyAt":"2026-08-12T00:00:00.000Z","title":"Beyond Pixels: From Video Priors to 4D Worlds","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.","upvotes":102,"discussionId":"6a7bdd081653ef87c6af1be2","projectPage":"https://hayd-zju.github.io/Beyond-Pixels/","ai_summary":"Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining.","ai_keywords":["4D generation","variational autoencoder","video diffusion transformers","latent-to-4D","spatiotemporal attention","4D decoder"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69bb527b98f36bb29435a9b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Cnwc2nQFw18HMetjC4fvJ.png","isPro":false,"fullname":"孙 子轩","user":"matthewonssd5s","type":"user"},{"_id":"663b6288e6257fa86a2de2ba","avatarUrl":"/avatars/6cf60e8a223425dd7062d06fce3b2a2c.svg","isPro":false,"fullname":"HAYD","user":"HAYD6","type":"user"},{"_id":"6425318d175bd2952281065e","avatarUrl":"/avatars/37deb6ceb1552dece43a1c8c13c1c871.svg","isPro":false,"fullname":"ZhenglinZhou","user":"zhenglin","type":"user"},{"_id":"65ef136cd7d63c2ed088255b","avatarUrl":"/avatars/2d8e2eb1e340060ed8391d47e747e0b8.svg","isPro":false,"fullname":"zhentao tan","user":"tztgo","type":"user"},{"_id":"673d4716cc1ef74a349cd2ad","avatarUrl":"/avatars/a88f1d461c199a2caa1d5e13b70921fe.svg","isPro":false,"fullname":"Yixuan Han","user":"yixuan7878","type":"user"},{"_id":"69638860aa6b67e1b631466a","avatarUrl":"/avatars/b850fb1eab5146ecaf77e6f18046caf8.svg","isPro":false,"fullname":"Canlin Luo","user":"luocanlin","type":"user"},{"_id":"6578102dee33d547aec44c8f","avatarUrl":"/avatars/8f1d0de6dbe3e7609f6006af3d84833b.svg","isPro":false,"fullname":"Wang Tong","user":"TommyIsNotHere","type":"user"},{"_id":"6a4a562c96095a21aaccc0f8","avatarUrl":"/avatars/f313fdf5215954b42b2ba7f7d854b6a1.svg","isPro":false,"fullname":"iABF","user":"iiiABF","type":"user"},{"_id":"6719db917a0d8c16a25807e0","avatarUrl":"/avatars/9262aaf0feb80cf70ab2b28f39da9bc3.svg","isPro":false,"fullname":"Hong Jiang","user":"123123aa123","type":"user"},{"_id":"68fc7cd0b3da2b6da7593015","avatarUrl":"/avatars/456a9d54689dbe020d7a60b3aff5f548.svg","isPro":false,"fullname":"sanity","user":"sanity2025","type":"user"},{"_id":"6a7c12090ecf195724f965ea","avatarUrl":"/avatars/1645d60fc4ca51c374d34ba89a17e2d8.svg","isPro":false,"fullname":"Alice Chen","user":"encountersunshine","type":"user"},{"_id":"65dc74d258ea0eec69e88e3a","avatarUrl":"/avatars/d72eda1dd0d324d6c0cdefb527f7b1d4.svg","isPro":false,"fullname":"JI YUAN HU","user":"little612pea","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.10744.md","query":{}}">
Beyond Pixels: From Video Priors to 4D Worlds
Abstract
Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining.
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.
Community
nice and efficient method!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.10744 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.10744 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.10744 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.