Hugging Face Daily Papers · · 3 min read

Beyond Pixels: From Video Priors to 4D Worlds

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Great method</p>\n","updatedAt":"2026-08-12T06:52:23.383Z","author":{"_id":"655b32bf7b098f9cb54571c5","avatarUrl":"/avatars/21c103e4db8d95a74935b2399a684cb4.svg","fullname":"jasca","name":"An-Jack","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7547126412391663},"editors":["An-Jack"],"editorAvatarUrls":["/avatars/21c103e4db8d95a74935b2399a684cb4.svg"],"reactions":[],"isReport":false}},{"id":"6a7c2aceccc1cd992a793150","author":{"_id":"682d3a381760873904707d3a","avatarUrl":"/avatars/c3cde49d1cf913d605641cf0f26d3dd1.svg","fullname":"Shu","name":"gdace829","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-12T08:11:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"good","html":"<p>good</p>\n","updatedAt":"2026-08-12T08:11:58.659Z","author":{"_id":"682d3a381760873904707d3a","avatarUrl":"/avatars/c3cde49d1cf913d605641cf0f26d3dd1.svg","fullname":"Shu","name":"gdace829","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8952815532684326},"editors":["gdace829"],"editorAvatarUrls":["/avatars/c3cde49d1cf913d605641cf0f26d3dd1.svg"],"reactions":[],"isReport":false}},{"id":"6a7c4571cf1ef357b0d48604","author":{"_id":"67864a46d02b5aadd99e41ea","avatarUrl":"/avatars/1774002e38a3f6e5f8ae0432e76a7066.svg","fullname":"Ruihang Xu","name":"ruihangxu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-08-12T10:05:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"nice and efficient method!","html":"<p>nice and efficient method!</p>\n","updatedAt":"2026-08-12T10:05:37.918Z","author":{"_id":"67864a46d02b5aadd99e41ea","avatarUrl":"/avatars/1774002e38a3f6e5f8ae0432e76a7066.svg","fullname":"Ruihang Xu","name":"ruihangxu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8343386054039001},"editors":["ruihangxu"],"editorAvatarUrls":["/avatars/1774002e38a3f6e5f8ae0432e76a7066.svg"],"reactions":[{"reaction":"🤗","users":["ruihangxu","HKX"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.10744","authors":[{"_id":"6a7bdd071653ef87c6af1bdd","name":"Zihao Liu","hidden":false},{"_id":"6a7bdd071653ef87c6af1bde","name":"Xiaolong Shen","hidden":false},{"_id":"6a7bdd071653ef87c6af1bdf","name":"Zhenglin Zhou","hidden":false},{"_id":"6a7bdd071653ef87c6af1be0","name":"Ruijie Quan","hidden":false},{"_id":"6a7bdd071653ef87c6af1be1","name":"Yi Yang","hidden":false}],"publishedAt":"2026-08-11T00:00:00.000Z","submittedOnDailyAt":"2026-08-12T00:00:00.000Z","title":"Beyond Pixels: From Video Priors to 4D Worlds","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.","upvotes":102,"discussionId":"6a7bdd081653ef87c6af1be2","projectPage":"https://hayd-zju.github.io/Beyond-Pixels/","ai_summary":"Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining.","ai_keywords":["4D generation","variational autoencoder","video diffusion transformers","latent-to-4D","spatiotemporal attention","4D decoder"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69bb527b98f36bb29435a9b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Cnwc2nQFw18HMetjC4fvJ.png","isPro":false,"fullname":"孙 子轩","user":"matthewonssd5s","type":"user"},{"_id":"663b6288e6257fa86a2de2ba","avatarUrl":"/avatars/6cf60e8a223425dd7062d06fce3b2a2c.svg","isPro":false,"fullname":"HAYD","user":"HAYD6","type":"user"},{"_id":"6425318d175bd2952281065e","avatarUrl":"/avatars/37deb6ceb1552dece43a1c8c13c1c871.svg","isPro":false,"fullname":"ZhenglinZhou","user":"zhenglin","type":"user"},{"_id":"65ef136cd7d63c2ed088255b","avatarUrl":"/avatars/2d8e2eb1e340060ed8391d47e747e0b8.svg","isPro":false,"fullname":"zhentao tan","user":"tztgo","type":"user"},{"_id":"673d4716cc1ef74a349cd2ad","avatarUrl":"/avatars/a88f1d461c199a2caa1d5e13b70921fe.svg","isPro":false,"fullname":"Yixuan Han","user":"yixuan7878","type":"user"},{"_id":"69638860aa6b67e1b631466a","avatarUrl":"/avatars/b850fb1eab5146ecaf77e6f18046caf8.svg","isPro":false,"fullname":"Canlin Luo","user":"luocanlin","type":"user"},{"_id":"6578102dee33d547aec44c8f","avatarUrl":"/avatars/8f1d0de6dbe3e7609f6006af3d84833b.svg","isPro":false,"fullname":"Wang Tong","user":"TommyIsNotHere","type":"user"},{"_id":"6a4a562c96095a21aaccc0f8","avatarUrl":"/avatars/f313fdf5215954b42b2ba7f7d854b6a1.svg","isPro":false,"fullname":"iABF","user":"iiiABF","type":"user"},{"_id":"6719db917a0d8c16a25807e0","avatarUrl":"/avatars/9262aaf0feb80cf70ab2b28f39da9bc3.svg","isPro":false,"fullname":"Hong Jiang","user":"123123aa123","type":"user"},{"_id":"68fc7cd0b3da2b6da7593015","avatarUrl":"/avatars/456a9d54689dbe020d7a60b3aff5f548.svg","isPro":false,"fullname":"sanity","user":"sanity2025","type":"user"},{"_id":"6a7c12090ecf195724f965ea","avatarUrl":"/avatars/1645d60fc4ca51c374d34ba89a17e2d8.svg","isPro":false,"fullname":"Alice Chen","user":"encountersunshine","type":"user"},{"_id":"65dc74d258ea0eec69e88e3a","avatarUrl":"/avatars/d72eda1dd0d324d6c0cdefb527f7b1d4.svg","isPro":false,"fullname":"JI YUAN HU","user":"little612pea","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.10744.md","query":{}}">
Papers
arxiv:2608.10744

Beyond Pixels: From Video Priors to 4D Worlds

Published on Aug 11
· Submitted by
taesiri
on Aug 12
#3 Paper of the day
Authors:
,

Abstract

Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining.

4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.

Community

nice and efficient method!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.10744
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.10744 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.10744 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.10744 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers