Hugging Face Daily Papers · · 4 min read

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques.</p>\n","updatedAt":"2026-07-07T08:38:24.432Z","author":{"_id":"63bc5552d8d676a229a1d553","avatarUrl":"/avatars/a911493e0138eab8db08fcddc44ea0c1.svg","fullname":"Gal Fiebelman","name":"galfiebelman","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6699864864349365},"editors":["galfiebelman"],"editorAvatarUrls":["/avatars/a911493e0138eab8db08fcddc44ea0c1.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05376","authors":[{"_id":"6a4c961225849b193a834347","user":{"_id":"63bc5552d8d676a229a1d553","avatarUrl":"/avatars/a911493e0138eab8db08fcddc44ea0c1.svg","isPro":false,"fullname":"Gal Fiebelman","user":"galfiebelman","type":"user","name":"galfiebelman"},"name":"Gal Fiebelman","status":"claimed_verified","statusLastChangedAt":"2026-07-07T08:30:32.521Z","hidden":false},{"_id":"6a4c961225849b193a834348","name":"Hadar Averbuch-Elor","hidden":false},{"_id":"6a4c961225849b193a834349","name":"Sagie Benaim","hidden":false}],"publishedAt":"2026-07-06T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing","submittedOnDailyBy":{"_id":"63bc5552d8d676a229a1d553","avatarUrl":"/avatars/a911493e0138eab8db08fcddc44ea0c1.svg","isPro":false,"fullname":"Gal Fiebelman","user":"galfiebelman","type":"user","name":"galfiebelman"},"summary":"Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.","upvotes":7,"discussionId":"6a4c961225849b193a83434a","projectPage":"https://galfiebelman.github.io/mv-forcing/","ai_summary":"A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques.","ai_keywords":["video diffusion models","temporal autoregression","multi-view synthesis","bidirectional attention","4D geometric bridge","autoregressive 3D reconstruction model","geometric prior","temporally unbounded generation","joint denoising regime","Distribution Matching Distillation","Spatio-Temporal Self-Forcing","student model"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63bc5552d8d676a229a1d553","avatarUrl":"/avatars/a911493e0138eab8db08fcddc44ea0c1.svg","isPro":false,"fullname":"Gal Fiebelman","user":"galfiebelman","type":"user"},{"_id":"6915e84a7ad296ad41f63ce7","avatarUrl":"/avatars/0c30f9892beafce7b6117b4f5612d6c0.svg","isPro":false,"fullname":"Itay Chachy","user":"itchachy","type":"user"},{"_id":"697b80f027319f7874015f8a","avatarUrl":"/avatars/1fe6cfea0be79a2901c334687512c01d.svg","isPro":false,"fullname":"Hadar Davidson","user":"HadarD","type":"user"},{"_id":"687363d49a81c7dcbcfa2d84","avatarUrl":"/avatars/5d943a5c811ed931c3fdcfee19253049.svg","isPro":false,"fullname":"jj","user":"realman123","type":"user"},{"_id":"69bcb2614df1e2c004b89a95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/obn93tUEz2RjyilQSGtPY.jpeg","isPro":false,"fullname":"宋 子豪","user":"lilybaker2026","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"699bc7f384bd5d3d91c502db","avatarUrl":"/avatars/4168c5018c0848a6ff0f24f6aea487da.svg","isPro":false,"fullname":"Graysonjohnson","user":"graysonjohnson3","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.05376.md","query":{}}">
Papers
arxiv:2607.05376

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

Published on Jul 6
· Submitted by
Gal Fiebelman
on Jul 7
Authors:
,

Abstract

A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques.

Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.

Community

Paper author Paper submitter about 13 hours ago

A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.05376
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.05376 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.05376 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.05376 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers