Hugging Face Daily Papers · · 3 min read

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/66eeda3676a8038cb448f11d/zNDXO2VqMJGpmowMv2mi0.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/66eeda3676a8038cb448f11d/zNDXO2VqMJGpmowMv2mi0.png\" alt=\"9ac820456b89a40298155000eb9e3514\"></a></p>\n","updatedAt":"2026-08-06T02:16:26.775Z","author":{"_id":"66eeda3676a8038cb448f11d","avatarUrl":"/avatars/8d6e61f4c9354c6720ccaa7be0fe1d9f.svg","fullname":"Haiyang Zhou","name":"Marblueocean","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3736467659473419},"editors":["Marblueocean"],"editorAvatarUrls":["/avatars/8d6e61f4c9354c6720ccaa7be0fe1d9f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04701","authors":[{"_id":"6a73ec72c5e410d076869a20","user":{"_id":"66eeda3676a8038cb448f11d","avatarUrl":"/avatars/8d6e61f4c9354c6720ccaa7be0fe1d9f.svg","isPro":false,"fullname":"Haiyang Zhou","user":"Marblueocean","type":"user","name":"Marblueocean"},"name":"Haiyang Zhou","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.492Z","hidden":false},{"_id":"6a73ec72c5e410d076869a21","name":"Wangbo Yu","hidden":false},{"_id":"6a73ec72c5e410d076869a22","name":"Chaoran Feng","hidden":false},{"_id":"6a73ec72c5e410d076869a23","name":"Xunyu Zhou","hidden":false},{"_id":"6a73ec72c5e410d076869a24","name":"Yonghong Tian","hidden":false},{"_id":"6a73ec72c5e410d076869a25","name":"Li Yuan","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models","submittedOnDailyBy":{"_id":"66eeda3676a8038cb448f11d","avatarUrl":"/avatars/8d6e61f4c9354c6720ccaa7be0fe1d9f.svg","isPro":false,"fullname":"Haiyang Zhou","user":"Marblueocean","type":"user","name":"Marblueocean"},"summary":"The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.","upvotes":5,"discussionId":"6a73ec72c5e410d076869a26","projectPage":"https://zhouhyocean.github.io/uniworld-view/","githubRepo":"https://github.com/PKU-YuanGroup/UniWorld-View","githubRepoAddedBy":"user","githubStars":92},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66eeda3676a8038cb448f11d","avatarUrl":"/avatars/8d6e61f4c9354c6720ccaa7be0fe1d9f.svg","isPro":false,"fullname":"Haiyang Zhou","user":"Marblueocean","type":"user"},{"_id":"6a6a9b87144c90a0d416e927","avatarUrl":"/avatars/fcabd41391321a12518a8b6b77f6db4d.svg","isPro":false,"fullname":"Robert Jackson","user":"granitereed","type":"user"},{"_id":"6a6c849d7827edbb08463d78","avatarUrl":"/avatars/fa58f398553193ff34c11959b184129a.svg","isPro":false,"fullname":"Jessica Anderson","user":"Jessica-Anderson","type":"user"},{"_id":"6a6c9c3fe34d1f023f2cf030","avatarUrl":"/avatars/d6a1bfc68f32352f04630041e0f289b2.svg","isPro":false,"fullname":"Barbara Moore","user":"Kestrel-Node","type":"user"},{"_id":"6a6dec834183641b5b7fc99a","avatarUrl":"/avatars/73dd8810b7be896a0209010a541086dd.svg","isPro":false,"fullname":"Thomas Johnson","user":"sam-0263298","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04701.md","query":{}}">
Papers
arxiv:2608.04701

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

Published on Aug 5
· Submitted by
Haiyang Zhou
on Aug 6
Authors:

Abstract

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.

Community

Paper author Paper submitter about 8 hours ago

9ac820456b89a40298155000eb9e3514

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04701
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.04701 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.04701 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.04701 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers