Hugging Face Daily Papers · · 4 min read

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this work, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization.</p>\n","updatedAt":"2026-07-16T03:16:34.504Z","author":{"_id":"68e741ea3edb0ff47e20084e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e741ea3edb0ff47e20084e/OyBgFqcU4QWyPF_K58Gt5.jpeg","fullname":"Wu Kai","name":"KaiiWuu1993","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8661251664161682},"editors":["KaiiWuu1993"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/68e741ea3edb0ff47e20084e/OyBgFqcU4QWyPF_K58Gt5.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.12000","authors":[{"_id":"6a56f852e548eb96f98ccb22","name":"Yufei Cai","hidden":false},{"_id":"6a56f852e548eb96f98ccb23","name":"Xuesong Niu","hidden":false},{"_id":"6a56f852e548eb96f98ccb24","name":"Hao Lu","hidden":false},{"_id":"6a56f852e548eb96f98ccb25","name":"Kun Gai","hidden":false},{"_id":"6a56f852e548eb96f98ccb26","name":"Kai Wu","hidden":false},{"_id":"6a56f852e548eb96f98ccb27","name":"Guosheng Lin","hidden":false}],"publishedAt":"2026-07-13T00:00:00.000Z","submittedOnDailyAt":"2026-07-16T00:00:00.000Z","title":"MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors","submittedOnDailyBy":{"_id":"68e741ea3edb0ff47e20084e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e741ea3edb0ff47e20084e/OyBgFqcU4QWyPF_K58Gt5.jpeg","isPro":false,"fullname":"Wu Kai","user":"KaiiWuu1993","type":"user","name":"KaiiWuu1993"},"summary":"Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.","upvotes":22,"discussionId":"6a56f852e548eb96f98ccb28","projectPage":"https://prototypenx.github.io/MetaView/","githubRepo":"https://github.com/KlingAIResearch/MetaView","githubRepoAddedBy":"user","githubStars":16,"organization":{"_id":"665f02ce9f9e5b38d0a256a8","name":"Kwai-Kolors","fullname":"Kolors Team, Kuaishou Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f0babaef9cc6810cec02ff/sVnELkcfVo5kxg5308rkr.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64620ca8a489ecb6b6b5b4e4","avatarUrl":"/avatars/9bf8060cb7bc59c0369ddf14c50aa480.svg","isPro":false,"fullname":"Xuesong Niu","user":"nxsEdson","type":"user"},{"_id":"649e83af47b275ac48466218","avatarUrl":"/avatars/5ce406848e467afc50ded4e8f7df5970.svg","isPro":false,"fullname":"dong she","user":"sd0809","type":"user"},{"_id":"64f9951e42f1c4a68c9882b3","avatarUrl":"/avatars/77fa010db659107a059d510e036025e0.svg","isPro":false,"fullname":"Huijie Liu","user":"liuhuijie6410","type":"user"},{"_id":"64c109616cf8e2b9f05486d4","avatarUrl":"/avatars/f6bb14d4e406b4315efda26752b17526.svg","isPro":false,"fullname":"Dong","user":"SiyaoDong","type":"user"},{"_id":"68e741ea3edb0ff47e20084e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e741ea3edb0ff47e20084e/OyBgFqcU4QWyPF_K58Gt5.jpeg","isPro":false,"fullname":"Wu Kai","user":"KaiiWuu1993","type":"user"},{"_id":"68fed13597cae9c0f484ebed","avatarUrl":"/avatars/8df380d81daab4fae8a52d0816bf8099.svg","isPro":false,"fullname":"libo","user":"libo31","type":"user"},{"_id":"6732cdfa07cf693a11536b88","avatarUrl":"/avatars/ba0f623f77baee34cbac1422570931da.svg","isPro":false,"fullname":"Zhou Xingchi","user":"hooccee","type":"user"},{"_id":"647076467fd7ecdbd0ea03b1","avatarUrl":"/avatars/6e090ea5f88977c6f70544175094c2a6.svg","isPro":false,"fullname":"Penghui Du","user":"eternaldolphin","type":"user"},{"_id":"63733e12c1142685da81d167","avatarUrl":"/avatars/20a1d0a0cc8ed180a5609c6fe7fe51a1.svg","isPro":false,"fullname":"Yu Xu","user":"YUXU915","type":"user"},{"_id":"66163a086dff9352ba4ec91b","avatarUrl":"/avatars/235e261e36549223e45b5750ae0860bd.svg","isPro":false,"fullname":"Xu Zhang","user":"xushu-me","type":"user"},{"_id":"66164b7a509f9e084f679ef4","avatarUrl":"/avatars/9c75fa2fd1d6ccd09748639a9a17bebc.svg","isPro":false,"fullname":"dacheng","user":"dcfucheng","type":"user"},{"_id":"66221f7a608fb4e7f46ecadb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yMCsWyTkL4Xdxe18uwMYR.jpeg","isPro":false,"fullname":"Tianlin Pan","user":"ldiex","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"665f02ce9f9e5b38d0a256a8","name":"Kwai-Kolors","fullname":"Kolors Team, Kuaishou Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f0babaef9cc6810cec02ff/sVnELkcfVo5kxg5308rkr.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.12000.md","query":{}}">
Papers
arxiv:2607.12000

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Published on Jul 13
· Submitted by
Wu Kai
on Jul 16
Authors:
,

Abstract

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.

Community

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this work, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.12000
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.12000 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers