SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion</p>\n","updatedAt":"2026-07-07T14:32:59.390Z","author":{"_id":"62262603a8218a9b862b18cc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648321672399-62262603a8218a9b862b18cc.jpeg","fullname":"Paul Engstler","name":"paulengstler","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"fr","probability":0.22952471673488617},"editors":["paulengstler"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1648321672399-62262603a8218a9b862b18cc.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05392","authors":[{"_id":"6a4d0dcc25849b193a834818","name":"Paul Engstler","hidden":false},{"_id":"6a4d0dcc25849b193a834819","name":"Iro Laina","hidden":false},{"_id":"6a4d0dcc25849b193a83481a","name":"Christian Rupprecht","hidden":false},{"_id":"6a4d0dcc25849b193a83481b","name":"Andrea Vedaldi","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/62262603a8218a9b862b18cc/oAC2INgwY-xQp1MA8FKg5.mp4"],"publishedAt":"2026-07-06T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion","submittedOnDailyBy":{"_id":"62262603a8218a9b862b18cc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648321672399-62262603a8218a9b862b18cc.jpeg","isPro":false,"fullname":"Paul Engstler","user":"paulengstler","type":"user","name":"paulengstler"},"summary":"We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.","upvotes":2,"discussionId":"6a4d0dcc25849b193a83481c","projectPage":"https://research.paulengstler.com/syncity-3k/","githubRepo":"https://github.com/paulengstler/syncity-3k","githubRepoAddedBy":"user","ai_summary":"SynCity 3000 generates large, coherent 3D scenes by adapting image-to-3D generators as convolutional operators through fine-tuning on synthetic scene data.","ai_keywords":["image-to-3D generators","convolutional operator","fine-tuning","synthetic data engine","dimetric image","3D scene generation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"627bbc28fbab61b048eba8b6","name":"Oxford","fullname":"University of Oxford","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/u0ey2LfYu6uG6iu8m_kH7.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62262603a8218a9b862b18cc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648321672399-62262603a8218a9b862b18cc.jpeg","isPro":false,"fullname":"Paul Engstler","user":"paulengstler","type":"user"},{"_id":"69bcd9f21ab1ebec4090a6d9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/RJHcPYAx6RqS3HHkyorza.png","isPro":false,"fullname":"Han Chenxi","user":"lilywilson","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"627bbc28fbab61b048eba8b6","name":"Oxford","fullname":"University of Oxford","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/u0ey2LfYu6uG6iu8m_kH7.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.05392.md","query":{}}">
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
Abstract
SynCity 3000 generates large, coherent 3D scenes by adapting image-to-3D generators as convolutional operators through fine-tuning on synthetic scene data.
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.
Community
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.05392 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.05392 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.05392 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.