PriorEdit3D learns instruction-guided, feed-forward 3D editing without paired 3D supervision by distilling visual, semantic, and geometric priors from pretrained foundation models.</p>\n","updatedAt":"2026-09-09T13:56:05.947Z","author":{"_id":"6356194ce1e6a03a1ed3e0b3","avatarUrl":"/avatars/92adf0c54549ac99315fa36da78b0448.svg","fullname":"wenhao","name":"costwen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7294015288352966},"editors":["costwen"],"editorAvatarUrls":["/avatars/92adf0c54549ac99315fa36da78b0448.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04942","authors":[{"_id":"6aa16527d8c54e38c0a369d7","user":{"_id":"6356194ce1e6a03a1ed3e0b3","avatarUrl":"/avatars/92adf0c54549ac99315fa36da78b0448.svg","isPro":false,"fullname":"wenhao","user":"costwen","type":"user","name":"costwen"},"name":"Hao Wen","status":"claimed_verified","statusLastChangedAt":"2026-09-09T16:45:04.638Z","hidden":false},{"_id":"6aa16527d8c54e38c0a369d8","name":"Weibin Yun","hidden":false},{"_id":"6aa16527d8c54e38c0a369d9","name":"Hongxing Fan","hidden":false},{"_id":"6aa16527d8c54e38c0a369da","name":"Haotian Lu","hidden":false},{"_id":"6aa16527d8c54e38c0a369db","name":"Rui Chen","hidden":false},{"_id":"6aa16527d8c54e38c0a369dc","name":"Zehuan Huang","hidden":false},{"_id":"6aa16527d8c54e38c0a369dd","name":"Lu Sheng","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"Learning 3D Editing without Paired Supervision via Generative Prior Distillation","submittedOnDailyBy":{"_id":"6356194ce1e6a03a1ed3e0b3","avatarUrl":"/avatars/92adf0c54549ac99315fa36da78b0448.svg","isPro":false,"fullname":"wenhao","user":"costwen","type":"user","name":"costwen"},"summary":"Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by either relying on slow test-time optimization or training on pseudo-pairs constructed via complex pipelines, which often introduce structural drift and geometric artifacts. In this paper, we propose a novel framework that learns feed-forward 3D editing without paired 3D supervision via Generative Prior Distillation. Instead of relying on ground-truth 3D pairs, our core idea is to distill visual, semantic, and geometric knowledge from powerful foundation models directly into a 3D editing model. Specifically, through a differentiable rendering pipeline, we supervise the 3D representation using two complementary signals: a 2D visual prior from an image editing model at the main editing view, and a semantic prior from a Vision-Language Model at novel views to ensure strict instruction following and source identity preservation. Crucially, to address the geometric collapse and multi-view inconsistencies inherent in 2D projection supervision, we introduce a 3D-aware Distribution Matching regularization. Acting as a geometric prior, this term operates in the 3D latent space, constraining the edited output to remain within the manifold of realistic 3D assets defined by a pretrained image to 3D teacher model. Extensive experiments demonstrate that our method achieves superior instruction fidelity and cross-view consistency, significantly outperforming state-of-the-art baselines. Our project is available at: https://github.com/thiamine128/PriorEdit3D.","upvotes":0,"discussionId":"6aa16528d8c54e38c0a369de","projectPage":"https://thiamine128.github.io/PriorEdit3D/","githubRepo":"https://github.com/thiamine128/PriorEdit3D","githubRepoAddedBy":"user","ai_summary":"A feed-forward 3D editing framework distills visual, semantic, and geometric priors from foundation models via differentiable rendering and 3D-aware distribution matching to avoid paired training data.","ai_keywords":["instruction-guided 3D editing","Generative Prior Distillation","differentiable rendering","2D visual prior","Vision-Language Model","3D-aware Distribution Matching","image to 3D teacher model","cross-view consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04942.md","query":{}}">
Learning 3D Editing without Paired Supervision via Generative Prior Distillation
Published on Sep 4
· Submitted by wenhao on Sep 9 Abstract
A feed-forward 3D editing framework distills visual, semantic, and geometric priors from foundation models via differentiable rendering and 3D-aware distribution matching to avoid paired training data.
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by either relying on slow test-time optimization or training on pseudo-pairs constructed via complex pipelines, which often introduce structural drift and geometric artifacts. In this paper, we propose a novel framework that learns feed-forward 3D editing without paired 3D supervision via Generative Prior Distillation. Instead of relying on ground-truth 3D pairs, our core idea is to distill visual, semantic, and geometric knowledge from powerful foundation models directly into a 3D editing model. Specifically, through a differentiable rendering pipeline, we supervise the 3D representation using two complementary signals: a 2D visual prior from an image editing model at the main editing view, and a semantic prior from a Vision-Language Model at novel views to ensure strict instruction following and source identity preservation. Crucially, to address the geometric collapse and multi-view inconsistencies inherent in 2D projection supervision, we introduce a 3D-aware Distribution Matching regularization. Acting as a geometric prior, this term operates in the 3D latent space, constraining the edited output to remain within the manifold of realistic 3D assets defined by a pretrained image to 3D teacher model. Extensive experiments demonstrate that our method achieves superior instruction fidelity and cross-view consistency, significantly outperforming state-of-the-art baselines. Our project is available at: https://github.com/thiamine128/PriorEdit3D.
Community
PriorEdit3D learns instruction-guided, feed-forward 3D editing without paired 3D supervision by distilling visual, semantic, and geometric priors from pretrained foundation models.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.04942 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.04942 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.04942 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.