TL;DR: TransNormal-2 estimates surface normal maps from a single RGB image in one deterministic rectified-flow step on the FLUX.2 [klein] 9B backbone (LoRA-adapted), followed by a lightweight Geometric Refinement Module (GRM) that applies bounded residual correction to boundary-localized decoding errors. It matches strong feed-forward baselines (e.g. MoGe-2) on general scenes and substantially outperforms prior art on transparent objects (glass, liquids).</p>\n","updatedAt":"2026-09-09T05:22:51.218Z","author":{"_id":"64d4890517fea7f4e76b3536","avatarUrl":"/avatars/9cf60c84b908d202fa19a8c9773f0011.svg","fullname":"Mingwei Li","name":"Longxiang-ai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8045266270637512},"editors":["Longxiang-ai"],"editorAvatarUrls":["/avatars/9cf60c84b908d202fa19a8c9773f0011.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06665","authors":[{"_id":"6aa0e756d0174964227bee07","name":"Mingwei Li","hidden":false},{"_id":"6aa0e756d0174964227bee08","name":"Yi Yang","hidden":false},{"_id":"6aa0e756d0174964227bee09","name":"Hehe Fan","hidden":false}],"publishedAt":"2026-09-06T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation","submittedOnDailyBy":{"_id":"64d4890517fea7f4e76b3536","avatarUrl":"/avatars/9cf60c84b908d202fa19a8c9773f0011.svg","isPro":false,"fullname":"Mingwei Li","user":"Longxiang-ai","type":"user","name":"Longxiang-ai"},"summary":"Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.","upvotes":17,"discussionId":"6aa0e756d0174964227bee0a","projectPage":"https://longxiang-ai.github.io/TransNormal-2/","githubRepo":"https://github.com/longxiang-ai/TransNormal-2","githubRepoAddedBy":"user","ai_summary":"TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors through geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations.","ai_keywords":["rectified-flow","VAE reconstruction","inverse rendering self-consistency","von Mises-Fisher angular loss","wavelet edge-aware regularization","Geometric Refinement Module","RGB-guided residual correction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64d4890517fea7f4e76b3536","avatarUrl":"/avatars/9cf60c84b908d202fa19a8c9773f0011.svg","isPro":false,"fullname":"Mingwei Li","user":"Longxiang-ai","type":"user"},{"_id":"641590e5486c7c9a5d13fe35","avatarUrl":"/avatars/2ef3432815e34a0eee45297fd99e5c40.svg","isPro":false,"fullname":"Ruisi Zhao","user":"zhaors00","type":"user"},{"_id":"6a69eb859ef85244dc4ca374","avatarUrl":"/avatars/21fc260ba33de7f03d663e84fba9199b.svg","isPro":false,"fullname":"Anthony Sanchez","user":"anthony-code","type":"user"},{"_id":"6a6c7d92b289c9e42b36715a","avatarUrl":"/avatars/cfeac81c781d73ff0db0523c551e16f9.svg","isPro":false,"fullname":"Edward Johnson","user":"Edward-Johnson","type":"user"},{"_id":"6a6dc7259ebe01e60754878f","avatarUrl":"/avatars/85ff5d222935ef19202882297a997f04.svg","isPro":false,"fullname":"Steven Williams","user":"Granite-Steven","type":"user"},{"_id":"6a6dee9ec2a26ede038d8358","avatarUrl":"/avatars/61b28fb945a54f7842219e34ecbe096f.svg","isPro":false,"fullname":"Linda Miller","user":"lindaMiller","type":"user"},{"_id":"6a9ae275c3a3d75ba18a0df8","avatarUrl":"/avatars/9a5fd2286fc550d4de7efcff6c2523f1.svg","isPro":false,"fullname":"Danielle Payne","user":"jennylester","type":"user"},{"_id":"6a9b55116c28a31b4e16f39f","avatarUrl":"/avatars/d20cb87e5bd204a5900087aa0c3c1944.svg","isPro":false,"fullname":"小林 零","user":"haruka87","type":"user"},{"_id":"6aa0f427899c3077e3fc90f6","avatarUrl":"/avatars/20051b3977c8a6b67ef3a324abc0b90f.svg","isPro":false,"fullname":"庄瑜","user":"shenmin91","type":"user"},{"_id":"6a6a95368acf46140bae93a3","avatarUrl":"/avatars/c64673a23a306cc73ed00382ec370e10.svg","isPro":false,"fullname":"Kevin Thompson","user":"rapidFox","type":"user"},{"_id":"6a6de4192a48f06370434ab2","avatarUrl":"/avatars/aa81799e69c954a6b8626a4dc775f1ec.svg","isPro":false,"fullname":"Kenneth Rodriguez","user":"AzureReed","type":"user"},{"_id":"6a701cb2665e43904874853c","avatarUrl":"/avatars/365a38e37c3d23fa7a9d1533c365075f.svg","isPro":false,"fullname":"Joseph Perez","user":"CedarPeak","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06665.md","query":{}}">
TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
Abstract
TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors through geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations.
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.
Community
TL;DR: TransNormal-2 estimates surface normal maps from a single RGB image in one deterministic rectified-flow step on the FLUX.2 [klein] 9B backbone (LoRA-adapted), followed by a lightweight Geometric Refinement Module (GRM) that applies bounded residual correction to boundary-localized decoding errors. It matches strong feed-forward baselines (e.g. MoGe-2) on general scenes and substantially outperforms prior art on transparent objects (glass, liquids).
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.06665 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.06665 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.