Hugging Face Daily Papers · · 4 min read

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

TL;DR: TransNormal-2 estimates surface normal maps from a single RGB image in one deterministic rectified-flow step on the FLUX.2 [klein] 9B backbone (LoRA-adapted), followed by a lightweight Geometric Refinement Module (GRM) that applies bounded residual correction to boundary-localized decoding errors. It matches strong feed-forward baselines (e.g. MoGe-2) on general scenes and substantially outperforms prior art on transparent objects (glass, liquids).</p>\n","updatedAt":"2026-09-09T05:22:51.218Z","author":{"_id":"64d4890517fea7f4e76b3536","avatarUrl":"/avatars/9cf60c84b908d202fa19a8c9773f0011.svg","fullname":"Mingwei Li","name":"Longxiang-ai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8045266270637512},"editors":["Longxiang-ai"],"editorAvatarUrls":["/avatars/9cf60c84b908d202fa19a8c9773f0011.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06665","authors":[{"_id":"6aa0e756d0174964227bee07","name":"Mingwei Li","hidden":false},{"_id":"6aa0e756d0174964227bee08","name":"Yi Yang","hidden":false},{"_id":"6aa0e756d0174964227bee09","name":"Hehe Fan","hidden":false}],"publishedAt":"2026-09-06T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation","submittedOnDailyBy":{"_id":"64d4890517fea7f4e76b3536","avatarUrl":"/avatars/9cf60c84b908d202fa19a8c9773f0011.svg","isPro":false,"fullname":"Mingwei Li","user":"Longxiang-ai","type":"user","name":"Longxiang-ai"},"summary":"Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.","upvotes":17,"discussionId":"6aa0e756d0174964227bee0a","projectPage":"https://longxiang-ai.github.io/TransNormal-2/","githubRepo":"https://github.com/longxiang-ai/TransNormal-2","githubRepoAddedBy":"user","ai_summary":"TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors through geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations.","ai_keywords":["rectified-flow","VAE reconstruction","inverse rendering self-consistency","von Mises-Fisher angular loss","wavelet edge-aware regularization","Geometric Refinement Module","RGB-guided residual correction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64d4890517fea7f4e76b3536","avatarUrl":"/avatars/9cf60c84b908d202fa19a8c9773f0011.svg","isPro":false,"fullname":"Mingwei Li","user":"Longxiang-ai","type":"user"},{"_id":"641590e5486c7c9a5d13fe35","avatarUrl":"/avatars/2ef3432815e34a0eee45297fd99e5c40.svg","isPro":false,"fullname":"Ruisi Zhao","user":"zhaors00","type":"user"},{"_id":"6a69eb859ef85244dc4ca374","avatarUrl":"/avatars/21fc260ba33de7f03d663e84fba9199b.svg","isPro":false,"fullname":"Anthony Sanchez","user":"anthony-code","type":"user"},{"_id":"6a6c7d92b289c9e42b36715a","avatarUrl":"/avatars/cfeac81c781d73ff0db0523c551e16f9.svg","isPro":false,"fullname":"Edward Johnson","user":"Edward-Johnson","type":"user"},{"_id":"6a6dc7259ebe01e60754878f","avatarUrl":"/avatars/85ff5d222935ef19202882297a997f04.svg","isPro":false,"fullname":"Steven Williams","user":"Granite-Steven","type":"user"},{"_id":"6a6dee9ec2a26ede038d8358","avatarUrl":"/avatars/61b28fb945a54f7842219e34ecbe096f.svg","isPro":false,"fullname":"Linda Miller","user":"lindaMiller","type":"user"},{"_id":"6a9ae275c3a3d75ba18a0df8","avatarUrl":"/avatars/9a5fd2286fc550d4de7efcff6c2523f1.svg","isPro":false,"fullname":"Danielle Payne","user":"jennylester","type":"user"},{"_id":"6a9b55116c28a31b4e16f39f","avatarUrl":"/avatars/d20cb87e5bd204a5900087aa0c3c1944.svg","isPro":false,"fullname":"小林 零","user":"haruka87","type":"user"},{"_id":"6aa0f427899c3077e3fc90f6","avatarUrl":"/avatars/20051b3977c8a6b67ef3a324abc0b90f.svg","isPro":false,"fullname":"庄瑜","user":"shenmin91","type":"user"},{"_id":"6a6a95368acf46140bae93a3","avatarUrl":"/avatars/c64673a23a306cc73ed00382ec370e10.svg","isPro":false,"fullname":"Kevin Thompson","user":"rapidFox","type":"user"},{"_id":"6a6de4192a48f06370434ab2","avatarUrl":"/avatars/aa81799e69c954a6b8626a4dc775f1ec.svg","isPro":false,"fullname":"Kenneth Rodriguez","user":"AzureReed","type":"user"},{"_id":"6a701cb2665e43904874853c","avatarUrl":"/avatars/365a38e37c3d23fa7a9d1533c365075f.svg","isPro":false,"fullname":"Joseph Perez","user":"CedarPeak","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06665.md","query":{}}">
Papers
arxiv:2609.06665

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation

Published on Sep 6
· Submitted by
Mingwei Li
on Sep 9
Authors:
,

Abstract

TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors through geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations.

Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.

Community

TL;DR: TransNormal-2 estimates surface normal maps from a single RGB image in one deterministic rectified-flow step on the FLUX.2 [klein] 9B backbone (LoRA-adapted), followed by a lightweight Geometric Refinement Module (GRM) that applies bounded residual correction to boundary-localized decoding errors. It matches strong feed-forward baselines (e.g. MoGe-2) on general scenes and substantially outperforms prior art on transparent objects (glass, liquids).

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.06665
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.06665 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.06665 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers