An adapter over DUNE for generating visual features grounded in 3D object geometry for accurate and fast zero-shot CAD-to-image alignment</p>\n","updatedAt":"2026-07-17T12:46:33.207Z","author":{"_id":"6568621b0e4b5ff9d55677ae","avatarUrl":"/avatars/a64a6791d4dcedd0786a6073de9bf371.svg","fullname":"Saad Ejaz","name":"saadejazz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8286172747612},"editors":["saadejazz"],"editorAvatarUrls":["/avatars/a64a6791d4dcedd0786a6073de9bf371.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.15058","authors":[{"_id":"6a5a23d30294a92d3d1ee42d","name":"Saad Ejaz","hidden":false},{"_id":"6a5a23d30294a92d3d1ee42e","name":"Miguel Fernandez-Cortizas","hidden":false},{"_id":"6a5a23d30294a92d3d1ee42f","name":"Javier Civera","hidden":false},{"_id":"6a5a23d30294a92d3d1ee430","name":"Holger Voos","hidden":false},{"_id":"6a5a23d30294a92d3d1ee431","name":"Jose Luis Sanchez-Lopez","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment","submittedOnDailyBy":{"_id":"6568621b0e4b5ff9d55677ae","avatarUrl":"/avatars/a64a6791d4dcedd0786a6073de9bf371.svg","isPro":false,"fullname":"Saad Ejaz","user":"saadejazz","type":"user","name":"saadejazz"},"summary":"CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to CAD models, yet typically their correspondences are appearance-driven and degrade under occlusion or sim-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a weakly-supervised framework for zero-shot CAD alignment with two key contributions. First, SUFLECA scales up geometry-grounded feature learning from pretrained visual representations through Normalized Object Coordinates (NOCs) supervision on 674K images spanning 12 real and synthetic datasets, learning compact geometry-aware features that generalize across domains. Second, we propose a geometrically consistent matching algorithm that establishes reliable one-to-one CAD-to-image correspondences. Together, these contributions enable accurate, sub-second alignment per object instance without iterative pose refinement. On ScanNet25k, SUFLECA achieves 33.4%/42.3% category/instance accuracy, outperforming, with a smaller computational footprint, the strongest zero-shot baseline by 10.3/12.2 percentage points and, for the first time on this benchmark, even surpassing fully supervised methods. Code is available at: https://github.com/snt-arg/SUFLECA","upvotes":2,"discussionId":"6a5a23d30294a92d3d1ee432","githubRepo":"https://github.com/snt-arg/SUFLECA","githubRepoAddedBy":"user","githubStars":1},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69bcee1fb0b4d685f7c21deb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/JiP6kPllc4jqCUgPcFM5o.png","isPro":false,"fullname":"山口大翔","user":"abigadams","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.15058.md","query":{}}">
SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
Abstract
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to CAD models, yet typically their correspondences are appearance-driven and degrade under occlusion or sim-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD Alignment), a weakly-supervised framework for zero-shot CAD alignment with two key contributions. First, SUFLECA scales up geometry-grounded feature learning from pretrained visual representations through Normalized Object Coordinates (NOCs) supervision on 674K images spanning 12 real and synthetic datasets, learning compact geometry-aware features that generalize across domains. Second, we propose a geometrically consistent matching algorithm that establishes reliable one-to-one CAD-to-image correspondences. Together, these contributions enable accurate, sub-second alignment per object instance without iterative pose refinement. On ScanNet25k, SUFLECA achieves 33.4%/42.3% category/instance accuracy, outperforming, with a smaller computational footprint, the strongest zero-shot baseline by 10.3/12.2 percentage points and, for the first time on this benchmark, even surpassing fully supervised methods. Code is available at: https://github.com/snt-arg/SUFLECA
Community
An adapter over DUNE for generating visual features grounded in 3D object geometry for accurate and fast zero-shot CAD-to-image alignment
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.15058 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.15058 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.15058 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.