Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.</p>\n","updatedAt":"2026-09-14T02:09:16.640Z","author":{"_id":"632b42626110e37dba3d5bcb","avatarUrl":"/avatars/ca70a15def71ee84f4f149db5e954843.svg","fullname":"Duan","name":"Jiafei1224","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8382601141929626},"editors":["Jiafei1224"],"editorAvatarUrls":["/avatars/ca70a15def71ee84f4f149db5e954843.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.12641","authors":[{"_id":"6aa756317ba345d44ad1488a","name":"Jianman Lin","hidden":false},{"_id":"6aa756317ba345d44ad1488b","user":{"_id":"6677c003a845e4470f3ed78d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6677c003a845e4470f3ed78d/izyZNLbvjhFSN40x0FcSF.png","isPro":false,"fullname":"Shailesh","user":"shailes-h","type":"user","name":"shailes-h"},"name":"Shailesh Shailesh","status":"claimed_verified","statusLastChangedAt":"2026-09-14T09:17:57.283Z","hidden":false},{"_id":"6aa756317ba345d44ad1488c","name":"Zhongyi Luo","hidden":false},{"_id":"6aa756317ba345d44ad1488d","name":"Jiafei Duan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/632b42626110e37dba3d5bcb/EXM_5UqH9QEmPH0nDWFQH.mp4"],"publishedAt":"2026-09-11T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models","submittedOnDailyBy":{"_id":"632b42626110e37dba3d5bcb","avatarUrl":"/avatars/ca70a15def71ee84f4f149db5e954843.svg","isPro":false,"fullname":"Duan","user":"Jiafei1224","type":"user","name":"Jiafei1224"},"summary":"Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.","upvotes":49,"discussionId":"6aa756317ba345d44ad1488e","projectPage":"https://magiclab-nus.github.io/LIT/?v=37955c5","githubRepo":"https://github.com/MAGICLAB-NUS/LIT","githubRepoAddedBy":"user","ai_summary":"LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information.","ai_keywords":["vision-action shortcuts","latent interface training","SE(3) end-effector pose","action chunks","vision-language-action models","world-action models","spatial-goal-conditioned action prior"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":24,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"632b42626110e37dba3d5bcb","avatarUrl":"/avatars/ca70a15def71ee84f4f149db5e954843.svg","isPro":false,"fullname":"Duan","user":"Jiafei1224","type":"user"},{"_id":"6aa2ae4fe397fbc5f42c847b","avatarUrl":"/avatars/83d5d73dc8645e1ea8bade8bfd5d3466.svg","isPro":false,"fullname":"MAGIC Lab@NUS","user":"NUSMAGIC","type":"user"},{"_id":"67421476e61fe2cad315dd18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/199r5VUQ6d2KcrUDLWnW1.png","isPro":false,"fullname":"Sang Nguyen","user":"tsangb34","type":"user"},{"_id":"6954d93ff46569e77f5c9138","avatarUrl":"/avatars/38817e195db9b3bd461b3bb82e6785cd.svg","isPro":false,"fullname":"lin","user":"miamia2","type":"user"},{"_id":"6721f7c5d3084146fa368888","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7c9k9BdSLUktlM3QrBIZ9.png","isPro":false,"fullname":"Furunemu","user":"jichuanman","type":"user"},{"_id":"64320be4034ecbefddd7e41d","avatarUrl":"/avatars/55581c93fcf4e2331c7100319105d5f5.svg","isPro":false,"fullname":"Zhihua Xu","user":"zizizihua","type":"user"},{"_id":"6aa7593461b5e464d7bd7e25","avatarUrl":"/avatars/7c39b2eee0180de76ea8085f514f714e.svg","isPro":false,"fullname":"tingtingwuwu","user":"tingtingwuwu1","type":"user"},{"_id":"6730625ad66bf1b63778faa8","avatarUrl":"/avatars/01f59923d8e29fc00b8f10612d2ef98a.svg","isPro":false,"fullname":"林浩诚","user":"match3966","type":"user"},{"_id":"6a16d918f35ca32e30f20dc1","avatarUrl":"/avatars/d05bacc6025c61ec1bbafbe7a357994d.svg","isPro":false,"fullname":"xinxi li","user":"lixinxii","type":"user"},{"_id":"6aa759b5e4f23f13c491b6c6","avatarUrl":"/avatars/2b5532ac64f5aeeffe3458353dd698e3.svg","isPro":false,"fullname":"Caiqijie","user":"Caigo1","type":"user"},{"_id":"6aa75a74bdc91e3c2eb7bead","avatarUrl":"/avatars/1b6a391abb4c238e0b80d53afec86337.svg","isPro":false,"fullname":"Keerthi Vasan Murugan","user":"k3rth1v5n","type":"user"},{"_id":"676e09030949a2ace95aa786","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/WUfYoEHdmSOEe_22fs-jJ.png","isPro":false,"fullname":"zhu","user":"Zhu777","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.12641.md","query":{}}">
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
Abstract
LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information.
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.
Community
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.12641 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.12641 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.12641 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.