Hugging Face Daily Papers · · 6 min read

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.</p>\n","updatedAt":"2026-09-14T02:09:16.640Z","author":{"_id":"632b42626110e37dba3d5bcb","avatarUrl":"/avatars/ca70a15def71ee84f4f149db5e954843.svg","fullname":"Duan","name":"Jiafei1224","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8382601141929626},"editors":["Jiafei1224"],"editorAvatarUrls":["/avatars/ca70a15def71ee84f4f149db5e954843.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.12641","authors":[{"_id":"6aa756317ba345d44ad1488a","name":"Jianman Lin","hidden":false},{"_id":"6aa756317ba345d44ad1488b","user":{"_id":"6677c003a845e4470f3ed78d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6677c003a845e4470f3ed78d/izyZNLbvjhFSN40x0FcSF.png","isPro":false,"fullname":"Shailesh","user":"shailes-h","type":"user","name":"shailes-h"},"name":"Shailesh Shailesh","status":"claimed_verified","statusLastChangedAt":"2026-09-14T09:17:57.283Z","hidden":false},{"_id":"6aa756317ba345d44ad1488c","name":"Zhongyi Luo","hidden":false},{"_id":"6aa756317ba345d44ad1488d","name":"Jiafei Duan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/632b42626110e37dba3d5bcb/EXM_5UqH9QEmPH0nDWFQH.mp4"],"publishedAt":"2026-09-11T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models","submittedOnDailyBy":{"_id":"632b42626110e37dba3d5bcb","avatarUrl":"/avatars/ca70a15def71ee84f4f149db5e954843.svg","isPro":false,"fullname":"Duan","user":"Jiafei1224","type":"user","name":"Jiafei1224"},"summary":"Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.","upvotes":49,"discussionId":"6aa756317ba345d44ad1488e","projectPage":"https://magiclab-nus.github.io/LIT/?v=37955c5","githubRepo":"https://github.com/MAGICLAB-NUS/LIT","githubRepoAddedBy":"user","ai_summary":"LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information.","ai_keywords":["vision-action shortcuts","latent interface training","SE(3) end-effector pose","action chunks","vision-language-action models","world-action models","spatial-goal-conditioned action prior"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":24,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"632b42626110e37dba3d5bcb","avatarUrl":"/avatars/ca70a15def71ee84f4f149db5e954843.svg","isPro":false,"fullname":"Duan","user":"Jiafei1224","type":"user"},{"_id":"6aa2ae4fe397fbc5f42c847b","avatarUrl":"/avatars/83d5d73dc8645e1ea8bade8bfd5d3466.svg","isPro":false,"fullname":"MAGIC Lab@NUS","user":"NUSMAGIC","type":"user"},{"_id":"67421476e61fe2cad315dd18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/199r5VUQ6d2KcrUDLWnW1.png","isPro":false,"fullname":"Sang Nguyen","user":"tsangb34","type":"user"},{"_id":"6954d93ff46569e77f5c9138","avatarUrl":"/avatars/38817e195db9b3bd461b3bb82e6785cd.svg","isPro":false,"fullname":"lin","user":"miamia2","type":"user"},{"_id":"6721f7c5d3084146fa368888","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7c9k9BdSLUktlM3QrBIZ9.png","isPro":false,"fullname":"Furunemu","user":"jichuanman","type":"user"},{"_id":"64320be4034ecbefddd7e41d","avatarUrl":"/avatars/55581c93fcf4e2331c7100319105d5f5.svg","isPro":false,"fullname":"Zhihua Xu","user":"zizizihua","type":"user"},{"_id":"6aa7593461b5e464d7bd7e25","avatarUrl":"/avatars/7c39b2eee0180de76ea8085f514f714e.svg","isPro":false,"fullname":"tingtingwuwu","user":"tingtingwuwu1","type":"user"},{"_id":"6730625ad66bf1b63778faa8","avatarUrl":"/avatars/01f59923d8e29fc00b8f10612d2ef98a.svg","isPro":false,"fullname":"林浩诚","user":"match3966","type":"user"},{"_id":"6a16d918f35ca32e30f20dc1","avatarUrl":"/avatars/d05bacc6025c61ec1bbafbe7a357994d.svg","isPro":false,"fullname":"xinxi li","user":"lixinxii","type":"user"},{"_id":"6aa759b5e4f23f13c491b6c6","avatarUrl":"/avatars/2b5532ac64f5aeeffe3458353dd698e3.svg","isPro":false,"fullname":"Caiqijie","user":"Caigo1","type":"user"},{"_id":"6aa75a74bdc91e3c2eb7bead","avatarUrl":"/avatars/1b6a391abb4c238e0b80d53afec86337.svg","isPro":false,"fullname":"Keerthi Vasan Murugan","user":"k3rth1v5n","type":"user"},{"_id":"676e09030949a2ace95aa786","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/WUfYoEHdmSOEe_22fs-jJ.png","isPro":false,"fullname":"zhu","user":"Zhu777","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.12641.md","query":{}}">
Papers
arxiv:2609.12641

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Published on Sep 11
· Submitted by
Duan
on Sep 14
#1 Paper of the day
Authors:
,

Abstract

LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information.

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.

Community

Paper submitter about 8 hours ago

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.12641
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.12641 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.12641 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.12641 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers