Vision-language-action (VLA) models commonly adopt an LLM-centric V→L→A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V→L→A pathway as a direct V+L→A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at this https URL.</p>\n","updatedAt":"2026-07-30T02:12:58.842Z","author":{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","fullname":"Dingkang Liang","name":"dkliang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.874606728553772},"editors":["dkliang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27205","authors":[{"_id":"6a6ab0d94463a8a84bdc3f6e","name":"Hengyi Xie","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f6f","name":"Chenfei Yao","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f70","name":"Xianjin Wu","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f71","name":"Xuanyang Xi","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f72","name":"Yiping Tang","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f73","name":"Di Xu","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f74","name":"Yingying Zhu","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f75","name":"Dingkang Liang","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f76","name":"Xiang Bai","hidden":false},{"_id":"6a6ab0d94463a8a84bdc3f77","name":"Han Ding","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/67467b5979406f42a14517e9/ZrwNVOXZaCzQvbewHyfqg.png"],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM","submittedOnDailyBy":{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user","name":"dkliang"},"summary":"Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V to L to A pathway as a direct V + L to A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.","upvotes":75,"discussionId":"6a6ab0d94463a8a84bdc3f78","projectPage":"https://h-embodvis.github.io/TurboVLA/","githubRepo":"https://github.com/H-EmbodVis/TurboVLA","githubRepoAddedBy":"user","githubStars":15,"organization":{"_id":"687cebf73858638f66e59f56","name":"H-EmbodVis","fullname":"H-EmbodVis","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/XkebfHCVngT9o82kTAVAV.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user"},{"_id":"67d2399318e5a69f350298f1","avatarUrl":"/avatars/77913329c5fb64019de40cae04c23f1a.svg","isPro":false,"fullname":"SifanTu","user":"SifanTu","type":"user"},{"_id":"668cb84b410a13fa3d3d1297","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/668cb84b410a13fa3d3d1297/65g-71-QYm4DrJAdUa4O2.png","isPro":false,"fullname":"Xianjin-Wu","user":"HyperbolicCurve","type":"user"},{"_id":"66744b07ab975c85911ed26e","avatarUrl":"/avatars/913221a99619e05ca4fcd178e625a098.svg","isPro":false,"fullname":"xu","user":"wxu2023","type":"user"},{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","isPro":false,"fullname":"Xin Zhou","user":"LMD0311","type":"user"},{"_id":"68d8da2b9541f51bd687b1e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d8da2b9541f51bd687b1e2/3vO1pqV1tnMLbWhxfjnMZ.png","isPro":false,"fullname":"Hengyi Xie","user":"yoloiwig","type":"user"},{"_id":"689caaca7183a6b1d6172461","avatarUrl":"/avatars/161b54e6884093548ea7388a6a59504c.svg","isPro":false,"fullname":"lyg","user":"lyg111","type":"user"},{"_id":"69d4e377984dc690fa30a0b9","avatarUrl":"/avatars/be597d31f10da27621447c2b56f3f24b.svg","isPro":false,"fullname":"Chaoqun Zheng","user":"ThisYQ","type":"user"},{"_id":"67bd32b530dee4dde50beb3c","avatarUrl":"/avatars/eadb6e4a6a9eb68236c60ce4d3ce58ec.svg","isPro":false,"fullname":"None","user":"IsolatedONE","type":"user"},{"_id":"68d5f4340abfe8b8120a56ca","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/zBxGi50dBL0o4OpudiFno.png","isPro":false,"fullname":"yangli_leo_00","user":"YangLi00","type":"user"},{"_id":"67bb4345489cb4dc98b873cb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mRIkBXm-SfdvMUvX7xqdl.png","isPro":false,"fullname":"Wang","user":"Ruzhuo","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"687cebf73858638f66e59f56","name":"H-EmbodVis","fullname":"H-EmbodVis","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/XkebfHCVngT9o82kTAVAV.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27205.md","query":{}}">
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Abstract
Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V to L to A pathway as a direct V + L to A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.
Community
Vision-language-action (VLA) models commonly adopt an LLM-centric V→L→A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V→L→A pathway as a direct V+L→A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at this https URL.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.27205 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.27205 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.27205 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.