Hugging Face Daily Papers · · 4 min read

From Foundation to Application: Improving VLA Models in Practice

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

LingBot-VLA 2.0<br>From Foundation to Application: Improving VLA Models in Practice</p>\n","updatedAt":"2026-07-08T07:46:49.624Z","author":{"_id":"69773d4ee6878183fb90a8c7","avatarUrl":"/avatars/3e9e4081e3beaf3f69c380387b8ee4c2.svg","fullname":"Wei Wu","name":"Weiww99","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6796033382415771},"editors":["Weiww99"],"editorAvatarUrls":["/avatars/3e9e4081e3beaf3f69c380387b8ee4c2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.06403","authors":[{"_id":"6a4dff5d25849b193a834c9b","name":"Wei Wu","hidden":false},{"_id":"6a4dff5d25849b193a834c9c","name":"Fangjing Wang","hidden":false},{"_id":"6a4dff5d25849b193a834c9d","name":"Fan Lu","hidden":false},{"_id":"6a4dff5d25849b193a834c9e","name":"He Sun","hidden":false},{"_id":"6a4dff5d25849b193a834c9f","name":"Shi Liu","hidden":false},{"_id":"6a4dff5d25849b193a834ca0","name":"Yunnan Wang","hidden":false},{"_id":"6a4dff5d25849b193a834ca1","name":"Yibin Yan","hidden":false},{"_id":"6a4dff5d25849b193a834ca2","name":"Yong Wang","hidden":false},{"_id":"6a4dff5d25849b193a834ca3","name":"Shuailei Ma","hidden":false},{"_id":"6a4dff5d25849b193a834ca4","name":"Xinyang Wang","hidden":false},{"_id":"6a4dff5d25849b193a834ca5","name":"Yibin Liu","hidden":false},{"_id":"6a4dff5d25849b193a834ca6","name":"Shuai Yang","hidden":false},{"_id":"6a4dff5d25849b193a834ca7","name":"Tianxiang Zhou","hidden":false},{"_id":"6a4dff5d25849b193a834ca8","name":"Kejia Zhang","hidden":false},{"_id":"6a4dff5d25849b193a834ca9","name":"Lei Zhou","hidden":false},{"_id":"6a4dff5d25849b193a834caa","name":"Cheng Su","hidden":false},{"_id":"6a4dff5d25849b193a834cab","name":"Nan Xue","hidden":false},{"_id":"6a4dff5d25849b193a834cac","name":"Bin Tan","hidden":false},{"_id":"6a4dff5d25849b193a834cad","name":"Han Zhang","hidden":false},{"_id":"6a4dff5d25849b193a834cae","name":"Youchao Zhang","hidden":false},{"_id":"6a4dff5d25849b193a834caf","name":"Fei Liao","hidden":false},{"_id":"6a4dff5d25849b193a834cb0","name":"Xing Zhu","hidden":false},{"_id":"6a4dff5d25849b193a834cb1","name":"Yujun Shen","hidden":false},{"_id":"6a4dff5d25849b193a834cb2","name":"Kecheng Zheng","hidden":false}],"publishedAt":"2026-07-07T00:00:00.000Z","submittedOnDailyAt":"2026-07-08T00:00:00.000Z","title":"From Foundation to Application: Improving VLA Models in Practice","submittedOnDailyBy":{"_id":"69773d4ee6878183fb90a8c7","avatarUrl":"/avatars/3e9e4081e3beaf3f69c380387b8ee4c2.svg","isPro":false,"fullname":"Wei Wu","user":"Weiww99","type":"user","name":"Weiww99"},"summary":"Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.","upvotes":10,"discussionId":"6a4dff5e25849b193a834cb3","projectPage":"https://technology.robbyant.com/lingbot-vla-v2","githubRepo":"https://github.com/robbyant/lingbot-vla-v2","githubRepoAddedBy":"user","ai_summary":"LingBot-VLA 2.0 enhances generalization across tasks and embodiments through expanded data preprocessing and training on diverse robot configurations, extends action space to include whole-body degrees of freedom for complex manipulation tasks, and incorporates predictive dynamics modeling using video representation and depth estimation for improved temporal reasoning.","ai_keywords":["VLA foundation models","data processing pipeline","robot trajectories","egocentric human videos","action space","degrees of freedom","video representation model","depth estimation model","predictive dynamics modeling","temporal reasoning","GM-100 benchmark","cross-embodiment long-horizon mobile manipulation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":197,"organization":{"_id":"69709f892cd08371c1011a2e","name":"robbyant","fullname":"Robbyant","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/ZTuImney4XzRmBHyUL47F.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67aeffda7330db26f93cd62f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/ktdFt5_F05qNJ-dnD_Fk3.jpeg","isPro":false,"fullname":"jiangbonadia","user":"NadiaJiang","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"687363d49a81c7dcbcfa2d84","avatarUrl":"/avatars/5d943a5c811ed931c3fdcfee19253049.svg","isPro":false,"fullname":"jj","user":"realman123","type":"user"},{"_id":"6186ddf6a7717cb375090c01","avatarUrl":"/avatars/716b6a7d1094c8036b2a8a7b9063e8aa.svg","isPro":true,"fullname":"Julien BLANCHON","user":"blanchon","type":"user"},{"_id":"6721acf7bbf9703bc285e840","avatarUrl":"/avatars/487f5cae476758d318465aa9ed103568.svg","isPro":false,"fullname":"Yujie Zhao","user":"HomieZ","type":"user"},{"_id":"61ca43f194afcb20ceb764d7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1640645587885-noauth.jpeg","isPro":false,"fullname":"Raul Garrido","user":"happybydefault","type":"user"},{"_id":"6881b53e8579760ec4199c3a","avatarUrl":"/avatars/94041eea0fe43b4fc056d6380e8f6633.svg","isPro":false,"fullname":"StreamFormer","user":"StreamFormer","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6395b7642adf062a0b252d4e","avatarUrl":"/avatars/aa9438c4ed337836a49e22ffd7187359.svg","isPro":false,"fullname":"Wangfangjing","user":"Yioutpi","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69709f892cd08371c1011a2e","name":"robbyant","fullname":"Robbyant","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/ZTuImney4XzRmBHyUL47F.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.06403.md","query":{}}">
Papers
arxiv:2607.06403

From Foundation to Application: Improving VLA Models in Practice

Published on Jul 7
· Submitted by
Wei Wu
on Jul 8
Authors:
,

Abstract

LingBot-VLA 2.0 enhances generalization across tasks and embodiments through expanded data preprocessing and training on diverse robot configurations, extends action space to include whole-body degrees of freedom for complex manipulation tasks, and incorporates predictive dynamics modeling using video representation and depth estimation for improved temporal reasoning.

Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.

Community

Paper submitter about 9 hours ago

LingBot-VLA 2.0
From Foundation to Application: Improving VLA Models in Practice

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.06403
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.06403 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.06403 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers