Hugging Face Daily Papers · · 4 min read

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<strong>Strong, open, and research-friendly.</strong> VLAct releases the data, models, and complete training/fine-tuning pipeline, with full continued pre-training requiring only <strong>16 GPUs</strong>. It achieves <strong>92.5% on RoboTwin 2.0</strong> and <strong>ranks #6 on RoboDojo</strong> by Success Rate, outperforming all World Action Models.</p>\n","updatedAt":"2026-08-31T02:01:21.258Z","author":{"_id":"6527b7280ae663e384eb8499","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg","fullname":"Senqiao Yang","name":"Senqiao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":16,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8025529980659485},"editors":["Senqiao"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.27550","authors":[{"_id":"6a94dc1e073195fee51571e2","name":"Senqiao Yang","hidden":false},{"_id":"6a94dc1e073195fee51571e3","name":"Chengyao Wang","hidden":false},{"_id":"6a94dc1e073195fee51571e4","name":"Yuxin Chen","hidden":false},{"_id":"6a94dc1e073195fee51571e5","name":"Zixuan Wang","hidden":false},{"_id":"6a94dc1e073195fee51571e6","name":"Longxiang Tang","hidden":false},{"_id":"6a94dc1e073195fee51571e7","name":"Haokun Gui","hidden":false},{"_id":"6a94dc1e073195fee51571e8","name":"Jinhui Ye","hidden":false},{"_id":"6a94dc1e073195fee51571e9","name":"Changsheng Lu","hidden":false},{"_id":"6a94dc1e073195fee51571ea","name":"Xiaoyang Wu","hidden":false},{"_id":"6a94dc1e073195fee51571eb","name":"Mingkang Zhu","hidden":false},{"_id":"6a94dc1e073195fee51571ec","name":"Pengguang Chen","hidden":false},{"_id":"6a94dc1e073195fee51571ed","name":"Shu Liu","hidden":false},{"_id":"6a94dc1e073195fee51571ee","name":"Zhuotao Tian","hidden":false},{"_id":"6a94dc1e073195fee51571ef","name":"Hengshuang Zhao","hidden":false},{"_id":"6a94dc1e073195fee51571f0","name":"Bei Yu","hidden":false},{"_id":"6a94dc1e073195fee51571f1","name":"Jiaya Jia","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6527b7280ae663e384eb8499/ph-Pvffp0gEKnfXl4MJTZ.mp4"],"publishedAt":"2026-08-27T00:00:00.000Z","submittedOnDailyAt":"2026-08-31T00:00:00.000Z","title":"Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models","submittedOnDailyBy":{"_id":"6527b7280ae663e384eb8499","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg","isPro":false,"fullname":"Senqiao Yang","user":"Senqiao","type":"user","name":"Senqiao"},"summary":"Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.","upvotes":23,"discussionId":"6a94dc1e073195fee51571f2","projectPage":"https://starvla.github.io/VLAct","githubRepo":"https://github.com/starVLA/VLAct","githubRepoAddedBy":"user","ai_summary":"VLAct improves vision-language-action model performance by pre-training on diverse robot data with preserved vision-language priors and shared action semantics, achieving strong results across simulations and unseen embodiments with limited compute.","ai_keywords":["Vision-Language-Action","VLA","VLM","multi-embodiment","cross-embodiment action layout","continuous action co-supervision","VLM-prior preservation","world-action model"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":6,"organization":{"_id":"6894537144aa3b85fc4a03fc","name":"StarVLA","fullname":"Star Vision Language Action Models","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662238bc9b0c7e78df0ffe35/k0CX50oOCJZ5C_DPjY2tG.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642e3bcb958faf258a40e89c","avatarUrl":"/avatars/dad142df2217f8eed1f45c9e7287d3ea.svg","isPro":false,"fullname":"Ruihang Chu","user":"Ruihang","type":"user"},{"_id":"6527b7280ae663e384eb8499","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6527b7280ae663e384eb8499/73yF3eu2cUx7jVZrhXnXx.jpeg","isPro":false,"fullname":"Senqiao Yang","user":"Senqiao","type":"user"},{"_id":"6761978dc3fabedfe434947a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/rdC16QDief2cfAQiUBOFe.png","isPro":false,"fullname":"vvcy","user":"vvcy1122","type":"user"},{"_id":"6423e35b30b0e4ab36dd1b16","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6423e35b30b0e4ab36dd1b16/pea6LVDS9PQAxvQt9GUiZ.jpeg","isPro":true,"fullname":"Wang Chengyao","user":"wcy1122","type":"user"},{"_id":"68e93364c913af54cf55f61a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/xE4hEpTo2C_oEmVVf5wv4.png","isPro":false,"fullname":"wang","user":"vvcy2233","type":"user"},{"_id":"66e1557c75d9226ba13b38d2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/vAwxAOynnsF2iqYeO3UBl.png","isPro":false,"fullname":"shenghe zheng","user":"desimfj","type":"user"},{"_id":"67640f123ebd56f692b00d9f","avatarUrl":"/avatars/db00af73849efa7a48639503389e1b96.svg","isPro":false,"fullname":"jack","user":"kimi000","type":"user"},{"_id":"6492a0d8d4ae24c933ace44d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6492a0d8d4ae24c933ace44d/FXYIucGnWkMDu4gHWy1qw.jpeg","isPro":false,"fullname":"Longxiang Tang","user":"lloong","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"674913ccf73677ad66dd950e","avatarUrl":"/avatars/4101c3ac384f797103dfc1523c14336d.svg","isPro":false,"fullname":"Yi Lin","user":"Yi-Fighter","type":"user"},{"_id":"67490cb036d71af194c4fc59","avatarUrl":"/avatars/aa7cea64511c2adb2b4a34031a76046d.svg","isPro":false,"fullname":"Steven Alex","user":"111Steven111","type":"user"},{"_id":"653e5d31ffd60206c8b64bb5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653e5d31ffd60206c8b64bb5/bgztraPC27L6culMlJw4s.png","isPro":false,"fullname":"Xinchen Zhang","user":"comin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6894537144aa3b85fc4a03fc","name":"StarVLA","fullname":"Star Vision Language Action Models","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662238bc9b0c7e78df0ffe35/k0CX50oOCJZ5C_DPjY2tG.png"},"query":{}}">
Papers
arxiv:2608.27550

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Published on Aug 27
· Submitted by
Senqiao Yang
on Aug 31
Authors:
,

Abstract

VLAct improves vision-language-action model performance by pre-training on diverse robot data with preserved vision-language priors and shared action semantics, achieving strong results across simulations and unseen embodiments with limited compute.

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.

Community

Paper submitter about 6 hours ago

Strong, open, and research-friendly. VLAct releases the data, models, and complete training/fine-tuning pipeline, with full continued pre-training requiring only 16 GPUs. It achieves 92.5% on RoboTwin 2.0 and ranks #6 on RoboDojo by Success Rate, outperforming all World Action Models.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.27550 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.27550 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.27550 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers