Hugging Face Daily Papers · · 5 min read

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🤖 <strong>What if a robot could compare possible futures before deciding what to do next?</strong></p>\n<p>We introduce <strong>τ₀-VLA</strong>, a hierarchical robot foundation model for long-horizon manipulation. Its high-level policy maintains execution memory and, when a decision is uncertain, allocates additional test-time computation to propose candidate subtasks, predict their visual consequences with a world model, and compare alternatives before committing. A generalist low-level VLA then executes the selected subtask across robot embodiments.</p>\n<p><strong>Highlights:</strong></p>\n<ul>\n<li>The low-level policy is trained on <strong>40,115 hours</strong> of heterogeneous real-world robot data with multimodal co-training.</li>\n<li>Selective test-time computation improves next-subtask prediction accuracy by <strong>15–24 percentage points</strong> across in-domain and distribution-shifted settings.</li>\n<li>We evaluate real-world manipulation tasks containing <strong>13–25 ordered steps</strong>, with episodes lasting up to <strong>12 minutes</strong>.</li>\n<li>Using the same low-level policy, hierarchical planning improves average closed-loop success from <strong>27.5% to 45.0%</strong> across four long-horizon tasks.</li>\n<li>We release the official code and pretrained low-level VLA checkpoint, with high-level policy on the way.</li>\n</ul>\n<p>🌐 <a href=\"https://tau0-vla.github.io/\" rel=\"nofollow\">Project page</a><br>💻 <a href=\"https://github.com/sii-research/tau-0-vla\" rel=\"nofollow\">Code</a><br>🤗 <a href=\"https://huggingface.co/sii-research/tau-0-vla\">Model checkpoint</a></p>\n<p>Questions and feedback are very welcome!</p>\n","updatedAt":"2026-08-21T15:57:22.416Z","author":{"_id":"64f09372f0a34dcdcecb5791","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f09372f0a34dcdcecb5791/55GPtTwHc8nTlev3si1HP.jpeg","fullname":"jrryzh(SII)","name":"J3rr1","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8214868903160095},"editors":["J3rr1"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64f09372f0a34dcdcecb5791/55GPtTwHc8nTlev3si1HP.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16885","authors":[{"_id":"6a841a35b153becad167707d","name":"Xiaowei Cai","hidden":false},{"_id":"6a841a35b153becad167707e","name":"Yunuo Cai","hidden":false},{"_id":"6a841a35b153becad167707f","name":"Bingao Chen","hidden":false},{"_id":"6a841a35b153becad1677080","name":"Jingxiao Chen","hidden":false},{"_id":"6a841a35b153becad1677081","name":"Zhi Chen","hidden":false},{"_id":"6a841a35b153becad1677082","name":"Siyuan Feng","hidden":false},{"_id":"6a841a35b153becad1677083","name":"Tengyu Hou","hidden":false},{"_id":"6a841a35b153becad1677084","name":"Jingshun Huang","hidden":false},{"_id":"6a841a35b153becad1677085","name":"Han Jiang","hidden":false},{"_id":"6a841a35b153becad1677086","name":"Runkun Ju","hidden":false},{"_id":"6a841a35b153becad1677087","name":"Dong Li","hidden":false},{"_id":"6a841a35b153becad1677088","name":"Mingxiang Li","hidden":false},{"_id":"6a841a35b153becad1677089","name":"Shaowei Li","hidden":false},{"_id":"6a841a35b153becad167708a","name":"Xinchen Li","hidden":false},{"_id":"6a841a35b153becad167708b","name":"Yifan Li","hidden":false},{"_id":"6a841a35b153becad167708c","name":"Yi Liu","hidden":false},{"_id":"6a841a35b153becad167708d","name":"Zhongyuan Liu","hidden":false},{"_id":"6a841a35b153becad167708e","name":"Jianlan Luo","hidden":false},{"_id":"6a841a35b153becad167708f","name":"Junwen Miao","hidden":false},{"_id":"6a841a35b153becad1677090","name":"Ruiqi Ni","hidden":false},{"_id":"6a841a35b153becad1677091","name":"Buqing Nie","hidden":false},{"_id":"6a841a35b153becad1677092","name":"Mingjie Pan","hidden":false},{"_id":"6a841a35b153becad1677093","name":"Xinlin Ren","hidden":false},{"_id":"6a841a35b153becad1677094","name":"Jianheng Song","hidden":false},{"_id":"6a841a35b153becad1677095","name":"Jiaxu Wang","hidden":false},{"_id":"6a841a35b153becad1677096","name":"Peiqi Wang","hidden":false},{"_id":"6a841a35b153becad1677097","name":"Sen Wang","hidden":false},{"_id":"6a841a35b153becad1677098","name":"Xiaoyan Wang","hidden":false},{"_id":"6a841a35b153becad1677099","name":"Dafeng Wei","hidden":false},{"_id":"6a841a35b153becad167709a","name":"Dongming Wu","hidden":false},{"_id":"6a841a35b153becad167709b","name":"Pengwei Xie","hidden":false},{"_id":"6a841a35b153becad167709c","name":"Pu Yang","hidden":false},{"_id":"6a841a35b153becad167709d","name":"Hangjian Ye","hidden":false},{"_id":"6a841a35b153becad167709e","name":"Xiangyu Yue","hidden":false},{"_id":"6a841a35b153becad167709f","user":{"_id":"64f09372f0a34dcdcecb5791","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f09372f0a34dcdcecb5791/55GPtTwHc8nTlev3si1HP.jpeg","isPro":false,"fullname":"jrryzh(SII)","user":"J3rr1","type":"user","name":"J3rr1"},"name":"Jinyu Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-21T13:19:27.798Z","hidden":false},{"_id":"6a841a35b153becad16770a0","name":"Qinglin Zhang","hidden":false},{"_id":"6a841a35b153becad16770a1","name":"Xueyong Zhao","hidden":false},{"_id":"6a841a35b153becad16770a2","name":"Pengfei Zhou","hidden":false},{"_id":"6a841a35b153becad16770a3","name":"Yue Zhou","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64f09372f0a34dcdcecb5791/MWnnwwK6Gr5AbMnvQL5Vj.mp4","https://cdn-uploads.huggingface.co/production/uploads/64f09372f0a34dcdcecb5791/KIhc2FQknI8N3CLLC0OkU.png","https://cdn-uploads.huggingface.co/production/uploads/64f09372f0a34dcdcecb5791/SW341izGJFsFT_E9jRsAN.png"],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation","submittedOnDailyBy":{"_id":"64f09372f0a34dcdcecb5791","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f09372f0a34dcdcecb5791/55GPtTwHc8nTlev3si1HP.jpeg","isPro":false,"fullname":"jrryzh(SII)","user":"J3rr1","type":"user","name":"J3rr1"},"summary":"Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.","upvotes":2,"discussionId":"6a841a35b153becad16770a4","projectPage":"https://tau0-vla.github.io/","githubRepo":"https://github.com/sii-research/tau-0-vla","githubRepoAddedBy":"user","ai_summary":"A hierarchical vision-language-action model improves long-horizon robot manipulation by using world-model-guided test-time search to scale computation for high-level subtask decisions.","ai_keywords":["vision-language-action","hierarchical VLA","world-model-guided test-time computation","execution memory","subtask generation","low-level policy","multimodal co-training"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":514,"organization":{"_id":"683ebd0d913d82e703e77286","name":"sii-research","fullname":"Shanghai Innovation Institute","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6144a0c4ff1146bbd84d9865/SQAtyVRxNjp9L0CUi0tgI.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64f09372f0a34dcdcecb5791","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f09372f0a34dcdcecb5791/55GPtTwHc8nTlev3si1HP.jpeg","isPro":false,"fullname":"jrryzh(SII)","user":"J3rr1","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"683ebd0d913d82e703e77286","name":"sii-research","fullname":"Shanghai Innovation Institute","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6144a0c4ff1146bbd84d9865/SQAtyVRxNjp9L0CUi0tgI.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16885.md","query":{}}">
Papers
arxiv:2608.16885

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Published on Aug 17
· Submitted by
jrryzh(SII)
on Aug 21
Authors:
,

Abstract

A hierarchical vision-language-action model improves long-horizon robot manipulation by using world-model-guided test-time search to scale computation for high-level subtask decisions.

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

Community

Paper author Paper submitter about 4 hours ago

🤖 What if a robot could compare possible futures before deciding what to do next?

We introduce τ₀-VLA, a hierarchical robot foundation model for long-horizon manipulation. Its high-level policy maintains execution memory and, when a decision is uncertain, allocates additional test-time computation to propose candidate subtasks, predict their visual consequences with a world model, and compare alternatives before committing. A generalist low-level VLA then executes the selected subtask across robot embodiments.

Highlights:

  • The low-level policy is trained on 40,115 hours of heterogeneous real-world robot data with multimodal co-training.
  • Selective test-time computation improves next-subtask prediction accuracy by 15–24 percentage points across in-domain and distribution-shifted settings.
  • We evaluate real-world manipulation tasks containing 13–25 ordered steps, with episodes lasting up to 12 minutes.
  • Using the same low-level policy, hierarchical planning improves average closed-loop success from 27.5% to 45.0% across four long-horizon tasks.
  • We release the official code and pretrained low-level VLA checkpoint, with high-level policy on the way.

🌐 Project page
💻 Code
🤗 Model checkpoint

Questions and feedback are very welcome!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16885
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16885 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16885 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers