Hugging Face Daily Papers · · 3 min read

Show-Harness: Just a VLM Agent Can Play Robots

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Project Page: <a href=\"https://showlab.github.io/Show-Harness/\" rel=\"nofollow\">https://showlab.github.io/Show-Harness/</a><br>GitHub: <a href=\"https://github.com/showlab/Show-Harness\" rel=\"nofollow\">https://github.com/showlab/Show-Harness</a></p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/64b7833aa5018e3c7c9b50d8/YlY2oyzajcLqlmyFRJ-fT.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>","updatedAt":"2026-09-10T02:04:20.621Z","author":{"_id":"64b7833aa5018e3c7c9b50d8","avatarUrl":"/avatars/782415605ed786b73f484fcc86a6384f.svg","fullname":"Zechen Bai","name":"ZechenBai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3919704258441925},"editors":["ZechenBai"],"editorAvatarUrls":["/avatars/782415605ed786b73f484fcc86a6384f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.10522","authors":[{"_id":"6aa20dbaa2aeb74440b1dd30","name":"Yanzhe Chen","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd31","name":"Zechen Bai","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd32","name":"Zhijun Cao","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd33","name":"Wenzheng Zeng","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd34","name":"Kevin Qinghong Lin","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd35","name":"Yiqi Lin","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd36","name":"Guoqiang Liang","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd37","name":"Kevin Yuchen Ma","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd38","name":"Qiming Huang","hidden":false},{"_id":"6aa20dbaa2aeb74440b1dd39","name":"Mike Zheng Shou","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64b7833aa5018e3c7c9b50d8/ZgOy5kpA3_e-hBTdppgRl.mp4"],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Show-Harness: Just a VLM Agent Can Play Robots","submittedOnDailyBy":{"_id":"64b7833aa5018e3c7c9b50d8","avatarUrl":"/avatars/782415605ed786b73f484fcc86a6384f.svg","isPro":false,"fullname":"Zechen Bai","user":"ZechenBai","type":"user","name":"ZechenBai"},"summary":"Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to \"play\" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to \"play\" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.","upvotes":45,"discussionId":"6aa20dbba2aeb74440b1dd3a","projectPage":"https://showlab.github.io/Show-Harness/","githubRepo":"https://github.com/showlab/Show-Harness","githubRepoAddedBy":"user","ai_summary":"Show-Harness links vision-language models to robot control via discrete semantic actions interpreted by embodiment-specific modules, enabling zero-shot and efficient fine-tuned deployment across robots and GUIs.","ai_keywords":["vision-language models","Show-Harness","Embodied Harness","semantic action units","embodiment-specific interpreters","zero-shot robot control","GUMI","GUI Manipulation Interface","VLA paradigms"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"63a553c4ce5763e06f78669c","name":"showlab","fullname":"Show Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671779505215-63a55320ce5763e06f78519c.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b7833aa5018e3c7c9b50d8","avatarUrl":"/avatars/782415605ed786b73f484fcc86a6384f.svg","isPro":false,"fullname":"Zechen Bai","user":"ZechenBai","type":"user"},{"_id":"6777819e38f9a731d4039535","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6777819e38f9a731d4039535/QzpGJhc_9Fzp_yM1uB-Zi.jpeg","isPro":false,"fullname":"CaoZhijun","user":"aaroncaozj","type":"user"},{"_id":"6a9679c4e51b9509ec5a3323","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/r8xZA14qqJuHGsr63dv_U.jpeg","isPro":false,"fullname":"Aurora CAO","user":"auroracaozj","type":"user"},{"_id":"6836a5aa14ebb7ff0c031871","avatarUrl":"/avatars/8bb04314e9b1cb3927b159f511796703.svg","isPro":false,"fullname":"Yang Pei","user":"yangpei-comp","type":"user"},{"_id":"657669f2b4379e65a8c6d5cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/657669f2b4379e65a8c6d5cf/9M7lS0Old5NqOrNx7W04N.jpeg","isPro":false,"fullname":"YanzheChen","user":"YanzheChen","type":"user"},{"_id":"69a2712b6b50254687e9eb62","avatarUrl":"/avatars/228e1691b144c996a268f83c15652686.svg","isPro":false,"fullname":"Shane","user":"Shane0823","type":"user"},{"_id":"672d7efba3cac338ea972881","avatarUrl":"/avatars/ff6ef1c4e08ef66c893e74aec8b5750c.svg","isPro":false,"fullname":"Xiaokang Liu","user":"hiskiv","type":"user"},{"_id":"66ec5888518ccb4e621fb419","avatarUrl":"/avatars/983be404e24f7679c7e7004123594ed9.svg","isPro":false,"fullname":"Hulingxiao He","user":"StevenHH2000","type":"user"},{"_id":"6996f56256426abbe6e6614a","avatarUrl":"/avatars/8a15d472223b6e538f823312566c2f6c.svg","isPro":false,"fullname":"Mue9rgobb4ay","user":"mue9rgobb4ay","type":"user"},{"_id":"66fbea5c205892e8bd828008","avatarUrl":"/avatars/4db9ac717cdc906ca3a03c19549007fc.svg","isPro":false,"fullname":"tianzl","user":"tianzl66","type":"user"},{"_id":"65c100779178d07d3a2c44d7","avatarUrl":"/avatars/a62c4e4a2908099b4cc0645af2c54879.svg","isPro":false,"fullname":"Kevin Ma","user":"Kevinskwk","type":"user"},{"_id":"674963d1818d522a4d44dc92","avatarUrl":"/avatars/5a2196620b110f9e9552e16554c7dcfb.svg","isPro":false,"fullname":"Chen Gao","user":"chen-gao","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"63a553c4ce5763e06f78669c","name":"showlab","fullname":"Show Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671779505215-63a55320ce5763e06f78519c.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.10522.md","query":{}}">
Papers
arxiv:2609.10522

Show-Harness: Just a VLM Agent Can Play Robots

Published on Sep 9
· Submitted by
Zechen Bai
on Sep 10
#1 Paper of the day
Authors:
,

Abstract

Show-Harness links vision-language models to robot control via discrete semantic actions interpreted by embodiment-specific modules, enabling zero-shot and efficient fine-tuned deployment across robots and GUIs.

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.

Community

This comment has been hidden (marked as Resolved)
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.10522
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.10522 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers