Hugging Face Daily Papers · · 4 min read

HumanCLAW: Can Vision-Language Models Act Through a Body?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A model may know where the sofa is, yet still not know how to move its body there and sit down. We call this ability <em>Action Intelligence</em>, and <strong>HumanCLAW</strong> makes it measurable without tying it to motor control. This is a first step toward models that can reason and act through a physical body.</p>\n","updatedAt":"2026-07-30T02:12:31.338Z","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9553734660148621},"editors":["kuvvi"],"editorAvatarUrls":["/avatars/e12efb8e030688a0afcc72176b453fb3.svg"],"reactions":[],"isReport":false}},{"id":"6a6abd110dfadb68854b893f","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false},"createdAt":"2026-07-30T02:55:13.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/MSBZV9moiVbvkRNwQr1W1.mp4\n","html":"<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/MSBZV9moiVbvkRNwQr1W1.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-30T02:55:13.223Z","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.504278838634491},"editors":["kuvvi"],"editorAvatarUrls":["/avatars/e12efb8e030688a0afcc72176b453fb3.svg"],"reactions":[],"isReport":false}},{"id":"6a6abdb6fdc6c38dcb9ba391","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false},"createdAt":"2026-07-30T02:57:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"![截屏2026-07-30 10.57.05](https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/lqmp8Xmx9scLetR5tmRiQ.png)\n","html":"<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/lqmp8Xmx9scLetR5tmRiQ.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/lqmp8Xmx9scLetR5tmRiQ.png\" alt=\"截屏2026-07-30 10.57.05\"></a></p>\n","updatedAt":"2026-07-30T02:57:58.230Z","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.45945918560028076},"editors":["kuvvi"],"editorAvatarUrls":["/avatars/e12efb8e030688a0afcc72176b453fb3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27180","authors":[{"_id":"6a6aacf34463a8a84bdc3f36","name":"Siyao Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f37","name":"Jiawei Gu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f38","name":"Shuai Liu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f39","name":"Kairui Hu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3a","name":"Zekun Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3b","name":"Linjie Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3c","name":"Chengcheng Tang","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3d","name":"Po-Chen Wu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3e","name":"Ivan Shugurov","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3f","name":"Lingni Ma","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f40","name":"Michael Zollhoefer","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f41","name":"Sizhe An","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f42","name":"Abhay Mittal","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f43","name":"Amy Zhao","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f44","name":"Ranjay Krishna","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f45","name":"Manling Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f46","name":"Ziwei Liu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f47","name":"Chuan Guo","hidden":false}],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"HumanCLAW: Can Vision-Language Models Act Through a Body?","submittedOnDailyBy":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","isPro":false,"fullname":"Jiawei Gu","user":"kuvvi","type":"user","name":"kuvvi"},"summary":"Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.","upvotes":55,"discussionId":"6a6aacf34463a8a84bdc3f48","projectPage":"https://human-claw.github.io/","githubRepo":"https://github.com/Human-CLAW/HumanCLAW","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"66b54027408752ae16404b05","name":"metaresearch","fullname":"Meta Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66b25f3f58babfaeb76112dc/2GmiaF075AZ7BcE538oPk.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","isPro":false,"fullname":"Jiawei Gu","user":"kuvvi","type":"user"},{"_id":"6400ba2b261cfa61f3a00555","avatarUrl":"/avatars/1311e0b5e21b1c94d73fcaf455d3c7f7.svg","isPro":false,"fullname":"Kairui","user":"KairuiHu","type":"user"},{"_id":"655fb67bc40c3a6a0d86f3ca","avatarUrl":"/avatars/242305c3d7106fa38aa9e4e20fcf191b.svg","isPro":false,"fullname":"Zekun Li","user":"kunkun0w0","type":"user"},{"_id":"642bfdbbc885078517154abb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bfdbbc885078517154abb/lLmc8_3VnTyk4qOvl-omI.jpeg","isPro":false,"fullname":"Chong Zhou","user":"chongzhou","type":"user"},{"_id":"645223fb01d7bd9555ea399a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645223fb01d7bd9555ea399a/fVR7XmGg6pMSRKx9sEdvT.png","isPro":false,"fullname":"Zhiyang Dou","user":"frankzydou","type":"user"},{"_id":"656e9db559afca5707c493c5","avatarUrl":"/avatars/ffe72a82aaf42c80eb988344415672ba.svg","isPro":false,"fullname":"Tao Lu","user":"Isaaclt","type":"user"},{"_id":"65fa40d26902f37123737aeb","avatarUrl":"/avatars/a72154efb760d69bdef4f20904ca0ec7.svg","isPro":false,"fullname":"Xueying","user":"XueyingJiang","type":"user"},{"_id":"65a7c0335e79abfa2ec30c52","avatarUrl":"/avatars/2f62f83f9c5c4cc9444571f067cd85b7.svg","isPro":false,"fullname":"Shuangrui Ding","user":"Mar2Ding","type":"user"},{"_id":"64bb77e786e7fb5b8a317a43","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bb77e786e7fb5b8a317a43/J0jOrlZJ9gazdYaeSH2Bo.png","isPro":false,"fullname":"kcz","user":"kcz358","type":"user"},{"_id":"62a993d80472c0b7f94027df","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62a993d80472c0b7f94027df/j5vp-IwLA2YBexylUHiQU.png","isPro":false,"fullname":"Zhang Yuanhan","user":"ZhangYuanhan","type":"user"},{"_id":"655c70d331c4978366d4b2e6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/655c70d331c4978366d4b2e6/X-KjTNkxtzeYu9ngBOh_C.jpeg","isPro":false,"fullname":"yiyexy","user":"yiyexy","type":"user"},{"_id":"64c38b3413dc689c2f12f03f","avatarUrl":"/avatars/d49bf90bb6860189a761b9f5773c09fc.svg","isPro":false,"fullname":"Zhenghai Xue","user":"ZhenghaiXue","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"66b54027408752ae16404b05","name":"metaresearch","fullname":"Meta Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66b25f3f58babfaeb76112dc/2GmiaF075AZ7BcE538oPk.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27180.md","query":{}}">
Papers
arxiv:2607.27180

HumanCLAW: Can Vision-Language Models Act Through a Body?

Published on Jul 29
· Submitted by
Jiawei Gu
on Jul 30
#2 Paper of the day
Authors:
,

Abstract

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.

Community

Paper submitter about 6 hours ago

A model may know where the sofa is, yet still not know how to move its body there and sit down. We call this ability Action Intelligence, and HumanCLAW makes it measurable without tying it to motor control. This is a first step toward models that can reason and act through a physical body.

Paper submitter about 5 hours ago

Paper submitter about 5 hours ago

截屏2026-07-30 10.57.05

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.27180
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.27180 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.27180 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.27180 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers