A model may know where the sofa is, yet still not know how to move its body there and sit down. We call this ability <em>Action Intelligence</em>, and <strong>HumanCLAW</strong> makes it measurable without tying it to motor control. This is a first step toward models that can reason and act through a physical body.</p>\n","updatedAt":"2026-07-30T02:12:31.338Z","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9553734660148621},"editors":["kuvvi"],"editorAvatarUrls":["/avatars/e12efb8e030688a0afcc72176b453fb3.svg"],"reactions":[],"isReport":false}},{"id":"6a6abd110dfadb68854b893f","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false},"createdAt":"2026-07-30T02:55:13.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/MSBZV9moiVbvkRNwQr1W1.mp4\n","html":"<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/MSBZV9moiVbvkRNwQr1W1.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-30T02:55:13.223Z","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.504278838634491},"editors":["kuvvi"],"editorAvatarUrls":["/avatars/e12efb8e030688a0afcc72176b453fb3.svg"],"reactions":[],"isReport":false}},{"id":"6a6abdb6fdc6c38dcb9ba391","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false},"createdAt":"2026-07-30T02:57:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\n","html":"<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/lqmp8Xmx9scLetR5tmRiQ.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/645b4819f9d4ec91fdd54852/lqmp8Xmx9scLetR5tmRiQ.png\" alt=\"截屏2026-07-30 10.57.05\"></a></p>\n","updatedAt":"2026-07-30T02:57:58.230Z","author":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","fullname":"Jiawei Gu","name":"kuvvi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.45945918560028076},"editors":["kuvvi"],"editorAvatarUrls":["/avatars/e12efb8e030688a0afcc72176b453fb3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27180","authors":[{"_id":"6a6aacf34463a8a84bdc3f36","name":"Siyao Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f37","name":"Jiawei Gu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f38","name":"Shuai Liu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f39","name":"Kairui Hu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3a","name":"Zekun Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3b","name":"Linjie Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3c","name":"Chengcheng Tang","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3d","name":"Po-Chen Wu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3e","name":"Ivan Shugurov","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f3f","name":"Lingni Ma","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f40","name":"Michael Zollhoefer","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f41","name":"Sizhe An","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f42","name":"Abhay Mittal","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f43","name":"Amy Zhao","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f44","name":"Ranjay Krishna","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f45","name":"Manling Li","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f46","name":"Ziwei Liu","hidden":false},{"_id":"6a6aacf34463a8a84bdc3f47","name":"Chuan Guo","hidden":false}],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"HumanCLAW: Can Vision-Language Models Act Through a Body?","submittedOnDailyBy":{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","isPro":false,"fullname":"Jiawei Gu","user":"kuvvi","type":"user","name":"kuvvi"},"summary":"Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.","upvotes":55,"discussionId":"6a6aacf34463a8a84bdc3f48","projectPage":"https://human-claw.github.io/","githubRepo":"https://github.com/Human-CLAW/HumanCLAW","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"66b54027408752ae16404b05","name":"metaresearch","fullname":"Meta Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66b25f3f58babfaeb76112dc/2GmiaF075AZ7BcE538oPk.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","isPro":false,"fullname":"Jiawei Gu","user":"kuvvi","type":"user"},{"_id":"6400ba2b261cfa61f3a00555","avatarUrl":"/avatars/1311e0b5e21b1c94d73fcaf455d3c7f7.svg","isPro":false,"fullname":"Kairui","user":"KairuiHu","type":"user"},{"_id":"655fb67bc40c3a6a0d86f3ca","avatarUrl":"/avatars/242305c3d7106fa38aa9e4e20fcf191b.svg","isPro":false,"fullname":"Zekun Li","user":"kunkun0w0","type":"user"},{"_id":"642bfdbbc885078517154abb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bfdbbc885078517154abb/lLmc8_3VnTyk4qOvl-omI.jpeg","isPro":false,"fullname":"Chong Zhou","user":"chongzhou","type":"user"},{"_id":"645223fb01d7bd9555ea399a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/645223fb01d7bd9555ea399a/fVR7XmGg6pMSRKx9sEdvT.png","isPro":false,"fullname":"Zhiyang Dou","user":"frankzydou","type":"user"},{"_id":"656e9db559afca5707c493c5","avatarUrl":"/avatars/ffe72a82aaf42c80eb988344415672ba.svg","isPro":false,"fullname":"Tao Lu","user":"Isaaclt","type":"user"},{"_id":"65fa40d26902f37123737aeb","avatarUrl":"/avatars/a72154efb760d69bdef4f20904ca0ec7.svg","isPro":false,"fullname":"Xueying","user":"XueyingJiang","type":"user"},{"_id":"65a7c0335e79abfa2ec30c52","avatarUrl":"/avatars/2f62f83f9c5c4cc9444571f067cd85b7.svg","isPro":false,"fullname":"Shuangrui Ding","user":"Mar2Ding","type":"user"},{"_id":"64bb77e786e7fb5b8a317a43","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bb77e786e7fb5b8a317a43/J0jOrlZJ9gazdYaeSH2Bo.png","isPro":false,"fullname":"kcz","user":"kcz358","type":"user"},{"_id":"62a993d80472c0b7f94027df","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62a993d80472c0b7f94027df/j5vp-IwLA2YBexylUHiQU.png","isPro":false,"fullname":"Zhang Yuanhan","user":"ZhangYuanhan","type":"user"},{"_id":"655c70d331c4978366d4b2e6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/655c70d331c4978366d4b2e6/X-KjTNkxtzeYu9ngBOh_C.jpeg","isPro":false,"fullname":"yiyexy","user":"yiyexy","type":"user"},{"_id":"64c38b3413dc689c2f12f03f","avatarUrl":"/avatars/d49bf90bb6860189a761b9f5773c09fc.svg","isPro":false,"fullname":"Zhenghai Xue","user":"ZhenghaiXue","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"66b54027408752ae16404b05","name":"metaresearch","fullname":"Meta Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66b25f3f58babfaeb76112dc/2GmiaF075AZ7BcE538oPk.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27180.md","query":{}}">
HumanCLAW: Can Vision-Language Models Act Through a Body?
Abstract
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
Community
A model may know where the sofa is, yet still not know how to move its body there and sit down. We call this ability Action Intelligence, and HumanCLAW makes it measurable without tying it to motor control. This is a first step toward models that can reason and act through a physical body.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.27180 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.27180 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.27180 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.