Hugging Face Daily Papers · · 6 min read

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven stateof-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-torobot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/d32q0jPQNf8mApagcDJ_N.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/d32q0jPQNf8mApagcDJ_N.png\" alt=\"Clipboard_Screenshot_1786680870\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/eMSQrbjhHLXiKjLGW7C0e.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/eMSQrbjhHLXiKjLGW7C0e.png\" alt=\"Clipboard_Screenshot_1786680893\"></a></p>\n","updatedAt":"2026-08-14T04:16:21.468Z","author":{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","fullname":"chaofanma","name":"chaofanma","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8105422258377075},"editors":["chaofanma"],"editorAvatarUrls":["/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.13049","authors":[{"_id":"6a7e93ae42823931a1f176cb","name":"Dingyi Rong","hidden":false},{"_id":"6a7e93ae42823931a1f176cc","name":"Yue Shi","hidden":false},{"_id":"6a7e93ae42823931a1f176cd","name":"Chaofan Ma","hidden":false},{"_id":"6a7e93ae42823931a1f176ce","name":"Jiezhang Cao","hidden":false},{"_id":"6a7e93ae42823931a1f176cf","name":"Zongrui Wang","hidden":false},{"_id":"6a7e93ae42823931a1f176d0","name":"Zeyu Zhang","hidden":false},{"_id":"6a7e93ae42823931a1f176d1","name":"Yao Mu","hidden":false},{"_id":"6a7e93ae42823931a1f176d2","name":"Guangtao Zhai","hidden":false},{"_id":"6a7e93ae42823931a1f176d3","name":"Ning Liu","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models","submittedOnDailyBy":{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","isPro":false,"fullname":"chaofanma","user":"chaofanma","type":"user","name":"chaofanma"},"summary":"Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.","upvotes":5,"discussionId":"6a7e93af42823931a1f176d4","projectPage":"https://rongdingyi.github.io/H2R-Bench/","githubRepo":"https://github.com/Rongdingyi/H2R-Bench","githubRepoAddedBy":"user","ai_summary":"H2R-Bench evaluates video generation models on transforming human manipulation videos into robot-centric demonstrations across embodiment constraints and interaction fidelity.","ai_keywords":["video world models","cross-embodiment transfer","human-to-robot manipulation","egocentric video","embodiment constraints","functional contact","action events","goal-state completion"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","isPro":false,"fullname":"chaofanma","user":"chaofanma","type":"user"},{"_id":"675a621105c46a17e6a229b3","avatarUrl":"/avatars/1b51f141df9d206fcbd2598b6e994aa6.svg","isPro":false,"fullname":"Dingyi Rong","user":"dingyi11","type":"user"},{"_id":"6a069177a01745697eb21189","avatarUrl":"/avatars/2cff7b99d89036a327c47e1cf3220617.svg","isPro":false,"fullname":"shiyue001","user":"shiyue0011","type":"user"},{"_id":"6655d5575b8ab1ed4f66265d","avatarUrl":"/avatars/1fd6da28eba1c804cad1cc490b374eac.svg","isPro":true,"fullname":"Chen Ye","user":"sjtuchenye","type":"user"},{"_id":"6731af65389aca4be7ce8a75","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6Ym2bfkiJzKOtDZ3LCdFg.png","isPro":false,"fullname":"Cumulus","user":"CumulusAlpha","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.13049.md","query":{}}">
Papers
arxiv:2608.13049

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Published on Aug 13
· Submitted by
chaofanma
on Aug 14
Authors:
,

Abstract

H2R-Bench evaluates video generation models on transforming human manipulation videos into robot-centric demonstrations across embodiment constraints and interaction fidelity.

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

Community

Paper submitter about 8 hours ago

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven stateof-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-torobot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

Clipboard_Screenshot_1786680870

Clipboard_Screenshot_1786680893

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.13049
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.13049 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.13049 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.13049 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers