Hugging Face Daily Papers · · 4 min read

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

SIGGRAPH Aisa 2026; Project page: <a href=\"https://cosmoh2g.github.io/\" rel=\"nofollow\">https://cosmoh2g.github.io/</a></p>\n","updatedAt":"2026-09-09T04:17:09.768Z","author":{"_id":"64749a0d5aba8edfb2eeaba7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png","fullname":"Mutian Xu","name":"Minoday","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4018838107585907},"editors":["Minoday"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07498","authors":[{"_id":"6aa0d4ccd0174964227bed26","name":"Hongxiang Zhao","hidden":false},{"_id":"6aa0d4ccd0174964227bed27","name":"Mutian Xu","hidden":false},{"_id":"6aa0d4ccd0174964227bed28","name":"Zeyu Jin","hidden":false},{"_id":"6aa0d4ccd0174964227bed29","name":"Yiming Hao","hidden":false},{"_id":"6aa0d4ccd0174964227bed2a","name":"Shuguang Cui","hidden":false},{"_id":"6aa0d4ccd0174964227bed2b","name":"Xiaoguang Han","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64749a0d5aba8edfb2eeaba7/dEFIMAOcVJ5FvJmExv4aS.mp4","https://cdn-uploads.huggingface.co/production/uploads/64749a0d5aba8edfb2eeaba7/6NfRkF4CgXZ0gCqEG4UXN.png"],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements","submittedOnDailyBy":{"_id":"64749a0d5aba8edfb2eeaba7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png","isPro":false,"fullname":"Mutian Xu","user":"Minoday","type":"user","name":"Minoday"},"summary":"Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.","upvotes":25,"discussionId":"6aa0d4cdd0174964227bed2c","projectPage":"https://cosmoh2g.github.io/","githubRepo":"https://github.com/GAP-LAB-CUHK-SZ/CosmoH2G","githubRepoAddedBy":"user","ai_summary":"A two-stage framework predicts sparse gripper keyframes and continuous actions to transfer complex hand demonstrations to robotic grippers while reducing drift via kinematic optimization.","ai_keywords":["hand-pose motions","hand-gripper paired demonstrations","sparse gripper keyframes","two-stage framework","grasping heuristic","kinematic consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64749a0d5aba8edfb2eeaba7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png","isPro":false,"fullname":"Mutian Xu","user":"Minoday","type":"user"},{"_id":"6865232943473cae5522b01d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/umCrG0k86JRCjU3tE4_6c.png","isPro":false,"fullname":"Liujia Ma","user":"Malacca345","type":"user"},{"_id":"65d7ed96c790f52e9373b00c","avatarUrl":"/avatars/b4ac04c20516198190da21688e8a16a0.svg","isPro":false,"fullname":"sunwanhu","user":"cvhadessun","type":"user"},{"_id":"65884a13d2ea3f32942f9303","avatarUrl":"/avatars/532cbbf5cf287afb80d465c08b844339.svg","isPro":false,"fullname":"YanxiDU","user":"Yokaze33","type":"user"},{"_id":"6448a44419538c015b2c0291","avatarUrl":"/avatars/3851ae76a2724245369c1a3660fa7727.svg","isPro":false,"fullname":"chesterzyan","user":"ChesterZYan","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"66027a9c21032a86a5d58220","avatarUrl":"/avatars/06c8db26e282cb1927c6408ad5397828.svg","isPro":false,"fullname":"anran lin","user":"picksh","type":"user"},{"_id":"66dfadac9f6980cf98570efa","avatarUrl":"/avatars/7d1035ed163e2c0a581771b1c723a2d4.svg","isPro":false,"fullname":"Xianzu Wu","user":"Xianzu","type":"user"},{"_id":"6a6aa2d8df1718450cef87ed","avatarUrl":"/avatars/37b86fdbfaa48d5e154f84d8a232eb15.svg","isPro":false,"fullname":"Joseph Williams","user":"LunarNico","type":"user"},{"_id":"6a6c858f217ea989952d4903","avatarUrl":"/avatars/3b91dcc188e129952493f3cad707e88a.svg","isPro":false,"fullname":"Michael Clark","user":"Quiet-Kai","type":"user"},{"_id":"6a6c9c7aa1aab08eb34a1057","avatarUrl":"/avatars/edc8e84ec7eb19de792c37266fe48176.svg","isPro":false,"fullname":"Patricia Wilson","user":"patricia-wilson","type":"user"},{"_id":"6a701bdc80a9e01789431dec","avatarUrl":"/avatars/923889d4773ad9e5b8e3af29f17e3897.svg","isPro":false,"fullname":"Mark Harris","user":"harborMark","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07498.md","query":{}}">
Papers
arxiv:2609.07498

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

Published on Sep 7
· Submitted by
Mutian Xu
on Sep 9
Authors:
,

Abstract

A two-stage framework predicts sparse gripper keyframes and continuous actions to transfer complex hand demonstrations to robotic grippers while reducing drift via kinematic optimization.

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.

Community

Paper submitter about 10 hours ago

SIGGRAPH Aisa 2026; Project page: https://cosmoh2g.github.io/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.07498
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.07498 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.07498 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.07498 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers