SIGGRAPH Aisa 2026; Project page: <a href=\"https://cosmoh2g.github.io/\" rel=\"nofollow\">https://cosmoh2g.github.io/</a></p>\n","updatedAt":"2026-09-09T04:17:09.768Z","author":{"_id":"64749a0d5aba8edfb2eeaba7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png","fullname":"Mutian Xu","name":"Minoday","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4018838107585907},"editors":["Minoday"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07498","authors":[{"_id":"6aa0d4ccd0174964227bed26","name":"Hongxiang Zhao","hidden":false},{"_id":"6aa0d4ccd0174964227bed27","name":"Mutian Xu","hidden":false},{"_id":"6aa0d4ccd0174964227bed28","name":"Zeyu Jin","hidden":false},{"_id":"6aa0d4ccd0174964227bed29","name":"Yiming Hao","hidden":false},{"_id":"6aa0d4ccd0174964227bed2a","name":"Shuguang Cui","hidden":false},{"_id":"6aa0d4ccd0174964227bed2b","name":"Xiaoguang Han","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64749a0d5aba8edfb2eeaba7/dEFIMAOcVJ5FvJmExv4aS.mp4","https://cdn-uploads.huggingface.co/production/uploads/64749a0d5aba8edfb2eeaba7/6NfRkF4CgXZ0gCqEG4UXN.png"],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements","submittedOnDailyBy":{"_id":"64749a0d5aba8edfb2eeaba7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png","isPro":false,"fullname":"Mutian Xu","user":"Minoday","type":"user","name":"Minoday"},"summary":"Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.","upvotes":25,"discussionId":"6aa0d4cdd0174964227bed2c","projectPage":"https://cosmoh2g.github.io/","githubRepo":"https://github.com/GAP-LAB-CUHK-SZ/CosmoH2G","githubRepoAddedBy":"user","ai_summary":"A two-stage framework predicts sparse gripper keyframes and continuous actions to transfer complex hand demonstrations to robotic grippers while reducing drift via kinematic optimization.","ai_keywords":["hand-pose motions","hand-gripper paired demonstrations","sparse gripper keyframes","two-stage framework","grasping heuristic","kinematic consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64749a0d5aba8edfb2eeaba7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64749a0d5aba8edfb2eeaba7/Tiy4DEdp3KQYh7Ij8Vmkn.png","isPro":false,"fullname":"Mutian Xu","user":"Minoday","type":"user"},{"_id":"6865232943473cae5522b01d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/umCrG0k86JRCjU3tE4_6c.png","isPro":false,"fullname":"Liujia Ma","user":"Malacca345","type":"user"},{"_id":"65d7ed96c790f52e9373b00c","avatarUrl":"/avatars/b4ac04c20516198190da21688e8a16a0.svg","isPro":false,"fullname":"sunwanhu","user":"cvhadessun","type":"user"},{"_id":"65884a13d2ea3f32942f9303","avatarUrl":"/avatars/532cbbf5cf287afb80d465c08b844339.svg","isPro":false,"fullname":"YanxiDU","user":"Yokaze33","type":"user"},{"_id":"6448a44419538c015b2c0291","avatarUrl":"/avatars/3851ae76a2724245369c1a3660fa7727.svg","isPro":false,"fullname":"chesterzyan","user":"ChesterZYan","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"66027a9c21032a86a5d58220","avatarUrl":"/avatars/06c8db26e282cb1927c6408ad5397828.svg","isPro":false,"fullname":"anran lin","user":"picksh","type":"user"},{"_id":"66dfadac9f6980cf98570efa","avatarUrl":"/avatars/7d1035ed163e2c0a581771b1c723a2d4.svg","isPro":false,"fullname":"Xianzu Wu","user":"Xianzu","type":"user"},{"_id":"6a6aa2d8df1718450cef87ed","avatarUrl":"/avatars/37b86fdbfaa48d5e154f84d8a232eb15.svg","isPro":false,"fullname":"Joseph Williams","user":"LunarNico","type":"user"},{"_id":"6a6c858f217ea989952d4903","avatarUrl":"/avatars/3b91dcc188e129952493f3cad707e88a.svg","isPro":false,"fullname":"Michael Clark","user":"Quiet-Kai","type":"user"},{"_id":"6a6c9c7aa1aab08eb34a1057","avatarUrl":"/avatars/edc8e84ec7eb19de792c37266fe48176.svg","isPro":false,"fullname":"Patricia Wilson","user":"patricia-wilson","type":"user"},{"_id":"6a701bdc80a9e01789431dec","avatarUrl":"/avatars/923889d4773ad9e5b8e3af29f17e3897.svg","isPro":false,"fullname":"Mark Harris","user":"harborMark","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07498.md","query":{}}">
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
Abstract
A two-stage framework predicts sparse gripper keyframes and continuous actions to transfer complex hand demonstrations to robotic grippers while reducing drift via kinematic optimization.
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.07498 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.07498 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.07498 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.