A scalable Ego2Robot pipeline</p>\n","updatedAt":"2026-08-06T03:19:39.450Z","author":{"_id":"63b45b4aa50cfcefda99d125","avatarUrl":"/avatars/664dbee742e44a403ca1e0cae0328a02.svg","fullname":"Ye Wang","name":"wwwyyy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5038469433784485},"editors":["wwwyyy"],"editorAvatarUrls":["/avatars/664dbee742e44a403ca1e0cae0328a02.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.02580","authors":[{"_id":"6a72b5a01a375f948521c479","name":"Ye Wang","hidden":false},{"_id":"6a72b5a01a375f948521c47a","name":"Pei Lin","hidden":false},{"_id":"6a72b5a01a375f948521c47b","name":"Xiong-Hui Chen","hidden":false},{"_id":"6a72b5a01a375f948521c47c","name":"Haoqi Yuan","hidden":false},{"_id":"6a72b5a01a375f948521c47d","name":"Zhixuan Liang","hidden":false},{"_id":"6a72b5a01a375f948521c47e","name":"Yiyang Huang","hidden":false},{"_id":"6a72b5a01a375f948521c47f","name":"Anzhe Chen","hidden":false},{"_id":"6a72b5a01a375f948521c480","name":"Zixing Lei","hidden":false},{"_id":"6a72b5a01a375f948521c481","name":"Jie Zhang","hidden":false},{"_id":"6a72b5a01a375f948521c482","name":"Tao Zhang","hidden":false},{"_id":"6a72b5a01a375f948521c483","name":"Haoyang Li","hidden":false},{"_id":"6a72b5a01a375f948521c484","name":"Tong Zhang","hidden":false},{"_id":"6a72b5a01a375f948521c485","name":"Chenxi Xiao","hidden":false},{"_id":"6a72b5a01a375f948521c486","name":"Ziyuan Jiao","hidden":false},{"_id":"6a72b5a01a375f948521c487","name":"Qin Jin","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/63b45b4aa50cfcefda99d125/7vIenpgwzyGTDkMwZ0OuD.mp4"],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data","submittedOnDailyBy":{"_id":"63b45b4aa50cfcefda99d125","avatarUrl":"/avatars/664dbee742e44a403ca1e0cae0328a02.svg","isPro":false,"fullname":"Ye Wang","user":"wwwyyy","type":"user","name":"wwwyyy"},"summary":"Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present Ego2Robot, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/","upvotes":13,"discussionId":"6a72b5a01a375f948521c488","projectPage":"https://www-ye.github.io/ego2robot_blog/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63b45b4aa50cfcefda99d125","avatarUrl":"/avatars/664dbee742e44a403ca1e0cae0328a02.svg","isPro":false,"fullname":"Ye Wang","user":"wwwyyy","type":"user"},{"_id":"65e2e60dc108de6a47f02a9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65e2e60dc108de6a47f02a9c/4_e7XHj8Ja3c-7gkhXwwU.jpeg","isPro":false,"fullname":"WEISHUAI ZENG","user":"WeishuaiZeng","type":"user"},{"_id":"64254247daa3502ee355f208","avatarUrl":"/avatars/0721c3ef77e7302aceb8295e342489fc.svg","isPro":false,"fullname":"tianyi yan","user":"yanty123","type":"user"},{"_id":"686f75cf64498736bc1ed44a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/X6OYS6rDm0va52pLxzCvw.png","isPro":false,"fullname":"Tao Zhang","user":"HH375356","type":"user"},{"_id":"6a740fcbe4b94ef9d7c8d18f","avatarUrl":"/avatars/551ebf310cbd8deb7584f5ef7ca4a5fa.svg","isPro":false,"fullname":"Subhadra Vaish","user":"subhadra-vaish1","type":"user"},{"_id":"614d69a8b06484609bf2292f","avatarUrl":"/avatars/917bf81b7854fab4b7b8ed429392cde2.svg","isPro":false,"fullname":"lihaoyang","user":"seeklhy","type":"user"},{"_id":"69d6ff6d2174044cf08b28c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/1mDaMIqg5xRXnCaZ-t3K_.png","isPro":false,"fullname":"Yeth","user":"yethdev","type":"user"},{"_id":"67f77ce6262969d32f3c5bbb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/yOLAtLna3aqAjCbkyY3zM.png","isPro":false,"fullname":"Pei Lin","user":"PerrinLin","type":"user"},{"_id":"6a6a82d4ce8b4ee20326608f","avatarUrl":"/avatars/a01c7ba7e8e706939874b22c4804665d.svg","isPro":false,"fullname":"George Martin","user":"georgemartin","type":"user"},{"_id":"6a6aa236144c90a0d41741b8","avatarUrl":"/avatars/3f1f7a14fa7466daf8161e65b0ecbeeb.svg","isPro":false,"fullname":"Linda Perez","user":"PixelEvan","type":"user"},{"_id":"6a6dc9adb34a88441964b81d","avatarUrl":"/avatars/1b426a084d1aa8f762b3ab7ab6844afc.svg","isPro":false,"fullname":"Joshua Rodriguez","user":"daniel-1133670","type":"user"},{"_id":"6a6dec339caa2ff7db7a720b","avatarUrl":"/avatars/8d103b3fb8475daf9ce93af90cdda06b.svg","isPro":false,"fullname":"Paul Gonzalez","user":"david-7490294","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.02580.md","query":{}}">
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Abstract
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present Ego2Robot, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/
Community
A scalable Ego2Robot pipeline
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.02580 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.02580 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.02580 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.