Hugging Face Daily Papers · · 4 min read

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/fDDptSe2sDrZC35MwTofV.jpeg\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/fDDptSe2sDrZC35MwTofV.jpeg\" alt=\"teaser\"></a></p>\n<p>Our <strong>full-stack system</strong> integrates <a href=\"https://github.com/egosteer/egosmith\" rel=\"nofollow\">EgoSmith</a>, <a href=\"https://github.com/egosteer/robot-stack\" rel=\"nofollow\">Robot Stack</a>, and <a href=\"https://github.com/egosteer/egosteer\" rel=\"nofollow\">EgoSteer</a> to learn from large-scale egocentric human videos and facilitate data-efficient real-robot post-training, enabling steerable dexterous manipulation across over 40 tasks alongside few-shot adaptation to complex, long-horizon tasks.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/q-xwpOnChCIquAlfGqjgZ.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/q-xwpOnChCIquAlfGqjgZ.png\" alt=\"egosmith\"></a></p>\n<p><strong>EgoSmith</strong> curates noisy egocentric videos into high-quality dexterous VLA data. It provides scalable human-hand interaction priors for language-guided robot pre-training.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/muZ1ZZitnps4lXoLfqHro.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/muZ1ZZitnps4lXoLfqHro.png\" alt=\"robot-stack\"></a></p>\n<p><strong>Unified Robot Stack</strong> unifies teleoperation, policy inference, and human-in-the-loop correction. It enables efficient robot data collection, deployment, and DAgger-style refinement within one shared control stack.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/CPKCqgnpEwscQwFNYf654.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/655d9f43b5da99edaf3f2f81/CPKCqgnpEwscQwFNYf654.png\" alt=\"egosteer\"></a></p>\n<p><strong>EgoSteer</strong> is a world-model-enhanced VLA for steerable dexterous manipulation. It predicts latency-aware wrist and fingertip actions while learning action-induced visual features.</p>\n","updatedAt":"2026-07-14T11:58:10.851Z","author":{"_id":"655d9f43b5da99edaf3f2f81","avatarUrl":"/avatars/c7225b3ed54d099a4fd87682427fb5bf.svg","fullname":"Yifan Zhong","name":"Yifan-Zhong","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"editors":["Yifan-Zhong"],"editorAvatarUrls":["/avatars/c7225b3ed54d099a4fd87682427fb5bf.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.09701","authors":[{"_id":"6a561d6da9d74d6e65bbdc16","name":"Yifan Zhong","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc17","name":"Zhang Chen","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc18","name":"Tianrui Guan","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc19","name":"Fanlian Zeng","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc1a","name":"Yuyao Ye","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc1b","name":"Tianjia He","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc1c","name":"Ka Nam Lui","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc1d","name":"Jiayi Li","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc1e","name":"Tingrui Zhang","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc1f","name":"Ruilin Yan","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc20","name":"Xinhao Ji","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc21","name":"Guangyu Zhao","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc22","name":"Wenjie Lou","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc23","name":"Jiayuan Zhang","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc24","name":"Yuanpei Chen","hidden":false},{"_id":"6a561d6da9d74d6e65bbdc25","name":"Yaodong Yang","hidden":false}],"publishedAt":"2026-06-21T16:16:02.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos","submittedOnDailyBy":{"_id":"655d9f43b5da99edaf3f2f81","avatarUrl":"/avatars/c7225b3ed54d099a4fd87682427fb5bf.svg","isPro":false,"fullname":"Yifan Zhong","user":"Yifan-Zhong","type":"user","name":"Yifan-Zhong"},"summary":"Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.","upvotes":8,"discussionId":"6a561d6da9d74d6e65bbdc26","projectPage":"https://egosteer.github.io/#/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"655d9f43b5da99edaf3f2f81","avatarUrl":"/avatars/c7225b3ed54d099a4fd87682427fb5bf.svg","isPro":false,"fullname":"Yifan Zhong","user":"Yifan-Zhong","type":"user"},{"_id":"67d677a4212701212b0c1dbb","avatarUrl":"/avatars/fa17aef57ae84c1c3a60c998e7dcfd0e.svg","isPro":false,"fullname":"Zhang Chen","user":"Achtoria","type":"user"},{"_id":"68ce14793fdd2afb06220669","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68ce14793fdd2afb06220669/sZ1DG-wpqcfvK7YrOA0wH.png","isPro":false,"fullname":"Andy zhang","user":"HelloWorldZTR","type":"user"},{"_id":"687efdb14ec2d31746e68f0f","avatarUrl":"/avatars/9f496c556186bff0330e96b060a2090e.svg","isPro":false,"fullname":"Jiayuan Zhang","user":"jyzhang096","type":"user"},{"_id":"637a2005bdf7309aa6d46c79","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1668947933107-noauth.jpeg","isPro":false,"fullname":"Guangyu Zhao","user":"frenzyfreeze","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6a363b399864eeeb5223439e","avatarUrl":"/avatars/cf3b9ef8a655427d21aa493dd51f87f3.svg","isPro":false,"fullname":"EgoSteer","user":"EgoSteer","type":"user"},{"_id":"69884e24ffe2121b2762554e","avatarUrl":"/avatars/9a4564b2a642dbcb80d5c5a2f8840793.svg","isPro":false,"fullname":"Yssrings","user":"Yssrings","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Papers
arxiv:2607.09701

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

Published on Jun 21
· Submitted by
Yifan Zhong
on Jul 14
Authors:
,

Abstract

Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.

Community

Paper submitter about 8 hours ago

teaser

Our full-stack system integrates EgoSmith, Robot Stack, and EgoSteer to learn from large-scale egocentric human videos and facilitate data-efficient real-robot post-training, enabling steerable dexterous manipulation across over 40 tasks alongside few-shot adaptation to complex, long-horizon tasks.

egosmith

EgoSmith curates noisy egocentric videos into high-quality dexterous VLA data. It provides scalable human-hand interaction priors for language-guided robot pre-training.

robot-stack

Unified Robot Stack unifies teleoperation, policy inference, and human-in-the-loop correction. It enables efficient robot data collection, deployment, and DAgger-style refinement within one shared control stack.

egosteer

EgoSteer is a world-model-enhanced VLA for steerable dexterous manipulation. It predicts latency-aware wrist and fingertip actions while learning action-induced visual features.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.09701 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.09701 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.09701 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers