Hi everyone! We’re excited to share HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone.</p>\n<p>The central question is simple: can we eliminate target-task robot teleoperation from post-training, rather than merely reduce it?</p>\n<p>HiFi-UMI is a portable, robot-free data-production system co-designed for action fidelity. It achieves 3 mm workspace-local end-effector accuracy, <40 μs cross-sensor synchronization, and ultra-wide six-view sensing, together with automated trajectory reconstruction, simulation replay, and quality validation.</p>\n<p>Our main findings:</p>\n<p>Across three VLA and WAM backbones—StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA—post-training using only HiFi-UMI demonstrations matches in-domain robot teleoperation, with success-rate differences of −2.5, +3.1, and −0.6 percentage points.<br>The strongest policy reaches 85% success on precision insertion, despite no HiFi-UMI demonstration being collected in the evaluation scene.<br>Pre-training on 4,000 hours reduces action error on ten unseen tasks by 41% and improves real-robot success by 18.1 percentage points.<br>We release HiFi-UMI-2K: 2,000 hours and 482K+ replayable demonstrations across 110+ scenes under CC BY 4.0.</p>\n<p>Our key takeaway: robot-free data can support deployment—not only pre-training—when it is sufficiently high-fidelity and action-aligned.</p>\n<p>📄 Paper:<a href=\"https://arxiv.org/abs/2607.25895\" rel=\"nofollow\">https://arxiv.org/abs/2607.25895</a><br>🌐 主页:<a href=\"https://cloud.simpleai.tech/simple-world-lab/hifi-umi/\" rel=\"nofollow\">https://cloud.simpleai.tech/simple-world-lab/hifi-umi/</a><br>🤗 数据集:<a href=\"https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K\">https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K</a><br>🔥 Daily Paper:<a href=\"https://huggingface.co/papers/2607.25895\">https://huggingface.co/papers/2607.25895</a></p>\n","updatedAt":"2026-07-29T06:26:01.639Z","author":{"_id":"64060b49a577649430bf6974","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64060b49a577649430bf6974/0YhJeunF5brCysnnEpLKG.jpeg","fullname":"Jiawei Wang","name":"Jarvis1111","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8548188805580139},"editors":["Jarvis1111"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64060b49a577649430bf6974/0YhJeunF5brCysnnEpLKG.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25895","authors":[{"_id":"6a6962c99d3a1231d492b7d4","name":"Simple AI","hidden":false},{"_id":"6a6962c99d3a1231d492b7d6","name":"Yuteng Wei","hidden":false},{"_id":"6a6962c99d3a1231d492b7d7","name":"Jinming Ma","hidden":false},{"_id":"6a6962c99d3a1231d492b7d8","user":{"_id":"64060b49a577649430bf6974","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64060b49a577649430bf6974/0YhJeunF5brCysnnEpLKG.jpeg","isPro":false,"fullname":"Jiawei Wang","user":"Jarvis1111","type":"user","name":"Jarvis1111"},"name":"Jiawei Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-29T08:45:04.449Z","hidden":false},{"_id":"6a6962c99d3a1231d492b7d9","name":"Weitao Zhou","hidden":false},{"_id":"6a6962c99d3a1231d492b7da","name":"Yushen Zuo","hidden":false},{"_id":"6a6962c99d3a1231d492b7db","name":"Ke Rui","hidden":false},{"_id":"6a6962c99d3a1231d492b7dc","name":"Minglei Li","hidden":false},{"_id":"6a6962c99d3a1231d492b7dd","name":"Jinhao Zhang","hidden":false},{"_id":"6a6962c99d3a1231d492b7de","name":"Zhikang Pan","hidden":false},{"_id":"6a6962c99d3a1231d492b7df","name":"Xiang Wang","hidden":false},{"_id":"6a6962c99d3a1231d492b7e0","name":"Haoran Jia","hidden":false},{"_id":"6a6962c99d3a1231d492b7e1","name":"Huan Du","hidden":false},{"_id":"6a6962c99d3a1231d492b7e2","name":"Zicheng Zeng","hidden":false},{"_id":"6a6962c99d3a1231d492b7e3","name":"Jun Ma","hidden":false},{"_id":"6a6962c99d3a1231d492b7e4","name":"Guiyu Qin","hidden":false},{"_id":"6a6962c99d3a1231d492b7e5","name":"Di Zhang","hidden":false},{"_id":"6a6962c99d3a1231d492b7e6","name":"Xiaofei Li","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64060b49a577649430bf6974/seT5Eiwx9StEGYtDsvzqS.webm"],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-29T00:00:00.000Z","title":"HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone","submittedOnDailyBy":{"_id":"64060b49a577649430bf6974","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64060b49a577649430bf6974/0YhJeunF5brCysnnEpLKG.jpeg","isPro":false,"fullname":"Jiawei Wang","user":"Jarvis1111","type":"user","name":"Jarvis1111"},"summary":"Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot \"anchor\" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.","upvotes":72,"discussionId":"6a6962c99d3a1231d492b7e7","projectPage":"https://cloud.simpleai.tech/simple-world-lab/hifi-umi/","organization":{"_id":"6a58b6bcf45a179ffad5c858","name":"simple-world-lab","fullname":"Simple World Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a58b4505c2ab1e637affd07/bKcJxgIPjCUkqyOOnX91d.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"672deba796a9d93d16f48e2b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/672deba796a9d93d16f48e2b/9miZZUnZ13GGk8OerCVUP.png","isPro":false,"fullname":"Minglei","user":"tkingcer","type":"user"},{"_id":"64060b49a577649430bf6974","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64060b49a577649430bf6974/0YhJeunF5brCysnnEpLKG.jpeg","isPro":false,"fullname":"Jiawei Wang","user":"Jarvis1111","type":"user"},{"_id":"643e9efa2263cdc630f88f5c","avatarUrl":"/avatars/96cea51f17e7d41ffb6a4b438e05f5cb.svg","isPro":false,"fullname":"Yushen Zuo","user":"YSZuo","type":"user"},{"_id":"68cba1393ca04e8e8effb140","avatarUrl":"/avatars/6e0c411390fa60b07ddaf2dbda971137.svg","isPro":false,"fullname":"wx","user":"Wangxxxx1111","type":"user"},{"_id":"64104b467a15af878ae6695d","avatarUrl":"/avatars/407983918c12411e5ed636bf7435522b.svg","isPro":false,"fullname":"Fangyu Lei","user":"FangyuLei","type":"user"},{"_id":"667d15bbdb6a8e980f05a0f7","avatarUrl":"/avatars/8f171fe6f473249fb245dd4f0f4af0c0.svg","isPro":false,"fullname":"yuan ma","user":"may22210297","type":"user"},{"_id":"69dcd5329277c5281e9f5889","avatarUrl":"/avatars/1706f887e45b791480f4fc8f86bf53b7.svg","isPro":false,"fullname":"Ma","user":"JinmingM","type":"user"},{"_id":"616648c84c0937d31946f21b","avatarUrl":"/avatars/7ca27de5c5116c91ff1db61ba6277ed5.svg","isPro":false,"fullname":"Ziyang","user":"hzy","type":"user"},{"_id":"67b6e2769fd5603fe15b84b9","avatarUrl":"/avatars/e6b42c49954a840518bc3e3aa2617912.svg","isPro":false,"fullname":"fengyichun","user":"fengyichun","type":"user"},{"_id":"662d166ba314b134a2e6dd89","avatarUrl":"/avatars/d1bacaa0caa5de609dd69a3328682859.svg","isPro":false,"fullname":"KaiHu","user":"KaiHuUTSC","type":"user"},{"_id":"685b5b5cd7a4335ed1fcaa69","avatarUrl":"/avatars/944aeb6a061a2feed61425fd4ab9e573.svg","isPro":false,"fullname":"ifzzh","user":"ifzzh","type":"user"},{"_id":"65976e856a30bf8860df97c7","avatarUrl":"/avatars/447377dc26a2f495551388baffc6c5c8.svg","isPro":false,"fullname":"Haibao Yu","user":"Haibao123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"6a58b6bcf45a179ffad5c858","name":"simple-world-lab","fullname":"Simple World Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a58b4505c2ab1e637affd07/bKcJxgIPjCUkqyOOnX91d.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25895.md","query":{}}">
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Abstract
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.
Community
Hi everyone! We’re excited to share HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone.
The central question is simple: can we eliminate target-task robot teleoperation from post-training, rather than merely reduce it?
HiFi-UMI is a portable, robot-free data-production system co-designed for action fidelity. It achieves 3 mm workspace-local end-effector accuracy, <40 μs cross-sensor synchronization, and ultra-wide six-view sensing, together with automated trajectory reconstruction, simulation replay, and quality validation.
Our main findings:
Across three VLA and WAM backbones—StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA—post-training using only HiFi-UMI demonstrations matches in-domain robot teleoperation, with success-rate differences of −2.5, +3.1, and −0.6 percentage points.
The strongest policy reaches 85% success on precision insertion, despite no HiFi-UMI demonstration being collected in the evaluation scene.
Pre-training on 4,000 hours reduces action error on ten unseen tasks by 41% and improves real-robot success by 18.1 percentage points.
We release HiFi-UMI-2K: 2,000 hours and 482K+ replayable demonstrations across 110+ scenes under CC BY 4.0.
Our key takeaway: robot-free data can support deployment—not only pre-training—when it is sufficiently high-fidelity and action-aligned.
📄 Paper:https://arxiv.org/abs/2607.25895
🌐 主页:https://cloud.simpleai.tech/simple-world-lab/hifi-umi/
🤗 数据集:https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K
🔥 Daily Paper:https://huggingface.co/papers/2607.25895
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.25895 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.25895 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.