RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.
RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
Abstract
Digital teleoperation replaces physical robot interaction with generative world models to create diverse training data for robotics, enabling efficient zero-shot Sim2Real transfer and improved real-world performance.
Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware and workspaces. We introduce digital teleoperation, a paradigm that decouples data collection from physical constraints by replacing the real robot with a generative world model. In this framework, an operator's hand-pose stream drives a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image. The recorded pose stream serves as an embodiment-agnostic action label transferable to any target robot via standard retargeting, yielding complete state-action trajectories for imitation learning independent of physical hardware. We instantiate this paradigm in RynnWorld-Teleop, a system that integrates depth-aware skeletal conditioning, progressive human-to-robot training on a video Diffusion Transformer, and streaming autoregressive distillation. This pipeline compresses the generative process into a single-pass inference, enabling 40+ FPS, real-time interactive generation on a single H100 GPU. Policies trained exclusively on RynnWorld-Teleop-generated data achieve effective zero-shot Sim2Real transfer across dexterous and diverse bimanual tasks. Moreover, augmenting real-world datasets with our digitally teleoperated data consistently improves success rates, demonstrating that RynnWorld-Teleop serves as a high-fidelity, scalable data engine for the next generation of robotic agents.
Get this paper in your agent:
hf papers read 2607.06558 curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper
No model linking this paper
Datasets citing this paper
Spaces citing this paper
No Space linking this paper
Collections including this paper
More from Hugging Face Daily Papers
-
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Aug 13
-
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Aug 13
-
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Aug 13
-
AVA-Encoder: Towards Agent-Native Video Representation Learning
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.