r/LocalLLaMA · · 1 min read

Xiaomi-Robotics-1: New robotics model released

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipulation in unseen environments and efficient adaptation to new tasks.

XR-1 follows a two-stage training paradigm inspired by large language models — pre-training for breadth, followed by post-training for alignment. It showcases that pre-training scaling behavior reliably transfers through post-training to real-world robot performance, with no signs of saturation.

XR-1 couples a pre-trained VLM (Qwen3-VL) with a Diffusion-Transformer (DiT) via a Mixture-of-Transformers (MoT) — the DiT matches the VLM in layer count but uses a smaller hidden size for faster inference.

HugginFace: https://huggingface.co/collections/XiaomiRobotics/xiaomi-robotics-1

GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-1

Paper: https://arxiv.org/abs/2607.15330

submitted by /u/121507090301
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA