Hugging Face Daily Papers · · 4 min read

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Code:<a href=\"https://github.com/open-gigaai/giga-world-1\" rel=\"nofollow\">https://github.com/open-gigaai/giga-world-1</a><br>Model:<a href=\"https://huggingface.co/open-gigaai/Giga-World-1\">https://huggingface.co/open-gigaai/Giga-World-1</a><br>Data:<a href=\"https://huggingface.co/datasets/open-gigaai/CVPR-2026-WorldModel-Track-Dataset\">https://huggingface.co/datasets/open-gigaai/CVPR-2026-WorldModel-Track-Dataset</a><br>Benchmark:<a href=\"https://huggingface.co/spaces/open-gigaai/CVPR-2026-WorldModel-Track-LeaderBoard\">https://huggingface.co/spaces/open-gigaai/CVPR-2026-WorldModel-Track-LeaderBoard</a></p>\n","updatedAt":"2026-07-07T02:44:55.893Z","author":{"_id":"6426616ea5ec4a5cbc535634","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6426616ea5ec4a5cbc535634/5IfSFYd9QOxz8K9QmBCst.png","fullname":"JeffWang","name":"Jeff-Wang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.2569692134857178},"editors":["Jeff-Wang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6426616ea5ec4a5cbc535634/5IfSFYd9QOxz8K9QmBCst.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.02642","authors":[{"_id":"6a4c64d625849b193a834046","name":"GigaWorld Team","hidden":false},{"_id":"6a4c64d625849b193a834047","user":{"_id":"64f020a42a7a94b68c0df99c","avatarUrl":"/avatars/39315929c3675039f9a7662e7b60ad04.svg","isPro":false,"fullname":"maangyuan","user":"maay21","type":"user","name":"maay21"},"name":"Angyuan Ma","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:12:09.489Z","hidden":false},{"_id":"6a4c64d625849b193a834048","name":"Boyuan Wang","hidden":false},{"_id":"6a4c64d625849b193a834049","name":"Bohan Li","hidden":false},{"_id":"6a4c64d625849b193a83404a","name":"Chaojun Ni","hidden":false},{"_id":"6a4c64d625849b193a83404b","name":"Guo Li","hidden":false},{"_id":"6a4c64d625849b193a83404c","name":"Guan Huang","hidden":false},{"_id":"6a4c64d625849b193a83404d","name":"Guosheng Zhao","hidden":false},{"_id":"6a4c64d625849b193a83404e","name":"Hao Li","hidden":false},{"_id":"6a4c64d625849b193a83404f","name":"Hengtao Li","hidden":false},{"_id":"6a4c64d625849b193a834050","name":"Jingyu Liu","hidden":false},{"_id":"6a4c64d625849b193a834051","name":"Jiwen Lu","hidden":false},{"_id":"6a4c64d625849b193a834052","name":"Qiuping Deng","hidden":false},{"_id":"6a4c64d625849b193a834053","name":"Tingdong Yu","hidden":false},{"_id":"6a4c64d625849b193a834054","name":"Xuancheng Xu","hidden":false},{"_id":"6a4c64d625849b193a834055","name":"Xinyu Zhou","hidden":false},{"_id":"6a4c64d625849b193a834056","name":"Xiuwei Xu","hidden":false},{"_id":"6a4c64d625849b193a834057","name":"Xinze Chen","hidden":false},{"_id":"6a4c64d625849b193a834058","user":{"_id":"6426616ea5ec4a5cbc535634","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6426616ea5ec4a5cbc535634/5IfSFYd9QOxz8K9QmBCst.png","isPro":false,"fullname":"JeffWang","user":"Jeff-Wang","type":"user","name":"Jeff-Wang"},"name":"Xiaofeng Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:12:13.375Z","hidden":false},{"_id":"6a4c64d625849b193a834059","name":"Xiaoyu Tian","hidden":false},{"_id":"6a4c64d625849b193a83405a","user":{"_id":"644012cf3e0374802e174f7c","avatarUrl":"/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg","isPro":false,"fullname":"Yang Wang","user":"supermodelteam","type":"user","name":"supermodelteam"},"name":"Yang Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:12:11.475Z","hidden":false},{"_id":"6a4c64d625849b193a83405b","name":"Yifan Chang","hidden":false},{"_id":"6a4c64d625849b193a83405c","name":"Yukun Zhou","hidden":false},{"_id":"6a4c64d625849b193a83405d","name":"Yun Ye","hidden":false},{"_id":"6a4c64d625849b193a83405e","name":"Zhenyu Wu","hidden":false},{"_id":"6a4c64d625849b193a83405f","name":"Zhanqian Wu","hidden":false},{"_id":"6a4c64d625849b193a834060","name":"Zheng Zhu","hidden":false}],"publishedAt":"2026-07-02T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation","submittedOnDailyBy":{"_id":"6426616ea5ec4a5cbc535634","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6426616ea5ec4a5cbc535634/5IfSFYd9QOxz8K9QmBCst.png","isPro":false,"fullname":"JeffWang","user":"Jeff-Wang","type":"user","name":"Jeff-Wang"},"summary":"Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in GigaWorld-1, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.","upvotes":31,"discussionId":"6a4c64d625849b193a834061","projectPage":"https://open-gigaai.github.io/giga-world-1/","githubRepo":"https://github.com/open-gigaai/giga-world-1","githubRepoAddedBy":"user","ai_summary":"World models for robotic policy evaluation are systematically studied through a new benchmark, revealing that long-horizon rollout consistency and robot-specific controllability are more important than short-term visual realism for reliable policy assessment.","ai_keywords":["world models","robotic policies","policy evaluation","real-robot teleoperation","video world models","action representation schemes","rollout horizons","simulation","real-world robot behavior","GigaWorld-1"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":145,"organization":{"_id":"68d6587936e2de9610d9f5f0","name":"open-gigaai","fullname":"GigaAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d6394328e169473e90e4a6/zUK7FKr_8XqrN0aFUgsD-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6426616ea5ec4a5cbc535634","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6426616ea5ec4a5cbc535634/5IfSFYd9QOxz8K9QmBCst.png","isPro":false,"fullname":"JeffWang","user":"Jeff-Wang","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"64739aa271f07ae738d2d088","avatarUrl":"/avatars/a91d5b18864f8c42abd542eaf1f38b17.svg","isPro":false,"fullname":"JiangnanShao","user":"Andy31415","type":"user"},{"_id":"654c7f32386fc5525c1ebc12","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/654c7f32386fc5525c1ebc12/RIUS_nkbt9n5iu46lLQkb.jpeg","isPro":false,"fullname":"zhao guosheng","user":"f1yfish","type":"user"},{"_id":"6466e66314e059dde8bbf35c","avatarUrl":"/avatars/f2a27bb4e04749ec51f3a9fc36d87fb0.svg","isPro":false,"fullname":"I0u0I","user":"I0u0I","type":"user"},{"_id":"66b03ee1e48856bb7197ed62","avatarUrl":"/avatars/fdfed15190951d4b88bf37e419623be7.svg","isPro":false,"fullname":"Yang","user":"Jian0227","type":"user"},{"_id":"68917b7bb4c8617704deaae6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Il651f9440ic7oKFeGm78.png","isPro":false,"fullname":"XuanchengXu","user":"XuanchengXu","type":"user"},{"_id":"6909ce9b391d11cef04e17ee","avatarUrl":"/avatars/d3d4274c7610e2e66b0319075b34c7f4.svg","isPro":false,"fullname":"Daniel Harris","user":"codecodecode123","type":"user"},{"_id":"673de6768fc2cb3dee787d9b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mqFdHZoaHAzkmiC-62rUk.png","isPro":false,"fullname":"JoeDeng","user":"cbtogu","type":"user"},{"_id":"644012cf3e0374802e174f7c","avatarUrl":"/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg","isPro":false,"fullname":"Yang Wang","user":"supermodelteam","type":"user"},{"_id":"66a1efaffbabdd7b0d2f1889","avatarUrl":"/avatars/c4ee6bb97a63f06515e30a088f9378e7.svg","isPro":false,"fullname":"zz","user":"zzhugging","type":"user"},{"_id":"686fc6d16ea5d5fb0a4a5b2d","avatarUrl":"/avatars/d7a793b03951c2a7682167e17deaf8db.svg","isPro":false,"fullname":"chong","user":"Phil1220","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68d6587936e2de9610d9f5f0","name":"open-gigaai","fullname":"GigaAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d6394328e169473e90e4a6/zUK7FKr_8XqrN0aFUgsD-.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.02642.md","query":{}}">
Papers
arxiv:2607.02642

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

Published on Jul 2
· Submitted by
JeffWang
on Jul 7
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

World models for robotic policy evaluation are systematically studied through a new benchmark, revealing that long-horizon rollout consistency and robot-specific controllability are more important than short-term visual realism for reliable policy assessment.

Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in GigaWorld-1, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.02642
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.02642 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.02642 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.02642 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers