Hugging Face Daily Papers · · 4 min read

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation</p>\n","updatedAt":"2026-09-09T03:55:18.721Z","author":{"_id":"66c4816e96583c59b09fec30","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66c4816e96583c59b09fec30/RLurCsmcgOfyuWVkpsv1L.jpeg","fullname":"Ryan Chen","name":"ryancll118","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5183228254318237},"editors":["ryancll118"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66c4816e96583c59b09fec30/RLurCsmcgOfyuWVkpsv1L.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.05588","authors":[{"_id":"6aa0cff3d0174964227beca6","name":"AgiBot Research Team","hidden":false},{"_id":"6aa0cff3d0174964227beca7","user":{"_id":"64734b15bf9b32c6bbc8c95e","avatarUrl":"/avatars/ff3fe30c3855ad261f90dd8aae470036.svg","isPro":false,"fullname":"Liu Renhang","user":"liu-hanghang","type":"user","name":"liu-hanghang"},"name":"Renhang Liu","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.592Z","hidden":false},{"_id":"6aa0cff3d0174964227beca8","name":"Wenzhi Zhao","hidden":false},{"_id":"6aa0cff3d0174964227beca9","name":"Zhuo Yang","hidden":false},{"_id":"6aa0cff3d0174964227becaa","name":"Liliang Chen","hidden":false},{"_id":"6aa0cff3d0174964227becab","name":"Pengfei Zhou","hidden":false},{"_id":"6aa0cff3d0174964227becac","name":"Shengcong Chen","hidden":false},{"_id":"6aa0cff3d0174964227becad","name":"Guanghui Ren","hidden":false},{"_id":"6aa0cff3d0174964227becae","name":"Youlun Peng","hidden":false},{"_id":"6aa0cff3d0174964227becaf","name":"Rongjun Jin","hidden":false},{"_id":"6aa0cff3d0174964227becb0","name":"Nan Wang","hidden":false},{"_id":"6aa0cff3d0174964227becb1","name":"Sukai Wang","hidden":false},{"_id":"6aa0cff3d0174964227becb2","name":"Xindong He","hidden":false},{"_id":"6aa0cff3d0174964227becb3","name":"Jinyuan Feng","hidden":false},{"_id":"6aa0cff3d0174964227becb4","name":"Ziyu Xiong","hidden":false},{"_id":"6aa0cff3d0174964227becb5","name":"Linqing Zhong","hidden":false},{"_id":"6aa0cff3d0174964227becb6","name":"Yifei Wei","hidden":false},{"_id":"6aa0cff3d0174964227becb7","name":"Feng Han","hidden":false},{"_id":"6aa0cff3d0174964227becb8","name":"Long Zhang","hidden":false},{"_id":"6aa0cff3d0174964227becb9","name":"Da Huang","hidden":false},{"_id":"6aa0cff3d0174964227becba","name":"Nanshu Zhao","hidden":false},{"_id":"6aa0cff3d0174964227becbb","name":"Chenghao Yin","hidden":false},{"_id":"6aa0cff3d0174964227becbc","name":"Mo Wu","hidden":false},{"_id":"6aa0cff3d0174964227becbd","name":"Zhaodong Yan","hidden":false},{"_id":"6aa0cff3d0174964227becbe","name":"Kongtao Hu","hidden":false},{"_id":"6aa0cff3d0174964227becbf","name":"Yuxiang Yan","hidden":false},{"_id":"6aa0cff3d0174964227becc0","name":"Aogelijiang Niyazi","hidden":false},{"_id":"6aa0cff3d0174964227becc1","name":"Yu Fang","hidden":false},{"_id":"6aa0cff3d0174964227becc2","name":"Jia Zeng","hidden":false},{"_id":"6aa0cff3d0174964227becc3","name":"Lizhu Meng","hidden":false},{"_id":"6aa0cff3d0174964227becc4","name":"Daizhen Lv","hidden":false},{"_id":"6aa0cff3d0174964227becc5","name":"Haoyu Cao","hidden":false},{"_id":"6aa0cff3d0174964227becc6","name":"Zhiwen Hou","hidden":false},{"_id":"6aa0cff3d0174964227becc7","name":"Lianjin Ye","hidden":false},{"_id":"6aa0cff3d0174964227becc8","name":"Yuehan Niu","hidden":false},{"_id":"6aa0cff3d0174964227becc9","name":"Zhikai Cai","hidden":false},{"_id":"6aa0cff3d0174964227becca","name":"Xuan Hu","hidden":false},{"_id":"6aa0cff3d0174964227beccb","name":"Hui Min","hidden":false},{"_id":"6aa0cff3d0174964227beccc","name":"Xiongfeng Cai","hidden":false},{"_id":"6aa0cff3d0174964227beccd","name":"Yue Liao","hidden":false},{"_id":"6aa0cff3d0174964227becce","name":"Jing Wu","hidden":false},{"_id":"6aa0cff3d0174964227beccf","name":"Soujanya Poria","hidden":false},{"_id":"6aa0cff3d0174964227becd0","name":"Ye Li","hidden":false},{"_id":"6aa0cff3d0174964227becd1","name":"Sanping Zhou","hidden":false},{"_id":"6aa0cff3d0174964227becd2","name":"Maoqing Yao","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation","submittedOnDailyBy":{"_id":"66c4816e96583c59b09fec30","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66c4816e96583c59b09fec30/RLurCsmcgOfyuWVkpsv1L.jpeg","isPro":false,"fullname":"Ryan Chen","user":"ryancll118","type":"user","name":"ryancll118"},"summary":"World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.","upvotes":45,"discussionId":"6aa0cff3d0174964227becd3","projectPage":"https://ge-act-v2.github.io/","ai_summary":"GE-Act 2.0 is a world-action model trained from scratch with a control-oriented autoencoder, single-step visual planner, and inverse dynamics model, using knowledge-aligned selective optimization to enable scalable zero-shot robot manipulation across diverse skills and conditions.","ai_keywords":["world-action models","control-oriented autoencoder","single-step visual planner","inverse dynamics model","knowledge-aligned selective optimization","cross-embodiment transfer","zero-shot out-of-distribution"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"676fc7c31c48eff17fac3135","name":"agibot-world","fullname":"AgiBot World","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e57309b78bc92221ce3b70/ewI1QvFVMDgSsShQeSvlX.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66c4816e96583c59b09fec30","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66c4816e96583c59b09fec30/RLurCsmcgOfyuWVkpsv1L.jpeg","isPro":false,"fullname":"Ryan Chen","user":"ryancll118","type":"user"},{"_id":"64ec90919e53684e6ebcea83","avatarUrl":"/avatars/e26439f925d851e318db4995c892a9f6.svg","isPro":false,"fullname":"bigcileng","user":"bigcileng","type":"user"},{"_id":"6763e2cfd3c85f9b6d828f6c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6763e2cfd3c85f9b6d828f6c/o6yqfHC-8Cs8zFGcqZZGW.png","isPro":false,"fullname":"AgiBot World","user":"AgiBotWorldAdmin","type":"user"},{"_id":"6870f8fa63bfb5e85ff2633c","avatarUrl":"/avatars/49cc5cd57f3258f828d2b635be864efa.svg","isPro":false,"fullname":"Zhuoxuan Li","user":"tigerblob99","type":"user"},{"_id":"6874b5d36d24c879a7d6362f","avatarUrl":"/avatars/701d619e1f9fe1a766fca3f41a86fcfe.svg","isPro":false,"fullname":"chinasmokers","user":"chinasmokers","type":"user"},{"_id":"646ec9b135f55eb49e405faa","avatarUrl":"/avatars/a17194be585d20e2a021e77a5a20e213.svg","isPro":false,"fullname":"Guanghui Ren","user":"sundrops","type":"user"},{"_id":"64734b15bf9b32c6bbc8c95e","avatarUrl":"/avatars/ff3fe30c3855ad261f90dd8aae470036.svg","isPro":false,"fullname":"Liu Renhang","user":"liu-hanghang","type":"user"},{"_id":"626b626405fe1cb65725aca1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/626b626405fe1cb65725aca1/E-uD9h3n0lN04MPDbgkoH.png","isPro":false,"fullname":"Soujanya Poria","user":"soujanyaporia","type":"user"},{"_id":"65f95363ca387c9d45a2d2ad","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f95363ca387c9d45a2d2ad/6VzOniOb4eFjOWuKLV32i.jpeg","isPro":false,"fullname":"Peilin Feng","user":"Sssunset","type":"user"},{"_id":"6a8225011806973d0d442c5d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8225011806973d0d442c5d/TffQWlLJucd8ElKuMHT4h.jpeg","isPro":false,"fullname":"Sergey M. Mikhailov","user":"snmik-hailov","type":"user"},{"_id":"69f30ce2379c89a39ae575fc","avatarUrl":"/avatars/414eee7272063dad5a7844c295a69ac3.svg","isPro":false,"fullname":"zhanglong","user":"zhanglong-agibot","type":"user"},{"_id":"681d09436cd8932baccd66dc","avatarUrl":"/avatars/68335666200ca64019a7f7062f2d0315.svg","isPro":false,"fullname":"Wangzheng Wang","user":"WWWWZZ","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"676fc7c31c48eff17fac3135","name":"agibot-world","fullname":"AgiBot World","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e57309b78bc92221ce3b70/ewI1QvFVMDgSsShQeSvlX.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.05588.md","query":{}}">
Papers
arxiv:2609.05588

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

Published on Sep 4
· Submitted by
Ryan Chen
on Sep 9
Authors:
,

Abstract

GE-Act 2.0 is a world-action model trained from scratch with a control-oriented autoencoder, single-step visual planner, and inverse dynamics model, using knowledge-aligned selective optimization to enable scalable zero-shot robot manipulation across diverse skills and conditions.

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

Community

Paper submitter about 10 hours ago

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.05588
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.05588 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.05588 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.05588 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers