Hugging Face Daily Papers · · 4 min read

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.</p>\n<p>Github: <a href=\"https://github.com/alibaba-damo-academy/RynnBrain\" rel=\"nofollow\">https://github.com/alibaba-damo-academy/RynnBrain</a></p>\n","updatedAt":"2026-07-21T03:09:00.312Z","author":{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","fullname":"Siteng Huang","name":"huangsiteng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg","fullname":"DAMO Academy","name":"Alibaba-DAMO-Academy","type":"org","isHf":false}}},"numEdits":2,"identifiedLanguage":{"language":"en","probability":0.9167981147766113},"editors":["huangsiteng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png"],"reactions":[],"isReport":false},"replies":[{"id":"6a5ee26749bc19ec22ad1fc1","author":{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","fullname":"Siteng Huang","name":"huangsiteng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg","fullname":"DAMO Academy","name":"Alibaba-DAMO-Academy","type":"org","isHf":false}},"createdAt":"2026-07-21T03:07:19.000Z","type":"comment","data":{"edited":true,"hidden":true,"hiddenBy":"","hiddenReason":"Resolved","latest":{"raw":"This comment has been hidden","html":"This comment has been hidden","updatedAt":"2026-07-21T03:08:30.515Z","author":{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","fullname":"Siteng Huang","name":"huangsiteng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg","fullname":"DAMO Academy","name":"Alibaba-DAMO-Academy","type":"org","isHf":false}}},"numEdits":0,"editors":[],"editorAvatarUrls":[],"reactions":[],"parentCommentId":"6a5ee151e2cbfb2c84c3248f"}}]}],"primaryEmailConfirmed":false,"paper":{"id":"2607.17977","authors":[{"_id":"6a5edcda4fe5d1d13e84ab1e","name":"Kehan Li","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab1f","name":"Bohan Hou","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab20","name":"Minghao Zhu","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab21","name":"Tianyi Zhang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab22","name":"Zesen Cheng","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab23","name":"Zhikai Wang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab24","name":"Sicong Leng","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab25","name":"Xin Li","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab26","name":"Xiao Lin","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab27","name":"Biying Yao","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab28","name":"Minghua Zeng","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab29","name":"Jiangpin Liu","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab2a","name":"Ronghao Dang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab2b","name":"Jiayan Guo","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab2c","name":"Siteng Huang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab2d","name":"Haoyu Zhao","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab2e","name":"Heng Ping","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab2f","name":"Yaxi Zhao","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab30","name":"Kexiang Wang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab31","name":"Tong Lu","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab32","name":"Shengke Xue","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab33","name":"Jiahao Tang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab34","name":"Yulei Wang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab35","name":"Zejing Wang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab36","name":"Jianwei Gao","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab37","name":"Shijian Lu","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab38","name":"Chengju Liu","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab39","name":"Jianfei Yang","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab3a","name":"Mingxiu Chen","hidden":false},{"_id":"6a5edcda4fe5d1d13e84ab3b","name":"Deli Zhao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65fd82762bf2cd20ddaa193f/AbDUUINXvYXZD23bI7a-A.mp4"],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-21T00:00:00.000Z","title":"RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model","submittedOnDailyBy":{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","isPro":false,"fullname":"Siteng Huang","user":"huangsiteng","type":"user","name":"huangsiteng"},"summary":"We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.","upvotes":21,"discussionId":"6a5edcdb4fe5d1d13e84ab3c","projectPage":"https://alibaba-damo-academy.github.io/RynnBrain/","organization":{"_id":"6808e7522a4d69d5111da55f","name":"Alibaba-DAMO-Academy","fullname":"DAMO Academy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65fd82762bf2cd20ddaa193f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/yBYbWp_mT7UusYdkqtAvw.png","isPro":false,"fullname":"Siteng Huang","user":"huangsiteng","type":"user"},{"_id":"63913b120cf6b11c487ca31d","avatarUrl":"/avatars/aec44edd5470dd6e767e0a25efd6fb5d.svg","isPro":false,"fullname":"Xin Li","user":"lixin4ever","type":"user"},{"_id":"609115c79a8bcaa437b234a9","avatarUrl":"/avatars/1631a91030703d8397133363cf82c863.svg","isPro":false,"fullname":"Leng Sicong","user":"Sicong","type":"user"},{"_id":"64c9e86a6a26cddbecd9bae2","avatarUrl":"/avatars/61a84989dbbc1898ebcba3236dbed039.svg","isPro":false,"fullname":"Tianyi Zhang","user":"TianyiZhang0213","type":"user"},{"_id":"65388514613fe158bd514e4c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65388514613fe158bd514e4c/hpCY_g2oxLd1Ruq8LTA29.jpeg","isPro":false,"fullname":"alterego238","user":"alterego238","type":"user"},{"_id":"692fa5d17ff1da99eb783dfb","avatarUrl":"/avatars/5477343d26250cbda7babb8f1fdee49d.svg","isPro":false,"fullname":"Qize Yu","user":"Skywalker0410","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6388af095a3d2a335622cb7c","avatarUrl":"/avatars/f548ce6a902cee8bdc74179bcd45534c.svg","isPro":false,"fullname":"Kehan Li","user":"lkhl","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6808e64de5dd22427c006e10","avatarUrl":"/avatars/834f8e232f49880bfe861b6c435c11ab.svg","isPro":false,"fullname":"DAMO_admin","user":"DAMO-Academy","type":"user"},{"_id":"63c3b12da361002ba0ae5381","avatarUrl":"/avatars/f505a612c7698b5975a1daa3cca6a1c3.svg","isPro":false,"fullname":"Han Zhang","user":"bibona","type":"user"},{"_id":"6571869cf3853f99bc558177","avatarUrl":"/avatars/505e3775831b50918a834258e04f7843.svg","isPro":false,"fullname":"Tian Bian","user":"TianB","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6808e7522a4d69d5111da55f","name":"Alibaba-DAMO-Academy","fullname":"DAMO Academy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.17977.md","query":{}}">
Papers
arxiv:2607.17977

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Published on Jul 20
· Submitted by
Siteng Huang
on Jul 21
Authors:
,

Abstract

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.

Community

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.

Github: https://github.com/alibaba-damo-academy/RynnBrain

This comment has been hidden (marked as Resolved)
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Upvote
21

Get this paper in your agent:

hf papers read 2607.17977
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.17977 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.17977 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.17977 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers