Excited to share Hy-Embodied-RxBrain, a foundation model for embodied cognition that jointly represents plans through language reasoning and visual imagination. RxBrain couples task decomposition, constraints, temporal logic, world-state prediction, and visual subgoals in a unified planning sequence. We also introduce an automatic text-visual supervision pipeline and RxBrain-Bench for evaluating joint embodied planning. The model further shows promising real-robot action generation performance without large-scale action-data pretraining.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/663237779292069aed584688/64h5kUi7zkT040ckgXJu9.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/663237779292069aed584688/64h5kUi7zkT040ckgXJu9.png\" alt=\"teaser\"></a></p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/663237779292069aed584688/sHpIz7GzhIlAJsfYKBBw8.qt\" controls=\"\" class=\"max-w-full!\"></video></p>","updatedAt":"2026-07-17T16:41:40.564Z","author":{"_id":"663237779292069aed584688","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663237779292069aed584688/eWblPrit8CqaP0Aq5jedB.jpeg","fullname":"Haotian Liang","name":"hhyhrhy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8253059983253479},"editors":["hhyhrhy"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/663237779292069aed584688/eWblPrit8CqaP0Aq5jedB.jpeg"],"reactions":[],"isReport":false}},{"id":"6a5af99b790a516fd7dc996b","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":376,"isUserFollowing":false},"createdAt":"2026-07-18T03:57:15.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [iFLYTEK-Embodied-Omni Technical Report](https://huggingface.co/papers/2607.02542) (2026)\n* [ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation](https://huggingface.co/papers/2606.17937) (2026)\n* [Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments](https://huggingface.co/papers/2605.30280) (2026)\n* [Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation](https://huggingface.co/papers/2606.17030) (2026)\n* [WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling](https://huggingface.co/papers/2607.03461) (2026)\n* [Hy-Embodied-VLM-1.0: Efficient Physical-World Agents](https://huggingface.co/papers/2607.12894) (2026)\n* [Native Video-Action Pretraining for Generalizable Robot Control](https://huggingface.co/papers/2607.08639) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.02542\">iFLYTEK-Embodied-Omni Technical Report</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.17937\">ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2605.30280\">Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.17030\">Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.03461\">WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.12894\">Hy-Embodied-VLM-1.0: Efficient Physical-World Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.08639\">Native Video-Action Pretraining for Generalizable Robot Control</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-07-18T03:57:15.063Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":376,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7229870557785034},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.14187","authors":[{"_id":"6a59dc336c2e371e6ca38383","name":"Haotian Liang","hidden":false},{"_id":"6a59dc336c2e371e6ca38384","name":"Mingkang Chen","hidden":false},{"_id":"6a59dc336c2e371e6ca38385","name":"Yufei Huang","hidden":false},{"_id":"6a59dc336c2e371e6ca38386","name":"Yuchun Guo","hidden":false},{"_id":"6a59dc336c2e371e6ca38387","name":"Xiaomeng Zhu","hidden":false},{"_id":"6a59dc336c2e371e6ca38388","name":"Xiangli Shi","hidden":false},{"_id":"6a59dc336c2e371e6ca38389","name":"Kaixuan Wang","hidden":false},{"_id":"6a59dc336c2e371e6ca3838a","name":"Yunxuan Mao","hidden":false},{"_id":"6a59dc336c2e371e6ca3838b","name":"Weijie Zhou","hidden":false},{"_id":"6a59dc336c2e371e6ca3838c","name":"Ling Chen","hidden":false},{"_id":"6a59dc336c2e371e6ca3838d","name":"Shirong Zeng","hidden":false},{"_id":"6a59dc336c2e371e6ca3838e","name":"Yueyu Long","hidden":false},{"_id":"6a59dc336c2e371e6ca3838f","name":"Yuchen Si","hidden":false},{"_id":"6a59dc336c2e371e6ca38390","name":"Yajuan Zhu","hidden":false},{"_id":"6a59dc336c2e371e6ca38391","name":"Xingyu Zhou","hidden":false},{"_id":"6a59dc336c2e371e6ca38392","name":"Minghui Wang","hidden":false},{"_id":"6a59dc336c2e371e6ca38393","name":"Wanjia He","hidden":false},{"_id":"6a59dc336c2e371e6ca38394","name":"Xin Yang","hidden":false},{"_id":"6a59dc336c2e371e6ca38395","name":"Lingzhu Xiang","hidden":false},{"_id":"6a59dc336c2e371e6ca38396","name":"Zhiqing Liu","hidden":false},{"_id":"6a59dc336c2e371e6ca38397","name":"Bohan Ma","hidden":false},{"_id":"6a59dc336c2e371e6ca38398","name":"Xiran Huang","hidden":false},{"_id":"6a59dc336c2e371e6ca38399","name":"Tianshuo Yang","hidden":false},{"_id":"6a59dc336c2e371e6ca3839a","name":"Zhiheng Liu","hidden":false},{"_id":"6a59dc336c2e371e6ca3839b","name":"Xuantang Xiong","hidden":false},{"_id":"6a59dc336c2e371e6ca3839c","name":"Zisheng Lu","hidden":false},{"_id":"6a59dc336c2e371e6ca3839d","name":"Ping Luo","hidden":false},{"_id":"6a59dc336c2e371e6ca3839e","name":"Yao Mu","hidden":false},{"_id":"6a59dc336c2e371e6ca3839f","name":"Han Hu","hidden":false},{"_id":"6a59dc336c2e371e6ca383a0","name":"Zhengyou Zhang","hidden":false}],"publishedAt":"2026-07-15T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination","submittedOnDailyBy":{"_id":"663237779292069aed584688","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663237779292069aed584688/eWblPrit8CqaP0Aq5jedB.jpeg","isPro":false,"fullname":"Haotian Liang","user":"hhyhrhy","type":"user","name":"hhyhrhy"},"summary":"Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.","upvotes":21,"discussionId":"6a59dc336c2e371e6ca383a1","projectPage":"https://tairos.tencent.com/openSourceModels/hy-embodied-rxbrain-1.0","githubRepo":"https://github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0","githubRepoAddedBy":"user","githubStars":81},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"663237779292069aed584688","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663237779292069aed584688/eWblPrit8CqaP0Aq5jedB.jpeg","isPro":false,"fullname":"Haotian Liang","user":"hhyhrhy","type":"user"},{"_id":"69e1d51ac035e98a57ebe6ca","avatarUrl":"/avatars/6c60346be92447593c07d2ce86d0ce0b.svg","isPro":false,"fullname":"TC","user":"Yuanbeval","type":"user"},{"_id":"67cb0302edbee142edd5e2db","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PqnUKiy-60f_2mmmlM7Wd.png","isPro":false,"fullname":"Violet Evergarden","user":"VioletEvergarden1","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6455fd90bfdf9c63ce2d32a9","avatarUrl":"/avatars/119255bddd172b862ac78333727a1842.svg","isPro":false,"fullname":"Xiaomeng Zhu","user":"Zhuxmmm","type":"user"},{"_id":"657010abf2c1c863093941cb","avatarUrl":"/avatars/25b8a1ecdea3b5ab687dc8331b5f8615.svg","isPro":false,"fullname":"Tianshuo","user":"Violin-Y","type":"user"},{"_id":"68590b26dfeeddc33dd2abd4","avatarUrl":"/avatars/395a247a98a06f7c6ec6f8965663bb20.svg","isPro":false,"fullname":"Shunyao Jiang","user":"yakamoz666","type":"user"},{"_id":"6a5b01e611bdcb9fbc4d42e2","avatarUrl":"/avatars/b50f887e56687f99ec8904a762355b96.svg","isPro":false,"fullname":"ShiinTaki","user":"ShiinaTaki2027","type":"user"},{"_id":"669338223cd89359e4d91f6b","avatarUrl":"/avatars/6ad38f17d9d19247eae91c50505d16c1.svg","isPro":false,"fullname":"Yunwei Li","user":"YunweiLi","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"641c04c3a6d3ff92426ef297","avatarUrl":"/avatars/37181e30486d8279c629899b545aee17.svg","isPro":false,"fullname":"Yifei Yang","user":"Yifei00","type":"user"},{"_id":"67fcc97cede5c434e0cc37e3","avatarUrl":"/avatars/b07e0a4744c1045828a621146ee6d3c2.svg","isPro":false,"fullname":"yunxuan mao","user":"maoyunxuan","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Abstract
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.
Community
Excited to share Hy-Embodied-RxBrain, a foundation model for embodied cognition that jointly represents plans through language reasoning and visual imagination. RxBrain couples task decomposition, constraints, temporal logic, world-state prediction, and visual subgoals in a unified planning sequence. We also introduce an automatic text-visual supervision pipeline and RxBrain-Bench for evaluating joint embodied planning. The model further shows promising real-robot action generation performance without large-scale action-data pretraining.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.14187 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.