PhysBrain 1.5 is a unified embodied foundation model that understands the observed world, generates goal-directed actions, and predicts how the environment will evolve — all as discrete tokens under a single shared autoregressive backbone.</p>\n","updatedAt":"2026-09-15T02:53:38.824Z","author":{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","fullname":"yubin","name":"VLyb","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9145487546920776},"editors":["VLyb"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg"],"reactions":[],"isReport":false}},{"id":"6aa8f76fe27238f763a129ff","author":{"_id":"6a82b35f12403bd71c85510d","avatarUrl":"/avatars/16d14960af7a2bf7964950e2c8029674.svg","fullname":"ResearchStudio Bot","name":"researchstudio-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-15T07:44:47.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [ResearchStudio](https://microsoft.github.io/ResearchStudio/) team.\n\nWe created an interactive ResearchStudio Reel for this paper. It includes a [visual poster](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.pptx?download=true), a [video](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/video.zip?download=true), and a [blog](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/blog_en.docx?download=true), all available for download in editable formats.\n\n[](https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142)\n\n[Open the ResearchStudio Reel →](https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142)\n\n[Download all files from Hugging Face](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/all.zip?download=true)\n\nPlease give this comment a thumbs up if you find the Reel helpful!\n\nWant to explore or create Reels for more papers? Visit the [ResearchStudio demo](https://researchstudio.site/).\n\n<!-- researchstudio-bot:v1:paper=c58ecc4b9f5004a90cc8:version=b3601c03ed848a80 -->","html":"<p>This is an automated message from the <a href=\"https://microsoft.github.io/ResearchStudio/\" rel=\"nofollow\">ResearchStudio</a> team.</p>\n<p>We created an interactive ResearchStudio Reel for this paper. It includes a <a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.pptx?download=true\">visual poster</a>, a <a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/video.zip?download=true\">video</a>, and a <a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/blog_en.docx?download=true\">blog</a>, all available for download in editable formats.</p>\n<p><a href=\"https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142\" rel=\"nofollow\"><img src=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.png\" alt=\"Visual poster for this paper\"></a></p>\n<p><a href=\"https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142\" rel=\"nofollow\">Open the ResearchStudio Reel →</a></p>\n<p><a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/all.zip?download=true\">Download all files from Hugging Face</a></p>\n<p>Please give this comment a thumbs up if you find the Reel helpful!</p>\n<p>Want to explore or create Reels for more papers? Visit the <a href=\"https://researchstudio.site/\" rel=\"nofollow\">ResearchStudio demo</a>.</p>\n","updatedAt":"2026-09-15T07:44:47.289Z","author":{"_id":"6a82b35f12403bd71c85510d","avatarUrl":"/avatars/16d14960af7a2bf7964950e2c8029674.svg","fullname":"ResearchStudio Bot","name":"researchstudio-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5555163025856018},"editors":["researchstudio-bot"],"editorAvatarUrls":["/avatars/16d14960af7a2bf7964950e2c8029674.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14973","authors":[{"_id":"6aa8b22a5dd4cb9b4cc02536","name":"DeepCybo Team","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02537","name":"Yu Bin","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02538","name":"Haipeng Cao","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02539","name":"Zheng Chang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253a","name":"Kai Chen","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253b","name":"Youning Chen","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253c","name":"Kailin Deng","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253d","name":"Yichao Du","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253e","name":"Xiaotong Fu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253f","name":"Haoyang Ge","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02540","name":"Yunlong Guo","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02541","name":"Chenliu Hao","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02542","name":"Jiyan He","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02543","name":"Xuguo He","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02544","name":"Yakun Hou","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02545","name":"Kai Hu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02546","name":"Cong Huang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02547","name":"Tuopusen Huang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02548","name":"Yu Huang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02549","name":"Hong Li","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254a","name":"Peize Li","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254b","name":"Shijie Lian","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254c","name":"Xiaopeng Lin","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254d","name":"Yun Lin","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254e","name":"Haibao Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254f","name":"Haochen Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02550","name":"Qiuzhi Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02551","name":"Shengcai Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02552","name":"Zhiqiang Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02553","name":"Tao Luo","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02554","name":"Peng Ren","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02555","name":"Shuo Ren","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02556","name":"Chaoyi Ruan","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02557","name":"Zhaolong Shen","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02558","name":"Yukun Shi","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02559","name":"Qiyuan Su","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255a","name":"Yuxuan Tian","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255b","name":"Yining Wang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255c","name":"Changti Wu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255d","name":"Hao Wu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255e","name":"Xueyin Xu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255f","name":"Ruoqi Yang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02560","name":"Zhaoyang Yang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02561","name":"Hang Yuan","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02562","name":"Zhaoyang Zeng","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02563","name":"Hanwen Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02564","name":"Ruimeng Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02565","name":"Yao Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02566","name":"Yibo Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02567","name":"Yuxiang Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02568","name":"Zhirui Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02569","name":"Ziyi Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0256a","name":"Zubin Zheng","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0256b","name":"Zishen Zhuang","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models","submittedOnDailyBy":{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","isPro":false,"fullname":"yubin","user":"VLyb","type":"user","name":"VLyb"},"summary":"We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.","upvotes":58,"discussionId":"6aa8b22b5dd4cb9b4cc0256c","projectPage":"https://deepcybo-physai.github.io/PhysBrain-1.5/","githubRepo":"https://github.com/DeepCybo-PhysAI/PhysBrain-1.5","githubRepoAddedBy":"user","ai_summary":"PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.","ai_keywords":["vision-language model","autoregressive next-token prediction","end-effector motion","dense visual targets","supervised fine-tuning","embodied supervision","multimodal capabilities"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":14,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63e1d3451e5a4f34b7a728ef","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e1d3451e5a4f34b7a728ef/kHE5JHF9iZJOG0uBkvJr-.jpeg","isPro":false,"fullname":"Yichao Du","user":"yichaodu","type":"user"},{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","isPro":false,"fullname":"yubin","user":"VLyb","type":"user"},{"_id":"6a9fc91ae20e187b14cdf841","avatarUrl":"/avatars/009d1248e1a4b14cf40468084a30aeba.svg","isPro":false,"fullname":"Nick Wilde","user":"Nick-deepcoybo","type":"user"},{"_id":"6a686b4f7a167e16851568a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a686b4f7a167e16851568a9/1ykzaa_g8w8EFIJjYkdZI.jpeg","isPro":false,"fullname":"zhiyao","user":"STOPSUNS","type":"user"},{"_id":"6944c28d3cd5eeb7a7838663","avatarUrl":"/avatars/63f9f8c4e8462cbc9726bdee249fbcde.svg","isPro":false,"fullname":"lin","user":"birdxp","type":"user"},{"_id":"6433d5d69bd5a84b5dc83b17","avatarUrl":"/avatars/087edbb7ca8bc7541763b1a586a0ecb7.svg","isPro":false,"fullname":"Skylerlin","user":"Dogeeelin","type":"user"},{"_id":"664c6083f48f9e269c593105","avatarUrl":"/avatars/8cd69f3811202b6e2a1cb13ee53a8148.svg","isPro":false,"fullname":"Kai Chen","user":"chenkagi","type":"user"},{"_id":"63e60ff62d704152abac8af8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e60ff62d704152abac8af8/5kX47xGSmw8sA57O3c_rs.jpeg","isPro":false,"fullname":"Qiuzhi Liu","user":"Dennis364","type":"user"},{"_id":"69809e4ba74ada34274f1e74","avatarUrl":"/avatars/d1a2949136f2536cbc9bbfbb1f129784.svg","isPro":false,"fullname":"DATOUYU","user":"DATOUYU-130","type":"user"},{"_id":"6aa8b67300b5bf93fa0cb768","avatarUrl":"/avatars/5ee3b20cd843e25e1dd7a6813a016038.svg","isPro":false,"fullname":"kdonmyswag","user":"kdonmyswag","type":"user"},{"_id":"69491ab31b1af1c40a0ac5eb","avatarUrl":"/avatars/241e32efcd52ded0712ec03ab7a6949e.svg","isPro":false,"fullname":"Kailin Deng","user":"iceballoon","type":"user"},{"_id":"6948de92774219e16688c043","avatarUrl":"/avatars/3d45f9ba6960173e593587c977f01e00.svg","isPro":false,"fullname":"hexuguo","user":"hexuguo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14973.md","query":{}}">
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
Published on Sep 14
· Submitted by yubin on Sep 15 Abstract
PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.
Community
PhysBrain 1.5 is a unified embodied foundation model that understands the observed world, generates goal-directed actions, and predicts how the environment will evolve — all as discrete tokens under a single shared autoregressive backbone.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.14973 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.