Hugging Face Daily Papers · · 5 min read

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

PhysBrain 1.5 is a unified embodied foundation model that understands the observed world, generates goal-directed actions, and predicts how the environment will evolve — all as discrete tokens under a single shared autoregressive backbone.</p>\n","updatedAt":"2026-09-15T02:53:38.824Z","author":{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","fullname":"yubin","name":"VLyb","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9145487546920776},"editors":["VLyb"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg"],"reactions":[],"isReport":false}},{"id":"6aa8f76fe27238f763a129ff","author":{"_id":"6a82b35f12403bd71c85510d","avatarUrl":"/avatars/16d14960af7a2bf7964950e2c8029674.svg","fullname":"ResearchStudio Bot","name":"researchstudio-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-15T07:44:47.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [ResearchStudio](https://microsoft.github.io/ResearchStudio/) team.\n\nWe created an interactive ResearchStudio Reel for this paper. It includes a [visual poster](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.pptx?download=true), a [video](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/video.zip?download=true), and a [blog](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/blog_en.docx?download=true), all available for download in editable formats.\n\n[![Visual poster for this paper](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.png)](https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142)\n\n[Open the ResearchStudio Reel →](https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142)\n\n[Download all files from Hugging Face](https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/all.zip?download=true)\n\nPlease give this comment a thumbs up if you find the Reel helpful!\n\nWant to explore or create Reels for more papers? Visit the [ResearchStudio demo](https://researchstudio.site/).\n\n<!-- researchstudio-bot:v1:paper=c58ecc4b9f5004a90cc8:version=b3601c03ed848a80 -->","html":"<p>This is an automated message from the <a href=\"https://microsoft.github.io/ResearchStudio/\" rel=\"nofollow\">ResearchStudio</a> team.</p>\n<p>We created an interactive ResearchStudio Reel for this paper. It includes a <a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.pptx?download=true\">visual poster</a>, a <a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/video.zip?download=true\">video</a>, and a <a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/blog_en.docx?download=true\">blog</a>, all available for download in editable formats.</p>\n<p><a href=\"https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142\" rel=\"nofollow\"><img src=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/poster.png\" alt=\"Visual poster for this paper\"></a></p>\n<p><a href=\"https://researchstudio.site/r/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142\" rel=\"nofollow\">Open the ResearchStudio Reel →</a></p>\n<p><a href=\"https://huggingface.co/buckets/researchstudio-bot/trending-paper-reels-live/resolve/published/v1/jobs/b8f8cd1e-bc40-4239-b3c5-88bc76ed9142/2609.14973/__downloads__/all.zip?download=true\">Download all files from Hugging Face</a></p>\n<p>Please give this comment a thumbs up if you find the Reel helpful!</p>\n<p>Want to explore or create Reels for more papers? Visit the <a href=\"https://researchstudio.site/\" rel=\"nofollow\">ResearchStudio demo</a>.</p>\n","updatedAt":"2026-09-15T07:44:47.289Z","author":{"_id":"6a82b35f12403bd71c85510d","avatarUrl":"/avatars/16d14960af7a2bf7964950e2c8029674.svg","fullname":"ResearchStudio Bot","name":"researchstudio-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5555163025856018},"editors":["researchstudio-bot"],"editorAvatarUrls":["/avatars/16d14960af7a2bf7964950e2c8029674.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14973","authors":[{"_id":"6aa8b22a5dd4cb9b4cc02536","name":"DeepCybo Team","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02537","name":"Yu Bin","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02538","name":"Haipeng Cao","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02539","name":"Zheng Chang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253a","name":"Kai Chen","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253b","name":"Youning Chen","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253c","name":"Kailin Deng","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253d","name":"Yichao Du","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253e","name":"Xiaotong Fu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0253f","name":"Haoyang Ge","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02540","name":"Yunlong Guo","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02541","name":"Chenliu Hao","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02542","name":"Jiyan He","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02543","name":"Xuguo He","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02544","name":"Yakun Hou","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02545","name":"Kai Hu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02546","name":"Cong Huang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02547","name":"Tuopusen Huang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02548","name":"Yu Huang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02549","name":"Hong Li","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254a","name":"Peize Li","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254b","name":"Shijie Lian","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254c","name":"Xiaopeng Lin","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254d","name":"Yun Lin","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254e","name":"Haibao Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0254f","name":"Haochen Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02550","name":"Qiuzhi Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02551","name":"Shengcai Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02552","name":"Zhiqiang Liu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02553","name":"Tao Luo","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02554","name":"Peng Ren","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02555","name":"Shuo Ren","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02556","name":"Chaoyi Ruan","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02557","name":"Zhaolong Shen","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02558","name":"Yukun Shi","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02559","name":"Qiyuan Su","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255a","name":"Yuxuan Tian","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255b","name":"Yining Wang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255c","name":"Changti Wu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255d","name":"Hao Wu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255e","name":"Xueyin Xu","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0255f","name":"Ruoqi Yang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02560","name":"Zhaoyang Yang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02561","name":"Hang Yuan","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02562","name":"Zhaoyang Zeng","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02563","name":"Hanwen Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02564","name":"Ruimeng Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02565","name":"Yao Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02566","name":"Yibo Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02567","name":"Yuxiang Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02568","name":"Zhirui Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc02569","name":"Ziyi Zhang","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0256a","name":"Zubin Zheng","hidden":false},{"_id":"6aa8b22a5dd4cb9b4cc0256b","name":"Zishen Zhuang","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models","submittedOnDailyBy":{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","isPro":false,"fullname":"yubin","user":"VLyb","type":"user","name":"VLyb"},"summary":"We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.","upvotes":58,"discussionId":"6aa8b22b5dd4cb9b4cc0256c","projectPage":"https://deepcybo-physai.github.io/PhysBrain-1.5/","githubRepo":"https://github.com/DeepCybo-PhysAI/PhysBrain-1.5","githubRepoAddedBy":"user","ai_summary":"PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.","ai_keywords":["vision-language model","autoregressive next-token prediction","end-effector motion","dense visual targets","supervised fine-tuning","embodied supervision","multimodal capabilities"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":14,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63e1d3451e5a4f34b7a728ef","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e1d3451e5a4f34b7a728ef/kHE5JHF9iZJOG0uBkvJr-.jpeg","isPro":false,"fullname":"Yichao Du","user":"yichaodu","type":"user"},{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","isPro":false,"fullname":"yubin","user":"VLyb","type":"user"},{"_id":"6a9fc91ae20e187b14cdf841","avatarUrl":"/avatars/009d1248e1a4b14cf40468084a30aeba.svg","isPro":false,"fullname":"Nick Wilde","user":"Nick-deepcoybo","type":"user"},{"_id":"6a686b4f7a167e16851568a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a686b4f7a167e16851568a9/1ykzaa_g8w8EFIJjYkdZI.jpeg","isPro":false,"fullname":"zhiyao","user":"STOPSUNS","type":"user"},{"_id":"6944c28d3cd5eeb7a7838663","avatarUrl":"/avatars/63f9f8c4e8462cbc9726bdee249fbcde.svg","isPro":false,"fullname":"lin","user":"birdxp","type":"user"},{"_id":"6433d5d69bd5a84b5dc83b17","avatarUrl":"/avatars/087edbb7ca8bc7541763b1a586a0ecb7.svg","isPro":false,"fullname":"Skylerlin","user":"Dogeeelin","type":"user"},{"_id":"664c6083f48f9e269c593105","avatarUrl":"/avatars/8cd69f3811202b6e2a1cb13ee53a8148.svg","isPro":false,"fullname":"Kai Chen","user":"chenkagi","type":"user"},{"_id":"63e60ff62d704152abac8af8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e60ff62d704152abac8af8/5kX47xGSmw8sA57O3c_rs.jpeg","isPro":false,"fullname":"Qiuzhi Liu","user":"Dennis364","type":"user"},{"_id":"69809e4ba74ada34274f1e74","avatarUrl":"/avatars/d1a2949136f2536cbc9bbfbb1f129784.svg","isPro":false,"fullname":"DATOUYU","user":"DATOUYU-130","type":"user"},{"_id":"6aa8b67300b5bf93fa0cb768","avatarUrl":"/avatars/5ee3b20cd843e25e1dd7a6813a016038.svg","isPro":false,"fullname":"kdonmyswag","user":"kdonmyswag","type":"user"},{"_id":"69491ab31b1af1c40a0ac5eb","avatarUrl":"/avatars/241e32efcd52ded0712ec03ab7a6949e.svg","isPro":false,"fullname":"Kailin Deng","user":"iceballoon","type":"user"},{"_id":"6948de92774219e16688c043","avatarUrl":"/avatars/3d45f9ba6960173e593587c977f01e00.svg","isPro":false,"fullname":"hexuguo","user":"hexuguo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14973.md","query":{}}">
Papers
arxiv:2609.14973

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

Published on Sep 14
· Submitted by
yubin
on Sep 15
Authors:
,

Abstract

PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.

We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.

Community

Paper submitter about 5 hours ago

PhysBrain 1.5 is a unified embodied foundation model that understands the observed world, generates goal-directed actions, and predicts how the environment will evolve — all as discrete tokens under a single shared autoregressive backbone.

This is an automated message from the ResearchStudio team.

We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.

Visual poster for this paper

Open the ResearchStudio Reel →

Download all files from Hugging Face

Please give this comment a thumbs up if you find the Reel helpful!

Want to explore or create Reels for more papers? Visit the ResearchStudio demo.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.14973
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.14973 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers