<video src=\"https://cdn-uploads.huggingface.co/production/uploads/65bc98383b879593a5a2f5e5/-Sqms5zHbYdvgn4XEcDbG.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-09-04T04:17:21.299Z","author":{"_id":"65bc98383b879593a5a2f5e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65bc98383b879593a5a2f5e5/p2ZtoTFN6tW-QkcPJf7YT.jpeg","fullname":"Kang Liao","name":"KangLiao","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":21,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.43339815735816956},"editors":["KangLiao"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65bc98383b879593a5a2f5e5/p2ZtoTFN6tW-QkcPJf7YT.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04196","authors":[{"_id":"6a9a44628f7c3b755723953c","user":{"_id":"65bc98383b879593a5a2f5e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65bc98383b879593a5a2f5e5/p2ZtoTFN6tW-QkcPJf7YT.jpeg","isPro":true,"fullname":"Kang Liao","user":"KangLiao","type":"user","name":"KangLiao"},"name":"Kang Liao","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.296Z","hidden":false},{"_id":"6a9a44628f7c3b755723953d","name":"Yihang Luo","hidden":false},{"_id":"6a9a44628f7c3b755723953e","name":"Xiao-Ming Wu","hidden":false},{"_id":"6a9a44628f7c3b755723953f","name":"Linyi Jin","hidden":false},{"_id":"6a9a44628f7c3b7557239540","name":"Size Wu","hidden":false},{"_id":"6a9a44628f7c3b7557239541","name":"Chunyu Lin","hidden":false},{"_id":"6a9a44628f7c3b7557239542","name":"Yao Zhao","hidden":false},{"_id":"6a9a44628f7c3b7557239543","name":"Fei Wang","hidden":false},{"_id":"6a9a44628f7c3b7557239544","name":"Wei Li","hidden":false},{"_id":"6a9a44628f7c3b7557239545","name":"Chen Change Loy","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States","submittedOnDailyBy":{"_id":"65bc98383b879593a5a2f5e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65bc98383b879593a5a2f5e5/p2ZtoTFN6tW-QkcPJf7YT.jpeg","isPro":true,"fullname":"Kang Liao","user":"KangLiao","type":"user","name":"KangLiao"},"summary":"We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.","upvotes":31,"discussionId":"6a9a44628f7c3b7557239546","projectPage":"https://kangliao929.github.io/projects/puffin-world/","ai_summary":"Puffin-World is a unified multimodal framework that jointly models physics, geometry, and appearance for physically consistent 3D world generation, reconstruction, and closed-loop exploration.","ai_keywords":["multimodal architecture","physical understanding","spatial simulation","3D world generation","3D reconstruction","gravity field","depth","Omni-Camera representation","physical dynamics propagation","generative process","vision-language-camera triplets"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6a4b593517d1bf38b7041d46","name":"ACERobotics","fullname":"ACE Robotics","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a4a45de4ba0a921a91a548a/MSX_HJLXaBYKy2ZByYTnN.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65bc98383b879593a5a2f5e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65bc98383b879593a5a2f5e5/p2ZtoTFN6tW-QkcPJf7YT.jpeg","isPro":true,"fullname":"Kang Liao","user":"KangLiao","type":"user"},{"_id":"69eb5cff11b847ca607c41ae","avatarUrl":"/avatars/fdd148677daf967b297dc5296e0e3569.svg","isPro":false,"fullname":"wu xiaoming","user":"draven-alg","type":"user"},{"_id":"61dd2b7389dddd97daead12f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61dd2b7389dddd97daead12f/xUL7YK3Mtvzz4rEhorFKF.jpeg","isPro":false,"fullname":"Xiao-Ming Wu","user":"DravenALG","type":"user"},{"_id":"66a4afec0c86556c158aee69","avatarUrl":"/avatars/1037edea5c3459d324d71c3388d5967d.svg","isPro":false,"fullname":"Jose","user":"chx7514","type":"user"},{"_id":"67459d997a49660f7f62452f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67459d997a49660f7f62452f/PLSAifYlCTOnrmAvimlUv.png","isPro":true,"fullname":"Chen Change Loy","user":"cavanloy","type":"user"},{"_id":"68ecfc9738815b886f969eb9","avatarUrl":"/avatars/473a9e2e70bcc4e47cd1fa86860d3724.svg","isPro":false,"fullname":"cyx","user":"butterfly009","type":"user"},{"_id":"68ecf820ef2baf48c6ef24c2","avatarUrl":"/avatars/5220687a5aada89ccc8e131b99a76f46.svg","isPro":false,"fullname":"sg","user":"ntusgjjll","type":"user"},{"_id":"690193714c7d26592314ceac","avatarUrl":"/avatars/94388f05aca1b792a80daf861038b2de.svg","isPro":false,"fullname":"Jonhson","user":"Jonhabc","type":"user"},{"_id":"68fa06422f3fdf3ca70db6aa","avatarUrl":"/avatars/c49319521b81e2816903b951cc01cb04.svg","isPro":false,"fullname":"Jenny","user":"JennyJulia","type":"user"},{"_id":"68ecf4c8f837dbc71215c1e9","avatarUrl":"/avatars/ae1e53c155ac7dd6d3be0e3cbfd37103.svg","isPro":false,"fullname":"fang","user":"yu789","type":"user"},{"_id":"64f059507fb910a323da5932","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f059507fb910a323da5932/vTAG_YVwwQ-lvdK2pHjuL.png","isPro":false,"fullname":"kaguramena","user":"kaguramena","type":"user"},{"_id":"68ba95c4a44514e1054952e1","avatarUrl":"/avatars/6bba5a1457a514bec32c34d20fbc7976.svg","isPro":false,"fullname":"Lang Nie","user":"lang96","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a4b593517d1bf38b7041d46","name":"ACERobotics","fullname":"ACE Robotics","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a4a45de4ba0a921a91a548a/MSX_HJLXaBYKy2ZByYTnN.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04196.md","query":{}}">
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Abstract
Puffin-World is a unified multimodal framework that jointly models physics, geometry, and appearance for physically consistent 3D world generation, reconstruction, and closed-loop exploration.
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.
Community
This comment has been hidden (marked as Resolved) Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.