arxiv: <a href=\"https://arxiv.org/abs/2608.30821\" rel=\"nofollow\">https://arxiv.org/abs/2608.30821</a><br>project page: <a href=\"https://lucida-r2s.github.io/\" rel=\"nofollow\">https://lucida-r2s.github.io/</a></p>\n","updatedAt":"2026-09-01T04:19:17.889Z","author":{"_id":"642002b51ccd411979d72b18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png","fullname":"Minghan Qin","name":"MinghanQin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.32916516065597534},"editors":["MinghanQin"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.30821","authors":[{"_id":"6a96517acd6ebc484732ed99","name":"Minghan Qin","hidden":false},{"_id":"6a96517acd6ebc484732ed9a","name":"Yuang Wang","hidden":false},{"_id":"6a96517acd6ebc484732ed9b","name":"Xiuyu Yang","hidden":false},{"_id":"6a96517acd6ebc484732ed9c","name":"Yushi Long","hidden":false},{"_id":"6a96517acd6ebc484732ed9d","name":"Yujian Zhang","hidden":false},{"_id":"6a96517acd6ebc484732ed9e","name":"Ruihuan Wang","hidden":false},{"_id":"6a96517acd6ebc484732ed9f","name":"Kai Ye","hidden":false},{"_id":"6a96517acd6ebc484732eda0","name":"Yangang Zhang","hidden":false},{"_id":"6a96517acd6ebc484732eda1","name":"Hang Li","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/642002b51ccd411979d72b18/5OGu7OkGTXPcGONQZhn8X.mp4"],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling","submittedOnDailyBy":{"_id":"642002b51ccd411979d72b18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png","isPro":false,"fullname":"Minghan Qin","user":"MinghanQin","type":"user","name":"MinghanQin"},"summary":"Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises
[email protected] from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.","upvotes":38,"discussionId":"6a96517bcd6ebc484732eda2","projectPage":"https://lucida-r2s.github.io/","ai_summary":"Lucida improves composable indoor scene reconstruction by distributing pipeline requirements across parsing, asset generation, and VLM-guided placement to achieve high-fidelity editable replicas from cluttered captures.","ai_keywords":["scene graph","multi-view evidence","VLM policy","GizmoAct","multi-turn GUI interaction","3D object detection","object pose estimation","scene reconstruction","ADD-SB","F-Score"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642002b51ccd411979d72b18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png","isPro":false,"fullname":"Minghan Qin","user":"MinghanQin","type":"user"},{"_id":"6422dd402f38c0a50cfd5405","avatarUrl":"/avatars/d3943b7b34c8026fa26eee3837590260.svg","isPro":false,"fullname":"Yuang Wang","user":"angshineee","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"6a15dabccfff5937535b56f1","avatarUrl":"/avatars/c673889a37f80cc19bf6bef0f60b2172.svg","isPro":false,"fullname":"Mateo Smith","user":"msmith25","type":"user"},{"_id":"6a1470b7b28ec6a2ad92193c","avatarUrl":"/avatars/6243b0a5d5d3c8bec740d07bdbb12950.svg","isPro":false,"fullname":"Grayson Scott","user":"grascott99","type":"user"},{"_id":"6a14692f2a9759cfbdfa9fc6","avatarUrl":"/avatars/cf44586b2b90564c128edcb6fb8cc1d0.svg","isPro":false,"fullname":"Yu Ziyi","user":"yuziyimh","type":"user"},{"_id":"6a147194d222ecc8fcefc506","avatarUrl":"/avatars/c6f23f083d4d05f4a6a05c7cc05bf99a.svg","isPro":false,"fullname":"Lily King","user":"lilyking27","type":"user"},{"_id":"6a146d3b486a5aab39d36238","avatarUrl":"/avatars/4dc1f205e97b39a2951be5300a9f3f37.svg","isPro":false,"fullname":"Oliver Lopez","user":"oliver-lopez","type":"user"},{"_id":"6a1467d8af3fe6cd43bdbad6","avatarUrl":"/avatars/ce26446e5a8452573162d27fee66a458.svg","isPro":false,"fullname":"Zhou Wenxuan","user":"zwenxuan","type":"user"},{"_id":"6a14682131430965c63f1bc3","avatarUrl":"/avatars/06ff73863b78bc69f313c8d2015ff241.svg","isPro":false,"fullname":"Hu Linxi","user":"hulinxi","type":"user"},{"_id":"6a15e69a8642d018701cf529","avatarUrl":"/avatars/d4239043b35c3c8552b4dce80d19084b.svg","isPro":false,"fullname":"Thomas ROBINSON","user":"thomasmk30","type":"user"},{"_id":"6a147f4f695d577a5249a9c8","avatarUrl":"/avatars/1194f0e6d9c809dd7767c413d64cd889.svg","isPro":false,"fullname":"Emily Brown","user":"emily-brown2025","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"},"query":{}}">
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
Abstract
Lucida improves composable indoor scene reconstruction by distributing pipeline requirements across parsing, asset generation, and VLM-guided placement to achieve high-fidelity editable replicas from cluttered captures.
Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises ADD-SB@0.05 from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.30821 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.30821 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.30821 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.