Hugging Face Daily Papers · · 4 min read

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

arxiv: <a href=\"https://arxiv.org/abs/2608.30821\" rel=\"nofollow\">https://arxiv.org/abs/2608.30821</a><br>project page: <a href=\"https://lucida-r2s.github.io/\" rel=\"nofollow\">https://lucida-r2s.github.io/</a></p>\n","updatedAt":"2026-09-01T04:19:17.889Z","author":{"_id":"642002b51ccd411979d72b18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png","fullname":"Minghan Qin","name":"MinghanQin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.32916516065597534},"editors":["MinghanQin"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.30821","authors":[{"_id":"6a96517acd6ebc484732ed99","name":"Minghan Qin","hidden":false},{"_id":"6a96517acd6ebc484732ed9a","name":"Yuang Wang","hidden":false},{"_id":"6a96517acd6ebc484732ed9b","name":"Xiuyu Yang","hidden":false},{"_id":"6a96517acd6ebc484732ed9c","name":"Yushi Long","hidden":false},{"_id":"6a96517acd6ebc484732ed9d","name":"Yujian Zhang","hidden":false},{"_id":"6a96517acd6ebc484732ed9e","name":"Ruihuan Wang","hidden":false},{"_id":"6a96517acd6ebc484732ed9f","name":"Kai Ye","hidden":false},{"_id":"6a96517acd6ebc484732eda0","name":"Yangang Zhang","hidden":false},{"_id":"6a96517acd6ebc484732eda1","name":"Hang Li","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/642002b51ccd411979d72b18/5OGu7OkGTXPcGONQZhn8X.mp4"],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling","submittedOnDailyBy":{"_id":"642002b51ccd411979d72b18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png","isPro":false,"fullname":"Minghan Qin","user":"MinghanQin","type":"user","name":"MinghanQin"},"summary":"Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises [email protected] from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.","upvotes":38,"discussionId":"6a96517bcd6ebc484732eda2","projectPage":"https://lucida-r2s.github.io/","ai_summary":"Lucida improves composable indoor scene reconstruction by distributing pipeline requirements across parsing, asset generation, and VLM-guided placement to achieve high-fidelity editable replicas from cluttered captures.","ai_keywords":["scene graph","multi-view evidence","VLM policy","GizmoAct","multi-turn GUI interaction","3D object detection","object pose estimation","scene reconstruction","ADD-SB","F-Score"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642002b51ccd411979d72b18","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642002b51ccd411979d72b18/JO9e0o8fAFNC-eYFbx_JW.png","isPro":false,"fullname":"Minghan Qin","user":"MinghanQin","type":"user"},{"_id":"6422dd402f38c0a50cfd5405","avatarUrl":"/avatars/d3943b7b34c8026fa26eee3837590260.svg","isPro":false,"fullname":"Yuang Wang","user":"angshineee","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"6a15dabccfff5937535b56f1","avatarUrl":"/avatars/c673889a37f80cc19bf6bef0f60b2172.svg","isPro":false,"fullname":"Mateo Smith","user":"msmith25","type":"user"},{"_id":"6a1470b7b28ec6a2ad92193c","avatarUrl":"/avatars/6243b0a5d5d3c8bec740d07bdbb12950.svg","isPro":false,"fullname":"Grayson Scott","user":"grascott99","type":"user"},{"_id":"6a14692f2a9759cfbdfa9fc6","avatarUrl":"/avatars/cf44586b2b90564c128edcb6fb8cc1d0.svg","isPro":false,"fullname":"Yu Ziyi","user":"yuziyimh","type":"user"},{"_id":"6a147194d222ecc8fcefc506","avatarUrl":"/avatars/c6f23f083d4d05f4a6a05c7cc05bf99a.svg","isPro":false,"fullname":"Lily King","user":"lilyking27","type":"user"},{"_id":"6a146d3b486a5aab39d36238","avatarUrl":"/avatars/4dc1f205e97b39a2951be5300a9f3f37.svg","isPro":false,"fullname":"Oliver Lopez","user":"oliver-lopez","type":"user"},{"_id":"6a1467d8af3fe6cd43bdbad6","avatarUrl":"/avatars/ce26446e5a8452573162d27fee66a458.svg","isPro":false,"fullname":"Zhou Wenxuan","user":"zwenxuan","type":"user"},{"_id":"6a14682131430965c63f1bc3","avatarUrl":"/avatars/06ff73863b78bc69f313c8d2015ff241.svg","isPro":false,"fullname":"Hu Linxi","user":"hulinxi","type":"user"},{"_id":"6a15e69a8642d018701cf529","avatarUrl":"/avatars/d4239043b35c3c8552b4dce80d19084b.svg","isPro":false,"fullname":"Thomas ROBINSON","user":"thomasmk30","type":"user"},{"_id":"6a147f4f695d577a5249a9c8","avatarUrl":"/avatars/1194f0e6d9c809dd7767c413d64cd889.svg","isPro":false,"fullname":"Emily Brown","user":"emily-brown2025","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"67d1140985ea0644e2f14b99","name":"ByteDance-Seed","fullname":"ByteDance Seed","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/flkDUqd_YEuFsjeNET3r-.png"},"query":{}}">
Papers
arxiv:2608.30821

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Published on Aug 31
· Submitted by
Minghan Qin
on Sep 1
#2 Paper of the day
Authors:
,

Abstract

Lucida improves composable indoor scene reconstruction by distributing pipeline requirements across parsing, asset generation, and VLM-guided placement to achieve high-fidelity editable replicas from cluttered captures.

Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises ADD-SB@0.05 from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.30821 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.30821 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.30821 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers