Large vision-language models describe what a scene contains well, yet still struggle to reconstruct the 3D structure behind a 2D image and reason about it — spatial intelligence. SpatialBlock takes the route human spatial cognition develops along — block play. We train LVLMs on SpatialBlock-15k, a fully synthetic, scalable set of 15,000 block-stacking problems spanning 3D-to-2D projection, viewpoint transformation and structural combination, extended with controlled color cues so that models learn to anchor their reasoning on task-relevant blocks. Two training strategies are released: a direct model that predicts the answer immediately, and a reason model that thinks before answering. Despite training only on synthetic data at a small scale, both transfer to real-scene spatial benchmarks.</p>\n","updatedAt":"2026-09-11T02:49:24.292Z","author":{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","fullname":"Soohyun Ryu","name":"rsoohyun","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9155384302139282},"editors":["rsoohyun"],"editorAvatarUrls":["/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07064","authors":[{"_id":"6aa21457a2aeb74440b1dd87","user":{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","isPro":false,"fullname":"Soohyun Ryu","user":"rsoohyun","type":"user","name":"rsoohyun"},"name":"Soohyun Ryu","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.222Z","hidden":false},{"_id":"6aa21457a2aeb74440b1dd88","user":{"_id":"6690a286181e2af45c742dd8","avatarUrl":"/avatars/511d0f86386e3b29a17b445d855b3aef.svg","isPro":false,"fullname":"Sohee Kim","user":"joyhee","type":"user","name":"joyhee"},"name":"Sohee Kim","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.229Z","hidden":false},{"_id":"6aa21457a2aeb74440b1dd89","name":"Eunho Yang","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem","submittedOnDailyBy":{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","isPro":false,"fullname":"Soohyun Ryu","user":"rsoohyun","type":"user","name":"rsoohyun"},"summary":"Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks. We introduce SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination. The dataset further incorporates controlled color modulation as visual cues to encourage anchor-based reasoning in visually complex conditions. Experiments demonstrate that LVLMs trained on our dataset through either direct answering or reasoning-based prediction significantly outperform baselines and generalize to real-world spatial tasks, despite the dataset's synthetic and compact nature. Code and data are available at https://github.com/rsoohyun/SpatialBlock.","upvotes":42,"discussionId":"6aa21458a2aeb74440b1dd8a","githubRepo":"https://github.com/rsoohyun/SpatialBlock","githubRepoAddedBy":"user","ai_summary":"Large vision-language models trained on synthetic block-manipulation tasks improve 3D spatial reasoning and generalize to real-world visual tasks.","ai_keywords":["Large Vision-Language Models","spatial intelligence","3D-to-2D projection","viewpoint transformation","structural combination","anchor-based reasoning","reasoning-based prediction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6690a286181e2af45c742dd8","avatarUrl":"/avatars/511d0f86386e3b29a17b445d855b3aef.svg","isPro":false,"fullname":"Sohee Kim","user":"joyhee","type":"user"},{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","isPro":false,"fullname":"Soohyun Ryu","user":"rsoohyun","type":"user"},{"_id":"66303ce3e1c93377db71efd5","avatarUrl":"/avatars/3d444dfb9799c9324c98cba893f4a10f.svg","isPro":false,"fullname":"Yoon Sik Park","user":"nooynoos","type":"user"},{"_id":"663073e4b108b8d71d4b6f32","avatarUrl":"/avatars/f66de2d81205397a29baf21169157bdd.svg","isPro":false,"fullname":"Inki Park","user":"seoharuss","type":"user"},{"_id":"62845957b410bd779033759c","avatarUrl":"/avatars/4feef73c06f2f7de6abf7a4789ac13f9.svg","isPro":false,"fullname":"Doohyuk Jang","user":"jadohu","type":"user"},{"_id":"657a64c91ccc3c2a5ea5cde4","avatarUrl":"/avatars/aee59f3650cc71424591340ec9f862b2.svg","isPro":false,"fullname":"Hangyeol Jung","user":"Hangyeol","type":"user"},{"_id":"66021a80561089b1a7ebfa01","avatarUrl":"/avatars/a61cf6f95fa0d6750297a00f897436dc.svg","isPro":false,"fullname":"JUNHYEOK CHOI","user":"junhyeokchoi","type":"user"},{"_id":"6666b61eaf95872a03a0a673","avatarUrl":"/avatars/fc0c144cf6307357d45d7ca2d6ba8d2f.svg","isPro":false,"fullname":"Joowon","user":"kjwispro","type":"user"},{"_id":"6371ce78789970f7bc673234","avatarUrl":"/avatars/ba363e0b8fddee143244934be7bc6db0.svg","isPro":false,"fullname":"Donghyeon Cho","user":"hyeon9698","type":"user"},{"_id":"62a4d58e81a4b10e93064ad6","avatarUrl":"/avatars/744d5cbc1745a26b816a458260aba050.svg","isPro":false,"fullname":"hangyulyoon","user":"hangyulmd","type":"user"},{"_id":"6aa36dde9d2d63a5b59fcc89","avatarUrl":"/avatars/e5424114acd61c3a510f13a589de6ee9.svg","isPro":false,"fullname":"Taeyong Choi","user":"tay-choi","type":"user"},{"_id":"664f558b6f16bbd9a1481b59","avatarUrl":"/avatars/581c88278e1d8bba66ee3b764f6dd3ed.svg","isPro":false,"fullname":"Jo sungmin","user":"JoJosmin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07064.md","query":{}}">
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
Abstract
Large vision-language models trained on synthetic block-manipulation tasks improve 3D spatial reasoning and generalize to real-world visual tasks.
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks. We introduce SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination. The dataset further incorporates controlled color modulation as visual cues to encourage anchor-based reasoning in visually complex conditions. Experiments demonstrate that LVLMs trained on our dataset through either direct answering or reasoning-based prediction significantly outperform baselines and generalize to real-world spatial tasks, despite the dataset's synthetic and compact nature. Code and data are available at https://github.com/rsoohyun/SpatialBlock.
Community
Large vision-language models describe what a scene contains well, yet still struggle to reconstruct the 3D structure behind a 2D image and reason about it — spatial intelligence. SpatialBlock takes the route human spatial cognition develops along — block play. We train LVLMs on SpatialBlock-15k, a fully synthetic, scalable set of 15,000 block-stacking problems spanning 3D-to-2D projection, viewpoint transformation and structural combination, extended with controlled color cues so that models learn to anchor their reasoning on task-relevant blocks. Two training strategies are released: a direct model that predicts the answer immediately, and a reason model that thinks before answering. Despite training only on synthetic data at a small scale, both transfer to real-scene spatial benchmarks.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.