Hugging Face Daily Papers · · 4 min read

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Large vision-language models describe what a scene contains well, yet still struggle to reconstruct the 3D structure behind a 2D image and reason about it — spatial intelligence. SpatialBlock takes the route human spatial cognition develops along — block play. We train LVLMs on SpatialBlock-15k, a fully synthetic, scalable set of 15,000 block-stacking problems spanning 3D-to-2D projection, viewpoint transformation and structural combination, extended with controlled color cues so that models learn to anchor their reasoning on task-relevant blocks. Two training strategies are released: a direct model that predicts the answer immediately, and a reason model that thinks before answering. Despite training only on synthetic data at a small scale, both transfer to real-scene spatial benchmarks.</p>\n","updatedAt":"2026-09-11T02:49:24.292Z","author":{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","fullname":"Soohyun Ryu","name":"rsoohyun","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9155384302139282},"editors":["rsoohyun"],"editorAvatarUrls":["/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07064","authors":[{"_id":"6aa21457a2aeb74440b1dd87","user":{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","isPro":false,"fullname":"Soohyun Ryu","user":"rsoohyun","type":"user","name":"rsoohyun"},"name":"Soohyun Ryu","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.222Z","hidden":false},{"_id":"6aa21457a2aeb74440b1dd88","user":{"_id":"6690a286181e2af45c742dd8","avatarUrl":"/avatars/511d0f86386e3b29a17b445d855b3aef.svg","isPro":false,"fullname":"Sohee Kim","user":"joyhee","type":"user","name":"joyhee"},"name":"Sohee Kim","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.229Z","hidden":false},{"_id":"6aa21457a2aeb74440b1dd89","name":"Eunho Yang","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem","submittedOnDailyBy":{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","isPro":false,"fullname":"Soohyun Ryu","user":"rsoohyun","type":"user","name":"rsoohyun"},"summary":"Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks. We introduce SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination. The dataset further incorporates controlled color modulation as visual cues to encourage anchor-based reasoning in visually complex conditions. Experiments demonstrate that LVLMs trained on our dataset through either direct answering or reasoning-based prediction significantly outperform baselines and generalize to real-world spatial tasks, despite the dataset's synthetic and compact nature. Code and data are available at https://github.com/rsoohyun/SpatialBlock.","upvotes":42,"discussionId":"6aa21458a2aeb74440b1dd8a","githubRepo":"https://github.com/rsoohyun/SpatialBlock","githubRepoAddedBy":"user","ai_summary":"Large vision-language models trained on synthetic block-manipulation tasks improve 3D spatial reasoning and generalize to real-world visual tasks.","ai_keywords":["Large Vision-Language Models","spatial intelligence","3D-to-2D projection","viewpoint transformation","structural combination","anchor-based reasoning","reasoning-based prediction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6690a286181e2af45c742dd8","avatarUrl":"/avatars/511d0f86386e3b29a17b445d855b3aef.svg","isPro":false,"fullname":"Sohee Kim","user":"joyhee","type":"user"},{"_id":"660beaac9b5015f91d6b4308","avatarUrl":"/avatars/38459c92f2bc1b95a3ee43ce0234c9c8.svg","isPro":false,"fullname":"Soohyun Ryu","user":"rsoohyun","type":"user"},{"_id":"66303ce3e1c93377db71efd5","avatarUrl":"/avatars/3d444dfb9799c9324c98cba893f4a10f.svg","isPro":false,"fullname":"Yoon Sik Park","user":"nooynoos","type":"user"},{"_id":"663073e4b108b8d71d4b6f32","avatarUrl":"/avatars/f66de2d81205397a29baf21169157bdd.svg","isPro":false,"fullname":"Inki Park","user":"seoharuss","type":"user"},{"_id":"62845957b410bd779033759c","avatarUrl":"/avatars/4feef73c06f2f7de6abf7a4789ac13f9.svg","isPro":false,"fullname":"Doohyuk Jang","user":"jadohu","type":"user"},{"_id":"657a64c91ccc3c2a5ea5cde4","avatarUrl":"/avatars/aee59f3650cc71424591340ec9f862b2.svg","isPro":false,"fullname":"Hangyeol Jung","user":"Hangyeol","type":"user"},{"_id":"66021a80561089b1a7ebfa01","avatarUrl":"/avatars/a61cf6f95fa0d6750297a00f897436dc.svg","isPro":false,"fullname":"JUNHYEOK CHOI","user":"junhyeokchoi","type":"user"},{"_id":"6666b61eaf95872a03a0a673","avatarUrl":"/avatars/fc0c144cf6307357d45d7ca2d6ba8d2f.svg","isPro":false,"fullname":"Joowon","user":"kjwispro","type":"user"},{"_id":"6371ce78789970f7bc673234","avatarUrl":"/avatars/ba363e0b8fddee143244934be7bc6db0.svg","isPro":false,"fullname":"Donghyeon Cho","user":"hyeon9698","type":"user"},{"_id":"62a4d58e81a4b10e93064ad6","avatarUrl":"/avatars/744d5cbc1745a26b816a458260aba050.svg","isPro":false,"fullname":"hangyulyoon","user":"hangyulmd","type":"user"},{"_id":"6aa36dde9d2d63a5b59fcc89","avatarUrl":"/avatars/e5424114acd61c3a510f13a589de6ee9.svg","isPro":false,"fullname":"Taeyong Choi","user":"tay-choi","type":"user"},{"_id":"664f558b6f16bbd9a1481b59","avatarUrl":"/avatars/581c88278e1d8bba66ee3b764f6dd3ed.svg","isPro":false,"fullname":"Jo sungmin","user":"JoJosmin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07064.md","query":{}}">
Papers
arxiv:2609.07064

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Published on Sep 7
· Submitted by
Soohyun Ryu
on Sep 11
#3 Paper of the day
Authors:

Abstract

Large vision-language models trained on synthetic block-manipulation tasks improve 3D spatial reasoning and generalize to real-world visual tasks.

Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks. We introduce SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination. The dataset further incorporates controlled color modulation as visual cues to encourage anchor-based reasoning in visually complex conditions. Experiments demonstrate that LVLMs trained on our dataset through either direct answering or reasoning-based prediction significantly outperform baselines and generalize to real-world spatial tasks, despite the dataset's synthetic and compact nature. Code and data are available at https://github.com/rsoohyun/SpatialBlock.

Community

Paper author Paper submitter about 11 hours ago

Large vision-language models describe what a scene contains well, yet still struggle to reconstruct the 3D structure behind a 2D image and reason about it — spatial intelligence. SpatialBlock takes the route human spatial cognition develops along — block play. We train LVLMs on SpatialBlock-15k, a fully synthetic, scalable set of 15,000 block-stacking problems spanning 3D-to-2D projection, viewpoint transformation and structural combination, extended with controlled color cues so that models learn to anchor their reasoning on task-relevant blocks. Two training strategies are released: a direct model that predicts the answer immediately, and a reason model that thinks before answering. Despite training only on synthetic data at a small scale, both transfer to real-scene spatial benchmarks.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.07064
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Browse 6 models citing this paper

Datasets citing this paper

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers