We propose a geometry-consistency-aware framework for video spatial reasoning.</p>\n","updatedAt":"2026-07-22T02:24:19.456Z","author":{"_id":"675aa0331b310ed01068857f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PtTMPhuHlfdui_CJOjgHV.png","fullname":"HuangTing","name":"Believe0029","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8645511865615845},"editors":["Believe0029"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PtTMPhuHlfdui_CJOjgHV.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.17599","authors":[{"_id":"6a60296f7e7f152167e470a4","user":{"_id":"675aa0331b310ed01068857f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PtTMPhuHlfdui_CJOjgHV.png","isPro":false,"fullname":"HuangTing","user":"Believe0029","type":"user","name":"Believe0029"},"name":"Ting Huang","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.292Z","hidden":false},{"_id":"6a60296f7e7f152167e470a5","name":"Zhenyu Zhang","hidden":false},{"_id":"6a60296f7e7f152167e470a6","name":"Wenyuan Huang","hidden":false},{"_id":"6a60296f7e7f152167e470a7","name":"Jian Yang","hidden":false},{"_id":"6a60296f7e7f152167e470a8","name":"Hao Tang","hidden":false}],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning","submittedOnDailyBy":{"_id":"675aa0331b310ed01068857f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PtTMPhuHlfdui_CJOjgHV.png","isPro":false,"fullname":"HuangTing","user":"Believe0029","type":"user","name":"Believe0029"},"summary":"Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.","upvotes":0,"discussionId":"6a60296f7e7f152167e470a9","projectPage":"https://believeht029.github.io/ConsiSpace/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.17599.md","query":{}}">
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Abstract
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.
Community
We propose a geometry-consistency-aware framework for video spatial reasoning.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.17599 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.17599 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.17599 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.