Hugging Face Daily Papers · · 3 min read

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We introduce SpatialCLI, a framework that teaches vision-language models to reason with spatial tools and internalize their capabilities for tool-free inference.</p>\n","updatedAt":"2026-07-31T01:49:37.955Z","author":{"_id":"670ff71f6b8de497472a81dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670ff71f6b8de497472a81dc/R2ydPAk6GU5Sk81Qja7E_.jpeg","fullname":"YANG ZHOU","name":"Yang-Zhou","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9144686460494995},"editors":["Yang-Zhou"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/670ff71f6b8de497472a81dc/R2ydPAk6GU5Sk81Qja7E_.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27703","authors":[{"_id":"6a6bfe2e7bd25d8874c07096","user":{"_id":"670ff71f6b8de497472a81dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670ff71f6b8de497472a81dc/R2ydPAk6GU5Sk81Qja7E_.jpeg","isPro":false,"fullname":"YANG ZHOU","user":"Yang-Zhou","type":"user","name":"Yang-Zhou"},"name":"Yang Zhou","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.405Z","hidden":false},{"_id":"6a6bfe2e7bd25d8874c07097","name":"Zixuan Huang","hidden":false},{"_id":"6a6bfe2e7bd25d8874c07098","user":{"_id":"641ba3407c21ab946bf5a1ff","avatarUrl":"/avatars/935d9196dcfc69f251c54de67928f42b.svg","isPro":false,"fullname":"Sunzhu Li","user":"sojuL","type":"user","name":"sojuL"},"name":"Sunzhu Li","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.413Z","hidden":false},{"_id":"6a6bfe2e7bd25d8874c07099","name":"Zhuo Yang","hidden":false},{"_id":"6a6bfe2e7bd25d8874c0709a","name":"Chen Zhang","hidden":false},{"_id":"6a6bfe2e7bd25d8874c0709b","user":{"_id":"623be9e1d1eb227788764959","avatarUrl":"/avatars/b6521b795a59754dbb40123fd4f63b8c.svg","isPro":false,"fullname":"Shunian Chen","user":"Shunian","type":"user","name":"Shunian"},"name":"Shunian Chen","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.419Z","hidden":false},{"_id":"6a6bfe2e7bd25d8874c0709c","name":"Caijun Yan","hidden":false},{"_id":"6a6bfe2e7bd25d8874c0709d","name":"Jianyao Xu","hidden":false},{"_id":"6a6bfe2e7bd25d8874c0709e","user":{"_id":"6713afea187a20dc579e121b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6713afea187a20dc579e121b/ELxVQLVF9ifuT-TCWPK22.jpeg","isPro":false,"fullname":"Shunyu Liu","user":"liushunyu","type":"user","name":"liushunyu"},"name":"Shunyu Liu","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.426Z","hidden":false},{"_id":"6a6bfe2e7bd25d8874c0709f","name":"Weijie Fu","hidden":false},{"_id":"6a6bfe2e7bd25d8874c070a0","name":"Peiliang Li","hidden":false},{"_id":"6a6bfe2e7bd25d8874c070a1","name":"Xiaozhi Chen","hidden":false},{"_id":"6a6bfe2e7bd25d8874c070a2","name":"Yuxiang Cai","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them","submittedOnDailyBy":{"_id":"670ff71f6b8de497472a81dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670ff71f6b8de497472a81dc/R2ydPAk6GU5Sk81Qja7E_.jpeg","isPro":false,"fullname":"YANG ZHOU","user":"Yang-Zhou","type":"user","name":"Yang-Zhou"},"summary":"Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.","upvotes":20,"discussionId":"6a6bfe2e7bd25d8874c070a3","githubRepo":"https://github.com/IANNXANG/SpatialCLI","githubRepoAddedBy":"user","githubStars":1},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"670ff71f6b8de497472a81dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670ff71f6b8de497472a81dc/R2ydPAk6GU5Sk81Qja7E_.jpeg","isPro":false,"fullname":"YANG ZHOU","user":"Yang-Zhou","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"68d9a8ace5ea71796005fa4b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d9a8ace5ea71796005fa4b/Hp9UWWX3BzZ2_MSCi3h3m.jpeg","isPro":false,"fullname":"Shubh","user":"shubhxho","type":"user"},{"_id":"6a688f08b1b7676c462495e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Fm4IteS5RgHzsOwBq-Z3Q.jpeg","isPro":false,"fullname":"Laurence Benoit","user":"mecalisse","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"68f99fce752b016555eedd62","avatarUrl":"/avatars/a56475cce8669c483ca2181f1fb870ae.svg","isPro":false,"fullname":"Yanglaodou","user":"Yanglaodou","type":"user"},{"_id":"641ba3407c21ab946bf5a1ff","avatarUrl":"/avatars/935d9196dcfc69f251c54de67928f42b.svg","isPro":false,"fullname":"Sunzhu Li","user":"sojuL","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64cc9e8fbb5d195b99459b27","avatarUrl":"/avatars/1648bb6b34315194d8627c9ac775caec.svg","isPro":false,"fullname":"Guiyong Zheng","user":"Versicles","type":"user"},{"_id":"6a5f749f674d10386c347af0","avatarUrl":"/avatars/6c41b4fec461bcffc2aee56b6ee5a59d.svg","isPro":false,"fullname":"Peiliang Li","user":"peiliangli","type":"user"},{"_id":"684161bddb77f5aa04c9e842","avatarUrl":"/avatars/acde945eca595d307358a4cd82a62772.svg","isPro":false,"fullname":"fu","user":"jamesfu2025","type":"user"},{"_id":"677f3181b5233456c1992bc1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/OGY5F3Z0-abmYx2hVicI_.png","isPro":false,"fullname":"Jason Hsu","user":"Jasonxxxyyy","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27703.md","query":{}}">
Papers
arxiv:2607.27703

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Published on Jul 30
· Submitted by
YANG ZHOU
on Jul 31

Abstract

Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.

Community

Paper author Paper submitter about 8 hours ago

We introduce SpatialCLI, a framework that teaches vision-language models to reason with spatial tools and internalize their capabilities for tool-free inference.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.27703
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

text answer)","featured":false}],"parentResourceAuthors":[{"name":"Yang-Zhou","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670ff71f6b8de497472a81dc/R2ydPAk6GU5Sk81Qja7E_.jpeg","type":"user"},{"name":"sojuL","avatarUrl":"/avatars/935d9196dcfc69f251c54de67928f42b.svg","type":"user"},{"name":"Shunian","avatarUrl":"/avatars/b6521b795a59754dbb40123fd4f63b8c.svg","type":"user"},{"name":"liushunyu","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6713afea187a20dc579e121b/ELxVQLVF9ifuT-TCWPK22.jpeg","type":"user"}]}">

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers