Thinking-with-images models interleave reasoning with visual operations such as crop-and-zoom. Aggregate accuracy gains often look like successful tool-use, but many rollouts do not actually rely on the returned visual evidence. This work formulates visual tool-use as a causal graph and audits it with three levels of intervention.</p>\n","updatedAt":"2026-08-13T03:03:07.258Z","author":{"_id":"64c0af1d32dd2d752e6e4874","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c0af1d32dd2d752e6e4874/iGmXL0mJRdtbhUb-kJAa7.png","fullname":"Zhiheng Wang","name":"zhwang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9543794989585876},"editors":["zhwang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64c0af1d32dd2d752e6e4874/iGmXL0mJRdtbhUb-kJAa7.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06270","authors":[{"_id":"6a7c5a121653ef87c6af1e81","user":{"_id":"64c0af1d32dd2d752e6e4874","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c0af1d32dd2d752e6e4874/iGmXL0mJRdtbhUb-kJAa7.png","isPro":false,"fullname":"Zhiheng Wang","user":"zhwang","type":"user","name":"zhwang"},"name":"Zhiheng Wang","status":"claimed_verified","statusLastChangedAt":"2026-08-12T16:45:05.203Z","hidden":false},{"_id":"6a7c5a121653ef87c6af1e82","name":"Bo Peng","hidden":false},{"_id":"6a7c5a121653ef87c6af1e83","name":"Lai Wei","hidden":false},{"_id":"6a7c5a121653ef87c6af1e84","name":"Chaochao Lu","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images","submittedOnDailyBy":{"_id":"64c0af1d32dd2d752e6e4874","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c0af1d32dd2d752e6e4874/iGmXL0mJRdtbhUb-kJAa7.png","isPro":false,"fullname":"Zhiheng Wang","user":"zhwang","type":"user","name":"zhwang"},"summary":"The \"thinking-with-images\" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.","upvotes":6,"discussionId":"6a7c5a131653ef87c6af1e85","githubRepo":"https://github.com/OpenCausaLab/CauAudit","githubRepoAddedBy":"user","ai_summary":"Visual tool-use in multimodal LLMs often lacks causal effectiveness, with returned observations frequently failing to influence answers or being used incoherently despite aggregate accuracy improvements.","ai_keywords":["multimodal LLMs","visual tool-use","causal graph","observation-mediated paths","action-induced shortcuts","Visual Evidence Gain","policy miscalibration","Calling Without Looking","Looking Without Planning","illusion of visual tool-use"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64c0af1d32dd2d752e6e4874","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c0af1d32dd2d752e6e4874/iGmXL0mJRdtbhUb-kJAa7.png","isPro":false,"fullname":"Zhiheng Wang","user":"zhwang","type":"user"},{"_id":"65d5b967eeb590ea7435ad07","avatarUrl":"/avatars/ca0a5e123d5da5aca97cfd8a2d07e60e.svg","isPro":true,"fullname":"Kaican Li","user":"m-Just","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"63a65b9161348381efe8f1c2","avatarUrl":"/avatars/0417475f8e0825cb6225d74c71fd727a.svg","isPro":false,"fullname":"Bo Peng (SII)","user":"MagicalChair","type":"user"},{"_id":"64e9f5bbf494f8b2a0670eee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/rM0q9i3DXc7nVY6MDjiCY.jpeg","isPro":false,"fullname":"Rijusmit Biswas","user":"Phantomcloak19","type":"user"},{"_id":"635bd117dc371b8f910218f1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/635bd117dc371b8f910218f1/-kaEISLWY2ooL0ulJut4p.jpeg","isPro":false,"fullname":"Kaguya-19","user":"Kaguya-19","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06270.md","query":{}}">
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
Abstract
Visual tool-use in multimodal LLMs often lacks causal effectiveness, with returned observations frequently failing to influence answers or being used incoherently despite aggregate accuracy improvements.
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.
Community
Thinking-with-images models interleave reasoning with visual operations such as crop-and-zoom. Aggregate accuracy gains often look like successful tool-use, but many rollouts do not actually rely on the returned visual evidence. This work formulates visual tool-use as a causal graph and audits it with three levels of intervention.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.06270 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.06270 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.06270 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.