Hugging Face Daily Papers · · 8 min read

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.</p>\n","updatedAt":"2026-08-06T16:03:48.409Z","author":{"_id":"67ab7826ab5ebf181a7f78d7","avatarUrl":"/avatars/d6baf414011d6df659da4eb58e9d8958.svg","fullname":"Sirun Li","name":"inorganicwriter","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8773412108421326},"editors":["inorganicwriter"],"editorAvatarUrls":["/avatars/d6baf414011d6df659da4eb58e9d8958.svg"],"reactions":[],"isReport":false}},{"id":"6a7539fffbce822596d1c018","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-07T01:50:55.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models](https://huggingface.co/papers/2608.04509) (2026)\n* [No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs](https://huggingface.co/papers/2606.31933) (2026)\n* [Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety](https://huggingface.co/papers/2606.25034) (2026)\n* [Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection](https://huggingface.co/papers/2607.10329) (2026)\n* [Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning](https://huggingface.co/papers/2606.24548) (2026)\n* [Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG](https://huggingface.co/papers/2607.10798) (2026)\n* [Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs](https://huggingface.co/papers/2607.03647) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2608.04509\">CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.31933\">No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.25034\">Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.10329\">Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.24548\">Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.10798\">Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.03647\">Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-07T01:50:55.612Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7412463426589966},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04244","authors":[{"_id":"6a73d234c5e410d0768699b7","user":{"_id":"67ab7826ab5ebf181a7f78d7","avatarUrl":"/avatars/d6baf414011d6df659da4eb58e9d8958.svg","isPro":false,"fullname":"Sirun Li","user":"inorganicwriter","type":"user","name":"inorganicwriter"},"name":"Sirun Li","status":"claimed_verified","statusLastChangedAt":"2026-08-07T00:45:04.140Z","hidden":false},{"_id":"6a73d234c5e410d0768699b8","name":"Minghao Liu","hidden":false},{"_id":"6a73d234c5e410d0768699b9","name":"Ling Dai","hidden":false},{"_id":"6a73d234c5e410d0768699ba","name":"Yong Li","hidden":false},{"_id":"6a73d234c5e410d0768699bb","name":"Haoxin Lyu","hidden":false},{"_id":"6a73d234c5e410d0768699bc","name":"Junting Zhou","hidden":false},{"_id":"6a73d234c5e410d0768699bd","name":"Fan Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/67ab7826ab5ebf181a7f78d7/pmBH-CF1vncXyCmakT7Ml.png","https://cdn-uploads.huggingface.co/production/uploads/67ab7826ab5ebf181a7f78d7/rG_mDgKQJuAlIU7mAIQln.png","https://cdn-uploads.huggingface.co/production/uploads/67ab7826ab5ebf181a7f78d7/eTdZqxNRmTpJ-TiPEm-Ja.png","https://cdn-uploads.huggingface.co/production/uploads/67ab7826ab5ebf181a7f78d7/qIkusVxdHS2dvEaxFt0XK.png"],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models","submittedOnDailyBy":{"_id":"67ab7826ab5ebf181a7f78d7","avatarUrl":"/avatars/d6baf414011d6df659da4eb58e9d8958.svg","isPro":false,"fullname":"Sirun Li","user":"inorganicwriter","type":"user","name":"inorganicwriter"},"summary":"Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.","upvotes":2,"discussionId":"6a73d234c5e410d0768699be","githubRepo":"https://github.com/inorganicwriter/SIGNPOST-Bench","githubRepoAddedBy":"user","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67ab7826ab5ebf181a7f78d7","avatarUrl":"/avatars/d6baf414011d6df659da4eb58e9d8958.svg","isPro":false,"fullname":"Sirun Li","user":"inorganicwriter","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04244.md","query":{}}">
Papers
arxiv:2608.04244

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Published on Aug 4
· Submitted by
Sirun Li
on Aug 6
Authors:

Abstract

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

Community

Paper author Paper submitter about 10 hours ago

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04244
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.04244 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.04244 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers