Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger.</p>\n","updatedAt":"2026-09-01T13:27:00.682Z","author":{"_id":"618c42e4f35289d4bcab9cf1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/618c42e4f35289d4bcab9cf1/wHmPhjRWUDdhyJMm8-jiJ.jpeg","fullname":"Yasmin Moslem","name":"ymoslem","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":27,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8543933033943176},"editors":["ymoslem"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/618c42e4f35289d4bcab9cf1/wHmPhjRWUDdhyJMm8-jiJ.jpeg"],"reactions":[],"isReport":false}},{"id":"6a977b34b9c98429e282c2b5","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-02T01:26:12.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models](https://huggingface.co/papers/2608.25662) (2026)\n* [Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models](https://huggingface.co/papers/2608.28058) (2026)\n* [UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations](https://huggingface.co/papers/2608.10835) (2026)\n* [Wiener Representation Filtering for VLM Hallucination Suppression](https://huggingface.co/papers/2608.08167) (2026)\n* [SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering](https://huggingface.co/papers/2607.04163) (2026)\n* [TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs](https://huggingface.co/papers/2608.05616) (2026)\n* [Hallucination Span Detection with Input-Side Evidence Alignment](https://huggingface.co/papers/2608.15804) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2608.25662\">Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.28058\">Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.10835\">UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.08167\">Wiener Representation Filtering for VLM Hallucination Suppression</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.04163\">SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.05616\">TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.15804\">Hallucination Span Detection with Input-Side Evidence Alignment</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-02T01:26:12.859Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7019367814064026},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.29974","authors":[{"_id":"6a96d25443140794820a7784","name":"Amanuel Gizachew Abebe","hidden":false},{"_id":"6a96d25443140794820a7785","name":"Yasmin Moslem","hidden":false}],"publishedAt":"2026-08-30T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models","submittedOnDailyBy":{"_id":"618c42e4f35289d4bcab9cf1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/618c42e4f35289d4bcab9cf1/wHmPhjRWUDdhyJMm8-jiJ.jpeg","isPro":true,"fullname":"Yasmin Moslem","user":"ymoslem","type":"user","name":"ymoslem"},"summary":"Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger. On the SHROOM-Visions English evaluation split, our ensemble achieves a Pearson calibration correlation of 0.41 and an overall IoU of 0.39, with a clean-response IoU of 0.91} and overall detection accuracy of 70.7%. We make our model weights and code publicly available.","upvotes":2,"discussionId":"6a96d25443140794820a7786","githubRepo":"https://github.com/Aman-byte1/Hallucination-Detection-in-LVLMs","githubRepoAddedBy":"user","ai_summary":"A hybrid system combining a multimodal sequence tagger and a generative vision-language model improves hallucination span detection and calibration through union-calibrated fusion.","ai_keywords":["Large Vision-Language Models","hallucination detection","span localization","sequence tagger","XLM-RoBERTa-Large","SigLIP","cross-attention","generative VLM","Qwen3.5-4B","Union-Calibrated Fusion","calibration correlation","IoU"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.29974.md","query":{}}">
SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
Abstract
A hybrid system combining a multimodal sequence tagger and a generative vision-language model improves hallucination span detection and calibration through union-calibrated fusion.
Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger. On the SHROOM-Visions English evaluation split, our ensemble achieves a Pearson calibration correlation of 0.41 and an overall IoU of 0.39, with a clean-response IoU of 0.91} and overall detection accuracy of 70.7%. We make our model weights and code publicly available.
Community
Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.29974 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.29974 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.29974 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.