Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.</p>\n","updatedAt":"2026-08-11T11:54:36.218Z","author":{"_id":"663f0f6550dd1c97a40391a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663f0f6550dd1c97a40391a4/xWmJWfmAAR3gVi4pw3ulU.jpeg","fullname":"ZhuWenjie","name":"ZhuWenjie98","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8817774057388306},"editors":["ZhuWenjie98"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/663f0f6550dd1c97a40391a4/xWmJWfmAAR3gVi4pw3ulU.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27637","authors":[{"_id":"6a7b0d32b7183653340c120d","name":"Wenjie Zhu","hidden":false},{"_id":"6a7b0d32b7183653340c120e","name":"Yabin Zhang","hidden":false},{"_id":"6a7b0d32b7183653340c120f","name":"Wenjun Zeng","hidden":false},{"_id":"6a7b0d32b7183653340c1210","name":"Lei Zhang","hidden":false}],"publishedAt":"2026-08-01T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models","submittedOnDailyBy":{"_id":"663f0f6550dd1c97a40391a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663f0f6550dd1c97a40391a4/xWmJWfmAAR3gVi4pw3ulU.jpeg","isPro":false,"fullname":"ZhuWenjie","user":"ZhuWenjie98","type":"user","name":"ZhuWenjie98"},"summary":"Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.","upvotes":2,"discussionId":"6a7b0d32b7183653340c1211","projectPage":"https://zhuwenjie98.github.io/MMOOC-project-page/","githubRepo":"https://github.com/ZhuWenjie98/MMOOC","githubRepoAddedBy":"user","ai_summary":"MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities.","ai_keywords":["Multimodal Large Language Models","out-of-context","Shifted IC","refusal rate","LLM-as-a-Judge"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":11,"organization":{"_id":"69dba0c8dc88214a5ddca3f2","name":"VCLab-HKPU","fullname":"VCLab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/69db8a673114e93a4dbdec28/WgIv-pv3Mt2HBSdoR35eW.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"663f0f6550dd1c97a40391a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663f0f6550dd1c97a40391a4/xWmJWfmAAR3gVi4pw3ulU.jpeg","isPro":false,"fullname":"ZhuWenjie","user":"ZhuWenjie98","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69dba0c8dc88214a5ddca3f2","name":"VCLab-HKPU","fullname":"VCLab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/69db8a673114e93a4dbdec28/WgIv-pv3Mt2HBSdoR35eW.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27637.md","query":{}}">
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Abstract
MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities.
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.
Community
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.27637 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.27637 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.27637 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.