Hugging Face Daily Papers · · 3 min read

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

HF: <a href=\"https://huggingface.co/datasets/matsant01/blind-spots-bench\">https://huggingface.co/datasets/matsant01/blind-spots-bench</a><br>GitHub: <a href=\"https://github.com/matteosantelmo/reasoning-blind-spots\" rel=\"nofollow\">https://github.com/matteosantelmo/reasoning-blind-spots</a></p>\n","updatedAt":"2026-07-15T13:20:53.948Z","author":{"_id":"64623812d9b49021880ebab1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64623812d9b49021880ebab1/iCGpUSxe5LmsuDTMn2LU_.jpeg","fullname":"Chengkun Li","name":"chengkunli","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.9145384430885315},"editors":["chengkunli"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64623812d9b49021880ebab1/iCGpUSxe5LmsuDTMn2LU_.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.08317","authors":[{"_id":"6a568e15e548eb96f98cc9b2","name":"Matteo Santelmo","hidden":false},{"_id":"6a568e15e548eb96f98cc9b3","name":"Xiuying Wei","hidden":false},{"_id":"6a568e15e548eb96f98cc9b4","name":"Israa Fakih","hidden":false},{"_id":"6a568e15e548eb96f98cc9b5","name":"Felix Bauer","hidden":false},{"_id":"6a568e15e548eb96f98cc9b6","name":"Juan Garcia Giraldo","hidden":false},{"_id":"6a568e15e548eb96f98cc9b7","name":"Chengkun Li","hidden":false},{"_id":"6a568e15e548eb96f98cc9b8","name":"Etienne Bamas","hidden":false},{"_id":"6a568e15e548eb96f98cc9b9","name":"Emmanuel Abbé","hidden":false}],"publishedAt":"2026-07-09T00:00:00.000Z","submittedOnDailyAt":"2026-07-15T00:00:00.000Z","title":"Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models","submittedOnDailyBy":{"_id":"64623812d9b49021880ebab1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64623812d9b49021880ebab1/iCGpUSxe5LmsuDTMn2LU_.jpeg","isPro":false,"fullname":"Chengkun Li","user":"chengkunli","type":"user","name":"chengkunli"},"summary":"Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce blind-spots-bench, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on blind-spots-bench reveals that closed-source frontier models can substantially outperform open-weight models with even approx10% gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of blind-spots-bench as a diagnostic stress test for identifying concrete weaknesses in current modern models.","upvotes":17,"discussionId":"6a568e16e548eb96f98cc9ba","githubRepo":"https://github.com/matteosantelmo/reasoning-blind-spots","githubRepoAddedBy":"user","githubStars":3,"organization":{"_id":"6a576b8869ba8959b34804a4","name":"epfl-ch","fullname":"EPFL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64623812d9b49021880ebab1/W0HJnlYYzjvwRFn8hDC5h.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"654df8932fdbbde41e809968","avatarUrl":"/avatars/4d43b91387428f4c267fe248039cfad5.svg","isPro":false,"fullname":"XiuyingWei","user":"barpitf","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"663b7d5873144935ec6516c4","avatarUrl":"/avatars/4ec89e6be3a3f5d5f743315cce33a9a9.svg","isPro":false,"fullname":"Juan Garcia Giraldo","user":"d23845jg","type":"user"},{"_id":"6556cff80e7a7067a934445f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6556cff80e7a7067a934445f/PoT7qQ6tVqLGGrYvGBbr8.jpeg","isPro":false,"fullname":"Valentina Pyatkin","user":"valpy","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"65c93ce8567e810c57cf41b3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65c93ce8567e810c57cf41b3/1bftRk56lFUEKBjypJaCt.png","isPro":false,"fullname":"Davit Melikidze","user":"davmel","type":"user"},{"_id":"658de04877105e6e408ca2ce","avatarUrl":"/avatars/6ae9582c586cea3f3377420440b6a955.svg","isPro":false,"fullname":"Jason O","user":"JayO-1","type":"user"},{"_id":"6616f2c02aadf440a229b037","avatarUrl":"/avatars/42d96af1a0cd7697e5e7e38eb4511bae.svg","isPro":false,"fullname":"Elena Lyulina","user":"elenal","type":"user"},{"_id":"650ed7adf141bc34f91a12ae","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/rZDwRiBcqeqlUQ7mThoyU.jpeg","isPro":true,"fullname":"Hanna Yukhymenko","user":"hannayukhymenko","type":"user"},{"_id":"633467e031a2be3938c35ba0","avatarUrl":"/avatars/b138102b22617e417811c73c6182eb3e.svg","isPro":false,"fullname":"William Li","user":"KitKatL","type":"user"},{"_id":"61b3576de49318df54457d8f","avatarUrl":"/avatars/5d425bea6ee6319d413e965b4499ec5c.svg","isPro":false,"fullname":"Biziel","user":"Grzegorz","type":"user"},{"_id":"6a579741e6a4bde191e98cfd","avatarUrl":"/avatars/b233c63e782639defbbec61f4e75a213.svg","isPro":false,"fullname":"Susana Cardoso","user":"susanacardoso","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6a576b8869ba8959b34804a4","name":"epfl-ch","fullname":"EPFL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64623812d9b49021880ebab1/W0HJnlYYzjvwRFn8hDC5h.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.08317.md","query":{}}">
Papers
arxiv:2607.08317

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Published on Jul 9
· Submitted by
Chengkun Li
on Jul 15
#3 Paper of the day
Authors:
,

Abstract

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce blind-spots-bench, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on blind-spots-bench reveals that closed-source frontier models can substantially outperform open-weight models with even approx10% gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of blind-spots-bench as a diagnostic stress test for identifying concrete weaknesses in current modern models.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.08317
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.08317 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.08317 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers