Hugging Face Daily Papers · · 4 min read

An Exam for Active Observers

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Vision is a loop, not a glance. </p>\n<p>We introduce ActiveVision, a benchmark testing whether models can repeatedly observe, reason, and seek new visual evidence. </p>\n<p>Humans solve 96.1%. The best frontier model: 10.6%. Fable 5, strong at reasoning and coding, scores just 3.5%.</p>\n<p>🌐 Website: <a href=\"https://activevision.dev\" rel=\"nofollow\">https://activevision.dev</a><br>💻 Code: <a href=\"https://github.com/saccharomycetes/ActiveVision\" rel=\"nofollow\">https://github.com/saccharomycetes/ActiveVision</a><br>🤗 Dataset: <a href=\"https://huggingface.co/datasets/activevisionai/ActiveVision\">https://huggingface.co/datasets/activevisionai/ActiveVision</a><br>📄 Paper: <a href=\"https://arxiv.org/abs/2607.16165\" rel=\"nofollow\">https://arxiv.org/abs/2607.16165</a></p>\n","updatedAt":"2026-07-23T05:46:44.446Z","author":{"_id":"640eabb7612e7f36f910b144","avatarUrl":"/avatars/9a6cf54db7cbc7af9466d60486931adb.svg","fullname":"Muzi Tao","name":"muzitao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7729053497314453},"editors":["muzitao"],"editorAvatarUrls":["/avatars/9a6cf54db7cbc7af9466d60486931adb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.16165","authors":[{"_id":"6a5db7b46a69ce099f4d6e9f","user":{"_id":"635b99d47a1656011516bff9","avatarUrl":"/avatars/7243c4171ff127ba90631f105881d9d7.svg","isPro":false,"fullname":"jiarui zhang","user":"jrzhang","type":"user","name":"jrzhang"},"name":"Jiarui Zhang","status":"claimed_verified","statusLastChangedAt":"2026-07-23T00:45:04.462Z","hidden":false},{"_id":"6a5db7b46a69ce099f4d6ea0","user":{"_id":"640eabb7612e7f36f910b144","avatarUrl":"/avatars/9a6cf54db7cbc7af9466d60486931adb.svg","isPro":false,"fullname":"Muzi Tao","user":"muzitao","type":"user","name":"muzitao"},"name":"Muzi Tao","status":"claimed_verified","statusLastChangedAt":"2026-07-23T00:45:04.451Z","hidden":false},{"_id":"6a5db7b46a69ce099f4d6ea1","name":"Shangshang Wang","hidden":false},{"_id":"6a5db7b46a69ce099f4d6ea2","name":"Ollie Liu","hidden":false},{"_id":"6a5db7b46a69ce099f4d6ea3","name":"Xuezhe Ma","hidden":false},{"_id":"6a5db7b46a69ce099f4d6ea4","name":"Willie Neiswanger","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/640eabb7612e7f36f910b144/b2GJaLtxSZ-UKjd__iBRp.mp4"],"publishedAt":"2026-07-17T00:00:00.000Z","submittedOnDailyAt":"2026-07-23T00:00:00.000Z","title":"An Exam for Active Observers","submittedOnDailyBy":{"_id":"640eabb7612e7f36f910b144","avatarUrl":"/avatars/9a6cf54db7cbc7af9466d60486931adb.svg","isPro":false,"fullname":"Muzi Tao","user":"muzitao","type":"user","name":"muzitao"},"summary":"Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.","upvotes":12,"discussionId":"6a5db7b56a69ce099f4d6ea5","projectPage":"https://activevision.dev","githubRepo":"https://github.com/saccharomycetes/ActiveVision","githubRepoAddedBy":"user","githubStars":7,"organization":{"_id":"66a403d0dcb5bbc6e98bb7d0","name":"UniversityofSouthernCalifornia","fullname":"University of Southern California","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66a403728069e3c30e0d8524/tkYCfeIJfF1FxtYiRZ8bf.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"640eabb7612e7f36f910b144","avatarUrl":"/avatars/9a6cf54db7cbc7af9466d60486931adb.svg","isPro":false,"fullname":"Muzi Tao","user":"muzitao","type":"user"},{"_id":"69fc2ee78d9e57872583d9d7","avatarUrl":"/avatars/c17dbf1c9e7183ae53cfd3ad274a0d3e.svg","isPro":false,"fullname":"hpXgvFBl7ZxO","user":"activevision","type":"user"},{"_id":"67602605418a4b3626c7aa4a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/04oL5G2MqBMZp0hom2KIu.png","isPro":false,"fullname":"Alex","user":"yjp0418","type":"user"},{"_id":"6374cbb7255276f3a22b4b35","avatarUrl":"/avatars/7cf1bbb83447441e5fa2e1e4fcf7617b.svg","isPro":true,"fullname":"Peter Tong","user":"tsbpp","type":"user"},{"_id":"635b99d47a1656011516bff9","avatarUrl":"/avatars/7243c4171ff127ba90631f105881d9d7.svg","isPro":false,"fullname":"jiarui zhang","user":"jrzhang","type":"user"},{"_id":"65f2c1107a5fad429d5fc2f1","avatarUrl":"/avatars/21269e78fd71f20c149e96a215f8ce96.svg","isPro":false,"fullname":"Nuan Wen","user":"nkw1234","type":"user"},{"_id":"6624517c204cf7d22a90306e","avatarUrl":"/avatars/82da7b7194f1d9ad49e8e1a470cf02ce.svg","isPro":false,"fullname":"Xuezhe Ma","user":"maxma1987","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6756f0dfc8fb1673dc1b2def","avatarUrl":"/avatars/b35c374fefbd219409b873ab09618572.svg","isPro":false,"fullname":"Tim van Engeland","user":"TvE1997","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"64d98ef7a4839890b25eb78b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64d98ef7a4839890b25eb78b/215-CSVLl81z6CAq0ECWU.jpeg","isPro":true,"fullname":"Fangyuan Yu","user":"Ksgk-fy","type":"user"},{"_id":"6570450a78d7aca0c361a177","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6570450a78d7aca0c361a177/MX7jHhTQwLs-BvYIu5rqb.jpeg","isPro":false,"fullname":"Harold Chen","user":"Harold328","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66a403d0dcb5bbc6e98bb7d0","name":"UniversityofSouthernCalifornia","fullname":"University of Southern California","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66a403728069e3c30e0d8524/tkYCfeIJfF1FxtYiRZ8bf.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.16165.md","query":{}}">
Papers
arxiv:2607.16165

An Exam for Active Observers

Published on Jul 17
· Submitted by
Muzi Tao
on Jul 23
Authors:

Abstract

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.

Community

Paper author Paper submitter about 8 hours ago

Vision is a loop, not a glance.

We introduce ActiveVision, a benchmark testing whether models can repeatedly observe, reason, and seek new visual evidence.

Humans solve 96.1%. The best frontier model: 10.6%. Fable 5, strong at reasoning and coding, scores just 3.5%.

🌐 Website: https://activevision.dev
💻 Code: https://github.com/saccharomycetes/ActiveVision
🤗 Dataset: https://huggingface.co/datasets/activevisionai/ActiveVision
📄 Paper: https://arxiv.org/abs/2607.16165

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.16165
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.16165 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.16165 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers