Hugging Face Daily Papers · · 4 min read

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We introduce <strong>VIABench</strong>, a comprehensive, time-aligned video benchmark for evaluating multimodal large language models (MLLMs) in real-world visual assistance scenarios for blind and visually impaired people.</p>\n<p>VIABench contains <strong>761 videos</strong>, <strong>14,526 manually curated annotations</strong>, and <strong>46.9 hours of footage</strong>. It covers three complementary tasks: <strong>Proactive Reminder</strong>, <strong>Visual Question Answering</strong>, and <strong>Vision-Guided Interaction</strong>. We also propose <strong>Token-Level Prompt Activation Decoding (TPAD)</strong>, a two-stage framework for evaluating proactive assistance in both online and offline settings.</p>\n<p>Our evaluation shows that current MLLMs still struggle with reliable real-world assistance, especially in anticipating navigation-critical events and responding in real time. We hope VIABench encourages progress toward safer and more useful visual assistants.</p>\n<p>Code and data: <a href=\"https://github.com/MCG-NJU/VIABench\" rel=\"nofollow\">https://github.com/MCG-NJU/VIABench</a></p>\n","updatedAt":"2026-07-17T07:39:06.595Z","author":{"_id":"660a7e1c3fbd33a1d0b0e233","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg","fullname":"Xiangyu Zeng","name":"Lanxingxuan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8469003438949585},"editors":["Lanxingxuan"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.14660","authors":[{"_id":"6a59983a6c2e371e6ca3810f","name":"Yunfeng Liu","hidden":false},{"_id":"6a59983a6c2e371e6ca38110","name":"Yuandong Yang","hidden":false},{"_id":"6a59983a6c2e371e6ca38111","name":"Jiarui Han","hidden":false},{"_id":"6a59983a6c2e371e6ca38112","name":"Zhenpeng Huang","hidden":false},{"_id":"6a59983a6c2e371e6ca38113","name":"Yuqing Tang","hidden":false},{"_id":"6a59983a6c2e371e6ca38114","name":"Xiangyu Zeng","hidden":false},{"_id":"6a59983a6c2e371e6ca38115","name":"Gangshan Wu","hidden":false},{"_id":"6a59983a6c2e371e6ca38116","name":"Limin Wang","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance","submittedOnDailyBy":{"_id":"660a7e1c3fbd33a1d0b0e233","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg","isPro":false,"fullname":"Xiangyu Zeng","user":"Lanxingxuan","type":"user","name":"Lanxingxuan"},"summary":"Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.","upvotes":6,"discussionId":"6a59983b6c2e371e6ca38117","githubRepo":"https://github.com/MCG-NJU/VIABench","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6314524a5f47a1896274d586","name":"NJU","fullname":"Nanjing University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662276136108-6314518e5f47a1896274d080.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"660a7e1c3fbd33a1d0b0e233","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg","isPro":false,"fullname":"Xiangyu Zeng","user":"Lanxingxuan","type":"user"},{"_id":"6a15ee7a785864c80e3bb4c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uzY_05mzVYiwWS3Ku7Jov.png","isPro":false,"fullname":"Robinson Lily","user":"lilyrobinson9","type":"user"},{"_id":"62c77f4352d8ae531f5511f9","avatarUrl":"/avatars/50198ccb02ccd286975a4613fbabee28.svg","isPro":false,"fullname":"Limin Wang","user":"lmwang","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"698301c931ae762ba620f33b","avatarUrl":"/avatars/39eb01bf3f7590061d9161a420e5b265.svg","isPro":false,"fullname":"Jean-Luc Moreau","user":"comment-king","type":"user"},{"_id":"670a9fe24d6a12dc2c6f0b02","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670a9fe24d6a12dc2c6f0b02/xdNoZRwUW8XiW7HqqtkRw.jpeg","isPro":false,"fullname":"Yuandong Yang","user":"yyd7","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6314524a5f47a1896274d586","name":"NJU","fullname":"Nanjing University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662276136108-6314518e5f47a1896274d080.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.14660.md","query":{}}">
Papers
arxiv:2607.14660

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

Published on Jul 16
· Submitted by
Xiangyu Zeng
on Jul 17
Authors:
,

Abstract

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.

Community

Paper submitter about 7 hours ago

We introduce VIABench, a comprehensive, time-aligned video benchmark for evaluating multimodal large language models (MLLMs) in real-world visual assistance scenarios for blind and visually impaired people.

VIABench contains 761 videos, 14,526 manually curated annotations, and 46.9 hours of footage. It covers three complementary tasks: Proactive Reminder, Visual Question Answering, and Vision-Guided Interaction. We also propose Token-Level Prompt Activation Decoding (TPAD), a two-stage framework for evaluating proactive assistance in both online and offline settings.

Our evaluation shows that current MLLMs still struggle with reliable real-world assistance, especially in anticipating navigation-critical events and responding in real time. We hope VIABench encourages progress toward safer and more useful visual assistants.

Code and data: https://github.com/MCG-NJU/VIABench

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.14660
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.14660 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.14660 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers