We introduce <strong>VIABench</strong>, a comprehensive, time-aligned video benchmark for evaluating multimodal large language models (MLLMs) in real-world visual assistance scenarios for blind and visually impaired people.</p>\n<p>VIABench contains <strong>761 videos</strong>, <strong>14,526 manually curated annotations</strong>, and <strong>46.9 hours of footage</strong>. It covers three complementary tasks: <strong>Proactive Reminder</strong>, <strong>Visual Question Answering</strong>, and <strong>Vision-Guided Interaction</strong>. We also propose <strong>Token-Level Prompt Activation Decoding (TPAD)</strong>, a two-stage framework for evaluating proactive assistance in both online and offline settings.</p>\n<p>Our evaluation shows that current MLLMs still struggle with reliable real-world assistance, especially in anticipating navigation-critical events and responding in real time. We hope VIABench encourages progress toward safer and more useful visual assistants.</p>\n<p>Code and data: <a href=\"https://github.com/MCG-NJU/VIABench\" rel=\"nofollow\">https://github.com/MCG-NJU/VIABench</a></p>\n","updatedAt":"2026-07-17T07:39:06.595Z","author":{"_id":"660a7e1c3fbd33a1d0b0e233","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg","fullname":"Xiangyu Zeng","name":"Lanxingxuan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8469003438949585},"editors":["Lanxingxuan"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.14660","authors":[{"_id":"6a59983a6c2e371e6ca3810f","name":"Yunfeng Liu","hidden":false},{"_id":"6a59983a6c2e371e6ca38110","name":"Yuandong Yang","hidden":false},{"_id":"6a59983a6c2e371e6ca38111","name":"Jiarui Han","hidden":false},{"_id":"6a59983a6c2e371e6ca38112","name":"Zhenpeng Huang","hidden":false},{"_id":"6a59983a6c2e371e6ca38113","name":"Yuqing Tang","hidden":false},{"_id":"6a59983a6c2e371e6ca38114","name":"Xiangyu Zeng","hidden":false},{"_id":"6a59983a6c2e371e6ca38115","name":"Gangshan Wu","hidden":false},{"_id":"6a59983a6c2e371e6ca38116","name":"Limin Wang","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance","submittedOnDailyBy":{"_id":"660a7e1c3fbd33a1d0b0e233","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg","isPro":false,"fullname":"Xiangyu Zeng","user":"Lanxingxuan","type":"user","name":"Lanxingxuan"},"summary":"Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.","upvotes":6,"discussionId":"6a59983b6c2e371e6ca38117","githubRepo":"https://github.com/MCG-NJU/VIABench","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6314524a5f47a1896274d586","name":"NJU","fullname":"Nanjing University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662276136108-6314518e5f47a1896274d080.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"660a7e1c3fbd33a1d0b0e233","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660a7e1c3fbd33a1d0b0e233/nYfQdyTWdxenIxE-5s0L-.jpeg","isPro":false,"fullname":"Xiangyu Zeng","user":"Lanxingxuan","type":"user"},{"_id":"6a15ee7a785864c80e3bb4c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uzY_05mzVYiwWS3Ku7Jov.png","isPro":false,"fullname":"Robinson Lily","user":"lilyrobinson9","type":"user"},{"_id":"62c77f4352d8ae531f5511f9","avatarUrl":"/avatars/50198ccb02ccd286975a4613fbabee28.svg","isPro":false,"fullname":"Limin Wang","user":"lmwang","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"698301c931ae762ba620f33b","avatarUrl":"/avatars/39eb01bf3f7590061d9161a420e5b265.svg","isPro":false,"fullname":"Jean-Luc Moreau","user":"comment-king","type":"user"},{"_id":"670a9fe24d6a12dc2c6f0b02","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670a9fe24d6a12dc2c6f0b02/xdNoZRwUW8XiW7HqqtkRw.jpeg","isPro":false,"fullname":"Yuandong Yang","user":"yyd7","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6314524a5f47a1896274d586","name":"NJU","fullname":"Nanjing University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662276136108-6314518e5f47a1896274d080.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.14660.md","query":{}}">
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Abstract
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.
Community
We introduce VIABench, a comprehensive, time-aligned video benchmark for evaluating multimodal large language models (MLLMs) in real-world visual assistance scenarios for blind and visually impaired people.
VIABench contains 761 videos, 14,526 manually curated annotations, and 46.9 hours of footage. It covers three complementary tasks: Proactive Reminder, Visual Question Answering, and Vision-Guided Interaction. We also propose Token-Level Prompt Activation Decoding (TPAD), a two-stage framework for evaluating proactive assistance in both online and offline settings.
Our evaluation shows that current MLLMs still struggle with reliable real-world assistance, especially in anticipating navigation-critical events and responding in real time. We hope VIABench encourages progress toward safer and more useful visual assistants.
Code and data: https://github.com/MCG-NJU/VIABench
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.14660 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.14660 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.