🎬 We introduce Video-DeepResearch (Video-DR) — a framework that redefines what a multimodal agent looks like when the input is not an image or a document, but the full<br> temporal stream of a video. The framework rests on three pillars: a decoupled perception→exploration paradigm with stage-wise tool unlocking (👁️ look first, 🌐 search<br> later) that structurally cures modality bias; a scalable video-grounded data engine producing multi-hop QA that cannot be shortcutted by parametric memory; and a two-stage<br> SFT → GRPO recipe that lets agents surpass their imitation ceiling and discover tool-use patterns rather than just copy them. The framework generalizes across model<br> scales and video categories — a shared blueprint for the next generation of vision-native research agents. 🔭 Our vision: streaming DR — know everything in vision. 🚀<br>code: <a href=\"https://github.com/Osilly/Vision-DeepResearch/tree/main/Video-DeepResearch\" rel=\"nofollow\">https://github.com/Osilly/Vision-DeepResearch/tree/main/Video-DeepResearch</a></p>\n","updatedAt":"2026-08-05T04:59:25.591Z","author":{"_id":"64b0a5037a475fba70a7260d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg","fullname":"Zhen Fang","name":"CostaliyA","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8142662048339844},"editors":["CostaliyA"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.03979","authors":[{"_id":"6a72c1e31a375f948521c4fa","name":"Zhen Fang","hidden":false},{"_id":"6a72c1e31a375f948521c4fb","name":"Yu Zeng","hidden":false},{"_id":"6a72c1e31a375f948521c4fc","name":"Wenxuan Huang","hidden":false},{"_id":"6a72c1e31a375f948521c4fd","name":"Yiming Zhao","hidden":false},{"_id":"6a72c1e31a375f948521c4fe","name":"Shiting Huang","hidden":false},{"_id":"6a72c1e31a375f948521c4ff","name":"Tianfei Ren","hidden":false},{"_id":"6a72c1e31a375f948521c500","name":"Qi Lu","hidden":false},{"_id":"6a72c1e31a375f948521c501","name":"Qingnan Ren","hidden":false},{"_id":"6a72c1e31a375f948521c502","name":"Qisheng Su","hidden":false},{"_id":"6a72c1e31a375f948521c503","name":"Lionel Z. Wang","hidden":false},{"_id":"6a72c1e31a375f948521c504","name":"Qingyu Yin","hidden":false},{"_id":"6a72c1e31a375f948521c505","name":"Shuang Chen","hidden":false},{"_id":"6a72c1e31a375f948521c506","name":"Zehui Chen","hidden":false},{"_id":"6a72c1e31a375f948521c507","name":"Lin Chen","hidden":false},{"_id":"6a72c1e31a375f948521c508","name":"Zhenfei Yin","hidden":false},{"_id":"6a72c1e31a375f948521c509","name":"Yao Hu","hidden":false},{"_id":"6a72c1e31a375f948521c50a","name":"Shaohui Lin","hidden":false},{"_id":"6a72c1e31a375f948521c50b","name":"Wanli Ouyang","hidden":false},{"_id":"6a72c1e31a375f948521c50c","name":"Shaosheng Cao","hidden":false},{"_id":"6a72c1e31a375f948521c50d","name":"Feng Zhao","hidden":false}],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent","submittedOnDailyBy":{"_id":"64b0a5037a475fba70a7260d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg","isPro":false,"fullname":"Zhen Fang","user":"CostaliyA","type":"user","name":"CostaliyA"},"summary":"We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.","upvotes":17,"discussionId":"6a72c1e41a375f948521c50e","projectPage":"https://costaliya.github.io/Video-DeepResearch/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b0a5037a475fba70a7260d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg","isPro":false,"fullname":"Zhen Fang","user":"CostaliyA","type":"user"},{"_id":"67d58d9e4d3c197e833ad92e","avatarUrl":"/avatars/244ec2218f8bd8cbc5a42cdb28c88eca.svg","isPro":false,"fullname":"Wei Yang","user":"wei-1-yang","type":"user"},{"_id":"66ae3fbf491b555fef3bac0c","avatarUrl":"/avatars/47353470d46097ce108d32792dbbf2a2.svg","isPro":false,"fullname":"Shiting Huang","user":"chocckaka","type":"user"},{"_id":"68a7d843ac4c72b9877b54b5","avatarUrl":"/avatars/7648ed4dbb4f3a209a8f62af7803b85f.svg","isPro":false,"fullname":"Kou Shi","user":"KouShi2","type":"user"},{"_id":"670a3bc3ada59c956f18cc17","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670a3bc3ada59c956f18cc17/57oBwS0V9m9SImYHtDb5f.jpeg","isPro":false,"fullname":"SII-sqs","user":"groundhogLLM","type":"user"},{"_id":"665d652e0f35c005de892108","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/665d652e0f35c005de892108/YxwwDwXSHFlVZJ0PCeUMZ.png","isPro":false,"fullname":"Yu Zeng","user":"YuZeng260","type":"user"},{"_id":"66cc19ad6f8945277c39cd86","avatarUrl":"/avatars/f031e77d7fb9dbebbc00fba9b5dd5357.svg","isPro":false,"fullname":"Kong","user":"csfufu","type":"user"},{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user"},{"_id":"65b44f8dd73977130cb7f80c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b44f8dd73977130cb7f80c/fna-FNtjS8hL1mJUs7MZu.jpeg","isPro":false,"fullname":"Guohui Zhang","user":"zghhui","type":"user"},{"_id":"661cf93620b47b0dad13ab89","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661cf93620b47b0dad13ab89/lwSsSabpkOEC9ADGehp64.png","isPro":false,"fullname":"Hao-Xuan Ma","user":"gh0stHunter","type":"user"},{"_id":"6901b520f8f20b9d7015a38d","avatarUrl":"/avatars/49e8d709818d1f0778758e39bd39f0a3.svg","isPro":false,"fullname":"zoushun","user":"shunzou1314","type":"user"},{"_id":"67dc162ec8c00778e8689f42","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67dc162ec8c00778e8689f42/_y_tO6W3ONOkOWbumAFXA.png","isPro":false,"fullname":"Wenxuan Huang","user":"Osilly","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.03979.md","query":{}}">
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Abstract
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.
Community
🎬 We introduce Video-DeepResearch (Video-DR) — a framework that redefines what a multimodal agent looks like when the input is not an image or a document, but the full
temporal stream of a video. The framework rests on three pillars: a decoupled perception→exploration paradigm with stage-wise tool unlocking (👁️ look first, 🌐 search
later) that structurally cures modality bias; a scalable video-grounded data engine producing multi-hop QA that cannot be shortcutted by parametric memory; and a two-stage
SFT → GRPO recipe that lets agents surpass their imitation ceiling and discover tool-use patterns rather than just copy them. The framework generalizes across model
scales and video categories — a shared blueprint for the next generation of vision-native research agents. 🔭 Our vision: streaming DR — know everything in vision. 🚀
code: https://github.com/Osilly/Vision-DeepResearch/tree/main/Video-DeepResearch
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.03979 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.03979 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.03979 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.