Hugging Face Daily Papers · · 3 min read

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A strong step toward multimodal deep research agents. DeepVoyager-VL scales vision-language reasoning to long-horizon trajectories with iterative visual and textual search, showing how agents can better solve complex real-world information-seeking tasks.</p>\n","updatedAt":"2026-08-04T06:48:34.643Z","author":{"_id":"634d51f37f1e86614c10ac8d","avatarUrl":"/avatars/3fbdc62c12cd99d3f8bfdb2339769bc4.svg","fullname":"HalcyonZhang","name":"Halcyon-Zhang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.798725426197052},"editors":["Halcyon-Zhang"],"editorAvatarUrls":["/avatars/3fbdc62c12cd99d3f8bfdb2339769bc4.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.01827","authors":[{"_id":"6a718b0dec5082b9f872cf71","name":"Huanyao Zhang","hidden":false},{"_id":"6a718b0dec5082b9f872cf72","name":"Jiepeng Zhou","hidden":false},{"_id":"6a718b0dec5082b9f872cf73","name":"Runhao Zhao","hidden":false},{"_id":"6a718b0dec5082b9f872cf74","name":"Yanzhe Shan","hidden":false},{"_id":"6a718b0dec5082b9f872cf75","name":"Jiaoyang Chen","hidden":false},{"_id":"6a718b0dec5082b9f872cf76","name":"Bowen Zhou","hidden":false},{"_id":"6a718b0dec5082b9f872cf77","name":"Bo Li","hidden":false},{"_id":"6a718b0dec5082b9f872cf78","name":"Fang Wang","hidden":false},{"_id":"6a718b0dec5082b9f872cf79","name":"Jialong Wu","hidden":false},{"_id":"6a718b0dec5082b9f872cf7a","name":"Zhengwei Tao","hidden":false},{"_id":"6a718b0dec5082b9f872cf7b","name":"Lang Mei","hidden":false},{"_id":"6a718b0dec5082b9f872cf7c","name":"Xiaohan Yu","hidden":false},{"_id":"6a718b0dec5082b9f872cf7d","name":"Liyan Liu","hidden":false},{"_id":"6a718b0dec5082b9f872cf7e","name":"Chong Chen","hidden":false},{"_id":"6a718b0dec5082b9f872cf7f","name":"Wentao Zhang","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents","submittedOnDailyBy":{"_id":"634d51f37f1e86614c10ac8d","avatarUrl":"/avatars/3fbdc62c12cd99d3f8bfdb2339769bc4.svg","isPro":false,"fullname":"HalcyonZhang","user":"Halcyon-Zhang","type":"user","name":"Halcyon-Zhang"},"summary":"Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.","upvotes":6,"discussionId":"6a718b0dec5082b9f872cf80","projectPage":"https://halcyon-zhang.github.io/DeepVoyager-VL/","githubRepo":"https://github.com/Halcyon-Zhang/DeepVoyager-VL","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"634d51f37f1e86614c10ac8d","avatarUrl":"/avatars/3fbdc62c12cd99d3f8bfdb2339769bc4.svg","isPro":false,"fullname":"HalcyonZhang","user":"Halcyon-Zhang","type":"user"},{"_id":"69b3e409fb40adc124b78c8a","avatarUrl":"/avatars/52faa4bf481a1bedb7a87f07e12c93b2.svg","isPro":false,"fullname":"Bowen Zhou","user":"bwzzz","type":"user"},{"_id":"698d7c21eadd3d43e294af7a","avatarUrl":"/avatars/a2ebf5495cb42965d1aca0213bc0ea1a.svg","isPro":false,"fullname":"Zhou","user":"Jaypern","type":"user"},{"_id":"675667afb8dae5992d8c233c","avatarUrl":"/avatars/ac804d19d004ec6b3cf2aa53674732b2.svg","isPro":false,"fullname":"LiBo","user":"tracks256","type":"user"},{"_id":"6935815ad3572b84072a32f0","avatarUrl":"/avatars/d0e4249f1410be57b2682bd4a48a68f0.svg","isPro":false,"fullname":"arthur klein","user":"arthurklein","type":"user"},{"_id":"6a3b9be682db1fb1ba73e9f5","avatarUrl":"/avatars/002f6096eec0d9341e3d3ee3e75c75a2.svg","isPro":false,"fullname":"LIYAN LIU","user":"liyan0228","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.01827.md","query":{}}">
Papers
arxiv:2608.01827

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Published on Aug 3
· Submitted by
HalcyonZhang
on Aug 4
Authors:
,

Abstract

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.

Community

A strong step toward multimodal deep research agents. DeepVoyager-VL scales vision-language reasoning to long-horizon trajectories with iterative visual and textual search, showing how agents can better solve complex real-world information-seeking tasks.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.01827
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.01827 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.01827 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.01827 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers