Hugging Face Daily Papers · · 5 min read

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.</p>\n","updatedAt":"2026-08-07T02:01:37.062Z","author":{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","fullname":"Kaicheng Yang","name":"Kaichengalex","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":12,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.828650176525116},"editors":["Kaichengalex"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06060","authors":[{"_id":"6a753857e1228e04b3238101","name":"Zelong Sun","hidden":false},{"_id":"6a753857e1228e04b3238102","name":"Jun Wang","hidden":false},{"_id":"6a753857e1228e04b3238103","user":{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user","name":"Kaichengalex"},"name":"Kaicheng Yang","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.481Z","hidden":false},{"_id":"6a753857e1228e04b3238104","name":"Tiancheng Gu","hidden":false},{"_id":"6a753857e1228e04b3238105","name":"Ziyong Feng","hidden":false},{"_id":"6a753857e1228e04b3238106","name":"Zhiwu Lu","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval","submittedOnDailyBy":{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user","name":"Kaichengalex"},"summary":"Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.","upvotes":26,"discussionId":"6a753857e1228e04b3238107","githubRepo":"https://github.com/deepglint/UniME-R1","githubRepoAddedBy":"user","githubStars":7,"organization":{"_id":"6679441eb61cd819a77c1439","name":"DeepGlint-AI","fullname":"DeepGlint","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/655c70d331c4978366d4b2e6/dBwgYKdU8VUX6p_OcQqIe.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user"},{"_id":"641030c77a15af878ae5bd8f","avatarUrl":"/avatars/8a5037edf55c78ebc317c8b191343671.svg","isPro":false,"fullname":"TianchengGu","user":"TianchengGu","type":"user"},{"_id":"623d7b1c19b08016c234411d","avatarUrl":"/avatars/cbadc8e39e60ddd152c636c81fe2c409.svg","isPro":false,"fullname":"JunWang","user":"JunWangSpace","type":"user"},{"_id":"668f88eab1d23de66e314095","avatarUrl":"/avatars/80fe3bf0037ecfe2416d29b405508a21.svg","isPro":false,"fullname":"LEEGG","user":"GGLEE1","type":"user"},{"_id":"6a69ec931b577a27e51d8379","avatarUrl":"/avatars/d2226010b481746cf7609a641aea8b77.svg","isPro":false,"fullname":"David Anderson","user":"david-anderson-research","type":"user"},{"_id":"6a6a925739a7f0911330c546","avatarUrl":"/avatars/21568a0081650e63921f1f70775f8bdd.svg","isPro":false,"fullname":"Timothy Lee","user":"timothylee","type":"user"},{"_id":"6a6aa11e9859d6ac81f379dd","avatarUrl":"/avatars/677da06f7762a169543cf98dfecfdbd6.svg","isPro":false,"fullname":"Andrew Rodriguez","user":"Velvet-Andrew","type":"user"},{"_id":"6a6aa169c125cc860a93b9f7","avatarUrl":"/avatars/4d7fa3c3bb83274d7ef3d3066621b801.svg","isPro":false,"fullname":"Karen Gonzalez","user":"vectorridge","type":"user"},{"_id":"6a6c7d92b289c9e42b36715a","avatarUrl":"/avatars/cfeac81c781d73ff0db0523c551e16f9.svg","isPro":false,"fullname":"Edward Johnson","user":"Edward-Johnson","type":"user"},{"_id":"6a6c8374ed2d5f6d8078cdd1","avatarUrl":"/avatars/81df2448657fcb21486b8fdd739ed5c6.svg","isPro":false,"fullname":"Karen Anderson","user":"Karen-Anderson","type":"user"},{"_id":"6a6c84f892fd458eaa9c7938","avatarUrl":"/avatars/6390b9321a236bf70ff45cf3b85a0ccf.svg","isPro":false,"fullname":"Jennifer Williams","user":"velvetscope","type":"user"},{"_id":"6a6c8761ae8ae6f800d03aea","avatarUrl":"/avatars/ad694019e16574bdc1bb7c998ec175e7.svg","isPro":false,"fullname":"Jennifer Davis","user":"Atlas-Arc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6679441eb61cd819a77c1439","name":"DeepGlint-AI","fullname":"DeepGlint","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/655c70d331c4978366d4b2e6/dBwgYKdU8VUX6p_OcQqIe.jpeg"},"query":{}}">
Papers
arxiv:2608.06060

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Published on Aug 6
· Submitted by
Kaicheng Yang
on Aug 7
Authors:
,

Abstract

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.

Community

Paper author Paper submitter about 15 hours ago

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.06060 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers