<a href=\"https://cdn-uploads.huggingface.co/production/uploads/65fba64c3bd9e8f8b78ba8cd/9PhuaczeXAIowYL72m5HS.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65fba64c3bd9e8f8b78ba8cd/9PhuaczeXAIowYL72m5HS.png\" alt=\"Screenshot 2026-08-11 at 12.54.12 AM\"></a><br>This paper formulates grounding as <strong>bidirectional concept correspondence</strong> over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.</p>\n<p>To address this task, this paper introduces <strong>ConCor-1</strong>, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score.</p>\n","updatedAt":"2026-08-11T07:54:43.590Z","author":{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","fullname":"Ziqi Gao","name":"UWGZQ","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8448707461357117},"editors":["UWGZQ"],"editorAvatarUrls":["/avatars/2586171de43ec80bd6107b25a6f42b29.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07886","authors":[{"_id":"6a7ad353019ce76dc7b3abc3","user":{"_id":"640131b08ba76abe4b71b5d0","avatarUrl":"/avatars/2288b96a9a0ae8f584768f54e098def1.svg","isPro":false,"fullname":"Jieyu Zhang","user":"jieyuz2","type":"user","name":"jieyuz2"},"name":"Jieyu Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.575Z","hidden":false},{"_id":"6a7ad353019ce76dc7b3abc4","user":{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","isPro":false,"fullname":"Ziqi Gao","user":"UWGZQ","type":"user","name":"UWGZQ"},"name":"Ziqi Gao","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.567Z","hidden":false},{"_id":"6a7ad353019ce76dc7b3abc5","name":"Luke Zettlemoyer","hidden":false},{"_id":"6a7ad353019ce76dc7b3abc6","name":"Ranjay Krishna","hidden":false}],"publishedAt":"2026-08-08T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Vision-Language Grounding as Bidirectional Concept Correspondence","submittedOnDailyBy":{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","isPro":false,"fullname":"Ziqi Gao","user":"UWGZQ","type":"user","name":"UWGZQ"},"summary":"Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce ConCor-1, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that ConCor-1 consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.","upvotes":1,"discussionId":"6a7ad354019ce76dc7b3abc7","projectPage":"https://uwgzq.github.io/papers/ConCor-1/","githubRepo":"https://github.com/RAIVNLab/ConCor","githubRepoAddedBy":"user","ai_summary":"ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases.","ai_keywords":["vision-language grounding","bidirectional concept correspondence","bridge tokens","cross-modal alignment","phrase grounding","referring expression grounding","open-vocabulary detection","ConCor-1"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"670851790e79a8b46f716948","name":"raivn","fullname":"RAIVN: Reasoning, AI, & Vision Lab @ UW","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6345a11743f4f2d2ed1057ca/vkXVAuaU_UoJmQL_ExpoH.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","isPro":false,"fullname":"Ziqi Gao","user":"UWGZQ","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"670851790e79a8b46f716948","name":"raivn","fullname":"RAIVN: Reasoning, AI, & Vision Lab @ UW","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6345a11743f4f2d2ed1057ca/vkXVAuaU_UoJmQL_ExpoH.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07886.md","query":{}}">
Vision-Language Grounding as Bidirectional Concept Correspondence
Abstract
ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases.
Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce ConCor-1, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that ConCor-1 consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.
Community

This paper formulates grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.
To address this task, this paper introduces ConCor-1, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.