Hugging Face Daily Papers · · 5 min read

Vision-Language Grounding as Bidirectional Concept Correspondence

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/65fba64c3bd9e8f8b78ba8cd/9PhuaczeXAIowYL72m5HS.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65fba64c3bd9e8f8b78ba8cd/9PhuaczeXAIowYL72m5HS.png\" alt=\"Screenshot 2026-08-11 at 12.54.12 AM\"></a><br>This paper formulates grounding as <strong>bidirectional concept correspondence</strong> over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.</p>\n<p>To address this task, this paper introduces <strong>ConCor-1</strong>, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score.</p>\n","updatedAt":"2026-08-11T07:54:43.590Z","author":{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","fullname":"Ziqi Gao","name":"UWGZQ","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8448707461357117},"editors":["UWGZQ"],"editorAvatarUrls":["/avatars/2586171de43ec80bd6107b25a6f42b29.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07886","authors":[{"_id":"6a7ad353019ce76dc7b3abc3","user":{"_id":"640131b08ba76abe4b71b5d0","avatarUrl":"/avatars/2288b96a9a0ae8f584768f54e098def1.svg","isPro":false,"fullname":"Jieyu Zhang","user":"jieyuz2","type":"user","name":"jieyuz2"},"name":"Jieyu Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.575Z","hidden":false},{"_id":"6a7ad353019ce76dc7b3abc4","user":{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","isPro":false,"fullname":"Ziqi Gao","user":"UWGZQ","type":"user","name":"UWGZQ"},"name":"Ziqi Gao","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.567Z","hidden":false},{"_id":"6a7ad353019ce76dc7b3abc5","name":"Luke Zettlemoyer","hidden":false},{"_id":"6a7ad353019ce76dc7b3abc6","name":"Ranjay Krishna","hidden":false}],"publishedAt":"2026-08-08T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Vision-Language Grounding as Bidirectional Concept Correspondence","submittedOnDailyBy":{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","isPro":false,"fullname":"Ziqi Gao","user":"UWGZQ","type":"user","name":"UWGZQ"},"summary":"Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce ConCor-1, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that ConCor-1 consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.","upvotes":1,"discussionId":"6a7ad354019ce76dc7b3abc7","projectPage":"https://uwgzq.github.io/papers/ConCor-1/","githubRepo":"https://github.com/RAIVNLab/ConCor","githubRepoAddedBy":"user","ai_summary":"ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases.","ai_keywords":["vision-language grounding","bidirectional concept correspondence","bridge tokens","cross-modal alignment","phrase grounding","referring expression grounding","open-vocabulary detection","ConCor-1"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"670851790e79a8b46f716948","name":"raivn","fullname":"RAIVN: Reasoning, AI, & Vision Lab @ UW","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6345a11743f4f2d2ed1057ca/vkXVAuaU_UoJmQL_ExpoH.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65fba64c3bd9e8f8b78ba8cd","avatarUrl":"/avatars/2586171de43ec80bd6107b25a6f42b29.svg","isPro":false,"fullname":"Ziqi Gao","user":"UWGZQ","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"670851790e79a8b46f716948","name":"raivn","fullname":"RAIVN: Reasoning, AI, & Vision Lab @ UW","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6345a11743f4f2d2ed1057ca/vkXVAuaU_UoJmQL_ExpoH.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07886.md","query":{}}">
Papers
arxiv:2608.07886

Vision-Language Grounding as Bidirectional Concept Correspondence

Published on Aug 8
· Submitted by
Ziqi Gao
on Aug 11
Authors:

Abstract

ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases.

Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce ConCor-1, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that ConCor-1 consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.

Community

Paper author Paper submitter about 11 hours ago

Screenshot 2026-08-11 at 12.54.12 AM
This paper formulates grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.

To address this task, this paper introduces ConCor-1, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.07886
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers