Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- and region-level evidence that future query tokens may need to select. We propose SaMer, an object-aware token merging framework that compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. SaMer uses object annotations only during training as a merge prior to discourage cross-instance mixing, requires no ground-truth bounding boxes or detectors at inference time, and adapts only the shared projection layer with frozen vision and language backbones. With K=64, SaMer removes more than 93% of image-side tokens and reduces ColPali storage by 16.09×, while improving R@1 on Flickr30K and MSCOCO. These gains arise because object-aware merging preserves query-selectable object evidence that pruning or feature-only pooling can remove or collapse. SaMer also outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.</p>\n","updatedAt":"2026-07-07T06:24:32.789Z","author":{"_id":"67bc90b961b6284817237bda","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ucJr5q-ub4RDkswy2Kw9Y.png","fullname":"Junha Jung","name":"JunhaJung","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8554505109786987},"editors":["JunhaJung"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ucJr5q-ub4RDkswy2Kw9Y.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04605","authors":[{"_id":"6a4c668b25849b193a834081","user":{"_id":"63f1de31f4e30ffd2bcd626b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f1de31f4e30ffd2bcd626b/aPHgcUj0NN68_fIEym3HE.jpeg","isPro":false,"fullname":"Suhyeong Park","user":"Codingchild","type":"user","name":"Codingchild"},"name":"Suhyeong Park","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:12:07.120Z","hidden":false},{"_id":"6a4c668b25849b193a834082","user":{"_id":"67bc90b961b6284817237bda","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ucJr5q-ub4RDkswy2Kw9Y.png","isPro":false,"fullname":"Junha Jung","user":"JunhaJung","type":"user","name":"JunhaJung"},"name":"Junha Jung","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:12:03.525Z","hidden":false},{"_id":"6a4c668b25849b193a834083","name":"Jungwoo Park","hidden":false},{"_id":"6a4c668b25849b193a834084","name":"Jaewoo Kang","hidden":false}],"publishedAt":"2026-07-06T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval","submittedOnDailyBy":{"_id":"67bc90b961b6284817237bda","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ucJr5q-ub4RDkswy2Kw9Y.png","isPro":false,"fullname":"Junha Jung","user":"JunhaJung","type":"user","name":"JunhaJung"},"summary":"Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- and region-level evidence that future query tokens may need to select. We propose SaMer, an object-aware token merging framework that compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. SaMer uses object annotations only during training as a merge prior to discourage cross-instance mixing, requires no ground-truth bounding boxes or detectors at inference time, and adapts only the shared projection layer with frozen vision and language backbones. With K=64, SaMer removes more than 93% of image-side tokens and reduces ColPali storage by 16.09times, while improving R@1 on Flickr30K and MSCOCO. These gains arise because object-aware merging preserves query-selectable object evidence that pruning or feature-only pooling can remove or collapse. SaMer also outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.","upvotes":15,"discussionId":"6a4c668b25849b193a834085","projectPage":"https://suhyeong10.github.io/samer-project-page/","githubRepo":"https://github.com/dmis-lab/SaMer","githubRepoAddedBy":"user","ai_summary":"Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retrieval performance.","ai_keywords":["vision-language retrieval","late interaction","token compression","object-aware merging","post-projector tokens","centroid compression","phrase-level grounding","multi-vector retrieval"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63f1de31f4e30ffd2bcd626b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f1de31f4e30ffd2bcd626b/aPHgcUj0NN68_fIEym3HE.jpeg","isPro":false,"fullname":"Suhyeong Park","user":"Codingchild","type":"user"},{"_id":"66d914d4c39b38d37fd46989","avatarUrl":"/avatars/8acf2b96957efe5c23119f97c54d06eb.svg","isPro":false,"fullname":"Dongyoung Lee","user":"GBEdge","type":"user"},{"_id":"67bc90b961b6284817237bda","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ucJr5q-ub4RDkswy2Kw9Y.png","isPro":false,"fullname":"Junha Jung","user":"JunhaJung","type":"user"},{"_id":"68121036912c2103dd3ba6cd","avatarUrl":"/avatars/2ac7eacb486bf1993f7465299282cd3c.svg","isPro":false,"fullname":"Taeyun Roh","user":"txxnrd","type":"user"},{"_id":"670f5c3f642f58673b1f435a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/wa_TulGokgVzMN0-_GlBB.png","isPro":false,"fullname":"YewonCho","user":"doldol330","type":"user"},{"_id":"68c057fe41131fad0b83ba3a","avatarUrl":"/avatars/de0abe8c7890d6ac6984f88a64a415df.svg","isPro":false,"fullname":"yoonki yoon","user":"dbsrl0218","type":"user"},{"_id":"5efbdc4ac3896117eab961a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602668910270-5efbdc4ac3896117eab961a9.png","isPro":false,"fullname":"Data Mining and Information Systems Lab","user":"dmis-lab","type":"user"},{"_id":"69bcf9ee903835978b371172","avatarUrl":"/avatars/6fd9e48ee186fab882aa79ddd3566a51.svg","isPro":false,"fullname":"Yumin Seol","user":"yumin-seol","type":"user"},{"_id":"63b62561bda8d44adf38e043","avatarUrl":"/avatars/d16b41d7c30dc3eb4da9c96bbf451a35.svg","isPro":false,"fullname":"Yumin Seol","user":"ymnseol","type":"user"},{"_id":"618a925429415f4b68c5ee78","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671626774917-618a925429415f4b68c5ee78.png","isPro":false,"fullname":"Chanran Kim","user":"seriousran","type":"user"},{"_id":"64b0eabd6c5beff5c5a9a713","avatarUrl":"/avatars/261fc032201ca7dde45cb5bea13e7e3d.svg","isPro":false,"fullname":"choseeun","user":"seny1004","type":"user"},{"_id":"6448d979de60e26710a6b692","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448d979de60e26710a6b692/hl8-JAC5AIn2JPtuyGgDb.jpeg","isPro":true,"fullname":"Woojun Jung","user":"woojun-jung","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04605.md","query":{}}">
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Abstract
Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retrieval performance.
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- and region-level evidence that future query tokens may need to select. We propose SaMer, an object-aware token merging framework that compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. SaMer uses object annotations only during training as a merge prior to discourage cross-instance mixing, requires no ground-truth bounding boxes or detectors at inference time, and adapts only the shared projection layer with frozen vision and language backbones. With K=64, SaMer removes more than 93% of image-side tokens and reduces ColPali storage by 16.09times, while improving R@1 on Flickr30K and MSCOCO. These gains arise because object-aware merging preserves query-selectable object evidence that pruning or feature-only pooling can remove or collapse. SaMer also outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.
Community
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- and region-level evidence that future query tokens may need to select. We propose SaMer, an object-aware token merging framework that compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. SaMer uses object annotations only during training as a merge prior to discourage cross-instance mixing, requires no ground-truth bounding boxes or detectors at inference time, and adapts only the shared projection layer with frozen vision and language backbones. With K=64, SaMer removes more than 93% of image-side tokens and reduces ColPali storage by 16.09×, while improving R@1 on Flickr30K and MSCOCO. These gains arise because object-aware merging preserves query-selectable object evidence that pruning or feature-only pooling can remove or collapse. SaMer also outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.04605 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.04605 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.04605 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.