Hugging Face Daily Papers · · 3 min read

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Hi, we publish new embedding models which perform well on fine-grained multimodal retrieval and compositional reasoning tasks.</p>\n","updatedAt":"2026-09-04T02:18:34.011Z","author":{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","fullname":"Tingyu Song","name":"songtingyu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.917273223400116},"editors":["songtingyu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04083","authors":[{"_id":"6a9a278c8f7c3b7557239421","name":"Tingyu Song","hidden":false},{"_id":"6a9a278c8f7c3b7557239422","name":"Mingxin Li","hidden":false},{"_id":"6a9a278c8f7c3b7557239423","name":"Yanzhao Zhang","hidden":false},{"_id":"6a9a278c8f7c3b7557239424","name":"Dingkun Long","hidden":false},{"_id":"6a9a278c8f7c3b7557239425","name":"Chu Liu","hidden":false},{"_id":"6a9a278c8f7c3b7557239426","name":"Pengjun Xie","hidden":false},{"_id":"6a9a278c8f7c3b7557239427","name":"Yilun Zhao","hidden":false},{"_id":"6a9a278c8f7c3b7557239428","name":"Shu Wu","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation","submittedOnDailyBy":{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user","name":"songtingyu"},"summary":"MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.","upvotes":10,"discussionId":"6a9a278c8f7c3b7557239429","ai_summary":"CORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.","ai_keywords":["MLLM-based embedding models","compositional retrieval","cross-attentive reranker","Rank-KL objective","CoSENT","contrastive learning","compositional reasoning benchmarks"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"661f98de142a51d630dbbcc4","name":"Alibaba-NLP","fullname":"Alibaba-NLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63fc4c00a3c067e62899d32b/dfd_EcIfylvu3sdc2WMqX.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66af69222f4c59963afc874f","avatarUrl":"/avatars/034ca7688282bdbeddbd4f03e54dead7.svg","isPro":false,"fullname":"Zheyuan Yang","user":"Raywithyou","type":"user"},{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user"},{"_id":"66e258bdc70c02b46dfed6e3","avatarUrl":"/avatars/ccc2d604616c018f45a268a610472cac.svg","isPro":false,"fullname":"Yuzheng Cai","user":"Ucreate","type":"user"},{"_id":"6434c530a5aed21dd119a393","avatarUrl":"/avatars/a5aa3d8dc8b3b987e6de39280a4c0765.svg","isPro":false,"fullname":"Bro","user":"H34lthy","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6746b1e2224b22ef67fbff11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6746b1e2224b22ef67fbff11/q6sN8lQjtLGCgen_41VxU.jpeg","isPro":false,"fullname":"Zhuoning Guo","user":"Zhuoning","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"65dfeee3d16fb170031df293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg","isPro":false,"fullname":"gan","user":"guo9","type":"user"},{"_id":"683c642b02c1a474a867964e","avatarUrl":"/avatars/63e44a9cf788ee7b3ad236407700ceca.svg","isPro":false,"fullname":"Jinbiao Wei","user":"mikeweii","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"661f98de142a51d630dbbcc4","name":"Alibaba-NLP","fullname":"Alibaba-NLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63fc4c00a3c067e62899d32b/dfd_EcIfylvu3sdc2WMqX.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04083.md","query":{}}">
Papers
arxiv:2609.04083

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

Published on Sep 3
· Submitted by
Tingyu Song
on Sep 4
Authors:
,

Abstract

CORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.

MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.

Community

Paper submitter about 7 hours ago

Hi, we publish new embedding models which perform well on fine-grained multimodal retrieval and compositional reasoning tasks.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.04083
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.04083 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.04083 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers