EMNLP 2026 main</p>\n","updatedAt":"2026-08-24T05:29:07.003Z","author":{"_id":"66decf61f9971122eec44dc8","avatarUrl":"/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg","fullname":"Enjun Du","name":"EnjunDu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7493515610694885},"editors":["EnjunDu"],"editorAvatarUrls":["/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.20886","authors":[{"_id":"6a8bd5fe3d26296ea30919eb","name":"Enjun Du","hidden":false},{"_id":"6a8bd5fe3d26296ea30919ec","name":"Siyi Liu","hidden":false},{"_id":"6a8bd5fe3d26296ea30919ed","name":"Zirong Chen","hidden":false},{"_id":"6a8bd5fe3d26296ea30919ee","name":"Xinyu Zuo","hidden":false},{"_id":"6a8bd5fe3d26296ea30919ef","name":"Jinwen Luo","hidden":false},{"_id":"6a8bd5fe3d26296ea30919f0","name":"Ruiwen Tao","hidden":false},{"_id":"6a8bd5fe3d26296ea30919f1","name":"Lisheng Duan","hidden":false},{"_id":"6a8bd5fe3d26296ea30919f2","name":"Haijin Liang","hidden":false},{"_id":"6a8bd5fe3d26296ea30919f3","name":"Jin Ma","hidden":false},{"_id":"6a8bd5fe3d26296ea30919f4","name":"Junfu Pu","hidden":false},{"_id":"6a8bd5fe3d26296ea30919f5","name":"Yongqi Zhang","hidden":false}],"publishedAt":"2026-08-21T00:00:00.000Z","submittedOnDailyAt":"2026-08-24T00:00:00.000Z","title":"EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking","submittedOnDailyBy":{"_id":"66decf61f9971122eec44dc8","avatarUrl":"/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg","isPro":false,"fullname":"Enjun Du","user":"EnjunDu","type":"user","name":"EnjunDu"},"summary":"Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.","upvotes":8,"discussionId":"6a8bd5fe3d26296ea30919f6","ai_summary":"EviRank reformulates multimodal image re-ranking as semantic constraint satisfaction by parsing queries into structured evidence packages and verifying candidates via rubric scoring and listwise comparison without training.","ai_keywords":["multimodal image re-ranking","compositional queries","semantic constraint satisfaction","evidence package","rubric scoring","listwise comparison","training-free","distilled student"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66decf61f9971122eec44dc8","avatarUrl":"/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg","isPro":false,"fullname":"Enjun Du","user":"EnjunDu","type":"user"},{"_id":"66d05337e62d6bbf50186c2f","avatarUrl":"/avatars/f5be15e754f0fbbb37d2cc5ea417f729.svg","isPro":false,"fullname":"Yijie Jin","user":"DrewJin0827","type":"user"},{"_id":"668fcd97495674543600db0b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/JfBeJ6gSXe379G9ejjrLR.jpeg","isPro":false,"fullname":"Gao Mingqi","user":"EchoMinkki","type":"user"},{"_id":"65df4f5af3efe60b06ea1409","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df4f5af3efe60b06ea1409/u-9mnMU8bNXvyQ8z1CG5m.png","isPro":false,"fullname":"Xinye Li","user":"asdfo123","type":"user"},{"_id":"687f853bb39262ba84f3eeff","avatarUrl":"/avatars/cdfc44fde8237f08f10192553fe5a075.svg","isPro":false,"fullname":"Junhao Shen","user":"shenjunhao","type":"user"},{"_id":"67d62da597767f49259a55d6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3bsqzf_J0Kv9tZ3qvZsE8.png","isPro":false,"fullname":"Xiangnan Wu","user":"XiangnanW","type":"user"},{"_id":"66699aa8a33847217b5a49c7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/u8Z-6U8U7ARXOpdBDI7Qm.png","isPro":false,"fullname":"Weijie Wang","user":"lhmd","type":"user"},{"_id":"6971ec48ba9e9917b8861c15","avatarUrl":"/avatars/cc68fd9e0d4b696cbedea41b96f71ddf.svg","isPro":false,"fullname":"Haomin Wan","user":"Nek0laos","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.20886.md","query":{}}">
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Abstract
EviRank reformulates multimodal image re-ranking as semantic constraint satisfaction by parsing queries into structured evidence packages and verifying candidates via rubric scoring and listwise comparison without training.
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.20886 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.20886 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.20886 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.