Please check out our project page at: <a href=\"https://huggingface.co/spaces/elaine1wan/vdiff-bench\">https://huggingface.co/spaces/elaine1wan/vdiff-bench</a><br>Cooking more interesting stuff! Stay tuned ;)</p>\n","updatedAt":"2026-09-09T18:32:04.558Z","author":{"_id":"645eba55b5c9a8666d0f36d6","avatarUrl":"/avatars/9fa64a29ba99bf7df540c40342969b95.svg","fullname":"Elaine Wan","name":"elaine1wan","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8904637098312378},"editors":["elaine1wan"],"editorAvatarUrls":["/avatars/9fa64a29ba99bf7df540c40342969b95.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06245","authors":[{"_id":"6aa1a524a2aeb74440b1dcb6","name":"Yixin Wan","hidden":false},{"_id":"6aa1a524a2aeb74440b1dcb7","name":"Tianle Zheng","hidden":false},{"_id":"6aa1a524a2aeb74440b1dcb8","name":"Kai-Wei Chang","hidden":false}],"publishedAt":"2026-09-05T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification","submittedOnDailyBy":{"_id":"645eba55b5c9a8666d0f36d6","avatarUrl":"/avatars/9fa64a29ba99bf7df540c40342969b95.svg","isPro":true,"fullname":"Elaine Wan","user":"elaine1wan","type":"user","name":"elaine1wan"},"summary":"Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, and illumination. Each question corresponds to two image inputs with 4 choices: the true difference, two hard negative descriptions, and a \"no difference\" distractor. To make the task challenging, we specifically curate ground-truth-conditioned negatives that require models to distinguish the actual change from nearby semantic alternatives. Experiments with 11 state-of-the-art open- and closed-source MLLMs show that fine-grained visual comparison remains brittle: models exhibit uneven performance across sources and change categories, with persistent failures on subtle low-level changes like noises and textures. For instance, three 7-8B-scale open-source MLLMs score 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level changes like noise and texture, falsely assuming no changes between two image inputs. Surprisingly, despite strong performance of other closed-source commercial models, Grok 4.3 demonstrate remarkable performance drop on identifying noise and texture differences between images, falling significantly behind large open-source models like Kimi K2.5 and K3. Overall, VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.","upvotes":9,"discussionId":"6aa1a525a2aeb74440b1dcb9","projectPage":"https://huggingface.co/spaces/elaine1wan/vdiff-bench","ai_summary":"VDiff-Bench evaluates multimodal language models on fine-grained image difference identification, revealing major weaknesses in detecting subtle low-level visual changes.","ai_keywords":["Multimodal Large Language Models","visual question answering","image difference identification","VDiff-Bench","hard negative descriptions","low-level changes","comparative visual understanding"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67784c39dac147922d8d09f0","name":"UCLA","fullname":"University of California, Los Angeles","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67784bd637dfa531fbce95a2/Nf0seEMEn66sPL3QsJXj4.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"645eba55b5c9a8666d0f36d6","avatarUrl":"/avatars/9fa64a29ba99bf7df540c40342969b95.svg","isPro":true,"fullname":"Elaine Wan","user":"elaine1wan","type":"user"},{"_id":"693b997f70c627f41b040637","avatarUrl":"/avatars/3b2663e21d9e5db4b9a4ba9c1ed935c7.svg","isPro":false,"fullname":"Chen","user":"schen8maomao","type":"user"},{"_id":"63f37af60be81bdc5d92eebb","avatarUrl":"/avatars/b8dfdff4ab36988ec9a8643e82a3d2db.svg","isPro":false,"fullname":"Huang","user":"Jinfa","type":"user"},{"_id":"66e477a352356419c4712065","avatarUrl":"/avatars/cd2969ce88e44735c3fe70312c0e64b3.svg","isPro":true,"fullname":"Ruiyao Xu","user":"ruiyaox","type":"user"},{"_id":"66e4e50a52356419c4a1ad14","avatarUrl":"/avatars/4be3ce17671785cbe7126b9c1141478b.svg","isPro":false,"fullname":"Tiankai Yang","user":"tiankaiy","type":"user"},{"_id":"674dfa5d362f943568420c56","avatarUrl":"/avatars/646348d5213128379d1a64b371d58c34.svg","isPro":false,"fullname":"Zilei Shao","user":"zoeshao0425","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"},{"_id":"6aa1bb76b5facad55cd59cf9","avatarUrl":"/avatars/2161d1e4073c8a1c806e98857dffeca2.svg","isPro":false,"fullname":"Yixin Wan","user":"ew728","type":"user"},{"_id":"62d9a2ac51e8289052c22e42","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d9a2ac51e8289052c22e42/d7QLmvyTnA4OvF4FTsNXh.jpeg","isPro":false,"fullname":"Yu Zhou","user":"bryanzhou008","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67784c39dac147922d8d09f0","name":"UCLA","fullname":"University of California, Los Angeles","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67784bd637dfa531fbce95a2/Nf0seEMEn66sPL3QsJXj4.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06245.md","query":{}}">
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Abstract
VDiff-Bench evaluates multimodal language models on fine-grained image difference identification, revealing major weaknesses in detecting subtle low-level visual changes.
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, and illumination. Each question corresponds to two image inputs with 4 choices: the true difference, two hard negative descriptions, and a "no difference" distractor. To make the task challenging, we specifically curate ground-truth-conditioned negatives that require models to distinguish the actual change from nearby semantic alternatives. Experiments with 11 state-of-the-art open- and closed-source MLLMs show that fine-grained visual comparison remains brittle: models exhibit uneven performance across sources and change categories, with persistent failures on subtle low-level changes like noises and textures. For instance, three 7-8B-scale open-source MLLMs score 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level changes like noise and texture, falsely assuming no changes between two image inputs. Surprisingly, despite strong performance of other closed-source commercial models, Grok 4.3 demonstrate remarkable performance drop on identifying noise and texture differences between images, falling significantly behind large open-source models like Kimi K2.5 and K3. Overall, VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.06245 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.06245 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.