Code: <a href=\"https://github.com/liuzhipenggg/CAPEval\" rel=\"nofollow\">https://github.com/liuzhipenggg/CAPEval</a></p>\n<p>Suppose an image contains 10 factual claims: Caption <strong>A references 9 of them but misstates 2</strong>, while Caption <strong>B only mentions 5, all of which are factually correct</strong>. Which caption is higher‑quality?</p>\n<p>You may find yourself stumped.</p>\n<p>This question is difficult to resolve with a single aggregated score, since each has its own merits. <em>Caption A’s strength lies in its comprehensiveness, whereas Caption B excels in factual accuracy</em>.</p>\n<p>Over the past several years, a growing body of work has used vision‑language models to‑recaption training images, aiming to improve dataset quality with richer and more accurate captions. The importance of captions in multimodal model training is now widely recognized.</p>\n<p>Yet, as the opening example demonstrates, a more fundamental question remains poorly understood: <strong>what kinds of captions genuinely benefit downstream model training?</strong></p>\n<p>To tackle this problem, we have proposed a new caption evaluation framework: <strong>CAPEval</strong> (<strong>C</strong>overage <strong>A</strong>nd <strong>P</strong>recision <strong>E</strong>valuation). Unlike previous benchmarks that evaluate captions in isolation, we <strong>further train</strong> vision‑language and text‑to‑image models on captions produced by different captioning models. Through controlled‑variable experiments, we examine how different quality attributes of captions ultimately influence downstream model performance.</p>\n<p>Our paper identifies a clear task‑dependent dissociation: <strong>Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance.</strong></p>\n","updatedAt":"2026-08-05T02:33:23.095Z","author":{"_id":"69b4400b169d46f0e76c8612","avatarUrl":"/avatars/be481d459e8da4494327ca7c3e582b04.svg","fullname":"Galaxy","name":"LiuzhipengUCAS","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8892461061477661},"editors":["LiuzhipengUCAS"],"editorAvatarUrls":["/avatars/be481d459e8da4494327ca7c3e582b04.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.02589","authors":[{"_id":"6a716536ec5082b9f872ceb6","user":{"_id":"69b4400b169d46f0e76c8612","avatarUrl":"/avatars/be481d459e8da4494327ca7c3e582b04.svg","isPro":false,"fullname":"Galaxy","user":"LiuzhipengUCAS","type":"user","name":"LiuzhipengUCAS"},"name":"Zhipeng Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-04T09:04:25.025Z","hidden":false},{"_id":"6a716536ec5082b9f872ceb7","name":"Haochen Wang","hidden":false},{"_id":"6a716536ec5082b9f872ceb8","name":"Zhaoxiang Zhang","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"CAPEval: A Decoupled Caption Evaluation across Understanding and Generation","submittedOnDailyBy":{"_id":"69b4400b169d46f0e76c8612","avatarUrl":"/avatars/be481d459e8da4494327ca7c3e582b04.svg","isPro":false,"fullname":"Galaxy","user":"LiuzhipengUCAS","type":"user","name":"LiuzhipengUCAS"},"summary":"Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.","upvotes":15,"discussionId":"6a716537ec5082b9f872ceb9","projectPage":"https://huggingface.co/datasets/LiuzhipengUCAS/CAPEval","githubRepo":"https://github.com/liuzhipenggg/CAPEval","githubRepoAddedBy":"user","githubStars":9},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69b4400b169d46f0e76c8612","avatarUrl":"/avatars/be481d459e8da4494327ca7c3e582b04.svg","isPro":false,"fullname":"Galaxy","user":"LiuzhipengUCAS","type":"user"},{"_id":"6a72a1b93a49b8bf5c63f487","avatarUrl":"/avatars/707432e609434c32f6e935cd5f4dc00b.svg","isPro":false,"fullname":"Yue Xiao","user":"Axiaoyue","type":"user"},{"_id":"6a72a42533b3688348136f61","avatarUrl":"/avatars/c3a4d8d646d7416a750d75f7511346e4.svg","isPro":false,"fullname":"Hua Zhu","user":"ZZbodyopye","type":"user"},{"_id":"662b47e0b771c8c663eb0de1","avatarUrl":"/avatars/1a709800107f1813701fe396a9b69928.svg","isPro":false,"fullname":"HaochenWang","user":"whc0926","type":"user"},{"_id":"69ebab9dc9af914d250c043d","avatarUrl":"/avatars/0cf5135636e85a22a944b674ff85e98f.svg","isPro":false,"fullname":"Galaxy","user":"ZhipengLiuUCAS","type":"user"},{"_id":"6a6a8260977fbfce4bac218d","avatarUrl":"/avatars/884d72707d0a24d2d62b991428d4fa4d.svg","isPro":false,"fullname":"Susan Thompson","user":"susanthompson","type":"user"},{"_id":"6a6a9a25a5b9c4c08baf22cb","avatarUrl":"/avatars/d3dbbe26cea0e3549d6051687e82fe4b.svg","isPro":false,"fullname":"Daniel Brown","user":"cobalttrail","type":"user"},{"_id":"6a6c7bb54c678c98fac02eb4","avatarUrl":"/avatars/cd68c63d1aecc36dbede25f2fa063602.svg","isPro":false,"fullname":"Linda Miller","user":"Indigo-Linda","type":"user"},{"_id":"6a6c8070ef16c968823e47e3","avatarUrl":"/avatars/fce97db08e578c90c99f8c06453ae1bb.svg","isPro":false,"fullname":"Matthew Moore","user":"orbitReed","type":"user"},{"_id":"6a6aa148ad5a6f2f63620596","avatarUrl":"/avatars/11497c4cb1c583ebe22f266fa8585e92.svg","isPro":false,"fullname":"David Williams","user":"timothy-4448888","type":"user"},{"_id":"6a6dc68dec5123fe9d7bb623","avatarUrl":"/avatars/1fb00ab7cb5eeb532824effa35ef137a.svg","isPro":false,"fullname":"Steven Smith","user":"james-6985730","type":"user"},{"_id":"6a6de7f10f8a360db52be308","avatarUrl":"/avatars/ff0b0c8b58aa8e165cf223cf59acc6dc.svg","isPro":false,"fullname":"Jennifer Miller","user":"jordan-5524794","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.02589.md","query":{}}">
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
Published on Aug 3
· Submitted by Galaxy on Aug 5 Abstract
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.
Community
Code: https://github.com/liuzhipenggg/CAPEval
Suppose an image contains 10 factual claims: Caption A references 9 of them but misstates 2, while Caption B only mentions 5, all of which are factually correct. Which caption is higher‑quality?
You may find yourself stumped.
This question is difficult to resolve with a single aggregated score, since each has its own merits. Caption A’s strength lies in its comprehensiveness, whereas Caption B excels in factual accuracy.
Over the past several years, a growing body of work has used vision‑language models to‑recaption training images, aiming to improve dataset quality with richer and more accurate captions. The importance of captions in multimodal model training is now widely recognized.
Yet, as the opening example demonstrates, a more fundamental question remains poorly understood: what kinds of captions genuinely benefit downstream model training?
To tackle this problem, we have proposed a new caption evaluation framework: CAPEval (Coverage And Precision Evaluation). Unlike previous benchmarks that evaluate captions in isolation, we further train vision‑language and text‑to‑image models on captions produced by different captioning models. Through controlled‑variable experiments, we examine how different quality attributes of captions ultimately influence downstream model performance.
Our paper identifies a clear task‑dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.02589 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.02589 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.02589 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.