r/MachineLearning · · 1 min read

VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed.

Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics. Also, clinically meaningful but rare words were erased leaving the generated report looking repetitive and boring. Importantly, of no clinical utility.

In the paper below, we discuss this behaviour of VLMs for RRG and introduce a framework to actually measure the erasure of terms and introduction of biased terms.

Paper: Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation

Link: Reference Paper

Url: https://arxiv.org/abs/2603.01625

submitted by /u/ade17_in
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning