VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed.
Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics. Also, clinically meaningful but rare words were erased leaving the generated report looking repetitive and boring. Importantly, of no clinical utility.
In the paper below, we discuss this behaviour of VLMs for RRG and introduce a framework to actually measure the erasure of terms and introduction of biased terms.
Paper: Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
Link: Reference Paper
[link] [comments]
More from r/MachineLearning
-
Revisiting the Efficient Channel Attention paper (2019, 12k citations) - the central hypothesis isn't quite right [D]
Aug 16
-
SSOG-Attention: Sum Of Separable Gaussians as a sub-quadratic and scalable alternative to SDPA. [R]
Aug 16
-
How can we solve long-range recall in linear attention? [D]
Aug 16
-
Survival of the Fitted: Qwen3.6-27B’s Jacobian lens reads and steers Qwen3.8-27B with zero refitting [R]
Aug 15
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.