Official submission of our paper introducing OmniMRG-Bench, a comprehensive benchmark for universal medical report evaluation across multiple imaging modalities. The paper also presents AtomiMed, a hierarchical atomic fact-checking framework that better aligns automatic evaluation with expert radiologist judgments.</p>\n","updatedAt":"2026-07-02T01:42:05.439Z","author":{"_id":"675521057ff406b0e7ff1ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png","fullname":"WANG","name":"Venn2024","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8591441512107849},"editors":["Venn2024"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.31292","authors":[{"_id":"6a45c0fd4f1dd35e48fb8e7f","name":"Yuan Wang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e80","name":"Wanxing Chang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e81","name":"Songtao Jiang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e82","name":"Shujian Gao","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e83","name":"Xiaotian Zhang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e84","name":"Ruifeng Yuan","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e85","name":"Weiwei Cao","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e86","name":"Bowen Shi","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e87","name":"Ling Zhang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e88","name":"Zuozhu Liu","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e89","name":"Jianpeng Zhang","hidden":false}],"publishedAt":"2026-06-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-02T00:00:00.000Z","title":"AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation","submittedOnDailyBy":{"_id":"675521057ff406b0e7ff1ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png","isPro":false,"fullname":"WANG","user":"Venn2024","type":"user","name":"Venn2024"},"summary":"Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks catastrophic diagnostic errors. We address this fundamental limitation by proposing AtomiMed, a universal, modality-agnostic evaluation framework that decomposes complex medical narratives into a standardized, multi-level hierarchy of Atomic Clinical Facts, encompassing Disease-level entities and Attribute-level descriptors, including location, morphology, and severity. By implementing an Agentic Cross-Verification loop between ground-truth and predicted reports, AtomiMed simulates a multi-radiologist peer-review process to verify clinical consistency, thus enabling the decoupled assessment of diagnostic detection and descriptive accuracy. To facilitate standardized evaluation, we introduce MRGEvalKit, an open-source toolkit for automated hierarchical extraction, and curate OmniMRG-Bench, a comprehensive multi-modal benchmark covering X-ray, CT, MRI, and Ultrasound. Extensive experiments on multiple expert-annotated reader studies demonstrate that AtomiMed achieves significantly higher correlation with human radiologist judgment compared to traditional and model-based metrics. Our code are release at https://github.com/Venn2336/MRGEvalkit","upvotes":4,"discussionId":"6a45c0fd4f1dd35e48fb8e8a","githubRepo":"https://github.com/Venn2336/MRGEvalkit","githubRepoAddedBy":"user","ai_summary":"AtomiMed presents a novel evaluation framework for medical report generation that decomposes clinical narratives into atomic facts and uses an agentic cross-verification process to improve accuracy assessment beyond traditional metrics.","ai_keywords":["Medical Report Generation","Atomic Clinical Facts","Agentic Cross-Verification","hierarchical extraction","multi-modal benchmark","radiologist judgment"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"6345aadf5efccdc07f1365a5","name":"ZhejiangUniversity","fullname":"Zhejiang University","avatar":"https://www.gravatar.com/avatar/d1d414628877bec2958f95ad283c15e7?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"675521057ff406b0e7ff1ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png","isPro":false,"fullname":"WANG","user":"Venn2024","type":"user"},{"_id":"68e4dbf1beab849e9baa6e26","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e4dbf1beab849e9baa6e26/It5fgZvTt0JO-gn5Wy81V.png","isPro":false,"fullname":"ZJU-AI4H","user":"ZJU-AI4H","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6345aadf5efccdc07f1365a5","name":"ZhejiangUniversity","fullname":"Zhejiang University","avatar":"https://www.gravatar.com/avatar/d1d414628877bec2958f95ad283c15e7?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.31292.md","query":{}}">
AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation
Published on Jun 30
· Submitted by WANG on Jul 2 Authors: ,
,
,
,
,
,
,
,
,
,
Abstract
AtomiMed presents a novel evaluation framework for medical report generation that decomposes clinical narratives into atomic facts and uses an agentic cross-verification process to improve accuracy assessment beyond traditional metrics.
Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks catastrophic diagnostic errors. We address this fundamental limitation by proposing AtomiMed, a universal, modality-agnostic evaluation framework that decomposes complex medical narratives into a standardized, multi-level hierarchy of Atomic Clinical Facts, encompassing Disease-level entities and Attribute-level descriptors, including location, morphology, and severity. By implementing an Agentic Cross-Verification loop between ground-truth and predicted reports, AtomiMed simulates a multi-radiologist peer-review process to verify clinical consistency, thus enabling the decoupled assessment of diagnostic detection and descriptive accuracy. To facilitate standardized evaluation, we introduce MRGEvalKit, an open-source toolkit for automated hierarchical extraction, and curate OmniMRG-Bench, a comprehensive multi-modal benchmark covering X-ray, CT, MRI, and Ultrasound. Extensive experiments on multiple expert-annotated reader studies demonstrate that AtomiMed achieves significantly higher correlation with human radiologist judgment compared to traditional and model-based metrics. Our code are release at https://github.com/Venn2336/MRGEvalkit
Community
Official submission of our paper introducing OmniMRG-Bench, a comprehensive benchmark for universal medical report evaluation across multiple imaging modalities. The paper also presents AtomiMed, a hierarchical atomic fact-checking framework that better aligns automatic evaluation with expert radiologist judgments.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2606.31292 in a model README.md to link it from this page.
Cite arxiv.org/abs/2606.31292 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2606.31292 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.