Hugging Face Daily Papers · · 3 min read

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Official submission of our paper introducing OmniMRG-Bench, a comprehensive benchmark for universal medical report evaluation across multiple imaging modalities. The paper also presents AtomiMed, a hierarchical atomic fact-checking framework that better aligns automatic evaluation with expert radiologist judgments.</p>\n","updatedAt":"2026-07-02T01:42:05.439Z","author":{"_id":"675521057ff406b0e7ff1ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png","fullname":"WANG","name":"Venn2024","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8591441512107849},"editors":["Venn2024"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.31292","authors":[{"_id":"6a45c0fd4f1dd35e48fb8e7f","name":"Yuan Wang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e80","name":"Wanxing Chang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e81","name":"Songtao Jiang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e82","name":"Shujian Gao","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e83","name":"Xiaotian Zhang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e84","name":"Ruifeng Yuan","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e85","name":"Weiwei Cao","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e86","name":"Bowen Shi","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e87","name":"Ling Zhang","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e88","name":"Zuozhu Liu","hidden":false},{"_id":"6a45c0fd4f1dd35e48fb8e89","name":"Jianpeng Zhang","hidden":false}],"publishedAt":"2026-06-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-02T00:00:00.000Z","title":"AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation","submittedOnDailyBy":{"_id":"675521057ff406b0e7ff1ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png","isPro":false,"fullname":"WANG","user":"Venn2024","type":"user","name":"Venn2024"},"summary":"Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks catastrophic diagnostic errors. We address this fundamental limitation by proposing AtomiMed, a universal, modality-agnostic evaluation framework that decomposes complex medical narratives into a standardized, multi-level hierarchy of Atomic Clinical Facts, encompassing Disease-level entities and Attribute-level descriptors, including location, morphology, and severity. By implementing an Agentic Cross-Verification loop between ground-truth and predicted reports, AtomiMed simulates a multi-radiologist peer-review process to verify clinical consistency, thus enabling the decoupled assessment of diagnostic detection and descriptive accuracy. To facilitate standardized evaluation, we introduce MRGEvalKit, an open-source toolkit for automated hierarchical extraction, and curate OmniMRG-Bench, a comprehensive multi-modal benchmark covering X-ray, CT, MRI, and Ultrasound. Extensive experiments on multiple expert-annotated reader studies demonstrate that AtomiMed achieves significantly higher correlation with human radiologist judgment compared to traditional and model-based metrics. Our code are release at https://github.com/Venn2336/MRGEvalkit","upvotes":4,"discussionId":"6a45c0fd4f1dd35e48fb8e8a","githubRepo":"https://github.com/Venn2336/MRGEvalkit","githubRepoAddedBy":"user","ai_summary":"AtomiMed presents a novel evaluation framework for medical report generation that decomposes clinical narratives into atomic facts and uses an agentic cross-verification process to improve accuracy assessment beyond traditional metrics.","ai_keywords":["Medical Report Generation","Atomic Clinical Facts","Agentic Cross-Verification","hierarchical extraction","multi-modal benchmark","radiologist judgment"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"6345aadf5efccdc07f1365a5","name":"ZhejiangUniversity","fullname":"Zhejiang University","avatar":"https://www.gravatar.com/avatar/d1d414628877bec2958f95ad283c15e7?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"675521057ff406b0e7ff1ae9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KKWQdnwh117nMIB45fs7t.png","isPro":false,"fullname":"WANG","user":"Venn2024","type":"user"},{"_id":"68e4dbf1beab849e9baa6e26","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e4dbf1beab849e9baa6e26/It5fgZvTt0JO-gn5Wy81V.png","isPro":false,"fullname":"ZJU-AI4H","user":"ZJU-AI4H","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6345aadf5efccdc07f1365a5","name":"ZhejiangUniversity","fullname":"Zhejiang University","avatar":"https://www.gravatar.com/avatar/d1d414628877bec2958f95ad283c15e7?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.31292.md","query":{}}">
Papers
arxiv:2606.31292

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

Published on Jun 30
· Submitted by
WANG
on Jul 2
Authors:
,
,
,
,
,
,
,
,
,
,

Abstract

AtomiMed presents a novel evaluation framework for medical report generation that decomposes clinical narratives into atomic facts and uses an agentic cross-verification process to improve accuracy assessment beyond traditional metrics.

Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks catastrophic diagnostic errors. We address this fundamental limitation by proposing AtomiMed, a universal, modality-agnostic evaluation framework that decomposes complex medical narratives into a standardized, multi-level hierarchy of Atomic Clinical Facts, encompassing Disease-level entities and Attribute-level descriptors, including location, morphology, and severity. By implementing an Agentic Cross-Verification loop between ground-truth and predicted reports, AtomiMed simulates a multi-radiologist peer-review process to verify clinical consistency, thus enabling the decoupled assessment of diagnostic detection and descriptive accuracy. To facilitate standardized evaluation, we introduce MRGEvalKit, an open-source toolkit for automated hierarchical extraction, and curate OmniMRG-Bench, a comprehensive multi-modal benchmark covering X-ray, CT, MRI, and Ultrasound. Extensive experiments on multiple expert-annotated reader studies demonstrate that AtomiMed achieves significantly higher correlation with human radiologist judgment compared to traditional and model-based metrics. Our code are release at https://github.com/Venn2336/MRGEvalkit

Community

Official submission of our paper introducing OmniMRG-Bench, a comprehensive benchmark for universal medical report evaluation across multiple imaging modalities. The paper also presents AtomiMed, a hierarchical atomic fact-checking framework that better aligns automatic evaluation with expert radiologist judgments.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.31292
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.31292 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.31292 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.31292 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers