Hugging Face Daily Papers · · 3 min read

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Code: <a href=\"https://github.com/alibaba-damo-academy/ClinFusion\" rel=\"nofollow\">https://github.com/alibaba-damo-academy/ClinFusion</a><br>Models: <a href=\"https://huggingface.co/collections/Alibaba-DAMO-Academy/clinfusion\">https://huggingface.co/collections/Alibaba-DAMO-Academy/clinfusion</a></p>\n","updatedAt":"2026-07-28T02:42:16.418Z","author":{"_id":"649d54b314afbb10ce2a9eeb","avatarUrl":"/avatars/15c325d8c2273ff63569f23015e98486.svg","fullname":"Hangjie Yuan","name":"JacobYuan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6702530980110168},"editors":["JacobYuan"],"editorAvatarUrls":["/avatars/15c325d8c2273ff63569f23015e98486.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.24743","authors":[{"_id":"6a6816d573f69d5af2bec5d9","name":"Hangjie Yuan","hidden":false},{"_id":"6a6816d573f69d5af2bec5da","name":"Yichen Qian","hidden":false},{"_id":"6a6816d573f69d5af2bec5db","name":"Zhiwei Tang","hidden":false},{"_id":"6a6816d573f69d5af2bec5dc","name":"Xianzhe Xu","hidden":false},{"_id":"6a6816d573f69d5af2bec5dd","name":"Lirong Wu","hidden":false},{"_id":"6a6816d573f69d5af2bec5de","name":"Sicheng Yang","hidden":false},{"_id":"6a6816d573f69d5af2bec5df","name":"Jinwang Wang","hidden":false},{"_id":"6a6816d573f69d5af2bec5e0","name":"Pengju Wang","hidden":false},{"_id":"6a6816d573f69d5af2bec5e1","name":"Zhitao Zeng","hidden":false},{"_id":"6a6816d573f69d5af2bec5e2","name":"Yizeng Han","hidden":false},{"_id":"6a6816d573f69d5af2bec5e3","name":"Yan Xing","hidden":false},{"_id":"6a6816d573f69d5af2bec5e4","name":"Shengxuan Luo","hidden":false},{"_id":"6a6816d573f69d5af2bec5e5","name":"Tao Feng","hidden":false},{"_id":"6a6816d573f69d5af2bec5e6","name":"Qing Xie","hidden":false},{"_id":"6a6816d573f69d5af2bec5e7","name":"Weigen Yao","hidden":false},{"_id":"6a6816d573f69d5af2bec5e8","name":"Yi Yang","hidden":false},{"_id":"6a6816d573f69d5af2bec5e9","name":"Zuozhu Liu","hidden":false},{"_id":"6a6816d573f69d5af2bec5ea","name":"Jiasheng Tang","hidden":false},{"_id":"6a6816d573f69d5af2bec5eb","name":"Shaocheng Wang","hidden":false},{"_id":"6a6816d573f69d5af2bec5ec","name":"Jitao Wang","hidden":false},{"_id":"6a6816d573f69d5af2bec5ed","name":"Jiahong Dong","hidden":false},{"_id":"6a6816d573f69d5af2bec5ee","name":"Weihua Chen","hidden":false},{"_id":"6a6816d573f69d5af2bec5ef","name":"Feng Xu","hidden":false},{"_id":"6a6816d573f69d5af2bec5f0","name":"Fan Wang","hidden":false}],"publishedAt":"2026-07-27T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding","submittedOnDailyBy":{"_id":"649d54b314afbb10ce2a9eeb","avatarUrl":"/avatars/15c325d8c2273ff63569f23015e98486.svg","isPro":false,"fullname":"Hangjie Yuan","user":"JacobYuan","type":"user","name":"JacobYuan"},"summary":"Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.","upvotes":4,"discussionId":"6a6816d573f69d5af2bec5f1","projectPage":"https://github.com/alibaba-damo-academy/ClinFusion","githubRepo":"https://github.com/alibaba-damo-academy/ClinFusion","githubRepoAddedBy":"user","githubStars":4,"organization":{"_id":"6808e7522a4d69d5111da55f","name":"Alibaba-DAMO-Academy","fullname":"DAMO Academy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"649d54b314afbb10ce2a9eeb","avatarUrl":"/avatars/15c325d8c2273ff63569f23015e98486.svg","isPro":false,"fullname":"Hangjie Yuan","user":"JacobYuan","type":"user"},{"_id":"64b047abb02b95456db10915","avatarUrl":"/avatars/1973377424ff5d586027b91c32f8675c.svg","isPro":false,"fullname":"David Yang","user":"yscript","type":"user"},{"_id":"6523e142d4b61d080773748f","avatarUrl":"/avatars/059774246cd671a9e8d24e3d3297043b.svg","isPro":false,"fullname":"Jankin","user":"jwwangchn","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6808e7522a4d69d5111da55f","name":"Alibaba-DAMO-Academy","fullname":"DAMO Academy","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6808e64de5dd22427c006e10/9J3vdB62CdeTOd_YrGh9w.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.24743.md","query":{}}">
Papers
arxiv:2607.24743

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Published on Jul 27
· Submitted by
Hangjie Yuan
on Jul 28
Authors:
,

Abstract

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.24743
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.24743 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.24743 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers