Hugging Face Daily Papers · · 4 min read

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

\n<li>MedPMC Corpus: <a href=\"https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline\">https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline</a></li>\n<li>MedPMC-CLIP: <a href=\"https://huggingface.co/Yale-BIDS-Chen/medpmc-clip-l-14_jun24_v1\">https://huggingface.co/Yale-BIDS-Chen/medpmc-clip-l-14_jun24_v1</a></li>\n</ul>\n","updatedAt":"2026-07-13T13:01:06.634Z","author":{"_id":"62c27106fc1922be7ba91270","avatarUrl":"/avatars/3a01709d62257fe19085753a9226cc2a.svg","fullname":"Hyunjae Kim","name":"Nowkim","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3382111191749573},"editors":["Nowkim"],"editorAvatarUrls":["/avatars/3a01709d62257fe19085753a9226cc2a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.07673","authors":[{"_id":"6a54e0f4a9d74d6e65bbce34","name":"Hyunjae Kim","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce35","name":"Dain Kim","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce36","name":"Pan Xiao","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce37","name":"Serina S. Applebaum","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce38","name":"Younjoon Chung","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce39","name":"Xuguang Ai","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce3a","name":"Yu Yin","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce3b","name":"Roy Jiang","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce3c","name":"Yuexi Du","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce3d","name":"Yawen Wei","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce3e","name":"Yiming Kong","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce3f","name":"Tuo Guo","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce40","name":"Zhiyuan Cao","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce41","name":"Mengmeng Du","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce42","name":"Yuelei Fu","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce43","name":"Yan Hu","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce44","name":"Rui Shi","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce45","name":"Gui Yang","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce46","name":"Kevin W. Jin","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce47","name":"Yuntian Liu","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce48","name":"Yuxuan Tian","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce49","name":"Jonathan Marquez","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce4a","name":"Zhen Chen","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce4b","name":"Sheng Zhang","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce4c","name":"Hoifung Poon","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce4d","name":"Hua Xu","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce4e","name":"Jaewoo Kang","hidden":false},{"_id":"6a54e0f4a9d74d6e65bbce4f","name":"Qingyu Chen","hidden":false}],"publishedAt":"2026-07-08T00:00:00.000Z","submittedOnDailyAt":"2026-07-13T00:00:00.000Z","title":"MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models","submittedOnDailyBy":{"_id":"62c27106fc1922be7ba91270","avatarUrl":"/avatars/3a01709d62257fe19085753a9226cc2a.svg","isPro":true,"fullname":"Hyunjae Kim","user":"Nowkim","type":"user","name":"Nowkim"},"summary":"Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.","upvotes":2,"discussionId":"6a54e0f5a9d74d6e65bbce50","githubRepo":"https://github.com/Yale-BIDS-Chen-Lab/MedPMC","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6898efc33c7d72d5dc83bd38","name":"Yale-BIDS-Chen","fullname":"Yale-BIDS-Chen-Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6898edbb4e9828e62a117e2f/8VBqM7qTAtK8RoAmAiozM.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"62c27106fc1922be7ba91270","avatarUrl":"/avatars/3a01709d62257fe19085753a9226cc2a.svg","isPro":true,"fullname":"Hyunjae Kim","user":"Nowkim","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6898efc33c7d72d5dc83bd38","name":"Yale-BIDS-Chen","fullname":"Yale-BIDS-Chen-Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6898edbb4e9828e62a117e2f/8VBqM7qTAtK8RoAmAiozM.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.07673.md","query":{}}">
Papers
arxiv:2607.07673

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Published on Jul 8
· Submitted by
Hyunjae Kim
on Jul 13
Authors:
,

Abstract

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.07673
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.07673 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.07673 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.07673 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers