Dual-IFM is a foundation model for retinal fundus images that is interpretable-by-design. Its BagNet backbone with small receptive fields produces class evidence maps faithful to the model's decision-making process, and a 2D projection layer learned during pretraining enables direct visualization of the representation space, revealing both meaningful clinical clusters and potential spurious correlations. Trained on 800,000+ color fundus photographs from multiple sources, Dual-IFM achieves performance comparable to RETFound (16x more parameters) while remaining interpretable.</p>\n","updatedAt":"2026-08-10T08:17:11.730Z","author":{"_id":"65020cc8bd92cd0a8f2da64e","avatarUrl":"/avatars/30fcefd92432899ea6a178dfe4990495.svg","fullname":"Camila Roa","name":"CamilaR20","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.831766664981842},"editors":["CamilaR20"],"editorAvatarUrls":["/avatars/30fcefd92432899ea6a178dfe4990495.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2603.18846","authors":[{"_id":"6a7702738e9301703eaa5a6e","name":"Samuel Ofosu Mensah","hidden":false},{"_id":"6a7702738e9301703eaa5a6f","user":{"_id":"65020cc8bd92cd0a8f2da64e","avatarUrl":"/avatars/30fcefd92432899ea6a178dfe4990495.svg","isPro":false,"fullname":"Camila Roa","user":"CamilaR20","type":"user","name":"CamilaR20"},"name":"Camila Roa","status":"claimed_verified","statusLastChangedAt":"2026-08-08T16:45:04.838Z","hidden":false},{"_id":"6a7702738e9301703eaa5a70","name":"Kerol Djoumessi","hidden":false},{"_id":"6a7702738e9301703eaa5a71","name":"Philipp Berens","hidden":false}],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"Towards Interpretable Foundation Models for Retinal Fundus Images","submittedOnDailyBy":{"_id":"65020cc8bd92cd0a8f2da64e","avatarUrl":"/avatars/30fcefd92432899ea6a178dfe4990495.svg","isPro":false,"fullname":"Camila Roa","user":"CamilaR20","type":"user","name":"CamilaR20"},"summary":"Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limited interpretability, a critical issue in high-stakes domains such as medical imaging. We propose DualIFM, a foundation model that is interpretable-by-design via a BagNet backbone whose small receptive fields generate class evidence maps that are faithful to the model's decision-making process. Additionally, DualIFM incorporates a 2D projection layer during pretraining that enables direct visualization of the representation space, providing a dataset-level view of the learned structure including meaningful clinical clusters as well as potential spurious correlations. We trained DualIFM on over 800,000 color fundus photographs from various sources to learn generalizable representations for different downstream tasks. Our model achieves performance comparable to RETFound, which has 16times more parameters, while providing interpretable predictions on out-of-distribution data. These results suggest that large-scale SSL pretraining paired with inherent interpretability can lead to robust representations for retinal imaging. Code and pretrained models are available at github.com/berenslab/interpretable_FM.","upvotes":0,"discussionId":"6a7702738e9301703eaa5a72","githubRepo":"https://github.com/berenslab/interpretable_FM","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"67efb00c6e51a83b10018b42","name":"UniTuebingen","fullname":"Eberhard Karls Universität Tübingen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67efaed96e51a83b10013635/IBHUSQLny8BmeCdJro55K.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"organization":{"_id":"67efb00c6e51a83b10018b42","name":"UniTuebingen","fullname":"Eberhard Karls Universität Tübingen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67efaed96e51a83b10013635/IBHUSQLny8BmeCdJro55K.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2603/2603.18846.md","query":{}}">
Towards Interpretable Foundation Models for Retinal Fundus Images
Abstract
Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limited interpretability, a critical issue in high-stakes domains such as medical imaging. We propose DualIFM, a foundation model that is interpretable-by-design via a BagNet backbone whose small receptive fields generate class evidence maps that are faithful to the model's decision-making process. Additionally, DualIFM incorporates a 2D projection layer during pretraining that enables direct visualization of the representation space, providing a dataset-level view of the learned structure including meaningful clinical clusters as well as potential spurious correlations. We trained DualIFM on over 800,000 color fundus photographs from various sources to learn generalizable representations for different downstream tasks. Our model achieves performance comparable to RETFound, which has 16times more parameters, while providing interpretable predictions on out-of-distribution data. These results suggest that large-scale SSL pretraining paired with inherent interpretability can lead to robust representations for retinal imaging. Code and pretrained models are available at github.com/berenslab/interpretable_FM.
Community
Dual-IFM is a foundation model for retinal fundus images that is interpretable-by-design. Its BagNet backbone with small receptive fields produces class evidence maps faithful to the model's decision-making process, and a 2D projection layer learned during pretraining enables direct visualization of the representation space, revealing both meaningful clinical clusters and potential spurious correlations. Trained on 800,000+ color fundus photographs from multiple sources, Dual-IFM achieves performance comparable to RETFound (16x more parameters) while remaining interpretable.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2603.18846 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2603.18846 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.