Hugging Face Daily Papers · · 4 min read

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Most clinical AI models are only tested when every scan is present. In the real world modalities go missing, and some models fail without ever raising a flag. This framework was built to catch those silent failures before deployment.</p>\n","updatedAt":"2026-08-04T14:33:46.437Z","author":{"_id":"628ddf04986ae70e823298f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png","fullname":"Sebastián Andres Cajas Ordóñez","name":"sebasmos","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9503255486488342},"editors":["sebasmos"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.01462","authors":[{"_id":"6a71f72a1a375f948521c04f","name":"Quang Bui","hidden":false},{"_id":"6a71f72a1a375f948521c050","user":{"_id":"6799a0d718cb282841d79267","avatarUrl":"/avatars/2197c339d7dbb9ecbb41d18acfdf172f.svg","isPro":false,"fullname":"Shlok Jaiswal","user":"gullyboyslok","type":"user","name":"gullyboyslok"},"name":"Shlok Jaiswal","status":"claimed_verified","statusLastChangedAt":"2026-08-04T16:45:04.783Z","hidden":false},{"_id":"6a71f72a1a375f948521c051","name":"Samuel Paik-Heintz","hidden":false},{"_id":"6a71f72a1a375f948521c052","name":"Kevin Zhou","hidden":false},{"_id":"6a71f72a1a375f948521c053","name":"Kaushik Madapati","hidden":false},{"_id":"6a71f72a1a375f948521c054","name":"Krittaphas Chaisutyakorn","hidden":false},{"_id":"6a71f72a1a375f948521c055","name":"Noah Dane Hebdon","hidden":false},{"_id":"6a71f72a1a375f948521c056","name":"Dimitrios Proios","hidden":false},{"_id":"6a71f72a1a375f948521c057","name":"Sebastián Andrés Cajas Ordóñez","hidden":false},{"_id":"6a71f72a1a375f948521c058","name":"Kacper Dobek","hidden":false},{"_id":"6a71f72a1a375f948521c059","name":"Boya Zhang","hidden":false},{"_id":"6a71f72a1a375f948521c05a","name":"Aly Dhedhi","hidden":false},{"_id":"6a71f72a1a375f948521c05b","name":"Ahram Han","hidden":false},{"_id":"6a71f72a1a375f948521c05c","name":"Kushul Reddy Palakala","hidden":false},{"_id":"6a71f72a1a375f948521c05d","name":"Rahul Gorijavolu","hidden":false},{"_id":"6a71f72a1a375f948521c05e","name":"Jacques Kpodonu","hidden":false},{"_id":"6a71f72a1a375f948521c05f","name":"Leo Anthony Celi","hidden":false}],"publishedAt":"2026-08-02T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI","submittedOnDailyBy":{"_id":"628ddf04986ae70e823298f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png","isPro":false,"fullname":"Sebastián Andres Cajas Ordóñez","user":"sebasmos","type":"user","name":"sebasmos"},"summary":"Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modality-level, and is separate from post-hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model-agnostic modality-failure framework: given N modality embeddings, any mask-aware probe, and labels, it returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals. We release it as a small, unit-tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per-modality loud-vs-silent rates, and scales to a three-modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per-example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT-ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC-IV cohort, where on the held-out test split (n = 245) dropping echo nearly doubles error. The narrow echo-to-ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED-AI.","upvotes":3,"discussionId":"6a71f72a1a375f948521c060","githubRepo":"https://github.com/criticaldata/PRIMED-AI","githubRepoAddedBy":"user","githubStars":5,"organization":{"_id":"63728bde14d543d507ae970d","name":"MIT","fullname":"Massachusetts Institute of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S90qoeEJeEYaYf-c7Zs8g.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"628ddf04986ae70e823298f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png","isPro":false,"fullname":"Sebastián Andres Cajas Ordóñez","user":"sebasmos","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6799a0d718cb282841d79267","avatarUrl":"/avatars/2197c339d7dbb9ecbb41d18acfdf172f.svg","isPro":false,"fullname":"Shlok Jaiswal","user":"gullyboyslok","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63728bde14d543d507ae970d","name":"MIT","fullname":"Massachusetts Institute of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S90qoeEJeEYaYf-c7Zs8g.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.01462.md","query":{}}">
Papers
arxiv:2608.01462

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

Authors:
,

Abstract

Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modality-level, and is separate from post-hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model-agnostic modality-failure framework: given N modality embeddings, any mask-aware probe, and labels, it returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals. We release it as a small, unit-tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per-modality loud-vs-silent rates, and scales to a three-modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per-example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT-ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC-IV cohort, where on the held-out test split (n = 245) dropping echo nearly doubles error. The narrow echo-to-ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED-AI.

Community

Paper submitter about 6 hours ago

Most clinical AI models are only tested when every scan is present. In the real world modalities go missing, and some models fail without ever raising a flag. This framework was built to catch those silent failures before deployment.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.01462
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.01462 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.01462 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.01462 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers