\n\t<a id=\"english\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#english\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tEnglish\n\t</span>\n</h3>\n<p>Existing AI systems can assist with coronary angiography interpretation, but most models still provide only the final prediction. The intermediate reasoning process remains largely a black box, making the output difficult to verify directly in clinical practice and thereby limiting trust in the model’s interpretation.</p>\n<p>CARDEA is designed to address this problem.</p>\n<p>Our goal is to use the reasoning trace produced by a large vision-language model before its final answer as a source of interpretability, while further guiding the model to indicate the image regions it attends to through Chain-of-Box (CoB), making it easier for clinicians to audit how the model reached its conclusions.</p>\n<p>CARDEA is trained in three stages. We first perform visual feature alignment to foundational CAG features, then use self-distillation to automatically synthesize CoB reasoning traces for cold-start training. Finally, during reinforcement learning with verifiable rewards, we introduce a reward mechanism that encourages both correct interpretation and CoB behavior.</p>\n<p>The model weights and inference code are publicly available.</p>\n<hr>\n<h3 class=\"relative group flex items-baseline\">\n\t<a id=\"中文\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#中文\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t中文\n\t</span>\n</h3>\n<p>冠狀動脈攝影目前已有人工智慧系統能輔助判讀,但多數模型仍只輸出最終預測,中間的推理過程仍像黑箱,臨床上不容易直接驗證,也因此降低了對模型判讀結果的信任。</p>\n<p>CARDEA 就是為了解決這個痛點。</p>\n<p>我們的目標是讓大型視覺語言模型在輸出最終結果前的推理過程成為可解釋性的來源,並進一步引導模型在推理時指出它所關注的多個影像區域,也就是 Chain-of-Box,讓臨床人員更容易檢查模型的判斷依據。</p>\n<p>CARDEA 採用三階段訓練。我們先進行基礎冠狀動脈攝影特徵對齊,再透過自我蒸餾自動合成帶有 CoB 的推理資料進行冷啟動訓練,最後在可驗證獎勵強化學習階段加入獎勵機制,同時鼓勵判讀正確性與 CoB 行為。</p>\n<p>目前模型權重與推論程式碼皆已公開。</p>\n","updatedAt":"2026-09-11T02:34:17.257Z","author":{"_id":"697035d2d974214e2ccd5c7b","avatarUrl":"/avatars/ade0fe033c0821ea8a8ed487a0c10db4.svg","fullname":"Jia-Jen Lee","name":"benbayibaurba","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6631733775138855},"editors":["benbayibaurba"],"editorAvatarUrls":["/avatars/ade0fe033c0821ea8a8ed487a0c10db4.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06931","authors":[{"_id":"6aa0b959d0174964227bebd5","user":{"_id":"697035d2d974214e2ccd5c7b","avatarUrl":"/avatars/ade0fe033c0821ea8a8ed487a0c10db4.svg","isPro":false,"fullname":"Jia-Jen Lee","user":"benbayibaurba","type":"user","name":"benbayibaurba"},"name":"Jia-Jen Lee","status":"admin_assigned","statusLastChangedAt":"2026-09-10T16:44:01.688Z","hidden":false},{"_id":"6aa0b959d0174964227bebd6","name":"Shih-Yen Hou","hidden":false},{"_id":"6aa0b959d0174964227bebd7","name":"Kee Koon Ng","hidden":false},{"_id":"6aa0b959d0174964227bebd8","name":"Wei-Chun Wang","hidden":false},{"_id":"6aa0b959d0174964227bebd9","name":"Shih-Sheng Chang","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation","submittedOnDailyBy":{"_id":"697035d2d974214e2ccd5c7b","avatarUrl":"/avatars/ade0fe033c0821ea8a8ed487a0c10db4.svg","isPro":false,"fullname":"Jia-Jen Lee","user":"benbayibaurba","type":"user","name":"benbayibaurba"},"summary":"Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation varies substantially among observers. Existing AI systems can improve consistency but lack auditable decision processes and are limited in comprehensive open-ended assessment, undermining clinician trust and clinical adoption readiness. We developed CARDEA, a unified large vision-language model that serves as the inference core of a CAG pipeline. It was trained solely on public datasets and closed-ended tasks in three stages: visual feature alignment, a self-distilled Chain-of-Box (CoB) cold start, and reinforcement learning with verifiable rewards (RLVR) with a CoB reward encouraging bounding-box use in the reasoning trace. We assessed its two study-level diagnoses, dominance classification and complexity assessment, against a dedicated classifier and two interventional cardiologists. Report generation was excluded from training and evaluated zero-shot across stages on an external cohort using vessel-severity macro-F_1. CARDEA trailed the classifier on in-distribution dominance but drew level under domain shift (accuracy, 0.91 [95% confidence interval (CI), 0.86 to 0.95]) and was comparable to the cardiologists on complexity assessment (accuracy, 0.90 [CI, 0.82 to 0.97]). Only RLVR improved zero-shot report generation, raising its vessel-severity macro-F_1 (0.686 [CI, 0.664 to 0.707]) above the untuned base model (0.513) and over twice the always-normal floor (0.312). CARDEA runs an end-to-end CAG pipeline from raw multi-view videos through keyframe selection to study-level diagnosis while exposing auditable spatial evidence behind its conclusions. RLVR on verifiable closed-ended tasks surfaced open-ended reporting ability that supervised imitation did not. Clinical use requires prospective validation against expert cardiologists.","upvotes":4,"discussionId":"6aa0b959d0174964227bebda","githubRepo":"https://github.com/benbayibaurba/cardea","githubRepoAddedBy":"user","ai_summary":"A unified vision-language model for coronary angiography uses chain-of-box reasoning and reinforcement learning with verifiable rewards to provide auditable diagnoses and improve zero-shot report generation.","ai_keywords":["large vision-language model","visual feature alignment","self-distilled Chain-of-Box","reinforcement learning with verifiable rewards","RLVR","bounding-box","zero-shot","macro-F1"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65efe26bffc396ef0ec79f89","avatarUrl":"/avatars/55b92db65836b37048772fc359d02030.svg","isPro":false,"fullname":"Ben Lee","user":"ben81828","type":"user"},{"_id":"697035d2d974214e2ccd5c7b","avatarUrl":"/avatars/ade0fe033c0821ea8a8ed487a0c10db4.svg","isPro":false,"fullname":"Jia-Jen Lee","user":"benbayibaurba","type":"user"},{"_id":"6aa37c870b0291cdd4e2a796","avatarUrl":"/avatars/416fd9ff4429d1cdc754926d5d3ac491.svg","isPro":false,"fullname":"Bochen-Jiang","user":"tavisJ","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06931.md","query":{}}">
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation
Abstract
A unified vision-language model for coronary angiography uses chain-of-box reasoning and reinforcement learning with verifiable rewards to provide auditable diagnoses and improve zero-shot report generation.
Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation varies substantially among observers. Existing AI systems can improve consistency but lack auditable decision processes and are limited in comprehensive open-ended assessment, undermining clinician trust and clinical adoption readiness. We developed CARDEA, a unified large vision-language model that serves as the inference core of a CAG pipeline. It was trained solely on public datasets and closed-ended tasks in three stages: visual feature alignment, a self-distilled Chain-of-Box (CoB) cold start, and reinforcement learning with verifiable rewards (RLVR) with a CoB reward encouraging bounding-box use in the reasoning trace. We assessed its two study-level diagnoses, dominance classification and complexity assessment, against a dedicated classifier and two interventional cardiologists. Report generation was excluded from training and evaluated zero-shot across stages on an external cohort using vessel-severity macro-F_1. CARDEA trailed the classifier on in-distribution dominance but drew level under domain shift (accuracy, 0.91 [95% confidence interval (CI), 0.86 to 0.95]) and was comparable to the cardiologists on complexity assessment (accuracy, 0.90 [CI, 0.82 to 0.97]). Only RLVR improved zero-shot report generation, raising its vessel-severity macro-F_1 (0.686 [CI, 0.664 to 0.707]) above the untuned base model (0.513) and over twice the always-normal floor (0.312). CARDEA runs an end-to-end CAG pipeline from raw multi-view videos through keyframe selection to study-level diagnosis while exposing auditable spatial evidence behind its conclusions. RLVR on verifiable closed-ended tasks surfaced open-ended reporting ability that supervised imitation did not. Clinical use requires prospective validation against expert cardiologists.
Community
English
Existing AI systems can assist with coronary angiography interpretation, but most models still provide only the final prediction. The intermediate reasoning process remains largely a black box, making the output difficult to verify directly in clinical practice and thereby limiting trust in the model’s interpretation.
CARDEA is designed to address this problem.
Our goal is to use the reasoning trace produced by a large vision-language model before its final answer as a source of interpretability, while further guiding the model to indicate the image regions it attends to through Chain-of-Box (CoB), making it easier for clinicians to audit how the model reached its conclusions.
CARDEA is trained in three stages. We first perform visual feature alignment to foundational CAG features, then use self-distillation to automatically synthesize CoB reasoning traces for cold-start training. Finally, during reinforcement learning with verifiable rewards, we introduce a reward mechanism that encourages both correct interpretation and CoB behavior.
The model weights and inference code are publicly available.
中文
冠狀動脈攝影目前已有人工智慧系統能輔助判讀,但多數模型仍只輸出最終預測,中間的推理過程仍像黑箱,臨床上不容易直接驗證,也因此降低了對模型判讀結果的信任。
CARDEA 就是為了解決這個痛點。
我們的目標是讓大型視覺語言模型在輸出最終結果前的推理過程成為可解釋性的來源,並進一步引導模型在推理時指出它所關注的多個影像區域,也就是 Chain-of-Box,讓臨床人員更容易檢查模型的判斷依據。
CARDEA 採用三階段訓練。我們先進行基礎冠狀動脈攝影特徵對齊,再透過自我蒸餾自動合成帶有 CoB 的推理資料進行冷啟動訓練,最後在可驗證獎勵強化學習階段加入獎勵機制,同時鼓勵判讀正確性與 CoB 行為。
目前模型權重與推論程式碼皆已公開。
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.06931 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.06931 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.