Cura 1T leads frontier models on 5 of 6 hardest healthcare benchmarks:</p>\n<ul>\n<li>HealthBench Hard: 36.8 (GPT-5.5: 31.5)</li>\n<li>HealthBench Professional: 66.2 (Claude Fable 5: 66.0)</li>\n<li>MedXpertQA-Text: 60.0 (GPT-5.5: 59.6)</li>\n<li>MedXpertQA-Multimodalt: 72.2 (GPT-5.5: 77.1)</li>\n<li>AgentClinic: 79.6 (Claude Opus 4.8: 79.4)</li>\n<li>MedAgentBench-v2: 94.0 (Claude Opus 4.8: 93.7)</li>\n</ul>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/632cdea254e2c512c8f95b12/sinoswJp-2Ar-K1BNzm0I.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/632cdea254e2c512c8f95b12/sinoswJp-2Ar-K1BNzm0I.png\" alt=\"main_comparison\"></a></p>\n","updatedAt":"2026-07-20T03:52:29.791Z","author":{"_id":"632cdea254e2c512c8f95b12","avatarUrl":"/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg","fullname":"Weiran Yao","name":"weirayao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.5801914930343628},"editors":["weirayao"],"editorAvatarUrls":["/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg"],"reactions":[{"reaction":"🔥","users":["actava-ai","moregpuspls","Emmawxx"],"count":3},{"reaction":"🚀","users":["actava-ai","moregpuspls","Emmawxx"],"count":3},{"reaction":"😎","users":["actava-ai"],"count":1},{"reaction":"👍","users":["Emmawxx"],"count":1}],"isReport":false}},{"id":"6a5d9ba5c028c3a08d69c4f6","author":{"_id":"632cdea254e2c512c8f95b12","avatarUrl":"/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg","fullname":"Weiran Yao","name":"weirayao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false},"createdAt":"2026-07-20T03:53:09.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"How we trained it: RSI (recursive self-improvement).\n\nEach iteration, a training agent plans a target capability, trains the model, evaluates the graded benchmark trajectories, and data agent synthesizes the next data mixture from the failure modes it finds.\n\nHumans gate every keep-or-revert decision. Reverted rounds stay in the record. One raised headline scores while quietly damaging a held-out subset, so we threw it out. The kept rounds add up: +14.6 on HealthBench Hard, +15.9 on HealthBench Professional, +9.3 on MedAgentBench.\n\nhttps://cdn-uploads.huggingface.co/production/uploads/632cdea254e2c512c8f95b12/BsiaKO-bgZdKp-Gc1koNj.mp4\n","html":"<p>How we trained it: RSI (recursive self-improvement).</p>\n<p>Each iteration, a training agent plans a target capability, trains the model, evaluates the graded benchmark trajectories, and data agent synthesizes the next data mixture from the failure modes it finds.</p>\n<p>Humans gate every keep-or-revert decision. Reverted rounds stay in the record. One raised headline scores while quietly damaging a held-out subset, so we threw it out. The kept rounds add up: +14.6 on HealthBench Hard, +15.9 on HealthBench Professional, +9.3 on MedAgentBench.</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/632cdea254e2c512c8f95b12/BsiaKO-bgZdKp-Gc1koNj.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-20T03:53:09.960Z","author":{"_id":"632cdea254e2c512c8f95b12","avatarUrl":"/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg","fullname":"Weiran Yao","name":"weirayao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8417842984199524},"editors":["weirayao"],"editorAvatarUrls":["/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg"],"reactions":[{"reaction":"🤯","users":["actava-ai","moregpuspls","Emmawxx"],"count":3},{"reaction":"🔥","users":["actava-ai","Emmawxx"],"count":2}],"isReport":false}},{"id":"6a5db5bdbeca2a1896f6b018","author":{"_id":"632cdea254e2c512c8f95b12","avatarUrl":"/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg","fullname":"Weiran Yao","name":"weirayao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false},"createdAt":"2026-07-20T05:44:29.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"@librarian-bot recommend","html":"<p><span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{"user":"librarian-bot"}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/librarian-bot\">@<span class=\"underline\">librarian-bot</span></a></span> </span></span> recommend</p>\n","updatedAt":"2026-07-20T05:44:29.779Z","author":{"_id":"632cdea254e2c512c8f95b12","avatarUrl":"/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg","fullname":"Weiran Yao","name":"weirayao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7918877601623535},"editors":["weirayao"],"editorAvatarUrls":["/avatars/a6d06cdd75861ae7d589f1343d81a5c5.svg"],"reactions":[],"isReport":false},"replies":[{"id":"6a5dbb8e970d16583af41530","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":376,"isUserFollowing":false},"createdAt":"2026-07-20T06:09:18.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [AutoMedBench: Towards Medical AutoResearch with Agentic AI Models](https://huggingface.co/papers/2606.01961) (2026)\n* [Towards Autonomous and Auditable Medical Imaging Model Development](https://huggingface.co/papers/2607.10522) (2026)\n* [HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents](https://huggingface.co/papers/2606.31179) (2026)\n* [Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance](https://huggingface.co/papers/2606.18613) (2026)\n* [EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning](https://huggingface.co/papers/2606.23301) (2026)\n* [MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning](https://huggingface.co/papers/2605.26567) (2026)\n* [CAREAgent: Clinical Agent with Structured Reasoning and Tool-Integrated for Order Generation](https://huggingface.co/papers/2606.01094) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.01961\">AutoMedBench: Towards Medical AutoResearch with Agentic AI Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.10522\">Towards Autonomous and Auditable Medical Imaging Model Development</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.31179\">HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.18613\">Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.23301\">EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2605.26567\">MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.01094\">CAREAgent: Clinical Agent with Structured Reasoning and Tool-Integrated for Order Generation</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-07-20T06:09:18.592Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":376,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7607210874557495},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false,"parentCommentId":"6a5db5bdbeca2a1896f6b018"}}]}],"primaryEmailConfirmed":false,"paper":{"id":"2607.15314","authors":[{"_id":"6a5d83ee6a69ce099f4d6d91","name":"actAVA AI","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d93","name":"Haolin Chen","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d94","name":"Leon Qi","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d95","name":"Steve Brown","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d96","name":"Deon Metelski","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d97","name":"Tao Xia","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d98","name":"Joonyul Lee","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d99","name":"Qixuan Wang","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d9a","name":"Kevin Riley","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d9b","name":"Frank Wang","hidden":false},{"_id":"6a5d83ee6a69ce099f4d6d9c","name":"Weiran Yao","hidden":false}],"publishedAt":"2026-07-15T00:00:00.000Z","submittedOnDailyAt":"2026-07-20T00:00:00.000Z","title":"Cura 1T: Specialized Model for Agentic Healthcare","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.","upvotes":32,"discussionId":"6a5d83ee6a69ce099f4d6d9d","projectPage":"https://actava.ai/cura","githubRepo":"https://github.com/actava-ai/Cura","githubRepoAddedBy":"user","githubStars":13},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"68edc310a4f606a8123967e7","avatarUrl":"/avatars/908302a1d06a5e0984af74f483b7ed69.svg","isPro":false,"fullname":"Weiran Yao","user":"weiran-actava","type":"user"},{"_id":"661573234c2f29635e93bb71","avatarUrl":"/avatars/fba95e382454485766b6349d6281b715.svg","isPro":false,"fullname":"Weiran Yao","user":"weiranyao","type":"user"},{"_id":"68edc7f6a17c9d2e3dba81b2","avatarUrl":"/avatars/e0cc4e300e9a2875f866ab04bfc7d4b9.svg","isPro":false,"fullname":"ggpd","user":"moregpuspls","type":"user"},{"_id":"6a14ed379224246011a04d87","avatarUrl":"/avatars/e3d1d655338bf264092bb1101e65f896.svg","isPro":false,"fullname":"Tao Xiong","user":"taoxiong-wch","type":"user"},{"_id":"6a14edbef05031c131840529","avatarUrl":"/avatars/862eae50cafb191b80b26ff8397aa3b7.svg","isPro":false,"fullname":"Yi Yang","user":"dryiyang","type":"user"},{"_id":"65489653cbe50f378d94f14a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65489653cbe50f378d94f14a/Bhl3yRJ9lYCc7HrYj9jJf.png","isPro":false,"fullname":"Leon Qi","user":"leonqi","type":"user"},{"_id":"64bb3e754b4ff0d509141938","avatarUrl":"/avatars/81617fd676a3d3133811348055e68fc7.svg","isPro":false,"fullname":"Zhepeng Cen","user":"czp16","type":"user"},{"_id":"6a0c6ae1d575c4c20ae1bae9","avatarUrl":"/avatars/27fee9022ef159138de77c1113f620cf.svg","isPro":false,"fullname":"Zhiwei Liu","user":"jimliu96","type":"user"},{"_id":"6553cd2296c5902fa386f57f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6553cd2296c5902fa386f57f/AEDbKbifdmDW4eJf6pUqw.png","isPro":false,"fullname":"Zeyu Tang","user":"zeyutang","type":"user"},{"_id":"5f12485c0c833276f61f1afb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1595033594228-noauth.jpeg","isPro":false,"fullname":"Xiangchen Song","user":"xiangchensong","type":"user"},{"_id":"64c49396bf1954890192bfbf","avatarUrl":"/avatars/b2dbd1d601911440788efcea2a2b77d3.svg","isPro":false,"fullname":"Yutong Dai","user":"UncleFish","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.15314.md","query":{}}">
Cura 1T: Specialized Model for Agentic Healthcare
Abstract
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.
Community
Cura 1T leads frontier models on 5 of 6 hardest healthcare benchmarks:
- HealthBench Hard: 36.8 (GPT-5.5: 31.5)
- HealthBench Professional: 66.2 (Claude Fable 5: 66.0)
- MedXpertQA-Text: 60.0 (GPT-5.5: 59.6)
- MedXpertQA-Multimodalt: 72.2 (GPT-5.5: 77.1)
- AgentClinic: 79.6 (Claude Opus 4.8: 79.4)
- MedAgentBench-v2: 94.0 (Claude Opus 4.8: 93.7)

How we trained it: RSI (recursive self-improvement).
Each iteration, a training agent plans a target capability, trains the model, evaluates the graded benchmark trajectories, and data agent synthesizes the next data mixture from the failure modes it finds.
Humans gate every keep-or-revert decision. Reverted rounds stay in the record. One raised headline scores while quietly damaging a held-out subset, so we threw it out. The kept rounds add up: +14.6 on HealthBench Hard, +15.9 on HealthBench Professional, +9.3 on MedAgentBench.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.