What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Ant Ling reports 83.83 on DiagnosisArena-MCQ for Ling-3.0-flash-Sante, its new medical reasoning model. The suffix matters: the task provides case information, examinations and tests, then asks the model to choose from four diagnoses.
That result tells us about selecting an answer when the candidate set and case evidence are supplied. It does not establish how the same model would generate an unrestricted differential, decide what history is missing, or choose which investigation to request next. Those would require different evaluations.
The release also reports two other medical results:
| Evaluation | Sante result | What the task adds |
|---|---|---|
| MedXpertQA-Text | 53.88 | Challenging medical questions in a text subset. |
| HealthBench Professional | 45.73 | Open-ended professional clinical chat, assessed with physician-written rubrics. |
The published HealthBench Professional definition includes care consultation, writing/documentation and medical research. Its score is not percentage accuracy. The Sante chart does not provide enough scoring detail to identify the reported value as length-adjusted or unadjusted, so a comparison with another published HBP result would need that checked first.
This is why the three results are useful together. They give Sante a broader medical-text evaluation profile than an exam score alone, while leaving specific questions open. For a case-answering application, the first decision is whether users supply the alternatives or expect the model to construct them. The release supports including Sante in that evaluation; the 83.83 figure applies to the supplied-options version.
[link] [comments]
More from r/MachineLearning
-
I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6x–4x faster than backprop)[R]
Sep 21
-
For NeurIPS: Is Paris or Syndey better for networking with U.S. tech companies? [D]
Sep 21
-
Systems for Machine Learning[D]
Sep 21
-
These Were NOT Rogue AI Escapes. Just SLOPPY Firewall Failures. [N]
Sep 21
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.