Jev's calibration was measured. The LLMs won [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| Source: Jev Benchmarks Its training method is literally called "Reinforcement Learning for Calibrated Decisions." Calibration gap vs human labels (lower = better): Yes/no: Jev 5.0, Gemini 3.8 Flash 2.0 Pick-one: Jev 9.8, DeepSeek V4.1 Flash 2.8 Rubric: Jev 19.7, GLM-5.3 12.9 It held to 95% accuracy, Jev still handles more decisions alone than any of them (86% of yes/no). Worse calibrated, better at knowing when it's right. [link] [comments] |
More from r/MachineLearning
-
I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6x–4x faster than backprop)[R]
Sep 21
-
For NeurIPS: Is Paris or Syndey better for networking with U.S. tech companies? [D]
Sep 21
-
Systems for Machine Learning[D]
Sep 21
-
These Were NOT Rogue AI Escapes. Just SLOPPY Firewall Failures. [N]
Sep 21
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.