r/MachineLearning · · 1 min read

Jev's calibration was measured. The LLMs won [D]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

Jev's calibration was measured. The LLMs won [D]

Source: Jev Benchmarks

Its training method is literally called "Reinforcement Learning for Calibrated Decisions."

Calibration gap vs human labels (lower = better): Yes/no: Jev 5.0, Gemini 3.8 Flash 2.0 Pick-one: Jev 9.8, DeepSeek V4.1 Flash 2.8 Rubric: Jev 19.7, GLM-5.3 12.9

It held to 95% accuracy, Jev still handles more decisions alone than any of them (86% of yes/no).

Worse calibrated, better at knowing when it's right.

submitted by /u/frappuccinoCoin
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning