[PAPER] GPQA, MMLU-Pro, and MMMU-Pro were audited for broken questions, and up to 12% of them had to be removed. New drop in clean versions released
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I was very curious why all the models were topping out on GPQA-Diamond around 92 or 93% (AA) and spent the last few weeks pouring over GPQA (Diamond and Extended), and then expanded to auditing MMLU-Pro and MMMU-Pro. It was quite frankly shocking just how many questions were malformed, had wrong answer keys, or questions with more than one realistic answer. In fact, on GPQA-Extended, MMLU-Pro and MMMU-Pro, ~12% of questions were verifiably broken! Once fixed, the top models hit around 98%. Full paper released here: https://github.com/adamallcock/answer-key-audit As part of this process, I have also shopped I would love any feedback you have on the paper, and what benchmarks I should look at next. PS: There are some verbatim examples of broken questions on page 28 if you want to take a look. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.