GoBench: Evaluating LLMs on the game of Go [R]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| GoBench evaluates LLMs on 9x9 Go games against a ladder of KataGo opponents, from random to superhuman. It measures general reasoning ability, strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated. GPT-6 Astra max achieves 2500 Elo, much lower than the best KataGo, which achieves 4400 Elo. With coding tools and two hours of preparation before evaluation, Codex with Astra achieves 3560 Elo. I will keep the leaderboard updated as long as it is not saturated. leaderboard: https://rolandgao.com/blog/gobench/ code: https://github.com/RolandGao/gobench paper: https://github.com/RolandGao/gobench/blob/main/paper/gobench2.pdf x: https://x.com/Roland65821498/status/2100253388562723298?s=20 [link] [comments] |
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.