GoBench: Evaluating LLMs on the game of Go [R]
What Is GoBench?
GoBench evaluates LLMs on 9x9 Go games against a ladder of KataGo opponents, ranging from random to superhuman. It measures general reasoning ability.
Key Findings
- GoBench strongly correlates with ARC-AGI 2 (r=0.83 correlation).
- The benchmark remains highly unsaturated.
- GPT-6 Astra max achieves 2500 Elo, much lower than the best KataGo, which achieves 4400 Elo.
- With coding tools and two hours of preparation before evaluation, Codex with Astra achieves 3560 Elo.
The author will keep the leaderboard updated as long as it is not saturated.
Comments
No comments yet. Start the discussion.