Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
May 25, 2026 · 20:03
Nicholas Kang and Michael Aaron from Google DeepMind's Kaggle team argue that AI evaluations are broken due to being scattered, stale, and lacking transparency, citing a competing lab publishing inflated results by using custom compaction settings. They introduce four solutions: hackathons to channel community expertise, a standardized agent exam that returned 500+ submissions in its first week without promotion, a Game Arena where models play poker, chess, and werewolf for an ELO rating that cannot saturate, and an open benchmarks platform. A wastewater treatment plant engineer in Turkey built a novel safety benchmark from 20 years of field experience. They note that on SWE-Bench Pro, six frontier models land within a couple of percentage points, but the harness shifts performance by 22%, complicating comparisons. Challenges include high cost (400,000 poker hands for statistical significance), maintaining community engagement, and dealing with fast model deprecation cycles.