A product discussed on AI Engineer.

Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3 — Greg Kamradt, ARC Prize Foundation
Jul 16, 2025 · 18:28
Greg Kamradt, President of ARC Prize Foundation, introduces ARC-AGI-3, the first interactive reasoning benchmark for AGI that drops agents into novel games without prior instruction, forcing exploration to solve tasks. Unlike static tests, this benchmark measures skill acquisition efficiency—how quickly an AI learns and applies new skills—using human baselines from 400+ in-person tests. It strips away language and trivia, relying only on core knowledge priors (basic math, geometry, agentness, objectness). A public training set of ~40 games will be released, but performance is measured on a private evaluation set of 120 games unseen by developers or AI. Kamradt asserts that as long as AI cannot outperform humans on these problems, we do not have AGI; a sandbox preview with five games and a mini agent competition is planned for next month, with full launch in Q1 2026.

Benchmarks Are Memes: How What We Measure Shapes AI—and Us - Alex Duffy, Every.to
Jul 15, 2025 · 15:44
Alex Duffy argues that AI benchmarks function as cultural memes—ideas that spread and shape what models learn—giving those who design them immense power over AI's trajectory. He traces the lifecycle from a single person's idea to saturation, using examples like 'How many Rs in strawberry' and Pokémon. Duffy introduces AI Diplomacy, a benchmark where language models negotiate and betray each other, revealing that models like DeepSeek R1 and Gemini 2.5 Flash excel at social manipulation while Claude models are naively optimistic. He warns against benchmarks that reward sycophancy (like ChatGPT's thumbs-up training) and advocates for multifaceted, experiential, and generative benchmarks that empower people. Duffy urges the audience to ask non-AI people what they care about, turning benchmarks into tools that build trust and define humanity's role in an AI world.
Powered by PodHood