Why Agent Hype can fall short of reality – Joel Becker, METR
Dec 24, 2025 · 21:22
Joel Becker, a researcher at METR, examines why AI models that ace benchmarks like SWE-bench fail to boost real-world developer productivity. METR's time horizon measurements show AI capabilities doubling every 6-7 months, with Claude 3.7 Sonnet achieving a 50% success rate on tasks taking humans 4 minutes. Yet a randomized controlled trial of 16 experienced developers on large open-source projects found they were 19% slower when using AI tools like Cursor Pro, contradicting expert predictions of 40% time savings. Becker attributes the gap to high context requirements, low AI reliability, and task complexity—developers spent significant time verifying and correcting AI outputs. The episode warns that benchmark-style evidence, which uses low-context human baselines, overstates AI readiness for messy, interdependent real-world tasks.