A product discussed on AI Engineer.

Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
May 25, 2026 · 20:03
Nicholas Kang and Michael Aaron from Google DeepMind's Kaggle team argue that AI evaluations are broken due to being scattered, stale, and lacking transparency, citing a competing lab publishing inflated results by using custom compaction settings. They introduce four solutions: hackathons to channel community expertise, a standardized agent exam that returned 500+ submissions in its first week without promotion, a Game Arena where models play poker, chess, and werewolf for an ELO rating that cannot saturate, and an open benchmarks platform. A wastewater treatment plant engineer in Turkey built a novel safety benchmark from 20 years of field experience. They note that on SWE-Bench Pro, six frontier models land within a couple of percentage points, but the harness shifts performance by 22%, complicating comparisons. Challenges include high cost (400,000 poker hands for statistical significance), maintaining community engagement, and dealing with fast model deprecation cycles.

Shipping complex AI applications — Braintrust & Trainline
May 1, 2026 · 1:38:34
Giran Moodley of Braintrust, joined by Trainline's Oussama Hafferssas and Mayank Soni, demonstrate how to ship production-grade multi-step AI agents by combining rigorous tracing, evaluation with golden datasets, and automated online scoring. They walk through building a support triage agent that progresses from a single prompt to a five-stage tool-calling pipeline, where tracing captures latency, tokens, and costs per step for deep debugging. Trainline shares using Braintrust to run offline evaluations before switching LLM models and to enable cross-functional self-service. The workshop covers identifying failure modes via production logs, tightening prompts, and re-evaluating to complete the feedback loop, emphasizing that observability and iterative evaluation are essential for moving from prototype to reliable production systems.

Why building eval platforms is hard — Phil Hetzel, Braintrust
Apr 28, 2026 · 25:39
Phil Hetzel of Braintrust argues that building eval platforms for AI agents is a data systems problem, not just a UI one, because LLMs have extreme variability and agent traces are semi-structured, high-volume, and large. He describes four maturity stages: simple spreadsheet with a for loop, a custom vibe-coded UI with a database, an experimentation playground for non-technical users, and finally a flywheel connecting offline evals with production observability. The core challenge is the data layer—traces can be 10-20 MB per span, requiring low-latency ingestion, aggregate analysis, and full-text search, which traditional databases cannot handle. Future platforms must surface unknown unknowns via topic modeling, support agent-to-agent interactions, and integrate automatic tracing through AI proxies. Braintrust addresses these with a custom data platform that separates hot and cold storage and enables SQL queries directly on traces.

The Future of Evals - Ankur Goyal, Braintrust
Aug 9, 2025 · 5:14
Ankur Goyal, CEO of Braintrust, argues that evals are being revolutionized by AI agents like Loop, which automatically optimizes prompts, datasets, and scorers. He notes that the average org runs 13 evals daily, with some exceeding 3,000, yet eval workflows remain painfully manual. Loop, powered by frontier models such as Claude 4—which Goyal says performs six times better than prior models—can now autonomously improve prompts and scoring. It runs inside Braintrust, allowing users to review suggested edits side-by-side or enable a fully automated mode. Goyal emphasizes that evals are critical for building reliable AI products, and Loop marks a shift from manual dashboards to AI-driven iteration.

Evals Are Not Unit Tests — Ido Pesok, Vercel v0
Aug 6, 2025 · 15:22
Ido Pesok, an engineer at Vercel working on v0, argues that evals—not unit tests—are the key to making LLM applications reliable, using the basketball court analogy of understanding your domain (the court), collecting user queries (data points), and scoring passes/fails. He illustrates with the Fruit Letter Counter app, which failed in production despite passing manual tests, because LLMs are inherently non-deterministic. Pesok advises collecting data from thumbs-up/down, logs, and forums, and structuring evals with constants in data and variables in task (e.g., system prompt). He emphasizes deterministic pass/fail scoring, adding extra prompt tags for easy parsing, and integrating evals into CI/CD to detect regressions when changing models or prompts. The episode stresses that improvement without measurement is limited, and evals enable systematic reliability gains.

[Evals Workshop] Mastering AI Evaluation: From Playground to Production
Jul 1, 2025 · 1:25:08
In this workshop, Braintrust Solutions Engineers Carlos Esteban and Doug guide participants through the complete AI evaluation lifecycle, from offline testing in the playground to production monitoring. They explain the three core ingredients of an eval—task, data set, and score—and demonstrate how to run evals both via the Braintrust UI and the SDK. The session covers LLM-as-judge vs. deterministic code scores, the importance of starting small with synthetic data, and how to use online scoring and logging to capture real user feedback. Human-in-the-loop review is highlighted as a way to establish ground truth and close the feedback loop. The presenters also address audience questions on bootstrapping data sets, non-determinism in LLM judges, and integrating evals into existing projects.

Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier
Jun 30, 2025 · 16:15
Zapier AI Tech Lead Rafal Willinski and Staff Engineer Vitor Balocco explain how Zapier's evaluation system turns agent failures into targeted improvements through a data flywheel. They detail collecting explicit feedback at critical moments and mining implicit signals like testing behavior, cursing, and user follow-ups. The pair advocates building unit test evals for specific failure modes, then trajectory evals and LLM-as-judge with rubrics to avoid overfitting and capture multi-turn criteria. They share that over-indexing on unit tests hurt model benchmarking, and reasoning models can compare model runs, revealing differences like Claude as a decisive executor versus Gemini's yapping. Ultimately, they argue that the goal is user satisfaction, so A/B testing on a small traffic fraction is the ultimate verification.

How to build world-class AI products — Sarah Sachs (AI lead @ Notion) & Carlos Esteban (Braintrust)
Jun 27, 2025 · 1:43:46
Sarah Sachs (AI lead at Notion) and Carlos Esteban (Braintrust) explain that building great AI products requires 10% prompting and 90% evals and observability, with Notion AI using Braintrust to iterate on prompts and models. Sachs details their cycle: curate small datasets, tie them to scoring functions (LLM-as-a-judge with per-sample prompts and heuristic checks), run evals before shipping, and use production logs to catch regressions. She notes Notion AI supports 100M+ users, switches models in under a day, and 60% of enterprise users are non-English—requiring multilingual eval rigor. Carlos and Doug then walk through Braintrust's framework: tasks (prompts, tools, agents), datasets, and scores (0-1). They demonstrate offline evals via playground and SDK, online scoring in production, user feedback capture, human review setup, and remote evals to bridge complex code with the playground. The workshop covers moving from pre-prod evals to production monitoring, closing the feedback loop by adding underperforming spans to datasets.
Powered by PodHood