A product discussed on AI Engineer.

Five hard earned lessons about Evals — Ankur Goyal, Braintrust
Aug 23, 2025 · 19:46
Ankur Goyal of Braintrust argues that successful AI applications depend on deliberately engineered evaluations (evals) to reflect real user feedback and drive product improvements. He details three signs of effective evals: launching updates within 24 hours (citing Notion), converting complaints into evals, and using evals offensively to assess use cases before shipping. Evals require custom scorers as specs—not off-the-shelf—and context engineering (optimizing tool definitions and outputs) is the new frontier; shifting outputs from JSON to YAML improves token efficiency. New models can upend everything, as shown by a benchmark that jumped from 10% to viable with Claude 4 Sonnet, so a model-agnostic architecture with proxies is key. Optimize the whole system (data, task, scoring)—Braintrust's Loop auto-optimizes prompts, data, and scorers. In Q&A, he advises human judgment when adding user feedback to evals to avoid overfitting.

The Future of Evals - Ankur Goyal, Braintrust
Aug 9, 2025 · 5:14
Ankur Goyal, CEO of Braintrust, argues that evals are being revolutionized by AI agents like Loop, which automatically optimizes prompts, datasets, and scorers. He notes that the average org runs 13 evals daily, with some exceeding 3,000, yet eval workflows remain painfully manual. Loop, powered by frontier models such as Claude 4—which Goyal says performs six times better than prior models—can now autonomously improve prompts and scoring. It runs inside Braintrust, allowing users to review suggested edits side-by-side or enable a fully automated mode. Goyal emphasizes that evals are critical for building reliable AI products, and Loop marks a shift from manual dashboards to AI-driven iteration.
Powered by PodHood