A company discussed on AI Engineer.

Judge the Judge: Building LLM Evaluators That Actually Work with GEPA — Mahmoud Mabrouk, Agenta AI
Apr 10, 2026 · 40:51
Mahmoud Mabrouk, co-founder of Agenta AI, demonstrates how to build calibrated LLM-as-a-judge evaluators using the GEPA prompt optimization algorithm, arguing that miscalibrated evals are worse than none. He walks through a practical workflow for a customer support agent using the TaoBench airline dataset, covering metric design, data annotation, and GEPA-based optimization. The seed judge achieved 61% accuracy; after optimization, accuracy rose to 74% with reduced bias, though the judge still struggled to fully learn the complex policy. Mabrouk shares key lessons: start with a seed prompt biased toward compliance, use larger models for refinement, overfit to training data first, and beware of high token costs.

Rise of the AI Architect — Clay Bavor, Cofounder, Sierra w/ Alessio Fanelli
Jul 24, 2025 · 18:55
Clay Bavor, cofounder of Sierra, and Alessio Fanelli discuss the rise of the AI Architect—a new role combining technology, brand, and business outcomes to build customer-facing AI agents. Sierra serves hundreds of millions of consumers this year. Bavor defines the AI Architect as wearing three hats: understanding AI capabilities, defining the agent's voice (e.g., Chubbies' irreverent Duncan Smothers), and driving business outcomes. Successful AI Architects embrace risk, start with narrow problems like processing a single return, and re-architect teams to coach the AI. On build vs. buy, Bavor warns of the "agent iceberg"—hundreds of hidden complexities like regression testing and model migration. He advises tracking model improvement in a Google Doc and anticipating future capabilities, predicting glasses as the ultimate interface for trusted personal AI.

The Agent Development Life Cycle — Zack Reneau-Wedeen, Sierra
Apr 11, 2025 · 18:40
Zack Reneau-Wedeen from Sierra explains the company's Agent Development Lifecycle for building reliable, testable AI agents at scale for brands like Chubbys and SiriusXM. He contrasts LLMs' nondeterministic, slow nature with traditional software, calling them a 'foundation of Jello.' Sierra treats every agent as a product, using an experience manager to review conversations, file issues, create tests, and release improvements—growing from hundreds of thousands of requests for Chubbys to tens of millions for larger customers. The lifecycle spans quality assurance, testing, and deployment, with reasoning models acting as a force multiplier. Voice agents launched generally in October 2024, handling calls with the same underlying platform, enabling responsive design across channels.
Powered by PodHood