Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents - Aparna Dhinakaran
Jun 10, 2025 · 14:25
Aparna Dhinakaran, co-founder of Arise, details how to build a self-improving evaluation stack for AI agents using her company's open-source tools Arise Phoenix and RiseX. She argues that production agents require three evaluation layers: tool-call correctness (right function and arguments), trajectory accuracy (correct order of steps), and multi-turn session consistency (context retention). Dhinakaran demonstrates with real traces from Arise's own Copilot, showing how a bottleneck in search Q&A correctness (only 50% accuracy) was traced to incorrect argument passing in a tool call. She stresses a dual iteration loop: improving agent prompts and simultaneously refining eval prompts (LLM-as-judge) to avoid static evaluation criteria. The talk concludes that continuous eval improvement is essential for agents to move from working once to working reliably in production.