# Why should anyone care about Evals? — Manu Goyal, Braintrust

AI Engineer · 2025-06-27

<https://aie.addtry.com/a2aa510c-76c3-477c-9b24-040dcd172a7a>

Manu Goyal, founding engineer at Braintrust, argues that evals are not just unit tests for AI but a critical tool for building a laboratory that enables 90% of product iteration before shipping to production. Drawing from his personal journey from a disappointed Nintendo-playing child to a self-driving car engineer at Nuro, he explains how evals provide the necessary signal to contextualize model improvements and reduce risk. He describes Braintrust's platform as a dev environment that combines tweaking prompts, logging data, and observability to create a data flywheel for AI development. Citing tech luminaries like Kevin Weil, Gary Tan, Mike Krieger, and Greg Brockman, he asserts that evals are the key to industry transformation and successful AI deployment. The talk concludes with an invitation to the Evals Track at the conference, emphasizing his message of 'evals, evals, evals'.

## Questions this episode answers

### Why should anyone care about evals in AI?

Manu Goyal explains that without evals, shipping AI to production is slow, expensive, and risky because you only get signal from real-world usage. Evals let you run experiments extensively before deploying, acting like a lab where you can do 90% of the iteration loop safely. This approach, used in self-driving car development, ensures models handle real-world scenarios like avoiding pedestrians, not just improving classification metrics.

[1:40](https://aie.addtry.com/a2aa510c-76c3-477c-9b24-040dcd172a7a?t=100000)

### How do evals speed up product iteration?

According to Manu, evals are more than regression tests; they build a laboratory for AI. Instead of shipping every change to production and waiting for feedback, you can test modifications extensively offline. This covers 90% of iteration before going live, allowing much faster and more confident releases. It transforms the development process from risky deployment cycles to rapid, data-driven experimentation.

[2:36](https://aie.addtry.com/a2aa510c-76c3-477c-9b24-040dcd172a7a?t=156000)

### How can evals create a data flywheel for AI improvement?

Manu and Braintrust advocate connecting offline evaluation metrics to online production data. By using the same metrics across both, you can pinpoint which production examples are most valuable for retraining or refinement. This creates a data flywheel: offline evals guide iteration, online performance flags key data to feed back into development, and the loop accelerates AI improvement over time.

[3:15](https://aie.addtry.com/a2aa510c-76c3-477c-9b24-040dcd172a7a?t=195000)

## Key moments

- **[0:00] Intro**
- **[0:25] Childhood Realization**
  - [0:44] Manu Goyal's childhood Nintendo 64 disappointment with rule-based systems inspired his AI career.
- **[1:20] Self-Driving Insight**
  - [1:40] Self-driving car development reveals that model tuning alone is insufficient without evals to assess real-world performance.
- **[2:19] Evals Lab**
  - [2:53] Evals build a laboratory enabling 90% of the product iteration loop before production, says Manu Goyal.
- **[3:42] Industry Validation**
- **[4:05] Braintrust**
- **[4:53] Evals Chant**
  - [4:53] "The key to industry transformation, the key to success, it's evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals. Woo!"

## Speakers

- **Manu Goyal** (guest)

## Topics

Evaluation Frameworks

## Mentioned

Braintrust (company)

## Transcript

### Intro

Whoo.

**Manu Goyal** [0:20]
Allright, who's excited about evals?

Whoo!

### Childhood Realization

**Manu Goyal** [0:25]
Allright, what can I do to get those juices flowing? Uh, I'm Manu, and I work at Braintrust, where we build a platform to do evals and a whole bunch of other stuff. Um, so I was thinking we could just start by, uh, talking a little bit my— about my own personal evals journey.

Now, you might see this picture and say, "Oh, what an adorable little boy absorbed in his Nintendo 64 video game." But if you look a little closer, you'll see a boy who's deeply disappointed with the state of technology in his society.

Because this boy, he knows that technology is not meant to be shackled to the constraints of rule-based systems doomed to do the same thing over and over and over. No, technology is meant to come alive, to grow and adapt, and really be a thought partner to mankind.

### Self-Driving Insight

**Manu Goyal** [1:20]
So I knew this in this moment, which is why I decided to devote my career to being a software engineer in the AI industry. And so I dropped the Nintendo and I started grinding away on LeetCode, and soon enough, I landed a job in the self-driving car industry.

Now, we can all learn a lot about self-driving cars, but the thing I took away was that, you know, you can spend all day tuning the model, changing the architecture, you know, adjusting the loss function, all good stuff, but it's never going to be enough for you to actually ship it to production,right?

I can't say, "Oh, my image classification rate went from 98% to 99%. Put it on the road." Right? We need some way to, you know, contextualize this model and understand if it actually works for our real-world application. You know, does it avoid pedestrians?

Does it negotiate traffic scenarios appropriately? Does it obey the law? All this stuff we actually need to understand. And how we're going to do that is with evals. Now, you know, the whole point here is, you know, evals aren't just unit tests for AI.

### Evals Lab

**Manu Goyal** [2:36]
They're not just for finding regressions,right? If I didn't have evals, the only way I can get any signal on my changes is by shipping it to prod and then getting signal, you know, in the real world. But that's expensive, it's slow, and ultimately it's pretty risky.

So what do evals do is, it's kind of like if you invest in good evals, you're kind of building a laboratory that lets you run experiments to your heart's content and do 90% of the product iteration loop before going to prod, and then now you can ship much more quickly, much more confidently.

Um, now, furthermore, if you actually apply the same metrics from offline to your online production data, you now have data-driven signal about which examples in prod are going to be most useful for that next iteration loop. And so with all of this knowledge, I was— my evals journey had completed, and I transformed from this guy to this guy.

So, success. Now, if this heartfelt childhood story isn't enough to do it for you, you don't have to take my word. You can take the words of all of these tech luminaries. We have Kevin Weil, Gary Tan, Mike Krieger, Greg Brockman, all extolling the virtues and the necessities of evals.

### Braintrust

**Manu Goyal** [4:05]
And surely, if they're all saying it, there's got to be something to it. It can't be a total scam. So there's got to be some— there's got to be something worth checking out here. So with all that buzz, I made my way to Braintrust, where our goal is to sort of build the dev platform to, of course, let you do evals, but also do all the things that go along with it.

So that involves, you know, tweaking prompts and experimenting in the playground. It involves logging data and sort of getting the observability component and kind of connecting all those together in this beautiful data flywheel so that we can— we can let you build the data flywheel to let your AI dreams come true.

### Evals Chant

**Manu Goyal** [4:53]
Because that's really what— what we're here for. Now, I know this was a dense and content-heavy presentation, so I'll try to distill it with one simple message, which is that the key to industry transformation, the key to success, it's evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals, evals.

Woo! Allright. Thank you. Please join the Evals Track Golden Gate Ballroom B. I'll see you there.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
