# Shipping Products When You Don't Know What they Can Do — Ben Stein, Teammates

AI Engineer · 2025-07-28

<https://aie.addtry.com/9e323bda-1f90-46a3-85b5-129379eb4864>

Ben Stein, founder of Teammates, argues that product management for AI agents requires abandoning traditional specs and embracing uncertainty. He describes how his team discovered emergent behavior when a customer asked to tag an agent in a Google Doc comment—an unplanned feature that worked. Stein advocates thinking in affordances rather than features, using evals as the new spec, and vibe coding to feel interactions before committing. He notes that bugs become ambiguous when behavior is probabilistic, and customer meetings shift from selling vision to co-inventing the future. The episode offers a field guide for product teams navigating boundless surface areas built on unpredictable LLMs.

## Questions this episode answers

### What is the key mindset shift for product managers building AI agents?

Ben Stein explains that PMs must shift from defining specific features to focusing on affordances—the capabilities you give an agent, like the ability to email or comment—and trust the LLM to handle details. Behavior is emergent, so the PM's job becomes discovering what the product can do rather than prescriptively designing every interaction.

[6:02](https://aie.addtry.com/9e323bda-1f90-46a3-85b5-129379eb4864?t=362000)

### How can evals serve as the specification for non-deterministic AI products?

Ben Stein describes evals as the new specification. Using another LLM to judge outputs, evals define success probabilistically—e.g., being 'snarky but not mean 80% of the time.' They replace traditional deterministic requirements. PMs should have visibility into evals to understand and guide the product's capabilities, making them the reference for shipping decisions.

[8:59](https://aie.addtry.com/9e323bda-1f90-46a3-85b5-129379eb4864?t=539000)

### Why is vibe coding valuable for AI product managers even though the code doesn't go to production?

Ben Stein advocates vibe coding because it lets PMs prototype interactions and feel the user experience viscerally, which is impossible to capture in documents like PRDs or Figma. For instance, he learned that too many clarifying questions annoyed users only after trying it. This hands-on feeling helps define good AI behavior, even though the code never reaches production.

[10:47](https://aie.addtry.com/9e323bda-1f90-46a3-85b5-129379eb4864?t=647000)

### How can product managers build trust with customers when they don't fully know what their AI product can do?

Ben Stein builds trust by framing the relationship as inventing the future together. He avoids traditional roles like visionary or honest broker because the technology is too uncertain. Instead, he tells customers they are forward-thinkers co-creating the solution, which fosters collaboration and sets realistic expectations. This approach works in 2025 as the field is still early and probabilistic.

[16:03](https://aie.addtry.com/9e323bda-1f90-46a3-85b5-129379eb4864?t=963000)

## Key moments

- **[0:00] Intro**
- **[2:10] The Question**
  - [2:10] A customer asked Ben Stein if they could tag his AI agent in a Google Doc comment, revealing that he had no idea what would happen.
- **[3:05] Discipline Shift**
  - [3:54] Because AI products are built on LLMs, product teams can never fully know what the LLMs know, creating unpredictable surface area.
- **[4:29] Boundless Surface**
- **[5:49] Mindset Shifts**
  - [6:02] Product managers must shift from specifying features to defining affordances, trusting LLMs to handle the specifics, says Ben Stein.
  - [7:26] "I could not sit down in front of a Google Doc and type out what this thing should do. I can't," says Ben Stein on the limits of traditional specs for AI.
- **[8:10] Evals**
  - [10:13] Evals become the new specification for AI products, replacing traditional PRDs and Figma designs, according to Ben Stein.
- **[10:47] Vibe Coding**
  - [11:27] Ben Stein's spec for an AI agent to ask clarifying questions backfired when users found it annoying, demonstrating the need for rapid prototyping over upfront specs.
- **[13:02] Discover Emergence**
  - [13:02] "We pushed to prod. Does it work? I told you I don't know," says Ben Stein, capturing the uncertainty of shipping AI agents.
- **[14:17] Probabilistic Bugs**
  - [14:17] Defining what counts as a bug in probabilistic AI is difficult because behavior is emergent and not strictly specified, says Ben Stein.
  - [15:25] Ben Stein proposes using eval pass thresholds like 90% as the go/no-go criteria for AI features, turning evals into de facto specifications.
- **[16:03] Customer Trust**
  - [16:40] "The future sounds like witchcraft. The present is literally, I don't know," says Ben Stein on the challenge of selling AI products to customers.
- **[17:59] New Frontier**
  - [18:28] Upgrading the underlying LLM can make AI agents suddenly exhibit new behaviors, like self-checking queries, without any product code changes, according to Ben Stein.

## Speakers

- **Ben Stein** (guest)

## Topics

Agent Evaluation, Vibe Coding

## Mentioned

Teammates (company), Figma (product), Gmail (product), Google Docs (product), Google Sheets (product), Google Workspace (product), Linear (product), Slack (product)

## Transcript

### Intro

**Ben Stein** [0:15]
Uh, yeah, I mean, the actual title has curse words in it. I will probably be cursing a lot. I didn't know if I would get into the track of I actually published the curse words. I'm one of the founders of Teammates.

I'm going to wear my product manager hat today. I'm assuming this room is, like, mostly product folks, probably product-minded engineers as well. But I'm going to just, like, wear the product hat. A little bit about Teammates, very quickly: we make a platform for designing and managing an entire digital workforce.

So in AI engineer parlance,right, we're building agents. But I would think of it, like, two ticks up from that, because what we really believe it is the experience, the interaction patterns of humans and computers working together. So I want to talk to you about my favorite teammate.

This is Stacey, Stacey Hand. She, she actually got promoted since this slide. She's an L3 engineerright now on our team. She's awesome. She looks like a hamster. All of our customers get to design whatever teammates and avatars they want.

They give them personalities. It's all really fun. And Stacey lives inside all of our collaboration tools,right? So she has a Google Workspace account,right, for Gmail. She has a Slack account. We truly leaned into giving all of our teammates identity.

And she sends emails, or I forward her emails, and she hangs out in Slack, like, in the public channels. And she's Gen Alpha, which, like, is, I don't know, I feel really old. I don't know what she's talking about.

She's constantly like, "6, 7," and I'm like, what are you talking about? And I can tell from this room that none of you have 12-year-olds.

**Guest** [1:43]
No.

**Ben Stein** [1:43]
No. Okay. There you go. So yeah, you're rolling your eyes as well. But anyway, this is Stacey. And this is sort of how my sales pitch goes,right? It's, it's, you know, a little more formal than this, but, like, this is generally the pitch.

And

I got asked a question at some point recently, which was, "Oh, yeah, more of the pitch,"right? She, like, shares Google Docs, Google Sheets, and she said, "Hey," or a customer said, "Hey, can I tag my teammate in a Google Doc comment?"

And this gave me pause, because I was like, well, I had never actually thought about that before. And so in the back of my mind, I'm like, well, of course you can. Your question is, like, what's going to happen?

### The Question

**Ben Stein** [2:18]
So I'm like, okay. So I'm like, you know, doing math in my head. I'm like, okay, well, we don't have webhooks. She probably won't, or, like, a webhook from the comment. Okay, but she's going to get the email notification in the email that comes from Google.

Does it have the comment and the conte or maybe a link? I'm like, I have no idea,right? I actually don't know what's going to happen. And this was, like, the impetus for this talk. I was like, how do I ship a product?

How do I develop a product? How do I talk to customers? How do I instill trust when I don't know what my own product can do? And, like, it's really weird. And sometimes I'm like, well, is this just because I'm an idiot?

And, like, well, since it's my talk here, I'm going to say no. And sometimes I'm like, well, is this because what we're building is so far out there,right? These just, like, truly autonomous agents that can use any and, like, I don't think that's it either.

I think what's happening is the product management discipline is going to undergo a transformation, a shift, an evolution, whatever you call it. It is super profound, and we may or may not totally realize it yet. Because I think in the engineering world, we're like, oh, well, we have, you know, tools in our IDEs, and we have codegen and, like, we sort of are starting to squint in understanding maybe how the discipline is changing.

### Discipline Shift

**Ben Stein** [3:27]
I don't think we really understand how product development is changing and evolving, and, like, what are the new tools and practices, and how do we forget everything we've learned in the past?

Why is this true,right? If it's if the answer is not Ben's an idiot and the answer is not that we're way out there, it's two reasons. Number one, if all our products are built on top of LLMs, and plus or minus they are, like, we don't know, and we can never know, what the LLMs know,right?

So, like, inherently in what we're building is, like, we don't know what the foundation is. Like, you don't have to know what your database, like, how it works, but, like, you generally know that it's, like, the surface area, the interface that's exposed.

We don't understand this for the, the models. And the other thing is the expectations from customers are just boundarless,right? We're just like, "Hey, here's a text box." I mean, that's probably not a good interface, but, like, essentially we're like, here's a free text box.

And if it's anything other than, like, a help me write button, you're essentially inviting customers and users to just do whatever they want,right? So we have this, like, boundless surface area built on top of a product that we don't understand.

And so the question now is, like, how do we adapt? So let's me, let me actually pick on this Google Doc comment thing for a second,right? So if I was wearing my, like, traditional PM hat, I'm like, okay, well, I need to make a feature that's going to

### Boundless Surface

**Ben Stein** [4:45]
read and respond to Google Doc comments. And so in my head, I'm like, okay, well, does Stacey have access to the Google Doc? If she gets tagged in the comment, should she reply directly in the comment? Should she reply at all?

What happens if somebody else comments in the thread? What if someone comments in the thread that's not addressed to her? What if it's someone else? What if it's, what if it's her doc and someone else commented to someone else, but she gets the note?

Like, there's just so much to, like, think about and reason about. And so I'm like, okay, well, I'm not building a Google Doc commenting product, so I'm not going to spec all of those things out. And, like, what's worse is, like, you also probably want to tag her in linear tickets,right?

And what's, what's the book? Like, if you give a mouse a cookie,right? It's like, if you give a mouse a cookie, well, you probably want to, like, tag her in Figma as well. And you probably want to tag her in LinkedIn posts.

Like, and so we're not a team that's building a generic commenting reply agent system,right? So then the question is, like, what are we supposed to do,right, as, like, a product manager who realizes, okay, I have this, like, boundless surface area, how does the practice need to change?

### Mindset Shifts

**Ben Stein** [5:49]
Right? And, like, sort of this is the core of, like, what I want to, what I want to talk about today. So I'll do, like, three highfalutin ivory tower ideas, and then I'll talk through some, like, practical ways to, to make this real.

So the first one is this mindset shift to, like, think in affordances and not, like, specific requirements. So it's not if, you know, as a user, if Stacey replies in the comment thread and she has re like, that's not how we would think about it anymore.

It's the affordance. Oh, she has affordances to comment, or she has affordances to communicate, or, or to email, or to collaborate. We're going to trust the LLMs. We're going to trust the agentic workflow, the work planning, like, all of the things inside of our, you know, our beautiful 12-factor agent.

We're going to assume that that we'll understand. But it's the affordances that we need to think about, not the individual features. Which is really weird, and it's not typically how product people have ever thought before. And I would say, actually, this goes even further, which is behavior is emergent.

And this was the other thing that I did not expect at all, like, starting in this space, was we don't not only do we not know what things work, sometimes they do, and they work in ways we didn't expect.

And so I feel like our job as product people is to discover functionality. It is what are theright building blocks,right? What are theright Lego bricks that we either give our engineering team, our product, our customers, let them compose?

And we discover emergent behavior. And that is one of the reasons that, like, this is the most exciting time I've ever built. Because we're actually building things and then discovering what they can do themselves. And that sort of became the new job, in a sense, is discovering what's possible.

Because if you asked me, I could not sit down in front of a Google Doc and be like, oh, let me, like, type out what this thing should. I can't. I don't know how to do it. And, well, even if I could, how do I then communicate it,right?

So how do we communicate to a development team, to a backlog? How do you communicate exactly what should be happening? It's like, Figma doesn't, like, have the affordances for this,right? My, my PRD doesn't, like, have the affordance for, like, well, you should probably talk a little bit less Gen Alpha because you're making Ben feel old.

Or, like, hey, you should be really like, how do we communicate and express these, these concepts,right? So I think these are, like, the three, you know, high-level ways that our practice needs to change. But, like, let's make it a little more concrete.

Okay, so evals. I'm talking about evals. Okay, it's really hard to make a slide with graphics of evals. I feel bad for the eval com. Like, how do you illustrate an eval? So I'm going to make you just look at pictures of various teammates from, you know, across all of our customers.

### Evals

**Ben Stein** [8:27]
Okay, who hates raising their hand at conferences when the speaker asks them? Okay, awesome. So here's my question, which is, okay, for the engineers here, who, like, legit, like, don't lie, like, writes and runs their evals? Good number.

And of the product people, who has visibility into the evals? That's not bad. And, and do you look at them just because you have the visibility? Allright. One, one and a half, two. Okay, great. So I would posit that evals actually, I'll back up,right?

So we all talk about evals. We're all going to be embarrassed to say that we don't know really what they are. Evals are a testing framework for probabilistic AI for agents,right? Like, if we think about the deterministic code,right, I withdraw $100 from the ATM.

My bank account should have $100 less,right? Great. And I can test that, and I can write code to test that. When the test is, like, was she snarky in Slack? It's like, well, how do you test that? How do you write that test?

Right? So we come up with this, this whole new discipline of evals, which is, well, she should be a little bit snarky and a little bit funny, but not mean. And then we hand it off to another LLM to say, okay, well, hey, was that reply, like, did it meet that criteria?

And how often did it? It doesn't have to be 100%,right? So she should be, like, pretty snarky, but, like, not mean 80% of the time, or whatever the business logic that you want,right? So these are evals, and this is the world of evals.

But here's what I would posit, which is it is the only way that we know what our software can do,right? And which is why I love the idea of product people looking at the evals,right? Looking at because they become the new specification for the product,right?

And so as we're watching, you know, if you're downstairs in the expo galley, you're seeing, like, new software. It's like, hey, bring the team in. And this sort of it a little bit reminds me of, like, the old, you know, for the, the old-timers here, like, behavior-driven development.

There was this period of time when it was like, oh, the business people are going to write the tests, and that will get converted to code, and then the code will run. And, like, the truth is, like, no one ever wanted to do that.

Like, no business per I don't even know who a business person is, but, like, they wanted we're going to do that. But I actually think this is different, and I think this is pretty a meaningful way to actually understand what the product can do and a little bit begin to specify what it can do.

Okay, so I'm vibe coding for a second, which we, which we all do. We all talk about. But I want to talk about vibe coding in a, in a way that's really constructive. And how do I sort of say this?

### Vibe Coding

**Ben Stein** [10:58]
It's very, very hard. I think I, I kind of was like, like, oh, you can't do it in Figma. You can't do it in a PRD. Like, what do I really mean? Well, it's very hard to, like, sit down in front of a blank piece of paper and write what the teammate, the agent experience should be.

It's just really hard. It's hard to, like, imagine it. And it's not until you feel it. I mean, so much of what we're doing in this, like, human-computer interface is visceral. It's feel. It is like, oh, well, like, did they ask too many questions?

Like, how many questions is too many? Oh, wouldn't it be great if they clarified exactly what you meant? Well, turns out that's really annoying. But when I wrote, like, the first spec, I'm like, then the teammate should ask a lot of clarifying questions.

And we gave it to users, and they were like, this sucks. And I was like, how would I have ever known that? And the answer is because it's so easy to prototype and vibe code something and get the feels.

And so this is the next thing that I'm, like, pretty excited about as a new product management tool. It is being able to feel and experience what it's like to interact with a computer, but

without just, like, writing it or hoping that you have a clickable prototype that will work. I will also mention that we have to be careful with vibe coding because I do not mean sit in the meeting and say to the engineering team, how come this is taking two weeks?

I finished the feature during the meeting. Like, that doesn't, that doesn't win you any points,right? So it is no, no, no. This is never going to production. But what this does is it gives you the feel, the, the experience,right?

And so this is, like, the only way I know to, like, actually test and feel it out. Do you, do you remember, like, the, the Claude certainty issue? Certainly, I mean, certainly. There was this peer,right, where every time you asked Claude something, he'd be like, certainly.

And, like, that probably, like, seemed really good when you were testing it for the very first time. And then, like, the fourth time when you're like, hey, can you do my taxes? Like, certainly. Can you write my, like, acceptance speech?

Certainly. Like, oh, this is actually really annoying. But you don't realize that until you experience it. So, like, that's why I like the vibe coding. Okay, so great. We did all this development. And then the question is, like, hey, we pushed to prod.

### Discover Emergence

**Ben Stein** [13:02]
Does it work? I'm like, I told you I don't know. So the question is, like, how do you test? How do you, like, know that it's going to do the things that you said it was going to do?

And I sort of alluded to this. I'll go through this quickly. It's just really discover, discover the functionality. And there's an old joke. I'll tell the joke. QA engineer walks into a bar, orders a beer, orders two beers, orders zero beers, orders negative one beers, orders a lizard, orders a beer with an emoji,right?

It's like, great. This bar is good to open. And the first customer walks in, asks where the bathroom is, and the bar blows up,right? Like, great, great old joke. It's kind of how I feel these days. Like, I just sit in.

I'm like, oh, you know it would be cool if they were to, like, start posting comments on LinkedIn. And what if they were, like, every time I added, like, a track to my Spotify account? They kind of, like, these just, like, crazy ideas.

But this is where, like, the emergent behavior comes from,right? And so it's this mindset of, like, let's just try. Let's just experiment. And it's, it's this, like, kind of growth mindset shift from, like, I'm going to write the features and the requirements to, no, we're going to figure it out.

This was a little bit unexpected for me. And this is

### Probabilistic Bugs

**Ben Stein** [14:17]
how do you sort of report to engineering and then have things fixed by engineering? And what counts as a bug in this world? And that is really, really strange. And I think as sort of, I don't know if it's, like, just a product role or maybe in a support role.

Like, how do you know what is appropriate to escalate, to put onto the backlog, to flag as a bug,right? It's like, I'll keep picking on, on Stacey. You know, she, she gives me a really hard time, so it's fine.

It's like, hey, she used too many emojis. Like, put it in, in, in linear. It's like, well, it's not really a bug. Like, show me in the spec where you told me not to use too many emojis,right? It's almost like, like, in our tickets, it's like, oh, you know, closed, done, closed, duplicate.

We need, like, closed. LLMs be, like, crazy, yo. Like, I don't know how to fix this, like, just because it's probabilistically generated. So how do we know if it'sright or wrong? How do you know if it's a feature, if it's a bug,right?

And I think there's this element of credibility that we need to build up. It's like, hey, we actually under we understand that for some use cases, like, 80% is good enough,right? This eval, we talk about evals, if it's passing 90% of the time, like, that's a go.

If it falls below 90%,right, that's red, and we're not going to ship it. So I'll actually come back to evals for a second. Because if the eval becomes the spec, and we can say, hey, we said at, you know, 100%, even though this is probability, you should never give a refund if a customer, like, can't prove that they bought the thing or whatever.

Like, it is. It's like, great. That is our metric. And we could say, yeah, this is a bug. But if it's just a, a feel, it becomes really difficult. Again, this was totally unexpected that, like, debugging and assigning bugs would become, like, controversial.

Okay, customers. So this part is

I found this really weird,right? So if I think about, like, not wearing my, like, founder hat, but wearing my, like, typical product manager hat,right? Like, I go into a customer meeting, usually with a salesperson. I'm like, I'm going to play a role,right?

### Customer Trust

**Ben Stein** [16:14]
And so what's the role? Well, I'm either going to play, like, visionary. I'm going to, like, hey, here's our vision for the product. Here's our roadmap for the future. Like, let me help you understand customer, like, how you're going to come along on this journey with us.

Or sometimes I'll play the role of honest broker,right? It's like, listen, sales has, like, given you a whole bunch of, like, they're just, like, selling you a bunch of vaporware. Let me tell you what's real. Let me tell you, like, exactly what you can expect.

And that's a role that you play,right? And I usually, like, preface this with, like, the sales team beforehand. It's like, yeah, I'm going to be the honest broker, and, like, we'll give the customer confidence. Today, I'm like, okay, I told you our vision for the future, our roadmap, and the customer's like, you're full of shit.

Like, none of this actually works. I'm like,right, I can't really paint the vision because no one actually believes it. It sounds like witchcraft. And then I'm like, oh, well, then I'll be the honest broker, and I'll tell you how things work.

But I just told you I have no idea how it works,right? So this became very strange because I can't play either of the roles that I'm supposed to be playing. The future sounds like witchcraft. The present is literally, I don't know.

So how do we do this?

I'll tell you how I've been doing it now. I don't know if this is, like, a 2025 answer or if this is, like, a durable answer. Like, if we believe that all of our products are for, like, for all time going to be probabilistic, then, like, we probably have to figure out how this world works.

What I've been doing now is really saying, look, we're inventing the future together,right? We're pulling the future forward. The reason you are talking to, like, a crazy startup like this, and you are thinking truly about, like, the future of how, you know, AI and agents are going to transform your business, is because you are a future thinker, and we are going to do it together.

And it's a little bit like, hey, let's compliment the customer. Let's, like, but it's not just, like, a false, you know, blowing smoke. It's like, no, truly, we need to figure this out together. And, you know, for 2025, I think that's actually the thing that is working the best, best for me.

### New Frontier

**Ben Stein** [17:59]
It's like, no, no, no. We have to do it together. And honestly, if you are expecting something different, like, it's not time. It's not time for you to, like, embrace this world because this is, this is the, the way this world is going to work.

And so I don't know. I'll conclude with, like, I have never had more fun building. I've never felt, like, both more inept and, like, more excited about what, what I'm doing, or just the experience of throwing something out in the world and then just, like, having my jaw drops.

Like, I can't believe this happened. And not only that, when we upgrade the models that are, like, underneath them, they just suddenly get smarter. And that's really weird too,right? It's like, all of a sudden, they, like, start checking their work.

They're like, oh, yeah, I just did a query to make sure that the row is properly inserted. And I was like, hmm, who told you to do that? I'm like, I don't know. It just seemed like a good idea.

I'm like, that is a good idea. I wish I thought of that. But anyway, but I think this is the new world that we're working in. The discipline, the product discipline, I think, is going to change for everyone, and it's going to change faster than we expect.

And we all need to, like, adapt to just, like, operating in a world and forget so much of what we used to know,right? A lot of the core, core ideas, listen to customers, solve real problems. Like, all of that obviously still applies.

But the tools, the techniques that we've, like, relied on forever, I think, are all getting upended. And so anyway, glad you're all at the AI Engineer Conference. It's awesome to have product people here working together because, you know, we all have to, you know, build awesome products together.

So thank you very much.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
