AIAI EngineerJul 2, 2025· 12:50

The Build-Operate Divide: Bridging Product Vision and AI Operational Reality

Jeremy Silva (Freeplay) and Chris Hernandez (Chime) argue that the biggest challenge in generative AI isn't building prototypes but crossing the 'quality chasm' from v1 to reliable v2 through operational iteration. They explain how the lower barrier to entry and faster iteration speed in Gen AI accentuate the need for high-quality ops, where product quality becomes a direct function of how fast teams move through monitoring, experimentation, and evaluation loops. Chris emphasizes that human-in-the-loop isn't just a safeguard but a feedback engine, and that existing QA and CX teams in operations are already equipped to become 'model shapers'—labeling data, testing prompts, and defining what good looks like. Jeremy introduces the emerging role of the 'AI quality lead,' a systems thinker who can run experiments and evaluations without writing production code. They conclude that scaling Gen AI is an operational and people challenge, not just a technical one, and that embedding quality and human feedback early is the key to building faster and better.

Transcript

Intro0:00

Jeremy Silva0:16

Um, we titled this talk "The Build-Operate Divide," because we've observed this troubling trend where, like, good AI product concepts fail to reach their full potential due to operational challenges. So today, Chris and I are going to be talking about some of the learnings we've gleaned from, like, our experiences on the front lines operationalizing AI products at scale.

By the end of this talk, we hope you're going to leave with an understanding of how you can kind of bridge that gap between product concept, um, and operational reality by understanding how to deliver quality through evals, human review, and how you might build your teams around that.

So, a little bit about ourselves: my name is Jeremy. I lead product at a company called Freeplay, um, which exists to help solve a lot of these operational problems and, and help companies ship great AI products.

Chris Hernandez1:00

Hi everyone, my name is Chris Hernandez. I lead our Speech Analytics team at Chime. Uh, I've been in the industry for about 10 years, uh, from a CX perspective, and, uh, about 9 years on the ML space. So excited to be here.

Iteration Speed1:12

Jeremy Silva1:12

So both Chris and I kind of came from the traditional ML world into the Gen AI world, and I thought it was, like, helpful to maybe take a note and talk about, like, what we feel like has changed in this transition.

Um, so I think the biggest difference, in my mind, is this decreased barrier to entry,right? In the traditional ML world, like, you need tons of data to even get started, there are these long model training cycle times, and that barrier to entry just kind of, like, goes down significantly, um, in a Gen AI world,right?

Like, you can now use the intelligence of the base models to be able to, like, make use of and leverage, uh, smaller data assets within your organization. And then what comes along with that is this, like, increased iteration speed, which is, like, as the barrier to entry goes down, the speed at which you can iterate starts to go up.

And this increased iteration speed starts to accentuate the need for, like, a high-quality ops function. Um, so we want to take a look at why.

Working across, like, dozens of different enterprise teams at Freeplay, we've noticed this common trend emerge, which is companies will build an initial prototype, maybe they'll even ship a v1 of that thing into production, but they inevitably hit this sort of quality chasm as they try and go from a v1 to a v2 that really drives true value for the customer.

Quality Chasm2:10

Jeremy Silva2:28

And there's this reliability problem in there. And what we've seen is, like, the only way that you really cross that quality chasm, that re and get to a reliable v2, is through iteration. And so we talk about this iteration loop as you move from monitoring to experimentation to testing of valuation.

And I, I use that word broadly speaking: human review, auto evaluation, like, all these things all kind of come into play. Um, but what happens is your product quality becomes a direct function of your ability to move through this loop,right?

The faster you can iterate, the more times you can move through this loop, the better your product quality becomes. So ops kind of sits at the foundation of this, and especially as you start to scale,right? Like, your ability to move through these loops is, like, ends up being an ops function.

Human Loop3:19

Jeremy Silva3:19

And the, like, not-so-subtle irony here is that to deliver high-quality AI products, you actually need a ton of human elbow grease. And I'll pass it to Chris here to talk a little bit more about the importance of human experts in this process.

Chris Hernandez3:34

Cool. Awesome. Um, you know, I made the joke that one day I was, like, you know, the first time I saw LLMs and we started playing around with it, you know, I typed and I'm like, uh, I wonder if it actually knows, like, basic things like, uh, who invented Wi-Fi, and then it spit out Abraham Lincoln, and I was like, oh my goodness, we're, we're in some trouble.

But, uh, all jokes aside, and maybe not that extreme, it is a good reminder that LLMs and even the smartest ones can make mistakes, uh, and they often do it with a lot of confidence. Um, so, uh, that's, of course, what we call hallucination.

And, uh, in the Gen AI space, um, it, it at times, like, you got to make sure that, you know, we don't leave it unchecked, and that these hallucinations can actually be tricky and risky. Uh, they could also mislead customers, and they could misinform decisions.

And when you think of, like, companies or industries like healthcare, uh, just a small hallucination could have a lot of risk in it. Um, so that's where human-in-the-loop comes in. That's why it's important to talk about this today.

And human-in-the-loop kind of ensures while Gen AI does the heavy lifting, uh, humans are still there to steer the ship and, and drive, drive in theright direction. Um, so today we're going to talk a little bit more about, like, human-in-the-loop, why it's important.

I don't think this is, like, groundbreaking, like, everyone understands human-in-the-loop, it's been around forever, but just reiterating on the importance of it. So LLMsright now are just, you know, we know that they're good at, uh, generating content and they do it really, really quickly.

But without humans, they often fall short of real-world reliability, especially when nuance, empathy, and context is required. Without human-in-the-loop, you're also not scaling productivity, you're scaling risk. And so a single hallucination in isolation might not seem like a big idea or big deal, but at scale it could be dangerous.

Um, so, uh, the last point I want to make on that is also, each time a, a human flags or corrects an output, it's a signal that we need to be using to help retrain and reinforce the models that we see.

And that's, of course, how, uh, you know, models are evolving over time. So I want everyone to think of human-in-the-loop not as, like, a safeguard, uh, but as a feedback mechanism or feedback engine, if you will. So every model output that goes through a human review, that feedback then gets fed for refinement and, and of course, improving the model.

And over time, this loop brings AI closer to real human expectations and behaviors. But there is a challenge, and, and speaking with many different teams, the challenge is that we don't have enough people to do all these reviews.

Like, model-graded evals and evals in general are great, but you still need to have the human element in there. So most teams don't have enough people to review thousands of the different outputs that you have. And that makes it really, really hard to improve on the models and even measure, uh, the current state of the models that you have.

Ops Teams6:06

Chris Hernandez6:06

But there is good news. And so this might not seem like, uh, a clear path in some of the organizations that you all might be, be part of, um, but we already have teams that are kind of trained to do this type of work, uh, of human-in-the-loop.

And those teams are, like, your quality folks or CX folks within the operations side of the business. And so especially when you look at call centers, they're already experts in doing the jobs that you need them to do, which is evaluate interactions at scale, spotting, spotting edge cases, and defining what really what good looks like.

So in this age of A uh, Gen AI, the roles, from my perspective, it's not shrinking, it's continuing to evolve. Um, and they're not justright, they're not just measuring outputs anymore, uh, or outcomes. Uh, I think they're really shaping, like, the future and what's next.

So as Gen AI becomes, like, more embedded in the operations, quality is also evolving. It's no longer about auditing what's already happened, uh, but it's also about shaping what's going to happen next. So the shift change, uh, in the QA teams who are in this ops space, they're not just scorekeepers anymore.

They're becoming model shapers, prompt testers, and AI performance monitors. And if there's any group that's already built for this kind of work, as I mentioned before, you might want to look at your ops team and see if there's anyone in the QA contact centers that could also help expand your projects, uh, even further.

Um, they already know how to evaluate the nuanced conversations. They also know how to surface these edge cases that I mentioned before. So talk a little bit about, like, the evolution of just quality in general or CX. Um, they've been, for the most part, focusing on auditing, coaching, compliance.

Uh, but as automation continues to scale, these roles aren't disappearing, but they're actually transforming. When you look back, you know, 25 years ago, um,right, you know, the main roles were just QA professionals were listening to phone calls and evaluating interactions.

But we see that as, like, continu like, a continued evolution where automation is now coming, uh, into play, and these folks are, like, transforming their skill sets to help solve for larger problems.

The diagram that we have here kind of just shows, like, where the QA scope is today and how it's expanding to the Gen AI space. At the core, we have what QA has always done, uh, which is auditing interactions, capturing quality behaviors.

But now, as you can see, we scope the, the scope itself is expanding, and QA professionals are testing prompts now, and they're tagging outputs, and they're really helping shape the model, uh, behavior that we're expecting. Uh, and there is a beauty in the Gen AI space, uh, which is it opens the doors to non-technical folks.

Uh, traditionally, uh, only the ML teams, engineers are the ones that are heavily involved in the outputs and deciding what's going to, you know, uh, come from it. Uh, but now, um, uh, you don't need to know how to build, uh, the model pipeline to know what, you know, a good output looks like.

Uh, just like you don't need to know how to, um, make wine to be a good wine connoisseur. Uh, the, the outputs or the expertise still matters and just as valuable. Uh, and then the contact center CX teams are, again, already equipped with this.

And so, um, it might be something to, to look into and just see how you could expand that reach even further.

Quality Lead9:09

Jeremy Silva9:09

So one of the things Chris is talking about here is, like, something we have observed, like, working across a number of customers as well, which is this stor sort of emerging role of, like, the AI quality lead. And importantly, like, I've almost never seen it actually called this, but it seems to exist at companies who are having, like, a lot of success in the Gen AI space.

And I expect this role to actually become more formalized and gain more traction,right? Like, Chris and his team are an example of this role. Um, and importantly, like, people in this role can come from a variety of different backgrounds: product, ops, engineering.

But, like, the key attributes of what makes someone a good AI quality lead is someone who first and foremost has a deep, deep understanding of the customer need and the domain, and then importantly is a systems thinker and is able to, like, systematically think about how to diagnose and solve these quality problems.

What that looks like day to day is this person is doing a lot of these just kind of, like, this new skill set: labeling data, writing evaluation criteria, running experiments and tests, and things like this, and prompt engineering.

And you'll notice that, like, there is something notably missing from that, which is, like, these are often not the people who are writing production code. But I think what has changed so significantly is writing production code is not the only way now that you can contribute in, like, a really hands-on way.

All of these things, like, given theright tool set and theright structure of your team, you can contribute to, like, this iteration loop, the prompt engineering, the evaluation, all this kind of stuff, without necessarily being the one, like, writing the production code day to day.

Chris is talking about, like, how you do this at scale, but we've also seen a lot of success where companies will have one person or two people in this role, um, especially when you have a smaller footprint, and that goes a long way, you know?

So I think what Chris is painting is, like, in larger enterprises, as you start to scale this stuff up, you really need, like, a meaningful quality team. But you can do this by, like, kind of empowering a single or couple individuals in this role.

Chris Hernandez11:14

Cool. So I'm just going to leave you all with, uh, a few last points. You know, I think while human-in-the-loop is important and, like, we just spent the last few minutes talking about it, uh, if you don't have the resources, I think, you know, looking at high-risk, high-trust areas is abs like, an absolute must.

Takeaways11:14

Chris Hernandez11:28

And so insert human-in-the-loop at decision points and not just for show. The next piece is, like, continue to bring your con like, your ops teams and CX teams into the lifecycle early to help what you know, define what good looks like, uh, so they can help build out goal insets and, uh, tests against real-world, real-world edge cases.

Um, and then this is another point I wanted to make too is just, like, the, the fact of, like, launch is not the finish line. Uh, track performance, flag hallucination, measure impact, and iterate. Uh, time and time again, I see teams, like, celebrating, which is great to see celebration of, like, a product actually launching or a solution being actually launched.

Uh, but the important part there that it's not the end, it's the beginning where you need to set up your teams to make sure that the outputs are performing how you expect them to perform. Um, and last but not least, scale is not just about tech anymore.

I think it's about people. And so leveraging QA, ops, uh, and support and frontline teams, I think, as a strategic partner in the Gen AI space is going to be what makes teams successful. And there's just one key takeaway today.

It's that scaling Gen AI isn't just a technical challenge anymore. It's an operational, uh, reliability and responsibility. Um, when you embed quality and then human feedback into that loop, theright people, uh, into your Gen AI systems, you're not just building faster, but you're also building better.

Um, so anyways, thank you guys. Appreciate it.