Intro0:00
Thanks, Ron. Uh, allright, nice to— nice to meet everyone. Fundamentally, what I want to talk about today is really similar to the previous speaker, in that I just want to try and share some examples of customers who have achieved significant real ROI from building with LLMs and generative AI products, and try and tease out some of the lessons that are common across all of those people.
What are they doing that's the same? And then hopefully, once I've done the kind of basics, the fundamentals, if there's time I'll try and do some more kind of tactical tips and tricks, things that are maybe less obvious.
I'm going to try and run through a lot of stuff, and so if I run out of time, maybe that tactical stuff will end up in Q&A, but we'll— we'll see, we'll see how it goes. But maybe to start with just a little bit of background on who I am and what Humanloop is, like why have I— what have I done to earn theright to come here and talk to you about these tips and tricks and what does and doesn't work.
Real ROI1:04
So fundamentally, we were probably the first LLM ops platform. We've been doing this for a couple of years now, since, you know, even before ChatGPT, and we've helped hundreds of companies, both startups and larger enterprises, to try and get AI into production.
We've seen a lot of people succeed. We've also seen a lot of people fail. And so what I'm going to try and tease out is, like, what are the things that the companies that are succeeding doingright? At a very high level, I'll try and go into a lot of detail about evaluations specifically, and then at the end, kind of open up for Q&A and maybe chat about a little bit more of tactical stuff.
And I also have a kind of— the team's also deeply technical, so sort of research experience in the past as well before doing this more hands-on product work. And the other core message that I want people to take away from this is, I think over the last year, year and a half, there was a lot of experimentation, a lot of testing stuff, there was the initial hype wave about LLMs.
And the other message that I want people to take away is that we are now at the stage that people are actually generating real revenue and real cost savings from this. It's no longer a stage of, like, some promised land in the future where you'll eventually get there.
I can give significant examples across this talk, but, you know, here's one concrete one: Filevine's a customer of ours. They're in the legal space, so a regulated industry somewhere where it's very sensitive, you might have thought it would be harder to succeed with LLM products.
They've been able to launch six products in the last year, and they've roughly doubled their revenue. And for a very late-stage, fast-growing, you know, Series D, Series E startup, that's a substantial revenue uplift from these new products. And so we're well past the stage of kind of, will this deliver value?
I think we have the evidence to suggest that it's already there. Okay, so before I get into the details of, like, what are the fundamental lessons, I want us just to have some— to be on the same page about what is the thing that we're trying to optimize, what are the components of an LLM application, and how does this tend to fit in together in practice.
Core Components2:35
And I like to try and simplify everything. So across the whole talk, I'm going to be trying to take complicated things and just make them seem significantly simpler. And I think fundamentally, most LLM applications are composed of just four key components that get chained together in various different ways.
And I completely agree with what people were saying in the discussion before, that what you're trying to do is very quickly put a pipeline together and then optimize each of these components towards making something that's sufficiently robust. And, you know, there's a lot of frameworks out there that would suggest that this is very complicated, but fundamentally, like, each block is actually very simple.
You have some base model, maybe it's the large model provider, maybe it's something small and fine-tuned. There's a prompt template, just a natural language instruction to the model, some select data selection strategy: am I using RAG, am I populating this from an API, and then maybe you also augment this with function calling.
And you chain these things together. But there really isn't much more to it than that. What makes it hard is not the complexity of the applications, it's like, how do I make each of these components actually good? And that's where most of the work lies.
One concrete example of, like, this framework in action, just so we have, like, one, you know, real application to think about, I think GitHub Copilot was the first really successful LLM app to drive real revenue in production. And same structure,right?
There's a base model, it's been fine-tuned in this case because they care about latency. They have a data selection strategy. So what they're trying to do is suggest code for you, so they're looking at the previous code just behind your cursor, the last 10 or so files that you touched, and they're just grabbing the most similar code from that and populating it into the context.
And they're very, very rigorous about evaluation, which we'll talk about in a moment, but fundamentally same structure: base model, prompt template, some data selection strategy, chained together. And there was also a comment in the last section about, like, systems and chains becoming more complicated over time.
I actually think we're going to see the opposite trend as the models get better. Like, a lot of the chaining and complexity that's being addedright now is a workaround around the fact that the models aren't that good at tool selection or aren't that good.
So I actually think that, uh, keep things simple, don't overcomplicate it, and just make these individual components good. Okay, that's all background. So, like, what are the fundamentals that I think you need to getright before we talk about the more tactical tricks and things?
People First4:58
And Colonel Boyd of Oodaloo fame used to famously go around the Pentagon shouting at people, people, ideas, machines, in that order. I think roughly the same thing applies if you're trying to build an LLM application. Fundamentally, you want theright people.
What's the skill set you need? What's theright mix of that? I think the next thing you need on the ideas front is, like, starting from clear evaluation criteria and thinking upfront about what feedback you're going to capture in your application.
Like, how are you going to measure whether the thing is actually working? And then finally, like, then you can think about the tooling and the infrastructure you need to make thatright. And the teams that have succeeded, I think, do all of these things in a particular way.
So I'm going to go through each one, and then I'll try and give concrete examples from either some of our customers or just people that I've spoken to and learned from in the space. Okay, so team composition. There's really two takeaways that I would be pushing quite hard here.
And the first is that you probably need less machine learning expertise than you think. So on the teams that have succeeded, they tend to be staffed more by generalist full-stack product engineers, you know, maybe the term AI engineer that this conference is about is starting to drift in that direction, and less by people who are fundamentally focused on model training.
So the people to kind of theright of the API line. They care about products, they do know about prompting, they know about the models, but they're not fundamentally machine learning people. And the second big takeaway that I think is the most fundamental one and the most underappreciated is how important domain experts are in getting to success here.
And I see someone waving me there. I totally agree. And domain experts—
It's good to be first.
It— yeah. I'm thinking about the numbers of people. So, like, you've got high volume the engineers, probably the most important being these domain experts. And the reason they're really important is, I think, traditionally in software, the role of the product manager or domain expert was they produce the spec, and then, you know, they figure out what's needed, and someone else goes and implements it.
And what LLMs have made possible is a much more direct contribution of those domain experts into the building of the application. They can be helping you create prompts, they can be helping define evaluations, providing feedback. You need to make sure that however you set this process up, and we'll come to tooling at the end, that those people can still be central.
And then finally, I do think you want some machine learning expertise. So it's possible to go too far the other way. And there are fundamental concepts like how do I build a representative test set and how do I think about evaluation that you want someone on the team to know about and be teaching everybody else, but they don't need to be doing hardcore machine learning model training.
So you don't need PhDs and kind of people who have a lot of experience training stuff, you just need people with good data science background and knowledge. Okay, so that's team composition. A few examples. So I think my favorite example on this one is Duolingo, because at Duolingo, so one of our customers, the linguists, do a ton of the prompt engineering.
In fact, one of the PMs told me, I don't know if this is still true, because about six months ago, that they don't let the engineers edit prompts. That actually the linguists do all the prompt engineering and then there's aright, there's sort of a one-way direction of travel from there into production code, because they're fundamentally the ones who know what good looks like, how to change it, how to look at the outputs and understand it.
At Filevine, there's, you know, there's another example I mentioned earlier, and Ironclad's a good one too. You have a lot of legal expertise being directly involved in the process. In Ironclad's case, they actually don't use legal experts through prompting, but in Filevine's case they do.
So they actually have legal professionals and people with legal expertise prompting the models and actually producing what is effectively production code, but it just happens to be in natural language. And the reason I put Fathom on here as an example is, so Fathom is a meeting note summarizer, smaller company, but I think it's a really good mental model for why domain expertise is so important.
So they're doing summarization, and you can think to yourself, like, what makes a good summary? They're summarizing meeting transcripts, and it's so context-dependent. There's no, like, answer to the correct summary for a meeting. It's like, who is it for, in what context?
And there's one product manager at Fathom who's done the majority of the prompting for their different meeting summaries. So if you're a salesperson, you get a different summary. If you're a product manager doing a one-on-one, you get a different summary.
But an engineer, how could you rely on them to have that domain knowledge? It wouldn't make sense. And so, you know, you really do want someone like the product manager to be deeply involved. So point number one, team composition.
Center domain experts. You don't need as much ML expertise as you probably think. The teams that we've seen succeed the most tend to have a balance of, like, lots of generalist engineers, lots of subject matter experts, a little bit of machine learning.
Okay, the next point is that you need to make evaluation, sort of baseline evaluation, the core to what you're doing. We spoke about, you know, there's being that simple block that you're optimizing over, decisions over model, over data selection strategy, prompt templates, and tools, but there's a combinatorially large number of decisions there very quickly.
Eval Strategy9:20
And if you don't have a good evaluation strategy in place, then it's really difficult to make those choices. A lot of teams spin their wheels making changes, eyeballing things, thinking they're improving them, or they just don't trust it enough to put it in production, especially if it's something that's reasonably high stakes.
So I think you have to start with evaluation. And I also think defining the evaluation is, in some sense, defining the spec. Like, you're saying what good looks like and what you ultimately care about. So how do the best companies do this?
The companies that I've seen that succeed really well have evaluation at every stage of development in different forms. So during prototyping, you know, you're just trying to validate a new idea, it's highly iterative, you're experimenting, like, is something even possible?
And here you're trying to evolve the evaluation criteria alongside your, the development of the application itself. So people will often put out a shitty prototype very quickly, internally, maybe even something that doesn't have the full UI wired up, and they're just trying to get a sense of, like, what does good look like?
And usually from that comes some evaluation criteria. And so there's this kind of back-and-forth evolution of, like, what should I be evaluating? And they tend to then distill those down into evaluations that will be used more rigorously as they get towards production.
And then once you're in production, then obviously you need to be able to monitor things, how is stuff behaving in the wild, and also to drill down and understand, like, if something goes wrong, why did it go wrong, and be able to fix that.
User Feedback10:51
And then finally, one concern that comes up a lot from people is, I'm going to go in and change a prompt because I noticed a problem, but how do I know that I'm not causing regressions elsewhere? Or a new model's come out and I want to shift to it, but I don't know whether I'm going to, like, introduce accidental mistakes.
If you've built evaluation well from the start, then a lot of these problems solve themselves. And so that's why I think it's really critical to think about evaluation at the beginning. There's various reasons why it's hard. I'm not going to have time to go into it in detail, and I think a lot of this now has become kind of consensus knowledge.
So I'm going to skip past this one, but ask me questions at the end if we care about it. But I would say that ultimately the ground truth answer to evaluation is, like, your users know whatright is, especially on the more subjective things, if you're doing summarization or question answering or whatever it might be.
So end user feedback's really priceless. I give the example of GitHub Copilot here. They use quite a complicated end user feedback mechanism to measure how good things are. So they're looking at both, like, was a suggestion accepted, but also did the code that they suggested stay in your codebase, and how much of it, at various different intervals.
So they have a really rich signal from their end users about whether or not it's working. But it's hard to get,right? So we do see lots of apps building this in. ChatGPT has it. We've seen it in others, thumbs up, thumbs down, copy-paste, regenerate,right, all of these different signals of end user feedback.
Really priceless, like, really important to try and build into your application. And the teams that succeed well think about this at the design stage. Like, how am I going to build these implicit signals of feedback into the application?
It tends to be lower volume, though, than you would like, and you can't get it during development, so it's not a panacea. I would say that we tend to see four different types of feedback that get collected from in-applications.
So one is actions, like, what did the user do when they received a generation? Issues is, like, someone actually flagging a specific issue, direct votes, and then corrections. If you're generating a summary, writing an email, doing things like that, it's actually very rich data to log the corrections or any edits that your users make.
It can be very helpful to improving things down the line. Okay, so that's, like, end user feedback, but you don't have it during development. So the other thing that we see teams doing a lot is trying to build a scorecard of different types of evaluators.
And the difference between the teams that are doing well here versus the ones that do less well is the extent to which they break down the subjective criteria that they're measuring into small individual components that can be independently tested.
Scoring Methods13:21
So, you know, we see teams using LLM as judge, and that can go really badly or it can go quite well. And the difference is sort of not expecting too much from the models. If you ask the model, is this a good piece of writing?
That's a very ambiguous subjective evaluation. You're going to get very noisy data. And if you ask the model if it prefers one of a few different options, there's lots of sort of biases that come into the ordering that you show things that you need to be aware of.
But you can break things down into much more specific questions. Is the tone of voice in this passage appropriate for a child, you know, if I'm doing a school-level project? Or is this piece of text, does it contain these five points that I always need to have in my structure?
Right, those kinds of questions LLM judge works well as. And then you always have your traditional codebase metrics, precision, recall, latency, that you would always have. We've not been able to see examples where people can get fully away from human evaluation.
Almost all of the best teams still have some amount of manual annotation that they augment with more scalable methods. And then you're optimizing on this predator frontier,right? So it's never the case that, like, one system is, like, just better across the board at all of these things.
It's usually a trade-off, which is why you want to have a scoreboard of different metrics that you can look at and then say, okay, this one's more expensive, but it's a significant lift in, you know, helpfulness or whatever it is that I most care about.
Like, am I happy with that trade-off? Which is a little bit different from traditional machine learning,right, where we would, like, try and have a single number that we're optimizing. Because here we're not caring about, like, how good is the model, we're caring how good is the product experience for the end users.
And that's much more multifaceted and has more trade-offs.
And, you know, some concrete examples of this, we did the GitHub one already. I think, like, Hex is a great example of this. I was speaking to Brian Bischoff, their head of AI, a couple of weeks ago, and he was talking about how they break down each of their evaluation criteria into small pieces that are essentially binary, that they can score independently of each other, and then take those in aggregate together to try and get an overall view.
And he was the one who said, kind of, you're seeking the single god metric, you're probably taking the wrong path. And Vanta's a really interesting example where they still rely, like, reasonably on a mixture of automated evaluation, but plenty of human feedback as well, because it's so high stakes and they're in a regulated place.
And so they need to be really confident of that, of those end results. I'm going to keep running because I'm very conscious of time, but people just chat out what you want questions about at the back and I can dig into things.
So the last point I want to talk about is, like, okay, if you've got the peopleright and you've got the ideasright in terms of building your evaluation criteria and getting the spec correctly, you've designed things that you will be able to capture end user feedback, you have a test set, a suite of tests that you can use for regression testing, then, like, how should you think about what tooling you either want to build or buy or kind of use for this process?
Tooling16:18
And I think there's three things that we've seen be really important. The first one is designing whatever system you're building to optimize for team collaboration. So there's, you know, you've got prompts which are natural language artifacts, they act like code.
If you store them in your codebase and just treat them as normal code, you alienate those domain experts who you want to be deeply involved in the process. And so try and design things in such a way that domain experts can be involved both in prompt engineering and critically in evaluation.
They may not know enough about how test sets work and metrics to drive the process themselves, but they're ultimately the ones who know what good looks like. The second thing is make sure that you're able to include evaluation at every stage of the process.
Soright from the beginning, during prototyping, you want lightweight evaluations, you want to be able to do evaluation for monitoring, and you also want it for regression testing. And then the last one is that I think you want really comprehensive logging.
Like, ideally you just want to be capturing inputs and outputs at every stage, and you want to be able to replay these things, and also to be able to take data points from your logs and put them into test sets of edge cases or things that you want to make sure that you succeed on in the future.
And those are, like, three fundamental, like, bits of tooling that almost everyone we've worked with has either bought or built themselves. And obviously, like, I'm biased because we're building tooling of this kind, but I'll try and give some examples of companies that have also built stuff themselves, and some of this is open source, so you can go and look at it.
So one very concrete example here is Rivet, which is an open source library that was built by Ironclad. And their CTO said to me that, like, they almost gave up on agents before they had this tooling. So they started to build agents, they added a whole bunch of function calls, it worked well with one, it worked well with two, and then once they added their third and fourth things, the whole system started sort of failing, and they were almost ready to give up on it.
And one of his engineers had gone and built this logging and re-running infrastructure, kind of as a weekend project secretly, to try and make it run. And it was only after they had that ability to debug these traces that they realized that actually they were able to get to performance that now is in production.
And I think for their biggest customers, something like 50% of their contracts are being auto-negotiated. But that wasn't possible without the tooling. Linus gave a talk yesterday about how Notion does this, and he was speaking in a lot of detail about the logging that they have, and in particular this ability to go and find any AI kind of run from production and re-run it and make changes to it.
And fundamentally, you know, that's the system that we've been trying to build at Humanloop as well, which is how do you take each of these components to your system, the prompts, the tool definitions, your evaluators, and the data sets, and then iterate on each of them with feedback very quickly whilst having everything logged.
And, you know, I mentioned Filevine at the beginning, having been able to sort of roughly double their ROI as, like, one really concrete example. Like, they're one of the people who've done this with us, and for them at least we've become their system of record for all of their prompts in production, and also the place where their domain experts, who are in this case legal professionals, are working with data scientists and PMs.
So obviously we're not the only ones out there doing this, but I think it shows a really concrete example of how this can drive actual either large cost savings or real revenue. It's not just hypothetical anymore. Okay, I'm going to end there.
Conclusion19:17
And then open up for questions. And if you scan this QR code, I think you can get the PowerPoint presentation and a bunch of other goodies as well.





