The Iceberg0:00
So, Echo AI. So Echo AI is bringing this fundamental new technology into the world of customer support and customer-facing teams. You've probably all seen this metaphor once before. You know, just the tip of the iceberg. There's so much that lies underneath.
And this is especially true with companies that are dealing with exceptionally high volumes of customer interactions. So whether that be customer support, or sales, or customer success teams, any kind of conversation that you're having with your customers is an opportunity to learn more about your business and what your customers need from it.
So, you know, at the surface, you probably get the signals around the things that are going wrong, or the things that are just routine to kind of handle. And most enterprises have a good idea of the approximate kind of set of categories that they generally have to deal with on a normal basis.
But there's just so much that lies underneath, and it's been virtually unlocked by just pure manpower that's required to look at these conversations. And there's this sort of cycle that happens with these types of companies. And, you know, everyone's kind of heard the startup mantra of just be obsessed about your customer,right?
Just do everything you can to listen to the customer, understand them to the best possible degree you're capable of at your scale. But as you grow and you get bigger and you get more customers, a lot of that becomes
this cycle where you just, like, you kind of lose touch with what they're saying, what they're telling you every day. And that only really comes with being successful,right? At, like, a, you know, a meaningful or, like, a manageable size of customers, you can typically use either, you know, staff or your sales team or your customer support team to sort of derive those insights and better, you know, navigate your company towards greater revenue growth.
But when you get too big to the point where that's, like, virtually impossible just due to the scale of conversations that you have to deal with, there becomes this issue where you're just, there's so much that is just happeningright underneath you without your knowing.
So the way we think about it, and just the first three columns here are effectively what every enterprise, every company at scale is trying to do, where they do manual reviews. So they do these, like, small sample sizes.
They try to collect, I don't know, 5% of the conversations. They run through an evaluation of it to check for, you know, maybe certain compliance things, like how well did the agent handle it, or perhaps are there any sort of common subjects or themes in these conversations.
At the end of the day, you're just sampling, and it leaves everyone unhappy at the end of the day. They understand that it's not a very kind of accurate process. So then you, like, involve engineers, and you start building these scripts.
You start doing retroactive analysis. You pull this data from all the different systems that your customers are interacting with you on, and you're, you know, doing these, like, kind of long-form analysis. And you're ultimately that can typically, like, derive insights that you know that you want to track in a forward-moving time.
So then you, like, move into this situation where you start building software to look for very specific things in every single conversation. And everything is just so retroactive. Like, everything is after the fact. It's already after fires had formed.
You have no sense of where the smoke is. And that's where generative AI, sorry, generative AI, can come in and really transform this process. Because generative AI unlocks this amazing capability of 100% coverage. So now, rather than doing that sampling or just, like, looking for the things that you know to look for, you can use generative AI to surface all the things that you didn't know to look for and look at everything all at once.
100% Coverage3:34
So here's a great example. This is just a small interaction between some random, you know, agent that's handling a chat from one of their customers. And in this one message that comes from the customer, you're able to find things like, why is this customer talking to me?
That'd be your intent. You know, what are the aspects of our business that are really kind of at the root of this? So, you know, routers and, you know, it was broken on deliveries. Maybe a supply chain issue.
You've got the basics, things like sentiment, understanding, like, not only the sentiment of your customers, but also the sentiment of your representatives, which is, you know, maybe more important. And then at the end of the day, with each of these messages, you can effectively extract so much more and so much depth that's ever been available to pull from just a single conversation.
And that's what our platform seeks to provide. Here's a great example from one of our customers. We don't have to read this, but wine enthusiasts, they shipped a new unit. It was a brand-new version of their they sell, like, really high-end wine refrigerators.
So their customers are spending a lot of money. It's a very special relationship that they have with their customers. They want those customers to buy more fridges. A lot of these customers are retail locations. So one of the insights we were able to surface for them was, you know, a defect in their manufacturing process.
Something that, you know, could have gone on for weeks and weeks and weeks and become a much bigger problem versus what our platform was able to surface in real time.
Echo AI Pipeline5:29
So, yeah, how does this all work? It all really starts with gathering all of those conversations. This is kind of like non-AI, boring stuff. We're, like, connecting to a bunch of different, like, contact systems, ticket systems, and so forth.
We pull all that in. We normalize it. We make it super clean, ready to go, and compressed to pass it into LLM prompts. And then we have dozens of these pipelines that are assessing these conversations in, you know, a variety of different ways, all of which are configurable by the user.
So the user, the customer can come in. They can tell us, like, what they care about, what they're looking for, and they'll actually work with us to write these prompts. And eventually, they write the prompts themselves and manage it over time.
And why this is, like, ultimately most important is when you are dealing at the accuracy sorry, when you're dealing at the, like, enterprise scale, they're ultimately most concerned around accuracy. There's a, I think, a huge hesitationright now in the market around accuracy.
Enterprise Accuracy6:09
I can deploy generative AI to try to understand these conversations, but do I really trust the insights? Is it going to be better than what my business analysts are doing or my CX leaders, my VP of sales? Does this system, does this technology, is it really capable of giving me insights that I ultimately trust?
So for us, it's important that we establish trust from the very beginning. So when we bring on a customer, within seven days, we try to introduce them to, like, okay, here are the insights. And then from there, we work towards a place where we're, you know, sort of a we like to say 95% accurate, but, you know, it's a lot of sampling, a lot of figuring out.
But virtually, we want to create that trust with the customer because that's ultimately what's going to get them to keep renewing and be a customer for a longer period of time. So Log10 plays a huge part in our ability to do this.
So not only have they created a huge amount of features and capabilities that allow for our engineering team to build faster and to understand, you know, the quality of the code that they're writing, but maybe more importantly, our solutions engineers, our implementation engineers who are actually working hand-in-hand with the customer who is telling us that, like, this isn't exactlyright.
And Log10 is kind of our go-to tool for managing that process. So.
Log10 Vision7:45
Thanks, Trey. So I'm really excited to share with you what we've built at Log10. So we basically an infra layer to improve LLM accuracy for your AI applications. We started with this vision of building self-improving systems. So having LLM applications that can improve prompts and models on themselves to ultimately drive accuracy improvements.
Obviously, we're not there as a field yet, but that's the vision we're driving towards. And I'm excited to share some of the work we've done along that path. So today, measuring and improving LLM accuracy is hard. You've probably had this experience if you've tried to deploy a prompt and kind of try to YOLO accuracy in prod.
And you've probably heard in the news of these instances where there was the Air Canada chatbot, which hallucinated a refund policy, and a judge in Canada forced them to honor the policy that was made up. So obviously, causing some financial problems as well as brand image problems.
There was the case of the Chevy Tahoe dealership chatbot, which was convinced to sell a truck for a dollar. We've run into similar issues where we were having some issues with a product, and a chatbot told us to play a game while we wait, kind of missing our emotional state in that moment.
And there's also issues with semantic search engines where, because they're doing this kind of RAG-based lookup, they sometimes just miss out on common sense, even though it may not be explicitly present in the source documents that they're pulling up.
And so people have tried to use human review as an alternative to as a way to sort of look at the output of the LLMs, and that's ultimately become the gold standard in terms of getting LLM apps into deployment, but it's time-consuming and expensive.
And as an alternative, we've tried to use AI-based review, so LLMs as a judge, for example. But obviously, this also has a lot of issues with accuracy. People have found that models tend to prefer their own output. They exhibit positional bias.
So just the ordering of if you present an option first, it might just prefer that over the second one, even if the second one's better. They have verbosity bias, a bias towards diversity of tokens and so forth. So kind of failing in these trivial ways.
Auto Feedback10:22
And so with Log10, we did a bunch of algorithms research work to try to address this issue and ask this question of, could we get the accuracy of human review with the speed and cost advantages of model-based review?
And that's exactly what we've solved for with our auto feedback system. So I think even in the previous talk, you saw a graph kind of like this where when you use LLMs as a judge, the predicted feedback can often be just set at one score regardless of what the ground truth is.
But with our auto feedback system, you get that much nicer, much better correlation between the predicted feedback and the actual feedback. And just to kind of motivate, what do you do once you have this measure of accuracy? So some of the downstream ways in which you can use the auto feedback is for things like monitoring.
So you get, like, an ongoing quality signal on how well your LLM application is doing. It can be used for triaging. So you make the most optimal use of the limited human resources you might have. And it can also be used for curating data sets, high-quality data sets for things like automated prompt improvement and fine-tuning, which can ultimately improve the accuracy of your LLM application.
Model Building11:45
So next, I'll say a little bit more about what the system is. So for some of the experiments we did as part of the research, we came up with these three different ways of building auto feedback models. So on the left, we have the ground truth data sets, which consist of the input and output, some kind of grading rubric, which you might give to a human for review, and their feedback.
And we had three variations where we could build these models with few-shot learning or with some kind of fine-tuning, and then finally creating these models with some bootstrap synthetic data and then fine-tuning the auto feedback model. And I won't have time to go into all the details.
We published some of this work in a blog post, which is available on our Substack, and there's a QR code there. But in summary, for a summary grading task where we used the TLDR data set, we were able to get a 45% improvement in evaluation accuracy by going from aggregate to annotator-specific models, going from GPT-3.5 to GPT-4 as the base model, going from few-shot learning to fine-tune models, and with the use of the bootstrap synthetic data.
Our approach was also very sample efficient. So by using the bootstrapping approach, we were able to achieve the accuracy of almost as if we had 1,000 ground truth labeled examples with just using 50 ground truth examples. So much faster to get started and not needing as much data to get to that high level of accuracy with the evaluation model.
We also extended this in a follow-up blog post to open-source models. So we were able to match the accuracy of GPT-4 and GPT-3.5 fine-tuned evaluation using Mistral 7B and LLaMA 70B chat. And that's in a follow-up blog post, which is accessible on this link.
Log10 Platform13:52
And so once we were able to show that we're able to get high confidence in the eval models that we were building in this way, we set it up for deployment within this auto feedback module on our platform.
And just kind of zooming out, we have, as part of the Log10 platform, fundamental LLM ops offering, which includes things like logging, debugging, and evaluation, as well as auto tuning features to do prompt optimizations and manage your fine-tunes.
And we have a seamless one-line integration which sits in between your LLM application and your LLM SDK. We have integrations into many of the common LLM SDKs, including OpenAI, Anthropic, Gemini, and a few of the open-source ones. And we integrate with frameworks as well.
So next, I'll hand it back to Trey for a demo.
Allrighty. Let me just make this slightly bigger. Well, okay. So here's Echo AI. So Echo AI is, like I said earlier, we connect into all of the different channels that your customer conversations are coming in. We transcribe them, we clean them, and then we allow you to basically, like, codify all the different things that you're looking to kind of analyze these conversations against, as well as offering a product that is purely generative and is kind of tasked with surfacing those insights in ways that you basically know the question, like, why are my customers canceling?
Live Demo14:53
And then we generate, like, you know, an ontology of different reasons for why that's happening. So let me give you an example of how we are making use of this feedback tool. So if I, like, jump into one of these this is a demo account.
So all these conversations are generated, so they're kind of silly. But nonetheless, you can see here we've got this example phone conversation that comes from a customer. Their TV was broken. One of the most, you know, simple, I think, on the surface features would be summarization of the transcript.
We actually rely on summarization for a variety of different downstream sort of analyses. So it's really important to us to kind of ensure high levels of accuracy. Not only that, this is also, you know, because of that kind of technical reason, we have, like, an immense number of prompts and throughput that has to get through LLMs.
So we do quite a bit of self-hosting and are constantly training new models to better handle different domains of our customer base. So here's an example. Like, you can see that this summary came in. It's pretty decent. All of these other sort of insights that you see here are all being generated by various different models and pipelines, all of which are being created from LLMs.
So each of these things are questions and specific sort of requirements that the customer is providing us. So it's really hard to stay on top of accuracy and quality as a result. Basically, every customer is different. So what we've done is we've leveraged Log10 not only for the ability to
very sorry, very quickly allow our engineers and solution engineers to go in and actually understand, like, what was the generated prompt for all of this. Pretty standard stuff, as many of you know. But what's maybe kind of more interesting is the ability to automatically generate feedback for these things.
So we've created a criteria that analyzes each of the summaries that we create against very specific user-defined, like, or defined by us criteria for how good of a quality the summarization was. So in this case, they actually look pretty good.
And there's one point deducted. You can actually read down here why it deducted that point. But let's say that, like, I want to provide kind of human override here. I could just come in here and change the point value and accept it.
And why that's useful for us is, like I said, it's critical that we continue a process that allows our solution engineers to this is an example, one against Mistral to kind of give us, like, really high fidelity human-provided feedback in a way that virtually is effortless for them to do so because what we're ultimately trying to do is collect as much of that as possible.
And Log10 has been, I think, a great tool in not only making that possible for the solution engineers and our engineers to do, but also automatically doing so behind the scenes. It's really changed our processes towards our fine-tuning data sets.
Thanks. Actually, I'll show one more thing to give an example of, like, how it's maybe more useful to an engineer. Here's another conversation where the summarization failed. We have just a reiteration of the instructions that looks like the system prompt.
No idea why. Let's we can actually kind of go in here and look at why. Okay. Well, there's a good reason why. And you can see that it's been graded accordingly, which is, you know, exactly what we would expect.
So we've been able to track hallucinations via this process. We've been able to see model drift in a meaningful way, in a data-driven way that we previously were unable to do. You know, a lot of it has been sort of requiring humans to sample these things and give feedback on them.
So this has been a huge tool in our, like, maintenance of achieving the utmost trust that we can kind of retain with our customers.
Thanks.
Results & SDK19:24
Great. So maybe just in the interest of time, I'll skip forward. So one of the big achievements we were able to get was using this auto feedback approach to get a 20 F1 point improvement in accuracy in one of the use cases with Echo AI.
And we published all of the details in a case study that's accessible there. And just in terms of getting started, obviously, we cannot share, you know, customer data here. So we created this new summarization app, which kind of shows an example of a summary grading application as well as
has a live version of the website that you could check out. If you want to take a picture of those QR codes, you can adapt this for your use cases, for your tasks, and try out auto feedback yourself.
We also have an SDK. So everything that was shown in the UI is available programmatically from Python and SDKs in other languages. The feedback type is pretty flexible. We have a bunch of recipes that you can get from our GitHub as well as notebooks to get started as well.
I'll skip over some of the mechanics, but it basically goes over how you create the task, create feedback, run the auto feedback locally on a simpler model, and then fetch feedback from a more complex model that runs on our cloud.
And maybe just in the interest of time, we'll I think, Trey, you already covered this, but yeah, invite you to use our platform. And for a limited time, you can use even the more advanced models on our platform for free.
Thanks a lot.
Thank you.





