Intro0:00
Well, thank you all for being here. It is a small but excited group of people around evals, and you know, evals may not be the most sexy talk at this conference, but it might be one of the most important.
And Nischal and I will spend time today speaking through why,right now, we think it's a seminal moment for evals and the import, sort of, of the changes that we're seeing in the market as we move into these agentic frameworks,right?
So as things are more automated, as things happen in a more agentic way, we need to make sure that we understand what's happening with evals. We're all here, obviously, because we're very long AI, we care a lot about AI, but we have to really understand what's going on with it.
And really, if you think about the crux of it, can we evaluate the performance of our systems in relationship to our goals for those systems? And that's what we'll be walking through today. And so, about us, my name is Sheila, and I'm joined by Nischal.
It's great. We're an investor-investor pair,right, a VC and portfolio company, and I think there should be more talks like this because some of the most fabulous portfolio companies are not great at giving shout-outs to themselves. So I get to give the shout-out to Klarity and to Nischal.
He's the co-founder and CTO of Klarity. He's sitting up here, but he'll be speaking very soon as well. And Klarity just announced their massive $70M Series B! financing on Monday. So we're going to hear about that journey. It hasn't been a short journey, so I think that can be quite inspirational for a lot of you as you think about what you're building and the size and the scope of what you're building.
And so that company builds what we call exponential organizations. Well, what does that mean,right? Software transformed how we all work,right? This internal systems is always synced, always on nature of software systems for internal work changed the nature of the efficiency and effectiveness of buildings that of businesses that we build.
Klarity is doing that for your external world,right? Most of your relationships with your customers, your partners, are dealt with through documents. Those documents are one-off negotiated, they're one-off pieces of paper,right? Klarity is automating all of that and then allowing you to build exponential organizations through that type of real-time relationship with those documents.
Cloud Era2:35
So you'll hear more about that and how Klarity has implemented a number of eval systems later in this presentation. Myself, I founded a venture firm called Tola Capital well over a decade ago now, and the reason I founded the firm was I was working at Microsoft.
I was running the database and developer platforms businesses at Microsoft. I co-led the company's enterprise strategy and was fighting the good fight against Team Windows to launch Azure. So being Team Cloud at a company that was based off of Windows was a really difficult thing, but it was fabulous as well, and obviously Azure has gone on to do pretty okay for itself.
So the genesis of the firm Tola Capital really was, hey, how do we think about this next generation of applications that would be cloud-based? And now we're even more excited by the opportunity to bring this next generation of AI-enabled applications to the fore.
So I'm going to walk down memory lane for one quick moment and say, you know, what we saw in the advent of the cloud was very clear. It would favor scale. The CapEx requirements, the physicality of building out those data centers was so expensive that you had to have a search business or an office business or a retail business to go fund that development, and then we would build on top of that,right?
The evaluation of those platforms was more straightforward. Am I offering you speeds and feeds? Am I offering you performance? And then, of course, you're layering on what functionality at what price I'm offering, but the evaluation of the physicality of that world was more straightforward.
Now, as we enter this AI world, we're saying, wow, there's, you know, some of the similar characteristics of large CapEx and large mega-cap participation,right, where we have obviously these systems running on top and models running on top of clouds and being trained by the clouds and that compute cost, that inference, the ever-so-difficult-to-get-your-hands-on chips, the talent of all of you in the room,right, the AI engineers that bring this to be.
AI Eval Challenges4:11
But in addition to that, you have the proliferation of open source and open source models and just the ability of those models to do an incredible job at delivering incredibly complicated scenarios that are, you know, really, really, really strong contenders.
And so you have a proliferation of players, a proliferation of models. You have a deep academic heritage of those open source and AI development, and you have great mega-cap partnerships for those models as well. So where are we going,right?
Where AI is just, it's so important to think about AI as more than just another tool,right? It is a reflection of us. It is a reflection of our understanding of the world, our intentions, our preferences, and at the end of the day, our society.
And we'll talk a little bit about what that means individually and collectively, especially in this world where agents will represent more of us as individuals. So let's talk about some of the shifts here. You know, we started with AI eating everything on the internet.
I like to call this all the garbage and all the gold on the internet was consumed by AI. Then we moved into, you know, doing that gave us emergent behaviors. That's, of course, a great fancy way of saying we're not exactly sure how it knows what it knows, but we know it knows it,right?
This era then, then we have trained data sets, curated data sets, but we're moving into the era of self-taught, self-learning, and self-sufficient models. And so if you pause and think about that for a second, before we get into a truly self-sufficient era, we really need to fix evals,right?
Because it kind of gets to be too late in that era. And so that transition to full automation is happening faster and more aggressively than any of us thought. So what does this mean,right? Narrow AI evaluation was, I'm a hammer, you're a nail, I'm hitting you, am I doing thatright?
You're an image, am I classifying that image? And I could tell whether I was performing on that task in theright manner in a pretty straightforward way. Then you get into broad AI,right? And this radar chart speaks a little bit to how this evaluation becomes more multifaceted.
What's my capability and intelligence? Am I serving a domain? Do I understand that domain? What are my values,right? What are the, and whose values do I care about? Mine, yours, the users, societies, today, tomorrow? All of these questions come together.
Safety, okay, do no harm. What does that look like? That could be different for people. And of course, the context and end-user awareness, which is often not discussed in an eval world,right? Your end-user is who you're delivering and developing these solutions for, but they're often an afterthought in terms of evals and evaluation.
And then we have to do that across all of the different modalities of image and text and video and kind of all of these together is creating a much larger problem around evals. So today, the tools are still simplistic.
Eval Pitfalls7:42
We're going to dive into kind of each of these areas of simplisticness and talk about some of the innovation happening to deliver this forward. And then what Nischal will do is show us how Klarity has dealt with each of these issues related to evals as well.
So benchmark hacking, I like to call this, you know, what am I solving for? Solve for X is how a lot of these benchmarks and leaderboards work today from an eval perspective. And it's interesting because scoring high on the benchmarks is pretty easy to do if you know what we're solving for.
If you understand X, you can do everything to solve for X, and you can look as intelligent as you want at solving for X, but the real reason is you may not understand anything about what's happening. You may be just solving for X.
And that's a pretty scary reality on some of these things. And so we say, wow, these ALMs passed an AP exam on a particular subject. Does that mean it was trained well on the questions that have heretofore come on that AP exam, or does that mean it actually understands the subject?
It's a really, really interesting question, but what we're seeing on a lot of the AI leaderboards is the best solvers for X are at the top of those leaderboards, and that's a problem. So one area where we're seeing research happen, this is a Microsoft research paper around dynamic benchmarks.
Rather than saying solve for X, can you identify this image, they're creating dynamic data sets. Basically with synthetic data where you can say, hey, we're moving objects around. These are not published. They are not public,right? So you don't know what the answer is coming into it.
But then I can test you on spatial reasoning, visual prompting, object recognition. The images are changing, so the models have no ability to memorize those benchmarks. This is an area, dynamic benchmarks in general, where I think we'll see a lot more work.
Benchmarks versus real-world scenarios is element two on this. You know, it's interesting. There's a lot of model creators that claim that their models perform very well and are very generalizable, and then you actually go and ask them a specific set of questions.
Even if I've done super well on MMLU and these things, and you say, okay, I'm going to go test this logic and test this reasoning, and the answers are simply wrong, and they're wrong much and most of the time.
And this great example from Finance Bench, which was Petronas AI's work and Stanford's work around saying, hey, you know, these basic financial questions were not answered when you actually benchmarked it on those real-world scenarios. And the user is not at the center of our evaluation universe,right?
We have to put the user at the center. We have to revamp the UX and feedback systems in order to really understand and capture more user needs. Evaluation of UX is super difficult,right? And Nischal will talk a lot more about this in the case of Klarity, but that's why we're all actually here, to deliver that end-user value.
And so if we're not doing that evaluation, what are we doing? Black box models,right? So this is a problem that's sort of hiding in plain sight, obviously. Evaluation is a proxy for a task. Evaluation is seeking truth. It is not truth.
And so how do we really understand what we're doing in a world where we can't mathematically represent what neural networks have learned and how they have learned that? And so the area of AI interpretability is not new, but it's super important as we think about the rise of these complex AI systems.
And so researchers are trying to open up these black box models, show us the how and the steps in this. And the transformer-based LLMs, we're really tracing information flowing through the network. And so this is a good example of, you know, sort of asking questions and seeing, hey, how am I answering it?
How are these pieces coming together? We'll see a lot more of this interpretability work in the reinforcement learning work that's happening with the current GPT models. It's super, super, super important that we get thisright now. Now, values,right? When we create AI models, we do instill our own values in them, whether we want to or not,right?
And these evaluations have to understand the technical capabilities, but also the underlying values that we're putting into the models. And there's a lot of questions on this, and I have way more questions than answers, as I think we all do, but it's, you know, are we creating values that benefit humanity?
As we go into this agentic world and you have your own model,right, does that just reinforce you? Is that a good thing,right? So the model of Sheila is going to believe more of what Sheila believes and get deeper and deeper into that Sheila-ness.
Is that a good thing,right? We've seen the echo chamber of ourselves in news. We've seen the polarization that this has caused in our society. We should be asking questions about where we're going to get to as we enter this agentic world.
And this is kind of a, you know, it's a question,right? This is a simple two-by-two to ask the question. Do we want to appease or challenge users? Do we want you to choose whether you are appeased or challenged?
Do we want to ground AI in present or future societal values, aspirational values versus present values? Do we want to encourage users to select their own values? Should model makers be responsible for this? Lots of questions, less answers.
Now we're going to turn it into a real-world example, leveraging Klarity to talk about how the company is replicating human cognition. Nischal?
Klarity Unveiled13:31
Thanks, Sheila. Thank you for having me here. My name's Nischal. I'm co-founder and CTO of Klarity. As Sheila mentioned, what we do is we automate back-office workflows, things that traditionally required large teams of offshore humans and throughout human history have been impossible to automate because they're cognitive and they're non-repetitive.
You can't just write down a simple set of steps and automate these workflows, and that's what we've been working on for the last eight-odd years. Predominantly, these are document-oriented workflows. So it's some kind of PDF that a human being is having to read as part of a company's back office.
One example of this is revenue recognition, typically part of an accounting team, matching invoices to purchase orders, so actually aligning two documents and matching them to each other, processing tax withholdings that often come in many different languages. But you probably get the sense.
You get PDFs that are completely unstructured. Somebody has to go through them because it's a tightly regulated process, part of a finance and accounting team. And today, this is done completely manually. There's actually a really great keynote yesterday by a gentleman who pointed out that document processing tasks are generally super tough for LLMs for a variety of reasons.
If the page is, like, rotated or the scan quality is bad, if you have, like, graphs or images, if you have tables inside of these, it is very, very tough to get this to work. And it's not the kind of thing you can just give to ChatGPT and it happens out of the box.
This is basically exactly what we do. And we spent, you know, a long, long time and tens of millions of dollars building a stack that's able to deal with these kinds of documents.
These are some of our customers, mostly B2B SaaS companies, mostly in a five-mile radius around where we areright now. And we predominantly serve their finance and accounting teams, although we're expanding quite a bit beyond that. Our journey has been a little bit unorthodox.
It's kind of like that, you've probably seen the startup curve of, like, the trough of disillusionment and you get the Tech Crunch article in the beginning. Something like that. We founded the company in 2016. We pivoted four times between 2016 and 2020.
And when we were kind of at our wits' end, about to give up, we did one final pivot and focused on finance and accounting teams and found pretty strong product-market fit there. In the last couple of years, we've completely re-platformed around generative AI, and that's really been a shot in the arm to the company and have gone on to raise 90 million plus off of that.
But before generative AI, we were in traditional ML for more than six years. We like to say that we started an AI company five years too late. And so a lot of the concepts around evaluations were, you know, they seemed pretty natural to us, and we didn't really see why generative AI had to be any different.
In supervised learning, which is a majority of what we did pre-GenAI, you have your train split, your test split, the metrics are pretty well-defined, F1, ROC curves. There are these, like, shared benchmark tasks, like sequence labeling and squad, and these were all very well understood.
So it wasn't immediately apparent to us why generative AI has to change any of these things. And in the two years since then, we've, as I mentioned, re-platformed to generative AI, and that's gotten us quite a bit more scale.
We process more than half a million documents for customers. We have more than 15 unique LLM use cases running in production today, and more than 10 LLMs under the hood. Oftentimes, a single use case has multiple LLMs working together.
GenAI Quirks16:54
But all this kind of begs the question, why does eval for generative AI have to be any different than traditional ML? And this is kind of the journey that we've been on in the last couple of years. The first, and a number of speakers have spoken about this, so I'm not going to go into too much detail, is non-deterministic performance.
Very challenging for us because it's a very cognitively demanding task. So if you upload the same PDF multiple times, very likely you'll get completely different responses. New user experiences, I think fundamentally, when you move from, like, discriminative models, classification, regression, random forests, to generative models, the types of experiences you can provide your users expand quite a bit.
And these new experiences are just much harder to evaluate. So a couple of examples from what we do. We have this tool called the Architect, where users will record their business workflow, like just them doing their job, upload it to Klarity, and then we'll create a business requirements document out of that.
It's typically like a 10-page Word document. It has a flowchart, images, very comprehensive, like what a McKinsey would build for you. It's not at all intuitive how to eval something like that. Another example, part of our product is you can do natural language analytics, so you don't need, like, a BI specialist.
You can just ask it questions like, hey, how's my contract population evolved over time? Give it to me in a stacked bar chart. It'll do that for you. Cool feature, but, like, how do you eval something like this?
And then a more traditional document extraction task. You are trying to find certain parts of a document that are consequential in some way. They could be tabular, they could be legalese buried inside of a document, and you need to know what accuracy you're doing that with.
So this is why new experiences, while very rewarding to our customers, have been very challenging to us from an eval perspective. The second is the rate of feature development. So in our previous deep neural net world, DNN world, a feature took like five to six months to build.
We literally had teams that would annotate data, dedicated teams to annotate data, GPUs, most of which are gathering dust now to train. And so end to end, it was like six months. So if it took like two to three weeks to build evals, thoughtful evals, that was completely acceptable.
It was a pretty small fraction of the feature development time. Now what we're seeing is we can get features out of the door in days. Within 12 hours of the ChatGPT launch, we launched Document Chat as a feature.
And so in that world, it's unacceptable that it takes a week, two weeks to build evals. It now becomes the bottleneck to feature development. And the third is benchmarks diverging from performance. Sheila talked about this a little bit.
What we've seen at Klarity is even slight differences in MMLU actually make a very big difference to us because we're kind of at the frontiers of human cognition, doing something that's very challenging for most human beings. We've seen many cases where, not going to name names, a model is supposed to be better in terms of MMLU, and then it's totally not on our internal benchmarks.
There's a variety of factors that go into it, but it's made testing very chaotic and challenging for us. And so the question then is, what can be done? And I'm pretty sure most application developers are running into these problems or something similar.
I'm not going to say that we have the silver bullet, and I would say we are nascent in our eval journey. But here's a couple of things that have worked for us. So the first is, like, really give yourself the gift of imperfection.
Practical Solutions19:48
Don't put the threshold of high-quality evals at the beginning of the feature development lifecycle. Good evals are not trivially cheap to build today. They probably are not going to be for the foreseeable future. And so we really try to front-load user testing.
What we found with these generative AI features is we have to think of each feature almost as its own product-market fit because we're delivering experiences that people have never had before, chatting with a document, natural language analytics, watching videos automatically.
And so there's a lot of user experience risk actually baked into each of these features. Compared to traditional machine learning, like a recommendation engine, where you have fairly high conviction that the form factor is correct. So what we try to do from a development perspective is front-load the UX risk, back-load building out evals.
Of course, I'm not advocating that you go into production and scale without building evals, but a lot of features die at the UX stage itself. Let them die before you build out evals. But once you decide to go into production, kind of our framework for this is you need to think backwards from the user experience, not forwards from what is easy to measure.
So let's not just say we want to measure F1 and hope that user experience correlates to that. We want to look at the end user value, move backwards from there. Not every eval can or should reflect the entirety of the experience, but you want your evals in aggregate to be reflective of the user experience.
A simple exercise that we do for this when we're building features is we kind of walk down the stack of what is the end user outcome, the business value we're driving, what is a good indicator of adoption slash utilization, and at the lowest level, how are we measuring health of this feature?
It could be something as simple as, like, JSON adherence, variability in the output, and so on. So in practice, this is what it looks like. Every customer is basically giving us their own set of labels. You could think of this as, like, bespoke enterprise AI.
And so we'll actually annotate data for that customer as part of our UAT process, build out use case-specific accuracy metrics for them. Are they trying to do matching? Are they trying to do extraction? And then there's still user feedback as kind of there to close the loop.
But we think it's very dangerous to assume that the absence of user feedback is positive feedback. So we try to put the majority of the onus on ourselves to be rigorous about metrics. And we have various tools to monitor data drift.
Are we getting documents that are very different from the population we've seen so far? We've invested a lot in our synthetic data generation stack. We actually surveyed, like, six-plus providers in the market, tried a bunch of them out.
We're not too happy, and so ended up building our own synthetic data generation stack. Everything that you see was synthetically generated. The way that we think about this is once a use case becomes large enough within the company, we have enough customers doing it, we want to invest in customer-agnostic evals.
That's when we go down the kind of path of synthetic data. And we have a team where part of their job is just monitoring that the synthetic data is distributionally similar to what we're seeing from customers. So we haven't yet cracked the problem of doing this in an automated, fully quantitative way.
The other kind of sanity check that we have is, of course, looking at accuracy scores and making sure that our models are not excessively or underperforming on synthetic data. The next little trick that we use is kind of reducing the degrees of freedom.
So this is kind of part of our architecture where we have numerous features, each of which require a custom prompt for each customer. So in this case, these are four features: free text, tabular extraction, matching, table composition. Each of these has a prompt for each customer.
So you can imagine over 100 customers, you could end up with, like, literally thousands or tens of thousands of prompts. So instead of manual prompt engineering, we have these APE things, automated prompt engineers. Now, the trouble is different LLMs have different levels of performance on different APE tasks and on different customers.
So you get, like, this exponential explosion in complexity. And what we did instead is we just said, well, we're seeing quite a bit of commonality in which LLMs do well on APE tasks. So let's just use one LLM for APE tasks, not permanently, and we'll continuously re-evaluate as new LLMs come out.
But if we can fix this dimension of freedom, it gets a lot easier to iterate. So could we eke out another percentage point of accuracy if we didn't do this? Yes. But building that grid search infrastructure is just too expensive, and we don't think it's a good investment of time.
So this is kind of another trick that we use to just, almost like dimensionality reduction at a project management level. And the last thing I'll mention is, like, identify future potential. I think people spend a lot of time, andrightly so, building evals for what their company does today.
But you should almost have a wish list of what are additional use cases that you want to grow into over time, three, four, five years down the line, because we are frequently surprised that technology is evolving faster than what we can see.
But it is too high of a bar to say we want to have, like, MMLU or BBH-level metrics for future workflows that nobody has asked us for yet. And so what we do is we have these very scrappy kind of future-facing evals.
For example, when GPTV came out, in about an hour, we were able to say, allright, this is our mental model of how it's going to do. It's good at check marks. It's maybe not so good at pie charts, et cetera, et cetera.
And so that, I think, muscle of just building this organizational mental model very quickly in a scrappy way is also something that's been helpful to us. I wouldn't be a startup founder if I didn't end with a shameless plug.
As Sheila said, we just raised a $70M Series B. We are hiring across the board, AI, backend, frontend, go-to-market roles. So if any of this sounds interesting to you, I'd love to chat. Thank you so much. And I'll hand it back to Sheila.
Thank you, Nischal. I have to say, working at Klarity is a dream, so anyone who's interested, please see either one of us afterwards. So this is my call to action. I've seen these paradigm shifts happen in the past.
Call to Action25:26
And I think one thing that's important is to remember the people that are early to this sort of AI revolution that's happening, these people matter disproportionately, and these people are all of you. And so understanding and owning your power as you think about what happens with the next generation of evals, of integrating values into things, it's actually more than just sort of words on a slide,right?
There is a real empowerment and opportunity to spend time thinking about this, to get thisright for the ecosystem and for the industry. So what do we do? We innovate, we reinvent benchmarks. We don't let the leaderboard benchmarks stick that we know are just kind of BS,right?
There's just too much happening today that isn't truly understanding how these systems work. We bring depth into the evaluations. We bring multifaceted nature of these evaluations together. We do so as a community,right? I think it's incredibly important that we bring curiosity, empathy, help to one another.
Like, you know, I love the fact that Klarity wanted to come and talk about their journey, what they did that wasright, what they did that was wrong. We share with one another on that. And I think that we introspect our own value systems,right?
I think that today we are looking at such a pace of innovation that we really need to think, what does it mean to drive this? How are we the trailblazers for this happening? You know, I don't know if folks have read this book, The Alignment Problem.
Brian Christensen wrote a great book, I think it was last year, and it was one of my favorite reads of the year. And he had this quote around sort of if we do see AGI, it will be an interesting mirror of society.
It will tell us which values are uniquely human in nature. Well, those change over time. How do we want to drive a society that has values that we can be proud of as these trailblazers? So feel empowered, be curious, stay empathetic, kind, fair, intelligent,right?
That's what we're doing. We're printing new intelligence for the world. And this is both your opportunity and your accountability. And so thank you all for leading the future of AI.





