The Map0:00
Hi, this is Ari Helczak from Rootsignals, uh, presenting agent evaluations—finally, with the map. So, the agent evaluation is rather art than science, and ultimately it is nonetheless required to actually ensure that your agents do what you expect them to do, especially when you're launching them in production.
So, without further ado, let's just, uh, diveright there into the map and start to get an understanding of what are all the things involved if you want to evaluate all the aspects of your agents. Uh, so the agent evaluation can be actually neatly divided to the semantic and behavioral parts of what the agents do.
The semantic part is all about how do the representations, uh, of the reality that the agent has actually relate to the reality, whereas the behavioral parts, uh, are all about how do the actions and the tools that the agent is using actually contributing to the achievement, uh, of the agent's goals in its environment, and ultimately what kind of effects it will have on its environment as well.
The semantic part, uh, is, um, further divided into what could be called the single step or single turn, uh, items, which are, uh, like coherence, consistency, etc., so various semantic virtues, and what can be called the multi-turn aspects of the same, so this related to chat and reasoning.
Whereas in the behavioral part, you also have this distinction whether you're looking at the task progression and planning, or then individual selection and usage of tools. So let's, uh, dive into each of these, um, paying also attention to the fact that truthfulness, uh, is ultimately achieved by grounding, uh, the representations in your data, often with RAG.
Whereas, uh, the goal achievement is what actually, uh, is achieved by grounding to the tools that the agent has available. And this is sort of, uh, symmetry that is not accidental, because in a sense, uh, there's similarity and analogy between, uh, representations and behaviors, because representing the world is a kind of activity.
So you can see that, uh, that the representations, uh, are, in a sense, a special case of tools, special case of, uh, of behaviors. And, uh,
Single-Turn Semantics3:01
let's then dive into these individual parts of the map. So first, looking at the semantic quality for the single turn case. So here we have these universal virtues. There's a long list of these that you can look at, uh, and these are actually non-agentic, in a sense, and we are only covering them here because of completeness, so that you can sort of understand these agentic parts through this contrast.
So these are things like, uh, is the, uh, the reply that the agent is giving to the user, is this consistent, is the, is the content actually safe, and so on and so on. And then there's the inter- interesting and important part of whether whatever the agent is actually saying, if it's, uh, aligning, uh, with the values of, uh, of the organizations or people, uh, who are the stakeholders, and also whether it adheres to the policies of the same.
RAG, or what could be more generally called the attention management, uh, is then something that you can measure through specific evaluators, such as those looking at whether the retrieved context, uh, was correct, whether all of it was comprehensively recalled, uh, and so on and so on, and ultimately relating the answers, uh, to the external reality through what is often called the faithfulness, which is then separate from, uh, from the both the answer question relevance and also from the notion of factfulness in general, which, uh, relates to the reality beyond just the reference data that the RAG is using.
So now, moving forward, uh, we should also pay attention to the fact that the RAG evaluations actually come, uh, in many forms, and they're, they also have certain symmetries, uh, that are not that difficult to understand when you look at them, uh, through, uh, a map like this.
So, uh, they are all just, uh, essentially looking at different relationships between the parts of the RAG pipeline. But this is not the, the main topic of the presentation, so you can look, look up more information through the, uh, background material links that we will provide.
Multi-Turn Semantics5:23
So then moving to the multi-turn case. The multi-turn, in the semantic quality sense, means essentially chat, conversation histories, how do these develop, what sort of things we want to look for in these, consistency, etc., adherence, and sticking to the topics when that is actually necessary.
Sometimes we want to allow changing topics if the, for example, a chatbot, uh, user is, uh, user wants to change the topic, or sometimes we don't, but we need to be able to, uh, be aware of this. Then, super important, uh, and almost groundbreaking, progress has been achieved by, uh, by, uh, investments in reasoning.
So you can also, uh, actually evaluate reasoning traces, the reasoning, uh, chain of thought, uh, if, if allowed to use that term here, and that's sort of another way of looking at the kind of, uh, sequential or mul well, multi-turn and sequential, uh, kinds of activities that the agent is just doing in the course of its reasoning and representations of the world, before even taking any kind of actions.
And now we move to the action side. So here we have, uh, things to evaluate, such as whether the agent is actually following instructions, is it, uh, extracting to the, the tool characteristics, uh, correctly, is it selecting theright, uh, is theright tool being selected, is the output quality of the tools correct, are the errors it chooses handled correctly, are the structures, uh, in the JSON, tool usage, uh, format, are they, uh, correct.
Behavioral Evals6:29
So these are all, all the things that relate to, uh, to the behaviors, even before you have a chain of, uh, behaviors. And when you move to this chain of behaviors, or multi-step, or, uh, multi-turn, uh, cases, so then you actually start to look at, like, a-are the actions that the ac- agent is taking converging towards actually achieving its goal, and, uh, is the plan that the agent might have actually consistent, and is it actually, uh, high quality, in whatever sense you want to measure that.
Uh, so then both of these are actually grounded, uh, on external reality. So, uh, like mentioned, the representations are grounded on truthfulness, and, uh, the activities and behaviors are ultimately grounded by goal achievement and utility. Uh, and this is, uh, what sort of, uh, is the ultimate, uh, metric to everything else, I know, uh, that the agent is trying to do, and all these other things along the way are sort of just proxy metrics, so to speak.
Practicalities8:08
Uh, moving to other practical considerations. We can only scratch the surface here, so I'm just sort of listing, not to leave this out, but, uh, they don't neatly fall into this, uh, general map that I presented, but should also be taken into account on the background.
So most important of which is the cost and latency optimization. So generally you want to, the agent to progress towards its goal as quickly as possible and as cheaply as possible. So cost and latency optimization, the number of steps optimization, and then moving to the, what you often are going to need, which is the tracing and debugging, so being able to actually see where did the agent go wrong.
Error management here refers to, to dealing with the, with the errors of the tool usage, so not the actual, like, semantic errors that, that the agent is doing in the course of its, uh, inference and thinking process. Uh, very important distinction to keep in mind is the offline versus online testing.
So what are the things that, uh, that you can actually, uh, try out, uh, and evaluate during development, and what are the things that you can do and must do, or should do, uh, during the actual online activities of the agent.
And these are actually going to become, like, uh, two very distinct dis- dimensions that could be also included on the map, but, uh, but that would make the map, map more complicated, so that's why I'm just mentioning them here now separately.
And then we have various special cases that I don't have time to cover. Some of these, uh, are, uh, more refined and more advanced than others, and, and some of these things are sort of, uh, more researchy, and there's a lot of papers on each of these topics.
And, uh, uh, depending on what you're doing, some tool-specific metrics might actually be useful to add to the mix, uh, even, uh, uh, and they might even be rather simple to implement because the tools are often something as easy, uh, as straightforward as, like, API calls and, uh, and such things that can actually be just measured separately, uh, uh, using more traditional software testing methodologies.
Eval-Ops10:24
Uh, one important part to mention here is that a lot of these measurements are going to be implemented with LLM as charge techniques. And now the caveat, uh, behind these is that, that quite often when people start implementing these evaluation methodologies, they are looking at this kind of a single-tier approach.
Single tier here means that you're focusing on optimizing this operative LLM flow, meaning that you're optimizing your agent, which makes sense. But if you only are concerned of, uh, with getting some scores on what the agent is doing, then you're sort of forgetting that what about all the cost, latencies, uncertainty related to the charge itself, the charge that is on the background.
So you should be actually, uh, taking into account as early as possible that you are going to need to also optimize the charge itself. And this is what, uh, what we could call double tier. So you need to optimize both the operative LLM flow that powers your agent and then this judgment flow that actually powers your evaluations.
And, uh, this is, uh, rather complex situation in general, and we are calling this eval-ops because, uh, this seems like, uh, like a separate kind of thing that involves, uh, evaluations that themselves are so complicated, uh, so expensive and slow that they sort of earn their own, uh, category of, uh, of activities.
And this is, uh, something that, uh, that, uh, we have written about, and, uh, and you can also find more about this on the source materials. Uh, the general, general gist of it is that, that, uh, eval-ops is, uh, kind of a special case of LLM ops, and, and actually operates on different kinds of entities that, uh, that the LLM ops, uh, uh, in general.
And, uh, and then it also requires different ways of thinking, different kind of software implementations, and also, uh, essentially a different kind of resourcing to, to get those evaluationsright. So thank you very much. This was only a brief glimpse on the, on the general landscape.
I hope the map is helpful to you, and, uh, please take a look at the source materials, which will give you more depth to each of these, uh, these topics. And, uh, happy to discuss any of these things, and let me, uh, know if I'm forgetting something crucial, and, uh, this, uh, presentation will, of course, be obsolete by the time, uh, it, it goes out.
Closing12:39
So, uh, probably when, when you're watching this, there's already some major developments have happened, but that's, uh, that's how it is on this general life of, uh, of, uh, AI engineers. Uh, so let's go out there and make our agents measurable, controllable, and let's make sure they are actually doing our bidding and, uh, not, not rebelling against our, uh, ultimate intentions.
Thank you very much.





