AIAI EngineerJun 4, 2026· 23:25

The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI

Vincent Chen, a research fellow at Snorkel AI, argues that the ability to measure AI has fallen behind the ability to build it, and benchmarks must shape future capabilities rather than just measure past progress. Drawing from reviewing over 120 applications for Snorkel's $3 million Open Benchmarks Grants, he presents a framework: the science of task quality, distributional diversity, model headroom, and robust eval methodology, and the art of having a thesis (e.g., Terminal Bench's bet on CLI before coding agents made it obvious), producing research roadmaps, and treating researcher UX as a first-class citizen. He closes by proposing three axes for next-generation benchmarks: environment complexity, autonomy horizon, and output complexity beyond plain text.

  1. 0:00Intro
  2. 1:39The Gap
  3. 5:09Framework
  4. 6:21Task Quality
  5. 8:27Distributional Diversity
  6. 9:53Model Headroom
  7. 11:28Robust Evaluation
  8. 13:24Thesis
  9. 14:37Roadmaps
  10. 15:59Researcher UX
  11. 18:12Next-Gen Benchmarks
  12. 22:41Closing

Powered by PodHood

Transcript

Intro0:00

Vincent Chen0:15

Hey everybody. How's it going? Lovely. Well, I'm very excited to be here, hailing from San Francisco. It's a little bit of a trek over, but I'm super excited to chat with you all today, after the talk and beyond.

My name is Vincent. I'm a research fellow and co-founder at Snorkel AI. And, you know, today I'm going to be talking about some meta-evaluations for building benchmarks, the art and science of what we've found to be really useful when building effective benchmarks.

I have the great privilege at Snorkel of working with both our researchers, great collaborators in academia, industry, and the open-source community to build great benchmarks. And I wanted to share some of the learnings that we've had over the, you know, last few years on what really makes for benchmarks that shape the field and move it forward.

So, a little bit about us. We're a frontier AI data lab. We have, you know, labs both at academic settings. You know, our co-founders have labs at Stanford, UW, Wisconsin. We have an internal team of four deployed engineers, applied engineers.

And I mention this because we get a lot of exposure to both, you know, the academic frontier, work with frontier labs, and also real enterprises and companies who are deploying in practice. So our focus as a company is on building the best data sets and environments to define and advance future AI capabilities.

The Gap1:39

Vincent Chen1:39

And we are in a fortunate spot where we get to play at this unique intersection of both, you know, the frontier academically, but also to contact reality with our deployments in enterprises. So today I wanted to talk about an asymmetry that we see in our real-world deployments.

There's real excitement around agents today. You know, we see it. Every person in this room, I'm sure, has played with these agents. And we see real progress marked by, you know, hill climbing on model cards. We see the vibes are improving,right, and truly shifting, especially in coding.

But when you ask individuals and enterprises or these, you know, large-scale organizations if they're fully ready to let these agents loose and, you know, deploy them in high-stakes environments, you get a little bit of hesitation. And that's not to say the capabilities aren't there, but our ability to actually measure these agents in practice, that is falling behind of where the capabilities actually are.

This is one of the challenges and research questions that I think are actually one of the most important in the field and one of the ones that we're very interested in here at Snorkel. So closing that gap, that evaluation gap, we believe requires a toolkit,right?

As I mentioned earlier, we're strong believers in field deployments,right? So this is, you know, actually deploying engineers, researchers on our teams to, again, contact reality and work with folks to deploy these models in real production settings where the stakes are high,right?

These are finance settings, insurance, you know, healthcare settings where it's not just about a number, it's about real outcomes. And we're also big fans of other eval tools,right? This is red teaming, private human evals, crowdsource labeling, a lot of the themes that we saw talked about today, which has been awesome to see.

But one of the things that we feel most strongly about is that

benchmarks, open benchmarks in particular, remain a really critical piece of the measurement toolkit. The best open benchmarks aren't just about, you know, taking a snapshot of progress looking backwards. They're actually about defining progress and shaping the field and setting a goalpost about where capabilities need to go.

And, you know, even looking at the last few months,right, benchmarks like Terminal Bench, Meters Long Horizon Benchmark, ARC AGI, these are really exciting and critical guideposts for where the field is going. And as a result, the path to safe, trustworthy agents will really depend on more of these benchmarks in practice.

So what are we doing at Snorkel? One of the things that, again, I'm very fortunate to be able to be a part of is the Open Benchmarks Grants. We recently, a few weeks ago, a month ago, deployed $3 million to commit to open benchmarks.

And this is a really fun job. I get to work with the best academic teams, you know, builders, to really accelerate and fund the next wave of benchmarks that's going to really steer and guide where the field is going.

We've had a wild reception so far. I'm a little, admittedly, behind on some reviews, but we've been really excited to see what the community has come up with so far. And in this talk in particular, you know, we've reviewed, I think, over 120 applications so far spanning academia and industry labs.

We wanted to share a few perspectives, a few learnings over the past few months about what we view as, one, table stakes for useful benchmarks,right? How do you actually build good empirical measuring sticks that are actually useful, you know, to measuring progress?

And two, what really separates, you know, those benchmarks that are shaping the frontier,right? What is the art and the science, if you will, of building really effective benchmarks at the end of the day? So as I'm doing this, I'll have a fun opportunity--maybe this is a little too American--but to pull out Timothy Chalamet and honor some of the greats, you know, some of the great benchmarks over the last few years that have really shaped the field in talking about some of these axes.

Framework5:09

Vincent Chen5:35

And I hope that these, you know, themes resonate with you and also inspire a bit of, you know, kind of new thinking about, hey, how can we actually deploy some of the learnings we're all kind of driving towards in our day-to-day work to, you know, shape the field and move it forward?

So two themes here, again, on the science side, you know, how do we actually build effective measuring sticks? We'll talk about task quality, distributional control, robust evals in general. And on the art side,right, really the differentiators for great benchmarks.

How do you build benchmarks with a thesis on where the field is going that inspire new roadmaps and that critically are built for this audience,right, a researcher audience, a builder audience, so that adoption is something that is way smoother and a first-class citizen for a bunch of these benchmarks?

So let's start with the science. This is, again, what makes for really effective measuring sticks as we've seen them in practice, in the deployments that we see in industry and academia and with frontier labs. So the first theme I want to talk about is individual task quality,right?

Task Quality6:21

Vincent Chen6:38

This is the idea that individual tasks need to be exceptionally rigorously validated,right? They need to represent real-world complexity. They need well-posed, well-structured instructions. They need verifiable solutions that ideally have been actually validated by real-world domain experts. One of the benchmarks here, GPQA, is one of my favorites, not just because it's been a very lasting and enduring benchmark that captures, you know, graduate-level and professional knowledge, even to this date,right?

You kind of still see this on model cards. But one of my favorite contributions is actually tucked away in the appendix. GPQA introduced, one, new adversarial quality control mechanisms. So the idea was that not only do these tasks need to be well-posed, they need to be tractable for other experts to solve.

So they had a very rigorous multi-reviewer protocol where there was an original author, you know, there were reviewers and adjudicators in the loop. There was opportunity for revision,right? These were tasks that were really pushing the frontier of knowledge, and it was non-trivial for any single expert to say, yeah, this is actually a good task or not.

And so developing this sort of rigorous adversarial quality control mechanism was one of the contributions I was most excited about here. And if you read the appendix, you also see that they introduced new incentive mechanisms,right? Payouts were actually based on whether there was certain agreement.

And, you know, coming from academia, you know, there's some inspiration here from the peer review process, as flawed as that is. But, you know, this type of innovation around how you actually get really rigorous, you know, multi-expert quality control leads to the type of outcomes that we see around individual task quality that we see as a key foundation for any benchmark that matters at the end of the day.

Distributional Diversity8:27

Vincent Chen8:27

Two is distributional diversity,right? This is the idea that for any benchmark that really matters, you want to define a clear taxonomy for the domain, for real-world tasks, and distribute those tasks intentionally. So this might be, hey, I captured a trace or a kind of real-world stream of traffic along, you know, how my agent is operating in the real world, and I want to really represent that distribution.

It could also mean, hey, I'm specifically characterizing and taxonomizing the failure modes that are, you know, paradoxically rare, but, you know, disproportionately important in production,right? If you take classic self-driving settings,right? Yellow lights or, you know, pedestrians or motorcyclists,right, might actually show up way less than other types of scenarios but are disproportionately important, you know, to getright in these settings.

And so defining that taxonomy, being really intentional about distributing tasks across it is one of the hallmarks of great benchmarks in our view. MMLU, a few years old now, constructed a quite ambitious taxonomy of, you know, 57 academic and professional domains across STEM, humanities, et cetera.

It's remained one of the lasting benchmarks for understanding graduate and professional-level knowledge. And again, a lot of this was, as we believe, a result of really thoughtful and intentional taxonomy design and building towards that.

The third axis here is around difficulty of individual tasks and model headroom,right? It's really important that the benchmark is unsaturated, that it exposes real soft spots in capabilities and reliably separates where models sit at the frontier. One of my favorite plots is the one on the topright.

Model Headroom9:53

Vincent Chen10:12

This was, you know, all credit to the ARC Prize Foundation team. ARC AGI 2, you know, for a very long time was unsaturated,right, for several months and years. And when there was the big reasoning push, you know, maybe 18, 24 months ago, we saw a massive leap in capabilities that actually corresponded to a real leap in model capabilities,right?

This was a benchmark that was intentionally designed to represent a type of efficiency or capability that humans have but models didn't have. And they really kind of captured well, hey, there's a lot of model headroom here. Humans can do this.

Where's that gap? And again, lo and behold, it correlated quite well with the recent, you know, O1-style reasoning push that has really dominated the field in the past 18, 24 months. Just a few weeks ago, the ARC team just launched ARC AGI 3.

And again, at launch, they had frontier models under 1%. Every single task was human-solvable to some degree. And so it remains one of, I think, the most meaningfully exciting benchmarks in the space where, you know, any new model, you know, people are kind of awaiting, hey, how does it do on ARC?

And I think they did this quite well,right? The kind of model headroom here is really, really exciting.

This last axis here I want to talk about on the empirical measurement side is all about robust eval methodologies. Now, this goes really deep. So just kind of capturing some of the high-level ideas. Benchmarks need to ideally go beyond accuracy to capture real-world dimensions that matter,right?

Robust Evaluation11:28

Vincent Chen11:44

This is everything from cost, latency, you know, the quality of the reasoning traces, some of the intermediate steps and tool use, whatever dimensions actually matter for the capability at hand. Capturing those as reward or supervision signals is really critical.

And measuring what it claims to is actually a non-trivial feat in, you know, building robust and reproducible benchmarks. So TaoBench is a benchmark that we're a big fan of. You know, it's had multiple evolutions over the year, but over the years, but it was a benchmark that was built to evaluate both task completion of these multi-turn agents.

They built a clever, you know, kind of user simulator, but also, you know, not just accuracy and completion, but adherence to policy constraints. So a model, for example, on theright-hand side, this was one of the examples from the paper, a model that books theright flight, but violates fair-class rules, still fails.

It's still a kind of no-go at the end of the day. So this notion of, hey, being intentional about what axes we actually care about, what do we actually want to measure, and measuring that rigorously is one of the hallmarks that matters when you're building these frontier evals.

So I want to shift a little bit now to the differentiators,right? What actually leads to the benchmarks that push the frontier? And that's not to say, you know, anything I mentioned on the last few slides have not pushed the frontier.

These are just special characteristics that I view as, you know, critical to the benchmarks that are real research contributions that are really shaping where's the field going, where are all the labs going to hill climb next? And this is the art, the special sauce that helps push us forward.

So one of the key hallmarks here is these benchmarks should have a thesis,right? They should have a research question about a subspace of capabilities, about where the field is going. It should revisit previous capabilities. And the most ambitious benchmarks are really a statement about where the world is going.

Thesis13:24

Vincent Chen13:41

Terminal Bench is one of these bets,right? It was a bet on the CLI, not just for coding agents, but for general-purpose computer use. And in many ways, I think this has turned out to be a largely correct and consequential bet,right?

As we're seeing teams at kind of Claude and Codex build their general-purpose, you know, enterprise capabilities on top of these coding and CLI-based tools, we're seeing this bet pan out. And again, Terminal Bench remains one of the most robust and kind of most important benchmarks that are measured on all the recent model cards.

So again, this was a bet early on to say, hey, we think the CLI is going to be really important as a core interface, a core abstraction and affordance for agents to interact with the real world in a general-purpose way.

And by measuring those capabilities, you know, I'd argue that it actually helped accelerate, you know, how the field is operating in this way.

The second piece here, I think, worth mentioning is the ability to kind of roadmap for the field. A great benchmark, you know, one that really shapes where all of us are going is producing new roadmaps,right? It's inspiring new attacks against research problems.

Roadmaps14:37

Vincent Chen14:52

It's helping folks ideate and come up with new ways for thinking about benchmarks and methods in general. And I think SweBench is a really phenomenal example of this,right? It was a simple idea. Often the best ones are quite simple,right?

How do you kind of leverage, you know, existing coding-type capabilities via PRs? And it spawned a new family of benchmarks,right? All the way from SweBench Lite, Verified, Pro, Multilingual, Multimodal, et cetera. And its evolution, I think, is still very relevant today.

It's evolved how we think about coding agents. And one of the things that's been awesome to see with the SweBench team is how many new research directions and kind of inspired benchmarks have come after it in this coding space.

And arguably, I'd say, you know, there's a lot more room to kind of innovate on top of this as well. What do the new ways of coding look like? How do these types of workflows apply to vibe coding and kind of this new layer of abstraction that software developers are applying?

I think it's been really exciting to see the foundation that the SweBench team sets and how that's going to shape, you know, how we think about coding agents moving forward.

Researcher UX15:59

Vincent Chen15:59

So this theme here, I think, is severely underrated. And this is the notion of researcher UX,right? I think the most prescient benchmark builders are committed to the researcher and builder experience. This is to say, it's really simple to run models and agents against your benchmark.

It's really simple to contribute new tasks to extend. And also, it's really simple to leverage some of the signals that you're getting for the benchmark for RL or kind of tuning post-hoc. I think this is really underrated. It's, you know, a classic product principle to make what you're building and putting out there easy to use by the community or by your core users.

And in this case, benchmarks have core users, which are other builders or researchers. And so really putting in time and attention to building those interfaces has been important for the adoption of some of the most important benchmarks,right? To call out, I think the Stanford team at CRFM built HELM, you know, several years ago, which I'd argue kind of pioneered a standardized modular harness for evaluating reproducible, you know, different scenarios as well as kind of models against a standard test bed of models.

Terminal Bench 2.0, just a few months ago, again, shipped with Harbor, which has been, in many ways, the de facto harness and kind of evaluation infrastructure for teams who are building agents more broadly. And so, you know, thankfully, we have a bunch of open-source software out there today, you know, based on this principle.

But as you're building your benchmarks,right, kind of considering, hey, how easy is this to extend? How easy is it for the community to adopt and eventually kind of hill climb against this? I think is a severely underrated factor for what makes for really high adoption of these frontier benchmarks.

So this is the full framework. Again, can go into more detail and please find me afterwards. But again, what makes for really empirically meaningful measuring sticks,right? It's task quality and attention to distributional control and diversity. It's difficulty and model headroom.

And of course, a robust eval methodology that measures the concrete axes that actually matter in practice and is intentional about it. And of course, on the ARC side,right, these great benchmarks really have a thesis on where the frontier is going.

They set roadmaps for the field, and they really prioritize researcher UX. Now, before I wrap up, I want to propose, you know, a few dimensions that we're really excited about at Snorkel that we think are really going to encapsulate, you know, the next wave of benchmarks.

Next-Gen Benchmarks18:12

Vincent Chen18:29

Tried to leave some more degrees of freedom here for creativity, but these are areas where we think there's a lot of room to push complexity, to push, you know, realism in benchmarks. And so I wanted to share a little bit of our internal roadmap and thinking around where the field is going and where we need more benchmarks and more contributions.

So this is our point of view. We think that the axes for the next great benchmarks are threefold. And I'll go into a little bit more detail about what I mean here in just a second. But one, it's environment complexity,right?

How complex, how realistic, how dynamic is the operating environment that these benchmarks are working? Are they representative of real-world settings that a professional, that, you know, a scientist, that someone using these tools could actually use? Two is autonomy horizon.

Do these benchmarks represent realistic and frontier horizon lengths that these agents are operating against? Are they capturing different points on the autonomy slider that are, again, representative of how users are using them as co-pilots versus fully autonomous agents?

Is this an intentional design in the benchmark? And three, capturing the wide range of output complexity. I think this is very under-explored today,right? Lots of chat-based or document-based outputs, not as much around nuanced, kind of differentiated reward signals,right?

Real artifacts that show up, you know, in day-to-day work that we represent. And critically, new artifacts,right? New types of form factors that we haven't even imagined about, you know, how agents interact with humans, how agents interact with each other.

So a little bit about each one. The first one here, I won't go into all of these in significant detail,right? Environment complexity is all about capturing the real-world complexity that is in our day-to-day working environments,right? And this gap is often where agents fail today.

Consider coding agents,right? A real codebase has org-specific policies, you know, lots of Slack context, screenshots, flaky tool chains, you know, CI that's kind of distributed. Human reviewers with knowledge in their heads about what they like and what they prefer.

Many contributors in parallel. Benchmarks today capture a fraction of this complexity and, you know, not just in coding, but other domains. There's a lot of excitement and opportunity to up the level of complexity and continue to drive what these models can do to represent real-world uses.

Two, again, on autonomy horizon,right? This is all about how long an agent can operate before reliability breaks down,right? Let's take a customer experience agent. You know, in many cases,right, these agents may lose track, you know, of context that was, you know, delivered a few weeks ago.

Different integrations or product specs might, you know, change the actual spec or requirements for a particular model. Reorgs can kind of shift, you know, priorities midstream. Real-world settings actually represent a lot more complexity that, again, is represented in these kind of long-term continual learning type settings that represent kind of changes in state and environment.

So we, again, think that there's a lot of room to contribute and build out new benchmarks that represent very, very long horizon and autonomous agents. And lastly, this axis is all about producing more complex work, more representative work, and also nuanced signals that can be used for not just evaluation, but reward signals during training.

This also has to be complex, and this gap is growing as well,right? Let's take, again, a software example,right, or a complex report for making strategic recommendations. It's non-trivial and subjective, you know, to define, hey, what is verifiable about a good recommendation, a good strategic proposal, a good roadmap in general?

The nuances of this need to be captured well. They need to capture organizational context, really good human judgment. And tomorrow's benchmarks really, you know, we're excited about signals that capture all of these settings. Trustworthy outputs,right? The ability for agents to actually capture their own uncertainty and define, hey, I'm actually not sure about this.

I actually need to stop or kind of ask for more information. Again, different types of outputs that aren't just, you know, a kind of plain text answer is something we're really excited about. So again, hopefully, you know, this inspired a little bit of thinking around, hey, what am I working on?

Closing22:41

Vincent Chen22:41

How can I turn this into a meaningful benchmark? We are still accepting, you know, benchmarks in the Open Benchmarks Grants. So if you're excited, please reach us at benchmarks.snorkel.ai. Feel free to reach me directly. And we're really excited to see where the field is going and, again, to use benchmarks to not just measure, you know, progress looking backwards, but really shape where things are going moving forward.

Thanks for your time and excited to catch up soon.