AIAI EngineerJul 8, 2026· 19:09

Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs

Nuno Campos from Witan Labs explains how his team spent four months teaching coding agents to master spreadsheets, raising accuracy from 50% to 92% on a financial analysis benchmark. Key failures included a rigid three-agent architecture and standalone representations like SQL and XML. The breakthrough was replacing 15 separate tools with a single JavaScript repo tool offering persistent state and code-mode semantics, enabling agents to combine multiple operations in one call and drastically reduce timeouts. The team built high-fidelity formula and rendering engines to close the verification loop, and added domain knowledge prompts to focus the model. For evaluation, they moved from LLM-as-judge to deterministic comparisons using golden spreadsheets. Campos urges practitioners to replace many tool calls with real scripting languages, invest in feedback loops, and regularly revisit interfaces as model capabilities evolve.

Transcript

Spreadsheet Challenge0:00

Nuno Campos0:17

Cool. Uh, hi everyone. My name is Nunu, and I want to talk to you about how we spent the last 4 months teaching coding agents to master spreadsheets. Um, so, essentially our goal was to get coding agents to be as good at spreadsheets as they are at, you know, Python, JavaScript, or whatever your favorite language is.

We started at around 50% accuracy on a financial analysis benchmark and got to 92%. So I'll chat about what actually moved the needle, and what didn't, and dead ends, and, uh, yeah, let's see.

Uh, so spreadsheets are a little bit harder for AI than you might think at first. If you, uh, if you think about how you would, you know, if you open Excel, how you'd find your way in an Excel file that you don't know, it's actually a very visual thing.

And you, you know, you just instantly see the structure: there's a revenue table here, assumptions in there, a chart in there, and it just, you know, feels intuitive and you don't even think about it. And, um, an LLM doesn't really see any of this.

It, you know, if you ask it, "What's the revenue?" then it has to figure out, "Which revenue do you mean?" You know, there's net revenue, gross revenue, revenue for this, revenue for that. Uh, which, which quarter, which year?

And then, uh, is the number found like an actual input? Is it a formula? And it's, so it's actually a deceptively hard task.

One, uh, thing we tried close to the beginning was to split the work into three agents. Um, so the, kind of, the central one was the edit agent that had like a five-step process, um, that, you know, you'd define the, the end state, you'd do a plan, you'd execute, you'd verify, all, all the things you're supposed to do.

Exploring Formats1:54

Nuno Campos2:14

And this, uh, kind of changed the kind of errors we got. Uh, without it, the, the agent would just make mistakes while actually building a financial model or something. And with this, it would maybe make those mistakes while planning, which was a lot easier to rectify.

Um, but this architecture in the end was too rigid because, you know, discovery ran once up front, uh, and then you couldn't revisit it, and, uh, the context wouldn't flow between the different agents. So it just turned out to be one dead end.

Uh, then some more dead ends. Uh, we, I think, ended up probably trying every conceivable way of representing a spreadsheet to an LLM. Uh, none, uh, really worked as a standalone representation, but two, uh, turned out to be useful as, uh, methods inside the, the repo that we ended up, uh, creating.

Uh, but they all had something going for them in theory, and that's why we tried it,right? So SQL has obviously been around for decades, so, you know, super popular in, uh, LLM training data, so agents are really good at it, supposed to be a great way to deal with structured data.

Uh, but it turns out that, you know, it doesn't quite work for this. Uh, XML is how Excel files are represented on disk, so, you know, maybe that was a good idea. It wa-it wasn't. Um, and, uh, you know, many others.

In the end, uh, we did get two, uh, useful things out of this. One was the concept of, uh, having these CSV or TSV, uh, views of part of a spreadsheet. Uh, this turned out to not be that great as the only way to interact with a spreadsheet, but as, uh, one piece of, uh, of the larger solution, uh, turns out to be used very, very often.

Uh, and HTML was also a step in theright direction as it, you know, introduced the idea of a layout and formatting and so on. Uh, so that ended up resulting in us building a rendering engine to let, um, uh, the agent see what the rendered spreadsheet looked like as an image.

Uh, and then, uh, you know, eventually we, uh, hit on, uh, what's was probably ended up being the biggest, uh, breakthrough, which was to replace the, you know, the many tools that we accumulated over time. I think at that time we had around 15 tools, uh, with a single, a single tool, which was a, a Node.js repo.

Single Repo4:22

Nuno Campos4:43

Um, and so, you know, all the, all the 15 tools that, that we had to start just became different JavaScript functions that, uh, the agent could combine in this one, uh, repo call. And why, why JavaScript? Uh, we needed a scripting language that's easy for, uh, to sandbox.

It's easy. LLMs are, you know, uh, very, uh, familiar with it. And, uh, Python would probably work equally well. We just went with JavaScript. Uh, but the actual implementation of the code that deals with the spreadsheet is actually in a completely different language, language in C#.

Um, uh, and that's kind of the advantage of this architecture. You just, you know, use the scripting language for what it's good at, which is letting the agents interact with it, and use theright language to then deal with the actual files.

And what it looked like before and after. So before, it would be you'd have, you know, 10 or 15 tool calls usually for an agent to, uh, to explore a spreadsheet and get to an answer. Um, and this would actually very often end up timing out and taking a long time because it was just, you know, doing things sequentially.

Even parallel tool calling didn't really help because you couldn't combine the results in any way. And after, you could, the agent would just combine, you know, the different things it wanted to do in a single tool call and get all the results at, at the same time.

Um, so some of you will be familiar with the idea of code mode. Uh, I think that's becoming more popular, uh, showed up in the Anthropic API, and Cloudflare has talked a bunch about it. Uh, a repo is actually, and the code mode is, you know, already super useful,right?

Because it's, that's that basic idea of combining multiple tools into a single, uh, into a single tool call. Uh, but the, a repo actually goes further. And the difference is it's basically code mode with persistent state so that, uh, you know, the agent calls the repo tool once, defines a few variables, and then sees the results, spends a few more reasoning tokens, and then the next time it calls the tool, those variables are, are still there.

So that means, uh, it actually can, uh, you know, build on its work. And what we observed with this is that just pure code mode without the repo semantics, uh, agents would very often write quite long scripts. Like 50 lines of JavaScript would be pretty common, um, which, which is great.

It means they're doing many things at the same time. But with a repo, they would actually write shorter scripts, um, which meant it could basically do more interleaving of putting reasoning in between each of the things that the, uh, agent was doing, uh, which, uh, many times resulted in, uh, the agent getting to a better answer, uh, faster because it was less, uh, static.

And the, another nice thing about this design is, um, in the previous way where we had separate tools, if we, you know, if we figured out, oh, there's a new, uh, a new method we need to give the agent access to, to, I don't know, explore the dependencies between formulas or something.

So that would mean creating several more tools that are going to go into the tool schema and, uh, we need to see how they play with each other. Uh, whereas with this approach, um, all it means is making a few more methods available in, in the JavaScript repo.

And making the agents aware of that is as simple as creating a TypeScript type, type definitions file and putting it into the prompt. And that.

Guest8:18

But.

Nuno Campos8:18

Well.

And.

Guest8:22

But.

Results8:24

Nuno Campos8:24

Uh, out of all of this was, uh, you know, we went from, as I said, uh, 50% before we had the repo, then 74, and then, you know, over time we made more changes, none as dramatic as the repo that eventually got us to 92% on this, uh, internal benchmark we have.

And these were changes like giving the agent better fuzzy search or formula tracing, uh, functions for dependencies or improving the system prompt or just fixing bugs. Um, uh, but, you know, it all adds up to a nice result.

And, uh, another thing I want to, uh, call out is, is the timeouts. So, uh, this approach ended up really, uh, you know, fixing the tasks that would time out. We usually ran tasks with a five-minute timeout because if it takes longer than five minutes to answer a question about a spreadsheet, that's not particularly useful.

Um, and this approach essentially resulted in zero timeouts because it just was a lot more efficient for the agent to, to do its thing.

Verification Loop9:28

Nuno Campos9:28

Um, there's a lot of parallels between, uh, spreadsheets and, and coding. Um, I'm sure you all use Claude Code or Codex or whatever coding agent, uh, every day. And it does a much better job when it can, say, run the compiler for your language or the linter or your tests and then iterate based on those results.

And, uh, you know, when we write code manually, that's true as well,right? If we're not allowed to compile or lint or test the code, then it's not going to produce a great result. And the same is true of spreadsheets.

And, but to, to enable that for, uh, for spreadsheet work, uh, we had to build, uh, uh, a couple of, uh, engines that, uh, can close that feedback loop. The two most important ones is, one is a formula engine to calculate the, the formulas.

And another one is a, a render engine to render the contents of a, of a range into like a, uh, a, uh, an image, uh, with all the formatting and layout and so on. And that's kind of the source of truth.

It's the verification loop that, uh, that makes the agent, uh, you know, confirm that it did theright thing, and when it didn't do theright thing, go and fix the formula or go and fix the formatting, uh, in order to, to make it correct.

But it only really works, uh, if the engine is actually high fidelity. Uh, so if you use, um, an incomplete engine that implements, say, 50% of the formulas in Excel, then what you end up with is actually worse results because the en the agent is going to write a formula that it thinks would work and in practice would work, and then it's going to try and compute it and it's going to get the wrong result or going to get an error because it's not implemented in the engine.

Uh, so the ver that verification loop is really only as good as the engines that power it. Uh, so this ends up with, you know, two, two different things. One is the repo, and that's an interface. It's how we present, uh, uh, our, uh, our tools to the agent.

Interfaces11:36

Nuno Campos11:36

And, uh, it, you know, a repo is the best interface that we could co-come up with today because coding is what, uh, the current state-of-the-art models are the best at. But that's not necessarily going to be true forever,right?

The, the labs are working on computer use a lot, so, you know, eventually maybe the models will be as good as at computer use with a mouse and keyboard as they are at, at coding. And at that point, maybe a repo is not going to be the best interface.

Um, but it's the best one today. What won't change is the need for, uh, for, you know, for that verification loop. And, uh, what's behind it is actually, uh, I think the, the more durable part because the more capable the models are, like there've, there've been, I don't know, four or five, uh, model releases while we've been doing this work.

And every time we've seen the more capable the model is, uh, the more they can get out of that verification loop.

Domain Prompts12:37

Nuno Campos12:37

And, uh, another thing we ended up doing is, um, uh, adding domain knowledge, uh, to, to the prompts. Um, a-and that actually ended up, you know, surviving, uh, uh, a-all of the different iterations of the tools. And, and it always, you know, produced, um, improved results.

Um, and this is not so much because the, the LLMs out of the box don't know what, I don't know, revenue or ARR means. It's more because they, you know, know many, many things and you kind of need to, uh, pigeonhole them a bit, a little bit into what you're, what you want them to focus on, um, for, for the specific task that you have.

And, uh, and it's actually super portable. Like this, almost the exact same, uh, prompt would work for the repo or the individual tools or, uh, or, or any of the other approaches.

Uh, I also want to touch a little bit on evaluation. Uh, it ended up being, uh, a lot of work, uh, to evaluate this stuff. And it was actually, you know, a really important part of what actually enabled us to be sure whether, you know, the CSV or the SQL representation were good is if we can actually evaluate it.

Evaluation13:30

Nuno Campos13:51

And evaluating it, um, uh, correctly, uh, turned out to be a bit of a journey as well. We started with LLM as LLM as a judge, um, only. And, you know, that works to some extent, and sometimes it's the only option you really have.

Uh, but, uh, the annoying part is sometimes you can't really tell if when a score changes, is it because the, you know, the agent changed something or the evaluator, uh, changed, uh, what it outputs. Uh, so we ended up doing a bunch of work to replace it with deterministic, uh, comparisons wherever that was possible, which is not always possible, um, but where we could, uh, for instance, you know, take a golden spreadsheet that had a set of inputs and a set of outputs and then use that as kind of a black, black box to test a spreadsheet that the model produced, saying, "Hey, if you put some numbers into these inputs, you get something out of these outputs."

Debugging14:51

Nuno Campos14:51

And then you put the same numbers into the spreadsheet the model produced, and you see if you get the same, uh, outputs. Uh, and that, uh, ends up being, uh, you know, sometimes more trustworthy than just using an LLM to grade that work.

And, uh, as with everything, there's, uh, bugs and infrastructure bugs, uh, when you're building agents, uh, many times end up looking like, uh, you know, reasoning failures. And it may seem like, oh, the model is doing something wrong.

Um, and, uh, but actually many times it turns out, you know, it's a, it's a bug where, uh, you're, you just have a bug in the code or the, the skill or the prompt has the wrong example and the model is, uh, following that very faithfully.

Or, uh, uh, there's actually, you know, a bug in the tools and they fail. Um, and then the model keeps retrying and it seems like the model is being dumb, but, you know, it's just trying to work around the issue.

Uh, so there's, you know, a lot of juice to get out of just really looking at those traces and seeing what is going wrong and trying to figure out, is this, you know, the model not getting it quiteright or is this something we can actually fix?

Oh.

Lessons16:08

Nuno Campos16:10

Sorry. Um, so I wanted to kind of end with, um, uh, a summary of what I think generalizes to other tasks. And, uh, I think the first thing is if your agent is making many sequential tool calls or even parallel tool calls, uh, then you've kind of invented a bad scripting language.

So you might as well just give the agent a real one. And that can be code mode or repo or whatever you want. Uh, the second one is I think feedback loops really matter. And if you happen to be working in a domain where those, uh, where you can build those feedback loops with existing tools, then great, uh, less work for you.

Uh, but if you, if you're not, if you're in a domain where those, uh, feedback loops don't actually exist, uh, I think it's actually really worth spending the time to build that rendering engine or calculation engine or whatever applies to your particular domain.

Uh, the third is, um, I think interfaces are super important. And as, you know, as I explained before, the, the repo really changed the results we got. Um, so you should really spend the time figuring out what the best interface is, but you should expect to have to revisit that.

Uh, because the, the models, the capability of the models is going to keep changing. And as they get better at other things, you may find that the best interface, uh, is something else and you need to, uh, find what that next one is.

Uh, next, I think, you know, we shouldn't really underestimate the power of planning and think before, uh, think before you act. Uh, and yeah, uh, sometimes the simple things really do make a difference. So don't, you know, actually spend the time on those as well.

Uh, domain knowledge I think is, uh, really important. And you need, you really need to spend a bunch of time thinking about what's the, what's the things that you need to re-remind the model about. It's not so much teaching the model, it's more reminding it to pay more attention to that than other things.

And, uh, lastly, yeah, evaluation. I think the more you can do deterministic evaluation, the better, which doesn't mean that you should, you know, avoid LLM as a judge. It just means, uh, if that's the only option you get, then that's exactly what you should do.

But if you can, uh, evaluate in some other way, do that. And, uh, always check your traces and your plumbing because, uh, sometimes agent confusion is just bugs and you should fix that. Thank you.