AIAI EngineerJun 5, 2026· 16:44

Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, OpenClaw

Vincent Koc, a core maintainer at OpenClaw, explains how the open-source project ships code faster than humans can read diffs by running 60–70 autonomous coding agents across swim lanes. In a single night, he and Peter Steinberger executed 2,700 commits touching 82% of the core codebase, launching a plugin architecture and changing close to a million lines of code. Koc describes managing agents as managing people: the key skill is reading reasoning tokens to detect when an agent is bullshitting, gained through sheer volume of token maxing. He argues 2025 was about token maxing; 2026 is about not wasting them, with agent-in-the-loop processes and opinionated swim lanes replacing blind Ralph looping. The episode details his agent development environment using .skills files, Git work trees, and a semantic graph for triaging 60K PRs, emphasizing that the bottleneck shifts from engineering to taste and process.

  1. 0:00Dark Factories
  2. 0:55The Jank
  3. 2:18Factory Shift
  4. 3:39Agent Scale
  5. 5:18Parallel Agents
  6. 7:05Great Refactor
  7. 9:23Swim Lanes
  8. 10:50Work Trees
  9. 11:56Reasoning Tokens
  10. 13:09Skills Framework
  11. 14:25PR Management
  12. 15:19Soft Skills

Powered by PodHood

Transcript

Dark Factories0:00

Vincent Koc0:16

They've got it. Cool. Amazing. So, welcome everyone. I'm Vincent—uh, what do I do? I'm one of the core maintainers at OpenClaw, working with Peter. And as you've heard before, I have a day job as well, same as Peter.

He has a day job at OpenAI. But, you know, it's an open source project. Amazing things have been happening. I'm going to talk about what I call Dark Factories, and how OpenClaw ships faster than you can read the diff.

This meme is absolutely hilarious. So, I think Peter posted this a week or two ago: "I wake up, there's a new technological advancement, I wake up." It's this joke that we're shipping at insane speed, and the velocity is just absolutely phenomenal.

And some of you might think, oh, this is some luck, or we're just, like, Ralph looping to the max. I think there's actual engineering work here, and I'm going to talk about that. Now, as I mentioned, I'm Vincent, I'm your friend in Clancar.

The Jank0:55

Vincent Koc1:10

This is me using VR goggles back in 2013. So, despite my accent that sounds somewhat Australian, I was born and raised in East London, not far from here. I actually went to college just down the road in Westminster, and, yeah, at some point decided to live in Australia and my accent changed.

But I used to love technology. I used to love being at the edge of technology. And this was, like, one of the first few, sort of, early VR goggles that came out. It came in this big box with a big warning sign on it saying, "Hey, use for 5 minutes at a time," because it didn't have, like, the anti-motion sickness built into it.

And the funny thing with this one here was that I didn't use it for 5 minutes, I used it for 3 hours. And I played Team Fortress 2, had an absolute blast, and then I vomited for 3 hours after that, because my vision turned into B-vision.

What I'm trying to say here is that, like, anything on the edge is going to be janky, it's going to be horrific, it's going to be uncharted territory. And working on OpenClaw and being part of the team that ships probably, you know, an insane velocity of commits to a point where I get rate limited by GitHub on an hourly basis is an interesting experience.

Factory Shift2:18

Vincent Koc2:18

And this experience, Britain's gone through before. We had the Industrial Revolution, when mills and cotton were being produced at extreme amounts of volume. And there's a lot of history here around production and productionization at scale in the UK and in Europe.

And I feel like we're going through this moment again. We're going through this moment of, how do we build at scale? And the ways we used to work before just don't work anymore. And it's kind of strange, because in my day job I kind of work in the space of evals, which everything is sort of structured, and there's telemetry, and it has to be all perfect.

And I work on a project where I'm—I have this blind faith in the harness. And it's this kind of two worlds, but they're starting to come together. We used to have handlooms in cottages, centralized mills everywhere. Craftsmen were the factory workers, but the bottleneck was the weaver's hands.

We're now switching to a world where engineers writing code in editors, not so much. Swarms across repos. Engineers are becoming factory managers, which I'm going to talk to. And the bottleneck becomes taste. You know, that lovely word. Italian mother's hands, yes.

So, in context, like, what does this mean? Like, are you talking absolute nonsense of people building things at absolute scale? They are. What happened was very similar to the ChatGPT era, where everyone denied it at scale that they were using ChatGPT.

Agent Scale3:39

Vincent Koc3:55

Everyone was in this absolute fearmongering sort of world. But what the reality was, was that everyone was using it. Everyone in secret was just like, "Oh my god, what's going on? I need to talk to it." And the same thing is happening with these autonomous agents at scale.

Some organizations have openly come out with it. So, for example, Anthropic, with their recent work they did on building a new C compiler. We had Spotify saying they're no longer writing code by hand, supposedly. Steve Yeager, which I absolutely love, saying he pushes about 50 PRs a day, total solo.

He calls himself a vibe maintainer. I can kind of relate to that. And OpenClaw, where we're pushing. At the peak, we were doing 800 commits a day. And realistically, like, there's about 10 to 15 core maintainers, all with day jobs.

It's kind of astronomical in terms of scale.

And for me, this was March 15. What was that, like, 2 or 3 weeks ago? Where I hit close to 3,000 commits per day. And if you actually look, my commits actually stop when I go to sleep. So if you want to see when I go to sleep and when I wake up and how many hours of sleep I have, you can just take a look at my commit history.

Yeah, it's astronomical. But the thing is, this is going to become the norm everywhere else. Like, this is, this is like me telling you, you need to wake up, that, you know, this scale of velocity is going to be normal.

And trying to review PRs and go through all this nonsense may not work. But somewhere in the mix is engineering. There is a form of engineering that's going to happen. So we did commit maxing. You know? Let's just go out there and smash as many commits as we can.

Parallel Agents5:18

Vincent Koc5:33

And this reminds me of Ralph looping,right? This, like, this guy, where you're like, "Hey, I'm just going to, like, give you a task, I'm going to burn tokens for, like, 8 to 9 hours." And you're waiting. You know?

You're waiting, you're hoping something happens. Maybe something happens. I don't know. But what if we had a bit more of an opinionated approach to this? What if we call it Bart looping? I don't know. One of the other maintainers gave me this idea.

Maybe we'll coin it. Do we need more than just tokens? What does that reward mechanism look like? How do we get a bit more opinionated? Yes, let's run loops, but let's be a bit more smart about how we do this.

So,right about the time you saw those 3,000 commits, this was the day before. I was at Nvidia with Peter, and the gentleman you see on the left is one of the other Nvidia gentlemen. And they were like, "Hey, we're building NemoClaw."

I'm like, "What? What's going on?" And, "Let's help you build it." And I was in the room. I was like, "I can't work on a laptop for, like, hours on end. Can you bring me a screen?" They brought me a screen.

Peter didn't have a screen, so that's his laptop on the left. He asked for a screen, so they gave him an even bigger screen than mine, because, you know, why not? And we just got to work. So he's running about, maybe, 15 codec sessions, and he's got his Mac Studio at home, his VPN into.

I'm running another, like, 10 or 15. And collectively, between us, we're probably running, with sub-agents included, maybe up to 60, 70 agents. But on the foreground, maybe 15 swim lanes, if you want to call it that. And we're just going for it.

Great Refactor7:05

Vincent Koc7:05

Funny thing is, we're working on NemoClaw on one side, but one maintainer decided, "I'm going to move some stuff around. I'm going to move a couple of folders around." And that was moving the entire channels. So, like, all our conversations with, like, MS Teams and Slack ended up moving to another location in the codebase.

And we were like, "Oh my goodness, we're going to have to change stuff." And I found a really nice place to put my drink as well. The Nvidia people don't like this. So what ended up happening is what we called the great refactor.

Essentially, where we were like, "Hey, we have lots of people raising PRs, and what they actually want is to build features." The thing is, we don't want to give everyone every single feature that they want, in which case it becomes bloat.

You heard Peter say earlier on, the challenge becomes, who do I say no to? It's not about saying yes. In a world where tokens are cheap, I can just say yes to absolutely everyone and merge everything in. But that's going to turn this codebase into an absolute fire dump.

So the vision was actually, we need to cut this codebase down. We need to rip it into pieces. And a plugin architecture somewhat made sense. Imagine if you're OpenAI or Mistral or Anthropic, what if you owned that piece of the provider code and it was handed to you, and it was separate from everything else?

So this code change that occurred was like a catalyst for us. It was 2 in the morning, we're tired, we thought, why not refactor the entire codebase? Sounds like a splendid idea. So, 2,700 commits later, close to a million lines of code change, touching 82% of the core codebase, plugins were launched.

The night before, I think it was like 1 in the morning, I'm trying to go to sleep, and the tests are not passing. And I was like, was I ecchymous, and did I fly too close to the sun?

As we like to call it, did I vibe too hard? I actually genuinely thought I vibed too hard. But as a team, we managed. We managed to bring this codebase back together again. But the saving grace was these awful sort of unit tests that AI code loves to generate that actually ended up overfitting on our code.

So when we completely ripped everything out, we still had these tests that were, like, extremely overfitting. And as long as they would go green, we knew we were kind of somewhat close.

Swim Lanes9:23

Vincent Koc9:23

So how do we do this? You know? In my case, I call it my factory. It's many codec sessions. Everyone asks me, like, what's this magic sauce? Like, how do you do this? What's this crazy insane thing? Like, how are you guys building this?

Very simple. I have swim lanes. It could be 5, it could be 10, it could be 20. But traditionally, they kind of cut themselves up into different pieces. So, like, if I—does this work, the laser? You can't really see it.

But imagine you're a factory manager and you have a production line blowing. Essentially, you might have a case where you have, let's just say, CI to one side, you might have features in one side, you might have bugs in another.

So when I'm refactoring and doing stuff,right now the codebase is quite stable. I want to refactor some tests. Well, that might be swim lanes 1 and 2. I don't need to really babysit them too much. I just tell them, take your time, make sure the tests pass, just commit.

Just push them through. Whereas with 3 and 4, I might be looking at specific features and issues around, say, Docker or one of our messaging channels. In which case, I'm having a conversation with those agents, they're going off investigating, doing the work, coming back.

And then maybe 5 is actually looking at new P0s and P1s. That might be using other data. It might be using GitHub. We have agents that run inside of a Discord channel, so when we do a release, we might be like, hey, what's happened in the last 2 hours that I need to be paying attention to?

And this will scale up and down. But what ends up becoming quite interesting is tokens are no longer the problem. Depends who you ask. What really ends up becoming the problem is just raw compute and my brain space in order to sort of keep an eye on all of these sessions.

Work Trees10:50

Vincent Koc11:05

So in harness, we trust. What ends up happening is I don't have this really insanely complicated process. The one thing I have complicated in my life is adopting Git work trees. And I kind of wish I hadn't. The only reason why I say this is when you're running an extremely heavy test harness, it ended up completely nuking my machine, because I ended up running, like, every PR I touch ends up becoming a new Git work tree.

I end up with, like, something close, like, 70 or 80 active Git work trees in any given day on my machine. And that's kind of hell. So I had to actually build some, like, magic sauce around my codec sessions.

So my codecs are aware of Git work trees. If I hit the escape key or it crashes, it will self-heal, self-recover, get, you know, spar stuff. But realistically, I should have adopted what Peter and other people do, is just, like, clone the repo 10 times and point 10 different, you know, codec sessions to each one.

Reasoning Tokens11:56

Vincent Koc11:56

But the trick here is that, like, I haven't done any magical sauce. I don't use plan mode or spec mode. I have a conversation with the agent, and we work through it, and we find a way to make it work.

So realistically, it looks a little bit like this from the matrix. And people go, "Oh, Vincent, like, how do you know it's kind of working?" And this is going to sound somewhat a little bit lunatic. If anyone's watched The Matrix and seen the scene where Neo goes over, he's like, "How do you know?

How do you read the text?" And the guy's like, "Oh, you know, I've been doing this for a while, so I can see, like, woman in red dress or guy walking dog." And you start to have this, like, relationship where you can feel the reasoning tokens.

I know it sounds somewhat ludicrous, but there's times where I'm looking at the swim lane and I'm like, this sounds off. It doesn't sound off because of what it's doing. It sounds off because of how it's explaining itself to me.

It's waffling. It's not making sense. It doesn't seem to know what it's doing. And this feels a lot like how I would manage people. If I had someone working for me and they started downright bullshitting, I'd be like, "Wait a minute, what's going on?"

So in these cases, I might just nuke the session and go, "You know what, I'm not going to deal with this section of code. I'm going to leave that to another maintainer." Or I might come back to it 4 or 5 days later.

But that experience feels very much like intuitive. And building that intuition, I've been able to get to because of the sheer volume of token maxing I've had to go through in the previous year. So there is engineering work.

Skills Framework13:09

Vincent Koc13:23

I call this the agent development environment. Essentially, the process goes, I have skills, I call it .skills, similar to .files. Both of my .skills and .files are available on GitHub. It's all open source. Go for it. Some of my skills are private.

But there's skills in there for, like, writing technical documentation, for example, that I've co-created with other developer experience and other engineers in the market. You can use a skills gym, something like a GEPA, which I'm also a contributor to.

Or you could just say, "Codecs, I've been using this skill in my last 2 weeks. Go through the codec sessions, read the logs, make improvements to the skill." I would then take that skill and deploy that into my OpenClaw, or take that into my, you know, personal environment.

And I'll use something like vercel skills.sh as, like, a mechanism to loop this. I've added some other testing and other elements on top of this, but there's a process to how I manage and maintain my skills as an engineer.

The way we manage PRs has some level of engineering work to it. There's this kind of running joke that every maintainer that joins the project decides to try and tackle, like, "Oh my God, we have 6,000 PRs. How are we going to solve it?

PR Management14:25

Vincent Koc14:35

I'm going to cluster everything and, like, figure this out."

Guest14:38

60K.

Vincent Koc14:39

How many?

Guest14:40

60K.

Vincent Koc14:40

There you go. There you go. Oh, no, thank you very much. So this was my flavor of, like, trying to solve this. It's like a semantic graphing, vector embedding on the entire GitHub stuff. This is one PR. It has 106 edges.

What ends up happening is that everyone else has the same problem, so they decide to send their flavor of the PR issue. It becomes utter noise. So there is even a process around, like, how we even consume what we're going to work on.

We might not call it a roadmap, but we have a way of kind of deduplicating and seeing what's out there. This might be a signal for me to say, "Okay, if there's enough pressure coming on one issue, it must be big enough that all these other clankers decided it's a big problem.

Maybe I should go and address it." There is evals, surprisingly. After all this refactoring work, we decided to make a fake Slack of sorts with both synthetic models and real models so we can run evaluation loops to check that each of the providers and the channels work.

Soft Skills15:19

Vincent Koc15:37

And this question was asked of me recently. "How do you manage 10-plus agents?" And this is something that you're thinking. I asked them back, "How do you manage 10-plus staff?" And they had no answer for me. I'd worked in large organizations like airlines and other places like that managing large AI teams.

I had experience managing up to 30, 40 people plus. So for me, it was not, like, a new paradigm. But I think for engineers and people working with these coding agents at scale, it's the soft skills that matter.

It's how do you ask your agent what's going on, how do you know when they're not bullshitting you, and how do you run that factory. So it's no longer about the model or the agent. It's about the process.

2025 was about token maxing. 2026 is about not wasting them. It's about token efficiency. It's about agent in the loop. Thank you.