AIAI EngineerJun 28, 2025· 23:48

Containing Agent Chaos — Solomon Hykes, Dagger

Solomon Hykes, creator of Docker and founder of Dagger, argues that containing agent chaos requires engineering reproducible execution workflows built on containerization applied to each step of an agent's workflow. He introduces 'container use'—agents developing inside fully isolated, customizable environments rather than just sandboxing outputs—and demonstrates a prototype using MCP to integrate with Claude Code and Goose. The system provides background work, rails, seamless human stepping in, and optionality by leveraging Dagger, Git-based state management, and ephemeral containers snapshotted per action. Hykes shows how agents can run parallel experiments, merge snapshots, and discard failed environments without pollution. The episode concludes with him open-sourcing the project as github.com/dagger/containeruse.

  1. 0:00Intro
  2. 0:57Agent Chaos
  3. 3:40Scaling Agents
  4. 5:46Four Needs
  5. 7:32Environment Design
  6. 8:24Containers
  7. 9:35Container Use
  8. 10:21Demo
  9. 17:09Parallel Agents
  10. 21:20Open Source

Powered by PodHood

Transcript

Intro0:00

Solomon Hykes0:22

Hello, hello.

Hmm.

Okay, my slides are up. You can see them,right? That's me. Okay. Well, this is a very special moment for me because I just realized yesterday walking in, this is the exact same spot, the same stage actually, that I stepped on almost exactly day-for-day 10 years ago to kick off Docker Con 2015.

I thought it was pretty funny. I don't know if anyone was there for that. Maybe this audience is too young, maybe? I don't know.

Agent Chaos0:57

Solomon Hykes0:57

Okay. Well, uh, I'm here to talk about chaos, specifically the kind of chaos that emerges when you try to use, uh, coding agents. Um, and I want to talk about chaos from the perspective of our community at Dagger, which is platform engineers.

Um, I don't know if there's any platform engineers in the room.

Okay. Just you and me, man. Okay. Well, it, it, it is known, uh, sometimes, uh, as other things, but basically platform engineers have a really tough job because they don't get to build and ship cool software. They get to enable all of you to build and ship cool software in the most productive way possible,right?

Uh, it's a really tough job. It takes range, it takes experience, it takes a lot of patience, but we do it for the endless gratification, you know, just the gratitude we get from developers. Just kidding. No one ever says thank you, but it's okay.

Someone has to do it. Tough job. Speaking of enabling, anyone here use coding agents?

We are outnumbered. Okay. Well, I, I want to say to you, congratulations and welcome to platform engineering. Yeah. I mean, your job now is to enable robots to ship awesome software while you spend more and more of your time enabling them to do that productively,right?

Tough job. I, I, I, I applaud you for giving up really the most fun and rewarding part of the job, you know? Very selfless. Uh, yeah. So, of course, this is not, uh, completely a reality yet. I mean, we're, we don't have quite yet the team of agents just kind of, you know, humming along, doing the, doing the job while we sit back and, um, fix environments for them.

But you can kind of see it coming,right? I mean, some of you are definitely doing that, hacking that together. There's a lot of cool posts out there and scripts and tools. Um, so we know it's coming. The question is, how do we enable this to, um, happen not just for this incredibly cool and, uh, bleeding-edge crowd, but for everyone else, uh, like everyone shipping software any-everywhere, just sort of creating maximum value by enabling agents to do the work for them, ultimately taking their jobs.

That is the dream,right?

Scaling Agents3:40

Solomon Hykes3:40

Okay. So, yeah. How do we do that and make it not too painful? Well, um, I want to go back to basics. What is an agent? Uh, the famous definition, of course, is it's an LLM that's wrecking everything in a loop on behalf of a human.

The diagram is from Anthropic. Thank you, Anthropic. I tweaked the explanation just a little bit. Uh, in the context of coding agents, it looks like this. Um, oh man, that was supposed to be animated. It's even better when it's animated.

It's okay. Yeah. You got one agent and it's doing stuff in the environment is your computer. Uh, and it can do great work. It can all do also do very crazy things. So you have to kind of watch it closely,right?

And approve, approve. No, no, don't do that. That's crazy. Yes, that's good. Um, that's kind of the status quo today. But of course, um, we want

to scale it,right? We want a team. So how do we do that? Well,right now, I would say there are two options, both equally wonderful and fun. The first one I call YOLO mode.

You know, I'll just run 10. What can happen? Uh, amazingly, this diagram is not the worst-case scenario.

But yeah, you know, you get the idea. So the, the whole methodology of watching it closely just kind of falls apart really quickly because they're all stepping on each other's toes. They're sharing an environment,right? Okay. Enter option two.

Oh, don't worry about that. We'll run the agents,right? We'll take care of everything. We've got background mode. We've got the, we've got the model. We've got the tools. We've got the environment. We've got the compute. We've got the secrets.

We've got everything. You know, just open an issue, wait for the PR, relax until, of course, it doesn't work. And then you're like, no, that's not what I meant. Um, but these, these actually work really well. I think like 10 of those launched just yet, just today and yesterday.

Um, and, and they're great. It's just that, um, you know, sometimes you just want to get in there, like, okay, give me the keyboard, you know, and sometimes you just want to run it on your machine or on your favorite compute provider,right?

Four Needs5:46

Solomon Hykes5:46

Use your favorite model. You want to mix and match. So there are limitations to this all-in-one model. So the question is, is there something better? Uh, is there just a scenario where I just got a team and they're working and, you know, I can step in or leave them alone and we're just kind of getting stuff done together?

So this is how I would summarize it. What I would want is really four things. First, I want background work. You know, I don't want to be in there just watching every action. That's obvious. Um, I want rails.

And that means I want to be able to constrain the agent to, to not just do things that I already know are not necessary. So obvious things like context of the project, what's, you know, what's our coding style, what's our, our what tools to use, but also here's how to build, here's how to test, here's the base image we, we use,right?

You can access this secret. You can access that. Just an easy way to do that because otherwise I'm going to waste so many tokens just correcting as I go,right? The third is inevitably when I do need to step in, I really, I want a really efficient and seamless way to do that.

And it can't be watch every action and it can't be just wait for the PR and do code review. You know, there's a, I need a middle ground here. And the fourth thing is I want optionality because like I was saying before, it's a crazy market.

You know, there's, there's awesome models, awesome compute, awesome infrastructure. Uh, agents are really cool. And as cool as they are now, I mean, one of you is probably like launching oneright now and then there's another one tomorrow. So I don't really want to lock myself into a whole package today and say no in advance to whatever's coming out tomorrow, not in this market.

Environment Design7:32

Solomon Hykes7:32

So to get that, um, I need an environment that has properties that match this. It needs to be isolated,right? So background work works. It needs to be customizable so I can set up those rails. It needs to be multiplayer so I can, you know, go, allright, give me that.

Let me fix this. Or let me check. Did you do it? You know, when the model says, I did it. Did you do it? And then, you know, it should be open. No, no shade on making money and scaling a huge cloud service.

That's great. You know, we have one. They're great, but I just want choice,right? I want to be able to choose and get the, the best commodity. Let's just use this word. It's okay. It's okay to use it. The best commodity component for each, uh, job.

Containers8:24

Solomon Hykes8:24

And, you know, could even be open source. Who knows? We could collaborate on this. Anyway, so unsurprisingly, maybe I'm going to talk about containers now. Someone actually said, you know, you should check that they know Docker. They know containers.

Uh, okay. Who knows what containers are? Who's used containers? Okay. Cool. Cool. Allright. Boost my confidence a little bit. But the point here is we have the technology and that's, it's not just about containers, but they do play a crucial role because it's a foundational technology and it is, it is underutilized.

We don't fully leverage what this technology can do because we're used to the first incarnation of the tools made for humans. Uh, same thing for Git. I see a lot of hacks involving Git work trees. Anyone playing with Git work trees to, to get stuff done?

Okay. You know what I'm talking about. So this is about that. Um, and of course we have models that are incredibly smart, getting smarter, and they, they can exercise these technologies, uh, really fully. We just need to integrate them in a native way so that we really, um, tackle the problem at hand, which is giving great environments to these agents.

Anyway, so if we built that native integration, what would it look like? Well, we have a take. Sorry, we at Dagger. I forgot completely to mention my company. That's okay.

Container Use9:35

Solomon Hykes9:46

Um, it's great. Check it out. Um, we, we have a take on that. Something we call container use. You know, there's computer use, browser use. Uh, these agents need container use. Um, they need a way to use containers to create environments and work inside of them.

This is not the same thing as sandboxing,right? There are a lot of ways to execute the output of the agent in a secure sandbox. Very useful, very cool, but that's not the same thing as the agent developing inside of containers entirely,right?

That's what we're talking about here. So

I asked my team, hey, we've been developing this thing. Oh, it's open source, but it's not yet open source. Like it's not finished. But I asked the team, I should show it,right? And they said, absolutely not. It's not ready.

Demo10:21

Solomon Hykes10:38

So anyway, you want a demo?

Okay.

Allright. Just so we're clear, this is you agreeing to watch me stumble through a broken demo of unfinished software. Yes? Okay.

So much could go wrongright now. Okay. This is my terminal. Can you see it? Okay. For, for technical reasons, I'm not going to go to full screen. You just got to stop me when I reach the edge. Oh, I actually, I can see it.

Never mind. Okay. Yeah. Old school. Okay. We used to do this all the time

in the old days. Okay. So, uh, here's what I'm going to do. I'm going to just, um, try to develop something very simple here. I got an empty directory. I'm going to try, try and make a little homepage for my awesome container use project.

And I'm going to use Claude, Claude Code. I'm going to try and use a bunch of them. Hopefully, I made something very clear. This is not a coding agent. It's environments that are portable that you can attach to any coding agent.

That's the idea. So you like Claude, use Claude. You like, you know, Codex, use Codex, et cetera, et cetera, et cetera. In an IDE, in the command line, whatever. And also in the cloud,right? In CI, lots of cool things you can do once you're async.

So, okay. One of the reasons the team said don't do a demo is I'm, I'm actually terrible at using Claude. So, uh, I have an alias for remembering the flag to disable all, you know, permissions. I got, I can never remember it.

And I have a prompt here. It's, yeah, I'll, I'll read it to you in a minute, but it's basically make me a homepage, uh, make it a Go web app so I can know what, what's going on because I'm not a cool kid writing TypeScript and run the app when you're done.

So while this runs, while this maybe runs, hopefully. Okay. Okay. Cool. So what's happening here is I configured Claude Code to use, to, you know, with container use, to use containers literally. Um, yeah, MCP. So it was an MCP integration.

There were other integrations that we're working on, but MCP is the obvious place to start. Um, and so now it has, you know, all its usual tools. This is vanilla, uh, Claude Code, but now it can create an environment for itself and now it's editing files in that environment like in a little sandbox.

And it can also run commands to build it and test it and of course run it in, uh, ephemeral containers. This is not one Docker container sitting there. Every time an action needs to be taken, there's an ephemeral container running and then being snapshotted and, and, uh, returning.

So it's just doing its thing. Um,

what would I want to show here? Okay. So here I'm going to first show that nothing has been polluting my workspace. It's happening in a little sandbox. And the way the sandbox works, the state of these files and the containers that are being run is, um, actually persisted, uh, in Git and a bunch of special Git objects that are kind of living alongside the repo.

So it'sright there if I need it. This is all local. Um, but it's not polluting my workspace by default. So hopefully it's going to produce something soon. Uh, while it does that, I'm going to use this little command line.

Is this readable? Okay. Little command line. See you. Like go work. See you later. But no, really, it's for container use. Um, and I can list environments and you can see there's a new environment that's been created here, uh, with a little random name here.

And so there's a few things I can do. One thing I can do is open a terminal.

And here, okay, this part is powered by Dagger,right? But we use Dagger as a sort of a toolbox. Just it has all the primitives you need. Um, and so here I can see exactly what the agent sees. Um, the files, but also the tools.

So I can see, okay, what, what Go version did you configure for yourself? Allright. Because the model, the, the agent is given the ability to figure out what environment it needs and then configure that, but in a repeatable containerized way.

Uh, so here I can see, okay, does it build?

Okay. It builds. Okay. So you're done. What's going on? Okay. While we do that, I'm also going to show you actually two more things to say. One, uh, a really cool feature of this that I'm not going to show is secrets.

So you can just plug in secrets from things like 1Password. I use 1Password. I don't want to use a separate password manager from an AI company. No offense. I just want to use my password manager. So I can just plug in and say, this environment gets this secret and boom, it can use it,right?

Um, and then the team said, please don't show that. That's just, that's going to break for sure. Um, so I won't. And the other thing I want to say is that because it's all powered by Dagger, um, and the point here it's containers and it's open source.

That's what you should know. Uh, it's running on my machine. Actually, no, it's not running on my machine because we're at a conference and there's a lot of things that can go wrong if you run containers and download images.

So instead, I, I just have it running on my home server in my basement about one mile this way. And it just kind of works seamlessly. It's streaming files up, streaming files down. It all just kind of works.

Um,

okay. This is the part that I cannot control, as you know. Um, okay. One more thing I'll show you. You can watch. So here I can see the history. So behind the scenes, every snapshot of the state is like a Git log.

It's actually using Git under the hood. So if I'm happy with a result, I can go and get it. Uh, so it's like a happy medium between the, it's like a loop, a collaboration loop that's justright. It's not watching every tool and wrecking a shared environment, but it's not waiting for a pull request and, you know, having these long back and forth.

It'sright in the middle. I can see everything going on and I can say, okay, give me the history of that. I want that. Okay. It says it's live. It's running. Ooh, pretty nice. Cool. Okay. So now,

okay, I appreciate it, but you guys are going to be honest, it's a little boring. So this design is boring.

Parallel Agents17:09

Solomon Hykes17:09

Make it really pop. Trying to impress AI Engineering World Fair audience.

Okay. Okay. So the reason I'm doing that is trying to create the circumstances where I would need a lot of parallel experiments,right? Make it pop. What does that mean? It can mean anything. What if I want to try several experiments in parallel,right?

So I'm just going to say, oh, well, hold on one second. Stop. Before I do that, I'm going to, um,

merge this,right? There's still nothing here, but I'm saying I like it. So I'm going to say merge that environment

and I have it. It's my history. I can open a pull request. I can clean it up, whatever. So that's, that's a loop that I can work with,right? Um, and now I can say, nah, boring. And then I can say, since the environment is now in this state, I can ask for help from a few other agents,right?

I can say, okay, hey, Claude YOLO. Uh, nope, that's notright. Claude YOLO, this web app looks a bit boring. Can you make it pop, please? Okay. And

go and go and go.

Okay. So this is where things start really going wrong. But

as the team pointed out, they said, they said, well, something's going to go wrong,right? They said, yeah, but you were kind of showing that if things go wrong, you can throw away the environment and you're good. You can restart.

I said, okay, that's cool. So, um, like, let's say I don't like this one. I'm like, nope, goodbye. That's it. I don't have to go clean up the mess,right? That's the whole point. Uh, okay. So this is getting a little messy.

Oh, I wanted to show Goose also. So Goose is a really cool open source agent. Whoops. Allright. Hold on one second. Goose YOLO. Same thing. Everyone has complicated flags for disabling all these safeties that I don't need anymore,right?

Because it's,

uh,

okay. Sorry.

Okay. Well, really taking a chance here. So while this is happening, uh, one thing we've been working on, but it's still work in progress, is there's a watch command. I always showed you that already, but as, so as, um, this is a Git command,right?

Thinly wrapped Git commands. Our UX is really, I cannot, words cannot express how unfinished this is, but, but it's, it'll evolve rapidly because the, the bones are strong. It's Git, it's Dagger, and you know, it's your existing agent,right?

So it's, and then a little bit of glue. Uh, so for example, here it is literally, it's a Git command that you can copy-paste. Uh, but as the agents work, you're going to see state snapshotting and you're going to see these branches just kind of, um, diverging.

And then I can diff them and apply them, merge them, whatever I want. Um, and what I really wanted to show, and then I'm done, is just, I just want to see one of them run. So you can see when the agent runs a service, like, and go, in this case, go run NPM, run whatever, it's doing it in its containerized environment.

And that's going to seamlessly be tunneled to my machine here on a different port without any conflicts,right? So if, when, when I say the environment's isolated, it's the, it's the files, it's context, it's configuration, and it's execution,right? Uh, and the cool, the cool extra thing is all of this is actually, technically, this here is running in my basement.

So you can go crazy on the infrastructure side. Like you can run this on a cluster. We like to run this stuff from CI. Uh, it's just a lot of fun stuff you can do. And okay, I'm getting at 30 seconds.

Come on. Oh, Goose is, oh, Goose is running. Great. Okay.

Open Source21:20

Solomon Hykes21:26

Allright. We did not solve prompt engineering. Do it.

Okay. Not done. Not done. Oh man. Okay. Well, just imagine.

Okay. Well, uh, while this happens, because I got 30 seconds left, I'm just going to say, um, thank you. And there's one last thing I, I want to say about DockerCon. Ten years ago, we used to open source stuff on stage all the time.

So if you want, I can go and open source itright now.

Okay. You have been warmed though about the not finished part,right? Okay. Okay. Oh, I think my, it would be funny if the demo failed at the clicking on GitHub part. Okay. Allright. Goodbye. Goodbye. Next time. I promise it works.

Okay. Haven't done this in a while.

Wait. Ooh.

I'm almost done. I promise. Come on. You did so well.

Change visibility. Yes. I want, yes. I have read and understand. Oh God.

Oh God.

Oh.

Yes.

At Dagger, we take security very seriously. Okay.

Allright. I think it's, wait, I think it's done.

Yes. Okay.

So yeah, thank you very much. And it's, uh, github.com/dagger/containeruse. Come say hi, come participate. And thank you so much for having me.