Mass Psychosis0:00
Right, hi there. Um, very nice to see you guys. It's like, it's so interesting to see, like, it's been such a long day and you guys still showed up. And, uh, I'm always very flattered when, like, I'm trying to present something and, like, people are there.
It just, uh, it makes you feel like what you do matters. So today the topic that I want to talk about is, like, "Don't build slop: 4 levels of AI agent maturity." And what I'm trying to do here is that I want to, I'm going to talk about something that, like, a lot of people brought this up to me, and there's like a mass psychosis problem.
So basically it's something like, it's like every person feels like, um, there's like this, like, giant set of robots around all of you that are just, like, kind of like breezing through, doing a lot of things, and, and you're like in the middle and you're so confused.
You're so confused of like, what is it, what should I do? Should I have like 15 agents ripping through all the time and I'm just like vibing? Or, or should I just be like a peasant and just like read through every line of code?
And it's, it's very hard to, like, make up your mind. I feel like you come to a place like this and you're always like, the formula just like gets to you, you get like a panic attack or something.
But like, I think, I think I want to, I kind of want to pull back out of this. And I want to be like, guys, let's slow down. Let's just slow down and let's think through like what are the problems that we can like actually solve.
And I will like actually help you build like really useful agents. And I want to take care of like every end of the spectrum of like your necessity to build things really fast, but at the same time, like, your necessity to, at a certain point, build things slow and take things to production.
So that, that's basically, um, the goal. To give an example of like mass psychosis, like, there's a lot of things that are very similar. So I'll give you an example of like three UIs. And these are three frontier labs.
And I guarantee you not one of you can predict which one is which. Uh, one of them is Factory, one of them is Codex, and one of them is Cursor. I, I am positive none of you know which one is which.
Even I don't know which one is which. But yeah, so that's basically the point. Like everything is same. You, you want to do your own thing. So to build agents in particular, uh, here's what, here's, here's what I would, would present here.
Four Levels2:13
That I want to break down this problem of building agents into like four different parts. The first part is just like trying to somehow figure out a way to like see if this even makes sense, like this even works.
That's like where you use a framework. The second part is like when you actually are serious, like, okay, I actually want to do something here. Now you build things by yourself, like as like a state machine, like building an actual agent.
The third part is like the UX workflow where you use Kanban, which I suggest is like a great, uh, form factor to be able to work with agents. And the fourth part would be like shipping to cloud. So this is like mostly an outline read of like how we want to be able to, uh, work with agents.
But I think it will give you a good heuristics of like how you, how you would, uh, solve this problem. So level one of building agents is like literally just like, just, just use a framework. And I, I think there's like a couple different frameworks, uh, like LangChain, LangGraph.
Frameworks2:56
Like I, I wouldn't, I don't use them. I, I, I work at like Cline. I, I, I'm supposed to like write all the agent experience stuff by myself. So why would I use a framework? So I won't be the best person to give you advice.
But I do think that if you're trying to find PMF, if you have a problem that's like, hey, I have this problem where I want to be able to, I don't know, aggregate emails or do something like rudimentary.
And I think that like an AI agent would probably be helpful here. I probably don't care about the best model. I just, I just want something that works. Any of the agent frameworks can give you something that just works in like half an hour.
You can just vibe code this. It, it, it's, it's a great thing to get started. It's a great thing to see that agents actually do work and that you can build them yourself. There are a lot of pitfalls with, uh, using frameworks.
And one of the biggest ones that I think is that like if you really want to take things to production, if you really want to build something serious, very quickly you learn that like the level of customizability, the level of like futuristicness, the level of, um, modularity that you need, you just, you won't find them in a framework.
DIY Agents4:09
I know a lot of people who disagree with me and a lot of those people are wrong. Anyway, um, so level two is like building, building your agents yourself,right? So here's how you, here's how you build. Like, I feel like, I feel like this is like a very intricate problem.
Like I can't sum this down. So I'll, I'll just like give you like five rules of like when you actually write code to build your agents. Like there, there are five rules that you could use. Um, and these rules will like give you a rough outline of like how to, how to build and write, write code for your own agents.
State Machine4:36
So the first one is like you always want to think of every agent as like a state machine. Um, state machine, it's like a sophomore year, freshman year topic. But like state machine is basically like every agent is, at the end of the day, it's like a recursive loop.
That's, that's basically like every, every, every hype cycle, whatever you think, it's at the end, it's all a recursive while loop. It's a while loop with a few conditions. And no matter what your agent is doing, it, it always, it doesn't matter if it's cursor, clock, or whatever, it is a while loop with a few conditions and a few end states.
And what you want to be able to do is that at any point of time, you want to be able to have like a mental model of which point in the state it is. So let's say you want to be able to like say, um, okay, so let's say you want to be able to, uh, I guess, read a few files and explain them, uh, through clock code.
So it starts from the user task at the top where you ask it like read a few files. It goes to the state of like reading a few files and like the action tool where it reads the file.
Then it realizes, oh, I have read the file. It makes sense. And then it will call the completion tool and just like complete. And the, the red thing at the bottom, task complete, is when the state machine finishes.
Uh, you can take this in a very complex way where you can like rip through the whole state machine for like 8 to 10 minutes or even like hours if you wanted to. Uh, but essentially every agent is a state machine.
If you can visualize that as a mental model, then every time you're building an agent, it will be so much easier for you to think through that. Um, the second rule is that every single thing you add to an agent risks making it worse.
Keep It Simple5:55
I think this is the hardest thing that we've learned. And this has been a very bitter lesson that a lot of agent builders have learned, which is that large system prompts, lots of different edge cases, lots of different like fancy if-else logic, all of that for frontier models just, just makes them worse.
Just like, just, just get out of the way of the model is the, is the lesson that we learned, which is that like frontier models are so good at their job that the less instructions you give them, they actually perform better.
A classic example is like if you go through the Codex repo, the prompt for GPT-5 versus the prompt for GPT-5.3 is one-third the size. Uh, part of the reason for that is that the newer models are so good at their job that giving them too many instructions and longer system prompts, it leads to a sensory overload where they get so many instructions that they get overwhelmed and can't figure out what theright thing to do.
Um, so I think the simpler it is, it's always better. And it's like you have to think that like every single thing I'm adding, I, I would be very careful that I'm, I, I hope to God I'm not making it worse.
And just always try to prune it down. We, we took it so far where we literally rewrote the entirety of Cline because we realized there was so much, uh, junk from the older, older versions of Cline. Claude Code has been written, I think, at, at least seven times from scratch, I think.
Um, but again, um, the, the, the people of the team might know better. Um, the third rule is that you want to be able to make agent an easy part of a pseudo-RL pipeline. And this is a tricky one, but basically what this means is that anytime you're building an agent, you want to be able to have some sort of like a CLI kind of a thing.
CLI Pipeline7:20
The reason for that is that as long as you have something that can build and test the agent really well in the form of a CLI, um, you want to be able to build things that are very easy to build and test with, um, other coding agents.
Soright now there's this thing, there's this like interactive dance that's happening between AI and humans where at, in the back in the day, humans used to like guide AI like do this. And I think at this point we're at a point where humans are being guided by AI.
And this is, I think, is a part of that where like if you're as a human, you want to be able to build things in such a way that like the AI can work very easily through it. So that might entail writing, writing, you know, agents theright way, using theright skills and building like a CL, CI, or, or CD, um, such that like the agent, other coding agents can easily build your agent, test it, make changes to it, and then test it end to end is a very critical part because that way if you want to make changes, you can just let, let a long running agent run through in a parallel thread.
It'll make all those changes tested and you have the whole thing running. But if it's harder to build and test, it, it would also be harder for you to like use agents to work on your agent. Um, super meta.
Um, rule number four, don't build slop. Guys, don't, please, for the love of God, don't, don't, don't build slop. Please. I think that, I think that there's, there's like, there's so much like throughput that you can get and so much, uh, tokens that can go through really, really fast.
Don't Build Slop8:43
I think that the best lessons that we've learned as like real engineers is that like it's really worth spending some time just like thinking through the architecture, thinking through the design and the outline of what your agent's supposed to do, making sure it actually makes sense.
And, uh, and just like actually at least spend some time reading the code, even if you don't write everything by hand. I think that that is like for us, we found that that to be super critical because I think that it, it's at least the architecture point of building an agent has to be done by a human and has to be done like very thoughtfully.
Even if you're using like an agent to have a conversation with it, spend a lot of time like trying to think through like what the architecture, what the state machine is going to be. Don't, don't just let like, uh, other models like rip through the code.
Um, and then the rule fifth is frontier labs kind of want to lock you down. And this is a tricky one. A lot of people will disagree, but basically what's happening is that a lot of times when you are working with the APIs of frontier labs, the APIs are trying to lock you down and they make the interchangeability harder.
Frontier APIs9:43
So to give a very precise example, the new set of models that have come out, say Opus 4.6, uh, Gemini 3.1 Pro, 5.3 Codex, they have this thing called reasoning traces. And reasoning traces are, um, a part of the cache and they're also the part of the reasoning, uh, uh, test time compute loop that the model does.
And when you have conversations and back and forth conversations with those models, you want to send the reasoning traces in the exact precise format that is expected. If you don't, the response would still work except that the performance would be degraded and you would have no way of knowing.
And a lot of people are kind of missing on the massive performance gains that are coming from the new models because they're just like not using the APIs in the exact precise way that they're supposed to be used.
Um, and there are asymmetries in the API. So some would argue that, hey, maybe you could use Open Router, but I don't think that's enough. I think you really want to be careful that like if you're using different frontier labs and different, uh, frontier lab APIs, like they, they have you very carefully thought through, uh, if the API is working correctly and you have actually tested it.
Um, so the form factor is like the next step is like the form factor of like how do you like, how do you visualize the agents,right? So I think originally I came back to like in the one of the previous slides, I tried to show you guys like the thing where like Codex and, uh, Cursor and others, they were all looking the same.
Kanban Boards11:03
And I think I have a different claim. So for this, uh, uh, on March 26, I made a tweet, uh, where I said we people should use Kanban. So Kanban boards, I think if anyone has used Linear, whatever, I'm sure all of you are familiar with Kanban boards.
So my argument is that Kanban boards are the thing, uh, to use. Um, Hansen in the audience was gracious enough to offer me his thoughts as well. So thank you so much. So Kanban boards are, are, are this idea that like if you're working through an agent, you're always, uh, inference bound.
A lot of you are working through like Codex or Opus and it's working for 8 to 10 minutes at a time. When one of the agents is working for 8 to 10 minutes, what do you do? You could doom scroll, but you can only doom scroll, scroll long.
So then you run another agent,right? So that way you have at least like two or three agents running in parallel at all times because you're inference bound. And they're all like mutating the same source potentially. So you want to isolate the thing that they're mutating, uh, to take care of that, like the, the, the isolation of state and then the inference bounding.
The best UX form factor to me is Kanban board, mainly because it takes care of, it gives you like, it makes you like an engineering manager which can look at all your agents and you gives you like a headline level view of what they are.
And it also helps you build like flows with them that like, okay, if these two tasks finish first, then I'll do the other task. And that helps you like become like you, you basically become like an engineering manager and all your agents are your ICs and, and you can, uh, walk, uh, look at them through this.
Um, so I was making this claim on March 26 and 10 hours ago, Claude Code came out with the same thing. So I believe I wasright. Um, um, okay. So yeah, you can use that through Claude Code, Cline, wherever you can use Cline as well, whatever, uh, whatever works for you.
Um, so yeah, so you want to think of like Kanban as basically like an engineering manager, which you would be. Then there's like a final step of like, okay, you have, you have like an agent, you've tested it, you've made it, and you, you have a good UX form factor to be able to like interface with it and look at it.
Cloud Agents13:12
Uh, what do you, what do you do then? Like how do you, how do you, how do you, how do you have an agent that's just like useful, it works well, but it, it scales, it really scales to like millions of tasks, millions of users.
If you're working for a company which has 8,000 people, how do you like, how do you make sure that like all of them can like very easily interface with this? And I think that rather than, rather than like making people install and like, you know, having these complex workflows in the machines and stuff, it's just, it's just so much jank and it's so hard.
Just like take it all, take it all out, put, put the hard work once, just take it all on the cloud. Um, so there are many benefits of cloud agents, but the primary one is that like you can completely paralyze and have like a separate machine for each of them.
There are like no local dependencies like, um, in the cloud, like the agent can set up the environment, do all the UX tasks. Becauseright now, one of the most missing pieces, the thing that a lot of people are not using that I use fairly extensively is cloud agents because they can really run, uh, really long.
So I, I on my phone often would send tasks that would run on a cloud agent for like 15 to 20 minutes. So let's say it's like a UX change of like, um, go build this, uh, VS Code extension.
And this VS Code extension, I want to sign in here. I want to click on settings. I want to change the settings to, uh, uh, pick this theme. And then I want you to test this thing in the terminal.
And the cloud agents are so good that they will actually manually do all the clicks of the Q&A testing that I just described on their own, figure out if they worked. And if they didn't work, it will just keep iterating, keep trying.
And this thing could easily take like 50 to 60 minutes. But if you send like a lot of these tasks in parallel through your phone or through your laptop or whatever, whatever, um, I think you, you have like, you have this like really easy customizable, extensible thing and it helps you scale really fast.
And then you could send all these tasks like running in like the, like the cloud machine, like, you know, like a messiah. And you could come back to your laptop and then it's like, oh, you just pull down the PR and you've got the whole thing going.
Um, the other aspect of like cloud, cloud agents is that like if you are working with a lot of different people, um, I think that it's just like it helps you build like a common setup that like so many other people can share and so many people can mutate.
Um, so I think bringing that together, I think would be like the final form factor. So my, my claim is that like there exists a future where, um, most of the UX of working with agents would be Kanban and then most of the actual compute, uh, that's involved with, uh, with agents would be on the, on the cloud.
Um, I think these four levels of frameworks are just like some the last part of like, you know, Kanban and shipping to cloud, like those things are very difficult and very intricate problems. And I think that if I were you, I would just like use this like as like rough heuristics and start with like, okay, let me just like, let me just do the bare minimum easy thing and then just like depending on how much effort I want to put in, I will like slide up and down, um, up and down these levels.
So yeah, so that's it for me. And yeah, so, um, I, I made a lot of hot takes. I feel like I think I left a lot of like open-ended questions here. If you have any questions or any thoughts, this might her, you're very welcome to reach out to me, send me any questions.
And, uh, yeah, um, it was very, very kind of you guys to give me your time. Uh, thank you so much.
May I take a photo of you? Allright.
Okay. Allright. Thank you. Uh, do you guys have any questions? It's like I've got a minute.
Allright. Oh, yeah. How do you think planning inside of a Kanban? Uh, like, I, I plan the sort of like back and forth requirements gathering, you know, the, the, uh, getting the agent to figure out what it wants for me to be the most useful part of.
Q&A17:15
Oh, yeah, yeah, yeah, yeah. Oh, yeah. Let me, let me show you. Let me show you. So basically like I think my interpretation of, of, of Kanban is that like it's just like if you see the screen, like you can go into any task and it will give you like the entire trace of the task.
Uh, and that point it's like the, so this, this interface that you're looking at is basically like, um, the actual CLI of Codex. And I think I would, I, what I, what I usually do is just like have a conversation here and then I know at a certain point that like, bro, you can go out on your own, do your thing.
At that point, I'll just pull out and just focus on other things. And then does it transition state when it either needs your review or asks you? Yes. Yes. Yes. So, so initially the taste is like, it's like let's say I would say like if I say read a few files or whatever,right?
Um, so it will be like initially it's in, in the in progress state and then when it's, uh, when it's like it needs my input, it will go to review. Um, yeah.
Allright. Thank you.





