Intro0:00
Yeah, well, thank you guys so much for having me. It's exciting to be back. I was last here at AI Engineer one year ago, and it's kind of crazy. I've always been—I've been telling Swix that we need to have these conferences way more often.
If it's going to be about AI software engineering, it probably should be like every two months or something like that, with the pace of everything's done. But it's going to be fun to talk a little bit about, you know, what we've seen in the space and what we've learned over the last 12 or 18 months building Devin over this time.
And I want to start this off with
Moore's Law0:47
Moore's Law for AI agents. And so you can kind of think of the capability or the capacity of an AI by how much work it can do uninterrupted until you have to come in and step in and intervene or steer it or whatever it is,right?
And, you know, in GPT-3, for example, it's—if you were to go and ask GPT-3 to do something, you know, it could probably get through a few words or so, and then it'll say something where it's like, okay, you know, this is probably not theright thing to say.
And GPT-3.5 was better, and GPT-4 was better,right? And so people talk about these lengths of tasks. And what you see in general is that that doubling time is about every seven months, which already is pretty crazy, actually. But in code, it's actually even faster.
It's every 70 days, which is two or three months. And so, you know, if you look at various software engineering tasks that start from the simplest single functions or single lines, and you go all the way to, you know, we're doing tasks now that take hours of humans' time, and an AI agent is able to just do all of that,right?
And if you think about doubling every 70 days, I mean, basically, you know, every two to three months means you get four to six doublings every year, which means that the amount of work that an AI agent can do in code goes something between 16 and 64x in a year, every year, at least for the last couple of years that we've seen.
And it's kind of crazy to think about, but that sounds aboutright, actually, for what we've seen. You know, 18 months ago, I would say the only really—the only product experience that had PMF in code was just tab completion,right?
It was just like, here's what I have so far, predict the next line for me. That was kind of all you really could do in a way that really worked. And we've gone from that, obviously, to full AI engineer that goes and just does all these tasks for you,right, and implements a ton of these things.
And people ask all the time, what is the, you know, what is the future interface, or what is theright way to do this, or what are the most important capabilities to solve for? And I think, funnily enough, the answer to all these questions actually is it changes every two or three months.
Like, every time you get to the next tier, the bottleneck that you're running into, or the most important capability, or theright way you should be interfacing with it, like, all of these actually change at each point. And so I wanted to talk a bit about some of those tiers for us over the last year or so.
And, you know, over the course of that time, obviously, you know, when we got started in the end of 2023, obviously, agents were not even a concept. And now everyone has, you know, everyone's talking about coding agents. People are doing more and more and more.
Repetitive Migrations3:25
And it's very cool to see. And each of these has kind of been almost a discrete tier for us. And soright around a year ago, when we were doing the last AI engineer talk, actually, the biggest use case that we really saw that was getting broad adoption was what I'll kind of call these repetitive migrations.
And so I'm talking like JavaScript to TypeScript, or like upgrading your Angular version from this one to that one, or going from this Java version to that Java version, or something like that. And those kinds of tasks in particular, what you typically see is you
have some massive code base that you want to apply this whole migration for. You have to go file by file and do every single one. And usually the set of steps is pretty clear,right? If you go to the Angular website or something like that, it'll tell you, allright, here's what you have to do.
This, this, this, this, this. And you want to go and execute each of these steps. It's not so routine that there's, you know, there's no classical deterministic program that solves that, but there's kind of a clear set of steps, and if you can follow those steps very well, then you can do the task.
And, you know, this was the thing for us because that was all you could really trust agents to do at the time. You know, you could do harder things once in a while, and you could do some really cool stuff occasionally.
But as far as something that was consistent enough that you could do it over and over and over, these kinds of like repetitive migrations that you would be doing for, you know, 10,000 files were, you know, in many ways, the easiest thing, which was cool, actually, because it was also kind of the most annoying thing for humans to do.
And I think that's generally been the trend where AI has always done these more boilerplate tasks and the more tedious stuff, the more repetitive stuff, and we get to do the more fun, creative stuff. And obviously, as time has gone on, it's taken on more and more of that boilerplate.
But for a problem like this one, a lot of what you need to do is you need Devin to be able to go and execute a set of steps super reliably. And so a lot of this was, you know, I would say the big capabilities problems to solve was mostly instruction following.
And so we built this system called Playbooks, where basically you could just outline a very clear set of steps, have it follow each of those step by step, and then do exactly what's said. Now, if you think about it, obviously, a lot of software engineering does not fall under the category of literally just follow 10 steps step by step and do exactly what it said.
But migration does. And it allowed us to go and actually do these, and this was kind of, I would say, the first big use case of Devin that really came up. I think one of the other big systems that got built around that time, which we've since rebuilt many times, is knowledge or memory,right?
Which is, you know, if you're doing the same task over and over and over again, then often the human will have feedback on, hey, by the way, you have to remember to do X thing, or you have to, you know, you need to do Y thing every time when you see this,right?
And so basically an ability to just maintain and understand the learnings from that and use that to improve the agent in every future one. And those were kind of the big problems of the time, you know, and that was summer of last year.
And around end of summer or fall or so, you know, I think the kind of big thing that started coming up was as these systems got more and more capable, instead of just doing the most routine migrations, you could do, you know, these more still pretty isolated, but a bit broader of these general kind of bugs or features where you can actually just tell it what you want to do and have it do it,right?
Isolated Tasks6:28
And so, for example, hey, Devin, in this repo select dropdown, can you please just list the currently selected ones at the top? Like having the checkboxes throughout just doesn't really. And Devin will just go and do that,right? And so if you think about it, it's, you know, it's something like the kind of level of task that you would give an intern.
And there are a few particular things that you have to solve for with this. First of all, usually these changes are pretty isolated and pretty contained. And so it's one, maybe two files that you really have to look at and change to do a task like this.
But at least you do still need to be able to set up the repo and work with the repo,right? And so you want to be able to run lint, you want to be able to run CI, all of these other things, so, you know, to at least have the basic checks of whether things work.
One of the big things that we built around then was the ability to really set up your repository ahead of time and build a snapshot that you could start off, that you could reload, that you could roll back, and all of these kinds of primitives as well,right?
So having this clean remote VM that could run all these things, it could run your CI, it could run your linter, and so on. But that's when we started to really see, I would say, a bit more broad of value,right?
I mean, migrations is one particular thing, and for that particular thing, we were showing a ton of value. And then we started to see where, you know, with these bug fixes or things like that, you would be able to just generally get value from Devin as almost like a junior buddy of yours.
Multi-File Changes8:14
And then in the fall, things really moved towards just much broader bugs and requests. And here it's, you know, most changes, again, you know, jumping another order of magnitude, most changes don't just contain themselves to one file,right? Often you have to go and look, see what's going on, you have to diagnose things, you have to figure out what's happening, you have to work across files and make theright changes.
Often these changes are, you know, hundreds of lines. If it's like, hey, I've got this bug, let's figure out what's going on, let's solve it,right? And, you know, there are a lot of things here that really started to make sense and really started to be important.
But one in particular I'll just point out was there's a lot of stuff that you can do with not just looking at the code as text, but thinking of it as this whole hierarchy,right? So understanding call hierarchies, running a language server is a big deal.
You have git commit history, which you can look at, which informs how these different files relate to one another. You have, obviously, you have like your linter and things like that, but you're able to kind of reference things across files.
And so like one of the big problems here, I think, was kind of working with the context of it and getting to the point where it could make changes across several files. It could be consistent across those changes.
It would be able to understand across the code base. And here was really the point, I would say, where you started to be able to just tag it and have it do an issue and just have it build it for you.
And so Slack was, you know, a huge part of the workflow then. And it was just, it made sense because it's where you discuss your issues and it's where you set these things up,right? So you would tag Devin in Slack and say, hey, by the way, we've got this bug, please take a look.
Or, you know, could you please go build this thing? This is especially fun part for us because this isright around when we went GA. And a lot of that was because it got to the point where you truly could just get set up with Devin and ask it a lot of these broad tasks and just have it do it.
But a lot of these, you know, a lot of the work that we did was around having Devin have better and better understanding of the code base,right? And if you think about it, you know, from the human lens, it's the same way where on your first day on the job, for example, being super fresh in the code base, it's kind of tough to know exactly what you're supposed to do.
Like a lot of these details are things that you understand over time or that are a representation of the code base that you build over time,right? And Devin had to do the same thing and had to understand, how do I plan this task out before I solve it?
How do I understand all the files that need to be changed? How do I go from there and make that diff?
Iterative Planning10:41
And around the spring of this year, again, every gap is like two or three months. You know, we got to an interesting point, which is once you start to get to harder and harder tasks, you as the human don't necessarily know everything that you want done at the time that you're giving the task,right?
If you're saying, hey, you know, I'd like to go and improve the architecture of this, or, you know, this function is slow, like let's profile it and look into it and see what needs to be done. Or, hey, like, you know, we really should handle this error case better, but like let's look at all the possibilities and see what we should, you know, what theright logic should be in each of these,right?
And basically what it meant is that this whole idea of taking a two-line prompt or a three-line prompt or something and then just having that result in a Devin task was not sufficient. And you wanted to really be able to work with Devin and specify a lot more.
And around this time, along with this kind of like better code base intelligence, we had a few different things that came up. And so we released DeepWiki, for example. And the whole idea of DeepWiki was, you know, funnily enough, is Devin had its own internal representation of the code base, but it turns out that for humans, it was great to look at that too, to be able to understand what was going on or to be able to ask questions quickly about the code base.
Closely related to that was search, which was the ability to really just ask questions about a code base and understand
some piece of this. And a lot of the workflow that really started to come up was actually basically this more iterative workflow where the first thing that you would do is you would ask a few questions, you would understand, you would basically have a more L2 experience where you can go and explore the code base with your agent, figure out what has to be done in the task, and then set your agent off to go do that because for these more complex tasks, you kind of needed that,right?
And so, you know, that was, I would say, kind of like a big paradigm shift for us then is understanding, you know, this is what also came along with Devin 2.0, for example, and the inIDE experience where often, yeah, you want to be able to have points where you closely monitor Devin for 10% of the task, 20% of the task, and then have it do work on its own for the other 80, 90%.
Autonomous Backlog12:51
And then lastly, most recently, in June, which is now, it was kind of, yeah, really the ability to just truly just kill your backlog and hand it a ton of tasks and have it do all these at once.
And, you know, if you think about this task, in many ways, I would say it's almost like a culmination of many of these different things that had to be done in the past. You have to work with all of these systems, obviously.
You have to integrate into all these. Certainly, you want to be able to work with linear or with Jira or systems like that, but you have to be able to scope out a task to understand what's meant by what's going on.
You have to decide when to go to the human for more approval or for questions or things like that. You have to work across several different files. Often, you have to understand even what repo is theright repo to make the change in if your org has multiple repos or what part of the code base is theright part of the code base that needs to change.
And to really get to the point where you can go and do this more autonomously, first of all, you have to have like a really great sense of confidence,right? And so, you know, rather than just going off and doing things immediately, you have to be able to say, okay, I am quite sure that this is the task and I'm going to go execute it now versus I don't understand what's going on, human, please give me help, basically,right?
But the other piece of it is this is, I think, the era where testing and this asynchronous testing gets really, really important,right? Which is if you want something to just deliver entire PRs for you for tasks that you do, especially for these larger tasks, you want to know that it can test it itself.
And often, the agent actually needs this iterative loop to be able to go and do that,right? So it needs to be able to run all the code locally. It needs to know what to test. It needs to know what to look for.
And in many ways, it's just a much higher context problem to solve for,right? Is this testing itself. And that brings us to now. And obviously, it's a pretty fun time to see because now what we're thinking about is, hey, maybe if instead of doing just one task, it's, you know, how do we think about tackling an entire project,right?
Looking Ahead14:46
And after we do a project, you know, what goes after that? And maybe one point that I would just make here is we talk about all these two X's, you know, that happen every couple of months. And I think from a kind of cosmic perspective, all the two X's look the same,right?
But in practice, every two X actually is a different one,right? And so when we were just doing, you know, tab completion, single line completion, it really was just a text problem. It is just like taking the single file so far and just predict what the line is next,right?
Over the last year or year and a half, we've had to think about so much more. How do you work with the human in linear or Slack or Jira? How do you take in feedback or steering? How do you help the human plan out and do all these things,right?
And moreover, obviously, there's a ton of the tooling and the capabilities work that have to be done of how does Devin test on its own? How does Devin, you know, make a lot of these longer-term decisions on its own?
How does it debug its own outputs or run theright shell commands to figure out what the feedback is and go from there? And so it's super exciting now that there's a lot more, there's a lot more coding agents in the space.
It's very fun to see. And I think that, you know, we're going to see another 16 to 64 X over the next 12 months as well. And so, yeah, super, super excited. Awesome. Well, that's all. Thank you guys so much for having me.





