Intro0:00
Um, okay, so we're starting five minutes early. Hey everyone, I'm Eric. I'm an engineer at Cursor, and I mostly work at, at the developer experience and product. And today I kind of wanted to talk to you about my experiences, like working at Cursor, dogfooding the product, and like getting to a place where you can build your own, like, software factory, and like what that kind of, like, takes, and the practical steps getting there.
To be honest, I don't think we're really there yet. Like, subparts of the product and subparts of the company are running, like, fairly autonomously. Um, but building a software factory takes a lot of work. I mean, like, look at, like, real-life, uh, factories producing, like, hardware.
There's a lot of assembly lines, there's a lot of people that goes into this, a lot of managing, observability, and all that. And there's a lot of concepts we can borrow from that world and put into the, uh, software world.
So anyway, here goes my observations from doing this. Um, but first, um, the agenda. I want to talk about, like, levels of autonomy. Uh, precursor to factory, pun intended. Um, building the factory, running the factory, and then scaling the factory.
And I want to finish with some Q&A for any kind of questions.
Autonomy1:26
Okay, so for the levels of autonomy, Dan Shapiro put out this blog post, I think in January or, uh, February, uh, explaining, like, six different stages of, uh, autonomy, uh, throughout, like, so automating software. Uh, Carpathia is also, like, previously used Cursor as an example of, like, going from tab to agent and all that.
But I think this kind of, like, encapsulates it really, really well. So we have the spicy autocomplete, uh, at the start. And this is kind of like where Cursor started in 22, 23, like, ages ago at this point.
Um, and we g kind of, like, gradually moved up the ladder in making the software creation more autonomous and letting the agents do more work. And I think most people, uh, adopting the AI tools, um, are, like, at somewhere between level two and level three, where you have a pair programmer who are essentially just going back and forth with the agent, asking questions, um, getting suggestions, asking the agent to do work, um, and eventually, like, finishing their tasks.
And the step above that would be, uh, having the AI generate the majority of the code, um, which we can see, like, here in the developer level three, um, where you, as a human, more kind of, like, reviews it, um, kind of, like, in the loop, following traces and all that.
But as you further progress, you're, like, becoming more and more of a manager. And we'll talk about, uh, this more later. But eventually, like, level four, I think this is where I'm mo at, at this point, like, for most, like, software projects, where I'm, like, delegating as much work as possible to agents and probably, like, reviewing the outputs before I actually review the code.
Um, because I still look at the code sometimes. Um, and lastly, we have the software factory, which is essentially, like, a black box. Um, Dan Shapiro calls it, like, the dark factory, where you don't really have an insight.
It's just, like, agents going around doing their thing, uh, shipping the code, testing the code, building the code, all that. And you, as a manager, just provides, like, the intent and the instructions, um, and, like, the goal, uh, from what you want out of the factory.
Okay. Um, yeah, so, like, why do you even want to create a factory? Um, first of all, like, throughput. You probably want to create more code with, like, less resources. Um, you can run agents 24/7. You don't have to, uh, rely on humans that need sleep and food and eat and all that.
Um, you can just, like, have more agents. Um, another, like, thing with the factory is, like, you have assembly lines. And assembly lines produces, um, consistent outputs. So if you build your factoryright, you can probably have very consistent output.
Um, but at some point, you initially, you feel like if you don't have a red setup, you might feel like the agents are getting more and more probabilistic and, like, you're losing a lot of determinism, um, because they just go off and do random things, um, which is probably a sign that you need to, like, build more guardrails for the factory.
And I think this is a function of the model capa capabilities as well. Like, as the model gets better, they can follow instructions better and just execute on whatever you want them to do. Um, and thirdly, um, you might want to have a factory because you can leverage your taste better.
Um, you can, like, get more out of your creativity out, um, instead of just, like, waiting for you, as a human, to create them and produce this, uh, software that you're creating. Um, and then obligatory, like, then and now.
This is what it used to look like. This is, like, a Tesla factory from a couple of years ago. Uh, and this is, like, kind of what we're getting after here. Okay, let's get straight into it. So to build a factory, what do you actually need?
Building Blocks5:04
Um, I like to think of this as primitives and patterns. So just, like, how do you structure the code? Um, is it, like, a modularized code base? Um, do you have this scattered all over the place? Is it co-located code, et cetera?
Um, because did, like, um, did distance in, um, in locating, like, if you have an agent, like, LSA-ing a folder, uh, it can, like, discover all the relevant files at once instead of having to grep and search all over the code base.
It can just, like, be very isolated to work within one, um, single part of the code base. And this goes the same with humans. Like, if you have an easy time, like, onboarding yourself to a new code base, an agent probably will have that too.
Um, the second thing is, like, usage patterns. Do you have specific, like, methods and services for authenticating a user? Do you have, like, startup scripts? Do you have a way to, like, write tests, et cetera? Do you have this boilerplate in place?
Um, because if you do, you can point the agent to, like, existing references and just asking this to reproduce, uh, over time. So those are, like, some of the, like, primitives and structures of the code base. Um, the second one would be guardrails.
So, like, you might you want to let the agents free, but not too free. Uh, so you want to have some rules and, and checks and, and hooks in place. Um, for example, um, a hook you might want to have is, uh, touching a specific part of the code base.
Uh, maybe the agent should not be able to change to, like, the most sensitive, like, encryption of sensitive data or authentication or anything like that, where, uh, a mistake could be, like, very, very, very costly, um, for the company or for you, as a human, et cetera.
Um, rules. Um, rules is probably the most misunderstood, um, concept since we launched Cursor rules. Um, there's, uh, Cursor directory, which launched a good collection of different rules. Um, and the assumption was usually that you should just install every rule that you can, depending on, like, what, um, uh, software stack you're using.
For example, if you're using Next.js, maybe you should have Next.js rules. Um, but what I found and what I'm seeing amongst our users and internally is that rules should just, like, emerge dynamically. Like, if you're finding agents going off the rails, you should probably create a rule for that.
And it should kind of, like, be sort of, like, an S SOP, uh, to showing, like, the agents what they can do and cannot do. And again, the models are getting so good at following specific rules that they usually don't go off the rails anymore.
And I think that just kind of, like, extrapolate over time as well. Um, and of course, tests. Like, can the agent, uh, verify its own work and can it run tests to know, like, oh, I messed something up, or, um, I made a change, depending in, like, in this specific area of the code, but it still passes.
I can I can still run the code and, uh, the check looks good. Um, and lastly, which I think is probably most exciting, is the enablers. Like, what can you allow the agents to do to actually let them be free?
Um, skills is good for this. Um, just giving the agents more capabilities, skills and MCPs. Um, accessing, like, external context. Um, getting, like, understanding of how to implement a certain thing. Uh, I'm going to show you some later, uh, in the Cursor code base, uh, what we are doing.
For example, like, feature flagging. Can we give, uh, the agents a skill to add a feature flag so when we launch them autonomously, they can just flag the actual changes made and merge to PR and come back to us like, "Hey, uh, if you want to try this, just turn on this flag.
Um, if you don't like it, we'll just revert to PR. If you like it, we can, like, expand it to more users." Um, and lastly, like, what kind of environment are you letting the agents run in? Um, can your agents, um, start your dev environment?
Can you just ask them to, like, "Hey, um, start my project," um, and let them do that without having to, like, have any human in the loop? Um, because if that's the case, you can probably, like, have them run, um, you can scale it up, like, infinitely on separate VMs.
Um, and then this checklist is, like, what I'm usually following, um, when thinking of, like, building the actual, like, uh, the factory. Um, and part of that is, like, is it runable? Um, there's a typo in here. I blame my Swedish.
Uh, there's is it accessible, like, the context that the agents need to have? Um, can they interface with Linear or Notion or Datadog or Slack, et cetera, just to understand and, like, see what's what is, like, the broader context of this the intent that the user would have?
And lastly, which I think people should be spending a lot more time is, like, building verifiable systems. How can, um, the agents themselves, like, verify their own work, whether that's through, um, unit tests or integration tests or, uh, UI tests, like, actually clicking around in the DOM and, like, trying to reproduce things that's actually happening for the end user?
Um, this is arguably easier for, uh, backend systems where there's, like, no UI really happening, and you can have, like, clearer contracts and boundaries of what should work and what shouldn't. Uh, whereas for, for web and UI and all that, you actually need to click around and making sure things work, the buttons actually have a loading spinner, et cetera.
Cursor Demo10:22
Okay, so this is, like, part of building the factory. So if we switch over to Cursor here, um, I'm not sure if you've seen this, but this is Cursor, uh, 3. Uh, we launched this a couple of weeks ago, and it's a complete rewrite of Cursor.
There's no VS Code anymore. Uh, most of you are probably familiar with this type of Cursor, um, where you have files and sidebars and, uh, a lot of different things, uh, whereas this is a bit more streamlined for, like, an agent-first, uh, workflow.
Uh, and we'll get to, like, why we created this as well, so, uh, at a later point. But I want to show you some parts of, um, some rules, et cetera. Let's see where I put them. Um, so for example, I built this music agent, uh, project.
And if you've used Ableton before, you probably recognize this.
Yeah. It's, like, really small. You can't.
Yeah, yeah, I'll expand it.
More? Good? Okay.
Um, yeah, so if you've used Ableton or any, like, music production software, um, you probably recognize this interface. Oops. Uh, it's not really working on this size.
Um, but what I, uh, essentially asked the agent to do here is, like, can you start a local dev server? And we can see that it worked for a while. Um, it explored some files, read package.json. And based on this, uh, there is a start script.
So, like, package.json and all these appendix files are so in distribution of the models that they know, like, we should immediately go to package.json, uh, if there's a JS project or if it exists, um, to look for a start script.
And this is, like, a good example of having, like, a pattern that is predefined and, like, making your code base more, like, in distribution, uh, in that way. Um, because now it's, like, it's super easy for the agents to understand, like, oh, I should just go in here and start the server.
So it started the server. Um, it's running on localhost:3000. Um, and let's see here. We can see that we have this agents md file. Um, so agents md is, like, Cursor rules. Uh, it's, uh, cross for many different harnesses.
Um, and what I wanted to accomplish with this project is essentially, like, building a factory around this idea of building, like, a online music, uh, creation tool. Um, and to do that, I, like, I forced myself never to write any code, um, myself, try not to look at the code that much either, and just, like, try to figure out, like, what is the systems and the structures I need around this.
Um, and, um, immediately, it became pretty clear that we need a way to start the project. Um, we need a way for the agent to, like, verify its own work. So the agent created this, uh, end-to-end tests, um, using, uh, Playwright so it can just spawn browsers, um, go to a route, et cetera, click around and get by test ID, uh, and making sure, like, for every different change I make, um, for example, the play button still works, or I can add notes to this project here, uh, without, um, anything breaking.
Um, so these are, like, some examples of, um, how you can create, like, verifiable outputs like that. Um, okay. Uh, we have vtest, we have this, et cetera. So let's see here. If you go back. Um, oh, yeah, another option here.
Oh, casual scrolling of Twitter. Um, a different way to verify the work is using, um, like, an automation to recode review. Um, you can ask the agent to just review the changes it made. Um, or you can use, like, um, a more, like, integrated tool like Bugbot that we have in Cursor that just looks at, uh, different PRs, uh, in GitHub and reviews them and comes back.
Um, and this is, like, also, like, one piece of the whole, like, factory that you should have multiple different stages where you pr you plan it, you produce it, you review it, and you essentially follow the whole, uh, SDLC, uh, but you, like, automate and codify, uh, this work.
Um, let's see here. Yes. Um, I di I want to show you this as well. Um, so we launched, uh, updated cloud agents, um, in the last couple of weeks where we gave, uh, each agent their separate VM, and you can have them, like, create this very reproducible environment in the cloud.
And this essentially allows you to scale, like, infinitely. Uh, but we also gave the agent a tool to test its own work, um, by controlling the computer. So, for example, we have Glass here, uh, which is the interface.
And I asked the agent to let's see here.
Uh, Glass agents still feel rough with the keyboard, control tab, et cetera, like, better, uh, accessibility and, um, using the keyboard to navigate the agents. Um, and I asked it to, uh, make the change and then record this with the full editor because the first one was just a sidebar.
So what we got back here is just a video of the agent actually testing its own work. So we can see that it has this highlighted row. I'm not sure if you can see that. Um, but just some context for me as a human to verify the work.
Um, and then it actually clicking around and using the keyboard to, to navigate. So with this, we're, like, we're getting kind of far in, like, the factory, like, where we're at. Like, a lot of the things are automated.
Like, review is automated. Um, the testing is automated. Um, uh, we have some rules to, like, steer the agents, et cetera. Um, but there's still a lot still a lot more to do. Um, so I think when you have this in place, the most important thing you can do is, like, shift your mindset.
Running Factory17:03
Like, you are going to look way less at code. So you are going to go from, like, worker to manager, um, where instead of just doing the work yourself, you're overseeing a lot of agents doing the work, uh, for you.
Um, so this also means going from sync to async because most of the work is going to happen in the background, and you can still tap in and see what's going on, uh, for different agents. But the more agents you spawn over time, the harder time you're going to have to, like, understand what's going on in each of them.
So then you need a way to aggregate these changes, like, upwards. Um, and it's just I think it's so interesting that it's just the same as, like, in human organization, like, all the same principles kind of follow. You still have, uh, you start with a very small team, and then you add more and more people because you need to get more throughput.
And all of a sudden, you need a manager to, like, oversee things. And then you add more managers, and then you need a manager of the manager. And this is essentially what's going to happen with agents too, but you are just going to, like, keep on going up the lev levels of abstraction.
So when you're a manager, you need to start thinking of, like, how do you scope and parallelize the work? Uh, because you want to get, like, higher throughput. Um, but some things are not necessarily, um it's not good to make all the changes at once.
For example, if you have two different tasks working on the same part of the code base, you're going to get merge conflicts. So you need to still, like, plan out, scope, and parallelize, uh, the work. Um, and one, like, one unit of work can always be one agent.
Um, so then, like, how do you take a long, long list of things you want to do and actually, like, make the most out of that, um, and run the most amount of agents that we can do? And to do this, I think it's important that you preserve, um, like, tribal knowledge of the code base.
Like, you still understand what's going on in the different systems. Um, you know, like, how data flows, uh, what users want, um, which part are critical, which part are don't. Um, so not outsourcing too much, uh, to the agents, but, like, very be very, like, direct, um, and, and managing, managing them pretty well.
And when you're going from sync to async, you are going to need to trust the agents a lot more, um, because you are going to send them off and doing longer and longer tasks. And when you do that, you need to, like, get more context up front.
So you kind of, like, front-load, uh, the context to the agents, either through, like, a plan or a long spec, and then you send them off, and then you let them go. And once you start doing this regularly, you're going to underst like, start to feel the agents.
You're going to, like, understand the models. You're going to see, like, these are the weaknesses, these are the strengths, and you are going to create, like, this alignment with the models. So you know, like, how to prompt them and what intent to give them.
And again, as the models keep getting better, you have to give them shorter or less and less prompts as you used to, uh, before, but you still got to provide the intent and be very clear, like, what, uh, with the change you want the agents to do.
Um, and there's, like, no there's no shortcut to this, uh, from what I've found and from what the team has found. You just got to, like, spawn a shitload of agents and just, like, let them do the work and see what happens.
And as long as you have good safety guardrails, you can just let them do that. Um, so you probably shouldn't let them push to prod, like, straight away.
Sorry, one question. Do you do you multiply the, the working environments as well, or do you let them all the agents work in parallel on the same development environment?
Um.
On the same code base at once?
Yeah. So this kind of comes down to, like, um, personally, I'm always using isolated environments, so in different VMs. Um, I just tweeted about this actually because on one hand, if you're sharing the workspace, you can have, like, Git work trees where you, like, have diff shallow copies essentially of the code base on the same machine, and you can reuse services.
But you're still going to have to branch every, like, database or cache or user management to have, like, reproducible and separate environments. Like, if you are going to make, uh, a lot of changes at once, you need to you want to know that they are pure and they're not, like, having side effects to the other branches.
And that's why I found, like, just using cloud agents, uh, where I spawn a VM and this VM can run a database, uh, uh, internal tooling, databases, other stuff, uh, and the Cursor app itself, and then have the agent just work in that isolated environment to be much better.
Um, it is more expensive. It's going to take a lot more work to set up your, like, factory or your environment to support this. But once you have it set it up properly, you can scale this to, like, 100 or 1,000 agents.
I'm not sure how many we are running today, but I bet it's, like, multiple thousands a day, um, just agents running in the same or, like, copies of the code base. Um, so that's what I would recommend.
Um, yeah. So when you're a manager, like, your job changes quite a bit. Um, so, um, you have to, like, look at your system as a whole. You got to, like, think of where is the human in the loop needed.
For example, do you have a log service like Datadog, and do you need to copy-paste the logs and go into the code base and paste them and, like, run the agents to identify and, and trace down issues? Or do you have user feedback that you need to copy-paste from Twitter into somewhere else and let the agents do something with that?
Um, do you have, like, a Notion thing, uh, where you have all your specs and you need to copy-paste the Notion or export them into Markdown and then to the agents? There's probably a way to, like, automate all these different things, uh, either it's, like, skills or MCPs or either or, or separate automations.
So think of, like, where is the human in the loop needed and try to, like, automate that away. Um, the second thing is, like, catch where how can you catch agents going, like, off, not doing what you actually want it to do?
Um, and this is, like, the this, this is, like, the perfect flywheel for improving your factory as well. If you can see agents, like, um, creating, like, wrong, uh, schemas in your database because they're not following naming conventions, et cetera, that's probably a rule somewhere.
Um, or if they are, um, just producing really ugly UI, there's probably a way for you to create a design system and let the agents be aware of the design systems where they can, uh, incorporate that and, uh, use it for the next kind of, like, iteration you do.
And yeah, then you take all these learnings and, uh, you use it to actually improve the factory.
Scaling24:03
And thirdly, uh, it comes to, like, scaling the factory. So now you have, like, your environment set up. You know how to, uh, be a manager, to, like, manage a fleet of agents. You scope the task, and you do all this.
Um, so how do you, like, actually take it from, like, 5 agents to 10 agents to 50 to 100, uh, agents? And, um, the thing is, again, um, not looking at code is going to be a real thing if the models get better, and they are getting better.
So observing the outcomes, um, kind of like the same thing as previously, like, where do they go off the rail? Uh, what are they producing? What are the artifacts, et cetera? Um, how can you make it so that the agents also can verify their own work and verify the outcome that they produce?
Um, you should set up automations. You should look again at the things you're doing repetitively. Um, so one thing we could do, for example, here is if we go to Cursor and we go to, uh, this music agent again, uh, I can ask, uh, looking at my, uh, chat history, what repetitive tasks, um, I doing?
Um, so we can ask the agent to, like, look at this and identify potential opportunities. Uh, so it's searching the agent transcripts, and it's producing
some kind of artifact of this. Um, yeah, we'll see how this goes. Um, I actually built this into a plugin. Oh, let's see here. Uh, plan execution loops, restarting the product direction. Um, let's see here. Ableton, like, UI iteration.
I should probably, like, put this in a rule saying, like, make it look like Ableton. Um, tooling, housekeeping, et cetera, et cetera. So this product is very short-lived, but if you're looking at an actual production thing where you have prompted a lot over time, you're probably going to find things that you are doing recurrently.
And I want to show you some things that we are doing at Cursor, um, that we are automating. Um, and some of these are not that obvious all the time, but one is, for example, let's see here.
Oh, not this one. Uh, let's see here. For example, daily review. So I have this, um, automation for checking my own daily review. So this is going to, um, look at Slack. It's going to look at GitHub. Um, and it's going to send me a summary of the things I've done, um, over the last day.
So I would previously have done this, like, writing down my notes maybe, um, thinking of, like, what did I get done today, uh, or, like, running an agent with access to MCP. But now I can just put this on a schedule and do this automatically for me.
Um, I want to show you a different one. Uh, for example, read merge PR comments. Um, this is also, like, a way for you to, uh, learn over time. So for all the PRs that we merge in our main repository, we can look at the comments, and we can look at what did humans actually review here, uh, and what did they say about the changes I made?
Uh, because if it's if a human actually goes in and reviews a PR and leaves a comment, there's probably, like, high, high value and high signal and high intent, uh, in that comment. And we can then store that later, uh, in order for the agents to actually learn over time.
Um, we have another one, uh, which I can show you here. Uh, this one. Yeah, again, the code owners. Um, so this one allows us to we essentially had this problem where, um, we had code owners in our code base, and they were kind ofright most of the time, like 80% of the time.
But for these 20% of the time, they caused a lot of bottlenecks for us internally. Like, we were blocked on merging the PR. We needed someone else to, um, to review it for us, and maybe they were in a different time zone, perhaps.
So what we started doing was building this agentic code owner thing. And what it essentially does is looking at PRs and checking, like, first of all, what's the risk of this? What's the risk level? Can we is it just, like, changing a variable name?
Or is it changing a constant that's changing, like, how long a trial subscription is or something like that? Um, and if it is low risk, it can just approve the PR because we don't really we don't want to block, uh, our own engineers, uh, on these things.
But if it is, um, we can see that it is a high-risk PR. And then we can find, like, okay, who, who made changes to this previously, and can we, like, pull in their feedback, um, and making the most out of this?
And, like, um, first of all, making the code safe, um, and not breaking any systems, but also for the user that actually did the initial change, keep them in the loop and, like, keeping them up to date on and refreshing their context of what's going on here.
So it kind of, like, it goes both ways. Um, and, and yeah, multiple, uh, value adds from doing this. Um, let's see if there's one more review. No, I think that was pretty much it. Um, or yeah, I have this one more thing called continual learning.
Um, so continual learning is another type of automation, um, that I created, um, a couple of weeks ago as well. And it essentially does what we did with the agent. We look at the previous transcripts we have, and we can then extract, like, memories and learnings from what we said previously.
Like, if we're correcting the agent to do, uh, a certain thing, like, um, use this component instead of that component, or, um, always, uh, refer to me as, uh, like, always, like, have very, like, verbose, uh, descriptions of things that you're doing.
Instead of me, like, every time going in and, and, um, asking the agent to do this, I can create a rule. But I'm kind of lazy, so I don't really, um, remember to create a rule. So instead, we can have this continual learning plugin that looks, looks through, uh, the transcripts and store this as a rule, uh, for you instead.
Um, so these are all examples of, like, systems to automate yourself away and to automate, like, things that the agent can do for you. Um, and I think that's the important part of, like, building these factories. Like, how can you identify, uh, the flywheels and loops where you can, uh, automate yourself away by building systems?
Um, okay. Um, and yeah, you are going to move up abstraction. So now you're managing 5 to 10 agents, but tomorrow you might be managing, uh, an agent managing other agents. Um, and that is just going to grow.
Like, you're going to have a lot of sub-agents, like, under you working for you.
Um, cool. So yeah, what I want you to take away from this is be very clear about the intent and, like, really think about what's the actual problem to solve here, what do we want to get out of this.
Um, don't outsource important decisions. Like, make sure you're staying in the loop for important decisions, um, whether this is, like, uh, safety or security or databases or payments and authentication. Um, some things are really important and should not be made, um, um, should not be decided by agents, but by humans.
Um, build tools and systems. Try to find these flywheels and, like, codify them and get them in, uh, your systems and let the agent have access to them.
Um, store context for later, whether that is, like, agent transcripts or artifacts of things you think look good. Um, because this is going to help the agent to, like, know what good and bad looks like over time. And this is going to change.
Um, so storing the context and building the tools and, like, keeping them up to date is, is more important than actually doing the work because this is going to provide, like, the framework and the guardrails, uh, for the agents.
Um, and lastly, like, let the agents be free. Like, think of what do they need? Um, I have a friend at Lovable. Uh, he mentioned that they set up a Slack channel or he gave the agent a tool, uh, vent tool so the agent could complain about things, uh, when it was running.
And the agent started complaining about, like, hey, um, I can't, like, access this image. Um, I'm, like, very frustrated about this. And then it posted straight into a Slack channel. And they, they set it up as a joke, but then they started scrolling through and, like, oh, this actually is very valuable.
Like, we should probably, like, give the agent access to reading images. And they did. And then the agent started complaining about something else. That was a problem with the harness. Um, so find ways to let the agents be free.
I think that's, um, very important thing.
Um, okay. That's kind of it. Uh, and that's kind of, like, a direction of and things we have found, like, building Cursor and, like, taking Cursor towards, uh, software factory. Um, I hope you learned a thing or two and can take away, um, some of this.
I'm happy to take any questions about anything Cursor. Yeah. Or actually, now we have the microphones coming here.
Q&A: Guardrails33:39
Uh, thank you very much. I have a question about, um, uh, code quality or architecture quality. So when agents ship, uh, tons of code, uh, and, uh, you barely can review them, uh, how do you, uh, ensure that the code is extensible and so on?
I mean, um, you can, uh, establish hooks, uh, or guardrails for measurable things like, I don't know, uh, number of lines in the file should not be more than something. But, uh, the architecture is not, uh, measured this way.
So, um, and agents, uh, they have this completion bias. They want to finish task, uh, as soon as possible. And, uh, they, uh, don't think ahead. They don't have their, uh, picture of the future, how code will evolve.
They just want to finish task now. And, uh, yeah. Thank you.
Yeah, it's a good question. Um, I think we as humans have the same problem, but it just takes a lot more time, uh, for us to, like, discover them. Um, one pattern like, the good thing about agents and models being, like, um, essentially, like, completion machines is that they will just look at existing references and just continue forward with that same path.
So if you have existing things you can point them to, I think that's very important. If you don't, I think there's a case where you let the agents do, um, one-off implementations here and there, and then eventually you have another agent, like, refactoring like we do as humans as well.
So, like, want to generalize, um, and build abstractions and all these things. Um, so, like, how can you build, like, a system to, like, detect this and verify that the abstractions that are getting built is also good and in line with what you want to do?
Um, but I think it's going to be, like, a lot more architectural review for humans. Um, and, and scoping and, like, planning of what the architecture should look like and system design. Um, but yeah, it's, it's a tough problem.
Thank you.
Hello, Eric. Thank you for the talk. Um, when it comes to the activities of building the factory, one thing that I observe, for example, when it comes to building things like rules in a in a team is that because it's so new, almost everybody feels, oh, this is a rule for me, and I don't want to inflict it on other people.
Mm.
And I notice this creation of silos where each engineer ends up having their own separate different factory. Do you have any advice on how to bring it to the point where the whole team is contributing to the creation of the factory?
Uh, it's a it's a great question. Um, I think it's hard. I think it's very cultural as well. Um, I mean, like, we developers have always created our own tools and, like, we want to have our own custom setup.
Uh, but at some points, like, we have to unify and, like, uh, on a certain structure. So I think historically we have, like, had PR reviews and all these kind of things as a ceremony to, like, align on the code that's being produced and making sure it's consistent.
I think we got to take the same principles and apply that to the tools we're building as well and, like, the, the guardrails and enablers and primitives. Um, so I think, I don't know, establishing some kind of, uh, a forum where you can discuss these things and, like, plan, like, what do we want the factory to look like?
What are the components we need? Like, what are the integrations we need? Do you have any examples of, like, specific things that people is it, like, flavor or is it more bigger changes that the agents are doing?
Um, what I notice when it comes to rules, they, they create, like, oh, I want to like, one person wants to write the test first, and they create the rule to write the test first, but they know that somebody else doesn't want to do it that way.
Mm.
So then they have the rules only on their machine. They, they don't share it because it is too unique to what they are. So they're collaborating. The whole team is collaborating on creating the code base.
Mm.
But the collaboration in creating the factory in thinking, well, are we deciding now that the factory writes the test first or not?
Mm.
That is a big decision that is hard to align everybody and accept that. Like, with all of these rules, not everybody's going to be completely on board. And in most cases, it doesn't matter when.
Yeah.
When you defer a little bit, but it is hard to.
Yeah, I guess it's, it's, it's a human problem and a human change that needs to be made.
Yeah. Yeah, this is a human.
But it's, uh, it's a good question. I'll think a little about it. Thank you.
Thanks for the talk. Um, a lot of the patterns, uh, resonate. Um, I was wondering what is needed? What kind of patterns can you suggest to take it to the next level if you work on enterprise brownfields, mission-critical systems that cannot fail, that cannot be insecure?
If you look at the recent supply chain attacks and you give your agents sandboxes, maybe that's not even enough.
Mm.
Um, so do humans remain accountable?
Yeah.
And we can't say, oh, uh, it's not my fault, my agent did that. So do you have any extra patterns that, um, um, or, or is it just inherently we have to keep reading the code, which may feel like reading assembly lines in the '80s or something?
Um, I think if you can spend a lot of compute and tokens upfront before you as a human actually, like, needs to be involved, I think that's a pattern that we found to be pretty successful. Um, so one thing is, like, manually writing tests for very critical parts of the systems, um, and then just letting the agents, like, run them, uh, a lot.
Um, the second part is, like, building automation to, like, um, our security team, they built, like, the Security Sentinel, which is an automation that, like, looks specifically for very, very, uh, specific invariants, um, of the system. And they run, like, 10 of these on, like, certain PRs that change certain files.
Um, and then, yeah, I think I think it's a bit contextual as well, but yeah, just spending a lot of tokens before, uh, and trying to, like, find different variants and, like, almost red teaming. Um.
So one thing I did is, um, instead of focusing on velocity and throughput, I focused on quality.
Sorry, what?
I used AI to focus on quality.
Mm-hmm.
And just improve the tests and just make it completely AI-ready.
Yeah, I think that's very good. Um, because if you as a human trust the tests, you probably are trusting the output even though you don't have to look at the code. Um, and that's kind of, like, where we're going.
Trying to think.
Hi. So, uh, thanks for a great presentation. Uh, I find myself kind of, like, lacking and slacking in using guideways, uh, especially, like, rules and hooks. Uh, partly because historically the, the, the knowledge of how to do that properly was very scattered and decentralized across, uh, whole web.
So you would have this, like, exotic GitHub repos would, uh, try to, like, centralize this knowledge. Or maybe you would have some, like, medium articles. Or maybe Cursor would Cursor company would do a blog post on this,right? But still, it was very evolving.
And also the capabilities of models themselves on especially on instruction following, uh, they are also evolving and they are getting better on that. And, and, and it always felt like kind of like duct taping.
Mm.
Uh, to me. So I'm wondering, uh, basically, can we have AI to help us with that? Meaning that could Cursor, for example, give us, like, proactive agents or maybe some new setup, uh, or maybe wizards kind of, uh, setups where we could identify our workflow and then help AI build us rules and, and guardrails and all those, like, rules artifacts, uh, for us.
So maybe just like a proactive agent. Uh, so, so maybe we would have, like, an agent that would scan our workflow globally.
Yeah.
And then help us build those artifacts. What do you think about it? And do, do you guys think about this in the company? Maybe do you work on that?
Yeah, totally. Um, I think now there's, like, two places where you can do this. One is, like, in the product itself with the whole, um, uh, with the, like, continual learning pro um, let's see here. Uh, oh, I don't have it installed.
Uh, we can go to marketplace. Yeah, with the continual learning kind of, uh, plugin to actually, like, look at your, um, transcripts and, like, extracting rules and memories and all that. That's, like, one way to do it. Then there's, like, another world where, um, you, like, change the weights of the model depending on, like, what your code base looks like and what, like, your engineers are doing, like, in a specific team.
Uh, and you, like, you reconcile them. And it's like, it's, like, true continual learning, uh, not this, uh, hacky plugin. Um, and you, like, actually bake that into the model, uh, so they actually know what your preferences are, etc.
Um, but totally, like, memory and rules and all that, I think that's going to become more and more important over time, um, because that's kind of, like, what's lacking. That's kind of, like, what's preventing me from having a lot of trust in the agents sometimes because, like, I say something and they forget about it, uh, but they're just, like, stateless machines.
So how do we capture this knowledge? Um, so I think we should put a lot more, like, time and effort into, um, building these systems.
If I may just follow up on that. So, so you, you say that you seem to first to, to start a project or, or to, to dive into the project that's already existing in the code base and then to build rules on top of that.
How about we, we first have rules and we want to start a new code base, new project.
Mm.
How to how to actually have those good rules for us? Do, do you think that humans should humans should still do that or can we also automate that? Can do we have a new best workflows for that?
I think it's hard because, like, my perspective of rules is, like, the bridge between, uh, the model behavior and the, like, the human behavior. And, like, how do we steer the models in a way that they follow me as a human, what I want to do?
Um, and in a new product, I'm not really sure what I want to do. Like, I kind of want to, like, outsource that to the model, like, see what are they doing here? Can I run different models? Do I want to, like, combine them or do I want to scrap everything?
Um, so I think it's, it's hard. Like, one like, the best example of a rule that I can think of internally is for BugBot. Um, so when we're doing database migrations, uh, we're not really using foreign keys on a database, um, for performance reasons.
And the models, like, theright way to do is to use foreign keys,right? Um, so they will always add a foreign key. Um, but when it, like, hits GitHub and there's a PR created, we have BugBot looking at this and reviewing is like, oh, I have this rule saying, like, we should never use foreign keys.
So then it flags this. Um, so that's, like, the gap between the human and the model and what we want the desired, like, intent we have versus what they have. So I think rules should, like, emerge dynamically over time.
Um, and before that, you should probably adjust to this ephemeral, like, specs and plans. Um, oh yeah, there's, like, one over here.
Yeah.
Oh yeah. Oh yeah.
Q&A: Workflows45:58
It doesn't work?
Oh, it, it worked. Uh, so thank you, Eric, um, for the talk. Um, as evolution evolution and trust is a big point, I'd like to know how you effectively do, uh, GUI testing and, uh, user acceptance testing.
Mm.
Automated.
Yeah.
If you could show, show, like, uh, something of your workflow.
Um, totally. The best or, like, the main way I do it is using let's see here. Oh yeah, I have this one, for example. Um, the main way I do it is using, uh, the Cursor cloud agents with the computer use that we have.
So I'm going to publish this. Oh no, that's bad.
Uh, I guess we're not doing that. Um, I have this website where it's running a, uh I have, like, seven components, um, like, a button, a dropdown, etc., etc. web components. And then I'm generating each of these components with a different model.
And because I want to, like, compare, like, what does, uh, Composer dropdown look like versus, uh, GPT-5.4 dropdown look like? And I put this in a grid. But when I created this website, there was an error where I had this, like, view code button so I could actually see the generated code.
It was not working because the model didn't, uh, bundle the actual code. So I went to Cursor and I clicked when clicking view code on a component, it says it cannot load a code. And it's like it's a very, like, short description.
So what the agent did, uh, you can see here, it spawned my local server. Um, it started, like, clicking around and pressing enter. Uh, we can see the cursor up here. And it's creating this, like, screen studio-esque, uh, recording where it's, like, um, chopping and speeding up and zooming in, etc.
Um, so here, it's taken a while because computer use is fairly slow. Um, it's consuming a lot of tokens. And we can see we have this view code button. And now we can actually see it's working too. Um, so since this is a very, like, much of a side product for me, I'm not really going to look at the code.
I'm just going to, like, see that this works and I'm going to merge it. Um, but you can keep on prompting the model to do very specific things for you, like, can you follow these, like, specific instructions? Um, like, a login flow, for example, you should click the button, you should log in.
Um, the models, this, like, login steps are probably so much in distribution that you can probably just prompt the model to say, like, go to this URL and click login or, like, log in. And it's going to, like, understand which steps it needs to take.
But then you can ask the model to, like, uh, input a wrong password or input a wrong email and see, uh, what are the results from the website? And maybe the, uh, website is giving, like, wrong credentials. And then the agent would understand, like, oh, I need to, like, put in theright credentials.
Um, so just like you would, um, like, hire a consultant, like a QA consultant, and giving them instructions, you would just give the same instructions, uh, to the agent. Um, so this is, like, one way to do it.
Um, I guess the other way would do, like, more, uh, playwright/puppeteer and just automating, like, a browser thing, uh, which is a bit more deterministic as you can review it, um, and check it in and, like, have other people reuse it.
Is that does that answer the question?
My question was going more into, uh, like, user acceptance testing to check does this thing actually lookright because that, uh, testing a login you can do you can automate that.
Mm-hmm.
You don't need an agent for that. But, like, does the does the website, uh, lookright? Is it consistent through all the pages that are generated and stuff like that?
Mm. Yeah. Yeah. Then I then I, I use cloud agents for that a lot. Um, there was one I can't remember now, but I think it was I did some changes in the docs and I just asked it to, like, open every single instance where this where it is referenced, uh, take a screenshot and give it back to me.
So then I could just, like, look at all the different screenshots, everything looked good, and then I could merge the code. Um, so letting, like, the agents do, uh, the navigation and clicking around and, uh, the testing for you.
Um, I think it works surprisingly well. Like, this was, like, a very much an AGI moment for me, uh, when we launched this in last year sometime internally. So have you have you had a chance to try cloud agents in Cursor?
No.
You should.
How is the.
Curious to get your feedback.
Around this. I know you are you have responsible.
Uh, which one? This one?
No, no, the agent, like, spawning all the VMs and the giving us the logs for.
Uh, well, what was the initial question? How long it took or?
Yeah. No, I, I see. But how expensive it would be for, like, a typical user?
Ah, um, very straightforward. Like, I for this one, I didn't know specific setup. Um, for, like, our own repository, uh, where we have, like when running Cursor, it's like we can actually, like, reproduce. Like, this demo here is running all the backend services for Cursor.
It's running all the frontend things. Um, and this is, like, a lot of a lot of different things. Um, so the VM is quite beefy. Um, but as long as you give theright instructions, it's working really well. What we did was creating this internal CLI that the agent could use to so, like, uh, we call it, um, like, Cursor Dev Tool.
Cursor Dev Tool backend start, Cursor Dev Tool frontend start. Um, and that is abstracting everything away, um, that actually needs to get to, like, uh, Orbstack to running, uh, ClickHouse and Postgres and Redis. And then the frontend running, like, um, Electron and then, uh, Glass here.
But then they just, like, coexist the two different processes. Um, and the agent have access to everything, like, just as a human would do. Um, and you can have, like, the agent be authenticated if you store, like, a snapshot where you are authenticated, etc.
Yeah. More, more I meant how is costing was in dollars, like.
Oh, sorry, sorry, sorry. Okay, okay, okay. Uh, my bad, my bad. Um, yeah, this one, I don't really have the I can probably look it up. I would guess this is, like, hmm, $1, something like that.
Yeah.
Um.
Just one term or?
There's, uh, like, for one ter probably, like, this initial one would be $1. Uh, and the other ones, I just asked them to re-record a bunch of different things. Um.
Hmm. Okay.
Something like I can look it up later. Totally.
But I jump back and forth between Cursor and other, like, like, cloud collaborative analytics.
Mm.
I, I don't know.
Yeah. And I guess depends on which model you're using too.
Yeah.
Okay. Okay. Uh, my question is about, uh, handoff between humans and.
Mm.
And agents whenever you are using different tools. So in my current setup, I have a product owner and a functional analyst that they, they work on cloud code and they prototype very fast, uh, with basically without, uh, um, so much thinking about, oh, the backend, the architectural choices or whatever.
And then they pass the, the control down to the delivery team that uses Cursor and has to make that stuff work, actually work. Uh, which best practices do you suggest in order to enforce a proper workflow between people just not knowing basically what they are doing?
Mm-hmm.
Uh, on a technical point of view, of course. Uh, and the people that needs to bring that thing that maybe has, uh, okay, some poor choices such as, okay, use that database or then, uh, cloud could change the idea and they move from a supabase to, uh, Turso, to, uh, any other kind of fancy database that actually is in that, uh, uh, in that environment and then bring that into some sound architectural choices.
Mm.
Moving from cloud code to, to Cursor.
Mm. Mm. I think what we're doing internally is, like, we have, like, one or two PMs and they are building a lot of different prototypes. Um, sometimes it's actually in the real, like, product itself. Uh, they're using maybe cloud agents and just prompting them.
They're getting, like, a video, like this back of the changes. And they say, like, oh, it kind of looks like I wanted to. And then they tweak the designs a bit. But the code might be really bad or, like, not following best practices.
Um, which if they had a if we had a good factory, then it probably would. Um, but if that's the case, uh, we hand off, like, a link to the cloud agents. We just copy the link and just send it to the, the, like, engineers, like, hey, this is, like, something that we want to build.
Um, does this make sense? Like, can we do this? Um, and then you have a lot of intent already expressed. Um, but the other case is, like, having the PMs, they have a separate repo called, like, prototypes. And it's just, like, an HTML file, like a mega HTML file, uh, reproducing, like, the Cursor UI or the dashboard.
Yeah. The, the problem is the migration. So, uh, just, uh, practical use case, I had my PO and functional team, uh, build out a very fancy demo using a Prisma and Turso.
Mm.
And whatever database and then storing data on Vercel, uh, Blob storage. And then my delivery team had to migrate that to use SQL Server and, uh, um, C# and Aspire for the backend. And the migration was really painful.
Yeah.
Even because, uh, when they use the agent freely with no constraint, uh, the agent sometimes decided to use, say, Next.js. Some other times decided to use Vite. Another time, it, it decided to use Svelte. And, uh, putting constraints in form of rules within that agent.
Mm.
Shape that down the path. But the problem is that, uh, we need to, uh, write a lot of rules and make them consistent. Uh, and it is not easy to, to manage all the workflow. So we, we are shifting a lot of effort from, uh, having people to write code to having people to write guardrails and the rules and whatsoever and make all the pieces talk to each other.
Mm. I see. Yeah, yeah, yeah. Yeah. I guess, um, if, if the POs and PMs can't have access to the actual codebase, just, like, handing off an artifact is, like, the minimum viable intent, uh, which could be, like, an interactive like, back in the days, it used to be, like, Figma prototypes,right?
You can click around and you get, like, a feeling for it. Now you can have them even higher fidelity where you have an interactive prototype using, like, web technology without, like, touching anything of the backend stuff or it doesn't have to be, like, a working thing for real if it's just a prototype internally.
Uh, but just enough to, like, your engineers can understand, like, oh, this is, like, the intended thing. If I click this thing, that should happen. Um, or if I, like, entered some text here and click send, a row should show up here.
Um, and I think all that can just be done, um, in the frontend. Kind of like a hackathon.
You don't think to migrate the, the prototype into something that becomes production really, but rather, uh, rewrite that?
Yeah, I think so. I think rewriting. Um, and I think, like, um, setting, like, clear expectations from, from the engineers to the PMs and POs, like, what engineers kind of want from the product organization and, like, what's most helpful for them.
So maybe not, like, vibe coding complete SaaS products is the most efficient thing.
Okay. Thank you.
Isaac, thank you for the presentation. Uh, my question as we're building more and more agent and it became part of our time-critical processes, how do you see the brownouts and blackouts as, as, uh, as a as a new risk?
And, um, what's your, uh, what's your view how it can be mitigated and, and the impact reduced?
Yeah, it's a great question. Um, it's a really good question. I think it comes down to what we talked. Like, the humans are still accountable for the things that's being shipped. Um, so the humans need to build, like, systems and observability and monitoring around the changes that's being made.
Um, and I think that still, like, comes down to understanding which are, like, system-critical areas of the codebase, making sure you have good, like, observability and understanding of everything that goes on. Maybe, like, every line should be human rewritten in these critical things or at least, like, always humanly reviewed by one or two people.
Um, and yeah, it's, it's close to vibe it's easy to vibe code close to the sun and fly too close. Um, so I think it's also, like, a cultural thing where you have to make sure that the humans are still, like, accountable for, for the things getting shipped.
Uh, but yeah, setting up good systems to understand, um, the changes being made, I think that's important. And tests.
Hey, Eric. Thanks a lot for the talk. And I'm assuming you're probably one of the people around the world that has the best understanding of how to use these technologies. So this question takes a step back about from the technology and things about processes and how do you manage yourself in your workdays.
And I wonder how long are these tasks or how, how long do you get to be away from your agents without babysitting them? And how do you actually invest this time? Uh, let's say you have 5, 10, 15 minutes.
How do you make the best out of your time? And maybe how many agents do you have in parallel, like, mental processes and how do you manage yourself? Thanks.
Yeah. It's a great question. And I think, like, once you like, there's, like, two levers to pull. Uh, one is, like, the scope of the of the change. Like, the larger the scope is, the longer the agents are gonna run.
And if you want them to run for a really long time, uh, you want to have, like, verifiable, um, systems so, like, they can check their own work, etc. Um, and the other thing is, like, how much can you parallelize?
Like, how many of these agents can you spawn off? Um, and I think the sad reality in some sense is that there's gonna be a lot of context switching. Um, I probably work in four different repos or, like, four different areas of the codebase at the same time.
Uh, whether that is, like, through a, like, single, like, feature that requires frontend, backend, database, um, testing, yada, yada. Uh, or if that's, like, five completely different things. It could be, like, docs. It could be, like, uh, side projects I'm exploring.
It could be fixing a bug from a Twitter user. Um, but I usually they range from, like, um, probably 5 to 10 agents five agents, like, asynchronously running in the cloud at all times. And while I'm waiting for these, I'm either, like, scrolling Twitter or
it's true. I also have the browser and Cursor now, so I can just stay in here and do it. Or I have, like, uh, synchronous task going where, like, I'm a bit back and forth. Uh, maybe that's, like, fixing a small thing in the codebase, or maybe that's, like, planning the next thing.
Maybe I'm, like, sourcing in Notion and Slack and just, like, creating a spec in Cursor using a model. So I love to, like, plan synchronously and then just execute the plans, like, asynchronously. And then once that is done, one of my cloud agents is probably done as well.
So I can come back and, like, review that, keep on prompting it a bit, maybe merging. Um, and some parts I still, like, need to test manually. Like, maybe I need to download a copy of Glass or Cursor 3, um, test it manually and, like, this looks good to me.
Uh, let's go ahead and merge.
Thank you. Uh, quick question. This factory building leaves us with a scattered ecosystem of a lot of markdown files. Is there an easy way to organize these files and to keep an overview of the factory you have actually built?
As maintaining a factory would require you to have an overview of the processes you want your coding agents to go through, what tools do you use? What methods do you recommend? How do you keep a mental map of the factory you have built and how do you maintain it?
Yeah. It's, it's a really good question. I think it's somewhat unsolved as well. Um, one of the reasons we rebuilt Cursor to look like this instead of, like, the traditional IDE is the fact that we are using more agents and we need, like, a better control panel where you can, like, see all the agents and manage them and spawn them, etc.
Um, so what's gonna happen with, like, Cursor 3, um, this is, like, the first step at, like, multi-agent orchestration. Uh, what's gonna happen is that these are gonna be, like, nested agents. So you're gonna have, like, opening this one up and you're gonna have, like, 10 agents in here.
Um, so you can still, like, introspect them and see what's going on and following the traces. But you're probably al also gonna have, like, somewhere here, like, some kind of project view where you can see, like, an aggregated status update of, like, here's what everyone is working on and here's, like, the latest.
Here's what you as a human need to review. Um, so I think these are product things that we are gonna build into Cursor. Um, but to, like, set the spec for the factory, I would probably, like, have a folder in your codebase, um, where you, like, outline how certain things should work.
Um, maybe that's, like, just markdown files of saying, "Here are some best practices." Uh, maybe it's probably rules. Um, and establishing some kind of council to decide on, like, what goes into the factory and what doesn't and, like, what are we lacking to, like, improve the factory.
Um, so as long as it's something that the agent can understand and read, which is files, um, that's probably what I would do and just store them as, uh, yeah, in your codebase that's checked in somewhere.
Thank you. Um, I-I'm just thinking about, like, teams of the future. Uh, so, you know, a, a year or two ago, it's like very reasonable to have, you know, an engineering team that might be several hundred people, several thousand people.
Um, what does this do for that? And, uh, what roles a-a-and kind of, like, roles in an engineering team,right? This is kind of akin to almost becoming somewhere between, like, a product manager and, like, a-an architect. Um, so what roles do engineers have?
Yeah, I think I think that's very accurate. Um, it's hard to predict, like, what are the, like, second, third, fourth order effects of, of this happening. And it's definitely, like, writing less code, looking at less code, um, spawning more agents.
Um, it's gonna be, like, how do you take 'cause, like, we're still building software for humans mostly. So, like, how do we know what other humans want? Like, how do we talk to our customers? How do we market what we're building?
How do we do all these things and bring them into the actual, like, factory? Um, who sets the direction? What's the intent? Um, all these things are coming from somewhere. Either it's, like, creativity from someone else's head or it's actually, like, a user demand.
Um, so having someone, like, doing that, it's gonna be very important. Having someone, like, s like, aligning that between the different humans in the org, I think, is gonna be important. Um, having people building, uh, the scaffolding for the other agents and, like, uh, just, uh, pr building the assembly lines where the agents can actually run.
I think that's also gonna be, uh, important. But, like, to what magnitude and how many people it's gonna be, like, in yeah, I don't know. It's really hard. Um, you can do a lot with the modelsright now with a very, like, small team if you have theright setups in place and, like, yeah, depending on the domain you're working in.
I don't know. Do you have any predictions?
Um, well, I, I, I s I see issues with kind of, like, uh, from, like, a labor perspective. Um, if you're if you're working in an incredibly agentic environment, what's your need to, like like, what happens to training new grads, hiring new grads, um, and kind of, like, the, the future from that perspective?
What happens with office politics and, like, land grabbing,right? Because basically, your, your value now becomes in your ability to configure and set up your own kind of, like, agentic team, not in your ability to kind of, well, like, program and be productive anymore.
Mm-hmm.
Right? The 10X engineer is no longer about, you know, words per minute. It's, like, prompting.
Yeah.
Yeah. Token to token usage.
Yeah.
Am I am I paid in tokens? Am I am I yeah. Leaderboard, you know. Um.
Gotta be token maxing.
Am I paid on amount and then, like, my token usage takes away from that? You know, how do you how do you optimize, you know, for that?
We gotta train the models to be more political, I think. That's the solution,right?
We need we need more, like, you know, water cooler talk.
I guess we're gonna have more of that if the agents are doing our work.
Q&A: Tools1:08:18
Hi, Eric.
Hey.
Thanks for your talk. Um, I was wondering, um, probably, uh, you are using, uh, a-at Cursor, uh, some kind of, uh, uh, issue tracking, uh, tools like Atlassian or Jira Jira. Okay. Um, are you using, uh I was wondering if you are using, uh, um, um, agents to check automatically check and, uh, um, read tasks directly from, uh, uh, Jira, for example, and spawn, um, uh, sub-agents to perform the work?
Or if there is always a, a human that, uh, um, start to work using Cursor.
Uh, so we're using Linear for issue management and, uh, we have this first-party integration as well. So for every ticket that's getting created in Linear, we spawn in the cloud agent. Um, so, like, one where I interface with this the most is, like, if we have a feature flag for a specific thing that's rolled out.
And if it's rolled out for two weeks with 100%, um, the system kind of, like, s signals us, like, "Hey, uh, you can it's a stale feature flag at this time. You can remove it." So then we have this to create an automatic issue in Linear.
And since that is hooked up with Cursor, it triggers a cloud agent to remove, uh, the feature flag. So it's kind of, like, completely automatic once the system knows that it's rolled out to everyone. And I can just, like I can probably look at a code and say, like, "Okay, we can merge this.
The feature is no longer active." Um, and we do this for, like, everything. So once you post something in Slack, uh, we either have a Linear Slack agent Slack agent look at it, or we have a Cursor automation to, like, um, look at the message that was posted and, uh, triage it and, like, look for duplicates or, like, if it's determined to be easy, like, start to implement the fix for it immediately.
Um, and this is, like, an example of where a human is, like, in the loop where it might not have to be. It could be, like, me going on Twitter and, like, seeing a tweet, like, "Something is broken with, um, the plan mode, uh, button dropdown."
I can copy that into Slack and then having the agent, uh, perform the work. But there's probably a way where we can just source this feedback immediately without me having to, like, scan it and triage it and copy-paste it.
Um, so that's kind of, like, a bit how we work with, uh, Linear and issue management. Um, but yeah, we we're also, like yeah, since we're spawning a cloud agent for every single thing, it provides a good way for us to dogfood the product and, like, test it out.
But I'm not sure if I would recommend that for, for everyone because it, it can be quite costly.
Thank you.
Um,
as cloud agents are a little expensive, uh, do you have something in roadmap to run something locally? Like, I'm, I'm just thinking of an alternative called dev containers and opening in that. But do you have something planned in the roadmap for that?
Um, what I think the closest thing you can do is probably just prompt the agent to run for a really long time. Um, it's kind of like the same thing with, like, running local models. Um, and the recent, like for I've tried it.
Like, I've probably tried it, like, once a month running, like, the best open source local model and, like, seeing how it works in Cursor. But it's never the same experience as running, like, um, GPT or Claude or Composer.
Um, and the same thing with, like, running really long things locally. I found it to not work that well as if it's running for a long time, it's probably gonna reuse your, your local database, your other local stuff, um, and it's gonna prevent you from doing other work locally unless you, like, create a VM on your own machine.
Um, um, and, and if you do, you could probably wait, never mind. Just re ignore everything I said. We launched Cursor workers. So Cursor worker is, um uh, we launched it, like, yesterday. Uh, it's a way for you to run the same, uh, infrastructure, uh, and orchestration layer as we do for cloud agents, but on any machine you might have.
Um, so you can do, like notright now. Uh, we can do CD dev, uh, or let me see a minute. Yeah. So you can do agents. So we have the agency UI and there's now a worker and you can call worker start.
Uh, so from here, we have a worker running. Um, and this worker is gonna show up in here. Let's see. So we can do self-hosted. Let's see here.
Oh, I don't think it's hooked up yet. There's a different, uh, account I'm running it on. But eventu essentially, you can run this on any kind of machine and you can get access to this, um, from, like, Cursor cloud.
So you can spawn multiple of these on your own machine or you can run, like, a Mac mini or you can have a VM, um, in any, like, cloud platform provider.
Uh, just to follow up on that. So you, you, you are saying that we can have isolated environments in the local itself using this command.
Yeah. So it's.
Can you call the open s still call the frontier models or composer models?
Yes, exactly.
Perfect.
So this is gonna, like, leverage the Cursor harness, um, but it's gonna run on wherever you're spawning this, uh, daemon.
Yeah, that's interesting. Thank you.
So I, like, I built this, like, Cursor claw thing, uh, where I have one running on my Mac mini and that has access to iMessage and calendar and all these kind of other things. And, um, yesterday we launched automations as well.
So I can get, like, um, like, a daily report or a weekly report or everything that's going on in my machine, uh, that I might, like, wanna know on a specific cadence. And since it's running, like, the agent daemon, you will get access to this in, like, Slack and the web and the mobile app that's coming, um, at some point, not too far out.
For iPad too?
Wait, sorry.
So iPads?
Sorry, what?
For iPads. A lot of a lot of time people wanted to have IDEs on iPads.
It's gonna use Swift UI, so it's probably gonna be compatible with, uh, iPads as well.
Cool.
I think that the two versions of iOS and iPadOS are two different things, actually.
Probably.
So we really get design for that's why we have you, you can use, like, GitHub, uh, workspaces.
Mm-hmm.
Uh, on, on iPadOS and kind, kinda works, but it's not really the same thing as Slack.
Got it. Yeah. Nice. Oh, yeah. One more.
I just wanna ask quite simple question. Like, when you have obviously more than one developer in your where you're working in your company and you're spawning hundred and hundred of agents to do a lot of different kind of work, how do you ensure you don't step on each other toes doing the same kind of work twice?
And even high like, you're running internally, do you still do you do you use a Scrum or still Agile ways of working or even that's already kind of gone out of the window already?
Mm-hmm. Um, yeah. What are we doing? We're not really following any, like, traditional methodologies in that sense. Uh, we do have, like, monthly goals and of things we wanna get shipped. Uh, but I think since everyone has so much, like, power at their fingertips with agents, uh, this, like, causes people to have, like, extreme ownership over certain things.
Um, so for the longest time, there was, like, one guy building, like, MCP and rules and, like, all kind of extensibility, uh, by himself. Um, and now we have, like, maybe one person focusing on MCP, uh, but they can own everything around MCP, and they don't really need to interact that much with other teams.
Um, but at some point, that's gonna break too. Um, and, like, so far in the, like, history of Cursor, we have, like, found ways to, like, going around this. The, like, the agentic code owner thing was probably one place where we stepped on each other toes where the code owners were, like, misconfigured.
So we could just, like, instead of having a deterministic thing, can we just pull in the relevant people at relevant time? Um, so, like, something like that is probably gonna happen with other, like, problems that we're gonna surface in the fu-future.
Thank you.
One question about the, the samples of agents. So do we get all the goodies that we get with cloud agents? These video walkthroughs, do we also have them?
I think computer use is the one thing that's, like, still in, uh, early access, I think. We're I think we haven't shipped to GA yet, but it's coming for sure. So this should be, like, completely on parity with the cloud agent.
Yeah.
Can you describe the, the profile of these?
Can you describe the profile of these kind of, like, uh, mix between product managers and engineers that, that take this, this ownership?
Mm-hmm. Yeah. So I guess the archetypes we have, it's like a PM. Um, they talk a lot internally in at Cursor. Like, they talk with, uh, go-to-market, with sales. Um, they talk with engineers, they talk with users. They just product manage and product manage and just keep everything together in a way and also, like, shield, like, engineers, uh, from various things.
Um, and then we have designers. Um, designers work, I would say, like, 50/50 in Figma and code at this point. Like, all of them do code. Uh, all of them, like, do push to production. Um, but it's a lot of, like, exploratory work.
Like, what should like, what does it look like when you have, like, 10 nested sub-agents? Um, and you can't really feel that in Figma. Like, you gotta actually, like, develop and prototype that. Um, and they work, um they work with PMs, and then we have engineers, of course.
Um, but I think Cursor is very fortunate to build, like, an developer product. So developers are building developer product, and it's kind of like they have good taste. They know what good and bad look like. They know, like, what developers want and don't want.
Um, and I think because of that, they can take such, like, ownership, and they can, like, go with a concept and go really, really far. Um, whereas so, like, the PM might be setting more of the business and, uh, like, the overarch overarching, like, direction.
And then the engineers and designers, like, collaborate on, like, what does this actually look like in code, but also, like, how should it feel and how should it look, um, for a developer?
Makes sense. Uh, are there, like, analysts in this mix as well, or is that done by the product managers?
Mm-hmm. Oh, yeah. That's a good yeah. So we have a data data team as well, uh, data scientists and analysts, and they are also working closely with, um, the PMs, of course, and, like, understanding, like, how users are using the product, where the bottlenecks are.
Uh, but also with, like, with engineers and, like, instrumenting the code in theright way and, like, understanding feature flags and why certain users hit certain paths and some don't. Um, so everyone is, like, just working together. Um, and we have, like I think the way we've structured the team is, like, pretty much, um, domain, like, extensibility might be one team.
Um, cloud might, might be one team. Um, and cloud should still be extensible, so then they have to also work together. Um, but we try to, like, keep it, like, um, modularized and not to ship our organization that much.
Thanks.
Cool. I guess one final question if there is one.
So, um, from time to time, I've messed up and started a cloud agent in a wrong repo or something where it just, like, went out on tangent, came back to it an hour later where it was desperately trying to get access to that repo.
Um, are there any ways to catch these agents that just don't provide any value? They just continue doing stuff, but they're not really making progress?
Mm-hmm. Yeah. I think that's, that's on us for sure. Um, over the last year, we have made a lot of improvements to the cloud agents where initially they were like, when they were, like, work, they were extremely useful, but most of the time they weren't.
Um, so, like, again, cloud agents also come from this, like, internal need of us just wanting to, like, run things asynchronously. Um, and because of that, we have also, like, put a lot, a lot of effort into making our own codebase work really well in cloud agents.
So maybe some like, we, like, have to sometimes, like, create new projects and jump into other projects and talk to our customers to understand, like, where do things fall short. And we try to have, like, instrumentation of, like, does the agent run for X amount of hours or minutes, and, like, does it touch any files at all, or, like, is it going in circles and loop detection and these kind of things.
Um, and this is, like, part of the observability I was talking about before. Um, most of that should happen on our side, uh, but there are always gonna be, like, very specific, uh, contextual things where, um, like, if you are the, uh, the codebase owner, need to, like, set up certain things.
Um, but yeah, we're, we're working on that, improving it. And if you have any examples, like, please come to me and I'll try to take a look.
I think the worst was when I started it on a wrong repo and it just, like.
Mm-hmm.
Called out to Slack MCP and tried to get access in 10, 10 different ways, and it failed.
Yeah. Yeah. Yeah. We could make that better.
Good that you're working on it.
Allright. Thanks, everyone, for coming. Um,
I'll be around for the next two days as well, so please grab me if you wanna discuss anything Cursor or anything at all, actually.





