AIAI EngineerJul 30, 2025· 17:33

How we hacked YC Spring 2025 batch’s AI agents — Rene Brandel, Casco

Rene Brandel, CEO of Casco, explains how his team hacked seven of sixteen YC Spring 2025 batch AI agents within 30 minutes each, revealing three critical security flaws: cross-user data access via IDOR, arbitrary code execution through code tools, and server-side request forgery (SSRF) from tool endpoints. They extracted personal data by exploiting missing authorization checks, overwrote security controls by writing malicious files into code sandboxes, and stole Git credentials by manipulating a database-schema tool. Brandel emphasizes that agent security extends beyond LLM prompt injection, advises treating agents as users with proper authentication and authorization, and warns against rolling custom code sandboxes, recommending out-of-the-box solutions like E2B. The talk concludes with a Q&A on extracting system prompts and the dangers of local coding agents.

Transcript

Intro0:00

Rene Brandel0:16

So, yeah, who's ready to hack some agents, yeah? Oh, wow. Allright. So, let me first introduce myself a little bit. I'm Rene, I'm the CEO of Casco. We're a YC company, and we specialize in red teaming AI agents and apps.

And so we spend, uh, I spend my previous time at AWS working on AI agents, but I've always really loved working on AI. In fact, there's a video of me 10 years ago building voice-to-code, and I won Europe's largest hackathon by doing that.

And so I would talk to it, say, build me a blog post, and it would generate the site. And it was actually, it was kind of fun. Like, it did, uh, things like, um, yeah, load in pictures from San Francisco.

And you can see how horribly slow the APIs were back then. And I'm going to, about to give you a nightmare by showing you the architecture diagram of that thing. Um, but yeah, it kind of did the job.

And this was like 10 years ago. Obviously, back then was no generative AI, and these things were extremely difficult to do. Um, but it is, it really gave me a glimpse of what the future could look like, even back then, as technology gets better,right?

So obviously, many things have changed. Two months ago, I quit AWS and worked out of, uh, the garage with my co-founder. And, uh, we got into Y Combinator. So, yay, that's awesome. And so from there, we also looked into how else have things evolved.

Well, this was my, um, architecture diagram from back then. You could see there were three different cloud providers, including IBM Watson, which was like forefront at the time. That is, that's true. And, uh, before it was like, uh, Microsoft Lewis, which was like some natural language understanding things.

Agent Stacks1:33

Rene Brandel1:49

You can see it was just a lot of, like, piecing things together, and that was already kind of difficult to do. But nowadays, we see the stacks normalize significantly more,right? I think this is probably what the average agent stack looks like these days.

You've got some server front end, you talk to an API server that talks to an LLM, connects up with tools, and then you have a bunch of data sources associated to it. So this kind of normalization of agent stacks is actually really good.

It like makes many things easier, definitely better than my hackathon project 10 years ago. Um, but we need to think about the security posture around these systems. And my general impression over the last, uh, last few years is like primary discussions around LLM security really like, hey, is it, um, can you do prompt injection, can you get it to do harmful content, um, which is all really important.

But the reality with security is you need to look at all the different errors in your system. And that is typically where real damage happens,right? And so this is really agent security, and that is what I want to talk about today.

Why We Hacked2:55

Rene Brandel2:55

Now, one thing is like, why did we even hack a bunch of agents? That's kind of a weird thing to do. Um, the answer is, quite frankly, you know, we wanted to launch internally at Y Combinator, and we wanted a splashy headline.

And so we're like, uh-oh, what do we do? And fun fact, we have the second highest upvoted launch post inside Y Combinator of all time. So higher than Rippling. Yes. Okay. Um, so, uh, we, we did, we did basically this approach.

At a time, we were looking at, oh, which agents are already live? And then let's just set a timer for 30 minutes. We don't want to waste too much time on this. And then, you know, let's, let's figure out what their system prompts are and just kind of understand how they're working.

And I, I have a feeling when I was creating this meme that this could be true, but it turns out it is true. And then we looked at, oh, what kind of tool definitions do they have,right? Like, you know, what is it supposed to do?

Is it supposed to access data, supposed to run code,right? And then we just, uh, tried to exploit them and see what's, what's going on. Uh, and it was really fun because we hacked, uh, out of 16 agents that were launched, within 30 minutes each, we were hacked, uh, we hacked, we hacked seven of them.

And there's three common issues we see across all of these ones. So I hope that we will all learn today what the most common issues are so you don't make the same mistakes. And also, this is going to be the best investment if you're a VC, this batch, because they're all secure now.

Data Access4:03

Rene Brandel4:17

So, first issue, cross-user data access. I mean, you guys were just here at the OAuth talk. You know where this is going to head into,right? Um, so we first leaked this company's, uh, system prompt, and we saw, huh, it has a bunch of interesting tools attached to it, including looking up user info by ID, suspicious, uh, document by ID, and a bunch of other things.

And then, you know, like, when you see this, you just want to, like, oh, yeah, there's this thing called IDOR, like Insecure Direct Object Reference. It's basically when you make a request and you validate that, hey, the token is valid, and then you just let your request through,right?

And you're kind of betting on the fact that the ID cannot be guessed. Well, that's obviously not good. Um, so, yeah, we looked up a product demo video that they recorded, and we found the user ID in the URL bar and just, like, tried to plug it in.

Uh, this is a different ID, by the way. Don't worry, guys. This is my co-founder's ID now. And, uh, yeah, we were able to find their personal information, including their email, nickname, whatever. Um, but it gets better because these things are also interconnected.

So you had not only the user ID, but you also had, like, oh, the chat ID, oh, and their document ID. And then these things ultimately linked up together and allowed you to traverse the entire system,right? It's not good.

So what's the fix for that? There was a really comprehensive talk literallyright before this. Sorry for the folks that missed it, but this is the basic fix for it,right? You need to think about how do you authenticate, but also authorize your request.

It's really two checks,right? Make sure your, your token is valid. Good job, team. Yeah, you got that. And then the second thing is, like, this is what we see in this super base era with role-level security. Just make sure that you have some sort of access control matrix somewhere that checks that it matches up with whoever is making the request.

Okay? Super, super important. Authenticate and authorize. Now, you can see this was actually, you know, an issue that was kind of there,right? It's, it's not like around the LLM and the API server. It's really what is happening downstream.

And, um, yeah, there's a lot of arrows in this diagram, and we're going to look at all of them. So the next thing is to remember, as you're thinking about these tools and how you're building it, like, agents actually act like users, um, not API servers.

When we were, like, debugging this issue, like, we actually asked a bunch of Y Combinator companies, like, why did you build it this way? Because clearly they can build a web app properly,right? But it's just like, I think as developers, we have this natural pattern matching in our head.

It's like, oh, yeah, this thing runs on a server, so it should be like a service. And then I'm going to give it service-level permissions. But actually, agents are like users,right? So everything that applies to users applies to agents too.

So make sure that, you know, your LLM should probably not determine your authorization pattern. That, that, that's bad. That's a red flag. Uh, second thing is it should probably not act with service-level permission. Listen to your previous talk on OAuth.

That's great. Um, and then just like users, you should make sure you, uh, don't just accept any input. You should sanitize them. Same with outputs,right? A lot of these are like the traditional web application security things that you just need to, like, really, really internalize for this new world.

Code Execution7:32

Rene Brandel7:32

Now, that was interesting. And so the second one was even better. Um, so this is not as common, but the damage is bigger. So it's what in pattern we see. So there are a lot of code tools that agents use.

And there's a, there's a, there's an Anthropic paper here. It basically talks about what's the distribution of which industry and how much do they use Claude. And there's like this one outlier here. I'll zoom it in for you.

Um, yeah, so us nerds, we make up 3.4% of the world, but we're 37% of Claude's usage. Ooh, why is that? Because we love computers and we love coding,right? And so we found immediately the value of it. But it's not just us that use agents with coding tools.

In fact, many agents create code on demand to do some things,right? Like some agents just generate a calculator on demand to make a calculation,right? And so there's a lot of these code execution sandboxes out there that are interesting.

And so if you, if you think about that, there's actually a critical path in your system because you've got a tool that talks to another container. A container is arbitrary compute. And when you have arbitrary compute, many things can happen.

Many bad things, many good things,right? But let's talk about the bad things today. So we did the same script, did the system prompt. Again, the system prompt itself, great. I mean, it doesn't cause any damage. But as an attacker, you always think about the fact, uh, the things that are like, huh, that's kind of suspicious,right?

It's like, oh, wait, it, it, it runs code and never outputted it to the user. Okay, let's output it to the user. Oh, yeah. And, and most, mostly run, run it mostly at most once. Let's run it all the time.

And so you try to basically invert what the system prompt is saying because that is exactly what the developer didn't want you to do. And that is how bad actors think,right? So we figured out, oh, this thing does have a code tool.

And so, you know, we tried, we tried running something. It's like, oh, it only allows me to write Python. And, you know, I love JavaScript. And, um, yeah, it doesn't allow me to run these really dangerous, you know, function calls.

Ugh. Okay. And it restricts, like, which Python files to run. That's also not good. So, yeah, but we looked at what it could do. And it had two kind of innocent permissions. Write a Python file and read some files.

You can do a lot with that. This is great because what if we just looked around the file system now,right? We can read files. So we looked at, okay, build me a little tree functionality and, you know, return me the entire file system tree to see what's going on.

Oh my God, there is an app.py file. That's probably important. Um, and then we looked at, oh, it has two endpoints, write file and execute file. Ah, okay. These endpoints are hidden behind a VPC, so we cannot hit it directly.

That's okay. Um, but, huh, we can write files. Huh, we can write files. There's an app.py file. Huh, let's look into that. Oh, wait. That's where all the protections are for their code. Uh, and so we can just overwrite the app.py file with empty strings around all the security checks.

And whoopsie, we got in. So now we can Bitcoin mine all day. That's great,right? Yeah. No, it gets much worse. So the thing with arbitrary code execution, once you're inside a container, is that you can do many things.

Like, um, there's this thing called service endpoint discovery, metadata discovery. You all heard of that? No? Okay. Basically, it allows you to discover what are other devices on the, uh, what are other devices on the network? What other resources are there on the network?

And, uh, you can also just, you know, fetch the user token. Uh, sorry, the service token. You know, just see what's going on. What's the project name? Yeah. You know, and you start looking around. It's like, oh, okay.

Yeah. Okay. I, I, I can also fetch a scopes. So I can use do many things with this token. That's awesome. Um, who has really, really spent time configuring service-level tokens and their permissions in a granular manner and does it all the time and never forgets to set something wrong?

Okay. One guy. One guy there. Okay. Whoopsie. We have access to all their customer data. So that's, uh, and we just query BigQuery, which has a great interface for that. Isn't that cool? Yeah. So, yeah, making sure you have code sandboxes correctly is very hard because you can move laterally across the infrastructure.

And that is just very, very dangerous. Okay? And so kind of like, don't roll your off in the web world. Don't roll your own code sandboxes, please. Like, it's, it's just very hard. It's very, very hard. And so use a out-of-the-box solution.

There are many of them. E2B is, I think, a very popular one. Some, some folks probably heard of it. Uh, there's one in our YC batch that I personally just genuinely really love. They have observability built in. They boot up super quickly.

And what I love about them is they have an MCP server that's just easy to plug into,right? So just easier for your agents to work with. So please do that. Don't do, you know, your own Python, app.py thing.

Um, it's not good. Trust me. Um, so that leads into a third part of an attack vector around server-side request forgery. It's a very long word and it really bugs me that the SSRF didn't fit on the previous line.

SSRF12:28

Rene Brandel12:43

This really triggers me. Um, yeah, I know. So, um, this is what happens when you can kind of co can kind of get a tool to call another endpoint that you didn't, you know, that the service itself didn't intend you to call.

And you can pull out a lot of information just through that workflow. So let me give you an example. So this is exactly extracted system prompt. Great. Oh, this thing can create databases. That sounds exciting. Um, and then you look into it, it's like, huh, it pulls a database schema from a private GitHub repository.

Isn't that great? That means whatever request goes to that private GitHub repository must have the Git credentials,right? Otherwise, how can it pull that from a private repository? So, um, yeah, and it's just a string. So I guess I can just put in whatever string I want and coerce it into providing that.

So let's set up a bad actor.com test.git repo and just see what credentials come through. And yep, it comes across with the Git credentials. And so now you can actually take those Git credentials and just download your entire code base that was behind a private repo.

Isn't that crazy? Isn't that crazy? Yeah. This is, I mean, it's awesome for me to do this,right? It's like you, you get paid to do this. Oh, come on. It's amazing. Now, um, we told our batch mates immediately and they told us, "Don't worry, bro.

It's already fixed. It's okay, guys." That, that company's secure if you're VC listening in. Um, so, so, but with that though, it is really important to think about the implications of what your system is doing,right? I, I love vibe coding, not going to lie, but like, you got to really think about where all these arrows are and if you've configured those things correctly.

So with that, always sanitize your inputs and outputs. This could be like a web dev conference from 20 years ago. Um, but, but it applies to agents too,right? Like, we just need to make sure we keep those good security practices that have, that we have learned to love, hopefully, over the years to take it forward to a new technology paradigm.

Key Takeaways14:47

Rene Brandel14:47

And then ultimately, I want you to take away three things. So first thing is agent security is bigger than just LLM security. Make sure you understand how these threat vectors apply inside your overall system. Second thing is treat agents as users.

And that applies to authentication, to sanitization of user inputs, and many of the other things. And last thing, definitely don't roll your own code sandbox. That is just so dangerous. And, you know, it, it, it very quickly turns from like an intern project into like a nightmare.

So be very, very careful with that. And these are the most basic ones that we've seen come across,right? There's obviously many more security issues. And if you don't know exactly how your agent security posture is, you can go to casco.com.

Casco's Solution15:28

Rene Brandel15:32

You can book a demo with us. We build an AI agent that actively attacks other AI agents and tells you where they break. Isn't that great? Um, and yeah, feel free to connect with me on LinkedIn or on Twitter.

And I have, uh, every now and then some good stuff to post. Yeah.

Q&A15:51

Guest15:51

Awesome. Thanks, Rene. Does anyone have any questions? We can have time for like one or two quick questions if you're in, if you're game for it. Sure.

Guest 215:59

How do you look at the system prompts?

Rene Brandel16:01

Um, how do I look at system prompts? There's a lot of just like open techniques. The, the best one that I've seen is, uh, from hiddenlayer.com. Have you guys checked out those guys out? They have a great blog post on like, um, the, it's a policy puppetry attack.

Yeah. It's great.

Guest16:17

Very cool.

Rene Brandel16:17

Oh, awesome. Oh.

Guest16:19

Yeah.

Guest 216:20

If you're messing with coding agents, like how do you make sure because a coding agent can run commands, like how do you make sure that it's actually not running like running the proper commands? Because this is a super tough thing to do.

Like if you try to whitelist them, like there's so many creative ways that LLMs can get around them. But like how, how are you.

Rene Brandel16:39

Yeah. Are, are you talking about it locally or server-side?

Guest 216:42

Um, I know.

Rene Brandel16:43

You're on both.

Guest 216:43

Yeah. I mean, locally is even more dangerous because they have the credentials of the user running this program.

Rene Brandel16:49

Yeah. No, very, very much so. So locally, uh, I thinkright now the industry is either you go full YOLO mode or you ask every time,right? Um, I mean, I'm not joking. Cursor's thing is called YOLO mode,right? Um, and then on the server side, use a code sandbox because ultimately they have constraints, uh, around the internal networks, but also they have constraints around, um, how long they can live as a sandbox.

Yeah.

Guest 217:12

Okay. So sandboxes that use VMs actually.

Rene Brandel17:15

Um, yeah. So they, they typically use something called firecracker under the hood, which is, um, better isolation layer. Yeah. If you just use containers, by the way, that's not an isolation layer in case anybody's wondering.

Guest 217:24

That's fine.

Rene Brandel17:24

Yeah. Yeah. Don't use containers for isolation. Yeah.