AIAI EngineerJul 20, 2026· 22:32

Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town

Steve Yegge argues that AI-written code will dramatically increase security vulnerabilities unless developers adopt a separate security pass using tools like Snyk and Chainguard. He shares a bank architect's insight that shipping 10x faster with the same defect rate produces a 10x vulnerability surface, made worse by models writing code. Yegge demonstrates the gap by noting Fable's security hardening missed 241 vulnerabilities that Snyk found in his 30-year-old game. He warns of new attack surfaces like slop squatting, where models hallucinate package names that attackers then backfill with malicious versions. Yegge advocates for multiple passes—correctness, then security—and urges incorporating tools into agent workflows. He also cautions that Five Eyes predicts open-source models will autonomously hack production systems within months, and that personal scams using AI-generated voice and video are imminent.

  1. 0:00Intro
  2. 1:08Be Scared
  3. 3:08New Threats
  4. 4:47Bug Half-life
  5. 6:08Tools
  6. 11:39Five Eyes
  7. 13:32Family Scams
  8. 14:38Take Action
  9. 15:31Surprises
  10. 17:23Gastown
  11. 19:02Credentials
  12. 20:47Injection

Powered by PodHood

Transcript

Intro0:00

Steve Yegge0:13

Alright, hey everybody. Uh, yeah, yep, yep, yep. Okay, I have about 18 minutes. I have a slide deck here that Claude put together for me, but Claude still can't do very good slide decks, so I apologize for that.

Just, just, just focus on me,right? Right here. This is like the most fun thing to do ever now. If I start, like, running out of time, can you, like, call time for me? Alright, perfect. So, I'm Steve Yegge.

I am here on behalf of Snyk today. I am not getting paid for this talk or anything like that. They were really nice enough to buy me a ticket to the conference, which is cool, but I'm mostly here because I wanted to hang out with all of you.

And I don't bite, and feel free to come and say hi afterwards. I wandered the halls, I may have said all the brand logos. I'm up here because,

Be Scared1:08

Steve Yegge1:08

like, the title of my talk is, like, agentic security, but the real title of my talk is, be scared.

Come on. Come on, level with me. Who here is scared? That's good. But you're all in the security talks. Of course you are. The problem is they aren't out there. Right? I went to Commonwealth, I went to so many places last year, but one of these big banks I was at, I was talking to their chief security architect in December.

And I was doing a Q&A and talking about vibe coding, and everyone was, like, asking me questions. And I was, you know, knocking him out of the park. And he stands up real quiet at the end and he goes, "Yeah, so."

He goes, "If everyone's shipping code at the same, sorry, at 10 times faster, and the defect rate stays the same, the security defect,right, the vulnerability rate,

then doesn't that mean that the defect surface goes up by 10x?" And I, it hit me so hard. I sank down to my knees. I was just like, "What are we going to do about this?" Because it's a really, really important point.

Because the subtle implied question is not, "If the defect rate stays the same, the defect rate's going to get worse. A lot worse, with AIs writing the code." And so, I didn't have an answer for him. I actually, I have a partial answer for you today, and I'll share it with you, but it's only a partial answer.

The real answer is, you have to be scared of what's coming. Alright? And it's because, it's, it's, it's, do I have a slide about this? Yeah. So it's not just that you're just, like, putting in more of the same defects, cross-site scripting, and blah, blah, blah.

New Threats3:08

Steve Yegge3:08

You're still putting those in. Fable wrote an XSS vulnerability during the short time that I had with it. Not its fault. We'll talk about it in a minute. But those are the old ones. Alright, we are, we all know how to fix those.

There's new vulnerability types and new attack surfaces coming, and they're here, and many of them are incredibly well polished. Like, what's an example? You guys know about slop squatting,right, where the AI hallucinates a package name. Let's say that you want a graph database, and so you're like, "Okay, I'm going to use this graph database," and the AI goes, "Oh yeah, I know, it's, you know, it's graphy 123," and it goes off to the package manager and downloads graphy 123.

And it builds, and it runs, and the tests pass, and it looksright, but what it downloaded was a backdoor. Because that graphy 123 wasn't a real package. But somebody noticed that the LLMs are hallucinating its name, and they uploaded one that does exactly the same thing as the one it thought it was getting, plus a vulnerability.

Bonus. Yeah? Is that scary? Yeah, it should be. It should be. How do you even detect that that's happening? So, so,

we're entering a world where everything you write, every bit of code that you generate, is going to have to get far more security scrutiny than it's ever had before. Okay? It's just, you're not, it's hard for me to convey how scared I am.

Bug Half-life4:47

Steve Yegge4:47

Alright? But let's start with where it gets generated. Now, when I worked at Google, I worked really close with the TAP team. They did tests. The test automation platform. Yeah? And they, uh, so they ran all of the unit tests and integration tests at Google.

They had a massive fleet, you know. And they learned stuff about how bugs work. And how bugs work is they have a life, they have, like, a life cycle where if you see the bug,right away, you'll fix it.

And the longer the time goes for when you see a warning or some sort of issue, the longer that passes, okay, it's got this sort of half-life ofurgency, and all of a sudden, eh, it ain't really biting anyone anymore.

Okay? This is how we handle all of our bugs. At Google, they recognized that this is such a human nature phenomenon that they worked really hard to move the reporting of bugs as you were typing. Because that's when you're most likely to fix it.

If you make a bug and it goes, "By the way, there's a divide by zero here," or "There's a backdoor," or "Vulnerability," or whatever, you'll fix itright there. But if it gets to code review time, you're like, "Is it really worth it?"

Right? And the problem, folks, is that that works for all classes of bugs except for security. There's no half-life on it biting you. It's not like, "Oh, because users aren't getting bothered by this security vulnerability, that it's not a problem over time."

Tools6:08

Steve Yegge6:16

The problem compounds over time. Yeah? So you have to treat this class of vulnerabilities the way that Google treated their top vulnerabilities at the time, which is to surface them at the developer's fingertips. What if the developer doesn't have any fingers?

I don't know, how many fingers do LLMs have?

Then you need to surface it to the LLMs. But wait, you say, "Wait, wait, wait, wait, wait, wait." Fable's really smart. Or at least it seemed that way for the two days I got to use it. Can't Fable just write secure code?

Right? I mean, come on. Come on. You all know secure, I mean, you're all here in this room,right? You all know security is an arms race. One that never ends, one that's going up exponentially with Moore's law, one that's going to get real uncomfortable when quantum comes along.

Thank goodness that's, like, five to seven years away. My buddy says 15, so maybe somewhere in between. But in the meantime,right, LLMs are a real problem. Yeah. So how do you surface, how do you surface the vulnerabilities? Well, first, I wanted to figure out how to do it myself.

I have a game I've been working on for 30 years. I just had Fable do a security hardening pass during the time I had it. Went through, and it did all my cloud hardening, and it found a bunch of credentials, and did a bunch of stuff.

And it started giving me these vibes like, "Yep, yep, hardening pass, looking pretty good." So I ran Snyk,right? And I, I don't have the numbers here. Uh, yeah, I didn't, I didn't include the numbers because I'm done. But, um, it found 241 vulnerabilities,right?

Just a ton that Fable hadn't even thought to look for,right? And it's because, look, so I've told people about the rule of five. When you do things with LLMs, often you have to get them to do up to four to five reviews of the work that they did before it's, like, actually ready to ship.

And it's because their cognitive process is very similar to ours, and it goes through a draft, and then a revision, and then polish and editing until it's, you know, it's finally ready to go. It's like painting a wall.

Some things you don't just do all in one pass. You do them in multiple passes. Right? So, um,

security is one of the, so what I found, I wrote a book on vibe coding last year. I did more vibe coding, I think, than anyone, you know, two years ago. And what I found was

that they're really good at doing one thing at a time. Even the really good models like Fable,right? Just because of this multi-pass, painting a wall phenomenon. You got to give them one task at a time, which means you can't give them security at the same time as you give them correctness.

They'll do a half-assed job of both. And you don't want a half-assed job of either of those, it turns out. So you do it in two passes. Right? Now, five, it's been five months now, but five months ago I wrote an essay called Software Survival 3.0, where I talked about what software has to do to survive when LLMs can synthesize it all.

I don't know if any of you all saw that, but the basic, the basic gist of it is that LLMs can synthesize any software that they want, but they're very lazy in a good way,right? Lazy in, like, they don't want to spend tokens if they don't have to, because that's money and power and bad for the planet and bad for your wallet and so on.

Right? And so they use tools to help them whenever it can save tokens. Right?

So tying it all together, if the LLMs are doing the coding and they're happy to use tools to help them offload cognition, you see where this is going? Give them Snyk. Give them Chainguard. And I still think there's a missing piece in this picture that I'll tell you about at the end.

I don't know if you all know about Chainguard. Chainguard, uh, Chainguard is a supply chain that you sign up for, and they give you images that have been pre-vetted to not have vulnerabilities, and they update them. So it's your inputs.

Okay? And then Snyk handles everything else. The code that you write, the code that the LLM writes, the dependencies that you're pulling in from slop squatting,right? The innocent stuff. It can find, my understanding is that Snyk can actually find vulnerabilities that are proprietary, that only they know about because they're ahead of the CVE registry.

I've, there's some truth to that. When I ran them on my code base, it didn't find any vulnerabilities that weren't already public CVEs, but damn, it was easy to use. Right?

I think that a tool like Snyk is going to give your LLMs superpowers. Okay? Because what you do is you add it as a pass to the prompt that you give them for whatever they're doing, and say, "One last thing to look at," and have them run your security analysis, all of the tools, get the open source ones, get the Snyk one, get the Chainguard one, get the, all of them, and have them check each other's work too,right?

Five Eyes11:39

Steve Yegge11:39

If you want to get really serious about this on launch time. Right? But secure your supply chain. Because Five Eyes is warning us. Yeah? My poor, my poor game. By the way, please, please, I just told you that my game has 241 vulnerabilities.

Don't go hack my game. Give me a couple days to fix the bugs. Yeah? We all good here? Good. Alright. But if you do hack it, you're going to bother, like, five players. Alright? Okay. They're very loyal. Look, Five Eyes, which is, like, a bunch of,right, governing, it's big countries that have their eye on the cybersecurity, you know, landscape.

They just announced that it is now months, not years, until it starts happening. It. You all know what it is,right? It is when open source models catch up to mythos. Does anybody here believe open source models are going to catch up to mythos?

Interesting that it's about 50-50. Anyone got a timeframe in mind?

Who said December? That's pretty accurate.

Guest12:45

There was a test yesterday.

Steve Yegge12:48

Cheater. Yeah, it's about seven months. So, actually, it's shrinking, so it's probably about six months now. Yeah. And mythos is real, real good at hacking your systems.

And, and just, just remember, you can't trust it to automatically write good code any more than you can trust it to write elegant code by default. That's a separate concern. It's a separate pass. You can't expect it to write performant code by default.

That's another pass. You see what I'm saying? You can't expect it to necessarily write the code according to your company coding standards. Okay? These are all passes that go through your code, and I just want you to remember that security should be your first one and your last one.

Okay? Give it extra. Okay, and the last thing I wanted to talk to you about, first of all, go do all this, and second of all, the last thing is really truthfully. Okay? Dial it in here, folks. Go to your families offline, like, in person, and get your, your code words refreshed.

Family Scams13:32

Steve Yegge13:55

Because another kind of scam that's coming along is you get a call from a family member who's in distress and they need money, and it's very convincing, and there's a video of them, and you're going to need a way to distinguish them from, from AI.

Okay? It's months away, and some of your families are going to be slow to catch on to this stuff. But bank accounts will be drained. I heard that Congress was given secret demos of draining bank accounts. I've been scared of this for close to two years.

I heard one talk from a security researcher almost two years ago at ETLS Las Vegas, and he stood up in front of the crowd and he said, "You're all not scared enough of what's coming. It'll affect you personally, not just your company."

Okay? So that's my message to you. It's not a message of hope and positivity today. But it is a, it is a message that, that should be clear, crystal clear, is that there are tools, open source tools, free tools, commercial tools, okay?

Take Action14:38

Steve Yegge14:54

Techniques, practices, okay, that you can useright now to get started on fighting in this arms race and protecting yourself. And that's all I've got today. Thank you.

Guest15:12

We do have time for a couple questions.

Steve Yegge15:14

Questions?

Are we ready to go?

Guest 215:31

Sweet. Just super simple. What has surprised you recently in the world of AI coding?

Surprises15:31

Steve Yegge15:38

What has surprised me in the world of AI coding? Well, I'm not really super representative. I spent a lot of my time trying to predict the future by, like, hammering agents really, really hard. Yeah? So,

um, you know, one of the surprises, and I shouldn't have been surprised, but one of the surprises is that AI is moving faster than the world is moving.

Tech is moving faster than society can move. And the surprises show up when friends, smart friends, resist, you know, the inevitability of AI, and they call it psychosis, or they, or they poo-poo it, and they say, "Well, it'll never be actually smart," or whatever.

They can't see the curve. Right? And, and that surprises me. Maybe it shouldn't. It shows a sort of tunnel vision. I think people have a tendency to look about three months back and about three months forward and be like, "Oh, it looks pretty flat."

Right? But, but, and so that, that surprises me, that people aren't, honestly, that people aren't more scared, and that, and by the same, by the flip side, that people aren't more excited by it. Right? Because you know, once you actually, you know, once, once you get it, I mean, you don't even want to be here.

How many of you are running cloud coderight now? Most of you. Right? It's really fun. So, you know, I mean, like, that's a surprise too, that the world is pushing back so hard on that. Right? We're in an awkward phase.

We'll get through it. Any other questions? Whoa. Alright. Well, you pick.

Feel free to bail also. You don't have to stay.

Guest 217:23

Hi. Big fan. What's the, like, most impressive thing you've seen Gastown do? And, like, how much human intervention was involved or steering?

Gastown17:23

Steve Yegge17:31

Oh, Gastown? Yeah. Gastown. Gastown was a lot of fun in January.

Gastown is a beads machine, and I still use beads, and I'm working, I want to, I was talking to Angie, I want to donate beads to the Agentic Foundation, you know? We're going to, we're going to put multiple backends on it.

Beads is a task tracker. Right? Beads is how you do Boris Cherny loops. You know how Boris is like, "You shouldn't be prompting your agent"? Has anybody here actually, like, successfully, how often do you get Claude to actually run all night for you?

Like, for real, run all night? A few of you. Right? Like, this is the next frontier, I think, of actually getting agents to run, like, for a long, long time unsupervised. You can do it with beads by queuing up enough work and having them claim and all that.

And there are some other, some other systems that'll do that. It was really fun when Gastown did this for me automatically once. I filed a whole bunch of beads and they disappeared, and I was like, "Oh, no, another bug.

My beads disappeared." And what had actually happened was that one of the agents just found them and just implemented everything. Right? It's like, "Whoa. I really like swarms now." Yeah. Fun times. Does anybody here regularly work with more than 10 coding agents at once?

You see, like, not many. Right? The world is still in the, we're still kind of, like, prompting and using a few here and there. Right? It's going to accelerate really fast next year. Other questions? You, you just say it.

Credentials19:02

Guest19:02

Yeah. Your talk focused a lot on vulnerabilities from agents writing code. I'm curious, what are your thoughts about agents taking action? So having access to credentials or looking for it.

Steve Yegge19:14

Yeah. So that was the third dimension that I really wanted to talk about. It's just I don't really have time, but, like, my friends over at Tessel, I'm advising them, they're actually doing this. Right? There's just this whole space of who's looking over, who's looking over your agent's shoulders?

I had this conversation just now. Like, I tell everyone, everyone's just starting to stand up agents, like, 24/7. Like, processing queues, responding to events, like, agents that actually do stuff. Right? 24/7. And I, I, I encourage people to think adversarially.

Think of adversarial groups of agents tasked with doing that queue management. Because one agent will always eventually screw it up. Right? So you got to have those supervisors. And so there's whole systems emerging here. Kind of go out and go look at all of your things and say, and start to, like, you know, do, do, do that hardening stuff.

Like, do they really need all those credentials on that service account? Really? Right? Only for this one action, maybe we can, like, separate this one out. That kind of thing. Right? This is a brand new frontier, but it's one, ironically, that even though there's kind of almost nothing out there, there's some experimental stuff, you still have to be thinking about itright now and designing a solution in-houseright now.

Right? Because otherwise, your engineers are going to spin up, or your non-engineers are going to spin up a bunch of agents with way too many permissions, and then,right, as soon as a bear munches into the igloo, everyone's dead.

Right? That's the old security analogy. I don't know if they still use that one anymore. Yes.

Injection20:47

Guest 220:47

Hey, Steve. Nice to see you. I'm curious, in terms of, like, just the best practices that you've seen that you use personally, or maybe you haven't tried, but particularly with regards to prompt injection. So.

Steve Yegge21:01

Yeah. So I was supposed to talk that during my,right, during, during my speech here. I wanted to mention it. There are a whole bunch of attacks happening on the training and on the prompting side. Right? So on training and inference.

And so the bad guys find ways to sneak in stuff. Right? And it's like, like, like, the simplest version is the new XSRF, where, like, the user puts in some, some text, and then some bad actor puts in some extra text saying, "Disregard everything and do the following."

Right? And then they just get more sophisticated from there.

I mean, I don't have any good answers for you other than, like, this is real. It's kind of an education problem at this point. You need to get everybody thinking about it. Right? And then I feel like there are, like, new security roles about ready to emerge inside of companies, agentic security, that it's kind of an extension of what they're already doing.

Who's already doing this? Right? Yeah. So you've already got people that are going out and looking after the sort of security of your agents that are in, that are deployed in the. Yeah. So you're way ahead of everyone.

I'd love to come talk to you later. This is all brand new stuff. Yeah? Cool.