AIAI EngineerJul 21, 2026· 19:47

2026 State of AI Engineering — Barr Yaron, Amplify Partners

In the 2026 State of AI Engineering survey presented by Amplify Partners' Barr Yaron, 1,048 AI engineers reveal that cost is now a first-class engineering constraint—40% say it regularly shapes how ambitiously they use AI. Agents have exploded: 95% of teams now use agents, and 89% of those agents have write access, tripling from last year. Image generation adoption doubled to 36%, while audio shows the strongest intent-to-adopt at 56%. Open-weight models augment rather than replace closed models—45% use open-weight, but over 90% of them also use closed models. Evals remain the top infrastructure challenge, and inference is the most bought layer, while prompt management (61% built in-house) stays close to product logic. Teams report 97% net positive impact, but 59% fear long-term liabilities from AI code, and over a third say non-developers now ship features.

Transcript

Intro0:00

Barr Yaron0:13

Now joining us on stage is the partner at Amplify, Barr Yaron.

Fantastic. You did a great job practicing. I feel very, very loved. Let's get started. So, like you just heard, my name is Barr. I run a survey every year on the state of AI engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides.

Just in the past week, we've had frontier releases treated like national security events. Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen. So if I miss a major announcement while I'm up here, please come find me after.

But that's exactly why we run the survey every year: to cut through the noise, take a moment, step back, and understand what AI engineers are actually doing. For the first time this year, we were thrilled to partner with Notion and Vercel to run this survey.

Very quickly on me. This is the least interesting slide. I'm an investment partner at Amplify, very lucky to invest in companies built by and for AI engineers. And I'll make the same promise that I make every single year, which is: short time on bar, long time on bar charts.

So let's getright into it with lots of bar charts.

First, let's talk about—well, maybe raise your hand. Did you fill out the survey? This is a very large group. Okay, yes, I see you in the front.

If the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because 1,000 of you gave your time, so thank you. We had 1,048 respondents this year, which is a lot of AI engineers.

Workforce2:26

Barr Yaron2:27

And to be precise, this is not just AI engineers, as I'm sure you see at the conference. Every year we see that AI engineering is more of a discipline than a job title. It touches founders, CTOs, engineers, product people, folks across company sizes and experience levels.

And that range shows up in experience too. For the third year running, we see the same pattern, which is skewed toward senior engineers but newer to AI. Of those with over 10 years of software experience, over half have 3 years or less of AI experience, which tracks.

These are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started engineering, the median new engineer has nearly as much AI experience as the median 10-year software veteran. So the newest engineers have never known software without this.

But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question, which is: when people say they're doing AI at work, what are they actually doing?

Modalities3:22

Barr Yaron3:37

So first up, I'd like to start with the modalities. We asked: which modalities are you actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause. But one piece of this chart that I always find very interesting and I always look at is the ratio of, nope, I'm not using this modality to, I'm not using it, but I do plan to.

I call this the intent-to-adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent-to-adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build.

And this is not a brand-new signal. Last year, audio also had the highest intent-to-adopt across modalities, but 37%. So audio continues to take the lead and have high interest, but that interest is accelerating. Now, there has been an audio swing, but if we look at what changed most from the last year in the survey, the biggest jump is actually in people using image generation.

The share of respondents using generative AI for images and feeling really good about it doubled, from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year plus survey time, we've had models: Nano Banana, Nano Banana 2, ChatGPT Images 2.0.

The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent-to-adopt, but image generation shows us what happens when a modality crosses that threshold.

So I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year.

Now, models. Who here spends time on Twitter? Allright, yes. I imagine this is a very Twitter-pilled crowd. If you spend any time on Twitter, in this circle, you've seen a lot written about open-weight models these past few months.

Models5:35

Barr Yaron5:50

And I think we'll see it even more in the next year. So we asked: what models are you actually using in production? 94% use closed models. 45% are using open-weight models. But here's the thing: open-weight models are not replacing closed models for the most part, at least not yet.

The respondents using open-weight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching.

We also asked, just to double-click on this, for the top three considerations when choosing a model. If you're choosing a model, what is important to you? And despite the airtime of the open versus closed, it's not what drives model choice.

It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward: it's quality. Quality dominates, followed by agentic capabilities like tool calling. And cost tiedright with it. Money, money, money. We'll get back to that.

One thing that I found very interesting is that reliability is not near the top. Only 1 in 5 named reliability. That doesn't mean teams stopped caring about reliability. There are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement, and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances: to quality, capability, cost.

But we could talk after. Allright, so here's where the model story all comes together. Like I said, teams are not choosing one model and calling it a day. Earlier, I showed that 87% of teams are using more than one model.

That's the opposite of standardization. And the way that they choose models for given tasks varies. Most popular is routing by task type. Some run multiple models to compare outputs. Some route based on cost. But models are good at different things.

What was interesting was that more than half of respondents said that their organization is starting to standardize on fewer AI tools. They're trading flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others.

But the headline here is that we're in the early, great standardization of the platform and tools, not the models. Allright. This is the slide where anyone who's opened an AI bill in the last year starts nodding. So it turns out that infinite intelligence still comes with a usage-based bill.

Cost Constraints8:20

Barr Yaron8:32

Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first-class engineering constraint. We see this in the data. 40% of respondents say that cost regularly shapes how ambitiously they use AI.

And another 36% say that it sometimes does. Well, this is pretty straightforward. So all in about 3 out of 4 respondents are adjusting their AI usage based on cost, and maybe the fourth has a company card.

That might be surprising, or maybe it's obvious, but 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing. But cost is becoming a real big part of the product decision today.

And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLAright under quality itself.

Agents9:36

Barr Yaron9:36

Which brings us to the biggest line item of them all: agents. We've been talking about agents for a while. This year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering, their escaping demo world.

So we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents.

This year, 95%—this seems high to me—95% say they're using agents roughly double last year. Second, amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data.

This year, that number is 89%. So when you combine these two shifts, more teams using agents, and more of those agents having write permissions, the share of all the respondents—and again, it's a survey—using write-enabled agents is up more than three times relative to last year.

So this is really the big shift. Agents are no longer reading, summarizing, drafting. They're taking actions inside of systems. And that raises the obvious question: how are we controlling all of this?

With pretty blunt instruments, there are many ways that folks are controlling agents today. The top two are human-in-the-loop approvals and gating permissions, which are theright instincts, but kind of the same toolkit you'd use to manage an intern.

Below that, the results scatter. Task decomposition, retrieval, memory, sandboxing. People are trying everything. Nobody has settled the control layer for agents. Memory and persistent context is one that I'm watching very carefullyright now. I think it's going to evolve a lot in the next year.

And when agents fail, or when people complain about agents failing, to be more precise, it's usually the thinking, not the plumbing. So like two-thirds say that hallucination or losing context mid-task is what frustrates them the most. Allright, so agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath.

Evals11:57

Barr Yaron12:04

So let's take a peek at the stack. We asked: what is the biggest challenge in your stack? Every single year that I ask this, the number one answer is evals. So evals lead here, same as always, but by a very thin margin.

That margin is getting smaller. And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. So if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map.

And the leading challenge, how to evaluate your AI outputs, requires many different methods. But as always, the vibe review is number one. So there are some consistent things that we'll see if they change over time, but they have not changed.

Build vs Buy12:53

Barr Yaron12:53

Okay, this is interesting. So across eight layers of the stack, we asked: what do people build versus buy? Again, maybe the corporate card is going to play a part in this. But there is a wide range and mix for every layer of the stack and a few clear takeaways.

So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure, and fair enough. Prompt management is the opposite. 61% build it themselves. Apparently, everyone's prompts are special.

And this is true of a lot of the product logic: prompts, RAG, evals. They tend to stay in-house on a relative basis. Fine-tuning is the clearest not yet. Most people don't have it at all. And folks are pretty locked in.

So those who bought aren't looking as much to build. Those who built aren't looking as much to buy. But those are the core takeaways from the usage in our stack.

So many of you work on teams. And like we said at the start, these range from solo founders to large enterprises. What is this doing to teams? And remember, this is a builder-heavy sample. But among builders, the vibes are good, which I'm sure if you look to your left and yourright, you're feeling that.

Culture Shift14:00

Barr Yaron14:19

The vibes are pretty good. 97% report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure, more experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free.

And so there's some happy campers as a result of that. But it's not free-free. There's no free lunch, as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true. And

over 9 in 10 respondents are feeling negative downstream effects in some way. The most common ones being widely discussed at this conference, online, and anywhere that you see AI engineers: erosion of deep technical skills and understanding of the code base.

And these are consequences of cheap code generation.

And the org chart is really feeling it. So many folks, 81%, are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Where you feel it the most is shipping software once exclusively the engineer's domain.

I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles. But today, over a third of teams have non-developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customer-facing features across the stack.

And even when non-developers aren't shipping, a third of teams see them building really useful things: prototypes, front-end mocks, and more. So shipping software is not gated on being an engineer. We knew this, but the extent to which it's being pushed is higher than I expected.

Allright, so where does all of this go? We always ask people to place bets rapid-fire. So let's talk about those results.

Future Bets16:28

Barr Yaron16:39

So present tense first. 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're, as Elphaba and Glinda say, I hope you're happy now. That's great. But 59% fear today's AI code creates long-term liabilities.

Only a third call software engineering a solved problem. Although when I have conversations with folks, sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Happier, faster, but embracing the maintenance build is the TL;DR.

And people are unsure what's going to happen with hiring. And for the five-year bets, we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We asked about the press release, not the achievement.

So will they declare it? Yes. What does that mean? Not sure.

Only 9% bet on transformers being state-of-the-art in five years. Most are unsure. But that was interesting. And then my favorite: will there be more AI compute in space or on land? 36% yes, 38% no. The most divisive question in the survey is about outer space.

I promised you a lot of bar charts, and that was a lot of information. So a review or our 2026 wrapped. Impact is overwhelmingly positive. Image gen doubled, or happy image gen doubled, while audio has the highest adoption intent, the same as last year.

Wrap Up18:03

Barr Yaron18:23

Cost really became a first-class constraint, and we see that everywhere in monitoring, in how ambitious folks that are going out and building AI products are behaving. Open weights augment, but they don't replace. So we're seeing a multi-model future with a consolidation of the stack.

Agents got write access more than ever before, tripling relative to last year, while the guardrails stayed pretty primitive. And inference is the buy market. Everything closer to product logic tends to relatively stay more in-house. It is a very exciting time to be an AI engineer.

I cannot wait to see how the next year unfolds. So you can find the full report in the link up here. Every chart, plus some cuts that we didn't have time for today. I won't ask you to fill out a survey about the survey.

But if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. We will see you next year, or per 36% of you, maybe in orbit.

Thank you.