AIAI EngineerDec 19, 2025· 18:11

Leadership in AI Assisted Engineering – Justin Reock, DX (acq. Atlassian)

Justin Reock, Deputy CTO at DX (acquired by Atlassian), argues that AI's impact on engineering productivity varies wildly and that leaders must move beyond top-down mandates to focus on psychological safety, measurement of actual outcomes, and targeted integration across the SDLC. He presents data showing a 2.6% average increase in change confidence but extreme variability across companies, with some seeing 20% drops. Emphasizes that writing code is rarely the bottleneck; instead, leaders should identify and fix bottlenecks like context switching, citing Morgan Stanley's DevGenAI saving 300,000 hours annually by converting legacy code specs and Zapier reducing engineer onboarding to two weeks via AI agents. Introduces DX's AI Measurement Framework covering utilization, impact, and cost, and stresses trust-building through system prompt feedback loops and temperature settings. The episode delivers actionable guidance on measuring AI's true impact and enabling engineers through education, time to learn, and creative unblocking of usage.

  1. 0:00Intro
  2. 0:44Mixed Signals
  3. 3:21Key Findings
  4. 4:47Quick Tips
  5. 6:00Reduce Fear
  6. 7:07Metrics
  7. 10:56Compliance
  8. 13:32Employee Success
  9. 14:33Unblock Usage
  10. 15:00SDLC Integration
  11. 17:36Next Steps

Powered by PodHood

Transcript

Intro0:00

Justin Reock0:21

Thanks for joining me in one of the later-day sessions. Looks like we kept a lot of people here. This is a nice full room; great to see it. We're going to go through a lot of content in a short amount of time, so I'm going to getright into it.

If you want to get deeper into any of this stuff, we have published this AI strategy playbook for senior executives, and a lot of the content that I'm going to go through, I'm not going to have time to get quite as deep, but this is just a nice PDF copy that you can come and refer to later.

Mixed Signals0:44

Justin Reock0:44

If you missed this QR code, don't worry, I'll show it again at the end. So, what is the current impact of Gen AI? Nobody knows,right? We've got Google on the one hand telling us that everyone's 10% more productive.

That's interesting. Now they're Google; they were already pretty productive to begin with. But we have this sort of now infamous METER, METR, study, which has some flaws in the way that study was put together, that showed actually a 19% decrease in productivity using codec assistance.

So there's a lot of volatility, a lot of variability. What was really interesting about this study, even though I mentioned there were some flaws, but every engineer that took part in this study felt more productive, but then the data actually bore out that they were less productive.

Kind of interesting,right? We've got this induced flow that makes us feel really good about what we're doing. So we need to address this. DORA has put out some really good research on this too, but this is based on industry averages.

This is impact based on what do we look at when we see a large sample and an average of how certain factors are being impacted by, in this case, 25% increase in AI adoption. We see these modest but positive leaning indicators: 7.5% increase in documentation quality and increase in code quality by about 3.4%.

At least that's not leaning in the other direction,right? And when we started digging through some of DX's data, we have, you know, we're the developer productivity measurement company, we have lots of aggregate data that we can look at with this.

We found the same thing when we looked at averages. We see about a 2.6% increase in overall change confidence, which is a percentage of people who answered positively that they feel confident in the changes that they're putting into production.

Similar positive leaning average when we looked at code maintainability, another qualitative metric, a 0.11% reduction in change failure rate, which when you think about the industry benchmark being 4%, it's not insignificant. But this is not the full story because this is what we saw when we broke the same studies down per company.

Every company here is a, every bar represents a company,right? We have some that are seeing 20% increases in change confidence, while others are seeing 20% decreases. We're seeing extreme volatility, which is why these averages look so innocuous, but they're belying the greater story of variability.

See the same thing with code maintainability, the same thing with change failure rate. So this is a 2% increase in change failure rate up here at the top. Again, with an industry benchmark of 4%, that means shipping as much as 50% more defects than we were shipping before,right?

Key Findings3:21

Justin Reock3:21

We want to make sure we're on the lower end of this, but how? Like, what should we be doing? Well, we found some patterns here. We see that some organizations are seeing positive impacts to KPIs, but others are struggling with adoption and even seeing some of these negative impacts.

Top-down mandates are not working,right? Driving towards, oh, we must have 100% adoption of AI. Great. I will update my read my file every morning and I will be compliant,right? We're not actually moving the needle anywhere when we do that.

We also find that lack of education and enablement has a big impact on sort of negatively impacting this. Some organizations just turn on the tech and expect it to just start working and everybody to know the best ways to use it.

And a difficulty measuring the impact or even knowing what we should be measuring. Like, what metrics should we be looking at? You know, does utilization really tell us much about the full story of Gen AI impact? This is another graph from DORA.

This is a Bayesian posterior distribution, which is an interesting way of representing data. Basically, you want your mass to be on the yellow side of this line, theright side of this line for the audience, yeah. And you want a sharp peak, which is telling you that we're pretty confident that this initiative will have this impact.

And if we look at some of the top-line initiatives here, these are things like clear AI policies,right? We want to make sure we have that. We want time to learn, not just giving people materials, but actually giving them space to experiment,right?

And so these types of factors are the ones that seem to be moving the needle the most. So we're going to go over some quick tips on how we can do all of these things, and again, the guide will go deeper into this.

Quick Tips4:47

Justin Reock4:57

We want to integrate across the SDLC,right? For most organizations, writing code has never been the bottleneck,right? We can increase productivity a bit by helping with code completion, but our biggest bottlenecks are elsewhere within the SDLC. There's a lot more to creating software than just writing code.

We want to unblock usage. We can't just say, well, we're worried about data exfiltration, so we can't try this thing. Like, no, get creative about it. We've got really good infrastructure out there now, like Bedrock and Fireworks AI, that can let us run powerful models in safe spaces.

We have to have open discussions about these metrics. We need to evangelize the wins, and we need to let our engineers know why we're gathering metrics and data. What is it that we're trying to improve? We have to reduce the fear of AI,right?

We have to make sure that people understand that this is not a technology that is ready to replace engineers. This is a technology that's really good at augmenting engineers and increasing the throughput of our business. We have to establish better compliance and trust, and we need to tie this stuff to employee success.

Reduce Fear6:00

Justin Reock6:00

These are new skill sets. AI is not coming for your job, but somebody really good at AI might take your job. And so as leaders, we have the opportunity to help our employees become more successful with this technology.

So how do we reduce the fear? Well, first of all, why do we need to do this? Well, there's a lot of good reasons, but I love to point to Google's project Aristotle. This was a 2012 study where Google wanted to figure out what are the characteristics of highly performant teams.

They thought that the recipe was just going to be what Google had, this combination of high performers, experienced managers, and basically unlimited resources, and they were dead wrong. Overwhelmingly, the biggest indicator of productivity was psychological safety, okay? And so that very much applies now.

We also have data, like this is SWE bench, I'm sure a lot of you have seen this, and there are some impressive benchmarks that the agents can do, like, a third of the things they're asked to do without any human intervention.

That means that they're not able to do two-thirds of them,right? Again, we are augmenting, we're not replacing, we're not ready, we may never be ready. So we need to be very transparent with what we're doing. We need to set very clear intents.

Metrics7:07

Justin Reock7:07

Why, you know, are we using this to augment, not to replace? We need to be proactive in the way that we communicate that and not just wait for people to get upset and possibly scared. We need to say, no, we are here to help you, to give you a better developer experience and to increase the throughput of the business.

And again, we have to have these discussions about metrics. Now, what metrics? What should we be looking at? Well, DX again, Developer Experience and Productivity Measurement Company. There are two sort of classes of metrics that we can be looking at, really two levers that matter here, and that's speed and quality,right?

We want to increase PR throughput, we want to increase our velocity, but not by just creating a bunch of slop that's going to give us a bunch of tech debt later that we're going to have to deal with.

And we've just kicked the bottleneck down the road if we do that,right? So we want to be looking at things like change failure rate, our overall perception of quality, change confidence, maintainability. And we have three types of metrics that we can be looking at here.

We have our telemetry metrics. These are the things coming out of the API, and they're good for some stuff, but they're not always accurate,right? We know, like, accept versus suggest was kind of like all the rage until we realized that engineers need to click accept in the IDE in order for the API to know about it.

Even if they do click accept, who's to say they didn't just go back and rewrite every line that was suggested,right? So that's providing us some context, but we also need to do some experience sampling. We need to, like, for instance, add a new field to a PR form that says, "I used AI to generate this PR," or "I enjoyed using AI to generate this PR," and get some data that way.

And then self-reported data or survey data. We are big on surveys, but let me underscore, we're big on effective surveys. 90% plus participation rates engineered against questions that treat developer experience as a systems problem, not a people problem, because that's what it is.

W. Edwards Deming, 90 to 95 percent of the productivity output of an organization is determined by the system and not the worker, okay? So foundational developer experience and developer productivity metrics still matter the most,right? Our AI metrics, like utilization and things, are telling us what's happening with the tech, but these core metrics that we've been able to trust are telling us whether these initiatives are actually working,right?

Are we actually moving the needle and having the outcomes that we want to see? So top companies are looking at different things,right? We are seeing, like, adoption metrics coming out of Microsoft. They've also got this great metric called a bad developer day.

I'm not going to go into it, but there's a really good white paper that shows, like, all the different telemetry that they can look at to determine what makes a bad developer day. Dropbox is looking at similar stuff, adoption, like weekly active users, daily active users, that sort of thing, but also looking at quality metrics, like change failure rate.

And booking is looking at similar stuff as well. And so we built a framework around this. We were first to market with what we call our DX AI Measurement Framework, and this is very much inspired by things like DORA, Space Framework, DevEx, just like our core four metrics set, which you can ask me about later.

And we take these metrics and we normalize them into these three dimensions of utilization, impact, and cost. And you can kind of think about this as a maturity curve too. A lot of people start just figuring out, okay, what's happening?

Who's using the tech? What's the percentage of pull requests that we're getting that are AI assisted, maybe through experience sampling? How many tasks are being assigned to agents? But then we can mature that perspective a little bit, and we can correlate that utilization to impact.

What is this actually doing to velocity? What is this actually doing to quality? And this is when we start getting more mature in our picture of our impact. And then finally, cost, although I like to joke that we're 15 years past the last hype cycle, which was cloud, and we still have new companies spinning up that are teaching us how to understand and optimize our cloud costs.

So we will see if we get there, although I also hear horror stories about people burning through 2,000 tokens, $2,000 worth of tokens a day, so we probably do need to hit that as well. What about compliance and trust?

Compliance10:56

Justin Reock11:07

What can we do to ensure that the output that's being generated is something that can be trusted by our engineers? We have a lot of levers to pull here, but one of the ones that I'd like to talk about is setting up a feedback loop for our system prompts.

So these could be called system prompts, cursor rules, agent markdown. Pretty much all of the mainstream solutions have something like this where you can go and provide a set of rules to control how these models behave. And I won't get too much into the technical details here.

We have an example where, like, the models have been providing outdated Spring Boot stuff. We want Spring Boot 3. It's been sending us Spring Boot 2 stuff. The big takeaway here is to have the feedback loop. Have a gatekeeper,right?

Have somebody or a group in the organization that can receive this feedback, that understand how to maintain and continuously improve these system prompts,right? And that way, we're always maintaining the way that these assistants or models or agents affect the whole business.

It also pays to understand the way that temperature works, especially when we're building agents,right? We do have some control over the determinism and non-determinism of these models. Again, like when a model is predicting a next token, it doesn't just have, like, one token.

It has a matrix of tokens, and those are associated with a certain probability of that being, like, theright token. And so we have this setting called temperature, which is heat, which is entropy, which is randomness, that can control the amount of randomness involved in actually picking that token.

This is sometimes called increasing the creativity of the model. And it's a number between zero and one. For those reasons I just mentioned, don't use zero or don't use one. Weird things will happen, but you want some decimal in between zero and one.

When we have a lower temperature, like we're seeing here, 0.0001, we give it the same task twice, and it gives us the exact same output character for character. When we set that temperature higher, this is an example of 0.9, I'm asking the agent to create a gradient for me, a simple task.

It's giving me two relatively valid solutions. I did ask it for a JavaScript method, and this is the only one that's giving me a JavaScript method. But the point is, there are wildly different approaches to the same problem when I've increased the creativity of that model.

So we need to think about, like, use case-wise, where should we have more creativity and where should we have more determinism? And temperature is another setting that we have that can help control this. You can experiment with all this using, like, Docker Model Runner, Ollama, LM Studio, that sort of thing.

How can we tie this to better employee success? We have to provide both education and adequate time to learn. So we put together a study where we sampled a bunch of developers that were saving at least an hour a day, excuse me, an hour a week, and we asked them to stack rank their top five most valuable use cases, and we built a guide around that.

Employee Success13:32

Justin Reock13:53

A guide that effectively goes through code examples, prompting examples of what we determined using this sort of data approach, where we should get more reflexive about our best practice and about the use cases that we're becoming reflexive in in our use of AI.

And so that's what this guide was about, and we've had this become required reading in certain engineering groups, and proud of that, and this is another way that we can help educate, but we need to give time. We don't have time to go through all of this.

I do think it's interesting that the number one use case for this was stack trace analysis,right? So not a generative use case, actually more of an interpretive use case. And we see some other ones here that are not too surprising, and there's examples of each of these.

Unblock Usage14:33

Justin Reock14:33

What about unblocking usage? How can we make sure that we can creatively ensure that engineers can take the most advantage of this? Well, leverage self-hosted and private models. That's getting easier and easier to do. Partner with compliance on day one,right?

Make sure that what you're doing is in line with your organization's compliance. You may find that you're making a lot of assumptions about things that you don't think you can do that you can actually do,right? And then think creatively around various barriers.

Finally, how can we integrate across the SDLC? What should we think about doing there? You know, and I'm a big Ellie Goldratt theory of constraints fan, probably have some others in the audience. An hour saved on something that isn't the bottleneck is worthless.

SDLC Integration15:00

Justin Reock15:14

And when we look at data across, in this case, almost 140,000 engineers, we find that there are definitely good, like, annualized time savings with AI that are being eclipsed by sources of context switching and interruption, meeting heavy days, these other things that it's like, yeah, we can save time here, but we're losing so much more time over there.

So find the bottleneck, fix the bottleneck,right? Morgan Stanley's been very public about building this thing called DevGenAI that looks at a bunch of legacy code, COBOL, mainframe natural. I hate to admit Perl because I'm an old-school Perl developer, but apparently that's legacy now too.

And basically creating specs for developers that can just be handed to developers to start modernizing the code without having to do all that reverse engineering,right? And they're saving about 300,000 hours annuallyright now doing this. There's a Wall Street Journal article about this, Business Insider article about it.

They're very public about that. Zapier. Zapier should be the example for everyone. They have a whole series of bots and agents that are doing things like assisting with onboarding. They can now make engineers effective in two weeks. Industry benchmark on the good side is, like, a month.

On the medium side is, like, 90 days. And because they're able to increase the effectiveness of the engineers that they're bringing into the organization, they realized that they should be hiring more,right? As opposed to trying to maintain status quo by, like, cutting headcount and trying to make individual engineers more productive, they said, no, we can get more value out of a single engineer.

We should be hiring faster than ever, and they are, and it's really increasing their competitive edge. I think that's theright attitude. Spotify's been helping out their SREs by pulling together context when incidents are detected and then taking things like runbook steps and other areas of context and documentation and pushing them directly into SRE channels so that those critical minutes of trying to get to the bottom of what's actually happening and what we should do to resolve the incident, they just eliminated that time,right?

It's significantly increased their MTTR. So let's get creative about areas in the SDLC that are our actual bottlenecks. Allright, next steps. Distribute this guide as a reference for integrating AI into the development workflows that you have. Determine a method for measuring and evaluating GenAI impact.

Next Steps17:36

Justin Reock17:36

It's really important to make sure that we're not on the bad sides of those graphs that I showed you earlier. And then track and measure AI adoption and see how that correlates to overall impact metrics and iterate on best practices and use cases.

And here's the guide again. Thank you so much.