AIAI EngineerApr 22, 2025· 16:18

The Devops Engineer Who Never Sleeps — Diamond Bishop, Datadog

Diamond Bishop, Director of AI Engineering at Datadog, explains how his team builds AI agents—the On-Call Engineer and Software Engineer—that automate DevOps tasks like incident investigation, remediation, and postmortem writing. He details how the On-Call Engineer wakes up for alerts, reads runbooks, queries logs and metrics, and suggests fixes to let human engineers sleep. The Software Engineer proactively fixes errors by generating code diffs and pull requests. Bishop shares four key lessons: scope tasks and evaluate rigorously, assemble teams of optimistic generalists and UX experts, adapt UX for human-agent collaboration, and treat observability as critical for debugging multi-step agent workflows. He predicts that within five years, AI agents will surpass humans as primary users of SaaS platforms like Datadog, urging builders to design for agent consumers.

  1. 0:00Intro
  2. 2:08AI Shift
  3. 3:45On-Call Engineer
  4. 6:57Software Engineer
  5. 7:47Lessons
  6. 8:32Scoping Tasks
  7. 10:13Building Team
  8. 10:58UX
  9. 11:35Observability
  10. 13:15Bitter Lesson
  11. 14:00Agents as Users
  12. 14:48Future

Powered by PodHood

Transcript

Intro0:00

Diamond Bishop0:17

Hey, I'm Diamond. I hope everyone's, you know, feeling the AGI today. Um, I'll be sharing our AI agents at Datadog and what we've learned building the DevOps engineer who never sleeps. I came all the way from the New York Times building, uh,right across here, to see all of you, so I hope it's worth it.

I've been working for my entire career, about 15 years or so, in AI, trying to build more AI friends and coworkers. Um, I wouldn't read too much into that; I have human ones too, um, I promise. Throughout the, kind of, AI winters and lulls of the last 15 years or so, I've managed to keep doing just that at Microsoft, Cortana, um, building out Alexa at Amazon, working on PyTorch at Meta, and building my own AI startup that was working on a DevOps assistant.

Now, at Datadog, we're building out BitsAI, which is the AI assistant who's there to help all of you with your DevOps problems. So today I'll talk a little bit about that, talk a little bit about the history of AI at Datadog, a little bit about how we think about AI agents today, and where we think things are going for the future.

Datadog is the observability and security platform for cloud applications. There's a lot that we do, um, but it kind of all boils down to being able to observe what's happening in your system and take action on that. Make it easier to understand, make it easier for us to, uh, simply understand and build out things to have a safer and more DevOps-friendly system.

We've been shipping AI for quite a while, actually. Um, it's not always in your face, it's not always out there saying, "Here's a big AI product," but things like proactive alerting, really understanding things like root cause analysis, impact analysis, and change tracking, and much more has been happening since 2015 or so.

But things are changing. This is a clear era shift. I think of this in kind of similar terms to the microprocessor or the shift to SaaS. Um, bigger, smarter models, reasoning and multimodal coming, uh, foundation model wars happening, this general shift where intelligence becomes too shift too cheap to meter.

AI Shift2:08

Diamond Bishop2:28

And what this means is products like Cursor are growing, you know, terribly fast, um, and really, people are expecting more and more from AI every day. Um, with these advancements at Datadog, we're really trying to rise to meet this shift as well.

The future's uncertain. This kind of ambiguity creates opportunity, but there's a lot of potential. For us, that's kind of the dawning of this intelligence age. We're working to move up the stack to leverage these advancements and give even more to our customers by making it so that you don't use Datadog as just the DevOps platform, but also as AI agents that use that platform for you.

This requires work in a few key areas that I'll talk about. Developing the actual agents, doing eval—you just heard a lot about eval. We think about that every day, for better or worse—um, and building out new types of observability.

There's a few agents that we're working onright now in private beta. The first is the AI Software Engineer. This kind of looks at problems for you, looks at errors, tries to recommend code, uh, that we can generate to help you improve your system.

The second is the AI On-Call Engineer. This wakes up for you in the middle of the night, does your work, hopefully makes it so you have to get paged less frequently. And then we have a lot more on the way.

On-Call Engineer3:45

Diamond Bishop3:45

So I'm going to talk a little bit about the AI On-Call Engineer first. This is the one that, you know, everyone wants to save them from that 2:00 a.m. alert. You don't want to have to wake up in the middle of the night, go and look through your runbook, go and figure out what's going on if you can help it.

Our On-Call Engineer is there to really make it so you can keep sleeping. This agent proactively kicks off when an alert occurs and works to first situationally orient, read things like your runbooks, grab context of the alert, and then goes and, you know, figures out the kind of common stuff that each of you would do on Datadog already.

Look through logs, look through metrics, look through traces, and kind of act in this loop to figure out what's going on.

The On-Call Agent's great for both automatically running investigations for me, but also, you know, being able to look through and find summaries and find information for me before I even get to my computer. So if I want to get insights into why an alert just occurred, or figure out why a trace might, uh, be showing an error, this agent can jump ahead, pull information for me, and show it to me.

We also have added a new page that makes it easy so that you can have human-AI collaboration. This is still something I'm thinking about a lot, is like, what kind of collaboration do we expect? We want our agents to act as humans, but we also need to be able to verify what they did and be able to kind of look over what they're doing and really learn from it.

It also helps you to kind of earn trust along the way. I can see the reason why, uh, this hypothesis, for example, was generated. I can see what the agent found, and I can make decisions about whether or not I agree along the way.

It also tells you things like, what steps did it actually take out of your runbook? And kind of like a junior engineer who does this work, I can go ask follow-up questions, find out why it did a certain thing.

A little more insight into how we're making this happen. Much like a human SRE or DevOps engineer, our agent works to put together hypotheses on what might be happening and reason over them. Coming up with ways to test them, use tools in the toolformer sense, to try out ideas, run queries against logs, metrics, etc., and work to validate or invalidate each hypothesis.

In the case that it does find a solid root cause, our agent, uh, can suggest remediations along the way. Again, just like a human might. Might say, "Hey, we should page in that other team that's involved here." Or it might offer to scale up or down your infrastructure.

Over time, we've planned to add more built-in actions and eventually discover new types of workflows based on what your team has done. But if you already have certain workflows that you've set up in Datadog, um, we can tie directly into them and make it so that our agent can understand those workflows and how they might map to helping you remediate a problem.

And if it's a real incident, the On-Call Engineer is not usually done once an issue is remediated. You usually go and write a postmortem, you go try to learn from it, you share it with your team. Our agent can do the same, write out your postmortem for you, look at what occurred during the entire time, what it did, what humans did, and put that together so that you have something ready in the morning.

So that was the On-Call Engineer. Um, that's the one that is, you know, trying to help you in the middle of the night, trying to help you every time alerts come on. Um, we also have this AI Software Engineer.

I think of this as the proactive developer, the DevOps or Software Engineering agent who observes and acts on things like errors coming through. This is kind of the error tracking assistant. It automatically analyzes these errors, identifies causes, and proposes solutions.

Software Engineer6:57

Diamond Bishop7:12

Those solutions can include generating a code fix and working to reduce the number of on-call incidents you have in the first place. So they can work in concert to make a better system over time. In this case, the assistant has caught a recursion issue, proposes a fix, and even creates a recursion test so that we can catch it if it happens again in the future.

We have the option to create a PR in GitHub or open the diff in VS Code for editing. This workflow significantly reduces the time spent by an engineer manually writing and testing code and greatly reduces human time spent overall.

So what have we learned building out these agents and some of the new ones that we're working on today? Well, we've learned quite a lot. Um, there's a lot of things that we started with that we kind of went back and, and redid, um, but a few areas I'll touch on that I hope help you as you develop your own.

Lessons7:47

Diamond Bishop8:01

First is scoping tasks for evaluation. It's very easy, you know, to build out demos quickly. Much harder sometimes to scope and eval what's occurring. Second is building theright team who's ready to move fast and deal with the ambiguity that comes with these kind of problems.

Third is that, you know, the UX of old is changing. Um, that's something that everyone needs to be comfortable with. And fourth is observability matters. You know, uh, I'm surprising for Datadog to say that, I'm sure, but observability is terribly important even in this new era.

So scoping the problems, scoping the work to be done. I like to think about this as defining jobs to be done and really kind of trying to clearly understand step by step what you'd like to do. Think about it from the human angle first and think about how another human might go and evaluate it.

Scoping Tasks8:32

Diamond Bishop8:48

Um, this is why we build out vertical task-specific agents rather than building out generalized agents. We also want, where possible, this to be measurable and verifiable and at each step. This has honestly been one of the biggest pain points for us, and I think this is true for many people working in agents, where you can quickly build out a demo, you can quickly build something that looks like it works, but then it's very hard to actually verify that over time and improve it.

Um, use your domain experts, but use them more like design partners or task verifiers. Don't use them as the people who will go and kind of write the code or rules for it, because there is a big difference in how these kind of stochastic models work versus how experts work.

You know, everyone kind of knows GNOME and his, uh, anti-NLP, um, rants, but that kind of stuff happens pretty frequently at domain experts. Eval, eval, eval. I can't stress this enough. Um, start by thinking deeply about your eval.

The number of mistakes we made by not thinking about eval first is, uh, frustrating and something that I think everyone should think about. It's very easy to build these demos, as I said, um, but everything in this fuzzy stochastic world requires good eval, even something small to start.

This means offline, online, and kind of living eval. Have end-to-end, uh, uh, uh, tasks, have end-to-end measurements. Um, make it so you also instrument appropriately the way to know if humans are using your productright and giving you feedback.

And then make this a living, breathing test set.

Building the team. Um, you don't have to have a bunch of ML experts. There aren't that many to go aroundright now. Um, what you really need is you want to seed it with one or two and then have a bunch of optimistic generalists who are very good at writing code and very willing to try things out fast.

Building Team10:13

Diamond Bishop10:27

Um, I'll also note that UX and front-end matters more than I'd like as a backend engineer myself, um, but it's terribly important as you collaborate with these agents and assistants. Um, and then you want teammates and people who are excited to be AI augmented themselves.

This is day-to-day AI use. This is explorer types who want to learn. This is a field that's changing fast. Um, and if you don't have people like that, you're going to kind of get stuck. You want folks who kind of, you know, yeah, yearn for the vast and endless AI capabilities,right?

Um, it's a big world out there and there's a lot going on. Ye old UX. Um, this is one of those things that I still, you know, we think about, we go back and forth every day. Um, it's an area that I didn't realize was quite so important initially when I started working in this field.

UX10:58

Diamond Bishop11:14

Um, despite my engineering sensibilities and lack of UX, it's terribly important. Um, this is such an early space of work. This is kind of one of the more important things here as you collaborate and work together. But the old UX patterns are changing.

Be comfortable with that. Um, and so far I'm partial to agents that work more and more like human teammates instead of building out a bunch of new pages or buttons.

Observability11:35

Diamond Bishop11:35

So who watches the watchman,right? Um, you have these agents running around. Um, observability is actually really important. And don't make it an afterthought. Um, these are complex workflows. You really need situational awareness to debug problems. And this has saved us time a lot as we start to work with, um, a new view that we're calling LLM Observability in the Datadog product.

Um, Datadog in general has a full observability stack. As many of you know, we can look at GPUs, um, we can look at LLM monitoring, we can look at really your system end-to-end. But tying in the LLM Observability has been very helpful because you have a wide variety of interactions and calls out to models you're hosting, models you're running, maybe models you're using through an API, and we can make them all, uh, kind of grouped together in the same pane of glass so you can look at them and debug what's occurring.

I will note though that this can get messy fast with agents. Our agent, for example, has very complex multi-step calls. You're not going to look at this and figure out what's going onright away. Um, this can be hundreds of calls.

This can be, uh, you know, uh, tons of different places where it's making decisions about tools, looping time and time again. And if you just look through a full list of these things, you'll never really figure out what's going on.

So here's a quick, you know, sneak peek into a more agent view of what's occurring inside of our observability tools. This is our agent graph. Um, really what this means is that I can kind of look at it just like our agent did and looking at workflows that are occurring.

You can see in this, even though it's a big graph, uh, there's a bright red node here. If we zoom into that, we can actually see where errors were occurring. This is very human-readable, something that makes it super easy to figure out what's going on when your complex workflow is running.

Bitter Lesson13:15

Diamond Bishop13:15

As an aside though, I do also want to note what I think of as kind of like the agent or application layer bitter lesson. Uh, general methods that can leverage new off-the-shelf models are ultimately the most effective. Um, by a large margin.

Um, I hate to say it, but like, you sit there, you fine-tune, you do all this work on like the specific, uh, you know, project, the specific task, and then all of a sudden, you know, OpenAI or someone comes out with a new model and it handles all this, you know, kind of quickly.

A lot of the reasoning is solved for you. Um, we're not quite there where it handles all of it very quickly, but you should be at a point where you can easily try out any of these models, um, and don't feel stuck to a particular model that you're, you've been working on for a while.

You know, a rising tide lifts all boats here.

Um, I also think a lot about not just building agents, but what it might mean for other agents to be users of Datadog and other SaaS products. Um, there's a good chance that agents surpass humans as users in the next five years.

Agents as Users14:00

Diamond Bishop14:11

Um, I'm probably somewhere in the middle on my estimate there. You know, there are people who will tell you that'll happen in the next year. There'll be people who will tell you, you know, it'll happen in 10 years.

I think we're somewhere around the five-year mark. Um, but this means that you shouldn't just be building for humans or building your own agents. You should really think about agents that might use your product as well. An example of this is like third-party agents like Claude might use, you know, Datadog directly.

I set this up with MCP relatively quickly. Um, but any type of agent that might be coming in and using your platform, you should think of the context you want to provide them, the information you want to provide about your APIs that agents would use more than humans.

So looking ahead, um, the future is going to be weird. It'll be fun. Uh, and AI is accelerating each and every day. I strongly believe that we'll be able to offer a team of DevSecOps agents for hire to each of you soon.

Future14:48

Diamond Bishop15:01

You don't have to go and use our platform directly and integrate directly. Ideally, our agents will do that for you and our agents will handle your on-call and everything like that for you. Um, I also do think that AI agents will be customers.

Many of you building out SRE agents and other types of agents, coding agents should use our platform, should use our tools, um, just like a human would. And, uh, we can't wait to see that. And generally, I think that small companies out there are going to be building, built by someone who can use automated developers like Cursor or Devin to get their ideas out into the real world and then agents like ours to handle operations and security in a way that lets, you know, an order of magnitude more ideas make it out into the real world.

Thank you so much. Um, please reach out if you're building any agents that want to use us, um, or if you'd like to check out our agents as well. Um, there's a lot to build here. And if you want to work in the space, we are hiring more AI engineers and people who are just excited about it.

But thank you very much.