Intro0:00
Thank you. Especially after lunch, that applause tells double. So, yeah, you need a platform team. Um, so usually I'd skip over this slide, but this time I'm going to explain to you my background, because it's important to understand what the next things is that I'm completely biased in this what I'm going to talk about.
So why am I biased? In 2009, I organized a small conference called DevOps Days. There were 60 people. And this is how the world kind of got into DevOps as a word. So my position has been privileged to be there from the beginning and seen 15 years of a new kind of phenomenon, discussion, transformation taking root across the whole enterprise.
So that's kind of like why a lot of what I'm talking about is things from the DevOps space, what we learned here, and I'm trying to apply this to whatever is the new thing we're doing. Um, when you apply your old paradigms to the new world, it could either be correct and it could be totally wrong.
So I'm explaining what I think that I'm seeing. Um, and one of the particular interests is, after seeing Dev, Sec, Ops, and kind of whatever Ops in the world, I especially like this part of the genAI, uh, in particular for that it's not like for me, it's automation intelligence.
So a lot of my work on DevOps was automating pipeline, making them more robust, and kind of making sure we're delivering the value on this. So that's why kind of I, I get excited about this new world.
So over the years, I have gray hair,right? So I've seen my fair share of new things coming in the industry. And the pattern is usually you have one team that's the pilot team when something new happens. We're going to try this out.
Scaling Patterns1:58
Then you scale this out into two or three more teams doing this, and then you extrapolate some learning, some patterns, and eventually you want to scale this out to all the teams. This is what happened with Cloud, the infrastructure.
You know, initially it was a small thing, uh, but then eventually this led into an abstraction layer of us kind of making this easier for other teams so they don't have to understand the whole space, but they can move faster within their domain.
Like DevOps, everybody was saying DevOps is a bad name. We, we kind of hate it, but in the end, the industry stuck with this. AI Engineer is a bad name. What does it mean? Nobody knows. But it's also the advantage that you can bring everything to the table and let it grow.
So if you make it too defined, then we are not evolving that way. That, so that, that has kind of evolved in the first, there was the Agile team that were like really proud as the change agent happened to it all, and we're kind of bringing this in.
I think what I learned over the years of DevOps is the name does not matter that much. In the beginning, it's really important that you have a label that you can search and find all the emerging stories of people doing this tech.
So that's how I look at AI Engineer as a term. If you search for this, it's going to be on your job post, it's going to be on your website. This is likely the term that kind of fits into us finding each other into the new space.
Bridging Friction3:59
And then with DevOps, we're trying to bridge something in the organization. Two things that had friction that did not work together. And I've witnessed firsthand in the company after, you know, now two years building genAI application. The friction was in the beginning.
We all want to do this genAI thing. It hit first the data science team, and then they were like screaming, "Hang on, we're not used to running things in production." Right? So all of a sudden you got friction.
They know their world, we know their other the other world. And we started slowly moving some engineers into the data science team. Eventually, the data science team became smaller of that part because it was less about the data science.
We heard it this morning. It was more about integration thing. And the ratio of the data science got smaller compared to the engineering. And then what we did in the organization is we started scaling this out to the different teams and eventually moved this into the platform.
So that kind of friction of the movement in an organization is something I, you know, that, that I see happen a lot. The data was left on their own in their own data lake, and then now we're actually binding and bonding back together in that way.
So in that way, that's the shiftright. Security was shift left. So we're always shifting around, like whatever the complexity is we can do.
Maybe first a question. Who is familiar with team topologies as a concept? Wow. Okay. Not that many. So first you had DevOps as a concept, bring two teams back in the same room, having them collaborate. That was one way of collaborating.
Team Topologies5:25
But that did not scale across the org. So slowly, uh, they became their own team that abstracted pieces away. The Kubernetes pipeline, all the infrastructure that, uh, developers needed to build, the security tooling, that kind of moved into a platform team.
Team topologies originally was actually called DevOps topologies because it was a m like a movement of the teams, uh, reacting to the you build it, you run it. So we eventually the teams you that are building and running it, we call them the, uh, kind of the feature teams, and they run on top of the platform.
And the platform team can work together with those teams either by just giving it an API,right? Run as a service and you're done. We can do more collaboration to see, hey, you know, what are you building? Let's build this, make sure that the platform is helping you.
So it's a little bit more like there are customers, we're building a product, you're building on top of that. And then we can do facilitation if they're having some issues and we're helping them. So kind of that interaction team.
Uh, there's a whole book on this, Team Topologies, and it helps people organize kind of their infrastructure teams, their platform teams, their SRE teams. So it's a known pattern how we deal with these kind of abstractions of new technology, bring it into the company.
So what I'm explaining here is this whole genAI thing. The impact is going to be like not just the data science team doing this, but what I tried to do in the last two years is bring this to the traditional application developers so they can scale this out into their own teams.
Platform Services7:11
We heard it this morning. What's the ideal team? It's the application developers, the traditional ones, and some of the data science on top of that. But there's a few hurdles to coming to that. First piece that I'm talking about is the platform.
So what service is going to an AI platform? And, you know, you're here at the fair. There's like, I don't know, how many vendors. There's lots of hardware. So I have no affiliation with any of the things that I'm showing.
I'm just going to show you some pieces that go into that platform that you're providing. First piece, access to models,right? It's this first thing they can do. Instead of every team going out to figure out what's the best model, where do we get access, that team kind of figures out what is appropriate for our company.
You pick your poison, what you like, what's your vendor, typically related to your cloud vendor. So there's a relationship with the cloud ops team and kind of bringing this in. Next piece that you hear it like RAG, let's have a vector database.
In the beginning, you had like very specific ones. Now you see they're all getting inside of the traditional vendors of kind of the databases and whatever. So it's not that special anymore, but you need to understand what it is, what is a vector, what is embeddings, how we're helping this.
So that's one of the other pieces you bring when you're having the RAG in. And having that, uh, vector database is not enough. So you bring in connectors to all your data sources. So that's another piece that that team builds in the infrastructure.
They connect whatever is out there in your company, and they expose that to all the other teams instead of them having to do this all the time one by one on teams and figuring out how that works. Some have called this RAG ops, you know, whatever ops,right?
There's always something to be, uh, managing and running on that. And then we had version control for code. Obviously, we want to have version control for models. So that's another piece they can provide. So instead of them figuring out on their own how we do this, there's a centralized repository.
Other teams can reuse, uh, some of the registr uh, some of the models. So it becomes like visible, like a library to all the other pieces as well.
And then especially the larger enterprises, uh, it's almost like they want to be cloud agnostic. They want to be model provider agnostic. So I'm not saying like in storage, we used to have S3 as being the standard protocol.
Now maybe that's the OpenAI protocol that being standardized on, like the one winner kind of takes it all. But it's also kind of building this in the access control, who gets access to what. So there's like a proxy that these kind of environments are building.
Uh, so not everybody can just go out and use whatever they want. So there's a little bit of that.
Then when you've been running this, you want to have similar to your observability and tracing. You want to capture whatever prompts is running in production. Uh, so that observability layer, you want to enhance your existing observability layer, but it's a little bit different.
Uh, you know, one prompt, it's rarely one prompt. Usually one prompt leads to five, six iterations, questions. So that's another piece of this. And you don't want every team to figure this out on their own. And then when you run it in production, you're going to have monitoring on your data quality.
And I know like, you know, LMM, MLOps had this before, but this is kind of a new thing that the traditional monitoring and metrics providers are not very capable of because like you had a health check for your API calls, now you have a health check for evals running all the time in production.
So you will notice if the model has changed, you will notice if the end user is doing strange things, you notice. So kind of that observability kind of is built in. And then lastly, you know, there's many things like caching services and you go on like feedback as a service been also discussed this morning.
Uh, not just thumbs up, thumbs down, try again, but also kind of in, uh, inline editing of the solution because you get better feedback. I mean, again, you don't want every team to build this service because it's quite expensive.
Uh, and you want this to be centrally, uh, managed in a good way. So these are just a few pieces, uh, to bring it on the board. Um, if you ever seen the Kubernetes ecosystem slide, this one is also expanding rapidly,right?
So there's similarities on whatever hardware we're running. It's just going to keep growing, uh, whatever solution we have out there. So that gives you a little bit of an idea is it's not just clouds and APIs, but this is a new set of infrastructure that you're running.
Enablement12:15
So next piece is you have all the infrastructure and you provide it to the teams. And one of the things we learned is that it's not just enough to say, "Here is a bunch of things you could use."
Like any good company, you guide your customers to use them. That's the enablement to kind of whatever you're providing that they're able to use. This is also the place where you get feedback, what is working and not working from your teams.
So what does the enablement look like for that team? You provide prototyping tools for them to easier do experimentation. And it's been mentioned before, don't forget the product owners. They want to learn, they want to experiment. So it's a simple thing you do to kind of get them excited, having them play around, find theright use case, uh, of their stuff.
And then you also connect it with the data of your company. So you do that in a secure way so they can experiment whatever data is there. And frameworks are great, and I'll come back to that later. For learning, how does it work?
What can we all do? And I learned a ton through these, uh, things. And there's a flavor for everybody, like large data, kind of like you like more like Microsoft stack. Definitely framework is, uh, one of the things, uh, to give them and hand them over, uh, to learn from.
And it also helps in education, documentation, uh, going from there. They want a local dev environment. They want to feel safe to do some coding, experimentation when they travel, kind of like they don't want to have the resources.
They're very like focused on having this local stack of development, uh, when working on this. There's pro and cons, but you know, to get them excited is one of the things that you can definitely do. And we've seen this morning, you can run more and more of these quality models on your laptop.
So that kind of helps them in the faster iteration. So things that I saw that actually go a lot bad is what is the actual use case? Uh, I've seen companies shout like, you know, we need to have the genAI.
Every part of the product needs to have genAI. But we're in the phase that we're still figuring out the real use case for a lot of things. And we're I'm not going to say we're running on the marketing budget, but sometimes it's a little bit like that,right?
And another pitfall is if you're very focused on kind of model training and fine-tuning, I can tell you that companies run by the data science and it might take a while until we actually get something in production as well.
Um, an overfocus on, uh, kind of, uh, cost, like run local, not use the perfect model. There's a lot of cost that, you know, goes down the drain that way. So we know that the models are going to get cheaper.
So focus on the first thing, get that business case focus, and the rest will kind of like we'll sort out later. If there's actually ROI, we'll reduce the cost on that. And then end user feedback. Um, the engineers often don't know what to do with the feedback because the product owner is the only one that knows the domain.
It was mentioned this morning. So that's kind of also like, yeah, we get all the feedback, but now what? Like, yeah, we're not used to handling this kind of feedback. We're used to handling error stacks as well. So those are a few of the observations.
And that brings me to there's a lot of emphasis on developer experience in general and improving the productivity. But what does it take to be building genAI applications? What kind of pains do you have? Like I've yet to see the developer that knows what to pick as a model.
Developer Pains15:35
They just pick one and it's it's okay, but you got to start somewhere. Access to the data in the test is hard. Uh, I hear a lot of stories that people use the framework. They were bitten by a framework.
They don't like the framework anymore because the frameworks were moving too fast. Uh, but and they start doing things themselves at the lower layer of APIs. I hope that's not the way we're going. I hope we eventually have an ecosystem of tools we're building on and not doing the DIY.
But it's kind of what I've seen people in the journey. They get excited with the framework and then it's like, okay, we we got this. It's just a prompt,right? Like it's not that difficult. And in essence, they'reright. But think about like the tracing, the monitoring, the observability, all the ecosystems you want to kind of bring together.
And we went through, I think, eight different models over two years for LLMs to get like improvement of quality. And me looking as a VP engineering, I was like, hang on. So every time we change the model, we have to rewrite and rework all the application prompts to make sure it works.
That's very costly. And if it wasn't costly enough, they were doing this manually. So I had no guarantee that was actually going to work. So we've talked enough about evals and working with evals, but it's definitely one of the pain points of you need to have testing evals before you do this refactoring.
And this brings me to the testing, which is every time that the engineers go like, what? Like how do I like write tests for this? Like it's, you know, this morning we're like, oh, it's a vibe check. Okay.
You know, I don't get it. It's like that's not engineering rigor. So I'm going to give you a like the buildup that they're saying like, well, I can do exact testing. That that's easy,right? Just does it need to be 20 characters, a regex?
I know that kind of stuff. Sentiment check. Okay. Well, you know, all of a sudden you have to run a model, a helper model, uh, to bring that in. We can see a simple thing like similar to the semantic, uh, distance is the question related to answer.
That's an easy one. But then you're telling me, how do you check an LLM? Use another LLM,right? So if the AI doesn't work, just use more AI. Like it blows my mind. Like it's like a vicious circle and it hasn't been really solved, but hey, it's the best thing we got.
And I got it. Like and eventually there's a human feedback. So it's it's a compensation of a bit of everything there. Um, so now I want to switching gears a bit. Like when you want to bring this to the engineers, a lot of the engineers are afraid of the AI in the beginning.
Testing Ironies18:23
So how do you get them excited? And the usual trick is we get them something on their coding pilot because then they see kind of the productivity that this brings. And so it goes hand in hand. Like it's a little bit weird that the same team that is building the genAI model is also being the evangelizing of kind of like take away the fear of developers using more AI.
And this one, uh, I know there's a lot of talk about like getting them more productive, but what I want to show here is yes, the co-pilot introduced a lot more code. It was a lot more things to do, but review times went up and the PRs sent by the actual AI were bigger,right?
So we got faster, but then we started slowing down again. So there's like I'm not saying it's like 100% compensating the other part, but there's definitely it's like there's friction on the line. And this brings me to a paper that where it was very instrumental in the beginning of, uh, DevOps.
It was called the Ironies of Automation. And now there's a new paper, the Ironies of GenAI Automation. And so you see, um, the role of the person using any of the AI is changing from the person producing to somebody who's managing and reviewing things,right?
So and you could say, well, you know, that that's good. Like, you know, but you see spending more time in the review and kind of faster of the generation, but it can only also lead to I don't understand anymore what this this is doing because I don't have the domain model.
I don't have the experience anymore. And then I'm just going to say accept. There's a funny story about GitHub co-pilot being used, uh, and having a higher acceptance rate for suggestions in the weekend because the developers like, whatever.
So, but we've seen this before, like the automation that acceptance. So we kind of have to go against this and you would gradually when you're not doing the producing job anymore, you lead, uh, like lose situational awareness of when things, uh, go, uh, failing.
And this brings me to this slide. So
DevOps automating things go away. I'll replace you with the shell script. That was kind of, you know, the left part. 20% automate. The other percent that grew into the industry was preparing for failure. First, we had CICD, we created more tests, evals.
Failure Prep21:01
Then we had monitoring because we needed to know what was going on in production. And then all of a sudden it's like, yeah, but if it goes down, like can we prevent the failure? We're going to assume failure.
That was the big model shift. So we started design for failure and then we needed to predict but we didn't know what was going on but we need situational awareness. That was the whole observability. And then in the end we brought in chaos engineering to kind of inject failures so we can keep training when failures happen.
Right? So you see kind of this paradigm of yeah we won a lot but kind of that spun up a whole new thing of doing. I'm not saying everybody's doing this. You can gladly skip all the tests if that's your risk appetite.
But that's kind of the shift that it was going to. And it's also reducing cognitive load in when failure happening, uh, and kind of work from there. So that's, you know, little piece on enablement. It's not just, you know, kind of making them genAI.
It's also putting them at ease. And then there's the governance piece. Uh,
Governance22:19
make a personal awareness program,right? Don't just copy things in. Uh, this is one that records your screen and looks everything on your screen,right? That's it's getting scary in that way of leaking things. Uh, but it also promises productivity.
So it's like it's a tension that you have to overcome in your company. Uh, you want to have them opt out on training. It's we've talked about this this morning. It's not always that clear where you opt out, when you opt out, uh, how to work from them.
So you have to make make them aware whenever they're putting a new service in there. You want to have them check the license, but that's overcome by just restricting the number of models you put in. Uh, it's not just all because it has the word open and it's open,right?
So we all learned that. Um, look at what the origin was for the model. Learn from that as well. So those are thing that typically goes into a governance, uh, uh, workflow, um, of approval. And then there's the European.
I come from Belgium. So risk levels assessment. What kind of workflow are you doing? Is it allowed? Should we do this? Kind of make this awareness a little bit of the law. I'm not making them specialist lawyers, but at least, you know, they kind of have to know the basics of that as well.
And then there's prompt injection, but we learn it's not a solved problem, but you got to have something,right, at least in place, uh, for the failures. Um, and then guardrails. Uh, but I don't know if you ever done like intrusion detection alerting or web application firewalls.
You put all the rules and after a while you say like the logs are so big you're like whatever. So it's it's not always a problem, but we kind of tell ourselves that there's guardrails in place, uh, in there.
So let's hope we are like on the path of improving those kind of tools eventually like for preventing failure and optimizing for failure. And then we want to like focus on the PII that goes over the wire, the metrics, the alerting because we're now sending a lot more text over the wire, uh, kind of uncontrolled things.
Platform Models24:26
So it needs a lot more scrutiny on there. So simple thing for me. These are the steps you bring in. You bring your infrastructure on one team. Uh, you kind of do the enablement on top of that and they also set the the rules for your company to use this.
And I briefly talked about the platform team which provides the internal focus and the internal enablement to the company. But, um, there's another model called the unfixed model and it's a little bit like alternative to the team topologies.
And they also talk about the experience crew. And what the experience crew is, they make AI look consistent in the product. So they're a team that goes on to the feature teams. It's like, so what are you doing?
How does it look like in the product? What's the UX? And kind of they do this. They have a whole lot of number of crews, but it this is also using companies and they call it, you know, the experience, the experience team that works kind of hand in hand.
So that leads me to cloudops, secops, developer experience, put the data platform in there, put the AI platform infrastructure in there, and on top of that do the AI, uh, kind of experiences in there. Um, I would not recommend to do this for a 10-person company.
Uh, don't do premature optimization, but if you're scaling out to 10 and more teams, that's probably a pattern that is known. And the nice thing is you have cross collaboration. The secops are really good at access control and governance.
So they help each other. The cloudops know what to spin up and know the cloud vendors. And so there's a lot of kind of working together. And if you bring that into one area, you have a better chance of pulling this in, uh, and being like a collaboration across.
Q&A26:15
So that was me and, you know, you can scan the link and connect and, uh, yeah, I don't know. What do you think? Want to have a platform team for this or are some people doing this already? Uh, let me know.
So, thank you.
Questions? Yeah, we have literally we do have a time for a question or two if anyone does want to ask anything. Uh, fantastic. On the slide. Will there be available? Oh, sorry. You want to get a copy of the slides?
Okay. He wants a copy of the slides. Yeah. Yeah. Sure. Okay. Cool. Oh, sorry. Yeah. Go. Go for it. So you you had mentioned guardrails as a service and this is one of the interesting paradigms that we're looking into as well and trying to, uh, sort of figure out how to best serve guardrails as a service.
Did you see or is is there a certain leaning towards you put guardrails in front of all models that you serve and then like in your model card you say this is the model these are all the guardrails in front or do you expect product teams to consume the guardrails but leave it in their remit to go and consume them properly?
Yeah. So what we've seen is that the central governance teams will put the generic rules in place and then depending on the use case and whatever they're putting the each of the teams put their own rules on top of that.
So there's a kind of self-servicing but also like a centralized component of rules. Not every team has to duplicate, uh, to go from there.
Cool. Thank you very much. Allright.
Help you get set up.





