AIAI EngineerApr 27, 2025· 23:00

How to Build Your Own AI Data Center in 2025 — Paul Gilbert, Arista Networks

Paul Gilbert, tech lead at Arista Networks, explains the key considerations for building AI data center networks in 2025, emphasizing the stark differences from traditional enterprise networks. He details the backend GPU network (eight 400G ports per H100 server, no over-subscription), the front-end storage network (calmer, 100-200G), and the need for lossless Ethernet with ECN and PFC flow control to prevent packet drops during synchronized GPU bursts. Gilbert highlights power challenges (10.2kW per GPU server, requiring 100-200kW water-cooled racks) and the importance of telemetry like RDMA error monitoring and an AI agent that correlates GPU and network issues. He also covers advanced load balancing (cluster-aware, up to 93% utilization) and smart system upgrades without downtime, while noting the upcoming Ultra Ethernet Consortium (v1.0 in 2025) that shifts congestion control to NICs.

  1. 0:00Intro
  2. 0:43AI Workloads
  3. 2:27Network Architecture
  4. 4:15GPU Hardware
  5. 5:51Traffic Challenges
  6. 9:29Resilience & Power
  7. 11:05Traffic & Congestion
  8. 12:58Network Isolation
  9. 14:48Visibility & Tools
  10. 19:07Design Guidelines
  11. 20:24Clusters & Future
  12. 22:05Summary

Powered by PodHood

Transcript

Intro0:00

Paul Gilbert0:17

My name is Paul Gilbert. I'm a tech lead for Arista Networks. I have an accent, but I'm actually based here in New York City, and I build, or design, or I help build and design enterprise networks. But what we do is plumbing, so I'm not going to talk about agents, but more kind of how you train models, what the infrastructure looks like, and how you do inferencing on the infrastructure.

AI Workloads0:43

Paul Gilbert0:43

I normally teach people the very basic stuff, so you guys probably know this already, but these are new terms for us. When we built computer networks, people would come to us and say job completion time, barrier—I'm pretty sure you guys know that—the inference.

And the question I get all the time is, you know, we can build a network to train a model, there's an algorithm, maybe you can use it to look at what you need, but then, you know, what's inference?

And, you know, it's changed a lot now because of chain of thought and reasoning models, the inference is a lot different. It used to be X, but now it's Y. I'm pretty sure you guys have seen this slide, but I use these just to talk to enterprises around kind of what they might be thinking in GPU size.

This Walid Sosa came up, Dr. Walid Sosa came up with this. On the left there is the training, and on theright there is the inference, and it's kind of, you know, on one you have times 18, times you have a times 2.

Again, I think that changes now with chain of thought and reasoning. I'm not too sure kind of which way it's going to go. And at the bottom there was a really interesting one, again, which I show customers, because most of the enterprises I talk to kind of don't understand models and how they work and training.

I know a little, but not a lot. But, you know, the model they trained down here was 2,048 GPUs for one to two months, and then when you go to inference, after fine-tuning and alignment, it's 4 H100s for inference.

So we talk to people about building different types of networks, which I'll speak about, but, you know, kind of I always start at the beginning, and, you know, this is, I got this slide, and I think it's really interesting.

You know, LLMs were kind of just trying a little bit of inference, but now with the next generation models, it's a lot. So this is what we build, or I build, and these are new terminologies for us from the networking world, backend networks, so this is where you connect GPUs to.

Network Architecture2:27

Paul Gilbert2:45

When we build these networks, they're completely isolated because GPUs are really, really expensive, they take a lot of power, and they're really hard to get hold of. So when people build AI networks in the enterprise, we don't connect nothing else to these networks.

The bottom part of that, the backend network, there's 8 GPUs per pool on those servers, and they can be Nvidia, they can be Super Micro, they can be whatever. They will go into a high-speed switch at the bottom there, you have a leaf switch and a spine switch, and nothing else attaches to that network.

And then on the front-end network is where you get storage from to train the model, obviously. You know, the GPUs synchronize, they do something, they calculate, they produce an algorithm, and they call for more data, and that's kind of the cycle.

The front-end network is not as intense as the backend. The backend network, depending on the model that you train, the GPUs will actually work at 400 gigabytes. And for us in the enterprise, you know, and I've built some big, big data centers, but I've never seen anything like that.

So this is, in the networking world, this is a completely new world to us. And we make the networks as simple as possible because, again, these are really expensive, and people want to get their money's worth, they want these running 24 by 7.

So we do IBGP or EBGP, just really simple protocols.

GPU Hardware4:15

Paul Gilbert4:15

I'm sure most of you have seen this, but again, I kind of teach this, so this was an infrastructure presentation, but, you know, that's kind of the back of an H100, probably the most popular, it is actually the most popular AI server out thereright now.

You can see in the middle there, there's four ports, but those four ports are broken out into two, so there's eight ports. Those are the GPU ports there, and then kind of over to the left there, there's the Ethernet ports.

So that's what we connect to. We've never seen anything like this before. You know, when you first speak to people about this, you know, I've seen servers with 400 gig, you know, and I do a lot of the big financial networks, but never before have we seen servers that can put this type of traffic onto a network.

You know, we always, they always ask me about this, and, you know, I got this from an Nvidia slide, it's down there, but, you know, there's this thing called scale up and scale out. I'm not really sure. Scale up, you know, when you buy, when my customers buy these servers, they always have eight GPUs in it.

You can't add anything to an Nvidia server, you get the DGX. If you go with the outsourced model, it's a HGX, so it's a third party. You don't really add things to it. But, so I don't see scale up, but scale out, you know, obviously you can build, we build the networks so that you can add more GPUs.

We can start very small, and we can go up to, you know, hundreds of thousands of GPUs. Not in the enterprise, but the cloud scale guys do. So what's different, you know, for us, again, it's, there's, it's hardware and software.

Traffic Challenges5:51

Paul Gilbert5:51

The hardware are those GPUs, we're not used to them. The first time I tried to configure one, it took me hours and hours, but I'd never seen them before, whereas other stuff I've seen pretty quickly. You know, and you have software, so CUDA and NICO are probably, you know, two of the biggest protocols, and you guys know more about that than me.

But we had to kind of understand, not CUDA, but NICO, because it has a collective, so we had to understand kind of how the collective works, because that will put traffic onto the network in a certain way. The hardware was completely different.

Again, you know, we had the eight 400 gig ports and the four 400 gig ports facing the front-end network. Totally new to us. The other thing was data center applications, kind of web app database. They're really easy. They go from one to the other, and then different parts of the network, and if one fails, you have some kind of load balancing or failure, and it fails over.

This is, AI networks are not like that. The GPUs all speak, they'll talk to each other, they'll get stuff, they'll send stuff, and if one fails, the job might fail, it might recover. But it's a different concept to us, so it's hard to imagine.

And traffic is bursty because all of these GPUs, if you have a thousand GPUs of 400 gig, they will all burst at the same time, and if they can, they will burst at 400 gig. So it's a lot of traffic on a network, and I've never seen anything like that.

So when we build these networks, we don't build them over subscribe, we build them one to one. In the data center world, we used to do one to ten, it went down to probably one to three, but never one to one, because it's just really expensive to build that kind of bandwidth.

But with AI networks, we need to, so we have no over subscription in the network. And from our point of view, if you look at what one of these servers can put on the network, you know, just a H100 is eight 400 gig GPUs, and four 400 gig is 4.8 terabytes, which is, and that's just one server.

The storage size, the front end, probably nowhere near that, but the back end is always wire rate. And then 800 gig is probably, you know, the Bs are just around the corner, I think in March they'll be released, I think there's some people have them.

And those are 800 gig. We support 800 gig today on the network, but each one of those servers, there's a possible 9.6 terabytes per server. And, you know, most people in my world, in the enterprise world, come from servers that maybe one, two, three, or four 100 gig Ethernet, but nothing like 9.6 terabytes per server.

So the other problem we have is the traffic patterns. When we load balance from kind of leaf to spine, we use a thing called Entropy, which is the five type of IP address, port, and MAC address, and we do pretty good load balancing.

But with GPUs, it's just one IP address, and it can sometimes match through a single uplink and over subscribe it, which would be really bad, because you'll start dropping an awful lot of packets. So we have to take a lot of care on how we load balance within the AI network, or how we build the back end and the front end.

So we have some pretty cool tools where we don't now look at the five tuples, we actually load balance on the percent of bandwidth that's being used on the uplink. And we can get up to about 93% utilization on all the uplinks to the downlinks, which is pretty good.

Resilience & Power9:29

Paul Gilbert9:29

You know, and again, one thing that's really new to us is a single GPU can, you know, or a set of GPUs, if they fail, sometimes the model will stop. I know it checkpoints, but a single GPU failure is a problem for us.

And if one of the big problems that we've always had is optics and transceivers and DOMs, which are the rates and the loss between them and the cables, etc. And when you start building these networks with thousands of GPUs, you will have a lot of cable problems, and you will have a lot of GPU problems.

So it's really hard for us because we, again, this world is new to us, the last year or so. Power, you know, power is, you could, you know, you've read the newspapers, you know, everyone's trying to buy nuclear power stations to power these things.

The average rack in a data center today is about 7 kW to 15 kW, and you can put, like, you know, 10 1RU racks into those, and you'll be fine. And when customers come to me say, yeah, we finally got GPUs, whatever, and I say to them, you know, what kind of racks have you got?

And they say, well, we're going to put them in, and then you can only put one of these servers in one of those racks, because they actually draw with eight GPUs, 10.2 kW, so you need new racks. Most enterprises now are waking up to this, and they're building racks between 100 and 200 kW, and they're water-cooled.

There's no way you could air-cool them in a data center. So that's a whole new concept to people as well, is water-cooled racks.

Traffic is both ways, which again is new to us. So north-south, you know, in a regular data center, you have users coming in, database app, web, whatever, and it comes in, it goes out. But in the AI world, when the GPUs speak, that traffic is east-west because it's speaking amongst each other.

Traffic & Congestion11:05

Paul Gilbert11:22

And then when they ask for more data from the storage network, it's north-south. So you have both traffic patterns. The east-west is really bad, that's kind of where they run wire rate. The front end to the storage is much more calmer, because most storage vendors can't put that kind of traffic on the networkright now.

I'm pretty sure they will one day, but they're more around 100, 200 gig.

And, you know, in a network, there's a certain amount of buffering on these switches, and buffering is bad because it means it can't send traffic somewhere because something else is not receiving the traffic. So you need a congestion control and feedback.

Andright now we use something called Rocky V2, which is two parts of Rocky V2. There's a PFC and an ECN. If you were building an AI network, your engineers, your network engineers will definitely know about this. ECN is an end-to-end flow control, where if traffic, if traffic is, if there's congestion somewhere in the network, packets are marked, they go to the receiver, the receiver sends back to the sender, you need to slow down because there's congestion.

And it goes through an algorithm, it pauses for a while, it slows down. And if it doesn't get any more ECN packets, it speeds up again. And PFC is basically stop, my buffers are full, I can't take any more, so it kind of is a dead stop.

So you have kind of a slow feedback mechanism with ECN, and then kind of an emergency stop with PFC.

Network Isolation12:58

Paul Gilbert12:58

The networks we build are really simple. We don't have things like, in regular data centers, we have DMZs with firewalls, load balancers, etc. We have connections to the internet, we have L4 through 7 service, a whole bunch of stuff.

When we build these networks, they're totally isolated. The GPU, the back end is completely isolated. The front end possibly could have connections to something, but even then it's so expensive to build, you don't want to take the chance.

On demand, the applications that we're used to, you know, if it fails or something fails, something will recover, and, you know, you may get a little skip or a jump, but if you've done theright thing, it's not going to be that bad.

In this world, if something fails, the model may fail, and the call that you get into the operations center is a different call than you get if your app kind of restarted and everything's good again.

The other thing is collectives. You know, obviously NICO will go out there and work out kind of where the GPUs are and what to do with it. But there's kind of different designs. So I tell my customers that speak to your data scientists and your programmers, developers, and find out kind of what they're doing and what kind of models they're building, because it can affect the network on kind of how you build it and how you design it.

So networks totally isolated, things are moving fast. We're at 800 gigright now, which is, you know, we have been for probably a year. We will see 1.6 terabytes on the network probably end of this year, early 2027. And it will just keep growing and growing and growing, and these models will get bigger and bigger and consume more and more and more, I'm pretty sure.

Visibility and telemetry, you know, all my customers, the call that they get when a model fails because the network is a problem, is a different call than they're used to. So we put different telemetry and visibility in there to make sure that if things are going wrong on the network, that, you know, they know about it hopefully before they get that call.

Visibility & Tools14:48

Paul Gilbert15:13

So, you know, I work for Arista, our operating system is called EOS, and we have a whole bunch of features there. So if you were building an AI network, I'm not sure that you guys speak to the engineers, but this is the type of things we talk about.

Lossless Ethernet, everyone thinks that you, you know, when you train the model, you can't drop packets. I've seen it, and you can, I think drop packets are okay. Consistent latency is okay, but if you drop so many packets, obviously it's a problem.

So flow control, lossless Ethernet is really key. ECN and PFC are part of that. As I said before, they're flow control mechanisms. One is a slow down, please, and the other one is a stop. And as you know, because GPUs are synchronized, if something slows down, you slow down one port, one GPU, everything slows down.

So you really got to be on top of kind of the over subscription, and if you are getting queuing, where is it? We have really good buffer. We can adjust buffers. We have different kinds of switches for different places in the network.

But we've found that models send and receive a particular size packet, and what we do is we adjust those buffers to accept those types of packets. Buffering is a really expensive commodity in switches in networking. And if you can find a way to allocate the buffers exactly tuned to the packet sizes, it's a win-win.

And we've worked out how to do that, which is good.

Yeah, monitoring is really key for us. I tell my customers, there's probably five things you want to do. One of them is RDMA, you know, these networks train using RDMA, you know, which is memory to memory writes rather than going CPU to memory.

And RDMA is a complex protocol, and it has 10 or 12, maybe more, kind of error codes. So if the network starts seeing problems and starts dropping packets, rather than just drop the packet on the floor, we can actually copy that packet to a buffer or send it somewhere, or just the headers and why we dropped that packet.

And if you think about it, it's really cool. Like most networks will, in congestion, your buffer fills up, you're going to drop the packets. We'll drop the packet, but we'll actually take a snapshot of the packet and the headers and the RME information in it, and we'll tell you why we dropped it.

Another thing we have, which is really good, we have an AI agent. You know, from the networking point of view, we can look at what's going on, but we don't really have any visibility into the GPU. So now we have an agent, which is an API and some code that we load on the GPUs in NVIDIA, and they will speak to the switch.

So that agent will say to the switch, how are you configured? So PFC and ECN, those flow control mechanisms have to be configured correctly, because if they're not, it will be a disaster. So the GPU will speak to the switch and say, this is how I'm configured.

The switch will say, yeah, you're good, we understand each other. And the second thing it does, it gives you a whole bunch of statistics about packets received, packets sent, RDMA errors, RDMA issues in there. So you can correlate now if the problem is the GPU or if it's the network, which is a huge step forward for us.

Another really cool feature we have is smart system upgrade. You know, if you're used to routers and switches, you know, you have to upgrade the software sometimes. Sometimes to get new features, sometimes to fix PCERTs, which are security vulnerabilities on that switch.

We've worked out a way now that we can do that. You can upgrade code without actually taking the switch offline. So, you know, if you have 1,024 GPUs with 64 switches in your network, you actually can upgrade those and the GPUs can keep working.

So it's a real big step forward for us.

Design Guidelines19:07

Paul Gilbert19:07

So, you know, for us, again, I don't know, but no over subscription on the back end, you can't because the GPUs use everything you give them. Address-wise, it's really important for us. It's a point-to-point connection, so it's slash 30, slash 31s.

You could use IPv6 if you have IPv4 problems, address space problems. All my customers, I tell BGP, because it's the best protocol out there, it's really simple and it's really quick. eVPN, VXLAN, if you have multi-tenancy, if you have a lot of different business units, lines of business using the network, you need things like advanced load balancing.

We have a couple of different, we actually look at the collective that you're running, load balance on that collective now, which we call cluster load balancing. You could deploy at Rocky, I tell all my customers, do it, because if you don't, your network's going to melt down, you're not going to know why.

These things will give you an early warning system that you need to do something with your network. So they're really key to have. And visibility and telemetry is really good at all times, because in the network, not the operation center, you always want to be aware of the problem before you get the call from the developers and the people that've paid a lot of money for that network.

I'm running out of time here, but this is kind of a 1,400 gig cluster, what it would look like, Spine and Leaf. Again, no over subscription, 800 gig links between the Leaf and Spine, 400 gig down to the GPUs.

Clusters & Future20:24

Paul Gilbert20:38

This is a 4,000 cluster. These ones, these are the bigger boxes, these are 16 slot. One of these boxes can take 576, 800 gig GPUs, so 1,150 to 400 gig GPUs. So if you're building clusters with thousands of GPUs, then this would be the box for you, the 7800 series.

And putting it all together, this is kind of what we would build, there's three networks here. There's a back end network where your GPUs live, there's a front end network where the storage live, and then there's the inference that you take the model, you put it somewhere else.

I don't have a matter of time. I did not come in, so I'm guessing I'm not.

So the other thing is, you know, there's Ultra Ethernet consortium, you know, I don't know if this interests you. Ethernet hasn't changed the way it's built probably 30 years. There's some things it could do better around congestion control, around packet spraying, around the NICs talking to each other.

So there's this thing called Ultra Ethernet consortium, version 1.0 will be ratified probably Q1 2025. And it's a kind of different way of building networks, and you probably won't see them till Q3, Q4. But most of the cloud scale guys were kind of really keen on this because it puts a lot more into the NICs and takes a lot more out of the network.

So we just get, we can do what we're good at, which is forwarding packets.

Summary22:05

Paul Gilbert22:05

So summary, you know, for us, we have the front end, which is the storage, the back end, which is the really important part for us. That part is really bursty. The GPUs are all synced, so they send and receive at the same time.

And if you have a slow GPU, that's a barrier because it stops everyone else. Job completion time is what matters to us. If, you know, we get the call that, you know, my job completion time was one hour yesterday, it's four days today, you know, it's probably our problem.

You know, models can checkpoint, but they're really expensive, you guys know that. And I'm done.