AIAI EngineerJul 19, 2025· 19:03

Dream Machine: Scaling to 1m users in 4 days — Keegan McCallum, Luma AI

Keegan McCallum, Head of ML Infrastructure at Luma AI, details how the company's Dream Machine model scaled from 500 to 9,000 H100 GPUs within hours to handle 1 million users in four days, outpacing ChatGPT's initial growth. He explains that their initial Triton inference server setup was brittle and ill-suited for multi-GPU, multi-node video models, prompting a re-architecture to a custom serving stack on vanilla PyTorch. To solve work starvation across user tiers, they implemented an SLO-based aging system that ranks jobs by the percentage of their worst-case wait time elapsed. For managing dozens of model versions, they store immutable full Python environments and checkpoints in object storage, with a YAML file controlling active deployments and enabling zero-downtime rollouts across thousands of GPUs. McCallum also discusses partnerships with Nvidia, AMD, and Grok, and how Luma's broader mission is to build general multimodal intelligence that generates, understands, and operates in the physical world.

Transcript

Launch Day0:00

Keegan McCallum0:16

So it's 9:00 a.m., June 11, 2024, and we send out the announcement and hold our breath, waiting to see the user sign-ups pour in. We're expecting significant traffic for the launch of Dream Machine, Luma's first video model, and we were woefully unprepared for what came next.

We'd allocated about 500 H100 GPUs; we thought that was a lot at the time. It wasn't. And over the next hour, we saw request after request pour in, and a giant queue of requests start to pile up. Luckily, we had a contingency plan for this, which was to take every GPU we had access to from every provider and start manually running SSH commands against them to spin up workers pulling in work from a global queue.

And we were able to get to about 5,000 H100s over the next six hours. And our queues of almost 100,000 finally started to drain around like 2:00 p.m. that day.

And that is when Amit, our CEO and self-appointed chaos monkey, decided to tweet that we had scaled up by 10x. "Come on in. It'll be faster now. And we'll see how it goes." So our queues were down to about 300 at this point.

And you can kind of see me here freaking out a little bit, going, "OK, about 10 minutes after the tweet goes out, they start to go up again, even though we've scaled up by 10 times." So now we're at 350.

We're at 400. We're at 1,400. And as this is going on, we're basically taking the entire training cluster--this is the only GPUs we hadn't used yet--taking the training cluster over. That's another like 4,000 H100 GPUs. And it barely made a dent; the queues were still going up.

So you'll see this. It's the KeckW emoji, which has become a cultural staple of Luma over time. And it was very much the mood at the time. Like, what am I supposed to do with this? So I'm here to talk to you guys today a little bit about Luma, who we are, what we do, and how we've managed to kind of scale our infrastructure up to initially a million users in four days, which was pretty nuts.

For some context, ChatGPT hit a million users in five days. And we processed about half a million videos over the course of this 12 hours. So yeah, we're going to talk a little bit about what we learned scaling up and how we did it.

About Luma2:51

Keegan McCallum3:06

But first, very quick, a little intro to Luma if you're not familiar with us. So we're not just a video model company. We're a foundation model lab, and we're aiming to build general multimodal intelligence that can generate, understand, and operate in the physical world just like a human can.

And to give you some sense of kind of what our models can do and where we're at today, this is a feature we dropped yesterday, a demo video for it.

This is a modify video. All these videos are initially taken on iPhones and basically uploaded to our platform with a text prompt. And you can turn that raw content into anything you want. And for the AI engineers out in the audience, shout out to Chakran.

He manages our public API. And you can integrate this functionality into your applications using it very easily. You don't need to do any kind of crazy prompt engineering. We've taken care of that for you. Just send us raw user prompts, generative media, and we'll send you back images and videos that meet your user's needs.

Host4:12

It could be used, it could be free.

Keegan McCallum4:15

So if you're interested in that, definitely hit me or Chakran up. He's done a wonderful job building out our SDK and our whole kind of DevRel side of things. So yeah, that is my little plug for our API in case anyone's interested.

Old Infra4:30

Keegan McCallum4:30

So back to infrastructure. This is what our serving stack basically looked like when we launched. So we had a bunch of just tightly coupled containers working together. One of the benefits of this is that we could kind of just launch these on raw machines with zero other dependencies.

So that worked out well for us at launch. But there were some challenges scaling this up. So like any good engineer, I didn't want to reinvent the wheel initially. So we reached for Triton inference server, which is kind of a classic general-purpose model serving server.

But there were some issues with it. This setup was brittle. If Triton went down, the CPU processes didn't necessarily know that it went down. And so you'd be kind of pulling jobs, and they'd fail. It was annoying. Worse, though, with these video models, you're needing to run these on multiple GPUs and actually multiple nodes in a lot of cases to get to the latency you need.

And Triton's just not built for that. Also, we run our inference now on multiple different chip sets. So Nvidia is the company that builds Triton. They don't have great support for things like AMD or Grok or any of those.

And finally, the biggest kind of hurdle was that this was really difficult to develop against for the researchers. It had a whole bunch of different idioms, a whole bunch of kind of incantations you needed to make to make it work well.

And the overall setup was just it just felt very janky. So what we ended up doing was re-architecting to address some of these things and building our own serving stack on top of vanilla PyTorch for all the GPU work.

Custom Stack5:59

Keegan McCallum6:14

That worked out really well because most of the vendors that are building these different chip sets, they make sure that PyTorch is fully supported. It's kind of this great substrate to build on top of. If you support very vanilla PyTorch things, you can typically make your model run anywhere.

You may need to optimize certain operations depending on your model to make things fast. But it's relatively easy to get started. And in terms of the decoupled architecture you see here, the CPU workers being decoupled is quite useful and important because you can use them to queue up work and pull in when you're dealing with videos, images, multimedia inputs on top of just text.

You want those to be in the cluster ready for the GPUs to pull. So you're not blocking the GPUs at all. And also, with this architecture, you can actually run the GPUs anywhere. As long as you can connect to Redis and our distributed storage, you can kind of see their CWIDFS.

You can have a GPU in any kind of random provider, any random VM, and use Tailscale, connect to the rest of this architecture, and then scale up without having to do all that other kind of crazy parallel SSH stuff.

So you want to run compute on your training cluster. Now you don't need to provision special machines or anything. You just run a command. There were a few more challenges that we hit after we got through kind of the initial hurdles of this infrastructure.

Scaling Hurdles7:32

Keegan McCallum7:49

So one of the big ones was back pressure. So because this is decoupled, you can get into a situation and because there's multiple clusters that are pulling in work from the same global queue, you can get into this situation where you have too many CPU workers pulling in work to one cluster, and they're kind of waiting there to be processed by the GPUs on that cluster when they could have been processed somewhere else.

So we came up with this dispatch limitation system. So we came up with a state where if a GPU has been pulled into a sorry, if a job has been pulled into a cluster and is waiting to be picked up by a GPU, you can put a limitation on how many jobs are in that state so that you can avoid this issue.

I'll talk more in depth about priorities and our fair scheduling woes. That was another one. We have multiple tiers of users, and deciding whose jobs get processed first is a constant optimization challenge. Also, handling different models. So these video models are big.

They are typically made up of 10, 20 different submodels. So you're pulling in a lot of weights. You're spending a lot of time compiling these things. So traditional auto scaling is super wasteful. You're wasting 10, 20 minutes of GPU time just warming things up.

And finally, handling bursts. So we basically built this system to handle bursts where you could scale up automatically on our training cluster, which our researchers hate me for. It makes them very upset. But it allows us to keep up with the demand from our users.

And so that decoupled architecture, plus a very simple scheduler that runs on top of Slurm and can just run these PyTorch workers, handles increasing the total pool of GPUs when we need it and scaling down when we don't.

Fair Scheduling9:34

Keegan McCallum9:50

So this system is a pull-based system, and there are queues that submit to queues that pull from queues. And it is, as Soroush and Vasu can both attest to, they hear from Luma, it is the least enjoyable part of working at Luma is managing these queues.

It is everybody's least favorite task. So initially, when we launched Ray 2, which is a much bigger, more resource-intensive model, we began to actually have to deal with the fact that we couldn't just keep scaling up more and more compute.

It just wasn't economical or feasible. So we had to deal with limited resources. And when you have a pull-based scheduler with a bunch of queues like this, there's this concept of work starvation that comes into play, where essentially we've got API tier jobs, which we try to process very quickly.

We've got enterprise. We've got unlimited, plus, light, free. So we've got all these different tiers that have different priority. And who gets to go first? So if you just naively process them in priority order, there may be enough enterprise or API jobs that the light jobs never get processed.

So we were having people waiting for like seven, eight, nine hours, and they were very unhappy. So also, shout out to Soroush. He actually, on a weekend after we designed this system, in a couple of hours, implemented this SLO-based system, which I'll go into, that allows us to more fairly schedule work across the limited resources that we have.

So how this system works is when you think about

the concept of work starvation, one of the typical approaches you can use to manage this is aging. So the idea is the longer something waits in the queue, the higher its priority goes. But what's the actual function to control that aging mechanism?

So we had the insight that this is a product problem, and it really comes down to how long, worst case, are you OK with different tiers waiting? So we have service level objectives that the product team defines, and it kind of controls this aging behavior.

And how that works is an API job, we may not want to wait in a queue more than a couple of minutes. But a light job, maybe we're OK with them waiting for 10 minutes. And you can configure these and then set a threshold.

Once the threshold gets hit, so say 50% of that worst case timing, then that job gets pulled to the front of the queue. And initially, that worked allright. But then you can actually hit another case of work starvation if you have a bunch of jobs that are potentially breaching the SLA, where a job with a long SLO, since it's been waiting in the queue for a longer time, will actually starve out the resources of, say, an API job that you don't want to wait as long.

So how we handled that was we ranked the jobs by the percentage of their SLO that they were at. So if an API job is waiting for a minute, that gets treated the same as a light job that's been waiting for 10 minutes.

And that actually works out really nicely in practice and results in kind of intuitive fair scheduling behaviors.

The last piece I wanted to talk about is kind of how we actually manage all these models. So this was something where I was glad that we kind of reached for Triton initially. They had this nice concept of a model repo that we've really leaned into.

Model Versioning13:24

Keegan McCallum13:38

So every model

has a folder in object storage somewhere with a bunch of subfolders that have different versions. And so as you develop the code and fine-tune these models, deploy new checkpoints, you create a bunch of these immutable versions. And then you have a simple YAML file in the root of the model folder that lets you define which one's active.

So that's really nice because you can kind of reproducibly in these versions, you store the full Python environment, so all the dependencies needed to run the model and the checkpoints. And this system's pretty nice because if you ever need to roll back or you want to run things in all these different desperate environments, you can make sure that you're running the exact same Python environment, the exact same checkpoint that you were before.

And very rarely are there issues that you can't really solve by just rolling back. So that's been quite useful for us. And we've actually built an automated kind of rollout system on top of this too, where when we update those YAML files, the workers will just kind of switch versions on the fly without restarting.

And that helps us kind of roll out new model changes to the whole fleet of these thousands of H100 and AMD GPUs all at once, which makes managing this much more sane than the early days of parallel SSH.

And yeah, if any of this seemed interesting to you up your alley, we're hiring. We're actively looking for cracked engineers, researchers, AI enthusiasts in general. Please hit us up. And yeah, I just wanted to thank everybody for taking the time today and thank the team at Luma because it's incredible the work that everyone there does.

Wrap-Up15:13

Host15:37

We've got about three minutes for questions. So let's take one or two, and let's see where we get. Any questions?

Q&A15:37

Guest15:48

Let's talk about the ability to deploy different datasets and PyTorch being your world's common denominator. Could you dive into that a little bit more? Because each of these chipsets have their own compilers. How does all that work out?

Keegan McCallum16:04

Yeah, so the kind of cheat code here is that the chip providers who we typically partner with really closely, they're always making sure that PyTorch, at least at a certain version, works for their chipset. So they're doing a lot of that work for us.

And then typically what'll happen is we'll try to run the models on whatever the new chipset is. Usually, it'll work. And then it'll work a bit more slowly. So we've got actually a team of like 10 guys. We call our Excel team that are optimizing the low-level operations within PyTorch using things like Triton and working with the chip providers to make sure that things are actually fast.

So you'll typically with PyTorch be able to run things anywhere. But a lot of the times, they won't be fast. And then we just work closely with the chipset providers to actually optimize the model over time.

Guest16:55

Are you able to share with what chipsets?

Keegan McCallum17:00

Yeah, yeah. We work really closely with Nvidia and AMD. And we are exploring some other providers, but nothing really deep yet. Amazon's got their own chips. Grok has some chips. We just announced a big partnership with Humane, who works really closely with Grok.

And so yeah, we're exploring some of these other chipsets through some of these partnerships.

Host17:24

Cool. Any other questions? Yeah?

Guest17:26

I'm just curious if you're comfortable saying if you run on bare metal within your use?

Keegan McCallum17:33

We do not. Well, I don't think so. But yeah, we work with cloud providers that basically kind of give us a working Kubernetes cluster at the very least, which is nice. We're not provisioning the nodes ourselves or anything.

But I actually would have to check with the individual providers if there's VMs or if they're actually metal machines. I know when we work with Amazon, they're not. But some of the other ones might be.

Guest18:01

With Amazon, you said a bunch of keywords.

Host18:11

Yeah?

Guest18:12

Yeah, your video model is super impressive. But I saw on your website you're exploring other use cases. Have you shipped any visual QA models yet?

Keegan McCallum18:22

Not yet. The way the general application works does involve video and image QA models. So essentially, the way the actual application works is we have an agent that you're interacting with. So when you upload a video or an image, that is actually being captioned by VLMs to enhance the overall prompt.

But no true VQA stuff quite yet. But there may or may not be some coming.