AIAI EngineerJun 27, 2025· 16:07

How fast are LLM inference engines anyway? — Charles Frye, Modal

Charles Frye presents benchmarks from hundreds of runs on Modal comparing open-source inference engines VLM, SGLang, and TensorRTLM across models like Qwen 3 and Gemma 27B, arguing open weights models have caught up to proprietary ones, making self-hosting viable. He shows, for example, that Qwen 3 (MoE) on VLM achieves ~1 request/sec with 128 input tokens and 1024 output tokens, while switching to a RAG-like workload (1024 in, 128 out) yields a 4x throughput improvement. Frye warns that optimizing for context over reasoning can improve latency without sacrificing quality, and notes that the engines' out-of-the-box performance varies by model—e.g., SGLang underperforms VLM on Gemma due to less optimization. He also highlights the gap between prefill (parallel) and decode (autoregressive) speeds, which a rationalist would expect from transformer architecture. The benchmarks, available at modal.com/llmalmanac, aim to help engineers choose hardware and engines, with contributions welcome for optimized configs like TensorRTLM's knobs.

Transcript

The Shift0:00

Charles Frye0:15

Thanks, everybody, for coming. Um, yeah, I wanted to talk about some work I've done recently on trying to figure out, uh, just how fast these inference engines are when you run open models on them. Uh, so the. Kind of been talking at AI Engineer since it was AI Engineer Summit two years ago, um, and the for a long time it's basically been the, like, OpenAI wrapper conference,right?

It's like, because just because, yeah, what am I going to do? Am I going to run an agent with BERT? Probably not. Um, and that was, like, it was exciting to talk about all these cool new technologies, see people building stuff like Cursor on top of them, or Devin, um, but for me, as somebody coming from, like, having trained my own models a lot, it was like, oh man, I want I want to I want to touch the weights.

I want to play with them. I want to, like, hack them. And the quality but the quality just, like, wasn't there yet to do, like, some of the interesting stuff. Um, and that's changed. You've got the Llama series, we've got the Qwen series, we've got DeepSeek, and so now we're, like, catching up to, like, maybe even literally catching up with the Frontier Labs, which would be pretty crazy.

Um, but we're at the very least, like, at the point where a lot of things people have been talking about at AI Engineer for years are possible with open weights models, where they weren't before. Um, and at the same time, there's also the development of the software stack on top of that.

Um, also, as somebody coming from, like, writing my writing and running my own PyTorch models, I was like, how hard could it be to run a language model? Like, you just, you torch nn.module, wrap a class around it.

And, like, yeah, I mean, having done stuff with Transformers before, it's like, oh yeah, there's definitely it's a little more complicated with Transformers. Training looks weird. Inference is different. Um, but, you know, pretty quickly, the state of play has advanced a lot on how to run a Transformer,right?

KV caching is, like, you know, just the first thing, and then it's this, you know, um, page detention. Now, multi-token prediction, speculative decoding, all this stuff. That's, like, pretty hard to write yourself. And so, you know, you want software to do that for you, probably.

Um, and so these open source engines are available now. You've got VLM and SGLang and TensorRTLM. So the combination of those two things has, like, kind of flipped the, uh, playing field around to where there's, like, you'd need a really good reason to run your own models.

Like, you're the U.S. government or something, um, or you want to run on an air-gapped system, um, or you, like or you had to, like, believe it in your heart. You know, you had to want open source models like the, um, like NuSaR Prime Intellect.

You had to be, like, a decentralized crypto bro or whatever to want to run your own open models. Um, but yeah, now the, the, like, uh, the situation has changed. It's really exciting. It sort of finally makes sense to self-host.

Um, so, uh, the just, like, want to do this really quickly. Does anybody here at this AI Engineer Summit 2023, the, like, first one? Anybody? Okay. Did anybody come to the AI Engineering 201 workshop the day before, um, uh, in the 500 Street Avenue?

Maybe not. I talked to one or two people who were there. Um, so I gave a talk on, on, like, you know, how to how to do your own AI stuff back then. Just wanted to pull out, like, a couple of slides.

Um, so, like, the one of the main this, like, 2023, we're like, who's going to win, open models or closed models? And the key statement, uh, in the talk was, like, if capabilities requirements saturate, open models will catch up to proprietary models and then dominate for those cases.

Inspired by what you see with operating systems, databases, programming languages, like, as soon as there's, like, a sort of, like, you know, um, capability level that everybody, like, that you don't need the absolute best thing. You just need something that's good enough.

Open, uh, like, collaborative projects tend to catch up and, and then have better properties. And that's happened with open models. So check. Um, uh, yeah, so that, um, that was mostly around, like, capabilities. Um, and, you know, at the time, there was only Llama, but now we've got a lot.

Um, the other one was, yeah, talked a little bit about LLM inference libraries at the time, and most of them are gone. TGI, RIP, um, uh, for example. But VLM was good then. It's stuck around now. Um, yeah, so, uh, I don't know what the slides in 2027's AI Engineering, uh, conference are going to look like, but, um, at the very least, like, three of my slides from two years ago wereright.

Um, so, yeah, um, hopefully it's not those three slides again in two years. Um, but yeah. Allright. So what does the LLM engine landscape look like two years later? Um, let's, uh, take a look. So what, uh, what, you know, we were advising a bunch of people on how to run like, people were coming to us, like, I want to run my own code completion editor.

LLM Almanac4:47

Charles Frye5:08

In editors, I want to run, like, big backfill jobs to, to, like, um, enrich data in databases with language models. Uh, it's too expensive or, uh, to run it on OpenAI, or I train my own model, so it's too expensive to run it on a, like, a provider like Fireworks.

They want to run it on, um, more generic infrastructure, like what we have at Modal. Um, and so they would people would come and they'd be like, allright, well, how fast can you run an 8 billion parameter Llama model with SGLang, uh, like, with 128 tokens in, 1,024 tokens out on a Tuesday when Mercury is in retrograde?

Um, and that would, like, take at first, it took a couple days to, like, re you know, figure out how to make sure all the packages are working, that we've got the fastest versions installed, and, um, like, uh, that we can give, like, a, you know, trustworthy number.

Got that down eventually to, like, you know, like an hour or two. Um, and then built some benchmarking software so we could get it done in about, like, 15, 20 minutes. But then, like, you know, the people ended up asking a lot of similar questions.

So we decided to just, uh, you know, fifth law of, uh, fifth mantra of performance is, uh, do it when they're not looking. Um, so, like, compute the thing ahead of time and store it. Uh, so we ran a giant benchmark over, um, like, 10 or so different models on VLM, SGLang, and TRTLM on about 10 different context lengths, um, and put that all up on the internet.

So let's take a look at that. See, I'll drop into this one. This is a live version of it. So you can find this at modal.com/llmalmanac. Um, the idea is that this is just one page in the, like, you know, your almanac, your little book that has the useful things you need to know to be an LLM engineer.

So to start, we've got our benchmarking results, uh, benchmarking methodology in detail, the open source code for it, and a little executive summary. Um, hope to put more stuff up there as we accumulate the things people need. Um, things about, like, speculative decoding and multi-token prediction and quantization.

Um, but yeah, so to start off, we got this little interface here. So, uh, yeah, anybody what's a model people want to see results for? Any if, uh, hopefully that's legible to people. Anybody got a favorite? Hmm? Qwen 3?

Live Demo7:27

Guest7:29

Ministral.

Charles Frye7:30

Ministral? Okay. Uh, excuse me. Uh, that oh man, that's that's my boss. I'm not going to do that one. Um, okay, we'll do any engine here. Oh yeah, we didn't do, uh so this was fun. SGLang's Qwen 3 support is a little buggy for the, um, for the 8-bit quant that we ran.

So I think we only have results for VLM. We'll stick with any engine there. Oh yeah, by the way, if you if you try this thing out, um, like, you'll see, like, if there's a giant tensor of configurations,right?

And we'd love to have that full tensor, but it doesn't always it's, like, either not always possible or, um, like, it's not clear how to do it. So there's a place where you can contribute configurations so you can build up a nice big database of, like, how to run these models.

We also haven't, like, carefully optimized any of these things. We started with out-of-the-box performance for all the engines, just because, like, optimizing 100 configurations is going to take some time. So we'd love, uh, like, contributions of optimized implementations, especially TensorRTLM, which has, like, a ton of knobs.

Um, and they have names like UserBuffer. Like, what is that? Um, yeah. Um, okay, yeah, so first token under 1 second. Let's say this is this is a pretty common SLO. Like, you want 1 second. It's a nice round number.

People feel like that's, like, a good amount of time to wait. I'd say 300 milliseconds is a tighter one. That's more, like, interactive. That's your, uh, Doherty threshold, if you're a fan of Holt and Catchfire. Um, made-up number, but, like, if somebody repeats a made-up number enough, it's a real number.

Um, so yeah, 300 milliseconds. Okay, so we can get a throughput of about 1 request per second on Qwen 3 in the Mixture of Experts model on VLM for 128 tokens in, 1,024 tokens out. Um, so that's what we got.

Over here, we got, like, a little, uh, code snippet here. So it should be the case that you can UVX modal run this, and if you have a token, that should just work immediately. Uh, and by immediately, I mean after 5 minutes of loading the model weights and, and spinning up the model server, but that's as immediate as it gets.

Um, cool. Allright, so that's one result. Uh, who's asking for Qwen? Are you satisfied? Yeah? Okay, great. Um, Gemma327b. Allright, on this one, let's do any engine here. Allright, this was allright, sorry, we got a tight filter on this.

I'm going to put the first token filter up. Oh yeah, this one, we're doing the BF16 quant is the only one that we could get working at first. So I think we eventually got the 8-bit quants working. I don't think that's a hard blocker, but it was the easiest one to get going was the BF16.

So these are definitely slower. Um, so you'll see 27 billion parameter models, so, like, 10x smaller in, uh, model weights. Roughly the same number of active parameters as Qwen 3, but we're getting, like, about the same, um, like, throughput, request per second on the same load, 128 in, 1024 out.

Um, so yeah, so it's interesting. You can see sort of which ones have had more optimization work on them. I think the Qwen 3 models and the Llama model series, you see a lot more optimization. Um, this is also one of the ones where you saw the biggest gap between, um, SGLang and VLM, the Gemma one.

So it looks like the VLM team spent a little bit more time, or Google's contributed a little bit more to VLM on, uh, getting good results. Oh yeah, let's go yeah, so the other thing you'll see is you generally like, so I just switched.

Sorry, I should say what I'm doing. So this is 128 tokens in, 1024 tokens out. Uh, and you can see we're getting about 1 request per second on this guy. Um, let's just and, and the first token comes back in 400 milliseconds.

Let's flip it to 1024 in and 128 out. Right? So this is going from, like, a reasoning workload to, like, a RAG workload. Big scare quotes on that, but it's just the difference between whether you're dominated by decode time, like, more of your tokens are decode, or more of your tokens, uh, prefill.

Workload Impact11:04

Charles Frye11:22

Um, and what you'll see very consistently in these results is that you get much higher throughput if you have more, like, tokens in the context as opposed to tokens being generated. Very straightforward. If I mean, if you know your Transformer architecture, it's, like, auto-regressive versus parallel.

Like, yeah, one of the first things you would learn if you looked at the, you know, kind of the implementation of the architecture. But it's nice to see it, like, nice and, you know, very cleanly. More of an empiricist than a rationalist myself, so I like to see data, um, and not, like, uh, chalkboard stuff.

Um, so yeah, so, uh, what I'm getting at here is that the, um, the request per second that we're seeing here is about 4 requests per second for VLM on the same workload, but with, like, context instead of, uh, um, generation as, uh, where the majority of the tokens are.

Um, so this is, uh, yeah, I gave a talk on GP like, GPUs a little bit earlier today, and, like, one of the big takeaways there is find things that, like, are throughput-oriented and evolve a lot of arithmetic and not, like, moving memory around or communication.

And that's exactly the difference here. You have, like, big matrix-matrix multiplications, load the weights one time, use them a bunch. Um, and that's exactly the difference here, and it's a 4x improvement. And that's using BF16, which does not have Tensor Core support?

No, no, BF16 has Tensor Core support, but it's the slow Tensor Core support compared to FP8 or FP4 on Hopper and Blackwell. And so you're, like, you're the real win there is the shorter, like, shorter numbers, faster multiplication.

It's actually quadratic in the bit width, so you get a big win as you go down. So this like, if we were to run some results with FP4 on Blackwells, you would see an even bigger gap than just this, like, 4x improvement.

4x is, like, barely enough to wake up for, you know? Um, but yeah, so that's a little like, not every application can you, like, change that. Um, like, your users might be bringing queries to you, so you don't have control.

Um, but the, um, uh, it can it's more the sort of thing where you, like, a product person is like, we should improve the quality. You're like, oh, how can I improve the quality without ter like, killing our latency?

Don't have it like, don't immediately reach for reasoning. Reach for context instead, because it's going to be cheaper and you're going to get better perform like, you're, you're going to find it easier to hit your latency SLAs. I forgot to point this out, but the latency is, like, almost identical in time in time to first token, even though we're doing 10 times as many tokens.

Basically a free lunch. Um, yeah, okay, so that's, uh, that's sort of, like, how I envision people using this interface and the data. Um, there is a there's a URL somewhere where you can just download the raw data.

Um, if you're interested in that, hit me up. The code is also open source if you want to run some of these benchmarks yourself. Um, I think I have I'll close there. Lots of other stuff to talk about, like, our benchmarking methodology, which is written up here, the, uh, executive summary, which you can share with your, um, with your leadership, um, uh, on, like, running open models.

Q&A14:07

Charles Frye14:27

Um, but I'll take a question or two before we close out. Yeah.

Guest 214:31

I mean, how far can you drive the, uh, max throughput? I said you had it set to 4?

Charles Frye14:36

Yeah.

Guest 214:37

Like, how far can you go out before it starts to inflect up?

Charles Frye14:39

Yeah, so that so one thing I'll say is, like, this is throughput per replica,right? So this is one GP is one GPU? Yeah, one H100. So, like, what you the way you solve your, like, total throughput is by scaling out rather than scaling up,right?

Um, but, um, so if you want if you want 400 QPS or whatever, like, eventually you're just going to have to scale out. But to your question of, like, yeah, how do you know, like, you know, where like, why are we saying this is the highest throughput you can get?

Guest 215:10

Is that.

Charles Frye15:11

Yeah. So the answer is, like, goes to our benchmarking methodology. What we do is first we dump, like, uh, you know, 1,000 requests and wait for them all to come back, calculate the, like, thou uh, seconds divided uh, requests divided by seconds, 1,000 divided by how long it took.

That's, like, a maximum throughput,right? Because we gave it maximum we exposed the maximum parallelism to the engine, so presumably they knew how to they were smart enough to handle that. Gives you a maximum RPS. Any more than that, you should expect from queuing theory that the latency will blow up,right?

Then the other side is, um, like, you send one request at a time, wait for it to come back, send another. And that gives us our, like, that's the fastest you could possibly run the server, and we sweep between to get the numbers that are here.

But yeah, cool. Allright, I'll, uh, got to move on to the next talk. I'll, I'll be outside if you, uh, have any questions. Thank you very much.