AIAI EngineerJul 21, 2026· 18:02

The Desktop Frontier — Ahmad Osman, Osmantic

Ahmad Osman, founder of Osmantic, argues that within roughly 18 months (by late 2027) a single RTX 5090 will run intelligence equivalent to GLM 5.2, driven by the Densing Law of increasing impact per parameter. He shows this trend through concrete examples: a 27B-parameter Qwen 3.5 now beats the 405B LLaMA 3, and the same eight RTX 3090s that once struggled with LLaMA 2 can now run 15 parallel Qwen 3.5 agents. Osman presents the Densing Law—every 3.5 months, 50% fewer parameters achieve the same capability—as a systematic pattern, not coincidence. He advocates for sovereign AI: owning your own hardware (like a DGX Station or RTX 5090) gives you control, avoids cloud limitations, and sees hardware appreciate in utility as models become more efficient. He asks why fund cloud data centers when local hardware can run frontier intelligence and grow more valuable over time.

Transcript

Intro0:00

Ahmad Osman0:13

Hey everyone. We are about to start this presentation, uh, it's called the Desktop Frontier, and

it's basically about where we started and how far we've come with local and open-source models. Um, how many—like, just a quick question, how many of you here follow me on X?

I'm amazing. Love you all. Love you all. So, you know, I sometimes, every now and then, I would say a prediction. Here is a new one: within roughly 18 months we are going to have the equivalent of GLM 5.2 class intelligence running on a single RTX 5090 with 32 gigabytes of VRAM.

Um, that's basically late 2027. This is conservative. We might actually get there faster.

So, you know, for a long time, the story has been: bigger models, bigger models, bigger models. How can we get to the next 5 trillion? How can we get to the 20 trillion? And I'm not saying that there won't ever be, like, a gap between frontier intelligence and, you know, open-source models.

Impact per Parameter1:18

Ahmad Osman1:37

There will always be a gap. But that gap will shrink, and the efficiency of the models will get exponentially better.

So the term that I like to think about is impact per parameter. You know, what capability are we talking about? What could a model do? What footprint, like, hardware footprint did it have last year in comparison to now?

And what hardware does that use, and what hardware did it need to use a year ago? And, you know, are we moving down for the same kind of quality on that hardware? Again, as I was saying earlier, I used to run LLaMA 2 on an RTX 3090.

It's now running Qwen 3.5, 3.6, 27 billion parameters. That's better than LLaMA 3 4.05. That's a 400 billion plus parameters model that you beat with a 27 billion parameter model a year and a half after.

So, yeah, as I was saying, similar capabilities are moving into smaller hardware footprint.

Benchmark scores are one thing, but also, you know, a year ago, this time a year ago, we didn't have any local models that were able to successfully run within closed code,right? It wasn't until GLM 4.5 that came out in late July, and GLM 4.5 Air required at least four RTX 3090s or an RTX Pro 6000.

Now that footprint for hardware is not needed anymore. All that you need is a single RTX 3090, 5090, and you have something much more capable, much more intelligent. So is this trend just random, or is there more to it?

Densing Law3:41

Ahmad Osman3:41

That's a question that everyone should ask.

Is it just by random chance that we've gotten this far from models that weren't able to sustain more than 4,000 tokens in terms of context lengths, and now we have things that are a million

tokens locally on your hardware that you own? It's not by chance. You know, it's not just a coincidence that we got here. There is research being done. There are efficiency gains to be made. There are architecture hacks that compound, and they will continue to compound.

And I think I like this line: "It's not that small models are beating big models. It's that newer, more efficient models are beating older, less efficient ones."

So, yeah, capability density is, you know, the literature I back this up with. Nature Machine Intelligence calls this pattern Densing Law. And basically, you know, every three and a half months we are having 50% fewer parameters. Whether that's in dense or activated, that's a different story, but we are getting way more intelligence out of the models that we're running.

So, you know,right now where we're at, it's GLM 5.2. That's our, you know, biggest player. And it's 744 billion parameters total, with only 40 billion parameters activated. And it supports up to 1 million context lengths. You can run this in NVF v4 on a machine, on a DGX Station, or on a server with 8 RTX Pro 6000s.

Desktop Frontier5:10

Ahmad Osman5:36

That's something that you, like, a DGX Station is something that you can set under your desk, and it's running this kind of frontier intelligence. Whether, you know, it's on one benchmark, it actually beats GPT 5.5 extra high. Doesn't that mean that we're getting somewhere with local and open-source models, that we can compete with the frontier, that we're not that far off from the best that you can get from the cloud?

We also have Nimotron 3 Ultra, which proved that NVF v4 training, more efficient training can be done on hardware,right? That's very important. That means that the footprint, even for training these models, for fine-tuning them, for making smaller specialized models, as I was talking earlier, could be more efficient, could be done cheaper, and could be, you know, could deliver you value in terms of economics way sooner or, you know, for much less money than you used to.

Yeah. So, you know, again, LLaMA 2, that was a 70 billion parameter model. If you tried to run thatright now, you'd laugh at it,right? That used to take 8 RTX 3090s to load up, and those same 8 RTX 3090s could run something like 15 parallel agentsright now with Qwen 3.5 27B.

That's a massive jump in terms of performance gains. So the Densing Law basically means that we have similar or better capabilities with significantly fewer parameters. That's the impact per parameter, as I was saying. I want everybody to leave here thinking about this term and, you know, thinking, where are you going to get a year from today?

Sovereign AI7:19

Ahmad Osman7:19

As I was saying earlier, everyone here has a phone, I'm assuming. Raise your hand if you have a phone. If you didn't raise your hand, we know you lied about other things as well. So come on, guys. So, you know, you can now run GPT 4.0 quality on your iPhone.

That's massive. That thing required data centers to serve. So why wouldn't you invest, you know, in sovereign AI? Why wouldn't you, as a consumer, as an individual, as a small-sized business, middle-sized business, enterprise, why wouldn't you want to be in control of the models that you run?

Why wouldn't you want to make sure that nothing gets taken away from you? That every little thing can be optimized for you later on. That the performance gains can be made specially and specifically for your use cases, and that you can save more money that way in the long run.

And, you know, ODS for consumers, it's basically the way that we support individuals, but enterprises also. And I think that there is something that we, like, as a community, we need to think about deeply. We need enterprises for open-source AI to run.

We need these people that are using the cloudright now, that are basically supporting data centers being built for cloud providers, to come on this side, to own their own hardware, to own the stack fully end to end, so that we can keep delivering open-source models.

So that there is an incentive for open-source providers to actually come out with models, so that we can come up with new licenses that allow open-source to thrive.

So again, open weight and the frontier, I think I—yeah, sorry, that was a missed click. You know, so smaller models started bunching above the weight after LLaMA 2 with Mistral 7B, one of my favorite models. If you try to build that modelright now in closed code or open code, it's not going to work.

Open Weights9:09

Ahmad Osman9:30

But it used to take so much in terms of hardware,right, that you'd now get from a 9B model that I can run with Telegram or with Hermes, for example, and do a lot of stuff with. So we've come a long way.

We had that. We had Mistral 8 by 7B, which, you know, everybody knows is an MOE. Then the progression went from that to LLaMA 3. You know, LLaMA 3 8B was one of my favorites still is. It had unique identity, in my opinion.

Then we had, like, the 70 billion, which was, like, the thing that I would run basically on my 8 RTX 3090s at home. Then there was, like, the 4.05, the 400 billion plus parameter LLaMA 3, which again required a lot of hardware.

And if you put it now against Qwen 3.5, the 27 billion parameter would lose against it. That's in the span of, what, two years? Three years and some? No, I think less than two years. That's summer 2024 to March 2026.

That's about 21 months. And the next big thing, in my opinion, GLM 2 27B, and then we had the Qwen 2.5, and that was the moment that I was like, okay, we actually are making progress, and the gap was shrinking between open-source models and the frontier.

Really, LLaMA 3 saved, like, you know, it really helped us a lot. And then Qwen 2.5 delivered a massive improvement, and there was a lot of fine-tuning and experiments that could be done on that one. There were amazing papers, and they helped the community immensely, in my opinion.

Reasoning11:12

Ahmad Osman11:12

Then the next big thing was DeepSeek R1, in my opinion, and reasoning becoming something that you can run at home. That was a massive MOE, almost 700 billion parameters. You know, you had to have, like, a very beefy server to actually get it up and running.

And then, you know, the improvements that came from just more training on that one and DeepSeek R1 that was released in May last year made a massive jump again. So it showed that post-training could deliver more improvements on the same checkpoints.

Then GPT open source, like, GPT-oss 120b. Anyone remembers that one from last summer? Yeah? Nobody here used it? Come on, guys. I need some help here.

It was one of the first open-source models that were able to successfully do tool calling, and it was a step forward. It showed us that we can do more with the hardware that we have running at home. That was a footprint shift,right, from, like, you know, that massive 700 billion parameters DeepSeek R1 that was, yeah, 671 billion parameters to something that was one-fifth, one-sixth of its size.

And GPT-oss was comparable, maybe better, more agentic performance.

Then the moment of Qwen 3.5, the 397, the 397 billion parameters. That's a beefy MOE. And, you know, what's funny is that about, it's about 15 times the size of the Qwen 3.6, and I'm here, I'm comparing 3.5 to 3.6 of the dense 27 billion parameter model.

And that dense model beats it. And that dense model has 40% higher number of activated parameters. So it's not that far off. That's a massive amount of performance gains in a very small amount of time with massively different footprint in terms of hardware requirements.

And that trend happened in, like, what, two, three months? So, you know, how far could we go from here? How far before we get to, you know, a recent model that there was some news about, you know, that is finally relaunched again?

How far before open-source delivers something of that quality that you could run on your own hardware and you can control and will not be taken away from you and will not refuse a request from you?

So again, these are just some benchmarks where you can see that an iteration on the 27 billion parameter model, a little bit more post-training, proved it across all benchmarks and made it run against a model that is almost 15 times its size.

And again, remember, this is 27 billion parameters activated versus 17 billion parameters activated. It's still massively the same amount. Like, you know, it's only 40% less in terms of the amount of time it would take to process things, but it's 15 times smaller.

That's a lot.

So again, how long until the prediction I made earlier becomes plausible when I said that we're going to have the equivalent of GLM 5.2 running on an RTX 5090? This is the math. 17 months. And this is a conservative math.

Prediction Check14:53

Ahmad Osman15:12

Earlier this year in December, I had a very viral post that I predicted that we're going to have the quality of OBS 4.5 running locally at home on a single RTX Pro 6000. That happened by March.

So a question. Hardware purchase today, does it get more valuable as models become more efficient and smaller in size? That's a good question. So why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens and then later on get.

Hardware Value15:33

Ahmad Osman15:57

Those subsidies are going to go away and you're not going to be able to run those models and they will have so many limitations. So might as well ask yourself, why not own the hardware yourself and be in control?

So yeah, the forward-looking question is basically, what will a DGX Station be able to run in 3, 6, 12, 18 months from now? That's something that there is a reason that I'm not selling any of my RTX 3090s if you follow me.

And I have a lot of hardware, guys. But I'm interested in seeing what I could do with them in a year or two from now, more than in the amount of money I would get for them today. This is not financial advice, by the way.

Let me make that very clear.

Own Your GPU16:37

Ahmad Osman16:37

So yeah, the desk-side frontier potential, you know, an NVIDIA DGX Station could run a lot of today. It could run GLM 5.2. What will it be able to run tomorrow? Six months, 18 months, two years from today? We know that, you know, RTX 3090s, the Amber architecture from 2020, sells at a higher value than MSRP today and is still being utilized for a lot of use cases.

So what will a DGX Station, the actively developed Blackwell architecture, will be able to run in a few months, a couple of years? That's a good question. So the question you have to ask yourself, if an RTX 3090 with 32 gigabytes of VRAM runs in the equivalent of a GLM 5.2 in 18 months?

And this is the question that everybody should be asking themselves. And I want you all to be looking at this screen, taking this very seriously. Okay?

Should you buy a GPU?

Thank you.