Intro0:00
Nice meeting you guys. Uh, great to be here, and I'm here to present Hyperbolic, which is an AI cloud for developers. And so my topic is: why we don't need more data centers. It's like a very eye-catching title, but what I want to clarify is, I still think building data centers is important, but just building data centers alone can't solve the problem.
So, uh, wait, before we get started, let me introduce myself. I'm Jasper, I'm the CEO and co-founder of Hyperbolic. I did my math PhD at UC Berkeley, finished my PhD in 2 years, which made me the fastest person in the history of Berkeley.
And then I also won a few Gold Medals. So after that, I worked at Citadel Securities, trying to use AI and machine learning to predict the market execute strategy. So I always have a passion about how to make things very efficient and how to help you to save money, because everyone knows that compute is actually one of the biggest costs for your companies or for your startups.
Usually, you need to, if you want to rent like 1,000 GPUs, we'll spend millions of dollars per year. And we think that these problems should be solved by not just building more data centers, but actually building a GPU marketplace.
The Problem1:19
So let's get started with the problem that we're facing. First, I think, so everyone knows that AI is going to integrate with everything in the future, and every company will be AI companies. So the demand for GPUs, as well as data centers, is exploding.
So by McKinsey, by 2030, we'll need 4x more data centers built in one quarter of the time that we built in the speed.
But what if I tell you that you actually don't need that many data centers, you actually need another solution. So we can break down the demand first. Right now, the current capacity for a data center is 55 gigawatts.
By the median scenario, we're going to see 22% annual growth rate for the demand. So in 2030, we're going to need 219 gigawatts.
Building Hurdles2:38
And however, it's like, there are a lot of challenges building data centers,right? So first, we everyone knows Stargate. So it takes, like, for the first Stargate data center, it takes like more than a billion dollars to build. And then also, it's very slow to connect data centers to the electrical grid.
For example,right now, the waitlist is like 7 years. So we need to wait 7 years to connect a 100 megawatts facility to the electrical grid in Northern Virginia. And then it also very consuming a lot of energy. So currently, we're spending 4% of the total electricity consumption in the US for just GPUs and data centers.
And also, it's not very environmentally sustainable. If you can look at the number, that's crazy CO2 emissions annually. And even say if we're going to deliver all the data centers on time, there's still a data center supply deficit of more than 15 gigawatts in the US alone by 2030.
And so it means that just building data centers can't solve the problem. On the other hand, we think the GPU utilization is actually pretty low. So according to Deloitte, GPUs sit idle 80% of the time for enterprises and companies.
Idle GPUs3:49
According to Sammy analysis, there exists 100-plus GPU clouds. So we can see, like, how fragmented the space is,right? A lot of you guys need GPUs, but you can't find them, or like, you're going to pay an extremely high price.
Marketplace4:28
On the other hand, there are a lot of GPUs sitting idle in data centers or in different clouds. And so naturally, a solution that we think we should build is actually build a GPU marketplace or like an aggregation layer that aggregates different data centers and GPU providers to solve the problem for GPU users.
It doesn't necessarily need to be Hyperbolic, but I just use Hyperbolic as an example to show here.
So I can just, like, share what we are trying to solve. So we're building this global orchestration layer. We invented a software called HyperDOS, which is short for Hyperbolic Distributed Operating System. So basically, it's like Kubernetes software. So any cluster, as long as it installs our software within 5 minutes, suddenly the data center becomes a cluster in your network.
And on the other side, users can rent GPUs in different ways that they want. Like, they can just do the spot instance, they can like on-demand, they can long-term reserve, or they can also, like, host models on top.
And so, like, we see that there are like several benefits. One, we kind of like solve the, like the matching problem of compute. And then second, like GPUs become commodities. So you're like, you don't need to spend too much time to wait for data centers, you just buy them on the marketplace.
And then third, you can have different options.
Cost Savings6:09
And so we do some math modeling. I mean, I don't have time to kind of put down the math in the slides, but this is our conclusion,right? Basically, we can save the cost by 50% to 75%. Even if you look at the current, we're running like some beta version of our marketplaceright now, and our GPU cost for H100 is $0.99 per hour.
But if you look at Google, for example, they have on-demand GPU, it's like $11. They're like Lambda, they have like $2 or $3. But on average, by aggregating more supply and then like have a uniform distribution channel, you can drastically reduce the price.
It's like the theory behind that is like the queuing theory, basically like is MMC theory. I probably next time if we're going to watch my talk, I will share more math behind that. But yeah, and then like you can just save time to vetting your suppliers because if you like think about, I mean, how many people here are founders or like need to acquire GPUs?
Yeah, so are you frustrated when you are trying to talk to, how many suppliers are you talking to? If you have talked to more than five, raise your hands. Are you frustrated when you're like trying to have like five sales calls and like try to like know which data GPUs are frustrated?
Yeah? Good? Yeah, that's good. Yeah, so basically by having like this uniform platform, like founders or like startups or companies no longer need to vet different data centers, they just like pick the one that they have high rating or like have the best price.
We're also going to do like benchmarking on the performance of the GPUs.
Allright, so,
oh, sorry.
Allright, so, sorry, somehow the graph didn't show.
Give me one sec.
Use Case8:29
Yeah, so basically we can think about a use case example. So let's say if you are a startup and you want like 1,000 GPUs at the beginning. So usually you will just reserve these 1,000 GPUs for a year,right?
You think like, I might need to use these GPUs for training and later on I want to do inference. And so you run some training jobs and then after 3 months, then you realize that, okay, now I have a good, a better idea by running those experiments and now I need 1,000 more GPUs just for a month,right?
And then after 6 months, a month takes, then you finish your training job. And then you realize that now I only need 500 GPUs for hosting my model, but I still have 500 GPUs left. So on the traditional, on Hyperbolic case, you basically can say, okay, I will rent 1,000 GPUs for a year at the beginning.
But then in month three, I can say, I just rent an extra 1,000 GPUs for just a month. And then a month, in month six, then I can say, okay, I can release my idle GPUs on Hyperbolic and try to sell them to other people that need them,right?
But if you just like use some traditional cloud, then you need to rent 1,000 GPUs at the beginning and then in month three you need to rent extra 1,000 GPUs for a year usually. And if you calculate the cost, compare that and then also like think about the price difference you will have, you can reduce the cost from 43.8 million to 6.9 million.
So it's like 6x saving. And you also help other people to get cheaper GPUs too because you can release those idle GPUs to other people. And so this is how we think that we're going to like increase the productivity.
Like people only think about saving, but actually this is not true for GPU,right? By scaling law, we know that the more compute you spend, the better quality your machine will be, your model will be. So it's not just about saving your cost by 6x, it's more about with the same budget, you will increase your productivity by 6x.
Productivity10:36
And imagine how many startups that they used only need to rely on OpenAI and Anthropic, those closed AI models, but now suddenly their money becomes more valuable and they can rent as many GPUs as they want for their training.
All-in-One11:19
And so the next step that we think usually the GPU marketplace will evolve into is that it will be an all-in-one platform for different AI workloads. Because what people really want is not just GPUs. They want to run their different AI jobs,right?
You will have AI inference, online inference, offline inference, and then you will also have training jobs.
Takeaways11:47
And so, yeah, so this is like to like some takeaway, like basically we don't think we need like just focus on building data centers. We also need to do like smarter allocation for the resources. And then second, we can reduce your cost by building GPU data marketplace.
And lastly, I think just focusing on building data centers is not very sustainable. We're costing a lot of energy, taking a lot of land. We should better reuse, recycle those idle compute by selling it to others. So if you're interested in trying out, you can come to our website.
The left QR code is the current product that we have, which is a marketplace. But then we're also launching our business card and enterprise card that gives you like production-ready GPUs with 99.5% reliability. Allright, thanks.
Awesome. So I actually got, I'm curious, can you tell us more about the kind of Hyperbolic OS? How exactly does that turn? Because I know a lot of times you have a data center, you have a cluster set of GPUs.
Q&A12:53
How does it actually work to connect it to Hyperbolic itself?
Yeah, so basically this is, HyperDOS is like a Kubernetes agent. So you just install that in your cluster as long as you have Kubernetes. I mean, most data centers have Kubernetes, but then even for your MacBook or for your PC, you can just install like microK8 to kind of become a Kubernetes-ready machine.
And so basically now you kind of have, we have terminology in-house. We call like our Hyperbolic server Monarch, and then we have different barons. So it's like a feudalism model. So different barons, they own different compute. And then every time when a user wants to rent GPU, they will talk to our Monarch server, and the Monarch server will send a request to like the baron, and then baron will just basically provision the machines and set up the SSH instance for customers to access.
Yeah.





