AIAI EngineerJul 18, 2026· 7:52

Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs

Thiyagarajan Maruthavanan of Kalmantic Labs argues that AI teams should stop renting inference from providers like Anthropic and build their own infrastructure, coining "rent to learn, own to earn." He recounts his app UltraSuno costing hundreds of thousands of dollars in inference, a stolen key hitting $10,000, and moving to his own DGX Sparks hardware. Three enterprises—a fund, hospital, and tax practice—each hit walls with renting: control, audit redlining, and reproducibility. He open-sourced JustTokenMax (benchmarked better than Netflix's Headroom) and wrote the book "Peak Inference" on building your own inference infra. He notes the market's conflicting pitches (Jensen's token factory, Nadella's unmetered, NeoClouds' endpoints) but concludes owning is essential post-PMF.

Transcript

Cost Crisis0:00

Thiyagarajan Maruthavanan0:00

One of the largest retailers in the country spent close to $200 million on inference with Anthropic, and decided that things got way out of hand and built their own infrastructure. I'm pretty sure most of you have read the news from Uber CTO on how they had planned a budget of their tokens for an entire year, and it got over in month 4.

I'm also confident that half of you in this room have come to a very similar conclusion: that as time goes by, the cost of intelligence really builds. Using inference feels like, you know, it's one of the most inexpensive things.

But then this is very different from using a phone, where you get a bill once every month and then you have, like, a specific set of amount that you can actually anchor your mind to. But in case of using these rented intelligence platform, they are, like, prepaid.

You load credits. It's almost as if you're loading credits inside a casino: you put some, and then you pull it, and then you are so addicted to it, then you end up doing more and more of it, and by some time you realize that you've blown past the threshold that you had mentally kept in mind.

UltraSuno1:00

Thiyagarajan Maruthavanan1:00

And I had this experience myself. I built an app called UltraSuno, and I experienced the inference cost ballooning here. Suno.com—has anybody heard about Suno.com? Yeah, Suno.com is this application that allows a user to turn a text prompt into music.

What I was interested in is doing the reverse, which is: given a particular song, what prompt could have actually generated it? This is something that I wanted, so I built this. And I was having a lot of fun using this application, shared it with a few friends, and it spread wide around, and it had hundreds of thousands of users.

But then the cost ballooned way more than what I had anticipated. Hundreds of thousands of dollars had to spend on inference. Now, this happens for many reasons. There are many talks that are there at AI itself, where people talk about how you need to manage your context better.

And many people forget about doing compression of their input token. And when there are agent loops, then there are many of these calls that are happening which are very, very wasteful. The inference endpoint that is consuming this is completely unaware of the shape of the workload, and which is why this happens.

And I have this other issue that had happened 3 weeks ago: my key got stolen. Someone in China got hold of it and then was sucking my endpoint dry. I could see the cost rise up from $7,000 to $7,500 to $8,000, and so on and so forth.

Thanks to my co-founder, who heads our research and technology, we were able to arrest it at $10,000, otherwise it could have been $100,000. Now, many people suggest that the alternative to rented intelligence platform is to use token factory.

Token Factory2:35

Thiyagarajan Maruthavanan2:35

Token factory is basically saying that why are you paying money to Anthropic and OpenAI? Instead, go open source, have these open source models that are already deployed somewhere on the cloud, and then they are provisioned as tokens per second.

There are NeoClouds, and then there are inference endpoint providers who actually do this. In fact, there is also an argument saying that, you know, you can build this token factory locally. There are AI Twitter influencers who actually talk about building inference in your garage, in your basement.

Buy GPU cards, rig them up together, and then you could actually run a local token factory. In fact, I was inspired by that a little bit. I bought my own DGX Sparks, and then I first moved UltraSuno from Anthropic to DGX Sparks.

DGX Sparks3:04

Thiyagarajan Maruthavanan3:19

It worked well. I ran into this one issue of memory being the bottleneck. And then it was good enough that I started building my next applications. I started having agents. I have some agents that I need for running my research lab, so these agents started shaping up inside the DGX Sparks.

You know, 2, 6, 8, 12. And it worked allright. The issue, though, is that, you know, it may not be reliable for enterprise, which is what I exactly faced. Three enterprises reached out to me to replicate the same setup for them.

Enterprise Walls3:47

Thiyagarajan Maruthavanan3:47

But for enterprises, renting and leasing don't cut it. Bill is a problem, but then there are a secondary set of problems that makes it an extremely ineffective approach. The enterprises that reached out to me: one was a fund, another a hospital, and the third a tax practice.

And each of them had a different wall that they had hit. The fund, it was an investment fund. They were running an investment analyst on a Nemo claw architecture, and they didn't want somebody else to dictate as to what the rate limit that they could consume.

So control became a big issue for them to actually go with token factories. Hospital had a different issue. They used the use case, it worked well, but later when they went through an audit, a third-party vendor dependency was redlined, and then they couldn't go forward.

The tax practice was a completely different issue. In a tax practice, what is happening is when an intelligence generates a recommendation, you want to be able to recreate it. And when you don't have access to the in-depth of the model, you will not be able to do this.

Rent vs. Own4:47

Thiyagarajan Maruthavanan4:47

And that became the issue. So that brings me to the most important point of this presentation: where do you sit? When do you stop renting infrastructure? If you're a startup, if you're a founder who is doing pre-product market fit work, so you're still figuring out that the use case that you have, if there is demand for it, you can get by by renting.

But if you're post-product market fit, you cannot afford to do it. And if you're an enterprise who's already budgeted a project, which means you're telling that you are assuming that this particular use case has product market fit, then again, you cannot ignore to build your own infrastructure.

Which is what I realized. And I said, like, this situation is like, if you're going to a new city, you may initially start with saying that I don't want to buy a house, let me actually rent and see.

Sometimes you might even Airbnb. You experience the environment, you experience the city, the neighborhood. But then eventually you have to buy the house. You cannot raise a family in an Airbnb. As I went through this experience, I decided, I came to the conclusion that I need to build my own inference infrastructure for the apps, the agents, and the scaling of the apps that I'm building.

And I call this just infer. And while I went through this exercise, I realized that there is optimization to be done at multiple layers. Even at the renting and the lease layer, you can do optimization around input cost of token management, and then, you know, context management, and so on and so forth.

Some of those experiences that I've had in the last couple of months, combined it into an open source project and published it as just token max. If you have used Headroom from Netflix, then this is an alternative to it.

We've benchmarked against Headroom, and on many parameters, just token max is far superior. If this is a thing that is of interest to you, give it a try, maybe a GitHub star if you like it. And I also wrote the book called Peak Inference: Infra Economics of AI Inference, when you have to think about building your own inference infrastructure.

The AI market is very different compared to the rest of the technology market that used to exist, because here the rules of the game change every 3 to 6 months. Which means it becomes a very noisy marketplace. You talk to someone like Jensen, he would say token factory is the future.

Market Noise6:41

Thiyagarajan Maruthavanan6:56

You hear someone like a Satya Nadella, he will say unmetered intelligence is the future, it is going to be local. And then when you hear NeoClouds and inference endpoint providers, they'll say, hey, inference endpoint providers are the ones that are going to capture the value in the marketplace.

Now, my experience walking from application to agents to scaling them, and then building my own inference infrastructure, taught me that if you want to learn, you can rent. But if you want to earn, then you have to own.

Own to Earn7:09

Thiyagarajan Maruthavanan7:22

And if there was the one sentence that you were to take away from this entire presentation, it is that: rent to learn, own to earn. But then you have to come to your own answers. Thank you. And if any of these topics are of interest to you, then I'm happy to talk to you about renting, about just token max, about how to build your own inference infrastructure.

I'm here at the AI Engineers Conference for the next 3 days. Hit me up on Twitter@MTRajan, or through my site, mtrajan.com.