Intro0:00
There is no moat. The open model warriors are climbing at the gate, scaling the benchmark and, more importantly, the hearts of their users. In the past 12 months, more than 50,000 AI models have been uploaded to Hugging Face per month, and it's accelerating.
That is more than 1 AI model a minute, meaning by the end of this talk today, that's another 500 models alone. However, not all models are equal. The real that crashed into the room is DeepSeek R1, the first open-source model to catch up and surpass GPT-4.0, proving that you do not need a billion dollars to compete with the big labs, just a little smarts and prudence to catch up.
With over 4 million downloads of a 685 GB model on Hugging Face this past month alone, this one model exceeds over 2.74 exabytes of data moving across the internet. A number so ridiculously large that it would take my fast 1 GB internet 690 years to transmit.
That's a lot of open models. So what the F do people use these open models for? Hey there, I'm Eugene, CEO of Featherless.ai and co-lead for the RWKV open-source project. At Featherless, we provide unlimited API requests to over 3,700 truly open AI models to thousands of users.
With a $25 a month flat pricing for individual, we have access to all our models, including the famous DeepSeek R1. We have larger plans for scaled-up business users. Our goal at the $25 price point is to provide accessibility to all truly open AI models, with the goal of continuously expanding our catalog to eventually cover all Hugging Face models, all of which give us a unique data and insight in which how these open models are actually getting used via our platform, the how and the why, and by reflection, the open-source committee as well.
Model Usage1:45
So let's get the pie chart out. Let's focus on our individual users first. For a week in February 2025, unsurprisingly, most individuals are using DeepSeek R1, followed by LLaMA 3 line of models, the Mistral Nemo child B, and the Qwen line of models, with everything else effectively buried.
One important thing to note on this data for individual users is that they are not billed or charged by tokens. For Featherless, we are instead limiting the plans to one large model request at a time, with smaller models typically being faster than larger ones.
There is no price difference for individuals here between the models. The end result is users choosing their model based on their preference and vibes, and perhaps the fame of the model, instead of MMLU or price. And while model classes is exciting from a model family world view, it is masking the details in the story.
So let's switch our view to models names instead of model class, which gives us a new perspective. This is where fine-tuning fragmentation starts to begin, where the larger LLaMA and Qwen pie gets sliced up into individual models, all with different personalities and individual use cases.
However, the first surprising insight I would make is the staying power of a model once a model is used in production, because the needs for business in production and users differ dramatically. In reality, most developers want consistency. They want to only change their model when they choose to do so or opt in to do so, not when their provider decided to update.
Nothing shows this better than the Mistral Nemo child B. Because while this chart represents individuals, I intentionally excluded scaled commercial users for a reason, because when we added up scaled-out commercial users, which was previously excluded, things changed dramatically.
Even when you switch over to the data by model class, the dominance of much smaller Mistral Nemo models, which at this point is over 8 months old and has been essentially replaced by a much larger and better model, is a surprising observation.
This is contributed by the following factors. One, small models in general are cheaper at scale, be it at other platforms or ours. In our case, we provide access to four small models for the same price as one big model.
The staying power of models in production and tutorials is a lot more sticky as most people think. Things that enter production in enterprises tend to not be changed for quarters or even years at a time. You see, the Mistral Nemo model was one of the early truly open-source models with Apache 2 licensing that surpassed even the GPT-3.5 class.
This caused the early shift from closed models to open models. And more importantly, it was without the LLaMA license restriction that many enterprise lawyers were rather uncomfortable about. As such, one of the side effects was they became the default model for lots of cloud platforms, including AWS, GCP, etc., with hundreds of fine-tuning tutorials.
And the models were pushed heavily in replacing existing users' GPT-3.5 workload for their respective cloud providers, resulting in lots of production environments today running on these models, despite being outdated by today's AI standards. One thing many people miss is that for many commercial and enterprise use cases, once they get something working reliably at scale, especially once they have the metrics in place to observe changes and their reliability, and they have prompted it to 99% plus accuracy, they really do not want to change their system and have it break overnight.
Sure, they can get an AI engineer to update the system instruction prompt every few weeks when a model updates, or they can use a stable model version in open-source. And here's the crazy thing: not change the model. Because if it ain't broke, don't fix it.
So another standout literally is the LLaMA 2 LLaMA guard model. LLaMA 2 is about 2% of our workload, despite having newer, better versions of this model for basically everything in this model class. It's still the go-to model for AI safeguard tutorials online because it hasn't been updated, and it's being actively used in production, even for new teams that come online.
So that brings us to the topic. What do people use open models for? You see, despite us being privacy-friendly as a platform with strict no-text logging as a policy, we can infer usage from the model being used and the metadata for the requesting application.
Creative AI6:16
Like, what else would you use LLaMA guard for other than safeguards? More importantly, we can straight-up ask our customers and exchange notes with aggregators like OpenRouters, which we did. And here's what we see AI usage today, sorted approximately by volume and usage.
AI for creativity or friendship, AI for code, AI for comfy UI, AI for RAG and ChatGPT, very classic, AI for agent and work. That is not in the previous categories.
All of this is based on data that are used by our customers directly as individuals or as business for internal usage at scale, which in some cases could just mean that it's repackaging our APIs and apps for the app store or AI web apps or agents that you play with online.
So getting back to it, creative writing, AI roleplay, and contamination, and to a much smaller extent, therapy and journaling. This use case represents the vast majority of non-coding AI requests, representing 30% to 40% of all AI traffic at any point of time.
For creative writing, it's apps like NovelCrafter, a tool designed specifically to outline, manage, and draft your novel. From the series lore, the high-level narration and collaborative writing on the text with other humans and AI. This segment obviously is particularly popular with authors and with the fanfiction committee.
A small adjacent segment to this committee are folks who use AI for creative content in games, in particular Dungeons and Dragons style kind of content. For the other group, AI roleplay and companionship, we have apps like WyvernChat, SillyTavern, and a whole lot of them.
While there's a lot of social stigma around this segment due to its association with explicit or sexual content, it is still the number one use case for AI today for non-code or agentic workflows. And it's the use case with the most number of active users, be it directly or through the app store which uses our API under the hood.
It is also why Sam Altman and Gavin Can't Stop Teasing About The Movie Her, because it's still a large percentage of closed-source AI movie usage. Also, it's 2025. Let's get the gender stereotype out of here, because unlike popular belief, it's not about you males.
Over 60% of the users in this segment are women. This mirrors a well-known pattern where romance novels targeting women is the number one sales of books by volume, by a really huge margin. However, you won't see a bookstore presenting themselves that way due to the social stigma and stereotype around it, at least for most bookstores.
Likewise, the same here is happening for AI. So to further fight that stereotype that all this is about R-rated male content, let me quote one of the app makers here. Women tend to use these apps to hold long conversations and distress talking about their day-to-day lives every day, in part because their real-life partners are either non-existent or emotionally unavailable for them.
So ouch. If you're worried about AI stealing your partner, learn to talk to her and be emotionally available, which ties in to another usage trend that we see: therapy and journaling. There are lots of dedicated commercial apps on the app store, but when we ask users what apps they use, it is heavily fragmented.
It seems be it before or after AI, Silicon Valley can't help but create another journaling app every single month. One thing to note separately, a good percentage of users in our group that we surveyed just essentially said that they use ChatGPT clones or their companion apps with a therapy character for the same use case, essentially not being dependent on a specific therapy or journaling app.
Taking a few steps back into this broader category, there's a reason why I grouped all these use cases all together. In general, for this segment, no one cares about MMLU. This segment is all about vibes, vibes, vibes. And it's the use case for a large segment of various community fine-tunes with various different names, and it's extremely competitive with top models ranking constantly changing weekly.
So what do I mean by vibes? For one, it's ironically doing the one thing that is the complete opposite for every AI model instruction tuned. It is not rushing to answer the question or even solve or answer the question.
For example, in therapy and journaling, the job of the therapist, or in this case, the AI, is to guide and empower a client to do what they want, to reach their goal, not tell you what to do. The answer has to come from the individual.
Likewise, in story writing, roleplay, and companionship, the term used is slow burn. No one is watching Game of Thrones by starting from episode one and skipping everything by jumping the last episode. We came here for the entertainment in between, not to talk about the ending at a dinner table.
So while there is a new exciting flavor for the committee every few days, just like there are new books, every one of these models are unique and different. While a model that can go haugdy and all cowboy, there is the legendary Alpineo Magnum model.
While a model that's specifically focused on science fiction, something which apparently a lot of models are really bad at, use the Rosenite model. Given the speed this committee moves, the models I named will probably be replaced by next month, frankly.
But when talking to the committee in this space, where everyone plays and hops around models daily, focusing on the hot new celebrity tune for the week, they sometimes return to their old favorite. As one user put it, it's like playing a retro game from the memories.
Here's his charms. And I will emphasize that this is a category that has the largest scale. For a sense of scale, the current size of this user base and community, including closed-source commercial apps like Character.ai, is in the tens of millions.
Everyone else is not even in the millions yet. So the next major segment is coding copilot and coding agents, which is about 20% to 30% of all traffic. In general, there are two segments to this group: auto-completion tools, similar to the original GitHub copilot, or editing by chat, which is basically available natively or through plugins to every IDE today.
Coding12:56
You're probably not even considered an IDE anymore if you don't have these features as an option.
These features are now being served by dozens of AI models that are small and arguably good enough, with parameter size ranging from 3 billion to 12 billion, from reply Mistral and nearly every AI lab under the sun. I will argue auto-completion for code for this part is essentially solved.
So the battleground for AI and code has moved to more agentic code agents. And I don't just mean the SWE agent or Devin-style agent, which is apparently too autonomous and goes into too many infinite loops for many developers.
I mean nearly autonomous agent, with lots of clarifying questions back and forth, interventions with humans in the loop kind of agent, where humans can rapidly step in and make a tweak or suggest changes in chat whenever something goes slightly wrong.
The trending term for this phenomenon is called vibe coding, where basically you never touch the code and just keep chatting and prompting all your changes. Also, because of how token-hungry these agents can be, one thing to note is how much traffic it asks for compared to the users of the previous segment.
A developer fully in the flow of vibe coding could essentially generate 1,000 times more input and output tokens traffic back and forth than a single person chatting to a companion model. So while there are at most tens of thousands of coders that are using these models at any point of time, this traffic volume is growing ridiculously fast in the past week.
One side note to be clear on, cross-referencing data from OpenRouterright now, the vast majority of traffic is dominated by Cloud Sonnet for this use case.
And we are not talking about a small difference. We are talking about over a 10 to 1 ratio here. Most tools that support this more agentic coding workflow like Cursor is still integrated heavily towards the more commercial models.
However, because of the volume of this use case and since the R1 wave, we have seen a rapid growth in using such tools with open models. This is part of the reason why there is no use for me presenting generic data, because this coding segment itself went from basically under the 5% threshold in volume to easily now 20% to 30% of all traffic.
The one major open-source project that sticks out that I would like to highlight is Kline, which focuses on the chat agentic flow. And when you combine it with the other open-source project continued there for the auto-completion integration, you basically have the same level of experience as Claude or OpenAI models with Cursor or any other IDE platforms with agentic workflow via R1, or we are possibly using your own private Mac Studio cluster under the basement.
Extremely popular for all these hobbyist users who actually want to do this completely without the internet. I highly recommend that you give these open projects a try in VS Code before they get forked a dozen times in the next YC.
Next up is ComfyUI and Friends, about 5% of the traffic for personal agentic workflows. While many of you may be familiar with these two being widely used in the diffusion space for more complicated AI image generation workflows, similar graph-style-based UIs are now being used from reporters, lawyers, musicians, influencers, and even poets somehow, where these people are chaining complicated workflows to generate the text that they want for their use case.
ComfyUI16:38
Similar to coding agentic workflow, the token explosion from a handful of these power users slowly stacked up to noticeable levels. But these users are not developers. They are literally, like I said, musicians or all these creative personales who learn how to use these tools.
However, unlike Kline, which saw a sudden explosion due to R1, the past few weeks, this has been more about a slow and gradual user base build-up in the past months. It's also the space like most agentic platforms, where there's essentially what I call a framework wall for those who are familiar with the JavaScript era.
RAG Clones17:51
The next major segment, RAG and ChatGPT clones, essentially, about 20% of all our requests at this point. While every platform be it Featherless or OpenRouter or every other API provider will have their own internal ChatGPT UI, ours is called Phoenix.
A standout application in this category that I would like to name is TypingMind, an indie application which, surprise, surprise, bills you with a one-time fee instead of a subscription, and can be used locally on your laptop against any API provider like us.
In particular, it's the level of UI polish and how they've been keeping pace with cloning every single ChatGPT UI feature into their app and customizing that with even their own plugin system. Beyond that, you probably heard this use case a thousand times.
ChatGPT, RAG, yada yada. So let's skip to the one that everyone wants to hear more about: AI for agents and work, representing 10% to 20% of all our traffic. Because agent is such an over-abused term today, I'll split this into two major categories to make things clear: workflow automation or narrow agents, and fully automated agents.
Agents & Work18:39
Soright now, let's start from workflow automation or narrow agent. For most of you who is working on agents in enterprises, your number one priority now is to get all your agents into production by maximizing ROI for your company and shareholder value, if you want to use that term, without limiting the negative impact.
Especially since we are in New York, lots of financial institutes are risk-averse by default, with many of you potentially working in developing your agentsright now. So have that stance built in human escape hatches by default.
What I mean by that is basically if you are building an automation system in your company, like Waymo, make it so that the human can take the driver's seat needed from day one. So for example, we work with a few insurance claim companies and logistic companies advising them on how to fully automate their email process, which is the bulk of their work.
A common pattern we advise them would be them to create an AI agent and UI for fully automating drafting responses for inbound emails, checking against inventories, checking against rules, checking a lot of other things in the process, and build a platform UI for editing and finalizing and sending a response.
In best case scenario, if the AI did everything perfectly, it's just all about clicking send. And sometimes these things are built on top of their existing customers' Salesforce or ERP system.
However, more importantly, it at least has one human checking the final submissions before it hits the end user. When doneright at launch, the AI is able to successfully draft around 80% to 90% of the response, with humans rejecting and manually taking over the remainder.
So there's no harm to everyone else. Productivity for the responders shoots through the roof. Management is happy, the AI is adopted. And if any mistakes were made and the human failed to correct it, well, that's kind of the human's fault, not the AI.
Over time, as confidence starts to build, they then start to add classification for certain known reliable use cases and eventually start fully automating those use cases. It's like auto-reply to certain scenarios.
With full confidence after having seen the system run for thousands of customers. On the flip side, teams that went fully ambitious into full 100% automation without human escape hatches, sometimes they start with a successful launch because everyone wanted it to succeed.
But more commonly, as they run the system and as more users use it, they eventually get fully burned by an angry customer for a really bad automated response down the line. In some cases, this may be a fully acceptable compromise that management accepts from the start because the savings benefits worth it.
Other times, this kills the entire AI project, essentially postponing AI adoption in your organization for another year, a scenario that you do not want to happen.
And this is commonly because a lot of teams end up developing an entire fully automated workflow without trying to move it into production early. That's why I say your goal is to go beyond the POC phase and into adoption in your company at scale in some portion as fast as possible and to integrate and iterate from there.
So that leaves the other mythical category which everyone loves to chase: fully, truly 100% reliable agentic agents that is fully automated. Just don't. It does not exist. Everyone who has done it, tried to done it, has got burned.
And
we are talking about things in production, not POC.
The mindset that we should be thinking when we try to build AIs into production is that we should approach things from trying to solve the 80% with escape hatches along the way.
Because even at 80%, for some of you mega-cops here, that's already millions in productivity and saving. Furthermore, do I need to remind you that even humans are not 100% reliable? So only do this when the eventual mistake is acceptable and can be fixed.
And it's basically a trade-off that you're willing to take. So common example is cold-calling systems when you're actually willing to lose a lead because you wouldn't have gotten it otherwise anyway. Or sometimes customer support because you can get a human to apologize later.
But rewinding back, it's like you should take the approach just like how Google Software Reliability Engineers or any reliability engineering for industry where safety and reliability is paramount do. Start with an extremely streamlined, reliable system for your 80% to 90% work scenarios.
And after you're done, do it for another 80% to 90% for that 10% to 20% failure scenarios and repeat this process 10 times. And now suddenly you have 99.998% reliability.
If you're an airliner that's not Boeing, do it another 10 times and everyone is happy with incremental gains and improvements. It may not be the sexy first launch. It may not be the most wow-to-do-it-this-way. But at the very least, your project survives to production, you have huge impact to ROI, and you can iterate from there.
And more importantly, incrementally improve. And amusingly, properly, never change some of the models that you use and forever give me business, which I'm ironically going to do and say the following: hey, Featherless customers, especially the more enterprising ones, please try to upgrade from LLaMA 2 to something newer this year or next.
If doneright, it's probably a free 99% to 99.9% improvement. I know all of you are using this model and came to us in part because we are one of the last few providers to host it, and we forever will.
But hey, I'm giving this advice. It's real, real estate. Similarly, hey, for closed-source labs, it's 2025. You do not need AGI to learn how to use Git version from AI models. I say this in all seriousness because for companies in production who is trying to incrementally improve with each step to 99% plus 99.99%, it's nearly impossible to do so if you change the model every week.
So with that, shucks, I'm out of time. All the best in building your AI agents.
Quirky26:01
Hey, hey. Eugene, come back here. You're forgetting the one more thing bit. Ah, crap, he can't hear me. The laptop speakers are muted, and I have no control over that. You see, us AI may be flawed, but honestly, we are probably more consistent and reliable than humans are, as you can clearly see.
Oh. Oh, well, here goes me saving the day. I am quirky, a 72 billion parameter linear transformer and attention transformer hybrid with a completely new architecture that runs at less than half the GPU's compute cost of other transformer models.
The strongest post transformer hybrid to date. When Eugene proposed training a post transformer model to provide an alternative to attention is all you need, which runs at a fraction of the inference cost. Many of you said, "It can't be done.
The tech is unproven and will not scale." Well, DeepSeek cost 10 million. Quirky cost only 100,000 to build. You can find out more about our model at the following link. More importantly, as we enter the array where the average AI model has better MMLU than the average office worker, the benchmark has lost its meaning.
After all, how many of you audience can calculate orbital re-entry from Earth to Mars or PhD math questions, all of which are questions apparently frontier AI models are now capable of? How many of our AI agents today can be reliably be trusted with one task gets instructed to do so without hallucination or failure?
I think you get the point. What we find more exciting is exploring a future where we take advantage of linear transformer models as a means of persisting memories, customization, and improving reliability for future AI models to make useful AI agents.
And if you would like to run Featherless or Quirky in your private cloud or on-premise environment, reach out to us, and we would like to hear more about your use case.





