AIAI EngineerOct 9, 2024· 18:55

Making Open Models 10x faster and better for Modern Application Innovation: Dmytro (Dima) Dzhulgakov

Dmytro Dzhulgakov, CTO of Fireworks AI, argues that open-source models are the future for GenAI applications because they offer lower latency, lower cost, and domain adaptability compared to proprietary models. He explains that open models can be up to 10x faster for narrow domains and cut costs significantly, using examples like fine-tuned Llama 3 for function calling. Fireworks addresses the challenges of setup, optimization, and production readiness with a custom serving stack that delivers the fastest inference for long prompts and image generation (e.g., SDXL). He highlights FireFunction V2, an open-source model for function calling that combines chat and tool use, and notes that Fireworks serves over 150 billion tokens per day for companies like Quora and Cursor. The talk emphasizes that platforms like Fireworks enable developers to start with serverless inference, fine-tune models, and scale to enterprise-grade deployment with dedicated hardware.

  1. 0:00Intro
  2. 0:48Our Roots
  3. 2:05Trade-offs
  4. 5:12Challenges
  5. 6:27Fireworks Stack
  6. 8:50Model Hub
  7. 10:33Compound AI
  8. 12:49FireFunction
  9. 15:25Platform
  10. 16:51At Scale

Powered by PodHood

Transcript

Intro0:00

Dmytro Dzhulgakov0:15

Hello everyone. So, my name is Dima. Uh, as mentioned, unfortunately my co-founder Linh, who was on the schedule, couldn't make it today because of some personal emergency. So you got me. And as you saw, we don't have yet AI to figure out videoer projection, but we have AI for a lot of other things.

Uh, so today I'm going to talk about Fireworks AI, and generally I'm going to con- continue the theme which Kathleen started about open models, uh, and how we basically focus on productionization and customization of open-source models in inference at Fireworks.

Our Roots0:48

Dmytro Dzhulgakov0:48

But first, as an introduction, what's our background? So the founding team of Fireworks comes from PyTorch leads at Meta and some veterans from Google AI. So we combined have, like, probably a decade of experience in productionizing AI in some of the biggest companies in the world.

And I myself personally have been core maintainer of PyTorch for the past 5 years. So the topic of open source is really close to my heart. And since we kind of led this revolution of open source toolchain for deep learning through our work on PyTorch and some of the Google technologies, we really believe that open source models are the future also for, like, for Gen AI application.

And our focus at Fireworks is precisely on that. Uh, so, I mean, how many people in the audience actually, like, use GPT and deploy it in production? And how many people, how many folks use open models, uh, or deploy the production?

Oh, okay. So I was about to convince you that share of open source models is going to grow over time, but it looks like in this audience it's already, already sizable. But nevertheless. Uh, so why, why basically this trade-off?

Why go big or why go small? Uh, currently still, like, bulk of production inference is still based on proprietary models. And, uh, the catch is that those are really good models and often frontier in many domains. Uh, however, the catch is that it's one model which is good in many, many sins.

Trade-offs2:05

Dmytro Dzhulgakov2:22

And it's often served in the same way regardless of the use case. Which means that maybe if you have batch inference on some narrow domain or you have some super real-time use case where you need to, you need to do, like, voice assistant or something, those are often served from the same infrastructure without customization.

In terms of model capabilities, it also means, yeah, like GPT-4 is great or Claude is great and can handle a lot of sins. But you are often paying a lot for additional capabilities which are not needed in particular use case.

You don't really need customer support chatbot to know about 150 Pokémons or be able to write your poetry. Uh, but you really want it to be really good in a particular narrow domain. So this kind of, uh, this kind of discrepancy for large models leads to several issues.

One, as I mentioned, is high latency because using a big model means longer response times, which is particularly important for real-time use cases like voice assistants. It gets more and more important with agentic stuff because for stuff like, for example, Next is going to be a Devin,right?

Like, you need to do a lot of steps for, like, something like agent-like application to do reasoning and call the model many times. So latency is really, really important. And often you see that you can pick smaller models like Llama or Gemma, which you just talked about, and achieve there for a narrow domain, same or better quality while being, you know, up to 10 times faster.

Uh, for example, for some of the function calling use cases, like externally benchmarked from, uh, from Berkeley, yeah, like the you get similar performance from fine-tuned Llama 3 at 10x speed. Cost is also, uh, is also an issue.

If you're running a big model for, uh, on a lot of traffic, you know, even if you have perhaps, I don't know, 5K tokens prompt and 10,000 users and each of them calls LLM 20 times per day, you know, on GPT-4, even on GPT-4o, it probably adds up to, like, 10K per day or something like several million per month.

Also, several million per year, which is a sizable cost of a startup. You can easily cut that with much smaller models. And that often we see as a, uh, as a kind of motivation for reaching out for smaller and more customizable models.

Uh, but really the, uh, like, where open models shine is domain adaptability. And that comes in two aspects. First, there are so many different fine-tunes and customizations. I think Kathleen was mentioning about, you know, Gemma-built Indian languages adaptation.

Like, there are models specialized for code or for medicine. If you had to hug and face, there are, like, tens of thousands of different model variants. And because the weights are open, you can always customize to your particular use case and tune, and, uh, tune quality specifically for, for what you need.

So open source models are great. So what are the challenges? Uh, the challenges really come from three areas. Uh, first, like, what we usually see when people try to use, you know, open model, something like Gemma or whatever, uh, or Llama 8B, uh, you run into complicated setup and maintenance,right?

Challenges5:12

Dmytro Dzhulgakov5:28

You need to go and find GPUs somewhere. You need to figure out which frameworks to run on those. You need to, like, download your models, maybe do some performance optimization and tuning. And you kind of have to repeat this process end to end every time the model gets updated or new version is released, etc.

Uh, on optimization itself, uh, there is, especially for LLMs, but generally for Gen AI models, there are many attributes and settings which are really, really dependent on your use case and requirements. Somebody needs low latency, somebody needs high throughput.

Prompts can be short, prompts can be long, etc. And choosing the optimal settings across the stack is actually not trivial. And as I show later, in many cases, you can get multiple X improvements from doing from doing this efficiently.

And finally, like, just getting it production ready is actually hard. Uh, if, as you kind of go from experimentation to production, even just babysitting GPUs on public clouds is not, not as easy because GPUs are finicky, not always reliable.

Uh, but getting to enterprise scale requires, you know, all the scalability technology, telemetry, observability, etc. So those are, uh, things which we focus on, uh, solving at Fireworks. So starting with efficiency, we built our own custom serving stack, which we believe is one of the fastest, if not the fastest.

Fireworks Stack6:27

Dmytro Dzhulgakov6:46

Uh, we did it, uh, did it from the ground up, meaning from writing our own, you know, CUDA kernels all the way to customizing how the stuff gets deployed and orchestrated on the service level. And that brings multiple optimizations.

But most importantly, we really focus on customizing the service stack to your needs, uh, which basically means for your custom workload and for your custom cost and latency, uh, latency requirements, we can, we can tune it for, uh, for those settings.

What does it mean in practice? What does customization mean in practice? Uh, for example, many use cases use RAG and, uh, use very long prompts. Uh, so there are many settings you can tune actually on the runtime level and the deployment level to optimize for long prompts, which often can be repeatable.

So caching is useful or just tuning settings so the throughput is higher while maintaining latency. So this is independently benchmarkable. If you go to, uh, you know, artificial analysis and select long prompt, where Fireworks actually is the fastest, even faster than some of the other providers which are over there at Expo Booth.

Uh, and, uh, we don't only focus, we don't only focus on LLM inference. Uh, we focus on many modalities. Uh, as an example, for image generation, we are the fastest provider serving SDXL. We're also the only provider serving SD3, uh, stability as a new model because their API actually routes to our servers.

Uh, and finally, as I mentioned, like, LLMs, like, customization, especially for LLMs, customization matters a lot. Uh, one requirement, like, one paradigm we have to think about performance of LLMs often is useful for use cases is to think about max, like, minimizing cost under a particular latency constraint.

We often have customers come in saying, like, "Hey, I need to, like, have this interactive implication. I need to generate that many tokens under 2 seconds." And that's where, that's really where, like, cross-stack optimizations shine. Uh, whereby, uh, tuning to a particular, like, latency cutoff and changing many settings, you can deliver much higher throughput, multiple times higher throughput, uh, which higher throughput basically means fewer GPUs and lower cost.

Model Hub8:50

Dmytro Dzhulgakov8:50

Uh, in terms of, in terms of, uh, model support, we support, support best quality open source models. Uh, you know, we heard about Gemma now, obviously, Llamas, uh, some of the ASR and text-to-speech models, uh, pretty much from, from many providers.

We also work with model developers. For example, uh, for example, YiLarge, uh, in the US is also served on, uh, on Fireworks, launched last week. And, uh, as a kind of platform capabilities, as I mentioned, we have a lot of open source models to, uh, to get you started or customized ones.

We do some of the fine-tuning of those models, uh, in-house. So I'm going to talk a little bit about function calling specialized models, uh, later on. Or we do some of the, uh, vision language models using, uh, ourselves, which we release as well.

And of course, the key for open source, uh, open model development is that it can be tuned for a particular use case. So we do provide a platform for fine-tuning, uh, whether you are bringing your data set collected elsewhere or collecting it live, uh, with the feedback when served on our platform.

Uh, specifically on customization, it's like one, uh, interesting, one interesting feature which a lot of people starting to experiment with models, uh, find interesting is if you try to fine-tune and deploy the resulting model, how, uh, how to serve it efficiently.

Uh, turns out if you do LoRA fine-tuning, which a lot of folks do, uh, you can do, uh, smart tricks and deploy multiple LoRA models on the same GPU, uh, actually thousands of them, which means that we can give you still serverless inference with paying per token, even if you have, like, thousands of model variants, uh, sitting and deployed there without having to pay any fixed cost.

Compound AI10:33

Dmytro Dzhulgakov10:33

Of course, single model is all great. Uh, but what we see increasingly more and more in applications is model is not the product,right, uh, uh, by itself. You need a kind of bigger system, uh, in order to solve target application.

And the reason for that is because, uh, models by themselves tend to hallucinate. So you need some grounding. And that's where, like, RAG or access to external knowledge bases comes in. Uh, also, we don't have, you know, yet an industry magical multimodal, uh, AI across all the modalities.

So often you have to kind of chain multiple types of models. And of course, you have all this, like, external tools and external actions which, uh, uh, kind of end-to-end applications might want to do in agentic form. Uh, so I think the term which I really like, which is popularized by Databricks, is, like, compound AI system.

Basically, increasingly seeing, like, transition from just the model being the product to kind of this combination of maybe, like, RAG and function calling and external tools, etc., built together as the product. And that's pretty much direction which we kind of see, uh, this field moving along, uh, over time.

So what does it mean from, uh, from kind of our perspective, what we, what we do in this case? Uh, so we see kind of as a function calling, like, agent as a, at the core of this, uh, emerging architecture, which might be connected to either domain specialized models, uh, served on our platform directly or maybe tuned for different needs and connected to external tools.

Uh, maybe it's a content interpreter or maybe it's like external APIs somewhere. Uh, with really, like, this kind of central agentic, uh, view, uh, kind of central, central model kind of coordinating and trying to triage the, uh, user, user requirements if it's, for example, a chatbot or something.

Uh, you probably all, uh, heard about, like, function calling, you know, popularized by OpenAI initially. That's, uh, that's basically the same idea. Uh, so yeah, the function calling is really like a how to, how to connect, uh, LLM to external tools and external, and external elements.

What does it mean in practice? So we actually, uh, focus on fine-tuning models specifically for function calling. So we released a series of models like that. Like, the latest one of FireFunction V2 was released two, uh, two weeks ago.

And, uh, what you can do with that is, uh, oh, okay, if I manage to click, if I manage to click on this button. Uh, what it means is, like, you can build, uh, applications which kind of combine free-form general chat, uh, capabilities with function calling.

FireFunction12:49

Dmytro Dzhulgakov13:08

So in this case, uh, this is, you know, this is FireFunction has some, uh, chat capabilities. So you could see you can, like, ask it what, what can you do? And it has, like, some self-reflection to tell you what it can do.

It's also connected in this demo app to a bunch of, uh, external tools. So it can, uh, query, uh, like, stock quotes. It can plot some charts, all those, like, external APIs. Uh, it can, uh, also generate, generate images.

But what it really needs to figure out is how to translate user query into do complex reasoning, translate it into function calls. So for example, if we ask it to generate a bar chart with top three, uh, like with stocks of top cloud providers, like the big three, it actually needs to do several steps,right?

It needs to understand that, like, top three cloud providers means, you know, AWS, GCP, and, uh, and Azure,right? And Azure is on Microsoft. It needs to, uh, then go do function calls, querying their stock prices. And finally, it needs to combine those information and send it to chat plotting API, which is what just happened, uh, in the, in the background.

Uh, another important aspect which you have to do for, like, efficient, uh, kind of function calling chat capabilities, you need to have contextual awareness. So if I ask it to add particular, uh, if I ask it to add Oracle to this graph, it needs to understand what I'm referring to and, like, still keep the previous context and regenerate the image.

And finally, you know, if I switch to a particular, to a different topic, it kind of needs to drop the previous context and understand that, like, hey, this is less, uh, this historical context is less important. I'm going to start from scratch.

So there is no, like, Oracle in that cat photo or whatever. Uh, so, you know, this particular demo is, uh, is actually open source. You can, like, go to our GitHub and try it out. It's built with FireFunction and it's built with, like, a few other mod, a few other models, including like SDXL, which are run on our platform.

Uh, the model itself for function calling is actually open source. Uh, it's on Hugging Face. I mean, you can, of course, call it on, uh, at Fireworks for optimal speeds, but you can also, uh, run it locally if you want.

It uses a bunch of, uh, you know, functionality on our platform. Uh, for example, like structure generation, uh, with, like, with JSON model grammar mode, which I think was similar to some of the previous talks from, like, Outline guys, uh, which we were talking here yesterday.

Platform15:25

Dmytro Dzhulgakov15:25

Uh, yeah, so finally, try, try it out. And generally, like, how to get started on Fireworks. So if you head out of Fireworks AI such models, you'll, you'll find a lot of, uh, open source, open base models which I mentioned about.

They're available in the playground. In terms of product offering, uh, we have this kind of range which can take you from early prototyping all the way to enterprise scale. So you can start with serverless inference, which is, you know, not different from, uh, you know, getting to open API, open AI playground or something where you pay per token.

Uh, it's a constant price. You don't need to worry about, like, hardware settings or anything. As I mentioned, you can still do fine-tuning. So you can, uh, you can do constant fine-tuning on our platform. You can bring your own LoRA adapter and still serve it serverless.

As you kind of graduate to, like, maybe like a startup and you graduate to a, uh, more production scale, uh, you might want to go to on-demand where it's more like dedicated hardware with more settings and modifications, uh, for your use case.

Uh, you can bring your own custom model fine-tuned from scratch or do it on our platform. And finally, as you kind of, if you scale up, uh, to a bigger volume and want to go to enterprise level where it's kind of discounted long-term, long-term contracts.

And we also will help you to kind of personalize hardware setup and do some of those tuning, uh, for performance, which I, uh, which I talked about earlier. And in terms of use cases, I mean, we are running production for, uh, for many, many companies ranging from small startups to big enterprises.

We are serving, like, last time I checked, like, more than 150 billion tokens per day. So, you know, companies like Quora built chatbots like Po, uh, Sourcetraft, and Cursor, which I think, I think Cursor had a talk here yesterday.

At Scale16:51

Dmytro Dzhulgakov17:04

They use us for, like, some of the code assistant functionality. And their, like, latency is really important. Uh, as you can imagine, you know, folks like Upstage and Liner are building, like, different assistants and agents, uh, on top of that.

So, uh, we are definitely production ready. Go try, try it out. Uh, finally, we care a lot about developers, you, uh, you guys. Uh, so actually this is external numbers from, uh, like last year BlendChain state of AI stuff where it turns out we are one of the, like, after Hugging Face, the most popular platform for where people pull models, which is great.

It was very nice to, uh, nice to hear. And again, for, like, for getting started, just, uh, you know, head out to, uh, head out to our website. You can go in the, go play in the playground, uh,right away.

So for example, you can run, you know, Llama or Gemma or whatever at the, at the top speeds, uh, and kind of go start building from there. I'm really excited to see what you can build with open models of FireFunction or some stuff which you, uh, which you can fine-tune on, uh, on your own.

And yeah, last point, we are, as I mentioned, OpenAI API compatible. So you can still use your, you know, your favorite tools, uh, the same clients, or you can use frameworks like BlendChain or LlamaIndex or etc. So yeah, really excited to, uh, to kind of, uh, to be here and talk a little bit about open source, uh, open source models and how we at Fireworks, uh, focusing on productionizing that and scaling it up.

Uh, go try it out and you can also find us at the booth, uh, at the expo. Thank you.