Intro0:00
Can you guys hear us?
Pretty good time.
Sound check? Allright. Give it up for Local AI, everyone. Woo!
I hope you guys are excited as we are. This is—woo—this is the Local AI Summit. So we're going to be here all day talking about Local AI. And the reason why is we hit an inflection point this year.
Not only did the models get really good, but the harnesses got really good. And this happened really fast. It's been, I think, a struggle for anyone here to keep up. I felt that most when I saw one of Andrej Karpathy's tweets.
In November he tweeted that you can't really trust these coding agents alone yet. You have to monitor them with an eye like a hawk. Three months later he tweets that he's struggling to keep up with the capabilities of how good this all has gone.
And the thing is, both times he wasright. This space is progressing really quickly, and it's wild how much you can do. And honestly, the only thing to do is just try to use it a little bit more today than you did yesterday.
That's how to keep up with this space. And that's exactly what you guys are doingright here. And so I'm really excited about this panel. The way that we use AI has also changed. I'm not just using chatbots. I'm not just asking simple questions.
When we got reasoning models, the profile of how the AI model responded changed. Not only is it bursty and responding to me, but before the burst it kind of plateaus for a bit. It's reasoning. It's churning on tokens that I'm not consuming.
And then we got agents. And suddenly I don't even want these agents to turn off. I want these always-on agents. They can always be productive if we set them up. And so we have enterprises that want to put a lot of their IP into this because it becomes more useful.
We have consumers with the same thing. I want to give it my health data, my medical records. I want to give it footage from my home camera. And both enterprises and consumers, we don't want that stuff to leak.
And as you also have a profile of tokens continuously generating, suddenly costs matter. And so local is amazing for both of those things. You get to make sure that you are plateaued on the costs for the tokens that you're generating.
Introductions2:20
And also, everything sits in that room. So we have amazing demos here after these talks where everything that's being run stays on those devices. It stays in this room. And so that's a really nice guarantee. So as we turn over to the panels, first, do you guys want to introduce yourselves?
Yeah. Should I start?
Please. You got the mic.
Yeah, your mic's up.
So yeah, I'm Alex. I'm the co-founder and CEO of EXO Labs and also the creator of Local.AI. So we've been working on Local AI for over two years, which feels like a lot longer in this space. And from the days of running Llama 4 FAB on two MacBooks to now where we have Demo Dev, Nimatron Ultra running on four Sparks, our mission is to make AI more accessible.
And it's crazy to see what is possible now.
Cool. Hey everybody, I'm Matt. I make content, videos. We have a newsletter all about artificial intelligence. I'm an AI enthusiast. And yeah, thanks for having me.
Hey everyone. My name is Ahmad Osman. I am the founder and CEO of Osmantic. Somebody jumped into my DMs asked, "Hey, what does OSMAN stand for? Open Source MAN?" So that became a company. And I also moderate LocalLLaMA at the subreddit.
I have been in this Local AI space since 2022. And yeah, let's make open source and Local AI win.
Love it. Woo!
Thank you. Open Source MAN gets the round of applause.
That's appropriate. Today is, I think, a pretty momentous day for the panel. Fable just came back, but it's a good reminder of why we need to have access to frontier intelligence to be able to build anything. I'm Joseph.
I'm co-founder and CEO at Roboflow. We do all things vision. I kind of like to joke that vision is like the original Local AI because everything needs to run probably concurrent with low latency alongside your video where the images and data is being captured.
Looking forward to the discussion today and all things why Local AI is the future of making AI useful for everyone.
Everyone, let's give Joseph and everyone a round of applause. Woo!
So one thing that I'm really excited about for all the panelists is you guys were all early to the space in your own way and very early. And I think as we all feel this inflection point now, when do you guys feel like you felt the inflection point?
Inflection Points4:38
Yeah, I can go first. Sure. I mean, look, when I first saw Llama,right, I got really excited. The idea that I can actually download this intelligence, run it from my local computer, it was very exciting. I'm a tinkerer at heart.
I'm a builder at heart. I've overclocked PCs for years. So this just felt so good. It just felt so interesting to be able to have this crazy alien intelligence running on my device in my office. And that was the first time where I didn't even think it was possible, to be honest.
So I saw Llama and I was like, "Oh, wow. This is incredible." So that was really the point at which I was very turned on to it.
I think it's kind of the same for me. Yeah, it's Llama 2, I think, that finally on my 4090 RTX 4090, I'm like, "Wow, I can actually understand this black box and customize it and play with parameters and sampling parameters and all those configurations and see how inference engines work."
That made me feel like something clicked. It's like, "Oh, magic. I can be a wizard now with this thing. I can control how it plays to my ways of thinking." And from there, the rest is history. Basically, I've been very vocal about why local and open source AI must coexist with cloud and how we are supposed to push for that and try to understand it as much as possible and teach people about it.
Totally. I actually felt this as well. I was on a plane, I think it was in 2023, and I had a model running on my phone and it was awful. It took 20 minutes to complete a sentence. But I didn't have internet.
And so I was able to ask, essentially, intelligence something there in the plane. And that was my first feeling of like, "Hmm, this is going to be something really cool."
Do you know that you can run the equivalent of GPT-4.0 on your iPhone now?
Really?
Yeah. It's Qwen 3.5, the four parameters, four billion parameters. It's basically the same quality as something that used to be served in data centers. And it's on a device in your pocket.
Yeah, that's crazy.
Yeah, I think if you just think about how crazy that GPT-4.0 moment was and now you can run it on a phone, it's like. I think a lot of this stuff was just you had to see the vision of where things were going because it was a toy,right?
It was like the first I think there were a few key points for me. One was running Llama 4 FAB locally,right? That was like the first really big open source model. And I remember the gap. It closed the gap quite significantly with the frontier closed models and open.
But it ran at two tokens per second. So it wasn't useful,right? And then another big moment was DeepSeek, both v3 and R1. DeepSeek v3 was like a massive MoE. So Llama 4 FAB was dense, so it was super slow.
MoE seemed to unlock the performance. So it was like, "Oh, wow. With the devices I already have, like with a Mac, Mac Studio, or Spark, I can run this massive model at actually decent performance, which is comparable with what you can run in the cloud."
And then I think just recently, to me, GLM 5.2 is a big moment as well because, again, it's closing the gap. And it's like Opus level. And we have it running over there on a device that literally can fit on your desk, the DGX station.
So this is, to me, just like a trend. And it's going to keep increasing,right? There's going to be smaller and smaller devices with less memory, better compression. We'll be able to run more capable models locally. And soon it will be the default.
I'll keep with the theme of airplane stories. So I agree that there's been multiple moments over time of what's been going on. But one person.
I have a presentation for that after this. So stay around, please.
Allright. One example moment where this came up for me was I was on a plane. I was sitting next to someone. Man, this is maybe two, three years ago. And they were hard of sight. And they were using if you on Apple, if you take a photo of something, you can use the accessibility settings and it'll describe the photo to you.
And we were seated on the plane together. And they were taking constant photos to understand where we were seated and how to get buckled in and these sorts of things. And I remember they took a photo of the seat back in front of them.
And the Apple accessibility described the photo, described it as like a printer or something. And it was like, "Oh, yeah, you're seated by a printer." And of course, the individual knew from the context that couldn't beright. And I was like, "You know, I wonder."
Llama had just come out, a multimodal model that kind of just described things with the pivot to text base rather than natively multimodal, but still kind of a useful example. And I was curious. I was like, "Well, let's see how Llama would do on the same thing."
So I took a photo of the seat tray in front of me. And it aptly described it. It was like, "Hey, you're on an airplane. That's the seat tray." And I showed the person sitting next to me who had never seen local models and never experienced Local AI.
And what that stood out to me is I was like, "Man, a company that's a trillion-dollar-plus market cap business shipping the latest intelligence on their phones for describing visual settings is inferior to something that's broadly accessible and available to anyone."
And that was such a watershed clear moment that even the largest companies don't have a monopoly on the frontier of intelligence. And so increasingly making that be accessible to others is, I think, going to be critical for understanding all the massive impacts the tech will have.
And that was Llama.
Tools & Vision10:13
Yeah. Yeah, it's interesting to hear.
That's like 2022 or 2023,right?
Around there. Yeah. Yeah, that was a while ago. It's interesting to hear you say about the accessibility of that,right? Taking this frontier intelligence, what made it useful was giving it peripherals, in this case, a camera so that it could access the data in front of you.
And I think that's where the inflection point this year was so much more than just models, but also these harnesses and what you can give it access to. It was CLIs so that you can give business systems and plug thoseright into your agent.
4.0 was really exciting because I was coding with it,right? I was taking code snippets from my code base. And then I'd go to ChatGPT and I'd paste it. I'd give it as much context as I thought I should.
And then I would take the result and I'd go back into my code base and paste it. And what was really great about tools like Cursor was it essentially was this harness that said, "What if I just had the full file system?"
And I can let these agents reason on what files they needed and so on. And so the way that you get more use out of all of these local models and systems is how can it interact with the real world?
I'm curious, Joseph, so you guys got the start in vision. And it seems like vision had to learn the hard way, what a lot of LLMs are discoveringright now. What do you think is a big lesson that the language world is discoveringright now that we could look to vision for?
I think one of the things that language has an advantage of is it's inherently a human construct. So language generally exists where people exist. And that also means usually you can use infinite amounts of computer, what is available in a data center.
Whereas a lot of vision, you're compute constrained. You're running maybe where there's low internet connectivity. You're running on a device. You're running on a robot. And the amount of compute you have available to you is what you get.
And so what did that do? That created, I think, an emphasis on specialized learners faster because it's like, "OK, I need this model to run in a limited context. And I'm not going to prioritize full world-scale generalizability. Instead, I'm going to have specialized in-domain context work."
And what's interesting is I think you're actually seeing language follow a similar thing. You're describing it in the context of harnesses for a given context. But tools like any sort of coding agent where you have a really good harness and is specialized for doing those sorts of tasks.
But you're also seeing that even in tax preparation or legal preparation of last-mile fine-tuning, adaptation, specialized learners. And so in some ways, I feel like the pendulum is swinging back to actually having specialized models, even in language, just as much as vision, I think, is one thing.
And then there's a whole different thing around how do you get the most out of compute optimization for what you're running on a niche device. But I know that probably the EXO folks can speak more about that.
Model Routing12:58
Absolutely.
But yeah, I think specialized models being you remember there was like one model will rule them all world was what everyone kind of thought. And I feel like the pendulum has swung back to people realizing specialized models.
I think we definitely at NVIDIA, we see that it's going to be a multimodal world. We definitely agree with that. There's actually one of the panels that's coming on later, I think, at 3:20 PM, is all about model routing.
And so that's really exciting. I guess as an enthusiast and as you use some of this, is that how your usage pattern looks? Are you using many models? What does that kind of look like?
Yeah, absolutely. Everything from obviously the top models, Fable, to kind of the more workhorse models with Sonnet, local models for things that maybe don't depend as much on low latency. Yeah, the multimodal world seems like a no-brainer at this point, especially as you're seeing all of these enterprise companies come out and look at their budgets and think, "OK, well, I need to continue to increase the total number of tokens that I'm consuming as a company.
But I also don't want to just completely blow out my budget." I know Coinbase and Brian Armstrong just came out with that just great post the other day talking about how their tokens are exploding, yet their costs are staying flat.
And that is because they're using a mixture of different models. You don't need the top model for every single use case. And in fact, most use cases, you don't. I think the most obvious application is let the top model plan, the architecture, whatever the kind of top-level plan is, and then the actual execution of the code can go to a more reasonably priced, smaller model.
Your most intelligent should provide you with the overall plan and then subtasks for your smaller execution, like executioner models. And that's exactly the future.
Yeah, especially we're talking about local. Local models are great at writing code. But maybe we offload the actual top-level planning to one of the frontier models. And you save a bunch. You control more of the workflow. It's a nice pattern.
Yeah, I think the market wants this. So it wasn't clear a few years ago when this stuff was just starting to play out. Would there just be one big model that everyone is using? But clearly, that is not what people want.
That is not what enterprises want. They don't want to be told what they can do by Dario. They don't want to be paying for the same model for all their workloads when some workloads don't actually need a gigantic model that costs $50 per million tokens.
They want control. They want sovereignty. They want the ability to switch out models. They don't want to get rug-pulled from one day to the other because of some safety risk or whatever. And that's really what's driving a lot of this progress.
So I'm more optimistic than ever, I think, about where things are going with local. I think that the market is basically pulling a lot of this stuff out of startups, out of enterprises that are building solutions around this.
And
yeah, it's happening,right?
Totally. I feel like this is a new frontier where as we go multimodal, that becomes really difficult. Of course, you need to route between the models. But then you have to figure out how to provide necessary context to whichever model you're routing to.
If you were making a plan and then breaking it off into sub-agents, what is the framework to do so? How should you do this? These are all open problems. And I think just being playful is kind of the best way to approach this.
But I like the way that you worded that of the market's kind of pulling for these answers. It's reaching in this direction. We're looking for startups and companies to fill this gap. Makes a ton of sense. You know what's funny?
As any enterprise consumes any piece of software, it matters a lot. If you're going to build a foundation, you want to know what the versions are. You want to know what your model is, whether that version changed in order to see changed behavior later.
So having also more control over all of this so that you know exactly what version it is, it can't be changed, not even because of any sort of regulation, but just simply if there are updates, you choose when you're opting into whatever it is on your entire stack.
Sovereignty16:48
Yeah. So when you think about
this wave that we're seeing here of this civilizational infrastructure that is called AI, you have to consider the potential of things being taken away from you and your sovereignty. And how can you be in charge of this thing full stack end to end?
That's hardware, software, and everything in between. That's the model weights. That's the specific version of the model, as you were saying. That's how you fine-tune it. And picking up on Joseph's statement here about small and specialized models, I
have been a proponent of that for quite some time. I have tweets about that in 2024 saying small and specialized models are the future. And it is going to be per user, per use case, per workflow for businesses to decide on their niche domains.
And you need to start onboarding them now because you want to collect the data points that raises. That's what you're asking about, Nader. You were asking about how do we get there? How do we decide which model gets routed to?
How do we decide which use case goes to which model or which endpoint? You need to collect data. And you need to be ready to basically break down
the traces that you've collected from your employees, from everybody in your organization, and decide how can we make the most optimal use case of these data points to which models and collect feedback as well. And you can actually automate that with agents as well.
So this is something that it's the whole thing about RSIright now, recursive self-improvement. It also applies to even agents and harnesses and to workflows and to use cases and to enterprises.
Do you want to describe that for everyone too? Maybe for folks who don't know RSI.
Recursive self-improvement. It means that a model would basically rent its own compute and then start training its own checkpoints and then deploy its next version so that it gets updated in certain ways, changes behavior in certain ways. So basically, a model is training itself.
Totally. For our product, Brev, which just makes it really easy to get a GPU, we've been focusing more on agents as a first-class audience because we're seeing this. So we're seeing growing usage from agents that just want to go grab a GPU directly.
Exactly.
Optimizations19:30
So yeah, you know it's funny. So the second thing you talked about was not just multimodal, but then all these optimizations to actually run it performant when you have less compute than in a, say, data center. I think, Alex, this is some fun work that we did.
So Alex has actually set up a second headquarters inside of NVIDIA. We got a conference room. And we were having dinner. And he mentioned that he really was motivated to get to squeeze out every drop of performance we could on the DGX Spark.
And so we said, "Hey, let's get a conference room at NVIDIA. Come bring your team. And anytime you need an expert across any pillar, let's just go and pull that person into the room." And yeah.
I actually have a follow-up here. So since I have been part of LocalLLaMA for quite some time, when I think about home labs, I think about the individuals that are actually within enterprises on boardrooms making decisions about the models that we're going to host or where we're going to get our AI from.
The things that these home labs focus on are optimizations. It's basically how can I extract the most economical value out of the hardware, the constraints that I have, the software that I have. Quantization came to be because of that.
When we're thinking about enterprises, it's the same thing. It's constraints. It's budget limitations. It's how can I make the most use of the number of tokens that I can get under the hardware that I have, whether that's a DGX station centrally in a small, middle-sized business or a DGX P300 cluster.
It's all about how can I optimize the software for the latest and greatest frontier open-source model out there and get as much value as possible out of it under the economical constraints that we have.
100%.
I think it's crucial to see home labs. And that's why this is local AI. It's local AI. But it could be on premises. It could be colocated hardware. It could be rented clusters. It could be from Brev, basically.
It's just the idea about controlling this thing end to end, from hardware to software to model endpoints.
Totally.
To model weights and collecting all of that data to train your specialized and smaller models that will be more efficient as you go. Yeah, the future is great. And the future is local.
Yeah, absolutely.
Making sure it's your dates, your weights, your compute. I think one thing that was really interesting, we did end up getting 10x performant improvements on the DGX Spark. And we sent an update to Jensen, to the executive staff, to the teams that we're helping.
And one line I really liked in that email was that we didn't solve any new computer science to do this. We actually took things that the experts at NVIDIA had already solved and was out there. And I think what we worked together to do really nicely was assemble it in a bouquet.
And I think that speaks to some of the usability here,right? We're talking about what is capable. What are the capabilities? But do you feel like it's just capabilities that are holding people back from adopting this?
Yeah. So just to tell the story a little bit of what happened here, so I think we had dinner on Thursday of the week. And then an email got sent out on Friday. So this idea came up, like, why don't we do a lab,right?
Why don't we this to me sounded crazy at the time. I didn't think this would be possible. But it was like Nader was like, oh, let's get a bunch of people from NVIDIA to help you guys and work with you guys to improve the performance on the Spark,right?
And I was like, OK, yeah, I mean, if we can, that would be great. Email got sent out, I think, pretty much that night. And then on Friday, Nader told me, be here on Monday at NVIDIA HQ. We're going to have teams of people here that are going to work together with you on this.
And turned up on Monday. And Jensen talks about this concept of swarming. It's basically the idea that the whole company will mobilize around something.
Have you ever seen little kids play soccer? There's a ball. And everyone just attacks it. It's like a tackle. It's just the motion of the soccer ball.
Yeah. And I had heard about this idea. But I'd never experienced it. And I can tell you, we were just back to back that whole day, people coming in and out of the room from all various teams, like the Nimatron team, people working on data center stuff, people working on VLLM.
I realized NVIDIA has a team for everything. So there's like a VLLM for Spark team, which is oddly specific. But they have all these teams. And we basically got to pull those resources in. And in those three weeks or so, as you mentioned, we basically got 10x performance versus what NVIDIA had running on the Spark in their existing playbook, which was using Hermes Agent.
So we did a bunch of optimizations there using VLLM as the sort of inference back end there, doing a lot of work with tuning the models, like quantizing the models to be fit for local. So what you'll find is I think one of the great things NVIDIA has done with the local hardware is it's the same architecture that's running in the data center as running on the Spark.
So it's Grace Blackwell. So the hardware is fundamentally the same, meaning you actually get a lot of things for free. So for example, the kernels are already really good. However, there is a lot of tuning and a lot of configuration that isright now designed specifically for the data center.
So a lot of the work that we did was not inventing anything new. But it was actually just tweaking things to work more performantly on the Spark. And the hardware is extremely capable. It's like it is data center-level hardware that's literally can sit on your desk.
So it's just about how do we activate that? And I think this has been one of my learnings from the last few years is we already have the hardware. The hardware is already really good. And the models are getting better.
They're getting a lot better at compression. So you can fit more on a smaller device. So I think it really is just a matter of more people looking at the space and working in it. And we will be able to do amazing things.
We have Nimatron 3 Ultra. It's a 550 billion parameter model running on four sparks over there, 30 tokens per second. So that's huge.
I want to say something here. It's so great to see the community being proactive. I think of Alex Cheemaright here with EXO as somebody who's a member of the community who's trying to push the frontier open-source intelligence to the next level.
I love how NVIDIA is working with them on that. When I think about this space, I think we're in the '90s of the Linux operating system. And we are just starting. The infrastructure is not there yet. We need so much more.
I was even telling Alex, I want to see how we can collaborate. We have an open-source deployment system called ODS that is a bunch of open-source tools that we deploy in each hardware. It configures stuff, configures agents, and basically gets you set up with the entire infrastructure end to end that you need for your agents locally.
This kind of stuff, we need that optimized for every piece of hardware out there so that we can get more and more people onboarded, have the open-source adoption just go to the moon. That's what I wantright now. This is why we're here.
User Experience26:50
We want everybody to know that local and open-source AI can run on anything, starting from your phone to your DGX Spark, as you were saying, Alex, to DGX stations, to the next level in the data centers. And it will deliver you frontier-level intelligence.
Yeah, absolutely. I think this is a really good question for you, Matt. You test everything for a mainstream audience. You're hearing about all of this technology. You're using this aggressively. Where does it still fall short for the average user for AI?
So I think about two things, the average kind of personal user and then the average enterprise using open-source. I mean, it needs to basically be as simple as opening Cursor. It needs to be maybe slightly more complicated than that or slightly more complicated than just installing Codex and opening it up.
Right now, to be fair, it is quite far from that. The stuff these guys are doing is incredible. But it is more sophisticated than what most people, including myself, are going to be capable of, let alone a business.
It's a full-time job to just do this kind of stuff. That's why we need to automate it.
And it really does need to be point and click. And once it gets there and there's a lot of great open-source projects. There's a lot of great projects in general that are getting there. But we're still not quite there.
The other thing to make it widely adopted is to allow people to better understand what use cases are appropriate for what type of model, for what type of harness, what type of hardware, knowing exactly the use case that I can get out of my home system or something that I'm renting from a service center.
That is incredibly important as well.
And I think it shouldn't be just in documentation. I think that's where open-source becomes so difficult for people. I think it needs to be a point and click. And it figures it out on its own. That's what ODS is about.
That's what I think EXO is about. That's what, as we grow more and more and building this infrastructure, we need to be thinking about the user experience for your average everyday user, not us as technical folks here. Because this is AI engineering summit.
And I'm pretty sure that we all can manage our way around this stuff. But your ChatGPT users, your Claude users, whatever out there, they want this to be an alternative that is just click, play, send a message, use an agent, and done.
Yeah, most people really don't want to know about the details. They just want it to work. And even if it worked seamlessly within Cursor or Codex or any of these other things, and it just worked, and they didn't have to think about it, that is number one for the vast majority of people.
So interesting on that. In a multimodal world where we have to pick a model, do you think that that's something that is a product? Or is that plumbing? Is that something?
Nader, you talked about the harness. It's exactly that. It should be some whatever the kind of the front entry point to somebody's AI experience is, that should be what is choosing which model at theright time for theright use case.
This is a very difficult problem. There are open-source projects, closed-source companies, building routing. That's only one piece of it,right? So go ahead.
No, like Carter from NVIDIA. He's going to be a moderator here. Hey, Carter. Looking good, man. So I shared ODS with him. And he told me he loved how it immediately downloaded that 2 billion parameter models, allowed me to start playing with it.
And then it started downloading the next model that would work perfectly on my device. That's the kind of experience we need to be giving these users when we're onboarding them. Don't make them sit down and have to think about all these quantizations and all these extensions and all these weird things.
That is too much work for your average user. That's how we lose them.
Specialized Models30:35
Yeah, it's interesting,right? When you're not having to deal with multimodal and you have just this generally good, big model, you can kind of ask it anything. You don't have to be as sophisticated or you don't have to understand your use case too well.
But as we talk about specialized models, if you're going to specialize in something, you need to know what that something is. And so the amount that I mean, can you guys speak to maybe what it's like to move into a specialized model or to build one?
I can start with that. On the cloud, you're basically getting the normal distribution. So every time that they are training something, they are taking the average feedback from everybody when they are being happy or sad or they are OK with the answer or, hey, no, this is not what I meant.
Go back or change the model and answer with a different model. That's the kind of feedback that allows cloud providers to fine-tune the next model. But if we're talking about small and specialized models for use cases, that's a lot of compute to train models.
That's a lot of storage for these weights. That does not happen without each use case or each business entity, et cetera, focusing on their own patterns and use cases and how they handle agents and how their employees message these agents and what workflows they are interested in.
So it's really it's per use case. It's per entity. It's not something that we can just generalize. And that's why continual learning is not being talked about by the frontier model as much. It's coming, though. It will happen.
And it needs to be running on local hardware for it to happen. That's how you can have something that is so optimized that is not just plotted markdown files sitting in one repo. And you think that that agent is not going to lose track of which skill to update or what to edit or which memories to update.
Because at some point, context lens becomes inefficient. And that's one issue. The current paradigm of agents is basically just saving to markdowns. The next one will be updating the weights. And that needs to happen locally.
Yeah, it's funny, too, because there was a tweet, I think, by Swix, that was like, why has fine-tuning as a service not taken off? It was like a couple months ago. And I think a big reason for that is that model customization itself is also a very hard problem.
And then what models can you hack on? I mean, that's why NVIDIA releases Nimatron. It's an open-source model. But everything from the dates sorry, the data to the weights to the recipes for how to do so. And then, of course, the final model, everything is open-sourced just so that it is a model that you know you can safely use and customize.
You know, it's interesting. I keep kind of flip-flopping on this point. I think people look at how good the generalized models are. And you give it theright context. Is that going to be better than having a fine-tuned model and all the work that comes with that?
And I keep going back and forth because there is a lot of value in being able to have kind of a smaller, very specialized model and maybe a bunch of them working in unison to accomplish whatever task you have.
But yeah, I'm not sure. What do you guys think about that?
I think if I can chime in on this one, I think the beauty of the open-source ecosystem is that all of these paths kind of get to be explored. And then whatever wins, it's like what happened with speculative decoding.
There's all these different ways of doing speculative decoding, which is this idea that you can use a smaller model to basically approximate a larger model to speed it up. And you've seen recently, literally, I think in the last week, there's been three different quite seem-to-be breakthroughs in speculative decoding that have just come from different places, one from DeepSeek.
There was some work that was done by Modal and the SG Lang team to build DFlash draft models for various Qwen models that are a big improvement on the previous ones. There's all this work being done. And so we kind of get to see all of it.
And basically, whatever ends up being the best thing will just be the thing that kind of wins,right?
That's a great point where maybe it's not users and consumers actually customizing their own models and using the specialized models. But someone who has a need does so and does so in an open-source fashion so that someone else can just adopt it,right?
If I'm going to start to do image gen for a particular use case, I can just go find a model. Actually, we saw this a lot also with early LoRAs on a lot of image gen models. I feel like that was really big in the 4.0 era as well.
I think ultimately, what the end user cares about is, does it solve my use case? And can it do so within my budget? And if there's another service that can manage the entire fine-tuning process and I don't have to think I know I keep saying I don't have to think about it.
Sometimes I do think about things. But in this specific use case.
I was thinking about too many things already to add one more thing to it.
I care about my business. I don't want to think about the I want it to be abstracted away from what I'm worrying about day to day.
Totally.
I could describe maybe a common flow that we see that's like distillation or general model-specific model. So if you think about it, in a lot of cases, it's like if the challenge that I like to pose that's like a thought experiment is like, if you tell someone, think of any object, then it's like, think of an object in your fridge.
The latter of those is actually easier. And the latter of those is actually where a lot of models kind of get deployed into real-world settings. So for example, you know the Monterey Bay and the Monterey Bay Aquarium Research Institute, MBARI, they're called.
They discovered a new fish species recently. And they built a Roboflow to process all of the underwater footage that they capture from their deep-sea submarines. And they use large models like Segment Anything 3 and LLMs as judge to basically say, hey, let's take all this video footage.
And let's build an auto-label pipeline to understand as many things as we can about all the video footage that we've collected. But then ultimately, Sam knows everything about things from fish to maybe architectural diagrams to
items in your fridge and everything in between. But they only care about things that are, in this case, underwater deep-sea exploration. And so a very common flow that we see for the fine-tune, last-mile distillation is, OK, let's take the large context of an array of models, have those maybe all have a pass at understanding something, use LLMs as judge to say, hey, we agree with some amount of consensus.
And now we have a specialized data set. Then we can use that specific model and actually run that on the submarine in real time and also post-processing for faster video. And if you think about it, a lot of problems are that shape, where it's like, yeah, I want general as much intelligence to understand the thing.
And then ultimately, the problem I'm solving is specific enough, even if not fully unbounded. And a mistake I see pretty frequently is people thinking like Sam 3, which is an awesome model and a great family of models. It's like, OK, well, I should take Sam 3 and then maybe just fine-tune Sam 3.
And in some ways, that actually doesn't make as much sense because you lose the thing that makes Sam 3 awesome, which is the open vocabulary capabilities. What might be better is if you know you're distilling down to a specific fixed class list, then you can actually drop the large, expensive autoencoder portion of Sam and use a specific, maybe like debtor or more specialized model depending on the task that you're solving and get all the benefits of speed up and accuracy while still having the general knowledge of preparing and curating your problem.
And I think a lot of problems are of that shape.
Are you managing that pipeline for your customer?
Yeah, the tooling makes it so they can do that.
Do you do it on their behalf? Or do you give them the tooling to do it?
Give them the platform to do that. And then there's recipes where someone can go and do that. And then like any good AI company, there's an FDE that if you want, I can sell you who to get around back here.
And they can do it for you.
That's a great example, basically, of the use cases and workflows that I was talking about when you're trying to
find collect the data as you go for your business, for your enterprise, and then decide how you're going to fine-tune a model, a small specialized model, on those use cases as you grow. That's how you become more token efficient.
Yeah.
Yeah.
I think this year and next year, you're going to see a lot of using these monster frontier models to bootstrap a more efficient setup that runs on open source. And I think that's great. I think this is how a lot of the open source models have been builtright now.
And it's proving to be quite hard to stop that. And I would just encourage people to just move away potentially from these frontier models, but use them. Use them for bootstrapping that,right?
I love that you said that, actually, because that's how the word AI engineer got coined,right? When Swix released that blog post that coined it, the way that we used to build AI products was we would start with the machine learning.
We would start by training a model. Then we would go try to discover a use case. And what these big models allowed us to do is actually flip it. We said we could start with discovering a use case and then get into ML if it makes sense.
And that was actually, honestly, an amazing foresight from Swix because he kind of defined that pattern for us a few years ago and gave this amazing conference as well.
Yeah, it's crazy how far this conference has come in just three years.
Open Problems39:40
Yeah, totally. I have a question. Do you guys have questions in the crowd? We have a few more minutes. And I know we were talking about opening this up if you guys want to ask the panelists directly. Just shout.
Yeah.
Please, go ahead.
What are the big open problems in local AI? I feel like you're the luminaries of the field. You know more. You have the lightest eye on the biggest open problems.
Do you want to repeat the question?
I'll repeat the one.
You just asked, what are the open problems in local AI?
I can repeat it. It's, what are the biggest open problems in local AI?
It continues to be optimizations. And for inference, it continues to be getting things easily kickstarted, which is what ODS says, XOS for specific hardware. It continues to be, how can I make the most out of my budget constraints and hardware constraints?
And how can I squeeze the most performance out of that? The models that I still have the same 3090s that I used to run LLaMA 2 and Dolphin fine-tunes on. And now they are running Qwen 3.5, Qwen 3.6, 27 billion parameters with excellent performance.
More on that in my next presentation. So it comes down to the optimization. And we are still very early that there is space for so many players, for so many contributors. We need all the help we can get to make this thing the success that it needs to be, that we need to make local AI the default.
This has been my stance for years now. I have been saying open source AI must win since forever. And the way we do that is by giving the people, whether that's individuals at home, middle-sized businesses, or enterprises, an easy way to use these models in a very efficient way that doesn't give them headaches more than solve their problems.
Totally. And if you look at the panels that are happening today, those are what we believe to be the biggest open-ended questions, which is why we tried to assemble this panel. And so we have quantization. So all about talking about the footprint so that the models do fit on these smaller hardware footprints.
There is model routing. There is models generally. So those are kind of what we feel like are the big problems now. And as we do future local AI summits, every time, the panels should change to be the topics du jour that are kind of holding the or are going to usher the next chapter in.
Can I add one thing to that? One of the biggest challenges so I think there's two big challenges. One is what we've been talking about of basically the trade-off of simplicity versus customizability. It's local. It's yours. You can do different things, optimization.
If it's hosted, it's built out of the box the way. That trade-off is always difficult. The second, which I actually encourage people in this room to help solve, is the importance of open models is becoming increasingly in question.
And I actually think that if you think local AI is important, then you think open source AI is important. And it's actually really important to be an advocate for being able to use, change, adapt, and toy with models.
And so I think that that's a problem that could increasingly be something that we feel less control over absent advocating.
That's a great point. I think everyone here feels very passionate about open source. That is why we actually have access to the space at all. That's why the space has progressed. It's a necessary competitive environment. It allows for the best ideas to make their way to everybody.
So definitely, when there's talks about that being a threat, we need to invest and advocate for open source. And open source is so much bigger than AI,right? There's a reason why compute was invented on the East Coast, but Silicon Valley happened here.
And it's because hippies realized that they could share ideas for free in software. I think Silicon Valley is this kind of tension between the capitalists that want to make money and hippies that want to give these ideas away for free.
And that tension is what creates such an incredible environment here.
They can coexist, by the way. You can have consumers, and you can sell to businesses.
And that's when the space works best.
Can I just say, if you care about this and you want to be more of an active participant, but maybe you don't want to get involved in the technical side or you're not technical, then there are ways to get involved.
So there's a website that just came out called writetoiintelligence.org. And this is a way for you to get involved and actually advocate for open source and to ensure that we maintain freedom of intelligence.
Closing43:49
Awesome. So we're out of time. I want to thank you guys so much. Local AI Summit's going to be a ton of fun. Thank you guys for this incredible way to kick this off. We're going to have amazing demos.
We're running foundational intelligence models here inside of this room. We're going to have a few more panels. It's going to be an exciting day. And we're all lingering here. So ask questions. And please participate. Thank you.





