Intro0:00
So, uh, why do I— well, first of all, who am I? Uh, do you know this thing called PyTorch? Um, a lot of people in AI used to know it, but now a lot of people just use high-level APIs and don't know what's powering things underneath.
But, like, PyTorch is the software probably powering your AI APIs. So I work on it, I co-founded the project, and it's a— it's a big project that is majority funded by Meta, where I work at. And so I'm talking about— so I'm not talking about Llama at all.
I work on Llama a little bit, but unfortunately, I am not in charge of Llama to try to sneak in some secrets for you guys. I'm not going to tell you when the next Llama is going to come or anything like that.
Agent Spark1:10
So, uh, why am I thinking about personal local agents? Well, as AI kind of started becoming more and more useful, one of the things that saved my time the most every single day was Swix's AI news. And it's basically, like, I have to keep up to date with all of what's going on in AI, that's my job.
And now, instead of basically spending 3, 4, 5 hours a day looking at a bunch of sources, like, the AI was aggregating a bunch of news for you. And I thought that was, like, one of the— like, one of the first applications I thought was, like, mind-blowingly personally, like, effective for my own productivity.
And that's when I started going into, like, hey, I'm going to, like, augment AI within my day-to-day in a deeper way. That's not an agent, though. AI news is more like an aggregator. But that's how it kind of started.
The other thing is I also work on robotics. And robots are essentially agents. They act in the world. So
my goal is to build home robots so that I don't need to do any errands. And so, as part of that journey as well, I've been, like, kind of getting into, like, okay, how do I get into understanding AI agents more deeply?
Context Matters2:33
The key takeaway I'm going to really, like, drill down to you today is agents, like, especially personal agents, have so much possible agency in taking actions on your behalf and stuff. And they have so much of your life context to actually be useful to you that you're better off keeping them local and private.
And I'm going to try to, like, sketch out a plan on how to do it, but I don't think I have a complete solution either. So first, like, agent. What is an agent? Like, and why did I say Swix's AI news is not an agent?
Well, an agent is something that can act in the world. Like, an agent is something that has agency. It can actually, like, take an action in the world. Anything that can only get context and do things, but then eventually can't act in the world is not an agent.
That's how I think about it. And what I think is, like, a highly intelligent agent without theright context is as good as a bag of rocks. It's, like, really useless. I'll give you a couple of examples very quickly.
Let's just say I build a personal agent. It has access to my Gmail, my WhatsApp, my calendar. And it's like, did I get my prescription renewed? And it's like, no, not yet. And, like, it's totally lying, except it didn't know because, like, I got the text from CVS on my iMessage and it didn't have access to that source.
And it was doing the best it can with the information it has, but if it didn't have the context, like, it's not going to, like, know how to do better. Similarly, I mean, you can make up, like, 100 examples like this where, like, you have access to one bank account, but, like, your money came into a Venmo and you're like, the agent lied to you.
What happens is, like, a personal agent that doesn't have theright context, it's largely going to be irritating to use. It's like, you don't know when it is useful and when it is not useful, so it's essentially not useful.
Like, even when it gives you some answer, you're like, hmm, is this actuallyright? I'm going to have to go dig in,right? So unless it hits a certain level of, like, reliability and predictability that you know it isright, it's not going to be actually useful to you.
Getting Context4:47
So now, like, why am I talking about personal agents specifically? And how do you, like, how do you get all this context to the agent? So let's just say you have, like, your OpenAI API or some other API or some local LLM.
What is all the context in the world that is personal to you, and how do you give it to the agent? Well, like, the number one thing that you possibly want to do is, like, just have variables,right? You're just, like, you can— your AI should see everything you see and listen to everything you hear.
And that is, like, obviously the best case of providing context to your AI agent, except, like, there's no battery life for any of these variable things, so that's not really practical. Maybe one day when you have, like, crazy batteries, but that's not really going to work.
The other thing could be, like, okay, like, most of my life is on my phone in the ways that I care about from, like, an agent perspective. What about just, like, running an agent on my phone? It's running the background and it's just, like, always, like, watching my screen or something.
Well, you know, that's where, like, Apple kicks you because, you know, they don't let you run a bunch of stuff, like, on your phone asynchronously. Even if you do, they have a lot of restrictions. So, like, the ecosystems kind of, like, kill you and not allowing you to do that.
And unfortunately, I use, like, Apple. So that's out. So the next one is, like, okay, actually, like, the thing that I found, like, relatively useful is, like, if you use, like, Apple in your daily life, you can actually get a Mac mini and, like, just put it somewhere in your home, connect it to the internet, and you can run your agents asynchronously.
There's no battery life issues. You can just log into all your services on your, like, Mac mini. And it also can access all the Android ecosystems because Android is actually open. So I work at neither of these companies, so I can say whatever I want.
So I think that's, like, what I think is a feasible
device to use to, like, run your AI agentright now. The next thing I want to talk about is, like, okay, why are you talking about local and private? Why can't you just, like, run this in the cloud, like, just subscribe to one of the large tech companies' agent services and run your life out of it?
Why Local7:05
Well, I want to give you, like, a few points here. First is, I want to talk about how this is different from you using other digital services. And I think it is different meaningfully, and I think it's also easy to understand.
So let's just think about, like, a lot of you in this room probably use, like, a cloud email service that is free for all of your life. All your taxes are going in there. Like, you know, everything personal is going in there.
Why do you trust it? The reason I think you trust it, at least the reason I trust it, is because it has a very simple mental model on how it will act on your behalf or how it will act in general.
Email in, reply out. It's not— it's basically not trying to do something sneaky under you that is unpredictable. It's a very simple mental model. Your trust of that service is correlated with whether you understand how it behaves on your behalf.
So imagine tomorrow if some unknown email service that you've been using forever says, "Oh, you know, for some of your emails that I have confidence in, I can auto-reply on your behalf." And you're like, "Okay, well, first of all, that might be true, but what is the worst-case action you can take?"
Maybe you'll, like, reply to my boss, like, something nasty. And, like, I don't want that to happen. And, like, that's— like, once the action space becomes powerful enough and unpredictable enough, you get uncomfortable with using a service that you're not fully in control.
And it can get, like, worse,right? Like, companies have to monetize in a million ways. And so what if, like, you're using, like, an online service and they suddenly are like, "Oh, you know, every time you ask for a shopping query, we're going to, like, start making the agent only buy from, like, stuff that gives us kickbacks or something."
So, like, I think, like, your personal agent is so personal to you and so intimate that I feel like ultimately you want to be in control on many aspects that you might not have control on eventually when you have to, like, trust an online service.
So that's, like, one of the biggest reasons, like, why I feel I want to build a personal agent that's local to myself. The second is decentralization. Like, I mean, you already see all these ecosystems that are walled gardens and, like, fighting with each other and don't allow each other to interoperate in various ways.
And if you build one of, like, your personal life, your personal agent around one ecosystem, like, is that something— like, it works fine for compartmentalized things like maps and email and various things, but, like, is that something that you want to really subscribe into for, like, an agent that can take so many different kinds of actions on your behalf in your day-to-day life?
That's, like, the other reason I feel like you should try to— we as a world should try to get to, like, local personalized agents as the norm.
And the third one is for various reasons. Okay, this is what I called— this is what I call, are you going to be punished for your thought crimes,right? Like, okay, you have a thought and it is not a good thought.
And, like, you know, should you be punished for it? And usually, like, the answer is no. Now, if you have a personal agent that is effectively augmenting you in such an intimate personal way, you might be asking it stuff that you generally wouldn't say out loud ever.
And in those cases, like, do you really want to take the risk of, like, putting it out into, like, some provider? Because, like, you know, you can actually ask Perplexity, like, enterprise-grade cloud API contracts that, like, are, like, enterprise-grade, not consumer-grade where they, like, get sloppy.
Even they have to, like, do a bunch of, like, legally mandated logging and then safety checks and stuff. So there is a possibility that, like, you might or might not want to take a risk on. But for me, I'm like, I don't want to ever get into a scenario where, like, I will be, like, prosecuted or persecuted for my thought crimes.
And, like, that I think is, like, another really powerful argument for myself at least to focus on, like, local agents for my most personal
Technical Challenges12:08
augmentation. So now coming to, well, I hope you're convinced that, yes, we actually, like, if you're going to build a personal AI agent, it has to be local and private. Well, okay, what's the problem? Well, let's go to the technical challenges first.
First, like, okay, you got to run this stuff,right? There are, like, great open-source projects that run a bunch of local models that are one of the key competencies of these agents. VLLM and SG-Lang are pretty great. They're both built on top of PyTorch.
So
effectively, there's this one time, like, we wrote a bug in PyTorch and a bunch of us had a Tesla car and Tesla uses PyTorch. And we were like, man, like, this is so scary because, like, are we writing bugs on ourselves?
That's an aside.
It was totally fine. The bug was not that bad.
So yeah, VLLM, SG-Lang are great. But local model inference is still, as of today, slow and limited. It's not as fast as, like, you know, if you just use, like, a cloud service. Even if you spend, like, enough money on a beefy machine, I think that's also rapidly changing.
Like, for example, locally, if you're using, like, a 20 billion or distilled model of some sort, it actually runs pretty fast. But if you want to use, like, the latest R1, like, full unquantized, then it runs, like, super duper damn slow.
I think this, like, is in a state of, like, it will fix itself. So you probably wouldn't get to run the latest and greatest.
And I think, like, the challenges are not so much the technical and infrastructural challenges. Like, they will kind of get to a place where they're fine. I think there's some challenges around, like, both the research and product that people need to think a bit more about.
I think there's a bit of a gap, and this is just an open challenge for this room for all of you AI engineers.
Model Gaps14:19
One is, like, the open multi-model models are good, but not great. I mean, they're not great in a couple areas. One is, like, just computer use. Even the closed models, like, the latest and greatest APIs that you can just pay money for, they're not that great for computer use.
They break all the time. So that needs to definitely get, like, into a better state. The other thing I can notice is, like, if I ask a model to do shopping for me, from clothes to shoes to furniture to whatever, it'll basically give me the most boring shit,right?
Like, it's like the basic stuff. And if I ask it, if I'm like, look, I'll tell you my tastes. And my tastes can get very, like, specific and fine. Like, the more specific I get, the more, like, bullshit it gives me.
Like, it's like, it's like the same, oh, you asked for, like, a red velvet sofa with oak wooden legs. Here's a green sofa that has velvet and it doesn't have, like, oak wooden legs. You know, like, they're not very good at identifying actually visually what you're asking for.
They mostly rely on, like, a bunch of text matching. The other thing you will notice, and this is a big one, is we don't have good catastrophic action classifiers. What do I mean by catastrophic actions? Is there's many actions an agent can take.
A lot of them are reversible or harmless. Like, even if it takes the action and it's not the action you wanted it to take, it's like, whatever. Oh, it had to go to, like, that particular Wikipedia link, but it went to this other one.
Okay, big deal, whatever. It'll just backtrack and go. But there's some actions that are actually catastrophic. For example, you ask it to go purchase, like, a renewal of your tide pods, and then it goes and, like, purchases a Tesla car.
You know, like, this is not the best thing for you to do. And some of these are called catastrophic actions. And I don't think there's a lot, like, there's some open research around, like, how to really get agents to get good at identifying catastrophic actions before taking them and then maybe, like, notifying the users instead.
But there's not enough. And so if you want to really trust your agents personal or in cloud, I think we got to get a bit better at these things.
So that's, like, a big one. And I think open-source voice mode is barely there. I feel like when I need a personal local agent, I definitely want voice mode because sometimes I want to talk to it and not actually type out everything I want to say.
Open Wins17:05
But still, why am I bullish about this whole thing? I am. Because one, I see open models are actually, like, compounding an intelligence, like, faster than closed models, like, based on how many resources are being put on them.
Like, what do I mean by that? Like, OpenAI is only improving their own model. Anthropic is only improving their own model with all the billions they have or whatever. But open models are improving themselves, like, in coordination across board.
And, you know, people didn't really believe it until Llama came out, and they didn't really believe it until Mistral came out, and then they didn't really believe it until Grok came out. And then they didn't really believe it until DeepSeek came out.
Like, basically, like, people are like, oh, you know, like, open models, you know, will not really win. But I think they are. Like, basically in open source, like, I've worked in open source, like, all my life. There's a starting coordination problem.
Like, initially you don't have enough of a critical mass to coordinate with each other. But once you have a critical coordinated mass, open source kind of starts winning in, like, an unprecedented way. And you see that with Linux.
You see that with a bunch of projects. So I am pretty bullish that open models will actually start getting, like, better than closed models,
like, per dollar of investment into open models.
And
that's what I said. Well, okay, I have some plugs. This is, like, gr.ink from my friend Ross Taylor, who worked on this model called Galactica, which got a lot of, like, criticism when it was released out of Meta.
Plugs18:39
It was this open science model before ChatGPT released. Now, like, doing science with, like, LLMs is pretty common, but, like, they got a lot of shit when they released. And he, like, quit, he, like, unreleased Galactica and he quit doing, like, a bunch of stuff publicly.
But then, like, he's working on, like, plugging the reasoning gap between open models and closed models. And there really is a bunch of open reasoning data that will help. So just a nice quick plug. The other quick plug is I work on PyTorch.
PyTorch is working on enabling local agents, especially the technical challenges that I talked about. And we're hiring. So if you are more than an AI engineer, if you're an AI engineer who's also, like, a systems engineer, then, like, PyTorch is hiring.
Well, that's what we got. The other thing, obviously, is I welcome you all to come to LlamaCon, which is happening on April 29th. And save the date. It's going to be very exciting. Lots of Llama stuff will happen there.
That's it. I think it's in California. I actually didn't look it up.





