AIAI EngineerFeb 6, 2025· 19:59

Which Jobs Can Be Replaced Today: Fryderyk Wiatrowski and Peter Albert

Fryderyk Wiatrowski and Peter Albert, co-founders of Zeta Labs, argue that autonomous browser agents will first replace reactive jobs—like customer support and scheduling—by automating low-leverage tasks while preserving human focus on high-leverage activities. They propose a trigger-pool system where agents react to emails, Slack, or events, requiring only approval for actions. Peter details building reliable agents: start with prompting optimized for model distribution, then add cognitive architectures (e.g., planning, scratchpads) to split tasks, and finally fine-tune with synthetic data or reinforcement learning. He advises minimizing noise in prompts, preferring text-based reasoning over images, and using language model judges to filter training data. The founders see continuous model improvement enabling agents to handle increasingly complex, proactive roles, moving toward full job replacement.

  1. 0:00Vision
  2. 1:11Autonomous Agents
  3. 2:04High Leverage
  4. 4:48Reactive Work
  5. 6:36Trigger Pool
  6. 9:20Browser Agents
  7. 10:33Prompting
  8. 13:59Task Splitting
  9. 17:09Fine-Tuning
  10. 19:17Reinforcement

Powered by PodHood

Transcript

Vision0:00

Fryderyk Wiatrowski0:13

My name is Fryderyk, co-founder of Zeta Labs, and me and Peter, my co-founder, today we'll speak about the job replacement, the future of job replacement, and how we see agents in there. I will start with a vision, and then Peter will tell you a bit more how you can contribute to that by building agents on your own.

Um, cool. So let me start with asking you a question. If you could hire a reliable autonomous browser agent today—let's assume that it's fully reliable and very fast—what tasks would you want to have replaced?

I need 3 tasks.

Guest0:53

The boring ones.

Guest 20:55

Travel organization.

Fryderyk Wiatrowski0:56

Yeah? Travel organization, yeah.

Guest0:58

Expenses.

Fryderyk Wiatrowski0:59

Expenses, that's cool. Yeah.

Guest 21:01

Calendar.

Fryderyk Wiatrowski1:03

Calendar. I love the calendar one. Um, cool. Um,

I think if you have a reliable agent and you want it to behave as your employee, um, our claim is that sometimes you don't want to prompt the agent. Just like when you hire a good employee to your company, you spend some time on the onboarding, initial training, um, but after that, if it's a good employee, it's very independent, um, you don't need to ask them to do things,right?

Autonomous Agents1:11

Fryderyk Wiatrowski1:41

So you can expect them to come up with the tasks on their own, and you probably shouldn't expect them to ask you every day, uh, what should be done. You don't want to overlook everything that's being done, and you don't you want the employee to come up with the tasks on their own.

Ideally, you just give them the vision and they will do the rest.

The question is how we can implement this into agents.

High Leverage2:04

Fryderyk Wiatrowski2:10

Probably this is not theright UI for this. Um, in this UI you just prompt the agent to do things for you, one-off. I think it's great for demos and for showing the capabilities of the agents, like demo drives, um, but I don't think this is what the future looks like.

So the question we can ask ourselves is how to embed agents into our workflows. Um, I think a good place to start is asking ourselves what are the high-leverage activities that we want to preserve in our daily life, and what's noise.

I will give you an example. In a founder's role, a very high-leverage activity, um, is hiring. Of course, if I hire a team of 10 people, they can do do the job for me 10 times better than me because I hired smarter people, and, uh, it's a huge leverage.

However, this huge leverage is surrounded by a lot of noise. For example, setting up meetings, or searching for them on LinkedIn. So I think the vision that we can, um, keep in mind when building agents and thinking about how we can embed them into our workflows is how to distill the high-leverage activities and preserve them for ourselves while outsourcing the low-leverage things for agents.

Um, one solutionright now is just to, like, if you're a, I don't know, a big company and a big founder, you can hire an executive assistant that will do the meetings for you and set up those things. Uh, you probably need to give them a big salary, like a 50, 70K, or something.

Um, but initially you always start with a problem. It's not like you look for an executive assistant. You start with a problem, "Hey, I need to set up meetings, and I don't have a very specific need for all the surrounding skills that they have."

So instead of hiring, how about we just hire agents to do just this one thing? Um, I think if we find a way to implement agents in this way, um, that would be truly revolutionary. Um, so for example, here, you know, if we build a simple agent today, uh, if you use our agent, Jace, go to Jace.ai and use our agent, you can just CC Jace into your emails and it will do the meeting setup for you.

Super simple, doesn't need to interact with browsers, tools, anything. Needs to know your calendar and, uh, and needs to have access to your availability. That's it. So as you can see, I CC'd Jace. Jace is replying to an investor interested in Zeta Labs and setting up a meeting for me, knowing my availability.

Um, that's sweet, um, but, you know, most of our tasks require, uh, integrations and tool access,right? Um, I think

Reactive Work4:48

Fryderyk Wiatrowski4:48

in order to think about how we can enable those integrations, we can distinguish two modes of human work.

One is reactive, and the other one is proactive.

A great example of a reactive job is a customer support. What happens in customer support is that when you get an email, for example, asking for a refund, you do a very well-defined task. So what you do is you go to a, uh, go to your logging system or whatever, you see whether someone used the app or not, and based on that, you can give them a refund or not.

Um, and that's pretty simple. Uh, you need to attach Stripe, the agent can do the refund and everything. And then there are those proactive jobs, like for example, a founder's job, where it's super difficult to describe what needs to be done by a simple set of rules,right?

Um, however, we still think that in every job, even in a founder's job, um, there is the reactive layer that still creates the noise and ideally would be handled for us. Um, so let's focus on the reactive part.

We think that the reactive jobs will go first. Um, the beauty of the reactive thing is that once you set the rules for the agent, once you go about the, uh, do the initial training, you don't need to prompt the agents anymore.

Uh, they will, um, they will just follow the rules and do everything for you. Um, but once the complexity comes in, it can be more difficult. Of course, agents as of today can perform reliably, I think, hundreds, uh, up to a thousand steps, um, but, uh, it's more we can't really trust them yet.

So we have to be careful.

Um, okay. So we talked about the reactive bit and the proactive bit. Uh, I think the first step to implementing the reactive, um, job replacement by agents is to create a pool of triggers, the triggers that cause the reactions of agents.

Trigger Pool6:36

Fryderyk Wiatrowski6:53

Once you have the pool and the agents know the rules by which they should pick up tasks and perform the actions,

the agents then can, like, the pool being, for example, you know, Slack messages, emails, or phone calls. The agents can pick up the tasks, suggest solutions for you, uh, do some upfront work in the browser, and then just show you, "Hey, do you do you want me to perform this action?"

And then you can just approve. So we went from prompting agents to do things to manage our calendars, um, to agents doing this on their own.

I think I think we all know that everything is a reaction. In particular, like, even founders' job is about reacting to things like macro market movements, things that are rather, um, unusual, and there are not many templates to being a founder as opposed to, for example, in customer support.

Um, however, you know, the more descriptive our rule set is, the more proactive jobs we'll be able to replace. Um, and as we extend the rule set for reacting to the triggers, I think we will go closer and closer to replacing the proactive jobs.

So assume that we have built the meta-aggregator of all the triggers that is being, you know, uh, macro movements of the market, as well as, you know, emails, phone calls, Slack messages, linear issues, whatever, whatever triggers our actions.

Um, I think the continuous improvement of the foundation models and their cognitive cognitive abilities going up will allow us to have a very complex rule set because those models will be able to reason just like humans in terms of performing their jobs and deciding on what's next.

Um, but the question is how to build this aggregator. And I think a very simple way to start would be, for example, in linear or in Slack, and then agents just picking up tasks and performing them. Um, but then once we build the aggregator, the next step is to make the agents act.

And, you know, browsers, as opposed to APIs, uh, allow us for very generic actions. Um, APIs are very often limited and not really well implemented. So if we can make the browser agents work, can we fully replace humans at their jobs?

Browser Agents9:20

Fryderyk Wiatrowski9:20

Um, I will pass now to Peter, and Peter will tell you how you can implement browser agents today to reliably perform your daily jobs and what are the challenges on the way.

Peter9:34

So, hi, I'm Peter. I previously worked on the Llama 2 models at Meta, and I'll, yeah, like Fryderyk said, I'll give you a bit more actionable insights on how to actually build agents and kind of what the steps are, uh, you need to go through there.

Um, so I think, like, if you want to build any kind of LLM part in your in your system, and especially for agents, you go through a few different steps of, uh, based on the complexity of your task and how much performance you want to have.

Um, and basically, this is kind of separate for every single system you have. Um, you usually can start up with prompting, and once things get more complex, you add cognitive architectures, you can add fine-tuning and, and reinforcement learning.

But I would only I would kind of go through each of the steps separately because each of these kind of reduces your iteration speed a lot. So with prompting cognitive architectures, you can do things within hours. Fine-tuning reinforcement learning is like more like a week to month, monthly projects, and any kind of changes you want to make basically slow everything down a lot.

Prompting10:33

Peter10:33

Um, but if you kind of need to go on a certain task to really high performance, you kind of have to go also through steps. Uh, so first steps, I'll, um, some general ideas, like, about how to improve your prompting.

So usually it's a good idea to kind of rewrite your prompts with language models itself because you kind of use a low perplexity text that you give to a model. So also you can have your own misunderstandings that you had initially will kind of be taken out.

We'll also try to use, like, XML syntax, like Anthropic kind of started with this for their prompts, uh, but I think it's useful for any model just to separate instructions from content. Um, also some useful, uh, mindset is to always try to match the fine-tuning and pre-trained distribution of your language model.

So if you use GPT-4, you kind of have to think about what did OpenAI probably fine-tune their models with, and, and also what is kind of in the wrap in there. So for example, you should probably prefer JSON or XML or Markdown, unified format, um, when trying to output text just because it's more, more, more frequent in pre-trained distribution.

So you shouldn't probably introduce your own format or your kind of own syntax to do things. Uh, you want to do things just like the most likely, uh, to appear in the wrap. Um, yeah, it's also similar for what you put in the system prompt versus user prompt, how you if you if you decide to split up things into multiple messages or just one, you just have to think about, like, what probably OpenAI got in their fine-tuning data set, what most people do.

And this is usually a good way to do, uh, to stay within the distribution of the fine-tuning and will give you better performance. Um, also you should kind of try to think of how to minimize computation at each token.

So if you have, like, a complex task, um, one thing is to keep in mind that you, if you want to output, uh, for example, mappings, like, like, uh, element ID five or four, you usually want to prefer text text values.

So if your model can output text, it's better than specific numbers, um, because basically you're skipping one mapping step. So you should try to let the model only reason about one thing at a time. Um, it's also relevant, like, if you do classification, for example.

Like, it's better to first classify things into broader categories and smaller ones instead of directly going to the smallest one. And if you do this, it will just improve performance in general. I would also, like, if you have some idea about how to think about a problem, I would try to give your if you do chain of thought, I wouldn't just say, say, uh, let's think step by step, but instead you should kind of think about how you would approach a problem and specifically add these things in your chain of thought.

So for example, first thinking about summarizing the problem you have and, and going for all the steps. Um, a few more ideas is that you should think about everything that you put in a prompt is basically noise, and, and some value will contain in there.

So if you can reduce the noise as much as possible and only have the most relevant context in there, the model doesn't need to, uh, find out what is noise and what is real context. Um, for future examples, some really useful tips is that you shouldn't probably, like, like, if you have really long inputs and really long, uh, future examples, you don't have to show the full ones.

You just can only focus on the few most important parts of them, and the model will still usually, um, capture kind of how, how things work. Um, also, like, even once you start prompting, you want to already set up some, maybe, like, 10, 20 examples to kind of always run each prompt through.

Additionally, to have like some better end-to-end evals as well. Um, yeah, so once you kind of, uh, exhausted all the gains you kind of get with prompting, you can try to split up tasks into multiple pieces. So basically, you can think about how much cognitive load you put on a on on a model with one prompt, and if you can split things up further, and also if you have some preconceived notion about how the problem should work, you can kind of improve this.

Task Splitting13:59

Peter14:22

So a few ideas about this is kind of where you add explicit state tracking of, like, if you have a longer task, you basically keep track of the state that the task currently is in. This basically increases the quality of context you're kind of giving your model.

Um, here's some ideas. For example, using kind of planning, you kind of generate plans, afterwards modify them, replan. You can also have the verification steps afterwards. Um, or you kind of take notes or have a scratch pad, uh, for intermediate work.

Um, but once you kind of split up things in multiple pieces, when latency becomes more of an issue. So they're often good ways to parallelize the work while still, um, kind of getting good performance. Also, for all of these kind of agents, it's really important to think about really natural interfaces for them to use tools.

Um, so for example, we found that if you want to update state, like, for example, uh, some notes you have, it's often good to address them, like, with key-value updates because basically you're kind of mimicking Python syntax. You kind of update some dictionary, and this kind of allows you to target elements we develop because you first the model only needs to think about the key, and then afterwards can think about what to update and not do both things at the same time.

Um, you can get even better performance usually with most models if you do full rewrites of what you kind of want to update because this gives the model even more time to think about things, but when you add more latency.

So I think that's, for example, why you're seeing cursor, like, all the text rewritten because currently even both best models are not that good at generating, like, small diffs. Um, also you should try to avoid any recursive nested structures if you can.

Uh, this will just, like, the more nested it is, will just be more out of distribution and make it more complex for a model. Um, yeah, also you can kind of use some reasoning templates. Like, if you know how the problems are structured, you should also kind of reflect this in a prompt.

Um, if you deal with images, then one good idea is that you try to move the reasoning into your text. So you first describe the key points in your image, like, let the model output the text form, what the image contains, and then afterwards reason about this.

Just because the models have been trained on trillions of tokens and text, and most of these, uh, like, image language pairs are usually much lesser data. So if you do reasoning about images, it usually performs worse than if you first convert it into text and then afterwards into, uh, more text.

Um, you should also think about how to design your components to correct for one part if you have one part in the system that creates an error that never handles it. And the more cognitive components you kind of add, the more brittleness you also add to the system.

So usually it'sright it's a good idea to kind of keep it to a minimum. Um, but on the other hand, like, the more components you add, you kind of more cognitive load you split into smaller pieces, so the model can, uh, can has less load on this.

Um, the next stage, kind of if you still don't get enough performance out of your model or if you want to get, uh, lower cost, um, you can kind of start for fine-tuning. And a simple way to kind of collect data, no matter what your use case is, is by simulating basically real interactions in your app.

Fine-Tuning17:09

Peter17:26

And you do this by creating templates about how a human, like, basically different roles of a human. And then you instruct one language model to act as a human, and basically the rest of your system acts as normal.

And it's basically synthetic data that you can kind of use to, to fine-tune on. And for this, like, prompt diversity and difficulty is really key. So I think one of the core issues of, like, Alpaca models, like, in the beginning, the first, uh, kind of fine-tuned models that went out there was that the prompt difficulty was way too low.

Basically, models learn much better the more difficult your prompt is. Even if your task is not that difficult, I would try to add more and more conditions and more and more complexity because then the model has more things to learn than just for a single question.

Um, yeah, also another thing you can kind of, uh, try is, like, if you have multiple steps in your pipeline, you can in the end, if you do fine-tune, you can distill to skip some of them. You basically have, like, some initial inputs and some final output, and you can directly distill a model to get these first inputs and out-generate your final outputs.

And this can increase a lot of, uh, decrease your latency, but will kind of decrease performance a bit. Um, then the next step that's really easy is that you, um, kind of filter your data. So this kind of gives you, like, usually with just fine-tuning, you can get to, like, GPT-4 performance on your specific task, but you can get much further if you simply do some rejection sampling or filtering of your data.

And for this, an easy way is if you don't have execution feedback in some way, is that you, um, basically use language model judges. You'll just judge the output of your model or, like, of a larger system or even the final output of your system.

And then basically you'll filter out whole swaths of your data that you know probably didn't work that well. Even if your judge is not perfect, this will kind of increase your performance a lot. Um, yeah, and finally, so last step that you can kind of approach is, like, reinforcement learning.

Reinforcement19:17

Peter19:17

So this even allows you to kind of optimize multiple steps in your system. So it's especially important for agents. And you good ways to get a signal for this for this is execution feedback or, like, we said, these language model judges of different parts of your system.

But I would usually consider reinforcement learning kind of a final step when other methods don't work, um, because it makes it kind of difficult to move to one, uh, to different models and also there's a lot of setup costs you have to do.

Um, yeah, so I think that's that. Awesome. Thanks.