Intro0:00
Hi, I'm Andy Treadman, partner at Theory Ventures, and today I'll be sharing our research on AI automation in the workplace. This is how we evaluate the best places for new startups to build, and the best places for executives to invest.
You can ask an LLM to do pretty much any job, and it'll give it a shot. Personally, I use AI assistants as my personal trainer, as my primary care doctor, and my editor, to name a few. And for businesses, we think it'll be the same: AI will support sales teams, customer teams, security, finance, engineering, and more.
But all jobs are not made equal; the nature of different work means that LLMs will be much more disruptive in some areas compared to others. So if you're a founder deciding what area to build in, or an executive deciding which function to invest in, what should you do?
Some quick context on Theory Ventures: we're based in San Francisco, and we invest in early-stage companies building on new innovations in data and machine learning, both in infrastructure and in application layer. We are very thematic and thesis-oriented investors; we spend most of our time doing deep research in areas like workflow automation, and we've talked to hundreds of buyers and builders across different job categories.
LLM Capabilities1:07
So first, let's think about what LLMs are good at. All an LLM is trained to do is to model the distribution of the data that it's trained on. That creates three key emergent properties that really matter for workflow automation.
The first is transformation: this is effectively taking information and converting it from one format to another. It could be from a PDF to a spreadsheet, from a spreadsheet to an email, or from an email to JSON. Integrating and transforming data has historically been the largest blocker for most workflow automation, and LLMs effectively solve this problem, so they're really powerful.
Second is synthesis: this is taking a lot of information and distilling it down, or answering a question. For example, at Theory we frequently use deep research platforms from Gemini and OpenAI, which can summarize hundreds of websites to answer a question we have about some technology or market.
And last but not least is reasoning. LLMs are pretty good at approximating human reasoning and decision-making. It's important to note that when an LLM is reasoning, it's just modeling the distribution of written reasoning data, and one big challenge is that we as humans, we usually don't write down our reasoning or all of the assumptions behind it.
So LLMs are really good at reasoning about what exists in training data. That can include basic common sense, stuff that's on Stack Overflow or other support forums, and increasingly lots of code in math and logic, which can be programmatically generated at large scale.
Domain-specific reasoning, like how a security analyst, or a lawyer, or an accountant might think through a complex case, will take more bespoke data collection, which we think is actually a great differentiating advantage for new companies. And now, like, what does it mean to do a job?
Job Breakdown2:28
Thinking about automation, you need to get really specific, and when we research a new space, we break down the workflows to a task or a subtask level. So here's an example in security operations: an analyst might get an alert, run a query, or do some research to get more information, transform or synthesize that result, analyze it, and then continue to iterate on those steps until they've reached a conclusion.
You can even break this down further, like within the querying step, they might choose a tool, then write a query, then debug errors. A couple notes: one, a job is not necessarily a single person; it could be done by multiple people throughout the team, or it could be done by automated systems set up by people.
We'll talk more about that later in the presentation. And second, a person's job isn't just the core tasks they work on; there are all of the interpersonal stuff, joining meetings, et cetera, which are really important in the context of an organization.
We'll also talk more about that later.
So where do LLMs add the most value? We see jobs generally existing on a spectrum of volume and complexity. There are very complex, low-volume jobs like strategic planning, where people spend most of their time just thinking and coordinating with others.
Volume & Complexity3:34
There are jobs in the middle, like customer support, who handle a number of cases that are generally pretty straightforward but might include reasoning with a customer, querying internal systems, et cetera. And then there are jobs who have massive amounts of relatively simple tasks.
Usually, these are jobs where the primary interface is a task queue, like a list of prospects in a CRM or a list of security alerts in a SIEM. Or there are ones where people are sending hundreds of repetitive emails or messages all the time, like in supply chain operations.
LLMs will impact all of these jobs, but differently. On the more complex side, we expect these systems will be mostly implemented as copilots. Even as the models improve, doing these jobs requires so much context, reasoning, and priorities, some written but most unwritten.
And so, for the foreseeable future, we imagine that these workflows will still be driven mostly by humans, who then delegate or accelerate some tasks with LLMs. In the middle is core workflow automation. This is where we think LLMs can automate substantial or end-to-end workflows, where humans are no longer in the driver's seat.
But they require more complex configuration with expert knowledge, and humans will still be helping out on a task-by-task basis, either as a reviewer, someone to escalate to, et cetera. This is a really great category for AI automation. Across different areas we've researched, we see a lot of jobs where 40% to 70% of day-to-day work can be automated.
Superhuman AI4:46
But today we're going to focus on this category at the top, where LLMs disrupt the job entirely. We're really excited about this area, so let's dig into what it means. High-volume, relatively low-complexity jobs will be the most transformed by LLMs because that's where they're already superhuman.
These are jobs that are hard because of scale. Teams just get overwhelmed by the volume. And so, in many cases, there's just too many things to handle, and so they build rules-based automations and workflows to do it for them.
And that's really awesome for people building with AI, because now your competition is no longer AI versus human; it's AI versus previous generation of rules-based software. And in many jobs where people think of LLM systems as being 80%, 90% as reliable as a human, in these kinds of jobs, an LLM system can be 10 times better than a human ever could be, because as long as they can do the task in the first place, there's no difference for them doing it 10 or 100 or 1,000 times each day.
I'll give two examples now from different companies in the Theory portfolio. The first is Dropzone AI in the security operations space. As companies grow, they buy more security products, each of which generates more alerts. And as the number of alerts grow, the companies then need teams to monitor them and investigate if they're real, to perform a remediation, or if they're a false positive.
Security Ops5:58
The problem is that there are just way too many alerts, and so security analysts might only look at 1% and then build rules-based systems to get rid of the rest. You can see an example on the left-hand side.
It works well for simple stuff, but all of the rules and workflows you need to maintain would really explode as you consider all of the edge cases. And then, from an analyst perspective, this is eye-bleedingly repetitive work. They typically leave the role after 12 to 18 months, and there's a shortage of about 4 million analysts globally.
Dropzone has built agentic systems that perform end-to-end investigations just like a human. In many ways, they're better than a typical human analyst. They're experts in every query language, they don't make typos, they have near-perfect memory. But the key thing is, they don't even need to be, because humans don't even have time to look through all these alerts.
They just need to be better than the rules-based systems, which is pretty easy. They then provide 24/7, 365 coverage, unlike a human analyst. And last but not least, they can share learnings across customers. So if there's a new kind of threat that comes in, all of the customers in the network can immediately be able to identify and block it.
Customer Engagement7:24
Another example is in customer engagement, like our portfolio company Amp. Any app, subscription service, or retail business wants to engage with its customers, whether by text, notifications, or through in-app personalization. And deciding what to send one user isn't that hard, but deciding what to send a million users really is.
So to handle the volume, marketing teams set up and manage different rules-based journeys. Here's a sequence: we'll send in new users, this is what we'll push to a customer who left something in their cart. There's an example of this interface on the left-hand side.
But we know here that everyone has different preferences. They might have different interests, they might respond to different types of messaging, maybe they prefer different channels, different times of the day. And today, marketing teams might be able to manage a few cohorts of customers, but designing the rules-based journeys across all of these variables would create a combinatorial explosion that would just be impossible for anyone to manage.
And they're also forced to evaluate the impact on a single metric, like message clicks, but they have no way to determine how a strategy might drive customer purchases a week later, or subscription retention a month down the line.
Amp, on theright-hand side, has built agentic automation that explores what messaging to send to whom, when, and how. Marketing teams turn more into experimentalists, where they craft hundreds of different variants and then let the AI system figure out how to distribute them and evaluate the impact on all of the users' activities over time.
They show massive uplift in customer satisfaction and in engagement metrics, whether transactional or retention-based, which makes sense because they're one-to-one personalized for each user versus bucketing that user along with 10 or 100,000 others. But even more than that, because the agents are constantly running experiments of what to show to whom, they help companies discover brand-new cohorts and insights on their customers.
One of their customers is a food delivery company, and they found a bunch of users who were only responding to 11:00 p.m. or midnight messages, which is later than they would ever typically message a customer, because they were this new cohort of late-night snackers that the company previously didn't know about.
That's a strategic asset that can be used by data science and product teams, and another example of the unique capabilities of agentic systems at scale.
So what does this mean for organizations? Obviously, people don't do their work on an island; they work with a team. And how will AI change how those teams work together? We think about it at two levels. First, on the role level, work shifts from completing tasks to reviewing the LLM outputs, whether as an approval workflow or escalation ones, to maintaining these LLM systems, maybe updating workflows, model instructions, data systems to improve the overall performance and reliability.
Org Impact9:34
And last, doing higher-order work that previously took a backseat to day-to-day operations, like in strategy. Of course, all of this work is higher level: reviewing, maintaining systems, strategy, all require expertise and experience in the role. But today, most functions doing this high-volume work are pyramid-shaped.
The largest groups of employees are the junior-level ICs who are completing the day-to-day tasks. So when an LLM system automates a substantial portion of this work, we expect these organizations will need to, one, shrink because fewer employees are needed overall; most of the work is automated.
And second, to invert, that the positions that do remain will be more managerial or advanced. So organizations will instead look like inverted pyramids or diamonds. This causes a lot of questions and challenges for businesses, which are today oriented around this pyramid structure.
For example, how do you hire and train new employees if there aren't a lot of entry-level roles? So it'll be an interesting type of challenge where we're looking to see how organizations respond over time.
Recap11:00
As a recap, when you're thinking about AI automating workflows, you need to break jobs down to their fundamental tasks, understand that jobs exist on a spectrum of complexity and volume, and that AI will be most disruptive in high-volume, low-complexity tasks.
And last but not least, surrounding context will determine how AI impacts teams and organizations, which is important both for founders looking to sell to these organizations and executives looking to think about how these will be transformed by AI.
Founder Note11:27
One last note for founders: here, we're exploring sort of technology problem fit, how well LLM systems can do different jobs. There are a whole variety of other factors that determine if an idea is good to build a business around, like the severity of the pain point, the incentive alignment, the market size and structure, et cetera.
Happy to chat more about that if anyone's thinking about building an AI workflow automation.
Thanks so much for taking the time. You can reach me at this email here. Have a nice day.





