AIAI EngineerMar 26, 2025· 15:15

How Deep Research Works - Mukund Sridhar & Aarush Selvan, Google DeepMind

Aarush Selvan and Mukund Sridhar from Google DeepMind explain how Gemini Deep Research works, a personal research agent that trades latency for comprehensiveness by taking up to five minutes to browse the web and synthesize reports. They discuss product challenges like building an asynchronous experience in a synchronous chatbot, setting user expectations, and presenting thousand-word outputs. Technical challenges covered include iterative planning with partial information, handling the fragmented web with entity resolution, managing growing context via recency-biased retrieval, and ensuring robustness to intermediate failures. The episode also explores future directions: going from aggregating information to providing expert-level insights, personalizing research to user roles, and combining web research with coding, data science, or video generation.

  1. 0:00Intro
  2. 2:11UX Design
  3. 5:29Under the Hood
  4. 10:44Context Crunch
  5. 12:00What's Next

Powered by PodHood

Transcript

Intro0:00

Aarush Selvan0:17

Hey everyone, I'm Aarush, I'm a Product Manager here at Google.

Mukund Sridhar0:20

Hey, I'm Mukund, I'm a Software Engineer at Google working on Deep Research.

Aarush Selvan0:23

Um, so, I don't know if people have had a chance to try Deep Research on Gemini, um, or are familiar with the product, but you can try it if you go to Gemini Advanced, and if you scroll past 2.0 Flash, 2.0 Flash Thinking Experimental, 2.0 Flash Thinking Experimental with Apps, 2.0 Pro Experimental, you will find, uh, 1.5 Pro with Deep Research, which is what we built.

Um, and if you have the chance to use it and you pay the 20 bucks, uh, you will see that it's a personal research agent that can browse the web for you to, to build the reports on your behalf.

And so our motivation and what we want to talk about today is kind of why we built it, some of the product challenges we overcame, and some of the technical challenges you'll face if building a web research agent.

Um, so our motivation was really we wanted to help people get smart fast. Um, we saw that research and learning queries are some of the top use cases in Gemini, but when you bring, like, really hard questions, uh, to chatbots in general, what we were finding is that it would often give you a blueprint for an answer rather than actually give you the answer itself,right?

So we had this query that we used to throw around of, like, tell me what does it take to get an athletic scholarship for shot put, and, like, how do I go get one? And often the answers would be things like, you should talk to coaches, you should find out how far you should be able to throw, and, you know, uh, you should make sure you have good grades.

But really what I want to know is, like, okay, what are the grade boundaries? Like, how far do I need to actually be able to throw? I want something super comprehensive, and, and that's where we saw a big opportunity.

Mukund Sridhar1:57

Yeah, so we said, what if we remove the constraints of compute and latency at inference time, let Gemini take as long as it wants, browse the web as much as it needs, and see if we can trade that off for a much comprehensive answer for the user.

Aarush Selvan2:11

But you've got to do it in five minutes, because beyond that, uh, we don't have the chips. Um, so, uh, this brought a bunch of product challenges for us. Um, Gemini, up to this point, is an inherently synchronous feature, it's a chatbot, um, and so you wanted to, we needed to figure out how do you sort of build asynchronous experiences in, in an inherently synchronous product.

UX Design2:11

Aarush Selvan2:33

Um, you also wanted to set expectations with users,right? Deep Research is good for, like, one very specific thing, but a lot of user queries to Gemini are things like, what's the weather, write me a joke, things like that, where waiting five minutes is not going to get you a good answer, and we wanted to set expectations.

Uh, and the last thing is, our answers can be thousands of words long, and we needed to figure out how do you make it easy for users to engage with really long outputs, and, um, uh, in, in a chat experience.

Um, so let's walk through kind of the UX and kind of think about how, how we solve some of these,right? So imagine you're a VC, uh, and everybody's talking about, you know, investing in nuclear in America, and so you come with this query like, hey, help me learn the latest technology breakthroughs in small nuclear reactors, and tell me interesting companies in the supply chain.

So the first step, um, when you bring this query to Deep Research is that Gemini will actually put together a research plan for you and present it in a card. And so this is the first way in which we're able to communicate with users, like, this is different.

This isn't your standard chatbot experience. Something's going to happen, you're going to hit start. But it's also an opportunity for us to actually show the user a research plan that they can edit and engage with, kind of like a good analyst,right?

They, they wouldn't just get to work, they'd actually show you, okay, here's how I'm going to approach this. And it's a way for users to, if they want, kind of engage and steer the direction of the research further.

Now,

uh, once you hit start, we actually try and show you, um, what Gemini is doing under the, under the hood in real time, uh, by showing you the, the websites it's browsing. And this is a feature that was built before thinking models, and thoughts are also a really great way of kind of showing transparency of what the model is thinking.

Um, but what's really nice here is while you wait, you can sort of click through the websites, dive into any of the content. Um, but what we also inadvertently saw is people trying to game that number to see how high it could go.

So we definitely saw people push that number into the, into the thousands, uh, to try and, um, you know, see how many websites Deep Research could read. Um, finally, we kind of get this report that's, you know, thousands of words long, and, um, we're really inspired by what kind of what Anthropic does with, um, artifacts, and so we thought that was a really great way of sort of being able to pin an artifact so that users can actually ask questions about the research while reading the material.

They don't have to scroll back and forth. And what's really neat about this is it means it's easy for you to engage in sort of changing the style of the report, adding sections, removing sections, asking follow-up questions, and, uh, and it sort of makes that really easy.

And the last part that's super important is kind of user trust and also doingright by the publishers. So we, we try and always show is all the sources we read as well as all the sources we used in the report, because not everything that we read is used, but it stays in context for follow-up questions.

And, and also sort of these are all things that, um, carry over to Google Docs as citations and things like that if you choose to export.

Under the Hood5:29

Mukund Sridhar5:29

Uh, so I thought today we can pick some of the challenges, uh, that one has to encounter while building a research agent and talk through some of them. So, uh, I picked four for today. So one is this, this long-running nature of tasks introduces a couple of things that we need to look into.

Second is the model has to plan iteratively and spend, uh, its time and compute during this time effectively. So what are those challenges there? And it has to do this, uh, while interacting with a very noisy environment that is the web.

And as you do this and, uh, read through information, very quickly you can start seeing your context grow and how, how do you effectively manage context.

So if, if you think about a job that runs for multiple minutes and something that can make many, many, uh, different LLM calls and calls to different services, they're bound to be failures,right? And today we're talking about, oh, of minutes, but you can very easily think in the future of, uh, these kind of research agents taking, like, multiple hours.

So it's important to be robust to intermediate failures of these various services of various reliabilities. And so being able to build a good state management solution, being able to recover from errors effectively so that you just don't drop the whole research, uh, task due to one failure.

That's one. The second aspect of doing this, what it enables us is to enable this feature, uh, cross-platform. So we believe more and more, uh, users will start kind of registering your asks, uh, or your research tasks and just, like, walk away, do their thing, and then you need to get notified.

And this can happen now across, uh, devices, and you can pick off, uh, uh, reading it, uh, uh, uh, once it's done.

So now what is the model doing at, uh, like, through these, you know, uh, few minutes? Uh, so let's take a, an example,right? So here, uh, we're looking for, uh, athletic scholarships, uh, for shot put. Uh, there are many facets to this query, and we kind of show this in a research plan like Aarush showed.

The first thing the model has to do is try to figure out which of these sub-problems it can start tackling in parallel versus things that are inherently sequential,right? So the model has to be able to reason to do that.

And, uh, the other challenge is, here you see, you're always going to land in this state where there's partial information, so it's important to look at all the information found so far before you decide what to do next.

So in this instance, the model found, hey, it's, it knows the qualifying standards, uh, for the D1 division, but in order to provide a complete report and answer the user's question, it has to go figure out what the equivalent for the D2 and D3 divisions are.

So this notion of being able to ground on information you find and then plan your next step is key.

Another example of partial information could be when you make searches. Uh, so in this case, we're trying to find the best roller coaster, uh, for kids. Uh, you might find results, uh, that provide partial information again. So here, uh, you end up at a link, uh, which talks about the top 10 roller coasters, but does not mention anything about them being suitable to kids.

Uh, so the planner has to recognize this fact and then go ahead and in the next steps of planning try to resolve this, uh, this ambiguity.

Um, another example of, uh, challenges in planning is information is often not found in one place. You find facets of information spread across different sources. So here, uh, we're trying to find, uh, what would, uh, what would it take to get a certification for a scuba dive, uh, in, in, in some dive centers nearby.

So you see, uh, one part or one source has, uh, the kind of the structure of, uh, what, what you have to go through to get a certification, but in a completely different source you have this notion of the pricing for this diving center.

So the model has to weave this together to figure out, uh, you know, what the cost structure for such a certification would look like.

Then there's the classic, uh, entity resolution problem. So you might find mentions of the same entity across different sources, so you need to be able to reason about some information indicators to kind of figure out if they're talking about the same entity or you need to explore more to verify such, uh, disambiguities.

Um, yeah, I think most people here have worked on some notion of a web problem, and we know, like, it's super fragmented. So, uh, here you see two different websites, uh, talking about the same thing, uh, about music festivals in Portugal this year.

Uh, on the left, uh, if you end up at such a website, it's easier and you get most of your information in one go. Uh, on theright, uh, the layout is different. So having a robust, uh, browsing mechanism if you want to navigate, uh, the web for your research tasks is another, uh, important challenge.

So like we saw, there is a lot of these intermediate outputs, and as you do this and you start getting streams of information during your planning, you can imagine your context size growing very quickly. Um, the other challenge that, uh, about context size is your research task doesn't typically end with your first query.

Context Crunch10:44

Mukund Sridhar11:05

People have follow-ups. People can say, hey, can you also do the same for this other topic? So there is, like, this kind of a follow-up, uh, deep research, and, uh, that also adds pressure on the context. Uh, we at Gemini have, uh, the liberty of really long context models, uh, but, uh, even then you have to design, uh, some way to make sure you, you effectively manage your context.

And there are multiple choices here. Each come with various different trade-offs. Uh, we're showing one here, uh, where we kind of have, like, this recency bias. So you have a lot more information about your current and your previous tasks, but as you get to older tasks, we kind of selectively pick out, uh, you know, things what we call as research notes and put it in a rag.

That way the model can still access it, but it's being selective. Uh, I'll hand it back to Aarush about, uh, to talk about what's next.

What's Next12:00

Aarush Selvan12:00

Yeah, so we were super excited to put this feature out in December. We weren't actually sure if anyone was going to use it, if anyone was going to care, um, to wait five minutes, uh, for something. And, uh, we were really positively surprised by the reception.

Um, and, and really what we, what we saw, um, was, hey, we've built something that's maybe as good as, like, a McKinsey analyst,right? And we give it away for 20 bucks. But, um, you know, that's, that's really great, and, um, but what it does is it just retrieves from the open web, and it's a text-in, text-out only system,right?

And so where we sort of, we sort of see a few different directions of where research agents are going to go next. And the first one is around expertise,right? So how do you go from McKinsey analyst to a McKinsey partner or Goldman Sachs partner or, like, a partner to law firm,right?

So that's really around not just being able to aggregate information and synthesize it, but also think through the so what of how do, like, what are the implications for what we're going to do and, and what are the most interesting insights and patterns that come out of it.

The, the other thing is, you know, there are plenty of domains beyond professional services like the sciences where you, you know, want to get really good. You know, you want something that can read many papers, form hypotheses, find really interesting patterns in, you know, what methods we used, uh, and, and come up with novel hypotheses to explore.

However, um, just because you build something that can be really smart doesn't mean that it's useful to someone,right? So, um, if we were thinking about a use case of running a due diligence on a company, the way you'd present that information to me would be very different to the way you'd present that information to, say, a Goldman Sachs banker,right?

Um, for me, you really want to talk through, like, what, like, what is this company and how's its position strategically, but a banker would want to know all the financial information, actually have a DCF that they could look at,right?

Actually, uh, have a, have a much more, like, fine-grained, uh, sort of, uh, finan uh, financial modeling and analysis. And, and that really should shape the way in which you browse the web,right? The way you browse the web, the way you frame your answer, the kind of questions they pursue should be very personalized to kind of meeting the user where they're at.

I think the last part is sort of something that goes across domains of what models can do,right? So not just being able to do web research with text, but being able to combine that with abilities in coding, data science, even video generation,right?

So coming back to this example, if you're doing a due diligence, what if it could go and do, like, a lot of statistical analysis and actually build financial models to inform the research output that it gives you,right? Telling you, hey, why is this a good company or not?

Um, I should say Google doesn't give financial advice, and, you know, it's not a financial advisor. Um, but yeah, and so we're really excited about the potential. We think there's a ton of headroom to make research agents better, and we are really glad we didn't call this Gemini Deep Dive, which was our best name before, uh, before launching this feature.

Um, that's it. Thank you so much.

Mukund Sridhar14:54

Thank you.