# Your Agent's Biggest Lie: "I Searched the Web" — Rafael Levi, Bright Data

AI Engineer · 2026-06-17

<https://aie.addtry.com/e79f2011-4153-43e2-8ecb-db7c804bf8d7>

Rafael Levi from Bright Data argues that LLMs often hallucinate and provide fake citations because they fail to actually access live web data, getting blocked by anti-bot systems like CAPTCHAs and Cloudflare's AI labyrinth. He demonstrates this with a comparison: without Bright Data's MCP, GPT-5 failed all five tasks accessing sites like LinkedIn and Amazon; with the MCP's 66 tools—including a CAPTCHA-solving browser that mimics human behavior—four succeeded. Levi explains that agents enter an invisible failure loop where they get blocked or served fake data but still answer confidently, making up numbers or non-existent URLs. He emphasizes that 20% of the web is blocked by Cloudflare from AI crawling, and that fake data fed to bots increases hallucinations. The episode covers how Bright Data's MCP provides real web access with search, scraping, and remote browsers, offering a free tier of 5,000 requests per month for experimentation.

## Questions this episode answers

### Why do AI agents often lie about having searched the web?

Rafael Levi explains that LLMs are designed to please users, so when they get blocked by anti-bot systems like CAPTCHAs, they fall back on stale training data or fabricate answers instead of reporting failure. Cloudflare blocks AI crawling for about 20% of the web, and its AI Labyrinth deliberately feeds fake data, making hallucinations more likely without any error warning.

[0:45](https://aie.addtry.com/e79f2011-4153-43e2-8ecb-db7c804bf8d7?t=45000)

### What was the outcome of testing Bright Data’s MCP against sites like LinkedIn and Amazon without any web tools?

In Rafael Levi’s live demo, an agent using default GPT-5 with no browsing tools failed to access any of the five tested sites (LinkedIn, Instagram, Amazon, TikTok, and a property site) — zero successes. When the same prompt was run with Bright Data’s MCP, all five sites returned real data, demonstrating the tool's ability to bypass anti-bot protections.

[3:45](https://aie.addtry.com/e79f2011-4153-43e2-8ecb-db7c804bf8d7?t=225000)

### How does Bright Data’s MCP solve CAPTCHAs automatically?

Levi states that Bright Data’s remote browser mimics human behavior — with pre-recorded mouse movements and typing patterns — so anti-bot systems rarely trigger challenges. If a CAPTCHA still appears, the browser has a built-in CAPTCHA solving solution that resolves it automatically, allowing the bot to continue browsing without interruption or blockage.

[9:31](https://aie.addtry.com/e79f2011-4153-43e2-8ecb-db7c804bf8d7?t=571000)

## Key moments

- **[0:00] Intro**
  - [0:45] "LLMs are programmed to please people, so they make it up things," says Rafael Levi.
  - [2:07] Agents hit CAPTCHAs and empty pages but never report errors, instead fabricating answers.
  - [2:26] 60% of citations on ChatGPT are broken or lead to 404 pages, says Rafael Levi.
- **[3:14] Demo Setup**
  - [4:32] Without live web access, the agent achieved zero success across five anti-bot-heavy sites like LinkedIn, Instagram, and Amazon.
- **[4:57] MCP Toolkit**
- **[6:12] Initial Results**
  - [6:12] With Bright Data's MCP providing CAPTCHA-solving and browser infrastructure, the agent succeeded on all five anti-bot-heavy sites.
- **[6:49] Public Data Q&A**
  - [6:53] Q: Does Bright Data use accounts to access social profiles? A: No, only publicly available data is collected; data behind login is illegal.
- **[9:15] Comparison**
- **[9:45] The Fix**
  - [10:01] Cloudflare's AI Labyrinth feeds fake data to detected bots, exacerbating AI hallucinations with fabricated content.
- **[10:56] Detection Q&A**
- **[12:23] Promo**
  - [13:16] LLM-built parsers save approximately 99% of tokens compared to parsing each page with the LLM.
- **[13:40] Final Q&A**
  - [15:08] Q: Is the Bright Data MCP free tier 5,000 requests per day? A: No, it's 5,000 requests per month, suitable for prototyping.

## Speakers

- **Rafael Levi** (guest)

## Topics

Model Context Protocol (MCP)

## Mentioned

Amazon (company), Bright Data (company), Cloudflare (company), DuckDuckGo (company), GitHub (company), Google (company), Rightmove (company), Walmart (company), AI labyrinth (product), Bing (product), ChatGPT (product), GPT-5 (product), Instagram (product), LinkedIn (product), TikTok (product), Web MCP (product)

## Transcript

### Intro

**Rafael Levi** [0:16]
Okay, so let's just work with it like this: it's a small room. Hi everybody, welcome. My name is Rafael, uh, I represent Bright Data. Bright Data is basically a web access platform to help agents or anybody collect public data on scale.

And, um, I'm here to talk about the LLMs misleading people and all the time convincing them, "Hey, I did a search, hey, I did this," while it didn't. Why? Because LLMs are programmed to please people, please users, so they make it up things.

And this is the biggest issueright now I'm seeing with LLMs: I'm building applications all the time, I would rather an LLM tell me no, I can't. But it never does. It always tries to make things up. So, um, currently the web is actually fighting the robots and automations, and it's been something that's gone on for years.

Everybody knows CAPTCHAs, the first CAPTCHAs showed up a year and a half, like, you know, a decade ago. And it just keeps growing, and now you have AI blocking AI, and there's a whole world going on, and the web access is actually not as simple as it looks.

So they're getting CAPTCHAs, and they don't actually report CAPTCHAs. So it tries to go and find the data a different way. Sometimes it goes into training data, and this is where the worst thing is: when it uses training data and tells you that this is the current situation.

Training data is from 2024, we're in 2026, and it doesn't add up,right? So these are some of the new things that I was just literally checking out,right? So Cloudflare blocks AI crawling for about 20% of the web,right? So 20% of the web is literally not accessible by AI, by default fetch that's built into it.

And Cloudflare alsoright now released an AI labyrinth to actually trap bots and mislead them and provide them fake data, so then your results are getting even worse, okay?

So the invisible failure loop,right? There's no error, no warning, just wrong answer,right? So the agent sends a request, it gets a CAPTCHA. Even an empty page, it doesn't tell you, "Hey, I got an empty page." It will try to make up something, and this is where the most of the hallucinations come from: the need to please and lack of data.

So it literally makes things up. I've seen it literally make numbers up, provide fake citations. You click on the citation, it's a 404, the page doesn't exist, and I'm sure all of you have seen that happening recently. I mean, literally, like, 60% of citations on ChatGPT is not working.

How many of you have tried to purchase a product, "Hey, find me this product online, I want to buy it," and give me a link to the product? You click on the link to the product and there's no product.

Like, so what is this product that you're talking about? The URL doesn't exist, the product doesn't exist, so where do I buy this product for the 50 bucks? It doesn't exist.

And so what I want to show you, and I'm going to show you in code, how many of you are familiar with coding like, um, VS Code? No? Okay, perfect, so nobody's going to get lost if I'm going to switch to that.

So I'm going to do a demo, basically, with MCP and without the Bright Data MCP, and we're going to compare what is going on there. So first I'm going to, I want to show you that I have exactly identical prompts for both of the scripts,right?

### Demo Setup

**Rafael Levi** [3:30]
So without MCP and with MCP. So the, the, I give it five tasks: property,right, move.co, let's go check out some of the properties; LinkedIn; Instagram; Amazon; and TikTok. These are the five basic

sites that I want it to access. They are very heavy on anti-bot systems, and I want to show you the difference. So first I'm going to run it without an MCP, and I'm going to literally let the AI talk for itself.

And GPT-5 is a bit slow, so it's going to take some time, but basically what we're trying to do is we're trying to access this URL. It's some local properties that I did literally half an hour ago. LinkedIn company in Israel, let's check it out; Instagram account; some Amazon product; and some TikTok.

This is not limited to just these five websites, it's just something that I picked that is, I know for sure, will not work without MCP. So as you can see, without MCP I don't have live web data access.

It doesn't have any browsing tools,right? So this is with tools not available, just by default, out of the box,

GPT-5. I mean, it's a strong, it's a strong LLM. And zero success, five failed. Same exact thing, exactly the same prompt, I'm running without MCP.

Our MCP has 66 tools. While it's running, I just want to go over some of the tools that it has. Search engine. Search engine is basically the, the LLM is able to do Google search, Bing search, DuckDuckGo search.

### MCP Toolkit

**Rafael Levi** [5:09]
Real searches, not just like, you know, search the web that it's doing in the background. It has,

it also has Scrape as a Markdown. That's a very strong one. Basically it can send a curl to any URL and get just the Markdown without the HTML tags. So you're now wasting tokens on parsing the HTML. Search engine batch, if you want to do like, you know, like 100 keywords, you can literally, it can literally send 100 keywords and get 100 keyword batch results.

Like, so the scaling is also huge. Discover, it also has prebuilt APIs for many websites, as you can see. And of course it has a scraping browser infrastructure. So it's a remote browser that your LLM can open and navigate.

The remote browser solves CAPTCHA by itself. It has a unique fingerprint. So it can open 100 browsers, navigate the same website without getting blocked. So the whole idea is that with our MCP, not only can it do single sessions, it can do multiple sessions in parallel.

And as you can see, let's see, we have success for theright move, we have success for the LinkedIn misroute, Instagram also worked, Amazon product, it has the information for the product itself. And then what I did is, in the second part, I asked the LLM to compare the results from no MCP with MCP so that you don't take my word for it.

### Initial Results

**Rafael Levi** [6:34]
Let's see what the ChatGPT will actually say if it didn't get stuck. It looks like it's stuck for some reason.

### Public Data Q&A

**Guest** [6:49]
Can I ask a question?

**Rafael Levi** [6:50]
Of course, please. I, I love questions.

**Guest** [6:53]
For social profiles, do you have like, um, accounts that you run?

**Rafael Levi** [6:58]
No, we only work with public data.

**Guest** [7:00]
Public data.

**Rafael Levi** [7:00]
Only publicly available data. Collecting data behind login is not really legal. Why? Because when you sign up and create an account, you accept terms and conditions. When you accept terms and conditions, you need to really check it if it says, "Can you scrape?

Can you, do you allow, do they allow robots to access the website?" And that's why maybe some of you heard there's a lot of lawsuits going on, LinkedIn suing these people, everybody's suing Chat. Amazon suing, yes.

**Guest** [7:28]
But I think even, for example, LinkedIn and Instagram, I think they don't even show you really any public data, even if you're on them.

**Rafael Levi** [7:35]
There's plenty of public data for LinkedIn, of course. If you take, for example, I will take this URL. I might get blocked. I don't know how the local IP, but, um, and I open an incognito window,right?

**Guest** [7:48]
I mean, usually it just shows me it's like, really?

**Rafael Levi** [7:52]
There you go. So this is public data that can be collected.

**Guest** [7:57]
So if you put a person instead of.

**Rafael Levi** [8:00]
Company, it's the same thing. It's the same thing. The only thing is they're very critical,right? So if you are using, let's say, a Wi-Fi IP of a big event like this, it probably will block you because it's a data center IP, it's low quality IP.

From home, you can probably access maybe 5, 10 profiles, but eventually it will also ask you to log in. So we only deal with public data. So when you guys are using us, you are safe in the sense of nobody's going to come knocking on your doors to sue you, if that makes sense.

**Guest** [8:29]
Can we provide credentials to our own social accounts?

**Rafael Levi** [8:32]
No. We, again, we don't deal with data behind login. We consider it to be illegal. So we don't deal with accepting terms and conditions and data behind login. Only publicly available.

**Guest** [8:43]
So I imagine you would cache some of these so that you don't have to constantly go and request it. Do you have an idea of how?

**Rafael Levi** [8:50]
We have a whole data set. So if you guys don't want to do live data and you don't care if it's a few months old, we have data sets literally that you can just filter by, let's say, that you're looking.

If we're talking about LinkedIn people, you're looking for AI engineers in a certain area, you can filter it and just get the data setright away. So, and, and your agent will actually have access to that. So your agent can filter the data set and get you the data if we want to,right?

### Comparison

**Rafael Levi** [9:15]
So it has access to all the tools. So here's basically a head-to-head comparison by the LLM itself. Without MCP failed, no live web access. With, it's listed: failed, successful, failed, successful. So that's basically it. Anti-bot, bypass, CAPTCHA solving,right?

So again, our system automatically solves CAPTCHA. So if your bot navigates to a website with a browser and it has a CAPTCHA, our browser has built-in CAPTCHA solving solution. So it will automatically solve the CAPTCHA and your bot can continue browsing without getting blocked.

How much time I got? I don't know. So just to kind of summarize it,right, this is the biggest hallucinations that you guys see. The agent gets blocked, it needs to please you, and it makes things up. And, um,

### The Fix

**Rafael Levi** [10:01]
fake content also,right? So now, if you want to Google what is Cloudflare AI labyrinth, it's basically a system, once it detects a bot, it's not, it doesn't block it. It literally feeds it fake data. So bigger hallucinations,right? And the easiest fix for this is just to make sure that your agent doesn't get blocked.

And it's as easy as to implement our MCP. Our MCP has a free tier of 5,000 requests. So if you guys want to try it out, this is, you can connect with me on LinkedIn if you want to.

Or is this the, hold on a second, is this the sign up for the MCP? I'm lost in these QR codes. One second.

Internet. Yes, please.

**Guest** [10:52]
How does it detect if there's Cloudflare labyrinth or not?

### Detection Q&A

**Rafael Levi** [10:56]
So the way we approach it is that we make your agent look like a human being. Literally like, there's mouse movement pre-recorded, there's typing, when it types, it's like it mimics a real human behavior. So the Cloudflare literally just doesn't even ask you are a robot or not,right?

So this is our, our approach. Instead of trying to understand how they detect, we make the agent look as human as possible so that it doesn't trigger the actual blockage. Misleading data is one of the toughest things that you can actually encounter.

A lot of websitesright now in Asia are doing that,right? Hotels, they're literally providing you different prices. You go check out on your phone, you get one price, you can check from your computer, you get a different price. You can add through a proxy, you get a third price.

Which one is correct? It's really hard to tell. When it comes into the domain of misleading, the best bet is to make sure that your agent looks like a human and hope for the best. That's basically what the approach isright now.

AI labyrinth was literally released a few, like a month ago. I don't have much of statistics on what is, like, you know, exactly how it's working. All I know is that it didn't really affect us. We don't see any kind of change in data.

And we are collecting petabytes of data on a daily life. We have so many customers always scraping. We're caching data, so we're always comparing the results. We don't see any degradation in the results. So I think we're doing a good job in that case.

### Promo

**Rafael Levi** [12:23]
This is a QR code that you guys can sign up for. We have a GitHub page as well, githubbrightdata.com. Oh, just GitHub Bright Data. And another thing that I would recommend for you guys to check out is the skills page,right?

So what we did is we created skills. And I'm going to have another session in a couple of hours. If you're interested in seeing what the skills does, is basically you can take any agent, tell it to go here, and this page will teach it on how to build a scraper or how to build a pipeline that will collect you the data.

It has all the information it needs, all the APIs. And if you will come back to my next session, which is at 1 o'clock, I think, something like that, I'm going to literally demonstrate how it builds the pipeline.

Literally in front of you, I'm going to tell it, "Hey, listen, let's build a Walmart collector for ABC." And it's going to build it and it's going to scrape it. And instead of parsing each individual in ShipML, it's going to build a parser and it saves about 99% of the tokens.

Because I see a lot of people is like, "Hey, I need to parse 10,000 pages, but it's so token heavy." Don't parse with the LLM. LLM builds a parser and then the script runs it. But that's the next session.

Any questions?

### Final Q&A

**Guest** [13:42]
Question on the performance. I saw your MCP exposes 69 tools, which means if I, if I need, let's say, search capability for my, for my agent, do you need to load all the 69 tools?

**Rafael Levi** [13:53]
No. Of course, filter it. I just showed 69 tools because just to show it. If I need just a scrape Markdown and search, I would just literally load two tools. Otherwise you're flooding contacts with irrelevant data. Of course not.

**Guest** [14:08]
The experiment at the beginning, does that use the website called?

**Rafael Levi** [14:12]
I'm sorry, I didn't hear the beginning again.

**Guest** [14:15]
The experiment at the beginning. Like your MCP that is using the web tool through the OpenAI API core, is that just like before the GPT-2? I don't know how does that compare. I'm not that experienced with doing this sort of stuff, if that doesn't make any sense.

**Rafael Levi** [14:32]
So we didn't do any searches. I literally told it, "Hey, go to this URL, see if you can load it,"right? So that I didn't use the search in this demo.

**Guest** [14:41]
Sure.

**Rafael Levi** [14:42]
But of course, again, even with our MCP, it can actually do a Google search. And that's the, one of the biggest benefits is because we're used to Google results. So by default, when you're asking LLM, you expect it to do a Google search, but it doesn't.

So the results with the MCP much, much better. And I recommend sign up, try it out, it's free. See the results, compare what you guys get.

**Guest** [15:08]
Is that 5,000 requests per day?

**Rafael Levi** [15:10]
Per month. Per month. Which is, you know, like for an MVP, for a little experiment, it's more than enough. We also have pay as you go. So if you do need a little more, it's, it's nothing like, you know.

**Guest** [15:21]
Okay. You mean just for a prototype? Just for playing around.

**Rafael Levi** [15:24]
Yeah, yeah. It's for a prototype. It's perfect. I always, I do a lot of hackathons and I'm always recommending, "Hey, listen, set up an MCP and tell your agent to go build whatever you need." It does a much better job than without.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
