Perplexity AI graphic logo
– Perplexity

Perplexity, the SoftBank-backed AI search app, is embroiled in a spat with Cloudflare over claims it employs “undeclared crawlers” to get around AI scraping restrictions.

Last month, Cloudflare brought in a block on AI crawlers – which visit web pages to capture information – making it so that such tools couldn’t access content on a user platform without their permission.

But in a scathing blog post published earlier this week, Cloudflare, one of the biggest internet architecture providers in the world, accused Perplexity of obfuscating its crawlers to access content.

“Although Perplexity initially crawls from their declared user agent, when they are presented with a network block, they appear to obscure their crawling identity in an attempt to circumvent the website’s preferences,” Cloudflare's blog post reads. “We see continued evidence that Perplexity is repeatedly modifying their user agent and changing their source ASNs to hide their crawling activity, as well as ignoring – or sometimes failing to even fetch – robots.txt files.”

Unlike traditional search engines, which simply provide a list of links relevant to a user query, Perplexity uses a foundation model to answer searches in natural language, providing conversational responses with links to sources.

Cloudflare claims it conducted a test that showed Perplexity was able to circumvent blocks designed to prevent crawlers from accessing its customers’ content for search responses without permission.

“We conducted an experiment by querying Perplexity AI with questions about these domains and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains,” the Cloudflare blog post reads. “This response was unexpected as we had taken all necessary precautions to prevent this data from being retrievable by their crawlers.”

Perplexity fights back

Perplexity has since been delisted by Cloudflare as a verified bot, but that didn’t stop the AI startup from coming out on the defensive.

A company spokesperson told The Verge that Cloudflare’s accusations were a “publicity stunt,” while a statement provided to TechCrunch described the post as a “sales pitch” for its blocking services and that the at-issue bot identified didn’t belong to the company.

In its own blog post published following the accusations, Perplexity argued that Cloudflare mischaracterized AI assistants as malicious: “They're arguing that any automated tool serving users should be suspect – a position that would criminalize email clients and web browsers, or any other service a would-be gatekeeper decided they don’t like.”

“Modern AI assistants work fundamentally differently from traditional web crawling,” Perplexity’s post reads. “When you ask Perplexity a question that requires current information – say, ‘what are the latest reviews for that new restaurant?’ – the AI doesn't already have that information sitting in a database somewhere. Instead, it goes to the relevant websites, reads the content, and brings back a summary tailored to your specific question.

“This is fundamentally different from traditional web crawling, in which crawlers systematically visit millions of pages to build massive databases, whether anyone asked for that specific information or not.”

The startup instead argued its “user-driven agents” only fetch content when a human user requests something specific, adding: “Perplexity’s user-driven agents do not store the information or train with it.”

The ire around such bots stems from the early days of generative AI, when model developers used bots to willfully scrape every corner of the internet for training data – a practice that’s become increasingly frowned upon by rights' holders unhappy with their intellectual property being used to train models.

OpenAI, for example, has GPTBot, which it uses to crawl content to train its foundation models to make them “more useful and safe.”

Perplexity, however, argues that its systems are not scrapers but AI agents that “work just like a human assistant. When you ask an AI assistant a question that requires current information, they don’t already know the answer. They look it up for you in order to complete whatever task you’ve asked.”

Instead, the AI firm claims Cloudflare confused Perplexity requests with unrelated traffic from BrowserBase, a third-party cloud browser service that Perplexity “only occasionally uses for highly specialized tasks (less than 45,000 daily requests).”

Perplexity contends that Cloudflare’s attack came either because it “needed a clever publicity moment” or it misattributed its traffic with BrowserBase's automated browser service in what it described as “embarrassing for a company whose core business is understanding and categorizing web traffic.”

The company further claims that Cloudflare hid its methodology behind its findings and “declined to answer questions” from Perplexity.

“If Cloudflare were truly interested in understanding the data they were seeing, how our systems work, or these fundamental concepts outlined above, they could have done what we encourage all Perplexity users to do. Just ask,” the startup’s blog post reads.