Perplexity AI has been caught in the crosshairs of another row over data scraping, this time with Reddit.
The forum giant filed legal action against the AI research app, accusing Perplexity and three web intelligence platforms of unlawfully harvesting data.
According to The Register, Reddit filed a lawsuit in the Southern District of New York against Perplexity, Oxylabs UAB, AWM Proxy, and SerpApi with the accusation of “industrial-scale, unlawful circumvention of data protections.” This was apparently achieved by the evasion of it and Google’s controls to scrape Reddit content directly from the search giant’s results.
Google, which is not a claimant in the lawsuit, has a $60 million deal with Reddit to access its data for AI training purposes, highlighted by Reddit’s lawyers with a note that Perplexity, in contrast, had no similar commercial agreement.
Perplexity was accused of securing Reddit data from SerpApi, Oxylabs, and AWMProxy, with the trio accused of “masking their identities, hiding their locations, and disguising their web scrapers as regular people” to bypass Reddit’s security measures across network traffic. An event in July of this year apparently saw a haul of close to three billion search engine results pages (“SERPs”), including text, URLs, images, and videos.
Oxylabs claims to operate “the world’s largest ethical proxy network," with an offering of more than 62,000 New York-based IP addresses.
In a statement provided to The Register, Perplexity said it was yet to receive the lawsuit, and that it “will always fight vigorously for users' rights to freely and fairly access public knowledge.”
“Our approach remains principled and responsible as we provide factual answers with accurate AI, and we will not tolerate threats against openness and the public interest," a spokesperson noted.
North Korean ‘flair’
Reddit in the lawsuit likened Perplexity to a “North Korean hacker," swiping a phrase from Cloudflare CEO Matthew Prince.
This summer saw Cloudflare accuse Perplexity, sans lawsuit nor mention of third-party tools, of obfuscating its crawlers to access content.
The web giant claimed it conducted a test that showed Perplexity was able to circumvent blocks designed to prevent crawlers from accessing its customers’ content for search responses without permission.
In response, Perplexity argued that Cloudflare mischaracterized its AI agents as malicious and scraper-like, instead of human-like assistants.
“They're arguing that any automated tool serving users should be suspect – a position that would criminalize email clients and web browsers, or any other service a would-be gatekeeper decided they don’t like,” Perplexity wrote in a blog post.
It also accused Cloudflare of conflating Perplexity requests with unrelated traffic from BrowserBase, a third-party cloud browser service that Perplexity “only occasionally uses for highly specialized tasks.”
Battle against the bots
Reddit’s lawsuit follows its blocking of the Internet Archive’s Wayback Machine this year, contending AI developers were using bots to scrape data from posts, comments, and profiles.
Reddit last year blocked access to web crawlers like OpenAI’s GPTBot, but implemented some non-commercial exemptions for researchers and archival organizations, including the Internet Archive.
Reddit also filed a lawsuit against the Claude developer Anthropic in June, contending the Amazon-backed AI developer unlawfully scraped data from its platform.
The moves echo increasing disgruntlement from creatives over their intellectual property being used to train AI models without consent. But while AI devs are in the spotlight, the likes of Google are not immune from current criticism.
Cloudflare’s Prince was reported this week to be pushing for increased regulation in the AI sector. The CEO argued that Google needed to unbundle its services, with the search giant claimed to be using the same web crawler to grab content for its AI products and services, alongside its search engine.
Comments