Reddit logo
– Brett Jordan/Unsplash

Reddit has reportedly claimed AI model builders hungry for training data have been scraping its platform using the Internet Archive’s Wayback Machine.

The Verge reports that Reddit has blocked the tool, designed to archive older websites and pages for preservation, contending AI developers were using bots to scrape data from posts, comments, and profiles.

Reddit has made it so that the Wayback Machine can only index its homepage, meaning it would only be able to capture trending posts.

”Internet Archive provides a service to the open web, but we’ve been made aware of instances where AI companies violate platform policies, including ours, and scrape data from the Wayback Machine,” a Reddit spokesperson told The Verge.

Reddit’s move to block what some consider to be a key archival tool for the internet comes as platforms look to shore up their data after the Wild West that was AI developers freely crawling the web for training data.

Google/Reddit AI search screenshot
– Reddit

For Reddit, shoring up its AI bot defenses is a business decision. The platform has penned deals worth tens of millions of dollars with Google to provide access to its data to train the latter’s AI models – though the partnership has resulted in some less than stellar responses.

Last June, the platform blocked access to web crawlers like OpenAI’s GPTBot, but implemented some non-commercial exemptions for researchers and archival organizations, including the Internet Archive.

The Internet Archive’s exemption appears to be over, with a Reddit spokesperson saying: “Until they’re able to defend their site and comply with platform policies (e.g., respecting user privacy, re: deleting removed content), we’re limiting some of their access to Reddit data to protect redditors.”

Reddit’s data protectionism adds to the growing swell of ire towards bots willfully scraping every corner of the internet for training data. Legal disputes are already being brought by rightsholders against AI developers from authors, musicians, and actors over their intellectual property being used to train models without consent.

The social media platform also joined those ranks, filing suit against the Claude developer Anthropic in June, contending the Amazon-backed AI developer scraped data from its platform without permission.

Reddit’s decision to block the Wayback Machine comes a week after Perplexity, the SoftBank-backed AI search app, was accused of using “undeclared crawlers” to get around Cloudflare’s AI scraping restrictions.

The biggest internet architecture providers in the world accused Perplexity of obfuscating its crawlers to access content, a feat the AI app denied, labeling the accusation as a “publicity stunt.”