By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Reddit Restricts Old Reddit Access to Combat AI Scraping
Reddit is intensifying its efforts to curb data scraping and automated traffic by imposing stricter access controls on its "Old Reddit" interface. The social media platform recently began requiring users to log in to access the classic version of the site. In the coming months, Reddit plans to further limit access, mandating that users not only be logged in but also meet additional criteria, which have not yet been fully specified by the company. This move is a direct response to the increasing prevalence of AI bots that scrape vast amounts of user-generated content for training large language models and other AI applications.
The decision to restrict "Old Reddit" access is part of a broader strategy by Reddit to protect its data and monetize its content more effectively. The platform has been a rich source of text-based data, making it a prime target for AI companies seeking to train their models. By making it harder for bots to access and scrape content indiscriminately, Reddit aims to gain more control over how its data is used and potentially negotiate licensing agreements with AI developers. This aligns with a growing trend among online platforms to assert ownership over their data in the face of widespread AI development.
Previously, "Old Reddit" offered a simpler, text-heavy browsing experience that many long-time users preferred. However, its less dynamic structure also made it more susceptible to automated scraping. The requirement to log in is a significant hurdle for bots, as it introduces an element of user authentication. The forthcoming additional criteria suggest Reddit is developing more sophisticated methods to distinguish between human users and automated agents. The company has not disclosed the specific metrics or conditions that will determine eligibility for "Old Reddit" access beyond being logged in, leaving users to anticipate further details.
This initiative by Reddit highlights the ongoing tension between AI development, which relies heavily on large datasets, and the rights of content creators and platforms to control their intellectual property. As AI models become more sophisticated, the methods used to protect data must also evolve. Reddit's strategy of restricting access to its platform, particularly to older, more easily scraped versions, represents a proactive stance in this evolving digital landscape. The company's objective is to ensure that its vast repository of discussions, posts, and comments is not freely exploited by entities that do not contribute to the platform's ecosystem or compensate its users and creators.
Original source — read the full reporting at the publisher:
Read on The VergeGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.