By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Meta, Alibaba AI Bots Threaten LGBT History Archive

The LGBT History Project, a volunteer-run online encyclopedia documenting the struggles and history of LGBTQ+ individuals, was nearly crippled by an influx of AI web crawlers from companies including Meta and Alibaba. Jonathan Harborne, the project's founder, established the site 15 years ago to preserve the memory of past generations' fight for recognition and respect, as well as the impact of the AIDS pandemic. The encyclopedia, which launched in 2011 and has since garnered over 50 million views, is archived by the British Library. However, a recent surge in automated traffic from AI bots attempting to scrape its content significantly slowed the site and increased operational costs. Harborne discovered the issue while modernizing the project's hosting infrastructure, migrating it to a new Amazon Web Services (AWS) server. Instead of improving performance, the migration led to a drastic slowdown. Upon analyzing his server logs with the assistance of Claude Code, Harborne identified an overwhelming volume of automated requests. On a single day, a crawler identified as belonging to Meta made approximately 26,000 requests, accessing not only published articles but also sensitive areas such as editing histories and login pages. This scraping activity resulted in the extraction of around 12 gigabytes of data. The increased server load necessitated a doubling of the server's capacity, directly increasing Harborne's personal expenses. He stated that he was "paying the cost for people like Meta to train their AIs" out of his own pocket. Meta was not the sole contributor to this problem; Harborne also reported that crawlers linked to Alibaba and other companies had been aggressively accessing the site over the preceding three weeks. This sustained pressure compelled Harborne to dedicate nights to learning advanced techniques for configuring Cloudflare, writing custom blocking rules, and implementing verification mechanisms to differentiate between human users and automated bots. The situation highlights a growing concern for digital archives and historical repositories facing the challenge of AI-driven data scraping, which can strain resources and threaten the accessibility of valuable online content.
Original source — read the full reporting at the publisher:
Read on Fast CompanyGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.