AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive
Jonathan Harborne founded his LGBT history encyclopedia 15 years ago to document some of the struggles his peers had encountered in their lives. He worried that even as changing social values made life better for LGBTQ+ people, the struggles past generations had endured for recognition and respect could be wiped from memory, along with the pain caused by the AIDS pandemic. The site, the LGBT History Project, recently passed 50 million views since its launch in 2011 and has been archived by the British Library for posterity. But an onslaught of AI bots seeking to scrape its content nearly took it offline, bringing the site to a crawl while also making it more expensive to run. Harborne only realized what was happening when he began modernizing the site and its hosting, migrating it from a platform that had been in place since the project’s founding. He moved it onto a new AWS server, expecting it to become faster and more secure. Instead, the opposite happened. Harborne, who lives in London, connected the server to Claude Code and asked it to help diagnose the problem. Its analysis of his server logs pointed to an enormous volume of automated traffic. On one day, Harborne says, a crawler belonging to Meta made roughly 26,000 requests, trawling not only published articles but years of editing histories, login pages, and other parts of the site. He says it pulled around 12GB of data. That matters because Harborne pays for the project himself and was forced to double the size of the server just to accommodate the scraping. “I was paying the cost for people like Meta to train their AIs,” he says. “there’s me that’s paying for all this stuff out of my back pocket”. Meta wasn’t the only one. Harborne says crawlers linked to Alibaba and other companies have hammered the site over the past three weeks, forcing him to spend nights learning how to configure Cloudflare, write blocking rules, and introduce checks to distinguish humans from bots. “This is quite advanced stuff that I’m having to learn,” he says. He alleges that Meta’s crawlers did not respect the site’s robots.txt file, which is meant to tell crawlers whether a site owner wants them accessing particular parts of a site. (Meta, Alibaba, and Tencent, another company Harborne mentioned had hit his server, didn’t respond to Fast Company‘s request for comment.)
Comments
No comments yet. Start the discussion.