Interestana
Home/News/AI Firms Acquire Thousands of Used Books for Training Data
Fast Company3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

AI Firms Acquire Thousands of Used Books for Training Data

AI Firms Acquire Thousands of Used Books for Training Data

Artificial intelligence companies are increasingly purchasing used books in bulk, a trend that has led to a resurgence in sales for some used bookstores. Barter Books, a bookstore in Northumberland, England, typically sells between 2,000 and 3,000 books weekly. However, the owner recently fulfilled a single-day order for that quantity from a company based in Canada. Similar spikes in demand are being observed by other bookshops in the region. While the exact purpose of these book acquisitions by AI companies remains undisclosed, a recent court ruling provides a potential insight into their use, suggesting that some of these books may ultimately be destroyed.

The AI industry has faced significant scrutiny from the literary world regarding copyright infringement. Anthropic recently settled a major copyright lawsuit, agreeing to pay $1.5 billion to over 300,000 writers who had filed a complaint two years prior against the AI giant. OpenAI is also currently involved in multiple copyright infringement lawsuits. These lawsuits, filed by entities such as The New York Times, Encyclopedia Britannica, comedian Sarah Silverman, and various nonfiction authors, allege that OpenAI copied their books to train its large language models and has profited significantly from this alleged exploitation of copyrighted material.

The Anthropic lawsuit's resolution may have established a precedent for AI companies. A court ruling determined that it was not unlawful for Anthropic to train its AI on copyrighted works, provided that the company compensated for the books it utilized. A similar judicial decision was made last year concerning Meta. The acquisition of used books presents a more cost-effective method for AI companies to obtain vast amounts of training data compared to other sources. Furthermore, some publishers may be hesitant to directly supply copies to AI companies if they are aware of the intended use, making the used book market a more accessible alternative for these firms.

This practice raises ethical and legal questions surrounding the use of copyrighted material for AI training and the potential destruction of physical books. The demand from AI companies for used books signifies a shift in the market, where literary works are being valued not just for their content but as raw material for technological advancement. The financial implications for authors and publishers, as well as the long-term impact on the availability and preservation of literature, are subjects of ongoing debate and legal challenges within the rapidly evolving landscape of artificial intelligence development.

Original source — read the full reporting at the publisher:

Read on Fast Company

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next