Interestana
Home/News/AI Labs Buy Pre-2022 Books, Develop Watermarking
Search Engine Journal2 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

AI Labs Buy Pre-2022 Books, Develop Watermarking

Leading artificial intelligence research laboratories are actively acquiring books published before the year 2022. This strategic procurement aims to circumvent the inclusion of "AI slop," or content generated by earlier AI models, in their training datasets. The practice highlights a growing concern within the AI development community regarding the quality and integrity of data used to train increasingly sophisticated artificial intelligence systems. By focusing on pre-2022 literature, these labs are attempting to ensure their models are learning from human-generated content that predates the widespread proliferation of AI-generated text, thereby potentially improving the accuracy and originality of their outputs.

Concurrently, these same AI labs are quietly developing and implementing text watermarking technologies. Watermarking involves embedding subtle, often imperceptible, signals within generated text that can later be detected to identify its AI origin. This dual approach—curating cleaner training data and developing detection mechanisms—suggests a proactive stance against the potential misuse and proliferation of AI-generated content. The development of watermarking is particularly significant as it offers a technical solution to distinguish between human and machine-created text, a challenge that has grown with the rapid advancement of generative AI.

The strategy of acquiring older books and developing watermarking indicates a fundamental shift away from "content-at-scale" strategies, which prioritize sheer volume of content. This approach is increasingly being recognized as a losing bet in the long term, especially as AI models become more adept at identifying and potentially devaluing mass-produced, lower-quality content. The focus is shifting towards quality, authenticity, and the ability to verify the source of information. This move by AI labs implies a recognition that the future of AI development and deployment will likely depend on the trustworthiness and provenance of the data and the outputs generated.

This trend, as reported by Search Engine Journal, points to a sophisticated understanding of the AI ecosystem's evolving dynamics. The ability to train models on high-quality, verified data and to ensure the outputs are distinguishable from human work are becoming critical competitive advantages. The investment in pre-2022 books and the development of watermarking are concrete steps taken by major AI players to address these challenges, signaling a move towards more responsible and robust AI development practices. The implications extend to content creators, publishers, and consumers of information, as the landscape of digital content authenticity is being reshaped by these technological advancements.

Original source — read the full reporting at the publisher:

Read on Search Engine Journal

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next