Amazon is systematically acquiring old, rare, and out-of-print books—particularly pre-2022 publications—to use as training data for AI language models. The strategy addresses a critical industry problem: as the internet becomes saturated with AI-generated content, companies need "pure" human-authored text to avoid model degradation (Model Collapse). This represents a significant shift in how AI companies source training data and highlights a growing competition for high-quality foundational datasets that guarantee human authorship.
← Back to all articles