The pipeline that fed creative work into AI systems
Researchers have traced, with increasing precision, how copyrighted work ended up training the AI models built to write, illustrate, and compete with the people who made it. The Pile, an 886-gigabyte collection of English text assembled by the AI research group EleutherAI in 2020, became one of the two most widely used training sets for large language models. One piece of it, a component called Books3, contained roughly 191,000 copyrighted book titles pulled from pirate website Bibliotik. A Danish anti-piracy group forced Books3 offline through legal takedown requests in July 2023, but by then, copies of the dataset had already spread across the internet.