1 min readfrom TechCrunch

Amazon, which started off selling books, is destroying rare texts to train AI

Our take

Amazon’s expansion into AI is raising critical questions about data sourcing. Reports indicate the company is destroying rare books—incredibly valuable resources for training Large Language Models—to feed its AI systems. This practice highlights a growing tension: while vast datasets are essential for LLM development, the reliance on irreplaceable historical materials presents a significant ethical and preservation concern.
Amazon, which started off selling books, is destroying rare texts to train AI

The recent reports detailing Amazon’s alleged destruction of rare books to train large language models (LLMs) have ignited a necessary, if uncomfortable, conversation about the ethics and resource demands of AI development. While the practice, if confirmed, is deeply troubling, it also underscores a critical reality: the insatiable appetite of these models for data, and the lengths to which organizations might go to satisfy it. The core issue isn't simply the loss of physical artifacts—though that's a significant concern in itself—but the implications for data provenance, intellectual property, and the long-term sustainability of AI innovation. We’ve consistently advocated for more thoughtful data management practices, and as highlighted in [5 Python Libraries That Make Data Cleaning More Enjoyable], efficient and responsible data handling is paramount, even before considering the ethical dimensions of source material. The reliance on scraping and, potentially, destructive acquisition of data reveals a need for more creative and sustainable approaches to model training.

The argument that rare books are invaluable for LLM training—because online data has already been extensively utilized—is a compelling one, particularly given the observed shifts in LLM behavior, as demonstrated by the rapid development of identity in models like Qwen2.5, as discussed in [It only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine"]. The nuances of language, historical context, and unique stylistic elements contained within these texts can contribute to richer, more accurate, and more human-like AI responses. However, this benefit cannot justify the destruction of irreplaceable cultural heritage. The current trajectory suggests a prioritization of immediate gains—enhanced model performance—over long-term considerations of preservation and ethical responsibility. The development of efficient techniques, such as those explored in [How to make any Sparse Attention / KV Compression look good?], offer potential pathways to reduce the data demands of LLMs without resorting to such drastic measures.

The situation highlights a broader challenge within the AI ecosystem: a lack of transparency and accountability regarding data sourcing. Many organizations are operating in a grey area, driven by the relentless pursuit of competitive advantage and the pressure to deploy increasingly powerful models. This isn’t about demonizing AI development; it's about advocating for a more responsible and sustainable approach. The reliance on potentially unethical data acquisition practices ultimately undermines the credibility and trustworthiness of AI systems. Furthermore, the narrative risks alienating the public and hindering the broader adoption of AI technologies. A focus on synthetic data generation, curated datasets, and innovative training techniques could mitigate the need for destructive data collection, fostering a more ethical and resilient AI ecosystem.

Ultimately, the Amazon controversy serves as a stark reminder that AI progress cannot come at any cost. It compels us to re-evaluate our data sourcing practices and prioritize the preservation of cultural heritage alongside the pursuit of technological advancement. The question moving forward is not *if* we can build more powerful AI models, but *how* we can do so in a way that aligns with our values and respects the legacy of human knowledge. What frameworks and regulatory mechanisms will emerge to ensure the responsible and ethical sourcing of data for AI training, and will organizations proactively adopt these standards before they are mandated?

Rare books are incredibly valuable for training LLMs, since these models have already trained on whatever's available online.

Read on the original site

Open the publisher's page for the full experience

View original article