AI companies destroy physical books – let's scan rare books before it's too late
In this article
**Marcus Webb** — Infrastructure engineer turned tech writer. Writes about AI, DevOps, and security.
> **Bottom line:** Aggressive data harvesting by AI companies for large language model (LLM) training is inadvertently accelerating the physical degradation and restricted access of rare and fragile books in major university and public libraries.
Reports from institutions like Oxford and Harvard indicate increased wear and damage to collections after repeated, often non-archival-standard scanning by private entities, particularly since late 2024.
Without a coordinated, publicly funded, open-source initiative to digitize these unique cultural artifacts *before* they're damaged or privatized, we risk losing irreplaceable knowledge to corporate black boxes and irreversible physical decay.
I tried to get my hands on a first-edition networking manual from the 80s last month. A specific, obscure text detailing early packet-switched network protocols.
What I found wasn't just a restricted archive; it was a physical casualty of the AI data gold rush.
The book, once pristine, showed clear signs of repeated, rough handling – creased pages, a visibly weakened spine, and a general wear that felt far beyond its age.
The experience left me convinced that we're actively sacrificing our most fragile historical records on the altar of large language models, and nobody's talking about the damage.
For years, as an infrastructure engineer, I’ve championed the idea of data permanence and robust archiving. The cloud, for all its ephemeral nature, promised a kind of digital immortality.
Yet, I’m seeing the very physical foundations of our collective knowledge being eroded by the insatiable appetite of AI models like ChatGPT 5, Claude 4.6, and Gemini 2.5.
These models, in their quest for ever more nuanced understanding, are consuming our libraries at an unprecedented rate, and the process is leaving behind a trail of physical destruction and intellectual privatization.
The Data Vampire at the Stacks
The core issue is simple: large language models require vast, diverse datasets to achieve their impressive feats of language generation and comprehension.
While much of this data comes from the open web, the truly valuable, high-quality, and often unique textual data lies within physical books, especially those published before the internet era.
Think historical documents, scientific journals, literary works, and, yes, obscure technical manuals. These are the texts that offer depth, nuance, and perspectives not easily found online.
AI companies, flush with venture capital and under immense pressure to deliver the next leap in model capability, have deployed armies of scanners into libraries worldwide.
Their objective is speed and volume, not necessarily archival preservation.
I’ve spoken with librarians who describe commercial scanning operations moving through their rare book rooms like data vacuums, often with equipment designed for mass digitization rather than careful handling of brittle, centuries-old paper.
These aren't the meticulous, white-gloved archivists of legend; they’re often contractors hitting quotas.
The physical toll is undeniable. Fragile bindings, already weakened by time, are stressed by repeated opening to 180 degrees. Brittle pages crack, especially at the spine.
Exposure to light, even for short periods under scanner lamps, accelerates degradation. Every scan is a micro-trauma to the book.
And when a single book might be scanned by multiple companies, each with its own proprietary process, the cumulative damage adds up fast.
The Black Box of Knowledge
Beyond the physical damage, there's a more insidious, systemic problem: the privatization of our shared heritage.
When an AI company scans a rare book, the resulting digital copy rarely, if ever, enters the public domain in an accessible, open format.
Instead, it becomes a proprietary asset, an input into a corporate black box.
The knowledge contained within that book is then re-processed, re-interpreted, and re-packaged by an LLM, its original form locked away behind API calls and usage fees.
Consider a unique 17th-century treatise on early mechanical engineering.
If Google scans it for training Gemini 2.5, that digitized text becomes a part of Google's intellectual property, not a publicly accessible resource for scholars or future AI researchers.
This creates a dangerous precedent where the digital representation of our most unique cultural artifacts exists only within the walled gardens of a few tech giants.
We’re effectively trading the physical decay of an artifact for its digital enclosure, often without any public benefit or right of access.
This isn't just about copyright, though that's a whole other minefield. This is about the fundamental principle of open access to knowledge.
As an infrastructure engineer, I understand the value of proprietary data.
But when that data is derived from unique, often publicly owned, or communally significant physical artifacts, the ethical calculus shifts.
We're creating a future where knowledge is mediated and controlled by algorithms whose training data, biases, and ultimate outputs are opaque.
The Reality Check: AI's Dual Nature
It's easy to sound like a luddite here, but I assure you, I'm not. AI has immense potential for knowledge preservation.
Imagine an LLM that can flawlessly transcribe centuries-old manuscripts, translate ancient languages, or cross-reference disparate historical texts in seconds.
These are powerful tools that could unlock vast archives of human knowledge that are currently inaccessible.
The problem isn't AI itself; it's the current commercial model of AI development that prioritizes speed and proprietary advantage over ethical sourcing, public good, and long-term preservation.
We're in a gold rush, and like any gold rush, the environment is taking a beating. The illusion of digital permanence is also a trap.
Just because something is scanned doesn't mean it's safe forever. Digital files require active management, migration, and robust infrastructure to survive.
A private company's internal copy, however well-managed, is not a substitute for a publicly accessible, openly licensed digital archive.
The current trajectory is unsustainable.
We cannot continue to sacrifice the physical integrity of our libraries and the open access of our intellectual heritage for the sake of incrementally better LLM performance.
The marginal gains in model capability are not worth the irreversible loss of unique artifacts and the creation of new knowledge silos.
Building a Digital Ark, Today
So, what do we do? We need a coordinated, global effort to build a "Digital Ark" for our most precious and fragile physical texts.
This isn't just a librarian's problem; it's an infrastructure problem, a data problem, and a societal problem.
1. The Open-Source Scanning Mandate
Any AI company, or any entity, wishing to scan books from public or university libraries for commercial AI training must be mandated to provide a high-resolution, archival-quality digital copy to a designated public repository under an open license (e.g., Creative Commons Zero or Public Domain).
This ensures that while they benefit from the data, humanity also benefits from its preservation and open access.
This mandate should have been in place since late 2024, but it’s not too late to implement it for future scanning efforts.
2. Publicly Funded, Preservation-First Digitization
We need significant government and philanthropic funding for libraries and archives to conduct their *own* digitization efforts, prioritizing preservation over speed.
This means using trained archivists, specialized equipment, and adherence to international archival standards (e.g., FADGI guidelines).
These efforts should focus on the most fragile, unique, and at-risk materials.
The output must be openly accessible. This is a massive infrastructure project, akin to building a national highway system for knowledge.
3. Developer-Led Archival Tools
As developers, we have a critical role to play.
We should be building and contributing to open-source tools for: * **High-accuracy OCR:** Especially for historical fonts, damaged pages, and diverse languages.
* **Decentralized storage solutions:** To ensure the long-term resilience and accessibility of digital archives, free from single points of failure or corporate control.
* **Metadata standards and validation:** To make these vast digital libraries searchable, discoverable, and usable for future generations, and to verify the provenance of digital copies.
* **AI for preservation:** Developing AI models that can *assist* archivists, for example, by identifying damage patterns or suggesting optimal preservation techniques, rather than models that are the cause of damage.
4. Ethical AI Data Sourcing Standards
The AI industry needs to establish and adhere to clear ethical guidelines for data acquisition.
This includes transparency about data sources, fair compensation for creators (where applicable), and respect for the physical integrity of source materials.
As consumers and developers, we should demand this transparency from the models we use.
The clock is ticking. Many of these rare books are already in a state of advanced decay. The AI gold rush is accelerating that process, often without a net benefit to public knowledge.
We have a narrow window — perhaps the next 5-10 years — to proactively capture and openly archive these unique artifacts before they are either physically lost or permanently locked behind corporate firewalls.
This isn't just about saving old books; it's about preserving the raw, unadulterated source code of human thought.
Are we truly comfortable with the idea that the definitive digital versions of our most precious texts might exist only within the proprietary datasets of a handful of corporations, or worse, cease to exist at all?
What's your take on this unseen cost of AI's rapid ascent?
---

