Books and Low-Background Steel
In 1919, in the natural harbour of Scapa Flow in Orkney, the German High Seas Fleet, interned there since the end of the First World War, was scuttled by its own crews. Despite the efforts of the Royal Navy to frustrate the process, more than fifty ships went to the bottom of the anchorage. For years afterwards, the site attracted salvage crews drawn by the ordinary scrap value of those sunken ships, and most of the fleet was raised and cut up between the wars.
That salvage took on a different significance after 1945 and the Trinity Test, the first nuclear explosion in New Mexico. Steel smelted after the test and the atmospheric explosions that followed was contaminated with fallout radionuclides. Steelmaking draws air through molten metal, and the air itself had changed, making the product unsuitable for radiation-sensitive instruments. Metal recovered from pre-war shipwrecks, so-called low-background metal, was uncontaminated. By then, only a handful of wrecks remained in Scapa Flow: most of the clean resource had been consumed before anyone knew it would matter.
Early in the development of generative AI, this became a popular metaphor for training data. In the same way that the Trinity Test marked a watershed for atmospheric radiation, the commercial release of models like ChatGPT is seen as a turning point for online content. Anything before that date was most likely written by humans; anything afterwards may have been generated or edited by AI. The analogy first circulated in late 2022, and early in 2023 Cloudflare’s then-CTO John Graham-Cumming began cataloguing pre-2022 digital data sources at lowbackgroundsteel.ai.
Broadly the metaphor makes sense, but there’s an important factor it ignores. Digital content from the web is a scarce but non-rivalrous resource: we can’t make any more of it, but it can be copied infinitely, and incremental use doesn’t affect the supply. There is, by contrast, a finite number of shipwrecks that can be harvested for low-background metal, and every use diminishes what remains. This is where books come into the picture, and where for me the metaphor really starts to work because, like ships, books are a scarce and rivalrous resource.
In the aftermath of the first wave of litigation over training AI on written content, particularly the Bartz v. Anthropic class action, it’s become clear that as well as ingesting large digital collections, AI platforms have been training on printed books. Earlier this year the book wholesale giant Ingram wrote to publishers allowing them to opt in or out of bulk sales of books to, presumably, AI companies or intermediaries acting for them. The Washington Post documented Anthropic’s Project Panama, a coordinated effort to buy millions of second-hand books and—in the words of the company’s own internal documents—“destructively scan all the books in the world” (a copy of each, that is, not every copy). Most recently, 404 Media revealed that industry data platform ISBNdb has pivoted from providing metadata to brokering bulk acquisitions of up to a million books for AI labs, with the pitch that pre-2022 publishing is structurally free of AI writing.
This is much closer to salvaging shipwrecks than anything in the digital version of the metaphor. The books procured are scanned destructively: spines guillotined, pages fed to scanners and residual matter sent for pulping. Everyone involved appears to understand how bad this looks. ISBNdb warns prospective clients about “the optics problem” and works under NDA, and Anthropic’s unsealed documents show the company anticipated a negative public reaction.
For modern books, available in a range of formats including print-on-demand and digital editions, the process seems merely wasteful. For scarcer books, it risks them being removed from circulation. ISBNdb makes a selling point of the fact that millions of valuable titles were never digitised, so the copies being destroyed are disproportionately those with the fewest substitutes.
There’s a legal irony in this. In Bartz v. Anthropic, Judge Alsup found that Anthropic’s format-shifting of purchased books was fair use partly because the originals were destroyed. The digital copy replaced the print copy, one for one, and was never shared outside the company. Whatever one makes of the legal reasoning, the incentive it creates is clear: it is currently legally safer to pulp a book than to preserve it.
There is also an important historical caveat. Demand for low-background steel has largely faded. Atmospheric radiation declined sharply after the 1963 Partial Test Ban Treaty, and steelmaking improved to the point that ordinary modern steel is now adequate for most uses. The premium on the scarce resource was temporary. The same could happen with content: better filtering, more sophisticated use of synthetic data, or better provenance standards for digital text might all reduce the value of pre-2022 print. Though the analogy runs the other way too: fallout decayed because nuclear testing ceased, while AI-generated text is accelerating. I don’t know which effect will prevail, and I’d be suspicious of undue certainty on that.
What seems clearer is that the status quo captures little value for the authors and publishers who made the original books. Every salvage purchase is a secondary-market transaction: the bookseller is paid and the author and publisher receive nothing—not even a clear sense of the size of the market (I wrote about some of the limitations in the data on the secondary market in a piece for the Society of Authors in 2021). ISBNdb presents this to its clients as a feature: the books it supplies have already discharged their financial obligations to their creators. Perhaps that is true in a narrow legal sense. But if AI labs will pay over the odds for physical copies purely because of when they were printed, that is a price signal publishers should be paying attention to. It suggests that the industry is allowing the second-hand market to capture a significant proportion of the commercial value in pre-2022 human-authored text. There is an obvious deal to be done: publishers owning their archives and providing licensed digital access to backlists would be cheaper and faster for AI platforms than industrial-scale salvage, better for authors, and would spare everyone headlines about the destruction of millions of books.
It also brings me to a practical suggestion. Fair use is a balance: society allows unlicensed copying because it expects some public benefit in return. When the copying destroys a scarce original, and the resulting file sits in an AI platform’s private corpus, the public side of that bargain looks thin. So, any AI platform destructively scanning books under cover of fair use should ensure that a digital master of each title is deposited with a national library or trusted archive, similar to statutory legal deposit or the HathiTrust model. From the perspective of an AI lab, nothing about this threatens the fair use defence in Bartz v. Anthropic, which depended on the copies not being shared. The mechanism and files already exist. The cost to the labs would be a rounding error, and it converts an optics problem into a social good. And, significantly, it holds whichever way the uncertainty resolves. If pre-2022 print holds its value, we will be glad the masters were preserved. If the premium evaporates, the deposit may be the only lasting good to come out of the salvage era.