Books scanned and shredded for AI
AI companies are reportedly being scanned and then shredded. (Image: The Spinoff)

Booksabout 11 hours ago

How AI is coming for second-hand books

Books scanned and shredded for AI
AI companies are reportedly being scanned and then shredded. (Image: The Spinoff)

Second-hand book sales are booming… but who’s buying the books and what in the hell are they doing with them?

What have book sales got to do with AI?

Media outlets across the world are reporting that business for second-hand booksellers is booming… but not because of the cost of living or a collective surge of interest in acquiring lovely old books for personal libraries. It seems more sinister than that. 

Second-hand booksellers in Wales, England, Ireland, India, the US, Europe, Australia and New Zealand have reported unusual – and unusually large – orders of books. Earlier this month RNZ reported that Wellington’s Book Haven started receiving bulk orders from a Canadian company in April. The titles requested were extremely specific and included New Zealand cook books, tourist pamphlets and history books. Book Haven’s owner Anne-Marie Thorby did some digging and discovered that the same buyer was scooping up tens of thousands of books from all over the world, and all orders were destined for the same American warehouse.

So, what’s the issue? Maybe someone who really loves books keeps them in an American warehouse?

The issue is Panama. In January 2026 The Washington Post (paywalled) revealed evidence that AI giant Anthropic was running a bold new training programme for Claude, its popular chat-bot based LLM. The programme, dubbed Project Panama, involves running millions of physical books through a hydraulic powered cutting machine that slices the spine clean off so the loose pages can be scanned, the information fed to Claude. 

The Washington Post unearthed the project among 4,000 pages of legal filings in a copyright case brought by book authors against Anthropic. An internal document stated “Project Panama is our effort to destructively scan all the books in the world”. It also said “We [Anthropic] don’t want it to be known that we are working on this”.

Book Haven’s Thorby, who was aware of Panama, put two and two together and came to the conclusion that all of those strange, bulk orders she was receiving were likely to be part of Project Panama, or something similar (like in this case (paywalled) where Nova Media tracked a shipment of rare books to an Amazon AI training facility). “Our role is to keep books alive, to keep them in the world,” she told RNZ’s Guyon Espiner. “We’re book lovers. We want books for the good of humanity, not just to transfer knowledge for private profit, going into building LLM machines.”

Book Haven in Wellington (Photos: Claire Mabey/The Spinoff)

How many books are we talking? 

We don’t have a definitive number, but given the sheer volume of orders booksellers are receiving, it’s in the tens of thousands at least. The Washington Post article reported that “a project proposal by a vendor that ultimately worked with Anthropic noted that the AI company was ‘seeking an experienced document scanning services vendor to convert from 500,000 to two million books over a six-month period'”.

The Guardian, however, reported an Anthropic spokesperson saying the company procures its books from mainstream commercial markets, that training LLM models on book was a widely used approach, but that “none of our data acquisition programs buy and destroy rare or antiquarian books”. 

Haven’t AI companies previously been accused of training models off digital books, regardless of copyright?

Yes, and copyright law is tied up in this latest development too. There are a few potential reasons why Anthropic would procure, cut, and scan physical books to feed to Claude: one reason is that older books are free from AI contamination (paywalled) and are therefore superior training material for an LLM. The other reason is that US copyright law enables AI companies to legally use purchased second-hand books for the purpose of training their LLMs without the copyright owner’s permission. 

Authors and publishers are taking AI companies to court over the use of their work. Anthropic, Meta, Google and OpenAI have all been accused of hoovering up copyrighted material to train their LLMs without the copyright holders’ consent. Meta has denied any wrongdoing in a pending case, Google has not commented publicly on another, while OpenAI says in response to legal action underway that it respects creators and content owners and is reviewing allegations. 

In July 2026 Anthropic settled a copyright case against it in the US and paid out $1.5bn to authors and publishers, but in many ways the outcome was a win for Anthropic. 

How so? 

The judge ruled that using books to train AI was fair use under copyright law, but building a permanent library of pirated copies of books was not. Anthropic paid out because of the latter. 

Shortly after that ruling, another US judge found Meta was also covered under fair use. In that case, however, the judge said he ruled in favour of Meta only because the authors did not make a strong enough case for their argument that Meta’s AI would cause market dilution by flooding the market with books similar to theirs. He wrote: “No matter how transformative LLM training may be, it’s hard to imagine that it can be fair use to use copyrighted books to develop a tool to make billions or trillions of dollars while enabling the creation of a potentially endless stream of competing works that could significantly harm the market for those books.” He said his ruling in favour of Meta was not because “Meta’s use of copyrighted materials to train its language models is lawful,” but because “the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one”.

Ah, I see, so avoid pirating books and use physical copies to train AI and then you’re sweet? 

So far, yes. Well, in the US, at least. Under US copyright law, as well as the doctrine of fair use, there’s a “first-sale doctrine”, which lets a buyer do certain things with an item after they’ve bought it (re-sell it, for example) without notifying the copyright holder. It’s this law that enables Anthropic and its rivals to buy and use physical books to train AI.

Legal expert Dr Alex Sims, however, says the law is different in New Zealand. “New Zealand does not recognise fair use, instead we have fair dealing, which is considerably narrower than fair use. Anthropic’s actions [meaning scanning hard copy books under fair use] would be extremely unlikely to amount to fair dealing in New Zealand,” she wrote on Newsroom yesterday.

What about all these books being destroyed? 

If you’re a book lover, the concept of slicing off the bindings to get to the leaves just to train a bot might sound akin to skinning puppies. There’s an ethical and moral argument to be made against the destruction of books, especially potentially very rare ones. No doubt second-hand booksellers, receiving bulk orders are finding themselves having to weigh up difficult decisions. 

All this makes it seem quaint to buy a book, read it and then give it a home on the bookshelf.

Nah, go buy that novel.