AI companies are chomping books, millions of books sacrificed to the ravening maws of the AI machines. And weirdly, I’m okay with it, but I’m going to start by making it sound hideous because that’s more fun.

VGT3 is a discreet Amazon operation that has been buying huge quantities of books and destroying them. The books are loaded into industrial slicers to cut off the bindings and spines; the pages are fed to high speed scanners; then the paper is pulped. See the cute picture above, the T Rex chomping into a book? It’s the VGT3 logo, painted on the outside of a giant Amazon building in North Las Vegas, and it’s the branding on employee materials for the VGT3 book executioners.

But it’s not just Amazon that is buying old books and tearing them to shreds. Anthropic’s secretive “Project Panama” began running books by the millions through its own guillotines in early 2024. Meta and SpaceXAI are known to have their own large-scale book scanning operations, as well as a number of other companies whose sole business is bookjacking. There’s a good chance that all the AI companies are scanning and destroying books.

Bookseller Scott Brown has recently done deep digging to learn as much as possible about the ongoing book destruction. His research turned up many shadowy companies purchasing and scanning books for the AI giants. One of the scanning companies, ARC Document Solutions,  claims to be scanning two million books every month. Brown estimates that “the total number of books destructively scanned to feed AI could easily be more than 50 million.”

Thanks to a tracking device, we know that the remains of at least one of the books (and perhaps millions more) wound up being turned into toilet paper.

The story swept the web a couple of weeks ago, feeding the tremendous demand for new reasons to hate AI companies. Spoiler alert: it’s not really a big deal; the hysterical reaction is another sign of moral panic in a world that can’t handle nuance and detail.

But it’s still a great story – and it’s a story that appeared first in one of my favorite science fiction novels, Rainbow’s End by Vernor Vinge, back in 2006.

Rainbows End

Vernor Vinge was a mathematician and computer science professor who wrote two award-winning SF novels in the 1990s with big space opera plots, followed in 2006 by Rainbows End, a near-future technothriller set in a world of ubiquitous augmented reality. In the novel everyone can see AR overlays on the real world – everything from maps and information to architecture makeovers and game skins.

The plot turns in part on public anger about a scanning project at the Geisel Library on the UCSD campus – an architectural marvel that looks like it was made specifically to feature in a science fiction story.

A giant tech corporation has invented a method to digitize the library’s entire book collection. Each book is shredded into confetti and sucked through a long tube lined with cameras that photograph the swirling scraps of paper from every angle. Smart software then matches up the unique rip marks on each scrap and reconstructs every book. The idea is that it’s faster to do it that way in a world of high speed computers and cameras than to carefully turn pages and scan them individually.

Today’s AI companies don’t shred each page but the result is the same – the act of scanning a book is also an act of destruction. It is startling to see how closely today’s reality mirrors a science fiction book written twenty years ago.

People have a visceral response to the idea of destroying books. The burning of the Alexandria library was an act of ultimate evil. In real life and in fiction, people are horrified when books are destroyed and they band together to fight villains who burn books – Nazis, Missouri Republicans, faceless state employees in Fahrenheit 451. In Rainbows End preservationists mobilize to oppose the book shredding at UCSD – staging sit-ins, sneaking into the building, and spreading the word online.

In 2026 there are no protestors holding signs in front of Amazon’s secret building in the desert, just a lot of very sternly written online articles and social media posts.

Granted, Rainbows End goes a little further – the protestors in the novel are being manipulated to cover up infiltration of a secret laboratory under the library developing mind-control technology. Obviously that’s crazy science fiction stuff! I mean, yes, the major tech companies are working on mind control through AI behavioral engineering and neurotechnology, and okay, key nuclear research for the Manhattan Project was done in basements and steam tunnels under Columbia University and, umm, maybe the novel isn’t that outlandish after all, come to think about it.

For now, let’s focus on the book destruction.

The very hungry algorithm

About a year ago secondhand booksellers noticed a huge spike in anonymous orders for books listed on marketplaces like Abebooks, Biblio, and Alibris. The buyers didn’t care about prices and titles were ordered indiscriminately – fiction, nonfiction, textbooks, children’s books, manuals, anything would do. It was obvious that no human being was involved in placing the orders; these were massive automated bulk orders placed by algorithms.

There was, however, one thing that all the books had in common: they were published between 1970 and 2022. And that unlocks the mystery.

The algorithms are searching the online bookselling databases for ISBN numbers, a unique number assigned to almost every book in print. You’ll find it above the bar code on the back cover and on the rights page in the front of most books. The ISBN system was standardized around 1970, so no books were purchased from before that date because the ISBN numbers make it so easy to automate the process and avoid duplicates.

But even more interesting is that there were no orders for books that had been published after late 2022. The first public release of ChatGPT was on November 30, 2022. More about that below.

404 Media published a story about the book destruction in July. Booksellers began trading stories about the orders they had processed and journalists did their own investigations. The secretive Amazon VGT3 facility was uncovered by an Apple Airtag hidden inside an obscure book.

When the books arrive, workers use hydraulic guillotines to slice off the spines and bindings, then feed the pages into high-speed scanners.

The destructive scanning has been driven by a legal loophole in a June 2025 summary judgment decision in a lawsuit against Anthropic. Judge William Alsup held that companies can use a fair use defense against copyright claims if they destroy the physical copies of books after scanning them to build an AI training library. If the AI companies kept a complete scan of a book and preserved the physical copy on a shelf, they would possess two copies of the copyrighted work, and that would expose them to copyright liability. The ruling creates an incentive to destroy the physical books. Kenneth Heafield, founder of the book-acquisition startup Last Token, openly posted that they ship books to partners and “importantly but sadly destroy the books to keep things legal per Bartz v Anthropic.”

Avoiding model collapse

The AI companies are not buying books to save them. The scanned books are not preserved. You can’t get ChatGPT to display page 163 of a scanned book.

Instead, the AI companies need more samples of highly-edited human-vetted prose to train their models. Books published before late 2022 are pristine and guaranteed not to include any AI content.

The training data teaches AIs about well-reasoned arguments and careful prose writing and storytelling. But the AI is not building a database of facts about each book and it is not storing a copy of the scanned pages.

The AI companies are trying desperately to avoid model collapse. Almost immediately after ChatGPT was introduced in November 2022, the web became contaminated with AI slop. If new AI models are trained on the billions of gallons of AI sludge already online, performance slowly degrades until all that is left is gray slurry. It’s AI inbreeding. If you’re old like me, you remember what happens if you repeatedly photocopy the same image over and over and watch it decay.

You’re already seeing the results when you read AI-produced news or posts online: the words are fine, just fine, but there’s a sameness spreading through our prose – it’s plausibly competent, relentlessly smooth, and statistically average.

Perspective

The online outrage machine has had a field day with this story. Stop the book chomping! Throw AI developers in jail! Make it illegal to scan and destroy books!

Let’s put copyright law aside for a moment and just look at the real world effects of the AI scanning so far.

There are zero credible reports of books being scanned that are exceedingly rare. There are zero reports of the last known copies of books disappearing, leaving a cultural void where they used to be. One book retailer involved in these sales, Half Moon Books, describes the titles that were ordered as “not what I would consider rare or antiquarian. They were more obscure academic commentaries on philosophy, politics and science. Nothing signed, no first editions and nothing that hadn’t been published in the 20th century or later.”

But – fifty million books! That’s a lot of books, right?

Sure, but let’s not get too weepy.

Exact numbers are hard to come by. The lowest estimate I found is that 320 million books are pulped, recycled, or discarded into landfills annually in the United States. Bookstores buy inventory with full return privileges; between 25% and 40% of all physical books shipped to bookstores are returned unsold. It’s far cheaper for publishers to pulp returned books than to put those books back into inventory, so between 65% and 95% of returned books are destroyed.

If you include discards from households, libraries, and thrift stores, estimates reach as high as 600-700 million volumes per year that are destroyed in the US. That’s just one country. This article claims that 140 million new books are pulped in France every year.

That’s the context when you think about how the evil machines are destroying the books: it’s true but it’s not all that significant. I can imagine a world where the AI companies are incentivized to pay authors instead of shredding books to fit through a legal loophole, but I can imagine a lot of worlds that are better than the cursed timeline we’re living through now.

There are many reasons to hate the giant AI companies. Many many reasons. Many many many reasons. Destroying books to make better AI is tone deaf and distasteful but personally I rank it below AI-driven loss of jobs, the looming global economic crash, the uneven and likely misguided construction of too many data centers, surveillance pricing, the poisonous effect of super-rich oligarchic tech CEOs, and the possibility that AI might destroy our society.

Maybe that’s just me. Read more books before they're gone!

Share This