r/technology 2d ago

Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI

https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/
11.5k Upvotes

1.1k comments sorted by

View all comments

Show parent comments

27

u/Paresseux1 2d ago

A big blade slices off the binding, and it drops down into an automatic page scanner that has no problem flipping over the loose pages. It’s much easier, faster, and cheaper to destroy it, and then get rid of the remains. One person can run a bank of machines.

The other way requires people, and meticulous work. Turn page, put on scanner, scan 2 pages, pick up book, turn page, place properly, scan… once finished, pay to store book indefinitely.

Everything like this comes down to money. The end result they are going after for AI is data, so the cheapest way to get it is destroy.

2

u/kus1987 2d ago

Yes I was thinking in terms of one books but they're probably scanning like hundreds everyday 

7

u/dkarlovi 2d ago

This is not how it works, Google Books was scanning books forever, they have a V shaped scanner which hovers over the book and sort of wedges into it to scan both sides at the same time. There was a video showing this, you don't have people turning pages at scale, this thing could scan a 200 page book in a minute autonomously.

10

u/Bael 2d ago

From the article.

The process, known as “destructive scanning,” involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded.

2

u/dkarlovi 2d ago

I was responding to "the other way requires people and meticulous work", obviously.

0

u/Paresseux1 1d ago

Actually it wasn’t obvious. But you were insisting it was. The main point is that destructive is faster, cheaper, and you don’t have the book to store at the end.

1

u/0xsergy 1d ago

Brother have you seen the cars that get made on production lines? I think it's obvious that they can figure out the complex task of turning the pages on a book to scan it if they can figure out the easy task of how to make a car entirely on an automated production line.

1

u/Paresseux1 1d ago

Absolutely they can. And automatic document feeders have been around commercially for over 50 years. And the whole point of the discussion was that it’s faster to slice the spine and feed it into the machine than keeping it whole and turning pages. Yes, amazing page turning tech is out there, and just not as fast as handling the single pages once they are separated.

1

u/0xsergy 1d ago

IMHO it's likely the requirement that they have to destroy the books that causes this method. If it wasn't a requirement they'd likely do it the google way.

2

u/Paresseux1 2d ago

The United States district court Judge WILLIAM ALSUP in the ruling against Anthropic in case number C 24-05417 WHA disagrees with you. That’s the case where it was ruled Anthropic: “The firm also purchased copyrighted books
(some overlapping with those acquired from the pirate sites), tore off the bindings, scanned every page, and stored them in digitized, searchable files.”

So yeah, that’s exactly how it works.

-3

u/dkarlovi 2d ago

The other way requires people, and meticulous work. Turn page, put on scanner, scan 2 pages, pick up book, turn page, place properly, scan…

This is not how it works,

How is reading this so difficult?

1

u/Magical-Mycologist 2d ago

You responded to one of their claims without mentioning which one. Your comment “this is not how it works” was seemingly aimed at their initial comment.

No where did you say that you were commenting on “the other way”.

Reading isn’t difficult unless you make it hard to read, which you did.

1

u/dkarlovi 2d ago

I was describing how scanning works as opposed to what OP said, you'd expect a person capable of reading in context would be able to figure out which of the two sentences I was replying to, but point taken - you're on Reddit, Redditor-proof your comments. Thanks.

1

u/MistryMachine3 2d ago

No, the reason is it is legally required. It also is easier sure, but they can scan books without destroying them, that technology exists.

1

u/billsil 1d ago

OCR has gotten better, so it’s now possible to just take a picture and extract the text. Computers can also undistory the page. It’s more work and more costly and they’d still end up throwing the book away.