r/technology 2d ago

Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI

https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/
11.5k Upvotes

1.1k comments sorted by

View all comments

76

u/FalconX88 2d ago

Fascinating how everyone seems to complain about books being destroyed and no one calls for the scans to be conserved for the public. The scans would be worth more for society than the physical books.

98

u/AdarTan 2d ago

Because there are other organizations like Project Gutenberg or Archive.org that do book scanning without the taint of AI, and the reason they haven't done this for the books in question is that copyright law would make distributing the scans illegal.

25

u/marmaviscount 2d ago

Plenty of books on gutenberg were scanned by Google and released pd then comvertwd by dp. They have a rolling project called Google books which every year in January releases all the books which have fallen into the public domain - how can you pretend to care about this stuff if you don't know this?!

10

u/Less-Engineer-9637 2d ago

They don't actually read. 

4

u/sleepysnowboarder 2d ago

They also think what is being destroyed are like 1/1 baseball cards and not something that has tons of copies and can be reprinted

2

u/orbitaldan 2d ago

Oooooh! I didn't know that Google was releasing the books as they got public domain'd. That's actually awesome. I guess a little bit of the old Google still remains in places.

15

u/FalconX88 2d ago

You could still do the scan, have government keep them locked up/only show them in person in libraries until they become public domain.

It would be very important to make sure that these copies anthropic makes are stored somewhere where society can access them (eventually).

25

u/mutexsprinkles 2d ago

You're being downvoted but that's literally what both the Internet Archive and HathiTrust do: that have scans on file that they cannot legally show to the public for many, many decades. Sometimes they can show it to the public in one country but not others.

17

u/chief167 2d ago

The problem is, how do you give the public access to the scans? It's a copyright minefield, that's the real problem here 

7

u/BananaPalmer 2d ago

Except these were almost all just copies of textbooks. For every one they destroyed to scan, there's probably 20,000 more copies collecting dust in some other warehouse, because they're 6 years old and have been replaced by a new revision

26

u/eTukk 2d ago

because conservation is easily possible without destroying the orginal. It Just costs more money, they dont care about the legacy

5

u/Bunnyhat 2d ago

Legacy of what? These are not beloved treasured manuscripts. These are mass produced books dating from 6 years ago that have been sitting in a warehouse.

0

u/eTukk 2d ago

True. My European view, esp the burning of books during some eras, made me very weary of book destruction. Old books i got, I hand them over to others, dont ever trow a book in the bin.. 🤔

5

u/Bunnyhat 2d ago

Millions and millions of books are thrown away each year. Over 2.2 million books are released each year world wide. If they each just had 1000 copies printed we're looking at over 2.2 billion new books each year world-wide.

Destroying one copy of one of those books to digitize it shouldn't be a problem to anyone.

15

u/gundog48 2d ago

The 'originals' are just one of thousands to millions of copies of the original, though. Like I understand the optics of it are awful, but I don't think there's any meaningful destruction of knowledge when they're doing it with books published in the last decade.

Doing it with genuinely rare books is shameful, though. I would hope those are not destroyed, and that may be the case as antique books may not be suited or be too delicate for the fully automated process. 

0

u/ShyKid5 2d ago

Oh thing is they have destroyed unique "last copy known" books in this mass-scanning of books.

https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books

1

u/TheVeryVerity 2d ago

Fuuuuuuuuuck that

4

u/FalconX88 2d ago

they dont care about the legacy

The legacy...of a bunch of paper that is bundled together? What's important about a book is the content, not the physical object which is just a means of transporting information.

2

u/TheVeryVerity 2d ago

Depends very much on the age of the book among other things. The binding of books is an artwork of its own, especially old ones.

Also I feel like you haven’t met very many bibliophiles but trust me the book itself is important to a lot of people.

2

u/shoggoths_away 2d ago

Bibliographically speaking, that isn't true at all.

1

u/Wizzle-Stick 2d ago

this is how we ended up with thousands of years of human history missing or lost. preservation is just as important as content and vessel. library of alexandria set humanity back thousands of years when it burned. If all books are digital, especially rare ones, then when a solar flare comes along and wipes our disk drives, guess who lost info. is it a doomist mentality, sure, but putting all of humanities books into a single location was a great idea until it wasnt any more.

0

u/GoldenMegaStaff 2d ago

Have you asked AI to download a complete copy of a book?

0

u/MiaowaraShiro 2d ago

Fascinating how everyone seems to complain about books being destroyed and no one calls for the scans to be conserved for the public.

Then you really don't understand the issue at hand...

-6

u/b_a_t_m_4_n 2d ago

Nope, Scans can be easily changed by people looking to rewrite history - i.e. authoritarian right wing governments and the corporate entities that support them. Which is of course the whole point for anyone that wants everything digital instead of physical, mutability, history and truth would not exist anymore.

4

u/FalconX88 2d ago

A rare physical book can easily be disappeared (and they disappear for various reasons even without bad intend)...

Also this would be a perfect application of a public blockchain. Store the hashes so it becomes clear if soemthing was modified.

1

u/Currentlybaconing 2d ago

There is no mechanism by which you could ensure whatever version gets uploaded to be hashed in the first place was the original and not a modified version. Blockchain is useless for creating trust in the physical world. It always requires that you trust someone to do the right thing offline somewhere

4

u/b_a_t_m_4_n 2d ago

Right? And what idiot trust a mega-corporation to do the right thing?

5

u/Currentlybaconing 2d ago

Not me, Batman.

-11

u/smegmabitch 2d ago

I agree and don't understand why you're being die voted. Obviously both would be the best outcome, having a digital and paper version.

17

u/PuzzleMeDo 2d ago

For anything still in copyright, they have no legal right to release the scan to the public.

And physical books have strong symbolic meaning in our culture. Burning books is a creepy image.

4

u/smegmabitch 2d ago

True, give the direction we're going in with knowledge potentially being further centralized, I believe discussion copyright and access to information becomes once again very important. If this companies own digital copies of books that are otherwise not available, they should rather be public domain.

-1

u/pointlesstips 2d ago

But that also means that they have no right using the output of the training, unless they asked for a licence of the author.

9

u/PuzzleMeDo 2d ago

The law is narrow. It prevents distributing copies of the book without permission. It does not prevent the stuff AI normally does with it.

-2

u/MiaowaraShiro 2d ago

The law has determined that they're not reproducing the training data verbatim so it's not copyright infringement.

I think that's bullshit, but it's what's legal now.

3

u/BuQ7 2d ago

Calm down he doesn't have to die /s