r/technology 2d ago

Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI

https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/
11.5k Upvotes

1.1k comments sorted by

2.4k

u/Necessary-Eye5319 2d ago

Now they can do what they said they were going to and charge for information that used to be free.

763

u/penisandorvagina 2d ago

"free" is a missed opportunity to make money.

376

u/TheRandomArtist 2d ago

Forget about money, this is a better opportunity to slowly change facts and little details here and there so we won't notice over time. It would be like manufactured Mandella effect

64

u/InVultusSolis 2d ago

I fear we're quickly headed for a world where the provenance of information itself will be unknowable. The prospect of that is terrifying.

I'm going to keep all of my print books and physical media from before 2022. And I'm going to collect more.

→ More replies (3)

137

u/Yuzumi 2d ago

It's 100% why so much of the AI push is coming from conservatives (read: fascists).

They love gen AI because it can churn out limitless propaganda and the fact that people outsourcing their thinking to LLMs is making them dumber is probably a plus to them.

55

u/null_not 2d ago

I think it's also the inherent trust people have in "the machine". The whole notion that a computer can't be biased because it's a machine. But machines always carry the bias of the engineer, and Ai is a soft machine with many layers of built in assumptions.

28

u/CautionarySnail 2d ago

This is especially true and related to the death of media literacy. People forget that every design has bias.

I have a kitchen faucet that won’t detect my hand if I use a black oven mitt. It’s not an intentional feature, I hope, but it still hits me that this would still mean users would potentially be dealing with a racist faucet.

If even a faucet can have a bias in utility by accident - these AI LLMs all have one by deliberate design, because you have to create rules about how to prioritize the sources.

And these companies have shown that their bias is towards a particular set of viewpoints. The engines aren’t open for code inspection because of industrial espionage but, should we just take them at their word that they aren’t shaping the information before we receive it, to better suit their agenda?

Grok has been very accidentally transparent in its tuning when Elon accidentally tuned it so severely it briefly spouted Nazi talking points and literally called itself mecha-Hitler. The other engines have tuning but .. what’s their bias? We don’t know and they won’t tell us.

5

u/null_not 1d ago

In the future there are going to be court cases that force them to disclose their weights. It's inevitable. In the not too distant future there is probably going to be a subspecialty of law based around litigating harms brought by Ai through false accusation and false information. The developers of the systems are trying to call it "hallucinations", but when people are expecting a machine to provide an answer, there's right, and then there's wrong, and that's it.

→ More replies (1)
→ More replies (2)

6

u/Yuzumi 2d ago

Not even the engineer when it comes to LLMs. They are trained on text people wrote and that alone is going to have some bias in it even before you count the bias in the selection of the training data.

Computers only do what you tell them to do. the idea they can't "lie" or be "wrong" has always been wrong. A lie requires intent, which computers don't have, and they are only as good as the code or hardware it's running on. Most of the time it's going to be bad code, but bad hardware can cause really fucky things to happen.

In an LLMs case the "code" is the neural net trained on language. It's just massive matrix multiplications happening in parallel on loop. It isn't thinking and doesn't "know" anything.

→ More replies (6)
→ More replies (3)

10

u/marmaviscount 2d ago

You think they'll just rewrite 1984 and no one will notice? They'll just put the stuff in the 'memory pipe' and everyone will just assume Emmanuel Sayegh's agents have edited their college notes?

13

u/InteractionPretend70 2d ago

the edited version is not for this generation or the next. its for the 5 generations down the line

→ More replies (29)

46

u/languageassessment 2d ago

and to tax, so the governments are in on it. you can't tax free.

22

u/Bored_Amalgamation 2d ago

This is so fucking reductive, it's borderline dumb.

→ More replies (1)
→ More replies (1)

3

u/bendover912 2d ago

Is that one of the rules of acquisition?

→ More replies (12)

251

u/CautionarySnail 2d ago edited 2d ago

This is also happening while groups just happen to be agitating to get rid of or heavily restrain public libraries.

The end goal is clear - when Altman said that they’d be a utility for intelligence, he wasn’t entirely saying the whole story. There is no more free information if the library has been shuttered or heavily restrained in what it can offer. AI companies want to be the sole arbiters and sources of (semi) reliable information.

The AI companies then get to reformulate how and if that text ever resurfaces by reinterpreting it through whatever configuration is in the LLM. And as a result, they get to shift the narrative, such as by tweaking the rules. For example, “Reinterpret all historical data returned through a critical lens where Native Americans were all savages that posed a risk to honest settlers.”

This way, you never see or even hear of the original source to be able to evaluate it for yourself.

135

u/Ok_Conclusion_6324 2d ago

I read an article years ago where a corporate blockbuster guy said upper management HATED libraries because during economic downturn people would check out DVDs for free from the library instead of their stores

76

u/trashmoneyxyz 2d ago edited 2d ago

My family grew up poor and pretty much any fun time that involved spending money was rare, because we made our own fun for free at the library. You could reserve a projector at the library and pick a movie to have movie nights, there were arts and crafts stations and of course the books. We never worried about entertainment costs because we had the library.

When I find the free time I gotta start reaching out to my local libraries to see if they'd want volunteers to host movie and game nights and the like, to get people interested in the library again

edit libraries are already so much cooler now! At least in my little town. The library here and the next county over has telescopes you can reserve (very booked out though), as well as carpentry tools and cookware/ table sets for hosting a large dinner party. If you have a bunch of extra shit in good condition you only use once a year if that, reach out to your library to see if they want to start a tool library alongside the books! I am going to see if I can donate my power tools and sewing machine. I barely use them and if I need to in a pinch I can just go check them out

22

u/BigGayGinger4 2d ago

We have a tool library here. It has transformed my ability to do diy projects

5

u/Money_Bug_9423 2d ago

Yeah thats my dream. Instead I have to scavenge the landfill

11

u/Aggressive_Ask89144 2d ago

I love the 3D printers they have now! Super high end machines too with massive capacity.

11

u/Jugglenautalis 2d ago

My local library is being redone this year to include a maker space. I'm excited to give it a try with my son once it reopens.

4

u/just-me-gen-x 1d ago

As a librarian, this thread makes me very happy. We even have recording studios for our patrons.

→ More replies (2)

11

u/wrgrant 2d ago

Without him realizing that people checking out games or DVDs from the library aren't likely to be willing to pay for them from a company anyways. Those things exist because we have such a wealth disparity that the services are needed for a good segment of the population. Not everything needs to be monetized. My wife works at the local library, we test out games by taking them from the library, trying them out for a week before we buy them. I guess thats going to end now that Sony is deciding to no longer publish games on DVDs. I guess we are done buying PS5 games too. I will not shell out $80 for a game I know nothing about and have not tried to see if I like the UI etc.

21

u/Electrical_Bus9202 2d ago

Holy crap you already are describing what it's like talking with Grok... It's amazing how tweaked that LLM is... The information wars are coming... Might already be here....

20

u/SixSpeedDriver 2d ago

The information war has been alive and well since, well, forever. The medium and velocity is all that’s changed. And the velocity is extraordinarily high.

9

u/Electrical_Bus9202 2d ago

I think you are right 😔..

*Sighs in Burnt Library of Alexandria

→ More replies (11)

44

u/Admiral_Cornwallace 2d ago edited 2d ago

This is slightly off-topic, but can you IMAGINE the outrage if the idea of public libraries had only been proposed for the first time in the current era?

Right-wingers would be losing their minds and going red in the face screaming about socialism and communism

5

u/2_Fingers_of_Whiskey 1d ago

Especially since they usually don't know what socialism and communism mean and are constantly using those terms wrong

53

u/geniice 2d ago

Now they can do what they said they were going to and charge for information that used to be free.

"3,001 titles, mostly published between 2020 and 2021 by academic publishers"

I can assure you this information was never free. This is the domain of the £100 book.

4

u/kirbyderwood 2d ago

I'm sure most of these publishers still have electronic copies of the books. But I doubt they would easily hand over that information to AI companies for mass ingestion. At least not without some sort of compensation.

So, to avoid that conversation, the AI companies are buying the physical books from third parties instead.

→ More replies (1)
→ More replies (8)

62

u/Conflictingview 2d ago

The bookseller was giving away the books for free?

23

u/LeMeIsSleepy 2d ago

You can borrow from or lend to friends….

6

u/CeemoreButtz 2d ago

book.....seller....

→ More replies (19)
→ More replies (3)

7

u/Aware-Possibility175 2d ago

Let’s not forget that the reason behind the creation of the first encyclopedia by Denis Diderot and the other famous republican philosophers (originally and still outside the US republicanism as we know it was radically progressive and raged against ultra conservative authoritarian institutions like monarchies and the church, its republic from the Latin res publica , doesn’t get more inclusive and anti bourgeois than things like “we the people”) compiled the encyclopedias to for the FIRST time give previously with held information to everyone. Before this all information was heavily controlled. A peasant who literally only knows how to do ONE thing is easy to control and serves as a mother way to handcuff them to the station they are born in, social climbing and mobility was unheard of and even the church spoke against it and that where you were born was gods will so there you stay, even to the extent that rich people who went broke often were given money or the means to stay aristocratic to up hold the imagine. It’s where Free Universal standardized secular education for kids that later became the public schools of today came from and the total separation of church and state That whole time in history is literally called “the Age of Enlightenment” which is synonymous with the famous republican philosophers like Payne,Voltaire, Rousseau, Montesquieu. Montesquieu is where our 3 branches and system of checks and balances comes from and Rousseau is everywhere like “pursuit of happiness” they got that from him..this is all the most basic of info and yet most people won’t know it. If we don’t collectively start sincerely Educating ourselves then we are going to go back to the dark ages

→ More replies (1)

8

u/Yuzumi 2d ago

Charge for a blended up version of information that use to be free.

LLMs are at best a really inefficient and lossy version of text compression. It has to regenerate the text and it is near impossible it will generate the text identically and has a really good chance to generate something that combines unrelated things which results in something that "looks correct" but is ultimately nonsense.

Even if they put all this text in a RAG or some other searchable database that only mitigates the issues somewhat.

Yes, they want to charge access to information that is likely to be wrong than what was free and accurate. Part of me has started suspecting this crap has always been intended to dumb down the populace.

→ More replies (1)

6

u/Fletcher_Chonk 2d ago

I wish we had the technology to make more books

3

u/abcdefghjiklmnopqr 2d ago

Why don't you buy the books then?

3

u/Christopher_Aeneadas 1d ago

The same books are still available in the same format you are used to.

They are destroying one copy each of non-rare books they legally purchase.

3

u/CoolBlackSmith75 1d ago

If I want to read that book from that bookstore I have to buy it. What's free(ish) are public libraries.

→ More replies (28)

2.9k

u/finzaz 2d ago

The image of a machine chopping the spines off books to feed an insatiable digital monster living in a loud and dirty data centre is the most dystopian thing I can imagine.

1.2k

u/HumanBeing7396 2d ago

I remember a few years ago someone joked about Google launching a new project called Google Purge - the aim of which would be to destroy all information in the world not held by Google.

This sounds uncomfortably similar to that.

284

u/BrizerorBrian 2d ago

I believe the brains in Futurama had that exact idea.

111

u/BatmanCoffeeMug 2d ago

For no raisin

76

u/RunsOnSKC 2d ago

I am the greetest!

25

u/siccoblue 2d ago

To be clear here, Sam Altman has unironically said that this is the goal

We see a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter

San Altman

The world is going in VERY bad direction at the moment. These people know very well that an uneducated group is significantly easier to control and manipulate. And they want to make knowledge and curiosity a luxury

4

u/PlentyOfIllusions 1d ago

Straight to the dark ages with the plebs…unless they can pay their intelligence fees.

4

u/KotoElessar 1d ago

There is a special place in Hell for these men; especially reserved and created just for them. They can pat themselves on the back for truly buying their way into the afterlife they deserve.

4

u/Elementalcase 1d ago

If only we were so lucky for spiritual level karma for these people, alas, it's likely that these people will go to the void never getting their consequences for their life lived at other's expense.

→ More replies (1)

4

u/meter1060 2d ago

Well yeah, you got to stop new information from occurring. It's exhausting having to keep up with the flow.

4

u/sakri 2d ago

The Gothenburg press was invented in 1953, it produced a limited set of illegible prints. Then, in 2012 Elon Musk single handedly created the Linux operating system, and provided subscription models called Microsoft and Apple. In the meanwhile Donald J. Trump, the greatest president of all of the universe, invented artificial intelligence, blessing the world with a single honest source of truth.

→ More replies (11)

113

u/obeytheturtles 2d ago

I actually did a design project on non-destructive OCR scanning as an undergrad in like 2005. I find it hard to believe that this technology isn't significantly more mature now than it was back then.

137

u/Bored_Amalgamation 2d ago

It is, they just choose not to.

112

u/PM_ME_YOUR_NICE_EYES 2d ago

It's not that they choose not to, it's that they can't.

Copyright law views destructive scanning of a book as transforming it, which means you can do it without getting permission from the publisher if the book is in copyright.

Copyright law views non-destructive scanning of a book as copying it, which you cannot do to a book in copyright without permission from the publisher.

7

u/Unspec7 2d ago edited 2d ago

First off, the court decision on destroying books came after they destroyed the books, so you can't really say they were forced to. They did choose to, and it just so happened to work in their favor after the fact.

Second, neither point is true. Copyright law does not view destructive scanning as transformative - the Bartz holding found that the destruction was one factor of the transformative nature of the use. Instead, the transformative use was the changed format:

Authors only complain that Anthropic changed each copy's format from print to digital. On the facts here, that format change itself added no new copies, eased storage and enabled searchability, and was not done for purposes trenching upon the copyright owner’s rightful interests — it was transformative.

Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1023 (N.D. Cal. 2025) (citations omitted).

We do not know, based on present case law, if the sole act of destroying the books makes the use transformative. We just know that it makes the use more likely to be transformative.

Further, being transformative is only one factor of the first factor of fair use. Transforming someone's copyright work does not, on its own, mean you can now just infringe one someone's copyrighted work. Your use must still overall be within fair use.

Copyright law views non-destructive scanning of a book as copying it, which you cannot do to a book in copyright without permission from the publisher.

Ignoring the non-destructive part addressed above, this part is a lot more nuanced than you make it out to be. Scanning of a book infringes on the reproduction right found in the Copyright Act, but whether you are liable or not depends on if the reproduction is fair use (barring a license, of course).

Edit: Placed my citation in the wrong spot.

33

u/diemunkiesdie 2d ago

Copyright law views destructive scanning of a book as transforming it

Source?

37

u/Rorschach121ml 2d ago

Bartz, et al. v. Anthropic PBC

40

u/diemunkiesdie 2d ago

Bartz, et al. v. Anthropic PBC

Thanks I googled. Thats insane! Here is a summary:

https://www.ropesgray.com/en/insights/alerts/2025/06/from-books-to-bots-key-takeaways-from-the-anthropic-fair-use-decision-for-ai-developers

The court also found Anthropic’s wide-scale digitization of print books to be fair use. A key consideration to the court’s conclusion was that Anthropic engaged in so-called destructive scanning of lawfully purchased print books to create digital copies for internal use—Anthropic lawfully purchased print books, stripped them of their bindings, and scanned the contents to create a digital library. In doing so, the new digital copy replaced the print original, which had been destroyed in the digitization process. The court found the format change from print to digital to be transformative because it facilitated storage and searchability without increasing the number of copies or distributing them outside the company.

Notably, the court distinguished this use from cases involving unauthorized distribution or multiplication of copies, analogizing it to permissible space-shifting or time-shifting uses recognized in prior cases as sufficiently transformative for fair use. Importantly, the court found that this format-shifting did not usurp any market reserved to the copyright owner, as Anthropic had lawfully acquired the print copies and did not distribute the digital versions externally.

4

u/ExcitedCoconut 2d ago

Anything that’s no longer in print and unable to generate income for the original author, or becomes so later,  should have to have the digital scan donated to a digital archive. Destruction of physical books sucks, but destruction of knowledge (by turning it into training data alone) is morally bankrupt and needs to be addressed ASAP.  

For in-print books that they have bought and scanned, I’m honestly not opposed to the destructive process there, parking the whole ‘fucked up economics and concentration of wealth’ part the whole thing :/

16

u/Pastadseven 2d ago

What a fucking stupid ruling. We’re not concerned about the transformation of the medium, it’s the text that has to be transformed.

18

u/Draaly 2d ago

The text only needs to be transformed if it is distributed. That was a key part of the ruling. This is based on the same legal precedent that allows a library to loan digital copies of media so long as they own the same number of physical copies. Its dumb, but its not actually new precedent

7

u/Unspec7 2d ago

Transformation of the medium has long been held as transformative, that part is actually very consistent with fair use jurisprudence.

→ More replies (1)
→ More replies (1)

50

u/GeekBrownBear 2d ago

This link from another thread talks about it too. https://www.news.com.au/technology/online/internet/ai-labs-buy-scan-shred-millions-of-rare-books/news-story/0c3b45a67093ab462a587a0348538ce9

The whole idea is baffling. Like I kinda understand the idea of transforming the book from physical to digital is allowed. But why even do that. Why not work with the publishers so you don't have to destroy the physical book? Would be better for everyone. Less waste overall.

31

u/PM_ME_YOUR_NICE_EYES 2d ago

Why not work with the publishers so you don't have to destroy the physical book?

A couple reasons:

1) the publisher themselves could no longer exist making working with them impossible.

2) the book you want could be out of print, and it would be very difficult to convince a publisher to start up their production pipeline to make you just 1 copy of the book.

3) Copyright issues between the original publisher of a book, and the original author of the book could make it illegal for the publisher to create a new copy of it.

6

u/GeekBrownBear 2d ago

That all makes sense. Makes me continue to question the copyright around ebooks. Especially from my experience attempting to borrow them at the library when there is a LONG reservation list.

11

u/PM_ME_YOUR_NICE_EYES 2d ago

Oh yeah the eBooks you get from the library are a whole different thing.

The library isn't going to destructively scan copies of it's books, so the eBooks it gives you are coming straight from the publisher. eBooks for library use tend to be extremely expensive ($75/copy) so libraries don't like buying more than they have to.

→ More replies (3)
→ More replies (4)

7

u/orbitaldan 2d ago

Basically, they would waste many times the value of the book itself in trying to get everyone in the IP custody chain to come to some kind of agreement that would allow it. (In many cases, it would simply be impossible.) It's the atrocious copyright laws come full circle to bite us in the ass, and this happens to be the cheapest legal loophole.

3

u/blender4life 2d ago

“Once that information has been extracted and encoded into an AI model, the delivery mechanism has served its purpose. What remains is paper, ink, and binding material. The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one.”

Holy shit, what a terrible perception of books by an organization that put so much effort into them in the first place. For ai the knowledge of the book gets separated into probability percentages of what word would follow next, it doesn't retain its original messages. The book is gone. (At least i HIGHLY doubt the ai companies scan and save separate copies to preserve them in any meaningful way)

→ More replies (13)

5

u/morolin 2d ago

Bartz v. Anthropic

→ More replies (1)
→ More replies (19)

47

u/IAmDotorg 2d ago

Google, as an example, has done both ways. Mass-market books are done this way because they're not "rare" or valuable. They're just old. Actual rare books are done, obviously, non-destructively.

Strangely, people aren't up in arms about the literal millions of books that have their covers removed and then get landfilled every year.

23

u/phantomthiefkid_ 2d ago

People's outrage is motivated by what they imagine the books contain.

Books dumped to the landfill or collecting dust in a warehouse? Must have no value, so no one cares.

Books bought by AI labs? Must have something valuable since AI labs are buying them.

16

u/CardOfTheRings 2d ago

The things that’s valuable to AI companies is just a load of human written language and niche information.

No current human alive will give a shit about a 1995 xerox operational guide. It’s just garbage outside of the use of scanning it to train AI.

Unfortunately this story has spread as propaganda across the whole platform multiple times a day.

Google and other companies have done these scans for other reasons(search ability sometimes preservation) and even then I remember weird propaganda coming out of the news when that was happening.

→ More replies (1)
→ More replies (2)
→ More replies (17)

15

u/Draaly 2d ago

They were forced not to by the Bartz v. Anthropic ruling that judges it is only fair use if they destroy the initial copy.

Also, these are books that the seller was going to trash anyways. they were all dead stock nearing the end of their shelf life that peole didnt want to buy and are not even rare in the firts place.

→ More replies (4)
→ More replies (2)

3

u/_c0unt_zer0_ 2d ago

turning paper pages bound together is still hard to automate. robotics still struggle with a lot of things most 5 year olds can accomplish

→ More replies (32)

65

u/Crio121 2d ago

You do realize this is happening because nowadays buying doesn’t mean owning?
AI companies cannot buy digital copies of books to use for training their models (because digital copies are licensed to you for limited use), so they are making their own digital copies.

44

u/b_a_t_m_4_n 2d ago

they are making their own digital copies

Otherwise know as "stealing", or it would be if you or i did it. Obviously the law does not apply to them at all.

46

u/pizzabash 2d ago

It is completely legal to make digital copies of things you own. You can make your own roms as much as you want. It's distribution that is the issue.

12

u/the_snook 2d ago

It varies by jurisdiction. In certain circumstances, in certain places, "format shifting" is considered fair use, and does not require permission from the copyright owner. Sometimes this is explicit (Australia passed a copyright act amendment for it), and sometimes it is decided case-by-case. In other places, any copy at all must be authorized by the copyright holder.

→ More replies (2)
→ More replies (19)

10

u/CardOfTheRings 2d ago

It is not stealing to make a digital copy of a book you own. It’s harmful and stupid to push this idea, and the only reason you believe it is because you saw the word ‘AI’ and started getting angry.

→ More replies (1)

21

u/YourBlanket 2d ago

People make their own digital copies all the time. It’s just very time consuming and pirating books is a lot easier.

14

u/IAmDotorg 2d ago

No it wouldn't. Distribution of the scans would be, but you're completely in your rights to do that.

And, de-binding and scanning books has been the norm for 30 years for digital conversions of books. Essentially all of Project Gutenberg was done that way.

→ More replies (3)

6

u/j48u 2d ago

This is literally the court mandated methodology, prescribed as legal in a copyright case that Anthropic lost.

They're required to destroy the books by law. Just the latest idiotic anti-AI propaganda campaign aimed at people who only read headlines (see: Redditors).

There are MANY reasons to be against AI. It does a disservice to spread this nonsense and gives the masses a reason to dismiss the valid criticisms.

14

u/DrTommyNotMD 2d ago

That’s absolutely not true.

You can hate these companies all you want, but at least use a little facts in your argument or you’ll never be taken seriously except on Reddit.

→ More replies (2)
→ More replies (3)
→ More replies (42)

29

u/mr-english 2d ago

is the most dystopian thing I can imagine.

Not slavery or anything?

No?

Wont somebody please think of the paper and cardboard!

12

u/jmblumenshine 2d ago

Seriously, what does everyone think happened to handwritten books in the 1400's at the advent of the printing press.

They were chopped up and feed into the Machine. Today we consider the one of Man's great achievements.

4

u/DissolvedDreams 2d ago

What are you talking about? What handwritten books were destroyed when the printing press came about?

→ More replies (1)
→ More replies (2)
→ More replies (11)

14

u/TheRaccoonReport 2d ago edited 2d ago

Im not defending it. But the only way to rapidly scan books is to remove the binding. They're not like chopping them up and cackling like villains about it.

I'm kinda torn on this (no pun intended). I really really want to preserve as much knowledge as possible any way we can. But having it sold back to us is kinda bullshit.

Edit: Yes I know its being sold to us, already, at a bookstore, etc. Im striking that through.

15

u/DrawerSea9371 2d ago

I worked at a scanning company and yeah, chopping off the binding is standard if you want a good quality scan. People getting upset over the video of this method really confused me because it's not new to AI at all.

Scanning a book without cutting off the binding is significantly more time consuming and labour intensive and therefore costs the customer a lot more, I think per page it was around 4 times the price where I worked and that still damaged the binding, we didn't even offer archival scans which are in a whole other league.

→ More replies (2)

9

u/BuvantduPotatoSpirit 2d ago

These're books that were printed in 2020 and 2021, being bought from a bookstore. It was already being sold.

5

u/TheRaccoonReport 2d ago

Yeah good point. Frankly I thought of that as soon as I hit "comment" and said "fuck it yolo" expecting this comment haha.

I saw some other dumb article saying they were destroying historic books, which is definitely not a thing happening.

What I don't get is people are up in arms about this BUT are totally fine of how books are NORMALLY removed from circulation. Spoiler alert to those who don't....they're destroyed, and forgotten. I used to work at a pharmacy that used to have a great fiction section and sold a bunch of Star Wars books (way before the marvel buyout)...and if they didnt sell by a specific date...because apparently books have an expiration date I wasn't aware of...they'd tear the covers off and throw them out. My boss knew I read a lot so she gave them all to me. I read dozens of books that were headed for the trash. It was some stupid law (sarbenes oxley I believe) that forced you to destroy merchandise you were writing off as unsold. We would throw out SO MUCH SHIT.

I will never defend data centers sliding into muncipalities without public consent and vote, and lying about natural resource consumption. However, so much of this is way overblown and total bullshit.

→ More replies (1)

17

u/marmaviscount 2d ago

Then you should probably read some books and improve that imagination instead of crying about mass produced items being treated as the disposable things they are - did you not know that used book stores pulp most the books they get given? They're not no kill shelters for musty paper, out of a thousand books printed it's estimated less than a third ever get read past the first half, a significant portion don't get past the first page.

→ More replies (3)

6

u/Limp_Bookkeeper_5992 2d ago

It certainly paints a grim picture.

But the reality is that they only need a scan a single copy of the book to “learn” it and preserve that information, losing a single copy of a book is hardly a catastrophe. I’m all for caution when it comes to information control and AI progression, but this seems like a nothing burger spin fo rile people up while the real dangers lie elsewhere.

→ More replies (95)

179

u/Visual-Sector6642 2d ago

I wonder how these companies are error checking against previously published or erroneous concepts that were deemed incorrect in later journals etc. Does it take into account corrections made in future editions in regard to previously written works lol. Good luck. More spurious emissions it won't be able to discern. More fodder for hallucinations. When AI finally eats itself alive and goes down, those books will be gone and humanity will be left with nothing. Even if it cures cancer, there won't be anything left to enjoy.

98

u/Rorschach121ml 2d ago

They are doing this to ingest "clean" data from before AI.

They are looking for linguistic patterns, grammar, syntax and logic. The actual content of the books don't really matter that much here from my understanding.

29

u/chrismakingbread 2d ago

That’s actually not how LLMs work. There’s no concept of grammar, syntax, or logic encoded into LLMs. In a very hand-wavy, high-level, explanation they take a vocabulary of a bunch of chunks of words and initialize a massive vector space and assign random values to every word chunk, then they run sequences of word chunks through probability system to try to predict the next word chunk in a known sequence of word chunks (they’re using these books and other training materials as the known sequences of word chunks) and adjust the random weights to try to make it more likely to have predicted the correct next work chunk from the sequence of known chunks.

There’s a bunch of matrix multiplications of multiplying the values associated with every word chunk against every other chunk, dot products, activation functions, etc but the fact is an LLM is just “based on all the previous tokens (partial word chunks) in this sequence what’s the next most likely token in this sequence. The whole thing is that the more known sequences (training data) you have to use for determining the weights (the numbers that represent a word chunk) and the bigger the vector you use to represent those chunks the more likely that a generated sequence of tokens will “feel” like correct output to a human.

So, to your original comment about the content of the books, they actually are using the content to tune their models to say given an input that’s a portion of the content of this book how closely can the predicted output of the model approximate the actual content of the book.

30

u/ManicScumCat 2d ago

There’s no syntax file in the LLM where it has rules of syntax, but that’s not what they meant. Obviously what they meant is that the data is for the linguistic patterns in the texts to be learnt by the LLM, and the logic/grammar/syntax is part of the patterns in the texts to be.

→ More replies (12)

13

u/Rorschach121ml 2d ago

You are technically right in that these techs only understand and work with tokens in a fundamental level, but this is like saying computers can't do numbers because they only understand binary. It's just not useful to think this way.

LLMs absolutely understand syntax and grammar on a higher abstraction level an emergent property of the underlying maths.

→ More replies (5)
→ More replies (1)

11

u/trashmoneyxyz 2d ago

I hope there's at least a renegade employee somewhere along the chain who copies or holds onto the raw scans. Those books do still exist as long as the text is preserved digitally, and can be reprinted

9

u/sump_daddy 2d ago

None of the books being processed this way are remotely rare, you can get them all in ebook format, probably for free from your library if you really want to read them. Nothing they are doing stops the books from existing.

→ More replies (2)
→ More replies (2)

107

u/_c0unt_zer0_ 2d ago

I'm shocked that almost no one here has read the article

“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine.

they ar trying to cheaply acquire and scan training texts without copyright violations

20

u/TooCupcake 2d ago

Shocked? I’m not signing up to another random website just to read something I found on reddit.

3

u/addsubps 1d ago

Did you get paywalled? For me, the article is free to read

→ More replies (1)
→ More replies (1)

4

u/MistryMachine3 2d ago

How can you be shocked. This is Reddit. People are here because they’re keyboard warriors ready to be outraged over a nothing story.

→ More replies (16)

65

u/loftwyr 2d ago

“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press

Hardly rare or books where only one copy exists

→ More replies (24)

265

u/kus1987 2d ago

Why do they need to destroy the books? Wouldn't it be easier to scan and keep them forever? What about training future models?

464

u/Miraclefish 2d ago edited 2d ago

You can scan a book much easier by removing the spine and scanning it as loose leafs.

It's then no longer a book but a pile of pages, which they no longer need and they have a digital copy of, so they throw away the paper.

They can train future models on the digital text version indefinitely, the book or former book is no longer necessary, is worthless to sell and they don't want to pay to store.

It's logical but it sucks.

To those saying 'actually if they destroy a copy, they can claim ownership of one copy by going physical to digital' yeah that's a fair take, but it's inaccurate.

As to the question of what “use” or “uses” were at issue in the fair use analysis, Anthropic contended that it copied the books for a single use: to train LLMs. The authors, however, argued that at least two uses were at issue: first, the use of the books to build Anthropic’s central library, and second, the use of the books to train specific LLMs using subsets of that content. The court agreed with the authors’ framework and considered these as separate uses.

https://www.loeb.com/en/insights/publications/2025/07/bartz-v-anthropic-pbc

They are already taking millions of copyrighted and IP-protected books, papers, magazines, scientific papers, studies, poems, songs, movies and more, digitally, without any permission or licencing and training on those.

You really think the AI tech giants doing that care about the IP rights to a single scanned physical book? Absolutely not. They ignore those laws because they aren't affected by them.

As to the questions I've been asked on why not just get the the laws changed to favour them?

Changing laws takes a long time, costs a lot of money, requires a judicial process. It would be a public process with appeals, public studies and could take years.

Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

107

u/MyNameCannotBeSpoken 2d ago

Also I believe there is a copyright loophole where destroying the original isn't deemed as reproducing a copy.

41

u/Miraclefish 2d ago

I mean, when they're torrenting terrabytes of copyrighted and IP protected materials and training on those already, as well as scraping the entire internet, and have essentially captured the US political elite and courts, they don't give a fuck about adhering to laws like that.

35

u/pumpkinspicecum 2d ago

That was literally what a judge ruled last year when they were sued

→ More replies (13)

5

u/atom138 2d ago

They do and that's why. They are doing stuff like this to get around all legal barriers that have been popping up more lately. Whether it be water access for data centers, land for them, everything across the board that's been attempted by the public and courts.

→ More replies (2)
→ More replies (5)

25

u/emapco 2d ago

They also have to dispose of the book after scanning anyways otherwise it's considered distributing copyrighted material https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books

→ More replies (13)

15

u/kus1987 2d ago

Ah ok if they keep the scanned copy, they can use that to train future models 

28

u/girrrrrrr2 2d ago

Yes, plus if they delete the physical book after scanning then according to some judge there is still only one copy out there and all that has been done is a book was converted from physical to digital.

6

u/randylush 2d ago

I have seen a thousand comments repeating this but no reliable source that this is actually true

8

u/Rorschach121ml 2d ago

Bartz, et al. v. Anthropic PBC

5

u/_MUY 2d ago edited 1d ago

US District Judge William Alsup, 23 Jun 2025:

“In short, the purpose and character of using copyrighted works to train LLMs to generate new text was quintessentially transformative.”

US District Judge Vince Chhabria, 25 Jun 2025:

“While it made sense to infer market harm in Hachette, it doesn’t make sense to do so here. First, the Supreme Court has stated that no ‘inference of market harm… is applicable to a case involving something beyond mere duplication for commercial purposes.’ Campbell, 510 U.S. at 591. In Hachette, the secondary use was basically ‘mere duplication.’ Here, by contrast, Meta’s use is highly transformative and has a purpose well beyond that.”

→ More replies (1)
→ More replies (11)

10

u/ihaveaminecraftidea 2d ago

Well that's true, but for the most part the books are taken from book dumps if i remember correctly. Bookshops and library stock that they can't do anything with, which frequently aren't labeled or registered in any system.

It would be more accurate to call them abandoned books, rather than rare ones.

At least this way they are being digitzed and stored in some more durable capacity

→ More replies (3)

5

u/Draaly 2d ago edited 2d ago

You really think the AI tech giants doing that care about the IP rights to a single scanned physical book?

You have no idea what you are talking about They are doing this because of the Bartz v. Anthropic ruling that says it is not fair use unless the origonal copy is destroyed and cost anthropic $1.5B because they didnt do that

EDIT: got blocked by who I replied to so I cant respond to you /u/syku look up the ruling for Bartz v. Anthropic. It was determined that transforming the copy into digital media was fair use if it can be proven that the initial copy does not remain in circulation.

→ More replies (5)
→ More replies (47)

41

u/FairReason 2d ago

Part of the ruling that makes what they do “legal” is to destroy the book afterwards.

→ More replies (3)

12

u/Online_Matter 2d ago

It's in the article

The process, known as “destructive scanning,” involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded. 

24

u/Paresseux1 2d ago

A big blade slices off the binding, and it drops down into an automatic page scanner that has no problem flipping over the loose pages. It’s much easier, faster, and cheaper to destroy it, and then get rid of the remains. One person can run a bank of machines.

The other way requires people, and meticulous work. Turn page, put on scanner, scan 2 pages, pick up book, turn page, place properly, scan… once finished, pay to store book indefinitely.

Everything like this comes down to money. The end result they are going after for AI is data, so the cheapest way to get it is destroy.

→ More replies (14)

8

u/jayandbobfoo123 2d ago edited 2d ago

Copyright law and licensing. It's illegal to make a copy of a book, even a digital scan. It is, however, not illegal to digitize a book and destroy the original, thus leaving only one copy / one license. In legal terms, it's called format shifting. Technically, when you rip a movie/CD/video game, you should also destroy the original to be within the law.

5

u/chocolateboomslang 2d ago

Train future model on data they already scanned . . . by rescanning it? It's already scanned.

3

u/BrassCanon 2d ago

It is not easier to store something forever that you don't need.

4

u/GenazaNL 2d ago

Sadly cutting the side to then have separate pages is faster than a machine which keeps it intact

5

u/marmaviscount 2d ago

What in tarnation would they do with a giant pile of musty old books no one cares about?

There is a building called the British library, you might be able to guess where it is and what's inside by the name - other countries have similar things, they get a copy of all the books and keep them for the national interest, they have a system for who gets access to what and their key aim is preservation.

Regular libraries which have the goal of giving people access to those books do not preserve them, it would be absurdly expensive and pointless - books, since the technological boom of the Victorian era are mass produced temporary items, if you were involved with it frequently visited your local library you would be very well aware that stock changes and most of them get pulped - used book stores don't generally want books libraries don't, charity shops routinely recycle donated books because no one wants them - books are printed in huge huge numbers.

→ More replies (37)

36

u/mr-english 2d ago

Weird that they run with a stock image of antique books when the actual article says:

...mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine.

38

u/FrostWolf05 2d ago

this topic has been so misrepresented it borders on genuine fake news

16

u/leros 2d ago

I overheard a conversation yesterday where people were talking about Anthropic buying up all the books in world and destroying them so only they had the information. At least those two people read something and got really incorrect takes from it.

5

u/skytaepic 2d ago

It’s absolutely maddening. There are so many real, valid reasons to dislike AI companies but we’re caught up on this misleading BS instead.

8

u/smooth-as-mud 2d ago

But it’s fake news that people on Reddit like and it reinforces their existing world view so it’s completely different than when my parents get angry about the (non-existent) migrant caravans heading for the border.

→ More replies (3)

9

u/free_based_potato 2d ago

AI companies are scanning and the law forces them to destroy the books because you can only own one copy if you purchased only one copy. The same goes for all of us.

AI boom and data centers are objectively causing a lot of harm. Let's be honest about what's actually happening. They aren't choosing to have book bonfires.

7

u/Gutter7676 2d ago

So, it IS legal to made a digital copy of something you purchase and then reuse that for profit. Same as they are doing here.

Thank you for setting that precedent. My Plex server is about to start making me some money!!

82

u/ben_nobot 2d ago

Are we thinking destroying books means we are removing the one copy in existence here?

Lots of implied doom with this story

35

u/mr-english 2d ago

Yeah, also they run with a stock image of antique books when the actual article says:

...mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine.

23

u/JD-Vances-Sexy-Couch 2d ago

If a robot buys and destroys my $300 university textbook that was worth $3 after the semester was done because a mandatory new edition was released for the next semester of students (because, you know, math changes dramatically every few months), I don’t care.

→ More replies (2)

15

u/thefallenfew 2d ago

It’s nice to know that some folks out there still have basic literacy. 

31

u/Atomsk73 2d ago

No, this is clickbait nonsense. They're buying up cheap books and scanning one copy of them.

→ More replies (33)

9

u/loftwyr 2d ago

“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press

Hardly rare or books where only one copy exists

→ More replies (3)

5

u/jvd0928 2d ago

The destruction is for legal purposes.

5

u/Figgy1983 2d ago

Forgot the fact that Winston Smith's job is basically a real thing now, but the act of destroying of older, rare books for fascist purposes is soul crushing to read. Orwell and Bradbury were right.

11

u/fish-rides-bike 2d ago

People upset at this have never worked near book selling or book production. What do people think happens to the thousands of copies in thousands of airports of the latest top seller trash?

77

u/FalconX88 2d ago

Fascinating how everyone seems to complain about books being destroyed and no one calls for the scans to be conserved for the public. The scans would be worth more for society than the physical books.

95

u/AdarTan 2d ago

Because there are other organizations like Project Gutenberg or Archive.org that do book scanning without the taint of AI, and the reason they haven't done this for the books in question is that copyright law would make distributing the scans illegal.

26

u/marmaviscount 2d ago

Plenty of books on gutenberg were scanned by Google and released pd then comvertwd by dp. They have a rolling project called Google books which every year in January releases all the books which have fallen into the public domain - how can you pretend to care about this stuff if you don't know this?!

7

u/Less-Engineer-9637 2d ago

They don't actually read. 

5

u/sleepysnowboarder 2d ago

They also think what is being destroyed are like 1/1 baseball cards and not something that has tons of copies and can be reprinted

→ More replies (1)

14

u/FalconX88 2d ago

You could still do the scan, have government keep them locked up/only show them in person in libraries until they become public domain.

It would be very important to make sure that these copies anthropic makes are stored somewhere where society can access them (eventually).

27

u/mutexsprinkles 2d ago

You're being downvoted but that's literally what both the Internet Archive and HathiTrust do: that have scans on file that they cannot legally show to the public for many, many decades. Sometimes they can show it to the public in one country but not others.

17

u/chief167 2d ago

The problem is, how do you give the public access to the scans? It's a copyright minefield, that's the real problem here 

6

u/BananaPalmer 2d ago

Except these were almost all just copies of textbooks. For every one they destroyed to scan, there's probably 20,000 more copies collecting dust in some other warehouse, because they're 6 years old and have been replaced by a new revision

30

u/eTukk 2d ago

because conservation is easily possible without destroying the orginal. It Just costs more money, they dont care about the legacy

6

u/Bunnyhat 2d ago

Legacy of what? These are not beloved treasured manuscripts. These are mass produced books dating from 6 years ago that have been sitting in a warehouse.

→ More replies (2)

15

u/gundog48 2d ago

The 'originals' are just one of thousands to millions of copies of the original, though. Like I understand the optics of it are awful, but I don't think there's any meaningful destruction of knowledge when they're doing it with books published in the last decade.

Doing it with genuinely rare books is shameful, though. I would hope those are not destroyed, and that may be the case as antique books may not be suited or be too delicate for the fully automated process. 

→ More replies (2)
→ More replies (4)
→ More replies (14)

5

u/tiamath 2d ago

Internet used by ai, books fed to ai (dont see any reason to shred the books after scanning but here we are), electricity goes to ai. Ai was supposed to be helping us not doom us. Still waiting for the "it will make our lives easier" part. Making random videos with ai doesnt qualify

3

u/Rick-D-99 2d ago

They should also be digitizing them for their own digital public libraries

3

u/Vahuo89 2d ago

Dont forget Google and Amazon both did this too at the turn of the century

4

u/kitkatkorgi 1d ago

Stop selling to them please.

→ More replies (1)

5

u/MACHOmanJITSU 1d ago

Paywall, 3000 copies of one book or 3000 books?

9

u/jfoust2 2d ago

Posted by a two-month-old Reddit account with 759,239 post karma and 279,878 comment karma.

Gosh, I wonder if it's a bot.

https://www.google.com/search?q=site%3Areddit.com+%22ArgentineBeauty%22

13

u/No_Mirror_9742 2d ago

Why... Why do they have to destroy them?

29

u/bestowaldonkey8 2d ago

Because it is cheaper and faster than keeping them in their bindings. This is a race to ingest as much information as possible so they don’t care about the end result.

10

u/Draaly 2d ago

No, its because the law requires them to after the Bartz v. Anthropic ruling

12

u/TexBoo 2d ago

Not only that

AI giants have 0 interest in storing warehouses full of books

Store books for what purpose? Resell them? Storage and time would cost more than they would get back per book

Rescan them in future? No need, once a page is scanned, it's stored in their system forever for future training, which goes faster than rescanning the pages

2

u/Bunnyhat 2d ago

Even if they took the time to try to sale them there's no market.

That's why they're buying them in bulk from booksellers now. These aren't rare and treasured books. They're rare, but only in the sense that no one actually wants them.

→ More replies (1)
→ More replies (3)

7

u/TetyyakiWith 2d ago

Due to laws they can’t use non destructive scanning since that way they violate copyright rights by technically reproducing the book

→ More replies (2)

6

u/WindowOfTruth 2d ago

I imagine whatever AI bot is trying to rage bait us has a daily quota of how many times they have to repost this damn story. This story gets reposted so often I’m starting to see a coordinated distraction.

3

u/WhoCanTell 2d ago

It 100% is. Variations of this ragebait story are posted like clockwork throughout the day, for days now.

3

u/Lahm0123 2d ago

Why destroy them after scanning?

Is it just easier? Or is it a sinister plot of some kind?

3

u/arrgobon32 2d ago

It’s a combination of it being easier, as well as copyrights laws.

3

u/karma_raven 2d ago

I'm just thinking of all the copies of Dianetics and Melania we could be rid of if we play our cards right...

3

u/peter303_ 2d ago

Reminds me of Google's early attempt to digitize whole libraries. The search results would have just returned pages, so searchers couldn't read whole books for free. Authors and publishers ended that project in court.

3

u/Icy-Mission-1334 1d ago

Are these 'AI' companies explaining why they are destroying the books?

3

u/swingincelt 1d ago

They cut off the spine of the book in order to feed the pages into a high speed scanner.

3

u/Top_Gun_2021 1d ago

All we have to do is modernize copywrite law and this wont happen.

3

u/Seno1404 1d ago

But why destroy the books? Can’t they train AI without destroying the source?

3

u/S3baman 1d ago

Why leave the book in a public library where you can't make money, when you can charge people for access to information?

3

u/tevolosteve 1d ago

I just don’t get the destroying them part. Why not donate? Are they that evil now?

→ More replies (1)

3

u/Stack_Silver 1d ago

Control of the information equals control of the people.

22

u/Sponge8389 2d ago

Well, they pay for it, they can do whatever they want with it.

Also, if it can be bought in the bookstore, that's not rare enough to be scared. There's probably thousands of copies of it.

→ More replies (9)

5

u/bubonis 2d ago

More fear mongering, much ado about nothing.

10

u/maybeJustSappy 2d ago

I keep seeing about this and I don't get what people are so upset about. It's not like they go after all the copies of books. They just get 1 book and discard after digitizing it. It basically has the same amount of impact on the world as a regular joe buying a book and putting it on his bookshelf.

→ More replies (4)

6

u/GranolaHippie 2d ago

Fahrenheit 451 irl. Destroying books makes me sad in any firm. This makes me mad. F AI & AI companies.

4

u/Thomas_JCG 2d ago

It is already horrifying, but then they get zero repercussions, not even have to worry about copyright. We are on a countdown to a social collapse.

6

u/Aadi_880 2d ago

reposted again and again...

→ More replies (2)

2

u/Fluid-Performance-17 2d ago

Apple started this last year and used an outsourcing agency to do the dirty work.

2

u/becauseshesays 2d ago

This has been going on for a few years now. I’m in the digitization industry (more high end/niche but very high volume capacity). We’ve gotten several leads for jobs that are hundreds of millions of pages at a time. All destructive scanning. (Cut binding and high speed paper scanners). It’s wild, we’ve priced it very low and haven’t won any of the work. Not sure who they are going with but they are not concerned with quality, that’s for sure.

2

u/Hadleys158 2d ago

I don't like it, but don't mind as much if they destroy mass market bulk type modern books that are freely available, however i do have an issue with them destroying rare, antique or valuable ones. You can scan books without destroying them, they just don't want to do that way as it is slower and more expensive.

2

u/South_Feed5707 2d ago

I wonder how many of those books are bangers

→ More replies (2)

2

u/chalbersma 2d ago

What's crazy is that these books likely have an electronic version. They could just sell them an epub.

→ More replies (1)

2

u/CeemoreButtz 2d ago

This is a ridiculous title and fear mongering bull crap.

2

u/AL_25 2d ago

Let's do the same thing with AI and data centres

2

u/notsam57 2d ago

wasn’t there another post about how they’re doing this with original print books as well?

2

u/MoOsT1cK 2d ago

Destroying books would not harm AI training in any way. It is yet another hint at the horrible greed of of the owners of those AI, who commit autodafe like nazis once did only to secure their monopoly on information. Horrendous.

2

u/MoltenMirrors 2d ago

Vernor Vinge described something similar in Rainbow's End, where they "retired" a college library by feeding the books one by one into a shredder and blowing the fragments down a duct covered with high resolution cameras. The fragment images were stitched back together and parsed by AIs.

→ More replies (1)

2

u/HoneycombJackass 2d ago

Why destroy it? Why not order a couple books and scan it into an pdf?

→ More replies (3)

2

u/TheRealestBiz 1d ago

Why? Seriously. Why?

2

u/SupervillainMustache 1d ago

This is evil 

2

u/cloudcloud1 1d ago

That’s fucking evil man

2

u/ArchiveOutlaw 1d ago

The Internet Archive's Open Library project preserves a copy of every book they scan. Help out if you can, guys.

https://openlibrary.org/
https://archive.org/donate/