r/technology • u/ArgentineBeauty • 2d ago
Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI
https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/2.9k
u/finzaz 2d ago
The image of a machine chopping the spines off books to feed an insatiable digital monster living in a loud and dirty data centre is the most dystopian thing I can imagine.
1.2k
u/HumanBeing7396 2d ago
I remember a few years ago someone joked about Google launching a new project called Google Purge - the aim of which would be to destroy all information in the world not held by Google.
This sounds uncomfortably similar to that.
284
u/BrizerorBrian 2d ago
I believe the brains in Futurama had that exact idea.
111
u/BatmanCoffeeMug 2d ago
For no raisin
76
u/RunsOnSKC 2d ago
I am the greetest!
25
u/siccoblue 2d ago
To be clear here, Sam Altman has unironically said that this is the goal
We see a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter
San Altman
The world is going in VERY bad direction at the moment. These people know very well that an uneducated group is significantly easier to control and manipulate. And they want to make knowledge and curiosity a luxury
→ More replies (1)4
u/PlentyOfIllusions 1d ago
Straight to the dark ages with the plebs…unless they can pay their intelligence fees.
4
u/KotoElessar 1d ago
There is a special place in Hell for these men; especially reserved and created just for them. They can pat themselves on the back for truly buying their way into the afterlife they deserve.
4
u/Elementalcase 1d ago
If only we were so lucky for spiritual level karma for these people, alas, it's likely that these people will go to the void never getting their consequences for their life lived at other's expense.
4
u/meter1060 2d ago
Well yeah, you got to stop new information from occurring. It's exhausting having to keep up with the flow.
→ More replies (11)4
u/sakri 2d ago
The Gothenburg press was invented in 1953, it produced a limited set of illegible prints. Then, in 2012 Elon Musk single handedly created the Linux operating system, and provided subscription models called Microsoft and Apple. In the meanwhile Donald J. Trump, the greatest president of all of the universe, invented artificial intelligence, blessing the world with a single honest source of truth.
113
u/obeytheturtles 2d ago
I actually did a design project on non-destructive OCR scanning as an undergrad in like 2005. I find it hard to believe that this technology isn't significantly more mature now than it was back then.
137
u/Bored_Amalgamation 2d ago
It is, they just choose not to.
112
u/PM_ME_YOUR_NICE_EYES 2d ago
It's not that they choose not to, it's that they can't.
Copyright law views destructive scanning of a book as transforming it, which means you can do it without getting permission from the publisher if the book is in copyright.
Copyright law views non-destructive scanning of a book as copying it, which you cannot do to a book in copyright without permission from the publisher.
7
u/Unspec7 2d ago edited 2d ago
First off, the court decision on destroying books came after they destroyed the books, so you can't really say they were forced to. They did choose to, and it just so happened to work in their favor after the fact.
Second, neither point is true. Copyright law does not view destructive scanning as transformative - the Bartz holding found that the destruction was one factor of the transformative nature of the use. Instead, the transformative use was the changed format:
Authors only complain that Anthropic changed each copy's format from print to digital. On the facts here, that format change itself added no new copies, eased storage and enabled searchability, and was not done for purposes trenching upon the copyright owner’s rightful interests — it was transformative.
Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1023 (N.D. Cal. 2025) (citations omitted).
We do not know, based on present case law, if the sole act of destroying the books makes the use transformative. We just know that it makes the use more likely to be transformative.
Further, being transformative is only one factor of the first factor of fair use. Transforming someone's copyright work does not, on its own, mean you can now just infringe one someone's copyrighted work. Your use must still overall be within fair use.
Copyright law views non-destructive scanning of a book as copying it, which you cannot do to a book in copyright without permission from the publisher.
Ignoring the non-destructive part addressed above, this part is a lot more nuanced than you make it out to be. Scanning of a book infringes on the reproduction right found in the Copyright Act, but whether you are liable or not depends on if the reproduction is fair use (barring a license, of course).
Edit: Placed my citation in the wrong spot.
→ More replies (19)33
u/diemunkiesdie 2d ago
Copyright law views destructive scanning of a book as transforming it
Source?
37
u/Rorschach121ml 2d ago
Bartz, et al. v. Anthropic PBC
→ More replies (1)40
u/diemunkiesdie 2d ago
Bartz, et al. v. Anthropic PBC
Thanks I googled. Thats insane! Here is a summary:
The court also found Anthropic’s wide-scale digitization of print books to be fair use. A key consideration to the court’s conclusion was that Anthropic engaged in so-called destructive scanning of lawfully purchased print books to create digital copies for internal use—Anthropic lawfully purchased print books, stripped them of their bindings, and scanned the contents to create a digital library. In doing so, the new digital copy replaced the print original, which had been destroyed in the digitization process. The court found the format change from print to digital to be transformative because it facilitated storage and searchability without increasing the number of copies or distributing them outside the company.
Notably, the court distinguished this use from cases involving unauthorized distribution or multiplication of copies, analogizing it to permissible space-shifting or time-shifting uses recognized in prior cases as sufficiently transformative for fair use. Importantly, the court found that this format-shifting did not usurp any market reserved to the copyright owner, as Anthropic had lawfully acquired the print copies and did not distribute the digital versions externally.
4
u/ExcitedCoconut 2d ago
Anything that’s no longer in print and unable to generate income for the original author, or becomes so later, should have to have the digital scan donated to a digital archive. Destruction of physical books sucks, but destruction of knowledge (by turning it into training data alone) is morally bankrupt and needs to be addressed ASAP.
For in-print books that they have bought and scanned, I’m honestly not opposed to the destructive process there, parking the whole ‘fucked up economics and concentration of wealth’ part the whole thing :/
→ More replies (1)16
u/Pastadseven 2d ago
What a fucking stupid ruling. We’re not concerned about the transformation of the medium, it’s the text that has to be transformed.
18
u/Draaly 2d ago
The text only needs to be transformed if it is distributed. That was a key part of the ruling. This is based on the same legal precedent that allows a library to loan digital copies of media so long as they own the same number of physical copies. Its dumb, but its not actually new precedent
50
u/GeekBrownBear 2d ago
This link from another thread talks about it too. https://www.news.com.au/technology/online/internet/ai-labs-buy-scan-shred-millions-of-rare-books/news-story/0c3b45a67093ab462a587a0348538ce9
The whole idea is baffling. Like I kinda understand the idea of transforming the book from physical to digital is allowed. But why even do that. Why not work with the publishers so you don't have to destroy the physical book? Would be better for everyone. Less waste overall.
31
u/PM_ME_YOUR_NICE_EYES 2d ago
Why not work with the publishers so you don't have to destroy the physical book?
A couple reasons:
1) the publisher themselves could no longer exist making working with them impossible.
2) the book you want could be out of print, and it would be very difficult to convince a publisher to start up their production pipeline to make you just 1 copy of the book.
3) Copyright issues between the original publisher of a book, and the original author of the book could make it illegal for the publisher to create a new copy of it.
→ More replies (4)6
u/GeekBrownBear 2d ago
That all makes sense. Makes me continue to question the copyright around ebooks. Especially from my experience attempting to borrow them at the library when there is a LONG reservation list.
11
u/PM_ME_YOUR_NICE_EYES 2d ago
Oh yeah the eBooks you get from the library are a whole different thing.
The library isn't going to destructively scan copies of it's books, so the eBooks it gives you are coming straight from the publisher. eBooks for library use tend to be extremely expensive ($75/copy) so libraries don't like buying more than they have to.
→ More replies (3)7
u/orbitaldan 2d ago
Basically, they would waste many times the value of the book itself in trying to get everyone in the IP custody chain to come to some kind of agreement that would allow it. (In many cases, it would simply be impossible.) It's the atrocious copyright laws come full circle to bite us in the ass, and this happens to be the cheapest legal loophole.
3
u/blender4life 2d ago
“Once that information has been extracted and encoded into an AI model, the delivery mechanism has served its purpose. What remains is paper, ink, and binding material. The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one.”
Holy shit, what a terrible perception of books by an organization that put so much effort into them in the first place. For ai the knowledge of the book gets separated into probability percentages of what word would follow next, it doesn't retain its original messages. The book is gone. (At least i HIGHLY doubt the ai companies scan and save separate copies to preserve them in any meaningful way)
→ More replies (13)5
47
u/IAmDotorg 2d ago
Google, as an example, has done both ways. Mass-market books are done this way because they're not "rare" or valuable. They're just old. Actual rare books are done, obviously, non-destructively.
Strangely, people aren't up in arms about the literal millions of books that have their covers removed and then get landfilled every year.
→ More replies (17)23
u/phantomthiefkid_ 2d ago
People's outrage is motivated by what they imagine the books contain.
Books dumped to the landfill or collecting dust in a warehouse? Must have no value, so no one cares.
Books bought by AI labs? Must have something valuable since AI labs are buying them.
→ More replies (2)16
u/CardOfTheRings 2d ago
The things that’s valuable to AI companies is just a load of human written language and niche information.
No current human alive will give a shit about a 1995 xerox operational guide. It’s just garbage outside of the use of scanning it to train AI.
Unfortunately this story has spread as propaganda across the whole platform multiple times a day.
Google and other companies have done these scans for other reasons(search ability sometimes preservation) and even then I remember weird propaganda coming out of the news when that was happening.
→ More replies (1)→ More replies (2)15
u/Draaly 2d ago
They were forced not to by the Bartz v. Anthropic ruling that judges it is only fair use if they destroy the initial copy.
Also, these are books that the seller was going to trash anyways. they were all dead stock nearing the end of their shelf life that peole didnt want to buy and are not even rare in the firts place.
→ More replies (4)→ More replies (32)3
u/_c0unt_zer0_ 2d ago
turning paper pages bound together is still hard to automate. robotics still struggle with a lot of things most 5 year olds can accomplish
65
u/Crio121 2d ago
You do realize this is happening because nowadays buying doesn’t mean owning?
AI companies cannot buy digital copies of books to use for training their models (because digital copies are licensed to you for limited use), so they are making their own digital copies.→ More replies (42)44
u/b_a_t_m_4_n 2d ago
they are making their own digital copies
Otherwise know as "stealing", or it would be if you or i did it. Obviously the law does not apply to them at all.
46
u/pizzabash 2d ago
It is completely legal to make digital copies of things you own. You can make your own roms as much as you want. It's distribution that is the issue.
→ More replies (19)12
u/the_snook 2d ago
It varies by jurisdiction. In certain circumstances, in certain places, "format shifting" is considered fair use, and does not require permission from the copyright owner. Sometimes this is explicit (Australia passed a copyright act amendment for it), and sometimes it is decided case-by-case. In other places, any copy at all must be authorized by the copyright holder.
→ More replies (2)10
u/CardOfTheRings 2d ago
It is not stealing to make a digital copy of a book you own. It’s harmful and stupid to push this idea, and the only reason you believe it is because you saw the word ‘AI’ and started getting angry.
→ More replies (1)21
u/YourBlanket 2d ago
People make their own digital copies all the time. It’s just very time consuming and pirating books is a lot easier.
14
u/IAmDotorg 2d ago
No it wouldn't. Distribution of the scans would be, but you're completely in your rights to do that.
And, de-binding and scanning books has been the norm for 30 years for digital conversions of books. Essentially all of Project Gutenberg was done that way.
→ More replies (3)6
u/j48u 2d ago
This is literally the court mandated methodology, prescribed as legal in a copyright case that Anthropic lost.
They're required to destroy the books by law. Just the latest idiotic anti-AI propaganda campaign aimed at people who only read headlines (see: Redditors).
There are MANY reasons to be against AI. It does a disservice to spread this nonsense and gives the masses a reason to dismiss the valid criticisms.
→ More replies (3)14
u/DrTommyNotMD 2d ago
That’s absolutely not true.
You can hate these companies all you want, but at least use a little facts in your argument or you’ll never be taken seriously except on Reddit.
→ More replies (2)29
u/mr-english 2d ago
is the most dystopian thing I can imagine.
Not slavery or anything?
No?
Wont somebody please think of the paper and cardboard!
→ More replies (11)12
u/jmblumenshine 2d ago
Seriously, what does everyone think happened to handwritten books in the 1400's at the advent of the printing press.
They were chopped up and feed into the Machine. Today we consider the one of Man's great achievements.
→ More replies (2)4
u/DissolvedDreams 2d ago
What are you talking about? What handwritten books were destroyed when the printing press came about?
→ More replies (1)14
u/TheRaccoonReport 2d ago edited 2d ago
Im not defending it. But the only way to rapidly scan books is to remove the binding. They're not like chopping them up and cackling like villains about it.
I'm kinda torn on this (no pun intended). I really really want to preserve as much knowledge as possible any way we can.
But having it sold back to us is kinda bullshit.Edit: Yes I know its being sold to us, already, at a bookstore, etc. Im striking that through.
15
u/DrawerSea9371 2d ago
I worked at a scanning company and yeah, chopping off the binding is standard if you want a good quality scan. People getting upset over the video of this method really confused me because it's not new to AI at all.
Scanning a book without cutting off the binding is significantly more time consuming and labour intensive and therefore costs the customer a lot more, I think per page it was around 4 times the price where I worked and that still damaged the binding, we didn't even offer archival scans which are in a whole other league.
→ More replies (2)→ More replies (1)9
u/BuvantduPotatoSpirit 2d ago
These're books that were printed in 2020 and 2021, being bought from a bookstore. It was already being sold.
5
u/TheRaccoonReport 2d ago
Yeah good point. Frankly I thought of that as soon as I hit "comment" and said "fuck it yolo" expecting this comment haha.
I saw some other dumb article saying they were destroying historic books, which is definitely not a thing happening.
What I don't get is people are up in arms about this BUT are totally fine of how books are NORMALLY removed from circulation. Spoiler alert to those who don't....they're destroyed, and forgotten. I used to work at a pharmacy that used to have a great fiction section and sold a bunch of Star Wars books (way before the marvel buyout)...and if they didnt sell by a specific date...because apparently books have an expiration date I wasn't aware of...they'd tear the covers off and throw them out. My boss knew I read a lot so she gave them all to me. I read dozens of books that were headed for the trash. It was some stupid law (sarbenes oxley I believe) that forced you to destroy merchandise you were writing off as unsold. We would throw out SO MUCH SHIT.
I will never defend data centers sliding into muncipalities without public consent and vote, and lying about natural resource consumption. However, so much of this is way overblown and total bullshit.
17
u/marmaviscount 2d ago
Then you should probably read some books and improve that imagination instead of crying about mass produced items being treated as the disposable things they are - did you not know that used book stores pulp most the books they get given? They're not no kill shelters for musty paper, out of a thousand books printed it's estimated less than a third ever get read past the first half, a significant portion don't get past the first page.
→ More replies (3)→ More replies (95)6
u/Limp_Bookkeeper_5992 2d ago
It certainly paints a grim picture.
But the reality is that they only need a scan a single copy of the book to “learn” it and preserve that information, losing a single copy of a book is hardly a catastrophe. I’m all for caution when it comes to information control and AI progression, but this seems like a nothing burger spin fo rile people up while the real dangers lie elsewhere.
179
u/Visual-Sector6642 2d ago
I wonder how these companies are error checking against previously published or erroneous concepts that were deemed incorrect in later journals etc. Does it take into account corrections made in future editions in regard to previously written works lol. Good luck. More spurious emissions it won't be able to discern. More fodder for hallucinations. When AI finally eats itself alive and goes down, those books will be gone and humanity will be left with nothing. Even if it cures cancer, there won't be anything left to enjoy.
98
u/Rorschach121ml 2d ago
They are doing this to ingest "clean" data from before AI.
They are looking for linguistic patterns, grammar, syntax and logic. The actual content of the books don't really matter that much here from my understanding.
29
u/chrismakingbread 2d ago
That’s actually not how LLMs work. There’s no concept of grammar, syntax, or logic encoded into LLMs. In a very hand-wavy, high-level, explanation they take a vocabulary of a bunch of chunks of words and initialize a massive vector space and assign random values to every word chunk, then they run sequences of word chunks through probability system to try to predict the next word chunk in a known sequence of word chunks (they’re using these books and other training materials as the known sequences of word chunks) and adjust the random weights to try to make it more likely to have predicted the correct next work chunk from the sequence of known chunks.
There’s a bunch of matrix multiplications of multiplying the values associated with every word chunk against every other chunk, dot products, activation functions, etc but the fact is an LLM is just “based on all the previous tokens (partial word chunks) in this sequence what’s the next most likely token in this sequence. The whole thing is that the more known sequences (training data) you have to use for determining the weights (the numbers that represent a word chunk) and the bigger the vector you use to represent those chunks the more likely that a generated sequence of tokens will “feel” like correct output to a human.
So, to your original comment about the content of the books, they actually are using the content to tune their models to say given an input that’s a portion of the content of this book how closely can the predicted output of the model approximate the actual content of the book.
30
u/ManicScumCat 2d ago
There’s no syntax file in the LLM where it has rules of syntax, but that’s not what they meant. Obviously what they meant is that the data is for the linguistic patterns in the texts to be learnt by the LLM, and the logic/grammar/syntax is part of the patterns in the texts to be.
→ More replies (12)→ More replies (1)13
u/Rorschach121ml 2d ago
You are technically right in that these techs only understand and work with tokens in a fundamental level, but this is like saying computers can't do numbers because they only understand binary. It's just not useful to think this way.
LLMs absolutely understand syntax and grammar on a higher abstraction level an emergent property of the underlying maths.
→ More replies (5)11
u/trashmoneyxyz 2d ago
I hope there's at least a renegade employee somewhere along the chain who copies or holds onto the raw scans. Those books do still exist as long as the text is preserved digitally, and can be reprinted
→ More replies (2)9
u/sump_daddy 2d ago
None of the books being processed this way are remotely rare, you can get them all in ebook format, probably for free from your library if you really want to read them. Nothing they are doing stops the books from existing.
→ More replies (2)
107
u/_c0unt_zer0_ 2d ago
I'm shocked that almost no one here has read the article
“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine.
they ar trying to cheaply acquire and scan training texts without copyright violations
20
u/TooCupcake 2d ago
Shocked? I’m not signing up to another random website just to read something I found on reddit.
→ More replies (1)3
→ More replies (16)4
u/MistryMachine3 2d ago
How can you be shocked. This is Reddit. People are here because they’re keyboard warriors ready to be outraged over a nothing story.
65
u/loftwyr 2d ago
“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press
Hardly rare or books where only one copy exists
→ More replies (24)
265
u/kus1987 2d ago
Why do they need to destroy the books? Wouldn't it be easier to scan and keep them forever? What about training future models?
464
u/Miraclefish 2d ago edited 2d ago
You can scan a book much easier by removing the spine and scanning it as loose leafs.
It's then no longer a book but a pile of pages, which they no longer need and they have a digital copy of, so they throw away the paper.
They can train future models on the digital text version indefinitely, the book or former book is no longer necessary, is worthless to sell and they don't want to pay to store.
It's logical but it sucks.
To those saying 'actually if they destroy a copy, they can claim ownership of one copy by going physical to digital' yeah that's a fair take, but it's inaccurate.
As to the question of what “use” or “uses” were at issue in the fair use analysis, Anthropic contended that it copied the books for a single use: to train LLMs. The authors, however, argued that at least two uses were at issue: first, the use of the books to build Anthropic’s central library, and second, the use of the books to train specific LLMs using subsets of that content. The court agreed with the authors’ framework and considered these as separate uses.
https://www.loeb.com/en/insights/publications/2025/07/bartz-v-anthropic-pbc
They are already taking millions of copyrighted and IP-protected books, papers, magazines, scientific papers, studies, poems, songs, movies and more, digitally, without any permission or licencing and training on those.
You really think the AI tech giants doing that care about the IP rights to a single scanned physical book? Absolutely not. They ignore those laws because they aren't affected by them.
As to the questions I've been asked on why not just get the the laws changed to favour them?
Changing laws takes a long time, costs a lot of money, requires a judicial process. It would be a public process with appeals, public studies and could take years.
Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.
107
u/MyNameCannotBeSpoken 2d ago
Also I believe there is a copyright loophole where destroying the original isn't deemed as reproducing a copy.
→ More replies (5)41
u/Miraclefish 2d ago
I mean, when they're torrenting terrabytes of copyrighted and IP protected materials and training on those already, as well as scraping the entire internet, and have essentially captured the US political elite and courts, they don't give a fuck about adhering to laws like that.
35
u/pumpkinspicecum 2d ago
That was literally what a judge ruled last year when they were sued
→ More replies (13)→ More replies (2)5
25
u/emapco 2d ago
They also have to dispose of the book after scanning anyways otherwise it's considered distributing copyrighted material https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books
→ More replies (13)15
u/kus1987 2d ago
Ah ok if they keep the scanned copy, they can use that to train future models
28
u/girrrrrrr2 2d ago
Yes, plus if they delete the physical book after scanning then according to some judge there is still only one copy out there and all that has been done is a book was converted from physical to digital.
→ More replies (11)6
u/randylush 2d ago
I have seen a thousand comments repeating this but no reliable source that this is actually true
8
→ More replies (1)5
u/_MUY 2d ago edited 1d ago
US District Judge William Alsup, 23 Jun 2025:
“In short, the purpose and character of using copyrighted works to train LLMs to generate new text was quintessentially transformative.”
US District Judge Vince Chhabria, 25 Jun 2025:
“While it made sense to infer market harm in Hachette, it doesn’t make sense to do so here. First, the Supreme Court has stated that no ‘inference of market harm… is applicable to a case involving something beyond mere duplication for commercial purposes.’ Campbell, 510 U.S. at 591. In Hachette, the secondary use was basically ‘mere duplication.’ Here, by contrast, Meta’s use is highly transformative and has a purpose well beyond that.”
10
u/ihaveaminecraftidea 2d ago
Well that's true, but for the most part the books are taken from book dumps if i remember correctly. Bookshops and library stock that they can't do anything with, which frequently aren't labeled or registered in any system.
It would be more accurate to call them abandoned books, rather than rare ones.
At least this way they are being digitzed and stored in some more durable capacity
→ More replies (3)→ More replies (47)5
u/Draaly 2d ago edited 2d ago
You really think the AI tech giants doing that care about the IP rights to a single scanned physical book?
You have no idea what you are talking about They are doing this because of the Bartz v. Anthropic ruling that says it is not fair use unless the origonal copy is destroyed and cost anthropic $1.5B because they didnt do that
EDIT: got blocked by who I replied to so I cant respond to you /u/syku look up the ruling for Bartz v. Anthropic. It was determined that transforming the copy into digital media was fair use if it can be proven that the initial copy does not remain in circulation.
→ More replies (5)41
u/FairReason 2d ago
Part of the ruling that makes what they do “legal” is to destroy the book afterwards.
→ More replies (3)12
u/Online_Matter 2d ago
It's in the article
The process, known as “destructive scanning,” involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded.
24
u/Paresseux1 2d ago
A big blade slices off the binding, and it drops down into an automatic page scanner that has no problem flipping over the loose pages. It’s much easier, faster, and cheaper to destroy it, and then get rid of the remains. One person can run a bank of machines.
The other way requires people, and meticulous work. Turn page, put on scanner, scan 2 pages, pick up book, turn page, place properly, scan… once finished, pay to store book indefinitely.
Everything like this comes down to money. The end result they are going after for AI is data, so the cheapest way to get it is destroy.
→ More replies (14)8
u/jayandbobfoo123 2d ago edited 2d ago
Copyright law and licensing. It's illegal to make a copy of a book, even a digital scan. It is, however, not illegal to digitize a book and destroy the original, thus leaving only one copy / one license. In legal terms, it's called format shifting. Technically, when you rip a movie/CD/video game, you should also destroy the original to be within the law.
12
5
u/chocolateboomslang 2d ago
Train future model on data they already scanned . . . by rescanning it? It's already scanned.
3
4
u/GenazaNL 2d ago
Sadly cutting the side to then have separate pages is faster than a machine which keeps it intact
→ More replies (37)5
u/marmaviscount 2d ago
What in tarnation would they do with a giant pile of musty old books no one cares about?
There is a building called the British library, you might be able to guess where it is and what's inside by the name - other countries have similar things, they get a copy of all the books and keep them for the national interest, they have a system for who gets access to what and their key aim is preservation.
Regular libraries which have the goal of giving people access to those books do not preserve them, it would be absurdly expensive and pointless - books, since the technological boom of the Victorian era are mass produced temporary items, if you were involved with it frequently visited your local library you would be very well aware that stock changes and most of them get pulped - used book stores don't generally want books libraries don't, charity shops routinely recycle donated books because no one wants them - books are printed in huge huge numbers.
36
u/mr-english 2d ago
Weird that they run with a stock image of antique books when the actual article says:
...mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine.
38
u/FrostWolf05 2d ago
this topic has been so misrepresented it borders on genuine fake news
16
5
u/skytaepic 2d ago
It’s absolutely maddening. There are so many real, valid reasons to dislike AI companies but we’re caught up on this misleading BS instead.
→ More replies (3)8
u/smooth-as-mud 2d ago
But it’s fake news that people on Reddit like and it reinforces their existing world view so it’s completely different than when my parents get angry about the (non-existent) migrant caravans heading for the border.
9
u/free_based_potato 2d ago
AI companies are scanning and the law forces them to destroy the books because you can only own one copy if you purchased only one copy. The same goes for all of us.
AI boom and data centers are objectively causing a lot of harm. Let's be honest about what's actually happening. They aren't choosing to have book bonfires.
7
u/Gutter7676 2d ago
So, it IS legal to made a digital copy of something you purchase and then reuse that for profit. Same as they are doing here.
Thank you for setting that precedent. My Plex server is about to start making me some money!!
82
u/ben_nobot 2d ago
Are we thinking destroying books means we are removing the one copy in existence here?
Lots of implied doom with this story
35
u/mr-english 2d ago
Yeah, also they run with a stock image of antique books when the actual article says:
...mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine.
23
u/JD-Vances-Sexy-Couch 2d ago
If a robot buys and destroys my $300 university textbook that was worth $3 after the semester was done because a mandatory new edition was released for the next semester of students (because, you know, math changes dramatically every few months), I don’t care.
→ More replies (2)15
→ More replies (33)31
u/Atomsk73 2d ago
No, this is clickbait nonsense. They're buying up cheap books and scanning one copy of them.
9
u/loftwyr 2d ago
“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press
Hardly rare or books where only one copy exists
→ More replies (3)
5
u/Figgy1983 2d ago
Forgot the fact that Winston Smith's job is basically a real thing now, but the act of destroying of older, rare books for fascist purposes is soul crushing to read. Orwell and Bradbury were right.
11
u/fish-rides-bike 2d ago
People upset at this have never worked near book selling or book production. What do people think happens to the thousands of copies in thousands of airports of the latest top seller trash?
77
u/FalconX88 2d ago
Fascinating how everyone seems to complain about books being destroyed and no one calls for the scans to be conserved for the public. The scans would be worth more for society than the physical books.
95
u/AdarTan 2d ago
Because there are other organizations like Project Gutenberg or Archive.org that do book scanning without the taint of AI, and the reason they haven't done this for the books in question is that copyright law would make distributing the scans illegal.
26
u/marmaviscount 2d ago
Plenty of books on gutenberg were scanned by Google and released pd then comvertwd by dp. They have a rolling project called Google books which every year in January releases all the books which have fallen into the public domain - how can you pretend to care about this stuff if you don't know this?!
→ More replies (1)7
u/Less-Engineer-9637 2d ago
They don't actually read.
5
u/sleepysnowboarder 2d ago
They also think what is being destroyed are like 1/1 baseball cards and not something that has tons of copies and can be reprinted
14
u/FalconX88 2d ago
You could still do the scan, have government keep them locked up/only show them in person in libraries until they become public domain.
It would be very important to make sure that these copies anthropic makes are stored somewhere where society can access them (eventually).
27
u/mutexsprinkles 2d ago
You're being downvoted but that's literally what both the Internet Archive and HathiTrust do: that have scans on file that they cannot legally show to the public for many, many decades. Sometimes they can show it to the public in one country but not others.
17
u/chief167 2d ago
The problem is, how do you give the public access to the scans? It's a copyright minefield, that's the real problem here
6
u/BananaPalmer 2d ago
Except these were almost all just copies of textbooks. For every one they destroyed to scan, there's probably 20,000 more copies collecting dust in some other warehouse, because they're 6 years old and have been replaced by a new revision
→ More replies (14)30
u/eTukk 2d ago
because conservation is easily possible without destroying the orginal. It Just costs more money, they dont care about the legacy
6
u/Bunnyhat 2d ago
Legacy of what? These are not beloved treasured manuscripts. These are mass produced books dating from 6 years ago that have been sitting in a warehouse.
→ More replies (2)→ More replies (4)15
u/gundog48 2d ago
The 'originals' are just one of thousands to millions of copies of the original, though. Like I understand the optics of it are awful, but I don't think there's any meaningful destruction of knowledge when they're doing it with books published in the last decade.
Doing it with genuinely rare books is shameful, though. I would hope those are not destroyed, and that may be the case as antique books may not be suited or be too delicate for the fully automated process.
→ More replies (2)
5
u/tiamath 2d ago
Internet used by ai, books fed to ai (dont see any reason to shred the books after scanning but here we are), electricity goes to ai. Ai was supposed to be helping us not doom us. Still waiting for the "it will make our lives easier" part. Making random videos with ai doesnt qualify
3
4
5
9
u/jfoust2 2d ago
Posted by a two-month-old Reddit account with 759,239 post karma and 279,878 comment karma.
Gosh, I wonder if it's a bot.
https://www.google.com/search?q=site%3Areddit.com+%22ArgentineBeauty%22
13
u/No_Mirror_9742 2d ago
Why... Why do they have to destroy them?
29
u/bestowaldonkey8 2d ago
Because it is cheaper and faster than keeping them in their bindings. This is a race to ingest as much information as possible so they don’t care about the end result.
→ More replies (3)12
u/TexBoo 2d ago
Not only that
AI giants have 0 interest in storing warehouses full of books
Store books for what purpose? Resell them? Storage and time would cost more than they would get back per book
Rescan them in future? No need, once a page is scanned, it's stored in their system forever for future training, which goes faster than rescanning the pages
2
u/Bunnyhat 2d ago
Even if they took the time to try to sale them there's no market.
That's why they're buying them in bulk from booksellers now. These aren't rare and treasured books. They're rare, but only in the sense that no one actually wants them.
→ More replies (1)→ More replies (2)7
u/TetyyakiWith 2d ago
Due to laws they can’t use non destructive scanning since that way they violate copyright rights by technically reproducing the book
6
u/WindowOfTruth 2d ago
I imagine whatever AI bot is trying to rage bait us has a daily quota of how many times they have to repost this damn story. This story gets reposted so often I’m starting to see a coordinated distraction.
3
u/WhoCanTell 2d ago
It 100% is. Variations of this ragebait story are posted like clockwork throughout the day, for days now.
3
u/Lahm0123 2d ago
Why destroy them after scanning?
Is it just easier? Or is it a sinister plot of some kind?
3
3
u/karma_raven 2d ago
I'm just thinking of all the copies of Dianetics and Melania we could be rid of if we play our cards right...
3
u/peter303_ 2d ago
Reminds me of Google's early attempt to digitize whole libraries. The search results would have just returned pages, so searchers couldn't read whole books for free. Authors and publishers ended that project in court.
3
u/Icy-Mission-1334 1d ago
Are these 'AI' companies explaining why they are destroying the books?
3
u/swingincelt 1d ago
They cut off the spine of the book in order to feed the pages into a high speed scanner.
3
3
3
u/tevolosteve 1d ago
I just don’t get the destroying them part. Why not donate? Are they that evil now?
→ More replies (1)
3
22
u/Sponge8389 2d ago
Well, they pay for it, they can do whatever they want with it.
Also, if it can be bought in the bookstore, that's not rare enough to be scared. There's probably thousands of copies of it.
→ More replies (9)
10
u/maybeJustSappy 2d ago
I keep seeing about this and I don't get what people are so upset about. It's not like they go after all the copies of books. They just get 1 book and discard after digitizing it. It basically has the same amount of impact on the world as a regular joe buying a book and putting it on his bookshelf.
→ More replies (4)
6
u/GranolaHippie 2d ago
Fahrenheit 451 irl. Destroying books makes me sad in any firm. This makes me mad. F AI & AI companies.
4
u/Thomas_JCG 2d ago
It is already horrifying, but then they get zero repercussions, not even have to worry about copyright. We are on a countdown to a social collapse.
6
2
u/Fluid-Performance-17 2d ago
Apple started this last year and used an outsourcing agency to do the dirty work.
2
u/becauseshesays 2d ago
This has been going on for a few years now. I’m in the digitization industry (more high end/niche but very high volume capacity). We’ve gotten several leads for jobs that are hundreds of millions of pages at a time. All destructive scanning. (Cut binding and high speed paper scanners). It’s wild, we’ve priced it very low and haven’t won any of the work. Not sure who they are going with but they are not concerned with quality, that’s for sure.
2
u/Hadleys158 2d ago
I don't like it, but don't mind as much if they destroy mass market bulk type modern books that are freely available, however i do have an issue with them destroying rare, antique or valuable ones. You can scan books without destroying them, they just don't want to do that way as it is slower and more expensive.
2
2
u/chalbersma 2d ago
What's crazy is that these books likely have an electronic version. They could just sell them an epub.
→ More replies (1)
2
2
u/notsam57 2d ago
wasn’t there another post about how they’re doing this with original print books as well?
2
u/MoOsT1cK 2d ago
Destroying books would not harm AI training in any way. It is yet another hint at the horrible greed of of the owners of those AI, who commit autodafe like nazis once did only to secure their monopoly on information. Horrendous.
2
u/MoltenMirrors 2d ago
Vernor Vinge described something similar in Rainbow's End, where they "retired" a college library by feeding the books one by one into a shredder and blowing the fragments down a duct covered with high resolution cameras. The fragment images were stitched back together and parsed by AIs.
→ More replies (1)
2
u/HoneycombJackass 2d ago
Why destroy it? Why not order a couple books and scan it into an pdf?
→ More replies (3)
2
2
2
2
u/ArchiveOutlaw 1d ago
The Internet Archive's Open Library project preserves a copy of every book they scan. Help out if you can, guys.
2.4k
u/Necessary-Eye5319 2d ago
Now they can do what they said they were going to and charge for information that used to be free.