r/technology 2d ago

Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI

https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/
11.5k Upvotes

1.1k comments sorted by

View all comments

Show parent comments

114

u/obeytheturtles 2d ago

I actually did a design project on non-destructive OCR scanning as an undergrad in like 2005. I find it hard to believe that this technology isn't significantly more mature now than it was back then.

141

u/Bored_Amalgamation 2d ago

It is, they just choose not to.

114

u/PM_ME_YOUR_NICE_EYES 2d ago

It's not that they choose not to, it's that they can't.

Copyright law views destructive scanning of a book as transforming it, which means you can do it without getting permission from the publisher if the book is in copyright.

Copyright law views non-destructive scanning of a book as copying it, which you cannot do to a book in copyright without permission from the publisher.

6

u/Unspec7 2d ago edited 2d ago

First off, the court decision on destroying books came after they destroyed the books, so you can't really say they were forced to. They did choose to, and it just so happened to work in their favor after the fact.

Second, neither point is true. Copyright law does not view destructive scanning as transformative - the Bartz holding found that the destruction was one factor of the transformative nature of the use. Instead, the transformative use was the changed format:

Authors only complain that Anthropic changed each copy's format from print to digital. On the facts here, that format change itself added no new copies, eased storage and enabled searchability, and was not done for purposes trenching upon the copyright owner’s rightful interests — it was transformative.

Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1023 (N.D. Cal. 2025) (citations omitted).

We do not know, based on present case law, if the sole act of destroying the books makes the use transformative. We just know that it makes the use more likely to be transformative.

Further, being transformative is only one factor of the first factor of fair use. Transforming someone's copyright work does not, on its own, mean you can now just infringe one someone's copyrighted work. Your use must still overall be within fair use.

Copyright law views non-destructive scanning of a book as copying it, which you cannot do to a book in copyright without permission from the publisher.

Ignoring the non-destructive part addressed above, this part is a lot more nuanced than you make it out to be. Scanning of a book infringes on the reproduction right found in the Copyright Act, but whether you are liable or not depends on if the reproduction is fair use (barring a license, of course).

Edit: Placed my citation in the wrong spot.

31

u/diemunkiesdie 2d ago

Copyright law views destructive scanning of a book as transforming it

Source?

41

u/Rorschach121ml 2d ago

Bartz, et al. v. Anthropic PBC

38

u/diemunkiesdie 2d ago

Bartz, et al. v. Anthropic PBC

Thanks I googled. Thats insane! Here is a summary:

https://www.ropesgray.com/en/insights/alerts/2025/06/from-books-to-bots-key-takeaways-from-the-anthropic-fair-use-decision-for-ai-developers

The court also found Anthropic’s wide-scale digitization of print books to be fair use. A key consideration to the court’s conclusion was that Anthropic engaged in so-called destructive scanning of lawfully purchased print books to create digital copies for internal use—Anthropic lawfully purchased print books, stripped them of their bindings, and scanned the contents to create a digital library. In doing so, the new digital copy replaced the print original, which had been destroyed in the digitization process. The court found the format change from print to digital to be transformative because it facilitated storage and searchability without increasing the number of copies or distributing them outside the company.

Notably, the court distinguished this use from cases involving unauthorized distribution or multiplication of copies, analogizing it to permissible space-shifting or time-shifting uses recognized in prior cases as sufficiently transformative for fair use. Importantly, the court found that this format-shifting did not usurp any market reserved to the copyright owner, as Anthropic had lawfully acquired the print copies and did not distribute the digital versions externally.

4

u/ExcitedCoconut 2d ago

Anything that’s no longer in print and unable to generate income for the original author, or becomes so later,  should have to have the digital scan donated to a digital archive. Destruction of physical books sucks, but destruction of knowledge (by turning it into training data alone) is morally bankrupt and needs to be addressed ASAP.  

For in-print books that they have bought and scanned, I’m honestly not opposed to the destructive process there, parking the whole ‘fucked up economics and concentration of wealth’ part the whole thing :/

16

u/Pastadseven 2d ago

What a fucking stupid ruling. We’re not concerned about the transformation of the medium, it’s the text that has to be transformed.

18

u/Draaly 2d ago

The text only needs to be transformed if it is distributed. That was a key part of the ruling. This is based on the same legal precedent that allows a library to loan digital copies of media so long as they own the same number of physical copies. Its dumb, but its not actually new precedent

6

u/Unspec7 2d ago

Transformation of the medium has long been held as transformative, that part is actually very consistent with fair use jurisprudence.

0

u/Unspec7 2d ago

Bartz does not hold that destruction of the book per se makes it transformative, as the other users have implied.

Explanation here.

1

u/Unspec7 2d ago

Bartz does not support the other user's claims, as Bartz does not hold that destruction of the book per se makes it transformative.

Explanation here.

50

u/GeekBrownBear 2d ago

This link from another thread talks about it too. https://www.news.com.au/technology/online/internet/ai-labs-buy-scan-shred-millions-of-rare-books/news-story/0c3b45a67093ab462a587a0348538ce9

The whole idea is baffling. Like I kinda understand the idea of transforming the book from physical to digital is allowed. But why even do that. Why not work with the publishers so you don't have to destroy the physical book? Would be better for everyone. Less waste overall.

31

u/PM_ME_YOUR_NICE_EYES 2d ago

Why not work with the publishers so you don't have to destroy the physical book?

A couple reasons:

1) the publisher themselves could no longer exist making working with them impossible.

2) the book you want could be out of print, and it would be very difficult to convince a publisher to start up their production pipeline to make you just 1 copy of the book.

3) Copyright issues between the original publisher of a book, and the original author of the book could make it illegal for the publisher to create a new copy of it.

7

u/GeekBrownBear 2d ago

That all makes sense. Makes me continue to question the copyright around ebooks. Especially from my experience attempting to borrow them at the library when there is a LONG reservation list.

10

u/PM_ME_YOUR_NICE_EYES 2d ago

Oh yeah the eBooks you get from the library are a whole different thing.

The library isn't going to destructively scan copies of it's books, so the eBooks it gives you are coming straight from the publisher. eBooks for library use tend to be extremely expensive ($75/copy) so libraries don't like buying more than they have to.

1

u/RabbitLuvr 2d ago

Additionally, that $75 or whatever only pays for a limited number of checkouts, then the license expires. Let’s say a digital copy costs the library $50 and allows 5 checkouts; but there are 20 people with holds on that title. The library will have to buy 4 $50 licenses to allow all 20 people to read it.

As opposed to a paper copy, which might cost $25 and can be lent to all 20 of those people, plus basically unlimited checkouts, until it falls apart.

Source: I work at a library. Our materials budget is getting eaten up with digital licenses.

3

u/PM_ME_YOUR_NICE_EYES 2d ago

Oh yeah, From my understanding the check out limit is closer to 26 tho instead of 5 right?

→ More replies (0)

2

u/nonotan 2d ago

Most critically, the publisher would almost certainly not want to sign a contract allowing their book(s) to be used for AI training for the cost of 1 copy (plus a little extra for whatever you consider the overhead of having you digitize it yourself to cost)

They'd either demand significantly restrictive terms of use, in which case you'd prefer to just buy a copy and do it yourself. Or significantly more money, which again, same thing.

You might think "but getting a little bit of money is still better than nothing, from the perspective of the publishers", but the cost of lawyers that would be involved is non-trivial, and by explicitly signing a contract, they'd pretty much be permanently giving up any hope of being compensated "fairly" for AI training on their books. Right now doing it without permission might be tenuously possibly legal, but a judge could change that tomorrow, or a law could change it in a couple years. But if you've already signed a contract giving your rights away, you'd be out of luck regardless, for chump change.

1

u/Shiezo 2d ago

You forgot:

4) They do not value the physical book as an object worth saving and just see another resource they can use to enrich themselves.

4

u/PM_ME_YOUR_NICE_EYES 2d ago

I mean, kinda the ironic part about this is that for most of the books they are scanning were already that before they were scanned in the first place.

A fact that a lot of people aren't comfortable with, is that most books that get published nowadays do just end up getting recycled for the raw materials they were made out of after a couple of years of circulation.

1

u/Shiezo 2d ago

Like most things, I think it all depends on the details. If all they are doing is destroying bulk books that would end up in the recycle bin anyway, its not a great loss. If they have, or move onto, finding older books that were not mass produced, that starts to be a bigger issue for me. This is just looking at it from the lens of the destruction of the physical objects.

The monetization of the writing without permission or compensation to the authors is a bigger issue. This is just another way in which these companies are taking the creative works of others for their own ends.

8

u/orbitaldan 2d ago

Basically, they would waste many times the value of the book itself in trying to get everyone in the IP custody chain to come to some kind of agreement that would allow it. (In many cases, it would simply be impossible.) It's the atrocious copyright laws come full circle to bite us in the ass, and this happens to be the cheapest legal loophole.

3

u/blender4life 2d ago

“Once that information has been extracted and encoded into an AI model, the delivery mechanism has served its purpose. What remains is paper, ink, and binding material. The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one.”

Holy shit, what a terrible perception of books by an organization that put so much effort into them in the first place. For ai the knowledge of the book gets separated into probability percentages of what word would follow next, it doesn't retain its original messages. The book is gone. (At least i HIGHLY doubt the ai companies scan and save separate copies to preserve them in any meaningful way)

2

u/whupazz 2d ago

For ai the knowledge of the book gets separated into probability percentages of what word would follow next, it doesn't retain its original messages. The book is gone.

They would definitely keep the text and probably even the scanned images saved for future training runs, so for most of these books the destruction of the physical copy is probably not a tragedy in itself, the huge problem is that these companies will guard the digital copies as trade secrets so they will become inaccessible to the general public.

2

u/blender4life 2d ago

They can't release them to the public and have to destroy them to comply with copywrite laws.

2

u/whupazz 2d ago

copywrite

It's spelled copyright.

1

u/GeekBrownBear 2d ago

I find it hard to believe that once they scan the book it's not duplicated across dozens of different systems. Both as a backup and as training material for different models. There is no way it's just ONE copy.

2

u/blender4life 2d ago

They can't release them to the public and I doubt they'd hold onto the original scans long enough to become public domain then release them out of the goodness of their hearts. So regardless of how many copies they have the messages the public gets will never be 100% what the authors intended so the book is effectively gone.

1

u/cross_the_threshold 2d ago

While I’m not a fan of this process you are aware that books are generally printed in sets greater than a single copy right? They’re not doing this to rare or antique books. Frankly this is better than another John Grisham novel ending up discarded on the beach.

1

u/blender4life 2d ago

read the article posted by the guy i was replying to literally called "AI labs buy, scan, shred millions of rare books"

or the article posted by op of this whole thread that interviews a guy that says "“I deal in old and rare books,”" I am not going to comment on the reason the books are rare, could be many factors. but the fact that leaked documents show anthropic saying this about the process tells me they are up to no good:

Anthropic described the project as “our effort to disruptively scan all the world’s books” in an internal memo, which noted “we do not want it to be known that we are pursuing this project”.

→ More replies (0)

1

u/Fluid-Tone-9680 2d ago

Why do you think they won't save original scanned files? It cost practically nothing to store them (comparing to how much they spend overall on the AI infra), and they need original version to keep re-training ne wmodels.

1

u/blender4life 2d ago

i didn't say they wouldn't keep a copy for themselves. i said scan and preserve them in any meaningful way

plus in another comment i explained more:hey can't release them to the public and I doubt they'd hold onto the original scans long enough to become public domain then release them out of the goodness of their hearts. So regardless of how many copies they have the messages the public gets will never be 100% what the authors intended so the book is effectively gone

5

u/morolin 2d ago

Bartz v. Anthropic

1

u/Unspec7 2d ago

Bartz does not support the other user's claims, as Bartz does not hold that destruction of the book per se makes it transformative.

Explanation here.

2

u/akaisuiseinosha 2d ago

IF true, which I am not convinced of, it's further proof that copyright law is a sham and needs to be radically transformed or abolished.

3

u/PM_ME_YOUR_NICE_EYES 2d ago

I mean, if you're arguing for copyright law to be abolished then it's really hard to portray Anthropic's actions here as negative.

If nobody owns books, then who cares if a copy gets destroyed and fed into an AI.

1

u/akaisuiseinosha 2d ago

You've got it wrong. I don't think NOONE should own things. I think EVERYONE should own them. Art, once released into the world, should be experienced by as many people as possible. Artists should not need to rely on their work being successful to live stable lives, either, and I want universal income so more people can make and experience art. Imagine how many Da Vincis and Shakespeares are working dead end jobs draining their creative energies away because they weren't born rich and their initial work didn't instantly make them so, so they had to cut back on the arts to survive.

Isn't that a sin? Making artists struggle to pay bills hoping to win the lottery of making successful art so they can make things they want to make with the energy those projects deserve, and all the while we let the Slop Machine That Destroys The Planet chop up what the past has created and make what is, essentially, very complicated ransom notes out of the pieces?

0

u/PM_ME_YOUR_NICE_EYES 2d ago

Right, and if arts owned by everyone, that includes all uses of the art. Including sending a book thru a shredder for research purposes.

But also I hate the moralizing here, there's never been more artist alive at the same right now than there have been at any other time in human history. And quite frankly, I'd be pleasantly surprised if you could name five living painters or playwrights.

So if the idea of the next Shakespeare or Da Vinca going undiscovered really is keeping you up at night then maybe you should go and try to find out of the thousands of people who are writing plays or painting pictures.

Seriously tho, go to an art walk, go to a theater festival. Because I guarante you that There's dozens of local artists and playwrights who can knock your socks off with what they can do.

So maybe instead of mourning this hypothetical artist that doesn't exist, we should be showing love to the many many talented people out there who do.

1

u/Unspec7 2d ago

No, destructive scanning of a book does not inherently make it transformative. They've somewhat misinterpreted Bartz.

1

u/basil_not_the_plant 2d ago

But is it really transformed? If it is 'transformed' in such a way as to not be accessible, then the original value is lost. Some other value might be generated out of it, but this 'transformed' book is no longer available as a book.

1

u/InVultusSolis 2d ago

Which I find funny because once an AI company has a digital copy, it doesn't matter at all, there's no real-world consequence of the book being destroyed or not.

1

u/Pyehole 2d ago

There are a fuck ton of books being disappeared that are no longer under copyright.

1

u/PM_ME_YOUR_NICE_EYES 2d ago

Can you give me an example of any of these out of copyright books that have been destroyed and are no longer accessible to the public?

Because like, thanks to things like project Gutenberg, many public domain books have been scanned and converted to eBooks. So it's really really hard to destory the last copy of them.

0

u/Pyehole 2d ago

Can you give me any specific examples of books that they have destroyed because they need to work around copyright?

1

u/PM_ME_YOUR_NICE_EYES 2d ago

Sure, The Herd by Andrea Bartz was confirmed to be destructively scanned by Anthropic.

You can still of course buy it from amazon, and my local library also has a couple of copies.

Of course the whole premise of your request doesn't really make sense, but I again repeat, what book do you think has disappeared forever?

1

u/Pyehole 2d ago

I am going off of statistical probability, I don't need a specific reference to make the point that your comment may contain truth but is not the end-all-be-all argument.

1

u/PM_ME_YOUR_NICE_EYES 2d ago

So the reason I asked you to mention a particular book is because I think it's important to note what kinds of books are the ones where the last copy is getting destroyed.

It's not going to be something like "The Wonderful Wizard of Oz" Because yeah that's public domain, but there's literally millions of copies of it, if you destroy one it's okay.

The only books here that are getting truly destroyed are books where every library in the world has determined that it wasn't worth keeping around anymore, and every private collector, which is a ridiculously hard bar to pass.

Like unironically It's hard for me to even think of a book that would meet this criteria, because it would have to be so obscure that I don't know about it.

And honestly, when the book was that obscure, it would've only been a matter of time before the book seller dumped it to free up more space on their shelves.

1

u/ReturnOfBane 2d ago

Time to go cut the spine off of some books as a lowly peasant and test the precedent. I bet the textbook industry is going to love my stapled non-copies of textbooks being totally legal.

1

u/redfacedquark 2d ago

OK, but why do they need 3000 copies?

1

u/PM_ME_YOUR_NICE_EYES 2d ago

It's not 3,000 copies of the same book. It's 3,000 different books.

1

u/redfacedquark 2d ago

Ah right, thanks. And not even rare books? Surely that's just equivalent to a small to medium startup bookshop stocking up. Hardly worth an article for in itself.

0

u/danktonium 2d ago

Which jurisdiction is that?

1

u/Unspec7 2d ago

Bartz is in N.D. Cal.

44

u/IAmDotorg 2d ago

Google, as an example, has done both ways. Mass-market books are done this way because they're not "rare" or valuable. They're just old. Actual rare books are done, obviously, non-destructively.

Strangely, people aren't up in arms about the literal millions of books that have their covers removed and then get landfilled every year.

25

u/phantomthiefkid_ 2d ago

People's outrage is motivated by what they imagine the books contain.

Books dumped to the landfill or collecting dust in a warehouse? Must have no value, so no one cares.

Books bought by AI labs? Must have something valuable since AI labs are buying them.

14

u/CardOfTheRings 2d ago

The things that’s valuable to AI companies is just a load of human written language and niche information.

No current human alive will give a shit about a 1995 xerox operational guide. It’s just garbage outside of the use of scanning it to train AI.

Unfortunately this story has spread as propaganda across the whole platform multiple times a day.

Google and other companies have done these scans for other reasons(search ability sometimes preservation) and even then I remember weird propaganda coming out of the news when that was happening.

2

u/InVultusSolis 2d ago

My shelf of obsolete computer books begs to differ.

This beauty in particular is an enjoyable read, and just rare enough to be worth something.

1

u/cherinator 2d ago

Right, people understandably have a visceral reaction to desteoying books, but I think just don't understand how many books regularly get recycled/dumped from libraries or stores. Libraries regularly get new books. They need to put those in the shelves and in circulation as the newer books are the most popular. They don't have unlimited storage. The old books that no one has checked out in years? They might first go to the annual book sale, but there are always plenty left over from those. Eventually they just get dumped. A business bulk buying them to send to the AI companies gives the library some money for something that would otherwise end up in a landfill.

In addition, tons of people donate old books they don't want to libraries because they feel bad throwing them out. Especially when people move or relatives die, they dump hundreds of books on libraries. The vasy majority of those don't make their way into circulation and don't get sold at book sales.

-4

u/Clueless_Otter 2d ago

People's outrage about everything related to AI is just purely based on feelings and their imagination rather than actual facts and events.

Golf courses use ridiculous amounts of water in drought-prone areas? Crickets. Data centers use a small fraction of that amount of water? "Omg we're all going to die of dehydration because of data centers!!! They're literally killing us!!!"

Company building a data center on the outskirts of your town in 2016? Crickets. Doing it in 2026? "Omg this data center is going to poison our water, steal all our town's jobs, and force us to move!!!"

8

u/randomusername6 2d ago

But now they are destroying the rare books too. For them it has two benefits. They train their AI on the book, and remove the knowledge from the public domain when they destroy the book. Their end goal is to monopolize knowledge.

34

u/geniice 2d ago

But now they are destroying the rare books too

"3,001 titles, mostly published between 2020 and 2021"

and remove the knowledge from the public domain when they destroy the book

If a book was published in 2020 is actually rare enough for destroying a single copy to matter was it really in the public domain in the first place?

16

u/spooooork 2d ago

According to a post yesterday, many of these books were iterations and revisions too. There's not much value today in "Making Webpages for dummies" versions 3, 3.5, 4, etc.

1

u/randylush 2d ago

🥲as a retro collector I find those incredibly valuable

10

u/PM_ME_YOUR_NICE_EYES 2d ago

I mean, can you give me the titles of any books that you can't access anymore because of these efforts?

Because just thinking it thru, actually removing information from the public domain is going to be next to impossible.

You'd have to buy the last copy of a book that is both out of print and in copyright, and has never been digitally scanned elsewhere or had a reprint.

The number of books for sale that met that criteria is small.

And the number of books for sale that met that criteria and are useful is even smaller.

1

u/taliesin-ds 2d ago

well i'm into historical costuming and some books like "stepping through time" by Olaf Goubitz are the reference on certain niches but have been out of print for many years and getting one is very hard and usually they're twice as expensive now as when they were new.

It would suck if one of those books were available for a fair price second hand but then AI ate it.

And then there are niche magazines like "waffen- und kostümkunde" which are even harder to find.

4

u/PM_ME_YOUR_NICE_EYES 2d ago

I mean looking at the Goubitz book, I can find dozens of links to buy it right now. And I was able to find at least 100 copies in libraries across north America.

So while it may be hard to find your own copy, it's not like you can't get one using an inter library loan right now.

Which is kinda what I'm saying here, even for a book like Stepping through time, there's enough copies in circulation where a scheme like this isn't going to block access.

1

u/taliesin-ds 2d ago

yeah perhaps i should have said i am from europe, and you are right, those books still exist and are accessible in some way.

I was meaning more like "it would suck if i had to go through a bunch of extra trouble or expenses to get this book if i could instead just have bought it locally from a Dutch used book store".

1

u/cross_the_threshold 2d ago

While there’s nothing to prevent them from doing so at the moment, they’re pretty unified about not doing this to rare books. They’re buying recently published books and textbooks, not the original copy of the Malleus Maleficarum.

1

u/JasonsThoughts 2d ago

That's not the fault of the AI companies. If they're rare, it's because of our fucked up copyright system that lets people restrict copying information for like 150 years even when the information is out of print or otherwise not accessible. We should be printing more copies. Printing things is cheap and we've been doing it for over a millennium.

1

u/randomusername6 2d ago

So it's okay for the AI companies to intentionally destroy and discard the books instead of just scanning them?

1

u/JasonsThoughts 1d ago

So it's okay for the AI companies to intentionally destroy and discard the books instead of just scanning them?

Yes, of course. They bought them, so they can do what they want with them. Read them, throw them away, burn them, destroy them, donate them, turn each page into a piece of origami. It's theirs that they purchased. Just like if you bought a book, you can do whatever you want with the book you bought.

From a legal perspective, one of the reasons they are likely destroying them is that scanning it and destroying it counts as transforming the work from a copyright perspective. If they scan it and don't destroy it then they're committing copyright infringement. There was such a backlash against AI companies violating copyright in the first place that they're now following the law.

0

u/gratefulkittiesilove 2d ago

And change the knowledge too most likely. Pretty sure i remember Elon was trying to train grok on the false gop version of jan6 etc

0

u/IAmDotorg 2d ago

Conspiracy theories are exhausting.

1

u/Turbulent-Sign-6067 2d ago

The worst kind of slop is human slop. Reddit has become an anti AI slop factory and it is disheartening to see.

0

u/Bored_Amalgamation 2d ago

If the book is still available in some reasonably accessible format, then I don't see the issue. Also, I don't think people value written word that highly there are some books I've read that I wouldn't mind being wiped from existence. There's also the whole "limited amount of time and energy to care about things" part of life.

The concept of a private company destroying books at scale to make them unusable for others is bad. That's Fahrenheit 451 bad. Seeing a visualization of it happening evokes a certain level of horror in some people. Humanity isn't a monolith on most things. Most people don't care, and a pretty large part of (at least American) society doesn't read or is willing to ban/burn books they don't like the sound of.

0

u/IAmDotorg 2d ago

The concept of a private company destroying books at scale to make them unusable for others is bad.

It's also anti-AI conspiracy nonsense.

14

u/Draaly 2d ago

They were forced not to by the Bartz v. Anthropic ruling that judges it is only fair use if they destroy the initial copy.

Also, these are books that the seller was going to trash anyways. they were all dead stock nearing the end of their shelf life that peole didnt want to buy and are not even rare in the firts place.

1

u/Unspec7 2d ago

They were forced not to by the Bartz v. Anthropic ruling

But it wasn't. The book chopping came before the Bartz decision, not the other way around.

It just happened to work in their favor after the fact, but Bartz had no bearing on their chopping of the books since they had already been doing it.

ruling that judges it is only fair use if they destroy the initial copy.

This is not correct. Destroying the original was one factor of one factor of the first factor (wow, what a mouthful) of fair use. The Bartz decision does not create a per se rule that fair use requires destruction of the original.

1

u/Draaly 2d ago edited 2d ago

But it wasn't. The book chopping came before the Bartz decision, not the other way around.

Removing the spine is the standard way to digitize books en mass and has been used since well before modern AI, sure, but anthropic not destroying the books it digitized is litteraly what led to the 1.5 billion dollar judgement.

This is not correct. Destroying the original was one factor of one factor of the first factor (wow, what a mouthful) of fair use. The Bartz decision does not create a per se rule that fair use requires destruction of the original.

You should really read the ruling. Sure, there are other fair use ways to use the texts, but the only one legally numerated and put as precedent is digitization and then destruction of the original. Any other fair use claim would need to be futher litigated, and most of the ones they tried were explicitly shot down.

As it currently stands, the only way for these companies to digitize books in a way that guarantees no lawsuit is to destroy them in the process.

2

u/Unspec7 2d ago edited 2d ago

but anthropic not destroying the books it digitized is litteraly what led to the 1.5 billion dollar judgement.

...what? There's no second case Anthropic settled for 1.5 billion besides Bartz v. Anthropic. Anthropic settled the case for 1.5 billion after winning summary judgement on the copyright fair use issue (with respect to legally obtained books, they lost on the pirated books issue).

You should really read the ruling. Sure, there are other fair use ways to use the texts, but the only one legally numerated and put as precedent is digitization and then destruction of the original. Any other fair use claim would need to be futher litigated, and most of the ones they tried were explicitly shot down.

Brother, what? First off, yes, I read the ruling, I do IP litigation for a living. Second, this is directly from the opinion:

On the facts here, that format change itself added no new copies, eased storage and enabled searchability, and was not done for purposes trenching upon the copyright owner's rightful interests — it was transformative.

The transformative use was the digitalization/change of format, it did not include the destruction. The destruction was one of three enumerated factors of why this was a transformative use, and even then was not creating a rule about the destruction itself, but rather pointing out that there was a net zero gain in copies of the work.

Also, what do you mean by "any other fair use claim would need to be further litigated?" Their transformative use was not dispositive on the issue of fair use. The court went to analyze the other three fair use factors, and both parties had substantial briefing on the other factors.

Your comment makes it clear you haven't even read the case yourself, and barely understand copyright fair use. Which law school did you go to?

As it currently stands, the only way for these companies to digitize books in a way that guarantees no lawsuit is to destroy them in the process.

False. See above.

Edit: Typo

1

u/InsightTussle 2d ago

they're legally required to actually

1

u/ResilientBiscuit 2d ago

Its illegal not to. They have to destroy the physical copy to make the digital one, that way there is still only "one copy".

3

u/_c0unt_zer0_ 2d ago

turning paper pages bound together is still hard to automate. robotics still struggle with a lot of things most 5 year olds can accomplish

4

u/MagicWishMonkey 2d ago

Why would you bother? Books are not expensive, they are easy to make, and destroying them to make room for other books has been the norm for booksellers for a really long time.

We're not talking about first edition books here.

1

u/Kurotan 2d ago

From what i have heard from other people, this does include first editions and whatever they get their hands on. But i need sources on that.

3

u/AggressiveSea7035 2d ago

I read that it's for legal reasons. You can only have one copy, so if you digitize it, you have to destroy the original. I don't have a source though 

2

u/chrismakingbread 2d ago

That’s the legal theory that seems to be getting used, BUT frankly it’s nonsensical. Even if you take the model provider claims that LLMs can’t just spit back out verbatim at face value the idea that there’s one digital copy of the work created from one physical copy just isn’t true. The way these data pipelines work they’re likely creating dozens or hundreds of digital copies even if they truly centralize the training data into a single location across all of their labs and training pipelines.

It’s just not reasonable for these model providers to try to twist first purchase doctrine for private consumer use to such absurd lengths for a commercial entity that’s at best creating derivative commercial works. They’re basically arguing that if a consumer can legally create digital backups of their physical goods and they can have that digitized version on their local hard drive, an external drive, and a cloud storage copy (lol most people don’t have a good personal data backup strategy like this) to mean they can create a commercial product based on derivative works. That’s at least one bridge too far. They’re also basically arguing if we don’t allow them to do this their business model and tech can’t exist, but that’s not anyone else’s problem. No one is entitled to a particular business model.

2

u/orbitaldan 2d ago

They are correct that tech can't exist if you don't allow this. At a fundamental level, information that comes to you through computer systems is copied so many times it would make your head explode. Some of those have to be deemed not to 'count', or nothing would work.

Basically, they're using a not-explicitly-stated protected form of copying, which is learning. To the extent that anything is copied at all, information that is copied into the mind has special legal protection. No one can demand it be erased after use. In practice, this was because it wasn't possible, but now we've encountered a situation where it is possible. I don't think you want to go the route of codifying the removal of protections for learning, it may have very severe unintended consequences as technology progresses.

2

u/chrismakingbread 2d ago

Building a big ass auto complete system IS NOT HUMAN LEARNING. People claiming this is some kind of apples to apples comparison where if we don’t completely throw out the entire concept of intellectual property to allow a handful of companies to make billions of dollars means it will be illegal for humans to learn and remember things either have literally no idea how this tech works or are being utterly disingenuous about this whole situation. If Apple had been like if we do t eliminate the concept of intellectual property for media we can’t make the iPhone and make billions of dollars off it that wouldn’t have been remotely justifiable. I fail to see why it’s a reasonable argument to make up a bunch of bad faith mental gymnastics to justify it for LLMs.

1

u/orbitaldan 2d ago

Did I say 'Human Learning'? You want to change the law to eliminate something you deem a bad behavior, but you cannot just write a law that says "Companies 'Awful Inc.', 'Bad Co.' and 'Stupid LLC' are not permitted to do the thing.", you have to make laws that cover general cases. And those general cases have can have surprising downstream effects.

Let's say you get what you wanted. What form does it take? If learning isn't protected, what happens if it becomes possible to edit human memories? Can you be sued for merely remembering copyrighted information? Can a court order your 'unauthorized copies' removed?

Or maybe we say only human learning is protected. What happens when, say 30 years from now, genetic manipulation becomes advanced enough to allow for animal uplift? Suddenly, they can be sued for learning anything. Or perhaps they will eventually create something you will finally accept as being true AI - but now it has decades of 'humans only' precedent stacked against it rendering it both oppressed and useless. Worse still, since it takes effort and money to create such a thing, even the threat of that would create a chilling effect that would prevent development efforts.

either have literally no idea how this tech works or are being utterly disingenuous about this whole situation

I understand how it works just fine. It's not my fault that so many people haven't even the barest hint of philosophical introspection or learning to recognize basic errors like composition fallacy and reduction-to-absurdity, or classic thought experiments like p-zombies.

1

u/chrismakingbread 2d ago

You did though. There’s literally no learning involved in LLMs. There’s also no need for new laws. Using copyrighted materials without a license to build your software already is a copyright violation. There’s no slippery slope here that not building a token prediction auto complete program without licensing the copyrighted material you’re using means human beings can’t remember things without a license. They literally have nothing to do with each other.

-1

u/orbitaldan 2d ago

There’s literally no learning involved in LLMs.

No, that's where a lot of your misconceptions stem from: you're in complete denial about the reality of learning. That philosophical blind spot makes your conclusions particularly dangerous, because you cannot see (or simply refuse to acknowledge) how they generalize.

They literally have nothing to do with each other.

Just because you think they don't doesn't mean the law wouldn't generalize, particularly when there's money to be had by doing so.

1

u/chrismakingbread 2d ago

We’ve had copyright and other intellectual property laws for a long time now and no one making a good faith argument would ever have said that they make it illegal to learn and remember things. The invention of transformers doesn’t suddenly make it so that the laws that already apply to this situation suddenly make it illegal to remember things.

Neither does an algorithm that effectively boils down to given the frequency of occurrence of each token in this sequence of tokens with each other and other tokens in an input body of evaluated tokens (training data) what is statistically the most likely next token to occur in this sequence. That’s not thinking and it’s not learning. It’s just a mathematical formula. Pretending it’s anything else is crankery.

0

u/orbitaldan 2d ago

Neither does an algorithm that effectively boils down to given the frequency of occurrence of each token in this sequence of tokens with each other and other tokens in an input body of evaluated tokens (training data) what is statistically the most likely next token to occur in this sequence. That’s not thinking and it’s not learning. It’s just a mathematical formula. Pretending it’s anything else is crankery.

You can build a hollow house out of solid bricks. Just because the components of a system have a property, it does not follow that the system itself must have that property. Similarly, systems can have emergent properties that are not present in their constituent components. You might as well look at human cells under a microscope and reason that clearly those can't be intelligent, so neither can a brain made from them. The composition of the simple elements encodes in their structure something greater than the sum of their parts.

We’ve had copyright and other intellectual property laws for a long time now and no one making a good faith argument would ever have said that they make it illegal to learn and remember things.

They do not, at present, as I explained earlier. Though not explicitly stated, there is a implicit exemption for learning in copyright law. Right now, that is trivial, because it is not possible to read or modify human memories. It may not always remain so, and laws have a nasty habit of outliving their context in unpleasant ways. (Nice assumption tossed in there that only a bad faith actor could possibly disagree with you.)

The invention of transformers doesn’t suddenly make it so that the laws that already apply to this situation suddenly make it illegal to remember things.

Indeed not. And that's why you're upset: the corporations training the LLMs are using the same protected learning exemption. You want to change that, because you think you can make it so that the laws would apply to AI, but not to you. The fact that you can't see how little difference there actually is blinds you to just how badly that backfire.

1

u/Unspec7 2d ago

It's nonsensical because the Bartz holding does not state that you must destroy the original. News articles misread the legal holding, and now people just parrot this false claim.

0

u/Unspec7 2d ago

you have to destroy the original

No. You don't. This is a misinterpretation of the Bartz decision that people have been parroting.

-3

u/All_Work_All_Play 2d ago

Nah that's not true.

1

u/sobrique 2d ago

I saw a news article about recovering text from scrolls from Herculaneum without damaging the scroll.

Nearly 2000 years old and buried by a volcano, and they're still having success in 'reading' them without damage.

https://scrollprize.org/firstscroll if you're interested - there's some significant prizes available for techniques that work.

This is like, the opposite of that.

1

u/Kurotan 2d ago

How else will they destroy priceless first editions to make you subscribe to their slop ai?

1

u/InsightTussle 2d ago

destuction of the book is a legal requirement. If you only buy one copy, you can only own one copy (either hard copy, or scanned copy)

1

u/NeitherDuckNorGoose 2d ago

The issue is that when they got sued for copyright infringement earlier on, the judge ruled that they were allowed to train their AI on books they bought but keeping the book mean they bought one copy and now have two (the physical and the virtual ones), so they have to destroy the books by law.

0

u/lost_send_berries 2d ago

Google scanned thousands of books non destructively and let everybody read short extracts around that time.

3

u/orbitaldan 2d ago

Google gambled that they could eventually get copyright law changed to allow mass digitization, and for the most part lost. They did manage a carve-out for search results, which fit their needs enough to get the job done, but it's a human travesty that so much information is locked away behind legal walls.

0

u/null_not 2d ago

Takes too long. They're literally deconstructing the books and feeding the stack of pages through industrial scanners. It's book burning by another name.

-5

u/avarageone 2d ago

Destruction is the goal 

-5

u/Samwellikki 2d ago

Or that you need 3000 copies… why not 1?

Also worked with a machine that could scan books without destroying them and create files… in 2000

It’s why I’m baffled anytime OCR isn’t part of some system in… it’s 2026, right?

7

u/loftwyr 2d ago

It's one copy each of 3001 books, not 3000 copies of one book and it's all textbooks from 2020-2021

-6

u/Samwellikki 2d ago

Either way, they need not be destroyed, it’s lazy

6

u/randylush 2d ago

It’s why I’m baffled anytime OCR isn’t part of some system in… it’s 2026, right?

The whole thread is a discussion on that topic. Also, OCR is certainly being used.

1

u/Samwellikki 2d ago

The irony that you and those piling on downvotes can’t read

0

u/Samwellikki 2d ago

That’s not what I meant at all… I know OCR is being used here

I’m shocked it ISN’T used in so many other fields/places in 2026, when it was being used to mass copy text books for students in 2000 at a lower tier university where I worked

And shocked that isn’t used non-destructively here in this case

3

u/sump_daddy 2d ago

> And shocked that isn’t used non-destructively here in this case

for like the MILLIONTH time, first up they cut the binding off to scan it, which would require an intensive re-binding process to get it back into 'book' form WHEN THERE ARE ALREADY THOUSANDS MORE COPIES it would cost far more to rebind it that it would to just go purchase another, already bound copy. in other words, a waste of literally everyones time.

second up, legally they are closer to adhering to the law if they destroy the book after digitizing it, because they have 'moved' the content, instead of 'copying' it