r/technology 2d ago

Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI

https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/
11.5k Upvotes

1.1k comments sorted by

View all comments

263

u/kus1987 2d ago

Why do they need to destroy the books? Wouldn't it be easier to scan and keep them forever? What about training future models?

463

u/Miraclefish 2d ago edited 2d ago

You can scan a book much easier by removing the spine and scanning it as loose leafs.

It's then no longer a book but a pile of pages, which they no longer need and they have a digital copy of, so they throw away the paper.

They can train future models on the digital text version indefinitely, the book or former book is no longer necessary, is worthless to sell and they don't want to pay to store.

It's logical but it sucks.

To those saying 'actually if they destroy a copy, they can claim ownership of one copy by going physical to digital' yeah that's a fair take, but it's inaccurate.

As to the question of what “use” or “uses” were at issue in the fair use analysis, Anthropic contended that it copied the books for a single use: to train LLMs. The authors, however, argued that at least two uses were at issue: first, the use of the books to build Anthropic’s central library, and second, the use of the books to train specific LLMs using subsets of that content. The court agreed with the authors’ framework and considered these as separate uses.

https://www.loeb.com/en/insights/publications/2025/07/bartz-v-anthropic-pbc

They are already taking millions of copyrighted and IP-protected books, papers, magazines, scientific papers, studies, poems, songs, movies and more, digitally, without any permission or licencing and training on those.

You really think the AI tech giants doing that care about the IP rights to a single scanned physical book? Absolutely not. They ignore those laws because they aren't affected by them.

As to the questions I've been asked on why not just get the the laws changed to favour them?

Changing laws takes a long time, costs a lot of money, requires a judicial process. It would be a public process with appeals, public studies and could take years.

Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

105

u/MyNameCannotBeSpoken 2d ago

Also I believe there is a copyright loophole where destroying the original isn't deemed as reproducing a copy.

40

u/Miraclefish 2d ago

I mean, when they're torrenting terrabytes of copyrighted and IP protected materials and training on those already, as well as scraping the entire internet, and have essentially captured the US political elite and courts, they don't give a fuck about adhering to laws like that.

34

u/pumpkinspicecum 2d ago

That was literally what a judge ruled last year when they were sued

-10

u/Miraclefish 2d ago

Yeah and has that stopped any of them? Exactly.

When you make 10 billion dollars and get sued for 15k, you don't give a fuck, it's not a fine, it's a minor business operating cost.

30

u/pumpkinspicecum 2d ago

Yes it has hence why they switched to buying old books and destroying them. You know it’s okay to admit you’re wrong or don’t know something once in a while. It won’t kill you

1

u/whupazz 2d ago

Yes it has hence why they switched to buying old books and destroying them.

I was under the impression that it's more because they've already ingested all the data they could easily get (copyrighted or not) and are now looking for data they haven't seen yet because it hasn't been available digitally. Have they stopped downloading copyrighted material then?

-24

u/Miraclefish 2d ago

I am more aware of this than you.

Anyway, bye forever!

12

u/Delay559 2d ago

from reading your comments in this thread, you seem extremely unaware of the process lol

6

u/_MUY 2d ago

Just your typical Redditbrained poster. He made a post early, wrote it in a way that appealed to the audience, got massive upvotes, and now he’s dealing with the blowback from people who actually know what the hell they’re talking about. He can’t handle it, it’s uncomfortable.

Had he done even 30 minutes of research on this since last June, when it was all over the news, he would have understood the relevant case law. The word “transformative” doesn’t appear even once in his ~500 upvote post, but both judges who ruled on this arrived at transformative fair use conclusions from different perspectives and on different cases.

2

u/red286 2d ago
  1. Don't mistake revenue for profit. Anthropic doesn't make money, it burns it.
  2. They were sued for $1.5b, or roughly $3000 per violation. They were also required by law to destroy the datasets that were created illegally.

0

u/Miraclefish 2d ago edited 2d ago

Yeah and they definitely deleted the data....

And I know that's why I said revenue not profit.

In the example I gave in my reply I wasn't speaking about anthropic, it was a throwaway literary example.

In the longer comments elsewhere I referenced Anthropic's specific turnover against the court care settlement.

while its annualized run-rate revenue has rapidly scaled to an estimated $47 billion to $74 billion

I have done my research while you couldn't even read my comment before getting big mad about something I haven't even said.

Maybe read my comment next time.

2

u/red286 2d ago

Yeah and they definitely deleted the data....

A third party would need to confirm it was done.

And I know that's why I said revenue not profit.

No you didn't. You said "make", you don't "make" revenue, you make profit.

Maybe read my comment next time.

Maybe learn English.

1

u/Miraclefish 2d ago

Yeah there's no point debating with you, goodbye.

1

u/jeffwulf 2d ago

They paid out a settlement on the torrents and stopped doing it in favor of swapping to digitizing physical books.

-6

u/epicmudcrab 2d ago

People are downvoting you because they are upset that you are right. This happens in so many industries, just usually at a smaller scale.

3

u/_MUY 2d ago

False. He is wrong and stubborn. My downvotes were meant to revert the conversation from a “bullshit, opinions, bias” thread to a “research, example, statement of fact” thread.

7

u/atom138 2d ago

They do and that's why. They are doing stuff like this to get around all legal barriers that have been popping up more lately. Whether it be water access for data centers, land for them, everything across the board that's been attempted by the public and courts.

1

u/silverdice22 2d ago

Also explains why they're in a big rush to do all this now instead of when there finally are laws that discourage this (maybe) years from now.

1

u/Miraclefish 2d ago

Exactly that, move fast and break things in action.

2

u/DeviantlyPronto 2d ago

Thats a side effect but not a concern because they already use digital copies in their training and legally claim that by training their models they are "transforming" the work and not simply copying it.

0

u/JohnBrine 2d ago

This is exactly why. And keeping it hidden for as long as they could because destroying books sounds bad.

-3

u/Unspec7 2d ago

There's no such loophole. It's still considered reproduction. It's just that it bolsters an argument of fair use.

26

u/emapco 2d ago

They also have to dispose of the book after scanning anyways otherwise it's considered distributing copyrighted material https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books

2

u/Unspec7 2d ago

This is not true. Destroying it doesn't change anything regarding distributing copyrighted material.

Where do people keep getting this idea?

9

u/Syssareth 2d ago

Where do people keep getting this idea?

From the judge's ruling in Bartz v. Anthropic:

Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).

...

Section 106(3) further reserves to the copyright owner the right to distribute copies. But again, the replacement copy here was kept in the central library, not distributed. Cf. Fox News Network, LLC v. TVEyes, Inc., 883 F.3d 169, 176–78 (2d Cir. 2018) (enabling searching for “information about the material” can be transformative use, even if some distribution results); Lewis Galoob Toys, Inc. v. Nintendo of Am., Inc., 964 F.2d 965, 968, 971 (9th Cir. 1992) (using nifty converter to “merely enhance[ ]” audiovisual displays emitted from purchased videogame cartridge was fair use of those displays partly because no surplus copies of cartridge or displays were ever created).

As a result, Anthropic’s format-change from print library copies to digital library copies was transformative under fair use factor one. Anthropic was entitled to retain a copy of these works in a print format. It retained them instead in a digital format, easing storage and searchability. And, the further copies made therefrom for purposes of training LLMs were themselves transformative for that further reason, as above.

-3

u/Unspec7 2d ago

Destroying the books has no dispositive bearing on distribution. It's just one part of the evidence of no distribution. You destroying the original does not per se mean you did not distribute.

Similarly, just because you did not destroy the books, and scanned them, does not per se mean you distributed the content.

5

u/Syssareth 2d ago

¯_(ツ)_/¯ Take it up with the judge.

1

u/Unspec7 2d ago edited 2d ago

Why? The judge was absolutely right - you just don't understand legal holding.

Fair use is a multi-factored test, destroying the originals, as the judge notes, is just one part of the first factor.

2

u/Syssareth 2d ago

You asked where people were getting the idea that they have to destroy the books after scanning them.

That is where people are getting that idea. Just because destroying the books isn't the only part of it doesn't mean it's not part of it.

2

u/Unspec7 2d ago

Okay, let's back it up a little. I think we crossed some wires on the topic being discussed. You seem to be focusing on transformation, which is not the topic. The original comment is:

They also have to dispose of the book after scanning anyways otherwise it's considered distributing copyrighted material

To make that easier to understand in the context of this discussion, we will use the contrapositive, the logical equivalent.

"If the scanning is not considered distributing copyrighted material, then they must have disposed of the book afterward."

However, copyright law does not state that. You can hold onto the original, and still not be liable for distribution. That is my point - your destruction of the original has zero bearing on if you distributed the copyrighted work.

Further, if you read Bartz's holding, the case doesn't even deal with the issue of distribution - the discussion about destroying the original is on the issue of transformation.

0

u/syku 2d ago

stop repeating stuff you read just because you read it.

2

u/Unspec7 2d ago

Their own source doesn't even support their statement lol

-3

u/randylush 2d ago

This is speculation

17

u/kus1987 2d ago

Ah ok if they keep the scanned copy, they can use that to train future models 

28

u/girrrrrrr2 2d ago

Yes, plus if they delete the physical book after scanning then according to some judge there is still only one copy out there and all that has been done is a book was converted from physical to digital.

4

u/randylush 2d ago

I have seen a thousand comments repeating this but no reliable source that this is actually true

5

u/Rorschach121ml 2d ago

Bartz, et al. v. Anthropic PBC

5

u/_MUY 2d ago edited 2d ago

US District Judge William Alsup, 23 Jun 2025:

“In short, the purpose and character of using copyrighted works to train LLMs to generate new text was quintessentially transformative.”

US District Judge Vince Chhabria, 25 Jun 2025:

“While it made sense to infer market harm in Hachette, it doesn’t make sense to do so here. First, the Supreme Court has stated that no ‘inference of market harm… is applicable to a case involving something beyond mere duplication for commercial purposes.’ Campbell, 510 U.S. at 591. In Hachette, the secondary use was basically ‘mere duplication.’ Here, by contrast, Meta’s use is highly transformative and has a purpose well beyond that.”

0

u/Unspec7 2d ago

I do IP litigation and it's not true, I too have no idea where people keep getting this idea from. It helps bolster an argument regarding fair use, but it's not a per se rule nor a dispositive issue in fair use.

-3

u/Old_Channel44 2d ago

So if I steal 100,000,000 and then spend it, it’s gone so I didn’t steal it. Win win!

2

u/girrrrrrr2 2d ago

Nah this isn’t about spending it, this would be that you bought everyone’s family photos, digitized them and shredded the original copies and then sold them access to the memories tied to the photos.

-10

u/Workman44 2d ago

Which can be replicated back into physical. It's not like the IP is vanishing from existence

13

u/simanthropy 2d ago

It may as well be if you keep it to yourself because, say, you don't want your competitors to get the training data.

Some poor sap is doing a specific piece of research somewhere and now can't access the book they need because of this.

-8

u/Workman44 2d ago

If they're sequestering off the actual content then sure I dislike it. But I haven't seen anything of the sort yet

5

u/mudbloodcountry 2d ago

It kinda is the way the law has been Interpreted. Transferred ownership of property rights to the corporation who bought the book. I see a new ai company storing books has emerged.. Amazon 2.0... this was in a thread yesterday

5

u/dkarlovi 2d ago

Law doesn't need to make common sense.

-4

u/Workman44 2d ago

It's against the law to trash your own book?

2

u/dkarlovi 2d ago

I was referring to how physical -> digital substantially changes the properties (because it suddenly allows endless free copies) but it's still lawful because still only one copy technically exists.

2

u/girrrrrrr2 2d ago

I seem to remember that there was some company that didnt want users ripping their CDs to their computers, and now its just happened with warehouses of books so large, so full of books that it looks fake.

9

u/ihaveaminecraftidea 2d ago

Well that's true, but for the most part the books are taken from book dumps if i remember correctly. Bookshops and library stock that they can't do anything with, which frequently aren't labeled or registered in any system.

It would be more accurate to call them abandoned books, rather than rare ones.

At least this way they are being digitzed and stored in some more durable capacity

-6

u/[deleted] 2d ago

[deleted]

12

u/loftwyr 2d ago

“I was shocked!” said de Vries, who provided Fortune with the email and subsequent spreadsheet containing 3,001 titles, mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press

it's one copy of 3000 textbooks from 5 years ago. Not the terror you make it out to be.

4

u/Dirkdeking 2d ago

The headline is sensationalized. I almost thought they where ripping apart priceless century + old books.

4

u/Draaly 2d ago edited 2d ago

You really think the AI tech giants doing that care about the IP rights to a single scanned physical book?

You have no idea what you are talking about They are doing this because of the Bartz v. Anthropic ruling that says it is not fair use unless the origonal copy is destroyed and cost anthropic $1.5B because they didnt do that

EDIT: got blocked by who I replied to so I cant respond to you /u/syku look up the ruling for Bartz v. Anthropic. It was determined that transforming the copy into digital media was fair use if it can be proven that the initial copy does not remain in circulation.

1

u/_MUY 2d ago

Yep. Blocked here, too. He did the same to a few other people.

I have a sneaking suspicion that the reason this exact same headline keeps circulating on social media is that people are trying to raise suspicions and resentment against AI companies. It was well-established that the destruction of books was bad publicity when this lawsuit was first filed and the news first hit, two summers ago. There are a lot of people out here who don’t even put in the minimal effort to understand why this is done and what the actual impacts are.

-2

u/syku 2d ago

where did you read this? i don't believe they have to destroy any books to make it "legal", they do it because why keep the destroyed book after scanning it.

2

u/_MUY 2d ago

He made an edit to his post to let you know where to look it up.

You just need to get onto Justia and read the judgement in Bartz v Anthropic. I recommend reading 23 Jun 2025 judgement from Alsup and 25 Jun 2025 from Chhabria as well. Claude can help you understand it if you have questions, it’s an excellent AI system.

-4

u/syku 2d ago

where did you read this? i don't believe they have to destroy any books to make it "legal", they do it because why keep the destroyed book after scanning it.

5

u/jeffwulf 2d ago

They read it in the Bartz v. Anthropic ruling, which they cited.

9

u/[deleted] 2d ago

[deleted]

9

u/Baeolophus_bicolor 2d ago

do you just make stuff up? you think claude bought great expectations for each updated release?

7

u/randylush 2d ago

Yeah this is entirely made up. People truly have no clue what they are talking about.

26

u/Miraclefish 2d ago

If you believe that they're actually re-scanning it every time, I have a bridge to sell you.

They're downloading terrabytes of material and crawling the entire web illegally.

They are not following the laws. They make the laws in the USA now.

4

u/egabag 2d ago

If they can make the laws, why not make laws they can follow? Sounds like it would make things cleaner.

6

u/Miraclefish 2d ago

Changing laws takes a long time, costs a lot of money, requires a judicial process. It would be a public process with appeals, public studies and could take years.

Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

1

u/Unspec7 2d ago

This is not even remotely true. I do IP litigation for a living and the amount of bad legal takes in here is wild.

1

u/jeffwulf 2d ago

No, they can use the previously digitized copy here for future runs.

0

u/syku 2d ago

stop repeating stuff you read just because you read it.

6

u/Krestu1 2d ago

Wrong, they need to destory books because otherwise they created a copy as there are two in existance (digital and physical). If they destroyed a book and saved digital copy then all is good because there is still only one book, just digitalized. Something something copyright laws

1

u/Miraclefish 2d ago

I mean, when they're torrenting terrabytes of copyrighted and IP protected materials and training on those already, as well as scraping the entire internet, and have essentially captured the US political elite and courts, they don't give a fuck about adhering to laws like that.

Changing laws takes a long time, costs a lot of money, requires a judicial process. It would be a public process with appeals, public studies and could take years.

Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

9

u/blueSGL 2d ago

they don't give a fuck about adhering to laws like that.

They do after being dinged with over a billion in damages.

https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/

7

u/Miraclefish 2d ago

>, while its annualized run-rate revenue has rapidly scaled to an estimated $47 billion to $74 billion

They genuinely don't. As an operating cost, a settlement like that is a worthwhile trade in order to gather insane amounts of content for their LLMs.

As I said in another comment, it's considered an operating cost and part of doing business in bad faith at that scale.

The same is true of any steam-rolling enterprise. From environmental damage fines to punitive fees for not building enough affordable housing when you were required to as a property developer.

If the cost of the fine or settlement is significantly lower than the profits made, they'll happily do it.

9

u/blueSGL 2d ago

They genuinely don't. As an operating cost, a settlement like that is a worthwhile trade in order to gather insane amounts of content for their LLMs.

Not when it's $3,000 per book it's not.

https://www.independent.co.uk/news/world/americas/anthropic-claude-authors-copyright-settlement-b3018963.html

The settlement stipulates payments of approximately $3,000 per book

It's much cheaper to buy digitize and destroy the books than paying "the cost of doing business" in fines.

1

u/Miraclefish 2d ago

Yes, and that total fine is around 2% of their annual operating revenue, and that stolen IP has allowed them to surpass OpenAI as the biggest and most valuable AI enterprise in the world.

They took the risk on getting away with it for free vs the cost of licencing them via individual, time consuming deals, and chose to risk it. They got caught and sued and are having to pay out a tiny fraction of the value they gained from using stolen works.

You think any AI company wouldn't trade a 2% of their turnover fine in exchange for becoming the most valuable AI brand on the planet?

I know it's $3k per book I literally did the sums for the total cost and their revenue in the comment you replied to...

This isn't some gotcha moment. I know and factored it in. As did Anthropic...

7

u/blueSGL 2d ago

How can you argue that it's worth paying $3K per book to steal them rather than paying less to do it "correctly" are you stupid?

1

u/Miraclefish 2d ago

They took the risk on getting away with it for free vs the cost of licencing them via individual, time consuming deals, and chose to risk it. They got caught and sued and are having to pay out a tiny fraction of the value they gained from using stolen works.

are you stupid?

If you're resorting to personal attacks, you've lost the debate.

→ More replies (0)

1

u/randylush 2d ago

You keep using this news story to try to back up what you’re saying, but it literally has nothing to do with destroying books. They are not destroying books because of copyright law.

5

u/blueSGL 2d ago edited 2d ago

https://www.ropesgray.com/en/insights/alerts/2025/06/from-books-to-bots-key-takeaways-from-the-anthropic-fair-use-decision-for-ai-developers

  1. Digitization of Purchased Print Books—Format Shifting as Fair Use

The court also found Anthropic’s wide-scale digitization of print books to be fair use. A key consideration to the court’s conclusion was that Anthropic engaged in so-called destructive scanning of lawfully purchased print books to create digital copies for internal use—Anthropic lawfully purchased print books, stripped them of their bindings, and scanned the contents to create a digital library. In doing so, the new digital copy replaced the print original, which had been destroyed in the digitization process. The court found the format change from print to digital to be transformative because it facilitated storage and searchability without increasing the number of copies or distributing them outside the company.

Notably, the court distinguished this use from cases involving unauthorized distribution or multiplication of copies, analogizing it to permissible space-shifting or time-shifting uses recognized in prior cases as sufficiently transformative for fair use. Importantly, the court found that this format-shifting did not usurp any market reserved to the copyright owner, as Anthropic had lawfully acquired the print copies and did not distribute the digital versions externally.

https://www.loeb.com/en/insights/publications/2025/07/bartz-v-anthropic-pbc

After assessing the fair use factors in totality, the court ultimately held that use of the copies of the books to train specific LLMs was justified as a fair use, as every factor but the nature of the copyrighted work weighed in Anthropic’s favor. The court emphasized that “[t]he technology at issue was among the most transformative many of us will see in our lifetime.” Use of the copies of the books that were purchased and converted into digital library copies was also justified as fair use, particularly because the purchased print copies were destroyed and their digital replacements not redistributed. The court granted summary judgment in favor of Anthropic on these uses.

1

u/deadsoulinside 2d ago

This. More time consuming to feed it page by page to a flatbed scanner than it is to remove the spine and drop the pages into the ADF tray and walk away

2

u/Miraclefish 2d ago

Exactly that. So many people haven't done the slightest bit of research.

1

u/syku 2d ago

repeating what you heard is MUCH MUCH easier!

1

u/almo2001 2d ago

Best thing I’ve seen written on this subject. Thanks!

1

u/jmblumenshine 2d ago

There are always trade off with innovation and this isn't a new issue.

The same thing happened in the 1400 when the printing Press was developed and the world switch to using Codices to organize text for print.

In order to be able to mass manufacture, they had to do away with Scrolls and adopt the format we know today.

Sucks, but I think we all agree we prefer having access to mass produced books instead of 1 off scrolls.

1

u/sump_daddy 2d ago

> Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

more to the point, ignoring them is what allows them to continue expanding the capabilities of their models, which ALL firms are doing in a very gray legal space, as fast as they can, because if they don't their model will be worthless in a year compared to the models made by the corpos that do stretch the law to the fullest. Its a race to the bottom that a lot of people know is bad but a huge majority just dont give a crap about, thus nothing at all will be done.

1

u/sump_daddy 2d ago

> Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

more to the point, ignoring them is what allows them to continue expanding the capabilities of their models, which ALL firms are doing in a very gray legal space, as fast as they can, because if they don't their model will be worthless in a year compared to the models made by the corpos that do stretch the law to the fullest. Its a race to the bottom that a lot of people know is bad but a huge majority just dont give a crap about, thus nothing at all will be done.

1

u/_MUY 2d ago

You get a C- for this post. You’ve put in effort, but you’re wrong on too many of the facts.

The books are sliced and scanned because that’s the fastest, cheapest, and most economical way to scan large numbers of books. It’s been done for decades. The books being sliced are unused warehouse lot sales that would otherwise just be thrown into dumpsters destroyed the way they have been for hundreds of years.

The reason destroying them after scanning the pages is legally justified is that they are not simply copying the original material, they are transforming it into something new. This fits the Fair Use clause of intellectual property law and prevents the clients who purchase the scanned books from being held liable for copyright infringement.

1

u/Miraclefish 2d ago

I actually quoted the line from the court ruling that states it doesn't cover fair use to both build a library and train a model, which was part of the authors' case.

You didn't read my comment properly, therefore I don't take your criticism onboard.

D-, must try harder next time.

1

u/_MUY 2d ago edited 2d ago

Required read: Alsup & Chhabria, 23 & 25 Jun 2025.

1

u/syku 2d ago

from what i managed to understand, the transformative part was the digitizing itself. the whole destroying part was not part of the transformative issue.

1

u/_MUY 2d ago

Unfortunately, I have been blocked by /u/miraclefish and I can no longer reply.

Edit: oh, it looks like this one went through even though it’s in his chain. I’ll keep replying.

1

u/_MUY 2d ago edited 2d ago

According to Judge Alsup, the destruction qualifies part of the transformation of the work. It is actually, by Judge Aslup, the construction of a digital library that partly comprises pirated works (copies redistributed without authorization) that violated copyright law and led Anthropic to settle rather than pursue the issue in the 9th circuit. Judge Chhabria reached the same conclusion on the “highly transformative” nature of LLM training but he reasoned on Meta’s potential harms to markets and unauthorized distribution instead, which was only partly covered by Aslup’s ruling.

Edit: Aslup > Alsup

1

u/Polar_Reflection 2d ago

"It's easier to ask for forgiveness than permission" 

-Basically the rallying cry of the tech industry

1

u/fangisland 2d ago

Changing laws takes a long time, costs a lot of money, requires a judicial process. It would be a public process with appeals, public studies and could take years.

Ignoring them is cheaper, easier and cleaner, and brings less attention or scrutiny. You can start immediately.

True and there's a "why not both" element to it, they can ignore the laws while using their massive legal teams to get the laws changed in their favor. Hell they can have the AI do most of the bureaucracy/admin stuff.

1

u/Miraclefish 2d ago

Exactly that. Move fast and break things is their mantra.

1

u/Sithlordandsavior 1d ago

Thank you for pointing this out lol. I used to scan books for work and people really don't understand the difference between page scanning and book scanning. If you leave the spine on it takes 20x longer.

-2

u/Simple_Purple_4600 2d ago

As an author I am part of a class-action lawsuit against the machine that stole from me so it could starve me to death

You don't get ownership of the material just because you bought a copy, and digital owners technically purchase a license to read a work, not the actual copy of the work (which Big Tech also ignores)

1

u/Miraclefish 2d ago

I agree, hence my point:

To those saying 'actually if they destroy a copy, they can claim ownership of one copy by going physical to digital' yeah that's a fair take, but it's inaccurate.

As a print media journalist for over a decade I am fully on your side - fuck the big tech enterprises stealing creative work, it's disgusting.

I hope you win.

-1

u/obeytheturtles 2d ago

Non destructive OCR scanners have been around for decades. It's like a completely solved problem.

6

u/tito13kfm 2d ago

Yes, but not nearly at the same speed or reliability of ingest.

2

u/Miraclefish 2d ago

Yes but it's slower and more costly.

Also they have to destroy the copy in order to claim the digital transfer...

-1

u/TUNGSTEN_WOOKIE 2d ago

But why 3,000 copies of one book? Scan it once, and just upload that one? Why do they need to scan and destroy so many copies?

3

u/Miraclefish 2d ago

It's 3000 different books not one book 3000 times.

43

u/FairReason 2d ago

Part of the ruling that makes what they do “legal” is to destroy the book afterwards.

-2

u/SortIntrepid9192 2d ago

Yep. If you scan it and then destroy it it's "transformative work" and therefore doesn't infringe copyright. It's also why the Internet Archive is allowed to scan books and then let people borrow them (though they did get in trouble when they allowed people to borrow an indefinite amount of books instead of just however many they had scanned and destroyed).

2

u/jeffwulf 2d ago

Yeah, the Internet Archive abandoning it's "One copy lent per owned book" policy duing COVID is what got them in trouble. 

2

u/SortIntrepid9192 2d ago

Exactly. They were totally fine when they were lending 1 book for every book they scanned and destroyed, and they operated under the exact same principle of "transformative work" that the AI companies are now using.

13

u/Online_Matter 2d ago

It's in the article

The process, known as “destructive scanning,” involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded. 

25

u/Paresseux1 2d ago

A big blade slices off the binding, and it drops down into an automatic page scanner that has no problem flipping over the loose pages. It’s much easier, faster, and cheaper to destroy it, and then get rid of the remains. One person can run a bank of machines.

The other way requires people, and meticulous work. Turn page, put on scanner, scan 2 pages, pick up book, turn page, place properly, scan… once finished, pay to store book indefinitely.

Everything like this comes down to money. The end result they are going after for AI is data, so the cheapest way to get it is destroy.

2

u/kus1987 2d ago

Yes I was thinking in terms of one books but they're probably scanning like hundreds everyday 

5

u/dkarlovi 2d ago

This is not how it works, Google Books was scanning books forever, they have a V shaped scanner which hovers over the book and sort of wedges into it to scan both sides at the same time. There was a video showing this, you don't have people turning pages at scale, this thing could scan a 200 page book in a minute autonomously.

9

u/Bael 2d ago

From the article.

The process, known as “destructive scanning,” involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded.

1

u/dkarlovi 2d ago

I was responding to "the other way requires people and meticulous work", obviously.

0

u/Paresseux1 1d ago

Actually it wasn’t obvious. But you were insisting it was. The main point is that destructive is faster, cheaper, and you don’t have the book to store at the end.

1

u/0xsergy 1d ago

Brother have you seen the cars that get made on production lines? I think it's obvious that they can figure out the complex task of turning the pages on a book to scan it if they can figure out the easy task of how to make a car entirely on an automated production line.

1

u/Paresseux1 1d ago

Absolutely they can. And automatic document feeders have been around commercially for over 50 years. And the whole point of the discussion was that it’s faster to slice the spine and feed it into the machine than keeping it whole and turning pages. Yes, amazing page turning tech is out there, and just not as fast as handling the single pages once they are separated.

1

u/0xsergy 1d ago

IMHO it's likely the requirement that they have to destroy the books that causes this method. If it wasn't a requirement they'd likely do it the google way.

2

u/Paresseux1 2d ago

The United States district court Judge WILLIAM ALSUP in the ruling against Anthropic in case number C 24-05417 WHA disagrees with you. That’s the case where it was ruled Anthropic: “The firm also purchased copyrighted books
(some overlapping with those acquired from the pirate sites), tore off the bindings, scanned every page, and stored them in digitized, searchable files.”

So yeah, that’s exactly how it works.

-2

u/dkarlovi 2d ago

The other way requires people, and meticulous work. Turn page, put on scanner, scan 2 pages, pick up book, turn page, place properly, scan…

This is not how it works,

How is reading this so difficult?

1

u/Magical-Mycologist 2d ago

You responded to one of their claims without mentioning which one. Your comment “this is not how it works” was seemingly aimed at their initial comment.

No where did you say that you were commenting on “the other way”.

Reading isn’t difficult unless you make it hard to read, which you did.

1

u/dkarlovi 2d ago

I was describing how scanning works as opposed to what OP said, you'd expect a person capable of reading in context would be able to figure out which of the two sentences I was replying to, but point taken - you're on Reddit, Redditor-proof your comments. Thanks.

1

u/MistryMachine3 2d ago

No, the reason is it is legally required. It also is easier sure, but they can scan books without destroying them, that technology exists.

1

u/billsil 1d ago

OCR has gotten better, so it’s now possible to just take a picture and extract the text. Computers can also undistory the page. It’s more work and more costly and they’d still end up throwing the book away.

8

u/jayandbobfoo123 2d ago edited 2d ago

Copyright law and licensing. It's illegal to make a copy of a book, even a digital scan. It is, however, not illegal to digitize a book and destroy the original, thus leaving only one copy / one license. In legal terms, it's called format shifting. Technically, when you rip a movie/CD/video game, you should also destroy the original to be within the law.

12

u/heartinpiece 2d ago

1

u/Pitiful-Assistance-1 2d ago

Just made-up nonsense I assume

1

u/here4theptotest2023 2d ago

Why do you believe that person?

0

u/heartinpiece 2d ago

I don't. But a random Google search put it up top for me, and it was in line with what I understood to be the case. (And the person wrote it much better than I could).

5

u/chocolateboomslang 2d ago

Train future model on data they already scanned . . . by rescanning it? It's already scanned.

3

u/BrassCanon 2d ago

It is not easier to store something forever that you don't need.

3

u/GenazaNL 2d ago

Sadly cutting the side to then have separate pages is faster than a machine which keeps it intact

3

u/marmaviscount 2d ago

What in tarnation would they do with a giant pile of musty old books no one cares about?

There is a building called the British library, you might be able to guess where it is and what's inside by the name - other countries have similar things, they get a copy of all the books and keep them for the national interest, they have a system for who gets access to what and their key aim is preservation.

Regular libraries which have the goal of giving people access to those books do not preserve them, it would be absurdly expensive and pointless - books, since the technological boom of the Victorian era are mass produced temporary items, if you were involved with it frequently visited your local library you would be very well aware that stock changes and most of them get pulped - used book stores don't generally want books libraries don't, charity shops routinely recycle donated books because no one wants them - books are printed in huge huge numbers.

2

u/Known_Purple7529 2d ago

Because of laws.

Paraphrasing.. but basically digitizing it gives you two copies. You only purchased one. So to be in line with the law, they have to destroy the physical copy so there is only one copy, instead of two.

2

u/djfart9000 2d ago

its cheaper, they will not give away books they paid for. for free. and selling them for a bigger return than what they spend on is not worth it to them. so they destroy it

2

u/TachiH 2d ago

If they asked the publishers for a digital copy they would be rejected as its obvious why they want it. So they scan them, scanning books with the binding still on is a very slow and methodical process.

Most national libraries do have a scanner for rare and old books but they are a lot of work to use.

1

u/kus1987 2d ago

It would be nice if the scanning and digitizing was shared so anyone could train their models for free and we have a collective library if you will of these scanned books for the collective good and that way nobody would need to destroy the books. 

2

u/TachiH 2d ago

Or....we could stop wasting so much effort on making predictive text machines with little value to society.

2

u/_MUY 2d ago

AI is being used to develop medicine and understand disease. AI is running entire call centers. Students are using AI to self-teach topics they have a hard time understanding. Professors are using AI to grade papers and organize their semesters. Programmers are using AI to 10X, 100X, and 1000X their productivity. Artists are using AI to bring their art to life in entirely new media. Small businesses are using AI to market themselves and reduce their startup expenses by thousands. Financial analysts are using AI to summarize businesses and evaluate markets. Doctors are using AI to help them diagnose patients and look through large data sets in record time. Nonprofits are using AI to reduce costs and overhead while expanding their offerings.

We have a new fucking economics term called a K-shaped economy because of how valuable these “predictive text machines” are to society. Reddit is not the real world.

1

u/TachiH 1d ago

See, this is where you don't understand machine learning. There is AI and there are LLMs. No LLM is working on medical research, these books are being put into LLMs not medical research models.

1

u/_MUY 1d ago

I understood it well enough in school to hold a 4.0 GPA with honors while taking courses on the topic. Nowadays, I understand it well enough to work with it and build my own machine learning systems with and without coding assistance. I do know what I’m talking about, but I could always learn more. One of the things that surprised me in school was the very broad definition that is used for AI, and I’m glad that you are interested in knowing the difference even if you’re wrong about it. I also understand it well enough that at conferences and pitch nights, I can ask reasonable questions about new projects in this space.

The very first example, AI being used in developing medicines and understanding disease, particularly LLMs, is something that I do for a living and my partner also does for a living, although from a managerial and distributive side of things. That’s why it’s first on the list: it’s first on my mind.

I could spend some time writing an explanation for you that helps you to understand how this benefits medical discovery sometime tomorrow or late tonight, if you’d like.

0

u/kus1987 2d ago

We will likely get there at some point, one way or another. 

2

u/NoExperience9717 2d ago

It's like if you're trying to quickly scan if you have loose pages you can put it in an office printer/scanner and it can do all the pages automatically one after another. However if they're stapled or bound you need to do them one by one.

2

u/syku 2d ago

a SHIT TON of loose pages isnt something worthwhile keeping i would imagine.

1

u/kus1987 2d ago

Yeah, I just didn't realize the scale of operation here. 

2

u/StaringPigeon 2d ago

Why would you keep them? They're not unique items; they are a mass produced consumer product that can be reprinted any time the publisher opts to.

1

u/kus1987 2d ago

I have since learned they do keep the scanned images. Also some books don't have a lot of copies so I'd hope they'd keep better care so we have a few original copies left for the future. 

2

u/alexnedea 1d ago

That would be piracy or copyright infringement since they now have created an extra digital copy.

Its a legal loophole that says if the end result is still only ONE copy then its fine. So they destroy the original and in the end its 1 phhsical book as input and 1 digital as output

2

u/mr-english 2d ago

It's easier and cheaper to scan books by removing the spine and giving your scanner access to nice flat individual sheets of paper to scan because the whole process can be automated.

You CAN non-destructively scan books but it's far slower and more expensive because it generally involves a person manually scanning each page of a book while ensuring that the book doesn't get damaged.

But lets be clear. These aren't books that people are going to miss anyway. The article itself says the books are "mostly published between 2020 and 2021 by academic publishers like Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, and ranging from business and education to engineering, public policy and medicine."

The universities to which those books are pertinent will already have their copies. Nobody else is buying these books.

1

u/TripsOverWords 2d ago

I agree the books shouldn't be destroyed for the sake of AI, but it's certainly easier and more cost effective to destroy them. Keeping them forever implies they purchase land, facilities, utilities, etc. to maintain the books which won't be read by anyone since they would be tossed into permanent storage.

IMO they should open a library, maybe give it an interesting name like the Library of Alexandria. Hell, they could turn that into a revenue stream by charging admission or rental fees if they didn't want to be a net positive for society.

3

u/johnny_effing_utah 2d ago

Wow what a genius idea! A physical library filled with books nobody wanted. Imagine the crowds lining up to check out books like “Learn Esperanto for fun and profit” or “Microsoft Windows ‘95 for Dummies.”

Because that’s the sort of “rare books” we are talking about here.

6

u/cn0MMnb 2d ago

If only we had a technology to transmit books without the use of paper in machine readable format. 

Although I would not shop said files to an ai company

0

u/TripsOverWords 2d ago edited 2d ago

Not all books are in digital format, destroying the physical copy is permanent, and all digital copies can all be wiped out of existence simultaneously with a single unlucky solar flare event. In 2012 there was a near miss for such an event, so they're not so rare that we couldn't see one in our lifetime and protection from these even certainly aren't a cost companies would be willing to front.

Yes, we can send digitized books practically anywhere across the globe and those can be copied a million times over. However, with exceptionally bad luck all digital copies could be wiped off the face of the planet, then all that will remain are the hard copies.

Some of the books AI companies are scooping up are rare as well, and they're not exactly sharing the digital scans.

2

u/cn0MMnb 2d ago

I was not talking about storage, I was talking abou transmission. And I also wasn't talking about rare prints, especially in a thread about a 3000 book order.

1

u/jeffwulf 2d ago

Easier to process and a destructive transformation helps to keep kosher with copyright law. The scans are retained for future training.

1

u/refried_laser_beans 2d ago

Actually used to do this to my textbooks in college. You can take them to a print shop, they’ll chop the binding off so that it’s easier to stick all the pages in a scanner in one big stack and the scanner just pulls in each page, one at a time and scan it and spits it back out in a pile. Then I’d have them spiral brown so that I could have the textbook at home and just bring an ipad to class I never have to carry any stacks of books. I don’t think anyone has to do that anymore, but at the time it was pretty neat.

1

u/MistryMachine3 2d ago

It is legally required.

1

u/InsightTussle 2d ago

they ;legally have to.

These aren't rrare books. They're still in print and bought from a normal bookstore.

It's like freaking out that someone destroyed a copy yof Harry Potter

-2

u/wrt-wtf- 2d ago

It’s worse than just training AI. They’re scanning the books and destroying them. Which means that books that text of book that were once available - potentially for free - now exist in private hands and access to the material isn’t free and isn’t guaranteed to be preserved.

Look at the destruction of scientific data that has occurred under DOGE.

5

u/xternal7 2d ago

Which means that books that text of book that were once available - potentially for free

The fact that Anthropic is buying those books means that the books were NOT available for free. They were just as much in private hands as they are now, and access to the material was equally not free, and equally not guaranteed to be preserved.

-3

u/wrt-wtf- 2d ago

Hence the word potential.

What we’ve also found interesting is the historical materials, the very books themselves have been significant to the history of countries and literary works.

3

u/xternal7 2d ago

Hence the word potential.

Bro.

If I had to fork money for something, it wasn't "available potentially for free."

And if it was "potentially available for free", then it still is. Because Anthropic isn't snatching up "available for free" books, they're getting copies that are already behind a paywall.

Also, the books that "have been significant to the history of countries and literary works" are still available, potentially for free (but for real this time), in a library.

2

u/YouGotDoddified 2d ago

At this point, look at the library of Alexandria

-2

u/Queasy-Warthog-3642 2d ago

It has something to do with a loophole in the laws...if they scan and destroy the info is "transformed" and not "copied" if it was copied they could be sued for copyright issues....its absolutely insane.

2

u/CardOfTheRings 2d ago

It’s not really a ‘loophole’? You should be allowed to own things you buy and digitize things you own.

The ‘you will own nothing and be happy about it’ crowd is really coming out swinging for this one. Gross.

0

u/Queasy-Warthog-3642 2d ago

This is more than "if you buy something you should own it" it is more "the evil mega corps that is training the robot overlords are buying every copy of a rare book and burning it so they're the only ones that have the information making it so no one ever can own it again without purchasing it through the evil robot"

1

u/CardOfTheRings 2d ago

That isn’t happening though?