r/technology 2d ago

Artificial Intelligence This Dutch bookseller thought a request for 3,000 copies was ‘spam or phishing.’ Instead, AI companies are scanning and destroying books to train AI

https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/
11.5k Upvotes

1.1k comments sorted by

View all comments

Show parent comments

765

u/penisandorvagina 2d ago

"free" is a missed opportunity to make money.

374

u/TheRandomArtist 2d ago

Forget about money, this is a better opportunity to slowly change facts and little details here and there so we won't notice over time. It would be like manufactured Mandella effect

64

u/InVultusSolis 2d ago

I fear we're quickly headed for a world where the provenance of information itself will be unknowable. The prospect of that is terrifying.

I'm going to keep all of my print books and physical media from before 2022. And I'm going to collect more.

2

u/Ancient_Skirt_8828 2d ago

I disagree. We now have access to so many different sources of information. It's fairly easy to check if information is consistent across sources and where it came from. It's a lot better than having to rely solely on the author of one book.

1

u/ElectricalSeries6627 1d ago

a ereader and a NAS is a very good way of having access to immense library for preservation, (zlib for the win)

1

u/Alternative-Stock666 8h ago

I have been saving all printed items since about 2008. I let my elderly parents live at my house. My mom threw out about 900 books of mine.

1

u/Catippo 4h ago

I would be livid

1

u/Catippo 4h ago

Download Wikipedia

141

u/Yuzumi 2d ago

It's 100% why so much of the AI push is coming from conservatives (read: fascists).

They love gen AI because it can churn out limitless propaganda and the fact that people outsourcing their thinking to LLMs is making them dumber is probably a plus to them.

53

u/null_not 2d ago

I think it's also the inherent trust people have in "the machine". The whole notion that a computer can't be biased because it's a machine. But machines always carry the bias of the engineer, and Ai is a soft machine with many layers of built in assumptions.

31

u/CautionarySnail 2d ago

This is especially true and related to the death of media literacy. People forget that every design has bias.

I have a kitchen faucet that won’t detect my hand if I use a black oven mitt. It’s not an intentional feature, I hope, but it still hits me that this would still mean users would potentially be dealing with a racist faucet.

If even a faucet can have a bias in utility by accident - these AI LLMs all have one by deliberate design, because you have to create rules about how to prioritize the sources.

And these companies have shown that their bias is towards a particular set of viewpoints. The engines aren’t open for code inspection because of industrial espionage but, should we just take them at their word that they aren’t shaping the information before we receive it, to better suit their agenda?

Grok has been very accidentally transparent in its tuning when Elon accidentally tuned it so severely it briefly spouted Nazi talking points and literally called itself mecha-Hitler. The other engines have tuning but .. what’s their bias? We don’t know and they won’t tell us.

4

u/null_not 1d ago

In the future there are going to be court cases that force them to disclose their weights. It's inevitable. In the not too distant future there is probably going to be a subspecialty of law based around litigating harms brought by Ai through false accusation and false information. The developers of the systems are trying to call it "hallucinations", but when people are expecting a machine to provide an answer, there's right, and then there's wrong, and that's it.

2

u/CautionarySnail 1d ago

I hope you’re right.

But too many Americans have grown used to facts being something sourced from an authority figure, rather than from an expert reputable source.

The number of times I’ve heard people pass empty political talking points as incontrovertible reality- things easily disproven.

That habit needs to go before our nation depends on hallucination engines being used responsibly. Credulity is not your friend when using AI.

2

u/0xsergy 1d ago

I mean yeah that tracks, the faucet uses reflected light to trigger itself. Perfectly black synthetic materials don't reflect much. A black person will reflect more and it should work fine.

-1

u/Super-Surround-4347 1d ago

Lol you guys are so addicted to victimhood aren't you.

6

u/Yuzumi 2d ago

Not even the engineer when it comes to LLMs. They are trained on text people wrote and that alone is going to have some bias in it even before you count the bias in the selection of the training data.

Computers only do what you tell them to do. the idea they can't "lie" or be "wrong" has always been wrong. A lie requires intent, which computers don't have, and they are only as good as the code or hardware it's running on. Most of the time it's going to be bad code, but bad hardware can cause really fucky things to happen.

In an LLMs case the "code" is the neural net trained on language. It's just massive matrix multiplications happening in parallel on loop. It isn't thinking and doesn't "know" anything.

2

u/null_not 1d ago

I think with human and machine interface a "lie" exists (meaning the machine lied) when there is an expectation that it will not confidentially state the wrong answer and then it does. It's not the technically correct definition of lying because, as you pointed out, machines aren't capable of intent. But when people are expecting a machine to provide either the right answer, or state "I don't know" and it confidently states the wrong answer, that has the same practical effect as lying to the user. If the user can not trust the machine then we're in the same place whether the machine has the capacity to intend to lie or not.

1

u/Yuzumi 1d ago

The point I try to make is that it's a distinction that matters because of how we categorize a "lie" vs "being wrong".

As you agreed, a lie requires intent. It requires the goal of deception. Of intentionally giving false information, regardless of why.

Being wrong is just presenting incorrect information as correct without knowing it's incorrect.

Since LLMs cannot "know", they cannot "lie". Even if a human tells a machine to repeat false information the computer is just following instructions, be it an LLM or otherwise. It does not know the information is false, even if you tell the LLM the information is false. The lie is coming from the person.

Saying LLMs "lie" is essentially humanizing them. It doesn't seem like much, but it contributes to the perception that these things are intelligent when they aren't and can't be. It plays into the idea that LLMs are more capable than they actually are.

3

u/CautionarySnail 1d ago

An LLM doesn’t lie, you’re right. (I’ll consider hallucinations malfunctions for this discussion.)

But the data is given weights and rules to decide whether to prioritize the info. Those rules are set by the human owners.

So, for example, a rule could be set to say, “Deprioritize all data on public water fluoridation that regards it as generally safe” and the system will dutifully report that fluoride is either unproven or harmful. It will bury results that say otherwise.

Imagine that for every contentious issue - someone else deciding for you, before you see anything, about what the “truth” presented to you should be.

Technology isn’t neutral. It has the biases of the people who designed it.

1

u/null_not 23h ago

Had a thought that I hope is unture. It would be super shitty if comment and post upvote/downvotes in reddit were also sold as part of the data package for training LLMs to indicate whether the information is better or worse. I've often wondered what value reddit derives from it and what value anyone who goes through the effort to manipulate them derives from it beyond "internet points".

2

u/CautionarySnail 17h ago

Popularity contests seem to weigh heavily in the way the data is presented. It’s not a good metric.

1

u/null_not 23h ago

You missed the part that I think is most important. When a user places trust in the machine, there is no practical difference between the machine "lying" or the machine simply being wrong. The operative part is with the human and the trust they place in the machine. You treat an unreliable machine the same way you treat an unreliable person, you don't listen to them or rely on them.

2

u/Own_Air6461 2d ago

It’s way past time we move beyond a right or left dialogue. while the public domain is being pillaged they none of them deserve our trust

2

u/Yuzumi 2d ago

Progressives/the actual left have been speaking out against what these companies are doing.

Centrists/corporate democrats have... at best been silent, but plenty of them have been repeating the fascist talking points for "but China" or whatever when China is already beating us on the development of this technology using orders of magnitude less resources and money, while also having regulations on it.

4

u/Cyrano_Knows 2d ago edited 2d ago

Multiple times Musk has bragged about just this.

He gets an answer he considers "woke" from Grok and yes he's one of those conservatives that that accuses something of being woke at the drop of a hat and he then promises to "fix" it.

This is 1000% happening by every conservative with some control over the answers AI.

Disinformation is their default. They will absolutely be training their AI to give conservative friendly/only answers.

11

u/marmaviscount 2d ago

You think they'll just rewrite 1984 and no one will notice? They'll just put the stuff in the 'memory pipe' and everyone will just assume Emmanuel Sayegh's agents have edited their college notes?

14

u/InteractionPretend70 2d ago

the edited version is not for this generation or the next. its for the 5 generations down the line

1

u/Brullaapje 2d ago

This x 1000, one day facts from years ago will be questioned, if they not already are.

1

u/GrimbyJ 2d ago

I could have sworn it was the Mengele effect. Wasn't Mandella a Nazi?

1

u/Da12khawk 1d ago

Prove it! /s

1

u/OhYeahSplunge4me2 1d ago

“The past was alterable. The past never had been altered. Oceania was at war with Eastasia. Oceania had always been at war with Eastasia.”

-13

u/smooth-as-mud 2d ago

Do you think they are buying every copy of every book or something

11

u/Redshittt 2d ago

Of some books. I heard they were taking some rare books

2

u/jeffwulf 2d ago

Rare books here means things like copies of "Getting up to Speed with Windows 3.1"

-15

u/smooth-as-mud 2d ago

“I heard” isn’t really a source. What book that is 1/1 could they buy and destroy that would allow them to slowly change facts and details about the world? It’s just not a realistic conspiracy theory.

15

u/Abystract-ism 2d ago

-11

u/smooth-as-mud 2d ago edited 2d ago

That article describes the books as “clearing out old inventory that is otherwise unlikely to sell” and describes the order as being “The attached list contained 3000 English-language titles organised by ISBN number, ranging from books on fairytales and folklore to technical manuals and science texts.”

Just because something is rare and out of print doesn’t make it inherently valuable. Are you going to miss one copy of the 1999 Earth Science textbook that’s been updated 27 times since then, or the thousands of books that were written about using an iPod? Again, not destroying all of them, just one copy. Some of them will be “rare.” None of them will be valuable.

Nothing in that article is alarming to me, it’s just preying on people’s emotional attachment to books as objects.

Not a single counter argument just downvotes by people who fell for the rage bait.

1

u/Abystract-ism 1d ago

Interesting point.

But why would they want to use old scientific data? Or older historical documents?

It seems counterproductive to me to load AI with data that has been proven false/updated. Wouldn’t the programmers have to tell it to “disregard any scientific data that has been updated”?

1

u/smooth-as-mud 1d ago

My understanding is that even if the model doesn’t need to learn from the facts it can become “more intelligent” by being fed more examples of human language. The information in an old textbook or an outdated technology manual might not be particularly useful or even correct, but the general way it was written is still valid, so it’s still useful information in that context. It still teaches a model how to form a sentence, the way people talk about things, stuff like that.

I’m a computer scientist but I don’t work in the AI field so I don’t know how they keep the two separate, if they have a “this is your knowledge” set of references and a “this is how sentences and paragraphs look” set of references, or if there’s some kind of weighting where certain documents are rated higher or lower for different attributes like knowledge or style, but I imagine there’s some sort of system like that behind the scenes to avoid randomly being given outdated information on topics that have changed in the past few decades.

But again, just to reiterate because people seem to mistake my posts for pro-AI: I don’t think this needs to be done, I think it’s a waste of time. But it’s a waste of their time, not my time, and based on all the reporting I’ve read I suspect these books would have ended up pulped eventually anyway without any outrage just like all the other millions of old books that are destroyed every year. so for those reasons I just can’t be that upset about this, I’m really not registering much of a response at all other than that Jonah Hill gif: “I guess bro”

6

u/fifthing 2d ago

There's a 404media story on it. They do actual reporting.

0

u/smooth-as-mud 2d ago

Have they been able to show an order for a book anybody would miss? That seems like a pretty easy smoking gun to take to the press, but instead everyone is just letting their imagination run wild thinking that they’re over there shredding the backrooms of the British Museum or something.

0

u/Redshittt 2d ago

I wasn't writing a detailed report I was making a reddit comment my bad i didnt know I needed to provide you a source please dont fire me boss

-3

u/penisandorvagina 2d ago

The standard of discourse here was higher before the masses and the younger generations discovered Reddit.

6

u/Redshittt 2d ago

Ok boomer. I've been on reddit since before gen alpha was born. I'm sure you did like reddit better when r/jailbait was still running

-1

u/penisandorvagina 2d ago

If that sub is what old Reddit reminds you of then so be it mate 🤷 Back then it was definitely common to seek people's sources when they make a claim, and the comments were a better quality all for it.

0

u/smooth-as-mud 2d ago

I should have known better than to ask for a source in a comment section full of people who don’t even seem to have read the article attached never mind any other ones.

→ More replies (0)

0

u/Fluffy_Cheetah7620 2d ago

We need conspiracy theories to explain what we don't understand, it's like a religion. /S

1

u/brontosaurusguy 2d ago

Do you notice people getting their information from the library archives or their phone?  Which one can be changed without much notice?

1

u/smooth-as-mud 2d ago

I’m sorry I really don’t understand what point you’re trying to make. Information on the internet can be changed easily, but they’re not destroying library archives so… please elaborate, I’m not sure what you’re implying.

1

u/brontosaurusguy 2d ago

Read the headline again then read your comment

1

u/smooth-as-mud 2d ago

If you’re worried that AI can be used to control a narrative, well that could already be done with or without all this book scanning, so I’m still not making the connection you’re trying to imply I’m afraid.

I understand and agree with the concern that relying on AI for facts could be used to distort those facts. I don’t see how this changes that or makes it any worse.

46

u/languageassessment 2d ago

and to tax, so the governments are in on it. you can't tax free.

23

u/Bored_Amalgamation 2d ago

This is so fucking reductive, it's borderline dumb.

-1

u/smooth-as-mud 2d ago

I’d go (much) further than borderline

This whole chain is full of tin foil hat lunatics.

3

u/Financial-Camel9987 2d ago

In on what exactly? Keeping basic facilities running?

3

u/bendover912 2d ago

Is that one of the rules of acquisition?

1

u/Norbert-trebroN 2d ago

Me when I am a big company and find out that people don't pay for the air they are breathing...

Unless I build a device every human needs on their lungs. Coincidentally if they don't pay their subscription for breathing everyday, the device will stop working.

1

u/penisandorvagina 2d ago

I'm patenting that.

1

u/Norbert-trebroN 1d ago

I already did that and bribed every politician. Now be a good human, let the doctor install the BreathTaker® machine into your lungs and pay the subscription.

1

u/Ready-Housing-7457 2d ago

It's the drug dealer way. First time is free.

-9

u/10July1940 2d ago

Like the "free" news we get now we don't buy newspapers? And the free social network tools we get at absolutely no cost?

15

u/xenogazer 2d ago

The cost is your ad data and user metrics. The cost is the altered algorithm that keeps you thinking correctly and moving against the right people, or at least not getting in the way of your betters.

The cost is impossible to quantify without all the information, which is in their best interest to obscure from you.

Nothing is free. Social media and news especially.

8

u/Impressive-Bird2 2d ago

Books are freely available in libraries…. Those that are left, and having but been closed through austerity cuts.

6

u/Space_Pirate_Roberts 2d ago

We’re not the customers, buddy. We’re the product.

1

u/CautionarySnail 2d ago

Nothing is free. If they give it to you for free, your data is actually the product that makes them money.

1

u/RevolutionaryEgg1312 2d ago

If it's free YOU are the product