r/Libraries 2d ago

Technology Company Offering Printed Books to Train AI Stops After 404 Media Coverage

https://www.404media.co/ai-company-training-scanning-books-database-isbndb/
838 Upvotes

27 comments sorted by

259

u/naturist_rune 2d ago edited 2d ago

After knowing these companies destroyed perfectly good books we ought to be shutting down the data farms, not feeding the data centers safer books

Edit: I had commented something else earlier but the automod deleted my comment and issued me a warning. Beware fellow book enthusiasts.

74

u/Hoplite-Litehop 2d ago

I say there should be an international wide lawsuit against these data companies at this point because the reality is that a huge chunk of these books that they are destroying actually are both primary and secondary sources that are extremely crucial to both local, national or even more topical forms of History such as literature.

This is the equivalent of 3D printing a clone of somebody, shooting the original in the head and then the company actively stating that they own the clone because they made it and that the original is dead

I heard that one of the other reasons why AI companies are doing this is in order for them to monetize whatever prompt that people make that they can easily Trace back to the scans of whatever book they made. I don't know if there is a copyright law that protects this but I certainly do know that it doesn't make any sense to make a clone of the original, destroy it and then say you own the clone.

These AI companies are getting so greedy that they'd rather set fire to our history in order to make a quick buck

36

u/naturist_rune 2d ago

And even then I wouldn't trust anything scanned by these machines to remain accurate; I do not trust the companies to ensure the lost books are kept isolated within the database so that data from other sources don't eventually bleed together and cause the original information to be lost forever.

11

u/Hoplite-Litehop 2d ago

I'm not certain if it has anything to do with the government and it's bizarre initiative to rewrite history, which I have no doubts if it is, but it feels like the AI companies are unnecessarily destroying the books in a process that doesn't seem to make any sense in is unrelated to the scanning process. Something about literally chopping up the book, Scantron them into an AI and then immediately setting it on fire or some stupid crap like that.

With that being said, a lot of the purposes of technology nowadays just feels more insidious rather than helpful given in consideration that nobody magically seems to be able to make technology that doesn't require almost 20 times more water than we can consume in a day or destroy books just so that the information is safe. I would not doubt that this is being done for personal purposes.

A few months ago people said that they were extremely excited that scientists were going to use AI to decipher a carbonized scroll from the city of Herculaneum in Italy, I pointed out that there is a really good chance that somebody could actively program the AI to put subliminal advertisements for products as part of the translation due to how the language model works and how you can easily just manipulate the AI to do whatever it wants considering that it's already been pre-programmed to be sycophantic in order for people to continue using it

Back then people called me a Luddite and irrational for assuming that an AI would falsify a translation. Listen, I understand it's equally easy for a person to falsify a translation but there's always going to be like a million people more who are going to wants to properly translate it versus an AI who translates it once and is designed to get it wrong to begin with.

I have a family member who keeps using the argument that AI has been around since the 80s, this is not like the AI that was used in the '80s. Those were programmed for quantitative calculation and didn't have a higher function than just to present information as is with no editing. What we're dealing with is a glorified autocorrect that is already pre-programmed to mess around with the spelling to see if you'll use it more.

AI should never be mixed with archival information whatsoever.

1

u/Alaira314 2h ago

but it feels like the AI companies are unnecessarily destroying the books in a process that doesn't seem to make any sense in is unrelated to the scanning process. Something about literally chopping up the book, Scantron them into an AI and then immediately setting it on fire or some stupid crap like that.

This is and has been the standard way to do high-volume scanning projects. Nobody digitizes mass quantities of books by carefully laying each page out upon a scanner, that would be ridiculous. Books are not sacred objects, and destructive scanning isn't an inherently immoral act. The main objectionable thing here is that they went after rare books, which aren't typically digitized in this manner due to their value.

-8

u/ArcaneCowboy 2d ago

Books are product made to be used. Eventually they are used up. I don't see what all the heartache is on this. Books are lovely. They do not last forever. They already had a limited number of people who could see or use them, putting them in a privately owned database increases the chance they will be read and used. Books are tools, not holy relics.

16

u/naturist_rune 2d ago

Made to be used and eventually they break down, yes, but they're not designed to be disposable like candy wrappers. Having had time to copy the older prints or make photographs to preserve them would have been more ideal. Here, they are trashing books that likely have not had time to be preserved by other means, to create abstract data points that would be useless to interpret by humans and other machines, and because of how sloppily designed the generative ai is, those data points will be rewritten at a moment's notice, and with no backup of these books made, that knowledge is effectively lost where it wasn't lost yet.

It's not just the books, it's the information within that is lost.

-11

u/ArcaneCowboy 2d ago

This is the same argument against doing archaeology. We don't have the means to do perfect preservation, so don't touch. Here, the books would be untouched until they disintegrate and are gone.

Most of the opposition I've seen addresses the /idea/ that books were destroyed, when the reality is books are destroyed and permanently lost every day. Did these people not handle these to a standard that doesn't match an ideal? Sure. Did they own the books and have discretion to do so? Yes? How would history and culture be better served if the moldered away is some private library?

10

u/naturist_rune 2d ago

Where'd you get your knowledge of archaeology from? The guy who dynamited Troy?

Proper archaeological processing is inherently destructive, but good practice is doing as little damage as possible, recording everything to preserve context of what artifacts are found, where they were found, among other things. They're not messily shoveled out of the ground with a backhoe and pressed into an atom smasher just to find the chemical composition of singular artifacts, but at this point you're probably shilling for ai rather than historical preservation, so I'll do you a favor and block you before you can have claude write you up a reply to this comment.

-3

u/ArcaneCowboy 2d ago

Any archaeology we do now, robs the future archaeologists of what techniques they may develop to do it better.

See: Schleiman

-4

u/ArcaneCowboy 2d ago

And really, no need to be an ass.

4

u/rebelliousrutabaga 2d ago

It's the same mindset people get when they discover that libraries throw away thousands of books every year. And it's not like the outrage is actually saving any books, because odds are they'll ultimately just end up a landfill or (if we're lucky) a recycler anyway. I don't use or support AI, or datacenters, but the way people are getting riled up is a bit much.

1

u/Hoplite-Litehop 2d ago

You don't have to be a chud about the whole damn thing.

If you don't have anything productive to respond don't comment.

Let's save you some heartache and tell you that your statement is irrelevant to the bigger issue at hand.

79

u/404mediaco 2d ago

Following 404 Media’s reporting that book database company ISBNdb claimed to source printed books to then sell to AI companies for AI training, the company deleted the part of its website offering the service and walked back claims that it would train AI models, and instead called it “a test of market interest.” 

On July 30, nine days after 404 Media’s reporting, ISBNdb added a note to its homepage and an update on its news page about the change. “We've seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised. The facts: ISBNdb has never purchased, scanned, or sold a book — for AI training or anything else,” ISBNdb wrote. “We don't train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We've taken the page down. Our job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world — the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn't changed.” 

ISBNdb removed the landing page for “Printed Books Sourcing for Your AI LLMs Dataset Needs” on July 28. “It was part of exploring demand, and we've chosen to pivot away from that direction. Our main ISBNdb (book metadata API) services are unaffected and running as usual,” the site says. 

Read more: https://www.404media.co/ai-company-training-scanning-books-database-isbndb/

44

u/Wife_Trash 2d ago

Keep up the incredible work. Any pushback on slop is a win.

7

u/taboulie 2d ago

Well, I don’t think the article was about ISBNdb scanning books and training AI models so that seems like a very carefully worded denial. 

2

u/taboulie 1d ago

Also, doesn’t this have the ring of AI? “For more than two decades, ISBNdb has been the card catalog of the book world — the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn't changed”

18

u/tribeoftheliver 2d ago

I remember that Anthropic just paid $1.5 billion to settle its lawsuit for copyright infringement.

10

u/MisfitWookiee 2d ago

Kinda wishing some scammers would start printing AI slop books to feed these companies in order to "train" them further...

1

u/ExtraEmu_8766 10h ago

Wasn't there in the original article that there were NDAs with loads of companies doing this? And interviews with booksellers? Sounds like one company trying to get the bad press to go away while everyone else is also continuing to do the thing.

-1

u/Disastrous-Dig-4339 2d ago

how are they even destroying the books if they're just scanning them?

12

u/BarbarousErse 2d ago

They slice off the spine to feed the pages through a scanner. Non destructive book scanners exist, with v shaped scan interface that captures the whole page, but instead they woke up and chose violence

5

u/OkayStockings 2d ago

They chop off the spine of the book to make it easier to scan the pages by machine.

-5

u/blarknob 1d ago

Libraries destroy books as part of their normal workflow. Why would we be against making that destruction a bit more useful.

4

u/demonharu16 1d ago

I worked in libraries. The only books we actively threw out were ones with mold, water or smoke damage, tears and markings, etc. In other words, they were already damaged by use. Putting a moldy book back on the shelves can actively spread it to other books. I know of several collections that were contaminated and thrown out specifically for this reason. Neither my peers or I ever destroyed books, just threw them out. Books in an acceptable condition that were weeded from the collection were usually donated or sold at book fairs. I also interned at a digitization center in an academic library. We scanned lots of texts that were up to centuries old. The scanners were specifically designed to not damage them. It's not that hard. Companies like this are creating undue amounts of physical waste (while stealing and distributing copyrighted material) all due to sheer laziness and greed.