this post was submitted on 28 Jul 2026
343 points (97.8% liked)

Technology

86701 readers
3560 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

you are viewing a single comment's thread
view the rest of the comments
[–] Goodlucksil@lemmy.dbzer0.com 25 points 22 hours ago* (last edited 22 hours ago) (9 children)

What's the point of destroying books after scanning them instead of reselling them apart from *hurr durr we are evil*?

[–] Jason2357@lemmy.ca 10 points 17 hours ago

Easier to scan if you cut the binding off, and it costs money to re-bind a book you purchased for 10 cents as part of a lot.

[–] chiliedogg@lemmy.world 14 points 20 hours ago

Have you ever tried scanning a book? It's a huge PITA.

They just cut the binding off so they can be run through a document feeder.

[–] MadameBisaster@lemmy.blahaj.zone 19 points 22 hours ago

Cause to easily scan them they remove the binding and rebinding is more work %han destroyibg so the point is capitalism

[–] daggermoon@lemmy.world 11 points 22 hours ago (1 children)

To keep their competitors from buying them and scanning them.

[–] Greyghoster@aussie.zone 1 points 21 hours ago (1 children)

That seems to be the most likely reason. Bastards.

[–] daggermoon@lemmy.world 1 points 20 hours ago

Sorry, but I did an oopsie. I watched the video linked, they destroy the books to scan them. It's really sad. Though I'm sure they would destroy them anyway.

[–] Bluescluestoothpaste@sh.itjust.works 3 points 19 hours ago* (last edited 19 hours ago)

I mean have you tried donating books? They just dump them into a storage room with thousands of other books donated that nobody wants.

[–] leds@feddit.dk 2 points 18 hours ago

Do make sure that knowledge is only available through their model and otherwise lost to humanity.

Same with bombarding small websites until they give up and pull the plug.

Everything is fucked

[–] altkey@lemmy.dbzer0.com 3 points 22 hours ago (1 children)

Fair use loophole, they claim they convert one physical book into one digital with no illicit copies, while also supporting the delusion of AI being an equivalent to a human reader. It is weird on so many levels but no one of noticeable weight asked them wtf.

[–] Jason2357@lemmy.ca 1 points 17 hours ago

That is not required by law. It is just easier to cut the bindings off. The other way is slow: https://archive.org/details/eliza-digitizing-book_202107

[–] General_Effort@lemmy.world 2 points 20 hours ago (1 children)

Copyright. They mustn't make a copy. There is legal precedent that confirms it's okay to transform the copy you bought into digital format. The Internet Archive relies a lot on that. I think they actually litigated it in the first place. So that's why the copyright heads are going so absolutely apeshit. If no one's charging you rent for using some data, then it's "unethical".

[–] Jason2357@lemmy.ca 5 points 17 hours ago (1 children)

The AI absolutely does not destroy books. They scan them the hard way with the binding still intact.

[–] General_Effort@lemmy.world 2 points 12 hours ago

I was misremembering. Google won the precedent, when they were suing over Google Books. IA had only 1 big lawsuit and were forced to settle. The non-destructive scanning is dicey.

[–] A_Random_Idiot@lemmy.world -1 points 18 hours ago

To avoid competitors also scanning the book to get some kind of edge.