
Technology
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
That's Mafiology, not innovation.
Great that they are digitalising, but can't they also find a non destructive way to do so? I wonder... Books have always been there and are inherent to our history and human existence. This should not be happening, ever. Books need to be well preserved.
But the destruction is the essential part, ironically, that makes this copying kosher for copyright purposes:
Destroying books, it turns out, isn't just cheaper than maintaining them: The presiding judge also ruled that it's transformative enough to constitute fair use under Section 107 of the Copyright Act.
So, somehow, destroying the books afterwards makes it all legal. In spite of the fact that no author who wrote a book in all of human history before about 3 years ago had to worry about a chatbot sucking up all of its work, and vomiting it back out, without compensation.
I think there's a categorical issue in how this whole story is framed, at least for the examples that have been cited by the booksellers themselves. These books are "rare" because they are low-print and have basically no demand, think obsolete technical books from the 60s or university theses printed at 100 copies. This kind of books are routinely destroyed by booksellers, and even publishers, because keeping inventory is expensive and nobody is buying or reading them. That's how AI companies are able to buy them for pennies on the dollar, and whatever they don't buy is probably eligible for recycling or burning by the seller/publisher anyway.
It's basically a way to slurp up a few billion tokens for cheap. They're destroying the copies by chopping off the spine and feeding the loose pages into page-fed scanners, which is way less expensive than those non-destructive scanning robots, but this doesn't work on really old books from previous centuries as the paper is different and fragile and will jam your scanner constantly.
What they're looking for is industrially-produced books with low circulation, not because they are culturally valuable but because they contain prose that hasn't yet been digitized and that's what they need for pre-training.
Did you read about it? The destruction is the part they are interested in
Don't sell books to anonymous people!
I wonder if there's anything positive to be said about ISBNDB. On one hand, they're supplying AI companies, yes, on the other hand, they seem to be digitalizing previously unavailable paper books, which is kinda useful for preservation? Except, I'm not entirely sure it's possible to properly get this data out of them on consumer scale: they seem to sell API access, but it's a big question of how usable and affordable it would be to retrieve a digital version of some rare ancient book that could just disappear into oblivion otherwise. Their cheapest subscription is 15$/month, but I have little idea of how what's listed there translates into volumes of books you can retrieve during this month.
Also what if the beancounters decided keeping this digital library is costing us too much so it's going bye bye.
Someone throw the book at 'em!
They’ll scan it and say, “please, sir, may I have some more.” Leeches.
