It's not secret, it was their defence when they got sued for copyright infringement. Instead of download all the books from Anna's archive like meta, they buy a copy, cut the binding, scan it, then destroy it. "We bought a copy for personal use then use the content for profit, it's not piracy"
Technology
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
we bought a copy for personao use, then use the content for profit, it's not privacy
So if I buy a song for personal use, then play that song all day in my club to thousands of people, it's not piracy, is what you're saying?
Because anthropic is full of shit and some weird ass mental gymnastics doesn't change anything
After this debacle, nobody can ever again shame me for piracy, let alone punish me for it
C'mon now. You're not nearly rich or influential enough to get away with that and you know it. Rules are for regular people, not the rich or mighty. Sheesh.
/s
Oh I know, but that why I'm getting more and more "Fuck the rules, fuck your laws, until they're the same for everybody"
If they reprinted those scanned books and sold them or even gave them away, they would be in more trouble than you would by sharing on limewire by dent of numbers. That isn't what they are doing with these books. In fact, they did get in trouble for using the books they didn't buy.
The two legal tiers are making themselves known again
“We bought a copy for personal use then use the content for profit, it’s not piracy”
That is an accurate view of how the court cases have ruled.
Downloading books without paying is illegal copyright infringement.
Using the data from the books to train an AI model is 'sufficiently transformative' and so falls under fair use exemptions for copyright protections.
Reminder, this includes "Morning Glory Milking Farm" and similar books.
I'm sure that will destroy any intelligence.
All of this, so some hustlebro can make his own AI slop blog polluting the internet, so instead of the actual information, you get an AI hallucinated one from googling.
People who are okay with this are absolutely disgusting. Some shitty AI company wastes a fuckton of our collective resources resources to build and run their AI data centers, and if that wasn't bad enough they generate a fuckton of unnecessary waste to train the goddamn thing. Fuck capitalism.
They make everything more expensive. Power, water, ram, storage, and now the used book market will shoot up in cost as millions of books are shredded.
AI data centers are cancer to our world - consumes massive energy and water, sucks all the processors and RAM from the market, and raises their price for us. Not to mention environmental impact.
I assume "destructively scan" means to cut the spine off so they lie flat, and that one copy of each book will be scanned? Isn't that a pretty normal way of doing it in cases where the prints aren't rare?
Is this an opportunity to self-publish my own book for $100k per copy and be guaranteed one sale?
No they will simply steal it, like they usually do.
How about 5000 $200 books written by their own AI (preferably for free, cheapest printing in existence) ?
Article is not available without registering. As for the title, "destructive" book scanning means you cut off the binding and put the pages in a scanner which easily flips through them and takes the pictures. If you're not scanning rare old books, this is a perfectly reasonable way to do it, because setting up a scanner for a normal book and manually turning each page to scan it takes a long time (Internet Archive has videos on how they do it, very nice and impressive, and logical since their original mission was scanning old public domain stuff, i.e. published before 1930 or so). If Anthropic will actually legally buy all those thousands upon thousands of books, that will be a pleasant precedent for an AI company.
Although I very much doubt that random uncritically gathered textual material can "teach their AI tool how to write well". They're still pushing for more and more training data, even though it's clear actual advancement will have to happen (if it can happen) through more refined usage of / training on the data.
Write a book where the spine is a required piece of the story for its understanding or completion.
Kind of like how House of Leaves is best enjoyed with the actual book.
I read one once where being able to slightly see through the pages was a key part of the plot
Which one, if you can recall? I love interactive books.
It was called 世界でいちばん透きとおった物語 by Hikaru Sugi, but I don’t think there’s an English translation because this kind of gimmick works a lot better in scripts where all characters are the same size, and a translation that ends up with a comparable arrangement of those letters would be a major pain too.
A slow-burn read by learning Japanese first. This one will take me while.
Fuck yeah, I can already reliably recognise like ⅓ of the hiragana set...if there's a multiple choice pick.
I'm taking a course, but if you want to just self study www.kanadojo.com is pretty good, and if you get anki there's a load of free resources to practice listening and reading. Anki is free on android and pc, but costs a bit on iOS. Www.Ankiweb.net