this post was submitted on 31 Aug 2026
57 points (72.1% liked)
Linux
14858 readers
917 users here now
A community for everything relating to the GNU/Linux operating system (except the memes!)
Also, check out:
Original icon base courtesy of lewing@isc.tamu.edu and The GIMP
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
By the way, there's like two open-source LLMs, which are large enough to be usable in some form. And one of them claims not to contain pirated training data. It's pretty much the opposite of "plenty".
Which ones are those two?
Yeah, good question! I bet nobody even knows their names. I had hoped to get a reply by @popcar2@programming.dev or @sims@lemmy.ml who seem to know plenty of them?!
I was talking about the RedPajama project who recreated the dataset of Meta's LLaMA and proceedingly trained a model on it. And Apertus by the ETH Zürich. They also factored in ethics when compiling the dataset.
I made a post asking on !localllama: https://programming.dev/post/55905594
In case anyone knows about the existence of such models, please educate me over there!
One more than i knew thank.