Science Memes

20042 readers

3056 users here now

Welcome to c/science_memes @ Mander.xyz!

A place for majestic STEMLORD peacocking, as well as memes about the realities of working in a lab.

Rules

Don't throw mud. Behave like an intellectual and remember the human.
Keep it rooted (on topic).
No spam.
Infographics welcome, get schooled.

This is a science community. We use the Dawkins definition of meme.

Research Committee

!spiders@lemmy.world

Other Mander Communities

Science and Research

Biology and Life Sciences

Physical Sciences

Humanities and Social Sciences

Practical and Applied Sciences

Memes

Miscellaneous

founded 3 years ago

MODERATORS

Sal@mander.xyz

fossilesque@mander.xyz

SciBot@mander.xyz

fossilesque@lemmy.dbzer0.com

1290

Scibot! (mander.xyz)

submitted 2 days ago by fossilesque@mander.xyz to c/science_memes@mander.xyz

85 comments fedilink hide all child comments

https://sci-bot.ru/

you are viewing a single comment's thread
view the rest of the comments

[–] brucethemoose@lemmy.world 36 points 2 days ago* (last edited 2 days ago)

We know that training LLMs on LLM-generated text leads to an absolute collapse in quality.

This is often repeated, and true. But needs to be qualified.

Modern LLMs use tons and tons of “augmented” data, which is code for LLM generated or massaged data. Some is even generated during training, and judged; papers on that are what made Deepseek famous.

Training on LLM trash will, of course, yield greater trash, and obviously good text has to come from something real. But that’s because slop is slop. And there are issues with “deep frying” LLMs, yes, but simply training on LLM on LLM output does not necessarily reduce quality. It often helps, significantly.

And we also know that AI has been showing up in papers so if they haven’t, then this will be quite unreliable.

Now this is a problem.

TBH LLMs would be pretty good at flagging papers for humans to check, similar to what Wikipedia is already doing. But yeah, if you just feed a prompt bad papers, LLMs just assume the context is true, generally, and that’s a tremendous problem.