this post was submitted on 28 Aug 2026
899 points (98.7% liked)

Technology

87646 readers
3440 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

you are viewing a single comment's thread
view the rest of the comments
[–] Lemming6969@lemmy.world 0 points 1 day ago (1 children)

Models do not store the original, they store the model, which is a huge graph of probabilities.

[–] Voroxpete@sh.itjust.works 2 points 23 hours ago* (last edited 23 hours ago) (1 children)

I'm aware. But if you read my explanation a little more carefully, you'll see that the argument being made is that this de facto constitutes a form of lossy compression.

The easy comparison is that a JPEG does not store an "original" image, but it contains information that can be used to almost perfectly recreate that image, with the help of a little math.

If the same argument can be said to hold true of LLMs - and yes, that is very much a load-bearing "if" - then they would constitute a form of lossy compression.

[–] Lemming6969@lemmy.world 2 points 22 hours ago (1 children)

How far does one take that? A picture of a horse could be transformed into Abraham Lincoln with the right algorithm. Is that lossy compression?

[–] Voroxpete@sh.itjust.works 0 points 13 hours ago (1 children)

That's what courts exist to decide.

[–] Rothe@piefed.social 0 points 12 hours ago

Not in the US though.