this post was submitted on 28 Aug 2026
796 points (98.7% liked)

Technology

87624 readers
4666 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

you are viewing a single comment's thread
view the rest of the comments
[–] Blaster_M@lemmy.world 40 points 22 hours ago (3 children)

amazing... I know of smaller models that specifically avoided that sort of thing because even negative training could go wrong so easily... they could have avoided that thing entirely, but nope...

I could understand if they were training a safety classifier (a model that learns what danger stuff is so it can recognize it when it sees it) but I doubt this is what they were doing here. That and handling such a radioactive dataset makes handing the demon core seem safe.

[–] uriel238@lemmy.blahaj.zone 2 points 8 hours ago

I thought that Google used a similar dataset, possibly borrowing images from the NCVIP in order to create an analytic rule-set for which to omit images from Google Image Search. (A larger, similar rule-set is used to omit legal pornography when safe-search is on.)

Mind you, this was before the hyperscale AI era, when LLMs were things like SIRI and Google Now. And Google search still focused on websearch hits and not AI summaries.

Fun story: This was an area of study of mine during the early 2010s, since every image search engine would filter porn hits whether or not you had safe-search on or off. If it was turned off, porn would be filtered to the end of the list unless the engine decided you were intentionally looking for porn in which case the porn hits would be shown at the top of the list. There was no way to get results that ignored the ID-as-porn status of the images. I would enter ambiguously risqué terms to see how explicit I needed to be before the engine decided I was looking for porn.

[–] Gullible@sh.itjust.works 32 points 22 hours ago (1 children)

"You don't understand! I need my 400 terabytes of child pornography to keep the kids safe!"

[–] DeathsEmbrace@lemmy.world 11 points 20 hours ago (1 children)

I bet you thats the same excuse the CIA used when they made all those honey pots for poor pedophiles.

[–] Gullible@sh.itjust.works 4 points 15 hours ago

for poor pedophiles

I think I know what you mean, but I strongly urge you to rephrase that in the future.

[–] Kaligalis@lemmy.world 6 points 19 hours ago

They probably let it just eat the internet without any humans looking at the training data. And yeah - there probably is some CSAM somewhere on the internet. Maybe, they used an agent swarm to specifically search for stuff not yet in the dataset and forgot to a blocklist. It's not like the tech bros are genuinely careful in what they do. Recklessness seems to be a common trait.