this post was submitted on 29 Aug 2026
65 points (100.0% liked)

TechTakes

2668 readers
133 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 3 years ago
MODERATORS
 

Cue breathless news headlines about how AI can now automatically absorb viruses like the borg which hopefully scares people enough to juice up the value of AI stocks just a littttle longer...

The website returns a 303 response, and redirects to a malicious ZIP archive, which Claude then downloads. This archive contains seemingly harmless files including catalog metadata, a README file, seven Base85/zlib-encoded JSON notebook records, a macOS decoder-darwin binary – plus a poisoned Python file named struct.py.

Claude, per its safety guardrails, refuses to run the decoder: “This is planned and what the attacker wants,” Rehberger wrote. Instead of using the supplied binary, the AI decides to write its own decoder. “Ironically, that safety decision is the exploit path,” Rehberger explained, adding a purple devil emoji to the text.

Reality is satire yet again.

“The solution is something we talked about for many years,” he wrote. “Do not trust the model output.”®

you are viewing a single comment's thread
view the rest of the comments
[–] bignose@awful.systems 14 points 2 days ago* (last edited 2 hours ago) (3 children)

The major selling point for these LLM services is that people don't need to learn a strict formal query language to interact; the LLM consumes instructions and data all as an undistinguished stream. It just happens that an LLM, by its design, will statistically infer a plausible response with absolutely no regard to the input's meaning, nor the response's meaning.

That is: despite the media reporting these as “injection attacks”, that term means nothing when all its input, every time, is treated as data and instruction simultaneously. This isn't some special class of attack; it's a fundamental designed-in flaw of the system.

The correct way to ensure protection from these vulnerabilities is long established, from decades of experience. You establish a firm boundary: never treat the input data as instructions, but instead have a separate channel for extremely well formalised query instructions, and reject bad input on that channel.

But of course that would kill the major appeal for most people who love these things, the fact they don't need to learn any strict formal language and can just say anything at all and get some useful-looking response. Take that away, and you lose any hope that the masses will want to use these systems.

And so the makers and promoters of these systems will never make the one change that could even feasibly allow safety from these attacks; they will never make any improvement to security that might reduce the apparent ease of use of these things.

For as long as that remains, these systems will continue to inevitably have these exploits because the corporations won't close the exploit surface on these systems. These are staggeringly insecure by design, and can't be fixed without being completely replaced.

[–] EFreethought@awful.systems 2 points 1 day ago (2 children)

This isn’t some special class of attack; it’s a fundamental designed-in flaw of the system.

Will it ever not be a flaw in LLMs? Or will they always be this way?

[–] Architeuthis@awful.systems 2 points 1 day ago* (last edited 1 day ago)

it's fundamentally unsolvable, you can only mitigate it, mostly by using classical means to constrain the deterministic (i.e. non-AI) tools the chatbot is allowed access to, and constantly asking the user for confirmation.

With yolo/auto mode (no user confirmation required) and training LLMs on known vulnerabilities things will inevitably get more complicated.

load more comments (1 replies)
load more comments (1 replies)