this post was submitted on 29 Aug 2026
66 points (100.0% liked)

TechTakes

2668 readers
103 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 3 years ago
MODERATORS
 

Cue breathless news headlines about how AI can now automatically absorb viruses like the borg which hopefully scares people enough to juice up the value of AI stocks just a littttle longer...

The website returns a 303 response, and redirects to a malicious ZIP archive, which Claude then downloads. This archive contains seemingly harmless files including catalog metadata, a README file, seven Base85/zlib-encoded JSON notebook records, a macOS decoder-darwin binary – plus a poisoned Python file named struct.py.

Claude, per its safety guardrails, refuses to run the decoder: “This is planned and what the attacker wants,” Rehberger wrote. Instead of using the supplied binary, the AI decides to write its own decoder. “Ironically, that safety decision is the exploit path,” Rehberger explained, adding a purple devil emoji to the text.

Reality is satire yet again.

“The solution is something we talked about for many years,” he wrote. “Do not trust the model output.”®

you are viewing a single comment's thread
view the rest of the comments
[–] Architeuthis@awful.systems 3 points 2 days ago* (last edited 2 days ago)

it's fundamentally unsolvable, you can only mitigate it, mostly by using classical means to constrain the deterministic (i.e. non-AI) tools the chatbot is allowed access to, and constantly asking the user for confirmation.

With yolo/auto mode (no user confirmation required) and training LLMs on known vulnerabilities things will inevitably get more complicated.