this post was submitted on 04 Sep 2026
48 points (91.4% liked)

TechTakes

2692 readers
114 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] lurker@awful.systems 9 points 5 days ago (3 children)

Someone pointed out that on some tests, Astra is actually a regression from previous models. Totally AGI guys

[–] diz@awful.systems 6 points 4 days ago

Yeah sloppers I know went from “see it proved new math! that totally clears us of plagiarism accusations we get when we repost that someone ai generated a frogger and a WoW clone” to “tried to use it for coding and it didn’t do a good job”.

The math stuff is mostly using Lean proof verifier and all that, by the way, the only use of LLM that they found that is actually legit because it doesn’t matter if its slopping, and you want your attempts randomized.

It must have been too good before. They needed it to be more general. /s

[–] scruiser@awful.systems 3 points 4 days ago

Benchmaxxing for one set of benchmarks can actually degrade performance on other benchmarks. They've plateaued for a while, all they can do is scale inference compute up and down (at logarithmically poor rates of exchange) and trade performance in one area for performance in others