this post was submitted on 07 Sep 2026
580 points (97.7% liked)

Technology

88253 readers
3519 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] panda_abyss@lemmy.ca 433 points 2 weeks ago (5 children)

Because they all use the illegal Colossus 2 data centre from SpaceX/XAi/fascism central and that data centre went down.

Google, Anthropic, and OpenAI all have contracts with them.

When that illegally running environmental disaster of data centre goes down, all those services all go over capacity and you get 502 rate limit errors.

[–] pelespirit@sh.itjust.works 122 points 2 weeks ago (2 children)

Does that mean Musk's businesses have access to everything people do on the systems that use their data centers?

[–] just_another_person@lemmy.world 160 points 2 weeks ago (2 children)

Generally no, because most hosting companies would have something baked into SLA/SLO contracts, but all of this shit is done so illegally and shadily now, I wouldn't put a hard "no" on that possibility.

[–] skvlp@feddit.nl 62 points 2 weeks ago (1 children)

I hope you’re right, but Elon don’t strike me as the guy whose most compliant to SLA, law, or anything else that might be an inconvenience to him.

[–] esc@piefed.social 18 points 2 weeks ago (3 children)

He can't micromanage everything and regular management will try to comply.

[–] bedwyr@piefed.ca 20 points 2 weeks ago

Executives are the biggest cheaters out there. They will only comply if it will hurt them if they don't, and it's won't here, as long as the protection money is produced they don't have to worry.

[–] skvlp@feddit.nl 8 points 2 weeks ago

I agree with that. But I think "the right" data scientists can infer too much from all those AI queries, and I think Elon can abuse that for his own gain.

[–] shartgargle@lemmy.zip 7 points 2 weeks ago

I highly doubt that he really knows that much, either.

[–] devfuuu@lemmy.world 42 points 2 weeks ago (1 children)

Nobody can verify that those "contracts" are actually being honored. We are all assuming the fascists will self police themselves out of good will.

[–] just_another_person@lemmy.world 27 points 2 weeks ago (2 children)

Not true. I've been involved in litigation with both AWS and Google over unsecured comms that were defined as being TLS secured in SLA/SLO contracts and found not to be. Not that anything nefarious was happening necessarily, but the expectation is clear.

Whether these asshats even check for such requirements with a Musk run company right now 🤷

It COULD possibly be that they are logging every exchange happening at the network fabric between the service layers, but nobody knows unless they intentionally take steps to investigate or accidentally prove it.

[–] expr@programming.dev 4 points 2 weeks ago (1 children)

Still means jack shit. Companies will lie through their teeth and violate any and all they can get away with.

They are not to be trusted.

[–] Cocodapuf@lemmy.world 2 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

You're missing the point, nobody is talking about trusting these companies.

You can tell if a connection is encrypted end to end or not. And if your paying for that service you can sue if you aren't receiving what your paying for.

In digital security nobody relies on trust if they can help it. And security matters if you want to keep a competitive edge on your competition, so even shitty companies care about that.

[–] expr@programming.dev 1 points 2 weeks ago (3 children)

I was not talking about your specific lawsuit, or TLS.

Companies trust other companies all the time, and it's foundational to most all SLA/SLOs. Any time I've voiced concerns around how AI companies are using the data we are giving them (like giving them access to our codebase), it's brushed off as "we have an agreement with them". It's just a load of hogwash. They can and will abuse all data they have access to, just as they have done thus far.

In this particular case, we are talking about a data center, and it is not at all reasonable to assume that the data that flows to said data center is in any way protected, especially one run by Musk.

load more comments (3 replies)
load more comments (1 replies)
[–] panda_abyss@lemmy.ca 22 points 2 weeks ago (2 children)

Impossible to know.

There are systems for doing things cryptographically secure, but I don’t know much about that.

[–] 314xel@lemmy.world 34 points 2 weeks ago (1 children)

Cryptography? Security? FROM A GENAI COMPANY?!

[–] Archer@lemmy.world 6 points 2 weeks ago

Well yeah, for themselves

Homomorphic encryption? Much slower than "trust me, bro".

[–] QuantumSparkles@sh.itjust.works 21 points 2 weeks ago (1 children)

And where is said data center? For research purposes

[–] bad1080@piefed.social 24 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

https://epoch.ai/data/ai-data-centers/directory/colossus-2 says "Memphis, Tennessee, United States" but also provides satellite imagery

[–] BeatTakeshi@lemmy.world 26 points 2 weeks ago (2 children)

The numbers of this madness are fucking scary, 1.5 GW of compute power, and huge arrays of cooler to keep it operating. It's basically a 1.5GW heater in the open. Oh and gas turbines to provide (part of) the electricity. Capitalism is burning this world down for a buck.

[–] NotASharkInAManSuit@lemmy.world 14 points 2 weeks ago (2 children)

Today I learned that “AI” data centers use more energy than time travel.

[–] shrugs@piefed.social 9 points 2 weeks ago

i bet they will also travel us to the apocalypse much faster

[–] BeatTakeshi@lemmy.world 2 points 2 weeks ago

I was about to say, slightly less according to Emmett Brown, then I checked and it's only in the French version that it says 2.21GW. TIL.

[–] filcuk@feddit.uk 1 points 2 weeks ago

We can't really produce enough heat through industry to affect the earth globally, if that's what you meant. It's insignificant in comparison to what the Sun provides.
However there is undeniably localised issues caused by these insane structures.

[–] BarbecueCowboy@lemmy.dbzer0.com 18 points 2 weeks ago (1 children)

It's kinda surprising,

I know specifically where one of the big ones hosts its models and its not there, but I guess they could have infrastructure in there.

[–] panda_abyss@lemmy.ca 22 points 2 weeks ago (2 children)

They’re oversold though, especially prompt caching and the parameter count war

The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.

[–] asbestos@lemmy.world 18 points 2 weeks ago (1 children)

Reads like American vs European/Asian cars

[–] dgriffith@aussie.zone 8 points 2 weeks ago* (last edited 2 weeks ago)

Always ready to try brute force first. And then some other, less palatable options if that doesn't work, like slightly less brute force.

[–] percent@infosec.pub 5 points 2 weeks ago (1 children)

The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.

(It's not super complex code, just some scripts that I would not have taken to time to write manually.)

I'd love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬

[–] isVeryLoud@lemmy.ca 3 points 2 weeks ago (1 children)
[–] percent@infosec.pub 4 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

There's not really anything interesting to show. It's just a home server in a 13 year old desktop ATX case.

There's no desk, monitor, keyboard, or mouse... But also no cool server rack.

Function over form, and it sits in a spare bedroom out of sight.

EDIT: I found the receipt for the case. It's a Cougar Volant Black Steel mid tower, purchased in 2013. So my server just looks like this:

[–] isVeryLoud@lemmy.ca 3 points 2 weeks ago (2 children)

I meant your LLM stack lol. I just have an RX 6800 XT in my main Linux PC for inference, but it has to share VRAM with the DE. Maybe I'll set it up for remote development from my laptop instead to free up VRAM.

What are you using? vLLM? llama.cpp? Which params? How much CPU offloading? Do you use draft models? Is it a MoE model? Have you tried llama-swap? Which agentic front-end are you using? I presume you set it up to access it without SSH'ing into the machine, did you do anything special or is it just a raw unsecured open port on the machine to the LAN?

[–] percent@infosec.pub 5 points 2 weeks ago (1 children)

Ohhh lol. Yeah it's Llama-swap, running llama.cpp for now, but might add vLLM to the llama-swap config to experiment with NVFP4.

I mainly use MoE models so I can get decent speed while using a 150-200k context window. My go-to model has been Qwen3.6 35B-A3B for a while. I tried Qwen3.8 27B, but it was too slow.

Gemma4 26B-A4B also runs nice and fast, but I generally get better results from Qwen3.6. I don't remember exactly how much CPU offloading is happening, but it's not much. As long as I can get like 40-50 tokens/sec, I'm usually satisfied enough.

For the coding harness, I've been running Pi in an Apple Container (sort of like Podman, but better isolation in a microvm). Though, I recently configured VS Code to use LLMs on my server, and it was actually pretty decent. Still need to explore a bit more, but so far VS Code's AI capabilities seem much better than they were a year ago (they seemed way behind, back then).

Also, I don't connect any harness directly to llama-swap. I have another container running Caddy, which acts as a gateway to AI providers. For other services (e.g. OpenRouter), the API key is injected in the Caddy container. I don't like having API keys or secrets anywhere where LLMs can read them. It's not so bad for my own self-hosted LLMs, but not cool to send secrets to a server owned by someone else.

[–] isVeryLoud@lemmy.ca 2 points 2 weeks ago (1 children)

How has tool use been for you? I struggled a lot with tool use with Gemma and Qwen, to the point where I needed to build a healing layer.

Regarding the coding harness, I was looking for something CLI-based or JetBrains-based, and I haven't had much luck getting my local llama.cpp models playing ball with OpenCode. They keep losing context and misusing tools.

I'm not too familiar with Apple containers as I'm running a full Linux stack, but I'll give Pi a try, seems interesting! Does it work for coding tasks or is it strictly an "orchestrator"?

[–] percent@infosec.pub 2 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

Tool use with Gemma has been hit or miss. I wouldn't rely on it for anything unsupervised.

Tool use for Qwen3.6 has been great lately, but I do remember seeing some issues with it too, a while back. I don't remember when/why the issues cleared up (I have tweaked configs a bit over time), but switching to Pi definitely helped.

I do remember having a lot more problems in OpenCode and it was practically unusable (which is why my recent experience with VS Code was surprising). I'd definitely recommend trying Pi.

A fresh Pi install is very minimal by design. The system prompt is tiny, so it's a pretty good fit for small LLMs like these. It's sort of like Neovim: Nothing fancy out of the box, but you can add lots of fancy things to it. I containerize it because I don't like giving LLMs (especially these small ones) unrestricted access to my host computer – though, I have not seen any signs of it accidentally doing something destructive, which is surprising.

There are similar alternatives to Apple Container for Linux (e.g. Docker Sandboxes, muvm, Firecracker). There's also this thing made specifically for Pi called Gondolin. I haven't tried it yet, but I may end up switching to that if it could simplify my stack.

Here's my current llama-swap/llama.cpp config for Qwen3.6 35B-A3B:

qwen3.6-35b-a3b:
    name: "Qwen3.6 35B-A3B (Coding)"
    proxy: "http://127.0.0.1/:$%7BPORT%7D" # If you're seeing a `/` after `127.0.0.1` here, don't include it. I think something in Lemmy is trying to "sanitize" this input by adding the `/`.
    cmd: |
      llama-server
      --port ${PORT}
      --no-webui
      -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL
      --jinja
      --parallel 1
      --flash-attn on
      --no-mmproj
      --load-mode none
      --reasoning-preserve
      --ctx-size 190000
      --temp 0.6
      --top-p 0.95
      --top-k 20
      --min-p 0.0
      --presence-penalty 0.0
      --repeat-penalty 1.01

A few notes about this config:

  • Now that I think of it, --reasoning-preserve might be another thing that helped with tool calls.
  • Note the -MTP part of the -hf param. MTP helps speed things up. Here's the Huggingface page for this model
  • You can also omit --no-mmproj if you need vision, but it might mean sacrificing speed or context size, so I usually just enable vision in a separate llama-swap model entry to use as needed.
  • Unsloth recommends --repeat-penalty 1.0, but I saw the LLM enter a thinking loop in VS Code, so I bumped it up just a tiny bit to 1.01. I have since seen it do something that resembled the same thought loop, but it was able to recover on its own. Not sure if it's a coincidence or if 1.01 was actually the solution, so worth some experimentation.
[–] isVeryLoud@lemmy.ca 1 points 2 weeks ago (1 children)

Cheers, I'll give this a try!

Regarding containerization, familiarize yourself with Dev Containers, they're super useful for limiting agents to your codebase, with the added bonus that any project you work on comes out of the box with the right version of the tools you need.

[–] percent@infosec.pub 2 points 2 weeks ago (1 children)

Yeah I used to use dev containers. Containers aren't generally a secure sandbox. They're a great guardrail for preventing accidents, but not so much with a malicious prompt injection. (I might be overly paranoid about these things.)

For tool version management, I usually set up a Nix Flake for each project (which also works inside dev containers).

load more comments (1 replies)
[–] Damage@feddit.it 2 points 2 weeks ago (1 children)

it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.

eh, just systemctl isolate multi-user.target

[–] isVeryLoud@lemmy.ca 1 points 2 weeks ago (1 children)

Correct, that's how I would do it, but then I need another machine to act as a head.

[–] Damage@feddit.it 3 points 2 weeks ago (1 children)

If your MB has onboard graphics, maybe you could mask the GPU and just pass it off to a container running the LLMs I guess

[–] isVeryLoud@lemmy.ca 1 points 2 weeks ago (1 children)

No onboard graphics unfortunately, that would have been the easy way out.

[–] Damage@feddit.it 2 points 2 weeks ago (1 children)

Well it works anyway even with a bit of occupied vram, but you could also buy a cheap videocard to use as an output. I have small intel card like that in my server for jellyfin transcoding, I think I paid 60€ for it, it hardly uses any power

[–] isVeryLoud@lemmy.ca 1 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

I actually did exactly that previously! I had both an RX 6800 XT and an RX 6600 in my system and I used the 6600 for video output. Unfortunately, this cuts my RX 6800 XT from PCIe 4 16x to PCIe 4 8x and severely slows down model loading for llama-swap. Joys of the X570!

And yes, I do have it running right now with a bit of occupied VRAM, but I need to limit my model to 14 GB to leave 2 GB free for GNOME Shell. I really want one of those 64 GB UMA Mac Mini, I heard they work really well because the GPU has direct access to system RAM.

[–] Damage@feddit.it 2 points 2 weeks ago (1 children)

So I have a framework laptop with ryzen ai cpu that uses 48gb of shared ram, and it does run Q4 llms fine enough, but I'm not sure it compares to a real GPU.

On my desktop I have an RX 7900 XTX but I've only dabbled in image generation so far, so right now I couldn't really tell you the difference.

[–] isVeryLoud@lemmy.ca 2 points 2 weeks ago (1 children)

24 GB VRAM. Damn, jealous! My 16 GB seems pitiful in comparison 😅

I do wonder if the Ryzen AI CPUs compare with Apple's UMA. I'm mostly interested in LLM inference for code generation and automation.

[–] Damage@feddit.it 2 points 2 weeks ago* (last edited 2 weeks ago)

Yeah I was lucky to buy a 6900XT when it was near the lowest price, so I sold that and added a couple hundred for the 7900, seemed like a good future-proofing move, given the times we're living in.

I think Apple silicon is faster than my generation of Ryzen, but the newest (Strix something?) with the LPDDRGGFASEWARGH5 memory should be faster yet. Of course buying all that memory right now would be quite painful.

[–] AnalogAllamma@lemmy.world 5 points 2 weeks ago (1 children)
[–] Zephyr@sh.itjust.works 1 points 2 weeks ago

I'm not sure that would effect inference on an already trained model. Data sets are primarily used in training.