this post was submitted on 12 Apr 2026
84 points (78.8% liked)

Selfhosted

59923 readers
560 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

  1. Be civil: we're here to support and learn from one another. Insults won't be tolerated. Flame wars are frowned upon.

  2. No spam.

  3. Posts here are to be centered around self-hosting. Please ensure it is clear in your post how it relates to self-hosting.

  4. Don't duplicate the full text of your blog or git here. Just post the link for folks to click.

  5. Submission headline should match the article title.

  6. No trolling.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

Quick post about a change I made that's worked out well.

I was using OpenAI API for automations in n8n — email summaries, content drafts, that kind of thing. Was spending ~$40/month.

Switched everything to Ollama running locally. The migration was pretty straightforward since n8n just hits an HTTP endpoint. Changed the URL from api.openai.com to localhost:11434 and updated the request format.

For most tasks (summarization, classification, drafting) the local models are good enough. Complex reasoning is worse but I don't need that for automation workflows.

Hardware: i7 with 16GB RAM, running Llama 3 8B. Plenty fast for async tasks.

you are viewing a single comment's thread
view the rest of the comments
[–] doodledup@lemmy.world -5 points 2 months ago (3 children)

Running a thousand watts and not running a thousand watts can be quiet a difference depending on where you live. And then consider buying all of the hardware. In many cases it's probably cheaper to just pay $40 al month.

[–] fuckwit_mcbumcrumble@lemmy.dbzer0.com 14 points 2 months ago (1 children)

What whack ass setup so you think OP has? Dual 5090s? They’re running it on an i7.

[–] T156@lemmy.world 7 points 2 months ago* (last edited 2 months ago)

It's also an 8 gigaparameter model. That's pretty tiny, even if they use it heaps.

[–] StripedMonkey@lemmy.zip 14 points 2 months ago

That would be true worst case, but you're never running inference 24/7. It's no crazier than gaming in that regard.

[–] semperverus@lemmy.world 9 points 2 months ago* (last edited 2 months ago)

Do you think it runs at 1000w continuously? On any decent GPU, the responses are nearly instantaneous to maybe a few seconds of runtime at maybe max GPU consumption.

Compare that to playing a few hours of cyberpunk 2077 with raytracing and maxed out settings at 4k.

Don't get me wrong, there's a lot to hate about AI/LLMs, but running one locally without data harvesting engines is pretty minimal. The creation of the larger models is where the consumption primarily comes in, and then the data centers that run them are servicing millions of inquiries a minute making the concentration of consumption at a single point significantly higher (plus they retrain the model there on current and user-fed data, including prompts, whereas your computer hosting ollama would not.)