this post was submitted on 29 Dec 2024
20 points (81.2% liked)

Selfhosted

60934 readers
490 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

Detailed Rules Post

  1. Be civil.

  2. No spam.

  3. Posts are to be related to self-hosting.

  4. Don't duplicate the full text of your blog or readme if you're providing a link.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.

  8. AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

I'm currently getting a lot of timeout errors and delays processing the analysis. What GPU can I add to this? Please advise.

you are viewing a single comment's thread
view the rest of the comments
[–] brucethemoose@lemmy.world 2 points 2 years ago* (last edited 2 years ago)

Unfortunately Nvidia is, by fair, the best choice for local LLM coder hosting, and there are basically two tiers:

  • Buy a used 3090, limit the clocks to like 1400 Mhz, and then host Qwen 2.5 coder 32B.

  • Buy a used 3060, host Arcee Medius 14B.

Both these will expose an OpenAI endpoint.

Run tabbyAPI instead of ollama, as it’s far faster and more vram efficient.

You can use AMD, but the setup is more involved. The kernel has to be compatible with the rocm package, and you need a 7000 card and some extra hoops for TabbyAPI compatibility.

Aside from that, an Arc B570 is not a terrible option for 14B coder models.