Partially. I started with hosting my own llama3.2 + granite4 models using Ollama for my Home Assistant smart home and for general chat with OpenWebUI. I also run whisper for speech-to-text locally on my 1080 Ti GPU. I like the privacy and ownership of my self-hosted models, but I started to run into limitations with the small weights. So I built some tools that allow me to selectively route traffic to larger models hosted on DeepInfra depending on my need. For example, to GLM/Kimi models for code reviews or for my custom harnesses or harder problems.
Selfhosted
A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.
Rules:
-
Be civil.
-
No spam.
-
Posts are to be related to self-hosting.
-
Don't duplicate the full text of your blog or readme if you're providing a link.
-
Submission headline should match the article title.
-
No trolling.
-
Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.
-
AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.
Resources:
- selfh.st Newsletter and index of selfhosted software and apps
- awesome-selfhosted software
- awesome-sysadmin resources
- Self-Hosted Podcast from Jupiter Broadcasting
Any issues on the community? Report it using the report flag.
Questions? DM the mods!
I'm using anythingllm. It's quite easy to setup and use. I'm impressed of the perf on comodity hardware.
I have the setup, never found a use for it though.
I don’t host it exactly, just use it when I don’t use my graphics card for gaming. I run Qwen3.6-35b on my 16gb vram RX 9700 xt with 34t/s. I use it as an IT advisor, admin and Linux teacher for my cachyOS gaming PC.
Hell naw my homelab is already sucking way too much power and running too hot.
No, I'm not interested in that topic
Yep.
Ollama + about 8 different models at the moment, hosted on a mac mini with open webui as a front end.
Predominantly for transcription, translation, an extra round of security checks on code, a more context friendly home assistant interface, and a daily run of context evaluation on property I'm looking for with a lot of specific needs (acreage, min elevation change, soil type, area, etc).
I have to recommend switching to llamacpp. It's SO much faster than ollama.
No I don't. Unforunetly using Claude (asking myself everyday why tf cuz I don't do crazy shit) but trying to move on to LumoAI even meaby will buy a premium version to check this out formyself.
I started out playing around with code generation using Ollama/open-webui and qwen 2.5 coder 14b on a 3060 12GB, but ended up on a winding journey with an ex datacenter card called the AMD V620. Its roughly equivalent to an RX 6800XT, but with double the VRAM. At this point i've really done nothing productive with it but learned a lot about bios settings, GPU/ROCm drivers, and custom fan solutions/PWM controls trying to get it setup and optimized haha.
It's pretty sick though, that amount of VRAM with 512GB/s bandwidth can run Qwen 3.6 27B dense with 100k context window at 20 tokens/sec in LM studio. Draws 300 watts at the wall on my ITX chassis (idling about 30w).
I've been dabbling in building an aviation weather and field condition report application using this, but my next step is to rebuild my VS Code environment into a new machine. I'm kinda enjoying just fucking around with building the hardware too though
Oh...i recognise this sickness :)
Jup. Ollama and OpenWebUI is a great stack to tinker with some LLM models. They're kinda useful for aggregating large datasets, translations, frontend development and gathering relevant sources for me to read into. Also, Qwen has been amazing in understanding frameworks without documentation and writing one for me. I had to use some self-developed PHP framework for a task once and without qwen, I would've taken probably two more weeks to get the task done.
MiniCPM has also been REALLY good at image detection, describing it as accurately as possible, feeding it into qwen who then searches what the object could be and returning the result. I always liked google lense and that stack gave me a TEMU-Version of google lense that isn't quite as reliable, but definitely very useful.
Yes. Currently using Gemma4:12b behind OpenWebUI and Hermes Agent plus a few lighter models for OCR and tagging in Paperless.
I have a simple slow model running on CPU in my cluster for karakeep. I've tried running a variety of models on my 7900XT but even with 16GB their performance just isn't there. My new work m5 Mac book with 48GB of ram is the first time I've seen usable performance for local models and it has been pretty impressive.
i don't use it at all, i do want some selfhosted speech to text model (whisper?) but my computer is ancient so it would be awfully slow. i have some multi hour audio recordings from presentations, would be nice to have them in text and searchable..
No, I have taste.
I use my gaming rig to serve up qwen3.6-coder to Open Web UI and that's been very successful in helping me refactor my home lab to be more effecient and easier to support. Over the years of building my server I got everything working, but lets just say it's a bot of a mess and a lot of shortcuts were taken.
I plan to look into ComfyUI soon but I do that have much of a use case for it at the moment.