Lights are on, server is on.
Selfhosted
A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.
Rules:
-
Be civil.
-
No spam.
-
Posts are to be related to self-hosting.
-
Don't duplicate the full text of your blog or readme if you're providing a link.
-
Submission headline should match the article title.
-
No trolling.
-
Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.
-
AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.
Resources:
- selfh.st Newsletter and index of selfhosted software and apps
- awesome-selfhosted software
- awesome-sysadmin resources
- Self-Hosted Podcast from Jupiter Broadcasting
Any issues on the community? Report it using the report flag.
Questions? DM the mods!
Whining from end users.
Zabbix
Prometheus + grafana. It's overkill tbh, most of the services restart automatically and I mostly ignore it. (it's an artifact from life where I cared about it)
Exactly this scenario at my place as well. 😄
I use nagios
I pay a group of schoolchildren to refresh a group of browser tabs and yell if anything isn't responding.
There are plenty of previous threads with recent answers if you search.
Well I used to use VGA for my monitors, now its mostly displayport.
Seriously though, prom+Loki+alloy+a few other things to Grafana, with alerting in grafana.
Uptime Kuma! For HDDs I have them report their “ping” as percentage full
I usually just walk down the hall and move the mouse. I don't really need remote monitoring. 😅
I use Alloy to collect Metrics of the host (Disk usage, CPU, Ram, etc.) and different logs. With Grafana the data is then displayed as a daschboard. If everything goes south, Alertmanager sends Mails to me
Grafana dashboard

You can find more details and the source code here: https://erasmus.works/
Ooooo, that looks nice.
I use naemon (nagios fork) and grafana Loki Prometheus influxdb
Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks
Cockpit. Total gamechanger for me. I don't have to ssh to my boxes individually to see what's going on. Instead it's a web UI AND you can connect basically every server you want to one instance.
https://github.com/cockpit-project/cockpit
That said. For individual servers?
htop/nmon (Get nmon straight from the source! https://nmon.sourceforge.io/pmwiki.php) for a birds-eye sysmon view of things
ps -ef | grep <whatever you're looking for> for investigating stuck/hung/zombie processes
dmseg
tail -f /var/log (or whatever log file you care about)
netstat for net statistics (i think ntop used to be a thing but idk anymore>...)
df -h to look at disk usage
those are all that come to mind right away.
EDIT: formatting and quick explanation of what each tool does
It's amazing the risk profile of that.
Lol I go to all of my server's specific cockpit addresses, sounds like I'm doing it wrong, I'll need to look at connecting them
So it turns out.... you can't connect them anymore.... The feature is deprecated
https://docs.cockpit-project.org/cockpit-guide/latest/guide/feature-machines.html
(also, I wanted to just link my response from another user who had asimilar question but for whateVER reasyon, Alx won't let me??)
How do you configure adding multiple servers to the same instance?
So I actually just went to make a guide on this, demonstrating on my ubuntu 26.04 lab vps, and it appears the connect option is not there.
However on my 24.04 lab vps, it's there. Same with my local 22.04 systems (yes I know I have to upgrade lmao)
I went to cockpit docks and sure enough, the feature is deprecated :( Dang.
https://docs.cockpit-project.org/cockpit-guide/latest/guide/feature-machines.html
It was really cool though while it worked.
I get a text pretty quickly.
When logged in locally, I use btop to see an overview of what's happening.
Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.
The phone I actually carry around has a ntfy client talking to ntfy server on the aforementioned monitoring solution, so I get buzzes when something goes down.
Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:
- if disk space > 80% consumed, send a low-priority alert
- if disk space > 95% consumed, send an urgent alert
- ... everything else, there's so much to check...
A script checks container status and accessibility, and pushes a ntfy alert if it goes down. Healthchecks.io reports if the whole Pi goes down.
I vibe coded a monitor that shows me all the logs and other stuff. And im using netbird cloud (free version) to setup dns so all my services have url "servicexyz.home.internal".
So i just connect to netbird vpn on my phone and go to that url.
You can pretty much use any of the other tools mentioned in this thread with this kind of setup, heck you can write a simple fastapi server that just runs "top" and return it (it would be like 20-30 lines of python) over a url that you can from anywhere as long as you are connected to your vpn
SNMP.
The old engine was nagios. The new engine will be telegraf -> mqtt -> brokerfest -> timescale -> Prometheus.
I guess. Not sure yet. The brokerfest is a series of broker pubsub between sites to both split out data from the stream for off-site CC, or pull it's own subscriptions in. Data could go host <- telegraf -> broker -> remote broker -> remote timescale -> remote Prometheus
I sync my podcasts hourly with Nextcloud and poddersync, so when my servers go down, my phone beeps at me.
I started with uptime kuma since it was simple and light on resources. I've switched to checkmk community (formerly raw) and really enjoying it. It's telling me so many things that are wrong that I missed. Like that one time I made a snapshot of a VM then forgot about it for 113 days...
Uptime Kuma, only really care if it's up or not. Got about 30 services monitored, sending notifications to Signal.
I use FreeBSD so I don't need to check it to make sure it's working properly