this post was submitted on 09 Oct 2026
88 points (97.8% liked)

Selfhosted

62826 readers
785 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

Detailed Rules Post

  1. Be civil.

  2. No spam.

  3. Posts are to be related to self-hosting.

  4. Don't duplicate the full text of your blog or readme if you're providing a link.

  5. Submission headline should match the article title.

  6. No trolling.

  7. Promotion posts require active participation, with an account that is at least 30 days old. F/LOSS without a paywall has exceptions, with requirements. See the rules link for details. Tags [CBH] or [AIP] are required, see the links in Rule 8 for details.

  8. AI-related discussions and AI-involved promotional posts have additional requirements for tagging, as noted in Rule 7 and the AI & Promotional Post Expanded Rules post, and find example disclosures here.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 3 years ago
MODERATORS
 

Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc...
So far my only method has been "hope and pray".

(page 2) 35 comments
sorted by: hot top controversial new old
[–] Carrots@quokk.au 26 points 21 hours ago (3 children)

Lights are on, server is on.

[–] Reannlegge@lemmy.ca 14 points 20 hours ago (1 children)

I recently found out that this is not always true.

load more comments (1 replies)
load more comments (2 replies)
[–] blarth@thelemmy.club 21 points 21 hours ago

Whining from end users.

[–] lazynooblet@lazysoci.al 4 points 15 hours ago
[–] esc@piefed.social 4 points 15 hours ago (1 children)

Prometheus + grafana. It's overkill tbh, most of the services restart automatically and I mostly ignore it. (it's an artifact from life where I cared about it)

[–] spacelord@sh.itjust.works 2 points 12 hours ago

Exactly this scenario at my place as well. 😄

[–] mhzawadi@lemmy.horwood.cloud 5 points 16 hours ago

I use nagios

[–] frongt@lemmy.zip 17 points 21 hours ago

I pay a group of schoolchildren to refresh a group of browser tabs and yell if anything isn't responding.

There are plenty of previous threads with recent answers if you search.

[–] curbstickle@anarchist.nexus 16 points 21 hours ago

Well I used to use VGA for my monitors, now its mostly displayport.

Seriously though, prom+Loki+alloy+a few other things to Grafana, with alerting in grafana.

[–] lIlIlIlIlIlIl@lemmy.world 14 points 21 hours ago (1 children)

Uptime Kuma! For HDDs I have them report their “ping” as percentage full

[–] WeirdGoesPro@lemmy.dbzer0.com 7 points 20 hours ago (2 children)

Can you teach me this wizardry?

load more comments (2 replies)
[–] popekingjoe@lemmy.world 5 points 18 hours ago

I usually just walk down the hall and move the mouse. I don't really need remote monitoring. 😅

[–] testgoofy@infosec.pub 5 points 18 hours ago

I use Alloy to collect Metrics of the host (Disk usage, CPU, Ram, etc.) and different logs. With Grafana the data is then displayed as a daschboard. If everything goes south, Alertmanager sends Mails to me

[–] Legion739@piefed.zip 1 points 14 hours ago (1 children)

Grafana dashboard

image
You can find more details and the source code here: https://erasmus.works/

[–] user224@lemmy.sdf.org 2 points 13 hours ago

Ooooo, that looks nice.

[–] I_Am_Jacks_____@sh.itjust.works 7 points 21 hours ago

I use naemon (nagios fork) and grafana Loki Prometheus influxdb

[–] truxnell@quokk.au 1 points 14 hours ago

Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks

[–] happy_wheels@lemmy.blahaj.zone 7 points 22 hours ago* (last edited 22 hours ago) (3 children)

Cockpit. Total gamechanger for me. I don't have to ssh to my boxes individually to see what's going on. Instead it's a web UI AND you can connect basically every server you want to one instance.

https://github.com/cockpit-project/cockpit

That said. For individual servers?

htop/nmon (Get nmon straight from the source! https://nmon.sourceforge.io/pmwiki.php) for a birds-eye sysmon view of things

ps -ef | grep <whatever you're looking for> for investigating stuck/hung/zombie processes

dmseg

tail -f /var/log (or whatever log file you care about)

netstat for net statistics (i think ntop used to be a thing but idk anymore>...)

df -h to look at disk usage

those are all that come to mind right away.

EDIT: formatting and quick explanation of what each tool does

[–] corsicanguppy@lemmy.ca 2 points 15 hours ago

It's amazing the risk profile of that.

[–] Owljfien@piefed.world 1 points 17 hours ago (1 children)

Lol I go to all of my server's specific cockpit addresses, sounds like I'm doing it wrong, I'll need to look at connecting them

[–] happy_wheels@lemmy.blahaj.zone 1 points 9 hours ago

So it turns out.... you can't connect them anymore.... The feature is deprecated

https://docs.cockpit-project.org/cockpit-guide/latest/guide/feature-machines.html

(also, I wanted to just link my response from another user who had asimilar question but for whateVER reasyon, Alx won't let me??)

[–] nosuchanon@lemmy.world 0 points 15 hours ago (1 children)

How do you configure adding multiple servers to the same instance?

[–] happy_wheels@lemmy.blahaj.zone 1 points 9 hours ago

So I actually just went to make a guide on this, demonstrating on my ubuntu 26.04 lab vps, and it appears the connect option is not there.

However on my 24.04 lab vps, it's there. Same with my local 22.04 systems (yes I know I have to upgrade lmao)

I went to cockpit docks and sure enough, the feature is deprecated :( Dang.

https://docs.cockpit-project.org/cockpit-guide/latest/guide/feature-machines.html

It was really cool though while it worked.

[–] i_stole_ur_taco@lemmy.ca 6 points 21 hours ago

I get a text pretty quickly.

[–] dihutenosa@piefed.social 1 points 15 hours ago* (last edited 15 hours ago)

When logged in locally, I use btop to see an overview of what's happening.

Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.

The phone I actually carry around has a ntfy client talking to ntfy server on the aforementioned monitoring solution, so I get buzzes when something goes down.

Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:

  • if disk space > 80% consumed, send a low-priority alert
  • if disk space > 95% consumed, send an urgent alert
  • ... everything else, there's so much to check...
[–] tigerhawkvok@startrek.website 1 points 15 hours ago

A script checks container status and accessibility, and pushes a ntfy alert if it goes down. Healthchecks.io reports if the whole Pi goes down.

I vibe coded a monitor that shows me all the logs and other stuff. And im using netbird cloud (free version) to setup dns so all my services have url "servicexyz.home.internal".

So i just connect to netbird vpn on my phone and go to that url.

You can pretty much use any of the other tools mentioned in this thread with this kind of setup, heck you can write a simple fastapi server that just runs "top" and return it (it would be like 20-30 lines of python) over a url that you can from anywhere as long as you are connected to your vpn

[–] corsicanguppy@lemmy.ca 1 points 15 hours ago

SNMP.

The old engine was nagios. The new engine will be telegraf -> mqtt -> brokerfest -> timescale -> Prometheus.

I guess. Not sure yet. The brokerfest is a series of broker pubsub between sites to both split out data from the stream for off-site CC, or pull it's own subscriptions in. Data could go host <- telegraf -> broker -> remote broker -> remote timescale -> remote Prometheus

[–] phanto@lemmy.ca 3 points 20 hours ago

I sync my podcasts hourly with Nextcloud and poddersync, so when my servers go down, my phone beeps at me.

[–] Limeade3425@lemmy.zip 4 points 21 hours ago

I started with uptime kuma since it was simple and light on resources. I've switched to checkmk community (formerly raw) and really enjoying it. It's telling me so many things that are wrong that I missed. Like that one time I made a snapshot of a VM then forgot about it for 113 days...

[–] CameronDev@programming.dev 4 points 22 hours ago

Uptime Kuma, only really care if it's up or not. Got about 30 services monitored, sending notifications to Signal.

[–] ryannathans@aussie.zone 0 points 14 hours ago (1 children)

I use FreeBSD so I don't need to check it to make sure it's working properly

load more comments (1 replies)
load more comments
view more: ‹ prev next ›