---
title: "How to monitor a VPS 24/7: metrics, alerts and uptime | StreetHosting"
description: "Set up 24/7 VPS monitoring: external uptime checks, CPU, RAM, disk and network metrics, alert thresholds, backup heartbeats and tools for every size."
url: "https://streethosting.com.br/en/guides/vps/monitor-vps-24-7"
type: "page"
language: "en-US"
---

VPS · 10 min · Intermediate

Published on Sep 28, 2026 · Updated on Sep 28, 2026

# 24/7 VPS monitoring strategy: what to measure and when to alert

An uptime monitor on its own will not tell you the disk is filling up or that backups stopped three weeks ago. Here is how to combine availability, resources, scheduled jobs and alerts into a routine that works day and night.

By [Equipe StreetHosting](https://streethosting.com.br/en/autores#equipe-streethosting) · StreetHosting infrastructure and support team

[Monitoring and alerts](https://streethosting.com.br/en/guides/topics/monitoring) [Hardware and datacenter](https://streethosting.com.br/en/guides/topics/hardware) [Linux administration](https://streethosting.com.br/en/guides/topics/linux)

Summarize with:

[](https://chat.openai.com/?q=Summarize%20the%20key%20points%20of%20this%20StreetHosting%20guide%3A%20https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7.%20Highlight%20the%20step-by-step%20instructions%2C%20the%20prerequisites%20and%20the%20most%20common%20mistakes. "ChatGPT") [](https://claude.ai/new?q=Summarize%20the%20key%20points%20of%20this%20StreetHosting%20guide%3A%20https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7.%20Highlight%20the%20step-by-step%20instructions%2C%20the%20prerequisites%20and%20the%20most%20common%20mistakes. "Claude") [](https://www.google.com/search?udm=50&aep=11&q=Summarize%20the%20key%20points%20of%20this%20StreetHosting%20guide%3A%20https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7.%20Highlight%20the%20step-by-step%20instructions%2C%20the%20prerequisites%20and%20the%20most%20common%20mistakes. "Google AI Mode") [](https://x.com/i/grok?text=Summarize%20the%20key%20points%20of%20this%20StreetHosting%20guide%3A%20https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7.%20Highlight%20the%20step-by-step%20instructions%2C%20the%20prerequisites%20and%20the%20most%20common%20mistakes. "Grok") [](https://www.perplexity.ai/search/new?q=Summarize%20the%20key%20points%20of%20this%20StreetHosting%20guide%3A%20https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7.%20Highlight%20the%20step-by-step%20instructions%2C%20the%20prerequisites%20and%20the%20most%20common%20mistakes. "Perplexity")

Share:

[](https://x.com/intent/tweet?text=How%20to%20monitor%20a%20VPS%2024%2F7%3A%20metrics%2C%20alerts%20and%20uptime&url=https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7 "Share on X") [](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7 "Share on Facebook") [](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7 "Share on LinkedIn") [](https://wa.me/?text=How%20to%20monitor%20a%20VPS%2024%2F7%3A%20metrics%2C%20alerts%20and%20uptime%20https%3A%2F%2Fstreethosting.com.br%2Fen%2Fguides%2Fvps%2Fmonitor-vps-24-7 "Share on WhatsApp")

For agents: Copy as Markdown [.md](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7.md)

In this guide 8 sections

* [The four layers of monitoring](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#camadas)
* [Metrics and alert thresholds](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#metricas-limites)
* [Availability seen from outside](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#monitor-externo)
* [Resources, services and history](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#metricas-internas)
* [Backups, cron and certificates](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#falhas-silenciosas)
* [Alerts that only fire when they should](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#alertas-sem-ruido)
* [Tools for every size](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#stack-por-tamanho)
* [Where to run the monitoring](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#onde-rodar)

Quick answer

To **monitor a VPS 24 hours a day**, combine four layers: an availability monitor running outside the VPS, internal CPU, RAM, disk and network metrics with history, heartbeats for scheduled jobs such as backups and cron, and alerts with clear thresholds sent to a channel someone actually reads. Uptime Kuma outside the machine, plus Netdata or Prometheus with Grafana inside it, cover most cases.

## The four layers of monitoring[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#camadas)

Monitoring is not installing a tool. It is answering four different questions, and each one calls for a different kind of check. Whoever covers only one of them finds out about the other failures from a customer complaint.

* **Availability:** does the service respond to someone on the outside? HTTP, TCP or ping check run from another machine, at short intervals.
* **Resources:** does the machine have headroom? CPU, memory, disk, I/O and network with history, so you know whether the problem started just now or has been growing for days.
* **Services:** are the processes that matter up? systemd units, containers, database, job queue.
* **Jobs and deadlines:** did what was supposed to happen on schedule actually happen? Backups, cron routines, certificate renewal.

The most common gap is covering the first layer and stopping there. The site has an uptime monitor, so nobody notices the disk crossed 85% the week before; it only shows up when the database stops writing and the site goes down. Or the backup fails every night for three weeks, silently, until the day someone needs it. The next sections cover each layer with what to measure, where to measure it and when to start alerting.

## Metrics and alert thresholds[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#metricas-limites)

The thresholds below are reasonable starting points for a Linux VPS running a web application, a bot or a game server. Use them for two weeks, compare against the history and adjust: the right threshold is the one that would have fired during past incidents and stayed quiet on normal days.

| Metric                    | Warning                                   | Critical                          | Why it matters                                               |
| ------------------------- | ----------------------------------------- | --------------------------------- | ------------------------------------------------------------ |
| HTTP response or TCP port | Response time three times above normal    | Three consecutive failures        | Confirming the failure avoids alarms over a one-second blip  |
| CPU usage                 | Above 85% for 15 minutes                  | Above 95% for 30 minutes          | A short spike is normal; sustained saturation is not         |
| 5-minute load average     | Above the number of vCPUs for 15 minutes  | Above twice the number of vCPUs   | Shows a queue of processes waiting for CPU or disk           |
| CPU steal                 | Above 5% for 15 minutes                   | Above 10% on a recurring basis    | Indicates contention for the host's physical processor       |
| Available memory          | Below 15%                                 | Below 5% or swap growing          | Near the end, the OOM killer terminates processes            |
| Disk, space used          | Above 80%                                 | Above 90%                         | Database and logs stop writing when the disk is full         |
| Disk, inodes used         | Above 80%                                 | Above 90%                         | Millions of small files exhaust inodes before space runs out |
| Disk latency              | Average wait above 20 ms                  | Above 100 ms for several minutes  | On NVMe, a few milliseconds is what to expect                |
| Network                   | Sustained traffic above 70% of the uplink | Packet loss on the external check | A large transfer or an attack saturates the link             |
| TLS certificate           | Less than 14 days to expiry               | Less than 7 days                  | Sign of an automatic renewal that failed                     |
| Backup                    | No success in 26 hours                    | No success in 48 hours            | Daily routine that stopped running                           |

Two metrics deserve special attention on a VPS. The first is load average, which only makes sense compared to the number of vCPUs: a load of 4 is fine on 8 vCPUs and a bottleneck on 2. The second is steal, the time the virtual machine wanted CPU and waited on the hypervisor. It does not show up on physical machines and is the first number to look at when the VPS gets slow with no heavy process in sight. The topic is covered in detail in [what CPU steal is on a VPS](https://streethosting.com.br/en/guides/vps/what-is-cpu-steal-vps).

## Availability seen from outside[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#monitor-externo)

The most ignored rule in monitoring: the availability check cannot live on the machine it watches. If the VPS freezes, loses its network or gets stuck in a reboot, the monitor goes down with it, and the silence looks like good news. The options, from simplest to most robust:

* **A small VPS just for monitoring:** runs Uptime Kuma and checks all the other machines. It costs little and stays under your control.
* **An external uptime service:** useful as a second opinion, because it checks from other networks.
* **Both together:** your own monitor covers the details (ports, keyword, push) and the external one confirms the problem is not a routing issue between the datacenter and the users.

The type of check matters too. An HTTP check that only looks at the 200 status code can pass while the application shows a friendly error page. Prefer checking for a word that only appears when the page really rendered, or a health endpoint that queries the database. For games, databases and SSH, use a TCP port check. Installation and each monitor type are covered in [monitoring with Uptime Kuma](https://streethosting.com.br/en/guides/vps/uptime-kuma-vps-monitoring).

Create a route in the application such as /health that tests the database connection and returns 200 only when everything responds. Nginx stays up when the application dies, so checking only the static page hides the outage that matters.

## Resources, services and history[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#metricas-internas)

For resources, there are two mature paths. Netdata installs in minutes, collects metrics every second and ships with ready-made alerts; it works very well for one to a few VPSs, and the starting point is [monitoring resources with htop and Netdata](https://streethosting.com.br/en/guides/vps/monitor-vps-resources-htop-netdata). Prometheus with node\_exporter and Grafana takes more work, but centralizes several machines, accepts custom queries and keeps long history; the step by step is in [installing Grafana and Prometheus on a VPS](https://streethosting.com.br/en/guides/vps/install-grafana-prometheus-vps).

Whatever the tool, keep at least 15 to 30 days of history. The most useful question in an incident is since when. Memory that climbs a little every day is a leak; memory that jumped at 3 AM is a scheduled job. Without history, the two look the same.

For services, the first step is to have systemd itself restart whatever goes down. In the units for your applications, include an automatic restart:

`[Service] Restart=on-failure RestartSec=5`

An automatic restart fixes the symptom, not the cause. A service that restarts twenty times a day is still broken, just silently. So keep an eye on failed units and journal errors:

`systemctl --failed journalctl -u minha-app -p err --since "1 hour ago"`

Filters by service, date and priority, and what to look for in each file under /var/log, are in [how to analyze logs on a Linux VPS](https://streethosting.com.br/en/guides/vps/analyze-linux-vps-logs). In Prometheus, the node\_exporter systemd collector is disabled by default; enable it with `--collector.systemd` if you want Grafana to alert on stopped units.

## Backups, cron and certificates[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#falhas-silenciosas)

The worst failures make no noise. The backup script that errors out every night, the certificate that does not renew because port 80 was closed in a firewall change, the cron job that vanished after a migration. None of them takes the site down today. All of them send the bill later.

The fix is to flip the logic with a heartbeat: instead of the monitor asking, the job itself reports that it ran. At the end of the script, a call to the monitor's push URL; if the signal does not arrive within the expected window, the alert fires. Uptime Kuma's Push monitor type does exactly that.

`#!/usr/bin/env bash set -euo pipefail /usr/local/sbin/backup-restic.sh # only reached if the backup finished without errors curl -fsS -m 10 --retry 3 "https://status.seu-dominio.com.br/api/push/SEU_TOKEN?status=up&msg=OK" > /dev/null`

With `set -e`, any error stops the script before curl. No signal, no heartbeat, alert fires. It is more reliable than sending a notice only on failure, because it also covers the scenario where the script never ran at all. The full backup setup with retention and restore testing is in [automated VPS backups with restic and cron](https://streethosting.com.br/en/guides/vps/vps-backup-restic-cron).

For certificates, turn on the expiry warning in the HTTP monitor and run `sudo certbot renew --dry-run` once in a while to confirm renewal still works. For the domain, note the expiration date and leave auto-renewal on at the registrar.

## Alerts that only fire when they should[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#alertas-sem-ruido)

Too many alerts are as bad as none. When the phone rings for every blip, the team learns to ignore it, and the real alert slips by. A few rules keep the signal clean:

* **Two severity levels:** critical wakes someone up (service down, disk above 90%); warning goes to a channel read during business hours (sustained high CPU, certificate at 14 days).
* **Confirmation before alerting:** two or three consecutive failures, with checks every 30 to 60 seconds on critical monitors.
* **Recovery notice:** notify when it comes back too, so nobody investigates something that already resolved itself.
* **Maintenance window:** pause alerts during planned updates, instead of turning the monitor off and forgetting to turn it back on.
* **The right channel:**the team's Discord or Telegram for day to day, email as a record. Never the public community channel. The setup is in [downtime alerts on Discord or Telegram](https://streethosting.com.br/en/guides/infrastructure/server-downtime-alerts-discord-telegram).

Finally, every alert needs an owner and a written first step. One line is enough: disk above 90%, run du on the root, check logs and Docker images. At 3 in the morning, nobody thinks straight without a runbook.

* Availability monitor running outside the VPS
* Metrics with at least 15 days of history
* Heartbeat on the backup and on critical cron jobs
* Certificate expiry warning turned on
* Critical alerts separated from warnings
* Owner and first step defined for every alert
* Alert test done by taking a service down on purpose

## Tools for every size[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#stack-por-tamanho)

There is no reason to build a large company's stack to watch a Discord bot. Start with the minimum that covers the four layers and grow when the number of machines calls for it.

| Scenario                                    | Availability                                  | Resources                                                  | Jobs and deadlines                            |
| ------------------------------------------- | --------------------------------------------- | ---------------------------------------------------------- | --------------------------------------------- |
| One VPS with a site, API or bot             | Uptime Kuma on another machine                | Netdata on the VPS itself                                  | Backup push and certificate warning           |
| Two to ten VPSs                             | Uptime Kuma on a monitoring VPS               | node\_exporter on each VPS, central Prometheus and Grafana | Push from every backup and every critical job |
| Operation with customers or an on-call team | Uptime Kuma plus a check from another network | Prometheus and Grafana with alerts by severity             | Push, certificates and a public status page   |

Test the alert before trusting it. Stop a service on purpose at a quiet time and time how long the notification takes to arrive. A monitor that was configured and never tested usually fails over a silly detail, such as a wrong webhook or email landing in spam.

## Where to run the monitoring[](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#onde-rodar)

The monitoring server has a simple profile: little CPU, stable memory, fast disk for the history and, above all, it needs to stay up when the other machines go down. A small VPS kept separate from the applications does this well, and the [Xeon VPS](https://streethosting.com.br/en/vps/xeon) line delivers more vCPUs for the money, which is what this kind of workload benefits from.

The R$ 26.00 per month plan, with 2 vCPUs, 2 GB of RAM and 20 GB of NVMe, runs Uptime Kuma and a Prometheus for a few machines. The R$ 43.00 one, with 3 vCPUs, 4 GB and 40 GB, fits Uptime Kuma, Prometheus and Grafana with 30 days of retention for a few dozen servers. If the operation grows, the R$ 60.00 plan has 4 vCPUs, 6 GB and 60 GB. All of them sit in São Paulo, with Enterprise Anti-DDoS, a 1 Gbps uplink, 20 ms average latency in Brazil and activation within 60 seconds. When the history needs more room, upgrading through the control panel increases memory, vCPUs and disk and charges only the prorated difference for the cycle.

An honest caveat: a monitoring VPS in the same datacenter as the applications detects service outages, freezes and resource shortages, but cannot see a routing problem between São Paulo and your users. For that layer, keep a check from another network. With both in place, the phone rings before the first customer complains.

In this guide

* [The four layers of monitoring](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#camadas)
* [Metrics and alert thresholds](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#metricas-limites)
* [Availability seen from outside](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#monitor-externo)
* [Resources, services and history](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#metricas-internas)
* [Backups, cron and certificates](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#falhas-silenciosas)
* [Alerts that only fire when they should](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#alertas-sem-ruido)
* [Tools for every size](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#stack-por-tamanho)
* [Where to run the monitoring](https://streethosting.com.br/en/guides/vps/monitor-vps-24-7#onde-rodar)

## Frequently asked questions

What is the best tool to monitor a VPS?

No single tool covers everything. For availability, use an external monitor such as Uptime Kuma running outside the VPS. For resources, Netdata handles one or a few machines and Prometheus with Grafana scales better across many. What matters is covering availability, resources and scheduled jobs.

Can I monitor the VPS from inside the VPS itself?

CPU, memory and disk metrics, yes. Availability, no: if the VPS freezes or loses its network, the monitor goes down with it and nothing gets sent. The uptime check has to run on another machine, preferably on another network.

At what disk usage percentage should I get an alert?

A good starting point is a warning at 80% and a critical alert at 90%, watching inodes as well. On small disks, such as 20 GB, prefer a threshold in free GB, because 10% is only 2 GB and a runaway log eats that in a few hours.

How do I know the backup ran without checking every day?

Use a heartbeat. At the end of the backup script, make a call to the monitor's push URL; if the signal does not arrive within the expected interval, you get an alert. With set -e in the script, the call only happens when the backup finished without errors.

Does monitoring slow down the VPS?

When configured properly, the impact is small. A node\_exporter uses a few MB of RAM and Netdata is light for what it delivers. The heavy part is storing long history on a small VPS; in that case, move Prometheus and Grafana to a separate machine.

Next step

See Xeon VPS

Xeon VPS for steady workloads, automation and long-running projects.

[See Xeon VPS](https://streethosting.com.br/en/vps/xeon)

[See VPS plans Root VPS in Brazil with NVMe and Anti-DDoS.](https://streethosting.com.br/en/vps)

## Related guides

[VPS Intermediate How to install Uptime Kuma on a VPS and monitor sites and ports Uptime Kuma checks sites, APIs, ports and scheduled jobs at set intervals and alerts you when something goes down. Here is how to install it with Docker, serve it over HTTPS, pick the right monitor type and turn on alerts. 8 min Read guide](https://streethosting.com.br/en/guides/vps/uptime-kuma-vps-monitoring) [VPS Intermediate Install Grafana and Prometheus on a VPS with node\_exporter node\_exporter publishes the machine's numbers, Prometheus scrapes them and keeps the history, and Grafana turns it all into graphs. Here is how to build the stack on Ubuntu 24.04 without leaving sensitive ports open to the internet. 10 min Read guide](https://streethosting.com.br/en/guides/vps/install-grafana-prometheus-vps) [VPS Intermediate How to analyze Linux VPS logs with journalctl and grep Almost every server problem leaves a trace in some log. See where each record lives on Ubuntu, how to slice by service, time and severity, and how to keep logs from filling the disk. 9 min Read guide](https://streethosting.com.br/en/guides/vps/analyze-linux-vps-logs)

[← Back to the Guide Center](https://streethosting.com.br/en/guides)
