Linux server monitoring means continuously watching a few key numbers (CPU, load average, memory and swap, disk space, disk I/O, network and service health) so you know something is wrong before a user tells you the site is slow or down. For a quick look, the built-in tools are enough: uptime for load, free -h for memory, df -h for disk, vmstat 1 and iostat -x 1 for CPU and I/O, and htop to see which process is to blame.
Real monitoring needs two more things, though: history (so you know what happened at 3 a.m. yesterday) and alerts (so you don't have to keep watching). For most small and medium servers, a lightweight dashboard such as Netdata, an uptime check from outside the server and a small health-check script run by cron cover both. Below we explain what each number means, compare the tools and write that script. The examples are for Ubuntu 24.04.
Short answer: Monitoring a Linux server means continuously tracking CPU and load average, memory and swap, disk space and I/O, network, and service health. For the current state, the built-in
uptime,free -h,df -h,vmstatandhtopare enough. For history and alerts you needsaror Netdata, a cron health-check script, and an uptime check from outside the server.
What to monitor
CPU
vmstat 1 5
In the vmstat output, the columns under cpu are the ones that matter:
us: time spent in user programs. If it's high, your application or database is genuinely doing compute work.sy: time spent in the kernel.wa(iowait): the CPU is idle because it's waiting on the disk. Highwameans the bottleneck is the disk, not the processor.st(steal): on a virtual server, time the host gave your CPU to other machines. If it's consistently high, the problem isn't your server; it's your provider.
What is load average?
uptime
nproc
The last three numbers from uptime are the load average over the past 1, 5 and 15 minutes. Load is the number of processes that are either running or waiting for the CPU; on Linux, processes waiting on the disk (state D) are counted too.
Compare it with the number of cores (nproc). On a 4-core server, a load of 4 means every core is busy and no queue has formed; a load of 12 means that on average eight processes are waiting their turn. Comparing the three numbers shows the trend: if the 1-minute figure is much higher than the 15-minute one, the problem has just started.
High load with low CPU usage usually means processes are stuck waiting on the disk; reach for iostat.
Memory and swap
free -h
Look at the available column, not free. Linux uses spare RAM for disk cache, so a low free value is normal. The real signs of memory pressure are:
availablestays close to zero.- In
vmstat 1, thesiandsocolumns (swap in and out) are constantly non-zero. A little swap in use is fine; constant swapping is the problem. - The OOM killer has killed a process:
journalctl -k | grep -i "out of memory"
Disk: space and I/O
df -h shows space and df -i shows inodes. If the disk is full, the full procedure is in Linux server disk full? How to find and free up space.
Running out of space isn't the only disk problem; slowness is another:
sudo apt install sysstat
iostat -x 1
A %util close to 100 means the disk is busy all the time, and r_await and w_await show the average wait per read and write request (in milliseconds). If these numbers climb, use sudo iotop (from the iotop package) to see which process is reading or writing so much.
Network
ss -s # connection summary
ss -tulpn # what is listening on which port
ip -s link # bytes, errors and dropped packets per interface
A large number of connections in TIME-WAIT, or a sudden jump in connections, can signal unusual traffic or an application that isn't closing its connections properly.
Service health
All the system numbers can look fine while the site is still broken:
systemctl --failed
curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' https://example.com/
If your application has a health endpoint (such as /up in Laravel), check that, since it tests the database connection as well as the web server. Running services under systemd supervision is covered in creating a systemd service.
Tools: from built-in commands to dashboards
Built-in tools and htop
uptime, free, vmstat, df and ss are on every server. htop brings them together on one live screen, and F6 lets you sort processes by CPU or memory. They all share the same weakness: they only show "right now". For more of the basics, see essential Linux commands for server administration.
sar: history without a dashboard
Besides iostat, the sysstat package includes sar, which records system statistics every few minutes. On Ubuntu you need to enable collection:
sudo sed -i 's/^ENABLED="false"/ENABLED="true"/' /etc/default/sysstat
sudo systemctl enable --now sysstat
After a few hours, sar -u shows today's CPU usage, sar -r memory and sar -q load. For earlier days: sar -u -f /var/log/sysstat/saDD, where DD is the day of the month.
Netdata: a lightweight, ready-made dashboard
Netdata is a web dashboard that, with no configuration, builds hundreds of per-second charts for CPU, memory, disk, network and well-known services (nginx, MySQL, Redis and others), and ships with default alerts.
sudo apt install netdata
The dashboard runs on port 19999. Don't open that port to everyone; if you need to view it remotely, allow only your own IP with ufw:
sudo ufw allow from 203.0.113.10 to any port 19999 proto tcp
(Replace 203.0.113.10 with your IP.) The Ubuntu repository version may lag behind the one on Netdata's site; for the latest, follow the install instructions in the Netdata documentation. Netdata itself uses some RAM and CPU, so bear that in mind on very small servers.
Prometheus and Grafana
Prometheus (collecting and storing metrics), node_exporter (per-server metrics) and Grafana (dashboards) together are the industry standard for multiple servers and services. The stack is powerful and flexible, but setting it up and maintaining it is real work, and it's usually overkill for a single server.
Checking from outside the server
Any tool running on the server goes silent when the server itself becomes unreachable. That's why at least one check has to run from somewhere else: an online uptime monitoring service that requests your site every few minutes and alerts you if it gets no response, or an open-source tool such as Uptime Kuma on another server.
Comparison
| Tool | What it shows | History | Alerts | Setup effort | Best for |
|---|---|---|---|---|---|
uptime, free, vmstat, df |
Current state | No | No | None | Troubleshooting right now |
htop and iotop |
Resource-hungry processes | No | No | One package | Finding the culprit |
sar (sysstat) |
CPU, memory, I/O, network | Yes (daily files) | No | Low | "What happened last night?" |
| Netdata | Hundreds of metrics with charts | Yes | Yes | Low | One to a few servers |
| Prometheus + Grafana | Any metric, many servers | Yes | Yes (Alertmanager) | High | Many servers and a team |
| External uptime check | Site availability | Depends on tool | Yes | Low | Every site, no exceptions |
| Cron health-check script | Whatever you write | In the journal | Yes | Low | A simple, custom start |
A small health-check script with cron
This script checks disk, memory, load, a few services and a health URL, and only sends a message when the state changes, so you don't get the same alert every five minutes:
#!/usr/bin/env bash
set -uo pipefail
DISK_LIMIT=85 # percent
MEM_MIN_AVAIL=10 # percent
SERVICES=(nginx php8.3-fpm mysql)
URL="https://example.com/up"
WEBHOOK_URL="${WEBHOOK_URL:-}"
STATE=/var/tmp/health-check.last
problems=()
# Disk space
while read -r use mount; do
use=${use%\%}
(( use >= DISK_LIMIT )) && problems+=("disk $mount ${use}%")
done < <(df -P -x tmpfs -x devtmpfs -x squashfs -x overlay | awk 'NR>1 {print $5, $6}')
# Available memory
mem=$(awk '/^MemTotal/ {t=$2} /^MemAvailable/ {a=$2} END {printf "%d", a*100/t}' /proc/meminfo)
(( mem < MEM_MIN_AVAIL )) && problems+=("memory available ${mem}%")
# Load above twice the number of cores
load=$(cut -d' ' -f1 /proc/loadavg)
cores=$(nproc)
awk -v l="$load" -v c="$cores" 'BEGIN {exit !(l > c * 2)}' && problems+=("load $load on $cores cores")
# Services
for s in "${SERVICES[@]}"; do
systemctl is-active --quiet "$s" || problems+=("service $s down")
done
# Health URL
code=$(curl -s -o /dev/null -w '%{http_code}' --max-time 10 "$URL")
[[ $code == 200 ]] || problems+=("$URL returned $code")
if (( ${#problems[@]} )); then
msg="[$(hostname)] $(printf '%s; ' "${problems[@]}")"
else
msg="[$(hostname)] OK"
fi
last=$(cat "$STATE" 2>/dev/null || true)
if [[ $msg != "$last" ]]; then
printf '%s\n' "$msg" > "$STATE"
logger -t health-check "$msg"
if [[ -n $WEBHOOK_URL ]]; then
curl -s --max-time 10 -H 'Content-Type: application/json' \
-d "{\"text\": \"$msg\"}" "$WEBHOOK_URL" > /dev/null || true
fi
fi
sudo chmod +x /usr/local/bin/health-check.sh
sudo /usr/local/bin/health-check.sh && journalctl -t health-check -n 5
And to run it automatically every five minutes:
WEBHOOK_URL=https://hooks.example.com/your-webhook
*/5 * * * * root /usr/local/bin/health-check.sh
Replace WEBHOOK_URL with the webhook of whatever chat app or service you use; without it, messages are only logged to the journal. Adjust the service list and thresholds to your own server. Cron gotchas, such as the restricted environment and preventing overlapping runs, are covered in cron jobs on Linux.
Keep in mind that this script doesn't replace an external check: if the server goes down, the script doesn't run either.
Where to start
- Today: set up an external uptime check for your site.
- This week: put the health-check script above in cron with your own thresholds, and enable
sysstat. - If you want more: install Netdata for ready-made charts and alerts.
- Once you have several servers: consider Prometheus and Grafana.
And don't confuse monitoring with backups: monitoring tells you something broke; a backup is what brings it back. If the server is new, start with the first hour on a new Linux server.
Frequently asked questions
What is a high load average on Linux?
Compare load with the number of cores (the output of nproc). A load equal to the core count means every core is busy but no queue has formed; if it's consistently above the core count, processes are waiting their turn. The script in this article alerts when load exceeds twice the number of cores.
Why is free memory always low on a Linux server?
Because Linux uses unused RAM for disk cache, which is normal. The right measure is the available column in free -h. A real memory shortage means available stays near zero, the si and so columns in vmstat are constantly non-zero, or the OOM killer has killed a process.
How do I tell if a slow server is disk-bound?
The signs are high wa (iowait) in vmstat, high load with low CPU usage, and in iostat -x 1 a %util near 100 with rising r_await and w_await. Use iotop to find the process doing all that reading or writing.
Is Netdata a good fit for a small server?
It's a good option for one to a few servers, since a single apt install gives you charts and alerts with no configuration. It does use some RAM and CPU itself, which matters on very small servers, and its port 19999 should not be left open to everyone.
Why monitor a server from the outside as well?
Any tool running on the server goes silent when the server is down or unreachable. An online uptime monitoring service, or Uptime Kuma on another server, requesting your site every few minutes reports exactly that case.
Wrap-up
Linux server monitoring starts with a few simple questions: how do CPU and load compare with the core count, how much memory is available and is swap churning, how full and how busy is the disk, and are the services and health URL responding? Built-in tools are enough for "right now", sar and Netdata give you history, and a small cron script plus an external check gets alerts to you. Start simple, and move to heavier tools only when you have a real need.