Server Monitoring Basics:What to Watch and Which Tools to Use
Which server metrics actually matter (CPU, memory, disk, I/O, network and services), how to set sensible alert thresholds, and the main kinds of monitoring tools for a VPS.

Table of Contents
Server monitoring means continuously measuring a server's health and alerting someone when something goes wrong. For a typical web server, watch CPU and load, memory and swap, disk space, disk I/O, network traffic, and whether key services (web server, PHP, database) and your website respond. Good monitoring tells you about problems before customers do, and gives you the history to understand why they happened.
What to monitor
| Metric | Why it matters | Example alert |
|---|---|---|
| CPU usage and load average | Sustained high load slows every request | Load above vCPU count for 10+ minutes |
| Memory and swap | Running out of RAM causes swapping and killed processes | Available memory below 10%; swap growing |
| Disk space | Full disks crash databases and stop logs and uploads | Any filesystem above 85% |
| Inodes | Running out of inodes stops file creation even with free space | Above 85% |
| Disk I/O wait | Storage bottlenecks | I/O wait high for several minutes |
| Network | Traffic spikes, possible attacks | Unusual inbound traffic |
| Services | Web server, PHP-FPM, database, cache running | Service down |
| HTTP checks | The website actually works | Non-200 response or timeout |
| SSL expiry | Expired certificates break the site | Under 14 days left |
| Backups | Silent backup failures | Last successful backup older than 24 hours |
Interpreting CPU, load and memory numbers is explained in Linux load average, CPU and memory explained.
Two kinds of monitoring
Internal (agent-based) monitoring runs on the server and collects detailed metrics such as CPU, memory, disk and per-process data. It gives depth, but if the server goes down, it goes down too.
External (synthetic) monitoring checks your website or ports from outside, like a visitor would. It catches network problems, DNS issues and full outages. See website uptime monitoring.
Use both: external checks to know that something is wrong, internal metrics to know why.
Types of tools
- Command-line tools for live investigation:
top,htop,free,df,iostat,ss. See how to check server resource usage. - Lightweight agents with dashboards, which install on the server and show real-time and historical metrics.
- Metrics stacks (for example Prometheus with Grafana) for teams running many servers.
- Hosted monitoring services, which combine agents, external checks and alerting.
- Your provider's graphs, which often show CPU, network and disk at the hypervisor level.
Choose based on how many servers you run and who will look at the data.
Setting good alerts
- Alert on symptoms that need action, not on every spike. A CPU spike for 30 seconds is normal; 15 minutes at 100% is not.
- Use durations, for example "for 5 minutes", to avoid flapping alerts.
- Warn early for slow problems such as disk space (warn at 80%, critical at 90%).
- Send alerts where someone will act, and decide who is on call.
- Review alerts monthly: delete noisy ones, add missing ones.
Logs are part of monitoring
Metrics show that something changed; logs often explain what. Keep and rotate:
- web server access and error logs;
- PHP-FPM and application logs;
- database error and slow query logs;
- system logs (
journalctlon systemd systems).
Sending logs to a separate system helps after a crash or a security incident.
Example alert rules for a web server
| Alert | Condition | Severity | First action |
|---|---|---|---|
| Website down | HTTP check fails from 2 locations for 2 minutes | Critical | Check server reachability, web server and PHP services |
| Disk filling | Any filesystem above 85% | Warning | Find and remove large logs, caches or old backups |
| Disk critical | Any filesystem above 95% | Critical | Free space immediately; investigate cause |
| Memory pressure | Available memory below 10% for 10 minutes | Warning | Check for runaway processes, swap growth |
| High load | Load above 2× vCPU count for 15 minutes | Warning | Identify the top processes and traffic |
| Service failed | systemd reports a failed unit | Critical | Read its logs, restart if safe |
| Backup stale | No successful backup in 26 hours | Warning | Check backup job logs |
| Certificate expiring | Less than 14 days | Warning | Check renewal configuration |
Tune thresholds after a few weeks based on what is normal for your server.
Writing a simple runbook
For each alert, write three to five lines: what it means, the first commands to run, how to fix the common causes, and when to escalate (for example to your hosting provider). During an incident at 3 am, a runbook is far more useful than memory. Command references are in how to check server resource usage.
A minimal setup for one VPS
- External HTTP check on the home page and one key page, every minute or two.
- An agent or provider graphs for CPU, memory, disk and network.
- Alerts for disk above 85%, memory pressure, service down and SSL expiry.
- A daily check that backups succeeded.
- A short written runbook: what to check first when each alert fires.
Frequently Asked Questions
How long should I keep monitoring data?
At least a few weeks of detailed data and several months of summaries, so you can compare against normal behaviour and see growth trends.
Does monitoring slow down the server?
A lightweight agent uses very little CPU and memory. Very frequent, heavy checks can add load, so keep intervals sensible.
What should I do when an alert fires?
Follow your runbook: confirm the problem, check recent changes, look at the relevant metric and logs, and fix or escalate. See troubleshooting a slow or unresponsive server.
Related reading
For the fundamentals, see what is a VPS. 24/7 server monitoring is part of a ServerNeed managed VPS.
Sources
Last updated 7 October 2026



