// Engineering Log

Monitoring: Part 1 — Why You Need It and What to Measure

Published on 2026-09-22

// Fast route

This article belongs to the topic Deploy and reliability.

Monitoring is the continuous collection and analysis of data about the operation of servers, networks, databases, and applications. Its goal is to detect a problem before users do: to see that disk space is running out, error rates are rising, or site response has slowed, and to intervene before the service stops working.

Why monitoring is needed

  • Early detection of problems. Most incidents develop gradually: a disk fills up over several days, memory leaks over several hours. Monitoring reveals such trends in advance.
  • Finding bottlenecks. Data on CPU, memory, disks, and database response times show what exactly limits performance.
  • Capacity planning. Metric history shows when a more powerful server or an additional disk will be needed.
  • Fast recovery. The earlier an alert arrives and the more detailed the failure data, the faster it is resolved.
  • Security. A sharp increase in traffic, failed logins, or outgoing connections can be a sign of an attack.

Types of monitoring

  • System — CPU load, memory, disks, network, uptime of servers and virtual machines.
  • Network — state of routers and switches, packet loss, latency, channel utilization.
  • Application monitoring (APM) — response time, database queries, exceptions inside the code.
  • User-side monitoring — real user interactions (RUM) and synthetic checks, where an external service regularly opens the site from different regions and measures availability and speed.

What to measure: three proven approaches

You can collect thousands of metrics, but you should alert on only a few. To avoid drowning in data, established methodologies are used.

Four golden signals. Google engineers in the SRE book highlight four metrics for services that users interact with:

  • latency — request processing time; successful and failed requests should be counted separately: a fast error response should not improve the average;
  • traffic — system load, for example the number of HTTP requests per second;
  • errors — the fraction of failed requests: explicit (500 status), implicit (200 status with incorrect content), and those that violate rules (response slower than the set threshold);
  • saturation — how “full” the system is, primarily by the most constrained resource.

The USE method (Brendan Gregg) is applied to resources — CPU, memory, disks, network. For each resource check three things: utilization (what fraction of time the resource was busy), saturation (how much work is waiting in queues), and errors. According to the author, this check finds most server problems with little time investment.

The RED method (Tom Wilkie, 2015) — for services that handle requests: rate of requests, errors, and duration of processing. If you build dashboards by RED for each service, they all look the same, and the on-duty engineer can understand even a service they didn’t write.

In practice USE describes servers, while RED and the golden signals describe the services running on them.

Alerting without unnecessary noise

The most common mistake when implementing monitoring is too many alerts. After a week they stop being read, and a real incident gets lost among false alarms. A few rules help avoid this.

  • Every alert should require action. Google SRE guidance says: if an alert can be responded to mechanically, it should not wake a person. Such cases are automated or turned into reports.
  • Alert on symptoms, not causes. It’s important for the user that the site returns errors or is slow, not that the CPU is at 90%. High utilization without service impact is a reason to check a graph, but not to wake the on-call engineer at night.
  • Add a firing delay. The condition should hold for several minutes before an alert is sent: short spikes should not trigger alarms.
  • Separate severity levels. Critical alerts — to a messenger or by phone 24/7; warnings — to a work chat or report.
  • Group alerts. If a switch fails, you don’t need twenty messages about every server behind it — one is enough.
  • Review rules. An alert that hasn’t led to action in a month should be disabled or reworked.

Most monitoring systems send alerts to Telegram, email, SMS, and on-call services.

Tools for the monitoring cycle

  • Munin — a simple system with ready-made graphs for a few servers.
  • Prometheus, Node Exporter, and Grafana — pull-based metric collection, a flexible query language, and alerting via Alertmanager.
  • Zabbix — an “all-in-one” system: agents, templates, alerts, and a web interface in one product.
  • VictoriaMetrics — a cost-effective metrics storage compatible with Prometheus, for long-term retention and large volumes.

An example of how monitoring looks in a large fleet can be found in the article about how to automate management of 366 servers.

Need help with infrastructure monitoring?

I will select and configure a stack for your situation. Write to me on Telegram — I will reply on a business day.

Написать в Telegram →

// Similar task

If you are dealing with something similar

This article belongs to one of the main working topics. You can keep reading on the topic, go to the homepage to understand what I do, or open the service pages directly.

Article topic

Deploy and reliability

Docker, CI/CD, releases, monitoring, observability, and incident handling.

Typical tasks behind this topic

  • Set up deployment without manual chaos
  • Add monitoring, alerts, and baseline observability
  • Investigate incidents and stabilize production

// Next step

If you need help with this topic, not just another article, it is better to go straight to the service page. The homepage and topic collection stay available as secondary routes.

Open services

// Contact

Need help?

Get in touch with me and I'll help solve the problem

I reply within one business day (03:00-13:00 GMT)

Или оставьте заявку здесь:

Confirm that you are not a bot.

Write and get a quick reply