// Engineering Log
Monitoring: Part 3 — Prometheus, Node Exporter and Grafana
Published on 2026-09-22
// Fast route
This article belongs to the topic Deploy and reliability.
Prometheus — an open-source monitoring system and time-series database (Apache 2.0 license). It was created at SoundCloud in 2012, and in 2016 the project joined the Cloud Native Computing Foundation, second after Kubernetes. Version 3.0 was released in November 2024; as of September 2026 the current version is 3.14, and 3.13 has been announced as a Long Term Support (LTS) release.
Prometheus is commonly used together with Node Exporter, which exposes Linux server metrics, Alertmanager, which sends alerts, and Grafana, where dashboards are built.
How Prometheus works
The main feature is the pull model: Prometheus scrapes targets over HTTP at a set interval and pulls metrics in a text format. Each metric is a name and a set of labels, for example node_cpu_seconds_total{cpu="0",mode="idle"}. Labels allow filtering and grouping of data in queries.
Ecosystem components:
- Prometheus Server — collects metrics, stores them on local disk, executes queries in PromQL, and evaluates alerting rules.
- Exporters — programs that translate system data into Prometheus format:
node_exporterfor servers,postgres_exporterandmysqld_exporterfor databases,blackbox_exporterfor checking websites and ports from the outside. - Alertmanager — receives fired alerts, groups them, suppresses duplicates, and sends to Telegram, email, or an on-call service.
- Pushgateway — an intermediary for short-lived jobs, such as nightly backups, which finish before Prometheus can scrape them.
Minimal installation
Node Exporter is run on every server; by default it exposes metrics on port 9100. Prometheus itself reads configuration from prometheus.yml:
global:
scrape_interval: 30s
evaluation_interval: 30s
rule_files:
- /etc/prometheus/rules/*.yml
alerting:
alertmanagers:
- static_configs:
- targets: ["localhost:9093"]
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["localhost:9090"]
- job_name: node
static_configs:
- targets:
- "192.0.2.21:9100"
- "192.0.2.22:9100"Port 9100 should be accessible only to the Prometheus server: metrics reveal a lot about the system’s configuration.
Alerting rule
Rules are stored in separate files. An example of two typical rules — instance is down and disk space is running out:
groups:
- name: node
rules:
- alert: InstanceDown
expr: up == 0
for: 5m
labels:
severity: critical
annotations:
summary: "{{ $labels.instance }} has not responded for 5 minutes"
- alert: DiskSpaceLow
expr: |
node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} < 0.10
for: 15m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }}: less than 10% free on {{ $labels.mountpoint }}"The for parameter sets how long a condition must hold before firing — this filters out short spikes. Rule files are checked before applying with the command promtool check rules.
Route in Alertmanager
Alertmanager decides who and how to send an alert. Critical alerts go to Telegram immediately, others go to email with grouping:
route:
receiver: email
group_by: ["alertname", "instance"]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- receiver: telegram
matchers:
- severity="critical"
group_wait: 10s
receivers:
- name: email
email_configs:
- to: "admin@example.ru"
- name: telegram
telegram_configs:
- bot_token: "ТОКЕН_БОТА"
chat_id: 123456789To send email, the global block should additionally specify the SMTP server (smtp_smarthost, smtp_from, and credentials).
Grafana
Prometheus has a simple built-in interface for queries, but dashboards are built in Grafana. It connects Prometheus as a data source, supports dozens of other sources, and allows building interactive dashboards with variables and alerts. There are pre-made dashboards for Node Exporter in the Grafana catalog; they can be imported by ID.
Since April 2021 Grafana is distributed under the AGPLv3 license (previously — Apache 2.0). For internal use this changes nothing, but if you modify Grafana’s code and provide it as a service to others, you will need to open your changes.
Advantages
- PromQL. A flexible query language: rates of change, percentiles, aggregation by labels.
- Exporters for almost everything. Databases, web servers, queues, network equipment via SNMP.
- Service discovery. Prometheus can discover targets in Kubernetes, Consul, clouds, and via DNS.
- Mature alerting. Alertmanager can group, suppress, and route notifications.
- Recording rules. Expensive queries can be precomputed and saved as new metrics.
Drawbacks
- Steep learning curve. You need to learn configuration, PromQL, and the label design.
- Local storage. By default data is kept for 15 days on the disk of a single server. For long-term storage and multiple servers, external storages are used: VictoriaMetrics, Thanos, Mimir.
- Pull model. Targets behind NAT or in isolated networks are hard to scrape; agents that send data (remote write) or Pushgateway help.
- High cardinality. Labels with unique values — user IDs, full URLs — sharply increase the number of series and memory usage.
Common mistakes
- Unique values in labels. The main cause of memory exhaustion. Put in labels only what you truly need to group by.
- No
fordelay. Alerts fire on every secondary spike. - Monitoring without checking the monitoring itself. If Prometheus goes down, there will be no alerts. You need an external availability check or a second instance.
- Exporters exposed to the internet. Port 9100 without access restrictions exposes versions, disk partitions, and network interfaces.
// Similar task
If you are dealing with something similar
This article belongs to one of the main working topics. You can keep reading on the topic, go to the homepage to understand what I do, or open the service pages directly.
Article topic
Deploy and reliability
Docker, CI/CD, releases, monitoring, observability, and incident handling.
Typical tasks behind this topic
- Set up deployment without manual chaos
- Add monitoring, alerts, and baseline observability
- Investigate incidents and stabilize production
// Next step
If you need help with this topic, not just another article, it is better to go straight to the service page. The homepage and topic collection stay available as secondary routes.
Open services// Contact
Need help?
Get in touch with me and I'll help solve the problem
I reply within one business day (03:00-13:00 GMT)
Или оставьте заявку здесь:
// Related