// Engineering Log
Monitoring: Part 5 — VictoriaMetrics
Published on 2026-09-22
// Fast route
This article belongs to the topic Deploy and reliability.
VictoriaMetrics — a time-series database compatible with Prometheus. It is used as a long-term storage for Prometheus metrics or as a complete replacement for it — together with the vmagent collector and the vmalert alerting module. The main edition is distributed under the Apache 2.0 license; some features (for example, downsampling of old data and different retention periods for different datasets) are available only in the paid Enterprise edition.
Why you need it
Prometheus stores data on the local disk of a single server, by default for 15 days. When you need to keep metrics for a year, aggregate data from multiple Prometheus instances in one place, or save disk space, a separate storage is used. VictoriaMetrics accepts data via the remote_write protocol and responds to queries through the same HTTP API as Prometheus, so Grafana connects to it as a normal Prometheus data source without changing dashboards.
Editions and operating modes
- Single-node — a single executable that receives, stores, and serves data. Suitable for most small and medium deployments and simplest to operate.
- Cluster — separate components
vminsert(ingest),vmstorage(storage), andvmselect(queries). Scales horizontally and supports multiple tenants on one deployment.
MetricsQL
VictoriaMetrics’ query language is called MetricsQL. It is backward compatible with PromQL: Grafana dashboards built for Prometheus work without changes. In addition, MetricsQL can:
- omit the window in square brackets —
rate(node_network_receive_bytes_total)will pick it based on the query step and the scrape interval; - extract repeated parts of a query into
WITH (...)templates; - use additional functions:
range_median,range_linear_regression,range_trim_outliers, and others.
Stack without Prometheus: vmagent, VictoriaMetrics and vmalert
A stack where all components are from VictoriaMetrics:
- vmagent on the collection server reads the usual
prometheus.yml(theglobalandscrape_configssections), scrapes targets and sends data to the storage. If the storage is unavailable, it stores data in a disk buffer and resends after connectivity is restored. - VictoriaMetrics single-node stores data and serves queries on port 8428.
- vmalert executes alerting and recording rules in Prometheus format and sends triggered alerts to Alertmanager.
- Grafana connects to VictoriaMetrics as a Prometheus data source.
Example of running components on one server:
# storage: data in /var/lib/victoria-metrics, keep for 12 months
victoria-metrics -storageDataPath=/var/lib/victoria-metrics -retentionPeriod=12
# scraper: targets from prometheus.yml, send to storage
vmagent -promscrape.config=/etc/vmagent/prometheus.yml \
-remoteWrite.url=http://localhost:8428/api/v1/write
# alerts: rules in Prometheus format, notifications via Alertmanager
vmalert -rule='/etc/vmalert/rules/*.yml' \
-datasource.url=http://localhost:8428 \
-notifier.url=http://localhost:9093 \
-remoteWrite.url=http://localhost:8428 \
-remoteRead.url=http://localhost:8428The -retentionPeriod parameter sets the retention period; a number without a suffix means months, the default is one month. You can also specify other units, for example -retentionPeriod=90d. The -remoteWrite.url and -remoteRead.url flags in vmalert are needed to persist alert state and recording rule results; for alerting only, -datasource.url and -notifier.url are sufficient.
Wrap the path pattern in -rule in quotes, otherwise the shell will expand it into a list of files. Alerting rules from Prometheus are transferred unchanged: vmalert understands the same file format.
If Prometheus is already running
You don’t have to migrate completely. Add a copy of data to be sent in prometheus.yml:
remote_write:
- url: http://192.0.2.30:8428/api/v1/writePrometheus continues to work as before, while VictoriaMetrics accumulates the long-term history. You can then shorten the local retention period in Prometheus.
Advantages
- Disk and memory savings. Developers and many users note that the same volume of metrics occupies noticeably less space than in Prometheus; the exact gain depends on the data, so test it on your own metrics.
- Simplicity of single-node. One process, one data directory, backups via snapshots (
/snapshot/create) and thevmbackuputility. - Compatibility. Exporters, Grafana dashboards, and Prometheus rules work without changes.
- Ingests data in various formats: Prometheus remote_write, InfluxDB line protocol, Graphite, OpenTSDB.
Disadvantages
- Separate components. VictoriaMetrics is the storage; collection, alerting, and visualization are vmagent, vmalert, Alertmanager, and Grafana, each with its own configuration.
- Some features are Enterprise-only. Before designing your architecture, compare required features with the list of Enterprise features in the documentation.
- Minor differences with PromQL. Compatibility is backward, but some functions behave differently; complex queries should be verified when migrating.
Common mistakes
- Retention period not set. By default data is kept for one month, which can be an unpleasant surprise when switching to a “long-term storage”.
- Port 8428 exposed to the outside. The storage API allows reading, writing, and deleting data. Restrict access by network or by reverse proxy with authentication.
- vmagent buffer on a small disk. With prolonged storage unavailability, the buffer fills the collection server disk; limit its size with the
-remoteWrite.maxDiskUsagePerURLparameter. - High cardinality. As with Prometheus, unique label values quickly increase the number of series and resource usage.
// Similar task
If you are dealing with something similar
This article belongs to one of the main working topics. You can keep reading on the topic, go to the homepage to understand what I do, or open the service pages directly.
Article topic
Deploy and reliability
Docker, CI/CD, releases, monitoring, observability, and incident handling.
Typical tasks behind this topic
- Set up deployment without manual chaos
- Add monitoring, alerts, and baseline observability
- Investigate incidents and stabilize production
// Next step
If you need help with this topic, not just another article, it is better to go straight to the service page. The homepage and topic collection stay available as secondary routes.
Open services// Contact
Need help?
Get in touch with me and I'll help solve the problem
I reply within one business day (03:00-13:00 GMT)
Или оставьте заявку здесь:
// Related