// Engineering Log

Monitoring: Part 4 — Zabbix

Published on 2026-09-22

// Fast route

This article belongs to the topic Deploy and reliability.

Zabbix — an “all-in-one” monitoring system: data collection, storage, triggers, alerts, dashboards and access control are combined in a single product. It’s chosen when you need to monitor servers, network equipment, databases and websites from a single interface without assembling the system from separate components.

Versions and license

  • Zabbix 7.0 LTS (June 2024) — long-term support branch: full support until June 30, 2027, security fixes until June 30, 2029. It’s sensible to choose this for production systems.
  • Zabbix 7.4 (June 2025) — current standard release with new features and a shorter support window.
  • Zabbix 6.0 LTS receives only security fixes until February 2027 — time to plan an upgrade.

Since version 7.0 Zabbix is distributed under the AGPLv3 license; before that GPLv2 was used since 2001. For a company that runs Zabbix internally, there is no difference. AGPL terms matter if you modify the code and provide Zabbix as a service to others.

What Zabbix consists of

  • Zabbix Server — the central process: accepts data, evaluates triggers, runs actions and sends alerts.
  • Database — stores configuration and history. Zabbix 7.4 supports MySQL and Percona (8.0.30 and newer), MariaDB (10.5 and newer), PostgreSQL (13 and newer) and the TimescaleDB extension for PostgreSQL. Oracle is no longer supported, SQLite is allowed only for proxies.
  • Web interface in PHP — configuration, dashboards, reports.
  • Agents on monitored hosts.
  • Proxies — intermediate collectors for remote sites.

Zabbix agent 2

There are two agents. The classic one is written in C. Zabbix agent 2 is a next-generation agent in Go that gradually replaces the first:

  • all checks are performed by plugins, and checks from different plugins run in parallel;
  • ready-made plugins for PostgreSQL, MySQL, Redis, MongoDB, Docker and other systems do not require additional scripts;
  • a persistent buffer (EnablePersistentBuffer) preserves data from active checks across restarts and connection losses.

Minimal configuration in /etc/zabbix/zabbix_agent2.conf:

Server=192.0.2.10
ServerActive=192.0.2.10
Hostname=web01.example.ru

Server — addresses allowed for passive checks (the server polls the agent). ServerActive — where the agent sends data in active mode. Hostname must match the host name in the web interface. Agent, server and web interface packages are installed from the official repository following the instructions on zabbix.com/download: the commands there are tailored to the Zabbix version, distribution and DBMS.

Templates

A template is a ready-made set of items, triggers, graphs and discovery rules. You attach, for example, the template “Linux by Zabbix agent” to a host and it immediately gets dozens of metrics and sensible triggers: low disk space, high load, reboots, agent unavailability.

Low-level discovery (LLD) rules automatically find filesystems, network interfaces, services and create metrics for each. Official templates exist for operating systems, DBMSs, web servers, SNMP network equipment and cloud services.

Custom templates should be built on top of the standard ones and stored as an export (YAML) to transfer between installations.

Proxy

Zabbix Proxy collects data on behalf of the server and forwards it in batches. It’s needed when:

  • hosts are located in another office or datacenter and the connection to the server is unreliable: the proxy accumulates data and sends it after the channel is restored;
  • there are thousands of monitored hosts and load needs to be distributed;
  • hosts are in an isolated network where the server cannot connect directly.

In version 7.0 proxy groups with load balancing and high availability appeared: hosts are assigned to a group, and if one proxy fails other proxies pick them up.

Other capabilities

  • Data collection by different methods: agent, SNMP, IPMI, HTTP requests, SSH, ODBC, JMX for Java applications.
  • Network discovery and auto-registration: new devices and agents are added automatically according to rules.
  • Web scenarios and browser checks: in 7.0 synthetic browser-based monitoring with screenshots was introduced.
  • Escalations: if a problem is not resolved within a set time, the alert is sent to the next level.
  • Two-factor authentication in the web interface (TOTP and Duo) since 7.0.

Advantages

  • Everything in one product, with a unified interface and access management.
  • Hundreds of ready-made templates, including for network equipment.
  • Trigger dependencies: when a switch fails you don’t get alerts for all servers behind it.
  • Mature documentation and a large community.

Disadvantages

  • Database load. On large installations history takes the most space and creates the main load. TimescaleDB, reasonable retention periods for history and trends, and increased collection intervals for infrequent data help.
  • Complexity of initial setup. You need to understand hosts, items, triggers and actions.
  • Less flexible data handling than PromQL: complex label-based analytics are more convenient in a stack with Prometheus and Grafana. Grafana can connect to Zabbix as a data source via a plugin.

Common mistakes

  • One host — manual set of checks. Checks are configured on each host instead of in a template; after six months configurations diverge.
  • History kept forever. The database size grows to hundreds of gigabytes. Keep detailed history for a week, and use trends for long periods.
  • All triggers at the same severity. The on-call receives a stream of “catastrophes” and stops reacting.
  • Upgrading without checks. Major version changes alter the database schema; before upgrading make a backup and test on a copy.

Need help with monitoring or DevOps infrastructure?

I'll set up Zabbix, Prometheus or another stack for your task. Write — we'll sort it out.

Написать в Telegram →

// Similar task

If you are dealing with something similar

This article belongs to one of the main working topics. You can keep reading on the topic, go to the homepage to understand what I do, or open the service pages directly.

Article topic

Deploy and reliability

Docker, CI/CD, releases, monitoring, observability, and incident handling.

Typical tasks behind this topic

  • Set up deployment without manual chaos
  • Add monitoring, alerts, and baseline observability
  • Investigate incidents and stabilize production

// Next step

If you need help with this topic, not just another article, it is better to go straight to the service page. The homepage and topic collection stay available as secondary routes.

Open services

// Contact

Need help?

Get in touch with me and I'll help solve the problem

I reply within one business day (03:00-13:00 GMT)

Или оставьте заявку здесь:

Confirm that you are not a bot.

Write and get a quick reply