// Engineering Log

Centralized logging: Part 1 — Why collect logs in one place

Published on 2026-09-22

// Fast route

This article belongs to the topic Deploy and reliability.

Metrics show something is wrong with the system: response time has increased, memory is exhausted, the share of successful requests has dropped. To understand why this happened, you need logs — records of events written by the operating system, web server, database, and the application itself. A log records what happened, when, on which server, and in which component.

While there is one server, logs can be read directly on it. When the number of servers and services grows, this approach stops working, and logs are collected in one place.

Why logs scattered across servers are bad

  • They are hard to collect. To investigate one failure you have to SSH into several machines and search for the required files.
  • Events cannot be correlated. A request passed through a load balancer, the application, and the database, but records about it are in three different places and in three different formats.
  • Search is slow. grep over files on a dozen servers works while logs are small; as volume grows it becomes impractical.
  • Logs get lost. Files are rotated and deleted, and if a server fails or is compromised, the records on it can disappear or be altered exactly when they are most needed.

What centralized collection provides

Centralized logging is the collection of logs from all servers and services into a common storage where you can search them, build graphs, and set up alerts.

  • One access point. All records are available in a web interface; you don’t need to access individual servers.
  • Fast search. Filters by time, host, service, severity level, and content.
  • Event correlation. Records from different components are visible on a single timeline; if the application forwards a request identifier, its path can be fully traced.
  • Durability. A copy of logs is stored separately from the servers that generated them.
  • Alerts. A notification arrives when the number of errors grows or a suspicious entry appears, for example a series of failed logins.

What the system consists of

Any centralized logging system performs the same steps:

  1. Collection. An agent on the server reads log files, the system journal, or container output. Examples of agents: Filebeat, Fluent Bit, Grafana Alloy, Vector.
  2. Transport. Records are sent to storage via syslog, HTTP, or the system’s native protocol.
  3. Processing. Lines are parsed into fields, unnecessary data is discarded, host name, environment, and service name are added.
  4. Storage and indexing. Records are stored so they can be searched quickly for the needed period.
  5. Search and visualization. Queries, dashboards, reports.
  6. Alerts. Rules by which the system notifies about a problem.

To collect logs from remote clients you can do without a full stack — such an option is discussed in the article “How to organize log collection from remote clients over HTTP”.

Structured logs

Most time during implementation is spent on parsing lines. A line like

2026-09-22 10:15:03 ERROR Payment failed for order 1842: timeout

has to be parsed with regular expressions, and any change in format breaks the parser. If the application writes logs in JSON from the start, the fields are already separated:

json
{"time":"2026-09-22T10:15:03Z","level":"error","service":"payments","order_id":1842,"msg":"payment failed","error":"timeout"}

With such a record you can filter all errors of the payments service or find everything related to order 1842 without extra configuration. Practical rules:

  • one event — one line;
  • time in UTC and in ISO 8601 format;
  • consistent field names across all services: level, service, msg, request_id;
  • the request identifier is passed between services and appears in every record.

Severity levels

The syslog standard (RFC 5424) defines eight levels — from 0 (emergency, system unusable) to 7 (debug, debugging messages). Applications usually use a shortened set: debug, info, warning, error, critical.

The debug level is better turned off in production: it increases log volume many times, and with it the storage cost. If debug entries are needed to investigate a specific problem, they are turned on temporarily.

How long to retain logs

Retention is determined by the task, not by free disk space:

  • Operational logs for incident investigation are usually needed for the last days or weeks.
  • Security logs — logins, permission changes, administrator actions — are kept longer to investigate an incident discovered months later.
  • Debug logs are sufficient to keep for a few days.

In all systems in this cycle the retention period is set by policy: old data is deleted automatically. Without such a policy the storage will eventually fill up.

Personal data in logs

Personal data easily ends up in logs: email addresses, phone numbers, names, form contents. If a log contains personal data, its collection and storage become processing of personal data, and the requirements of Federal Law No. 152-FZ apply. In particular, according to part 7 of Article 5 of the law, personal data are stored no longer than the purposes of processing require, unless the period is established by federal law or contract.

Practically this means:

  • do not write passwords, tokens, card numbers, or full form contents to logs;
  • mask or remove personal data during processing, before it reaches storage;
  • restrict access to logs to the people who really need it;
  • set a retention period and delete data after it expires.

Common mistakes

  • Logs without a unified format. Each service writes in its own way, and you have to write a custom parser for each.
  • No retention policy. Storage grows until the disk is full, after which it stops accepting new records.
  • Debug level enabled in production. Log volume grows many times with no benefit.
  • Secrets in logs. Tokens and passwords in logs are accessible to anyone who has access to the logging system.
  • Storage on the same server. If the server fails, logs needed for investigation disappear with it.

Which system to choose

This series covers four common systems:

  • ELK Stack (Elasticsearch, Logstash, Kibana) — full-text search and advanced analytics, but high resource requirements.
  • OpenSearch — an open fork of Elasticsearch and Kibana with similar capabilities.
  • Graylog — a ready-made platform for logs with streams, processing rules, and access control.
  • Loki and Grafana — economical storage by indexing only labels, convenient alongside Prometheus.

Need to collect logs in one place?

I will choose a system based on volume and requirements, configure collection, storage, and alerts. Contact me on Telegram.

Написать в Telegram →

// Similar task

If you are dealing with something similar

This article belongs to one of the main working topics. You can keep reading on the topic, go to the homepage to understand what I do, or open the service pages directly.

Article topic

Deploy and reliability

Docker, CI/CD, releases, monitoring, observability, and incident handling.

Typical tasks behind this topic

  • Set up deployment without manual chaos
  • Add monitoring, alerts, and baseline observability
  • Investigate incidents and stabilize production

// Next step

If you need help with this topic, not just another article, it is better to go straight to the service page. The homepage and topic collection stay available as secondary routes.

Open services

// Contact

Need help?

Get in touch with me and I'll help solve the problem

I reply within one business day (03:00-13:00 GMT)

Или оставьте заявку здесь:

Confirm that you are not a bot.

Write and get a quick reply