// Engineering Log
Redundancy of communication channels: Part 4 — Internet connectivity: BGP, DNS failover and CDN
Published on 2026-09-22
// Fast route
This article belongs to the topic Networking and routing.
For a website, online store or API the main risk is loss of Internet connectivity: customers stop seeing the service. For an office this is a halt in working with cloud services. Redundancy for Internet egress is built differently depending on what you need to protect: outbound access for employees or incoming client connections to your servers.
Two providers without BGP
The most affordable scheme: two providers, each with its own address, and a router with automatic failover.
The router checks not just the presence of the link to the provider, but Internet reachability beyond it. For this they ping an external address reachable only via the specific provider. In MikroTik RouterOS this is done with recursive routes:
/ip route
add dst-address=8.8.8.8/32 gateway=203.0.113.1 scope=10
add dst-address=0.0.0.0/0 gateway=8.8.8.8 target-scope=11 check-gateway=ping distance=1
add dst-address=0.0.0.0/0 gateway=198.51.100.1 distance=2The first line sends the check address only via the primary provider (203.0.113.1). The default route with distance=1 points to that address and is checked with ping: RouterOS sends the check once every 10 seconds and after two failed replies considers the gateway unreachable. Then traffic moves to the route via the backup provider (distance=2). Details are in the MikroTik documentation, section Failover (WAN Backup).
Limitations of this scheme:
- Switchover takes tens of seconds, and open connections are dropped: the second provider has a different external address.
- Incoming connections use two different addresses. Replies must go out via the provider through which the request arrived — see “MikroTik: return traffic via the same gateway”. To ensure clients can find the service when one provider fails you need DNS with health checks (below).
For an office this is usually enough. For public services — not always.
BGP multihoming: same addresses through two providers
To keep your addresses reachable if any provider fails you need your own Autonomous System (AS) and your own block of addresses, which you announce to both providers via BGP. The whole Internet sees two paths to your network and will use the other if one fails.
What is required
- Autonomous System number. In the RIPE NCC region (Europe, the Middle East, Russia) policy requires a network to be connected to at least two providers and to have its own routing policy. An AS number is obtained via RIPE NCC membership or via a sponsoring LIR — an intermediary organization, for example a provider.
- An IPv4 address block. RIPE NCC has run out of free IPv4 addresses. New members who have not previously received IPv4 can join a waiting list for one /24 block (256 addresses); the wait lasts more than a year. Applications are checked against EU sanctions lists. In practice addresses are leased or bought on the secondary market, or you get a block from one of the providers with its consent to announce it via the second — this is discussed with providers in advance.
- A router with BGP support and an agreement with both providers for a BGP session.
How quickly BGP switches
A common misconception is “BGP switches in seconds.” If the failure is not visible at the link level (for example, equipment failed beyond the peering point), BGP learns about it via the hold timer. RFC 4271 suggests a default of 90 seconds, and some vendors use even larger values. In addition, changes need to propagate across the Internet.
Detection can be sped up with BFD (RFC 5880) — a lightweight protocol that exchanges packets at intervals of hundreds of milliseconds and notifies BGP of failures almost immediately. It must be supported/negotiated with the provider. On your side you can reduce BGP timers, but too aggressive values lead to false positives.
DNS failover
If there are multiple servers (or one server with two providers and different addresses), failover can be done at the DNS level: a DNS service checks address health and returns only working addresses to clients.
- Health checks are required. Simple delivery of several A records in round-robin does not track failures: some clients will continue to get the address of a non-working server.
- Short TTL (60–300 seconds), otherwise clients will keep the old address until cache expiry. Some resolvers and browsers cache longer, so switching is not instantaneous.
- DNS itself is also redundant: the zone is served by multiple servers in different networks.
DNS failover does not require your own AS and is suitable when minute-level failover is acceptable.
CDN
A content delivery network stores copies of static files (and sometimes pages) on servers worldwide and serves them to users from the nearest node. When the origin server is briefly unavailable the CDN can continue serving cached content, and it also absorbs part of DDoS attacks.
There is an important limitation for a Russian audience. As of 9 June 2025 Russian providers, including Rostelecom, MegaFon, Beeline and MTS, limit traffic to sites behind Cloudflare: users receive only the first 16 KB of each resource, and sites stop working. Cloudflare confirmed this and stated that it cannot circumvent the restriction on its side. If your audience is in Russia, choose a CDN with nodes in Russia.
Clouds and multiple locations
For services with high availability requirements they reserve not only channels but also the servers themselves: application copies run in different availability zones of a cloud or in different data centers, and traffic between them is distributed by a load balancer or DNS. This is a question of application architecture and data replication, not just connectivity.
What to choose
| Task | Reasonable solution |
|---|---|
| Office, outbound access | Two providers using different technologies, router with health checks |
| Small site on a single server | Hosting in a data center with redundant channels, DNS with health checks |
| Service on multiple servers | DNS failover or a load balancer, CDN for static content |
| Service where seconds are critical | Own AS and addresses, BGP with two providers, BFD |
Common mistakes
- Checking the link instead of Internet reachability. The link to the provider exists, but there is no Internet beyond it, and the router does not fail over.
- Round-robin DNS instead of failover.
- Long TTL on records for a service that needs fast failover.
- BGP without BFD expecting instantaneous switching.
- CDN unavailable to the target audience.
- Backup channel without monitoring: its failure is discovered on the day of the main outage.
// Similar task
If you are dealing with something similar
This article belongs to one of the main working topics. You can keep reading on the topic, go to the homepage to understand what I do, or open the service pages directly.
Article topic
Networking and routing
MikroTik, VPN, routing, DNS, BGP, connectivity, and access troubleshooting.
Typical tasks behind this topic
- Set up VPN and secure access to office or cloud
- Fix routing, DNS, or unstable connectivity
- Configure MikroTik, firewall, and external links
// Next step
If you need help with this topic, not just another article, it is better to go straight to the service page. The homepage and topic collection stay available as secondary routes.
Open services// Contact
Need help?
Get in touch with me and I'll help solve the problem
I reply within one business day (03:00-13:00 GMT)
Или оставьте заявку здесь:
// Related