// Engineering Log
Communication Channel Redundancy: Part 2 — Inside the Building: Network Interface Cards, Switches, bonding and LACP
Published on 2026-09-22
Most common network failures happen not at the provider but inside the building: a pulled patch cable, a burned switch port, a failed power supply, a faulty SFP module. Protecting against these is the cheapest, and you should start redundancy at this level.
Two NICs and two switches
Basic scheme for a fault-tolerant server connection:
- the server has at least two network ports, preferably on different network interface cards (NICs), so that the failure of one card does not disable both ports;
- each port is connected to its own switch;
- switches are powered from different UPS units or different power lines, and in the rack — from different power supplies if there are two.
If both server cables go to the same switch, this protects against a cable break and a port failure, but not against the failure of the switch itself.
Bonding: combining interfaces in Linux
In Linux several physical interfaces are combined into a single logical interface — a bond. Applications and the IP address operate with the bond interface, while the driver distributes traffic across physical ports and switches it over on failure. Different vendors call the same thing teaming, EtherChannel or link aggregation.
Bonding driver modes according to the Linux kernel documentation:
| Mode | Name | What it does | Is switch configuration needed |
|---|---|---|---|
| 0 | balance-rr | Sends packets over interfaces in round-robin | Yes (static aggregation) |
| 1 | active-backup | One interface active, others in backup | No |
| 2 | balance-xor | Chooses interface by hash | Yes (static aggregation) |
| 3 | broadcast | Sends everything over all interfaces | Depends on topology |
| 4 | 802.3ad | Dynamic aggregation per IEEE 802.3ad (LACP) | Yes, LACP |
| 5 | balance-tlb | Balances outgoing traffic | No |
| 6 | balance-alb | Balances outgoing and incoming IPv4 traffic | No |
In practice two modes are used.
- active-backup (1) — the simplest and most compatible. Works with any switches, including when ports are connected to two different switches without shared management. Throughput equals one port.
- 802.3ad (4), LACP — the switch and server negotiate aggregation, traffic is distributed across all ports and a failed port is excluded automatically. Requires LACP support on the switch, and when connected to two switches — they must be combined into a stack or MLAG (see below).
Why a single flow doesn’t get faster
A common misconception: “two 1 Gbit/s channels give 2 Gbit/s.” In 802.3ad mode this is true only for the sum of many connections. The interface for each packet is chosen by a hash of addresses, and all packets of one connection go through one port. The kernel documentation explicitly states that no single connection can use more bandwidth than one interface.
By default the hash is computed by MAC addresses (layer2), and all traffic between the server and the router may go through one port. The layer3+4 policy considers IP addresses and ports, so different connections to the same address are distributed across different interfaces. Copying one large file will still go at the speed of a single port.
How bond detects failures
The miimon parameter sets how often the driver checks the link status on a physical interface, in milliseconds. The documentation recommends starting at 100 ms. MII monitoring only sees the state of its own port: if the port is present but the fault is further down the chain, it will not detect it. In 802.3ad mode such cases are additionally caught by the LACP protocol itself, and for other modes there is ARP monitoring.
Example: LACP in Netplan
On Ubuntu the network is configured via Netplan. Example of a bond of two ports in 802.3ad mode with a static address:
network:
version: 2
renderer: networkd
ethernets:
enp1s0: {}
enp2s0: {}
bonds:
bond0:
interfaces: [enp1s0, enp2s0]
addresses: [192.0.2.10/24]
routes:
- to: default
via: 192.0.2.1
nameservers:
addresses: [192.0.2.53]
parameters:
mode: 802.3ad
lacp-rate: fast
mii-monitor-interval: 100
transmit-hash-policy: layer3+4Before applying on a remote server use sudo netplan try: if the connection drops, the settings will be rolled back automatically. On the switch the corresponding ports are combined into an LACP group (port-channel), otherwise a bond in 802.3ad mode will not come up. For setups without switch support replace mode: 802.3ad with mode: active-backup and remove lacp-rate.
VLANs, bridges and tunnels over bond are discussed in detail in the article “Netplan: advanced network configuration”.
Stack and MLAG: aggregation across two switches
Classic LACP works between two devices. To connect a server to two switches and still use both ports, switches are combined in one of the following ways:
- stack — several switches are managed as one; if one stack member fails the others continue to operate, but upgrading software sometimes requires rebooting the whole stack;
- MLAG (multi-chassis link aggregation) — two independent switches present themselves to the server as one LACP peer. Implementations differ between vendors and are not interoperable.
If switches do not support either stack or MLAG, use active-backup: it protects against switch failure without cooperation from the switches.
Inter-switch links: loops and STP
When switches are connected with multiple cables for redundancy, loops are formed, and without protection the network will go down from a broadcast storm. Protocols STP, RSTP and MSTP disable extra links and enable them when the primary fails. RSTP restores connectivity faster than original STP. It is more reliable to use link aggregation between switches, and keep STP as protection against switching errors.
Cables, modules and power
- Copper and fiber. Twisted-pair copper is cheaper and convenient inside the rack; fiber is needed for long distances and is immune to electromagnetic interference. For redundant links between floors or buildings it’s convenient if the lines follow different paths.
- Keep spare SFP modules on site: their failure is a common cause of an optical line break.
- Storage. If the server uses network storage via iSCSI or Fibre Channel, configure multipath access (multipath): the server sees a single disk via multiple paths and survives the failure of one of them.
- Power. A switch without a UPS will shut down on the first power glitch along with all redundant links.
Common mistakes
- LACP on the server but not on the switch (or vice versa). Bond doesn’t come up or works on a single port, and the problem only appears on failure.
- Expecting doubled speed for a single copy or a single database.
- Two server ports in one switch while being sure the server is “redundant.”
- Configuring the network remotely without a rollback.
netplan applywith an error in the bond cuts off access to the server;netplan tryprevents this. - No testing. During a maintenance window disconnect each cable and each switch in turn and make sure the server remains reachable.
// Contact
Need help?
Get in touch with me and I'll help solve the problem
I reply within one business day (03:00-13:00 GMT)
Или оставьте заявку здесь:
// Related