// Engineering Log

Communication Channel Redundancy: Part 2 — Inside the Building: Network Interface Cards, Switches, bonding and LACP

Published on 2026-09-22

Most common network failures happen not at the provider but inside the building: a pulled patch cable, a burned switch port, a failed power supply, a faulty SFP module. Protecting against these is the cheapest, and you should start redundancy at this level.

Two NICs and two switches

Basic scheme for a fault-tolerant server connection:

  • the server has at least two network ports, preferably on different network interface cards (NICs), so that the failure of one card does not disable both ports;
  • each port is connected to its own switch;
  • switches are powered from different UPS units or different power lines, and in the rack — from different power supplies if there are two.

If both server cables go to the same switch, this protects against a cable break and a port failure, but not against the failure of the switch itself.

Bonding: combining interfaces in Linux

In Linux several physical interfaces are combined into a single logical interface — a bond. Applications and the IP address operate with the bond interface, while the driver distributes traffic across physical ports and switches it over on failure. Different vendors call the same thing teaming, EtherChannel or link aggregation.

Bonding driver modes according to the Linux kernel documentation:

ModeNameWhat it doesIs switch configuration needed
0balance-rrSends packets over interfaces in round-robinYes (static aggregation)
1active-backupOne interface active, others in backupNo
2balance-xorChooses interface by hashYes (static aggregation)
3broadcastSends everything over all interfacesDepends on topology
4802.3adDynamic aggregation per IEEE 802.3ad (LACP)Yes, LACP
5balance-tlbBalances outgoing trafficNo
6balance-albBalances outgoing and incoming IPv4 trafficNo

In practice two modes are used.

  • active-backup (1) — the simplest and most compatible. Works with any switches, including when ports are connected to two different switches without shared management. Throughput equals one port.
  • 802.3ad (4), LACP — the switch and server negotiate aggregation, traffic is distributed across all ports and a failed port is excluded automatically. Requires LACP support on the switch, and when connected to two switches — they must be combined into a stack or MLAG (see below).

Why a single flow doesn’t get faster

A common misconception: “two 1 Gbit/s channels give 2 Gbit/s.” In 802.3ad mode this is true only for the sum of many connections. The interface for each packet is chosen by a hash of addresses, and all packets of one connection go through one port. The kernel documentation explicitly states that no single connection can use more bandwidth than one interface.

By default the hash is computed by MAC addresses (layer2), and all traffic between the server and the router may go through one port. The layer3+4 policy considers IP addresses and ports, so different connections to the same address are distributed across different interfaces. Copying one large file will still go at the speed of a single port.

How bond detects failures

The miimon parameter sets how often the driver checks the link status on a physical interface, in milliseconds. The documentation recommends starting at 100 ms. MII monitoring only sees the state of its own port: if the port is present but the fault is further down the chain, it will not detect it. In 802.3ad mode such cases are additionally caught by the LACP protocol itself, and for other modes there is ARP monitoring.

Example: LACP in Netplan

On Ubuntu the network is configured via Netplan. Example of a bond of two ports in 802.3ad mode with a static address:

yaml
network:
  version: 2
  renderer: networkd
  ethernets:
    enp1s0: {}
    enp2s0: {}
  bonds:
    bond0:
      interfaces: [enp1s0, enp2s0]
      addresses: [192.0.2.10/24]
      routes:
        - to: default
          via: 192.0.2.1
      nameservers:
        addresses: [192.0.2.53]
      parameters:
        mode: 802.3ad
        lacp-rate: fast
        mii-monitor-interval: 100
        transmit-hash-policy: layer3+4

Before applying on a remote server use sudo netplan try: if the connection drops, the settings will be rolled back automatically. On the switch the corresponding ports are combined into an LACP group (port-channel), otherwise a bond in 802.3ad mode will not come up. For setups without switch support replace mode: 802.3ad with mode: active-backup and remove lacp-rate.

VLANs, bridges and tunnels over bond are discussed in detail in the article “Netplan: advanced network configuration”.

Stack and MLAG: aggregation across two switches

Classic LACP works between two devices. To connect a server to two switches and still use both ports, switches are combined in one of the following ways:

  • stack — several switches are managed as one; if one stack member fails the others continue to operate, but upgrading software sometimes requires rebooting the whole stack;
  • MLAG (multi-chassis link aggregation) — two independent switches present themselves to the server as one LACP peer. Implementations differ between vendors and are not interoperable.

If switches do not support either stack or MLAG, use active-backup: it protects against switch failure without cooperation from the switches.

When switches are connected with multiple cables for redundancy, loops are formed, and without protection the network will go down from a broadcast storm. Protocols STP, RSTP and MSTP disable extra links and enable them when the primary fails. RSTP restores connectivity faster than original STP. It is more reliable to use link aggregation between switches, and keep STP as protection against switching errors.

Cables, modules and power

  • Copper and fiber. Twisted-pair copper is cheaper and convenient inside the rack; fiber is needed for long distances and is immune to electromagnetic interference. For redundant links between floors or buildings it’s convenient if the lines follow different paths.
  • Keep spare SFP modules on site: their failure is a common cause of an optical line break.
  • Storage. If the server uses network storage via iSCSI or Fibre Channel, configure multipath access (multipath): the server sees a single disk via multiple paths and survives the failure of one of them.
  • Power. A switch without a UPS will shut down on the first power glitch along with all redundant links.

Common mistakes

  • LACP on the server but not on the switch (or vice versa). Bond doesn’t come up or works on a single port, and the problem only appears on failure.
  • Expecting doubled speed for a single copy or a single database.
  • Two server ports in one switch while being sure the server is “redundant.”
  • Configuring the network remotely without a rollback. netplan apply with an error in the bond cuts off access to the server; netplan try prevents this.
  • No testing. During a maintenance window disconnect each cable and each switch in turn and make sure the server remains reachable.

// Contact

Need help?

Get in touch with me and I'll help solve the problem

I reply within one business day (03:00-13:00 GMT)

Или оставьте заявку здесь:

Confirm that you are not a bot.

Write and get a quick reply