Slowyslowyapp.com

03 Bufferbloat

Active queue management

Dropping or marking a packet early is the signal that makes senders back off, which is the whole fix.

Entry 03 of 04Where the delay actually comes from.All of Bufferbloat

A screen showing a latency graph with a sharp drop after a change
The drop after a change is the signal arriving early enough to be useful.slowyapp.com picture kit

The signal that travels backward

A queue is a waiting room, and every waiting room has a failure mode: when it fills completely, nothing else can enter. The simplest queues do exactly one thing when they hit capacity — they drop the next arriving packet. This is called tail-drop, and it works in the narrowest sense: the queue does not overflow. But tail-drop is a bad signal. It fires too late, it fires on everything at once, and by the time it fires the queue has already been full for long enough to impose significant delay on every packet that got in before the cutoff. The queue has become the problem it was supposed to prevent.

Active queue management (AQM) is the family of techniques that replace the passive tail-drop reflex with deliberate, early intervention — dropping or marking packets before the queue fills, in a way that sends a usable signal back to the sender. The insight is simple: a packet drop is not a disaster to be deferred; it is a communication. The sender, running congestion control, interprets a drop as evidence of a congested path and reduces its transmission rate. AQM turns that mechanism from an emergency brake into a continuous feedback loop.

How early dropping became a specification

The first formal treatment of the idea came from Sally Floyd and Van Jacobson, then both at Lawrence Berkeley National Laboratory in Berkeley, California. Their 1993 paper introduced Random Early Detection — RED — which probabilistically drops packets as a queue's average occupancy rises, well before the queue is full. The probability of dropping increases with measured queue depth, so heavily loaded flows receive more drops and back off more, while lightly loaded flows are rarely touched. The mechanism uses a weighted moving average of queue length rather than instantaneous depth, which smooths out short bursts and avoids penalising flows that happen to arrive during a momentary spike.

A long queue of identical objects on a conveyor in a plain industrial setting
A line of identical objects waiting for a narrow exit is the whole mechanism.slowyapp.com picture kit

RED was immediately influential and became part of the recommended architecture for Internet routers, eventually codified by the Internet Engineering Task Force. But deployment revealed a difficulty: RED requires careful parameter tuning. The target queue range, the drop probability slope, and the weight applied to the moving average all interact, and the right values depend on link speed, expected load, and round-trip times — quantities that vary between deployments and change over time. In practice, many operators either misconfigured RED or simply left it disabled, falling back to tail-drop on enormous buffers. This is one thread that leads to what Jim Gettys later named bufferbloat: the condition in which cheap memory and long queues produce multi-second latency spikes on otherwise fast links, because the queue absorbs everything rather than signalling back.

CoDel and the target-delay approach

The answer to RED's tuning problem came in work by Kathleen Nichols and Van Jacobson, published in 2012 in ACM Queue. Their algorithm, Controlled Delay — CoDel — replaced queue-length measurement with sojourn time: how long an individual packet actually waited inside the queue. The key observation is that a queue doing its job correctly should be able to drain almost immediately; a packet that sits in a queue for more than a small target interval is evidence that the queue has grown beyond what the path needs. CoDel measures that per-packet waiting time and begins dropping when the minimum sojourn time in a sliding window exceeds a fixed target — typically five milliseconds — for longer than a fixed interval. Once dropping begins, the drop rate increases as a function of time, so persistent congestion triggers progressively stronger signals without requiring any per-flow state.

The appeal of CoDel is its near-absence of operator-adjustable parameters. The target and interval have sensible defaults that work across a wide range of link speeds, and the algorithm adapts to whatever the actual path looks like rather than requiring advance knowledge of it. Nichols and Jacobson described the design in a queue-centric paper that is unusually accessible — the argument builds from physical intuition before arriving at the mathematics, which is characteristic of Jacobson's published style.

Marking instead of dropping: ECN

Dropping a packet is an unambiguous signal, but it is also wasteful. Explicit Congestion Notification — ECN — provides an alternative: rather than discarding a packet, a router sets a bit in the packet header, and the receiving endpoint reflects that mark back to the sender in an acknowledgement. The sender interprets the mark as it would a drop, and reduces its rate, but the packet itself arrives and is used. ECN is defined by the IETF and requires cooperation from both endpoints; it cannot be used unless the sender and receiver have negotiated its use at the start of a connection. Where it is available, AQM algorithms can mark rather than drop, preserving throughput while still delivering the congestion signal. CoDel supports ECN marking and uses it by default when endpoints have negotiated it.

FQ-CoDel and the flow-isolation layer

CoDel operates on a single queue, which means a single heavy flow can still crowd out smaller ones before the algorithm responds. The practical deployment of CoDel in Linux and in open-source router firmware combines it with fair queuing — specifically, stochastic fair queuing — to produce FQ-CoDel. Traffic arriving at the shaper is sorted into per-flow buckets using a hash of the packet's five-tuple, and each bucket is served in turn. CoDel runs independently inside each bucket. The result is that a large file transfer and a voice call are isolated from each other; the file transfer's queue can grow and receive drops while the voice call's queue stays short. Luigi Rizzo's work on fair queuing in the dummynet shaper, developed in Pisa, Italy, is part of the intellectual lineage here — the idea of per-flow isolation predates CoDel by more than a decade.

Dave Taht and the community around Bufferbloat.net were central to getting FQ-CoDel into the router firmware and kernel paths that most users actually run. The algorithm is now present in the Linux traffic control subsystem and in OpenWrt, which powers a large fraction of home and small-office routers. Apple integrated similar ideas into its own networking stack, adopting L4S-adjacent approaches as the standard evolved.

Every AQM algorithm is ultimately doing one thing: making the congestion signal arrive earlier and more selectively than tail-drop would deliver it.

What the signal actually does

Every AQM algorithm is ultimately doing one thing: making the congestion signal arrive earlier and more selectively than tail-drop would deliver it. The sender does not need to know the queue's depth, the link's speed, or even that a queue exists — it only needs to see that packets are being dropped or marked, and to reduce its rate accordingly. The congestion-control logic in TCP and its successors handles the rest. AQM is the part of the system that ensures that signal is timely and proportionate rather than catastrophic and late. A queue that never fills is a queue that never imposes unnecessary delay, which is the whole point: the fix is not a faster link, it is a smarter waiting room.

Several parallel cable runs of different colours entering a switch
Parallel runs into one switch: several flows, one departure point, one decision about order.slowyapp.com picture kit