03 Bufferbloat
Where the delay comes from
The wire is fast. The wait is not.

The wire is not the problem
A copper pair or a strand of glass is, by almost any everyday measure, instant. Light travels roughly 200,000 kilometres per second through fibre — about two-thirds of its speed in vacuum — meaning a packet moving from London to New York crosses the Atlantic in under 40 milliseconds of pure propagation time. For traffic that stays within a city, the physics contribute almost nothing to the delay a user feels. And yet video calls stutter, gaming pings spike, and web pages hang. The wire is not where the time goes.
The time goes into queues. Somewhere between a keyboard and a server, packets are waiting in line, and that wait — not the physics, not the medium — is what makes a connection feel slow.
What a queue actually does
Every device that forwards packets — a home router, a switch, a wireless access point, a server's network interface — has a buffer attached to each outgoing link. That buffer is a queue, and its purpose is to absorb the difference between the rate at which packets arrive and the rate at which the link can send them. If data arrives in a burst faster than the link can drain it, the buffer holds the excess rather than discarding it immediately. Packets are released in order, one by one, as the link becomes free.

This is not a flaw; it is the mechanism. Without buffers, any momentary mismatch between arrival rate and link rate would mean a dropped packet, and the sender would have to retransmit. Buffers trade space for time, smoothing bursts so that individual packets are not lost to a transient surge. The question is never whether to buffer at all, but how much.
The delay a packet accumulates while sitting in a queue is called queuing latency, and it is the dominant variable in most real network paths. Propagation delay is fixed by geography. Transmission delay — the time to actually push bits onto the wire — is fixed by link speed. Queuing delay is the variable, and it is the one that expands when a network is under load.
How buffers got too big
For most of the internet's early history, memory was expensive enough that buffers were modest by necessity, which meant queues stayed short and queuing latency stayed low. As DRAM prices fell through the 1990s and into the 2000s, equipment manufacturers began shipping routers and home gateways with far larger buffers than the network's traffic patterns required. The reasoning was straightforward and well-intentioned: more buffer means fewer drops, and fewer drops feels like better performance.
The problem, which Jim Gettys named bufferbloat around 2010 and which quickly attracted the attention of researchers at Bufferbloat.net, is that oversized buffers do not improve performance — they defer the signal that would have prompted senders to slow down. TCP's congestion-control mechanism, designed by Van Jacobson at Lawrence Berkeley National Laboratory and described in his landmark 1988 paper, relies on detecting loss (or, in later variants, rising delay) as a signal that the network is congested. When a buffer is large enough to absorb seconds of traffic, loss arrives so late — if at all — that the sender has long since filled the queue entirely. Latency balloons. The connection's throughput may look reasonable on a graph while every interactive packet is waiting behind a long train of bulk data.
The arithmetic is not complicated. A home broadband gateway with a 500-millisecond buffer on a 10 Mbit/s uplink can hold around 625 kilobytes. A large file transfer, or a software update running in the background, can fill that buffer and pin it full. Every DNS query, every VoIP packet, every video-call acknowledgement then waits behind the transfer, adding up to half a second of delay at the queue alone, on top of whatever the wire itself contributes.
Where the queues actually sit
The instinct, when diagnosing a slow connection, is to blame the internet at large — the distant server, the undersea cable, the backbone router. In practice, the queue that causes most of the experienced delay is almost always the one closest to the user: the outbound queue on the home router's uplink interface, or the queue inside the wireless access point that sits between a laptop and the router.
Wireless links are particularly prone to deep queuing delays because the medium is shared and variable.
Wireless links are particularly prone to deep queuing delays because the medium is shared and variable. A wireless access point cannot send to all clients simultaneously; it queues for each, and it may also queue in its own driver before the packet reaches the radio. The effective queue can span several layers, and the cumulative delay compounds. Luigi Rizzo's work on dummynet, developed at the University of Pisa, modelled exactly these kinds of multi-stage delay paths, which made it possible to reproduce in a lab what real users were experiencing in the field.
The asymmetry of a typical broadband connection worsens the problem. Upload capacity is usually much narrower than download capacity. A filled upload queue does not just delay outbound data — it delays the TCP acknowledgements that the receiving end sends back to pace the download. Fill the uplink queue and both directions slow down.
The signal that fixes it
Active queue management exists to break this cycle. Rather than allowing a buffer to fill until it drops packets at the tail — too late for congestion control to respond without a significant latency spike — an AQM algorithm can drop or mark packets earlier, when the queue is growing but not yet full. That early signal reaches the sender in time for it to back off before the buffer overflows.
Kathleen Nichols and Van Jacobson developed the CoDel algorithm — Controlled Delay — specifically to target queuing latency rather than queue length. Published through the IETF and described in a 2012 paper for ACM Queue, CoDel measures how long individual packets spend waiting and acts when that time exceeds a threshold, regardless of how many bytes are in the buffer. It distinguishes between a queue that is momentarily busy and one that is chronically full, and it signals only the latter. The result is that a buffer can still absorb legitimate bursts while being prevented from becoming a reservoir of permanent delay. Sally Floyd's earlier work on Random Early Detection had pointed at the same problem from a slightly different angle — mark before the drop, and mark early enough to matter.
The delay is not a mystery and it is not in the wire. It is sitting in a buffer somewhere along the path, usually close to home, waiting for a queue discipline that is smart enough to ask how long it has been there.
