03 Bufferbloat
Cheap memory, long queues
Memory got cheap, so equipment was given enormous buffers, and rather than dropping a packet it now queues it for seconds.

When the buffer became the problem
For most of networking's early history, memory was expensive enough to be a genuine constraint. A router had a small buffer, and when traffic arrived faster than it could leave, packets were dropped — quickly, and without ceremony. That seems brutal, but it was also useful: a dropped packet told the sender to slow down, and the whole system responded. The queue stayed short. Delay stayed low.
Then memory prices collapsed.
The fall was not subtle. Through the 1990s and into the 2000s, the cost of DRAM dropped by orders of magnitude, following a curve that made it economically rational to stuff every piece of networked equipment with far more buffer than it had any engineering reason to need. Router vendors, modem makers, and chip designers all responded the same way: they gave their hardware enormous queues, because the memory cost almost nothing and because — on the surface — a device that drops fewer packets looks better than one that drops more. A packet queued is a packet not dropped. Marketing logic pointed straight at maximum buffer depth.

The engineering consequences took a while to surface. A queue that is deep enough never needs to drop — but it also never sends the congestion signal that TCP congestion control depends on. Van Jacobson's work at Lawrence Berkeley National Laboratory in the late 1980s had established that TCP would respond to loss by backing off: the sender detects a missing acknowledgement and cuts its sending rate. That feedback loop only works if the loss signal actually arrives. A buffer that queues packets for hundreds of milliseconds, or several seconds, delays the signal so long that by the time the sender gets it, it has already pushed far more data into the network than the link can carry. The queue grows; the delay grows; everyone on that link suffers.
Jim Gettys, working at home on a slow DSL connection around 2010, noticed that his connection felt terrible even when his throughput meter showed it was full. Video calls broke apart, interactive traffic stalled, and yet bytes were clearly moving. He started measuring — carefully, methodically — and found that the round-trip times on his connection were ballooning to values that had nothing to do with propagation delay. Packets were sitting in a queue inside his DSL modem for seconds at a time. The modem had been given a large buffer, the buffer was full, and every interactive packet had to wait behind a large backlog of bulk transfer data before it could get out.
Gettys named the problem bufferbloat, and the name stuck. He wrote it up publicly and pushed it into conversations at the Internet Engineering Task Force, where it eventually attracted serious attention. The scale turned out to be systemic: the same over-buffering was present in home routers, in DSL and cable modems, in cellular base stations, in the switches inside data centres, and in end-host network stacks. Equipment built on the assumption that more memory was always better had quietly made latency across the consumer internet far worse than the raw link speeds would predict.
The queue's inner life
To understand why a large buffer hurts, it helps to think carefully about what a queue actually is. A queue exists at any point where packets arrive faster than they can depart. The outbound link has a fixed capacity — a rate at which it can drain the queue — and everything that arrives above that rate waits. A shallow queue fills quickly, drops a packet, and lets the sender know something is wrong. A deep queue absorbs the excess silently, holding packets in order, draining them at line rate when it can.
The problem is time. A packet sitting at the back of a long queue is waiting not for transmission to complete — that is microseconds — but for every packet ahead of it to be transmitted first. On a slow link with a deep buffer, the wait can be substantial. A home broadband upload link running at a few megabits per second, filled with a large file transfer and backed by a buffer sized for high-speed core routing, can impose queuing delays that dwarf propagation delay entirely. The wire is fast; the queue is the bottleneck.
What makes this particularly damaging is the interaction with congestion control. TCP assumes that a lost packet means the network is congested. Bufferbloat converts congestion into delay instead of loss: the network is congested, but no packets are dropped, so TCP never backs off. Multiple flows pile into the same buffer. The delay climbs for all of them. Interactive traffic — a DNS lookup, a keystroke over SSH, a voice packet — is indistinguishable inside the queue from bulk data. It waits its turn.
If packets are spending too long in the queue, CoDel drops; if they move through quickly, it does not.
The response the field developed is active queue management: algorithms that intervene inside the queue before it is full, deliberately dropping or ECN-marking packets to signal congestion early. Kathleen Nichols and Van Jacobson published CoDel — Controlled Delay — in 2012, a scheme that targets not queue length but queue sojourn time, the time a packet actually spends waiting. If packets are spending too long in the queue, CoDel drops; if they move through quickly, it does not. The algorithm was described in a paper at ACM Queue and represents a clean break from threshold-based AQM that had come before: it adapts to the link rate automatically, without needing to know what that rate is.
Dave Taht and the community that coalesced around Bufferbloat.net drove much of the practical deployment work, combining CoDel with fair-queuing disciplines to produce fq_codel and CAKE — algorithms that both manage queue depth and separate traffic into flows so that one bulk transfer cannot monopolise the queue at the expense of interactive packets. Luigi Rizzo's work in Pisa contributed foundational ideas about queue scheduling that underpin much of this. Sally Floyd's earlier research on random early detection had established the principle that dropping before a queue is full is better than tail-drop; the CoDel work built on that foundation while solving the problem that RED required manual tuning that almost nobody got right.
The deeper irony of bufferbloat is that the extra memory achieved essentially nothing useful. A buffer large enough to queue packets for seconds cannot recover from a lost packet more reliably than one that queues them for milliseconds — TCP will retransmit either way. What the oversized buffer accomplishes is hiding congestion, punishing interactive traffic, and making the network feel worse to the person sitting at the keyboard while delivering identical or worse aggregate throughput. The cheap memory made the queue longer without making the link faster. It just changed where the delay comes from.
The problem is not fully solved. Cellular networks, home-router firmware, and cable modem chipsets remain inconsistent about whether they deploy any AQM at all. Buffer sizes inside end-host stacks still vary widely. But the mechanism is now understood, the literature is clear, and the tools to fix it exist — which is a different situation from 2010, when the problem did not yet have a name.
