Writing

Cold caches fail at the worst possible moment

Sep 30, 2026

A cache is a promise that the expensive work has already been done. A cold cache is that same promise, still unkept. Nothing has been fetched yet, so every request pays full price. And the trouble is that the promise comes due at the exact moment you can least afford it, the instant a flood of traffic arrives all at once, asking for things nothing has fetched yet.

We saw this cleanly during live-event testing. When a partner shifted a large audience over to us, response times spiked, up toward half a second, before settling back to where they belonged, under two hundred milliseconds, once the caches filled. The system wasn’t broken. It was cold. Every layer that normally serves from a warm cache was, for those first moments, doing the full expensive work for every request simultaneously, and the latency reflected exactly that.

The spike lands at the worst possible time

Here’s why this is more dangerous than it looks. The cold-cache penalty doesn’t show up during a quiet ramp where you’d barely notice it. It shows up precisely when traffic arrives in a wall, because that’s when the cache is both empty and slammed at the same time. The one scenario where you most need fast responses, the big surge, is the exact scenario that guarantees the caches are cold. The failure mode and the load are perfectly correlated, which is the worst kind of correlation.

Warm the path before the surge

The fix is unglamorous, and it works. Don’t let the surge warm your caches..warm them yourself first. In our case the recommendation was to keep a steady trickle of real traffic flowing through the path continuously, rather than letting it go idle and cold between bursts. A path that’s always lightly used is a path that’s always warm, and a warm path absorbs the surge instead of choking on it.

Generalize it and the principle holds well beyond streaming. Any system with an expensive first request and a cheap cached one has a cold-start cliff hiding in it. Pre-warm before a known spike. Keep critical paths exercised so they never go fully cold. Treat “the cache will fill up once traffic arrives” as the optimistic lie it is, because by the time it fills, your users have already felt the cold version.

Not a footnote, a latent outage

A cold cache isn’t a performance footnote. It’s a latent outage with a timer set to “whenever the most traffic arrives.” Warm the path before the surge, keep it warm between surges, and the spike that would have been your worst moment turns into a non-event nobody notices. The best compliment a cache can get is that no one ever felt it being empty.