How Web Caching and CDNs Work
The fastest request is the one you never make. Almost every technique for making the web feel quick comes back to that idea: storing a copy of something so you do not have to compute it, fetch it, or send it across the world again. Caching sounds simple, and the basic mechanics are. What makes it interesting is that it contains one of the genuinely hard problems in all of computing, and the whole system is built to manage that problem carefully.
Caching happens in layers
When your browser shows a page, the data for it may have been served from any of several caches, each closer to you than the last. There is the cache inside your own browser. There are caches run by content delivery networks spread around the world. There can be caches sitting in front of the origin server, and caches inside the application itself. A request travels outward only as far as it has to before something can answer it.
The goal at every layer is the same. Avoid repeating work that has already been done, and avoid moving bytes that have already been moved. The closer the answer lives to the user, the faster it arrives.
The browser cache
The first and cheapest cache is the one on your own device. After the browser downloads an image, a stylesheet, or a script, it can keep a copy and reuse it on the next page instead of downloading it again. The server controls this behavior with response headers, and the main one is Cache-Control.[1]
Cache-Control: public, max-age=31536000, immutable
That header tells the browser it may keep the file for a year and that the file will never change, so it should not even bother checking. This looks reckless until you see the trick that makes it safe. Files that can be cached forever are given names that include a hash of their contents, something like app.9f2a1c.js. When the file changes, its name changes, so the browser treats it as a brand new file and fetches it. The old cached copy is simply never requested again. This is why build tools rename your assets on every deploy.
HTML is usually handled differently, with a short or zero cache lifetime, because the HTML is what points at those hashed filenames. You want the document itself to be fresh so it can reference the latest assets.
Validation: checking without re-downloading
Not everything can be cached for a year, and for those cases there is a middle ground. Instead of blindly reusing a copy or blindly re-downloading, the browser can ask the server a cheap question: "has this changed since I last saw it?"
The server supports this with a validator, usually an ETag, which is a short fingerprint of the content.[2]
ETag: "a1b2c3"
On the next request, the browser sends that fingerprint back with If-None-Match. If the content has not changed, the server replies with a 304 Not Modified and no body at all[3]. The browser reuses its stored copy, and only a few bytes crossed the network instead of the whole file. It is a small round trip that avoids a large download.
CDNs and the edge
Your origin server lives in one place. Your users do not. Physics sets a floor on how fast data can travel, and a request crossing an ocean pays for that distance on every trip. A content delivery network solves this by keeping copies of your content on servers in many locations around the world[4], often called edge nodes.
When a user requests a file, they are routed to the nearest edge node. If that node already has a copy, it serves it immediately, and the request never reaches your origin. If it does not, the edge node fetches it from the origin once, stores it, and serves everyone else in that region from the local copy. A CDN turns one slow trip to a faraway server into many fast trips to a nearby one, and it absorbs huge amounts of traffic that would otherwise hammer your origin.
Modern CDNs cache more than static files. They can cache full HTML pages and even run small amounts of code at the edge, which blurs the line between a cache and an application server.
Cache invalidation, the hard part
There is an old joke in software that there are only two hard problems, and cache invalidation is one of them. The joke survives because it is true. Storing a copy is easy. Knowing when that copy has gone stale, and getting rid of it at the right moment, is not.
The danger runs in both directions. Hold a cached copy too long and users see outdated prices, old articles, or a logged-out version of a page they are logged into. Clear caches too aggressively and you lose the entire benefit, sending every request back to the origin. A bug in invalidation logic can serve one user's private data to another, which is how caching mistakes turn into security incidents.
The strategies are a spectrum. Content-hashed filenames sidestep the problem entirely for assets, since a changed file is a new URL. For content that changes on its own schedule, you set a time-to-live and accept that the cache may be a little behind for that window. When you need something gone right now, you issue an explicit purge to the CDN, telling it to drop a specific URL or a tagged group of URLs.
Stale-while-revalidate
One pattern deserves a mention because it resolves a common tension nicely. stale-while-revalidate lets a cache serve a slightly out-of-date copy immediately[5] while it fetches a fresh one in the background.
Cache-Control: max-age=60, stale-while-revalidate=600
For ten minutes after the first minute of freshness, the user gets the cached version instantly, and the cache quietly updates itself for the next visitor. Nobody waits on a slow fetch, and the content is never badly out of date. It trades a small, bounded amount of staleness for a large gain in perceived speed, which is often exactly the right deal.
The mindset
Good caching is less about any single header and more about a habit of thinking. For every piece of content, ask how often it really changes and who is allowed to see it, then cache as aggressively as those two answers allow. Assets that never change can live forever. Personalized or sensitive responses need care. Most things sit in between and do well with a short lifetime and background revalidation. Get that reasoning right and the layers do the rest, quietly making the web faster by doing less.