Serve from near the user. The fastest request never reaches you.
The clean case is content-addressed. When a filename encodes a hash of its own content, a script bundle, a font file, an image export, the edge can hold it forever, because the one property that would ever force a change, the content itself, is baked into the address it is stored under. There is no invalidation decision to make and no window where the edge is serving something the origin has since changed, because a new version lives at a new address rather than overwriting the old one. I treat this shape, not the file's role in the application, as the actual test for whether an asset belongs at the edge with an unlimited lifetime.
The next case is public and uniform: a marketing page, a published article, a product listing that renders the same markup for anyone who requests it. Nothing about the response depends on who is asking, so a single cached copy genuinely answers every request the same way, and the only decision left is how stale the edge is allowed to let that copy get before checking the origin again. That decision belongs with the content, not with the CDN's default: a page that changes rarely can sit at the edge for a long stretch, while one that changes constantly needs a much shorter leash even though nothing else about it looks different from the outside.
The honest limit starts the moment a response depends on the person requesting it rather than on the URL alone, and a personalised page still has something to offer the edge even though the whole response cannot be cached as one piece. A page can be built so the parts that are genuinely the same for everyone, the layout, the navigation, the bulk of the markup, are served from the edge as a shell, while the parts that depend on identity are fetched separately once the shell has already loaded. Composing the page this way keeps the expensive, identical work at the edge and pushes only the small personalised fragment back to the origin, rather than pretending the whole response can be cached and then discovering, later, exactly whose data ended up in whose cache entry.
The Vary header is the mechanism meant to stop one visitor's response from being handed to another: it tells every cache between the origin and the browser that the response differs by the value of a named header, so each distinct value gets its own stored copy instead of one shared across all of them. The mechanism only works when every layer in the path honours it consistently. A response that carries Vary on a header that changes what is rendered in one code path, and stops carrying it once a later change forgets to set it, produces a cache that is fragmented correctly for some requests and not at all for others, and that gap behaves like the identity folded into the key that this lab's caching article covers, just discovered at the header level instead of the design level.
Cache-Control's public and private directives are the explicit instruction the origin gives about who may hold a copy of a response at all. Public tells any shared cache along the way, the CDN, a corporate proxy, an internet provider's own cache, that it may store the response and serve it to whoever asks next. Private restricts storage to the requesting browser alone, the only cache that already belongs to the one person the response was meant for. I read public on a response as a claim that anyone requesting that exact URL is entitled to see exactly what is stored under it, and that claim needs to still be true, not inherited from a framework default set before the route ever returned anything personalised.
Most CDNs decide by default which parts of a request feed the cache key, typically the path and perhaps the query string, while cookies are excluded unless a rule says otherwise. That default is a reasonable starting point and a real hazard once a route starts reading a cookie to decide what to render. Including every cookie in the key to be safe multiplies the key space to roughly one entry per visitor, since a session cookie or a tracking identifier differs for nearly everyone regardless of whether it changes the response, and a cache with that many single-use entries spends its capacity storing almost nothing anyone else will ever ask for again. Including none of them is how a route that has quietly become identity-dependent keeps answering from a copy meant for someone else. The correct key includes only the values the response depends on, named explicitly in the CDN's configuration, rather than left to whichever default the platform shipped with.
This is what being deliberate about the cache key, the instruction the concept's own caution leaves open instead of spelling it out, means in practice: naming the Vary headers that matter, setting private wherever a shared cache should never hold a copy, and including in the key only what the response provably depends on. None of that is configured once and forgotten. A route that starts reading a new cookie or header a year after launch inherits none of these decisions automatically, and the same review has to run again against whatever the route now depends on.
The cheapest purge strategy is the one that needs no purge at all. When a route's output is versioned the same way a static asset is, a new build identifier or a new content hash folded into the path, a deploy that changes the output produces a new address rather than overwriting the old one. The previous copy stays valid, simply unreferenced, and ages out of the edge on its own schedule without anyone having to remove it. This is the starting point worth reaching for wherever a route's shape allows it, since a purge that is never issued cannot also be issued late or issued against the wrong pattern.
Content that must keep a stable address, an article whose URL is meant to survive years of links pointing at it, a product page indexed by search engines, cannot use this trick, because changing the address on every edit defeats the reason the URL stays stable in the first place. What it needs instead is a purge triggered by the edit itself: the cache entry is tagged, at the moment it is stored, with an identifier tied to the record that produced it, so editing that record purges precisely the entries it produced rather than guessing at which cached paths might be affected. I tag by the record that produced the entry, since a page is often assembled from several records at once, and each one needs to be able to trigger a purge of every page it appears on, not only the page whose own URL matches it most obviously.
The purge a deploy pipeline remembers is usually the one attached to whatever changed directly: the route whose template was edited gets invalidated, and the pipeline calls that finished. What it forgets is everything that depends on the changed piece without being the piece itself. A shared header, a footer, a pricing table rendered into dozens of otherwise unrelated pages, changes once, in one file, and every page that composed it needs the same purge the direct route received, or visitors land on a mix of the new shared fragment and whatever the rest of that page's cached copy still holds from before the deploy. Purge coverage has to follow the dependency graph a change touches, not the list of files a commit happened to modify, because those two lists agree only when nothing shared ever changes, which is precisely the case a cache is least likely to need help with.
Time to first byte measures how long a browser waits for the response to begin arriving, and it is the number most often quoted as proof that a CDN made a page fast. For a genuinely dynamic response, one that must reach the origin, run the application and query the database on every single request, the edge changes very little about where that time goes. A closer network hop can shorten the connection setup, the DNS lookup, the handshake, and those savings are real. None of them touch the time the origin itself spends computing the answer, which is usually most of the number anyone is looking at.
Edge compute narrows that gap for some workloads by letting a small amount of logic run at the same location as the cache, close enough to the user that a request never has to reach the origin at all for simple decisions: a redirect, a header rewrite, an authentication check against a signed token. What it does not remove is the constraint that made the origin necessary in the first place for anything heavier. A primary database usually lives in one region, and code running at many edge locations still has to cross the same distance to reach it that the origin always did, so moving logic to the edge only pays off for the slice of work that never needed the database at all. I check what a request still has to reach back to a single region for before deciding that moving its logic to the edge will make it faster, since a function that runs at the edge and then waits on a distant database has added a hop rather than removed one.
The pages where this matters most are the ones that look handled because they carry a sensible caching policy, when what they deliver is headers and hope: a policy a fully personalised response can never make use of, sitting in front of an origin that still has to do the same work it always did. A dashboard rendered fresh for every authenticated visitor is not made faster by anything sitting at the edge beyond the transport savings already described, and treating the presence of a CDN as evidence that a page is fast is how a slow origin query stays uninvestigated behind infrastructure that was never positioned to fix it.
Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.
Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.
Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.