Hiring BarSupport

Design a CDN & Edge Caching — Staff-Level Case Study

Case study69 min read7 diagrams

Technologies referenced in this case study: API Gateways & Edge Proxies · Redis

Related case studies: Distributed Caching · Blob Storage · Load Balancer · API Gateway · Rate Limiting · Caching Fundamentals · Scaling Reads · Large Blobs

How to Use This Case Study#

Organized for interview use first, reference second.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 2, 5
Targeted Study1–2 hrsExecutive Summary → Walkthrough → §3 Fault Lines → §4 Failure Modes → Deep Dives 2 and 4
Deep Dive3+ hrsEverything, including §11 Principal Lens and appendices
What is a CDN & Edge Caching? — Why interviewers pick this topic

A content delivery network is a globally distributed fleet of caching proxies (points of presence, PoPs) placed close to users. Requests are routed to a nearby PoP (via anycast or DNS); the PoP serves from cache on a hit, or fetches from the origin (often through an intermediate shield tier) on a miss. Beyond caching, the edge terminates TLS close to users, reuses warm connections to origin, absorbs DDoS traffic, and increasingly runs code.

Before vs After — the product launch:

Without a well-designed edge:
t=0:      Launch email to 20M users; product page and images served from us-east origin
t=+30s:   Traffic 40x normal; users in Sydney see 1.8s TTFB (4 RTTs at ~220ms each)
t=+1min:  Origin egress saturates at 20 Gbps; image requests queue
t=+2min:  Origin app servers at 100% CPU rendering the same page 50,000 times per second
t=+5min:  Errors spike; launch trends on social media for the wrong reason

With a well-designed edge:
t=0:      Same email; page cached at edge with max-age=60, stale-while-revalidate=300
t=+30s:   97% of requests served from 300+ PoPs; Sydney TTFB ~40ms
t=+30s:   Shield tier collapses misses: origin sees ~1 request per object per shield per minute
t=+1min:  Origin at 15% CPU; egress ~0.5 Gbps
t=+10min: Price change published → surrogate-key purge → global in < 1s
Result:   A launch, not an incident.

Why interviewers reach for this question: Every candidate knows "put a CDN in front of it." Few can reason about what's cacheable, what the cache key is, how fast a change must propagate, what happens when the CDN itself fails, or how a caching misconfiguration can leak one user's private page to another. It tests caching judgment at global scale, security, cost, and vendor strategy — all in one prompt.

Mechanics Refresher: Edge Caching Building Blocks
MechanismHow It WorksProsCons
TTL expiry (Cache-Control: max-age, s-maxage)Object served until age exceeds TTL, then revalidated or refetchedSimple, predictableFreshness bounded by TTL; long TTL = stale, short TTL = low hit ratio
Versioned URLs (app.3f9a2c.js)Content hash in the URL; new content = new URL; TTL = 1 year, immutablePerfect freshness and hit ratioOnly for assets referenced by something else (HTML, manifest)
Purge / invalidationExplicit removal by URL, prefix, or tagFreshness on demandPropagation time varies (sub-second to minutes); purge storms hit origin
Surrogate keys / cache tagsTag objects (product:123); purge by tagPrecise invalidation of every page that shows product 123Tagging discipline across teams
stale-while-revalidate / stale-if-error (RFC 5861)Serve stale while fetching fresh in background; serve stale if origin errorsHides origin latency and outagesUsers see stale content for a bounded window
Request collapsingConcurrent misses for the same key wait for one origin fetchPrevents thundering herd on popular missesHead-of-line waiting on slow origin
Tiered cache / origin shieldEdge misses go to a regional shield before originOrigin sees 1 request per shield, not per PoPExtra hop on miss (~10–50ms)
Anycast / GeoDNS routingUsers routed to a nearby PoP by BGP or DNSLow latency, DDoS absorptionAnycast shifts with BGP; DNS needs resolver-location accuracy

For most production systems: Versioned immutable URLs for static assets, short TTLs + stale-while-revalidate + surrogate-key purge for semi-dynamic HTML/API responses, a shield tier in front of origin, request collapsing, and explicit Cache-Control: private, no-store on anything personalized.


Executive Summary

If you only read one section, read this.

What This Interview Actually Tests#

A CDN is not a "put CloudFront in front of it" question. Everyone knows edges cache things.

This is a freshness, correctness and blast-radius question that tests:

  • Whether you decide what is cacheable and for how long based on who is harmed by staleness
  • Whether you design the cache key — the single most consequential and most error-prone decision in edge caching
  • Whether you protect the origin from misses, purges, and cold caches — not just from users
  • Whether you treat the CDN as a third-party dependency that can fail, leak, and lock you in

The key insight: The hit ratio is a business metric and the cache key is a security boundary. A 95% → 99% hit ratio cuts origin load 5×; a cache key missing one header can serve a logged-in user's account page to strangers. Staff engineers optimize the first without ever risking the second.

The L5 vs L6 vs L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Put a CDN in front of static assets"Classifies content by freshness need and personalization: immutable, semi-dynamic, personalized, never-cacheAsks what the edge should own for the whole company: caching, TLS, WAF, bot defense, compute — and what it must never own
FreshnessPicks TTLs per file typeVersioned URLs for assets; short TTL + SWR + tag purge for pages; freshness SLO agreed with productMakes purge latency and freshness contracts part of the platform SLA across all teams
Cache keyURLNormalized URL + explicit allowlist of headers/cookies/query params; everything else strippedOwns a cache-key policy standard; key changes reviewed like security changes
Origin protection"The CDN reduces origin load"Shield tier, request collapsing, purge rate limits, stale-if-error; origin sized for cold-cache scenariosPrices origin capacity for worst-case cache loss vs CDN spend; decides on multi-CDN
Failure"CDNs are highly available"CDN outage plan: DNS failover to a second CDN or direct-to-origin with capacity mathTreats the CDN as a correlated dependency for the entire internet presence; vendor risk and exit strategy
OwnershipFrontend team configures the CDNPlatform owns edge config as code; product teams own cache headers for their responsesEdge platform team, cost allocation per team by egress, contract negotiation
Why "cache key" separates levels

L5: The URL is the cache key; the CDN handles it. Reasonable, and dangerous in both directions. Include too little (ignore Accept-Language when the page varies by language) and users see the wrong content — or someone else's content. Include too much (every query parameter, including utm_source and fbclid) and the hit ratio collapses because each marketing link is a separate object.

L6: "The cache key is an allowlist, not a denylist. Path, a normalized set of query params the page actually uses, and at most one or two headers, like device class or language — normalized into a small number of values. Cookies are stripped for cacheable routes. Any response that depends on the user is private, no-store and never enters the shared cache."

L7: Recognizes that cache-key mistakes are among the most damaging incidents an edge can cause — data leakage — and makes key changes subject to security review and automated tests that request the same URL as two different users.

Why "origin protection" separates levels

L5: Assumes the CDN always shields the origin. It does — until a global purge, a new deploy with new URLs, a cache-key change, or a CDN node failure causes a wave of misses that all arrive at once.

L6: Sizes the origin for the miss scenarios: "With 300 PoPs and no shield, a purge of our homepage sends up to 300 simultaneous fetches; with a 6-region shield and collapsing, it sends 6. I'll also serve stale-if-error for 24 hours so an origin outage degrades freshness, not availability."

L7: Prices it: origin capacity to survive a full cold cache vs the CDN's committed spend, and decides which is cheaper insurance.

Why "failure" separates levels

L5: "The CDN is globally distributed, so it's highly available."

L6: "CDNs fail globally — a bad config push or software bug at the provider can take out most PoPs at once, as the Fastly (June 2021) and Cloudflare (July 2019) incidents showed. My plan: DNS-level failover to a second CDN or to origin, with low DNS TTLs on the edge hostname, and origin capacity to take some fraction of traffic directly."

L7: Decides whether multi-CDN is worth it: roughly doubles integration and config effort, dilutes volume discounts, but converts a total outage into a partial one. That's a business decision priced in revenue per minute.

The Staff Positions#

PositionRationale
Version your static assets; never purge themContent-hashed URLs with 1-year immutable TTL give ~100% hit ratio and perfect freshness
Cache key is an allowlistEvery extra dimension fragments the cache; every missing one risks wrong or leaked content
Personalized = private, no-store by defaultOpt in to shared caching, never opt out; leakage is worse than a miss
Short TTL + SWR beats long TTL + purge for pagesBounded staleness without depending on purge reliability
Always shield the originTiered cache + collapsing turns 300 PoP misses into a handful of origin fetches
stale-if-error on everything cacheableOrigin outages degrade freshness, not availability
Plan for the CDN failingIt's a single global dependency; DNS failover and origin capacity are the plan

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Static asset delivery (JS, CSS, images, fonts)Latency and hit ratio; huge fan-outVersioned immutable URLs, long TTL, image optimization at edgeDeploy references unpublished asset (404 cached)~99%+ hit ratio; zero staleness by construction
Semi-dynamic content & API acceleration (catalog pages, public APIs, feeds)Freshness vs origin loadShort TTL, SWR, surrogate-key purge, micro-caching (1–10s), shield, collapsingStale price/stock; purge storms; cache-key leakageFreshness SLO (e.g., ≤ 60s) signed off by product; zero cross-user leakage
Large media & video (VOD segments, downloads)Bandwidth cost, long tail, byte rangesSegment caching, tiered cache, pre-positioning, ISP-embedded cachesLong-tail misses flood origin; cache pollution by one-hit wondersRebuffer ratio and startup time targets; origin egress budget

🎯 Staff Move: "I'll design for semi-dynamic content — catalog pages and public API responses — because that's where freshness, cache keys and origin protection collide. Static assets are nearly solved by versioned URLs, and video is a bandwidth-economics problem I'd treat separately."

The Five Fault Lines#

#Fault LineThe Tension
1Freshness vs Hit RatioShort TTLs are fresh and miss often; long TTLs hit and go stale. Purge closes the gap — if it works
2Cache Key Precision vs FragmentationInclude more in the key (correct, fragmented) or less (efficient, risk of wrong/leaked content)?
3Flat vs Tiered CachingEvery PoP goes to origin (lower miss latency) or through shields (origin protection, extra hop)?
4Single CDN vs Multi-CDNOne vendor (simple, cheaper, correlated failure) or several (resilient, complex, config drift)?
5Logic at the Edge vs at the OriginRun personalization, auth and routing at the edge (fast) or keep it central (simple, consistent, debuggable)?

In the Wild: Real Production Systems#

Netflix — Open Connect#

Netflix built its own CDN, Open Connect, and gives caching appliances (OCAs) to ISPs to install inside their networks at no charge, plus deploying them at internet exchange points. Content is pre-positioned during off-peak hours based on predicted popularity rather than pulled on demand, so the overwhelming majority of Netflix video traffic is served from an appliance close to — often inside — the viewer's ISP. The control plane (steering which OCA serves a client) runs in AWS; the bytes never touch it.

Staff insight: For large media with predictable demand, push beats pull. And it's a build-vs-buy decision driven by scale: at a double-digit share of peak internet traffic, owning the edge is cheaper and better than renting it. Almost nobody else is at that scale.

Fastly — Instant Purge and Surrogate Keys#

Fastly built its platform around fast, tag-based invalidation: responses carry a Surrogate-Key header with tags (e.g., product-123 category-9), and a purge by tag propagates globally in roughly 150ms by their public description. That makes it practical to cache semi-dynamic content — news articles, product pages, API responses — with long TTLs and purge on write, rather than short TTLs. On June 8, 2021, a customer's valid configuration change triggered a latent software bug and caused most of Fastly's network to return errors for under an hour, taking many major sites down simultaneously.

Staff insight: Instant purge changes the freshness tradeoff — "cache long, purge on change" becomes viable. The 2021 outage is the canonical example that a CDN is a correlated dependency across a large share of the web.

Cloudflare — Anycast, Tiered Cache, and a Global Regex#

Cloudflare routes users via anycast to 300+ cities, offers Tiered Cache (lower-tier PoPs fetch from upper-tier PoPs before origin), and runs WAF and code (Workers, on V8 isolates with very low startup cost) at every PoP. On July 2, 2019, a new WAF managed rule containing a regular expression with catastrophic backtracking was deployed globally at once and drove CPU to 100% across the network for about 27 minutes; the public post-mortem led to staged rollouts for WAF rules.

Staff insight: Every piece of logic you run at the edge runs everywhere at once. Edge config and code need staged, canaried rollouts even more than origin code does, because the blast radius is your entire internet presence.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Cache it at the CDN""What's the cache key? What if the page varies by language or login state?"Cache-key judgment and leakage awareness
"TTL of 5 minutes""Price changed. The business wants it visible in 10 seconds."Freshness mechanisms: purge, tags, SWR
"Purge on change""A bulk import changes 2M products. What happens to origin?"Purge storms and origin protection
"The CDN absorbs load""The cache is empty after a config change. How much traffic hits origin?"Cold-cache sizing, shield, collapsing
"CDNs are highly available""Your CDN is returning 503s globally. What now?"Vendor failure plan
"Run it at the edge""How do you debug and roll back logic running in 300 cities?"Edge compute operational maturity

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Requests reach the nearest PoP via anycast/DNS. The PoP terminates TLS (saving ~2 round trips to a distant origin), normalizes the cache key, and serves hits. Misses go to a regional shield, which collapses concurrent misses so origin sees roughly one fetch per object per shield. The origin tells the edge what to do through headers — Cache-Control for freshness and privacy, Surrogate-Key for purge targeting — so cache policy lives with the team that owns the content. A publish event triggers a tag purge. The origin only accepts traffic from the CDN, so attackers can't bypass the edge. The DNS layer is the escape hatch if the CDN itself fails.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Static assets"Cache with a long TTL, purge on deploy""Content-hashed URLs, 1-year immutable. Never purge; new content is a new URL."
Pages / APIs"5-minute TTL""Short TTL (30–60s) + stale-while-revalidate + tag purge on write; freshness SLO agreed with product."
Cache key"The URL""Allowlist: path, normalized params, one or two normalized headers. Cookies stripped. Personalized = private, no-store."
Origin protection"CDN reduces load""Shield tier + request collapsing + purge rate limits + stale-if-error. Origin sized for the cold-cache case."
CDN outage"CDNs don't go down""They do, globally. DNS failover to a second CDN or origin; low TTL on the edge hostname."
Edge logic"Move it to the edge for speed""Only stateless, latency-critical logic; staged rollout by PoP group; origin remains the source of truth."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Hit ratio 95% vs 99%Origin sees 5% vs 1% — a 5× differenceLast few points of hit ratio are worth the most
Speed of light in fiber~200 km/ms → ~1ms RTT per 100 kmPhysics sets the floor; only proximity beats it
Typical RTTsIntra-region 1–5ms; US coast-to-coast ~60–70ms; NY–London ~70–80ms; US–Australia ~150–200msNew-connection cost = RTT × handshakes
New HTTPS connectionTCP 1 RTT + TLS 1.3 1 RTT (TLS 1.2: 2 RTT) before the first requestEdge TLS termination saves 2–3 long RTTs per new connection
Fastly tag purge (public figure)~150ms globalMakes "cache long, purge on change" viable
CloudFront invalidation pricingFirst 1,000 paths/month free, then ~$0.005/pathPurge-heavy designs cost money on some CDNs
Versioned asset TTLmax-age=31536000, immutable (1 year)Zero revalidation for static assets
Micro-cache TTL1–10sAt 10K RPS on one URL, a 1s TTL cuts origin to ~1 RPS per PoP/shield
Shield tiers3–6 regions typical300 PoPs → a handful of origin fetches per object
One-hit wondersA large fraction of objects requested only once (Akamai researchers reported roughly three-quarters in one study)Admission filters protect cache space for popular content
Video segment length2–6s; HD ~5 MbpsSegments are the cache unit for streaming
CDN egress list price (order of magnitude)~$0.02–0.09/GB, much lower at committed PB volumesEgress is the dominant CDN cost line

Interview Walkthrough

Common prompts: "Design a CDN", "Design the edge layer for a global e-commerce site", "Our origin can't handle launch traffic — design caching." If the prompt is literally "design a CDN" (i.e., you are the CDN provider), spend proportionally more time on PoP architecture, routing, and purge propagation; if it's "use a CDN for our product", spend it on cacheability, keys, freshness and failure. Ask which.

Phase 1: Requirements & Framing (2–3 min)#

"Two framings are possible — building a CDN as a provider, or designing our product's edge on top of one. I'll assume the second unless you'd prefer the first. Then: what content are we serving — static assets, HTML pages, API responses, video? How fresh must each be — can a price be 60 seconds stale? Is any of it personalized? And where are users — global or one region? I'll assume a global e-commerce site: 50M daily users, 70% outside the origin's region, static assets plus catalog pages plus a public product API, with prices that must update within 60 seconds and nothing personalized in the shared cache."

Established:

  • Content classes drive strategy
  • Freshness SLO is a product decision you'll get signed off
  • Personalization boundary is explicit
  • Geography justifies the edge

Phase 2: Core Entities & API (1–2 min)#

EntityFields That MatterOwner
Cached objectcache key, body, headers, stored_at, TTL, SWR/SIE windows, surrogate keys, sizeEdge
Cache policy (per route)cacheable?, key allowlist (params, headers), TTL, SWR, SIE, private?Product team (via response headers) + platform (defaults, guards)
Purge requesttags or URLs, soft/hard, requester, reasonPublishing systems; rate-limited by platform
Edge configroutes, origins, shield mapping, WAF rules, edge functions, versionPlatform (config as code)

The contract between origin and edge is HTTP headers:

Static asset:   Cache-Control: public, max-age=31536000, immutable
Catalog page:   Cache-Control: public, s-maxage=60, stale-while-revalidate=300, stale-if-error=86400
                Surrogate-Key: product-123 category-9 page-template-v4
                Vary: Accept-Encoding
Account page:   Cache-Control: private, no-store
Public API:     Cache-Control: public, s-maxage=10, stale-while-revalidate=30
Purge API:      POST /purge  { "tags": ["product-123"], "soft": true }

🎯 Staff Move: "The origin owns cache policy through headers, so the team that knows the content decides its freshness. The platform owns guardrails — like refusing to cache any response that sets a cookie, regardless of what the headers say."

Phase 3: High-Level Architecture (≤5 min)#

Draw the overview: users → DNS/anycast → PoPs → shields → origin, with a purge path and a config path. Name owners:

  1. Routing (DNS/anycast, failover) — platform/networking
  2. PoPs (TLS, WAF, cache, collapsing) — CDN vendor + platform config
  3. Shield tier — CDN feature, platform-configured
  4. Origin (emits headers, accepts only CDN traffic) — product teams + platform
  5. Purge pipeline (publish events → tag purge) — platform, fed by product systems

Phase 4: Transition to Depth#

"The architecture is standard. Where designs actually fail is in three places: the cache key — which determines both hit ratio and whether we ever leak content — freshness and purge at scale, and what happens when the cache is cold or the CDN itself fails. I'll go in that order."

Phase 5: Deep Dives (25–30 min)#

Deep dive 1 — Cache key (7 min).

Incoming:  GET /p/red-shoes?utm_source=email&color=red&size=9&fbclid=abc
           Cookie: session=...; ab_test=B     Accept-Language: de-DE,de;q=0.9
Normalize: path lowercase, strip trailing slash
           query: keep allowlist {color, size}, sorted → color=red&size=9
           headers: lang = first supported of Accept-Language → 'de'
           cookies: none in key; stripped from origin request for cacheable routes
Key:       /p/red-shoes|color=red&size=9|lang=de|enc=br
Variants:  ~8 languages × 2 encodings = 16 per URL (not thousands)

Deep dive 2 — Freshness (7 min). Prices must be visible within 60s. Options: TTL 60s (simple, origin load = 1 fetch per object per shield per minute), or TTL 1 day + tag purge on change (better hit ratio, depends on purge reliability). Staff default: TTL 60s + SWR 300s as the floor guarantee, plus tag purge for faster updates — purge makes it faster, TTL makes it bounded even if purge fails.

Deep dive 3 — Origin protection (6 min).

Worst case: global purge of homepage at 50K RPS
  No shield, no collapsing:  ~300 PoPs × concurrent misses → thousands of origin requests in 1s
  Shield (6) + collapsing:    ~6 origin requests
  Soft purge (mark stale):    0 blocking origin requests — served stale while 6 revalidate
Bulk purge (2M products):     rate-limit purges to ~5K tags/s; prefer soft purge; stagger

Deep dive 4 — Failure (5 min). Origin down → stale-if-error for 24h keeps catalog up. CDN down → DNS failover (edge hostname TTL 60s) to CDN B or to origin direct with origin capacity for ~20–30% of traffic plus aggressive shedding.

Deep dive 5 — Security (3 min). Never cache Set-Cookie responses; private, no-store on personalized routes; origin accepts only CDN traffic (IP allowlist + secret header or mTLS); unkeyed-header poisoning defenses.

Phase 6: Wrap-Up (2–3 min)#

"To summarize: versioned immutable assets; semi-dynamic pages with short TTL, stale-while-revalidate and tag purge; an allowlist cache key with personalization excluded by default; a shield tier and collapsing so the origin sees a handful of requests per object; stale-if-error for origin outages; and DNS failover for CDN outages. Next I'd add hit-ratio and freshness SLOs per route, an automated leakage test in CI, and evaluate multi-CDN once revenue per minute justifies it. Biggest risk: a cache-key or header mistake that caches private content — that's where I'd put the automated guardrails first."

Common Timing Mistakes#

MistakeTime LostWhat to Do Instead
Explaining anycast/BGP in depth8 minOne sentence unless you're building the CDN
Listing CDN vendors and features5 minVendor comparison belongs in the technology guide; say what you need
Only discussing static assets—They're solved by versioning; spend time on semi-dynamic
Never mentioning the cache key—It's the most important design decision
Ignoring CDN failure—Interviewers will ask; have the DNS failover answer ready

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Edge caching is deceptively cross-cutting. The CDN configuration is owned by one team; the headers that control it are emitted by dozens; the content's freshness requirements come from product; the costs land in finance; and the failures — stale prices, leaked pages, global outages — land on everyone. A Senior engineer can configure a CDN correctly for one application. A Staff engineer designs the contract that lets fifty teams cache safely without each one learning the hard way.

1.2 The L5 vs L6 vs L7 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 vs L7 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"If this response were served to a different user, or two minutes late, who would be harmed — and how badly?"

Ask it per route. The answers sort every response into a strategy:

  • Nobody, ever → immutable, cache forever with versioned URLs
  • Nobody, if within N seconds → shared cache with TTL N, SWR, purge
  • The user, if another user sees it → private, no-store; never in the shared cache
  • The business, if late (prices, inventory, legal notices) → short TTL + purge + product-owned freshness SLO

2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1: Static asset delivery. Solved problem if done right: content-hashed filenames, 1-year TTL, immutable, served from edge with Brotli/gzip and image format negotiation (AVIF/WebP). The remaining risks are deploy ordering (HTML references a new asset before it's uploaded, and the 404 gets cached) and cache pollution. Hit ratio should be 98–99%+.

Intent 2: Semi-dynamic content and API acceleration. The hard case. Catalog pages, search result pages for anonymous users, public API responses, article pages. They change — sometimes often — and the business cares how fast changes appear. The toolkit: short TTLs, SWR, micro-caching for extremely hot URLs (1–10s TTLs flatten spikes), tag purge on change, request collapsing, and a strict personalization boundary. Even uncacheable dynamic requests benefit from the edge: TLS termination near the user and warm, reused connections from edge to origin save 100–300ms for distant users.

Intent 3: Large media and video. Different economics. Objects are large (MB–GB), the catalog has a long tail, and egress cost dominates. Segment-level caching (2–6s chunks), tiered caches with large disk tiers, byte-range support, admission policies so one-time views don't evict popular content, and — at very large scale — pre-positioning and ISP-embedded caches (Netflix Open Connect). Metrics are rebuffer ratio, startup time and origin egress.

🎯 Staff Move: "These three share a CDN but not a strategy. Static assets want versioning, semi-dynamic wants purge and freshness contracts, video wants bandwidth economics. I'll focus on semi-dynamic, and I'll give static a one-line answer so we don't spend time on a solved problem."

2.2 When NOT to Use (Shared) Edge Caching#

SituationWhy It's WrongWhat to Do Instead
Personalized responses (cart, account, feed)Shared cache + personalization = leakage risk; hit ratio near zero anywayprivate, no-store; still use edge for TLS + connection reuse; cache fragments via edge-side includes only with care
Strong consistency required (balances, inventory decrements, auth decisions)Any TTL is a correctness bugServe from origin; cache only non-authoritative displays
Write-heavy APIsNothing to cache; edge adds a hopDirect to regional API endpoints (API Gateway)
Users all in one region near originEdge adds cost with little latency benefitRegional LB + in-region cache (Distributed Caching)
Regulated data with residency requirementsEdge caches may store data in other jurisdictionsRegion-restricted edges or no caching; legal sign-off
Very low traffic long tailObjects expire before a second request; pure miss overheadServe from object storage directly or cache only at shield

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It MattersWhat to Say
Freshness requirement per content typeDrives TTL vs purge"Prices ≤ 60s, descriptions ≤ 1 hour, legal notices immediate — product signs off"
PersonalizationLeakage risk"Nothing personalized enters the shared cache"
GeographyLatency value of edge"70% of users outside origin region"
Origin capacityCold-cache survival"Origin sized for ~10% of peak without CDN"
Build vs buyProvider vs customer design"Using a commercial CDN; I'll note where building differs"
BudgetEgress dominates"Egress is the main cost; hit ratio and compression are the levers"

2.4 Precise Terminology#

TermPrecise MeaningCommon Confusion
PoPA CDN location with caching serversAssumed to be one server
OriginThe authoritative source the CDN fetches fromConfused with shield
Shield / tiered cacheIntermediate cache layer between PoPs and originThought of as a separate CDN
Hit ratio (request)% of requests served from cacheConfused with byte hit ratio
Byte hit ratio / offload% of bytes served from cacheDiffers sharply for mixed small/large objects
Cache keyThe identity of a cached objectAssumed to be the URL
VaryResponse header telling caches which request headers change the responseOverused (Vary: Cookie, Vary: User-Agent) → fragmentation
Purge (hard)Remove object; next request is a missCauses origin stampede
Soft purgeMark object stale; serve stale while revalidatingUnknown to many candidates
SWR / SIEstale-while-revalidate / stale-if-error windowsConfused with TTL
Request collapsingCoalesce concurrent misses into one origin fetchConfused with caching itself
Cache poisoningAttacker causes a harmful response to be cached under a key others requestAssumed impossible behind a CDN
Cache deceptionAttacker tricks the cache into storing a victim's private page under a cacheable-looking URLUnknown to most candidates

3. The Five Fault Lines#

3.1 Fault Line 1: Freshness vs Hit Ratio#

StrategyWhat WorksWhat BreaksWho Pays
Long TTL, no purgeHighest hit ratioStale content for up to TTLBusiness (wrong prices), users
Short TTL (30–60s)Bounded staleness; no purge dependencyLower hit ratio; origin load = objects × shields / TTLOrigin (capacity), finance (egress from origin)
Long TTL + purge on changeHigh hit ratio and fast updatesDepends on purge reliability and publishers remembering to purge; purge stormsPlatform (purge pipeline), origin (stampedes)
Short TTL + SWR + tag purge (Staff default)Bounded staleness even if purge fails; purge makes it faster; SWR hides origin latencyMore moving partsPlatform
Micro-caching (1–10s) for hot dynamic URLsFlattens spikes; origin load ≈ 1 request/s per shield per URLUp to 10s stalenessProduct signs off on seconds
Origin load for one hot page, 6 shields, TTL 60s:           6 requests/min
Same page, TTL 1s (micro-cache):                             6 requests/s
Same page uncached at 50K RPS:                               50,000 requests/s
Staleness bound with TTL 60s + SWR 300s:                    ≤ 60s fresh; up to 360s if origin slow

🎯 Staff Move: "Purge makes updates fast; TTL makes them bounded. I want both, because a purge pipeline that silently drops messages will happen, and I'd rather be 60 seconds stale than 3 days stale."

When to deviate: News and live events: purge-first with instant purge and long TTLs, because a headline update matters in seconds and traffic is enormous. Legal/compliance content: purge + verification (fetch from several PoPs after purge and assert new content).

3.2 Fault Line 2: Cache Key Precision vs Fragmentation#

StrategyWhat WorksWhat BreaksWho Pays
Full URL + all headers via VaryNever wrongHit ratio collapses (Vary: User-Agent → thousands of variants)Origin, finance
URL onlySimple; high hit ratioWrong content when response varies by language/device/cookie; possible leakageUsers; security
Allowlist key (path + selected params + normalized headers)Correct and efficientRequires per-route definition and disciplinePlatform + product teams
Edge-computed variant (device class, country → small enum)Controlled variantsLogic at edge must match origin'sPlatform

The asymmetric risk: an under-specified key serves wrong content; an over-specified key serves slow content. Wrong is worse — especially when "wrong" means another user's data.

Leakage guardrails (non-negotiable):

  • Never cache responses containing Set-Cookie
  • Default all routes to non-cacheable unless the route is explicitly cacheable
  • Strip Cookie and Authorization from origin requests on cacheable routes, so the origin cannot personalize them
  • CI test: request each cacheable route as two different users; assert identical bodies

When to deviate: A/B tests on cacheable pages — include the experiment bucket (not the user ID) in the key, with a small number of buckets.

3.3 Fault Line 3: Flat vs Tiered Caching#

StrategyWhat WorksWhat BreaksWho Pays
Flat (every PoP → origin)Lowest miss latencyOrigin sees N PoPs × misses; long-tail hit ratio poor (each PoP sees little traffic per object)Origin
Single shieldOrigin sees 1 fetch per objectShield is a hotspot and a single point of failure; distant PoPs pay extra RTTShield capacity; users far from shield
Regional shields (3–6)Balance of origin protection and miss latencyMore complex; objects stored at more layersPlatform config
Shield + origin-side cacheProtects origin compute even on shield missAnother layer to invalidatePlatform
Diagram: 3.3 Fault Line 3: Flat vs Tiered Caching

The Staff default: Regional shields with collapsing at both edge and shield. The shield also improves hit ratio for long-tail content, because it aggregates demand from many PoPs.

When to deviate: Latency-critical dynamic API calls that aren't cacheable — bypass the shield and go edge → origin directly over warm connections.

3.4 Fault Line 4: Single CDN vs Multi-CDN#

StrategyWhat WorksWhat BreaksWho Pays
Single CDNSimplest; best volume pricing; one config dialectProvider outage = your outageBusiness, during provider incidents
Primary + cold standbyEscape hatch via DNSStandby cold cache → origin stampede on failover; config driftOrigin, platform
Active multi-CDN (traffic split by performance/cost)Resilience; performance steering; pricing leverage2× config, purge must fan out to both, lowest-common-denominator featuresPlatform (≈2× effort), finance (split volume)
Build your ownFull control; cheapest at extreme scaleEnormous capex/opex; only at Netflix-like scaleCompany

The Staff default: Single CDN with a tested DNS failover path to origin (or to a warm secondary for revenue-critical hostnames), low TTL on the edge hostname, and config expressed in a vendor-neutral format where possible. Move to active multi-CDN when revenue per minute of outage justifies roughly doubling edge operational effort.

🎯 Staff Move: "A CDN is a single global dependency, even though it's physically distributed. The question isn't whether it'll fail, it's whether we can move traffic in minutes when it does — and whether our origin survives the cold cache."

3.5 Fault Line 5: Logic at the Edge vs at the Origin#

StrategyWhat WorksWhat BreaksWho Pays
Dumb edge (cache + TLS only)Simple, debuggable, origin is the source of truthPersonalization and auth need origin round tripsDistant users (latency)
Edge config logic (redirects, header rewrites, geo routing)Low latency; no code deploysConfig sprawl; a bad rule deploys globallyPlatform
Edge compute (isolates/functions)Auth token validation, A/B assignment, personalization of cached shells at edgeGlobal blast radius; limited state; observability and debugging harder; vendor lock-inPlatform + product teams

The Staff default: Keep the edge stateless and small: token validation (not issuance), bucket assignment, redirects, header normalization, image transforms. Anything that needs strong state stays at origin. All edge logic ships through staged rollout by PoP group with automatic rollback — Cloudflare's 2019 incident is the reason.

When to deviate: Latency-critical global products (ads, bidding) where 100ms matters — push more to the edge, with the operational investment that implies.


4. Failure Modes & Operational Reality#

The lifecycle of a cached object — most failures happen at a transition:

Diagram: 4. Failure Modes & Operational Reality

4.1 Purge Stampede#

Setup: homepage + 40K category pages tagged 'layout-v7'. 300 PoPs, no shield,
       hard purge. Normal origin load 2K RPS (98% hit ratio at 100K RPS).

t=0:      Design team ships a header change; publisher hard-purges tag 'layout-v7'
t=+0.2s:  40K objects evicted in every PoP
t=+0.5s:  100K RPS of requests → ~all miss; collapsing per PoP helps, but 300 PoPs ×
          40K hot objects = up to 12M distinct origin fetches over the next minute
t=+5s:    Origin at 60K RPS (30× normal); render latency 8s; timeouts
t=+10s:   Edges receive 5xx; nothing stale to serve (hard purge) → users see errors
t=+4min:  Caches refill slowly; origin recovers

Detection: origin_rps spike correlated with purge_events_total; edge_hit_ratio collapse; origin_5xx_rate.

Mitigation: Stop further purges; enable/extend stale serving at the edge if the vendor allows; shed at origin.

Prevention: Shield tier (origin fetches ÷ ~50); soft purge (mark stale, serve stale while revalidating) as the default; purge rate limits per publisher; broad tags (layout, template) require approval; staggered purges for bulk changes.

Owner: Platform owns the purge pipeline and its limits; the publishing team owns what they purge.

Diagram: 4.1 Purge Stampede

4.2 Cache Leakage — Private Content in the Shared Cache#

t=0:      A refactor moves the account page behind a new route; the framework's
          default Cache-Control is 'public, max-age=300' for GET routes
t=+1min:  User A loads /account; response (with A's name, address, last orders) cached
t=+1–6min: Every user requesting /account in that PoP sees User A's page
t=+20min: Support tickets; security incident declared
Scale:    Every PoP caches its own first visitor's page. Hundreds of users exposed.

This is not hypothetical: on December 25, 2015, a caching configuration change at Valve caused Steam store pages containing other users' account information to be served to some users for a period, per Valve's public statement.

Detection: private_marker_in_cached_response_total (origin injects a response header like X-Private: 1 on personalized responses; the edge alarms if it ever caches one); synthetic two-user checks every minute.

Mitigation: Purge everything under the route; disable caching for the route via emergency config; assess exposure from logs.

Prevention: Default non-cacheable (opt in per route); strip cookies/authorization on cacheable routes; never cache Set-Cookie; CI two-user test; edge rule that refuses to cache responses carrying the private marker.

Owner: Security + platform (guardrails); product team (route headers). Legal/privacy for disclosure decisions.

4.3 Cache Poisoning via Unkeyed Inputs#

An attacker sends GET / with X-Forwarded-Host: evil.example. The origin uses that header to build absolute URLs for script tags. The response — now loading scripts from the attacker's domain — is cached under the key for /, because X-Forwarded-Host isn't in the key. Every visitor to that PoP gets the poisoned page until TTL.

Detection: Content integrity monitors (fetch key pages from many PoPs, diff against expected); CSP violation reports spike.

Prevention: Edge strips all non-allowlisted request headers before forwarding to origin; origin never trusts forwarding headers for URL generation; everything the origin uses to vary the response must be in the key (or stripped).

Owner: Platform (header normalization); security (testing).

4.4 CDN Provider Outage#

t=0:      Provider pushes a bad config/software update; 85% of PoPs return 503
t=+1min:  Synthetic monitors fail from all regions; revenue drops to near zero
t=+3min:  Incident declared; decision: fail over?
t=+5min:  DNS change: edge hostname → secondary CDN (or origin LB); TTL 60s
t=+6–10min: Traffic shifts as resolvers honor TTL (some take longer)
t=+10min: Secondary CDN has a cold cache → origin sees 20–40% of total traffic
t=+15min: Origin sheds low-priority routes; catalog serves; checkout prioritized

Detection: External synthetic monitoring from multiple networks (not from the CDN itself); RUM error beacons.

Mitigation: Pre-approved failover runbook with a decision threshold (e.g., > 25% global error rate for > 3 min).

Prevention: Low TTL on edge hostnames; tested failover quarterly; origin capacity and shedding plan for cold-cache load; secondary CDN warmed with top objects if active multi-CDN.

Owner: Platform/SRE; incident commander makes the call per runbook.

4.5 Cached Errors and Deploy Ordering#

A deploy publishes HTML referencing app.9f3e.js before the asset is uploaded. The first requests get 404, which the CDN caches for its default negative TTL (sometimes minutes). The site is broken for everyone at those PoPs even after the asset appears.

Prevention: Upload assets before publishing HTML (and keep old assets for ≥ 1 week for clients with old HTML); negative-cache TTL ≤ 10s for asset paths; never cache 5xx beyond a few seconds.

Owner: Frontend platform (deploy ordering); edge platform (negative-caching defaults).

4.6 Hit-Ratio Erosion — The Slow Failure#

Over months, hit ratio drifts from 95% to 78%. Causes: new marketing query parameters, a Vary: User-Agent added by a framework upgrade, cookies from a new analytics vendor forwarded to origin and used in keys, an A/B test framework keyed on user ID. Origin cost and latency creep up; nobody notices until capacity planning.

Detection: hit_ratio per route with weekly trend alerts; top-N cache-key cardinality by route.

Prevention: Key allowlists; periodic cardinality reports; hit ratio as an SLO per route owner.

Owner: Route owners (their headers); platform (reporting).

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Purge stampedeorigin_rps spike after purge_eventsOrigin, then all usersSoft purge, shield, purge rate limitPlatform + publisher
Private content cachedprivate_marker_in_cached_response_total > 0; two-user syntheticEvery user of affected route/PoPs; privacy incidentPurge route; disable cachingSecurity + platform + route owner
Cache poisoningIntegrity monitor diffs; CSP reportsAll users at poisoned PoPsPurge; strip headerPlatform + security
CDN provider outageExternal synthetics, RUM errorsEntire web presenceDNS failover; origin sheddingSRE / platform
Cached 404/5xxedge_status_404_rate for new asset pathsUsers at affected PoPsPurge path; short negative TTLFrontend platform
Hit-ratio erosionWeekly hit_ratio trend per routeOrigin cost and latencyKey allowlist fixesRoute owners
Origin outageorigin_5xx_rate; served_stale_total ↑Freshness (if SIE configured); availability otherwisestale-if-errorOrigin team
Bad edge config/code pushError rate spike correlated with config versionGlobalAutomatic rollback; staged rolloutPlatform

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Framing"Add a CDN"Classifies content by freshness and personalizationDefines the edge platform's scope for the company
Cache keyURLAllowlist + normalization + leakage guardrailsKey policy standard with security review
FreshnessTTLsTTL + SWR + tag purge; freshness SLO with productFreshness contracts and purge SLA as platform commitments
Origin protectionImplicitShield, collapsing, soft purge, SIE, cold-cache sizingPrices origin headroom vs CDN spend
Failure"CDNs are reliable"DNS failover, cold-cache plan, staged edge configVendor risk, multi-CDN economics, exit strategy
SecurityTLS at edgePoisoning, deception, private-content guardrails, origin lockdownEdge as a security control plane (WAF, bot) with its own governance
CostNot discussedEgress-driven; hit ratio and compression as leversContract strategy, cost allocation per team

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Starts with content classes"Static, semi-dynamic, private — three strategies."
Treats the key as a security boundary"Allowlist the key; personalized routes are private, no-store by default."
Bounds freshness twice"Purge makes it fast, TTL makes it bounded."
Protects origin from its own CDN"Soft purge plus shields: 6 origin fetches, not 300."
Plans for vendor failure"Edge hostname TTL 60s, tested failover, origin sized for 30% cold."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Cache everything for 1 hour"Ignores freshness requirements and personalization
"Purge on every change" with hard purge and no shieldDesigns a stampede
Vary: Cookie or Vary: User-Agent as the personalization answerFragments the cache to near-zero hit ratio, or leaks
No failure plan for the CDNTreats a global dependency as infallible
Moves business logic to the edge without rollout disciplineGlobal blast radius

5.4 Common False Positives#

  • Knowing anycast and BGP details ≠ edge design judgment. Useful if you're building a CDN; rarely the question.
  • Vendor feature fluency ≠ design. Knowing every CloudFront behavior setting doesn't answer who owns freshness.
  • High hit ratio ≠ good design. A 99% hit ratio achieved by caching personalized pages is a data breach.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minProvider vs customer; content classes; freshness; personalization
Header contract & entities3–6 minCache-Control, Surrogate-Key, purge API
Architecture6–11 minRouting, PoPs, shields, origin, purge, config
Cache key11–18 minAllowlist, normalization, leakage guardrails
Freshness & purge18–25 minTTL + SWR + tags; purge at scale
Origin protection & failure25–34 minShield, collapsing, SIE, CDN outage
Edge compute / security / cost34–42 minWhatever the interviewer pushes
Wrap-up42–45 minNext steps, biggest risk

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"You're the CDN provider — design it"Systems depthPoP architecture (L4 LB → cache servers with consistent hashing), anycast, purge fan-out, config distribution
"Personalize the homepage"Leakage risk managementCache the shell, fetch personalized fragments client-side or via edge includes with per-user private fragments
"Serve video"Media economicsSegments, tiered disk caches, admission policies, pre-positioning
"Price changes must appear in 1 second"Freshness limitsInstant tag purge + TTL backstop; or don't cache price, cache everything else
"CDN costs doubled"Cost leversHit ratio by route, compression, image formats, commit contracts, multi-CDN leverage
"Protect against DDoS"Edge as securityAnycast absorption, WAF, rate limits at edge (Rate Limiting), origin lockdown

6.3 What to Deliberately Skip#

  • BGP and anycast internals (unless building a CDN)
  • HTTP caching spec corner cases (must-revalidate vs proxy-revalidate)
  • Vendor feature lists
  • Image optimization details beyond "format negotiation and resizing at edge"

6.4 Follow-Up Questions to Expect#

  1. "How does request collapsing interact with a slow origin?" — Waiters queue behind one fetch; cap wait time and serve stale if available.
  2. "What's the difference between max-age and s-maxage?" — s-maxage applies to shared caches; lets the CDN cache longer than browsers.
  3. "How do you cache an API that requires an auth token?" — Only if the response doesn't depend on the caller; validate the token at the edge, strip it from the key and origin request.
  4. "How do you know the purge worked?" — Purge acknowledgments plus sampled fetches from multiple PoPs asserting the new version header.
  5. "How do you pick the TTL?" — From the freshness SLO product signs off on; purge as the fast path.
  6. "Anycast or DNS-based routing?" — Anycast for simplicity and DDoS absorption; DNS for finer control; many CDNs use both.
  7. "How do you prevent the origin from being hit directly?" — IP allowlist of CDN ranges plus a secret header or mTLS from edge to origin.

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design the CDN strategy for our global e-commerce site."

Staff Answer

"I'll start by classifying content, because each class gets a different strategy. Static assets — JS, CSS, images — get content-hashed URLs and a one-year immutable TTL; that's essentially solved. Catalog pages and the public product API are semi-dynamic: short TTL, stale-while-revalidate, and tag-based purge when products change, with a freshness SLO product signs off on — say 60 seconds for price. Cart, account and checkout are personalized: private, no-store, never in the shared cache, but still routed through the edge for TLS termination and warm origin connections.

Then the three things I'd go deep on: the cache key, because it determines both hit ratio and whether we can ever leak someone's data; freshness and purge at scale, because purge storms are how CDNs hurt origins; and failure — origin outages and CDN outages."

Why this is L6:

  • Classifies before configuring
  • Gets product sign-off on freshness
  • Names the cache key as a security boundary up front

What L7 adds:

  • Frames the edge as a company platform with a contract all teams use, not a per-app config
  • Mentions cost allocation and vendor strategy as later decisions
❌ Common L5 Trap

"Put CloudFront in front of everything with a 1-hour TTL and invalidate on deploy."

Why this misses: No distinction between static and dynamic, a 1-hour TTL on prices, and "everything" includes personalized pages — a leakage risk.


Drill 2: Cache Key Design#

Prompt: "Product pages vary by language and currency. Marketing adds UTM parameters to every link. How do you design the cache key?"

Staff Answer

"Allowlist. The key is the normalized path, the query parameters the page actually uses — say color and size, sorted — and two normalized dimensions: language (mapped from Accept-Language to one of our ~8 supported languages) and currency (derived from a country or explicit preference, mapped to ~10 currencies). UTM parameters, fbclid, gclid and every other unknown parameter are stripped from the key — and from the origin request, so the origin can't accidentally vary on them. Cookies are stripped on this route.

That bounds variants to about 8 × 10 = 80 per product URL, rather than one per marketing link. And I'd make the page emit a header echoing the language and currency it rendered, so a monitor can detect mismatches between key and content."

Why this is L6:

  • Allowlist over denylist
  • Normalizes headers into small enums
  • Strips unkeyed inputs from origin requests so they can't vary the response

What L7 adds:

  • Proposes a platform-owned key policy per route type, so 50 teams don't each reinvent normalization
  • Tracks key cardinality per route as a cost metric

Drill 3: Make It Concrete — Freshness#

Prompt: "A price change must be visible globally within 10 seconds. How?"

Staff Answer

"Two mechanisms. The fast path: the pricing service publishes a 'product updated' event; a purge worker issues a soft purge for tag product-123, which on a CDN with fast tag purge propagates in roughly a second or less. The backstop: TTL on the page. If I need 10 seconds as a guarantee, not just typical behavior, TTL must be ≤ 10s — micro-caching — which at 6 shields is ~0.6 origin requests per second per product page, fine for popular products.

Alternatively, split the page: cache the product page for 5 minutes but fetch the price from a tiny, uncacheable or 1-second-cached price endpoint. That puts the freshness requirement only where it matters. And checkout always re-prices from the source of truth — the displayed price is never authoritative."

Why this is L6:

  • Distinguishes typical freshness (purge) from guaranteed freshness (TTL)
  • Offers splitting the freshness-sensitive fragment out
  • Keeps the transaction authoritative

What L7 adds:

  • Makes 'freshness SLO per content type' a published table owned by product, with the platform guaranteeing purge latency as an SLA

Drill 4: Origin Is Down#

Prompt: "Our origin is down for 20 minutes. What do users see?"

Staff Answer

"For cacheable content: the last good version, because every cacheable response carries stale-if-error=86400. The catalog stays browsable; prices might be up to 20 minutes stale, which product accepted for this scenario since checkout re-prices anyway. For uncacheable content — cart, checkout — errors, unless we have a static maintenance experience at the edge for those routes.

What makes this work: SIE configured on all cacheable routes, objects not evicted too aggressively, and the edge not converting origin 5xx into cached errors. I'd verify quarterly by blocking origin from the CDN for 10 minutes in a controlled game day."

Why this is L6:

  • Uses SIE to turn an outage into staleness
  • Names what still breaks and what users see
  • Proves it with a game day

What L7 adds:

  • Designs a static 'degraded shell' for transactional routes served from edge storage, with a product-approved message and status page link

Drill 5: Purge at Scale#

Prompt: "A nightly bulk import updates prices for 2M products. How do you invalidate?"

Staff Answer

"Not with 2M hard purges at once. Options in order of preference: If prices are split into a small price fragment with short TTL, the import needs no purge at all — the TTL catches it. If pages include prices, soft-purge by tag, rate-limited to a few thousand tags per second and ordered by popularity so hot products refresh first; the long tail refreshes on TTL. Soft purge means requests keep being served stale while shields revalidate, so the origin sees at most ~6 fetches per product, spread over the purge window.

And I'd ask whether the import should purge at all: if 95% of those products receive fewer than one view per hour, their cached copies will expire before anyone sees them."

Why this is L6:

  • Rejects the naive bulk hard purge
  • Uses popularity ordering, soft purge and rate limits
  • Questions whether purge is needed

What L7 adds:

  • Sets per-publisher purge quotas and a review process for broad tags

Drill 6: Hot Object#

Prompt: "A single video segment gets 2M requests per second during a live event. What breaks?"

Staff Answer

"The cache hit ratio is fine; the problem is per-server load. Inside a PoP, objects are typically sharded to cache servers by consistent hashing of the key, so one hot key lands on one server. At 2M RPS globally across 300 PoPs, the busiest PoPs may see 50K RPS for that object — beyond a single server's NIC for multi-MB segments.

Fixes: replicate hot objects across multiple servers in the PoP (hot-key detection triggers replication to N servers, and the PoP load balancer spreads requests); serve from memory rather than disk; for live, use short segments with predictable names so the edge can prefetch the next segment; and collapse misses for the next segment so origin sees one fetch per shield."

Why this is L6:

  • Identifies intra-PoP hot-key concentration, not hit ratio, as the bottleneck
  • Offers replication, memory tiering, and prefetch

What L7 adds:

  • For scheduled mega-events, pre-provisions capacity with the CDN and tests a rehearsal event — the contract, not the architecture, is the lever

Drill 7: Build vs Buy#

Prompt: "Should we build our own CDN?"

Staff Answer

"Almost certainly not. A commercial CDN amortizes hundreds of PoPs, peering relationships, DDoS capacity and 24/7 operations across thousands of customers. Building means capex for servers in dozens of locations, peering negotiations, a network operations team — tens of engineers — and years to reach parity.

The exceptions are companies whose traffic is a meaningful share of the internet and highly predictable — Netflix built Open Connect because at their volume, owning the edge and placing boxes inside ISPs is cheaper and better. For us, the better levers are negotiating committed-use pricing, adding a second CDN for leverage and resilience once spend justifies it, and keeping our configuration portable. See the technology guide for how edge proxies compare."

Why this is L6:

  • Names the specific conditions that justify building
  • Offers realistic cost levers instead

What L7 adds:

  • Models the crossover point: at what monthly egress does a hybrid (own caches in top 5 metros + commercial CDN elsewhere) break even?

Drill 8: Policy Change Without an Outage#

Prompt: "You need to change the cache key for all product pages to add a currency dimension. How do you roll it out?"

Staff Answer

"A key change effectively empties the cache for those pages, because every new key is a miss. So: stage it by PoP group — 5% of PoPs first — and watch origin load, hit ratio and error rate. Before starting, warm the shield: shields can use the new key while edges are still on the old one, so the shield fills gradually. Pick a low-traffic window. Make sure origin can absorb the miss rate for the canary PoPs — roughly the canary fraction × the per-object miss rate.

Validate correctness in canary: two synthetic users with different currencies request the same product; assert different currencies in the body and distinct cache entries. Rollback is a config revert, which is also a key change — so rollback has its own miss cost; plan origin capacity for both directions."

Why this is L6:

  • Recognizes a key change as a cache flush
  • Stages by PoP group with origin capacity checks
  • Tests correctness and plans rollback cost

What L7 adds:

  • Makes edge config changes flow through a pipeline with automated canary analysis and PoP-group staging for every team

Drill 9: Cost#

Prompt: "Our CDN bill is $400K/month and growing 8% a month. What do you do?"

Staff Answer

"First, attribute it: bytes by route, content type and team. Usually a few items dominate — unoptimized images, video, large JS bundles, or API responses that bypass cache. Then pull the levers by impact:

  • Bytes: image format negotiation (AVIF/WebP typically 25–50% smaller than JPEG), resizing to device, Brotli for text. Often 20–40% of the bill.
  • Hit ratio: fixes to cache keys raise offload and cut origin egress, which is often billed separately by the cloud provider.
  • Contract: committed-use pricing at our volume; a second CDN as leverage.
  • Architecture: long-tail media served from cheaper tiers.

Then set a unit metric — CDN $ per 1,000 sessions — so growth in the bill is measured against growth in the business."

Why this is L6:

  • Attributes before optimizing
  • Ranks levers by impact with numbers
  • Introduces a unit-cost metric

What L7 adds:

  • Allocates CDN cost to teams by route so owners see their own egress
  • Negotiates multi-year contracts with volume tiers aligned to the growth forecast

Drill 10: Multi-Region Origin#

Prompt: "We're adding a second origin region. How does the CDN use it?"

Staff Answer

"Each shield gets a preferred origin — EU shield to EU origin, US shields to US origin — with failover to the other origin on health-check failure. That halves the long-haul miss latency and gives origin redundancy.

Care points: both origins must produce identical cacheable responses for the same key — same data version — or users flip between versions. Purges must hit both origins' caches if they have their own. And failover means one origin takes 2× miss traffic, so each origin must handle the full miss load, or shed. CDN-level origin failover must be faster than DNS — health checks every ~5–10s."

Why this is L6:

  • Maps shields to origins
  • Flags version consistency across origins and failover capacity

What L7 adds:

  • Decides whether the second region is active-active for writes too, and aligns edge routing with data-residency requirements

8. Deep Dive Scenarios#

Deep Dive 1: Launch-Day Origin Meltdown Behind a CDN#

Context: A limited-edition product launches at 10:00. The product page is behind the CDN with s-maxage=60. At 10:00:05, origin CPU hits 100% and error rates reach 40%. The CDN dashboard shows a 35% hit ratio for the product page, far below the expected 99%. You're pulled in.

Questions to Surface First:

  • Why is the hit ratio 35% on a single hot URL — is the cache key fragmented?
  • Is request collapsing enabled, and is there a shield?
  • Is the page actually cacheable — is the origin emitting Set-Cookie or private under load (e.g., a session started for the queue)?
  • Is the origin serving errors that are being cached, or bypassing cache?

Typical L5 Approach: Scales origin, raises the TTL. Scaling takes minutes; raising TTL doesn't help if requests are missing for another reason.

Staff Approach: Diagnoses fragmentation first. Finds the launch email links carried a unique ref parameter per recipient, which wasn't stripped — every click was a distinct cache key. Pushes an emergency edge rule stripping unknown query parameters for that route (staged to 10% of PoPs for 60s, then global), confirms hit ratio climbs above 95% within two minutes, and enables stale-if-error so remaining origin errors aren't shown.

Principal Approach: Treats "a marketing link format can take down the origin" as a platform defect. Makes query-parameter allowlists the default for every cacheable route, adds a pre-launch checklist that includes a synthetic load test through the CDN using the real campaign URLs, and establishes a joint marketing–engineering launch calendar.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Check hit_ratio and top cache keys for the route. Enable SIE and micro-cache fallback.
TriageKey cardinality spike → unstripped ref parameter.
Quick fixEdge rule: allowlist params for /p/*; staged rollout in minutes.
GuardrailsOrigin shedding on non-critical routes until hit ratio > 95%.
Post-mortemDefault param allowlist; launch load test through the CDN with real links.

Metrics to Watch: hit_ratio{route}, cache_key_cardinality{route}, origin_rps, origin_cpu, edge_5xx_rate.

Organizational Follow-up: Launch checklist owned by the platform team with marketing sign-off on link formats.

Ownership Question: "Who owns the fact that marketing links bypass the cache?" Staff answer: The platform. Marketing will always add tracking parameters; the edge's job is to ignore what the page doesn't use. Making marketing responsible for cache keys is asking the wrong team to understand the wrong system.

Key Takeaway: "When a single hot URL has a low hit ratio, the cache key is fragmented until proven otherwise."

What clears the Staff bar:

  • Looks at key cardinality before scaling
  • Fixes at the edge in minutes with staging
  • Makes the safe behavior the platform default

Deep Dive 2: The Silent Leak#

Context: A security researcher emails that loading /orders/recent.css on your site shows the page of whoever last requested it. Your team confirms: a new framework treats unknown extensions as the base route, so /orders/recent.css renders the user's order page — and the CDN caches anything ending in .css for 1 day. This is web cache deception.

Questions to Surface First:

  • How long has this been possible, and which routes are affected?
  • Do logs show exploitation (unusual extensions on personalized routes)?
  • Why does the CDN cache by extension rather than by the origin's Cache-Control?
  • Does the origin mark personalized responses as private?

Typical L5 Approach: Adds a rule to not cache /orders/*. Fixes one route; the pattern applies to every personalized route.

Staff Approach: Fixes the class: the edge must never override the origin's Cache-Control by extension; personalized responses always carry Cache-Control: private, no-store plus a private marker the edge refuses to cache; the framework must 404 unknown extensions on application routes. Purges all cached content on personalized paths; reviews logs for exploitation; engages security and privacy for disclosure assessment.

Principal Approach: Establishes that caching behavior is a security control. Edge caching rules go through security review; a quarterly external test for cache deception and poisoning is part of the security program; the "private by default" posture is written into the edge standard and verified in CI for every route.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateDisable extension-based caching rules; purge personalized paths globally.
TriageLog search for suspicious extensions on personalized routes; exposure estimate.
Quick fixOrigin: private, no-store + private marker on all authenticated responses.
GuardrailsEdge refuses to cache marked responses; CI two-user test; framework rejects unknown extensions.
Post-mortemEdge rules overrode origin intent; no defense in depth.

Metrics to Watch: private_marker_in_cached_response_total, cached_responses_on_auth_routes_total, suspicious-extension request rate.

Organizational Follow-up: Security review for edge rules; bug bounty scope includes caching.

Ownership Question: "Who owns this vulnerability — the framework team, the edge team, or the orders team?" Staff answer: All three had a gap, but the edge platform owns the fix that makes the class impossible: never cache a response the origin marked private, regardless of path. Defense in depth means no single team's mistake leaks data.

Key Takeaway: "The edge should never be smarter than the origin about what's private."

What clears the Staff bar:

  • Fixes the class, not the route
  • Defense in depth across framework, origin and edge
  • Engages security and privacy immediately

Deep Dive 3: Onboarding a Large Customer Surface — The Public API#

Context: The product API team wants to put its public read API (1.2M RPS peak, 40% of which are identical catalog queries) behind the CDN. Responses depend on the API key's plan tier (some fields hidden for free tier). The API team is worried about stale data and about the CDN seeing API keys.

Questions to Surface First:

  • Does the response vary by API key, or only by plan tier?
  • What freshness does the API contract promise?
  • Where are rate limits enforced — would caching bypass them?
  • Is the API key a secret that must not appear in cache keys or logs?

Typical L5 Approach: Adds the Authorization header to the cache key. Correct but fragments the cache by customer — hit ratio collapses to near zero for the long tail of keys.

Staff Approach: Validates the API key at the edge (edge function checks a signed token or looks up key → tier in a replicated edge KV), then keys the cache by tier (3 values), not by key. Rate limiting still runs per key at the edge before cache lookup, so cached responses count against quotas. Cache TTL matches the API's documented freshness (e.g., 30s) with Age header exposed to clients. The raw key is stripped from the origin request and never logged at edge.

Principal Approach: Frames it as a product decision: caching changes the API's freshness contract, which should be documented and versioned. Aligns API gateway and CDN responsibilities so there's one place for authentication and rate limiting (API Gateway), and one for caching — avoiding two teams implementing auth.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignEdge auth → tier; key = path + normalized params + tier.
Rate limitingPer-key at edge, before cache lookup.
FreshnessTTL 30s + SWR 60s; Age exposed; documented in API reference.
RolloutShadow: edge computes cache decisions, compares responses to origin for 1 week; then 5% → 100%.
SuccessOrigin RPS −35%; p75 latency outside origin region −150ms; zero tier-leak incidents.

Metrics to Watch: hit_ratio{route=api}, tier_mismatch_detected_total, rate_limit_rejections{cache_status}.

Organizational Follow-up: API docs updated with freshness semantics; gateway and edge teams agree on auth ownership.

Ownership Question: "If a free-tier customer receives a pro-tier field from cache, whose bug is it?" Staff answer: The edge platform's key design. Tier must be in the key, and a synthetic test with keys from each tier runs continuously. The API team owns the tier-to-fields mapping; the platform owns guaranteeing it's honored in cache.

Key Takeaway: "Key the cache by the smallest dimension the response actually depends on — the tier, not the customer."

What clears the Staff bar:

  • Keys by tier, not identity
  • Keeps rate limiting ahead of cache
  • Documents freshness as part of the API contract

Deep Dive 4: Post-Mortem — The Provider Outage#

Context: Your sole CDN provider had a 45-minute global outage. Your site was fully down. Leadership asks for a post-mortem and a plan so "this never happens again."

Questions to Surface First:

  • Did we have a failover plan? Was it executed, and if not, why?
  • What's the DNS TTL on the edge hostname, and who can change DNS in an incident?
  • Could the origin have served traffic directly? At what fraction?
  • What did 45 minutes cost?

Typical L5 Approach: Proposes adding a second CDN.

Staff Approach: Finds the real gaps: edge hostname had a 1-hour DNS TTL (failover would have taken an hour anyway), no runbook with a decision threshold, and origin only locked to CDN IPs — so direct traffic was impossible without a firewall change. Fixes: TTL 60s, runbook with decision criteria and pre-approved authority, origin able to accept direct traffic via a pre-provisioned LB behind a separate hostname, and shedding rules to cope with the load. Evaluates a secondary CDN as a separate decision with costs.

Principal Approach: Answers "never again" honestly: no plan makes a dependency infallible; the question is how much to spend on reducing expected downtime. Quantifies: expected provider outage time per year × revenue per minute versus the cost of active multi-CDN (config duplication, split volume discounts, ~1–2 additional engineers). Brings leadership a decision with numbers rather than a reflexive "add another vendor."

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (future incidents)Decision threshold: > 25% global error from external synthetics for 3 min → failover.
TriageWhy no failover: TTL, runbook, origin lockdown.
Quick fixDNS TTL 60s; direct-origin hostname and LB pre-provisioned; runbook drill.
GuardrailsQuarterly failover drill at low traffic; origin shedding plan for 30% of traffic.
Post-mortemThe plan existed as an idea, not as tested infrastructure.

Metrics to Watch: External synthetic success rate by provider; dns_ttl{hostname}; time-to-failover in drills.

Organizational Follow-up: Named failover owner and on-call authority; multi-CDN business case.

Ownership Question: "Who has the authority to move all traffic off the CDN?" Staff answer: The incident commander, per runbook criteria, without needing VP approval at 3 AM. The VP approves the criteria in advance.

Key Takeaway: "A failover plan that's never been executed is a hypothesis."

What clears the Staff bar:

  • Finds that the existing plan was unexecutable (TTL, lockdown)
  • Separates quick fixes from the multi-CDN business decision
  • Pre-approves authority

Deep Dive 5: Multi-Region Expansion with Data Residency#

Context: The company is entering the EU and must keep EU users' personal data in the EU. The CDN serves global traffic from all PoPs. Legal asks: "Does the CDN store personal data outside the EU?"

Questions to Surface First:

  • What does the CDN cache — is any of it personal data? (Should be none, if personalized = no-store.)
  • What about logs — edge logs include IP addresses, which are personal data under GDPR?
  • Do EU users ever get routed to non-EU PoPs (anycast doesn't respect borders precisely)?
  • Do edge functions process personal data (e.g., tokens with user IDs)?

Typical L5 Approach: Adds an EU origin and assumes the CDN is fine because it's "just a cache."

Staff Approach: Audits all three data flows: cache content (verify no personal data is cached — the private-by-default guardrail is now a compliance control), logs (configure regional log delivery and IP truncation), and edge compute (tokens validated at edge contain only pseudonymous IDs). For personalized routes, uses the CDN's regional services feature (where available) to restrict TLS termination and processing to EU PoPs for EU hostnames.

Principal Approach: Treats the edge as part of the company's data-processing footprint — it needs a data-processing agreement, a documented data map, and a review whenever new edge features (logging, compute, KV) are adopted. Establishes that the edge platform team is accountable to the privacy program, not only to engineering.

Staff Approach — Full Reasoning
PhaseWhat to Do
AuditCache content, logs, edge compute — each checked for personal data.
ControlsPrivate-by-default verified; regional log pipelines; IP truncation.
RoutingEU hostnames with EU-only processing for personalized routes.
ValidationSample cached objects; confirm no personal data.
DocumentationData map and DPA updated; legal sign-off.

Metrics to Watch: private_marker_in_cached_response_total, log delivery region, EU-hostname PoP distribution.

Organizational Follow-up: Privacy review for new edge features.

Ownership Question: "Who signs off that the edge is compliant?" Staff answer: Legal/privacy signs off, based on a data map the edge platform team maintains. Engineering can't self-certify compliance.

Key Takeaway: "'It's just a cache' is not a compliance position. Know what the edge stores, logs and processes."

What clears the Staff bar:

  • Audits cache, logs and compute separately
  • Reuses private-by-default as a compliance control
  • Involves legal for sign-off

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Classify content into immutable, semi-dynamic, personalized and never-cache, and give each a strategy
  • Design an allowlist cache key with normalization, and explain its hit-ratio and security consequences
  • Combine TTL, stale-while-revalidate, stale-if-error and tag purge to bound freshness twice
  • Quantify origin protection from shields, collapsing and soft purge
  • Identify cache poisoning, cache deception and private-content leakage and prevent each class
  • Plan for CDN provider failure with DNS TTLs, runbooks and origin capacity
  • Reason about CDN cost as egress × (1 − offload) plus contract strategy
  • Decide what logic belongs at the edge and how to roll it out safely

The Bar for This Question#

Mid-level (L4): Knows CDNs cache static content near users; sets TTLs; mentions invalidation.

Senior (L5): Configures a CDN competently: versioned assets, TTLs, purge on deploy, TLS at edge. Misses cache-key design as a security boundary, purge stampedes, stale-if-error, vendor failure and cost.

Staff+ (L6): Classifies content, designs allowlist keys with leakage guardrails, bounds freshness with TTL plus purge, protects origin with shields and soft purge, plans for origin and CDN failure, treats edge config as code with staged rollout, and names owners for freshness and headers. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "The Cache Key Is a Security Boundary, Not a Performance Setting"#

Key Too BroadKey Too Narrow
Lower hit ratio, higher costWrong content served — possibly another user's
Recoverable, measurablePrivacy incident, possibly reportable

The Staff position: Design the key with security review, default routes to non-cacheable, and test with two users in CI.

Why this matters in interviews: Candidates discuss keys only for hit ratio. Saying "leakage" first is a strong signal.

10.2 "Purge Is an Optimization; TTL Is the Guarantee"#

Purge pipelines fail silently — dropped events, rate limits, publishers who forget. Designs that rely on purge for correctness with multi-day TTLs are one lost message away from serving a 3-day-old price.

The Staff position: TTL bounds staleness; purge accelerates freshness. The exception is providers with reliable instant purge and purge verification — then "cache long, purge on change" is sound.

Why this matters in interviews: It demonstrates you think about failure of the freshness mechanism itself.

10.3 "Most Teams Should Cache Less at the Edge, Not More"#

Aggressive edge caching of semi-dynamic and personalized-adjacent content produces the most dangerous edge incidents (leaks, stale prices, poisoned pages), while the latency benefit for many of those routes is achieved by TLS termination and connection reuse alone.

The Staff position: Cache static aggressively; cache semi-dynamic deliberately with contracts; for dynamic routes, use the edge for TLS and connection reuse without caching.

Why this matters in interviews: It shows you distinguish the edge's two value propositions: caching and proximity.

10.4 "Multi-CDN Is Usually a Negotiating Tactic, Not a Reliability Strategy"#

Multi-CDN BenefitMulti-CDN Cost
Survives a single provider outage~2× config, testing and purge fan-out; features limited to common denominator
Pricing leverageSplit volume reduces tier discounts
Performance steeringRequires RUM-based steering infrastructure

The Staff position: A tested DNS failover path gets most of the reliability benefit for a fraction of the cost. Active multi-CDN is justified at high revenue per minute or when the pricing leverage pays for the effort.

Why this matters in interviews: Reflexively proposing multi-CDN is an L5 signal; pricing it is L6/L7.

10.5 "Edge Compute Is a Distributed System You Deploy to 300 Cities at Once"#

Running code at the edge moves you from one deploy target to hundreds, with limited observability, constrained state and a global blast radius. The latency win is real — tens to hundreds of milliseconds for distant users — but only for logic that is stateless and stable.

The Staff position: Token validation, routing, A/B bucketing, header normalization: yes. Business logic with state: no, until you have staged rollout, per-PoP observability and fast rollback.

Why this matters in interviews: It shows you evaluate edge compute as an operational commitment, not a feature.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

At L6, the edge is a well-configured caching layer with safe keys and bounded freshness. At L7, the edge is the company's front door: the first code every customer request touches, a security control plane (TLS, WAF, bot defense, DDoS), a data processor subject to privacy law, a six- or seven-figure monthly line item, and a single vendor relationship whose outage is a total outage. The Principal decisions are about scope (what the edge owns), governance (who can change a rule that runs in 300 cities), money (egress contracts, cost allocation, build-vs-buy crossover), and vendor risk (lock-in via proprietary edge compute and config dialects).

The Org-Level Fault Line#

A central edge platform vs teams configuring the CDN directly.

OptionWhat It BuysWhat It CostsWho Pays
Teams configure the CDN console directlySpeed; no bottleneckInconsistent keys, leak risk, no staged rollout, surprise costsSecurity (incidents), finance (bills), everyone (outages from a rule change)
Central edge team owns all configConsistency and safetyEvery change queues behind one team; product velocity suffersProduct teams (waiting)
Edge platform with self-service contract (the L7 default)Teams control cache policy via response headers and a reviewed route manifest; platform owns guardrails, pipeline, and vendorPlatform must build the pipeline, guardrails and testsPlatform headcount

🧭 Principal Move: "Teams own their headers; the platform owns the guardrails and the blast radius. No one — including the platform team — pushes a rule to all PoPs at once."

Cost Model#

Assumptions: blended CDN egress ~$0.01–0.05/GB depending on commitment and geography; origin/cloud egress ~$0.05–0.09/GB; fully loaded engineer ~$300K/year; figures are order-of-magnitude for planning, not quotes.

ScaleTrafficMonthly CDNOrigin Egress SavedHeadcountOn-call
Startup~50 TB/month~$2–5K (often pay-as-you-go)~$2–4K~0.2 FTEShared; CDN outages are rare pages
Growth~2 PB/month~$40–100K with commit~$100–150K2–3 FTE edge/platformPlatform rotation; purge and config incidents dominate
Large~50 PB/month, video-heavy~$0.5–1.5M with negotiated tiersMillions8–15 FTE (edge, security, contracts)Dedicated edge rotation; multi-CDN steering

The strategic numbers: every point of byte offload above 95% saves roughly 1% of total bytes from origin egress, and image/format optimization often cuts total bytes 20–40%. Negotiated commits typically matter more than any single engineering optimization at large scale.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to ReverseWhy
TTLs, purge rules, key allowlistsTwo-wayMinutes (with miss cost on key changes)Tune continuously
Asset URL versioning schemeMostly one-wayEvery build pipeline and HTML templateGet it right early
Proprietary edge compute (vendor-specific runtime/KV)Expensive two-wayRewrite of edge logic per vendorLock-in; keep logic small and portable
Multi-year CDN commitOne-way for the termContract penaltiesAlign with growth forecast
Building own PoPs / ISP cachesOne-wayCapex, peering, teamOnly at extreme, predictable volume
Public hostnames and URL structureOne-waySEO, links, clientsURLs are forever

The Standard I'd Write#

RFC: Edge Caching & Delivery Standard v1

Scope: All public hostnames served through the company edge.

Requirements:

  • Routes MUST be non-cacheable in the shared cache unless declared cacheable in the route manifest.
  • Responses that depend on user identity MUST send Cache-Control: private, no-store and the private marker header; the edge MUST refuse to cache marked responses.
  • Cacheable routes MUST declare a cache-key allowlist; unknown query parameters, cookies and Authorization MUST be stripped before cache lookup and origin fetch.
  • Static assets MUST use content-hashed URLs with max-age=31536000, immutable; they MUST NOT be purged.
  • Cacheable routes SHOULD set stale-if-error ≥ 1 hour and stale-while-revalidate ≥ TTL.
  • Purges SHOULD be soft by default; broad tags require platform approval; per-publisher purge rate limits apply.
  • Edge config and code MUST roll out by PoP group with automated health gates; no global instantaneous deploys.
  • Edge hostnames MUST have DNS TTL ≤ 60s and a tested failover target.

Exceptions: Filed with the edge platform; security review required for any exception to private-by-default.

Success metrics: Zero private-content cache incidents; byte offload ≥ 95% for static, ≥ 80% for semi-dynamic routes; purge p99 ≤ 5s; failover drill completed quarterly with time-to-shift ≤ 10 min; CDN $ per 1,000 sessions trending flat or down.

What I'd Tell the VP#

"Almost every customer request enters through our CDN, so it's both our biggest performance lever and a single point of failure. We're putting three things in place: rules that make it impossible for a team to accidentally cache one customer's private page for another, a way to update prices and content globally within seconds without overloading our servers, and a tested plan to move traffic elsewhere if the CDN provider has an outage like the ones that took down large parts of the web in 2019 and 2021. We'll also charge each team for the bandwidth their pages use, which historically cuts the bill 20–30% once people can see it. A second CDN is on the table — I'll bring you the cost versus the expected downtime it avoids so we can decide with numbers."

Principal Interview Signals#

SignalWhat It Sounds Like
Defines edge scope"The edge owns TLS, caching, WAF and stateless routing — never business state."
Governs blast radius"Every edge change rolls out by PoP group; nobody pushes globally, including us."
Prices vendor risk"Multi-CDN costs ~2 engineers and our volume discount; one provider outage a year costs X. Here's the break-even."
Manages lock-in"Edge logic stays small and portable; proprietary KV only for caches we can rebuild."
Treats edge as a data processor"Logs contain IPs — the edge is in our privacy data map."

Staff answers that L7 interviewers find insufficient:

  • A perfectly configured CDN with no governance for the 30 teams who will change it next year.
  • "Add a second CDN" without cost, feature-parity and purge fan-out analysis.
  • No view of the contract — commit levels and pricing tiers are often the biggest cost lever.

Appendices

Appendix A: Mechanics in Depth#

A.1 Edge Request Handling#

handle(req):
  key = normalize(req)                         # allowlist params, normalized headers
  if route_not_cacheable(req): return proxy_to_origin(strip(req))
  obj = cache.get(key)
  if obj and obj.fresh():                      return obj                      # HIT
  if obj and obj.within_swr():                 async revalidate(key); return obj   # STALE, revalidating
  resp = collapse(key, lambda: fetch_from_shield(strip(req)))   # one fetch per key in flight
  if resp.is_error() and obj and obj.within_sie(): return obj   # serve stale on error
  if cacheable(resp) and not resp.has('Set-Cookie') and not resp.has('X-Private'):
      cache.put(key, resp, ttl=resp.s_maxage)
  return resp

A.2 Why Versioned URLs Beat Purge for Assets#

ApproachFreshnessHit RatioFailure Mode
Fixed URL + purge on deployDepends on purge + browser caches (which you can't purge)Dips after each deployBrowsers keep old asset for its max-age
Content-hashed URLNew URL = instantly fresh everywhere, including browsers~100%Old HTML referencing removed assets → keep old assets ≥ 1 week

A.3 Admission and Eviction#

  • LRU is the common baseline; video and large-object caches use segmented LRU or size-aware policies.
  • Admission filters (e.g., only cache on second request within a window, tracked with a Bloom filter) keep one-hit wonders from evicting popular content — valuable because a large share of objects are requested once.
  • Tiered storage: RAM for the hottest objects, SSD for the warm set, larger disk tiers at shields.

A.4 PoP Internals (If You're Building the CDN)#

Diagram: A.4 PoP Internals (If You're Building the CDN)

Front servers terminate TLS and route by consistent hash of the cache key to the cache server that owns it, so each object is stored once per PoP. Hot objects are replicated to multiple cache servers to avoid single-server saturation. See Consistent Hashing.


Appendix B: Cache Keys and Variants#

B.1 Key Construction#

ComponentRuleExample
HostLowercase; canonical host onlywww.example.com
PathLowercase if case-insensitive routes; collapse slashes/p/red-shoes
QueryAllowlist, sorted, deduplicatedcolor=red&size=9
HeadersNormalized to small enumslang=de, dev=mobile
EncodingNegotiated separately (br, gzip, identity)enc=br
CookiesExcluded; stripped for cacheable routes—

B.2 Vary Discipline#

Vary ValueVariantsVerdict
Accept-Encoding2–3Fine (CDN normalizes)
Accept-Language (raw)Thousands of stringsNormalize to supported languages first
User-AgentTens of thousandsNever; derive device class at edge
CookieOne per userNever on shared cache

B.3 Route Classes#

ClassExamplesPolicy
Immutable/assets/*.{hash}.js1 year, immutable
Semi-dynamic/p/*, /c/*, public API GETs30–300s TTL, SWR, SIE, tag purge
Micro-cachedHot dynamic endpoints1–10s TTL, collapsing
Private/account, /cart, /checkoutprivate, no-store; edge for TLS only
Never through edgeAdmin, internalSeparate hostname, not exposed

Appendix C: Invalidation Mechanisms#

MechanismPrecisionSpeedOrigin ImpactBest For
TTL expiryPer objectBounded by TTLSmoothBaseline freshness
URL purgeSingle URLSeconds–minutes depending on providerStampede if hotIndividual corrections
Prefix/wildcard purgeMany URLsVariesLarge stampede riskEmergency
Tag / surrogate-key purgeEvery object taggedSub-second on some providersProportional to tag breadthSemi-dynamic content
Soft purgeAs above, marks staleSameMinimal (revalidation)Default for all purges
Versioned URLNew objectInstantNoneStatic assets

C.1 Quick Comparison#

QuestionTTLTag PurgeVersioned URL
Works if messages are lost?YesNoYes
Freshness≤ TTL~secondsInstant
Needs publisher discipline?NoYes (tagging, purging)Build pipeline only

Appendix D: HTTP Contract and Client Behavior#

D.1 Headers Cheat Sheet#

HeaderPurposeTypical Value
Cache-Control: s-maxageShared-cache TTLs-maxage=60
Cache-Control: max-ageBrowser TTLmax-age=0 for HTML, 1 year for hashed assets
stale-while-revalidateServe stale during background refresh300
stale-if-errorServe stale on origin error86400
Surrogate-Key / Cache-TagPurge targeting (provider-specific)product-123 cat-9
Surrogate-ControlEdge-only directives (some providers)max-age=3600
VaryDeclares request headers that change the responseAccept-Encoding
AgeSeconds object has been in cacheExposed for debugging and API freshness
ETag / Last-ModifiedConditional revalidation (304)Saves bytes on revalidation

D.2 Browser vs Edge TTL#

Set HTML's browser max-age=0 (or short) with a longer s-maxage — you can purge the edge but not browsers. Hashed assets can be long in both.

D.3 Origin Lockdown#

Accept origin traffic only from CDN egress IP ranges and require a secret header or mTLS certificate from the edge, rotated regularly. IP ranges alone are shared by every customer of the same CDN.


Appendix E: Observability#

E.1 Core Metrics — Non-Negotiable#

# Effectiveness
edge_hit_ratio{route}              byte_offload_pct{route}
shield_hit_ratio                   origin_rps / origin_egress_gbps
cache_key_cardinality{route}       # fragmentation detector

# User experience (RUM + synthetics)
ttfb_p75{country}                  lcp_p75{country}
edge_error_rate{pop}               external_synthetic_success{provider}

# Freshness
purge_latency_p99                  purge_requests_total{publisher}
served_stale_total{reason=swr|sie} content_version_mismatch_total   # sampled checks

# Safety
private_marker_in_cached_response_total   two_user_check_failures_total

E.2 Critical Alerts#

AlertConditionAction
Private content cachedAny private_marker_in_cached_response_totalPage security + platform; purge
Two-user check failureAnyPage platform
Hit ratio collapseRoute hit ratio drops > 20 points in 10 minPage platform; check key cardinality
Origin overload via missesorigin_rps > 3× baseline with hit ratio dropPage origin + platform
Provider degradationExternal synthetic success < 90% for 3 minPage SRE; evaluate failover
Purge latencypurge_latency_p99 > 30sTicket; page if freshness SLO at risk

E.3 Control Plane vs Data Plane#

The CDN's config/purge APIs (control plane) can fail independently of serving (data plane). Keep an out-of-band way to change DNS and to disable caching rules, and never make publishing depend synchronously on purge success — queue and retry.

E.4 Debugging "Users See Old Content"#

  1. Check Age and version headers from several PoPs.
  2. Was a purge issued? Did it succeed? (purge logs, latency)
  3. Is the browser caching it (max-age on HTML)?
  4. Is the shield serving stale (SWR/SIE because origin is erroring)?
  5. Is the cache key including something that splits old and new?

Appendix F: Scale Evolution#

F.1 What Works at Each Scale#

ScaleEnoughAdd When
EarlyCDN for static assets with versioned URLsGlobal users or launch traffic
GrowthSemi-dynamic caching, shield, SWR/SIE, tag purge, allowlist keysMultiple teams changing edge config
ScaleEdge platform contract, staged config rollout, cost allocation, failover drillsProvider outage cost or bill growth
Very largeMulti-CDN steering, own caches in top metros, pre-positioning for mediaTraffic and predictability justify capex

F.2 Multi-Region Path#

Single origin → regional origins with shield mapping → data-residency-aware hostnames → per-region edge policies.

F.3 What You Don't Build on Day One#

  • Multi-CDN steering
  • Edge compute beyond redirects and header normalization
  • Your own PoPs
  • Custom admission policies
  • Real-time purge verification infrastructure (start with sampled checks)

Appendix G: Multi-Tenancy, Fairness and Cost#

G.1 Shared Cache Capacity#

In a shared edge, one team's large objects (video, downloads) can evict another team's hot pages. Separate cache namespaces or storage tiers per content class; admission filters for large objects.

G.2 Purge Fairness#

A single publisher issuing broad purges can degrade freshness for everyone by consuming purge capacity. Per-publisher rate limits; priority lanes for corrections and legal takedowns.

G.3 Cost Allocation#

Attribute egress by route prefix or hostname to owning teams; publish monthly. Visibility alone typically drives meaningful reductions (image optimization, dropping unused bundles).

G.4 Tradeoff Summary#

ChoiceDefaultWho Pays If Wrong
CacheabilityOpt-in per routeSecurity, if opt-out
Cache keyAllowlistUsers (wrong/leaked content) or finance (fragmentation)
FreshnessTTL + SWR + tag purgeBusiness (stale data)
PurgeSoft, rate-limitedOrigin (stampede)
VendorSingle + tested failoverBusiness during provider outage
Edge logicStateless, staged rolloutEveryone, on a global bad push
  1. Loading the index…