Technologies referenced in this case study: API Gateways & Edge Proxies · Redis
Related case studies: Distributed Caching · Blob Storage · Load Balancer · API Gateway · Rate Limiting · Caching Fundamentals · Scaling Reads · Large Blobs
How to Use This Case Study#
Organized for interview use first, reference second.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 2, 5 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → §3 Fault Lines → §4 Failure Modes → Deep Dives 2 and 4 |
| Deep Dive | 3+ hrs | Everything, including §11 Principal Lens and appendices |
What is a CDN & Edge Caching? — Why interviewers pick this topic
A content delivery network is a globally distributed fleet of caching proxies (points of presence, PoPs) placed close to users. Requests are routed to a nearby PoP (via anycast or DNS); the PoP serves from cache on a hit, or fetches from the origin (often through an intermediate shield tier) on a miss. Beyond caching, the edge terminates TLS close to users, reuses warm connections to origin, absorbs DDoS traffic, and increasingly runs code.
Before vs After — the product launch:
Without a well-designed edge:
t=0: Launch email to 20M users; product page and images served from us-east origin
t=+30s: Traffic 40x normal; users in Sydney see 1.8s TTFB (4 RTTs at ~220ms each)
t=+1min: Origin egress saturates at 20 Gbps; image requests queue
t=+2min: Origin app servers at 100% CPU rendering the same page 50,000 times per second
t=+5min: Errors spike; launch trends on social media for the wrong reason
With a well-designed edge:
t=0: Same email; page cached at edge with max-age=60, stale-while-revalidate=300
t=+30s: 97% of requests served from 300+ PoPs; Sydney TTFB ~40ms
t=+30s: Shield tier collapses misses: origin sees ~1 request per object per shield per minute
t=+1min: Origin at 15% CPU; egress ~0.5 Gbps
t=+10min: Price change published → surrogate-key purge → global in < 1s
Result: A launch, not an incident.
Why interviewers reach for this question: Every candidate knows "put a CDN in front of it." Few can reason about what's cacheable, what the cache key is, how fast a change must propagate, what happens when the CDN itself fails, or how a caching misconfiguration can leak one user's private page to another. It tests caching judgment at global scale, security, cost, and vendor strategy — all in one prompt.
Mechanics Refresher: Edge Caching Building Blocks
| Mechanism | How It Works | Pros | Cons |
|---|---|---|---|
TTL expiry (Cache-Control: max-age, s-maxage) | Object served until age exceeds TTL, then revalidated or refetched | Simple, predictable | Freshness bounded by TTL; long TTL = stale, short TTL = low hit ratio |
Versioned URLs (app.3f9a2c.js) | Content hash in the URL; new content = new URL; TTL = 1 year, immutable | Perfect freshness and hit ratio | Only for assets referenced by something else (HTML, manifest) |
| Purge / invalidation | Explicit removal by URL, prefix, or tag | Freshness on demand | Propagation time varies (sub-second to minutes); purge storms hit origin |
| Surrogate keys / cache tags | Tag objects (product:123); purge by tag | Precise invalidation of every page that shows product 123 | Tagging discipline across teams |
| stale-while-revalidate / stale-if-error (RFC 5861) | Serve stale while fetching fresh in background; serve stale if origin errors | Hides origin latency and outages | Users see stale content for a bounded window |
| Request collapsing | Concurrent misses for the same key wait for one origin fetch | Prevents thundering herd on popular misses | Head-of-line waiting on slow origin |
| Tiered cache / origin shield | Edge misses go to a regional shield before origin | Origin sees 1 request per shield, not per PoP | Extra hop on miss (~10–50ms) |
| Anycast / GeoDNS routing | Users routed to a nearby PoP by BGP or DNS | Low latency, DDoS absorption | Anycast shifts with BGP; DNS needs resolver-location accuracy |
For most production systems: Versioned immutable URLs for static assets, short TTLs + stale-while-revalidate + surrogate-key purge for semi-dynamic HTML/API responses, a shield tier in front of origin, request collapsing, and explicit Cache-Control: private, no-store on anything personalized.
Executive Summary
If you only read one section, read this.
What This Interview Actually Tests#
A CDN is not a "put CloudFront in front of it" question. Everyone knows edges cache things.
This is a freshness, correctness and blast-radius question that tests:
- Whether you decide what is cacheable and for how long based on who is harmed by staleness
- Whether you design the cache key — the single most consequential and most error-prone decision in edge caching
- Whether you protect the origin from misses, purges, and cold caches — not just from users
- Whether you treat the CDN as a third-party dependency that can fail, leak, and lock you in
The key insight: The hit ratio is a business metric and the cache key is a security boundary. A 95% → 99% hit ratio cuts origin load 5×; a cache key missing one header can serve a logged-in user's account page to strangers. Staff engineers optimize the first without ever risking the second.
The L5 vs L6 vs L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Put a CDN in front of static assets" | Classifies content by freshness need and personalization: immutable, semi-dynamic, personalized, never-cache | Asks what the edge should own for the whole company: caching, TLS, WAF, bot defense, compute — and what it must never own |
| Freshness | Picks TTLs per file type | Versioned URLs for assets; short TTL + SWR + tag purge for pages; freshness SLO agreed with product | Makes purge latency and freshness contracts part of the platform SLA across all teams |
| Cache key | URL | Normalized URL + explicit allowlist of headers/cookies/query params; everything else stripped | Owns a cache-key policy standard; key changes reviewed like security changes |
| Origin protection | "The CDN reduces origin load" | Shield tier, request collapsing, purge rate limits, stale-if-error; origin sized for cold-cache scenarios | Prices origin capacity for worst-case cache loss vs CDN spend; decides on multi-CDN |
| Failure | "CDNs are highly available" | CDN outage plan: DNS failover to a second CDN or direct-to-origin with capacity math | Treats the CDN as a correlated dependency for the entire internet presence; vendor risk and exit strategy |
| Ownership | Frontend team configures the CDN | Platform owns edge config as code; product teams own cache headers for their responses | Edge platform team, cost allocation per team by egress, contract negotiation |
Why "cache key" separates levels
L5: The URL is the cache key; the CDN handles it. Reasonable, and dangerous in both directions. Include too little (ignore Accept-Language when the page varies by language) and users see the wrong content — or someone else's content. Include too much (every query parameter, including utm_source and fbclid) and the hit ratio collapses because each marketing link is a separate object.
L6: "The cache key is an allowlist, not a denylist. Path, a normalized set of query params the page actually uses, and at most one or two headers, like device class or language — normalized into a small number of values. Cookies are stripped for cacheable routes. Any response that depends on the user is private, no-store and never enters the shared cache."
L7: Recognizes that cache-key mistakes are among the most damaging incidents an edge can cause — data leakage — and makes key changes subject to security review and automated tests that request the same URL as two different users.
Why "origin protection" separates levels
L5: Assumes the CDN always shields the origin. It does — until a global purge, a new deploy with new URLs, a cache-key change, or a CDN node failure causes a wave of misses that all arrive at once.
L6: Sizes the origin for the miss scenarios: "With 300 PoPs and no shield, a purge of our homepage sends up to 300 simultaneous fetches; with a 6-region shield and collapsing, it sends 6. I'll also serve stale-if-error for 24 hours so an origin outage degrades freshness, not availability."
L7: Prices it: origin capacity to survive a full cold cache vs the CDN's committed spend, and decides which is cheaper insurance.
Why "failure" separates levels
L5: "The CDN is globally distributed, so it's highly available."
L6: "CDNs fail globally — a bad config push or software bug at the provider can take out most PoPs at once, as the Fastly (June 2021) and Cloudflare (July 2019) incidents showed. My plan: DNS-level failover to a second CDN or to origin, with low DNS TTLs on the edge hostname, and origin capacity to take some fraction of traffic directly."
L7: Decides whether multi-CDN is worth it: roughly doubles integration and config effort, dilutes volume discounts, but converts a total outage into a partial one. That's a business decision priced in revenue per minute.
The Staff Positions#
| Position | Rationale |
|---|---|
| Version your static assets; never purge them | Content-hashed URLs with 1-year immutable TTL give ~100% hit ratio and perfect freshness |
| Cache key is an allowlist | Every extra dimension fragments the cache; every missing one risks wrong or leaked content |
Personalized = private, no-store by default | Opt in to shared caching, never opt out; leakage is worse than a miss |
| Short TTL + SWR beats long TTL + purge for pages | Bounded staleness without depending on purge reliability |
| Always shield the origin | Tiered cache + collapsing turns 300 PoP misses into a handful of origin fetches |
| stale-if-error on everything cacheable | Origin outages degrade freshness, not availability |
| Plan for the CDN failing | It's a single global dependency; DNS failover and origin capacity are the plan |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Static asset delivery (JS, CSS, images, fonts) | Latency and hit ratio; huge fan-out | Versioned immutable URLs, long TTL, image optimization at edge | Deploy references unpublished asset (404 cached) | ~99%+ hit ratio; zero staleness by construction |
| Semi-dynamic content & API acceleration (catalog pages, public APIs, feeds) | Freshness vs origin load | Short TTL, SWR, surrogate-key purge, micro-caching (1–10s), shield, collapsing | Stale price/stock; purge storms; cache-key leakage | Freshness SLO (e.g., ≤ 60s) signed off by product; zero cross-user leakage |
| Large media & video (VOD segments, downloads) | Bandwidth cost, long tail, byte ranges | Segment caching, tiered cache, pre-positioning, ISP-embedded caches | Long-tail misses flood origin; cache pollution by one-hit wonders | Rebuffer ratio and startup time targets; origin egress budget |
🎯 Staff Move: "I'll design for semi-dynamic content — catalog pages and public API responses — because that's where freshness, cache keys and origin protection collide. Static assets are nearly solved by versioned URLs, and video is a bandwidth-economics problem I'd treat separately."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Freshness vs Hit Ratio | Short TTLs are fresh and miss often; long TTLs hit and go stale. Purge closes the gap — if it works |
| 2 | Cache Key Precision vs Fragmentation | Include more in the key (correct, fragmented) or less (efficient, risk of wrong/leaked content)? |
| 3 | Flat vs Tiered Caching | Every PoP goes to origin (lower miss latency) or through shields (origin protection, extra hop)? |
| 4 | Single CDN vs Multi-CDN | One vendor (simple, cheaper, correlated failure) or several (resilient, complex, config drift)? |
| 5 | Logic at the Edge vs at the Origin | Run personalization, auth and routing at the edge (fast) or keep it central (simple, consistent, debuggable)? |
In the Wild: Real Production Systems#
Netflix — Open Connect#
Netflix built its own CDN, Open Connect, and gives caching appliances (OCAs) to ISPs to install inside their networks at no charge, plus deploying them at internet exchange points. Content is pre-positioned during off-peak hours based on predicted popularity rather than pulled on demand, so the overwhelming majority of Netflix video traffic is served from an appliance close to — often inside — the viewer's ISP. The control plane (steering which OCA serves a client) runs in AWS; the bytes never touch it.
Staff insight: For large media with predictable demand, push beats pull. And it's a build-vs-buy decision driven by scale: at a double-digit share of peak internet traffic, owning the edge is cheaper and better than renting it. Almost nobody else is at that scale.
Fastly — Instant Purge and Surrogate Keys#
Fastly built its platform around fast, tag-based invalidation: responses carry a Surrogate-Key header with tags (e.g., product-123 category-9), and a purge by tag propagates globally in roughly 150ms by their public description. That makes it practical to cache semi-dynamic content — news articles, product pages, API responses — with long TTLs and purge on write, rather than short TTLs. On June 8, 2021, a customer's valid configuration change triggered a latent software bug and caused most of Fastly's network to return errors for under an hour, taking many major sites down simultaneously.
Staff insight: Instant purge changes the freshness tradeoff — "cache long, purge on change" becomes viable. The 2021 outage is the canonical example that a CDN is a correlated dependency across a large share of the web.
Cloudflare — Anycast, Tiered Cache, and a Global Regex#
Cloudflare routes users via anycast to 300+ cities, offers Tiered Cache (lower-tier PoPs fetch from upper-tier PoPs before origin), and runs WAF and code (Workers, on V8 isolates with very low startup cost) at every PoP. On July 2, 2019, a new WAF managed rule containing a regular expression with catastrophic backtracking was deployed globally at once and drove CPU to 100% across the network for about 27 minutes; the public post-mortem led to staged rollouts for WAF rules.
Staff insight: Every piece of logic you run at the edge runs everywhere at once. Edge config and code need staged, canaried rollouts even more than origin code does, because the blast radius is your entire internet presence.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Cache it at the CDN" | "What's the cache key? What if the page varies by language or login state?" | Cache-key judgment and leakage awareness |
| "TTL of 5 minutes" | "Price changed. The business wants it visible in 10 seconds." | Freshness mechanisms: purge, tags, SWR |
| "Purge on change" | "A bulk import changes 2M products. What happens to origin?" | Purge storms and origin protection |
| "The CDN absorbs load" | "The cache is empty after a config change. How much traffic hits origin?" | Cold-cache sizing, shield, collapsing |
| "CDNs are highly available" | "Your CDN is returning 503s globally. What now?" | Vendor failure plan |
| "Run it at the edge" | "How do you debug and roll back logic running in 300 cities?" | Edge compute operational maturity |
System Architecture Overview#
Reading the diagram: Requests reach the nearest PoP via anycast/DNS. The PoP terminates TLS (saving ~2 round trips to a distant origin), normalizes the cache key, and serves hits. Misses go to a regional shield, which collapses concurrent misses so origin sees roughly one fetch per object per shield. The origin tells the edge what to do through headers —
Cache-Controlfor freshness and privacy,Surrogate-Keyfor purge targeting — so cache policy lives with the team that owns the content. A publish event triggers a tag purge. The origin only accepts traffic from the CDN, so attackers can't bypass the edge. The DNS layer is the escape hatch if the CDN itself fails.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Static assets | "Cache with a long TTL, purge on deploy" | "Content-hashed URLs, 1-year immutable. Never purge; new content is a new URL." |
| Pages / APIs | "5-minute TTL" | "Short TTL (30–60s) + stale-while-revalidate + tag purge on write; freshness SLO agreed with product." |
| Cache key | "The URL" | "Allowlist: path, normalized params, one or two normalized headers. Cookies stripped. Personalized = private, no-store." |
| Origin protection | "CDN reduces load" | "Shield tier + request collapsing + purge rate limits + stale-if-error. Origin sized for the cold-cache case." |
| CDN outage | "CDNs don't go down" | "They do, globally. DNS failover to a second CDN or origin; low TTL on the edge hostname." |
| Edge logic | "Move it to the edge for speed" | "Only stateless, latency-critical logic; staged rollout by PoP group; origin remains the source of truth." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Hit ratio 95% vs 99% | Origin sees 5% vs 1% — a 5× difference | Last few points of hit ratio are worth the most |
| Speed of light in fiber | ~200 km/ms → ~1ms RTT per 100 km | Physics sets the floor; only proximity beats it |
| Typical RTTs | Intra-region 1–5ms; US coast-to-coast ~60–70ms; NY–London ~70–80ms; US–Australia ~150–200ms | New-connection cost = RTT × handshakes |
| New HTTPS connection | TCP 1 RTT + TLS 1.3 1 RTT (TLS 1.2: 2 RTT) before the first request | Edge TLS termination saves 2–3 long RTTs per new connection |
| Fastly tag purge (public figure) | ~150ms global | Makes "cache long, purge on change" viable |
| CloudFront invalidation pricing | First 1,000 paths/month free, then ~$0.005/path | Purge-heavy designs cost money on some CDNs |
| Versioned asset TTL | max-age=31536000, immutable (1 year) | Zero revalidation for static assets |
| Micro-cache TTL | 1–10s | At 10K RPS on one URL, a 1s TTL cuts origin to ~1 RPS per PoP/shield |
| Shield tiers | 3–6 regions typical | 300 PoPs → a handful of origin fetches per object |
| One-hit wonders | A large fraction of objects requested only once (Akamai researchers reported roughly three-quarters in one study) | Admission filters protect cache space for popular content |
| Video segment length | 2–6s; HD ~5 Mbps | Segments are the cache unit for streaming |
| CDN egress list price (order of magnitude) | ~$0.02–0.09/GB, much lower at committed PB volumes | Egress is the dominant CDN cost line |
Interview Walkthrough
Common prompts: "Design a CDN", "Design the edge layer for a global e-commerce site", "Our origin can't handle launch traffic — design caching." If the prompt is literally "design a CDN" (i.e., you are the CDN provider), spend proportionally more time on PoP architecture, routing, and purge propagation; if it's "use a CDN for our product", spend it on cacheability, keys, freshness and failure. Ask which.
Phase 1: Requirements & Framing (2–3 min)#
"Two framings are possible — building a CDN as a provider, or designing our product's edge on top of one. I'll assume the second unless you'd prefer the first. Then: what content are we serving — static assets, HTML pages, API responses, video? How fresh must each be — can a price be 60 seconds stale? Is any of it personalized? And where are users — global or one region? I'll assume a global e-commerce site: 50M daily users, 70% outside the origin's region, static assets plus catalog pages plus a public product API, with prices that must update within 60 seconds and nothing personalized in the shared cache."
Established:
- Content classes drive strategy
- Freshness SLO is a product decision you'll get signed off
- Personalization boundary is explicit
- Geography justifies the edge
Phase 2: Core Entities & API (1–2 min)#
| Entity | Fields That Matter | Owner |
|---|---|---|
| Cached object | cache key, body, headers, stored_at, TTL, SWR/SIE windows, surrogate keys, size | Edge |
| Cache policy (per route) | cacheable?, key allowlist (params, headers), TTL, SWR, SIE, private? | Product team (via response headers) + platform (defaults, guards) |
| Purge request | tags or URLs, soft/hard, requester, reason | Publishing systems; rate-limited by platform |
| Edge config | routes, origins, shield mapping, WAF rules, edge functions, version | Platform (config as code) |
The contract between origin and edge is HTTP headers:
Static asset: Cache-Control: public, max-age=31536000, immutable
Catalog page: Cache-Control: public, s-maxage=60, stale-while-revalidate=300, stale-if-error=86400
Surrogate-Key: product-123 category-9 page-template-v4
Vary: Accept-Encoding
Account page: Cache-Control: private, no-store
Public API: Cache-Control: public, s-maxage=10, stale-while-revalidate=30
Purge API: POST /purge { "tags": ["product-123"], "soft": true }
🎯 Staff Move: "The origin owns cache policy through headers, so the team that knows the content decides its freshness. The platform owns guardrails — like refusing to cache any response that sets a cookie, regardless of what the headers say."
Phase 3: High-Level Architecture (≤5 min)#
Draw the overview: users → DNS/anycast → PoPs → shields → origin, with a purge path and a config path. Name owners:
- Routing (DNS/anycast, failover) — platform/networking
- PoPs (TLS, WAF, cache, collapsing) — CDN vendor + platform config
- Shield tier — CDN feature, platform-configured
- Origin (emits headers, accepts only CDN traffic) — product teams + platform
- Purge pipeline (publish events → tag purge) — platform, fed by product systems
Phase 4: Transition to Depth#
"The architecture is standard. Where designs actually fail is in three places: the cache key — which determines both hit ratio and whether we ever leak content — freshness and purge at scale, and what happens when the cache is cold or the CDN itself fails. I'll go in that order."
Phase 5: Deep Dives (25–30 min)#
Deep dive 1 — Cache key (7 min).
Incoming: GET /p/red-shoes?utm_source=email&color=red&size=9&fbclid=abc
Cookie: session=...; ab_test=B Accept-Language: de-DE,de;q=0.9
Normalize: path lowercase, strip trailing slash
query: keep allowlist {color, size}, sorted → color=red&size=9
headers: lang = first supported of Accept-Language → 'de'
cookies: none in key; stripped from origin request for cacheable routes
Key: /p/red-shoes|color=red&size=9|lang=de|enc=br
Variants: ~8 languages × 2 encodings = 16 per URL (not thousands)
Deep dive 2 — Freshness (7 min). Prices must be visible within 60s. Options: TTL 60s (simple, origin load = 1 fetch per object per shield per minute), or TTL 1 day + tag purge on change (better hit ratio, depends on purge reliability). Staff default: TTL 60s + SWR 300s as the floor guarantee, plus tag purge for faster updates — purge makes it faster, TTL makes it bounded even if purge fails.
Deep dive 3 — Origin protection (6 min).
Worst case: global purge of homepage at 50K RPS
No shield, no collapsing: ~300 PoPs × concurrent misses → thousands of origin requests in 1s
Shield (6) + collapsing: ~6 origin requests
Soft purge (mark stale): 0 blocking origin requests — served stale while 6 revalidate
Bulk purge (2M products): rate-limit purges to ~5K tags/s; prefer soft purge; stagger
Deep dive 4 — Failure (5 min). Origin down → stale-if-error for 24h keeps catalog up. CDN down → DNS failover (edge hostname TTL 60s) to CDN B or to origin direct with origin capacity for ~20–30% of traffic plus aggressive shedding.
Deep dive 5 — Security (3 min). Never cache Set-Cookie responses; private, no-store on personalized routes; origin accepts only CDN traffic (IP allowlist + secret header or mTLS); unkeyed-header poisoning defenses.
Phase 6: Wrap-Up (2–3 min)#
"To summarize: versioned immutable assets; semi-dynamic pages with short TTL, stale-while-revalidate and tag purge; an allowlist cache key with personalization excluded by default; a shield tier and collapsing so the origin sees a handful of requests per object; stale-if-error for origin outages; and DNS failover for CDN outages. Next I'd add hit-ratio and freshness SLOs per route, an automated leakage test in CI, and evaluate multi-CDN once revenue per minute justifies it. Biggest risk: a cache-key or header mistake that caches private content — that's where I'd put the automated guardrails first."
Common Timing Mistakes#
| Mistake | Time Lost | What to Do Instead |
|---|---|---|
| Explaining anycast/BGP in depth | 8 min | One sentence unless you're building the CDN |
| Listing CDN vendors and features | 5 min | Vendor comparison belongs in the technology guide; say what you need |
| Only discussing static assets | — | They're solved by versioning; spend time on semi-dynamic |
| Never mentioning the cache key | — | It's the most important design decision |
| Ignoring CDN failure | — | Interviewers will ask; have the DNS failover answer ready |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Edge caching is deceptively cross-cutting. The CDN configuration is owned by one team; the headers that control it are emitted by dozens; the content's freshness requirements come from product; the costs land in finance; and the failures — stale prices, leaked pages, global outages — land on everyone. A Senior engineer can configure a CDN correctly for one application. A Staff engineer designs the contract that lets fifty teams cache safely without each one learning the hard way.
1.2 The L5 vs L6 vs L7 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"If this response were served to a different user, or two minutes late, who would be harmed — and how badly?"
Ask it per route. The answers sort every response into a strategy:
- Nobody, ever → immutable, cache forever with versioned URLs
- Nobody, if within N seconds → shared cache with TTL N, SWR, purge
- The user, if another user sees it →
private, no-store; never in the shared cache - The business, if late (prices, inventory, legal notices) → short TTL + purge + product-owned freshness SLO
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1: Static asset delivery. Solved problem if done right: content-hashed filenames, 1-year TTL, immutable, served from edge with Brotli/gzip and image format negotiation (AVIF/WebP). The remaining risks are deploy ordering (HTML references a new asset before it's uploaded, and the 404 gets cached) and cache pollution. Hit ratio should be 98–99%+.
Intent 2: Semi-dynamic content and API acceleration. The hard case. Catalog pages, search result pages for anonymous users, public API responses, article pages. They change — sometimes often — and the business cares how fast changes appear. The toolkit: short TTLs, SWR, micro-caching for extremely hot URLs (1–10s TTLs flatten spikes), tag purge on change, request collapsing, and a strict personalization boundary. Even uncacheable dynamic requests benefit from the edge: TLS termination near the user and warm, reused connections from edge to origin save 100–300ms for distant users.
Intent 3: Large media and video. Different economics. Objects are large (MB–GB), the catalog has a long tail, and egress cost dominates. Segment-level caching (2–6s chunks), tiered caches with large disk tiers, byte-range support, admission policies so one-time views don't evict popular content, and — at very large scale — pre-positioning and ISP-embedded caches (Netflix Open Connect). Metrics are rebuffer ratio, startup time and origin egress.
🎯 Staff Move: "These three share a CDN but not a strategy. Static assets want versioning, semi-dynamic wants purge and freshness contracts, video wants bandwidth economics. I'll focus on semi-dynamic, and I'll give static a one-line answer so we don't spend time on a solved problem."
2.2 When NOT to Use (Shared) Edge Caching#
| Situation | Why It's Wrong | What to Do Instead |
|---|---|---|
| Personalized responses (cart, account, feed) | Shared cache + personalization = leakage risk; hit ratio near zero anyway | private, no-store; still use edge for TLS + connection reuse; cache fragments via edge-side includes only with care |
| Strong consistency required (balances, inventory decrements, auth decisions) | Any TTL is a correctness bug | Serve from origin; cache only non-authoritative displays |
| Write-heavy APIs | Nothing to cache; edge adds a hop | Direct to regional API endpoints (API Gateway) |
| Users all in one region near origin | Edge adds cost with little latency benefit | Regional LB + in-region cache (Distributed Caching) |
| Regulated data with residency requirements | Edge caches may store data in other jurisdictions | Region-restricted edges or no caching; legal sign-off |
| Very low traffic long tail | Objects expire before a second request; pure miss overhead | Serve from object storage directly or cache only at shield |
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It Matters | What to Say |
|---|---|---|
| Freshness requirement per content type | Drives TTL vs purge | "Prices ≤ 60s, descriptions ≤ 1 hour, legal notices immediate — product signs off" |
| Personalization | Leakage risk | "Nothing personalized enters the shared cache" |
| Geography | Latency value of edge | "70% of users outside origin region" |
| Origin capacity | Cold-cache survival | "Origin sized for ~10% of peak without CDN" |
| Build vs buy | Provider vs customer design | "Using a commercial CDN; I'll note where building differs" |
| Budget | Egress dominates | "Egress is the main cost; hit ratio and compression are the levers" |
2.4 Precise Terminology#
| Term | Precise Meaning | Common Confusion |
|---|---|---|
| PoP | A CDN location with caching servers | Assumed to be one server |
| Origin | The authoritative source the CDN fetches from | Confused with shield |
| Shield / tiered cache | Intermediate cache layer between PoPs and origin | Thought of as a separate CDN |
| Hit ratio (request) | % of requests served from cache | Confused with byte hit ratio |
| Byte hit ratio / offload | % of bytes served from cache | Differs sharply for mixed small/large objects |
| Cache key | The identity of a cached object | Assumed to be the URL |
| Vary | Response header telling caches which request headers change the response | Overused (Vary: Cookie, Vary: User-Agent) → fragmentation |
| Purge (hard) | Remove object; next request is a miss | Causes origin stampede |
| Soft purge | Mark object stale; serve stale while revalidating | Unknown to many candidates |
| SWR / SIE | stale-while-revalidate / stale-if-error windows | Confused with TTL |
| Request collapsing | Coalesce concurrent misses into one origin fetch | Confused with caching itself |
| Cache poisoning | Attacker causes a harmful response to be cached under a key others request | Assumed impossible behind a CDN |
| Cache deception | Attacker tricks the cache into storing a victim's private page under a cacheable-looking URL | Unknown to most candidates |
3. The Five Fault Lines#
3.1 Fault Line 1: Freshness vs Hit Ratio#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Long TTL, no purge | Highest hit ratio | Stale content for up to TTL | Business (wrong prices), users |
| Short TTL (30–60s) | Bounded staleness; no purge dependency | Lower hit ratio; origin load = objects × shields / TTL | Origin (capacity), finance (egress from origin) |
| Long TTL + purge on change | High hit ratio and fast updates | Depends on purge reliability and publishers remembering to purge; purge storms | Platform (purge pipeline), origin (stampedes) |
| Short TTL + SWR + tag purge (Staff default) | Bounded staleness even if purge fails; purge makes it faster; SWR hides origin latency | More moving parts | Platform |
| Micro-caching (1–10s) for hot dynamic URLs | Flattens spikes; origin load ≈ 1 request/s per shield per URL | Up to 10s staleness | Product signs off on seconds |
Origin load for one hot page, 6 shields, TTL 60s: 6 requests/min
Same page, TTL 1s (micro-cache): 6 requests/s
Same page uncached at 50K RPS: 50,000 requests/s
Staleness bound with TTL 60s + SWR 300s: ≤ 60s fresh; up to 360s if origin slow
🎯 Staff Move: "Purge makes updates fast; TTL makes them bounded. I want both, because a purge pipeline that silently drops messages will happen, and I'd rather be 60 seconds stale than 3 days stale."
When to deviate: News and live events: purge-first with instant purge and long TTLs, because a headline update matters in seconds and traffic is enormous. Legal/compliance content: purge + verification (fetch from several PoPs after purge and assert new content).
3.2 Fault Line 2: Cache Key Precision vs Fragmentation#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Full URL + all headers via Vary | Never wrong | Hit ratio collapses (Vary: User-Agent → thousands of variants) | Origin, finance |
| URL only | Simple; high hit ratio | Wrong content when response varies by language/device/cookie; possible leakage | Users; security |
| Allowlist key (path + selected params + normalized headers) | Correct and efficient | Requires per-route definition and discipline | Platform + product teams |
| Edge-computed variant (device class, country → small enum) | Controlled variants | Logic at edge must match origin's | Platform |
The asymmetric risk: an under-specified key serves wrong content; an over-specified key serves slow content. Wrong is worse — especially when "wrong" means another user's data.
Leakage guardrails (non-negotiable):
- Never cache responses containing
Set-Cookie - Default all routes to non-cacheable unless the route is explicitly cacheable
- Strip
CookieandAuthorizationfrom origin requests on cacheable routes, so the origin cannot personalize them - CI test: request each cacheable route as two different users; assert identical bodies
When to deviate: A/B tests on cacheable pages — include the experiment bucket (not the user ID) in the key, with a small number of buckets.
3.3 Fault Line 3: Flat vs Tiered Caching#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Flat (every PoP → origin) | Lowest miss latency | Origin sees N PoPs × misses; long-tail hit ratio poor (each PoP sees little traffic per object) | Origin |
| Single shield | Origin sees 1 fetch per object | Shield is a hotspot and a single point of failure; distant PoPs pay extra RTT | Shield capacity; users far from shield |
| Regional shields (3–6) | Balance of origin protection and miss latency | More complex; objects stored at more layers | Platform config |
| Shield + origin-side cache | Protects origin compute even on shield miss | Another layer to invalidate | Platform |
The Staff default: Regional shields with collapsing at both edge and shield. The shield also improves hit ratio for long-tail content, because it aggregates demand from many PoPs.
When to deviate: Latency-critical dynamic API calls that aren't cacheable — bypass the shield and go edge → origin directly over warm connections.
3.4 Fault Line 4: Single CDN vs Multi-CDN#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Single CDN | Simplest; best volume pricing; one config dialect | Provider outage = your outage | Business, during provider incidents |
| Primary + cold standby | Escape hatch via DNS | Standby cold cache → origin stampede on failover; config drift | Origin, platform |
| Active multi-CDN (traffic split by performance/cost) | Resilience; performance steering; pricing leverage | 2× config, purge must fan out to both, lowest-common-denominator features | Platform (≈2× effort), finance (split volume) |
| Build your own | Full control; cheapest at extreme scale | Enormous capex/opex; only at Netflix-like scale | Company |
The Staff default: Single CDN with a tested DNS failover path to origin (or to a warm secondary for revenue-critical hostnames), low TTL on the edge hostname, and config expressed in a vendor-neutral format where possible. Move to active multi-CDN when revenue per minute of outage justifies roughly doubling edge operational effort.
🎯 Staff Move: "A CDN is a single global dependency, even though it's physically distributed. The question isn't whether it'll fail, it's whether we can move traffic in minutes when it does — and whether our origin survives the cold cache."
3.5 Fault Line 5: Logic at the Edge vs at the Origin#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Dumb edge (cache + TLS only) | Simple, debuggable, origin is the source of truth | Personalization and auth need origin round trips | Distant users (latency) |
| Edge config logic (redirects, header rewrites, geo routing) | Low latency; no code deploys | Config sprawl; a bad rule deploys globally | Platform |
| Edge compute (isolates/functions) | Auth token validation, A/B assignment, personalization of cached shells at edge | Global blast radius; limited state; observability and debugging harder; vendor lock-in | Platform + product teams |
The Staff default: Keep the edge stateless and small: token validation (not issuance), bucket assignment, redirects, header normalization, image transforms. Anything that needs strong state stays at origin. All edge logic ships through staged rollout by PoP group with automatic rollback — Cloudflare's 2019 incident is the reason.
When to deviate: Latency-critical global products (ads, bidding) where 100ms matters — push more to the edge, with the operational investment that implies.
4. Failure Modes & Operational Reality#
The lifecycle of a cached object — most failures happen at a transition:
4.1 Purge Stampede#
Setup: homepage + 40K category pages tagged 'layout-v7'. 300 PoPs, no shield,
hard purge. Normal origin load 2K RPS (98% hit ratio at 100K RPS).
t=0: Design team ships a header change; publisher hard-purges tag 'layout-v7'
t=+0.2s: 40K objects evicted in every PoP
t=+0.5s: 100K RPS of requests → ~all miss; collapsing per PoP helps, but 300 PoPs ×
40K hot objects = up to 12M distinct origin fetches over the next minute
t=+5s: Origin at 60K RPS (30× normal); render latency 8s; timeouts
t=+10s: Edges receive 5xx; nothing stale to serve (hard purge) → users see errors
t=+4min: Caches refill slowly; origin recovers
Detection: origin_rps spike correlated with purge_events_total; edge_hit_ratio collapse; origin_5xx_rate.
Mitigation: Stop further purges; enable/extend stale serving at the edge if the vendor allows; shed at origin.
Prevention: Shield tier (origin fetches ÷ ~50); soft purge (mark stale, serve stale while revalidating) as the default; purge rate limits per publisher; broad tags (layout, template) require approval; staggered purges for bulk changes.
Owner: Platform owns the purge pipeline and its limits; the publishing team owns what they purge.
4.2 Cache Leakage — Private Content in the Shared Cache#
t=0: A refactor moves the account page behind a new route; the framework's
default Cache-Control is 'public, max-age=300' for GET routes
t=+1min: User A loads /account; response (with A's name, address, last orders) cached
t=+1–6min: Every user requesting /account in that PoP sees User A's page
t=+20min: Support tickets; security incident declared
Scale: Every PoP caches its own first visitor's page. Hundreds of users exposed.
This is not hypothetical: on December 25, 2015, a caching configuration change at Valve caused Steam store pages containing other users' account information to be served to some users for a period, per Valve's public statement.
Detection: private_marker_in_cached_response_total (origin injects a response header like X-Private: 1 on personalized responses; the edge alarms if it ever caches one); synthetic two-user checks every minute.
Mitigation: Purge everything under the route; disable caching for the route via emergency config; assess exposure from logs.
Prevention: Default non-cacheable (opt in per route); strip cookies/authorization on cacheable routes; never cache Set-Cookie; CI two-user test; edge rule that refuses to cache responses carrying the private marker.
Owner: Security + platform (guardrails); product team (route headers). Legal/privacy for disclosure decisions.
4.3 Cache Poisoning via Unkeyed Inputs#
An attacker sends GET / with X-Forwarded-Host: evil.example. The origin uses that header to build absolute URLs for script tags. The response — now loading scripts from the attacker's domain — is cached under the key for /, because X-Forwarded-Host isn't in the key. Every visitor to that PoP gets the poisoned page until TTL.
Detection: Content integrity monitors (fetch key pages from many PoPs, diff against expected); CSP violation reports spike.
Prevention: Edge strips all non-allowlisted request headers before forwarding to origin; origin never trusts forwarding headers for URL generation; everything the origin uses to vary the response must be in the key (or stripped).
Owner: Platform (header normalization); security (testing).
4.4 CDN Provider Outage#
t=0: Provider pushes a bad config/software update; 85% of PoPs return 503
t=+1min: Synthetic monitors fail from all regions; revenue drops to near zero
t=+3min: Incident declared; decision: fail over?
t=+5min: DNS change: edge hostname → secondary CDN (or origin LB); TTL 60s
t=+6–10min: Traffic shifts as resolvers honor TTL (some take longer)
t=+10min: Secondary CDN has a cold cache → origin sees 20–40% of total traffic
t=+15min: Origin sheds low-priority routes; catalog serves; checkout prioritized
Detection: External synthetic monitoring from multiple networks (not from the CDN itself); RUM error beacons.
Mitigation: Pre-approved failover runbook with a decision threshold (e.g., > 25% global error rate for > 3 min).
Prevention: Low TTL on edge hostnames; tested failover quarterly; origin capacity and shedding plan for cold-cache load; secondary CDN warmed with top objects if active multi-CDN.
Owner: Platform/SRE; incident commander makes the call per runbook.
4.5 Cached Errors and Deploy Ordering#
A deploy publishes HTML referencing app.9f3e.js before the asset is uploaded. The first requests get 404, which the CDN caches for its default negative TTL (sometimes minutes). The site is broken for everyone at those PoPs even after the asset appears.
Prevention: Upload assets before publishing HTML (and keep old assets for ≥ 1 week for clients with old HTML); negative-cache TTL ≤ 10s for asset paths; never cache 5xx beyond a few seconds.
Owner: Frontend platform (deploy ordering); edge platform (negative-caching defaults).
4.6 Hit-Ratio Erosion — The Slow Failure#
Over months, hit ratio drifts from 95% to 78%. Causes: new marketing query parameters, a Vary: User-Agent added by a framework upgrade, cookies from a new analytics vendor forwarded to origin and used in keys, an A/B test framework keyed on user ID. Origin cost and latency creep up; nobody notices until capacity planning.
Detection: hit_ratio per route with weekly trend alerts; top-N cache-key cardinality by route.
Prevention: Key allowlists; periodic cardinality reports; hit ratio as an SLO per route owner.
Owner: Route owners (their headers); platform (reporting).
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Purge stampede | origin_rps spike after purge_events | Origin, then all users | Soft purge, shield, purge rate limit | Platform + publisher |
| Private content cached | private_marker_in_cached_response_total > 0; two-user synthetic | Every user of affected route/PoPs; privacy incident | Purge route; disable caching | Security + platform + route owner |
| Cache poisoning | Integrity monitor diffs; CSP reports | All users at poisoned PoPs | Purge; strip header | Platform + security |
| CDN provider outage | External synthetics, RUM errors | Entire web presence | DNS failover; origin shedding | SRE / platform |
| Cached 404/5xx | edge_status_404_rate for new asset paths | Users at affected PoPs | Purge path; short negative TTL | Frontend platform |
| Hit-ratio erosion | Weekly hit_ratio trend per route | Origin cost and latency | Key allowlist fixes | Route owners |
| Origin outage | origin_5xx_rate; served_stale_total ↑ | Freshness (if SIE configured); availability otherwise | stale-if-error | Origin team |
| Bad edge config/code push | Error rate spike correlated with config version | Global | Automatic rollback; staged rollout | Platform |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | "Add a CDN" | Classifies content by freshness and personalization | Defines the edge platform's scope for the company |
| Cache key | URL | Allowlist + normalization + leakage guardrails | Key policy standard with security review |
| Freshness | TTLs | TTL + SWR + tag purge; freshness SLO with product | Freshness contracts and purge SLA as platform commitments |
| Origin protection | Implicit | Shield, collapsing, soft purge, SIE, cold-cache sizing | Prices origin headroom vs CDN spend |
| Failure | "CDNs are reliable" | DNS failover, cold-cache plan, staged edge config | Vendor risk, multi-CDN economics, exit strategy |
| Security | TLS at edge | Poisoning, deception, private-content guardrails, origin lockdown | Edge as a security control plane (WAF, bot) with its own governance |
| Cost | Not discussed | Egress-driven; hit ratio and compression as levers | Contract strategy, cost allocation per team |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Starts with content classes | "Static, semi-dynamic, private — three strategies." |
| Treats the key as a security boundary | "Allowlist the key; personalized routes are private, no-store by default." |
| Bounds freshness twice | "Purge makes it fast, TTL makes it bounded." |
| Protects origin from its own CDN | "Soft purge plus shields: 6 origin fetches, not 300." |
| Plans for vendor failure | "Edge hostname TTL 60s, tested failover, origin sized for 30% cold." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Cache everything for 1 hour" | Ignores freshness requirements and personalization |
| "Purge on every change" with hard purge and no shield | Designs a stampede |
Vary: Cookie or Vary: User-Agent as the personalization answer | Fragments the cache to near-zero hit ratio, or leaks |
| No failure plan for the CDN | Treats a global dependency as infallible |
| Moves business logic to the edge without rollout discipline | Global blast radius |
5.4 Common False Positives#
- Knowing anycast and BGP details ≠ edge design judgment. Useful if you're building a CDN; rarely the question.
- Vendor feature fluency ≠ design. Knowing every CloudFront behavior setting doesn't answer who owns freshness.
- High hit ratio ≠ good design. A 99% hit ratio achieved by caching personalized pages is a data breach.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Provider vs customer; content classes; freshness; personalization |
| Header contract & entities | 3–6 min | Cache-Control, Surrogate-Key, purge API |
| Architecture | 6–11 min | Routing, PoPs, shields, origin, purge, config |
| Cache key | 11–18 min | Allowlist, normalization, leakage guardrails |
| Freshness & purge | 18–25 min | TTL + SWR + tags; purge at scale |
| Origin protection & failure | 25–34 min | Shield, collapsing, SIE, CDN outage |
| Edge compute / security / cost | 34–42 min | Whatever the interviewer pushes |
| Wrap-up | 42–45 min | Next steps, biggest risk |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "You're the CDN provider — design it" | Systems depth | PoP architecture (L4 LB → cache servers with consistent hashing), anycast, purge fan-out, config distribution |
| "Personalize the homepage" | Leakage risk management | Cache the shell, fetch personalized fragments client-side or via edge includes with per-user private fragments |
| "Serve video" | Media economics | Segments, tiered disk caches, admission policies, pre-positioning |
| "Price changes must appear in 1 second" | Freshness limits | Instant tag purge + TTL backstop; or don't cache price, cache everything else |
| "CDN costs doubled" | Cost levers | Hit ratio by route, compression, image formats, commit contracts, multi-CDN leverage |
| "Protect against DDoS" | Edge as security | Anycast absorption, WAF, rate limits at edge (Rate Limiting), origin lockdown |
6.3 What to Deliberately Skip#
- BGP and anycast internals (unless building a CDN)
- HTTP caching spec corner cases (
must-revalidatevsproxy-revalidate) - Vendor feature lists
- Image optimization details beyond "format negotiation and resizing at edge"
6.4 Follow-Up Questions to Expect#
- "How does request collapsing interact with a slow origin?" — Waiters queue behind one fetch; cap wait time and serve stale if available.
- "What's the difference between
max-ageands-maxage?" —s-maxageapplies to shared caches; lets the CDN cache longer than browsers. - "How do you cache an API that requires an auth token?" — Only if the response doesn't depend on the caller; validate the token at the edge, strip it from the key and origin request.
- "How do you know the purge worked?" — Purge acknowledgments plus sampled fetches from multiple PoPs asserting the new version header.
- "How do you pick the TTL?" — From the freshness SLO product signs off on; purge as the fast path.
- "Anycast or DNS-based routing?" — Anycast for simplicity and DDoS absorption; DNS for finer control; many CDNs use both.
- "How do you prevent the origin from being hit directly?" — IP allowlist of CDN ranges plus a secret header or mTLS from edge to origin.
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design the CDN strategy for our global e-commerce site."
Staff Answer
"I'll start by classifying content, because each class gets a different strategy. Static assets — JS, CSS, images — get content-hashed URLs and a one-year immutable TTL; that's essentially solved. Catalog pages and the public product API are semi-dynamic: short TTL, stale-while-revalidate, and tag-based purge when products change, with a freshness SLO product signs off on — say 60 seconds for price. Cart, account and checkout are personalized: private, no-store, never in the shared cache, but still routed through the edge for TLS termination and warm origin connections.
Then the three things I'd go deep on: the cache key, because it determines both hit ratio and whether we can ever leak someone's data; freshness and purge at scale, because purge storms are how CDNs hurt origins; and failure — origin outages and CDN outages."
Why this is L6:
- Classifies before configuring
- Gets product sign-off on freshness
- Names the cache key as a security boundary up front
What L7 adds:
- Frames the edge as a company platform with a contract all teams use, not a per-app config
- Mentions cost allocation and vendor strategy as later decisions
❌ Common L5 Trap
"Put CloudFront in front of everything with a 1-hour TTL and invalidate on deploy."
Why this misses: No distinction between static and dynamic, a 1-hour TTL on prices, and "everything" includes personalized pages — a leakage risk.
Drill 2: Cache Key Design#
Prompt: "Product pages vary by language and currency. Marketing adds UTM parameters to every link. How do you design the cache key?"
Staff Answer
"Allowlist. The key is the normalized path, the query parameters the page actually uses — say color and size, sorted — and two normalized dimensions: language (mapped from Accept-Language to one of our ~8 supported languages) and currency (derived from a country or explicit preference, mapped to ~10 currencies). UTM parameters, fbclid, gclid and every other unknown parameter are stripped from the key — and from the origin request, so the origin can't accidentally vary on them. Cookies are stripped on this route.
That bounds variants to about 8 × 10 = 80 per product URL, rather than one per marketing link. And I'd make the page emit a header echoing the language and currency it rendered, so a monitor can detect mismatches between key and content."
Why this is L6:
- Allowlist over denylist
- Normalizes headers into small enums
- Strips unkeyed inputs from origin requests so they can't vary the response
What L7 adds:
- Proposes a platform-owned key policy per route type, so 50 teams don't each reinvent normalization
- Tracks key cardinality per route as a cost metric
Drill 3: Make It Concrete — Freshness#
Prompt: "A price change must be visible globally within 10 seconds. How?"
Staff Answer
"Two mechanisms. The fast path: the pricing service publishes a 'product updated' event; a purge worker issues a soft purge for tag product-123, which on a CDN with fast tag purge propagates in roughly a second or less. The backstop: TTL on the page. If I need 10 seconds as a guarantee, not just typical behavior, TTL must be ≤ 10s — micro-caching — which at 6 shields is ~0.6 origin requests per second per product page, fine for popular products.
Alternatively, split the page: cache the product page for 5 minutes but fetch the price from a tiny, uncacheable or 1-second-cached price endpoint. That puts the freshness requirement only where it matters. And checkout always re-prices from the source of truth — the displayed price is never authoritative."
Why this is L6:
- Distinguishes typical freshness (purge) from guaranteed freshness (TTL)
- Offers splitting the freshness-sensitive fragment out
- Keeps the transaction authoritative
What L7 adds:
- Makes 'freshness SLO per content type' a published table owned by product, with the platform guaranteeing purge latency as an SLA
Drill 4: Origin Is Down#
Prompt: "Our origin is down for 20 minutes. What do users see?"
Staff Answer
"For cacheable content: the last good version, because every cacheable response carries stale-if-error=86400. The catalog stays browsable; prices might be up to 20 minutes stale, which product accepted for this scenario since checkout re-prices anyway. For uncacheable content — cart, checkout — errors, unless we have a static maintenance experience at the edge for those routes.
What makes this work: SIE configured on all cacheable routes, objects not evicted too aggressively, and the edge not converting origin 5xx into cached errors. I'd verify quarterly by blocking origin from the CDN for 10 minutes in a controlled game day."
Why this is L6:
- Uses SIE to turn an outage into staleness
- Names what still breaks and what users see
- Proves it with a game day
What L7 adds:
- Designs a static 'degraded shell' for transactional routes served from edge storage, with a product-approved message and status page link
Drill 5: Purge at Scale#
Prompt: "A nightly bulk import updates prices for 2M products. How do you invalidate?"
Staff Answer
"Not with 2M hard purges at once. Options in order of preference: If prices are split into a small price fragment with short TTL, the import needs no purge at all — the TTL catches it. If pages include prices, soft-purge by tag, rate-limited to a few thousand tags per second and ordered by popularity so hot products refresh first; the long tail refreshes on TTL. Soft purge means requests keep being served stale while shields revalidate, so the origin sees at most ~6 fetches per product, spread over the purge window.
And I'd ask whether the import should purge at all: if 95% of those products receive fewer than one view per hour, their cached copies will expire before anyone sees them."
Why this is L6:
- Rejects the naive bulk hard purge
- Uses popularity ordering, soft purge and rate limits
- Questions whether purge is needed
What L7 adds:
- Sets per-publisher purge quotas and a review process for broad tags
Drill 6: Hot Object#
Prompt: "A single video segment gets 2M requests per second during a live event. What breaks?"
Staff Answer
"The cache hit ratio is fine; the problem is per-server load. Inside a PoP, objects are typically sharded to cache servers by consistent hashing of the key, so one hot key lands on one server. At 2M RPS globally across 300 PoPs, the busiest PoPs may see 50K RPS for that object — beyond a single server's NIC for multi-MB segments.
Fixes: replicate hot objects across multiple servers in the PoP (hot-key detection triggers replication to N servers, and the PoP load balancer spreads requests); serve from memory rather than disk; for live, use short segments with predictable names so the edge can prefetch the next segment; and collapse misses for the next segment so origin sees one fetch per shield."
Why this is L6:
- Identifies intra-PoP hot-key concentration, not hit ratio, as the bottleneck
- Offers replication, memory tiering, and prefetch
What L7 adds:
- For scheduled mega-events, pre-provisions capacity with the CDN and tests a rehearsal event — the contract, not the architecture, is the lever
Drill 7: Build vs Buy#
Prompt: "Should we build our own CDN?"
Staff Answer
"Almost certainly not. A commercial CDN amortizes hundreds of PoPs, peering relationships, DDoS capacity and 24/7 operations across thousands of customers. Building means capex for servers in dozens of locations, peering negotiations, a network operations team — tens of engineers — and years to reach parity.
The exceptions are companies whose traffic is a meaningful share of the internet and highly predictable — Netflix built Open Connect because at their volume, owning the edge and placing boxes inside ISPs is cheaper and better. For us, the better levers are negotiating committed-use pricing, adding a second CDN for leverage and resilience once spend justifies it, and keeping our configuration portable. See the technology guide for how edge proxies compare."
Why this is L6:
- Names the specific conditions that justify building
- Offers realistic cost levers instead
What L7 adds:
- Models the crossover point: at what monthly egress does a hybrid (own caches in top 5 metros + commercial CDN elsewhere) break even?
Drill 8: Policy Change Without an Outage#
Prompt: "You need to change the cache key for all product pages to add a currency dimension. How do you roll it out?"
Staff Answer
"A key change effectively empties the cache for those pages, because every new key is a miss. So: stage it by PoP group — 5% of PoPs first — and watch origin load, hit ratio and error rate. Before starting, warm the shield: shields can use the new key while edges are still on the old one, so the shield fills gradually. Pick a low-traffic window. Make sure origin can absorb the miss rate for the canary PoPs — roughly the canary fraction × the per-object miss rate.
Validate correctness in canary: two synthetic users with different currencies request the same product; assert different currencies in the body and distinct cache entries. Rollback is a config revert, which is also a key change — so rollback has its own miss cost; plan origin capacity for both directions."
Why this is L6:
- Recognizes a key change as a cache flush
- Stages by PoP group with origin capacity checks
- Tests correctness and plans rollback cost
What L7 adds:
- Makes edge config changes flow through a pipeline with automated canary analysis and PoP-group staging for every team
Drill 9: Cost#
Prompt: "Our CDN bill is $400K/month and growing 8% a month. What do you do?"
Staff Answer
"First, attribute it: bytes by route, content type and team. Usually a few items dominate — unoptimized images, video, large JS bundles, or API responses that bypass cache. Then pull the levers by impact:
- Bytes: image format negotiation (AVIF/WebP typically 25–50% smaller than JPEG), resizing to device, Brotli for text. Often 20–40% of the bill.
- Hit ratio: fixes to cache keys raise offload and cut origin egress, which is often billed separately by the cloud provider.
- Contract: committed-use pricing at our volume; a second CDN as leverage.
- Architecture: long-tail media served from cheaper tiers.
Then set a unit metric — CDN $ per 1,000 sessions — so growth in the bill is measured against growth in the business."
Why this is L6:
- Attributes before optimizing
- Ranks levers by impact with numbers
- Introduces a unit-cost metric
What L7 adds:
- Allocates CDN cost to teams by route so owners see their own egress
- Negotiates multi-year contracts with volume tiers aligned to the growth forecast
Drill 10: Multi-Region Origin#
Prompt: "We're adding a second origin region. How does the CDN use it?"
Staff Answer
"Each shield gets a preferred origin — EU shield to EU origin, US shields to US origin — with failover to the other origin on health-check failure. That halves the long-haul miss latency and gives origin redundancy.
Care points: both origins must produce identical cacheable responses for the same key — same data version — or users flip between versions. Purges must hit both origins' caches if they have their own. And failover means one origin takes 2× miss traffic, so each origin must handle the full miss load, or shed. CDN-level origin failover must be faster than DNS — health checks every ~5–10s."
Why this is L6:
- Maps shields to origins
- Flags version consistency across origins and failover capacity
What L7 adds:
- Decides whether the second region is active-active for writes too, and aligns edge routing with data-residency requirements
8. Deep Dive Scenarios#
Deep Dive 1: Launch-Day Origin Meltdown Behind a CDN#
Context: A limited-edition product launches at 10:00. The product page is behind the CDN with s-maxage=60. At 10:00:05, origin CPU hits 100% and error rates reach 40%. The CDN dashboard shows a 35% hit ratio for the product page, far below the expected 99%. You're pulled in.
Questions to Surface First:
- Why is the hit ratio 35% on a single hot URL — is the cache key fragmented?
- Is request collapsing enabled, and is there a shield?
- Is the page actually cacheable — is the origin emitting
Set-Cookieorprivateunder load (e.g., a session started for the queue)? - Is the origin serving errors that are being cached, or bypassing cache?
Typical L5 Approach: Scales origin, raises the TTL. Scaling takes minutes; raising TTL doesn't help if requests are missing for another reason.
Staff Approach: Diagnoses fragmentation first. Finds the launch email links carried a unique
refparameter per recipient, which wasn't stripped — every click was a distinct cache key. Pushes an emergency edge rule stripping unknown query parameters for that route (staged to 10% of PoPs for 60s, then global), confirms hit ratio climbs above 95% within two minutes, and enables stale-if-error so remaining origin errors aren't shown.
Principal Approach: Treats "a marketing link format can take down the origin" as a platform defect. Makes query-parameter allowlists the default for every cacheable route, adds a pre-launch checklist that includes a synthetic load test through the CDN using the real campaign URLs, and establishes a joint marketing–engineering launch calendar.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check hit_ratio and top cache keys for the route. Enable SIE and micro-cache fallback. |
| Triage | Key cardinality spike → unstripped ref parameter. |
| Quick fix | Edge rule: allowlist params for /p/*; staged rollout in minutes. |
| Guardrails | Origin shedding on non-critical routes until hit ratio > 95%. |
| Post-mortem | Default param allowlist; launch load test through the CDN with real links. |
Metrics to Watch: hit_ratio{route}, cache_key_cardinality{route}, origin_rps, origin_cpu, edge_5xx_rate.
Organizational Follow-up: Launch checklist owned by the platform team with marketing sign-off on link formats.
Ownership Question: "Who owns the fact that marketing links bypass the cache?" Staff answer: The platform. Marketing will always add tracking parameters; the edge's job is to ignore what the page doesn't use. Making marketing responsible for cache keys is asking the wrong team to understand the wrong system.
Key Takeaway: "When a single hot URL has a low hit ratio, the cache key is fragmented until proven otherwise."
What clears the Staff bar:
- Looks at key cardinality before scaling
- Fixes at the edge in minutes with staging
- Makes the safe behavior the platform default
Deep Dive 2: The Silent Leak#
Context: A security researcher emails that loading /orders/recent.css on your site shows the page of whoever last requested it. Your team confirms: a new framework treats unknown extensions as the base route, so /orders/recent.css renders the user's order page — and the CDN caches anything ending in .css for 1 day. This is web cache deception.
Questions to Surface First:
- How long has this been possible, and which routes are affected?
- Do logs show exploitation (unusual extensions on personalized routes)?
- Why does the CDN cache by extension rather than by the origin's
Cache-Control? - Does the origin mark personalized responses as
private?
Typical L5 Approach: Adds a rule to not cache
/orders/*. Fixes one route; the pattern applies to every personalized route.
Staff Approach: Fixes the class: the edge must never override the origin's
Cache-Controlby extension; personalized responses always carryCache-Control: private, no-storeplus a private marker the edge refuses to cache; the framework must 404 unknown extensions on application routes. Purges all cached content on personalized paths; reviews logs for exploitation; engages security and privacy for disclosure assessment.
Principal Approach: Establishes that caching behavior is a security control. Edge caching rules go through security review; a quarterly external test for cache deception and poisoning is part of the security program; the "private by default" posture is written into the edge standard and verified in CI for every route.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Disable extension-based caching rules; purge personalized paths globally. |
| Triage | Log search for suspicious extensions on personalized routes; exposure estimate. |
| Quick fix | Origin: private, no-store + private marker on all authenticated responses. |
| Guardrails | Edge refuses to cache marked responses; CI two-user test; framework rejects unknown extensions. |
| Post-mortem | Edge rules overrode origin intent; no defense in depth. |
Metrics to Watch: private_marker_in_cached_response_total, cached_responses_on_auth_routes_total, suspicious-extension request rate.
Organizational Follow-up: Security review for edge rules; bug bounty scope includes caching.
Ownership Question: "Who owns this vulnerability — the framework team, the edge team, or the orders team?" Staff answer: All three had a gap, but the edge platform owns the fix that makes the class impossible: never cache a response the origin marked private, regardless of path. Defense in depth means no single team's mistake leaks data.
Key Takeaway: "The edge should never be smarter than the origin about what's private."
What clears the Staff bar:
- Fixes the class, not the route
- Defense in depth across framework, origin and edge
- Engages security and privacy immediately
Deep Dive 3: Onboarding a Large Customer Surface — The Public API#
Context: The product API team wants to put its public read API (1.2M RPS peak, 40% of which are identical catalog queries) behind the CDN. Responses depend on the API key's plan tier (some fields hidden for free tier). The API team is worried about stale data and about the CDN seeing API keys.
Questions to Surface First:
- Does the response vary by API key, or only by plan tier?
- What freshness does the API contract promise?
- Where are rate limits enforced — would caching bypass them?
- Is the API key a secret that must not appear in cache keys or logs?
Typical L5 Approach: Adds the
Authorizationheader to the cache key. Correct but fragments the cache by customer — hit ratio collapses to near zero for the long tail of keys.
Staff Approach: Validates the API key at the edge (edge function checks a signed token or looks up key → tier in a replicated edge KV), then keys the cache by tier (3 values), not by key. Rate limiting still runs per key at the edge before cache lookup, so cached responses count against quotas. Cache TTL matches the API's documented freshness (e.g., 30s) with
Ageheader exposed to clients. The raw key is stripped from the origin request and never logged at edge.
Principal Approach: Frames it as a product decision: caching changes the API's freshness contract, which should be documented and versioned. Aligns API gateway and CDN responsibilities so there's one place for authentication and rate limiting (API Gateway), and one for caching — avoiding two teams implementing auth.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Edge auth → tier; key = path + normalized params + tier. |
| Rate limiting | Per-key at edge, before cache lookup. |
| Freshness | TTL 30s + SWR 60s; Age exposed; documented in API reference. |
| Rollout | Shadow: edge computes cache decisions, compares responses to origin for 1 week; then 5% → 100%. |
| Success | Origin RPS −35%; p75 latency outside origin region −150ms; zero tier-leak incidents. |
Metrics to Watch: hit_ratio{route=api}, tier_mismatch_detected_total, rate_limit_rejections{cache_status}.
Organizational Follow-up: API docs updated with freshness semantics; gateway and edge teams agree on auth ownership.
Ownership Question: "If a free-tier customer receives a pro-tier field from cache, whose bug is it?" Staff answer: The edge platform's key design. Tier must be in the key, and a synthetic test with keys from each tier runs continuously. The API team owns the tier-to-fields mapping; the platform owns guaranteeing it's honored in cache.
Key Takeaway: "Key the cache by the smallest dimension the response actually depends on — the tier, not the customer."
What clears the Staff bar:
- Keys by tier, not identity
- Keeps rate limiting ahead of cache
- Documents freshness as part of the API contract
Deep Dive 4: Post-Mortem — The Provider Outage#
Context: Your sole CDN provider had a 45-minute global outage. Your site was fully down. Leadership asks for a post-mortem and a plan so "this never happens again."
Questions to Surface First:
- Did we have a failover plan? Was it executed, and if not, why?
- What's the DNS TTL on the edge hostname, and who can change DNS in an incident?
- Could the origin have served traffic directly? At what fraction?
- What did 45 minutes cost?
Typical L5 Approach: Proposes adding a second CDN.
Staff Approach: Finds the real gaps: edge hostname had a 1-hour DNS TTL (failover would have taken an hour anyway), no runbook with a decision threshold, and origin only locked to CDN IPs — so direct traffic was impossible without a firewall change. Fixes: TTL 60s, runbook with decision criteria and pre-approved authority, origin able to accept direct traffic via a pre-provisioned LB behind a separate hostname, and shedding rules to cope with the load. Evaluates a secondary CDN as a separate decision with costs.
Principal Approach: Answers "never again" honestly: no plan makes a dependency infallible; the question is how much to spend on reducing expected downtime. Quantifies: expected provider outage time per year × revenue per minute versus the cost of active multi-CDN (config duplication, split volume discounts, ~1–2 additional engineers). Brings leadership a decision with numbers rather than a reflexive "add another vendor."
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (future incidents) | Decision threshold: > 25% global error from external synthetics for 3 min → failover. |
| Triage | Why no failover: TTL, runbook, origin lockdown. |
| Quick fix | DNS TTL 60s; direct-origin hostname and LB pre-provisioned; runbook drill. |
| Guardrails | Quarterly failover drill at low traffic; origin shedding plan for 30% of traffic. |
| Post-mortem | The plan existed as an idea, not as tested infrastructure. |
Metrics to Watch: External synthetic success rate by provider; dns_ttl{hostname}; time-to-failover in drills.
Organizational Follow-up: Named failover owner and on-call authority; multi-CDN business case.
Ownership Question: "Who has the authority to move all traffic off the CDN?" Staff answer: The incident commander, per runbook criteria, without needing VP approval at 3 AM. The VP approves the criteria in advance.
Key Takeaway: "A failover plan that's never been executed is a hypothesis."
What clears the Staff bar:
- Finds that the existing plan was unexecutable (TTL, lockdown)
- Separates quick fixes from the multi-CDN business decision
- Pre-approves authority
Deep Dive 5: Multi-Region Expansion with Data Residency#
Context: The company is entering the EU and must keep EU users' personal data in the EU. The CDN serves global traffic from all PoPs. Legal asks: "Does the CDN store personal data outside the EU?"
Questions to Surface First:
- What does the CDN cache — is any of it personal data? (Should be none, if personalized =
no-store.) - What about logs — edge logs include IP addresses, which are personal data under GDPR?
- Do EU users ever get routed to non-EU PoPs (anycast doesn't respect borders precisely)?
- Do edge functions process personal data (e.g., tokens with user IDs)?
Typical L5 Approach: Adds an EU origin and assumes the CDN is fine because it's "just a cache."
Staff Approach: Audits all three data flows: cache content (verify no personal data is cached — the private-by-default guardrail is now a compliance control), logs (configure regional log delivery and IP truncation), and edge compute (tokens validated at edge contain only pseudonymous IDs). For personalized routes, uses the CDN's regional services feature (where available) to restrict TLS termination and processing to EU PoPs for EU hostnames.
Principal Approach: Treats the edge as part of the company's data-processing footprint — it needs a data-processing agreement, a documented data map, and a review whenever new edge features (logging, compute, KV) are adopted. Establishes that the edge platform team is accountable to the privacy program, not only to engineering.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Audit | Cache content, logs, edge compute — each checked for personal data. |
| Controls | Private-by-default verified; regional log pipelines; IP truncation. |
| Routing | EU hostnames with EU-only processing for personalized routes. |
| Validation | Sample cached objects; confirm no personal data. |
| Documentation | Data map and DPA updated; legal sign-off. |
Metrics to Watch: private_marker_in_cached_response_total, log delivery region, EU-hostname PoP distribution.
Organizational Follow-up: Privacy review for new edge features.
Ownership Question: "Who signs off that the edge is compliant?" Staff answer: Legal/privacy signs off, based on a data map the edge platform team maintains. Engineering can't self-certify compliance.
Key Takeaway: "'It's just a cache' is not a compliance position. Know what the edge stores, logs and processes."
What clears the Staff bar:
- Audits cache, logs and compute separately
- Reuses private-by-default as a compliance control
- Involves legal for sign-off
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Classify content into immutable, semi-dynamic, personalized and never-cache, and give each a strategy
- Design an allowlist cache key with normalization, and explain its hit-ratio and security consequences
- Combine TTL, stale-while-revalidate, stale-if-error and tag purge to bound freshness twice
- Quantify origin protection from shields, collapsing and soft purge
- Identify cache poisoning, cache deception and private-content leakage and prevent each class
- Plan for CDN provider failure with DNS TTLs, runbooks and origin capacity
- Reason about CDN cost as egress × (1 − offload) plus contract strategy
- Decide what logic belongs at the edge and how to roll it out safely
The Bar for This Question#
Mid-level (L4): Knows CDNs cache static content near users; sets TTLs; mentions invalidation.
Senior (L5): Configures a CDN competently: versioned assets, TTLs, purge on deploy, TLS at edge. Misses cache-key design as a security boundary, purge stampedes, stale-if-error, vendor failure and cost.
Staff+ (L6): Classifies content, designs allowlist keys with leakage guardrails, bounds freshness with TTL plus purge, protects origin with shields and soft purge, plans for origin and CDN failure, treats edge config as code with staged rollout, and names owners for freshness and headers. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "The Cache Key Is a Security Boundary, Not a Performance Setting"#
| Key Too Broad | Key Too Narrow |
|---|---|
| Lower hit ratio, higher cost | Wrong content served — possibly another user's |
| Recoverable, measurable | Privacy incident, possibly reportable |
The Staff position: Design the key with security review, default routes to non-cacheable, and test with two users in CI.
Why this matters in interviews: Candidates discuss keys only for hit ratio. Saying "leakage" first is a strong signal.
10.2 "Purge Is an Optimization; TTL Is the Guarantee"#
Purge pipelines fail silently — dropped events, rate limits, publishers who forget. Designs that rely on purge for correctness with multi-day TTLs are one lost message away from serving a 3-day-old price.
The Staff position: TTL bounds staleness; purge accelerates freshness. The exception is providers with reliable instant purge and purge verification — then "cache long, purge on change" is sound.
Why this matters in interviews: It demonstrates you think about failure of the freshness mechanism itself.
10.3 "Most Teams Should Cache Less at the Edge, Not More"#
Aggressive edge caching of semi-dynamic and personalized-adjacent content produces the most dangerous edge incidents (leaks, stale prices, poisoned pages), while the latency benefit for many of those routes is achieved by TLS termination and connection reuse alone.
The Staff position: Cache static aggressively; cache semi-dynamic deliberately with contracts; for dynamic routes, use the edge for TLS and connection reuse without caching.
Why this matters in interviews: It shows you distinguish the edge's two value propositions: caching and proximity.
10.4 "Multi-CDN Is Usually a Negotiating Tactic, Not a Reliability Strategy"#
| Multi-CDN Benefit | Multi-CDN Cost |
|---|---|
| Survives a single provider outage | ~2× config, testing and purge fan-out; features limited to common denominator |
| Pricing leverage | Split volume reduces tier discounts |
| Performance steering | Requires RUM-based steering infrastructure |
The Staff position: A tested DNS failover path gets most of the reliability benefit for a fraction of the cost. Active multi-CDN is justified at high revenue per minute or when the pricing leverage pays for the effort.
Why this matters in interviews: Reflexively proposing multi-CDN is an L5 signal; pricing it is L6/L7.
10.5 "Edge Compute Is a Distributed System You Deploy to 300 Cities at Once"#
Running code at the edge moves you from one deploy target to hundreds, with limited observability, constrained state and a global blast radius. The latency win is real — tens to hundreds of milliseconds for distant users — but only for logic that is stateless and stable.
The Staff position: Token validation, routing, A/B bucketing, header normalization: yes. Business logic with state: no, until you have staged rollout, per-PoP observability and fast rollback.
Why this matters in interviews: It shows you evaluate edge compute as an operational commitment, not a feature.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
At L6, the edge is a well-configured caching layer with safe keys and bounded freshness. At L7, the edge is the company's front door: the first code every customer request touches, a security control plane (TLS, WAF, bot defense, DDoS), a data processor subject to privacy law, a six- or seven-figure monthly line item, and a single vendor relationship whose outage is a total outage. The Principal decisions are about scope (what the edge owns), governance (who can change a rule that runs in 300 cities), money (egress contracts, cost allocation, build-vs-buy crossover), and vendor risk (lock-in via proprietary edge compute and config dialects).
The Org-Level Fault Line#
A central edge platform vs teams configuring the CDN directly.
| Option | What It Buys | What It Costs | Who Pays |
|---|---|---|---|
| Teams configure the CDN console directly | Speed; no bottleneck | Inconsistent keys, leak risk, no staged rollout, surprise costs | Security (incidents), finance (bills), everyone (outages from a rule change) |
| Central edge team owns all config | Consistency and safety | Every change queues behind one team; product velocity suffers | Product teams (waiting) |
| Edge platform with self-service contract (the L7 default) | Teams control cache policy via response headers and a reviewed route manifest; platform owns guardrails, pipeline, and vendor | Platform must build the pipeline, guardrails and tests | Platform headcount |
🧭 Principal Move: "Teams own their headers; the platform owns the guardrails and the blast radius. No one — including the platform team — pushes a rule to all PoPs at once."
Cost Model#
Assumptions: blended CDN egress ~$0.01–0.05/GB depending on commitment and geography; origin/cloud egress ~$0.05–0.09/GB; fully loaded engineer ~$300K/year; figures are order-of-magnitude for planning, not quotes.
| Scale | Traffic | Monthly CDN | Origin Egress Saved | Headcount | On-call |
|---|---|---|---|---|---|
| Startup | ~50 TB/month | ~$2–5K (often pay-as-you-go) | ~$2–4K | ~0.2 FTE | Shared; CDN outages are rare pages |
| Growth | ~2 PB/month | ~$40–100K with commit | ~$100–150K | 2–3 FTE edge/platform | Platform rotation; purge and config incidents dominate |
| Large | ~50 PB/month, video-heavy | ~$0.5–1.5M with negotiated tiers | Millions | 8–15 FTE (edge, security, contracts) | Dedicated edge rotation; multi-CDN steering |
The strategic numbers: every point of byte offload above 95% saves roughly 1% of total bytes from origin egress, and image/format optimization often cuts total bytes 20–40%. Negotiated commits typically matter more than any single engineering optimization at large scale.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse | Why |
|---|---|---|---|
| TTLs, purge rules, key allowlists | Two-way | Minutes (with miss cost on key changes) | Tune continuously |
| Asset URL versioning scheme | Mostly one-way | Every build pipeline and HTML template | Get it right early |
| Proprietary edge compute (vendor-specific runtime/KV) | Expensive two-way | Rewrite of edge logic per vendor | Lock-in; keep logic small and portable |
| Multi-year CDN commit | One-way for the term | Contract penalties | Align with growth forecast |
| Building own PoPs / ISP caches | One-way | Capex, peering, team | Only at extreme, predictable volume |
| Public hostnames and URL structure | One-way | SEO, links, clients | URLs are forever |
The Standard I'd Write#
RFC: Edge Caching & Delivery Standard v1
Scope: All public hostnames served through the company edge.
Requirements:
- Routes MUST be non-cacheable in the shared cache unless declared cacheable in the route manifest.
- Responses that depend on user identity MUST send
Cache-Control: private, no-storeand the private marker header; the edge MUST refuse to cache marked responses.- Cacheable routes MUST declare a cache-key allowlist; unknown query parameters, cookies and
AuthorizationMUST be stripped before cache lookup and origin fetch.- Static assets MUST use content-hashed URLs with
max-age=31536000, immutable; they MUST NOT be purged.- Cacheable routes SHOULD set
stale-if-error≥ 1 hour andstale-while-revalidate≥ TTL.- Purges SHOULD be soft by default; broad tags require platform approval; per-publisher purge rate limits apply.
- Edge config and code MUST roll out by PoP group with automated health gates; no global instantaneous deploys.
- Edge hostnames MUST have DNS TTL ≤ 60s and a tested failover target.
Exceptions: Filed with the edge platform; security review required for any exception to private-by-default.
Success metrics: Zero private-content cache incidents; byte offload ≥ 95% for static, ≥ 80% for semi-dynamic routes; purge p99 ≤ 5s; failover drill completed quarterly with time-to-shift ≤ 10 min; CDN $ per 1,000 sessions trending flat or down.
What I'd Tell the VP#
"Almost every customer request enters through our CDN, so it's both our biggest performance lever and a single point of failure. We're putting three things in place: rules that make it impossible for a team to accidentally cache one customer's private page for another, a way to update prices and content globally within seconds without overloading our servers, and a tested plan to move traffic elsewhere if the CDN provider has an outage like the ones that took down large parts of the web in 2019 and 2021. We'll also charge each team for the bandwidth their pages use, which historically cuts the bill 20–30% once people can see it. A second CDN is on the table — I'll bring you the cost versus the expected downtime it avoids so we can decide with numbers."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Defines edge scope | "The edge owns TLS, caching, WAF and stateless routing — never business state." |
| Governs blast radius | "Every edge change rolls out by PoP group; nobody pushes globally, including us." |
| Prices vendor risk | "Multi-CDN costs ~2 engineers and our volume discount; one provider outage a year costs X. Here's the break-even." |
| Manages lock-in | "Edge logic stays small and portable; proprietary KV only for caches we can rebuild." |
| Treats edge as a data processor | "Logs contain IPs — the edge is in our privacy data map." |
Staff answers that L7 interviewers find insufficient:
- A perfectly configured CDN with no governance for the 30 teams who will change it next year.
- "Add a second CDN" without cost, feature-parity and purge fan-out analysis.
- No view of the contract — commit levels and pricing tiers are often the biggest cost lever.
Appendices
Appendix A: Mechanics in Depth#
A.1 Edge Request Handling#
handle(req):
key = normalize(req) # allowlist params, normalized headers
if route_not_cacheable(req): return proxy_to_origin(strip(req))
obj = cache.get(key)
if obj and obj.fresh(): return obj # HIT
if obj and obj.within_swr(): async revalidate(key); return obj # STALE, revalidating
resp = collapse(key, lambda: fetch_from_shield(strip(req))) # one fetch per key in flight
if resp.is_error() and obj and obj.within_sie(): return obj # serve stale on error
if cacheable(resp) and not resp.has('Set-Cookie') and not resp.has('X-Private'):
cache.put(key, resp, ttl=resp.s_maxage)
return resp
A.2 Why Versioned URLs Beat Purge for Assets#
| Approach | Freshness | Hit Ratio | Failure Mode |
|---|---|---|---|
| Fixed URL + purge on deploy | Depends on purge + browser caches (which you can't purge) | Dips after each deploy | Browsers keep old asset for its max-age |
| Content-hashed URL | New URL = instantly fresh everywhere, including browsers | ~100% | Old HTML referencing removed assets → keep old assets ≥ 1 week |
A.3 Admission and Eviction#
- LRU is the common baseline; video and large-object caches use segmented LRU or size-aware policies.
- Admission filters (e.g., only cache on second request within a window, tracked with a Bloom filter) keep one-hit wonders from evicting popular content — valuable because a large share of objects are requested once.
- Tiered storage: RAM for the hottest objects, SSD for the warm set, larger disk tiers at shields.
A.4 PoP Internals (If You're Building the CDN)#
Front servers terminate TLS and route by consistent hash of the cache key to the cache server that owns it, so each object is stored once per PoP. Hot objects are replicated to multiple cache servers to avoid single-server saturation. See Consistent Hashing.
Appendix B: Cache Keys and Variants#
B.1 Key Construction#
| Component | Rule | Example |
|---|---|---|
| Host | Lowercase; canonical host only | www.example.com |
| Path | Lowercase if case-insensitive routes; collapse slashes | /p/red-shoes |
| Query | Allowlist, sorted, deduplicated | color=red&size=9 |
| Headers | Normalized to small enums | lang=de, dev=mobile |
| Encoding | Negotiated separately (br, gzip, identity) | enc=br |
| Cookies | Excluded; stripped for cacheable routes | — |
B.2 Vary Discipline#
| Vary Value | Variants | Verdict |
|---|---|---|
Accept-Encoding | 2–3 | Fine (CDN normalizes) |
Accept-Language (raw) | Thousands of strings | Normalize to supported languages first |
User-Agent | Tens of thousands | Never; derive device class at edge |
Cookie | One per user | Never on shared cache |
B.3 Route Classes#
| Class | Examples | Policy |
|---|---|---|
| Immutable | /assets/*.{hash}.js | 1 year, immutable |
| Semi-dynamic | /p/*, /c/*, public API GETs | 30–300s TTL, SWR, SIE, tag purge |
| Micro-cached | Hot dynamic endpoints | 1–10s TTL, collapsing |
| Private | /account, /cart, /checkout | private, no-store; edge for TLS only |
| Never through edge | Admin, internal | Separate hostname, not exposed |
Appendix C: Invalidation Mechanisms#
| Mechanism | Precision | Speed | Origin Impact | Best For |
|---|---|---|---|---|
| TTL expiry | Per object | Bounded by TTL | Smooth | Baseline freshness |
| URL purge | Single URL | Seconds–minutes depending on provider | Stampede if hot | Individual corrections |
| Prefix/wildcard purge | Many URLs | Varies | Large stampede risk | Emergency |
| Tag / surrogate-key purge | Every object tagged | Sub-second on some providers | Proportional to tag breadth | Semi-dynamic content |
| Soft purge | As above, marks stale | Same | Minimal (revalidation) | Default for all purges |
| Versioned URL | New object | Instant | None | Static assets |
C.1 Quick Comparison#
| Question | TTL | Tag Purge | Versioned URL |
|---|---|---|---|
| Works if messages are lost? | Yes | No | Yes |
| Freshness | ≤ TTL | ~seconds | Instant |
| Needs publisher discipline? | No | Yes (tagging, purging) | Build pipeline only |
Appendix D: HTTP Contract and Client Behavior#
D.1 Headers Cheat Sheet#
| Header | Purpose | Typical Value |
|---|---|---|
Cache-Control: s-maxage | Shared-cache TTL | s-maxage=60 |
Cache-Control: max-age | Browser TTL | max-age=0 for HTML, 1 year for hashed assets |
stale-while-revalidate | Serve stale during background refresh | 300 |
stale-if-error | Serve stale on origin error | 86400 |
Surrogate-Key / Cache-Tag | Purge targeting (provider-specific) | product-123 cat-9 |
Surrogate-Control | Edge-only directives (some providers) | max-age=3600 |
Vary | Declares request headers that change the response | Accept-Encoding |
Age | Seconds object has been in cache | Exposed for debugging and API freshness |
ETag / Last-Modified | Conditional revalidation (304) | Saves bytes on revalidation |
D.2 Browser vs Edge TTL#
Set HTML's browser max-age=0 (or short) with a longer s-maxage — you can purge the edge but not browsers. Hashed assets can be long in both.
D.3 Origin Lockdown#
Accept origin traffic only from CDN egress IP ranges and require a secret header or mTLS certificate from the edge, rotated regularly. IP ranges alone are shared by every customer of the same CDN.
Appendix E: Observability#
E.1 Core Metrics — Non-Negotiable#
# Effectiveness
edge_hit_ratio{route} byte_offload_pct{route}
shield_hit_ratio origin_rps / origin_egress_gbps
cache_key_cardinality{route} # fragmentation detector
# User experience (RUM + synthetics)
ttfb_p75{country} lcp_p75{country}
edge_error_rate{pop} external_synthetic_success{provider}
# Freshness
purge_latency_p99 purge_requests_total{publisher}
served_stale_total{reason=swr|sie} content_version_mismatch_total # sampled checks
# Safety
private_marker_in_cached_response_total two_user_check_failures_total
E.2 Critical Alerts#
| Alert | Condition | Action |
|---|---|---|
| Private content cached | Any private_marker_in_cached_response_total | Page security + platform; purge |
| Two-user check failure | Any | Page platform |
| Hit ratio collapse | Route hit ratio drops > 20 points in 10 min | Page platform; check key cardinality |
| Origin overload via misses | origin_rps > 3× baseline with hit ratio drop | Page origin + platform |
| Provider degradation | External synthetic success < 90% for 3 min | Page SRE; evaluate failover |
| Purge latency | purge_latency_p99 > 30s | Ticket; page if freshness SLO at risk |
E.3 Control Plane vs Data Plane#
The CDN's config/purge APIs (control plane) can fail independently of serving (data plane). Keep an out-of-band way to change DNS and to disable caching rules, and never make publishing depend synchronously on purge success — queue and retry.
E.4 Debugging "Users See Old Content"#
- Check
Ageand version headers from several PoPs. - Was a purge issued? Did it succeed? (purge logs, latency)
- Is the browser caching it (
max-ageon HTML)? - Is the shield serving stale (SWR/SIE because origin is erroring)?
- Is the cache key including something that splits old and new?
Appendix F: Scale Evolution#
F.1 What Works at Each Scale#
| Scale | Enough | Add When |
|---|---|---|
| Early | CDN for static assets with versioned URLs | Global users or launch traffic |
| Growth | Semi-dynamic caching, shield, SWR/SIE, tag purge, allowlist keys | Multiple teams changing edge config |
| Scale | Edge platform contract, staged config rollout, cost allocation, failover drills | Provider outage cost or bill growth |
| Very large | Multi-CDN steering, own caches in top metros, pre-positioning for media | Traffic and predictability justify capex |
F.2 Multi-Region Path#
Single origin → regional origins with shield mapping → data-residency-aware hostnames → per-region edge policies.
F.3 What You Don't Build on Day One#
- Multi-CDN steering
- Edge compute beyond redirects and header normalization
- Your own PoPs
- Custom admission policies
- Real-time purge verification infrastructure (start with sampled checks)
Appendix G: Multi-Tenancy, Fairness and Cost#
G.1 Shared Cache Capacity#
In a shared edge, one team's large objects (video, downloads) can evict another team's hot pages. Separate cache namespaces or storage tiers per content class; admission filters for large objects.
G.2 Purge Fairness#
A single publisher issuing broad purges can degrade freshness for everyone by consuming purge capacity. Per-publisher rate limits; priority lanes for corrections and legal takedowns.
G.3 Cost Allocation#
Attribute egress by route prefix or hostname to owning teams; publish monthly. Visibility alone typically drives meaningful reductions (image optimization, dropping unused bundles).
G.4 Tradeoff Summary#
| Choice | Default | Who Pays If Wrong |
|---|---|---|
| Cacheability | Opt-in per route | Security, if opt-out |
| Cache key | Allowlist | Users (wrong/leaked content) or finance (fragmentation) |
| Freshness | TTL + SWR + tag purge | Business (stale data) |
| Purge | Soft, rate-limited | Origin (stampede) |
| Vendor | Single + tested failover | Business during provider outage |
| Edge logic | Stateless, staged rollout | Everyone, on a global bad push |