Technologies referenced in this case study: Redis · DynamoDB · Cassandra · Kafka · Flink & Stream Processing · API Gateways
Related: CDN & Edge Caching · Distributed Caching · Rate Limiting · Stream Processing · Scaling Reads · Caching Fundamentals · Back-of-Envelope Estimation
How to Use This Case Study#
This is the Staff/Principal version of the most over-practiced interview question in the industry. Everyone can produce the base62 answer; this case study is about everything after it.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 5 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failures) → Deep Dives 2 and 4 |
| Deep Dive | 3+ hrs | Everything, including The Principal Lens and appendices (ID schemes, caching, abuse pipeline, deletion) |
What is a URL Shortener? — Why interviewers pick this topic
A URL shortener maps a short code (sho.rt/aZ3k9Qx) to a long destination URL and redirects anyone who visits the short link. Products built on top of it add custom aliases, branded domains, QR codes, click analytics, link expiry and link management for marketing teams.
The mechanics fit on an index card. The hard parts are what the index card leaves out: you are operating an open redirector on a trusted domain, which attackers will use to disguise phishing and malware; your links get printed on billboards and packaging, so they must resolve for a decade; your analytics are billed to customers, so bots and prefetchers must be filtered; and legal will eventually ask you to delete something that is cached in 300 edge locations.
Before vs After — the "phishing on our domain" incident:
Without an abuse posture:
t=0: Attacker scripts 40,000 short links to a credential-phishing page
t=+10min: Links sent in an SMS campaign; our domain looks trustworthy
t=+2h: Browser safe-browsing lists flag our shortener domain itself
t=+2h: Every link on the domain shows a red interstitial — including
millions of legitimate customer links on packaging and billboards
t=+3 days: Delisting requested; enterprise customers escalate; churn begins
With an abuse posture:
t=0: Same script. Account is 4 minutes old, no verified email,
creation rate 60/min → throttled at link 20, remaining queued for review
t=+30s: Destination scanned asynchronously against threat feeds → flagged
t=+31s: All 20 created links switched to an interstitial/blocked state;
edge cache purged for those 20 codes
t=+5min: Account suspended. Domain reputation untouched.
Why interviewers reach for this question: It is simple enough that there's no hiding behind complexity. At Senior level it tests estimation and a clean read path. At Staff level it tests whether you notice that the real system is a trust, abuse and lifecycle system with a trivially simple data model.
Mechanics Refresher: Short Code Generation
| Scheme | How It Works | Pros | Cons |
|---|---|---|---|
| Hash of URL (MD5/SHA → base62, truncate) | Deterministic code per URL | Dedup for free | Collisions after truncation need handling; same URL for two users shares analytics and ownership |
| Global counter → base62 | Increment a counter, encode | No collisions, shortest codes | Codes are enumerable (scrape all links); single counter is a coordination point |
| Counter ranges per server | Each server leases a block of 1M IDs | No per-request coordination | Still enumerable; gaps on crash (harmless) |
| Random code, check-and-insert | Random 7–8 chars, conditional insert | Unguessable; no coordination | Collision retries grow as keyspace fills |
| Snowflake-style IDs → base62 | Time + worker + sequence | Sortable, coordination-free | Longer codes (~11 chars); leaks creation time |
| Counter + bijective scramble (e.g., Feistel) | Counter mapped through a keyed permutation | Unique and non-enumerable, short | Key management; a leaked key makes codes predictable again |
For most production systems: random 7-character base62 codes with a conditional insert (or a counter mapped through a keyed permutation). Enumerability is a privacy bug — many shortened URLs point to documents people assumed were unlisted.
Executive Summary
If you only read one section, read this. Every decision in the case study follows from the contrast and the fault lines below.
What This Interview Actually Tests#
A URL shortener is not a hashing question. Base62 is solved; the interviewer has heard it 500 times.
It is a trust and lifecycle question with a read-path SLO, testing:
- Whether you see that you're running an open redirector that attackers will weaponize
- Whether you choose redirect semantics (301 vs 302) knowing they decide your analytics and your ability to revoke
- Whether you can keep a 100:1–1000:1 read path fast at the edge while still counting clicks accurately
- Whether you design for a link's full life: creation, abuse takedown, legal deletion, expiry, and the promise that it still works in 10 years
The key insight: Once a short link is printed, you can't take it back. Every decision — code format, domain, redirect type, caching, deletion — is a promise to people you'll never talk to. Staff engineers design the promise; Senior engineers design the table.
The L5 vs L6 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Estimates QPS and storage, picks base62 | Asks who creates links (anonymous public vs authenticated businesses) — abuse and analytics depend on it | Asks what the company is promising: link permanence, domain reputation, and whether this is a product or an internal platform |
| ID generation | Counter or MD5 truncation | Random or permuted codes (non-enumerable), conditional insert, collision math stated | Treats code length and domain as one-way doors; reserves namespaces and plans for 10+ years of growth |
| Redirect | "301 because it's permanent" | 302/307 by default to keep analytics and revocation; 301 only for opt-in static links | Decides org-wide which domains are allowed to be cached permanently by browsers — a revocation-capability decision |
| Abuse | "Rate limit the create API" | Layered: creation friction by trust tier, async destination scanning, click-time checks, interstitials, domain-reputation monitoring | Owns domain reputation as a company asset; separate domains per trust tier so one abuse wave can't blocklist enterprise links |
| Analytics | Increment a counter per click | Async click stream, bot/prefetch filtering, dedup, approximate uniques; never on the redirect's critical path | Defines analytics as a billed product with accuracy SLAs and privacy constraints (retention, IP truncation) |
| Deletion | "Delete the row" | Tombstone, purge edge caches, never reuse codes, legal-hold aware, analytics erasure | Designs the compliance pipeline (GDPR, DMCA, court orders) with SLAs and audit, and the deprecation policy if the product ever shuts down |
Why "first move" separates levels
L5: "100M new URLs per day, 10B redirects per day, 500 bytes per row → 18 TB per year. Base62, 7 characters." All correct. But every design decision later depends on a question never asked: who can create links?
L6: "Is creation anonymous and public, like the classic bit.ly homepage, or authenticated for business customers? Anonymous public creation means I'm primarily designing an abuse system. Authenticated business creation means I'm primarily designing an analytics and branded-domain system. I'll assume both exist, with different trust tiers — which drives a lot of the design."
L7: "Before the tiers: what are we promising? A printed link on a cereal box is a 10-year obligation. If we can't commit to that, we shouldn't let customers print our links — and if we can, the domain, code format and storage are one-way doors I want decided deliberately."
Why "redirect" separates levels
L5: "301 Moved Permanently, because the mapping never changes, and it lets browsers cache it so we get less load." True — and it means the browser never asks you again. You lose click analytics from that browser and you lose the ability to redirect it to a warning page if the destination turns malicious.
L6: "302 (or 307) by default, with a short Cache-Control: max-age like 5 minutes at the CDN and private, max-age=0 for browsers. I keep the ability to count, to revoke, and to show an interstitial. 301 is an opt-in for customers who want zero-latency repeat visits and accept losing repeat-visit analytics and revocation."
L7: "The redirect code is a revocation-capability policy. A 301 cached in a browser can persist indefinitely. If legal needs a link dead in 24 hours, 301s are a commitment we can't honor. That goes into the platform standard."
Why "abuse" separates levels
L5: "Rate limit link creation per IP, and check URLs against a blocklist."
L6: "Abuse happens at three times: at creation (friction and scanning), after creation (destinations change — a clean page turns into phishing after the link is shared), and at click time (the last chance). So I scan synchronously against a fast local blocklist at creation (< 20ms), asynchronously against threat-intel APIs and a crawler within 1 minute, re-scan on a schedule and on click-velocity anomalies, and serve an interstitial for flagged links. The metric that matters is time-from-creation-to-block for malicious links."
L7: "The asset at risk isn't the link; it's the domain. If safe-browsing lists flag our domain, every customer's link breaks. I'd put anonymous links and enterprise branded links on different domains so an anonymous abuse wave has a bounded blast radius."
The Staff Positions#
| Position | Rationale |
|---|---|
| Non-enumerable codes (random or permuted), never a raw counter | Enumerable codes let anyone scrape every link, many of which point to "unlisted" documents |
| 302/307 by default; 301 only as an opt-in | Keeps analytics, revocation, and interstitials possible |
| Redirect path never waits on analytics | Clicks are emitted asynchronously; the redirect SLO is p99 < 20ms at the edge |
| Abuse is a lifecycle, not a create-time check | Destinations change after creation; scanning must continue |
| Codes are never reused | A reused code sends old printed links to a new destination — a trust and security failure |
| Deletion = tombstone + cache purge + analytics erasure | A row delete leaves the link alive in CDN caches and in analytics |
| Separate domains by trust tier | Contains domain-reputation blast radius |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Public consumer shortener (anonymous or free accounts) | Abuse dominates; domain reputation is everything | Trust tiers, creation friction, async scanning, interstitials, 302 | Phishing wave blocklists the domain | Malicious link blocked within minutes; legitimate links never broken |
| Branded business links (marketing, custom domains, analytics) | Analytics accuracy billed to customers; links printed for years | Custom domains with managed TLS, click pipeline with bot filtering, link management APIs | Analytics disputes; broken printed links | Click counts within ±1–2% after filtering; 10-year resolvability |
| Platform link wrapping (every link in a messaging/social product, like Twitter's t.co) | Latency and safety at huge scale; links not user-managed | Automatic wrapping, click-time safety checks, very high cache hit rates | Wrapping service outage breaks every link in the product | Redirect p99 < 10–20ms; safety check on every click |
🎯 Staff Move: "I'll design a public shortener that also serves business customers with branded domains, because that's where abuse, analytics and permanence all collide. The anonymous tier is primarily an abuse problem; the business tier is primarily an analytics and permanence problem. They share the redirect path but not the domain or the trust model."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | ID Generation: Short vs Unguessable vs Coordination-Free | Shorter codes, non-enumerability, and zero coordination can't all be maximized at once |
| 2 | Redirect Semantics: 301 vs 302 | Browser caching and speed vs analytics, revocability and safety |
| 3 | Abuse: Friction at Creation vs Detection After | Blocking bad links up front slows everyone; detecting later leaves a window |
| 4 | Analytics Accuracy vs Redirect Latency | Counting precisely on the hot path vs async counting with filtering and approximation |
| 5 | Permanence vs Deletion | "Links never break" vs GDPR, DMCA, court orders and product shutdowns |
In the Wild: Real Production Systems#
Why this section belongs here: Each of these teaches a lesson the textbook design misses.
Bitly — The Business Is Analytics and Branded Links#
Bitly began as the canonical public shortener and evolved into a link-management and analytics product for businesses, with custom branded domains, QR codes and click analytics as the paid features. Public shortening is effectively the free acquisition funnel.
Staff insight: The redirect is the cheap part; the product is what you learn from clicks and the trust of a custom domain. A design that treats analytics as an afterthought counter misses what customers pay for.
Twitter's t.co — Wrapping Every Link for Safety and Measurement#
Twitter wraps links posted in tweets with its t.co shortener. This lets it check destinations against known-malicious lists when users click and measure link engagement, regardless of the original URL's length.
Staff insight: Platform link wrapping inverts the intent: users don't choose to shorten; the platform does it for safety. The redirect path becomes a safety-check path, and its availability is the availability of every link in the product.
Google's goo.gl — The Permanence Promise, Tested#
Google stopped allowing new goo.gl links in 2019 and later announced that existing links would stop redirecting in August 2025 — then, after pushback, narrowed that plan to preserve links that were still actively used.
Staff insight: Even a company with near-unlimited infrastructure found that "every link forever" is an ongoing cost and a governance decision. Principal-level design includes the deprecation policy: what you promise, what a shutdown looks like, and how much it costs to keep the lights on for a read-only archive.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Auto-increment counter + base62" | "Can someone scrape all your links?" | Privacy and enumerability |
| "MD5 the URL and take 7 characters" | "What's your collision rate? Two users shorten the same URL — who owns the analytics?" | Collision math, ownership semantics |
| "301 redirect" | "How do you count clicks? How do you disable a phishing link?" | Revocation and analytics |
| "Cache in Redis" | "What's the hit rate? What happens on a viral link? On a deleted link?" | Caching literacy, purge |
| "Rate limit creation" | "Attacker uses 10,000 residential IPs. Now what?" | Layered abuse thinking |
| "Increment click count in the DB" | "Ten million clicks in an hour on one link?" | Hot keys, async aggregation |
| "Delete the row" | "It's cached in 300 PoPs and 50M browsers" | Lifecycle, cache purge, 301 caching |
System Architecture Overview#
Reading the diagram: Most redirects are served at the edge from a short-TTL cache without touching the origin. Cache misses hit the redirect service, which checks an in-memory/Redis cache and then the link store. Click events leave the hot path asynchronously — from edge logs or the service — into a stream processor that filters bots and builds per-link rollups. Creation is a separate, low-QPS path with trust-tier checks and a fast local scan; deeper scanning continues asynchronously and can block and purge a link at any point in its life. Legal takedowns use the same block-and-purge path.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| IDs | "Counter + base62" | "Random 7-char base62 with conditional insert, or a counter through a keyed permutation. Never enumerable." |
| Redirect | "301" | "302 by default: keeps analytics, revocation and interstitials. 301 opt-in." |
| Read path | "Redis cache" | "Edge cache with 5-min TTL, origin cache, KV store. Negative caching for unknown codes. p99 < 20ms." |
| Analytics | "Counter in DB" | "Async click stream, bot/prefetch filtering, HyperLogLog uniques, rollups. Never on the redirect path." |
| Abuse | "Rate limit + blocklist" | "Trust tiers, sync + async scanning, rescans, click-time checks, interstitials, domain separation." |
| Deletion | "Delete row" | "Tombstone, purge CDN, erase analytics per policy, never reuse the code, honor legal holds." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| 62^6 | ~56.8 billion | Exhausted in ~1.5 years at 100M/day — too small for a public shortener |
| 62^7 | ~3.52 trillion | Default length; decades of headroom at 100M/day |
| 62^8 | ~218 trillion | For very high volume or extra sparseness against guessing |
| Collision probability, random 7-char at 10B links | ~0.28% per insert | One retry in ~350 inserts — fine with conditional insert |
| Guess success rate, 10B links in 62^7 space | ~1 in 350 | Random codes are sparse but NOT secret; don't put secrets behind short links |
| Read:write ratio | 100:1 to 1000:1 | Optimize the redirect path; creation can be slow and careful |
| Redirect p99 target at edge | < 20ms | Users perceive the shortener as part of page load |
| Row size (code, URL, owner, status, timestamps) | ~300–600 bytes | 10B links ≈ 3–6 TB before replication |
| Edge cache TTL | 1–5 min typical | Bounds revocation delay without purge |
| Cache hit rate (edge + origin) | 95–99% | Power-law click distribution; top 1% of links take most clicks |
| Bot / prefetch share of raw clicks | often 20–50%+ | Link previews, scanners, crawlers inflate raw counts |
| Sync creation scan budget | < 20ms | Local blocklist / bloom filter only |
| Async scan target | < 1 min from creation | Most phishing damage happens in the first hours |
| Viral link peak | 10K–100K+ clicks/sec on one code | Hot key — must be served from edge |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on estimation and base62 math — the part every interviewer has heard hundreds of times — and never reach abuse, redirect semantics, or deletion. Do the estimation in 2 minutes, say "the core is a key-value lookup," and spend the time on the parts that decide your level.
Phase 1: Requirements & Framing (2–3 minutes)#
Functional scope in one sentence:
"Create a short link for a long URL — optionally with a custom alias and expiry — redirect visitors, and give link owners click analytics."
Then the questions that change the design:
"Who creates links — anonymous users, free accounts, paying businesses with branded domains, or all three? How long must a link work — is this for printed media? Do customers pay for analytics accuracy? And what's our legal posture on takedowns? I'll assume all three creator tiers, links that must work for 10+ years unless taken down, analytics as a paid feature, and that we must honor GDPR erasure and DMCA-style takedowns."
Estimation in 60 seconds:
"100M new links/day ≈ 1.2K writes/sec average, ~5K peak. Redirects at 100:1 → 10B/day ≈ 115K/sec average, ~500K/sec peak, with viral links at 50K+/sec on a single code. Storage: 500 bytes × 100M/day × 365 ≈ 18 TB/year before replication. Clicks: 10B events/day × ~200 bytes ≈ 2 TB/day raw — analytics is the storage problem, not links."
🎯 Staff Move: "Notice analytics raw data is ~40× the link data. The link table is small; the click stream and the abuse system are where the engineering goes."
Phase 2: Core Entities & API (1–2 minutes)#
- Link:
code,domain,long_url,owner_id,created_at,expires_at,status(active / flagged / blocked / deleted / expired),redirect_type(302 default),trust_tier - Click event:
code,ts,edge_pop,country,referrer_host,ua_class,is_bot,visitor_hash(salted, rotating) - Domain:
domain,owner_id,tls_cert_state,trust_tier
POST /v1/links { long_url, custom_alias?, domain?, expires_at? } → 201 { short_url, code }
GET /{code} → 302 Location: long_url | 404 | 410 | interstitial
GET /v1/links/{code}/stats?from&to&granularity → { clicks, uniques, by_country, by_referrer }
DELETE /v1/links/{code} → 204 (tombstone + purge)
PATCH /v1/links/{code} { long_url } → 200 (paid tier only; re-scanned)
🎯 Staff Move: "Status is the most important column. A link isn't just 'exists or not' — it's active, flagged (interstitial), blocked, deleted (410 Gone), or expired. Every read path must respect it, and every status change must purge caches."
Phase 3: High-Level Architecture (≤5 minutes)#
Walk the flow in 90 seconds:
- Visitor hits the edge. 90%+ of redirects are answered from edge cache (short TTL).
- On miss, the redirect service checks an in-memory/Redis cache, then the link store (single-key lookup, ~2–5ms).
- It checks status, returns 302 (or 410 / interstitial), and emits a click event asynchronously.
- Creation goes through trust-tier checks, a fast local scan, code generation with conditional insert, and triggers async deep scanning.
- Any status change (abuse block, deletion, edit) writes the store, then purges the code at the edge and origin caches.
🎯 Staff Move: "That's the base design; it would pass an L5 bar. The Staff-level questions are: which ID scheme and why; 301 vs 302; how we stop this becoming a phishing service; how analytics stays accurate without touching the redirect latency; and what 'delete' means across 300 caches. Which first?"
Phase 4: Transition to Depth (1 minute)#
"The lookup is trivial. What makes this system hard is that it's a public redirector on a trusted domain with a 10-year promise. I'd go deep on abuse first — it's the failure that takes down every customer's links at once — then redirect semantics and caching, then analytics, then deletion."
Phase 5: Deep Dives (25–30 minutes)#
Deep dive A: ID generation (4–5 min)
"Requirements: short, unique, non-enumerable, no per-request coordination. A raw counter is enumerable — anyone can walk aaaaaa1, aaaaaa2… and scrape every link, including ones to 'anyone with the link' documents. MD5-truncate collides and ties identity to the URL, so two users shortening the same URL would share a code and its analytics. I'll use random 7-char base62 — 3.5 trillion space — with a conditional insert (PutItem if attribute_not_exists(code)). At 10B links, collision chance per insert is ~0.3%; retry once. Alternative with no retries: each server leases counter blocks of 1M and maps counter values through a keyed 42-bit Feistel permutation, then base62 — unique by construction, non-sequential to outsiders."
"Custom aliases share the same namespace but go through a reserved-word list (brand names, 'login', 'admin', profanity), a minimum length for free tiers (e.g., ≥ 8 chars, to preserve short codes for random allocation), and a trademark dispute process."
Deep dive B: Redirect semantics and caching (5–7 min)
"302 by default. 301 makes browsers cache the redirect indefinitely — we'd lose repeat-click analytics and, worse, the ability to put an interstitial in front of a link that turns malicious. At the edge: cache the redirect for 5 minutes (s-maxage=300), browsers no-store or max-age=0. Revocation bound without purge is 5 minutes; with an active purge API it's ~seconds across PoPs. Negative-cache unknown codes for 60s so scanners probing random codes don't hammer the store — but a newly created code must purge its negative entry, or a creator who tests their link instantly sees 404 for a minute."
"Viral links: one code at 100K clicks/sec. The edge absorbs it — each PoP fetches once per TTL. Request coalescing at the edge (one origin fetch per PoP on miss) prevents a stampede when the TTL expires."
Deep dive C: Abuse lifecycle (6–8 min)
"Three layers. Creation: trust tier sets friction — anonymous: 10 links/hour per IP+device fingerprint, CAPTCHA after 3; free verified accounts: 100/day; paid: contractual limits. A synchronous check against a local blocklist and bloom filter of known-bad domains in < 20ms. After creation: async scanner within 1 minute — threat-intel APIs such as Google Safe Browsing, a headless crawler following the redirect chain (attackers chain through other shorteners), and ML scoring on URL features. Rescan on schedule and when click velocity spikes — a link getting 10K clicks in its first hour from SMS referrers is suspicious. Click time: flagged links show an interstitial; blocked links show a block page. Metric: abuse.time_to_block_p50 for confirmed-malicious links, target < 5 minutes."
Deep dive D: Analytics without slowing redirects (4–5 min)
"The redirect emits an event to a local buffer and returns; events flow to Kafka and a stream processor. The processor filters bots (known UA lists, datacenter ASNs, link-preview fetchers like messaging-app unfurlers, prefetch headers), deduplicates double-clicks within a few seconds, and writes per-link minute rollups. Uniques via HyperLogLog (~1.6% error at 1.5 KB per sketch). Edge-served clicks come from CDN logs, delivered in seconds to minutes. Customers see 'near real-time' (< 1 min lag) and a daily-finalized number that's what we bill on."
Deep dive E: Deletion and compliance (3–4 min)
"Delete is a status change to a tombstone, never a row removal: the code must never be reallocated, or old printed links send people to a new owner's destination. Tombstone returns 410 Gone. Purge edge and origin caches. For GDPR erasure of a user, we delete or anonymize their links' destinations and their analytics identifiers — keeping the code reserved. Legal holds override deletion for links under investigation. 301s already cached in browsers are the one thing we can't recall — another reason 302 is the default."
Phase 6: Wrap-Up (2–3 minutes)#
"The core is a key-value lookup behind an edge cache. The system is really three others around it: an abuse system that protects domain reputation for every customer, an analytics pipeline that customers pay for and must be accurate after bot filtering, and a lifecycle system that makes delete, takedown and expiry actually work across caches while keeping the 10-year promise for everything else. Ownership: redirect platform, trust & safety, analytics, and legal each own a piece, with trust & safety authorized to block any link without an engineering deploy."
"Next: per-customer branded domains with automated TLS, link-in-bio pages, and a read-only 'archive mode' plan in case the product is ever deprecated."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| 10 min of estimation | Derives every number in detail | 60 seconds of numbers, highlights that analytics dwarfs link storage |
| Base62 lecture | Explains encoding math | "Random 7 chars, 3.5T space, conditional insert. Moving on." |
| Database bake-off | SQL vs NoSQL debate | "Single-key lookups; any KV store. DynamoDB or Cassandra." |
| No abuse discussion | Stops at "rate limit" | Lifecycle abuse with time-to-block metric |
| 301 by reflex | "Permanent, cacheable" | Explains the revocation and analytics cost |
| Delete as row delete | "DELETE FROM links" | Tombstone, purge, never reuse, legal holds |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Because it's familiar, the URL shortener exposes level differences brutally. Everyone gets the lookup right. The Staff-level signal is noticing the problems that emerge only in production: a public redirect service becomes a phishing tool; a counter-based ID leaks private documents; a 301 makes a malicious link unrevokable; bots inflate billed analytics; a code reused after deletion sends a printed QR code to a stranger's site. None of these are hard to fix — they're hard to notice.
What interviewers listen for:
- Creator trust tiers asked about in the first 2 minutes
- Non-enumerable ID scheme with a reason
- Redirect status code chosen for revocation and analytics
- Abuse as a lifecycle with a time-to-block metric
- Deletion that includes caches, analytics and code non-reuse
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"A link created yesterday by a free account is now pointing — through two other redirectors — to a bank-login phishing page, and it's being clicked 2,000 times a minute from SMS. How long until your system stops sending people there, and what does the visitor see?"
A Staff answer names the detection path (click-velocity trigger → rescan following the redirect chain), the action (status → blocked, purge edge within seconds), the visitor experience (block page), the metric (time-to-block), and the follow-up (account suspended, sibling links by the same account/destination blocked).
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Public consumer shortener → abuse-first
- Constraint: anyone can create; attackers will
- Mechanism: trust tiers, creation friction, sync + async scanning, interstitials, separate domain
- Failure mode: domain blocklisted by browsers/email providers → every link on it breaks
- Who pays: every customer on the domain; trust & safety on-call
Branded business links → analytics- and permanence-first
- Constraint: customers pay for analytics and print links for years
- Mechanism: custom domains with managed TLS, click pipeline with bot filtering, link editing (paid), SLAs
- Failure mode: inaccurate analytics → billing disputes; expired TLS cert → brand's links break
- Who pays: customer success, finance (credits), the platform on-call for cert renewals
Platform link wrapping → latency- and safety-first
- Constraint: every outbound link in a product passes through; no user management
- Mechanism: extreme cache hit rates, click-time safety checks, fail-open to the original URL for availability (or fail-closed for known-bad)
- Failure mode: wrapper outage = all links in the product broken
- Who pays: the whole product's users; the wrapper team carries the product's availability
2.2 When NOT to Build a URL Shortener#
- You need it for one marketing campaign. Use an existing link-management product; building one means inheriting abuse and permanence obligations.
- You want short links to private content. Short codes are guessable at scale (~1 in 350 random guesses hits some link at 10B links). Use long, high-entropy capability URLs or authenticated access instead.
- Internal "go/" links. An internal go-links service is a much simpler problem — authenticated users, no public abuse, human-readable names. Don't over-engineer it with the public system's machinery.
- Deep links in your own app. Use your own domain with path routing; a shortener adds a hop and a dependency.
2.3 What the Interviewer Leaves Underspecified#
| Omitted Detail | Why It Matters | What to Ask |
|---|---|---|
| Who creates links | Drives abuse posture | "Anonymous, free accounts, paying businesses?" |
| Link lifetime | Drives storage, ID length, deprecation plan | "Must links work for years? Can they expire?" |
| Editability | Editing destinations enables bait-and-switch abuse | "Can owners change where a link points?" |
| Analytics as product | Accuracy bar, bot filtering, retention | "Is analytics billed? What accuracy is promised?" |
| Custom domains | TLS at scale, domain ownership verification | "Do customers bring their own domains?" |
| Legal posture | Takedown SLAs, GDPR, legal holds | "What's our obligation on takedown requests?" |
Staff engineers surface these. Senior engineers assume a hash table.
2.4 Precise Terminology#
| Term | Precise Meaning | Common Misuse |
|---|---|---|
| 301 Moved Permanently | Browsers and proxies may cache indefinitely | Chosen "for performance" without considering revocation |
| 302 Found / 307 Temporary Redirect | Not cached by default; 307 preserves method | Assumed to be slower (edge caching still applies) |
| 308 Permanent Redirect | Like 301 but preserves method | Rarely relevant for GET redirects |
| 410 Gone | Resource intentionally removed | 404 used for deleted links (loses meaning) |
| Open redirect | A redirector that sends users anywhere | Treated as harmless |
| Interstitial | Warning page before redirecting | Treated as a block |
| Enumeration | Iterating the code space to discover links | Dismissed because "codes look random" |
| Negative caching | Caching "not found" results | Forgotten, so random-code probing hits the DB |
| Tombstone | Record marking a deleted code as permanently reserved | Row deleted, code reused |
| Unique visitor | Deduplicated visitor within a window, usually approximate | Promised as exact |
🎯 Staff Insight: "A short code is not a secret. At 10 billion links in a 3.5-trillion space, a random guess hits some live link 1 time in 350. Anyone shortening a private URL should know that; we should say it in our docs."
3. The Fault Lines#
Every fault line here looks like a small technical choice and turns out to be a promise to people outside the company — visitors, customers, courts.
3.1 Fault Line 1: ID Generation — Short vs Unguessable vs Coordination-Free#
The tension: The shortest codes come from dense counters, which are enumerable. Unguessable codes need sparse random allocation, which means longer codes or collision retries. Coordination-free generation needs either randomness or pre-leased ranges.
| Scheme | Length at 10B links | Enumerable? | Coordination | Who Pays |
|---|---|---|---|---|
| Global counter | 6 chars | Yes — trivially | Single counter (hot) | Link owners (scraped private URLs) |
| Leased counter ranges | 6 chars | Yes, with gaps | Lease every 1M IDs | Link owners |
| Hash + truncate | 7 chars | No | None, but collisions | Users sharing a URL (merged ownership/analytics) |
| Random + conditional insert | 7 chars | No (sparse) | None; ~0.3% retries at 10B | Platform (retry logic, negligible) |
| Counter + keyed permutation | 7 chars | No, unless key leaks | Range leases | Security (key management) |
| Snowflake → base62 | ~11 chars | Partially (time-ordered) | None | Users (longer links) |
L6 answer: "Random 7-char with a conditional insert. It's coordination-free, non-enumerable, and the collision math is benign — at 10B links a retry happens once per ~350 inserts. I'd monitor create.collision_retries and move new codes to 8 characters around 50B links, when density starts to matter for both retries and guessability. Existing codes never change length."
When to deviate: Internal go-links (enumeration is fine, humans pick names); extreme write rates with strict latency where retries are unwelcome → counter + permutation.
🧭 Principal Move: "Code length and alphabet are a one-way door — they're printed on packaging. I'd reserve prefixes now: e.g., codes starting with a particular character reserved for future formats, and exclude ambiguous characters (0/O, 1/l/I) in codes intended for print and QR use."
3.2 Fault Line 2: Redirect Semantics — 301 vs 302#
The tension: 301s are cached by browsers, so repeat visits skip you entirely — fast and cheap. That same property means you can no longer count, redirect to a warning, or revoke for that browser.
| Choice | Latency (repeat visit) | Analytics | Revocation | Who Pays |
|---|---|---|---|---|
| 301, long cache | ~0ms (browser-local) | First visit only | Impossible for cached browsers | Visitors (if link turns malicious), customers (lost analytics) |
| 302/307 + edge cache 5 min | ~10–30ms (edge hop) | Every click | ≤ 5 min, seconds with purge | Platform (edge cost per click) |
| 302 + no caching | 30–100ms (origin) | Every click | Immediate | Origin capacity |
| 200 + HTML/JS redirect | +1 round trip + render | Every click, plus client signals | Immediate | Visitors (slower), SEO |
L6 answer: "302 with a 5-minute edge TTL is the default. The edge makes the cost of 302 nearly identical to 301 for users — tens of milliseconds — while keeping every click visible and every link revocable. I'd offer 301 only as an opt-in for paying customers who explicitly trade analytics for speed, and never on the anonymous tier where abuse risk is highest."
🎯 Staff Insight: "The argument for 301 is usually cost. At $0.50–1.00 per million edge requests, 10B redirects/day is ~$5–10K/day. That's a real number — and it's exactly what the analytics product is priced to cover."
3.3 Fault Line 3: Abuse — Friction at Creation vs Detection After#
The tension: Every check at creation slows or blocks legitimate users; every check after creation leaves a window where a malicious link works.
| Layer | When | Budget | Catches | Misses | Who Pays |
|---|---|---|---|---|---|
| Trust-tier rate limits | Creation | < 5ms | Bulk scripted abuse | Distributed low-rate abuse | Anonymous users (CAPTCHAs) |
| Sync local blocklist | Creation | < 20ms | Known-bad domains | New domains, redirect chains | Nobody (cheap) |
| Async scan (threat intel + crawler) | 0–60s after | Seconds | Known and newly flagged destinations, chains | Cloaked pages that show bots clean content | Platform (crawler fleet) |
| Rescan on schedule / velocity | Hours–days | Background | Bait-and-switch destinations | Low-traffic long-tail | Platform |
| Click-time check | Every click (edge flag lookup) | < 1ms | Anything already flagged | Unflagged | Visitors (interstitial) |
| Domain separation | Architecture | — | Limits reputation blast radius | — | Platform (multiple domains, certs) |
L6 answer: "Friction scaled by trust: anonymous creation gets rate limits, CAPTCHAs and optionally a preview interstitial on every link for the first hour; verified paying customers get near-zero friction. Detection is continuous. The number I'd publish internally is time-to-block p50 and p95 for confirmed-malicious links, plus 'victim clicks before block'."
When to deviate: A platform link-wrapper (t.co-style) has no creator friction at all — the safety burden moves entirely to click time.
🧭 Principal Move: "I'd give anonymous links a separate domain from branded business links. When — not if — the anonymous domain gets flagged by a browser or email provider, enterprise customers' printed links keep working. That's a domain-portfolio decision, owned above any one team."
3.4 Fault Line 4: Analytics Accuracy vs Redirect Latency#
The tension: Precise counting wants a synchronous, transactional increment per click. The redirect wants zero dependencies. And raw clicks are polluted: link previews, security scanners, prefetching browsers and bots can be 20–50%+ of hits.
| Approach | Redirect Impact | Accuracy | Who Pays |
|---|---|---|---|
| Sync DB increment | +2–10ms, hot-row contention on viral links | Exact raw count, includes bots | Visitors (latency), DB on-call |
| Async stream + rollups | ~0ms (buffered emit) | Exact after pipeline; lag seconds–minutes | Analytics team |
| Async + bot filtering + dedup | ~0ms | "Human clicks" within ±1–2% | Analytics team (classifier upkeep) |
| Sampling | ~0ms | ±x% statistically | Customers (can't bill on samples) |
L6 answer: "Async stream with filtering. Raw events from the edge and origin go to Kafka; the stream processor tags bots (UA lists, datacenter ASNs, known preview fetchers, headless signatures), dedups repeated clicks from the same visitor hash within 5 seconds, and writes per-minute rollups. Uniques use HyperLogLog. The customer dashboard shows 'provisional' numbers within a minute and 'final' numbers after a daily reconciliation, which is what we bill on."
🎯 Staff Insight: "Always report both raw and filtered counts. When a customer's own ad platform shows 30% more clicks than we do, the conversation is easy if we can show exactly which were link-preview bots."
3.5 Fault Line 5: Permanence vs Deletion#
The tension: The product promise is that links don't break. Law, safety and business reality require that some links stop working — quickly, everywhere, and provably.
| Trigger | Required Outcome | SLA (typical target) | Who Decides |
|---|---|---|---|
| Owner deletes link | Stops resolving; code reserved | Minutes | Owner |
| Abuse block | Block page; code reserved | Seconds–minutes | Trust & safety (automated + human) |
| Copyright / DMCA-style notice | Disable pending review | Hours–days (legal process) | Legal |
| GDPR erasure of a user | Destinations and personal analytics removed | Within statutory window (~30 days) | Privacy team |
| Court order / legal hold | Preserve data, possibly disable | Per order | Legal |
| Expiry | 410 Gone after expires_at | At expiry ± cache TTL | Owner at creation |
| Product deprecation | Read-only or shut down | Months of notice | Executive |
L6 answer: "Every terminal state becomes a tombstone that reserves the code forever — reuse would send old printed links to a new destination. Every status change writes the store first, then purges edge and origin caches, and the purge is verified by probing a sample of PoPs. Analytics tied to a deleted user's personal data is erased or anonymized per policy; aggregate counts may remain."
4. Failure Modes & Operational Reality#
4.1 Domain Blocklisted — Every Link Breaks#
t=0: Phishing campaign uses 3,000 links created over 6 hours via rotating residential IPs
t=+4h: Email providers start junking all mail containing sho.rt links
t=+6h: A major browser's safe-browsing list flags sho.rt domain-wide
t=+6h: Every link on sho.rt — including enterprise links on printed packaging — shows
a full-page red warning
t=+6h10m: Support volume 20× baseline; enterprise escalations
t=+2 days: Delisting after cleanup evidence; reputation damage persists for weeks
Detection: domain.reputation_status{provider} (poll public safe-browsing/transparency APIs hourly), abuse.links_blocked_per_hour, email.deliverability_rate for links in our own notifications.
Blast radius: Every link on the domain.
Mitigation: Mass-block the campaign (by destination, account cluster, creation fingerprint); file delisting with evidence.
Prevention: Domain separation by trust tier; creation velocity anomaly detection across accounts; preview interstitials for new anonymous links.
Owner: Trust & safety (response), platform (domain architecture), comms (customer notice).
4.2 Viral Link Stampede at TTL Expiry#
t=0: Celebrity posts a link: 120K clicks/sec globally within 2 minutes
t=+0s: Edge caches warm; origin sees ~1 request per PoP per 5 min. Fine.
t=+5m: TTL expires simultaneously in 300 PoPs; each PoP gets thousands of concurrent misses
Without coalescing: 300 PoPs × 2,000 concurrent misses = 600K origin requests in 1s
Redis shard owning the key: 1 hot key → CPU 100% → timeouts → 5xx at edge
With coalescing + stale-while-revalidate:
Each PoP sends 1 origin request and serves stale for up to 60s meanwhile.
Origin sees ~300 requests. No blip.
Detection: edge.origin_fetch_rate, cache.hot_key_qps{code}, redirect.5xx_rate.
Mitigation: Request coalescing at the edge; stale-while-revalidate=60; origin-side in-process cache for hot codes.
Prevention: TTL jitter (±20%) so PoPs don't expire simultaneously. See Distributed Caching.
Owner: Redirect platform team.
4.3 Deletion That Didn't Delete#
t=0: Legal takedown: link must stop resolving within 24h
t=+10min: Engineer sets status=deleted in the store. Ticket closed.
t=+3 days: Complainant reports link still redirects for them.
Root cause: link was created as 301 by a paying customer; complainant's browser
cached the 301 indefinitely. Also, two PoPs missed the purge
(purge API returned 200 but PoP was mid-deploy).
Detection: Post-purge verification probes (purge.verification_failures); synthetic checks of a sample of deleted codes from multiple regions.
Blast radius: Legal exposure per affected link.
Mitigation: Re-purge; for 301-cached browsers, nothing server-side works — document it.
Prevention: 301 disallowed on tiers where takedown risk is high; purge verification loop; takedown SLA defined as "edge-verified."
Owner: Platform (purge reliability), legal (SLA definition), product (301 policy).
4.4 Analytics Inflation Dispute#
A customer's campaign shows 1.8M clicks in our dashboard but 900K sessions in their web analytics. Investigation: a messaging app's link-preview fetcher changed its user-agent string, bypassing the bot filter; security gateways at enterprise recipients pre-click every link in email.
Detection: analytics.bot_ratio{code} anomaly vs baseline; analytics.clicks_without_js_followup (if the destination is instrumented); ASN distribution shift.
Mitigation: Update classifiers; recompute affected rollups from raw events (keep raw 30–90 days for exactly this).
Prevention: Classifier regression tests; report raw and filtered counts; publish methodology.
Owner: Analytics team; customer success for communication and credits.
4.5 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Link store unavailable | store.error_rate, redirect.5xx_rate | Cache misses only (~1–5% of clicks) | Serve stale from cache; extend TTL | Redirect platform |
| Viral hot key | cache.hot_key_qps | One code; possibly its cache shard | Coalescing, SWR, local cache | Redirect platform |
| Domain blocklisted | domain.reputation_status | All links on the domain | Mass block, delisting | Trust & safety |
| Scanner backlog | abuse.scan_queue_age_s | Malicious links live longer | Scale scanners; stricter creation friction | Trust & safety |
| Purge failures | purge.verification_failures | Deleted links still resolve | Re-purge, alert CDN | Platform |
| Click pipeline lag | analytics.pipeline_lag_s | Stale dashboards | Scale consumers; no redirect impact | Analytics |
| Bot filter drift | analytics.bot_ratio anomaly | Billing accuracy | Recompute from raw | Analytics |
| TLS cert expiry on custom domain | tls.cert_days_remaining{domain} | All links on that brand's domain | Automated renewal; alert at 21 days | Platform |
| Code collision spike | create.collision_retries | Creation latency | Lengthen new codes | Platform |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| IDs | Counter or hash + base62 | Non-enumerable scheme with collision math and growth plan | Code format as one-way door; reserved prefixes; print-safe alphabet |
| Redirects | 301 for performance | 302 default + edge TTL; 301 opt-in with stated tradeoff | Revocation capability as org policy across domains |
| Read path | Redis cache | Edge + origin cache, negative cache, coalescing, SWR | Cost per billion redirects tied to pricing |
| Abuse | Rate limit + blocklist | Lifecycle: tiers, sync/async scan, rescans, click-time, time-to-block | Domain portfolio and reputation as company asset; T&S authority model |
| Analytics | Counter per click | Async pipeline, bot filtering, HLL, provisional vs final | Analytics as billed product with published methodology and privacy limits |
| Lifecycle | Delete row | Tombstones, purge verification, no reuse, legal holds | Compliance pipeline with SLAs; deprecation and archive policy |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Trust tiers early | "Who creates links? Anonymous creation means this is mostly an abuse system." |
| Enumeration named | "A counter lets anyone scrape every link — including unlisted docs." |
| 302 with a reason | "302 keeps analytics and revocation; the edge makes it nearly as fast as 301." |
| Time-to-block metric | "I measure minutes from creation to block for confirmed-malicious links." |
| Never reuse codes | "A reused code sends a printed QR code to someone else's site." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| 15 minutes on base62 and estimation | Spends the interview on what everyone knows |
| Synchronous click counting on a hot row | Viral link takes down the database |
| No abuse discussion | Ignores the dominant real-world failure |
| "Delete the row" | Leaves caches serving it and enables code reuse |
| Sharding debate for a 20 TB KV dataset | Complexity where none is needed |
5.4 Common False Positives#
- Elaborate consistent-hashing ring for the link store: A managed KV store handles this; it's not the interesting part.
- Bloom filters for collision checks: Clever, but a conditional insert already solves it.
- "Use a distributed ID service like Snowflake": Valid, but produces longer codes and doesn't address enumeration.
- Perfect capacity math: Necessary, not sufficient.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Creator tiers, lifetime, analytics, legal; 60-second estimate |
| Entities + API | 3–5 min | Link with status; click event; create/redirect/stats/delete |
| High-level design | 5–12 min | Edge → redirect → cache → store; async clicks; create path with scan |
| Transition | 12 min | Offer abuse, redirect/caching, analytics, deletion |
| Deep dives | 12–40 min | Abuse lifecycle → 302 + caching → IDs → analytics → deletion |
| Wrap-up | 40–45 min | Ownership across platform, T&S, analytics, legal; permanence promise |
6.2 How Interviewers Pivot — And What They're Testing#
| Interviewer Signal | What They Care About | Where to Go Deep |
|---|---|---|
| "How do you generate codes?" | Enumeration awareness | Random vs permuted; collision math |
| "It went viral" | Hot key handling | Edge, coalescing, SWR, TTL jitter |
| "Someone's using it for phishing" | Abuse maturity | Lifecycle, time-to-block, domain separation |
| "Customer says clicks are wrong" | Analytics rigor | Bot filtering, raw vs filtered, recompute |
| "Legal wants it gone" | Lifecycle | Tombstones, purge verification, 301 caveat |
| "Custom domains" | Operational scale | TLS automation, domain verification |
6.3 What to Deliberately Skip#
| Topic | Why L5 Goes Here | What L6 Says Instead |
|---|---|---|
| Base62 encoding math | Easy points | "Standard base62; 7 chars = 3.5T." |
| SQL vs NoSQL | Familiar debate | "Single-key lookups; any KV store works." |
| Detailed sharding | Feels scalable | "20 TB over years; a managed KV store partitions it." |
| Frontend link UI | Visible | "CRUD over the API." |
6.4 Follow-Up Questions to Expect#
- "Two users shorten the same URL. Same code or different?"
- "How do you stop someone from scraping all your links?"
- "A link gets 100K clicks/sec. Walk me through the cache behavior at TTL expiry."
- "How quickly can you disable a malicious link everywhere?"
- "How do you count unique visitors without storing IPs?"
- "A customer wants to change where their printed QR code points. Allowed?"
- "The company decides to shut this product down. What happens to the links?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design a URL shortener."
Staff Answer
"Quick framing: the lookup is a key-value get behind a cache — I'll cover it in five minutes. The design is really decided by four questions: who can create links, how long they must work, whether analytics is a paid product, and what our takedown obligations are. I'll assume anonymous plus business tiers, 10-year links, billed analytics, and GDPR/DMCA obligations. Numbers: ~1.2K creates/sec, ~115K redirects/sec average, 500K peak, viral single codes at 50K+/sec; link data ~18 TB/year, click data ~2 TB/day raw. I'll go: base design, then abuse, then 302 and caching, then IDs, analytics and deletion."
Why this is L6:
- Compresses the familiar part and says so
- Names the four questions that actually change the design
- Notices analytics dwarfs link storage
What L7 adds:
- Asks what promise the company is making (link permanence) and whether it's funded for 10 years
- Raises domain strategy (separate trust tiers) as a first-day decision
Drill 2: Same URL, Two Users#
Prompt: "Two users shorten the same long URL. Do they get the same code?"
Staff Answer
"No — different codes. A link is owned: it has an owner, analytics, expiry, and potentially a takedown status. Sharing a code would merge two customers' analytics and let one user's deletion break the other's printed link. Dedup is only useful for the same owner: if a user shortens the same URL twice, I can return their existing code via a (owner_id, url_hash) index — saving storage and keeping their analytics in one place. For anonymous users, I don't dedupe, because dedup by URL across anonymous users leaks that someone else shortened this URL."
Why this is L6:
- Treats ownership semantics, not storage, as the deciding factor
- Offers per-owner dedup as a scoped optimization
- Spots the privacy leak in global dedup
What L7 adds:
- Notes storage saving from global dedup is small (~10–20%) vs the permanent ownership complexity
- Writes ownership semantics into the API contract so no future team "optimizes" it
Drill 3: Make 302 Concrete#
Prompt: "Why not 301? It's permanent and faster."
Staff Answer
"A 301 can be cached by the browser indefinitely. From then on, that browser never asks us — so we lose repeat-click analytics and, more importantly, we can't intervene if the destination becomes malicious or legal requires a takedown. With a 302 and a 5-minute edge cache, repeat visits cost a ~10–30ms edge hop instead of 0ms, we see every click, and revocation takes at most 5 minutes (seconds with purge). The cost is edge requests: ~$0.50–1 per million, so 10B/day is ~$5–10K/day. I'd let paying customers opt into 301 for specific links, with a warning that they can't be revoked from browsers."
Why this is L6:
- Frames 301 vs 302 as revocability and analytics, not just caching
- Quantifies both latency and cost
- Provides a scoped opt-in rather than a blanket rule
What L7 adds:
- Makes redirect type part of the takedown SLA policy: anything with legal takedown obligations cannot be 301
- Ties the edge cost to analytics pricing
Drill 4: The Store Is Down#
Prompt: "The link store is unavailable for 10 minutes. What happens?"
Staff Answer
"Redirects mostly keep working: 95–99% hit the edge or origin cache. For misses, serve stale: if the origin cache has an expired entry, serve it (links rarely change) — I'd keep origin cache entries with a soft TTL of 5 min and a hard TTL of 24h. True misses (cold links) fail with a 503 and a retry-after — maybe 1% of clicks. Creation fails closed: we can't guarantee uniqueness or scanning without the store, so we return 503. Status changes (blocks, deletes) queue and apply when the store returns; the edge purge can still be issued immediately for abuse blocks. Alerts: redirect.stale_served_rate and store.error_rate."
Why this is L6:
- Distinguishes read and write degraded modes
- Uses stale-serve with a hard ceiling
- Keeps the abuse block path working even with the store down
What L7 adds:
- Designs the cache layer as a deliberate availability tier with an SLO independent of the store
- Considers a multi-region read replica for the store vs cost
Drill 5: Viral Hot Key#
Prompt: "A single link gets 150K clicks per second. What breaks?"
Staff Answer
"Redirects: nothing, if the edge is set up correctly. Each PoP fetches once per TTL; with request coalescing and stale-while-revalidate, origin sees ~one request per PoP per TTL. Add TTL jitter so 300 PoPs don't expire together. Analytics: 150K events/sec on one key would be a hot partition if the stream is keyed by code — so pre-aggregate at the edge or in the first stream stage (per-PoP per-second counts), and key the raw stream by random partition with the rollup keyed by code. Abuse: a sudden spike is also a signal; the velocity trigger rescans the destination — benign here, but it's the same signal as a phishing blast."
Why this is L6:
- Handles the hot key on both the redirect and the analytics path
- Names coalescing, SWR and TTL jitter precisely
- Connects velocity to abuse detection
What L7 adds:
- Uses viral events as capacity test data; publishes a "max clicks/sec per link" platform guarantee
- Budgets CDN cost spikes and negotiates commit-based pricing
Drill 6: Multi-Tenant Custom Domains#
Prompt: "Enterprise customers want their own short domains, like brand.co/sale."
Staff Answer
"Three pieces. Domain verification: the customer adds a DNS record (TXT for ownership, CNAME or A to our edge). TLS: automated certificates via ACME for each domain, issued on verification and renewed at 30 days before expiry — at 50K custom domains, that's ~1,700 renewals a day and a hard dependency on the CA's rate limits, so I'd batch and monitor tls.cert_days_remaining. Namespace: codes are unique per domain, so the key becomes (domain, code); brand domains can use short custom aliases freely. Isolation: a brand domain's links are only created by that customer, so abuse risk is lower and reputation is theirs. Failure to watch: a customer lets their DNS lapse — the domain gets re-registered by someone else, and our records are now pointing at a domain we don't control; detect via periodic DNS re-verification."
Why this is L6:
- Concrete verification, TLS automation, and renewal math
- Keys by
(domain, code) - Anticipates the domain-lapse failure
What L7 adds:
- Prices custom domains to cover TLS and support load; defines the offboarding path when a customer leaves
- Owns the CA dependency as a vendor risk (multiple CAs)
Drill 7: Build vs Buy#
Prompt: "Should our company build its own shortener for marketing links?"
Staff Answer
"Buy, unless links are core to our product. A link-management vendor gives branded domains, analytics, QR codes and — critically — an abuse team for a few hundred to a few thousand dollars a month. Building means inheriting a 10-year permanence obligation and an abuse surface. Exceptions: we're a platform wrapping user-generated links (safety at click time is our responsibility), we need data residency the vendors can't meet, or click data is strategically sensitive. If we buy, we must own the domain — then we can migrate vendors by pointing DNS elsewhere and importing mappings. The mapping export must be in the contract."
Why this is L6:
- Crisp criteria for building
- Identifies the domain as the lock-in lever
- Contractual export of mappings as the exit path
What L7 adds:
- Frames the 10-year permanence cost as a liability on the balance sheet of the decision
- Standardizes one link vendor across marketing teams to avoid domain sprawl
Drill 8: Policy Change Without an Outage#
Prompt: "We want to start showing interstitials for all anonymous links. How do we roll it out?"
Staff Answer
"It changes visitor experience for millions of existing links, so stage it. Shadow first: compute the decision (would-show-interstitial) for all anonymous clicks for a week and measure — what fraction of clicks, which referrers, how many from our largest free users. Then apply to new anonymous links only, then to anonymous links created in the last 30 days, then all. Each step behind a flag with a 1-hour rollback. Measure interstitial click-through (legitimate users continue at 80–90%; phishing traffic drops sharply) and complaints. Communicate to free-tier users before applying to old links, with an upgrade path to verified accounts that skip it."
Why this is L6:
- Shadow mode to quantify impact
- Newest-first rollout where abuse concentrates
- Measures both safety and user impact
What L7 adds:
- Uses the data to set trust-tier policy company-wide
- Coordinates with product and growth, since interstitials affect the free funnel
Drill 9: Cost#
Prompt: "The shortener costs $400K/month. Where does it go and how do we cut it?"
Staff Answer
"Decompose: CDN requests (10B/day at ~$0.75/M ≈ $225K/month) usually dominate; click-stream and analytics storage second (2 TB/day raw, hot 30 days, rollups forever); link store and caches are small (tens of TB). Levers: edge cost — longer TTL for links from trusted tiers (e.g., 1 hour for verified business links: fewer origin fetches, same per-request CDN fee, though), negotiate committed-use CDN pricing (often 30–50% off list), or serve redirects from our own edge in the top regions. Analytics — keep raw events 30 days not 90, store rollups at minute granularity for 7 days then hourly, compress. Don't cut: scanning — it protects the domain."
Why this is L6:
- Knows CDN requests dominate, not storage
- Concrete retention and rollup levers
- Names what not to cut
What L7 adds:
- Unit economics: cost per million redirects vs revenue per customer; free-tier cost as customer-acquisition spend
- Considers owning edge infrastructure at scale vs vendor pricing
Drill 10: Multi-Region#
Prompt: "Make it global."
Staff Answer
"Reads are already global via the edge. For misses, replicate the link store to 3 regions (e.g., DynamoDB global tables or Cassandra multi-DC) — reads local, p99 < 10ms. Writes: random codes make multi-region writes safe — two regions creating the same code simultaneously is a ~0.3% × tiny-window event, but I'd still avoid it by giving each region a disjoint code prefix bit or by routing creation to a home region (creation latency doesn't matter much). The tricky part is status changes: a block must reach every region and every PoP fast — so the block path writes the store and issues a global edge purge directly, not waiting for replication. Analytics: regional ingestion, global rollups."
Why this is L6:
- Recognizes the edge already solves global reads
- Handles cross-region code collision explicitly
- Makes the safety path independent of replication lag
What L7 adds:
- Data residency for analytics (EU visitor data processed in EU)
- Evaluates whether multi-region store is needed at all given edge hit rates — often a cost that buys little
8. Deep Dive Scenarios#
Deep Dive 1: Peak Traffic — Super Bowl QR Code#
Context: A customer runs a QR code in a TV ad seen by 100M people. In the 30 seconds it's on screen, the link receives 2M clicks. Your dashboards show a spike in 5xx errors for about 20 seconds. The customer's CMO is on the phone.
Questions to Surface First:
- Where did the 5xx originate — edge, origin, or the destination site?
- Was the link cached at the edge before the ad aired?
- Did analytics keep up, and will the customer's dashboard be accurate?
- Was the customer's destination site itself overwhelmed (not our problem, but their perception)?
Typical L5 Approach: Scales the redirect service and Redis, adds capacity headroom. Misses that the link was created 10 minutes before airtime with a negative-cache entry from the customer's test before activation — PoPs served cached 404s then stampeded the origin.
Staff Approach: Traces: the customer tested the code before the link was activated; edges negative-cached the 404 for 60s; at airtime, the negative entries expired unevenly and PoPs stampeded the origin without coalescing on the 404→200 transition. Fix: creating or activating a link purges negative-cache entries; coalescing applies to negative entries too; pre-warm edges for scheduled campaigns.
Principal Approach: Creates a "scheduled high-traffic event" product feature: customers register big campaigns, which triggers pre-warming, a capacity review, and a war-room contact. Turns an incident into a premium offering and a capacity-planning input.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm the link resolves globally now; check destination health separately. |
| Triage | Per-PoP 5xx and origin fetch timeline; negative-cache hit rate for the code. |
| Quick fix | Purge negative entries on activation; coalesce all origin fetches. |
| Guardrails | Synthetic multi-PoP check on activation of any link flagged as campaign. |
| Post-mortem | Negative caching interaction with creation; pre-warm API. |
Metrics to Watch: edge.negative_cache_hits{code}, edge.origin_fetch_rate, redirect.5xx_rate{pop}, analytics.pipeline_lag_s.
Organizational Follow-up: Campaign registration workflow with customer success; incident credit policy.
Ownership Question: "Who talks to the CMO?" Staff answer: Customer success owns the conversation, with a written incident summary from the redirect platform team within 24 hours — including the 20-second impact window and click counts reconciled.
Key Takeaway: "Negative caching protects you from scanners and bites you on launches. Every activation must purge its 'not found'."
What clears the Staff bar:
- Finds the negative-cache interaction
- Separates our failure from the destination's
- Pre-warming as a product capability
Deep Dive 2: Silent Failure — The Scanner Stopped#
Context: Your async scanner's threat-intel API key expired 9 days ago. The scanner has been marking every link "clean" on API errors. Nobody noticed until a security researcher tweets a list of 2,000 phishing links on your domain.
Questions to Surface First:
- How many links were created in the 9-day window, and what share would normally be flagged?
- Is our domain already on any blocklist?
- Why did an error default to "clean"?
- What else in the pipeline fails open silently?
Typical L5 Approach: Renews the key, rescans the window's links, blocks what's found. Adds an alert for API errors. Fixes this failure mode, not the class.
Staff Approach: Rescans all 9 days of links immediately with priority on high-click ones; blocks and purges. Changes the scanner to treat errors as "unknown" — anonymous links in "unknown" get the interstitial. Adds a canary: a known-bad test URL is created every 5 minutes and must be blocked within 2 minutes, or the on-call is paged.
Principal Approach: Establishes that every safety control must have a positive-proof canary — not "no errors," but "caught the known-bad thing." Audits T&S, rate-limiting and takedown pipelines for fail-open-without-alert. Reviews secret expiry management across the org.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Renew key; rescan window prioritized by clicks; block and purge. |
| Triage | Estimate victims (clicks on subsequently-blocked links); check domain reputation. |
| Quick fix | Errors → "unknown" → interstitial for low-trust tiers. |
| Guardrails | Known-bad canary every 5 min; abuse.canary_time_to_block alert. |
| Post-mortem | Fail-open default; secret expiry; researcher disclosure response. |
Metrics to Watch: abuse.scanner_error_rate, abuse.canary_time_to_block_s, abuse.flag_rate (a flag rate dropping to zero is itself an alert), domain.reputation_status.
Organizational Follow-up: Security researcher acknowledgement; secrets expiry inventory.
Ownership Question: "Who owned knowing the scanner was working?" Staff answer: Trust & safety engineering. The gap was measuring the scanner's health by errors instead of by outcomes. The canary makes the outcome observable.
Key Takeaway: "A safety system that reports 'no threats found' when it's broken is worse than one that's visibly down. Measure detection with a canary, not with the absence of errors."
What clears the Staff bar:
- Fixes the fail-open default, not just the key
- Outcome-based canary
- Alert on detection rate dropping to zero
Deep Dive 3: Large Customer Onboarding#
Context: A global retailer wants to migrate 400M existing short links from a competitor onto their branded domain on your platform, with zero broken links, in 6 weeks.
Questions to Surface First:
- Do we get the full mapping export, including status (deleted/blocked) and expiry?
- Will the codes keep their exact values? Any conflicts with our code format?
- What's their click volume, and when will DNS cut over?
- Do they need historical analytics imported?
Typical L5 Approach: Bulk-imports the mappings and flips DNS. Doesn't scan 400M destinations (some are now malicious or dead) and doesn't plan for DNS propagation during which both platforms serve.
Staff Approach: Imports into
(brand_domain, code)namespace, preserving codes exactly; imports tombstones so deleted codes stay dead; bulk-scans destinations at a throttled rate before cutover (flagging malicious ones); pre-warms edges for the top 1M links by historical clicks; cuts DNS with a low TTL set a week earlier; runs a synthetic sample of 100K codes against both platforms and compares.
Principal Approach: Productizes migration ("bring your links") as a competitive capability — with a standard mapping format, verification tooling and a published zero-broken-link guarantee — and prices it. Also negotiates that our own mapping export exists for customers leaving us, because fairness in exit builds trust in entry.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Weeks 1–2 | Mapping export spec; import including tombstones; conflict report. |
| Weeks 3–4 | Throttled destination scan; flagged list reviewed with customer. |
| Week 5 | Lower DNS TTL; pre-warm edges; dual-platform comparison on 100K sample. |
| Week 6 | DNS cutover; monitor 404 rate on the brand domain (target: equal to pre-migration). |
Metrics to Watch: import.rows_processed, import.conflicts, brand.404_rate{domain}, migration.sample_mismatch.
Organizational Follow-up: Migration playbook reused for future customers.
Ownership Question: "Who decides what to do with the 30K destinations our scanner flags?" Staff answer: Trust & safety decides blocking policy; the customer is notified and can appeal. We don't import known-malicious destinations as active links, regardless of contract size.
Key Takeaway: "Importing links imports their liabilities. Scan before you serve."
What clears the Staff bar:
- Preserves codes and tombstones
- Scans imported destinations
- DNS cutover plan with verification
Deep Dive 4: Post-Mortem — Reused Code Incident#
Context: A customer finds that an old printed brochure's QR code now leads to an adult website. Investigation: 18 months ago, a cleanup job hard-deleted expired links to reclaim space; the random generator later re-issued that code to a new user.
Questions to Surface First:
- How many codes were hard-deleted, and how many have been reissued?
- Which reissued codes still receive traffic from old sources (referrers, QR scans)?
- Why did the cleanup job exist? Who approved it?
- What does our terms of service promise about expired links?
Typical L5 Approach: Blocks the reissued code, restores the old mapping, fixes the cleanup job. Doesn't find the other reissued codes.
Staff Approach: Reconstructs the set of hard-deleted codes from backups and logs; finds 1.2M hard-deleted, 14K reissued; for each reissued code, checks traffic for signals of old-source clicks (QR scans, old referrers). Works with legal on which new owners must be moved to new codes. Replaces the cleanup job with tombstoning — a tombstone is ~40 bytes; 1B tombstones is 40 GB, trivial versus the risk.
Principal Approach: Writes "codes are never reused" into the platform standard and the customer terms. Adds a design-review rule: any job that deletes data in a namespace promised to be permanent needs sign-off from the owning product and legal. Reviews other "permanent identifier" namespaces in the company (user handles, invoice numbers) for the same risk.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Block the reported code; notify the customer. |
| Triage | Reconstruct hard-deleted set; intersect with issued codes; rank by old-source traffic. |
| Fix | Tombstone-only deletion; generator checks tombstones (they're just rows). |
| Remediation | Migrate affected new owners to new codes with notice; restore old mappings as tombstones or redirects per original owner's wish. |
| Post-mortem | Storage-saving job with no product review; permanence not documented as an invariant. |
Metrics to Watch: create.tombstone_conflicts (should be nonzero — proves tombstones block reuse), links.hard_deletes (should be zero).
Organizational Follow-up: Invariant registry for permanent namespaces.
Ownership Question: "Who approved deleting the rows?" Staff answer: An infra engineer optimizing storage, with no review — because the invariant "codes are never reused" existed only in people's heads. The fix is writing it down where job authors will see it.
Key Takeaway: "A tombstone costs 40 bytes. A reused code costs a customer's trust. Never trade the second for the first."
What clears the Staff bar:
- Finds all affected codes, not just the reported one
- Quantifies tombstone cost to kill the storage argument
- Makes the invariant explicit
Deep Dive 5: Multi-Region Expansion — EU Data Residency#
Context: An EU regulator inquiry asks where visitor click data is processed. Today all analytics run in us-east, including IP-derived data for EU visitors. Leadership wants EU residency within 6 months.
Questions to Surface First:
- What personal data do we collect per click (IP, user agent, visitor hash)?
- Can we reduce what we collect rather than relocate it?
- Do customers in the EU need their dashboards served from the EU?
- What about historical data?
Typical L5 Approach: Stands up an EU analytics pipeline and routes EU clicks there. Correct but heavy; misses minimization.
Staff Approach: Minimizes first: truncate IPs at the edge (derive country, then drop the IP), replace visitor identifiers with a daily-rotating salted hash, and never store raw IPs centrally. Then region-route what remains: EU-edge click events go to an EU Kafka cluster and EU rollups; global customers see merged aggregates (aggregates without personal data can cross regions). Historical data: re-process and delete raw EU events older than the retention window.
Principal Approach: Sets a company-wide telemetry minimization standard — "derive at the edge, drop the identifier" — which makes most residency requirements trivial for all products, not just the shortener. Engages legal to define which aggregates are non-personal and can be global.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Months 0–1 | Data inventory per click; legal classification. |
| Months 1–2 | Edge minimization: IP → country/ASN, rotate visitor salt daily. |
| Months 2–4 | EU ingestion and rollups; merged aggregate API. |
| Months 4–6 | Historical re-processing; deletion of raw EU data in us-east; audit evidence. |
Metrics to Watch: analytics.raw_ip_fields_stored (target 0), analytics.eu_events_processed_outside_eu (target 0), analytics.pipeline_lag_s{region}.
Organizational Follow-up: Privacy review for any new click field.
Ownership Question: "Who owns proving residency to the regulator?" Staff answer: The privacy team owns the response; the analytics team owns the evidence — pipeline configs, data inventories and deletion logs — produced from the system, not assembled by hand.
Key Takeaway: "The cheapest data to relocate is the data you never stored. Minimize at the edge before you build a region."
What clears the Staff bar:
- Minimization before replication
- Aggregates vs personal data distinction
- Evidence generated by the system
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Deliver the base design (estimate, API, edge → cache → store) in under 10 minutes
- Choose a non-enumerable code scheme and state its collision math and growth plan
- Defend 302 as the default with the analytics, revocation and cost tradeoffs quantified
- Handle a viral hot key on both the redirect and the analytics paths
- Design abuse protection as a lifecycle, with time-to-block as the headline metric
- Build an analytics pipeline with bot filtering, provisional vs final counts, and privacy minimization
- Define deletion as tombstone + purge + verification + no reuse, including legal flows
- Speak to domain strategy, permanence promises and deprecation at the org level
The Bar for This Question#
Mid-level (L4): Produces the hash table: base62 encoding, a database, a cache, a redirect. Estimation mostly correct. Little discussion of what happens after creation.
Senior (L5): Clean design with reasonable capacity numbers, a KV store, Redis caching, and a counter-based or hash-based ID scheme. Mentions rate limiting on creation and a click counter. Works on the happy path. Gaps: enumerable codes, 301 by reflex, synchronous or naive click counting, abuse as a create-time check, and deletion as a row delete.
Staff+ (L6): Treats the lookup as solved and spends the time on the trust and lifecycle system: creator tiers, non-enumerable codes, 302 with edge caching and purge, hot-key handling, async analytics with bot filtering, abuse detection across the link's life with a time-to-block metric, and deletion that actually removes links from every cache while never reusing codes. Ownership across platform, trust & safety, analytics and legal is explicit. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 A URL Shortener Is a Trust & Safety Product With a Redirect Attached#
| Evidence | Detail |
|---|---|
| Open redirector | Attackers use trusted domains to disguise destinations |
| Domain-wide blast radius | One abuse wave can blocklist every customer's links |
| Destination drift | Clean links turn malicious after sharing |
The Staff position: Staff the abuse system like a core product, with its own metrics, on-call and authority to block without deploys.
Why this matters in interviews: It reframes the most over-practiced question in a way interviewers rarely hear.
10.2 301 Redirects Are a Trap for Almost Everyone#
| Evidence | Detail |
|---|---|
| Browser caching | Can't count, warn or revoke |
| Edge makes 302 cheap | ~10–30ms vs 0ms on repeat visits |
| Legal takedowns | Can't honor for cached browsers |
The Staff position: 302 by default; 301 as a documented opt-in.
Why this matters in interviews: It's the fastest way to show you think about the link's lifecycle, not just its first click.
10.3 Sequential IDs Are a Privacy Bug#
| Evidence | Detail |
|---|---|
| Enumeration | Walk the code space, collect every destination |
| "Unlisted" documents | Commonly shared via short links |
| Security research | Public research has shown large-scale scanning of shortener code spaces exposing private links |
The Staff position: Random or permuted codes; never raw counters for public links; and tell users short links are not secrets.
Why this matters in interviews: It's a concrete, verifiable reason to reject the "textbook" answer.
10.4 "Links Never Break" Is a Financial Commitment, Not a Technical One#
| Evidence | Detail |
|---|---|
| goo.gl | Google announced shutdown of existing links, then narrowed it to inactive ones |
| Ongoing cost | Serving, TLS, abuse handling for links nobody pays for |
| Printed media | Packaging and signage outlive product roadmaps |
The Staff position: Decide the permanence promise explicitly, fund it, and publish a deprecation policy (e.g., read-only archive mode with years of notice).
Why this matters in interviews: Turns a trivia question into a lifecycle and ownership discussion.
10.5 Analytics Should Report Fewer Clicks Than Customers Expect#
| Evidence | Detail |
|---|---|
| Bots and previews | 20–50%+ of raw hits in many channels |
| Email security scanners | Pre-click every link in enterprise mail |
| Customer trust | Inflated numbers get discovered and destroy credibility |
The Staff position: Filter aggressively, show raw alongside filtered, publish methodology, bill on filtered.
Why this matters in interviews: Shows you think about analytics as a product with a truth obligation.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
At Staff level, the shortener is a system with a trivial core and a hard perimeter. At Principal level, it's a portfolio of long-lived public promises: every printed link is an obligation to resolve for years, every domain is a reputation asset shared by all customers, every click record is personal data subject to law. The Principal questions are how many such promises the company should make, how they are funded, which domains carry which risks, who has authority to break a link, and what happens to all of it if the product is ever retired.
The Org-Level Fault Line#
One company link platform vs every team running its own shortener. Marketing, product, support and partner teams each want short links; each tends to buy a tool or build a script.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| One internal link platform with branded domains | One abuse posture, one domain portfolio, unified analytics | Platform team must serve varied needs | Platform headcount (~2–4 engineers) |
| Every team picks a vendor | Fast for each team | 6 vendors, 10 domains, inconsistent takedown, links die when contracts lapse | Company reputation; legal during takedowns |
| One vendor, centrally contracted, company-owned domains | Vendor handles abuse; company keeps exit | Vendor feature limits | Procurement + a small ownership function |
🧭 Principal Move: "Company-owned domains are non-negotiable, whoever serves the redirects. If the domain is ours, we can always move the links. If it's the vendor's, every printed link is hostage to a contract renewal."
Cost Model#
Assumptions: CDN at ~$0.50–1.00 per million requests (list; committed pricing lower), KV storage ~$0.25/GB-month with replication, stream processing and analytics storage at cloud list rates, fully loaded engineer ~$300K/year. Order-of-magnitude.
| Scale | Profile | Infra $/month | Headcount | On-call Load |
|---|---|---|---|---|
| Small (internal/company links) | 1M links, 50M clicks/month | $500–2K | 0.5 FTE | Low; business hours |
| Medium (SaaS product) | 1B links, 30B clicks/month | $40K–100K (CDN ~50%, analytics ~35%) | 6–10 engineers + 2–4 T&S analysts | Rotation; abuse spikes weekly |
| Large (public shortener / platform wrapper) | 50B links, 300B+ clicks/month | $300K–1M+ | 30–60 engineers + T&S team of 10–30 | 24×7; T&S follow-the-sun |
The surprise at every scale: trust & safety people cost as much as infrastructure once the service is public, and analytics storage — not links — is the largest data line.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost |
|---|---|---|
| Domain name(s) links are printed on | One-way | Printed media can't be recalled |
| Code alphabet and length for existing links | One-way | Every existing link depends on it |
| Code reuse policy | One-way | Once reused, trust is gone |
| Promise of permanence in terms of service | One-way | Breaking it is a public, reputational event |
| 301 for any link | One-way per browser | Cached indefinitely in visitors' browsers |
| Edge TTL, scanning thresholds | Two-way | Config change |
| KV store vendor | Two-way-ish | Mapping migration; days to weeks |
| Analytics retention periods | Two-way going shorter, one-way for data already deleted | — |
The Standard I'd Write#
RFC: Public Link Standard (v1)
Scope: Any service that issues short or redirect links on company-owned domains, including vendor-operated ones.
MUST:
- Links MUST be issued only on company-owned domains, with the mapping exportable at any time.
- Codes MUST be non-enumerable for publicly creatable links and MUST NEVER be reused; deletions are tombstones.
- Redirects MUST default to 302/307; 301 requires explicit opt-in and MUST NOT be used where takedown obligations apply.
- Every status change (block, delete, expire) MUST purge edge caches and be verified at a sample of PoPs within 10 minutes.
- Publicly creatable links MUST pass asynchronous destination scanning with a published time-to-block SLO (p95 < 15 min) and a known-bad canary.
- Click data MUST minimize personal data at the edge (no raw IP storage) and honor regional residency.
SHOULD: separate domains per trust tier; report raw and filtered clicks; register high-traffic campaigns for pre-warming.
Exceptions: Reviewed by platform, T&S and legal; time-limited to 12 months.
Success metrics: zero code reuses; takedown SLA met ≥ 99%; domain reputation incidents per year; time-to-block p95; number of company link domains (target: down, not up).
What I'd Tell the VP#
Short links look like a tiny feature, but every one we issue is a promise that a printed QR code or a shared link will keep working for years — and a risk that criminals will use our trusted domain to hide scams. I'd consolidate the company on one link platform on domains we own, so we can always move providers without breaking anything. The main investments are an abuse-response function and automatic safety scanning, because a single phishing wave can make browsers warn on every link we've ever issued, including our biggest customers'. Analytics is where the product value is, and we'll report conservative, bot-filtered numbers customers can trust. We should also decide now, in writing, what happens to these links if we ever retire the product.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Domains as assets | "Domain reputation is shared by every customer; I'd segment it by trust tier." |
| Permanence is funded | "A 10-year link promise is ~$X/year of serving and T&S for links that don't pay." |
| Exit design | "Company-owned domains and exportable mappings make vendors replaceable." |
| Minimization over relocation | "Drop the IP at the edge and most residency problems disappear." |
| Knows when not to standardize | "Internal go-links are a different, simpler product; don't force them onto the public platform." |
Staff answers that L7 interviewers find insufficient:
- A thorough abuse design with no view on domain portfolio or reputation blast radius
- Deletion handled perfectly for one service while five other teams issue links on vendor domains
- No position on what the company promises if the product is shut down
Appendices
Appendix A: Code Generation in Depth#
A.1 Random Codes With Conditional Insert#
ALPHABET = "23456789abcdefghjkmnpqrstuvwxyzABCDEFGHJKLMNPQRSTUVWXYZ" # print-safe variant (55 chars)
# use full base62 for non-print links
create(long_url, owner):
for attempt in 1..5:
code = random_string(ALPHABET, length=current_length()) # CSPRNG
if reserved(code) or offensive(code): continue
ok = store.put_if_absent(key=(domain, code),
value={url, owner, status: ACTIVE, created_at})
if ok: return code
metrics.inc("create.collision_retries")
raise Retryable()
Collision probability per insert = N / S (N = existing codes incl. tombstones, S = keyspace). At N = 10B, S = 62^7 ≈ 3.52T → 0.28%. At N = 50B → 1.4% — time to move new codes to 8 characters.
A.2 Counter + Keyed Permutation#
lease = counter_service.lease(block=1_000_000) # e.g. [5_000_000_000, 5_001_000_000)
next_id():
n = lease.next()
x = feistel_permute(n, key=K, bits=42) # bijection on [0, 2^42)
return base62(x).rjust(7, '0') # 2^42 ≈ 4.4T fits in 7 base62 chars
Unique by construction, no retries, not sequential to outsiders. Risk: if K leaks, codes become predictable — rotate by switching to a new key for a new code range.
A.3 Custom Aliases#
| Rule | Reason |
|---|---|
| Reserved words list (brand, system paths, profanity) | Impersonation and embarrassment |
| Minimum length on shared domains (≥ 8 for free tier) | Keep short space for random codes; limit squatting |
| Unlimited on customer's own branded domain | Their namespace |
| Trademark dispute workflow | Legal |
Appendix B: Data Model#
| Table | Key | Fields | Notes |
|---|---|---|---|
links | (domain, code) | long_url, owner_id, status, redirect_type, created_at, expires_at, trust_tier, scan_state | ~500 bytes; tombstones ~40 bytes |
links_by_owner | (owner_id, created_at) | code | Link management UI |
owner_url_dedup | (owner_id, url_hash) | code | Per-owner dedup only |
domains | domain | owner_id, verified_at, cert_state, trust_tier | Custom domains |
click_rollups | (domain, code, minute) | clicks_raw, clicks_filtered, hll_uniques, by_country | Minute → hourly after 7 days |
Appendix C: Caching Mechanisms — Quick Comparison#
| Layer | TTL | Hit Rate Contribution | Invalidation | Notes |
|---|---|---|---|---|
| Browser | 0 (302) or indefinite (301) | Only for 301 | None | Why 302 is default |
| Edge PoP | 5 min ± 20% jitter, SWR 60s | 90–97% | Purge API + verify | Coalescing mandatory |
| Edge negative | 60s | Absorbs scanners | Purge on create/activate | Launch pitfall |
| Origin in-process | 30s | Hot codes | Pub/sub invalidation | Protects Redis from hot keys |
| Redis | 24h soft / hard | 60–80% of edge misses | Delete on status change | Serve-stale on store outage |
| Link store | — | Remainder | Source of truth | Single-key reads 2–5ms |
Appendix D: API Contract and Client Behavior#
| Response | When | Headers |
|---|---|---|
302 Found | Active link | Location, Cache-Control: private, max-age=0; edge uses s-maxage=300 |
301 Moved Permanently | Opt-in links | Cache-Control: max-age=86400 (bounded, not infinite) |
200 interstitial | Flagged link | Cache-Control: no-store; "continue" link |
404 Not Found | Unknown code | Negative-cached 60s at edge |
410 Gone | Deleted / expired / blocked | Short cache; block page for abuse |
429 Too Many Requests | Creation limits | Retry-After |
Referrer handling: set Referrer-Policy on redirects deliberately — many shorteners strip or reduce referrer to avoid leaking the short link context to destinations.
Appendix E: Observability#
E.1 Core Metrics#
redirect.latency_ms{p50,p99,pop}
redirect.5xx_rate{pop}
edge.hit_rate / edge.origin_fetch_rate / edge.negative_cache_hits
cache.hot_key_qps{code}
create.rate{tier} / create.collision_retries
abuse.time_to_block_s{p50,p95}
abuse.canary_time_to_block_s
abuse.scan_queue_age_s / abuse.flag_rate
domain.reputation_status{provider}
purge.verification_failures
analytics.pipeline_lag_s / analytics.bot_ratio{code}
tls.cert_days_remaining{domain}
links.hard_deletes # must be 0
E.2 Critical Alerts#
| Alert | Threshold | Action |
|---|---|---|
| Redirect p99 | > 50ms for 5 min | Page platform |
| Domain reputation change | any provider flags a domain | Page T&S immediately |
| Canary not blocked | > 2 min | Page T&S |
| Flag rate zero | 0 flags for 1 hour on anonymous tier | Page T&S (scanner likely broken) |
| Purge verification failure | any takedown not edge-verified in 10 min | Page platform |
| Cert expiry | < 14 days on any custom domain | Page platform |
| Hard deletes | > 0 | Page; invariant violation |
E.3 Control Plane vs Data Plane#
The redirect data plane (edge + redirect service + caches) must keep serving if the control plane (creation, scanning, dashboards) is down. The one control-plane action that must stay available is block + purge — keep it on an independent path so an outage of creation or the main store never prevents stopping a live phishing link.
Appendix F: Scale Evolution#
F.1 What Works at Each Scale#
| Scale | Design |
|---|---|
| < 10M links | Single PostgreSQL + CDN; async click log to a warehouse |
| 10M – 10B links | Managed KV store, Redis, CDN edge logic, Kafka click stream, T&S scanning |
| > 10B links, platform wrapping | Edge-native redirect logic, regional stores, dedicated T&S org, custom edge where CDN cost dominates |
F.2 What You Don't Build on Day One#
- Your own edge network
- ML abuse models (start with threat feeds, rules and velocity triggers)
- Exact unique-visitor counting
- Multi-region write path for creation
- Link editing for free-tier users (bait-and-switch risk)
Appendix G: Multi-Tenancy, Fairness and Cost#
| Concern | Mechanism |
|---|---|
| Abuse blast radius | Domain per trust tier; account-level suspension across links |
| Noisy creator | Per-account and per-tier creation limits |
| Analytics cost per tenant | Retention and granularity by plan |
| Custom domain cost | TLS and support priced into enterprise plans |
| Free-tier cost | Treated as acquisition spend; capped by interstitials and limits |