Technologies referenced in this case study: Redis · Apache Kafka · Apache Flink · Cassandra · DynamoDB · OLAP Databases
Related: Ad Click Aggregator · Rate Limiter · Distributed Cache · Ledger and Digital Wallet · News Feed · Feature Flags · Hot Keys · Graceful Degradation · Latency & Protocols
Reading Guide#
Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal View) and the appendices on pacing control, auction math and the event pipeline |
What is an Ad Serving and Auction System? — Why interviewers pick this topic
An ad server decides, in the time it takes a page or a feed to render, which ad a specific person sees in a specific slot, and what the advertiser pays for it. A request arrives with a placement, a user (or a device, or nothing), a page context and a deadline. The system must find the few thousand ads that are eligible out of millions, estimate how likely this person is to click or convert on each, run an auction that converts bids and predictions into a ranking and a price, check that the winner still has budget and hasn't already been shown to this person five times today, and return the ad — all inside roughly 100 ms. Then it must record what was actually shown and clicked so the advertiser can be billed exactly.
The hard part is not ranking. The hard part is that money moves on every decision under a deadline you cannot extend: a budget that must not be blown by a factor of two when a campaign goes viral, an auction whose pricing rule changes what advertisers bid, a frequency counter read a million times a second, and an event pipeline whose duplicates become invoices.
Before vs After — the "viral campaign" scenario:
Without distributed budget control:
t=0: A $50K daily-budget campaign matches a trending topic. Win rate jumps 30×.
t=+1min: Ad servers check spend from a counter updated by a batch job every 5 minutes.
The campaign wins ~180K impressions/min at $9 CPM → $1,620/min.
t=+5min: Batch job lands: spend $8.1K. Counter still says "under budget" for the next window.
t=+31min: Spend crosses $50K. Batch lags; servers keep serving for another 5 minutes.
t=+36min: Campaign stops at $58K. $8K overdelivered. Advertiser contract: pay only up to budget.
The $8K is revenue the platform gave away — and impressions other advertisers lost.
With leased budget slices and pacing:
t=0: Same viral match. Each ad server holds a small budget lease (~$5) for this campaign.
t=+1min: Leases are consumed in seconds; servers ask the budget service for more.
Pacing controller sees spend running 30× ahead of plan → throttle to 4% participation.
t=+8h: Campaign spends evenly through the day, ends at $50,040.
Overdelivery bounded by outstanding leases: ~$40 across 8 regions.
Why interviewers reach for this question: It looks like "filter ads, rank by bid × CTR, return the top one" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: a hard latency budget split across a fan-out, an auction rule that is an economic contract with advertisers, budget enforcement that is a distributed counter problem with money on the line, frequency caps that read per-user state at enormous QPS, and a billing pipeline where served, rendered and charged are three different numbers.
Mechanics Refresher: Auction and Delivery Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| eCPM ranking | Normalize bids to expected value per 1,000 impressions: CPC bid × pCTR × 1000, CPA bid × pCVR × 1000, CPM bid as-is | Compares advertisers who pay on different events | Only as good as the prediction's calibration |
| Second-price auction | Winner pays just enough to beat the runner-up (adjusted by quality) | Truthful bidding is a sensible strategy in the simple case; stable prices | Opaque to buyers when the seller adds floors or multiple auction layers |
| First-price auction | Winner pays its bid | Transparent; simple to explain | Buyers shade bids; prices noisier; the burden moves to the bidder |
| Reserve price / floor | Minimum clearing price per slot | Protects inventory value | Too high → unfilled slots; must be tested, not guessed |
| Budget pacing | Spread a daily budget across the day by throttling participation or bid | Avoids spending everything at 9 a.m. | Controller tuning; lags on spikes |
| Frequency cap | Limit impressions per user per ad (or campaign) per window | Better user experience; less wasted spend | Per-user state at serving QPS |
| Impression beacon | Client fires a pixel/event when the ad actually renders | Billing on real exposure | Fires late, fires twice, sometimes never |
| Click redirect | Click passes through a tracking endpoint before the landing page | Reliable click capture | Adds a hop; must be fast and fraud-filtered |
For most production systems: a two-stage candidate funnel (cheap retrieval, then model ranking) under a hard internal deadline, eCPM ranking with a quality term, a clearly documented pricing rule (first-price is now common in programmatic; owned-and-operated platforms often keep a generalized second-price or a variant), pacing by probabilistic throttling fed by near-real-time spend, budget enforcement through leased slices with a bounded overdelivery, approximate frequency caps in a low-latency store that fail open, and billing on rendered impressions and deduplicated clicks. The primitives are not the interview — the latency budget, the budget-overspend bound and what you bill on are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
Ad serving is not a ranking question. Everyone can sort by bid × pCTR.
It is a money-under-a-deadline question that tests:
- Whether you split a fixed ~100 ms budget across stages and decide what to drop when a stage is slow
- Whether you know the auction rule is an economic contract — it changes how advertisers bid, not just who wins
- Whether you can bound budget overspend with a number, and say who absorbs it
- Whether you distinguish served, rendered and billable impressions, and dedupe the events that become invoices
The key insight: Every ad decision spends someone's money with incomplete, slightly stale information. A Senior design optimizes the ranking. A Staff design bounds the error: how stale the spend counter may be, how much a campaign may overdeliver, how many duplicate impressions reach billing, and what the system serves when the deadline hits. Revenue is a function of the ranking; trust is a function of the bounds.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Ad index → filter → rank by bid × CTR → return top ad | Asks "Is this our own inventory with our own advertisers, or an exchange with external bidders? What's the latency budget, and what pays — impressions, clicks or conversions?" | Asks "What's the marketplace's health metric — advertiser ROI, fill rate, user experience — and who arbitrates when they conflict?" |
| Latency | "Make each service fast" | "100 ms end to end; ~60 ms internal; per-stage deadlines; when ranking runs late, serve the best already-scored candidate rather than time out" | Treats the latency budget as a revenue curve — each 10 ms of serving latency has a measured fill and CTR cost |
| Auction | "Highest eCPM wins, pays its bid" | "Rank by eCPM with a quality term; pricing rule is documented; reserve prices tested by experiment; pricing changes shipped like contract changes" | Decides first- vs second-price at the marketplace level, prices the transition cost for advertisers and their tooling |
| Budget | "Check spend before serving" | "Leased budget slices per server, refreshed in seconds; pacing by throttling; overdelivery bounded to the sum of outstanding leases, written down" | Sets the overdelivery policy with finance: who eats it, how it's reported, how it's audited |
| Frequency caps | "Store counts in a database" | "Approximate counters in a low-latency store keyed by user and campaign, local caching, fail-open on store outage with a ceiling" | Treats caps as a user-experience policy owned by product, not a serving detail |
| Billing | "Count served ads" | "Bill on rendered impressions and deduplicated clicks; served ≠ rendered ≠ billable; idempotent event IDs; fraud filtering before invoicing" | Makes the billing pipeline auditable by advertisers and third-party measurement; reconciles discrepancies as a product surface |
Why "budget" separates levels
L5: "Before serving an ad, check the campaign's spend against its budget in the database." At 1M requests/s, with the top campaigns eligible for a large share of them, that's hundreds of thousands of reads per second on a handful of hot rows — and the spend number is already stale because impressions are counted asynchronously after rendering. The check either melts the store or reads a value minutes old.
L6: "Budget is a distributed counter with an explicit error bound. The budget service hands each ad server a small lease — a slice of remaining budget, sized to about 10 seconds of that server's expected spend. Servers decrement locally with no network call. When a lease runs low they request another; when the campaign is nearly exhausted, leases shrink. The worst-case overdelivery is the sum of outstanding leases plus in-flight impressions that render after the stop — I'd target under 1% of daily budget and alert above that."
L7: "Overdelivery is a finance policy. If advertisers are billed only up to budget, every overdelivered impression is revenue we gave away and inventory another advertiser didn't get. I'd agree the bound with finance, report it per campaign, and audit it monthly. Pacing quality — spending evenly to the end of the day — is a product promise sales makes."
Why "auction" separates levels
L5: "The highest bidder wins and pays what they bid." That's a first-price auction — which is a legitimate choice, but the candidate doesn't know it's a choice, doesn't account for predicted click rate when bids are per-click, and can't explain how advertisers will respond.
L6: "Advertisers bid on different events, so I rank by eCPM — bid × predicted action rate × 1000 — times a quality factor. Pricing: on our own inventory I'd use a generalized second-price rule: the winner pays the minimum needed to keep its rank, so advertisers aren't punished for bidding their true value. If we sell into programmatic exchanges, first-price is now the norm there, and bidders shade. Whichever we choose, it's documented, and reserve prices are set by experiment."
L7: "Changing the pricing rule changes every advertiser's optimal strategy. Moving from second- to first-price shifts work to the buyers — their bid shading tools — and creates a transition period where revenue can dip or spike. I'd stage it with holdouts, publish it months ahead and measure advertiser ROI, not just platform revenue."
Why "billing" separates levels
L5: "Each time we serve an ad, increment the impression count." But a served ad isn't seen: the page might close, the slot might be below the fold, the app might crash. Charging on serve overbills advertisers by whatever fraction never rendered — often a double-digit percentage on mobile.
L6: "Three numbers: served (we returned it), rendered (the client fired the impression beacon) and billable (rendered, deduplicated, not filtered as invalid traffic). Every served ad gets an impression ID signed into the beacon URL; billing counts unique IDs once. Clicks are deduplicated per impression and filtered for bots before they become charges."
L7: "Advertisers reconcile our numbers against third-party measurement. The discrepancy rate is a product metric. I'd make the billing pipeline auditable — event-level logs available to large advertisers — and agree a tolerance, so disputes are resolved with data rather than credits."
Positions to Commit To#
| Position | Rationale |
|---|---|
| Hard per-stage deadlines inside a ~100 ms budget; degrade, don't time out | A late ad is a lost impression; a slightly worse ad is still revenue |
| Two-stage funnel: cheap retrieval to ~1–5K, model ranking on ~100–500 | Scoring millions of ads per request with a heavy model is unaffordable |
| eCPM with a quality term; pricing rule documented and changed only with notice | Bids on clicks, conversions and impressions must be comparable; the rule is a contract |
| Budget enforced by leased slices; overdelivery bounded and alerted | No strong global counter survives serving QPS on hot campaigns |
| Pacing by throttling participation, fed by near-real-time spend | Spending the day's budget by 10 a.m. is a pacing failure even if it's within budget |
| Frequency caps approximate and fail-open with a ceiling | Showing an ad a fourth time is a UX miss; failing every request is a revenue outage |
| Bill on rendered, deduplicated, fraud-filtered events — never on serve | The invoice must survive an advertiser's audit |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Owned-and-operated ad serving (ads in our own feed, search or app) | We own users, inventory and advertisers; one internal auction; ~100 ms inside our page render | Retrieval + ranking + internal auction + pacing + caps + billing pipeline | Overspend; miscalibrated predictions silently cost revenue; latency hurts the host page | Overdelivery < 1% of budget; billable counts reconcile within ~0.5% of event logs |
| Ad exchange (we sell a publisher's slot to external bidders) | Bidders on the internet; OpenRTB-style requests with a strict tmax; we hold no advertiser budgets | Fan-out to bidders with timeouts, floor prices, first-price clearing, bid-response validation | Slow bidders eat the deadline; bad creatives; auction transparency disputes | Bidder timeout rate < ~5%; auction log reproducible for disputes |
| Demand-side bidder (we bid on others' exchanges for our advertisers) | We answer inside the exchange's tmax minus network time — often 10–30 ms of compute; huge request volume, most of which we decline | Fast no-bid filters, bid shading, budget and pacing on our side | Bidding on everything burns compute; overbidding burns budget | p99 response well under tmax; no-bid rate tuned to cost per bid |
🎯 Staff Move: "I'll design the owned-and-operated case — ads inside our own feed and search, our advertisers, our auction. That's where budget, pacing, auction rules and billing all live in one system. If we were an exchange, the core problem would be fanning out to external bidders under a deadline; if we were a bidder, it would be deciding what not to bid on. I'll mention where those differ."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Second-Price vs First-Price Pricing | Truthful-bidding incentives and stable prices, or transparency with bid shading pushed to buyers? |
| 2 | Budget Enforcement: Strong Counter vs Leased Slices vs Async | Exact spend at a hot global counter, bounded overspend with local leases, or cheap async counting with unbounded lag? |
| 3 | Funnel Depth vs the Latency Budget | Score more candidates with a bigger model, or keep a tail-latency margin and serve on time? |
| 4 | Frequency Caps: Exact vs Approximate, Fail-Open vs Fail-Closed | Per-user state at serving QPS — how correct must it be, and what happens when the store is down? |
| 5 | What You Bill On: Served, Rendered or Viewable | Simple server-side counting, or client-side events that are late, duplicated and fraud-prone but true? |
How Real Companies Built It#
Why this section belongs here: The auction rule, the request deadline and the budget contract are all things large platforms have documented publicly. Citing them shows you know these are deliberate product decisions, not implementation details.
Google Ad Manager — Moving the Whole Platform to First-Price Auctions#
In September 2019 Google announced it was completing the rollout of a unified first-price auction across Google Ad Manager, after tests that it described as showing a neutral to positive impact on publishers' total revenue and a more competitive market in which third-party buyers won an increased share of impressions (Google Ad Manager blog).
Staff insight: A pricing-rule change at that scale was announced, staged and justified with test results — because it changes every buyer's bidding strategy. In an interview, say: "The pricing rule is a marketplace contract. Changing it is a migration for every advertiser's bidding tools, so it ships with notice and holdouts, not a config flip."
OpenRTB — The Deadline and the Auction Type Are Fields in the Request#
The IAB Tech Lab's OpenRTB 2.6 specification puts both decisions in the bid request itself: tmax is the maximum time in milliseconds the exchange allows for bids to be received, including internet latency, and at is the auction type, where 1 means first price and 2 means "second price plus" (the default). A bidder signals no-bid with an empty HTTP 204 response, and the spec's example request carries "tmax": 120 (OpenRTB 2.6).
Staff insight: The industry standard treats the latency budget and the pricing rule as explicit per-request parameters, not implicit behavior. Do the same internally: every stage receives an absolute deadline, and the auction type is a documented property of the placement.
Google Ads — A Daily Budget That Is an Average, With a Monthly Ceiling#
Google Ads documents that a campaign might spend up to twice its average daily budget on a given day to take advantage of traffic fluctuations, and that over a month it will spend no more than 30.4 times the average daily budget (Google Ads Help).
Staff insight: "Budget" is a contract with a defined tolerance, not an exact per-request limit. That's exactly the framing a Staff candidate should bring: "I'd define the budget promise first — daily, monthly, with what tolerance — and then design enforcement to that bound. An exact per-request global counter is neither necessary nor affordable."
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "We rank by bid × pCTR" | "Your pCTR model is miscalibrated by 20% for one advertiser vertical. Who notices, and how?" | Prediction quality as a revenue and fairness risk |
| "We check the budget before serving" | "A campaign goes viral and wins 30× its usual rate. How much does it overspend?" | Distributed counters, bounded error |
| "Winner pays second price" | "Why not first price? What do advertisers do differently?" | Auction incentives, contract thinking |
| "We cap frequency at 3 per day" | "The cap store is down. Do you serve or not?" | Fail-open vs fail-closed with a ceiling |
| "Ranking takes 30 ms" | "p99 of the feature store just went from 5 ms to 60 ms. What does the user see?" | Per-stage deadlines, graceful degradation |
| "We count impressions" | "Served, rendered or viewed — which one do you bill, and how do you dedupe?" | Billing correctness |
| "Budgets reset at midnight" | "Whose midnight? What happens to load at 00:00?" | Time zones, synchronized-reset herds |
System Architecture Overview#
Reading the diagram: The gateway stamps an absolute deadline on every request. Retrieval narrows millions of ads to a couple of thousand using an in-memory targeting index and drops campaigns the pacing controller is currently throttling. Ranking scores the survivors in two stages. The auction converts bids and predictions into a winner and a price. The guard stage checks a local budget lease and the frequency counter — no synchronous call to a global budget. Served, rendered and click events flow through a log to a stream aggregator that dedupes by impression ID, filters invalid traffic and feeds both billing and the control loop. The metric that tells you money is safe is
campaign.overdelivery_pct; the one that tells you revenue is safe isserve.deadline_miss_rate.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Latency | "Make every service fast" | "~100 ms end to end, per-stage deadlines, serve the best scored candidate when a stage runs late." |
| Candidates | "Filter all ads, rank them" | "Retrieval to ~2K from an in-memory index, light model to ~300, heavy model on those." |
| Auction | "Highest bid wins" | "eCPM × quality; pricing rule documented; reserve prices set by experiment." |
| Budget | "Check spend in the DB" | "Leased slices per server, refreshed in seconds; overdelivery bounded by outstanding leases; < 1% target." |
| Pacing | "Stop when the budget is gone" | "Throttle participation probability to track a spend plan through the day." |
| Frequency caps | "Exact count per user" | "Approximate counters, TTL'd, fail-open with a ceiling on store outage." |
| Billing | "Count served ads" | "Rendered, deduplicated by impression ID, invalid traffic filtered, then billed." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| End-to-end ad latency budget | ~100 ms typical target | Everything else is subdivided from this |
OpenRTB example tmax | 120 ms (spec example; exchanges set their own) | External bidders get the deadline in the request |
| Internal serving budget | ~50–70 ms after network and client overhead (illustrative) | What retrieval, ranking and auction actually share |
| Ad requests | ~100K–1M/s at large consumer scale (illustrative) | Drives fan-out and hot-campaign load |
| Active ads in index | ~1–10M (illustrative) | Retrieval must be index-driven, not scan-driven |
| Candidates after retrieval | ~1–5K | Light model budget |
| Candidates for heavy ranking | ~100–500 | Heavy model budget: ~10–25 ms batched |
| Feature store read | p99 ~2–5 ms target | One slow dependency eats the ranking budget |
| Budget lease size | ~5–30 s of a server's expected spend | Sets the overdelivery bound |
| Pacing control interval | ~10–60 s | Faster reacts to spikes; slower is more stable |
| Spend feedback lag | seconds to ~1 min target | Pacing and leases depend on it |
| Google Ads daily budget tolerance | up to 2× average daily on a given day; ≤ 30.4× per month | A documented budget contract with tolerance |
| Frequency cap state | ~16–32 bytes per (user, campaign) counter | 500M users × 20 active counters ≈ 160–320 GB |
| Served vs rendered gap | commonly 5–20% on mobile (illustrative) | Billing on serve overbills by that much |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on targeting criteria, the campaign schema and the ML model architecture, then run out of time before the interviewer asks the questions that decide the level: "A campaign goes viral. How much does it overspend?" and "What do you bill on?" Compress the funnel to ~8 minutes and spend the rest on the latency budget, the auction rule, budget and pacing, frequency caps and the event pipeline.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"When a client renders an ad slot, it asks us for an ad. We pick eligible ads by targeting, predict how likely this user is to click or convert, run an auction to choose a winner and a price, enforce budgets, pacing and frequency caps, return the ad, and record impressions and clicks so advertisers are billed correctly."
Then the non-functional requirements, which is where the design lives:
"Four constraints drive everything. One: a hard deadline — about 100 ms end to end, so roughly 60 ms for our own work, and a late answer is worth nothing. Two: money moves on every decision, so budget overspend must be bounded with a number we can defend. Three: the auction rule is a contract advertisers bid against. Four: billing must be correct and auditable — served is not billed. I'll assume 500K ad requests per second at peak, 5 million active ads from 200K advertisers, 500 million daily users, and a mix of CPM, CPC and CPA bids."
Then name the underspecified parts:
"I'd confirm: is this our own inventory or an exchange? Do we bill on impressions rendered or viewable? What's the budget promise — exact daily, or average with a tolerance? I'll assume owned-and-operated inventory, billing on rendered impressions and clicks, and a daily budget with overdelivery bounded under 1%."
🎯 Staff Move: Saying "a late answer is worth nothing, and overspend needs a number" reframes the problem from "rank ads" to "spend other people's money correctly under a deadline" — that's the sentence that sets the level.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Campaign:
campaign_id,advertiser_id,daily_budget,budget_tz,bid_type(CPM/CPC/CPA),bid_amount,targeting,freq_cap(e.g. 3/day),state - Ad (creative):
ad_id,campaign_id,format,asset_refs,review_state - Ad request:
request_id,placement_id,user_key(or none),context(page, query, geo, device),deadline_ms - Impression:
impression_id(signed),request_id,ad_id,price,auction_type,served_at,rendered_at? - Click:
click_id,impression_id,ts,ip_hash,ua_hash
Serving API:
POST /v1/ads:select
{ placement_id, slots: 1, user_key, context: { geo, device, page, query }, deadline_ms: 80 }
→ 200 { ads: [ { ad_id, creative, price_micros, impression_url, click_url } ] }
→ 204 No ad (nothing eligible or deadline reached with no candidate)
GET /i?imp=<signed impression_id> (impression beacon, fires on render)
GET /c?imp=<signed impression_id>&dest=… (click redirect, 302 to landing page)
Advertiser API (control plane, not latency-critical):
POST /v1/campaigns { budget, bid, targeting, freq_cap }
PATCH /v1/campaigns/{id} { budget?, bid?, state? } → index refresh ≤ 30s
GET /v1/campaigns/{id}/delivery spend, impressions, pacing status
🎯 Staff Move: "The impression ID is generated at serve time and signed into the beacon and click URLs. Billing counts each signed ID once. A replayed or forged beacon either fails the signature or dedupes away."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk one request in 90 seconds:
- The gateway resolves the user and context, stamps
deadline = now + 80 ms, and calls retrieval. - Retrieval queries an in-memory inverted index over targeting attributes — geo, language, interests, placement — intersecting posting lists to ~2,000 eligible ads, dropping campaigns the pacing controller has throttled out of this request.
- Ranking fetches features in one batched call, runs a light model on 2,000, then a heavy model on the top 300 to predict pCTR and pCVR.
- The auction computes
eCPM = bid × p(action) × 1000 × quality, applies the placement's reserve price and computes the winner's price under the documented rule. - The guard checks the winner's frequency counter for this user and decrements the local budget lease. If either fails, it falls to the next candidate.
- We return the ad with a signed impression ID. The client fires the beacon on render; clicks pass through the redirect. Events stream to an aggregator that dedupes, filters invalid traffic, bills, and feeds spend back to pacing and budget within seconds.
🎯 Staff Move: Say out loud: "Nothing on the hot path makes a synchronous call to a global counter. Budget is a local lease, caps are a single low-latency read, and the index is in memory. That's how the 60 ms holds at p99." You've now spent ~8 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes this hard is that we're spending advertisers' money with stale information under a deadline. I'd like to go deep on the latency budget and degradation, the auction and pricing rule, budget enforcement and pacing, frequency caps, and the billing pipeline. Where would you like to start?"
If no preference: start with budget. It's the question with money attached, and it decides the level.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: Budget enforcement and pacing (7–8 min)
"Budget is two problems. Enforcement — don't exceed the budget by more than a bound — and pacing — spend it evenly across the day. For enforcement, the budget service owns remaining spend per campaign and hands each ad server a lease sized to ~10 seconds of that server's expected spend for that campaign. Servers decrement locally. When a lease is 70% used, they ask for more; when remaining budget is under, say, 5%, leases shrink to a few impressions each. For pacing, a controller compares actual spend to a plan curve shaped like the day's traffic and sets a participation probability per campaign every 30 seconds: ahead of plan → lower, behind → higher."
Quantify: "Worst-case overdelivery is outstanding leases plus impressions served but not yet rendered. 400 servers × $0.50 lease ≈ $200 on a $50K campaign — 0.4%. I'd alert per campaign above 1%."
Who pays: "If advertisers are billed only up to budget, overdelivery is free inventory — the platform pays. That's why the bound is a finance number, not just an engineering one."
Deep dive 2: The latency budget and degradation (6–7 min)
"80 ms deadline at the gateway. Retrieval gets 15, features and ranking 25, auction and guard 5, leaving slack for serialization and the network. Each stage receives the absolute deadline, not a relative timeout, so stages can't each assume they have their full share. If the heavy model hasn't returned by its cutoff, the auction runs on light-model scores. If retrieval is slow, we serve from a per-placement cache of recent winners. A late response is a lost impression — so the degradation ladder always ends in a valid ad or a 204, never a timeout."
Quantify: "Fan-out to 50 ranking shards means p99 of the request is dominated by the slowest shard. At 50 shards, a 1-in-100 slow shard hits ~40% of requests — so we hedge or cut off the tail rather than wait."
Deep dive 3: The auction and pricing rule (5–6 min)
"Rank by eCPM × quality, where quality folds in predicted user experience — hide rate, landing-page quality. Pricing on our own inventory: generalized second price — the winner pays the smallest bid that would still have kept its position, plus a cent, never below the reserve. That keeps advertisers from needing to shade. Reserve prices per placement come from experiments, not intuition. The rule and its changes are published; a change ships behind a holdout with advertiser-ROI guardrails."
Deep dive 4: Frequency caps (4–5 min)
"Counters keyed by (user, campaign, day) in a low-latency KV store, ~20 bytes each, TTL 26 hours. One batched read per request for the top candidates' campaigns. Increments happen on render events, not on serve, so the counter can lag by seconds — we accept ±1 impression. If the store is down, we fail open, but with a ceiling: we stop serving campaigns with caps of 1 per day, because those advertisers bought exactly one exposure."
Deep dive 5: Billing pipeline and ownership (3–4 min)
"Billable impression = rendered beacon with a valid signature, first occurrence of that impression ID, not filtered as invalid traffic, within 1 hour of serve. Clicks: first click per impression, fraud-filtered. The aggregator writes billable events to a ledger with the impression ID as the idempotency key, so replaying the event log never double-bills. The ads serving team owns serving latency and pacing quality; the billing team owns the ledger and invoices; the trust team owns invalid-traffic rules; the ML team owns model calibration with alerts on drift."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: every ad decision spends someone's money with slightly stale information under a hard deadline, so the design is about bounds. Per-stage deadlines with a degradation ladder; an eCPM auction with a documented pricing rule; budget enforced by leases with a written overdelivery bound and pacing by throttling; approximate frequency caps that fail open with a ceiling; and billing on rendered, deduplicated, filtered events."
The evolution closer:
"What I'd build later: auto-bidding — advertisers set a target cost per action and we bid for them, which makes pacing and bidding one control loop; budget leases across regions; and an advertiser-facing event-level log for audit. What I'd not build: a globally consistent spend counter on the hot path — I'd invest in shrinking feedback lag instead."
🎯 Staff Move: End on who absorbs the error. "Overdelivery is bounded at under 1% and the platform eats it; billing never charges for an unrendered or duplicated impression; and when the deadline hits, the user sees a slightly less relevant ad, not a blank slot."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Model obsession | 10 min on the CTR model architecture | "Light model then heavy model; calibration matters more than architecture" |
| Targeting schema | Enumerates every targeting attribute | "Inverted index over targeting attributes; intersect posting lists" |
| Exact budget | Synchronous DB check per request | Leases, a stated overdelivery bound |
| No deadline story | Assumes every stage returns | Absolute deadline propagated; degradation ladder |
| Billing on serve | "Count impressions when we return the ad" | Served ≠ rendered ≠ billable; dedupe by impression ID |
| No pricing rule | "Highest bid wins" | Names the rule and its effect on bidding |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Ad serving is the interview where every component has a price tag attached to its error. A stale spend counter is overspend. A miscalibrated model is mispriced inventory. A duplicated beacon is an overcharge. A slow feature store is a blank slot. Senior candidates design each component to work; Staff candidates design each component's error to be bounded, measured and owned — and can say who pays when the bound is exceeded.
It also has an unusual feedback structure. The serving path depends on numbers computed by the event pipeline — spend, frequency counts, click rates — and those numbers arrive late. The system is a control loop with lag: serve, observe, adjust. Treating it as a request/response service misses that most incidents are control-loop incidents — pacing that oscillates, budgets that react too slowly, models trained on yesterday's distribution.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"It's 9:05 a.m. A retailer's $50,000 daily campaign just matched a trending topic and its win rate jumped 30×. Walk me through, minute by minute, what your system does with its budget, what the advertiser sees, and how much it overspends by the end of the day."
A candidate who answers with lease drain (seconds), the pacing controller's next tick (30 s) dropping participation, the spend-feedback lag that bounds how far ahead it can get, the overdelivery bound (sum of leases, ~$200), the advertiser's dashboard showing "pacing limited", and the end-of-day result (budget spent evenly, ±0.5%) has run an ad system. A candidate who says "we check the budget before serving" has written a filter.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Owned-and-operated serving → bounded money under a deadline
- Constraint: our inventory, our advertisers, our auction; latency inside our own page render
- Strategy: index retrieval, two-stage ranking, eCPM auction, leases and pacing, caps, billing pipeline
- Failure mode: overspend, miscalibration, deadline misses, overbilling
- Who pays for imperfection: the platform (overdelivery, credits), advertisers (overbilling until disputed), users (bad or repetitive ads), host-page teams (latency)
Ad exchange → fan-out under someone else's deadline
- Constraint: external bidders across the internet; we hold no budgets;
tmaxis a hard wall - Strategy: per-bidder timeouts and QPS throttles, floor prices, first-price clearing, creative validation, reproducible auction logs
- Failure mode: slow bidders consume the deadline; bidders sent traffic they never bid on waste everyone's money
- Who pays: publishers (unfilled slots), bidders (wasted compute), the exchange (disputes)
Demand-side bidder → deciding what not to bid on
- Constraint: answer inside
tmaxminus round-trip network time; most requests should be declined cheaply - Strategy: pre-filters before any model; bid shading for first-price; pacing and budget on our side
- Failure mode: bidding on everything burns compute and budget; slow responses count as no-bid
- Who pays: our advertisers (overpaying without shading), our infra budget
2.2 When NOT to Build an Ad Serving System#
- You have a handful of direct-sold sponsorships. A scheduling table and a rotation rule — "sponsor A on the homepage this week" — is enough. Auctions earn their complexity with many competing advertisers.
- You want to monetize inventory but advertising is not your business. Join an ad network or exchange and let them run demand, auctions and billing. Your engineering problem becomes an SDK integration and a revenue report.
- You need contextual placements only, no user data. A retrieval-plus-rules system without user-level frequency state and with simpler billing may be enough; the per-user state machinery is the expensive part.
- Your volume is tiny. At 100 requests/s, a database-backed budget check is fine. The lease machinery earns its keep when hot campaigns see thousands of requests per second.
🎯 Staff Insight: "An ad server is a marketplace with a deadline. If there's no marketplace — few advertisers, no competition for slots — I'd use a schedule, not an auction."
2.3 What the Interviewer Leaves Underspecified#
Interviewers deliberately omit:
- Owned inventory vs exchange — the systems barely overlap
- What the advertiser pays for — impressions, clicks or conversions change ranking, billing and fraud exposure
- The budget promise — exact daily cap vs average daily with monthly ceiling
- What counts as an impression — served, rendered or viewable
- Who owns pricing changes — engineering, ads product or finance
- Time zones — budgets "per day" in the advertiser's zone create synchronized resets
Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.
2.4 Precise Terminology#
| Term | What It Means | Why It Matters in the Interview |
|---|---|---|
| eCPM | Expected revenue per 1,000 impressions | Normalizes CPM, CPC and CPA bids for ranking |
| pCTR / pCVR | Predicted click / conversion probability | Mispredictions misprice inventory |
| Calibration | Predicted rates match observed rates on average | A ranking can be right and the prices still wrong |
| Reserve price (floor) | Minimum clearing price for a slot | Revenue lever; set by experiment |
| Second price (generalized) | Winner pays what it needed to keep its rank | Encourages truthful bids |
| First price | Winner pays its bid | Transparent; buyers shade |
| Bid shading | Bidder lowers bids below value in first-price auctions | Shifts optimization work to buyers |
| Pacing | Spreading spend across the budget period | Prevents front-loaded spend |
| Lease (budget slice) | Pre-allocated spend a server may use without asking | Bounds overspend without a hot counter |
| Overdelivery | Spend or impressions beyond the budget | Usually free inventory — the platform pays |
| Fill rate | Fraction of requests that return an ad | Falls when deadlines are missed |
| Invalid traffic | Bot or fraudulent impressions and clicks | Must be filtered before billing |
🎯 Staff Insight: If the interviewer says "enforce the budget exactly", ask: "Exactly per request, or exactly on the invoice? We can always bill no more than the budget. Serving no more than the budget, per request, at this QPS, needs a global lock on hot campaigns. I'd bound serving overspend and cap the invoice."
3. Where the Design Splits#
Every ad-serving decision has a technical side (index layout, counter consistency, timeouts) and an economic side (what advertisers pay, what the platform gives away, what users tolerate). Interviewers grade the second side.
3.1 Fault Line 1: Second-Price vs First-Price Pricing#
The tension: In a second-price rule the winner pays roughly what was needed to beat the next bidder, which lets advertisers bid their true value. In first-price the winner pays its bid, which is easier to explain and audit but pushes bidders to shade, making bids noisier and the bidder's tooling more complex.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Generalized second price (with quality) | Advertisers can bid value; prices stable; familiar on owned inventory | Clearing price is opaque to buyers; floors and layered auctions muddy it | Advertisers (trust in an opaque price) |
| First price | Transparent: you pay what you bid; matches programmatic norms | Buyers must shade; small advertisers without tooling overpay | Small advertisers (overpaying), buyers (tooling) |
| First price + platform auto-bidding | Advertisers set targets; platform shades for them | Platform is both auctioneer and bidder's agent — trust issue | Platform (credibility) |
| Fixed price (direct deals) | Predictable for both sides | No price discovery; unsold inventory | Publisher (revenue left on the table) |
Staff default: "On owned inventory with many small advertisers, generalized second price with a quality term and experiment-set reserves. It's forgiving to advertisers without bidding tools. If we sell into programmatic exchanges, we follow first price there and offer auto-bidding to advertisers who don't want to shade. Either way the rule is written down per placement."
When to deviate:
- Selling through exchanges: first price is the norm; fighting it costs participation.
- Mostly auto-bid advertisers: the pricing rule matters less because the platform bids on their behalf; focus on target-cost accuracy.
- Guaranteed contracts: fixed-price reserved delivery, served ahead of the auction with its own pacing.
🧭 Principal Move: "A pricing-rule change rewrites every advertiser's strategy. I'd stage it per placement with a holdout, publish it a quarter ahead, and judge it on advertiser ROI and retention — a revenue bump that churns small advertisers isn't a win."
❌ Common L5 Trap: "Second price is always better because it's truthful." Truthfulness holds in the simplest single-slot setting; with multiple slots, quality scores, reserves and auction layers, it isn't strictly true, and the industry largely moved programmatic to first price for transparency. The Staff answer picks a rule for a reason and names who it favors.
3.2 Fault Line 2: Budget Enforcement — Strong Counter vs Leased Slices vs Async#
The tension: An exact budget needs a single, strongly consistent counter per campaign touched on every win. A hot campaign can win tens of thousands of auctions per second, and spend is only known after render. Strong counters melt; async counters lag by minutes.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Strong central counter per win | Exact serving-time spend | Hot-row contention; a network call on the hot path; still blind to render lag | Users (latency), platform (store cost) |
| Leased slices to servers | No hot-path network call; bounded overspend = outstanding leases | Lease sizing and reclaim; unused leases strand budget near the end | Platform (bounded overdelivery) |
| Async counting, periodic check | Cheapest; simplest | Overspend unbounded during spikes — minutes of lag × spike rate | Platform (overdelivery), finance (disputes) |
| Bill-capped only, serve freely | Advertiser never pays over budget | Platform gives away unbounded inventory | Platform + other advertisers (lost slots) |
Staff default: "Leases. The budget service holds the authoritative remaining budget per campaign and grants leases of roughly 10 seconds of a server's expected spend. Servers spend locally. Near exhaustion, lease size shrinks so the tail of the budget is spent in small pieces. Leases expire after 60 seconds and unused amounts return. Overdelivery bound = sum of outstanding leases + impressions in flight between serve and render; I'd report it per campaign and alert above 1%."
When to deviate:
- Low-QPS campaigns: a strong counter is fine; leases are for hot ones. A hybrid routes by campaign QPS.
- Guaranteed delivery contracts: underdelivery is the risk, not overspend — pacing matters more than enforcement.
🧭 Principal Move: "The overdelivery bound is a line item. At our volume, 0.5% overdelivery on $10M daily spend is $50K a day of free inventory. I'd show finance that number next to the cost of tightening it — smaller leases mean more budget-service QPS — and let them choose."
❌ Common L5 Trap: "Use Redis INCRBY on the campaign's spend key for every impression and stop when it crosses the budget." Correct-looking, but it puts a network round trip on the hot path, turns the top campaigns into hot keys at tens of thousands of writes per second, and still counts serves rather than renders. See Hot Keys.
3.3 Fault Line 3: Funnel Depth vs the Latency Budget#
The tension: Scoring more candidates with a heavier model raises relevance and revenue. Every extra millisecond eats a fixed budget, and fan-out amplifies tail latency. Past some point, more scoring loses more impressions to deadline misses than it gains in relevance.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Heavy model on all retrieved candidates | Best relevance | Blows the budget; GPU/CPU cost scales with candidates | Host page (latency), infra |
| Two-stage: light on ~2K, heavy on ~300 | Most of the relevance at a fraction of the cost | Light-model misses can't be recovered | ML team (two models to maintain) |
| Fixed depth, wait for all shards | Deterministic | Tail latency from the slowest shard | Fill rate |
| Deadline-aware: score until cutoff, use what's ready | Never misses the deadline | Variable quality under load | Revenue (slightly lower eCPM under stress) |
Staff default: "Two-stage ranking under absolute deadlines. Each stage knows the time remaining. Heavy-model scoring stops at its cutoff and the auction uses light-model scores for the rest. Retrieval shards are hedged — send to a second replica after the shard's p95 — so one slow shard doesn't set request latency."
When to deviate:
- Search ads with few slots and high value: spend more of the budget on heavy scoring; each slot is worth more.
- Video pre-roll with seconds of buffer: the budget is looser; deeper ranking is affordable.
🧭 Principal Move: "I'd measure the revenue curve: eCPM gain per extra millisecond of ranking against fill loss per millisecond of p99. The optimal depth is where those cross — and it moves every time the model changes."
❌ Common L5 Trap: "Set a 100 ms timeout on each service call." Sequential stages each with their own 100 ms timeout can take 400 ms. Without an absolute deadline propagated through the request, timeouts compound and the client gives up before you answer.
3.4 Fault Line 4: Frequency Caps — Exact vs Approximate, Fail-Open vs Fail-Closed#
The tension: Caps require per-user, per-campaign state read on the hot path for every candidate that might win. Exactness requires synchronous increments on serve; availability requires deciding what to do when the store is slow or down.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Exact: increment on serve, synchronously | Never exceeds the cap | Write on the hot path; counts unrendered serves | Users (latency), advertisers (undercounted reach) |
| Approximate: increment on render, read on serve | No hot-path write; counts real exposures | Can exceed cap by in-flight impressions (±1–2) | Advertisers (occasional extra exposure) |
| Client-side counters | No server state | Lost on new device, cleared storage, users without IDs | Advertisers (weak caps) |
| Fail-closed on store outage | Caps never violated | Revenue outage for every capped campaign | Platform (revenue) |
| Fail-open with a ceiling | Revenue continues | Some users see an ad more than intended | Users (repetition), bounded |
Staff default: "Approximate counters, keyed by (user, campaign, day) in a replicated low-latency store, incremented from render events, read in one batch per request for the top ~20 campaigns, with a 1-second local cache per user on each server. On store timeout (5 ms), fail open — but exclude campaigns with a cap of 1, and add a per-server ceiling so no user sees the same campaign more than 5 times from that server during the outage."
When to deviate:
- Regulated categories (e.g., alcohol, gambling) with legal exposure limits: fail closed for those campaigns.
- Logged-out traffic without stable IDs: caps are best-effort; say so in the advertiser docs.
🎯 Staff Insight: "A cap is a user-experience promise and a reach promise. Product owns the policy for what happens on outage; I own making that policy cheap to execute."
❌ Common L5 Trap: "Store the count in the user's database row and increment it when serving." That's a write on every serve to a store sized for profile reads, counting serves that never rendered — and when it's slow, every ad request is slow.
3.5 Fault Line 5: What You Bill On — Served, Rendered or Viewable#
The tension: Server-side counts are complete and immediate but overcount exposure. Client-side events are true exposure but late, lossy, duplicable and forgeable.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Bill on serve | Simple; complete; immediate | Overbills by unrendered rate (5–20% on mobile, illustrative) | Advertisers (overbilled), then platform (disputes) |
| Bill on rendered beacon | Real exposure; industry-standard for many formats | Beacons late, duplicated, sometimes lost; needs dedupe and fraud filtering | Platform (pipeline complexity) |
| Bill on viewable (e.g., in view for N seconds) | Strongest exposure signal | More client logic; measurement disputes | Platform (measurement), publishers (lower billable volume) |
| Bill on click/conversion only | Advertiser pays for outcomes | Fraud pressure concentrates on clicks; conversion attribution lag | Trust team (fraud), advertisers (attribution disputes) |
Staff default: "Bill on rendered impressions for display and feed formats, on valid first clicks for CPC. Each served ad carries a signed impression ID; the aggregator accepts the first valid beacon per ID within an hour of serve, runs invalid-traffic filters, and writes to the billing ledger keyed by impression ID. Replays of the event log are idempotent."
When to deviate:
- Brand campaigns sold on viewability: add viewability measurement and bill on it.
- Performance campaigns: bill on conversions with a defined attribution window, accepting days of lag before final spend.
🧭 Principal Move: "Billing is the contract advertisers audit. I'd publish the definition of a billable impression, give large advertisers event-level logs, and track the discrepancy rate against third-party measurement as a product metric."
❌ Common L5 Trap: "We log each served ad and charge from that log." It's complete and simple, and it charges for ads nobody saw. The first large advertiser to compare against their own measurement opens a dispute — and the platform refunds the gap.
4. When It Breaks#
4.1 The Budget Blowout — Feedback Lag Meets a Viral Campaign#
t=0: Campaign C ($40K/day) matches a trending news topic. Win rate 25× baseline.
t=+30s: Spend aggregator falls behind: Kafka consumer lag 0s → 6 min (partition skew,
C's events all hash to one partition).
t=+1min: Budget service grants leases from stale spend. Pacing controller sees
spend "on plan" because its input is 6 minutes old.
t=+12min: Real spend $31K. Reported spend $9K.
t=+18min: Lag recovers. Budget service sees $46K spent. Revokes leases.
t=+19min: Final spend $46.8K. Overdelivery 17% — $6.8K of free inventory.
Detection: spend.feedback_lag_s (alert > 60 s); campaign.overdelivery_pct (alert > 1%); aggregator.partition_lag_max.
Mitigation: when feedback lag exceeds a threshold, the budget service shrinks all leases by a factor proportional to lag and pacing caps participation for campaigns whose serve-side win rate jumped — the serving tier knows its own win counts even when billing data is late.
Prevention: partition spend events by impression ID, not campaign, then aggregate per campaign in a second stage; budget leases fed by serve-side spend estimates (price × expected render rate) so enforcement doesn't depend on the slow billing path.
Owner: ads serving on-call (budget service), data platform (aggregator).
4.2 Silent Miscalibration — The Model That Cost 6% of Revenue#
t=0: New pCTR model ships after an offline AUC gain of +0.4%.
t=+1 day: Ranking order looks fine. But predicted CTR for one vertical is 1.3× observed.
t=+3 days: CPC advertisers in that vertical win more auctions, pay more per impression,
click rates below prediction → their cost per click is fine, but their win
rate crowds out CPM advertisers paying real money.
t=+9 days: Finance flags revenue per 1,000 requests down 6% in that vertical.
Detection: model.calibration_ratio{vertical} = sum(predicted clicks) / sum(observed clicks), alert outside 0.9–1.1; revenue per 1,000 requests by vertical against a holdout.
Mitigation: roll back to the prior model; apply a per-segment calibration layer that rescales predictions to match observed rates daily.
Prevention: online launch criteria include calibration by segment and revenue against a holdout, not just offline AUC. A ranking metric can improve while the prices it implies get worse.
Owner: ads ML team (calibration); ads serving (holdout infrastructure).
4.3 Frequency Store Outage — Fail Open or Go Dark?#
t=0: Frequency store cluster in one region loses a third of its nodes.
t=+5s: Cap reads p99: 2ms → 400ms. Every ad request blocks on the batched read.
t=+20s: serve.deadline_miss_rate: 0.3% → 61%. Fill rate collapses in that region.
t=+3min: On-call flips cap checks to fail-open via feature flag. Fill recovers.
t=+40min: Users report seeing the same ad 11 times in a session.
Detection: freqstore.read_p99, serve.deadline_miss_rate, serve.capcheck_failopen_total.
Mitigation: cap reads get their own 5 ms deadline and fail open automatically; a per-server, per-user in-memory ceiling (5 impressions per campaign per hour) applies during fail-open; campaigns with cap = 1 are excluded.
Prevention: the fail-open policy is decided in advance by product and coded as the default, not improvised during the incident; chaos tests kill a third of the store monthly.
Owner: ads serving on-call; product owns the fail-open policy.
4.4 The Slow Dependency — Feature Store Tail Eats the Budget#
t=0: A feature store compaction job runs during peak. p99 5ms → 70ms.
t=+10s: Ranking waits for features; heavy model starts at t=+75ms of an 80ms budget.
t=+30s: Deadline misses 0.3% → 28%. Clients render blank slots or house ads.
t=+2min: Revenue per minute down 25%.
Detection: featurestore.read_p99, rank.stage_budget_exceeded_total, serve.fill_rate.
Mitigation: feature reads get a hard 8 ms cutoff; missing features use defaults so the light model still runs; heavy model skipped if under 10 ms remain.
Prevention: compaction scheduled off-peak with rate limits; hedged reads to a second replica after p95; the degradation ladder tested by injecting latency weekly.
Owner: feature store team (compaction), ads serving (degradation ladder).
4.5 Double Billing After a Pipeline Replay#
t=0: Aggregator bug drops 2 hours of click events. Fix deployed.
t=+1h: Data engineer replays the event log from 2 hours back to recover clicks.
t=+1h: Aggregator's dedupe state only holds 30 minutes of impression IDs.
t=+2h: 1.5 hours of impressions are billed twice. 4,100 advertisers overcharged.
t=+3 days: Advertisers' third-party numbers disagree; disputes open.
Detection: billable impressions per hour vs served-and-rendered from the raw log (should never exceed); billing.duplicate_rejections drops to zero during a replay (suspicious).
Mitigation: recompute billing for the window from raw events with full dedupe; issue credits.
Prevention: the billing ledger enforces uniqueness on impression ID itself — an idempotency key at the sink — so replays are safe regardless of the stream processor's dedupe window. See Idempotency & Exactly-Once.
Owner: billing team (ledger), data platform (replay runbook).
4.6 The Midnight Herd — Budgets Reset at Once#
At 00:00 in a major time zone, thousands of campaigns that exhausted yesterday's budget reset simultaneously. Every ad server requests fresh leases in the same second; the budget service sees 400 servers × 30K campaigns of lease requests, and its p99 jumps to seconds. Campaigns serve nothing for the first few minutes of the day, then pacing — seeing them behind plan — overcorrects and spends aggressively.
Detection: budget.lease_request_rate spike at time-zone boundaries; campaign.zero_serve_minutes_after_reset.
Prevention: pre-issue first leases for the new day a few minutes before reset with a "valid from" timestamp; jitter lease refreshes; pacing plans start from the first minute, not from the first observed spend.
Owner: ads serving (budget service).
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Budget blowout | spend.feedback_lag_s > 60, campaign.overdelivery_pct > 1 | Hot campaigns; platform revenue | Shrink leases on lag; serve-side spend estimates | Ads serving |
| Miscalibration | model.calibration_ratio outside 0.9–1.1 | A vertical's pricing | Rollback; calibration layer | Ads ML |
| Cap store outage | freqstore.read_p99, deadline misses | Region's fill rate or UX | Auto fail-open with ceiling | Ads serving + product |
| Slow dependency | rank.stage_budget_exceeded_total | Fill rate | Hard cutoffs, defaults, hedging | Ads serving + feature store |
| Double billing | Billable > rendered unique | Advertiser invoices | Ledger uniqueness on impression ID | Billing |
| Midnight herd | Lease request spike at reset | Campaigns in that zone | Pre-issued leases, jitter | Ads serving |
| Click fraud burst | Click rate per IP/device anomaly | CPC advertisers' spend | Invalid-traffic filters before billing | Trust & safety |
| Index staleness | index.refresh_age_s > 120 | Paused campaigns still serving | Alert; push pause events directly | Ads serving |
🎯 Staff Insight: The two numbers that matter are on different sides of the ledger. "
serve.deadline_miss_ratetells me we're losing revenue;campaign.overdelivery_pcttells me we're giving it away. I'd page on both, and neither is visible in average latency."
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | "Pick the best ad for the user" | Names owned inventory vs exchange vs bidder; frames it as money under a deadline | Frames the marketplace's health metrics and who arbitrates among advertisers, users and revenue |
| Latency | Fast services, per-call timeouts | Absolute deadline propagated, per-stage budgets, degradation ladder, hedged fan-out | Prices latency as a revenue curve; sets the org's serving SLO with host-page teams |
| Auction | Highest bid wins | eCPM × quality; names the pricing rule and its incentive effect; reserves by experiment | Owns pricing-rule changes as marketplace migrations with advertiser-ROI guardrails |
| Budget & pacing | Check spend before serving | Leases with a stated overdelivery bound; throttling-based pacing; serve-side spend estimates | Sets overdelivery policy with finance; reports and audits it |
| Frequency caps | Exact counts in a DB | Approximate render-fed counters; fail-open with ceiling; per-category exceptions | Treats caps as a UX policy owned by product with measured user impact |
| Billing | Count served ads | Rendered + deduped + filtered; ledger idempotent on impression ID | Auditable billing; discrepancy rate as a product KPI |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Bounds the money error | "Overdelivery is at most the sum of outstanding leases — about 0.4% here — and I'd alert per campaign above 1%." |
| Deadline-first design | "Every stage gets the absolute deadline; when ranking runs late, the auction uses light-model scores." |
| Knows pricing is a contract | "Changing to first price changes every advertiser's bidding strategy; it ships with notice and holdouts." |
| Separates served from billable | "Served, rendered, billable — three numbers; we bill the third." |
| Fail-open with a ceiling | "Cap store down: serve, but exclude cap-of-one campaigns and enforce a local ceiling." |
| Calibration over AUC | "A model can rank better and price worse; I'd gate launches on calibration by segment." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Synchronous global spend counter per request | Hot keys, latency, and still blind to render lag |
| "Highest bid wins" with mixed bid types | Ignores eCPM normalization and predicted action rates |
| No deadline propagation | Compounding timeouts; blank slots under load |
| Billing on serve | Overbills; first audit becomes a dispute |
| Exact frequency caps on the hot path | Write per serve; outage takes down serving |
| No feedback-lag awareness | Pacing and budgets act on stale data during the spikes that matter |
5.4 Common False Positives#
- Deep ML knowledge ≠ ad serving design. A beautiful ranking model doesn't bound overspend or dedupe beacons.
- Auction theory ≠ marketplace judgment. Reciting incentive properties without naming who the rule favors and how to change it is incomplete.
- "We use Redis" ≠ budget enforcement. A fast counter is still a hot key and still counts the wrong event.
- Low average latency ≠ meeting the deadline. The fill rate lives at p99 across a fan-out.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Owned inventory; deadline, bounded money, pricing contract, billing correctness; numbers |
| Entities & API | 3–5 min | Campaign, ad, request, impression, click; signed impression ID; 204 for no ad |
| Architecture | 5–10 min | ≤ 8 boxes; retrieval → ranking → auction → guard; event log → aggregator |
| Budget & pacing | 10–18 min | Leases, overdelivery bound, throttling-based pacing, feedback lag |
| Latency & degradation | 18–24 min | Absolute deadlines, stage budgets, fallback ladder, hedging |
| Auction | 24–30 min | eCPM × quality, pricing rule, reserves, change management |
| Caps + billing | 30–38 min | Approximate caps, fail-open with ceiling; rendered, deduped, filtered billing |
| Pivot / wrap | 38–45 min | Multi-region budgets, auto-bidding, cost; close on who absorbs the error |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Make the budget exact" | Do you know the cost? | Exact invoice is easy; exact serving needs a hot lock; propose a bound |
| "Add external bidders" | Exchange mechanics | Per-bidder timeouts within tmax, first-price, bidder throttling |
| "A model regresses revenue silently" | Calibration and holdouts | Calibration by segment, revenue vs holdout as launch gate |
| "Go multi-region" | Budgets across regions | Per-region budget shares rebalanced every few seconds; bound grows with regions |
| "An advertiser disputes the bill" | Auditability | Event-level logs, signed impression IDs, dedupe evidence |
| "How do you test the auction?" | Determinism | Auction logged with inputs; replayable offline; shadow auctions for rule changes |
6.3 What to Deliberately Skip#
- Model architecture — "light model then heavy model; I'd focus on calibration."
- Creative review and policy — one line: ads are reviewed before entering the index.
- Targeting taxonomy — "inverted index over attributes."
- Advertiser UI — say what the delivery report shows, not how.
- Attribution modeling — mention the window; skip the math.
6.4 Follow-Up Questions to Expect#
- "A campaign's win rate jumps 30×. Minute by minute, what happens to its budget?"
- "Split the 100 ms. What happens when one stage is slow?"
- "Second price or first price? What do advertisers do differently under each?"
- "The frequency store is down. Serve or not? Who decided?"
- "What exactly is a billable impression, and how do you prevent double billing on replay?"
- "How do you know a new model isn't quietly losing money?"
- "Budgets are per day in the advertiser's time zone. What happens at midnight?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design an ad serving system for our app."
Staff Answer
"Is this our own inventory with our own advertisers, or are we selling slots to external bidders? I'll assume our own — ads in our feed and search, our advertisers, our auction. Numbers: 500K ad requests per second at peak, 5 million active ads, 200K advertisers, a mix of CPM, CPC and CPA bids, and about 100 ms end to end, which leaves ~60 ms for our own work.
Four constraints shape everything: a hard deadline where a late ad is worth nothing; money on every decision, so overspend needs a bound; a pricing rule advertisers bid against; and billing that survives an audit. I'll go: entities and the impression lifecycle → the serving funnel → budget and pacing → latency and degradation → the auction → frequency caps → billing and ownership."
Why this is L6:
- Distinguishes owned inventory from exchange and bidder designs before drawing
- Frames the problem as bounded money under a deadline, with numbers
- Previews an outline that ends with billing and ownership
What L7 adds:
- Asks what the marketplace optimizes — revenue, advertiser ROI, user experience — and who arbitrates
- Asks whether other surfaces already run their own ad logic — the consolidation question
- Frames advertiser retention as a success metric alongside revenue
❌ Common L5 Trap
"We store ads in a database with their targeting. On each request we filter ads matching the user, rank them by bid times predicted CTR with an ML model, return the top one, and log the impression for billing."
Why this misses: Each step is reasonable; together they ignore the deadline, the budget and the difference between served and billed. The next questions — "a campaign goes viral" and "what do you bill on?" — have no answer.
Drill 2: Split the Latency Budget#
Prompt: "You have 100 ms end to end. Walk me through where it goes and what happens when a stage is slow."
Staff Answer
"Roughly 20 ms is network and client overhead, so the gateway stamps an absolute deadline of now + 80 ms. Retrieval gets 15 ms against an in-memory targeting index; feature fetch and the light model about 10; the heavy model 15; auction and guard 5; leaving ~15 ms of slack for serialization and tail. Every downstream call carries the deadline itself, not a relative timeout, so stages can't each assume a full share.
Degradation ladder: heavy model late → auction uses light scores; features late → defaults and light model only; retrieval late → per-placement cache of recent winners; nothing ready → 204 so the client can render a house ad. Retrieval is sharded, so I hedge: if a shard hasn't answered by its p95, send the same query to a replica. The metric I page on is serve.deadline_miss_rate, not average latency."
Why this is L6:
- Absolute deadline propagation instead of per-call timeouts
- A degradation ladder that always produces a valid response
- Names the fan-out tail problem and hedging
What L7 adds:
- Measures the revenue curve of latency — eCPM gained per ms of ranking vs fill lost per ms of p99
- Negotiates the budget with the host-page team as a shared SLO
❌ Common L5 Trap
"Each service gets a 100 ms timeout and we make them all fast with caching."
Why this misses: Four sequential stages with independent 100 ms timeouts can take 400 ms. Without a shared deadline and a fallback at each stage, one slow dependency turns into blank slots.
Drill 3: Make It Concrete — Size the Funnel and the Counters#
Prompt: "Size the serving tier, the frequency store and the event pipeline for 500K requests per second."
Staff Answer
"Serving: 500K requests/s × ~2,000 light-model scorings = 1 billion light scorings/s and 150 million heavy scorings/s. If a server does ~2,000 requests/s with batched inference, that's ~250 servers plus headroom, call it 400 across zones — the heavy model dominates, so it may run on accelerators.
Frequency store: 500M daily users × ~20 active capped campaigns × ~24 bytes ≈ 240 GB, replicated ×3 ≈ 720 GB — fits in memory across a modest cluster. Reads: one batched read per request = 500K/s. Writes: one per rendered impression, ~400K/s at a 80% render rate.
Events: served + rendered + clicks ≈ 1M events/s at ~300 bytes ≈ 300 MB/s, ~26 TB/day raw. Billing needs dedupe state for impression IDs over a 1-hour window: 400K/s × 3,600 ≈ 1.4 billion IDs; at ~16 bytes that's ~23 GB, partitioned across the aggregator."
Why this is L6:
- Sizes scoring work by funnel stage, which is what drives hardware
- Sizes per-user state and its read/write rates separately
- Sizes the dedupe window, which is the billing correctness mechanism
What L7 adds:
- Prices inference: heavy-model compute is likely the largest cost line, and depth is a revenue-vs-cost decision
- Uses the cost estimator framing to compare accelerator vs CPU serving at 3-year volume
❌ Common L5 Trap
"500K requests per second, each takes 50 ms, so we need 25,000 concurrent requests — about 250 servers."
Why this misses: Sizes the request, not the work. The scoring count per request, the per-user state and the billing dedupe state are what size the system; the candidate skipped all three.
Drill 4: The Viral Campaign#
Prompt: "A $50,000 daily campaign starts winning 30× more auctions at 9 a.m. How much does it overspend, and why?"
Staff Answer
"Two mechanisms act. Enforcement: each of ~400 servers holds a lease sized to ~10 seconds of its expected spend for this campaign — maybe $0.50 normally. At 30× win rate the lease drains in a third of a second, and servers request refills. The budget service sees refill rate spike and, because remaining budget is still large, grants them — so enforcement alone doesn't stop the spending; it just bounds the error at the end.
Pacing does the real work: the controller compares spend to a plan curve every 30 seconds. At 30× ahead of plan it drops participation probability — say to 3–4% — so the campaign enters only that fraction of auctions it's eligible for. With serve-side spend estimates feeding the controller, it reacts in one or two ticks.
End-of-day overspend: bounded by outstanding leases at exhaustion plus serves not yet rendered — with shrinking leases near the end, ~$50–200, under 0.5%. The risk is feedback lag: if spend data is minutes old, the bound grows with lag × spike rate, so I alert on spend.feedback_lag_s and shrink leases when it rises."
Why this is L6:
- Separates enforcement from pacing and explains which one acts on a spike
- Gives a numeric overdelivery bound with its derivation
- Identifies feedback lag as the risk multiplier
What L7 adds:
- Treats the bound as a finance policy and reports overdelivery per campaign
- Notes that pacing quality — spending evenly to end of day — is a sales promise and measures it
❌ Common L5 Trap
"We check the remaining budget in the database before serving each ad, so it can't overspend."
Why this misses: At 30× win rate the check becomes a hot-row read at thousands per second, and spend is recorded after render, so the number checked is already stale. It either melts the store or overspends anyway.
Drill 5: Dependency Down — The Frequency Store#
Prompt: "The frequency-cap store in one region just became slow — p99 400 ms. What happens, and what should happen?"
Staff Answer
"What should happen: the cap read has its own 5 ms deadline inside the guard stage. On timeout we fail open automatically, increment serve.capcheck_failopen_total, and apply a local ceiling — per server, per user, no more than 5 impressions of a campaign per hour, held in a small in-memory LRU. Campaigns with a cap of 1 and regulated categories are excluded while caps are unavailable, because those advertisers bought a specific exposure limit.
What must not happen: every request blocking on a 400 ms read and blowing the 80 ms deadline — a revenue outage from a UX feature. The fail-open policy is product's decision, made in advance and encoded as default behavior, so the on-call doesn't have to choose during an incident. Afterward, I'd report the number of over-cap impressions to affected advertisers if contracts require it."
Why this is L6:
- Bounded wait with automatic fail-open, not a manual flag flip
- A ceiling and category exclusions limit the UX damage
- The policy is owned by product and decided ahead of time
What L7 adds:
- Defines a class of 'soft dependencies' across ads with a uniform fail-open contract and chaos tests
- Measures the user-experience cost of fail-open (hide rate) to inform the ceiling
❌ Common L5 Trap
"If the cap store is unavailable, we don't serve capped campaigns, to respect the advertiser's settings."
Why this misses: Most campaigns are capped, so fail-closed is effectively a regional revenue outage — and the latency of slow reads still breaks the deadline for every request before it fails.
Drill 6: The Hot Campaign#
Prompt: "One advertiser's campaign is eligible for 40% of all requests. What breaks?"
Staff Answer
"Three hot spots. Budget: with leases there's no per-request write, but refill requests concentrate on one budget-service key — I'd shard the campaign's remaining budget into sub-budgets owned by different budget-service partitions and lease from those. Frequency caps: the campaign's counters are spread across users, so the store is fine — but the batched read always includes it, which is fine. Event pipeline: if spend events are partitioned by campaign, one partition gets 40% of traffic — partition by impression ID and pre-aggregate before rolling up per campaign.
Auction quality: a campaign in 40% of auctions makes pacing errors expensive — I'd run its pacing loop at a tighter interval, 10 s, and watch its participation probability. And I'd check whether targeting is too broad by mistake; 'eligible for 40%' is often a targeting misconfiguration."
Why this is L6:
- Finds the hot key in each component, not just one
- Splits counters and partitions by a high-cardinality key with two-stage aggregation
- Questions whether the input is a bug
What L7 adds:
- Adds an onboarding guardrail that flags campaigns whose eligibility exceeds a threshold before launch
- Links this to Hot Keys as an org-wide pattern for counters
❌ Common L5 Trap
"Put that campaign's counter on a bigger Redis instance."
Why this misses: A bigger node raises the ceiling but keeps one key on one thread. The fix is splitting the key and aggregating, and the candidate didn't look at the event pipeline's partitioning at all.
Drill 7: Build vs Buy#
Prompt: "Should we build our own ad serving stack or use a third-party ad server and exchange?"
Staff Answer
"It depends on whether ads are the business. If we're a content product monetizing spare inventory, buy: an ad network or exchange brings demand, auctions, billing, fraud filtering and advertiser relationships. Building gives us none of the demand, which is the hard part. Our job is a solid SDK integration, a placement policy and revenue reporting.
Build when we have first-party signals that make our inventory worth more — logged-in users, intent from search — and enough advertisers to run our own marketplace, because then the targeting and ranking are our moat and the platform take rate would exceed the cost of a team. A common path is hybrid: our own auction for direct advertisers, with an exchange as backfill for unsold slots, compared in one auction on eCPM. What I'd never outsource is the user-data policy — what signals leave our systems."
Why this is L6:
- Identifies demand, not technology, as the hard part to build
- Gives criteria: first-party signals and advertiser volume
- Proposes the hybrid with exchange backfill
What L7 adds:
- Prices the take rate against a team's fully loaded cost over 3 years
- Treats user-data sharing with third parties as a privacy and regulatory decision owned above engineering
❌ Common L5 Trap
"Build it — we have the engineers and it's a filter, a model and an auction."
Why this misses: The serving stack is the smaller problem. Advertiser demand, billing, fraud and measurement are the business, and none of them are costed.
Drill 8: Changing a Reserve Price Without an Outage#
Prompt: "Product wants to raise reserve prices on the home feed by 30%. How do you ship it?"
Staff Answer
"A reserve change trades fill for price, so the question is whether total revenue rises and whether advertisers stay. I'd ship it as an experiment: reserves are a per-placement config with a version, and a request-level holdout keeps 5–10% of traffic on the current value. We watch revenue per 1,000 requests, fill rate, average clearing price, advertiser win rate by spend tier and user metrics — fewer ads might also mean more engagement.
Rollout: 1% → 10% → 50% over two weeks, with guardrails that auto-revert if fill drops more than a set amount or small-advertiser win rate falls past a threshold. Auctions log the reserve version so we can attribute changes and answer disputes. If it wins, communicate it to advertisers, because their delivery and pacing will shift."
Why this is L6:
- Treats a pricing change as an experiment with a holdout and guardrails
- Watches advertiser-side and user-side metrics, not just revenue
- Versions the config so auctions are reproducible
What L7 adds:
- Considers whether small advertisers are priced out — marketplace health, not just revenue
- Establishes a pricing change process for every placement, with sign-off from ads product and finance
❌ Common L5 Trap
"Update the reserve price in config and deploy it."
Why this misses: Fill drops, pacing for every campaign shifts, and nobody can tell whether revenue rose or fell because there's no holdout. If it hurt, the damage is discovered by finance a week later.
Drill 9: The Cost Question#
Prompt: "Where does the money go in this system, and how would you cut infrastructure cost by 30%?"
Staff Answer
"Three buckets. Ranking inference, usually the largest: heavy-model scoring of hundreds of candidates per request. Levers: cut heavy-model depth from 300 to 150 if the revenue curve says the marginal 150 add little; distill the model; cache scores for logged-out traffic by context. That's likely 20–40% of inference cost with a measured eCPM impact.
Event pipeline and storage: a million events per second, raw and aggregated, kept for billing disputes and training. Keep raw events in object storage with compression, aggregates in an OLAP store; shorten hot retention.
Serving fleet headroom: provisioned for peak. Autoscale on request rate with a deadline-miss guardrail. I'd measure every cut against revenue, because a 1% eCPM drop usually costs more than the whole infra saving."
Why this is L6:
- Identifies inference as the dominant cost and gives specific levers
- Ties cost cuts to a revenue measurement
- Doesn't optimize the cheap parts first
What L7 adds:
- Frames cost per 1,000 requests against revenue per 1,000 requests — margin, not cost
- Makes model owners accountable for inference cost in launch reviews
❌ Common L5 Trap
"Use spot instances and reduce the number of replicas."
Why this misses: It ignores where the cost is and what each cut does to revenue. In ads, a cost cut that lowers eCPM by a fraction of a percent can lose more than it saves.
Drill 10: Multi-Region Budgets#
Prompt: "We serve from four regions. How do budgets work across them?"
Staff Answer
"Serving is region-local; budgets are global. I'd give each region a share of each campaign's remaining budget — initially proportional to that region's recent spend for the campaign — and have a global budget coordinator rebalance shares every few seconds based on reported spend. Within a region, the regional budget service issues leases from its share as before.
Overdelivery bound grows: it's now outstanding leases plus each region's unreported spend since its last sync. With a 5-second sync, four regions and shrinking shares near exhaustion, it stays well under 1%. If a region is partitioned from the coordinator, it spends only its current share and then stops — fail-closed for budget, because overspend during a partition is unbounded otherwise. Frequency caps stay regional with users pinned to a home region; cross-region cap drift is accepted and documented."
Why this is L6:
- Region-local serving with a global budget split into rebalanced shares
- States how the bound changes and what happens during a partition
- Accepts and documents regional cap drift
What L7 adds:
- Decides per-region data residency for user-level state as a legal requirement
- Prices the overdelivery bound change of adding a region before launching it
❌ Common L5 Trap
"Use a globally replicated database for the spend counter so all regions see the same value."
Why this misses: A synchronous cross-region write per win adds tens of milliseconds to a 60 ms budget and creates a global hot key. It trades a bounded overspend for an unbounded latency and availability problem.
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — The Championship Game#
Context: During a major sports final, ad requests jump from 300K/s to 1.4M/s in the 90 seconds after kickoff. Fill rate drops from 96% to 58%, and the top brand campaigns — which paid for this moment — are underdelivering. The on-call escalates to you.
Questions to Surface First:
- Are we missing deadlines (latency) or declining to serve (budget, pacing, caps)?
- Which stage is over its budget — retrieval, features, ranking or guard?
- Are leases being refused because the budget service is overloaded?
- Is pacing throttling brand campaigns that are behind plan because the plan didn't expect this spike?
Typical L5 Approach: Scales the serving fleet. Autoscaling takes 6 minutes; by then the spike has moved to halftime and the brand campaigns have underdelivered.
Staff Approach: Reads stage metrics: ranking is fine, but the budget service's lease-refill p99 is 900 ms under the refill surge, so guards fall through to lower candidates. Raises lease sizes for the top campaigns (a config already built for this), which cuts refill QPS ~5×. Pacing plans for brand campaigns flagged as event-driven get a spike-shaped curve.
Principal Approach: Treats scheduled marquee events as capacity events with a playbook: pre-scaling, pre-sized leases, pacing curves for event-driven campaigns, and a sales process that tells the platform which campaigns bought the moment.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Stage budgets: ranking 22 ms (ok), guard 1 ms but fall-through rate 38%. Budget service refill p99 900 ms. Raise lease multiplier ×5 for top 500 campaigns. |
| Triage | Refill QPS proportional to request QPS ÷ lease size; spike ×4.6 with fixed lease size → budget-service saturation. |
| Quick fix | Leases sized from a short-window spend rate, not a daily average; pre-scaled budget service. |
| Guardrails | Alert on budget.refill_p99 > 50ms and guard.fallthrough_rate > 10%. |
| Post-mortem | Pacing plans for event-driven campaigns; event calendar fed to capacity planning. |
Metrics to Watch: serve.fill_rate, guard.fallthrough_rate, budget.refill_p99, campaign.pacing_ratio{brand}
Organizational Follow-up: sales flags campaigns sold against specific events; ads serving runs a load test at 5× the previous event's peak before each one.
Ownership Question: "Who owns underdelivery on a sold moment?" Staff answer: Ads serving owns the capacity and pacing behavior; sales owns flagging the commitment in advance — without the flag, pacing treats it like any other campaign.
Key Takeaway: "Lease size must scale with spend rate, or the budget service becomes the bottleneck exactly when the money is largest."
What clears the Staff bar:
- Distinguishes deadline misses from declines before acting
- Fixes lease sizing rather than just adding servers
- Connects the incident to how campaigns are sold
Deep Dive 2: Silent Failure — Revenue Down 4%, Nothing Alerting#
Context: Finance reports that revenue per 1,000 requests has drifted down 4% over two weeks. Latency, fill rate and error rates are all normal. Nobody changed the auction.
Questions to Surface First:
- Did a model, feature or reserve-price config change in the window?
- Is the drop uniform or concentrated in a segment — vertical, device, region?
- Are predictions still calibrated by segment?
- Did invalid-traffic filtering change, shifting what counts as billable?
Typical L5 Approach: Checks dashboards, sees all-green service health, and attributes it to seasonality.
Staff Approach: Splits revenue per 1,000 requests by segment and finds the drop is concentrated in Android. A feature pipeline change two weeks earlier started sending a device feature as null for a new OS version; the model's predictions for those users regressed toward the mean, compressing eCPM and clearing prices. Backfills the feature and adds per-feature null-rate monitoring.
Principal Approach: Establishes that revenue per 1,000 requests by segment, against a long-running holdout, is a first-class production signal — the same rank as latency — with an owner and an alert.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Segment revenue per 1,000 requests by vertical, device, region, placement. Drop isolated to new Android OS version. |
| Triage | Feature device_model null for that OS since parser change; model calibration ratio 0.82 for the segment. |
| Quick fix | Fix parser, backfill features; temporarily apply a segment calibration multiplier. |
| Guardrails | Alert on feature null-rate changes > 5 points; calibration ratio by segment outside 0.9–1.1. |
| Post-mortem | Feature pipeline changes go through shadow comparison against the previous version. |
Metrics to Watch: feature.null_rate{feature}, model.calibration_ratio{segment}, revenue_per_mille{segment} vs holdout
Organizational Follow-up: feature pipeline team adds model owners as reviewers for schema changes to features in production models.
Ownership Question: "Who owns a revenue drop with all systems green?" Staff answer: The ads ML team owns prediction quality signals; the ads serving team owns the holdout and segment revenue dashboards that make the drop visible within a day, not two weeks.
Key Takeaway: "In ads, the most expensive failures don't throw errors. They show up as prices."
What clears the Staff bar:
- Segments the revenue signal instead of accepting averages
- Traces the drop to a feature, then to calibration
- Adds monitoring on inputs, not just outputs
Deep Dive 3: Large-Customer Onboarding — 80,000 Campaigns and a 100M-User Audience#
Context: A large retailer is onboarding with 80,000 campaigns (one per product category and region), each targeting a custom audience list of up to 100 million user IDs, with tight frequency caps. Sales has signed. You're asked to make it work.
Questions to Surface First:
- How are custom audiences represented in the targeting index — and how big are they?
- What does 80,000 more campaigns do to retrieval posting lists and pacing controller load?
- How many of these campaigns compete with each other for the same users?
- What frequency-cap level does the retailer want — per campaign or per advertiser?
Typical L5 Approach: Loads audiences as user-ID lists attached to each campaign. Retrieval now intersects 100M-entry lists per request; p99 retrieval doubles.
Staff Approach: Inverts the representation: user → audience-segment IDs, stored with user features, so retrieval reads the user's segments (tens) and looks up campaigns targeting those segments. Proposes advertiser-level frequency caps so the retailer's own campaigns don't saturate one user. Groups campaigns under a shared budget so pacing runs on 50 budget groups, not 80,000 campaigns.
Principal Approach: Recognizes a product gap: large advertisers need campaign groups with shared budgets and advertiser-level caps. Productizes both, with limits on audience-list size per plan.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (design review) | Estimate: audience membership as user→segments costs ~8 bytes × segments per user; as campaign→users it costs 100M × 8 bytes × campaigns. |
| Triage | Self-competition: 200 of their campaigns eligible for the same user; their own auctions raise their own prices. |
| Quick fix | User→segment index; advertiser-level dedupe in the auction (best ad per advertiser per slot). |
| Guardrails | Per-advertiser limits on eligible-campaigns-per-request; alert on retrieval candidate count growth. |
| Post-mortem (pre-mortem) | What if they update all audiences daily? Batch audience loads off-peak with versioned swaps. |
Metrics to Watch: retrieval.candidates_per_request{advertiser}, retrieval.p99, pacing.controller_loop_s, auction.same_advertiser_competition_rate
Organizational Follow-up: sales engineering gets an onboarding questionnaire covering campaign count, audience sizes and cap level.
Ownership Question: "Who owns self-competition?" Staff answer: The platform does — an advertiser shouldn't bid against itself, so the auction keeps one ad per advertiser per slot by default.
Key Takeaway: "Invert large audiences to user-side segments, and pace groups, not campaigns."
What clears the Staff bar:
- Finds the representation problem in targeting before it hits retrieval latency
- Prevents an advertiser from competing with itself
- Reduces control-loop load with budget groups
Deep Dive 4: Post-Mortem — $410,000 in Duplicate Charges#
Context: After a stream-processor upgrade, a replay of 3 hours of events was used to recover a gap. Billing double-counted 2.5 hours of impressions for 4,100 advertisers. Advertisers' third-party measurement disagreed by ~14%, and disputes followed. You're leading the post-mortem.
Questions to Surface First:
- Where was deduplication enforced — the stream processor's state, or the ledger?
- How long was the dedupe window, and how long was the replay?
- Did any alert compare billable counts to raw unique impression IDs?
- Who approved the replay, and was there a runbook?
Typical L5 Approach: Increases the stream processor's dedupe window to 6 hours.
Staff Approach: Moves the uniqueness guarantee to the sink: the billing ledger rejects a second write with the same impression ID, regardless of what the stream processor does. Adds a reconciliation job comparing billable impressions per hour to unique rendered impression IDs from raw events; any excess pages. Issues credits computed from the raw log.
Principal Approach: Establishes that every pipeline writing money has idempotency at the sink and an independent reconciliation, and that replays of money-bearing streams require a reviewed runbook. Publishes the definition of billable impressions to advertisers.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Freeze invoicing for affected window; compute true counts from raw events with full dedupe. |
| Triage | Dedupe state 30 min; replay 3 h; ledger had no uniqueness constraint on impression ID. |
| Quick fix | Unique key on impression ID in the ledger; reprocess the window; credits issued. |
| Guardrails | Hourly reconciliation: billable ≤ unique rendered; page on any excess. |
| Post-mortem | Replay runbook requires billing sign-off; replays target a shadow ledger first. |
Metrics to Watch: billing.reconciliation_excess, billing.duplicate_rejections, pipeline.replay_in_progress
Organizational Follow-up: account management contacts top advertisers proactively with corrected reports.
Ownership Question: "Who owns billing correctness during a replay?" Staff answer: The billing team owns the ledger's uniqueness and the reconciliation; data platform owns the replay mechanics; neither should be able to double-bill alone.
Key Takeaway: "Dedupe in the stream is an optimization. Uniqueness in the ledger is the guarantee."
What clears the Staff bar:
- Moves idempotency to the sink rather than tuning a window
- Adds independent reconciliation against raw events
- Treats replays of money streams as controlled operations
Deep Dive 5: Multi-Region Expansion — Serving EU Users From the EU#
Context: The company is moving EU ad serving into an EU region for latency and data-protection reasons. User-level data for EU users — features, frequency counters, event logs — must stay in the EU. Advertisers run global campaigns with a single budget.
Questions to Surface First:
- Which state is user-level (must stay) and which is campaign-level (can be global)?
- How do global budgets work when one region's spend must be reported without user data?
- How does consent affect what signals can be used for ranking?
- What happens to models trained on global data?
Typical L5 Approach: Deploys a copy of the serving stack in the EU and replicates all databases between regions.
Staff Approach: Splits state: campaign data and budgets are global, replicated to the EU; user features, frequency counters and raw events for EU users live only in the EU. The EU region reports aggregated spend per campaign — no user IDs — to the global budget coordinator. Ranking uses only signals permitted by each user's consent state, with a contextual-only model path when consent is absent.
Principal Approach: Defines data classification — user-level vs aggregate — as a company standard, with automated checks that no user-level ads data crosses regions, and a model governance process for training on regional data.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Inventory state by class: campaign (global), user (regional), events (regional raw, global aggregates). |
| Triage | Billing aggregates contain user IDs for dedupe → dedupe must run in-region; only deduped spend leaves. |
| Quick fix | EU aggregator dedupes and reports spend per campaign per minute to the global coordinator. |
| Guardrails | Marker users with synthetic IDs; scan non-EU stores for markers weekly. |
| Post-mortem (pre-launch) | Budget shares for EU rebalanced by spend; overdelivery bound recalculated with an extra region. |
Metrics to Watch: residency.marker_found_outside_region (must be 0), eu.serve.deadline_miss_rate, budget.region_sync_lag_s{eu}
Organizational Follow-up: legal and privacy approve the data classification; ML team documents which models may train on EU data.
Ownership Question: "Who proves user data stayed in the EU?" Staff answer: Ads platform proves it for ads stores with marker tests; privacy owns the company-level attestation.
Key Takeaway: "Budgets are global, users are local. Only aggregates cross the border."
What clears the Staff bar:
- Classifies state by residency requirement before moving anything
- Keeps dedupe in-region so only aggregates cross
- Handles consent as a ranking-path decision
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Split a ~100 ms budget into per-stage deadlines and define a degradation ladder that always returns a valid response
- Design a two-stage retrieval and ranking funnel and size it by scoring work
- Rank by eCPM with a quality term and explain how second- and first-price rules change advertiser behavior
- Enforce budgets with leased slices and derive the overdelivery bound
- Pace spend with a throttling controller and explain why feedback lag is the main risk
- Implement approximate frequency caps that fail open with a ceiling
- Define billable impressions and make billing idempotent at the ledger
- Detect silent revenue regressions through calibration and holdouts
The Bar for This Question#
Mid-level (L4): Builds a service that filters ads by targeting, sorts by bid, returns the top one and logs it. Works in a demo. No deadline, no budget story, bills on serve.
Senior (L5): Adds an ML model for CTR, ranks by bid × pCTR, checks budget in a database, caps frequency with counters, and logs impressions to Kafka. The gap: synchronous budget checks on hot campaigns, no deadline propagation, no named pricing rule, billing on served ads, and exact caps on the hot path. The design works at average load and fails on the first viral campaign or slow dependency.
Staff+ (L6): Frames the problem as spending other people's money under a deadline within the first five minutes. Propagates an absolute deadline with a degradation ladder. Names and justifies the pricing rule. Bounds overdelivery with leases and paces with a controller, aware of feedback lag. Caps frequency approximately with an owned fail-open policy. Bills on rendered, deduplicated, filtered events with sink-level idempotency. Names who pays — the platform absorbs bounded overdelivery; advertisers never pay for unrendered or duplicate impressions. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 "Exact Budgets" Are a Billing Feature, Not a Serving Feature#
| Claim | Reality |
|---|---|
| "We never exceed budget" | The invoice never exceeds budget; serving overshoots by a bound |
| "Check before every serve" | Hot keys and stale reads — still overshoots |
| "Leases are approximate" | Leases are exact about their error: the bound is computable |
The Staff position: Cap the invoice exactly; bound serving overdelivery and report it.
Why this matters in interviews: Promising exact serving-time budgets is a Senior signal; computing the bound is Staff.
10.2 Calibration Matters More Than AUC#
| Model Change | Ranking | Prices |
|---|---|---|
| Better AUC, calibrated | Better | Correct |
| Better AUC, miscalibrated | Better | Wrong — revenue leaks |
| Same AUC, recalibrated | Same | Fixed |
The Staff position: Gate launches on calibration by segment and revenue against a holdout.
Why this matters in interviews: It shows you know predictions are prices, not just scores.
10.3 Billing on Serve Is Overbilling With Extra Steps#
| Count | What It Measures |
|---|---|
| Served | What we returned |
| Rendered | What the client showed |
| Billable | Rendered, unique, valid |
The Staff position: Bill only the third number, with idempotency at the ledger.
Why this matters in interviews: Interviewers ask "what do you bill on?" to see if you know there's a choice.
10.4 A Late Ad Is Worth Less Than a Worse Ad#
| Response | Revenue |
|---|---|
| Best ad, 30 ms late | Zero — the slot rendered without it |
| Second-best ad, on time | Most of the value |
| House ad, on time | Some user value |
The Staff position: Design the degradation ladder before tuning the model.
Why this matters in interviews: It reorders priorities from relevance to deadline, which is how production ad systems actually behave.
10.5 The Pricing Rule Is a Product Decision#
The Staff position: Second- vs first-price isn't an algorithm choice. It decides whether advertisers or the platform does the bid optimization, and changing it is a migration for every buyer. Engineers implement it; ads product and finance own it.
Why this matters in interviews: Treating pricing as a config flag is the clearest sign of missing marketplace judgment.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer builds a serving system that meets its deadline, bounds overspend and bills correctly. The Principal engineer notices that feed ads, search ads and video ads each run their own auction, their own pacing and their own billing — and that an advertiser running on all three gets three budgets that don't know about each other, three definitions of an impression and three dashboards. The L7 problem is one marketplace: shared budget and pacing across surfaces, one billing definition, one auction framework with per-surface parameters, and governance for pricing changes.
The Org-Level Fault Line#
One ads platform vs per-surface ad stacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each surface runs its own ad stack | Surface teams move fast; tuned to their format | N budgets per advertiser; inconsistent billing; duplicated infra | Advertisers (fragmentation), finance (reconciliation) |
| One platform, surfaces as placements | One budget, one billing definition, cross-surface pacing | Platform must serve very different latency and format needs | Platform team (scope) |
| Shared budget, billing and events; surface-owned ranking | Consistency where money is involved; surface teams own relevance | Interface discipline between ranking and auction | Platform (contracts), surfaces (integration) |
🧭 Principal Move: "The platform owns everything that touches money: budgets, pacing, auction framework, billing and invalid-traffic filtering. Surfaces own retrieval and ranking for their format and plug in via a scored-candidate interface. No surface bills on its own."
Cost Model#
Assumptions: CPU serving with batched inference for light models, accelerators for heavy models at large scale; event pipeline on a managed log; fully loaded engineer ~$250K/year. Illustrative ranges.
| Scale | Volume | Infra ($/month) | Headcount | On-call Load | Notes |
|---|---|---|---|---|---|
| Startup | 1K req/s, 2K campaigns | ~$3–10K | 2–3 eng | Shared rotation | DB-backed budgets are fine; buy demand via an exchange |
| Growth | 100K req/s, 100K campaigns | ~$150–400K (serving, inference, event pipeline, OLAP) | 15–30 eng across serving, ML, billing | Dedicated rotations for serving and billing | Inference and event storage dominate |
| Large | 1M+ req/s, millions of ads, 4 regions | ~$2–6M | 100+ eng | Per-region serving rotations; billing rotation | Every 1% of eCPM outweighs most infra savings |
The pricing insight: at scale, infrastructure is a small fraction of revenue, so the binding costs are revenue errors — miscalibration, deadline misses, overdelivery and disputes. A 0.5% overdelivery rate on $10M/day of spend is ~$1.8M a year; a 1% eCPM loss from an over-aggressive cost cut is larger than most infra budgets.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Pricing rule (first vs second price) | One-way-ish | Every advertiser's bidding strategy and tooling adapts |
| Definition of a billable impression | One-way | Contracts, reports and third-party reconciliation depend on it |
| Impression ID format and signing | One-way | Clients, beacons and billing dedupe depend on it |
| Budget semantics (daily average vs exact) | One-way-ish | Advertisers plan spend against it |
| Lease sizes, pacing interval | Two-way | Config |
| Ranking model | Two-way | Experiment |
| Event store engine | Two-way | Internal migration |
The Standard I'd Write#
RFC-ADS-001: Money-Bearing Ad Decisions
Status: Approved Owners: Ads Platform + Finance
Scope
Every system that selects, prices or bills an ad on any surface.
MUST
1. Price through the shared auction framework; pricing rules are documented
per placement and changed only through the pricing-change process.
2. Enforce budgets via the platform budget service; report overdelivery per
campaign; alert above 1% of daily budget.
3. Bill only rendered, deduplicated, invalid-traffic-filtered events, with
uniqueness on impression ID enforced at the ledger.
4. Propagate an absolute request deadline; return a valid response or a no-ad
signal before it.
5. Gate model launches on calibration by segment and revenue vs holdout.
SHOULD
1. Use approximate frequency caps with a product-approved fail-open policy.
2. Size budget leases from short-window spend rates.
3. Keep a 1-5% long-running holdout per surface.
Exceptions
Filed with Ads Platform; finance sign-off required for any billing exception;
time-boxed to one quarter.
Success metrics
- Overdelivery: under 1% of spend platform-wide
- Billing discrepancy vs third-party measurement: within agreed tolerance
- Deadline miss rate: under 0.5% at p99 traffic
- Surfaces billing outside the platform: 0
What I'd Tell the VP#
"Our three ad surfaces each run their own budgets and billing, so an advertiser spending on all three gets three invoices with three definitions of an impression, and we can't pace their spend across surfaces. I'm proposing a shared platform for everything that touches money — budgets, auctions, billing — while surface teams keep ownership of ranking. It's roughly a year of a 10-person team. The return is fewer billing disputes, cross-surface campaigns that sales can sell as one product, and overdelivery we can measure and cap instead of discovering. The main risk is migration; I'd start with billing, where inconsistency costs us most."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices revenue errors | "Half a percent of overdelivery is $1.8M a year; that's the number finance needs, not lease sizes." |
| Identifies one-way doors | "The billable-impression definition is forever; the ranking model changes weekly." |
| Redraws ownership | "Platform owns money; surfaces own relevance." |
| Treats marketplace health as a metric | "I'd judge a pricing change on small-advertiser retention, not just revenue." |
| Knows when not to build | "Without first-party signals and advertiser demand, an exchange beats our own stack." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a great auction for the feed" — correct, but ignores the other surfaces with their own budgets.
- "Overdelivery is bounded under 1%" — good, but not translated into dollars or agreed with finance.
- "We'll add an EU region" — no data classification separating user-level from aggregate state.
Appendices
Appendix A: Mechanics in Depth#
A.1 The Serving Path With Deadlines#
handle(request):
deadline = now() + 80ms
user = resolve_user(request, deadline - 70ms) # cached, ~1ms
cands = retrieval.query(targeting(user, request.context),
exclude=pacing.throttled_sample(),
deadline=deadline - 60ms) # hedged across shards
if cands.empty: cands = placement_cache.recent_winners(request.placement)
feats = features.batch_get(user, cands, deadline=deadline - 50ms) or defaults
light = light_model.score(cands, feats)
top = top_k(light, 300)
heavy = heavy_model.score(top, feats, deadline=deadline - 15ms) # partial ok
scored = merge(heavy, light) # light where heavy missing
for winner, price in auction.run(scored, request.placement):
if not freq.allows(user, winner.campaign, deadline=deadline - 5ms): continue
if not leases.try_spend(winner.campaign, price): continue
imp = sign(impression_id(), winner, price)
log.served(imp)
return ad(winner, imp)
return NO_AD # HTTP 204
A.2 Pricing Rules#
eCPM(ad) = bid_cpm for CPM
= bid_cpc * pCTR * 1000 for CPC
= bid_cpa * pCTR * pCVR * 1000 for CPA
score(ad) = eCPM(ad) * quality(ad)
Generalized second price (one slot):
winner = argmax score
price_eCPM = max(reserve, score(runner_up) / quality(winner)) + 0.01
CPC advertisers are charged price_eCPM / (pCTR * 1000) per click
First price:
price_eCPM = eCPM(winner) (bidder shades; platform may offer auto-bidding)
A.3 Pacing Controller#
every 30s for each campaign c:
planned = budget(c) * traffic_curve_fraction(now, c.tz) # share of day elapsed
actual = spend_estimate(c) # serve-side estimate, corrected by billing
error = (actual - planned) / budget(c)
p(c) = clamp(p(c) * exp(-k * error), p_min, 1.0) # multiplicative update
if remaining(c) < 5% of budget: lease_size(c) = small
serving: campaign c enters an auction with probability p(c)
Appendix B: Data Model#
CREATE TABLE campaigns (
campaign_id TEXT PRIMARY KEY,
advertiser_id TEXT NOT NULL,
budget_group TEXT, -- shared budget across campaigns
daily_budget BIGINT NOT NULL, -- micros
budget_tz TEXT NOT NULL,
bid_type TEXT NOT NULL, -- cpm | cpc | cpa
bid_micros BIGINT NOT NULL,
freq_cap INT, -- impressions per user per day
targeting JSONB NOT NULL,
state TEXT NOT NULL -- active | paused | exhausted
);
-- billing ledger: uniqueness is the double-billing guard
CREATE TABLE billable_events (
impression_id TEXT PRIMARY KEY,
campaign_id TEXT NOT NULL,
event_type TEXT NOT NULL, -- impression | click
price_micros BIGINT NOT NULL,
rendered_at TIMESTAMPTZ NOT NULL
);
-- frequency store (KV): key = user_id|campaign_id|yyyymmdd, value = count, TTL 26h
Appendix C: Coordination Mechanisms#
C.1 Budget Leases#
C.2 Quick Comparison#
| Mechanism | Guarantees | Failure Mode | Use For |
|---|---|---|---|
| Absolute deadline | Response before the client gives up | Stages ignore it | Every request |
| Budget lease | Overdelivery ≤ outstanding leases | Lag inflates the bound | Hot campaigns |
| Pacing throttle | Even spend over the period | Oscillation with lag | Every campaign |
| Approximate cap | Bounded over-exposure | Store outage | Capped campaigns |
| Signed impression ID | Forged beacons rejected | Key leak | Every served ad |
| Ledger uniqueness | No double billing | Missing constraint | Billing sink |
| Holdout | Revenue changes attributable | Holdout too small | Every surface |
Appendix D: API Contract & Client Behavior#
- Clients send
deadline_msand render a house ad on 204 or on timeout; they never retry an ad request synchronously. - The impression beacon fires once on render; retries on network failure are allowed because billing dedupes by impression ID.
- Beacons older than 1 hour after serve are logged but not billed.
- Click URLs redirect with a 302 in under 20 ms; click logging is async.
- Exchange-style integrations follow the OpenRTB contract: honor
tmax, return HTTP 204 for no-bid.
Appendix E: Observability#
Core metrics:
serve.deadline_miss_rate,serve.fill_rate,rank.stage_budget_exceeded_total{stage}campaign.overdelivery_pct,spend.feedback_lag_s,budget.refill_p99,guard.fallthrough_ratemodel.calibration_ratio{segment},revenue_per_mille{segment}vs holdoutfreqstore.read_p99,serve.capcheck_failopen_totalbilling.reconciliation_excess,billing.duplicate_rejections,ivt.filtered_rate
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| Deadline miss rate | > 2% for 3 min | Page |
| Spend feedback lag | > 60 s for 2 min | Page |
| Campaign overdelivery | > 1% of daily budget | Page (business hours ticket if < 3%) |
| Billing reconciliation excess | > 0 | Page (billing) |
| Calibration ratio by segment | outside 0.9–1.1 for 1 h | Ticket → page at 6 h |
| Frequency store read p99 | > 5 ms for 5 min | Ticket (auto fail-open active) |
Debugging the silent failure: service health is useless for revenue regressions. Watch revenue per 1,000 requests by segment against a holdout, calibration by segment and feature null rates.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 1K req/s | DB-backed budget check, single model, bill on rendered | First viral campaign |
| 1K–100K req/s | In-memory index, two-stage ranking, leases, pacing controller | Event-pipeline lag; hot campaigns |
| 100K–1M req/s | Sharded retrieval with hedging, accelerator inference, sink-idempotent billing | Cross-surface budgets |
| > 1M req/s, multi-region | Regional serving, global budget coordinator, regional user data | Org governance of pricing and billing |
What you don't build on day one: auto-bidding, cross-surface budgets, multi-region budget coordination, viewability billing, advertiser audit logs. Each has a trigger in Section 11.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Advertiser fairness: one ad per advertiser per slot by default, so large advertisers don't bid against themselves or crowd out small ones with near-duplicate campaigns.
- Budget groups: large advertisers pace groups, not individual campaigns, to keep controller load bounded.
- Small-advertiser health: track win rate and retention by spend tier; pricing changes are judged on it.
- Cost attribution: inference cost per surface and per model; model launches report cost per 1,000 requests alongside revenue impact. Use the back-of-envelope calculator and latency budget tool to rehearse the sizing.