Technologies referenced in this case study: Apache Kafka · PostgreSQL · Redis · DynamoDB · Envoy, Kong & NGINX · OLAP Databases
Related: Notifications · Message Broker · Job Scheduler · Idempotency & Exactly-Once · Rate Limiter · Backpressure & Load Shedding · Push vs Poll · API Contracts That Age Well
Go deeper:
- Email delivery faces the same untrusted-receiver problem with an extra twist, since mailbox providers score your reputation and throttle accordingly; see Email Delivery Service.
Reading Guide#
Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal View) and the appendices on scheduling, signing and SSRF |
What is a Webhook Delivery Platform? — Why interviewers pick this topic
A webhook platform tells other people's servers that something happened in yours. When an invoice is paid, a repository is pushed or an order ships, the platform sends an HTTP POST with a signed JSON payload to a URL the customer registered. The customer's code reacts — fulfills the order, kicks off CI, updates a CRM.
The hard part is not making an HTTP request. The hard part is that the receiving endpoints are tens of thousands of servers you don't control: some are fast, some time out after 30 seconds, some return 500 for three days, some are a misconfigured URL pointing at your own cloud metadata service. You must deliver every event at least once, never let one customer's broken endpoint delay anyone else's, prove to a customer what you sent and when, and let them replay what they missed.
Before vs After — the "one slow customer" scenario:
Without per-endpoint isolation:
t=0: Customer A's endpoint starts responding in 29s (their DB is locked).
t=+30s: Shared worker pool of 2,000: each A delivery holds a worker for 30s timeout.
A produces 300 events/s → needs 9,000 worker-slots. Pool exhausted.
t=+1min: Deliveries to all 40,000 other endpoints queue behind A. p99 latency: 2s → 14 min.
t=+5min: Customers B–Z open tickets: "your webhooks are delayed". Status page goes amber.
t=+40min: On-call manually pauses A. Backlog of 6M events drains over 25 minutes.
With per-endpoint concurrency limits and fair scheduling:
t=0: Customer A's endpoint slows to 29s.
t=+30s: A is capped at 20 concurrent deliveries; its queue grows, nobody else's does.
t=+2min: A's circuit opens after 50% failures over 1 min → deliveries spaced by backoff.
t=+2min: Other endpoints: p99 unchanged at 2s.
t=+3h: A recovers. Its backlog drains at its own concurrency cap, oldest first.
A's dashboard shows every attempt, status code and latency. Zero tickets from B–Z.
Why interviewers reach for this question: It looks like "put events on a queue and POST them" — a Senior answer in five minutes. The Staff answer lives in what the queue picture hides: multi-tenant isolation against hostile and broken receivers, a retry policy that is really a product contract, ordering that you probably shouldn't promise, signing and SSRF defense, and observability you expose to customers, not just to your on-call.
Mechanics Refresher: Delivery Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| At-least-once delivery | Retry until a 2xx is received or the retry budget is spent | Simple; no lost events within the window | Receivers must dedupe by event ID |
| Exponential backoff with jitter | Attempt n waits ~base × 2ⁿ ± random | Spreads load from recovering endpoints | Long tail: late attempts hours apart |
| Per-endpoint queue / concurrency cap | Each endpoint gets a bounded number of in-flight requests | One slow endpoint can't starve others | Many queues to schedule fairly |
| Circuit breaker per endpoint | Stop sending after a failure threshold; probe periodically | Saves workers and the receiver | Delays recovery by up to one probe interval |
| HMAC signature with timestamp | sig = HMAC(secret, timestamp + "." + body) in a header | Authenticity + replay protection | Shared secret management and rotation |
| Thin vs snapshot payload | Send only IDs and type, or the full object at event time | Thin: always fresh, small. Snapshot: no callback needed | Thin: receivers call your API. Snapshot: stale on arrival, larger |
| Replay / redelivery | Re-send stored events to an endpoint on request | Customers recover from their outages | Must keep event payloads for the replay window |
| Delivery log | Every attempt: time, status, latency, response excerpt | Customer self-service debugging | Storage at billions of rows/month |
For most production systems: At-least-once, unordered delivery with a stable event ID, per-endpoint concurrency caps and circuit breakers, exponential backoff for roughly three days, HMAC-SHA256 signatures with a timestamp and a 5-minute tolerance, SSRF-safe egress through dedicated proxies, and a customer-facing delivery log with self-service replay. The primitives are not the interview — isolation, the retry contract and who owns a failed delivery are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
A webhook platform is not a queueing question. Everyone can put events on Kafka and POST them.
It is a multi-tenant isolation and contract question that tests:
- Whether you design for the worst receiver, not the median one — slow, failing, malicious or misconfigured
- Whether your retry policy is a deliberate product promise with a defined end, not "retry with backoff"
- Whether you refuse to promise ordering you can't deliver — and know what to offer instead
- Whether customers can see and fix their own delivery problems without opening a ticket
The key insight: The platform's reliability is bounded by its least reliable receiver unless you isolate per endpoint. The queue is easy; the scheduler that keeps 40,000 independently failing endpoints from affecting each other is the system. Staff candidates design the scheduler; Senior candidates design the queue.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Event → Kafka → workers → HTTP POST | Asks "Who are the receivers — third-party customers or our own services? How many endpoints, and what's our promise when they're down?" | Asks "Is webhooks a product surface we version and support, or plumbing? Who owns the customer contract when deliveries fail?" |
| Isolation | Shared worker pool, more workers when slow | "Per-endpoint concurrency caps, circuit breakers and a fair scheduler, so one endpoint at 30s latency can't consume shared capacity" | Sets tenant tiers with reserved delivery capacity; prices isolation into plans |
| Retries | "Exponential backoff, 5 retries" | "Backoff with jitter for ~3 days, then the event is marked failed, visible in the dashboard, replayable for 30 days; endpoints failing continuously for days are disabled with notice" | Makes the retry schedule a published, versioned contract; measures support tickets per million failed deliveries as the product KPI |
| Ordering | "Kafka partition per customer for ordering" | "Unordered by default; events carry IDs, timestamps and resource versions so receivers can reconcile; per-resource ordering only as an opt-in with head-of-line blocking made explicit" | Decides org-wide that event schemas carry versions so ordering is never a platform promise |
| Security | "HTTPS" | "HMAC-SHA256 over timestamp + body, 5-minute tolerance, dual secrets during rotation, egress through proxies that block private and metadata IPs" | Treats customer-supplied URLs as an attack surface with its own threat model and pen tests |
| Observability | Internal dashboards | "Customer-facing delivery log: every attempt, status, latency, response excerpt, one-click replay" | Uses delivery logs as the support contract — tickets resolved by link, not investigation |
Why "isolation" separates levels
L5: "Workers pull from the queue and POST. If deliveries slow down, autoscale the workers." This works until one large customer's endpoint starts taking 30 seconds. Each of its deliveries holds a worker for the full timeout, autoscaling adds workers that also get consumed by the slow endpoint, and every other customer's deliveries wait behind it. The shared pool converts one customer's outage into a platform outage.
L6: "Every endpoint has a concurrency cap — say 20 in-flight by default, configurable per plan. The scheduler only dispatches to an endpoint with free slots, so a slow endpoint queues its own work and nobody else's. Behind that, a circuit breaker per endpoint: over 50% failures in a minute opens it, and we probe every 30–60 seconds. Worker capacity is spent on endpoints that can actually accept deliveries."
L7: "Isolation is a product tier. Enterprise customers get dedicated delivery capacity and higher caps; free-tier endpoints share a pool with stricter caps. I'd price the reserved capacity and make sure the free tier can't, in aggregate, starve paying customers — fair share by tenant, not by endpoint, because one tenant can register 500 endpoints."
Why "retries" separates levels
L5: "Retry five times with exponential backoff." Five attempts with base 1s finish within about a minute — useless against a receiver that's down for a 2-hour deploy. And what happens after the fifth? The event is silently dropped, and the customer finds out when their reconciliation fails weeks later.
L6: "The schedule spans about three days — roughly 15–20 attempts with jitter: seconds, minutes, then hours. After the last attempt the event is marked failed in the customer's delivery log, and it's replayable for 30 days. An endpoint that has failed every attempt for 3–5 days gets disabled, with emails at the first failure and before disabling. Each of those numbers is a promise customers build against, so they're documented."
L7: "The retry window is a cost and support decision. Three days of retries at our failure rate means storing and re-attempting ~X billion deliveries a month; shortening to 24 hours saves infra and raises support tickets from customers with weekend outages. I'd measure tickets per million failed deliveries and set the window where the sum of infra and support cost is lowest."
Why "ordering" separates levels
L5: "We'll guarantee ordering by putting each customer's events on one Kafka partition." Ordering in Kafka doesn't survive the HTTP hop: retries reorder events, parallel delivery reorders them, and if you enforce strict order per customer, one failing event blocks every later event for that customer for up to three days.
L6: "Unordered at-least-once is the default, and I'd say so in the docs. Each event carries a unique ID, a creation timestamp and — more usefully — the resource's version number, so a receiver can ignore an invoice.updated with version 7 if it already processed version 8. For customers who truly need order, opt-in per-resource ordering: events for the same resource ID are delivered serially, and a stuck event blocks only that resource, with a visible 'blocked' state."
L7: "Ordering is a schema decision more than a delivery decision. I'd require every event type to carry a monotonic resource version — an org-wide standard for every producing team — so the platform never has to promise order and receivers can always converge."
Positions to Commit To#
| Position | Rationale |
|---|---|
| At-least-once, unordered, with a stable event ID | Exactly-once over HTTP to a server you don't control doesn't exist; dedupe is the receiver's job and you must make it easy |
| Per-endpoint concurrency caps + circuit breakers | The slowest receiver must not set the platform's latency |
| Fair scheduling by tenant, not by endpoint or by event | One tenant can register hundreds of endpoints or emit 1,000× more events |
| A finite, documented retry schedule (~3 days), then failed + replayable | Infinite retries are infinite cost; silent drops are broken trust |
| Sign every delivery: HMAC-SHA256 over timestamp + body, 5-min tolerance | Authenticity and replay protection are table stakes |
| Egress through SSRF-safe proxies; resolve and pin DNS; block private ranges | Customer-supplied URLs are an attack surface pointed at your network |
| Customer-facing delivery log and replay | Every delivery question a customer can answer themselves is a ticket you don't get |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Outbound webhooks to third-party customers | Untrusted, heterogeneous receivers; multi-tenant fairness; public contract | Per-endpoint isolation, long retry window, signing, SSRF defense, delivery log, replay | One customer starves others; SSRF; silent drops | Every event delivered or visibly failed; p99 first-attempt latency < 5s for healthy endpoints |
| Internal service-to-service event fanout | Trusted receivers; high volume; schema evolution | A message broker with consumer groups — not webhooks | Using HTTP push where pull fits better | Consumer lag SLOs |
| Inbound webhook ingestion (receiving from providers) | Bursty, unordered, duplicated input from vendors | Fast-ack, persist raw, verify signature, process async, reconcile with provider API | Dropped events during deploys; processing before verification | Every provider event persisted within 1s of receipt |
🎯 Staff Move: "I'll design outbound webhooks to third-party customers — tens of thousands of endpoints we don't control. If the receivers were our own services, I'd use a broker with consumer groups and skip HTTP push entirely. If we're receiving webhooks, the design is a fast-ack ingestion buffer, which I can cover at the end. The interesting problems are isolation and the retry contract."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Ordered vs Unordered Delivery | Promise order per customer or resource (head-of-line blocking, slow) or deliver unordered (fast, receivers reconcile)? |
| 2 | Shared Pool vs Per-Endpoint Isolation | Simple shared workers (one slow endpoint hurts everyone) or per-endpoint caps and fair scheduling (millions of logical queues)? |
| 3 | Retry Window and Giving Up | How long to retry, when to disable an endpoint, and who decides — cost vs trust vs support load |
| 4 | Thin vs Snapshot Payloads | Send the full object (stale, large, sensitive) or just a reference (fresh, small, receivers call back)? |
| 5 | Trust Boundary: Signing and Egress | Shared-secret HMAC vs asymmetric signatures; how to stop customer URLs from reaching your internal network |
How Real Companies Built It#
Why this section belongs here: Three well-known webhook providers made visibly different choices about retries and redelivery. Naming the differences shows you see the retry policy as a product decision.
Stripe — Three Days of Retries, Signed Timestamps, No Ordering Promise#
Stripe's documentation states that in live mode it attempts delivery for up to three days with exponential backoff, that it does not guarantee events arrive in the order they were generated, and that endpoints may receive the same event more than once — receivers should log processed event IDs. Each delivery carries a Stripe-Signature header containing a timestamp and an HMAC-SHA256 signature over the timestamp and raw body; Stripe's libraries default to a 5-minute tolerance to reject replays, retries get a fresh timestamp and signature, and during secret rotation the old secret can stay valid for up to 24 hours with one signature generated per active secret. Manual resends are available for up to 15 days in the dashboard (Stripe docs).
Staff insight: Every one of those numbers is a contract — 3 days, 5 minutes, 24 hours, 15 days. In an interview, state your platform's equivalents explicitly; it shows you know the retry policy is something customers build against, not an internal tuning knob.
Shopify — Short Window, Fast Timeout, Subscriptions Removed on Persistent Failure#
Shopify's documentation says the receiving app must respond within five seconds or the delivery fails; failed deliveries are retried up to eight times over four hours, and if failures persist the subscription is removed, after which the app must recreate it and recover missed data from the outage period (Shopify docs).
Staff insight: A much shorter window than three days, paired with an aggressive response timeout, is a deliberate choice: it protects the platform's delivery capacity at the cost of pushing recovery onto receivers. Say: "Retry window length is a tradeoff between our infra and our customers' operational burden — and whatever we pick, receivers need a reconciliation path via the API."
GitHub — No Automatic Redelivery; Customers Redeliver Themselves#
GitHub's documentation states that it does not automatically redeliver failed webhook deliveries; instead it provides manual and API-driven redelivery, and its best-practices guide asks receivers to respond with a 2xx within 10 seconds. Each delivery carries an X-GitHub-Delivery header that is unique per event and preserved on redelivery, which receivers can use to detect duplicates (GitHub docs: failed deliveries, best practices).
Staff insight: "No automatic retries plus great redelivery tooling" is a legitimate design point — it puts control in the receiver's hands and keeps the platform's load predictable. The lesson for the interview: replay and an inspectable delivery log matter as much as the retry schedule, and a stable delivery ID is what makes redelivery safe.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Workers consume from Kafka and POST" | "One customer's endpoint takes 30s to respond. What happens to everyone else?" | Isolation, head-of-line blocking |
| "We retry with exponential backoff" | "For how long? Then what? Does the customer know?" | Retry contract, visibility, giving up |
| "We guarantee ordering" | "Event 3 fails for two days. What happens to events 4 through 40,000?" | Head-of-line blocking, opt-in ordering |
| "We sign payloads" | "How does a customer rotate the secret without dropping deliveries?" | Dual-secret rotation |
| "Customers register any URL" | "What if the URL is http://169.254.169.254/latest/meta-data/?" | SSRF defense |
| "We store delivery attempts" | "At 50K events/s with retries, how big is that log, and for how long?" | Storage sizing, retention |
| "Customers can replay" | "A customer replays 30 days of events into an endpoint that handles 10/s. Now what?" | Replay rate limiting, isolation for replays |
System Architecture Overview#
Reading the diagram: Producers write events to a durable log; the log is the replay source of truth. A matcher fans each event out to delivery tasks, one per subscribed endpoint. The scheduler is the heart of the system: it holds per-endpoint ready lists and a due-time index for retries, and only dispatches to endpoints with free concurrency slots and closed circuits, sharing capacity fairly across tenants. Workers sign and send through an egress proxy that blocks private address space. Every attempt lands in a customer-visible delivery log. The metric that tells you the platform is healthy is
scheduler.oldest_ready_agefor healthy endpoints — it should stay in seconds no matter how many endpoints are failing.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Delivery guarantee | "Exactly once" | "At-least-once with a stable event ID. Receivers dedupe; we make that trivial." |
| Slow endpoints | "Autoscale workers" | "Per-endpoint concurrency cap + circuit breaker + fair scheduling by tenant." |
| Retries | "5 retries with backoff" | "~3 days, ~16 attempts, jittered; then failed, visible, replayable 30 days; disable after persistent failure with notice." |
| Ordering | "Partition per customer" | "Unordered default; resource version on every event; opt-in per-resource ordering with visible blocking." |
| Security | "HTTPS" | "HMAC-SHA256 over timestamp.body, 5-min tolerance, dual secrets for rotation, SSRF-safe egress." |
| Payload | "Send the object" | "Snapshot for small, non-sensitive types; thin event + fetch for large or sensitive ones." |
| Customer support | "Check our logs" | "Customer-facing delivery log with response codes and one-click replay." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| Stripe live-mode retry window | up to 3 days, exponential backoff | A common long-window reference point |
| Shopify retry policy | 8 attempts over 4 hours; 5s response timeout | A common short-window reference point |
| GitHub response expectation | 2xx within 10s; no automatic redelivery | Redelivery tooling as the recovery path |
| Signature timestamp tolerance | ~5 minutes (Stripe library default) | Replay protection without false rejects from clock skew |
| Delivery timeout | 5–15s typical | Shorter protects capacity; longer reduces false failures |
| Healthy endpoint latency | p50 50–300ms, p99 1–3s | Sizes worker concurrency |
| Endpoints failing at any moment | commonly ~1–5% of active endpoints | Isolation must assume hundreds of failing endpoints at all times |
| Worker sizing (async I/O) | ~1–5K concurrent requests per worker process | 50K events/s × 300ms ≈ 15K in flight → a handful of workers if async, thousands of threads if not |
| Delivery log row | ~300–600 bytes (with truncated response excerpt) | 50K/s × 1.3 attempts × 30d ≈ 170B rows; ~50–100TB raw — columnar compression matters |
| Fan-out ratio | ~1–3 deliveries per event | Size delivery throughput, not event throughput |
| Default per-endpoint concurrency | ~10–50 in flight | Bounds blast radius of one slow endpoint |
| Replay window | 15–30 days of retained events | Sets event-log retention and storage |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on the event schema, the subscription API and Kafka partitioning, then run out of time before the interviewer asks the only question that matters: "One of your biggest customers' endpoints just started timing out. What happens to everyone else?" Compress the plumbing to ~8 minutes and spend the rest on isolation, the retry contract, security and customer-facing observability.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"Customers register HTTPS endpoints and subscribe to event types. When an event happens in our platform, we POST a signed payload to every subscribed endpoint, retry on failure, and give customers a delivery log and replay. Receivers are third-party servers we don't control."
Then the non-functional requirements, which is where the design lives:
"Three constraints drive everything. One: isolation — at any moment a few percent of endpoints are slow or down, and they must not affect healthy endpoints' latency. Two: no silent loss — every event is either delivered or visibly failed after a documented retry window. Three: receivers are untrusted and our requests go to arbitrary URLs, so signing and egress security are requirements, not features. I'll assume 40,000 active endpoints across 10,000 tenants, 20K events per second average and 50K at peak, with a fan-out of about 1.4 deliveries per event."
Then name the underspecified parts:
"I'd confirm: do customers need ordering? How long should we retry? Are payloads allowed to contain personal data? I'll assume unordered by default with opt-in per-resource ordering, a three-day retry window, and thin payloads for sensitive event types."
🎯 Staff Move: Saying "at any moment a few percent of endpoints are down" reframes the problem from "deliver events" to "isolate failures" — that's the sentence that sets the level.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Event:
event_id(globally unique),tenant_id,type,resource_id,resource_version,created_at,payload_ref - Endpoint:
endpoint_id,tenant_id,url,secrets[](active + rotating),event_types[],max_concurrency,ordering(none / per_resource),state(active / circuit_open / disabled) - Delivery:
delivery_id,event_id,endpoint_id,attempt_count,next_attempt_at,state(pending / in_flight / delivered / failed_final),last_status - Attempt:
delivery_id,attempt_no,started_at,latency_ms,status_code,error_class,response_excerpt(first 1KB)
Customer-facing API:
POST /v1/endpoints { url, event_types[], ordering? }
POST /v1/endpoints/{id}/secrets:rotate { expire_old_in: "24h" }
GET /v1/endpoints/{id}/deliveries?status=failed&since=…
POST /v1/deliveries/{id}/retry
POST /v1/endpoints/{id}/replay { since, until, event_types? } → replay job
POST /v1/endpoints/{id}/test { event_type }
Delivery request (what the customer receives):
POST https://customer.example/hooks
Webhook-Id: evt_01HV… (stable across retries)
Webhook-Timestamp: 1767225600
Webhook-Signature: v1=5257a8… (HMAC-SHA256 of "timestamp.body"; one per active secret)
Content-Type: application/json
{ "id": "evt_01HV…", "type": "invoice.paid", "created": …, "data": { "object": { "id": "in_9", "version": 8, … } } }
🎯 Staff Move: "The event ID is stable across retries and replays; the signature and timestamp are fresh on every attempt. That combination lets receivers dedupe by ID while still rejecting captured-and-replayed requests by timestamp."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk one event in 90 seconds:
- Billing commits an invoice payment and an outbox row in one transaction; the relay publishes
invoice.paidto the event log withevent_id. - The matcher looks up subscriptions for
(tenant, invoice.paid)— 2 endpoints — and creates 2 delivery tasks. - Each task lands on its endpoint's ready list. The scheduler picks endpoints round-robin within tenant shares, dispatching only if the endpoint has a free concurrency slot and its circuit is closed.
- A worker signs the body with each active secret, sends through the egress proxy with a 10s timeout.
- 2xx → delivered. Timeout, connection error, 5xx or 429 → schedule retry at
now + backoff(attempt) ± jitter. Most 4xx → also retry (customers misconfigure and fix), but 410 Gone disables the endpoint. - Every attempt writes a row to the delivery log, visible to the customer within seconds.
🎯 Staff Move: Say out loud: "The scheduler never blocks on a slow endpoint — it skips endpoints without free slots. Worker capacity is only ever spent on endpoints that can accept work." You've now spent ~8 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes this hard is that thousands of endpoints are failing at any moment, customers want ordering we can't cheaply give, and we're making HTTP requests to arbitrary URLs. I'd like to go deep on isolation and scheduling, the retry contract and giving up, ordering, and security. Where would you like to start?"
If no preference: start with isolation. It's the question that decides the level.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: Isolation and fair scheduling (7–8 min)
"Each endpoint has a ready list and a concurrency cap, default 20. The scheduler keeps a set of 'dispatchable' endpoints — ready work, free slots, closed circuit — and picks among them with weighted fair sharing by tenant: each tenant gets a share proportional to plan, split among its endpoints. A slow endpoint fills its 20 slots and drops out of the dispatchable set; its queue grows, nobody else's does. Circuit breaker per endpoint: over 50% failures across at least 20 attempts in a minute opens it; while open, deliveries wait and one probe goes out every 30–60 seconds."
Quantify: "50K deliveries/s at 300ms median is ~15K in flight. A 30s endpoint capped at 20 slots consumes 20 of those, not 9,000. With async I/O, that's a handful of worker processes."
Who pays: "The failing customer's events get delayed — which is right; their endpoint is the cause. Healthy customers pay nothing."
Deep dive 2: The retry contract (6–7 min)
"Schedule: immediate, 5s, 30s, 2m, 10m, 30m, 1h, then every 2–3h until 3 days — about 16 attempts, each jittered ±20%. After the last, failed_final, visible in the log and replayable for 30 days. If an endpoint has had zero successes for 3 days, we disable it — emailing at first failure, at 24h and at disable. Disabling stops us retrying into the void and tells the customer something is structurally wrong."
Quantify: "With ~3% of endpoints failing at any time and their traffic share, retries add maybe 10–30% to attempt volume normally — and a burst when a large endpoint recovers, which the per-endpoint cap absorbs."
Deep dive 3: Ordering (4–5 min)
"Default unordered. Every event carries resource_id and resource_version; receivers keep the highest version seen per resource and drop older ones. Opt-in per-resource ordering: the scheduler dispatches only the oldest pending delivery per (endpoint, resource) — a failing event blocks that resource only, and the dashboard shows it as blocked with the reason."
Deep dive 4: Security (4–5 min)
"Signing: HMAC-SHA256 over timestamp.body, a header carrying one signature per active secret so rotation has an overlap window of up to 24 hours. Egress: every request goes through proxies that resolve DNS themselves, refuse private, loopback, link-local and metadata ranges — including after redirects, which we don't follow anyway — and send from a published IP range customers can allowlist."
Deep dive 5: Customer observability and ownership (3–4 min)
"Every attempt is visible to the customer within ~5 seconds: status, latency, error class, first 1KB of response. They can retry one delivery, replay a time range, or send a test event. The webhook platform team owns delivery health and pages on scheduler.oldest_ready_age for healthy endpoints. Customer endpoint failures are not our incidents — they're surfaced to the customer. Producing teams own event schemas and versions."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: webhook reliability is bounded by the worst receiver unless you isolate per endpoint. So per-endpoint caps, circuits and a fair scheduler; a finite, documented retry contract with visible failure and replay; unordered delivery with versions instead of an ordering promise; and signing plus SSRF-safe egress because we're sending requests to arbitrary URLs."
The evolution closer:
"What I'd build later: dedicated delivery capacity for enterprise tenants; alternative destinations like delivering events into a customer's own queue or event bus for high-volume customers; and a schema registry with versioned event types. What I'd not build: exactly-once delivery — I'd invest in making receiver dedupe trivial instead."
🎯 Staff Move: End on who owns a failed delivery. "When an event fails after three days, the customer sees it, gets an email, and can replay it in one click. Our on-call never investigates a single customer's endpoint — that's what the delivery log is for."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Schema obsession | 10 min on event JSON shape | Names event fields in 30s; spends time on scheduling |
| Kafka deep dive | Explains partitions, consumer groups, offsets | "The log is the replay source; delivery state lives in the scheduler" |
| Ordering promise | Guarantees per-customer order via partitions | Unordered default; versions; opt-in per-resource |
| No isolation story | Autoscaling as the answer to slow endpoints | Concurrency caps and circuits volunteered before asked |
| No end to retries | "Retry with backoff" | States the window, the final state and the disable rule |
| No security | HTTPS only | Signing, rotation and SSRF in the first 20 minutes |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Webhooks are the interview where your system's quality is graded by other people's code. You can't fix a customer's endpoint, you can't make it faster, and you can't stop it from returning 500 for three days. The candidate must design a system whose reliability for each tenant is independent of every other tenant's behavior — and whose failure modes are explained to customers by the product, not by your on-call. That is Staff work: designing blast-radius boundaries and operational contracts with parties you don't control.
It also has a deceptive cost structure. Healthy deliveries are cheap. Failing deliveries are expensive: they hold connections for the full timeout, they retry for days, they fill the delivery log, and they generate support tickets. A platform where 3% of endpoints fail can spend more than half its capacity on them unless the scheduler is designed to make failure cheap.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"Your largest customer's endpoint is returning 503 for every request and has been for six hours. They emit 2,000 events per second. Walk me through what your system is doing right now, what it costs, and what the customer sees."
A candidate who answers with the circuit state (open, probing every minute), the backlog (43M deliveries queued for that endpoint, retained, not consuming workers), the cost (storage, not compute), the customer view (dashboard banner, failure email at first failure and 24h, disable warning before day 3) and the recovery plan (drain at the endpoint's concurrency cap when probes succeed, oldest first or newest first by customer choice) has run a webhook platform. A candidate who says "we retry with backoff" has written a webhook sender.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Outbound webhooks to third-party customers → isolation and contract
- Constraint: untrusted, heterogeneous receivers; public, documented behavior; multi-tenant fairness
- Strategy: per-endpoint scheduling, finite retries, signing, SSRF-safe egress, delivery log, replay
- Failure mode: noisy-neighbor starvation, SSRF, silent drops, signature breakage during rotation
- Who pays for imperfection: healthy customers (delays caused by others), the platform's security team (SSRF), support (tickets)
Internal event fanout → use a broker
- Constraint: trusted consumers, high volume, need for replay and consumer-controlled pace
- Strategy: a log-based broker with consumer groups; consumers pull at their own rate
- Failure mode: building HTTP push with retries for consumers that could simply pull
- Who pays: the team maintaining a bespoke push system that duplicates the broker
Inbound webhook ingestion → fast-ack buffer
- Constraint: providers expect a 2xx within 5–10s; events arrive duplicated and unordered
- Strategy: verify signature, persist raw event keyed by provider event ID, return 200 in < 100ms, process asynchronously, reconcile against provider API
- Failure mode: processing inline and timing out → provider retries → duplicates; dropping events during deploys
- Who pays: the business, when a payment or order event is missed
2.2 When NOT to Build a Webhook Platform#
- Your receivers are your own services. Use a broker. Pull beats push when you control the consumer: backpressure is natural, replay is an offset reset, and there's no HTTP retry policy to design.
- You have fewer than ~100 endpoints and low volume. A job queue with a retry policy and a delivery table is enough. The per-endpoint scheduler earns its keep when noisy neighbors exist.
- Your customers need high throughput and strict order. Offer delivery into their own queue or event bus, or a pull-based events API with a cursor. Webhooks are the wrong transport for 5,000 ordered events per second to one consumer.
- Buying is viable. Managed webhook-sending services exist; if webhooks aren't core to your product, a vendor with per-endpoint isolation and a customer portal is likely cheaper than two engineers for a year — but the delivery contract you publish is still yours.
🎯 Staff Insight: "Webhooks are push over a network you don't control to code you don't control. I'd use them only when the receiver is external. For internal consumers, or high-volume external ones, a pull API with a cursor is simpler and more honest."
2.3 What the Interviewer Leaves Underspecified#
Interviewers deliberately omit:
- Ordering expectations — many candidates promise it; few can afford it
- Retry window and what happens after — "retry with backoff" isn't a policy until it ends
- Payload content — snapshot vs thin, and whether personal data may be pushed to customer URLs
- Endpoint validation — who checks that a URL isn't pointing at your internal network
- Customer visibility — whether customers can see attempts and replay, or must file tickets
- Peak shape — month-start billing runs or flash sales can produce 10–50× bursts for one tenant
Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.
2.4 Precise Terminology#
| Term | What It Means | Why It Matters in the Interview |
|---|---|---|
| Event | Something that happened, with a stable ID | Unit of dedupe for receivers |
| Delivery | One event to one endpoint | Unit of scheduling and retry |
| Attempt | One HTTP request for a delivery | Unit of the delivery log |
| Endpoint | Customer URL + subscription + secrets | Unit of isolation |
| Concurrency cap | Max in-flight attempts per endpoint | Bounds blast radius |
| Circuit breaker | Pause deliveries after failure threshold; probe | Saves capacity during endpoint outages |
| Head-of-line blocking | A stuck item blocks those behind it | The cost of ordering |
| Thin event | Type + IDs; receiver fetches details | Fresh data, smaller payload, extra API call |
| Snapshot event | Full object at event time | No callback; can be stale on arrival |
| Replay | Re-sending stored events on request | Requires event retention |
| SSRF | Server-side request forgery — tricking your server into calling internal addresses | Customer-supplied URLs are SSRF by design |
🎯 Staff Insight: If the interviewer says "guarantee delivery", ask: "Guarantee it within what window, and what does the customer see when the window closes? 'Guaranteed' without a window is a promise to retry forever."
3. Where the Design Splits#
Every webhook decision has a technical side (queue topology, retry math, signatures) and an organizational side (what we promise customers, who handles a failed delivery, who approves a retry-window change). Interviewers grade the second side.
3.1 Fault Line 1: Ordered vs Unordered Delivery#
The tension: Customers often ask for ordering — invoice.created before invoice.paid. Delivering in order requires serializing deliveries per key, which means one failing event blocks everything behind it, for as long as the retry window.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Strict order per tenant | Simplest mental model for receivers | One bad event blocks the tenant for up to 3 days; throughput = 1 / endpoint latency | Customer (stalled integration) |
| Strict order per resource (opt-in) | Blocks only one resource; parallel across resources | Scheduler tracks per-(endpoint, resource) heads; blocked state must be visible | Platform team (scheduler complexity) |
| Unordered + resource versions | Full parallelism; no blocking | Receivers must compare versions | Receivers (a few lines of code) |
| Unordered, no versions | Simplest for the platform | Receivers can't tell stale from fresh | Receivers (subtle bugs) |
Staff default: "Unordered at-least-once is the documented default. Every event carries resource_id and a monotonic resource_version, and our docs show the five-line receiver pattern for ignoring stale versions. Customers who need order can opt into per-resource ordering per endpoint; the dashboard shows blocked resources and the event blocking them."
When to deviate:
- State-transfer events (the payload is the full current state): ordering barely matters — the latest version wins.
- Command-like events that can't be reconciled by version (rare): per-resource ordering opt-in, with a shorter retry window so blocks don't last days.
- High-volume ordered streams: steer the customer to a pull API with a cursor or delivery into their own queue.
🧭 Principal Move: "I'd make
resource_versionmandatory in the event schema standard for every producing team. Then ordering is never a platform guarantee — it's a property of the data, and every receiver can converge."
❌ Common L5 Trap: "We'll put each customer's events in one Kafka partition, so they're delivered in order." Order in the log is lost at the first HTTP retry or the first parallel worker. To keep it, you'd deliver one event at a time per customer and block on failures — at 200ms per delivery, 5 events/s per customer, and a three-day stall on the first bad event.
3.2 Fault Line 2: Shared Pool vs Per-Endpoint Isolation#
The tension: A single shared queue and worker pool is simple and efficient when all receivers behave. With 40,000 receivers, some always misbehave, and a shared pool lets the slowest one set everyone's latency.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Shared queue, shared pool | Simple; one queue to monitor | Slow endpoints consume workers; head-of-line blocking across tenants | Healthy customers (delays) |
| Queue per tenant | Tenant isolation | 10K+ queues; one tenant's slow endpoint still blocks its fast ones | Platform (queue sprawl) |
| Per-endpoint ready lists + caps + fair scheduler | Isolation at the right grain; capacity only spent on dispatchable endpoints | Scheduler is a real component to build and operate | Platform team (scheduler ownership) |
| Dedicated workers per large tenant | Strong isolation for VIPs | Underutilized capacity; manual placement | Finance (cost) |
Staff default: "Per-endpoint ready lists with a concurrency cap (default 20, plan-dependent), a circuit breaker per endpoint, and a scheduler that does weighted fair sharing across tenants — work-conserving, so idle capacity goes to whoever has work. Workers use async I/O so in-flight requests are cheap; a slow endpoint costs 20 sockets, not 20 threads."
When to deviate:
- Small platform (< 100 endpoints): a shared queue with per-endpoint in-flight counters in Redis is enough.
- Enterprise tenants with contractual latency: dedicated scheduler shards or worker pools, priced into the contract.
🧭 Principal Move: "Fairness by endpoint is gameable — a tenant registers 500 endpoints. Fairness by event is gameable — a tenant emits 1,000× more events. Fair share by tenant, weighted by plan, is the policy; I'd publish it so sales doesn't promise something the scheduler can't do."
❌ Common L5 Trap: "When deliveries slow down, we autoscale workers." Autoscaling on queue depth adds workers that the slow endpoint immediately consumes, because every new worker picks up the next item — which is likely to be for the endpoint with the biggest backlog. You've scaled the problem.
3.3 Fault Line 3: Retry Window and Giving Up#
The tension: Longer retry windows recover more events across longer customer outages but cost storage and attempts and delay the signal to the customer that something is broken. Shorter windows are cheap and push recovery onto customers.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Short (hours, ~8 attempts) | Low cost; fast signal | Weekend outages lose events; customers must reconcile via API | Customers (recovery work) |
| Long (~3 days, ~16 attempts) | Covers most outages automatically | Large backlogs; recovery bursts; storage | Platform (infra) |
| No automatic retry, great redelivery | Predictable load; customer control | Customers must build redelivery automation | Customers (tooling) |
| Infinite retry | Nothing ever "fails" | Unbounded cost; zombie endpoints forever | Platform (cost), nobody notices |
Staff default: "About three days: attempts at 0, 5s, 30s, 2m, 10m, 30m, 1h, 2h, then every 3h — 16 or so attempts, ±20% jitter. Then failed_final, visible, replayable for 30 days. An endpoint with no 2xx for 3 consecutive days is disabled after two warning emails. Respect Retry-After on 429 and 503 up to a cap of 1 hour."
When to deviate:
- Time-sensitive events (OTP delivery, real-time alerts): short window — a 6-hour-old alert is worse than none. Product signs off.
- Financial events (payments, payouts): long window plus a reconciliation API, because a missed event costs money.
🧭 Principal Move: "The retry window is priced: attempts and storage on one side, support tickets on the other. I'd measure tickets per million
failed_finaldeliveries for a quarter before changing the window, and publish a change with 90 days' notice — it's a contract."
❌ Common L5 Trap: "We retry 5 times and then put it on a dead-letter queue." Whose DLQ? If it's ours, nobody watches it. If it's invisible to the customer, the event is effectively dropped. The Staff answer makes final failure customer-visible and replayable — the DLQ is the customer's dashboard.
3.4 Fault Line 4: Thin vs Snapshot Payloads#
The tension: Snapshot payloads let receivers act without calling back — and are stale on arrival, large, and push potentially sensitive data to customer URLs. Thin payloads are small and always force a fresh read — and add API load and latency on the receiver side.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Snapshot (full object) | One request; receivers stay simple | Stale under reordering; large; sensitive data at rest in customer logs | Customers (stale state), security (data spread) |
| Thin (type + IDs + version) | Fresh read; tiny; no sensitive data in transit to logs | Receivers call your API → API load spike after bursts | API team (read capacity) |
| Hybrid by event type | Snapshot where small and safe; thin where large or sensitive | Two receiver patterns to document | Docs/DX team |
Staff default: "Hybrid, decided per event type in the schema. Small, non-sensitive types are snapshots with a resource version. Large or sensitive types — anything with personal data or exceeding ~64KB — are thin, and our API's read path is sized for the callback burst: if we emit 50K thin events/s at peak, assume up to 50K extra reads/s within seconds."
When to deviate:
- Regulated data (health, financial account numbers): thin only, always.
- Receivers on constrained networks or with no API credentials (simple integrations): snapshot.
🎯 Staff Insight: "Thin events convert webhook load into API load. If I switch a high-volume event type to thin, I'm telling the API team to absorb a synchronized read burst — that's a cross-team capacity conversation, not a payload format choice."
❌ Common L5 Trap: "Always send the full object so receivers don't have to call back." Then a retried
subscription.updatedfrom an hour ago arrives after a newer one and overwrites the receiver's current state with stale data — and the payload contains a customer's address sitting in a third party's request logs.
3.5 Fault Line 5: Trust Boundary — Signing and Egress#
The tension: Receivers must verify that a request came from you; you must ensure that a customer-supplied URL can't make your infrastructure attack itself. Both are security requirements that interact with operations: secret rotation must not drop deliveries, and egress controls must not break legitimate endpoints.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| HMAC shared secret per endpoint | Simple, fast, widely understood | Secret must be stored by both sides; rotation needs overlap | Customers (secret handling) |
| Asymmetric signatures (publish public key) | No shared secret; one key for all receivers | Key rotation for everyone at once; slower verification | Platform (key management) |
| mTLS to receivers | Strong mutual auth | Customer must manage client CA trust; operationally heavy | Customers (setup) |
| Direct egress from workers | Simple | SSRF to metadata, internal admin, other tenants' services | Security (breach) |
| Egress via SSRF-safe proxy | DNS pinned, private ranges blocked, fixed IPs | Another hop to operate | Platform (proxy fleet) |
Staff default: "HMAC-SHA256 per endpoint, signing timestamp.body, a signature header with one entry per active secret so rotation has a configurable overlap up to 24 hours. Egress only through proxies in an isolated network segment with no route to internal services; the proxy resolves DNS itself, pins the IP for the connection, blocks private and metadata ranges and doesn't follow redirects. We publish our egress IP ranges so customers can allowlist."
When to deviate:
- Regulated customers may require mTLS — offer it as an enterprise option.
- A single public key for all receivers can be preferable when customers dislike secret handling, at the cost of coordinated key rotation.
🧭 Principal Move: "Customer-supplied URLs are an attack surface that the webhook team will under-invest in because it's 'just HTTP'. I'd have security own the egress proxy's threat model and pen-test it twice a year, with DNS-rebinding and redirect cases in the test suite."
❌ Common L5 Trap: "We validate the URL when the customer registers it — we reject private IPs." DNS can change after validation (rebinding): the hostname resolved to a public IP at registration and resolves to
169.254.169.254at delivery time. The check must happen at connect time on the resolved IP, every time.
4. When It Breaks#
4.1 The Slow Giant — One Endpoint Eats the Pool#
t=0: Tenant T7 (12% of all events) deploys a handler that holds a DB lock.
Endpoint latency: 120ms → 29s. No errors yet — just slow 200s.
t=+30s: Shared pool (pre-isolation design): T7 deliveries occupy 3,600 of 4,000 slots.
t=+1min: scheduler.oldest_ready_age for healthy endpoints: 2s → 90s.
t=+3min: Autoscaler adds 50% workers. T7 absorbs them within 40s.
t=+6min: Page: delivery.first_attempt_p99{healthy} > 60s for 5m.
t=+9min: On-call sets T7 endpoint max_concurrency=20 by hand. Healthy p99 recovers in 2 min.
t=+4h: T7 fixes handler. Its 21M-delivery backlog drains at 20 in-flight over ~6h.
Detection: delivery.first_attempt_p99 split by endpoint health; endpoint.inflight{endpoint} top-N; worker.slots_used_by_top_endpoint_pct.
Mitigation: enforce per-endpoint concurrency caps (default, not opt-in); latency-aware circuit — slow successes count against the endpoint too once p99 exceeds 10s.
Prevention: caps on by default for every endpoint; load test with 5% of endpoints at 30s latency.
Owner: webhook platform on-call.
4.2 The Recovery Stampede — 3 Days of Backlog in 3 Minutes#
t=0: Tenant T3's endpoint (load balancer misconfig) has failed for 60 hours.
Backlog: 9.4M deliveries in RETRY_WAIT, all due within the next hour.
t=+0: T3 fixes the LB. Next probe succeeds. Circuit closes.
t=+10s: Scheduler dispatches T3 at cap 200 (enterprise plan). T3's handler does 40/s.
t=+40s: T3's servers fall over again. 502s. Circuit reopens.
t=+2min: Cycle repeats every ~3 minutes. T3 opens a Sev-1 ticket against us.
Detection: endpoint.circuit_flaps{endpoint} (open→closed→open within 10 min); per-endpoint success rate after close.
Mitigation: slow-start on circuit close — begin at 2 in-flight, double every 30s while success rate > 95% (like TCP slow start); customer-configurable max delivery rate.
Prevention: slow-start is the default recovery behavior; dashboard lets customers choose drain order (oldest-first or newest-first) and pause/resume.
Owner: webhook platform.
4.3 The Silent 200 — "You Never Sent It"#
t=0: A customer's new framework middleware returns 200 for every webhook
before their handler runs, then the handler throws.
t=+5 days: Customer: "40,000 order events are missing. Your platform dropped them."
t=+5 days: Support has no visibility into customer handlers; escalates to engineering.
t=+5 days: Delivery log shows 40,000 attempts, all 200 in ~4ms, from our egress IPs.
Detection: this is the customer's failure — the platform's job is evidence. Delivery log with timestamps, status codes, response latency and response-body excerpts. A latency anomaly (p50 4ms vs historical 180ms) can be surfaced to the customer as a hint.
Mitigation: customer replays the window from the dashboard.
Prevention: customer-facing delivery log, self-serve replay, and documentation recommending receivers persist before acknowledging.
Owner: customer (root cause); platform owns the evidence and replay.
4.4 SSRF — The Metadata Endpoint#
t=0: Attacker registers https://hooks.attacker.example as an endpoint. Resolves public.
t=+1h: Attacker changes DNS: hooks.attacker.example → 169.254.169.254, TTL 1s.
t=+1h: Worker (direct egress) resolves at send time, connects to cloud metadata service.
t=+1h: Response excerpt (first 1KB) shows instance credentials in the attacker's delivery log.
Detection: egress.blocked_destination_total (should be nonzero — attempts happen); alerts on any successful connection to a non-public range (should be impossible).
Mitigation: rotate exposed credentials; disable the endpoint and tenant.
Prevention: egress proxy in an isolated segment; resolve-once-and-pin; block private, link-local, loopback, metadata and own VPC ranges at connect time; no redirects; instance metadata requires session tokens; response excerpts never shown for blocked or non-public destinations.
Owner: security (threat model), webhook platform (proxy).
4.5 Duplicate Storm After a Scheduler Failover#
t=0: Scheduler primary loses its lease. In-flight state for 14K deliveries not yet persisted.
t=+15s: Standby takes over; treats those 14K as pending; re-dispatches.
t=+16s: Customers receive ~14K duplicates within seconds. Most dedupe by Webhook-Id.
t=+2h: Two customers without dedupe report double-fulfilled orders.
Detection: delivery.redispatch_after_failover_total; customer reports.
Mitigation: this is at-least-once working as designed; the event ID made it safe for receivers that follow the docs.
Prevention: persist IN_FLIGHT with a lease before dispatch so the standby only re-dispatches expired leases; keep the docs' "dedupe by ID" guidance prominent with code samples.
Owner: webhook platform (scheduler), customers (dedupe).
4.6 Secret Rotation Breaks Verification#
A customer rotates their signing secret with "expire old immediately", but their receiver fleet takes 20 minutes to roll out the new secret. Every delivery in that window fails verification — the receiver returns 401, we retry, and the backlog clears once the rollout finishes. If they had chosen the 24-hour overlap, both signatures would have been in the header.
Detection: per-endpoint spike in 401/403 right after a rotation event — surface it in the dashboard with a hint.
Prevention: default overlap of 24 hours; dashboard warns when choosing immediate expiry.
Owner: customer (rollout), platform (safe default).
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Slow endpoint | endpoint.inflight at cap, latency p99 > 10s | That endpoint (with caps); platform (without) | Caps, latency-aware circuit | Webhook on-call |
| Recovery stampede | endpoint.circuit_flaps > 3/10min | That customer | Slow-start, customer rate limit | Webhook platform |
| Silent 200 | Latency anomaly; customer report | That customer's data | Delivery log evidence, replay | Customer (root cause) |
| SSRF attempt | egress.blocked_destination_total | Internal network if unblocked | Egress proxy, pin DNS | Security + webhook platform |
| Scheduler failover duplicates | delivery.redispatch_after_failover_total | Customers without dedupe | Leased in-flight state | Webhook platform |
| Event log lag | matcher.consumer_lag_seconds > 30 | All tenants' first attempts | Scale matcher; partition rebalance | Webhook on-call |
| Delivery log ingestion lag | dlog.ingest_lag_seconds > 60 | Customer visibility | Buffer; degrade excerpt capture | Webhook platform |
| Tenant burst | tenant.events_per_sec > 10× baseline | That tenant's latency (fair share) | Fair scheduler; burst credits | Webhook platform |
| Signature failures post-rotation | 401/403 spike after rotation | That endpoint | Overlap window, dashboard hint | Customer |
🎯 Staff Insight: The platform's health metric is healthy-endpoint latency, not overall latency. "Overall p99 will always look bad, because some endpoints are always down. I'd page on first-attempt p99 for endpoints whose last 20 attempts succeeded — that's the number that only moves when we're the problem."
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | "Send events to URLs with retries" | Names untrusted receivers and noisy neighbors as the core problem; commits to outbound third-party | Asks whether webhooks are a versioned product surface and who owns the customer contract |
| Isolation | Shared pool, autoscaling | Per-endpoint caps, circuits, tenant fair share, async I/O | Plan-tiered reserved capacity; fairness policy published and priced |
| Retries | "Backoff, N retries, DLQ" | ~3-day jittered schedule; failed_final visible; replay; disable with notice; slow-start on recovery | Retry window set by infra cost vs support tickets; changes announced as contract changes |
| Ordering | Kafka partition per customer | Unordered + resource versions; opt-in per-resource ordering with visible blocking | Resource versions mandated in the org's event schema standard |
| Security | HTTPS | HMAC over timestamp.body, rotation overlap, SSRF-safe egress with connect-time checks | Egress threat model owned by security; pen tests; published IP ranges |
| Operations | Internal dashboards | Healthy-endpoint latency SLO; customer-facing delivery log; on-call never debugs a customer endpoint | Support tickets per million deliveries as the product KPI; self-serve as the support model |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Designs for the worst receiver | "At any moment a few percent of endpoints are down. The design has to make that cheap." |
| Isolation at the right grain | "Concurrency cap per endpoint, fair share per tenant." |
| Retry policy with an end | "Three days, then failed, visible, replayable for 30 days." |
| Refuses a false ordering promise | "Unordered by default; resource versions let receivers converge." |
| Connect-time SSRF checks | "The proxy resolves and pins the IP and checks it at connect time — registration-time checks are bypassed by DNS rebinding." |
| Measures the right latency | "I page on first-attempt latency for healthy endpoints." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Exactly-once delivery" | Impossible across an HTTP hop to an uncontrolled receiver |
| Autoscaling as the slow-endpoint answer | Scales the problem; no isolation |
| Strict per-customer ordering via partitions | Head-of-line blocking for days; order lost on retry anyway |
| No end to retries, or silent DLQ | Infinite cost or silent loss |
| No SSRF consideration | Customer URLs are a direct path into the network |
| No customer visibility | Every failure becomes a support ticket and an engineering investigation |
5.4 Common False Positives#
- Deep Kafka knowledge ≠ webhook design. Partition counts and consumer-group tuning don't address the HTTP hop, where every interesting failure lives.
- Elaborate retry math ≠ a retry contract. A beautiful backoff formula without a stated end, final state and customer notification is incomplete.
- "We use a circuit breaker" ≠ isolation. A global circuit breaker on "the webhook service" protects nothing; it must be per endpoint.
- Security buzzwords ≠ SSRF defense. "We use a WAF" doesn't stop your own workers from calling the metadata service.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Outbound to third parties; isolation, no silent loss, untrusted receivers; numbers |
| Entities & API | 3–5 min | Event, endpoint, delivery, attempt; stable event ID; signature header |
| Architecture | 5–10 min | ≤ 8 boxes; event log → matcher → scheduler → workers → egress proxy |
| Isolation | 10–18 min | Per-endpoint caps, circuits, tenant fair share, async workers |
| Retry contract | 18–25 min | Schedule, final state, disable rule, slow-start recovery, replay |
| Ordering + payloads | 25–31 min | Unordered + versions; opt-in per-resource; thin vs snapshot |
| Security + observability | 31–38 min | HMAC + rotation; SSRF proxy; delivery log; healthy-endpoint SLO |
| Pivot / wrap | 38–45 min | Multi-region, enterprise tiers, cost; close on ownership of failed deliveries |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "A customer needs strict ordering" | Do you know the cost? | Per-resource opt-in, visible blocking, shorter window — or a pull API |
| "A customer replays 30 days" | Replay isolation | Replay as a separate low-priority lane with customer-set rate; doesn't share the live cap |
| "Make it multi-region" | Event locality, duplicates | Deliver from the region where the event was produced; failover accepts duplicates |
| "10× traffic for one tenant on Black Friday" | Fairness | Tenant fair share + burst credits; their latency rises, others' doesn't |
| "Customer says we never sent it" | Evidence | Delivery log with status, latency, excerpt; replay |
| "How do you test this?" | Chaos at the edge | Fake endpoint fleet with injected latency, 5xx, resets, slow bodies, DNS rebinding |
6.3 What to Deliberately Skip#
- Kafka internals — "the log is durable and partitioned by tenant; delivery state lives in the scheduler."
- Subscription CRUD — one line.
- Event schema design per product — producers own it; mention versions.
- Dashboard UI — say what it shows, not how it looks.
- Exactly-once — say why you won't offer it, in one sentence.
6.4 Follow-Up Questions to Expect#
- "One endpoint takes 30 seconds per request and gets 12% of all events. Walk me through the next 10 minutes."
- "How long do you retry, what's the schedule, and what happens at the end?"
- "A customer insists on ordering. What do you offer, and what does it cost them?"
- "How does a customer rotate their secret without dropping deliveries?"
- "Show me how a customer-registered URL could reach your internal network, and how you stop it."
- "An endpoint recovers after 3 days with 9M queued deliveries. What happens when the circuit closes?"
- "How big is the delivery log, and how do customers query it?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design a webhook system for our platform."
Staff Answer
"Who receives them — third-party customers or our own services? If they're our own services, I'd steer to a broker with consumer groups; HTTP push isn't the right tool. I'll assume third-party customers: about 40,000 endpoints across 10,000 tenants, 20K events per second average, 50K peak, about 1.4 deliveries per event.
Three constraints shape everything. Isolation: at any moment a few percent of endpoints are slow or down, and they mustn't affect anyone else. No silent loss: every event is delivered or visibly failed after a documented window. Untrusted receivers: we sign every request and we treat customer URLs as an attack surface. I'll go: entities and the delivery contract → architecture with a scheduler at the center → isolation → retry contract → ordering → security → customer observability and ownership."
Why this is L6:
- Distinguishes external webhooks from internal fanout and redirects the latter
- Frames the problem as isolation against failing receivers, with numbers
- Previews an outline that ends with ownership and customer visibility
What L7 adds:
- Asks whether webhooks are a versioned product surface with a support contract
- Asks whether other teams already send webhooks independently — the consolidation question
- Frames support tickets per million deliveries as the success metric
❌ Common L5 Trap
"Services publish events to Kafka, a consumer group of workers reads them, looks up subscribers and POSTs to each URL with retries and exponential backoff. Failed events go to a DLQ."
Why this misses: Each piece is reasonable; together they let one slow endpoint stall everyone, never say when retries end or who watches the DLQ, and ignore signing and SSRF. The interviewer's next question — "one endpoint takes 30 seconds" — has no answer.
Drill 2: The Slow Endpoint#
Prompt: "Your largest customer's endpoint starts taking 29 seconds to respond. They're 12% of traffic. What happens?"
Staff Answer
"Their endpoint fills its concurrency cap — 20 by default, maybe 200 on an enterprise plan — and drops out of the dispatchable set. The scheduler keeps dispatching to everyone else. Their deliveries queue; nobody else's do. Because workers use async I/O, 200 slow requests cost 200 sockets and some memory, not 200 threads.
29-second 200s aren't failures, so the error-based circuit won't open. I'd make the circuit latency-aware: if p99 for an endpoint exceeds 10s over a minute, treat slow successes as partial failures and reduce its cap toward a floor. The customer's dashboard shows rising latency and backlog; we email them once latency crosses the threshold for 10 minutes. Healthy-endpoint first-attempt p99 stays at ~2s — that's the metric we page on, and it shouldn't move."
Why this is L6:
- Isolation via caps and the dispatchable set, not autoscaling
- Notices that slow successes evade error-based circuits and fixes it
- Names the metric that proves the platform isn't the problem
What L7 adds:
- Notes that enterprise caps are a priced resource, and a 200-slot cap at 29s is a capacity commitment sales must understand
- Adds the customer-facing latency alert as a product feature that reduces tickets
❌ Common L5 Trap
"We'll autoscale the workers based on queue depth, so throughput stays up."
Why this misses: New workers pull the next deliveries, which are disproportionately for the endpoint with the biggest backlog — the slow one. Autoscaling feeds the slow endpoint more capacity and doesn't help anyone else.
Drill 3: Make It Concrete — Size the Workers and the Log#
Prompt: "Size the delivery tier and the delivery log for 50K events/s peak."
Staff Answer
"Deliveries: 50K events × 1.4 fan-out = 70K deliveries/s, plus ~20% retries ≈ 85K attempts/s. Healthy median latency ~250ms → ~21K in flight; add failing endpoints at their caps — say 1,000 failing endpoints × 20 = 20K — so ~40K concurrent requests. An async worker comfortably handles 2–5K concurrent requests, so ~10–20 worker processes plus headroom; across three zones, 30. With thread-per-request, it would be 40K threads — the reason async matters.
Delivery log: 85K attempts/s × 86,400 ≈ 7.3B rows/day. At ~400 bytes with a 1KB-truncated excerpt compressed, ~3TB/day raw, maybe 300–600GB/day in a columnar store. 30-day retention ≈ 10–20TB compressed. Queries are by endpoint and time, so partition by day and sort by (tenant, endpoint, time)."
Why this is L6:
- Sizes attempts, not events, and includes failing endpoints at their caps
- Shows why async I/O changes the hardware by orders of magnitude
- Sizes the log with compression and a query-shaped layout
What L7 adds:
- Prices it: the delivery log likely costs more than the workers; tiered retention by plan (7 vs 30 days) is a pricing lever
- Considers storing excerpts only for non-2xx attempts — a 70% storage cut
❌ Common L5 Trap
"50K events per second, each takes about 200ms, so we need 10,000 workers."
Why this misses: Confuses workers with concurrent requests, ignores fan-out and retries, ignores failing endpoints holding connections for full timeouts, and skips the delivery log — often the largest storage cost in the system.
Drill 4: The Retry Contract#
Prompt: "Define your retry policy. All of it."
Staff Answer
"Retryable: connection errors, timeouts (10s), 5xx, 429, 408, and most 4xx — customers misconfigure auth and fix it within hours. Not retryable: 410 Gone, which disables the endpoint. Schedule: 0, 5s, 30s, 2m, 10m, 30m, 1h, 2h, then every 3h until 72 hours — about 16 attempts, ±20% jitter. Retry-After on 429 or 503 is honored up to an hour.
At 72 hours: failed_final, visible in the log, replayable for 30 days. Emails at the first failure (debounced to one per endpoint per day), at 24h of continuous failure and 24h before disabling. An endpoint with zero successes for 72 hours is disabled — events for it still accumulate as failed and replayable, so re-enabling plus a replay recovers everything in the 30-day window. All of this is in the public docs."
Why this is L6:
- Covers classification, schedule, jitter, server hints, the end state and notifications
- Disabling doesn't lose data — replay recovers it
- Treats the policy as a documented contract
What L7 adds:
- Sets the window by measuring support tickets per million failed deliveries versus infra cost
- Commits to 90 days' notice for changes to the published policy
❌ Common L5 Trap
"Retry with exponential backoff, max 5 attempts, then dead-letter queue."
Why this misses: Five attempts span a minute or two — useless against a deploy outage. The DLQ is invisible to the customer and unowned internally, so the event is effectively lost.
Drill 5: The Recovering Giant#
Prompt: "A customer's endpoint was down for 3 days. It just came back. 9 million deliveries are waiting. What happens?"
Staff Answer
"The next probe succeeds and the circuit closes — but into slow-start, not full cap. Start at 2 in-flight, double every 30s while the success rate stays above 95%, halve on failure. That finds their real capacity without knocking them over. If they've set a max rate in the dashboard, that's the ceiling.
Ordering of the drain is the customer's choice: oldest-first (default, preserves rough chronology) or newest-first (current state sooner, then backfill). Events past the 72-hour window are already failed_final; the dashboard offers 'replay all failed since outage start' as one action, which runs in a separate replay lane at the endpoint's rate. Meanwhile, new live events share the endpoint's slots with the drain — I'd reserve ~25% of its slots for live traffic so their current integration isn't stuck behind three days of history."
Why this is L6:
- Slow-start avoids the recovery stampede
- Gives the customer control over drain order and rate
- Reserves capacity for live events during a backlog drain
What L7 adds:
- Tracks
endpoint.circuit_flapsplatform-wide as a signal that recovery defaults need tuning - Productizes 'outage recovery' as a guided flow in the dashboard
❌ Common L5 Trap
"The retries resume and the backlog drains."
Why this misses: At full cap, 9M deliveries arriving at a freshly recovered service usually knock it over again — the circuit flaps, the customer's outage extends, and they blame the platform.
Drill 6: The Noisy Tenant#
Prompt: "One tenant emits 40% of all events during their billing run on the 1st. Other tenants complain about delays. Fix it."
Staff Answer
"Per-endpoint caps don't help here — the tenant has 300 healthy endpoints, each under its cap, and they're legitimately consuming capacity. The fix is fair share by tenant: the scheduler allocates dispatch capacity across tenants with pending work, weighted by plan, and work-conserving — if others are idle, the big tenant gets everything. When others have work, the big tenant gets its weighted share and its own latency rises, which is the right victim.
I'd add burst credits so short spikes don't feel throttled, and show the tenant a 'throttled by fair share' indicator so they understand why their p99 rose. If they need guaranteed throughput at that volume, that's a reserved-capacity plan."
Why this is L6:
- Recognizes endpoint caps don't address tenant-level skew
- Work-conserving weighted fair share, with the right victim named
- Turns the problem into a visible state and a product option
What L7 adds:
- Uses the billing-run pattern as input to capacity planning — first-of-month peaks are predictable
- Prices reserved capacity so heavy tenants fund the headroom they need
❌ Common L5 Trap
"Rate-limit the tenant to a fixed number of events per second."
Why this misses: A fixed limit wastes capacity when others are idle and still doesn't guarantee fairness when many tenants are busy. It also turns a delivery-latency problem into dropped or rejected events at ingest.
Drill 7: Build vs Buy#
Prompt: "Should we build this or use a managed webhook-sending service?"
Staff Answer
"If webhooks are a feature rather than the product, buying is often right: a credible vendor gives per-endpoint isolation, signing, retries and a customer portal — roughly two engineers for two to three quarters of build, plus ongoing on-call, that we don't spend. I'd evaluate: isolation semantics under a slow giant, retry contract configurability, SSRF defense, data residency, delivery-log retention, and how replay works. The things I can't outsource are our event schemas, the retry contract we publish, and the customer relationship when deliveries fail.
I'd build when webhooks are core to the product, volume makes per-message vendor pricing exceed ~2–3 engineers' cost, or we need deep integration — e.g., per-resource ordering tied to our data model. Either way, producers publish to our own event log, so the sender is swappable."
Why this is L6:
- Gives criteria and thresholds, not a preference
- Separates what can be outsourced (delivery mechanics) from what can't (contract, schemas)
- Keeps the sender swappable behind our own event log
What L7 adds:
- Prices vendor cost at 3-year projected volume, not today's
- Requires exportable delivery logs and an exit plan before signing
❌ Common L5 Trap
"It's just HTTP requests with retries; we should build it — it's simple."
Why this misses: The HTTP request is the simple part. Isolation, fair scheduling, SSRF defense, rotation, delivery logs and replay are the system, and the candidate hasn't costed any of them.
Drill 8: Changing the Retry Policy Without an Outage#
Prompt: "Finance wants to cut the retry window from 3 days to 24 hours to save cost. How do you ship it?"
Staff Answer
"First, it's a contract change, so the question is whether it should ship at all. I'd measure: what fraction of deliveries eventually succeed between hour 24 and hour 72? If it's 0.5% of failed deliveries and those cluster in a few tenants with weekend outages, cutting the window shifts that recovery onto them as support tickets and replays. I'd present cost saved versus projected tickets.
If we proceed: announce with 90 days' notice, make the window a per-plan setting so enterprise keeps 72 hours, ship it as config keyed by endpoint plan with the current schedule as default, then flip free and standard plans in waves — 5%, 25%, 100% — watching failed_final rate, replay usage and tickets. Deliveries already in flight keep the schedule they started with; only new deliveries get the new window."
Why this is L6:
- Treats the change as a contract change backed by data
- Plan-dependent window as a middle ground
- Staged rollout with metrics; in-flight deliveries aren't retroactively cut
What L7 adds:
- Makes the window a pricing lever across plans rather than a global cost cut
- Establishes a change-management process for all public delivery contracts
❌ Common L5 Trap
"Change the max-retry config and deploy."
Why this misses: Silently breaking a documented contract turns into an incident when a customer's weekend outage loses two days of events that would previously have been delivered — and in-flight deliveries get cut mid-schedule.
Drill 9: The Cost of Failure#
Prompt: "Where does the money go in this system, and how would you cut it by 30%?"
Staff Answer
"Three buckets. Delivery log storage, usually the largest: billions of attempt rows a day. Cut by storing response excerpts only for non-2xx attempts, compressing in a columnar store, and tiering retention by plan — 7 days for free, 30 for paid. That alone can cut log cost 50–70%.
Egress and compute for failing endpoints: retries to endpoints that have failed for days. Disabling after 72 hours and lengthening late-stage backoff caps it. Circuit-open endpoints should cost storage, not attempts.
Event log retention for replay: 30 days of payloads. Keep thin references for large payloads and store snapshots in object storage at ~$0.02/GB-month rather than in the broker. Healthy deliveries are cheap — don't optimize those first."
Why this is L6:
- Identifies failure handling and observability as the real cost drivers
- Gives specific levers with estimated impact
- Avoids optimizing the cheap path
What L7 adds:
- Attributes cost per tenant and per plan so pricing reflects delivery-log and retry costs
- Treats retention tiers as product packaging, coordinated with sales
❌ Common L5 Trap
"Use cheaper instances and reduce the number of workers."
Why this misses: Workers are a small slice of cost in an async design. The money is in storing attempts and retrying into dead endpoints — the candidate is optimizing the wrong bucket.
Drill 10: Multi-Region#
Prompt: "We're going multi-region. How do webhooks work?"
Staff Answer
"Deliver from the region where the event was produced, through that region's egress IPs, which we publish per region so customers can allowlist all of them. Each region has its own event log, scheduler and workers; there's no cross-region coordination on the hot path. For EU data residency, EU tenants' events are produced, stored and delivered only from the EU region, and their delivery logs stay there.
On regional failure, the standby region takes over the failed region's event log from replication. Deliveries whose state hadn't replicated are re-dispatched — duplicates, which is at-least-once doing its job; receivers dedupe by event ID. Per-resource ordering for opted-in endpoints can break across a failover — I'd document that explicitly rather than pretend otherwise."
Why this is L6:
- Region-local delivery, no cross-region coordination
- Data residency carried through to delivery logs
- Honest about duplicates and ordering on failover
What L7 adds:
- Plans egress IP changes as a customer-facing migration with long notice — allowlists break silently
- Decides residency per tenant contract, not per feature
❌ Common L5 Trap
"Run active-active with a global queue so any region can deliver any event."
Why this misses: Adds cross-region latency and coordination for no benefit, breaks data residency, and makes egress IPs unpredictable for customers who allowlist.
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — First-of-Month Billing Run#
Context: At 00:00 UTC on the 1st, event volume jumps from 20K/s to 140K/s as subscription renewals fire across thousands of tenants. Healthy-endpoint first-attempt p99 goes from 2s to 11 minutes. The on-call escalates to you.
Questions to Surface First:
- Is the delay in ingest (event log lag), fan-out (matcher lag) or dispatch (scheduler ready age)?
- Are workers saturated, or is the scheduler not dispatching?
- Is one tenant dominating, or is it broad?
- Did any endpoint group start failing at the same time (customers' own billing handlers overloaded)?
Typical L5 Approach: Doubles workers. Matcher lag was the bottleneck, so nothing improves for 20 minutes.
Staff Approach: Reads the pipeline stage by stage, finds matcher consumer lag at 9 minutes because subscription lookups miss the cache under the burst, scales the matcher and pre-warms the subscription cache. Confirms the scheduler's tenant fair share kept small tenants' latency bounded once events reached it.
Principal Approach: Treats the 1st of the month as a known, scheduled peak and turns it into a capacity plan: pre-scaling on a calendar, load tests at 10× baseline before each month-end, and a conversation with billing about spreading renewals across the first hour.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Stage metrics: ingest lag 1s, matcher lag 9 min, scheduler ready age 3s. Bottleneck = matcher. Scale matcher partitions' consumers ×4. |
| Triage | Subscription cache TTL 30s; burst of cold tenants → DB lookups at 140K/s → DB saturation → matcher stalls. |
| Quick fix | Raise cache TTL to 5 min with invalidation on subscription change; pre-warm cache at 23:50 on month-end. |
| Guardrails | Alert on matcher.consumer_lag_seconds > 30; dashboard of lag per pipeline stage. |
| Post-mortem | Why did a predictable peak surprise us? Add calendar-based pre-scaling and a monthly load test. |
Metrics to Watch: ingest.lag_seconds, matcher.consumer_lag_seconds, scheduler.oldest_ready_age{healthy}, subs_cache.hit_rate
Organizational Follow-up: billing agrees to jitter renewals across 60 minutes, cutting the peak ~5×.
Ownership Question: "Who owns the month-end peak?" Staff answer: The webhook platform owns capacity for it; billing owns the shape of the burst it generates, agreed via a producer contract.
Key Takeaway: "Find the stage that's slow before adding capacity. Webhook pipelines have three queues, and only one of them is usually the problem."
What clears the Staff bar:
- Localizes the bottleneck to a pipeline stage before acting
- Fixes the cache behavior, not just capacity
- Turns a predictable peak into a plan with the producing team
Deep Dive 2: Silent Failure — Deliveries "Succeeded" With Nothing Sent#
Context: A large customer reports that no order.shipped events have arrived for 36 hours. Your dashboard shows all deliveries to their endpoint as delivered with 200s.
Questions to Surface First:
- Are 200s coming from the customer's server, or from something in between?
- Did anything change in our egress path — proxy, DNS, TLS?
- What do the response excerpts and latencies look like compared to last week?
- Is this one customer or every customer behind a particular proxy?
Typical L5 Approach: Tells the customer the log shows 200s, so the problem is on their side.
Staff Approach: Notices latency dropped from 180ms to 3ms and the response excerpt is an HTML captive page. A proxy config change routed a subset of destinations through a corporate filtering service that returns 200 with a block page. Rolls back, replays 36 hours for affected endpoints.
Principal Approach: Establishes that "200" isn't sufficient evidence of delivery when the path includes infrastructure we operate. Adds end-to-end canaries — synthetic endpoints that verify signatures and report receipt — and treats egress as a production dependency with change review.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Compare latency and excerpt for the customer's attempts this week vs last. 3ms + HTML body → not their server. |
| Triage | Egress proxy change 36h ago added an upstream filter for one destination ASN; 410 endpoints affected. |
| Quick fix | Roll back proxy config. Bulk-replay deliveries for affected endpoints since change time. |
| Guardrails | Canary endpoints in every major cloud and ASN, verifying signature and body; alert on canary receipt gaps > 2 min. |
| Post-mortem | Egress config changes now go through staged rollout with canary verification. |
Metrics to Watch: canary.receipt_gap_seconds, endpoint.latency_p50 shift detector, egress.config_version
Organizational Follow-up: proactive notice to the 410 affected customers with the replay status.
Ownership Question: "Who owns a false 200?" Staff answer: If the 200 came from infrastructure in our path, we do — the delivery log must reflect what the customer's server said, not what a middlebox said.
Key Takeaway: "A 200 is evidence only if you know who sent it. Canary receivers are how you know."
What clears the Staff bar:
- Uses latency and response excerpts as forensic evidence
- Doesn't reflexively blame the customer
- Adds end-to-end verification, not just more logging
Deep Dive 3: Large-Customer Onboarding — 15,000 Events/s to One Endpoint#
Context: A new enterprise customer wants all events — about 15,000 per second at peak — delivered to a single endpoint behind their API gateway, in order per account. Sales has signed. You're asked to make it work.
Questions to Surface First:
- What can their endpoint actually sustain? Latency at that rate?
- Do they need order per account, or will resource versions suffice?
- Would a different transport — delivery into their queue, or a pull API — serve them better?
- What's the blast radius on our shared capacity?
Typical L5 Approach: Raises their endpoint cap to 5,000 and enables strict ordering. Ordering serializes per account, throughput collapses on hot accounts, and the cap consumes shared capacity.
Staff Approach: Measures their endpoint: 300ms p50 at 2,000 concurrent. Proposes per-resource ordering (account-level) with a cap of ~5,000, on dedicated scheduler and worker capacity, plus batching — up to 100 events per request — which cuts request count 50–100×. Offers delivery into their own queue as the long-term option.
Principal Approach: Recognizes a product gap: high-volume customers need a different delivery product. Adds batched delivery and queue destinations as an enterprise tier, priced on reserved capacity.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (design review) | Load test their endpoint with our fake event generator. Capacity: ~6–7K events/s unbatched. |
| Triage | Unbatched, 15K/s at 300ms needs ~4,500 concurrent; feasible but fragile. Batched (100/request), ~150 requests/s. |
| Quick fix | Batched delivery with per-account ordering within and across batches; dedicated worker pool. |
| Guardrails | Their lane is isolated; shared pool impact zero; alert on their lane's backlog separately. |
| Post-mortem (pre-mortem) | What happens when their endpoint fails for an hour? 54M backlog — slow-start drain at batch level. |
Metrics to Watch: lane.enterprise_x.backlog, batch.size_avg, endpoint.latency_p99{x}, resource.blocked_count{x}
Organizational Follow-up: sales engineering gets a capacity questionnaire for future large integrations.
Ownership Question: "Who owns their throughput?" Staff answer: We own delivering at their contracted rate on dedicated capacity; they own an endpoint that can absorb it, verified by a joint load test before go-live.
Key Takeaway: "When one customer needs 15K/s in order, webhooks are the wrong transport unless you batch — and they probably want a queue."
What clears the Staff bar:
- Measures the receiver before committing
- Uses batching and dedicated capacity instead of raising shared caps
- Offers a better-suited transport
Deep Dive 4: Post-Mortem — Credentials Exposed via a Customer Endpoint#
Context: A security researcher reports that registering an endpoint with a hostname that later resolves to a link-local address returned cloud instance metadata in the delivery log's response excerpt. You're leading the post-mortem.
Questions to Surface First:
- Which egress path did these requests take? Why wasn't the proxy blocking them?
- Were credentials actually exposed? Which roles, which permissions?
- How many endpoints have ever resolved to non-public ranges?
- Do workers have any route to internal networks?
Typical L5 Approach: Adds a check at registration that rejects private IPs.
Staff Approach: Finds that a retry path bypassed the proxy for "faster retries" and resolved DNS at send time. Rotates credentials, removes the bypass, moves IP checks to connect time on the pinned resolved address, and hides response excerpts for any non-public destination.
Principal Approach: Treats egress as a security boundary owned by security, with network-level enforcement — workers in a segment with no route except via the proxy — so a code path can't bypass it. Adds the scenario to the pen-test suite and requires session tokens for instance metadata everywhere.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Rotate exposed credentials. Disable the retry fast path. Block the reporter's endpoint pending review. |
| Triage | Retry fast path used a direct HTTP client; DNS rebinding resolved to 169.254.169.254 at send time. |
| Quick fix | All traffic via proxy; proxy pins IPs and blocks private, link-local, loopback, metadata and own VPC ranges. |
| Guardrails | Network policy: workers can reach only the proxy. Alert on any non-proxy egress attempt. |
| Post-mortem | Why did a code-level optimization bypass a security boundary? Make the boundary network-enforced. |
Metrics to Watch: egress.blocked_destination_total, egress.non_proxy_attempts (must be 0), metadata.unauthenticated_requests
Organizational Follow-up: disclosure to affected parties per policy; researcher credited through the bug bounty program.
Ownership Question: "Who owns the egress boundary?" Staff answer: Security owns the threat model and network policy; the webhook platform owns the proxy implementation and its tests.
Key Takeaway: "SSRF defenses in application code get bypassed by the next optimization. Enforce the boundary in the network."
What clears the Staff bar:
- Recognizes registration-time checks are bypassed by DNS rebinding
- Moves enforcement to connect time and to the network layer
- Hides excerpts for non-public destinations
Deep Dive 5: Multi-Region Expansion — EU Data Residency#
Context: The company is launching EU data residency. EU tenants' data must not leave the EU, including event payloads and delivery logs. Webhooks currently deliver from one US region.
Questions to Surface First:
- Where are EU tenants' events produced? Already in an EU region?
- Do delivery logs and response excerpts count as customer data? (Yes — excerpts can contain anything.)
- Will egress IPs change, and how do customers with allowlists find out?
- How are cross-region tenants (US company, EU subsidiary) handled?
Typical L5 Approach: Deploys a second copy of the webhook stack in the EU and routes EU tenants there. Forgets the delivery log is replicated to the US analytics warehouse.
Staff Approach: Region-local everything for EU tenants: event log, scheduler, workers, egress, delivery log and replay storage. Analytics gets aggregated metrics only. New EU egress IPs are published 90 days ahead; tenants migrate with a dual-region delivery window where both IP ranges are valid.
Principal Approach: Defines residency as a tenant attribute honored by every platform that touches customer data, with an audit that proves no EU payload or excerpt exists outside the EU. Webhooks is one of many systems; the standard is company-wide.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Inventory data: payloads, excerpts, delivery logs, replay storage, metrics with tenant IDs. |
| Triage | Delivery log export to US warehouse contains excerpts → residency violation. |
| Quick fix | EU delivery logs stay in EU; warehouse receives aggregates without excerpts. |
| Guardrails | Residency test: synthetic EU tenant with marker payloads; scan US stores for markers weekly. |
| Post-mortem (pre-launch) | Customer comms plan for egress IP change; dual-IP window of 30 days. |
Metrics to Watch: residency.marker_found_outside_region (must be 0), eu.delivery.first_attempt_p99, allowlist.blocked_rate{region}
Organizational Follow-up: legal approves the data inventory; customer success contacts tenants with known allowlists.
Ownership Question: "Who proves residency?" Staff answer: The webhook platform proves it for its data stores with automated marker tests; compliance owns the company-level attestation.
Key Takeaway: "Residency includes the logs. Response excerpts are customer data you didn't know you were storing."
What clears the Staff bar:
- Includes delivery logs and excerpts in the residency scope
- Plans egress IP changes as a customer migration
- Proves residency with tests, not assertions
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why webhook reliability is bounded by the worst receiver without per-endpoint isolation
- Design a scheduler with per-endpoint concurrency caps, circuit breakers (error- and latency-aware) and tenant fair share
- State a complete retry contract: classification, schedule, jitter, final state, disable rule, notifications and replay
- Choose unordered delivery with resource versions, and price opt-in per-resource ordering
- Decide thin vs snapshot payloads per event type, including the API load thin events create
- Sign deliveries with HMAC over timestamp and body, support overlapping secrets during rotation, and enforce SSRF defenses at connect time
- Size workers by concurrent requests (async), and size the delivery log as the largest storage cost
- Handle recovery from long outages with slow-start and customer-controlled drain
The Bar for This Question#
Mid-level (L4): Builds a queue and workers that POST events with retries. Works for well-behaved receivers. No isolation, no end to retries, no signing beyond HTTPS.
Senior (L5): Adds exponential backoff, a DLQ, HMAC signing, Kafka partitioning and autoscaling. The gap: shared worker pool lets one slow endpoint stall everyone; promises ordering via partitions; retries end in an unowned DLQ; registration-time URL checks only. The design works in a demo and fails on the first slow giant.
Staff+ (L6): Frames the problem as isolation against untrusted receivers within the first five minutes. Designs per-endpoint caps, circuits and tenant fair share. States a finite, documented retry contract with visible failure and replay. Refuses a false ordering promise and offers versions. Enforces SSRF defenses at connect time. Gives customers the delivery log so the on-call never debugs a customer's endpoint. Names who pays — failing customers absorb their own delays; healthy customers pay nothing. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 "Guaranteed Delivery" Means "Guaranteed Within a Window — and Then We Tell You"#
| Claim | Reality |
|---|---|
| "We guarantee delivery" | Only within the retry window, to an endpoint that eventually returns 2xx |
| "We never drop events" | You drop them at window end unless failure is visible and replayable |
| "Exactly-once webhooks" | Impossible over HTTP to an uncontrolled receiver |
The Staff position: Promise at-least-once within a stated window, then visible failure and replay. Anything stronger is untestable.
Why this matters in interviews: Accepting "guaranteed delivery" uncritically is a Senior signal; defining its window and end state is Staff.
10.2 Ordering Is Almost Always the Wrong Promise#
| Ordering Mode | Cost |
|---|---|
| Per tenant | One bad event stalls the tenant for days |
| Per resource | Scheduler complexity; still blocks |
| Unordered + versions | Five lines of receiver code |
The Staff position: Put versions in the data and let receivers converge. Offer per-resource ordering only to customers who prove they need it.
Why this matters in interviews: Interviewers ask for ordering to see if you'll pay for it without asking the price.
10.3 The Delivery Log Is the Product#
| Without It | With It |
|---|---|
| Every failure is a ticket | Customers self-diagnose |
| Engineers investigate customer endpoints | On-call never does |
| "You never sent it" is unanswerable | Answered with a link |
The Staff position: A webhook platform without a customer-facing delivery log is half-built.
Why this matters in interviews: It shows you think about operational cost beyond your own pager.
10.4 Autoscaling Makes Slow Endpoints Worse#
| Response | Effect |
|---|---|
| Add workers | Slow endpoint absorbs them |
| Per-endpoint cap | Slow endpoint bounded |
| Latency-aware circuit | Slow endpoint backs off |
The Staff position: Capacity is not the fix for a misbehaving tenant. Isolation is.
Why this matters in interviews: It's the most common wrong answer to the most common follow-up.
10.5 For Internal Consumers, Webhooks Are an Anti-Pattern#
The Staff position: If you control the receiver, let it pull from a log. Push-with-retries is a workaround for not controlling the other side; using it internally reinvents a broker badly, with HTTP timeouts instead of offsets.
Why this matters in interviews: Redirecting an internal-fanout framing to a broker in the first minute is a strong scoping signal.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer builds an isolated, secure, observable webhook platform. The Principal engineer notices that four product teams send webhooks independently — billing with its own retry loop, the repository service with no signing, the marketplace team with a different header format — and that customers integrating with two products see two contracts. The L7 problem is one event delivery contract for the company: a shared platform, a schema standard with resource versions, a published retry and signing contract, and a support model where customers self-serve.
The Org-Level Fault Line#
One delivery platform vs per-product senders.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each product sends its own webhooks | Fast to ship per team | N contracts, N signing schemes, N SSRF surfaces, N support playbooks | Customers (inconsistency), security (surface area), support |
| One platform, products publish events | One contract, one security boundary, one delivery log | Platform must support every product's needs; schema governance | Platform team (scope), producers (schema discipline) |
| Platform plus product-owned schemas with a standard | Consistent delivery; product teams own their event content | Requires schema review and versioning tooling | Platform (tooling), producers (reviews) |
🧭 Principal Move: "The platform owns delivery, signing, egress, retries and the delivery log. Products own event schemas, which must follow the standard — resource ID, monotonic version, no secrets in snapshots. Direct outbound HTTP to customer URLs from product services is blocked at the network after a two-quarter migration."
Cost Model#
Assumptions: async workers; columnar delivery log with compression; object storage for replay payloads at ~$0.02/GB-month; fully loaded engineer ~$250K/year.
| Scale | Volume | Infra ($/month) | Headcount | On-call Load | Notes |
|---|---|---|---|---|---|
| Startup | 500 endpoints, 50 events/s | ~$500–1.5K (job queue, Postgres) | 1 eng part-time | Shared rotation | Buying is often cheaper |
| Growth | 40K endpoints, 20–50K events/s | ~$25–60K (log, scheduler, ~30 workers, proxies, columnar log ~15TB) | 4–6 eng | Dedicated rotation, 2–4 pages/month | Delivery log is the largest line |
| Enterprise | 500K endpoints, 300K events/s, 3 regions | ~$250–500K | 12–20 eng (platform, egress/security, customer tooling) | Per-region rotations | Reserved-capacity tiers fund headroom |
The pricing insight: failure costs more than success. At growth scale, deliveries to failing endpoints and their log rows can be 40–60% of spend while being ~3% of endpoints. Disabling dead endpoints after 72 hours and storing excerpts only for non-2xx attempts are the two highest-ROI changes available.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Delivery semantics promised publicly (at-least-once, unordered) | One-way | Customers build against it; tightening later is expensive, loosening breaks them |
| Signature scheme and header format | One-way | Every receiver's verification code changes |
| Event ID format and stability across retries | One-way | Receivers' dedupe depends on it |
| Published egress IP ranges | One-way-ish | Allowlists break silently; 90-day migrations |
| Retry schedule | Two-way (with notice) | Contract change, announced |
| Scheduler implementation | Two-way | Internal |
| Delivery log storage engine | Two-way | Internal migration |
| Per-plan concurrency caps | Two-way | Config |
The Standard I'd Write#
RFC-EVENTS-002: Outbound Event Delivery Standard
Status: Approved Owners: Event Platform + Security
Scope
Every system that sends HTTP requests to customer-supplied URLs.
MUST
1. Send only through the Event Delivery Platform; direct egress to customer
URLs is blocked at the network.
2. Include a stable event ID, resource ID and monotonic resource version.
3. Exclude secrets and credentials from snapshot payloads; regulated data
uses thin events only.
4. Accept the platform's published delivery contract: at-least-once,
unordered by default, documented retry window, signed requests.
5. Register event types in the schema registry with backward-compatible evolution.
SHOULD
1. Prefer thin events for payloads over 64KB.
2. Coordinate scheduled bursts (billing runs) with the platform team.
3. Provide a reconciliation API for every event type.
Exceptions
Filed with Event Platform; security review required for any egress exception;
time-boxed to two quarters.
Success metrics
- Healthy-endpoint first-attempt p99: < 5s
- Support tickets per million failed deliveries: tracked, trending down
- Products sending webhooks outside the platform: 0 by end of year 2
- SSRF blocks bypassed: 0
What I'd Tell the VP#
"Four teams send webhooks four different ways, and customers who integrate with two of our products get two contracts, two signature schemes and two support experiences. I'm proposing one event delivery platform: it costs about five engineers for a year, eliminates a class of security risk — servers making requests to arbitrary customer URLs — and lets customers debug their own integrations, which should cut webhook-related support tickets substantially. Product teams keep ownership of what their events say; the platform owns how they're delivered. The main risk is migration effort; I'd move the highest-volume product first."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices failure, not success | "Three percent of endpoints drive half the spend; dead-endpoint disabling is the top ROI item." |
| Identifies one-way doors | "The signature format and event ID semantics are forever; the retry schedule isn't." |
| Redraws ownership | "Platform owns delivery; products own schemas; security owns egress." |
| Makes support a design input | "Tickets per million failed deliveries is the KPI I'd hold the platform to." |
| Knows when not to standardize | "Internal consumers don't get webhooks; they get the broker." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a great webhook platform for billing" — correct, but ignores the three other senders.
- "Customers can see failed deliveries in the dashboard" — good, but no metric connects it to support cost.
- "We'll add a second region" — no mention of egress IP migration or residency for delivery logs.
Appendices
Appendix A: Mechanics in Depth#
A.1 The Scheduler Loop#
state:
ready[endpoint] : FIFO of delivery_ids (or per-resource heads if ordered)
due_index : min-heap of (next_attempt_at, delivery_id)
inflight[endpoint] : count, cap[endpoint]
circuit[endpoint] : CLOSED | OPEN(until) | HALF_OPEN
deficit[tenant] : weighted fair-share counter
every 10ms:
move due retries from due_index into ready[endpoint]
for tenant in round_robin(tenants_with_ready_work):
deficit[tenant] += weight[tenant] * quantum
for endpoint in tenant.endpoints_with_ready_work:
while deficit[tenant] > 0 and inflight[endpoint] < cap[endpoint]
and circuit[endpoint] != OPEN and ready[endpoint]:
d = ready[endpoint].pop()
lease(d, ttl=timeout+5s) # persisted before dispatch
dispatch(d); inflight[endpoint] += 1; deficit[tenant] -= 1
on_result(d, outcome):
inflight[d.endpoint] -= 1
record_attempt(d, outcome) # delivery log
update_circuit(d.endpoint, outcome, latency)
if outcome.success: mark_delivered(d)
elif outcome.gone: disable(d.endpoint); mark_failed_final(d)
elif d.age > window: mark_failed_final(d)
else: due_index.push(now + backoff(d.attempts) * jitter(0.8, 1.2), d)
A.2 Backoff Schedule#
attempt: 1 2 3 4 5 6 7 8 9..16
delay: 0 5s 30s 2m 10m 30m 1h 2h every 3h until 72h
jitter: ±20% on each delay; Retry-After honored up to 1h on 429/503
A.3 Circuit Breaker With Latency Awareness and Slow-Start#
window: last 60s, min 20 attempts
open if failure_rate > 50% or p99_latency > 10s
open → wait 30s (doubling to max 10m) → HALF_OPEN: send 1 probe
probe ok → CLOSED with cap = 2, double every 30s while success > 95%, up to configured cap
probe fail → OPEN with longer wait
Appendix B: Data Model#
CREATE TABLE endpoints (
endpoint_id TEXT PRIMARY KEY,
tenant_id TEXT NOT NULL,
url TEXT NOT NULL,
event_types TEXT[] NOT NULL,
ordering TEXT NOT NULL DEFAULT 'none', -- none | per_resource
max_concurrency INT NOT NULL DEFAULT 20,
max_rate_per_s INT, -- customer-set ceiling
state TEXT NOT NULL DEFAULT 'active',-- active | disabled
disabled_reason TEXT,
region TEXT NOT NULL
);
CREATE TABLE endpoint_secrets (
endpoint_id TEXT NOT NULL REFERENCES endpoints,
secret_id TEXT NOT NULL,
secret_enc BYTEA NOT NULL, -- encrypted with KMS
expires_at TIMESTAMPTZ, -- NULL = active indefinitely
PRIMARY KEY (endpoint_id, secret_id)
);
-- deliveries: hot state in the scheduler's store (sharded by endpoint),
-- persisted for durability; attempts go to the columnar delivery log:
-- attempts(tenant_id, endpoint_id, ts, delivery_id, event_id, attempt_no,
-- status_code, error_class, latency_ms, response_excerpt) PARTITION BY day
Appendix C: Coordination Mechanisms#
C.1 Signing and Verification#
C.2 Quick Comparison#
| Mechanism | Guarantees | Failure Mode | Use For |
|---|---|---|---|
| Stable event ID | Receiver can dedupe | Receiver doesn't dedupe | Every event |
| Per-endpoint cap | Bounded blast radius | Cap too high | Every endpoint |
| Latency-aware circuit | Slow endpoints back off | Thresholds too loose | Every endpoint |
| Tenant fair share | No tenant starves others | Weights misconfigured | Scheduler |
| Leased in-flight state | Bounded duplicates on failover | Lease TTL too short | Scheduler HA |
| Resource version | Receivers converge regardless of order | Producers omit it | Event schemas |
| HMAC + timestamp | Authenticity + replay protection | Clock skew; secret leak | Every delivery |
| Egress proxy with pinning | No SSRF | Bypass paths | All egress |
Appendix D: API Contract & Client Behavior#
- Receivers should verify the signature, check the timestamp within 5 minutes, dedupe on the event ID, persist, return 2xx within 10 seconds, and process asynchronously.
- Receivers should treat events as unordered and use
resource_versionto ignore stale updates. - 410 Gone disables the endpoint; any other non-2xx is retried per the schedule.
- Redirects are not followed; register the final URL.
- Secret rotation:
POST /secrets:rotatewith an overlap of up to 24 hours; both signatures are sent during the overlap. - Replay:
POST /endpoints/{id}/replaycreates a job in a separate lane, rate-limited to the endpoint's configured max rate; replayed deliveries carry the original event ID and aWebhook-Replay: trueheader.
Appendix E: Observability#
Core metrics:
delivery.first_attempt_p99{endpoint_health},scheduler.oldest_ready_age{healthy}endpoint.inflight{endpoint},endpoint.circuit_state,endpoint.circuit_flapsdelivery.attempts_per_s{outcome},events.failed_final_total{tenant}matcher.consumer_lag_seconds,ingest.lag_secondsegress.blocked_destination_total,egress.non_proxy_attemptsdlog.ingest_lag_seconds,canary.receipt_gap_seconds
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| Healthy-endpoint first-attempt p99 | > 60s for 5 min | Page |
| Matcher consumer lag | > 30s for 5 min | Page |
| Canary receipt gap | > 2 min | Page |
| Non-proxy egress attempt | > 0 | Sev-1 page (security) |
| Delivery log ingest lag | > 5 min | Ticket → page at 30 min |
| Circuit flaps platform-wide | > 3× baseline | Ticket |
Debugging the silent failure: overall success rate and overall latency are useless — they're dominated by customer endpoints that are always failing somewhere. Watch healthy-endpoint latency, canary receipt and per-endpoint latency shifts (a sudden drop to single-digit milliseconds usually means something other than the customer's server is answering).
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 1K endpoints | Job queue, shared workers, per-endpoint in-flight counter in Redis | First slow giant |
| 1K–50K endpoints | Dedicated scheduler with per-endpoint lists, fair share, async workers, egress proxy | Scheduler state size; delivery log cost |
| 50K–500K endpoints | Scheduler sharded by endpoint hash; columnar log with tiered retention | Cross-shard tenant fairness; multi-region |
| > 500K endpoints | Regional platforms; batched and queue destinations for heavy tenants | Org coordination and contract governance |
What you don't build on day one: per-resource ordering, multi-region delivery, batched delivery, queue destinations, mTLS, a schema registry. Each has a trigger in Section 11.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Fair share by tenant, weighted by plan, work-conserving; fairness by endpoint or by event is gameable.
- Per-endpoint caps by plan: free 5, standard 20, enterprise configurable with reserved capacity.
- Replay lane: replays never share the live lane's capacity; rate-limited to the endpoint's max rate.
- Cost attribution: attempts, retries and delivery-log bytes metered per tenant; heavy-failure tenants are visible in cost reports.
- Dead endpoint policy: disable after 72 hours of zero success; events remain replayable for 30 days — the single largest cost control.