Hiring BarSupport

Design a Notification System — Staff-Level Case Study

Case study60 min read6 diagrams

Technologies referenced in this case study: Kafka · Redis · Cassandra · PostgreSQL · DynamoDB

Related: Message Queues · Rate Limiting · Distributed Job Scheduler · Real-time WebSockets · Chat & Messaging · Degraded Mode · Build vs Buy

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once. Return to individual sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Active Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → two Deep Dives
Deep Dive3+ hrsEverything, including the Principal Lens (Section 11) and Appendices
What is a Notification System? — Why interviewers pick this topic

A notification system takes an event from some product service ("order shipped", "someone replied to you", "your login code is 482913", "spring sale starts now"), decides whether, to whom, on which channel, and when to tell a person about it, renders the message, and hands it to a delivery provider — APNs, FCM, an SMS aggregator, an email service — then tracks what happened.

Before vs After — the campaign-starves-OTP incident:

Without a designed notification platform (one shared queue, one worker pool):
t=0:       Marketing launches a 40M-user push + email campaign at 10:00
t=+30s:    Shared queue depth goes from 2K to 40M; workers drain at 60K/s
t=+1min:   A user requests a login OTP by SMS; it sits behind ~36M campaign messages
t=+10min:  OTP finally sends; it expired after 5 minutes. Login fails.
t=+15min:  Support volume 6× normal: "I can't log in." Password reset emails also delayed.
t=+40min:  Campaign completes. Login success rate recovers. Revenue lost on checkout logins.

With priority-isolated lanes:
t=0:       Same campaign enters the bulk lane, rate-limited to 50K/s
t=+1min:   OTP enters the transactional lane — separate queue, reserved workers, reserved provider quota
t=+1min+3s: SMS handed to provider; p95 OTP delivery 6s as usual
t=+14min:  Campaign completes. No one outside marketing noticed.

Why interviewers reach for this question: The naive design — queue, workers, call APNs — is simple and works in the demo. The real content is everything the demo hides: duplicate sends, stale notifications that arrive after they're useful, third-party providers that throttle and fail, users who get 40 pings a day and uninstall, and fifteen teams who all believe their notification is the important one. It's a question about priority, idempotency, and governance of a shared resource — the user's attention.

Mechanics Refresher: Channels and Providers
ChannelProviderDelivery SemanticsCostLatency
iOS pushAPNs (HTTP/2)Best effort; if device offline, APNs stores one notification per app (the latest)FreeSub-second to seconds
Android pushFCMBest effort; stores messages for offline devices up to a TTL (default 4 weeks)FreeSub-second to seconds
SMSAggregators (e.g., Twilio, Vonage) → carriersCarrier-dependent; delivery receipts partial~$0.005–0.01+ per US segment; far higher internationally1–30s, tail minutes
EmailSES, SendGrid, MailgunAccepted ≠ inboxed; spam filtering; bounces async~$0.10 per 1,000 on SESSeconds to minutes
In-app inboxYour own store + WebSocket/pollYou control it fullyYour inframs when online
Web pushBrowser push servicesBest effort, subscription churn highFreeSeconds

For most production systems: in-app inbox as the durable system of record, push as the primary interruptive channel, email for records and digests, SMS only for high-value transactional messages (OTP, fraud alerts) because it costs real money per message. The channel list is almost never the interview question — priority isolation, dedupe, and who decides whether to interrupt a human are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

A notification system is not a fan-out question. Everyone can put messages in a queue and call APNs.

It is an "attention is a shared, finite resource — who is allowed to spend it, and what happens when delivery fails" question that tests:

  • Whether you isolate priorities so a marketing blast can never delay a login code
  • Whether you understand that a duplicate notification is a user-visible bug and design idempotency end to end
  • Whether you treat freshness as correctness — a "your driver is here" push that arrives 10 minutes late is worse than none
  • Whether you see third-party providers as rate-limited, partially failing dependencies you don't control
  • Whether you define who governs the user's attention across fifteen product teams

The key insight: The system's hardest decision isn't how to send — it's whether to send. Staff candidates design the decision layer (preferences, frequency caps, dedupe, collapse, expiry, quiet hours) as a first-class platform component, and isolate delivery by priority so the rare important message is never behind the common unimportant one.

The L5 vs L6 vs L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws queue → workers → APNs/FCM/SMS/emailAsks "transactional, engagement, or campaign?" and designs priority lanesAsks who governs the user's attention budget across teams and how it's measured
Delivery guarantee"Guaranteed delivery with retries"At-least-once to providers + idempotency key per (user, notification); expiry TTL per typeSets org rule: every notification type declares priority, TTL, dedupe key, and collapse key at registration
PriorityOne queue, maybe a priority fieldSeparate lanes with reserved workers and reserved provider quota; campaigns rate-limitedTreats provider quota as a budget allocated across business units, with finance-visible SMS spend
Failure"Retry on provider error"Per-provider circuit breakers, failover (SMS provider A → B), backoff honoring provider throttles, stale dropMulti-provider contracts negotiated for redundancy; game days for provider outages
User experienceSends every eventPreferences, quiet hours, frequency caps, aggregation ("5 people liked your post")Attention budget as an org metric: notifications/user/day, opt-out and uninstall rate as guardrails every team is judged on
OwnershipPlatform sends what teams askPlatform owns delivery SLO; teams own content, targeting, and their type's opt-out rateCentral review for new notification types; kill switch per type; teams lose sending rights above opt-out thresholds
Why "priority" separates levels

L5: One queue with a priority field, or a priority queue. When 40M campaign messages arrive, the transactional message jumps the queue — but shares the same workers, the same connection pools, and the same provider rate limit. If the provider throttles because of campaign volume, the OTP gets throttled too.

L6: "Priority must be isolated at every shared resource: queue, workers, and provider quota. Transactional has its own topic, its own consumer pool sized for 10× its peak, and a reserved slice of our SMS and push provider throughput. Campaigns run through a token bucket at 50K/s and are the first thing shed. A campaign should be physically incapable of delaying an OTP."

L7: Recognizes provider quota and SMS spend as org resources. "Our SMS contract gives us N messages/s; the transactional lane owns 40% of it permanently. Business units draw from the rest by budget."

Why "delivery guarantee" separates levels

L5: Promises "guaranteed delivery." But APNs is best-effort, FCM is best-effort, SMS delivery receipts are partial, and email "accepted" means nothing about the inbox. The honest guarantee stops at the provider handoff.

L6: "We guarantee at-least-once handoff to a provider within the type's TTL, or an explicit expiry. Retries use the same idempotency key, and APNs apns-collapse-id / FCM collapse_key make on-device duplicates collapse. Beyond the provider, we measure delivery and engagement, but we don't claim to control it. For messages that must be seen, the in-app inbox is the durable record."

L7: Makes the per-type contract explicit and auditable: priority, TTL, dedupe window, collapse key, allowed channels, and fallback channel.

Why "user experience" separates levels

L5: Treats every event as a notification. Users get "Alex liked your photo" 14 times in an hour.

L6: Designs the decision layer: aggregation windows ("Alex and 13 others liked your photo"), per-user frequency caps (e.g., max 3 non-transactional pushes/day), quiet hours by local time, and suppression if the user is currently active in-app.

L7: Owns the org-level tradeoff: every team optimizes its own click-through, and the sum of locally-rational teams is a user who uninstalls. The fix is a shared budget and a central arbiter — like LinkedIn's publicly described notification-volume control system.

The Staff Positions#

PositionRationale
Priority isolation at queue, workers, and provider quotaA shared resource anywhere in the path re-couples campaigns and OTPs
At-least-once handoff + idempotency key per notificationDuplicates are user-visible; retries are unavoidable
Every notification type has a TTLA stale notification is worse than none; expired messages are dropped, not sent
Decision layer before deliveryPreferences, caps, quiet hours, and collapse cost less than sending and apologizing
In-app inbox as the durable recordPush/SMS/email are best-effort; the inbox is where "must be seen" lives
Multi-provider for SMS and email; circuit breaker per providerProviders throttle and fail; the platform must degrade, not stall
Platform owns delivery; teams own content and opt-out rateAttention is shared; teams must be accountable for how they spend it

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Transactional (OTP, password reset, receipts, fraud alerts, ride arrival)Latency and reliability; low volumeDedicated lane, reserved quota, multi-provider failover, short TTLLate OTP = failed login; duplicate receipt = confusionp95 handoff < 2s; p99 delivery < 30s; never duplicated user-visibly
Social / activity ("X replied", "Y liked")Volume spikes, relevance, fatigueAggregation windows, collapse keys, frequency caps, suppress-if-activeSpam → opt-out, uninstallRelevant, not repeated; minutes of delay acceptable
Campaigns / marketingMassive fan-out (10–100M), provider limits, legal complianceAudience streaming, rate-limited send, send-time optimization, unsubscribeStarving other lanes; compliance violationsThroughput and compliance; individual loss acceptable

🎯 Staff Move: "These three intents can't share a pipeline without one starving another. I'll design the platform around priority lanes and a shared decision layer, and I'll go deep on transactional first because it has the strictest bar — then show that campaigns are physically incapable of delaying it."

The Five Fault Lines#

#Fault LineThe Tension
1Duplicate vs Lost NotificationsRetry aggressively (risk double-send) or give up (risk silence)?
2Shared Pipeline vs Priority IsolationOne efficient pipeline or separate lanes with reserved (sometimes idle) capacity?
3Deliver Late vs Drop StaleSend everything eventually, or expire notifications whose moment has passed?
4Send Everything vs Protect AttentionMaximize each team's engagement or cap, aggregate, and suppress for the user's sake?
5Single Provider vs Multi-ProviderSimplicity and volume discounts vs resilience and routing complexity?

In the Wild: Real Production Systems#

Why this section belongs here: Citing specific production systems demonstrates you've studied operational reality, not textbook designs.

LinkedIn — Centralized Notification Volume Control#

LinkedIn has publicly described "Air Traffic Controller" (ATC), a centralized system that decides whether, when, and through which channel to send each notification to a member, applying volume limits and relevance scoring across the many product teams that generate notifications.

Staff insight: ATC is the org-level answer to Fault Line 4: individual teams don't decide to interrupt a member; a shared arbiter does, with a global view of what that member has already received. Cite it when the interviewer asks "who decides?"

Slack — The "Should We Notify?" Decision Tree#

Slack has published a well-known flowchart showing the logic that decides whether a message triggers a notification: channel mute state, user preferences, do-not-disturb, whether the user is active on another device, thread subscriptions, mentions, and more.

Staff insight: The complexity of "should we send?" dwarfs "how do we send?" A candidate who designs the decision layer as a first-class component — with suppress-if-active and DND — is designing the part users actually experience.

Apple APNs and Google FCM — Best-Effort Delivery With Collapse#

APNs documents that when a device is offline it stores only one notification per app (the most recent) and that apns-collapse-id lets a newer notification replace an older one. FCM supports collapse keys and a TTL (default 4 weeks) for offline devices, plus priority levels that affect delivery timing under device power management.

Staff insight: The platforms you depend on are explicitly best-effort and explicitly support collapse and expiry. Designing your own system around "guaranteed delivery" contradicts the providers' contracts; designing around TTLs and collapse keys aligns with them.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Kafka queue and workers""Marketing sends to 50M users. Where is the OTP?"Priority isolation
"Retry on failure""The provider timed out but actually delivered. Now what?"Idempotency and collapse
"Guaranteed delivery""APNs is best-effort. What exactly do you guarantee?"Precise guarantee boundaries
"Send when the event happens""The worker was down for 20 minutes. Do we send 'your driver is here' now?"Freshness / TTL
"Users can set preferences""Fifteen teams each send 'just one' push a day."Attention governance
"Twilio for SMS""Twilio is degraded in one country. OTPs failing."Multi-provider failover

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Every producer goes through one API that enforces the type registry (priority, TTL, dedupe key). The decision layer answers whether and when — preferences, caps, quiet hours, aggregation — before anything is queued for delivery. Lanes are physically separate topics and worker pools with reserved provider quota, so P2 volume cannot touch P0 latency. Channel senders speak provider protocols with collapse IDs and per-provider circuit breakers. The in-app inbox is the durable record; the delivery log tracks each notification's state machine for dedupe, analytics, and debugging.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Architecture"Queue + workers + providers""Decision layer, then isolated priority lanes with reserved workers and reserved provider quota."
Guarantee"Guaranteed delivery""At-least-once handoff within TTL; idempotency key per notification; collapse IDs on device. Inbox is the durable record."
Freshness"Retry until sent""Every type has a TTL. Expired = dropped and counted, never sent late."
Fatigue"User preferences""Preferences + frequency caps + aggregation + quiet hours + suppress-if-active."
Providers"Use Twilio / SES""Two providers per paid channel, per-country routing, circuit breaker per provider, honor 429 backoff."
Ownership"Platform sends for everyone""Platform owns handoff SLO per lane. Teams own content, targeting, and opt-out rate for their types."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
OTP validitytypically 5–10 minSets P0 TTL; late OTP = failed login
P0 handoff SLOp95 < 2s, p99 < 5sThe transactional promise
SMS cost (US)~$0.005–0.01+ per segment; international often 5–10×SMS is the budget line; never use it for engagement at scale
Email cost (SES)~$0.10 per 1,000Email is cheap; deliverability is the constraint
Push costFree at the providerCost is attention, not dollars
APNs offline storage1 notification per app (latest)Collapse is built in; don't rely on backlog delivery
FCM default TTL4 weeksSet shorter TTLs explicitly for time-sensitive messages
Campaign send rate20–100K/s typical, provider-limited50M audience ≈ 10–40 minutes
Aggregation window2–10 min for socialTurns 14 pushes into 1
Frequency cap (non-transactional)~2–5 pushes/user/dayAbove this, opt-out rates climb sharply for most apps
Email bounce rate thresholdkeep hard bounces < ~2%ESPs throttle or suspend senders with high bounce/complaint rates

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on templates, channel adapters, and the preferences schema, then have no time for the only two questions that matter: "Marketing is sending to 50M users — where is the login code?" and "The provider timed out but delivered — did we just send it twice?"


Phase 1: Requirements & Framing (2–3 minutes)#

Functional requirements in 30 seconds:

"Product services send notifications to users over push, SMS, email, and an in-app inbox. Users control preferences. We track delivery state."

Then intent and non-functionals:

"Three different workloads hide under 'notifications': transactional — OTPs, receipts, ride arrivals — low volume, latency-critical; social activity — high volume, relevance and fatigue matter; and campaigns — 50M-user fan-outs where throughput and compliance matter. They can't share a pipeline. I'll design a platform with isolated priority lanes and a shared decision layer, and go deep on transactional correctness first."

Commit to numbers:

"500M users, 100M DAU. ~150K notification-worthy events/s peak, most social; after aggregation and caps, ~30K sends/s; campaigns up to 100M recipients. P0 handoff p95 under 2 seconds regardless of campaign load. No user-visible duplicates. Every type has a TTL — expired notifications are dropped, not sent late."

🎯 Staff Move: "Regardless of campaign load" is the load-bearing phrase. It commits you to physical isolation — not a priority field — before you draw a box.


Phase 2: Core Entities & API (1–2 minutes)#

  • NotificationType (registry): type_id, owner_team, priority (P0/P1/P2), ttl, channels[], fallback_channel, dedupe_window, collapse_key_template, aggregation, counts_toward_cap
  • Notification: notification_id, idempotency_key, user_id, type_id, payload, created_at, expires_at
  • Delivery: notification_id, channel, provider, state (queued → handed_off → delivered | failed | expired | suppressed), attempts
POST /notify  { type_id, user_id | audience_id, idempotency_key, payload, event_time }
  → 202 { notification_id, decision: queued | suppressed(reason) | aggregated }
GET  /users/{id}/inbox?cursor=…
PUT  /users/{id}/preferences  { type_id | category, channels, quiet_hours }

🎯 Staff Move: "idempotency_key is required and supplied by the producer — for an OTP it's the auth request ID; for a reply notification it's the reply's ID. The platform dedupes on (user_id, idempotency_key) within the type's window. A producer retry can't double-send."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the flow in 90 seconds:

  1. Producer calls the API with type and idempotency key; API rejects unregistered types and dedupes
  2. Decision layer loads user preferences and cap counters (Redis), applies quiet hours in the user's time zone, suppresses if the user is active in-app, and either queues, aggregates (holds for a window), or suppresses with a reason
  3. Queued notifications land in the lane for their type's priority — separate Kafka topics and consumer groups
  4. Channel senders resolve device tokens / phone / email, render templates, check TTL, and call providers with collapse IDs, behind per-provider circuit breakers
  5. Delivery log records every state transition; provider receipts and bounces update it asynchronously
  6. The in-app inbox is written for types that have an inbox representation — it's the durable record

Key points to state:

  1. Decide before deliver — the cheapest notification is the one you don't send
  2. Isolation at every shared resource — topics, workers, and provider quota
  3. Idempotency at the API, collapse at the device
  4. TTL checked at dequeue and before provider call
  5. Two SLOs: P0 handoff latency (platform) and opt-out rate per type (owning team)

🎯 Staff Move: Then say: "That's the happy path. The design really gets tested by three things: a campaign during a login spike, a provider that's timing out but still delivering, and fifteen teams competing for the same user's lock screen."


Phase 4: Transition to Depth (1 minute)#

"Three places to go deep: priority isolation under campaign load, duplicate-vs-lost with flaky providers, and attention governance across teams. Which do you prefer?"

Default: priority isolation — it's the most common real incident and the clearest architectural decision.


Phase 5: Deep Dives (25–30 minutes)#

Deep dive A: Priority isolation (7–9 min)

"Isolation must hold at four places. Queue: P0 is its own topic, so its lag is independent of P2's backlog. Workers: P0 has its own consumer pool, sized at 10× its peak — it's ~2K/s, so that's cheap. Provider quota: if our SMS aggregator allows 1,000 msg/s on our account, P0 gets a reserved 400/s through a separate sender token bucket; campaigns can use at most 600/s. Same for APNs connections — P0 has its own HTTP/2 connection pool. Downstream dependencies: template rendering and the device-token store are shared, so P0 calls them with a separate client pool and higher timeout priority. And the campaign lane has a global token bucket at 50K/s — 50M users in ~17 minutes — which is also what our providers can absorb."

Deep dive B: Duplicate vs lost (7–9 min)

"Our request to APNs times out after 10s. Did it deliver? Unknown. If we retry, we might double-send; if we don't, we might lose it. For push I retry with the same apns-collapse-id = notification_id, so if both arrive the second replaces the first on the device — the user sees one. For SMS there's no collapse, so for OTPs I retry only via a status check where the provider supports it, or I send once more with the same code — a duplicate OTP with the same code is harmless, a lost one blocks login. For email, providers accept a message ID; duplicates are rare and cheap. The platform-level dedupe is on (user, idempotency_key) in Redis with a TTL of the type's dedupe window — say 24h for receipts, 10 min for OTP."

Deep dive C: Attention governance (6–8 min)

"Fifteen teams each think their notification is the one that matters. If each sends one push a day, the user gets 15 and turns notifications off — which kills the OTP channel too. So the decision layer enforces a per-user daily cap for non-transactional pushes, say 3, and ranks candidates within the cap by predicted value. Social events aggregate in 5-minute windows. Quiet hours 22:00–08:00 local defer P1/P2. Each type's owner is accountable for its opt-out and disable rate; if a type drives opt-outs above a threshold, it's automatically throttled and the owner is notified."

🎯 Staff Move: Close each dive with ownership: "Platform owns P0 latency and provider failover. The growth team owns the cap values, signed off with product leadership. Each team owns its type's opt-out rate."


Phase 6: Wrap-Up (2–3 minutes)#

"Summary: registry-enforced types with priority, TTL, and dedupe; a decision layer for preferences, caps, aggregation, and quiet hours; isolated lanes with reserved provider quota; idempotency at ingress and collapse at the device; multi-provider SMS/email with circuit breakers; inbox as the durable record. Next I'd build send-time optimization for campaigns and a learned ranking for which notifications win the daily cap. I'd deliberately not build our own SMS carrier connections — aggregators exist for a reason."

Common Timing Mistakes#

MistakeTime LostFix
Template engine and localization design5–8 min"Templates versioned in a registry; rendered in the sender. Standard."
Device token management details5 min"Token table per user; prune on APNs 410 / FCM unregistered."
Detailed preferences UI4 min"Category × channel matrix. Moving on."
Comparing push providers3 min"APNs for iOS, FCM for Android — not a choice."
Never reaching campaign-vs-OTPInterview-endingTransition by minute 12

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Notifications are where product ambition meets user tolerance. Every product team has a growth metric that goes up when they send more notifications — in the short term. Every user has a tolerance that goes down with each one — and when it breaks, they disable notifications for the whole app, including the ones that mattered.

That's the Staff-level content hiding in a queue-and-workers problem: a shared resource with no natural owner (the user's attention and your provider quota), external dependencies you don't control (APNs, carriers, ESPs), and correctness defined by time (a notification's value decays, sometimes to negative). The interviewer is watching for whether you see those three things or just build a fast pipe.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

The L5 path partitions by channel (push queue, SMS queue). The L6 path partitions by priority first, then channel — because the incident that matters is a campaign on the same channel as an OTP.

1.3 The Staff Question That Cuts Through Everything#

"If we could only send one notification to this user today, which one — and who decided?"

It forces:

  • Priority: some notifications are categorically more important
  • Budget: attention is finite and must be allocated
  • Governance: someone other than the sending team decides

2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1: Transactional. Triggered by a user's own action or an event about their account/money/safety: OTP, password reset, payment receipt, fraud alert, "your driver has arrived." Low volume (~1–5% of sends), very high value per message, strict freshness. The correctness bar is fast, exactly-one-visible, fallback on failure. A late OTP is a failed login; a duplicate "payment received" is a support ticket; a missed fraud alert is a loss. These bypass frequency caps and quiet hours (with a narrow allowlist), and may fall back from push to SMS.

Intent 2: Social / activity. Triggered by other users: replies, mentions, likes, follows. High volume, bursty, highly skewed (a popular post generates thousands of like events). The correctness bar is relevant and not repetitive. Minutes of delay are fine — often desired, to aggregate. Frequency caps and quiet hours apply. Suppress when the user is actively in the app.

Intent 3: Campaigns / marketing. Triggered by the business: promotions, re-engagement, announcements. Enormous fan-out, scheduled, legally constrained (unsubscribe links, consent, CAN-SPAM, GDPR, TCPA for SMS in the US). The correctness bar is throughput within provider limits, compliance, and not harming the other lanes. Individual losses are acceptable; sending to someone who unsubscribed is not.

🎯 Staff Move: "Transactional and campaign share channels and nothing else. Transactional needs reserved capacity that sits idle 95% of the time; campaigns need to be throttled even when capacity is available. If I put them in one pipeline, I'll either over-provision everything or starve logins during every campaign."

2.2 When NOT to Build a Notification Platform#

SituationUse InsteadWhy
< 1M users, one product teamManaged services (a customer-engagement platform, Firebase + an ESP)Preferences, templates, and analytics come built in
Only transactional emailAn ESP's transactional API directlyA platform adds latency and failure modes for no benefit
Real-time presence or chat deliveryReal-time WebSockets / ChatIn-session delivery is a connection problem, not a notification problem
Marketing-only needsA marketing automation vendorSegmentation, A/B, and compliance tooling are the product
Internal alerting (on-call pages)PagerDuty / OpsgenieEscalation policies, acknowledgment, and phone-call fallback are the hard part

"I'd build a platform when three or more teams send notifications and we've had our first 'campaign delayed OTPs' or 'users disabled push' incident. Before that, buy."

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It MattersWhat to Say
Mix of intentsDrives isolation design"Transactional, social, and campaigns — isolated lanes."
Delivery guaranteeProviders are best-effort"At-least-once handoff within TTL; inbox is durable."
FreshnessStale can be harmful"Every type has a TTL; expired = dropped."
Duplicate toleranceUsers see duplicates"Zero user-visible duplicates for transactional."
Fatigue controlsOpt-outs kill the channel"Caps, aggregation, quiet hours, suppress-if-active."
ComplianceLegal exposure"Consent and unsubscribe enforced in the decision layer, not by each team."
Multi-deviceOne user, 3 devices"Send to all active tokens; collapse and read-sync via inbox."
Time zonesQuiet hours, send-time"User's local time zone from profile; default from device."

2.4 Precise Terminology#

TermPrecise Meaning
HandoffProvider accepted the request (HTTP 200 from APNs/FCM/SMS/ESP). The platform's controllable boundary.
DeliveryDevice or inbox received it — known only partially (FCM delivery data, SMS DLRs, email non-bounce)
Idempotency keyProducer-supplied ID; platform dedupes (user, key) within the type's dedupe window
Collapse key / IDProvider-level key so a newer notification replaces an older one on device
TTL / expires_atAfter this time the notification is dropped and counted as expired
LaneAn isolated queue + worker pool + provider quota for one priority class
Frequency capMax non-transactional notifications per user per period, across all teams
Aggregation windowTime the platform holds similar notifications to merge them ("Alex and 4 others…")
SuppressionDecision not to send, with a recorded reason (capped, quiet_hours, user_active, opted_out, expired)
Fallback channelAlternate channel if the primary fails within a deadline (push → SMS for OTP)

3. The Fault Lines#

3.1 Fault Line 1: Duplicate vs Lost Notifications#

The tension: Provider calls fail ambiguously. A timeout doesn't tell you whether the message was delivered.

Diagram: 3.1 Fault Line 1: Duplicate vs Lost Notifications
StrategyWhat WorksWhat BreaksWho Pays
At-most-once (never retry)Zero duplicatesEvery provider blip loses messages; OTP failuresUsers who can't log in
Retry blindlyNothing lostDuplicate SMS, duplicate receipts; SMS costs doubleUsers; finance (SMS spend)
Retry with idempotency + collapseDuplicates collapse on device (push) or are harmless (same OTP code)Needs per-channel strategies; SMS has no collapsePlatform — per-channel logic
Status check before retryAvoids duplicates where providers expose statusAdds latency; not all providers support itPlatform

Staff default: Ingress dedupe on (user, idempotency_key); push retries reuse collapse-id; SMS retries reuse the same content (same OTP code) and prefer status-check-then-failover; email retries reuse a stable Message-ID.

"For transactional, a harmless duplicate beats a lost message. I design the duplicate to be harmless — same code, same collapse ID — rather than trying to prevent it."

3.2 Fault Line 2: Shared Pipeline vs Priority Isolation#

The tension: Isolation means reserved capacity that sits idle most of the time. Sharing is efficient until the day it isn't.

StrategyWhat WorksWhat BreaksWho Pays
One queue, FIFOSimplest; max utilizationCampaign starves OTPUsers trying to log in; revenue
One queue with priority fieldTransactional jumps the queueStill shares workers, connection pools, provider rate limitsSame, less often
Separate topics, shared workersQueue lag isolatedWorkers busy on P2 can't pick up P0Same, under load
Full lane isolation (topic + workers + provider quota)P0 latency independent of P2~10–20% extra idle capacity; quota fragmentationPlatform budget — cheap relative to one OTP outage

Staff default: Full isolation for P0; P1 and P2 may share workers with weighted fair scheduling, but P2 has a hard global token bucket and is shed first.

When to deviate: Small systems (< 1K sends/s total) can use one pipeline with a priority queue and a reserved worker slice — just never let P2 exhaust provider quota.

3.3 Fault Line 3: Deliver Late vs Drop Stale#

The tension: "Never lose a notification" sounds right until the 20-minute-late "Your driver has arrived" push lands after the ride started.

TypeValue DecayTTLLate Behavior
OTPUseless after code expiry5 minDrop; user requests a new code
Ride / delivery arrivalHarmful after ~2 min2–5 minDrop; in-app state shows truth
"X replied to you"Declines over hours6–24 hSend; or fold into digest
Payment receiptStable (record)7 daysSend; also in inbox and email
Flash sale startingZero after sale endsSale endDrop
CampaignDeclines over a day24 hDrop after TTL

Staff default: Every type declares a TTL at registration. Senders check now < expires_at at dequeue and immediately before the provider call; provider-side TTL (APNs apns-expiration, FCM ttl) is set to the remaining TTL. Expired messages are counted in notif.expired_total{type} — a rising count is an incident signal, not noise.

"A late notification is not a delivered notification. It's a new bug — the user now sees wrong information on their lock screen."

3.4 Fault Line 4: Send Everything vs Protect Attention#

The tension: Each team's notifications increase its engagement metric; together they decrease everyone's.

Diagram: 3.4 Fault Line 4: Send Everything vs Protect Attention
StrategyWhat WorksWhat BreaksWho Pays
Each team sends freelyTeam autonomy; local metrics upUsers get 15+/day; disable all notificationsEvery team, including transactional
Per-type rate limitsSome protectionNo cross-team view; sum still too highUsers
Global per-user cap + rankingBudget enforced; best candidates winTeams lose notifications to others; ranking disputesTeams with low-value types
Learned send decisions (relevance models, bandits)Optimizes long-term engagementModel ops; explainability to teamsPlatform ML team

Staff default: Global per-user cap for non-transactional push (e.g., 3/day), aggregation windows, suppress-if-active, quiet hours, and a simple value ranking; move to learned ranking when volume justifies it. Duolingo has published research on using bandit algorithms to choose notification content — a public example of treating "which notification" as an optimization problem.

3.5 Fault Line 5: Single Provider vs Multi-Provider#

The tension: One provider per channel is simpler and gets better volume pricing. Providers have regional outages, carrier-specific failures, and throttle limits.

StrategyWhat WorksWhat BreaksWho Pays
Single SMS providerSimple; volume discountProvider or country route outage = OTP outageUsers; revenue
Primary + cold standbyFailover existsStandby untested; sender registration (e.g., US 10DLC, sender IDs) not set up when neededIncident responders
Active-active with per-country routingBest delivery rates per country; instant failoverTwo contracts, two integrations, delivery-rate monitoring per routePlatform team maintains routing

Staff default: Active-active for SMS with per-country routing by measured delivery rate and cost, at least 5–10% of traffic on the secondary to keep it warm. Email: primary ESP + warm secondary with dedicated IPs warmed in advance. Push: APNs and FCM are the only options — isolate by connection pool and honor their backoff.

"A standby provider you haven't sent traffic through in six months is a hypothesis, not a failover."


4. Failure Modes & Operational Reality#

4.1 The Campaign That Starved Logins#

t=0:       Growth launches "spring sale" to 48M users; enters shared push+SMS pipeline
t=+20s:    SMS provider account limit (1,000 msg/s) fully consumed by campaign SMS segment (2M users)
t=+40s:    OTP SMS requests receive 429 from provider; sender retries with 1s backoff
t=+2min:   OTP p95 latency 4 min; OTPs expire at 5 min; login success rate drops from 97% to 71%
t=+8min:   Auth on-call paged for login failures; suspects auth service
t=+22min:  Notification on-call identifies provider quota exhaustion; pauses campaign
t=+25min:  OTP latency recovers

Detection: notif.handoff_latency_p95{lane="P0"} > 5s; provider.throttled_total{provider,lane}; auth.otp_success_rate correlated.

Blast radius: All SMS-dependent logins and password resets.

Mitigation: Reserved provider quota for P0 via separate sender token buckets; campaign SMS on a separate sender ID / account where the provider supports it.

Prevention: Lane isolation to the provider; campaign preflight check that computes required provider throughput and refuses launches that exceed P2's share.

Owner: Notification platform (isolation); growth team (campaign preflight).

4.2 The Duplicate Storm After a Provider Timeout#

APNs responses slow to 12s due to a network issue between our region and Apple; our client timeout is 10s. Every push "fails" at 10s and is retried — but most actually delivered.

t=0:       APNs latency rises to 12s p50
t=+10s:    Timeouts; retry policy re-sends (new request, no collapse-id configured on this type)
t=+30s:    Users receive each social push 2–3 times
t=+5min:   "Why am I getting everything 3 times?" trends on social media

Detection: provider.timeout_rate{provider} spike; notif.duplicate_suspected (same notification_id handed off > 1 time); app-side dedupe telemetry.

Mitigation: Collapse-id = notification_id for every push type (mandatory at registration); retry budget (≤ 10% of requests); circuit breaker opens at 50% timeout rate for 30s instead of retrying into a slow dependency.

Owner: Platform.

4.3 Stale Notifications After a Backlog#

A consumer deploy bug stops the P1 lane for 35 minutes. On recovery, 60M queued social notifications flood out — including "Sam is live now!" for streams that ended 30 minutes ago.

Detection: lane.consumer_lag_seconds{lane} > 60; notif.expired_total should spike on recovery — if it doesn't, TTLs aren't being enforced.

Mitigation: TTL enforcement at dequeue; drain backlog at a controlled rate (backlog replay is itself a herd for providers and for the app's backend when users tap).

Owner: Platform (TTL enforcement); each type owner (setting a correct TTL).

4.4 Token Rot — Pushing Into the Void#

Device tokens become invalid when users uninstall or reinstall. APNs returns 410 Unregistered and FCM returns UNREGISTERED; if we don't prune, 30% of push requests go to dead tokens, inflating cost, skewing delivery metrics, and in some cases hurting sender reputation.

Detection: push.invalid_token_rate > 5%; "sent" rising while opens flat.

Mitigation: Delete tokens on 410/UNREGISTERED; expire tokens not refreshed by the app in 60 days.

Owner: Platform.

4.5 Email Reputation Collapse#

A team imports an old mailing list and sends a campaign; 9% hard bounces and elevated spam complaints. The ESP throttles the whole account — including password reset emails.

Detection: email.bounce_rate{type} > 2%; email.complaint_rate > 0.1%; ESP account status.

Mitigation: Separate sending domains/IPs (and ideally separate ESP subaccounts) for transactional vs marketing; list hygiene required before campaigns; auto-pause campaign at bounce > 3%.

Owner: Growth (list hygiene); platform (domain separation and auto-pause).

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Campaign starves P0handoff_latency_p95{P0} > 5s; provider 429sLogins, resetsReserved quota; campaign preflightPlatform + growth
Provider timeouts → duplicatesduplicate_suspected; timeout rateAll push usersCollapse IDs; retry budget; breakerPlatform
Backlog → stale floodLane lag; expired_total flat on recoveryUsers, app backendTTL at dequeue; controlled drainPlatform + type owners
Dead tokensinvalid_token_rate > 5%Cost, metricsPrune on 410; 60-day expiryPlatform
Email reputation hitBounce > 2%, complaints > 0.1%All email incl. transactionalSeparate domains/IPs; auto-pauseGrowth + platform
SMS provider country outagesms.delivery_rate{country,provider} dropOTPs in that countryPer-country failover routingPlatform
Preferences store downDecision layer errorsAll non-transactionalFail-closed for campaigns (don't send); fail-open to defaults for transactionalPlatform
Over-notification by one teamoptout_rate{type} > thresholdApp-wide notification opt-outAuto-throttle type; reviewOwning team
Quiet-hours timezone bugComplaints; sends at 03:00 localUsers in one regionTimezone tests; send-time auditPlatform

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
ArchitectureQueue per channel + workersDecision layer + priority lanes isolated through provider quotaPlatform with type registry, governance, and chargeback across business units
Guarantee"Guaranteed delivery"At-least-once handoff within TTL; idempotency + collapse; inbox as durable recordPer-type contracts audited; transactional delivery SLO reported to leadership
FreshnessRetry until sentTTL per type enforced at dequeue and providerValue-decay model per type informs TTL defaults org-wide
FatiguePreferencesCaps, aggregation, quiet hours, suppress-if-activeAttention budget as org metric; teams accountable for opt-out rate; central arbiter
ProvidersOne per channelMulti-provider SMS/email, per-country routing, breakers, warm standbyContract strategy, sender registration, cost per delivered message by country
ComplianceUnsubscribe linkConsent enforced centrally in decision layerLegal-owned consent model; audit trail; regional regulation map

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Isolation to the provider"Separate topic isn't enough — P0 needs reserved provider quota or the campaign still throttles OTPs."
Honest guarantee"We control handoff, not delivery. APNs is best-effort; the inbox is the durable record."
Freshness as correctness"A late 'driver arrived' is a new bug. Every type has a TTL."
Harmless duplicates"Retry the OTP with the same code; push retries reuse the collapse ID."
Attention governance"Fifteen teams × one push a day = uninstall. The cap is global and owned by product leadership."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Partition by channel onlyCampaign and OTP still share the push and SMS paths
"Guaranteed delivery"Contradicts provider contracts; no guarantee boundary
Unbounded retriesDuplicates, cost, and stale sends
No decision layerDesigns a pipe, not a notification system
Each team handles its own unsubscribeCompliance risk; inconsistent enforcement

5.4 Common False Positives#

  • Deep APNs/FCM API knowledge ≠ system design. Knowing HTTP/2 headers doesn't answer priority isolation.
  • Kafka partition math ≠ Staff. Throughput is rarely the bottleneck; providers and attention are.
  • Fancy ML ranking ≠ governance. A model that picks the best notification still needs an owner for the cap and an override for transactional.
  • "We'll use WebSockets" ≠ notifications. In-session delivery doesn't reach users who aren't in the app — which is the point of notifications.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minThree intents; P0 isolated "regardless of campaign load"; TTL per type
Entities + API3–5 minType registry, Notification, Delivery; required idempotency key
High-level design5–12 minAPI → decision layer → lanes → senders → providers → delivery log
Transition12 minOffer: isolation, duplicates, governance
Deep dives12–38 minIsolation → duplicates/providers → fatigue/governance → freshness
Org & evolution38–43 minOwnership split; compliance; build vs buy
Wrap-up43–45 minSummary; what's next; what you won't build

6.2 How Interviewers Pivot — And What They're Testing#

Interviewer PivotWhat They're TestingWhere to Go
"Send to all 100M users now."Fan-out and provider limitsAudience streaming, token bucket, send-time spread
"The SMS provider is down in India."Dependency failurePer-country routing; failover; warm secondary
"Users complain about too many notifications."Product judgmentCaps, aggregation, governance
"User has 3 devices."Multi-device semanticsSend to all active tokens; inbox read-sync clears badges
"How do you know it was delivered?"Guarantee honestyHandoff vs delivery; receipts; engagement proxies
"A user unsubscribed yesterday and got a campaign today."ComplianceCentral consent check at send time, not audience build time
"What if the notification service is down?"Degraded modeProducers buffer; P0 fallback path; inbox writes

6.3 What to Deliberately Skip#

TopicWhy L5 Goes HereWhat L6 Says Instead
Template engineConcrete"Versioned templates, rendered in senders, localized by user locale."
Rich push mediaFeature-y"Payloads ≤ 4 KB; media via URL."
Detailed analytics pipelineFamiliar"Delivery log to the warehouse; standard."
Preference UI matrixEasy"Category × channel. Moving on."

6.4 Follow-Up Questions to Expect#

  1. "How do you implement 'Alex and 4 others liked your post' without delaying the first notification too long?"
  2. "How do you send at 9 a.m. local time to 100M users across 30 time zones?"
  3. "What happens to queued notifications for a user who deletes their account?"
  4. "How would you roll out a new notification type safely?"
  5. "What's the cost per month of SMS OTPs at 20M logins/day, and how would you cut it?"
  6. "How do you handle badge counts across devices?"
  7. "How do you ensure a fraud alert reaches the user if push fails?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a notification system."

Staff Answer

"Three workloads hide under that name: transactional (OTP, receipts, arrivals), social activity (replies, likes), and campaigns (50M-user blasts). They need physical isolation — a campaign must be incapable of delaying an OTP. And the hardest part isn't sending, it's deciding whether to send: preferences, frequency caps, aggregation, quiet hours, and TTLs so we never send stale notifications. I'll design a type registry, a decision layer, isolated priority lanes with reserved provider quota, and idempotent channel senders, then go deep on isolation, duplicates, and governance."

Why this is L6:

  • Separates intents by correctness bar
  • Commits to isolation and the decision layer before drawing
  • Names freshness as correctness

What L7 adds:

  • Frames attention as an org-level budget and asks who owns the cap
  • Asks about SMS spend and provider contracts as business constraints
❌ Common L5 Trap

"Services publish to Kafka; we have a push worker, an SMS worker, and an email worker that call APNs/FCM, Twilio, and SES, with retries and a preferences check."

Why this misses: Partitioned by channel, not priority — the campaign SMS and the OTP SMS share a worker and a provider quota. No TTL, no dedupe, no caps.


Drill 2: Core Mechanic — Aggregation#

Prompt: "A post gets 500 likes in 10 minutes. How many notifications does the author get?"

Staff Answer

"Ideally 1–2. First like: send immediately if the user hasn't received a like notification for this post in the last hour — speed matters for the first one. Subsequent likes enter a 10-minute aggregation window keyed by (user, post, 'like'). At window close, send one push: 'Alex, Sam and 498 others liked your post,' with collapse-id = user:post:likes so it replaces the earlier push on the lock screen. Windows back off exponentially — 10 min, 30 min, 2 h — while activity continues. The inbox shows a single aggregated item updated in place. State lives in Redis: a hash per aggregation key with count and first N actors, TTL = window + slack."

Why this is L6:

  • Balances immediacy (first like) with fatigue (rest aggregated)
  • Uses collapse IDs to replace, not stack
  • Concrete state and keys

What L7 adds:

  • Makes aggregation policies part of the type registry so every social product gets them by default, not by reinvention

Drill 3: "Priority Isolation" — Make It Concrete#

Prompt: "You said P0 is isolated. Show me exactly where campaign traffic could still interfere."

Staff Answer

"Walk the path: API — shared, so P0 has a reserved concurrency slice and campaigns enter through a separate bulk endpoint with its own rate limit. Decision layer — P0 transactional types skip caps and use a fast path with only an opt-out check. Kafka — separate topic, separate consumer group. Workers — separate deployment, autoscaled on P0 lag only. Template rendering and device-token lookups — shared services, so P0 uses a separate client pool with a 50ms timeout and cached tokens. Provider — reserved quota via a separate token bucket; for SMS a separate sender ID/account; for APNs a separate HTTP/2 connection pool. Remaining shared risk: the provider itself degrading globally — covered by failover to provider B for SMS."

Why this is L6:

  • Audits every shared resource, not just the queue
  • Names the remaining residual risk and its mitigation

What L7 adds:

  • Runs a quarterly "campaign during login peak" game day and publishes P0 latency during it

Drill 4: Dependency Down — Preferences Store Unavailable#

Prompt: "The preferences store is down. Do you send notifications?"

Staff Answer

"It depends on the lane, and I'd decide it now, not during the incident. Campaigns: fail-closed — sending marketing to someone who opted out is a legal problem; campaigns pause. Social: fail-closed for push (hold in queue within TTL), still write to inbox. Transactional: fail-open to the default policy — OTPs and fraud alerts are allowed regardless of marketing preferences, and a user can't opt out of their own login code. Decision layer keeps a local cache of preferences for recently active users (5-minute TTL) to cover short blips."

Why this is L6:

  • Different fail modes per intent, justified by harm
  • Legal constraint drives campaigns to fail-closed

What L7 adds:

  • Has legal pre-approve the degraded-mode matrix and records it in the runbook

Drill 5: Hot Key — The Celebrity Fan-Out#

Prompt: "A celebrity with 80M followers posts. Every follower gets a notification."

Staff Answer

"First, should every follower get one? Probably only followers who opted into 'posts from X' — maybe 5–10%. For those, one event fans out to millions of per-user notifications. Don't expand synchronously: the fan-out service streams the follower list in pages of 10K into the P1 lane at a bounded rate (e.g., 100K/s), so 8M notifications take ~80s. Each still goes through the decision layer — cap, quiet hours. TTL of ~2 hours: if the pipeline is behind, late followers are dropped rather than notified hours after. The app backend also needs protection: 8M notification taps in 2 minutes hit the post page, so the post is pre-warmed in cache."

Why this is L6:

  • Questions whether fan-out is desired
  • Bounded streaming fan-out, TTL, and downstream herd awareness

What L7 adds:

  • Coordinates with the feed team — Feed Generation faces the same celebrity fan-out and should share the follower-streaming infrastructure

Drill 6: Multi-Tenant — Fifteen Teams, One Lock Screen#

Prompt: "Fifteen teams send notifications. Opt-out rates are climbing. What do you do?"

Staff Answer

"Measure attribution first: opt-out and disable events are joined to the last N notifications the user received, giving an opt-out rate per type. Then: a global per-user daily cap on non-transactional push (3/day), with candidates ranked by type priority and predicted engagement; aggregation defaults; and a per-type guardrail — a type whose opt-out rate exceeds 2× the median is auto-throttled to 50% and its owner is notified. New types launch at 1% of eligible users and graduate after a week of healthy metrics. Product leadership owns the cap value; the platform enforces it."

Why this is L6:

  • Attribution before policy
  • Mechanisms that make teams accountable
  • Cap owned by leadership, enforced by platform

What L7 adds:

  • Makes "notification opt-out rate" a company-level health metric reviewed alongside retention

Drill 7: Build vs Buy#

Prompt: "Why not buy a notification platform (a customer engagement vendor) instead?"

Staff Answer

"For campaigns, I'd strongly consider buying — segmentation, journeys, A/B testing, and compliance tooling are the vendor's product. For transactional, the calculus is different: OTP latency and reliability are core to login, and I want control over provider routing and failover. A common split: build a thin transactional + social platform (decision layer, lanes, senders), buy campaign orchestration, and have the vendor send through our platform's P2 lane — or at least through separate provider accounts so their volume can't hurt our P0 quota. Build cost is ~4–6 engineers ongoing; vendor cost scales with MAUs or message volume and can reach six figures a year."

Why this is L6:

  • Splits the decision by intent
  • Protects P0 isolation even with a vendor in the path

What L7 adds:

  • Negotiates data-export and exit terms so audience and preference data isn't locked into the vendor

Drill 8: Policy Change Without Outage — Changing the Frequency Cap#

Prompt: "Leadership wants to lower the daily push cap from 5 to 3."

Staff Answer

"Shadow first: compute what would have been suppressed under cap 3 for a week and show each team its projected loss. Then A/B: cap 3 for 5% of users, measure DAU, sessions, opt-out rate, and each type's downstream conversion for 2 weeks. Roll out 5 → 25 → 100% with a flag. Teams who lose volume get the ranking signals so they can improve relevance. The cap is config with versioning; rollback is instant. Transactional types are unaffected."

Why this is L6:

  • Shadow → experiment → staged rollout
  • Transparent per-team impact

What L7 adds:

  • Uses the experiment to set a long-term attention budget policy, not a one-off number

Drill 9: Cost#

Prompt: "SMS costs $1.2M a year. Cut it."

Staff Answer

"Break down by type and country first. Typically OTPs dominate, and international routes cost 5–10× US. Levers: (1) move OTP to push-based approval or app-generated codes for users with the app installed — SMS only as fallback; (2) WhatsApp or other OTT channels in countries where they're cheaper and more reliable; (3) route per country by cost-adjusted delivery rate; (4) kill SMS for any non-transactional type; (5) fraud: block SMS pumping — attackers trigger OTPs to premium numbers — with per-country and per-prefix rate limits. SMS pumping alone can be a large share of spend. Realistic saving 40–60%."

Why this is L6:

  • Breaks cost down before cutting
  • Names SMS pumping fraud, a real and commonly missed driver

What L7 adds:

  • Moves auth strategy toward passkeys/push approval as the long-term cost and security fix — coordinating with the identity team

Drill 10: Multi-Region#

Prompt: "We're going multi-region. How does the notification platform change?"

Staff Answer

"Notifications are user-homed: process in the user's home region, where their preferences, tokens, and inbox live. Producers in any region route to the user's home region. APNs and FCM are global endpoints, so any region can send. Dedupe must be in the home region to be consistent. On home-region failure: P0 fails over to a secondary region with a replicated subset — tokens and phone numbers for transactional — accepting that caps and aggregation state may be stale (it's fine to exceed a cap during a regional outage; it's not fine to lose an OTP). P1/P2 wait for recovery within TTL."

Why this is L6:

  • Home-region processing keeps dedupe and caps consistent
  • Different failover posture per lane

What L7 adds:

  • Checks residency constraints for phone numbers and emails before replicating them

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — Black Friday#

Context: Black Friday, 09:00. Three campaigns launched at once (18M, 22M, 30M recipients). Order-confirmation emails are 25 minutes late; customer support is flooded with "did my order go through?"

Questions to Surface First:

  • Are order confirmations in the P0 lane, and is email P0 isolated at the ESP level?
  • Which resource is saturated — our workers, the ESP account throughput, or ESP-side throttling?
  • Who approved three simultaneous campaigns?

Typical L5 Approach: Scales email workers. The bottleneck is the ESP account rate, so more workers produce more 429s.

Staff Approach: Finds order confirmations and campaigns share one ESP account; the ESP is throttling the account. Pauses two campaigns, reserves throughput for transactional, and moves order confirmations to the warm secondary ESP immediately. Afterward: separate ESP subaccounts/IPs for transactional vs marketing, and a campaign scheduler that enforces aggregate P2 throughput.

Principal Approach: Establishes a campaign calendar with a throughput budget per hour on peak days, owned by marketing ops, and makes "transactional isolated at the provider account level" a platform invariant.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Check handoff_latency{lane,channel} and provider.throttled_total{provider}.
TriageEmail P0 lag 25 min; ESP 429s at account level; three campaigns active.
Quick fixPause two campaigns; failover order confirmations to secondary ESP.
GuardrailsSeparate transactional subaccount; aggregate P2 throughput cap; campaign preflight.
Post-mortemWhy was provider-level isolation missing for email? Why three simultaneous launches?

Metrics to Watch: notif.handoff_latency_p95{lane,channel}, provider.throttled_total, campaign.active_count, support.contacts{topic="order_status"}

Organizational Follow-up: Marketing ops owns a peak-day calendar; platform owns isolation and preflight.

Ownership Question: "Who can pause a revenue campaign on Black Friday?" Staff answer: The notification on-call, under a runbook that says transactional latency beats campaign throughput, with automatic notification to marketing ops. Pre-agreed, not negotiated live.

Key Takeaway: "Isolation that stops at your queue isn't isolation. The provider account is a shared resource."

What clears the Staff bar:

  • Finds the shared provider account, not the worker count
  • Uses a warm secondary provider
  • Pre-agreed authority to pause campaigns

Deep Dive 2: Silent Failure — Android Pushes Not Arriving#

Context: Engagement from Android push has dropped 30% over 3 weeks. No errors; FCM returns success for every request.

Questions to Surface First:

  • Did app code change how notifications are displayed or how tokens refresh?
  • Did we change message priority or TTL?
  • Are opens down uniformly or on specific OS versions/manufacturers?

Typical L5 Approach: Checks error logs, finds none, concludes FCM is working and the drop is product-side.

Staff Approach: Treats "success but no engagement" as a delivery issue until proven otherwise. Finds a change 3 weeks ago that set all social pushes to normal priority with a 60-second TTL; devices in battery-optimization states batch normal-priority messages, and a short TTL expires them before they're delivered. Fixes TTL to type-appropriate values and uses high priority only for time-sensitive user-visible types (per FCM guidance, overuse of high priority can be deprioritized).

Principal Approach: Adds end-to-end delivery observability — client-side receipt telemetry sampled at 1% — so "provider accepted" is never mistaken for "delivered" again, and puts send-parameter changes under review.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateCompare open rate by platform, OS version, priority, and TTL setting.
TriageDrop isolated to Android, correlates with config change on day 0.
Quick fixRevert TTL; set priority per type.
GuardrailsClient receipt telemetry; alert on push.receipt_rate{platform} drop > 10%.
Post-mortemWhy was a send-parameter change unreviewed? Why no delivery-level metric?

Metrics to Watch: push.receipt_rate{platform,os}, push.open_rate{type,platform}, notif.handoff_success_rate

Organizational Follow-up: Send parameters (priority, TTL) owned in the type registry with review.

Ownership Question: "Who owns 'delivered' when the provider says success?" Staff answer: The platform owns measuring it end to end via client receipts. Provider success is an input, not the SLO.

Key Takeaway: "Provider success is handoff, not delivery. Measure from the device, or you'll find out from the engagement graph three weeks later."

What clears the Staff bar:

  • Distrusts provider success as a proxy
  • Segments metrics to isolate the cause
  • Adds device-side measurement

Deep Dive 3: Large-Customer Onboarding — B2B Tenant With Compliance Requirements#

Context: A bank signs up to use your notification platform (you're a B2B SaaS) to send 5M transaction alerts/day by SMS and push. Requirements: 99.9% of alerts handed off within 10s, full audit trail for 7 years, and no alerts during their customers' quiet hours unless fraud-related.

Questions to Surface First:

  • Which alerts are fraud (bypass quiet hours) vs informational?
  • Does "audit trail" mean content, or metadata only?
  • Whose SMS sender identity and registration — ours or theirs?

Typical L5 Approach: Adds the tenant to existing lanes and raises provider limits.

Staff Approach: Puts the bank in a dedicated P0 tenant lane with reserved provider quota and its own sender registration; types classified as fraud vs informational with quiet-hour bypass only for fraud; audit log of metadata (not content) to immutable storage with 7-year retention; per-tenant SLO dashboards.

Principal Approach: Uses the bank as the forcing function for tenant isolation tiers (shared / reserved / dedicated) with pricing, and for a compliance package (audit retention, residency) sold as a product tier.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignTenant lane; reserved SMS throughput 100/s + burst; separate APNs/FCM credentials (theirs).
ComplianceImmutable audit log (object lock), metadata only; content hash for dispute verification.
SLO99.9% handoff < 10s; measured per tenant; credits in contract.
TestingLoad test at 3× daily peak; provider failover drill before go-live.
CostReserved capacity priced into contract; SMS passed through at cost plus margin.

Metrics to Watch: notif.handoff_latency{tenant}, audit.write_failures{tenant}, quiet_hours.bypass_total{tenant,type}

Organizational Follow-up: Legal review of audit/retention; security review of credential handling.

Ownership Question: "Who classifies an alert as fraud to bypass quiet hours?" Staff answer: The bank, via type registration. We enforce the classification and log every bypass; we don't decide what's fraud.

Key Takeaway: "Large tenants turn implicit shared resources into explicit contracts. Price the isolation."

What clears the Staff bar:

  • Tenant-level isolation and SLOs
  • Separates metadata audit from content retention
  • Tenant owns classification

Deep Dive 4: Post-Mortem — 2.3M Users Got the Wrong Notification#

Context: A template bug sent 2.3M users a push saying "Your payment of $0.00 failed." Support is overwhelmed; the brand team is involved.

Questions to Surface First:

  • How did a broken template reach production at full volume?
  • Was there a kill switch, and how fast was it used?
  • Can we send a correction, and should we?

Typical L5 Approach: Fixes the template, adds a unit test.

Staff Approach: Identifies three missing controls: templates deployed without preview against real payload samples; new/changed types sent at 100% immediately; no per-type kill switch (took 18 minutes to stop). Adds template validation with required-field checks, staged rollout for template changes (1% → 10% → 100% over an hour with anomaly checks on opens and complaints), and a one-click per-type kill switch that stops dequeue within 10s.

Principal Approach: Makes notification content changes follow the same change-management as code deploys, and defines a correction-message policy (who approves apology sends — comms, not engineering).

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateKill switch the type; purge queued messages of that type.
TriageTemplate variable amount missing from new payload schema; rendered default 0.
CorrectionComms decides on a follow-up; in-app inbox item updated in place with correct info.
GuardrailsRender-validation against sampled payloads; staged rollout; kill switch SLO 10s.
Post-mortemSchema change in producer wasn't contract-tested against the template.

Metrics to Watch: notif.sent{type,template_version}, render.missing_field_total, killswitch.time_to_stop_seconds

Organizational Follow-up: Producer-template contract tests; comms runbook for correction messages.

Ownership Question: "Who decides whether to send a correction to 2.3M users?" Staff answer: Comms and the product owner of the type. A correction is itself a notification with fatigue and brand costs — it's not an engineering decision.

Key Takeaway: "Notification content is a production deploy to millions of lock screens. It needs staged rollout and a kill switch."

What clears the Staff bar:

  • Finds missing controls, not just the bug
  • Kill switch with a time SLO
  • Correction decision routed to comms

Deep Dive 5: Multi-Region Expansion — Launching in the EU and India#

Context: The product launches in the EU and India. EU requires consent tracking and residency for personal data; India's SMS routes require sender registration and templates pre-approved under local regulation (DLT). OTP delivery rates in India on the current provider are 82%.

Questions to Surface First:

  • Where do EU users' tokens, phone numbers, and inbox live?
  • Who owns the regulatory registration work per country?
  • What's an acceptable OTP delivery rate, and what's the fallback?

Typical L5 Approach: Deploys the same stack in two new regions and uses the existing SMS provider.

Staff Approach: EU users homed in EU region (preferences, tokens, inbox, delivery log); consent stored with provenance and enforced in the decision layer. India: add a local SMS provider with registered templates; route per operator by measured delivery rate; add a WhatsApp or voice fallback for OTP if SMS fails in 30s. Target OTP delivery ≥ 95%.

Principal Approach: Creates a country-launch checklist owned jointly by legal, platform, and growth (sender registration, consent model, local providers, quiet-hour norms) so each new country isn't a bespoke project.

Staff Approach — Full Reasoning
PhaseWhat to Do
ResidencyEU home region; only aggregates leave the EU.
RegulationTemplate pre-registration for India; consent provenance for EU.
ProvidersLocal provider in India; per-operator routing; fallback channel.
Measuresms.delivery_rate{country,operator,provider}; OTP completion rate by country.
RolloutLaunch with fallback enabled; shadow-route 10% to compare providers for 2 weeks.

Metrics to Watch: otp.completion_rate{country}, sms.delivery_rate{country,provider}, consent.missing_total

Organizational Follow-up: Legal owns the regulatory map; platform owns routing; growth owns local campaign rules.

Ownership Question: "Who is accountable for OTP delivery rate in India?" Staff answer: The platform owns the rate and the routing. Legal owns the registration dependency. Both are on the launch checklist with named owners.

Key Takeaway: "Notification delivery is local: carriers, regulations, and norms differ by country. Build a launch checklist, not a one-off."

What clears the Staff bar:

  • Residency and consent as design inputs
  • Per-country provider routing with measured delivery
  • Fallback channel for critical OTP

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Separate transactional, social, and campaign intents and commit to isolating them
  • Design a decision layer (preferences, caps, aggregation, quiet hours, suppress-if-active) as a first-class component
  • Isolate priority lanes at every shared resource — queue, workers, dependencies, and provider quota
  • State an honest guarantee: at-least-once handoff within TTL, with idempotency and collapse making duplicates harmless
  • Treat freshness as correctness with per-type TTLs enforced at dequeue and at the provider
  • Design multi-provider routing with circuit breakers and warm failover
  • Assign ownership: platform owns delivery SLOs; teams own content, targeting, and opt-out rates; legal owns consent

The Bar for This Question#

Mid-level (L4): A queue and workers per channel calling providers, with retries and a preferences table.

Senior (L5): Adds Kafka, templates, device token management, a priority field, and delivery tracking. Competent and scalable, but partitions by channel, promises guaranteed delivery, and leaves fatigue and freshness implicit.

Staff+ (L6): Designs the decision of whether to send, isolates priorities down to provider quota, makes duplicates harmless and stale messages impossible, plans for provider failure, and makes attention a governed resource with named owners. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 The Best Notification System Sends Fewer Notifications#

MetricShort-term Effect of More SendsLong-term Effect
SessionsUpDown after opt-outs
Push opt-in rateFlatDown — the channel degrades for everyone
UninstallsFlatUp

The Staff position: Suppression is a feature. A platform should be judged partly by what it didn't send.

Why this matters in interviews: It shows you understand the system's real objective, not just its throughput.

10.2 "Guaranteed Delivery" Is a Lie Your Providers Won't Tell You#

The Staff position: Every push provider is best-effort; SMS depends on carriers; email "accepted" says nothing about inboxes. The honest contract is handoff-within-TTL, measured delivery, and a durable inbox.

Why this matters in interviews: Precision about guarantees is a Staff marker.

10.3 Partition by Priority, Not by Channel#

The Staff position: The canonical diagram has one queue per channel. The incident that matters is a campaign and an OTP on the same channel. Priority is the primary partition; channel is secondary.

Why this matters in interviews: It reframes the architecture in one sentence.

10.4 SMS Is a Security and Fraud Surface, Not Just a Channel#

The Staff position: SMS OTP is vulnerable to SIM-swap and is a target for SMS-pumping fraud that runs up costs. It should be the fallback, not the default, for users with the app installed.

Why this matters in interviews: It connects notifications to security, cost, and identity strategy.

10.5 Every Notification Type Needs an Owner Who Can Be Embarrassed#

The Staff position: Types without an owning team and an opt-out metric attached accumulate forever. Registration requires an owner; types with no sends in 90 days are retired.

Why this matters in interviews: Lifecycle governance is what makes a platform sustainable over years.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

At Staff level, the notification system is a platform with lanes and a decision layer. At Principal level, it's the arbiter of a shared, depletable company asset: users' willingness to be interrupted. Every product team draws on it, no team pays for depleting it, and once a user disables notifications, every team — including security and payments — loses the channel. The L7 job is to make that asset visible, budgeted, and governed; to treat provider contracts and SMS spend as financial decisions; and to connect notifications to identity strategy (passkeys vs SMS OTP), compliance (consent across jurisdictions), and brand.

The Org-Level Fault Line#

Central arbitration of attention vs team autonomy over sending.

OptionWhat It BuysWhat It Costs
Central arbiter (global caps, ranking, review of new types)Protects the channel; consistent UX; complianceTeams feel throttled; ranking disputes escalate to leadership
Team autonomy with guardrails (per-type opt-out thresholds, no global cap)Speed for teamsSum-of-locally-rational sends still degrades the channel
Full autonomyMaximum team velocityTragedy of the commons; opt-out spiral

The Principal position: Central arbiter for push and SMS (interruptive, scarce, costly), guardrails-only for inbox and email (non-interruptive). Product leadership owns the cap; the platform enforces; teams compete on relevance.

Cost Model#

Assumptions: US-heavy user base, SMS ~$0.008/segment average (international higher), email ~$0.10/1,000, engineer fully loaded ~$250K/year, shared Kafka/Redis allocated by share.

ScaleVolumeInfra + Provider $/monthHeadcountOn-call Load
Small1M users; 5M pushes, 100K SMS, 2M emails/month~$2–4K (SMS ~$800)0.5 FTE or vendorRare
Medium50M users; 1.5B pushes, 30M SMS, 300M emails/month~$300–350K (SMS ~$240K, infra ~$40K, email ~$30K)5–7 engineers2–4 pages/month
Large500M users; 15B pushes, 300M SMS, 3B emails/month~$3M+ (SMS dominates)15–25 engineers across platform, deliverability, ML rankingDedicated 24×7; provider war rooms

The line item to manage is SMS. Moving 70% of OTPs to push-approval or passkeys at Medium scale saves ~$150K/month — more than the platform team's entire headcount cost.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
Users disabling push at OS levelOne-way (for you)Re-prompting is limited by the OS; lost opt-ins rarely come back
Consent data modelOne-way-ishRetroactively proving consent is impossible
Sender identities / domains / IP reputationOne-way-ishBurned reputation takes weeks to months to rebuild
Idempotency and type registry contractOne-way-ishEvery producer integrates against it
Queue technology, lane countTwo-wayInternal
SMS provider choiceTwo-way if multi-provider from the startSingle-provider lock-in makes it one-way
Cap valueTwo-wayConfig with experiment

The Standard I'd Write#

RFC: User Notification Standard (v1)

Scope: Any push, SMS, email, or in-app notification sent to a user of any product.

MUST:

  1. Send only through the notification platform; direct provider calls from product services are prohibited.
  2. Register each type with owner team, priority, TTL, channels, idempotency key source, collapse key, and category for consent.
  3. Supply an idempotency key per notification.
  4. Marketing notifications MUST pass central consent and unsubscribe checks at send time.
  5. New types launch to ≤ 1% of eligible users for 7 days before general availability.

SHOULD:

  1. Use aggregation for any type expected to fire more than once per user per hour.
  2. Prefer push or in-app over SMS for anything not security-critical.
  3. Set TTLs that reflect the notification's value decay.

Exceptions: Approved by the notification platform lead and, for cap bypass, the VP of product.

Success metrics: P0 handoff p95 < 2s during campaigns; push opt-in rate stable or rising quarter over quarter; zero consent violations; SMS spend per MAU down 30% year over year.

What I'd Tell the VP#

Our users' willingness to receive notifications is a shared company asset, and right now every team spends it without a budget. Push opt-in has dropped for three straight quarters, and last month a marketing campaign delayed login codes for 20 minutes. I'm proposing a central notification platform with isolated lanes, so logins and payments are never delayed by campaigns, and a daily attention budget per user that product leadership owns. Separately, SMS is costing us roughly $3M a year and is our weakest login factor; shifting to push approval and passkeys cuts about half of that while improving security. The team cost is 6 engineers, less than the SMS saving alone.

Principal Interview Signals#

SignalWhat It Sounds Like
Attention as an asset"Opt-in rate is a depletable company asset; every team draws on it."
Prices the channels"SMS is 80% of the notification bill; the fix is an identity strategy, not a cheaper provider."
Governance design"Leadership owns the cap, the platform enforces, teams compete on relevance."
One-way doors"OS-level push opt-outs rarely come back. That's the irreversible cost of over-sending."
Cross-org connection"Notifications, identity, and fraud are one conversation about SMS."

Staff answers that L7 interviewers find insufficient:

  • "We'll add a global frequency cap" — with no owner for the value, no dispute process, and no measurement of long-term effect.
  • "Multi-provider SMS for resilience" — without addressing cost, fraud, or replacing SMS as the primary factor.
  • "Teams own their opt-out rate" — without a mechanism that changes behavior (throttling, review, scorecards).

Appendices

Appendix A: Mechanics in Depth#

A.1 Ingress Dedupe#

key = "dedupe:" + user_id + ":" + idempotency_key
if SET key notification_id NX EX type.dedupe_window_s:
    proceed
else:
    return 202 { notification_id: GET key, decision: "duplicate" }

A.2 Aggregation Window#

agg_key = user_id + ":" + type.aggregation_key(payload)      # e.g. user:post:like
HINCRBY agg:{agg_key} count 1; LPUSH agg:{agg_key}:actors actor (LTRIM 0 4)
if first event and no recent send for agg_key: send now, mark window start
else if no timer: schedule flush at now + window (backoff 10m, 30m, 2h)
flush: render "A, B and N others…", collapse_id = agg_key, send, reset

Aggregation timers use the Distributed Job Scheduler or a delay queue.

A.3 Sender With TTL and Breaker#

on_dequeue(n):
  if now >= n.expires_at: record(expired); return
  provider = router.pick(n.channel, n.country, breaker_state)
  resp = provider.send(n, collapse_id=n.collapse_key, ttl=n.expires_at - now, timeout=3s)
  match resp:
    ok        → record(handed_off)
    throttled → requeue with backoff honoring Retry-After (still check TTL)
    invalid_token → delete token; record(failed_permanent)
    timeout   → breaker.record_failure(); retry same collapse_id / same content, ≤ 2 attempts

Appendix B: Data Model#

RecordStoreKeyNotes
Type registryPostgrestype_idOwner, priority, TTL, channels, keys
PreferencesPostgres + Redis cacheuser_idCategory × channel, quiet hours, timezone
Cap countersRediscap:{user}:{yyyy-mm-dd}TTL 48h
Device tokensCassandra/DynamoDBuser_id → tokensPlatform, app version, last_refreshed
InboxCassandra(user_id) clustered by created_at DESC90-day retention; read state
Delivery logCassandra → warehousenotification_idState transitions, provider IDs
Consent logAppend-only storeuser_id, categoryProvenance, timestamp, source

Appendix C: Delivery Strategies — Quick Comparison#

StrategyLatencyDuplicate RiskCostUse For
Push immediateSecondsLow with collapseFreeMost user-visible types
Push aggregatedMinutesVery lowFreeSocial activity
SMSSeconds–minutesMedium (no collapse)$$$OTP fallback, fraud
EmailSeconds–minutesLow$Receipts, digests, campaigns
In-app inboxInstant on openNone (upsert)Infra onlyEverything with a record
Digest (daily/weekly)HoursNone$Low-urgency updates

Appendix D: Provider Contract and Client Behavior#

  • APNs: HTTP/2; apns-collapse-id (≤ 64 bytes); apns-expiration; apns-priority 10 (immediate) vs 5 (power-considerate); 410 → prune token.
  • FCM: collapse_key; ttl; priority high vs normal; UNREGISTERED → prune token; overuse of high priority for non-user-visible messages can be deprioritized.
  • SMS: honor provider throttles; per-country sender registration; delivery receipts are partial — treat as signals, not truth.
  • Email: separate transactional and marketing domains/IPs; SPF, DKIM, DMARC; List-Unsubscribe header for marketing.
  • Client: dedupe by notification_id on device; inbox read-sync clears badges across devices.

Appendix E: Observability#

notif.handoff_latency_seconds{lane,channel}   # platform SLO
notif.expired_total{type}
notif.suppressed_total{type,reason}
notif.duplicate_suspected_total{type}
provider.error_rate{provider,country}, provider.throttled_total{provider}
push.receipt_rate{platform}                   # client-side sampled
optout_rate{type}, push_disable_rate
sms.cost_per_day{country,type}
AlertThresholdRoutes To
P0 handoff latencyp95 > 5s for 3 minPlatform
Provider errors> 5% for 5 min per providerPlatform
Lane lagP1 > 5 min, P2 > 30 minPlatform
Opt-out spiketype opt-out > 2× medianOwning team
SMS spend anomalydaily spend > 150% of 7-day avgPlatform + fraud
Bounce rate> 3% on any campaignGrowth; auto-pause

Appendix F: Scale Evolution#

ScaleDesign
< 1M usersVendor, or a queue with priority + ESP/SMS APIs
1–50MType registry, decision layer, P0/P1/P2 lanes, multi-provider SMS
50–500MLearned ranking, attention budget, per-country routing, client receipts
Multi-region / B2BUser-homed processing, tenant tiers, compliance packages

What you don't build on day one: learned ranking, send-time optimization, own carrier connections, multi-region failover, tenant tiers.

Appendix G: Multi-Tenancy, Fairness, and Cost#

  • Lanes and quotas: P0 reserved; P1/P2 weighted fair across teams; P2 global token bucket.
  • Attention budget: per-user daily cap across teams; ranking by type priority × predicted value.
  • Chargeback: SMS and email at provider cost per team/type; push by volume share of infrastructure.
  • Fairness signal: a team taking > 40% of the P2 lane for > 1 hour needs a scheduled campaign window.
  • Lifecycle: types with no sends in 90 days are retired; types without an active owner are disabled.
  1. Loading the index…