Technologies referenced in this case study: Kafka · Redis · Cassandra · PostgreSQL · DynamoDB
Related: Message Queues · Rate Limiting · Distributed Job Scheduler · Real-time WebSockets · Chat & Messaging · Degraded Mode · Build vs Buy
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once. Return to individual sections for targeted review.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Active Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → two Deep Dives |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens (Section 11) and Appendices |
What is a Notification System? — Why interviewers pick this topic
A notification system takes an event from some product service ("order shipped", "someone replied to you", "your login code is 482913", "spring sale starts now"), decides whether, to whom, on which channel, and when to tell a person about it, renders the message, and hands it to a delivery provider — APNs, FCM, an SMS aggregator, an email service — then tracks what happened.
Before vs After — the campaign-starves-OTP incident:
Without a designed notification platform (one shared queue, one worker pool):
t=0: Marketing launches a 40M-user push + email campaign at 10:00
t=+30s: Shared queue depth goes from 2K to 40M; workers drain at 60K/s
t=+1min: A user requests a login OTP by SMS; it sits behind ~36M campaign messages
t=+10min: OTP finally sends; it expired after 5 minutes. Login fails.
t=+15min: Support volume 6× normal: "I can't log in." Password reset emails also delayed.
t=+40min: Campaign completes. Login success rate recovers. Revenue lost on checkout logins.
With priority-isolated lanes:
t=0: Same campaign enters the bulk lane, rate-limited to 50K/s
t=+1min: OTP enters the transactional lane — separate queue, reserved workers, reserved provider quota
t=+1min+3s: SMS handed to provider; p95 OTP delivery 6s as usual
t=+14min: Campaign completes. No one outside marketing noticed.
Why interviewers reach for this question: The naive design — queue, workers, call APNs — is simple and works in the demo. The real content is everything the demo hides: duplicate sends, stale notifications that arrive after they're useful, third-party providers that throttle and fail, users who get 40 pings a day and uninstall, and fifteen teams who all believe their notification is the important one. It's a question about priority, idempotency, and governance of a shared resource — the user's attention.
Mechanics Refresher: Channels and Providers
| Channel | Provider | Delivery Semantics | Cost | Latency |
|---|---|---|---|---|
| iOS push | APNs (HTTP/2) | Best effort; if device offline, APNs stores one notification per app (the latest) | Free | Sub-second to seconds |
| Android push | FCM | Best effort; stores messages for offline devices up to a TTL (default 4 weeks) | Free | Sub-second to seconds |
| SMS | Aggregators (e.g., Twilio, Vonage) → carriers | Carrier-dependent; delivery receipts partial | ~$0.005–0.01+ per US segment; far higher internationally | 1–30s, tail minutes |
| SES, SendGrid, Mailgun | Accepted ≠ inboxed; spam filtering; bounces async | ~$0.10 per 1,000 on SES | Seconds to minutes | |
| In-app inbox | Your own store + WebSocket/poll | You control it fully | Your infra | ms when online |
| Web push | Browser push services | Best effort, subscription churn high | Free | Seconds |
For most production systems: in-app inbox as the durable system of record, push as the primary interruptive channel, email for records and digests, SMS only for high-value transactional messages (OTP, fraud alerts) because it costs real money per message. The channel list is almost never the interview question — priority isolation, dedupe, and who decides whether to interrupt a human are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What This Interview Actually Tests#
A notification system is not a fan-out question. Everyone can put messages in a queue and call APNs.
It is an "attention is a shared, finite resource — who is allowed to spend it, and what happens when delivery fails" question that tests:
- Whether you isolate priorities so a marketing blast can never delay a login code
- Whether you understand that a duplicate notification is a user-visible bug and design idempotency end to end
- Whether you treat freshness as correctness — a "your driver is here" push that arrives 10 minutes late is worse than none
- Whether you see third-party providers as rate-limited, partially failing dependencies you don't control
- Whether you define who governs the user's attention across fifteen product teams
The key insight: The system's hardest decision isn't how to send — it's whether to send. Staff candidates design the decision layer (preferences, frequency caps, dedupe, collapse, expiry, quiet hours) as a first-class platform component, and isolate delivery by priority so the rare important message is never behind the common unimportant one.
The L5 vs L6 vs L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws queue → workers → APNs/FCM/SMS/email | Asks "transactional, engagement, or campaign?" and designs priority lanes | Asks who governs the user's attention budget across teams and how it's measured |
| Delivery guarantee | "Guaranteed delivery with retries" | At-least-once to providers + idempotency key per (user, notification); expiry TTL per type | Sets org rule: every notification type declares priority, TTL, dedupe key, and collapse key at registration |
| Priority | One queue, maybe a priority field | Separate lanes with reserved workers and reserved provider quota; campaigns rate-limited | Treats provider quota as a budget allocated across business units, with finance-visible SMS spend |
| Failure | "Retry on provider error" | Per-provider circuit breakers, failover (SMS provider A → B), backoff honoring provider throttles, stale drop | Multi-provider contracts negotiated for redundancy; game days for provider outages |
| User experience | Sends every event | Preferences, quiet hours, frequency caps, aggregation ("5 people liked your post") | Attention budget as an org metric: notifications/user/day, opt-out and uninstall rate as guardrails every team is judged on |
| Ownership | Platform sends what teams ask | Platform owns delivery SLO; teams own content, targeting, and their type's opt-out rate | Central review for new notification types; kill switch per type; teams lose sending rights above opt-out thresholds |
Why "priority" separates levels
L5: One queue with a priority field, or a priority queue. When 40M campaign messages arrive, the transactional message jumps the queue — but shares the same workers, the same connection pools, and the same provider rate limit. If the provider throttles because of campaign volume, the OTP gets throttled too.
L6: "Priority must be isolated at every shared resource: queue, workers, and provider quota. Transactional has its own topic, its own consumer pool sized for 10× its peak, and a reserved slice of our SMS and push provider throughput. Campaigns run through a token bucket at 50K/s and are the first thing shed. A campaign should be physically incapable of delaying an OTP."
L7: Recognizes provider quota and SMS spend as org resources. "Our SMS contract gives us N messages/s; the transactional lane owns 40% of it permanently. Business units draw from the rest by budget."
Why "delivery guarantee" separates levels
L5: Promises "guaranteed delivery." But APNs is best-effort, FCM is best-effort, SMS delivery receipts are partial, and email "accepted" means nothing about the inbox. The honest guarantee stops at the provider handoff.
L6: "We guarantee at-least-once handoff to a provider within the type's TTL, or an explicit expiry. Retries use the same idempotency key, and APNs apns-collapse-id / FCM collapse_key make on-device duplicates collapse. Beyond the provider, we measure delivery and engagement, but we don't claim to control it. For messages that must be seen, the in-app inbox is the durable record."
L7: Makes the per-type contract explicit and auditable: priority, TTL, dedupe window, collapse key, allowed channels, and fallback channel.
Why "user experience" separates levels
L5: Treats every event as a notification. Users get "Alex liked your photo" 14 times in an hour.
L6: Designs the decision layer: aggregation windows ("Alex and 13 others liked your photo"), per-user frequency caps (e.g., max 3 non-transactional pushes/day), quiet hours by local time, and suppression if the user is currently active in-app.
L7: Owns the org-level tradeoff: every team optimizes its own click-through, and the sum of locally-rational teams is a user who uninstalls. The fix is a shared budget and a central arbiter — like LinkedIn's publicly described notification-volume control system.
The Staff Positions#
| Position | Rationale |
|---|---|
| Priority isolation at queue, workers, and provider quota | A shared resource anywhere in the path re-couples campaigns and OTPs |
| At-least-once handoff + idempotency key per notification | Duplicates are user-visible; retries are unavoidable |
| Every notification type has a TTL | A stale notification is worse than none; expired messages are dropped, not sent |
| Decision layer before delivery | Preferences, caps, quiet hours, and collapse cost less than sending and apologizing |
| In-app inbox as the durable record | Push/SMS/email are best-effort; the inbox is where "must be seen" lives |
| Multi-provider for SMS and email; circuit breaker per provider | Providers throttle and fail; the platform must degrade, not stall |
| Platform owns delivery; teams own content and opt-out rate | Attention is shared; teams must be accountable for how they spend it |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Transactional (OTP, password reset, receipts, fraud alerts, ride arrival) | Latency and reliability; low volume | Dedicated lane, reserved quota, multi-provider failover, short TTL | Late OTP = failed login; duplicate receipt = confusion | p95 handoff < 2s; p99 delivery < 30s; never duplicated user-visibly |
| Social / activity ("X replied", "Y liked") | Volume spikes, relevance, fatigue | Aggregation windows, collapse keys, frequency caps, suppress-if-active | Spam → opt-out, uninstall | Relevant, not repeated; minutes of delay acceptable |
| Campaigns / marketing | Massive fan-out (10–100M), provider limits, legal compliance | Audience streaming, rate-limited send, send-time optimization, unsubscribe | Starving other lanes; compliance violations | Throughput and compliance; individual loss acceptable |
🎯 Staff Move: "These three intents can't share a pipeline without one starving another. I'll design the platform around priority lanes and a shared decision layer, and I'll go deep on transactional first because it has the strictest bar — then show that campaigns are physically incapable of delaying it."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Duplicate vs Lost Notifications | Retry aggressively (risk double-send) or give up (risk silence)? |
| 2 | Shared Pipeline vs Priority Isolation | One efficient pipeline or separate lanes with reserved (sometimes idle) capacity? |
| 3 | Deliver Late vs Drop Stale | Send everything eventually, or expire notifications whose moment has passed? |
| 4 | Send Everything vs Protect Attention | Maximize each team's engagement or cap, aggregate, and suppress for the user's sake? |
| 5 | Single Provider vs Multi-Provider | Simplicity and volume discounts vs resilience and routing complexity? |
In the Wild: Real Production Systems#
Why this section belongs here: Citing specific production systems demonstrates you've studied operational reality, not textbook designs.
LinkedIn — Centralized Notification Volume Control#
LinkedIn has publicly described "Air Traffic Controller" (ATC), a centralized system that decides whether, when, and through which channel to send each notification to a member, applying volume limits and relevance scoring across the many product teams that generate notifications.
Staff insight: ATC is the org-level answer to Fault Line 4: individual teams don't decide to interrupt a member; a shared arbiter does, with a global view of what that member has already received. Cite it when the interviewer asks "who decides?"
Slack — The "Should We Notify?" Decision Tree#
Slack has published a well-known flowchart showing the logic that decides whether a message triggers a notification: channel mute state, user preferences, do-not-disturb, whether the user is active on another device, thread subscriptions, mentions, and more.
Staff insight: The complexity of "should we send?" dwarfs "how do we send?" A candidate who designs the decision layer as a first-class component — with suppress-if-active and DND — is designing the part users actually experience.
Apple APNs and Google FCM — Best-Effort Delivery With Collapse#
APNs documents that when a device is offline it stores only one notification per app (the most recent) and that apns-collapse-id lets a newer notification replace an older one. FCM supports collapse keys and a TTL (default 4 weeks) for offline devices, plus priority levels that affect delivery timing under device power management.
Staff insight: The platforms you depend on are explicitly best-effort and explicitly support collapse and expiry. Designing your own system around "guaranteed delivery" contradicts the providers' contracts; designing around TTLs and collapse keys aligns with them.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Kafka queue and workers" | "Marketing sends to 50M users. Where is the OTP?" | Priority isolation |
| "Retry on failure" | "The provider timed out but actually delivered. Now what?" | Idempotency and collapse |
| "Guaranteed delivery" | "APNs is best-effort. What exactly do you guarantee?" | Precise guarantee boundaries |
| "Send when the event happens" | "The worker was down for 20 minutes. Do we send 'your driver is here' now?" | Freshness / TTL |
| "Users can set preferences" | "Fifteen teams each send 'just one' push a day." | Attention governance |
| "Twilio for SMS" | "Twilio is degraded in one country. OTPs failing." | Multi-provider failover |
System Architecture Overview#
Reading the diagram: Every producer goes through one API that enforces the type registry (priority, TTL, dedupe key). The decision layer answers whether and when — preferences, caps, quiet hours, aggregation — before anything is queued for delivery. Lanes are physically separate topics and worker pools with reserved provider quota, so P2 volume cannot touch P0 latency. Channel senders speak provider protocols with collapse IDs and per-provider circuit breakers. The in-app inbox is the durable record; the delivery log tracks each notification's state machine for dedupe, analytics, and debugging.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Architecture | "Queue + workers + providers" | "Decision layer, then isolated priority lanes with reserved workers and reserved provider quota." |
| Guarantee | "Guaranteed delivery" | "At-least-once handoff within TTL; idempotency key per notification; collapse IDs on device. Inbox is the durable record." |
| Freshness | "Retry until sent" | "Every type has a TTL. Expired = dropped and counted, never sent late." |
| Fatigue | "User preferences" | "Preferences + frequency caps + aggregation + quiet hours + suppress-if-active." |
| Providers | "Use Twilio / SES" | "Two providers per paid channel, per-country routing, circuit breaker per provider, honor 429 backoff." |
| Ownership | "Platform sends for everyone" | "Platform owns handoff SLO per lane. Teams own content, targeting, and opt-out rate for their types." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| OTP validity | typically 5–10 min | Sets P0 TTL; late OTP = failed login |
| P0 handoff SLO | p95 < 2s, p99 < 5s | The transactional promise |
| SMS cost (US) | ~$0.005–0.01+ per segment; international often 5–10× | SMS is the budget line; never use it for engagement at scale |
| Email cost (SES) | ~$0.10 per 1,000 | Email is cheap; deliverability is the constraint |
| Push cost | Free at the provider | Cost is attention, not dollars |
| APNs offline storage | 1 notification per app (latest) | Collapse is built in; don't rely on backlog delivery |
| FCM default TTL | 4 weeks | Set shorter TTLs explicitly for time-sensitive messages |
| Campaign send rate | 20–100K/s typical, provider-limited | 50M audience ≈ 10–40 minutes |
| Aggregation window | 2–10 min for social | Turns 14 pushes into 1 |
| Frequency cap (non-transactional) | ~2–5 pushes/user/day | Above this, opt-out rates climb sharply for most apps |
| Email bounce rate threshold | keep hard bounces < ~2% | ESPs throttle or suspend senders with high bounce/complaint rates |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on templates, channel adapters, and the preferences schema, then have no time for the only two questions that matter: "Marketing is sending to 50M users — where is the login code?" and "The provider timed out but delivered — did we just send it twice?"
Phase 1: Requirements & Framing (2–3 minutes)#
Functional requirements in 30 seconds:
"Product services send notifications to users over push, SMS, email, and an in-app inbox. Users control preferences. We track delivery state."
Then intent and non-functionals:
"Three different workloads hide under 'notifications': transactional — OTPs, receipts, ride arrivals — low volume, latency-critical; social activity — high volume, relevance and fatigue matter; and campaigns — 50M-user fan-outs where throughput and compliance matter. They can't share a pipeline. I'll design a platform with isolated priority lanes and a shared decision layer, and go deep on transactional correctness first."
Commit to numbers:
"500M users, 100M DAU. ~150K notification-worthy events/s peak, most social; after aggregation and caps, ~30K sends/s; campaigns up to 100M recipients. P0 handoff p95 under 2 seconds regardless of campaign load. No user-visible duplicates. Every type has a TTL — expired notifications are dropped, not sent late."
🎯 Staff Move: "Regardless of campaign load" is the load-bearing phrase. It commits you to physical isolation — not a priority field — before you draw a box.
Phase 2: Core Entities & API (1–2 minutes)#
- NotificationType (registry):
type_id,owner_team,priority(P0/P1/P2),ttl,channels[],fallback_channel,dedupe_window,collapse_key_template,aggregation,counts_toward_cap - Notification:
notification_id,idempotency_key,user_id,type_id,payload,created_at,expires_at - Delivery:
notification_id,channel,provider,state(queued → handed_off → delivered | failed | expired | suppressed),attempts
POST /notify { type_id, user_id | audience_id, idempotency_key, payload, event_time }
→ 202 { notification_id, decision: queued | suppressed(reason) | aggregated }
GET /users/{id}/inbox?cursor=…
PUT /users/{id}/preferences { type_id | category, channels, quiet_hours }
🎯 Staff Move: "
idempotency_keyis required and supplied by the producer — for an OTP it's the auth request ID; for a reply notification it's the reply's ID. The platform dedupes on(user_id, idempotency_key)within the type's window. A producer retry can't double-send."
Phase 3: High-Level Architecture (≤5 minutes)#
Walk the flow in 90 seconds:
- Producer calls the API with type and idempotency key; API rejects unregistered types and dedupes
- Decision layer loads user preferences and cap counters (Redis), applies quiet hours in the user's time zone, suppresses if the user is active in-app, and either queues, aggregates (holds for a window), or suppresses with a reason
- Queued notifications land in the lane for their type's priority — separate Kafka topics and consumer groups
- Channel senders resolve device tokens / phone / email, render templates, check TTL, and call providers with collapse IDs, behind per-provider circuit breakers
- Delivery log records every state transition; provider receipts and bounces update it asynchronously
- The in-app inbox is written for types that have an inbox representation — it's the durable record
Key points to state:
- Decide before deliver — the cheapest notification is the one you don't send
- Isolation at every shared resource — topics, workers, and provider quota
- Idempotency at the API, collapse at the device
- TTL checked at dequeue and before provider call
- Two SLOs: P0 handoff latency (platform) and opt-out rate per type (owning team)
🎯 Staff Move: Then say: "That's the happy path. The design really gets tested by three things: a campaign during a login spike, a provider that's timing out but still delivering, and fifteen teams competing for the same user's lock screen."
Phase 4: Transition to Depth (1 minute)#
"Three places to go deep: priority isolation under campaign load, duplicate-vs-lost with flaky providers, and attention governance across teams. Which do you prefer?"
Default: priority isolation — it's the most common real incident and the clearest architectural decision.
Phase 5: Deep Dives (25–30 minutes)#
Deep dive A: Priority isolation (7–9 min)
"Isolation must hold at four places. Queue: P0 is its own topic, so its lag is independent of P2's backlog. Workers: P0 has its own consumer pool, sized at 10× its peak — it's ~2K/s, so that's cheap. Provider quota: if our SMS aggregator allows 1,000 msg/s on our account, P0 gets a reserved 400/s through a separate sender token bucket; campaigns can use at most 600/s. Same for APNs connections — P0 has its own HTTP/2 connection pool. Downstream dependencies: template rendering and the device-token store are shared, so P0 calls them with a separate client pool and higher timeout priority. And the campaign lane has a global token bucket at 50K/s — 50M users in ~17 minutes — which is also what our providers can absorb."
Deep dive B: Duplicate vs lost (7–9 min)
"Our request to APNs times out after 10s. Did it deliver? Unknown. If we retry, we might double-send; if we don't, we might lose it. For push I retry with the same apns-collapse-id = notification_id, so if both arrive the second replaces the first on the device — the user sees one. For SMS there's no collapse, so for OTPs I retry only via a status check where the provider supports it, or I send once more with the same code — a duplicate OTP with the same code is harmless, a lost one blocks login. For email, providers accept a message ID; duplicates are rare and cheap. The platform-level dedupe is on (user, idempotency_key) in Redis with a TTL of the type's dedupe window — say 24h for receipts, 10 min for OTP."
Deep dive C: Attention governance (6–8 min)
"Fifteen teams each think their notification is the one that matters. If each sends one push a day, the user gets 15 and turns notifications off — which kills the OTP channel too. So the decision layer enforces a per-user daily cap for non-transactional pushes, say 3, and ranks candidates within the cap by predicted value. Social events aggregate in 5-minute windows. Quiet hours 22:00–08:00 local defer P1/P2. Each type's owner is accountable for its opt-out and disable rate; if a type drives opt-outs above a threshold, it's automatically throttled and the owner is notified."
🎯 Staff Move: Close each dive with ownership: "Platform owns P0 latency and provider failover. The growth team owns the cap values, signed off with product leadership. Each team owns its type's opt-out rate."
Phase 6: Wrap-Up (2–3 minutes)#
"Summary: registry-enforced types with priority, TTL, and dedupe; a decision layer for preferences, caps, aggregation, and quiet hours; isolated lanes with reserved provider quota; idempotency at ingress and collapse at the device; multi-provider SMS/email with circuit breakers; inbox as the durable record. Next I'd build send-time optimization for campaigns and a learned ranking for which notifications win the daily cap. I'd deliberately not build our own SMS carrier connections — aggregators exist for a reason."
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| Template engine and localization design | 5–8 min | "Templates versioned in a registry; rendered in the sender. Standard." |
| Device token management details | 5 min | "Token table per user; prune on APNs 410 / FCM unregistered." |
| Detailed preferences UI | 4 min | "Category × channel matrix. Moving on." |
| Comparing push providers | 3 min | "APNs for iOS, FCM for Android — not a choice." |
| Never reaching campaign-vs-OTP | Interview-ending | Transition by minute 12 |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Notifications are where product ambition meets user tolerance. Every product team has a growth metric that goes up when they send more notifications — in the short term. Every user has a tolerance that goes down with each one — and when it breaks, they disable notifications for the whole app, including the ones that mattered.
That's the Staff-level content hiding in a queue-and-workers problem: a shared resource with no natural owner (the user's attention and your provider quota), external dependencies you don't control (APNs, carriers, ESPs), and correctness defined by time (a notification's value decays, sometimes to negative). The interviewer is watching for whether you see those three things or just build a fast pipe.
1.2 The L5 vs L6 Contrast — Visual#
The L5 path partitions by channel (push queue, SMS queue). The L6 path partitions by priority first, then channel — because the incident that matters is a campaign on the same channel as an OTP.
1.3 The Staff Question That Cuts Through Everything#
"If we could only send one notification to this user today, which one — and who decided?"
It forces:
- Priority: some notifications are categorically more important
- Budget: attention is finite and must be allocated
- Governance: someone other than the sending team decides
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1: Transactional. Triggered by a user's own action or an event about their account/money/safety: OTP, password reset, payment receipt, fraud alert, "your driver has arrived." Low volume (~1–5% of sends), very high value per message, strict freshness. The correctness bar is fast, exactly-one-visible, fallback on failure. A late OTP is a failed login; a duplicate "payment received" is a support ticket; a missed fraud alert is a loss. These bypass frequency caps and quiet hours (with a narrow allowlist), and may fall back from push to SMS.
Intent 2: Social / activity. Triggered by other users: replies, mentions, likes, follows. High volume, bursty, highly skewed (a popular post generates thousands of like events). The correctness bar is relevant and not repetitive. Minutes of delay are fine — often desired, to aggregate. Frequency caps and quiet hours apply. Suppress when the user is actively in the app.
Intent 3: Campaigns / marketing. Triggered by the business: promotions, re-engagement, announcements. Enormous fan-out, scheduled, legally constrained (unsubscribe links, consent, CAN-SPAM, GDPR, TCPA for SMS in the US). The correctness bar is throughput within provider limits, compliance, and not harming the other lanes. Individual losses are acceptable; sending to someone who unsubscribed is not.
🎯 Staff Move: "Transactional and campaign share channels and nothing else. Transactional needs reserved capacity that sits idle 95% of the time; campaigns need to be throttled even when capacity is available. If I put them in one pipeline, I'll either over-provision everything or starve logins during every campaign."
2.2 When NOT to Build a Notification Platform#
| Situation | Use Instead | Why |
|---|---|---|
| < 1M users, one product team | Managed services (a customer-engagement platform, Firebase + an ESP) | Preferences, templates, and analytics come built in |
| Only transactional email | An ESP's transactional API directly | A platform adds latency and failure modes for no benefit |
| Real-time presence or chat delivery | Real-time WebSockets / Chat | In-session delivery is a connection problem, not a notification problem |
| Marketing-only needs | A marketing automation vendor | Segmentation, A/B, and compliance tooling are the product |
| Internal alerting (on-call pages) | PagerDuty / Opsgenie | Escalation policies, acknowledgment, and phone-call fallback are the hard part |
"I'd build a platform when three or more teams send notifications and we've had our first 'campaign delayed OTPs' or 'users disabled push' incident. Before that, buy."
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It Matters | What to Say |
|---|---|---|
| Mix of intents | Drives isolation design | "Transactional, social, and campaigns — isolated lanes." |
| Delivery guarantee | Providers are best-effort | "At-least-once handoff within TTL; inbox is durable." |
| Freshness | Stale can be harmful | "Every type has a TTL; expired = dropped." |
| Duplicate tolerance | Users see duplicates | "Zero user-visible duplicates for transactional." |
| Fatigue controls | Opt-outs kill the channel | "Caps, aggregation, quiet hours, suppress-if-active." |
| Compliance | Legal exposure | "Consent and unsubscribe enforced in the decision layer, not by each team." |
| Multi-device | One user, 3 devices | "Send to all active tokens; collapse and read-sync via inbox." |
| Time zones | Quiet hours, send-time | "User's local time zone from profile; default from device." |
2.4 Precise Terminology#
| Term | Precise Meaning |
|---|---|
| Handoff | Provider accepted the request (HTTP 200 from APNs/FCM/SMS/ESP). The platform's controllable boundary. |
| Delivery | Device or inbox received it — known only partially (FCM delivery data, SMS DLRs, email non-bounce) |
| Idempotency key | Producer-supplied ID; platform dedupes (user, key) within the type's dedupe window |
| Collapse key / ID | Provider-level key so a newer notification replaces an older one on device |
| TTL / expires_at | After this time the notification is dropped and counted as expired |
| Lane | An isolated queue + worker pool + provider quota for one priority class |
| Frequency cap | Max non-transactional notifications per user per period, across all teams |
| Aggregation window | Time the platform holds similar notifications to merge them ("Alex and 4 others…") |
| Suppression | Decision not to send, with a recorded reason (capped, quiet_hours, user_active, opted_out, expired) |
| Fallback channel | Alternate channel if the primary fails within a deadline (push → SMS for OTP) |
3. The Fault Lines#
3.1 Fault Line 1: Duplicate vs Lost Notifications#
The tension: Provider calls fail ambiguously. A timeout doesn't tell you whether the message was delivered.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| At-most-once (never retry) | Zero duplicates | Every provider blip loses messages; OTP failures | Users who can't log in |
| Retry blindly | Nothing lost | Duplicate SMS, duplicate receipts; SMS costs double | Users; finance (SMS spend) |
| Retry with idempotency + collapse | Duplicates collapse on device (push) or are harmless (same OTP code) | Needs per-channel strategies; SMS has no collapse | Platform — per-channel logic |
| Status check before retry | Avoids duplicates where providers expose status | Adds latency; not all providers support it | Platform |
Staff default: Ingress dedupe on (user, idempotency_key); push retries reuse collapse-id; SMS retries reuse the same content (same OTP code) and prefer status-check-then-failover; email retries reuse a stable Message-ID.
"For transactional, a harmless duplicate beats a lost message. I design the duplicate to be harmless — same code, same collapse ID — rather than trying to prevent it."
3.2 Fault Line 2: Shared Pipeline vs Priority Isolation#
The tension: Isolation means reserved capacity that sits idle most of the time. Sharing is efficient until the day it isn't.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| One queue, FIFO | Simplest; max utilization | Campaign starves OTP | Users trying to log in; revenue |
| One queue with priority field | Transactional jumps the queue | Still shares workers, connection pools, provider rate limits | Same, less often |
| Separate topics, shared workers | Queue lag isolated | Workers busy on P2 can't pick up P0 | Same, under load |
| Full lane isolation (topic + workers + provider quota) | P0 latency independent of P2 | ~10–20% extra idle capacity; quota fragmentation | Platform budget — cheap relative to one OTP outage |
Staff default: Full isolation for P0; P1 and P2 may share workers with weighted fair scheduling, but P2 has a hard global token bucket and is shed first.
When to deviate: Small systems (< 1K sends/s total) can use one pipeline with a priority queue and a reserved worker slice — just never let P2 exhaust provider quota.
3.3 Fault Line 3: Deliver Late vs Drop Stale#
The tension: "Never lose a notification" sounds right until the 20-minute-late "Your driver has arrived" push lands after the ride started.
| Type | Value Decay | TTL | Late Behavior |
|---|---|---|---|
| OTP | Useless after code expiry | 5 min | Drop; user requests a new code |
| Ride / delivery arrival | Harmful after ~2 min | 2–5 min | Drop; in-app state shows truth |
| "X replied to you" | Declines over hours | 6–24 h | Send; or fold into digest |
| Payment receipt | Stable (record) | 7 days | Send; also in inbox and email |
| Flash sale starting | Zero after sale ends | Sale end | Drop |
| Campaign | Declines over a day | 24 h | Drop after TTL |
Staff default: Every type declares a TTL at registration. Senders check now < expires_at at dequeue and immediately before the provider call; provider-side TTL (APNs apns-expiration, FCM ttl) is set to the remaining TTL. Expired messages are counted in notif.expired_total{type} — a rising count is an incident signal, not noise.
"A late notification is not a delivered notification. It's a new bug — the user now sees wrong information on their lock screen."
3.4 Fault Line 4: Send Everything vs Protect Attention#
The tension: Each team's notifications increase its engagement metric; together they decrease everyone's.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each team sends freely | Team autonomy; local metrics up | Users get 15+/day; disable all notifications | Every team, including transactional |
| Per-type rate limits | Some protection | No cross-team view; sum still too high | Users |
| Global per-user cap + ranking | Budget enforced; best candidates win | Teams lose notifications to others; ranking disputes | Teams with low-value types |
| Learned send decisions (relevance models, bandits) | Optimizes long-term engagement | Model ops; explainability to teams | Platform ML team |
Staff default: Global per-user cap for non-transactional push (e.g., 3/day), aggregation windows, suppress-if-active, quiet hours, and a simple value ranking; move to learned ranking when volume justifies it. Duolingo has published research on using bandit algorithms to choose notification content — a public example of treating "which notification" as an optimization problem.
3.5 Fault Line 5: Single Provider vs Multi-Provider#
The tension: One provider per channel is simpler and gets better volume pricing. Providers have regional outages, carrier-specific failures, and throttle limits.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Single SMS provider | Simple; volume discount | Provider or country route outage = OTP outage | Users; revenue |
| Primary + cold standby | Failover exists | Standby untested; sender registration (e.g., US 10DLC, sender IDs) not set up when needed | Incident responders |
| Active-active with per-country routing | Best delivery rates per country; instant failover | Two contracts, two integrations, delivery-rate monitoring per route | Platform team maintains routing |
Staff default: Active-active for SMS with per-country routing by measured delivery rate and cost, at least 5–10% of traffic on the secondary to keep it warm. Email: primary ESP + warm secondary with dedicated IPs warmed in advance. Push: APNs and FCM are the only options — isolate by connection pool and honor their backoff.
"A standby provider you haven't sent traffic through in six months is a hypothesis, not a failover."
4. Failure Modes & Operational Reality#
4.1 The Campaign That Starved Logins#
t=0: Growth launches "spring sale" to 48M users; enters shared push+SMS pipeline
t=+20s: SMS provider account limit (1,000 msg/s) fully consumed by campaign SMS segment (2M users)
t=+40s: OTP SMS requests receive 429 from provider; sender retries with 1s backoff
t=+2min: OTP p95 latency 4 min; OTPs expire at 5 min; login success rate drops from 97% to 71%
t=+8min: Auth on-call paged for login failures; suspects auth service
t=+22min: Notification on-call identifies provider quota exhaustion; pauses campaign
t=+25min: OTP latency recovers
Detection: notif.handoff_latency_p95{lane="P0"} > 5s; provider.throttled_total{provider,lane}; auth.otp_success_rate correlated.
Blast radius: All SMS-dependent logins and password resets.
Mitigation: Reserved provider quota for P0 via separate sender token buckets; campaign SMS on a separate sender ID / account where the provider supports it.
Prevention: Lane isolation to the provider; campaign preflight check that computes required provider throughput and refuses launches that exceed P2's share.
Owner: Notification platform (isolation); growth team (campaign preflight).
4.2 The Duplicate Storm After a Provider Timeout#
APNs responses slow to 12s due to a network issue between our region and Apple; our client timeout is 10s. Every push "fails" at 10s and is retried — but most actually delivered.
t=0: APNs latency rises to 12s p50
t=+10s: Timeouts; retry policy re-sends (new request, no collapse-id configured on this type)
t=+30s: Users receive each social push 2–3 times
t=+5min: "Why am I getting everything 3 times?" trends on social media
Detection: provider.timeout_rate{provider} spike; notif.duplicate_suspected (same notification_id handed off > 1 time); app-side dedupe telemetry.
Mitigation: Collapse-id = notification_id for every push type (mandatory at registration); retry budget (≤ 10% of requests); circuit breaker opens at 50% timeout rate for 30s instead of retrying into a slow dependency.
Owner: Platform.
4.3 Stale Notifications After a Backlog#
A consumer deploy bug stops the P1 lane for 35 minutes. On recovery, 60M queued social notifications flood out — including "Sam is live now!" for streams that ended 30 minutes ago.
Detection: lane.consumer_lag_seconds{lane} > 60; notif.expired_total should spike on recovery — if it doesn't, TTLs aren't being enforced.
Mitigation: TTL enforcement at dequeue; drain backlog at a controlled rate (backlog replay is itself a herd for providers and for the app's backend when users tap).
Owner: Platform (TTL enforcement); each type owner (setting a correct TTL).
4.4 Token Rot — Pushing Into the Void#
Device tokens become invalid when users uninstall or reinstall. APNs returns 410 Unregistered and FCM returns UNREGISTERED; if we don't prune, 30% of push requests go to dead tokens, inflating cost, skewing delivery metrics, and in some cases hurting sender reputation.
Detection: push.invalid_token_rate > 5%; "sent" rising while opens flat.
Mitigation: Delete tokens on 410/UNREGISTERED; expire tokens not refreshed by the app in 60 days.
Owner: Platform.
4.5 Email Reputation Collapse#
A team imports an old mailing list and sends a campaign; 9% hard bounces and elevated spam complaints. The ESP throttles the whole account — including password reset emails.
Detection: email.bounce_rate{type} > 2%; email.complaint_rate > 0.1%; ESP account status.
Mitigation: Separate sending domains/IPs (and ideally separate ESP subaccounts) for transactional vs marketing; list hygiene required before campaigns; auto-pause campaign at bounce > 3%.
Owner: Growth (list hygiene); platform (domain separation and auto-pause).
4.6 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Campaign starves P0 | handoff_latency_p95{P0} > 5s; provider 429s | Logins, resets | Reserved quota; campaign preflight | Platform + growth |
| Provider timeouts → duplicates | duplicate_suspected; timeout rate | All push users | Collapse IDs; retry budget; breaker | Platform |
| Backlog → stale flood | Lane lag; expired_total flat on recovery | Users, app backend | TTL at dequeue; controlled drain | Platform + type owners |
| Dead tokens | invalid_token_rate > 5% | Cost, metrics | Prune on 410; 60-day expiry | Platform |
| Email reputation hit | Bounce > 2%, complaints > 0.1% | All email incl. transactional | Separate domains/IPs; auto-pause | Growth + platform |
| SMS provider country outage | sms.delivery_rate{country,provider} drop | OTPs in that country | Per-country failover routing | Platform |
| Preferences store down | Decision layer errors | All non-transactional | Fail-closed for campaigns (don't send); fail-open to defaults for transactional | Platform |
| Over-notification by one team | optout_rate{type} > threshold | App-wide notification opt-out | Auto-throttle type; review | Owning team |
| Quiet-hours timezone bug | Complaints; sends at 03:00 local | Users in one region | Timezone tests; send-time audit | Platform |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Architecture | Queue per channel + workers | Decision layer + priority lanes isolated through provider quota | Platform with type registry, governance, and chargeback across business units |
| Guarantee | "Guaranteed delivery" | At-least-once handoff within TTL; idempotency + collapse; inbox as durable record | Per-type contracts audited; transactional delivery SLO reported to leadership |
| Freshness | Retry until sent | TTL per type enforced at dequeue and provider | Value-decay model per type informs TTL defaults org-wide |
| Fatigue | Preferences | Caps, aggregation, quiet hours, suppress-if-active | Attention budget as org metric; teams accountable for opt-out rate; central arbiter |
| Providers | One per channel | Multi-provider SMS/email, per-country routing, breakers, warm standby | Contract strategy, sender registration, cost per delivered message by country |
| Compliance | Unsubscribe link | Consent enforced centrally in decision layer | Legal-owned consent model; audit trail; regional regulation map |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Isolation to the provider | "Separate topic isn't enough — P0 needs reserved provider quota or the campaign still throttles OTPs." |
| Honest guarantee | "We control handoff, not delivery. APNs is best-effort; the inbox is the durable record." |
| Freshness as correctness | "A late 'driver arrived' is a new bug. Every type has a TTL." |
| Harmless duplicates | "Retry the OTP with the same code; push retries reuse the collapse ID." |
| Attention governance | "Fifteen teams × one push a day = uninstall. The cap is global and owned by product leadership." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Partition by channel only | Campaign and OTP still share the push and SMS paths |
| "Guaranteed delivery" | Contradicts provider contracts; no guarantee boundary |
| Unbounded retries | Duplicates, cost, and stale sends |
| No decision layer | Designs a pipe, not a notification system |
| Each team handles its own unsubscribe | Compliance risk; inconsistent enforcement |
5.4 Common False Positives#
- Deep APNs/FCM API knowledge ≠ system design. Knowing HTTP/2 headers doesn't answer priority isolation.
- Kafka partition math ≠ Staff. Throughput is rarely the bottleneck; providers and attention are.
- Fancy ML ranking ≠ governance. A model that picks the best notification still needs an owner for the cap and an override for transactional.
- "We'll use WebSockets" ≠ notifications. In-session delivery doesn't reach users who aren't in the app — which is the point of notifications.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Three intents; P0 isolated "regardless of campaign load"; TTL per type |
| Entities + API | 3–5 min | Type registry, Notification, Delivery; required idempotency key |
| High-level design | 5–12 min | API → decision layer → lanes → senders → providers → delivery log |
| Transition | 12 min | Offer: isolation, duplicates, governance |
| Deep dives | 12–38 min | Isolation → duplicates/providers → fatigue/governance → freshness |
| Org & evolution | 38–43 min | Ownership split; compliance; build vs buy |
| Wrap-up | 43–45 min | Summary; what's next; what you won't build |
6.2 How Interviewers Pivot — And What They're Testing#
| Interviewer Pivot | What They're Testing | Where to Go |
|---|---|---|
| "Send to all 100M users now." | Fan-out and provider limits | Audience streaming, token bucket, send-time spread |
| "The SMS provider is down in India." | Dependency failure | Per-country routing; failover; warm secondary |
| "Users complain about too many notifications." | Product judgment | Caps, aggregation, governance |
| "User has 3 devices." | Multi-device semantics | Send to all active tokens; inbox read-sync clears badges |
| "How do you know it was delivered?" | Guarantee honesty | Handoff vs delivery; receipts; engagement proxies |
| "A user unsubscribed yesterday and got a campaign today." | Compliance | Central consent check at send time, not audience build time |
| "What if the notification service is down?" | Degraded mode | Producers buffer; P0 fallback path; inbox writes |
6.3 What to Deliberately Skip#
| Topic | Why L5 Goes Here | What L6 Says Instead |
|---|---|---|
| Template engine | Concrete | "Versioned templates, rendered in senders, localized by user locale." |
| Rich push media | Feature-y | "Payloads ≤ 4 KB; media via URL." |
| Detailed analytics pipeline | Familiar | "Delivery log to the warehouse; standard." |
| Preference UI matrix | Easy | "Category × channel. Moving on." |
6.4 Follow-Up Questions to Expect#
- "How do you implement 'Alex and 4 others liked your post' without delaying the first notification too long?"
- "How do you send at 9 a.m. local time to 100M users across 30 time zones?"
- "What happens to queued notifications for a user who deletes their account?"
- "How would you roll out a new notification type safely?"
- "What's the cost per month of SMS OTPs at 20M logins/day, and how would you cut it?"
- "How do you handle badge counts across devices?"
- "How do you ensure a fraud alert reaches the user if push fails?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design a notification system."
Staff Answer
"Three workloads hide under that name: transactional (OTP, receipts, arrivals), social activity (replies, likes), and campaigns (50M-user blasts). They need physical isolation — a campaign must be incapable of delaying an OTP. And the hardest part isn't sending, it's deciding whether to send: preferences, frequency caps, aggregation, quiet hours, and TTLs so we never send stale notifications. I'll design a type registry, a decision layer, isolated priority lanes with reserved provider quota, and idempotent channel senders, then go deep on isolation, duplicates, and governance."
Why this is L6:
- Separates intents by correctness bar
- Commits to isolation and the decision layer before drawing
- Names freshness as correctness
What L7 adds:
- Frames attention as an org-level budget and asks who owns the cap
- Asks about SMS spend and provider contracts as business constraints
❌ Common L5 Trap
"Services publish to Kafka; we have a push worker, an SMS worker, and an email worker that call APNs/FCM, Twilio, and SES, with retries and a preferences check."
Why this misses: Partitioned by channel, not priority — the campaign SMS and the OTP SMS share a worker and a provider quota. No TTL, no dedupe, no caps.
Drill 2: Core Mechanic — Aggregation#
Prompt: "A post gets 500 likes in 10 minutes. How many notifications does the author get?"
Staff Answer
"Ideally 1–2. First like: send immediately if the user hasn't received a like notification for this post in the last hour — speed matters for the first one. Subsequent likes enter a 10-minute aggregation window keyed by (user, post, 'like'). At window close, send one push: 'Alex, Sam and 498 others liked your post,' with collapse-id = user:post:likes so it replaces the earlier push on the lock screen. Windows back off exponentially — 10 min, 30 min, 2 h — while activity continues. The inbox shows a single aggregated item updated in place. State lives in Redis: a hash per aggregation key with count and first N actors, TTL = window + slack."
Why this is L6:
- Balances immediacy (first like) with fatigue (rest aggregated)
- Uses collapse IDs to replace, not stack
- Concrete state and keys
What L7 adds:
- Makes aggregation policies part of the type registry so every social product gets them by default, not by reinvention
Drill 3: "Priority Isolation" — Make It Concrete#
Prompt: "You said P0 is isolated. Show me exactly where campaign traffic could still interfere."
Staff Answer
"Walk the path: API — shared, so P0 has a reserved concurrency slice and campaigns enter through a separate bulk endpoint with its own rate limit. Decision layer — P0 transactional types skip caps and use a fast path with only an opt-out check. Kafka — separate topic, separate consumer group. Workers — separate deployment, autoscaled on P0 lag only. Template rendering and device-token lookups — shared services, so P0 uses a separate client pool with a 50ms timeout and cached tokens. Provider — reserved quota via a separate token bucket; for SMS a separate sender ID/account; for APNs a separate HTTP/2 connection pool. Remaining shared risk: the provider itself degrading globally — covered by failover to provider B for SMS."
Why this is L6:
- Audits every shared resource, not just the queue
- Names the remaining residual risk and its mitigation
What L7 adds:
- Runs a quarterly "campaign during login peak" game day and publishes P0 latency during it
Drill 4: Dependency Down — Preferences Store Unavailable#
Prompt: "The preferences store is down. Do you send notifications?"
Staff Answer
"It depends on the lane, and I'd decide it now, not during the incident. Campaigns: fail-closed — sending marketing to someone who opted out is a legal problem; campaigns pause. Social: fail-closed for push (hold in queue within TTL), still write to inbox. Transactional: fail-open to the default policy — OTPs and fraud alerts are allowed regardless of marketing preferences, and a user can't opt out of their own login code. Decision layer keeps a local cache of preferences for recently active users (5-minute TTL) to cover short blips."
Why this is L6:
- Different fail modes per intent, justified by harm
- Legal constraint drives campaigns to fail-closed
What L7 adds:
- Has legal pre-approve the degraded-mode matrix and records it in the runbook
Drill 5: Hot Key — The Celebrity Fan-Out#
Prompt: "A celebrity with 80M followers posts. Every follower gets a notification."
Staff Answer
"First, should every follower get one? Probably only followers who opted into 'posts from X' — maybe 5–10%. For those, one event fans out to millions of per-user notifications. Don't expand synchronously: the fan-out service streams the follower list in pages of 10K into the P1 lane at a bounded rate (e.g., 100K/s), so 8M notifications take ~80s. Each still goes through the decision layer — cap, quiet hours. TTL of ~2 hours: if the pipeline is behind, late followers are dropped rather than notified hours after. The app backend also needs protection: 8M notification taps in 2 minutes hit the post page, so the post is pre-warmed in cache."
Why this is L6:
- Questions whether fan-out is desired
- Bounded streaming fan-out, TTL, and downstream herd awareness
What L7 adds:
- Coordinates with the feed team — Feed Generation faces the same celebrity fan-out and should share the follower-streaming infrastructure
Drill 6: Multi-Tenant — Fifteen Teams, One Lock Screen#
Prompt: "Fifteen teams send notifications. Opt-out rates are climbing. What do you do?"
Staff Answer
"Measure attribution first: opt-out and disable events are joined to the last N notifications the user received, giving an opt-out rate per type. Then: a global per-user daily cap on non-transactional push (3/day), with candidates ranked by type priority and predicted engagement; aggregation defaults; and a per-type guardrail — a type whose opt-out rate exceeds 2× the median is auto-throttled to 50% and its owner is notified. New types launch at 1% of eligible users and graduate after a week of healthy metrics. Product leadership owns the cap value; the platform enforces it."
Why this is L6:
- Attribution before policy
- Mechanisms that make teams accountable
- Cap owned by leadership, enforced by platform
What L7 adds:
- Makes "notification opt-out rate" a company-level health metric reviewed alongside retention
Drill 7: Build vs Buy#
Prompt: "Why not buy a notification platform (a customer engagement vendor) instead?"
Staff Answer
"For campaigns, I'd strongly consider buying — segmentation, journeys, A/B testing, and compliance tooling are the vendor's product. For transactional, the calculus is different: OTP latency and reliability are core to login, and I want control over provider routing and failover. A common split: build a thin transactional + social platform (decision layer, lanes, senders), buy campaign orchestration, and have the vendor send through our platform's P2 lane — or at least through separate provider accounts so their volume can't hurt our P0 quota. Build cost is ~4–6 engineers ongoing; vendor cost scales with MAUs or message volume and can reach six figures a year."
Why this is L6:
- Splits the decision by intent
- Protects P0 isolation even with a vendor in the path
What L7 adds:
- Negotiates data-export and exit terms so audience and preference data isn't locked into the vendor
Drill 8: Policy Change Without Outage — Changing the Frequency Cap#
Prompt: "Leadership wants to lower the daily push cap from 5 to 3."
Staff Answer
"Shadow first: compute what would have been suppressed under cap 3 for a week and show each team its projected loss. Then A/B: cap 3 for 5% of users, measure DAU, sessions, opt-out rate, and each type's downstream conversion for 2 weeks. Roll out 5 → 25 → 100% with a flag. Teams who lose volume get the ranking signals so they can improve relevance. The cap is config with versioning; rollback is instant. Transactional types are unaffected."
Why this is L6:
- Shadow → experiment → staged rollout
- Transparent per-team impact
What L7 adds:
- Uses the experiment to set a long-term attention budget policy, not a one-off number
Drill 9: Cost#
Prompt: "SMS costs $1.2M a year. Cut it."
Staff Answer
"Break down by type and country first. Typically OTPs dominate, and international routes cost 5–10× US. Levers: (1) move OTP to push-based approval or app-generated codes for users with the app installed — SMS only as fallback; (2) WhatsApp or other OTT channels in countries where they're cheaper and more reliable; (3) route per country by cost-adjusted delivery rate; (4) kill SMS for any non-transactional type; (5) fraud: block SMS pumping — attackers trigger OTPs to premium numbers — with per-country and per-prefix rate limits. SMS pumping alone can be a large share of spend. Realistic saving 40–60%."
Why this is L6:
- Breaks cost down before cutting
- Names SMS pumping fraud, a real and commonly missed driver
What L7 adds:
- Moves auth strategy toward passkeys/push approval as the long-term cost and security fix — coordinating with the identity team
Drill 10: Multi-Region#
Prompt: "We're going multi-region. How does the notification platform change?"
Staff Answer
"Notifications are user-homed: process in the user's home region, where their preferences, tokens, and inbox live. Producers in any region route to the user's home region. APNs and FCM are global endpoints, so any region can send. Dedupe must be in the home region to be consistent. On home-region failure: P0 fails over to a secondary region with a replicated subset — tokens and phone numbers for transactional — accepting that caps and aggregation state may be stale (it's fine to exceed a cap during a regional outage; it's not fine to lose an OTP). P1/P2 wait for recovery within TTL."
Why this is L6:
- Home-region processing keeps dedupe and caps consistent
- Different failover posture per lane
What L7 adds:
- Checks residency constraints for phone numbers and emails before replicating them
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — Black Friday#
Context: Black Friday, 09:00. Three campaigns launched at once (18M, 22M, 30M recipients). Order-confirmation emails are 25 minutes late; customer support is flooded with "did my order go through?"
Questions to Surface First:
- Are order confirmations in the P0 lane, and is email P0 isolated at the ESP level?
- Which resource is saturated — our workers, the ESP account throughput, or ESP-side throttling?
- Who approved three simultaneous campaigns?
Typical L5 Approach: Scales email workers. The bottleneck is the ESP account rate, so more workers produce more 429s.
Staff Approach: Finds order confirmations and campaigns share one ESP account; the ESP is throttling the account. Pauses two campaigns, reserves throughput for transactional, and moves order confirmations to the warm secondary ESP immediately. Afterward: separate ESP subaccounts/IPs for transactional vs marketing, and a campaign scheduler that enforces aggregate P2 throughput.
Principal Approach: Establishes a campaign calendar with a throughput budget per hour on peak days, owned by marketing ops, and makes "transactional isolated at the provider account level" a platform invariant.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check handoff_latency{lane,channel} and provider.throttled_total{provider}. |
| Triage | Email P0 lag 25 min; ESP 429s at account level; three campaigns active. |
| Quick fix | Pause two campaigns; failover order confirmations to secondary ESP. |
| Guardrails | Separate transactional subaccount; aggregate P2 throughput cap; campaign preflight. |
| Post-mortem | Why was provider-level isolation missing for email? Why three simultaneous launches? |
Metrics to Watch: notif.handoff_latency_p95{lane,channel}, provider.throttled_total, campaign.active_count, support.contacts{topic="order_status"}
Organizational Follow-up: Marketing ops owns a peak-day calendar; platform owns isolation and preflight.
Ownership Question: "Who can pause a revenue campaign on Black Friday?" Staff answer: The notification on-call, under a runbook that says transactional latency beats campaign throughput, with automatic notification to marketing ops. Pre-agreed, not negotiated live.
Key Takeaway: "Isolation that stops at your queue isn't isolation. The provider account is a shared resource."
What clears the Staff bar:
- Finds the shared provider account, not the worker count
- Uses a warm secondary provider
- Pre-agreed authority to pause campaigns
Deep Dive 2: Silent Failure — Android Pushes Not Arriving#
Context: Engagement from Android push has dropped 30% over 3 weeks. No errors; FCM returns success for every request.
Questions to Surface First:
- Did app code change how notifications are displayed or how tokens refresh?
- Did we change message priority or TTL?
- Are opens down uniformly or on specific OS versions/manufacturers?
Typical L5 Approach: Checks error logs, finds none, concludes FCM is working and the drop is product-side.
Staff Approach: Treats "success but no engagement" as a delivery issue until proven otherwise. Finds a change 3 weeks ago that set all social pushes to normal priority with a 60-second TTL; devices in battery-optimization states batch normal-priority messages, and a short TTL expires them before they're delivered. Fixes TTL to type-appropriate values and uses high priority only for time-sensitive user-visible types (per FCM guidance, overuse of high priority can be deprioritized).
Principal Approach: Adds end-to-end delivery observability — client-side receipt telemetry sampled at 1% — so "provider accepted" is never mistaken for "delivered" again, and puts send-parameter changes under review.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Compare open rate by platform, OS version, priority, and TTL setting. |
| Triage | Drop isolated to Android, correlates with config change on day 0. |
| Quick fix | Revert TTL; set priority per type. |
| Guardrails | Client receipt telemetry; alert on push.receipt_rate{platform} drop > 10%. |
| Post-mortem | Why was a send-parameter change unreviewed? Why no delivery-level metric? |
Metrics to Watch: push.receipt_rate{platform,os}, push.open_rate{type,platform}, notif.handoff_success_rate
Organizational Follow-up: Send parameters (priority, TTL) owned in the type registry with review.
Ownership Question: "Who owns 'delivered' when the provider says success?" Staff answer: The platform owns measuring it end to end via client receipts. Provider success is an input, not the SLO.
Key Takeaway: "Provider success is handoff, not delivery. Measure from the device, or you'll find out from the engagement graph three weeks later."
What clears the Staff bar:
- Distrusts provider success as a proxy
- Segments metrics to isolate the cause
- Adds device-side measurement
Deep Dive 3: Large-Customer Onboarding — B2B Tenant With Compliance Requirements#
Context: A bank signs up to use your notification platform (you're a B2B SaaS) to send 5M transaction alerts/day by SMS and push. Requirements: 99.9% of alerts handed off within 10s, full audit trail for 7 years, and no alerts during their customers' quiet hours unless fraud-related.
Questions to Surface First:
- Which alerts are fraud (bypass quiet hours) vs informational?
- Does "audit trail" mean content, or metadata only?
- Whose SMS sender identity and registration — ours or theirs?
Typical L5 Approach: Adds the tenant to existing lanes and raises provider limits.
Staff Approach: Puts the bank in a dedicated P0 tenant lane with reserved provider quota and its own sender registration; types classified as fraud vs informational with quiet-hour bypass only for fraud; audit log of metadata (not content) to immutable storage with 7-year retention; per-tenant SLO dashboards.
Principal Approach: Uses the bank as the forcing function for tenant isolation tiers (shared / reserved / dedicated) with pricing, and for a compliance package (audit retention, residency) sold as a product tier.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Tenant lane; reserved SMS throughput 100/s + burst; separate APNs/FCM credentials (theirs). |
| Compliance | Immutable audit log (object lock), metadata only; content hash for dispute verification. |
| SLO | 99.9% handoff < 10s; measured per tenant; credits in contract. |
| Testing | Load test at 3× daily peak; provider failover drill before go-live. |
| Cost | Reserved capacity priced into contract; SMS passed through at cost plus margin. |
Metrics to Watch: notif.handoff_latency{tenant}, audit.write_failures{tenant}, quiet_hours.bypass_total{tenant,type}
Organizational Follow-up: Legal review of audit/retention; security review of credential handling.
Ownership Question: "Who classifies an alert as fraud to bypass quiet hours?" Staff answer: The bank, via type registration. We enforce the classification and log every bypass; we don't decide what's fraud.
Key Takeaway: "Large tenants turn implicit shared resources into explicit contracts. Price the isolation."
What clears the Staff bar:
- Tenant-level isolation and SLOs
- Separates metadata audit from content retention
- Tenant owns classification
Deep Dive 4: Post-Mortem — 2.3M Users Got the Wrong Notification#
Context: A template bug sent 2.3M users a push saying "Your payment of $0.00 failed." Support is overwhelmed; the brand team is involved.
Questions to Surface First:
- How did a broken template reach production at full volume?
- Was there a kill switch, and how fast was it used?
- Can we send a correction, and should we?
Typical L5 Approach: Fixes the template, adds a unit test.
Staff Approach: Identifies three missing controls: templates deployed without preview against real payload samples; new/changed types sent at 100% immediately; no per-type kill switch (took 18 minutes to stop). Adds template validation with required-field checks, staged rollout for template changes (1% → 10% → 100% over an hour with anomaly checks on opens and complaints), and a one-click per-type kill switch that stops dequeue within 10s.
Principal Approach: Makes notification content changes follow the same change-management as code deploys, and defines a correction-message policy (who approves apology sends — comms, not engineering).
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Kill switch the type; purge queued messages of that type. |
| Triage | Template variable amount missing from new payload schema; rendered default 0. |
| Correction | Comms decides on a follow-up; in-app inbox item updated in place with correct info. |
| Guardrails | Render-validation against sampled payloads; staged rollout; kill switch SLO 10s. |
| Post-mortem | Schema change in producer wasn't contract-tested against the template. |
Metrics to Watch: notif.sent{type,template_version}, render.missing_field_total, killswitch.time_to_stop_seconds
Organizational Follow-up: Producer-template contract tests; comms runbook for correction messages.
Ownership Question: "Who decides whether to send a correction to 2.3M users?" Staff answer: Comms and the product owner of the type. A correction is itself a notification with fatigue and brand costs — it's not an engineering decision.
Key Takeaway: "Notification content is a production deploy to millions of lock screens. It needs staged rollout and a kill switch."
What clears the Staff bar:
- Finds missing controls, not just the bug
- Kill switch with a time SLO
- Correction decision routed to comms
Deep Dive 5: Multi-Region Expansion — Launching in the EU and India#
Context: The product launches in the EU and India. EU requires consent tracking and residency for personal data; India's SMS routes require sender registration and templates pre-approved under local regulation (DLT). OTP delivery rates in India on the current provider are 82%.
Questions to Surface First:
- Where do EU users' tokens, phone numbers, and inbox live?
- Who owns the regulatory registration work per country?
- What's an acceptable OTP delivery rate, and what's the fallback?
Typical L5 Approach: Deploys the same stack in two new regions and uses the existing SMS provider.
Staff Approach: EU users homed in EU region (preferences, tokens, inbox, delivery log); consent stored with provenance and enforced in the decision layer. India: add a local SMS provider with registered templates; route per operator by measured delivery rate; add a WhatsApp or voice fallback for OTP if SMS fails in 30s. Target OTP delivery ≥ 95%.
Principal Approach: Creates a country-launch checklist owned jointly by legal, platform, and growth (sender registration, consent model, local providers, quiet-hour norms) so each new country isn't a bespoke project.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Residency | EU home region; only aggregates leave the EU. |
| Regulation | Template pre-registration for India; consent provenance for EU. |
| Providers | Local provider in India; per-operator routing; fallback channel. |
| Measure | sms.delivery_rate{country,operator,provider}; OTP completion rate by country. |
| Rollout | Launch with fallback enabled; shadow-route 10% to compare providers for 2 weeks. |
Metrics to Watch: otp.completion_rate{country}, sms.delivery_rate{country,provider}, consent.missing_total
Organizational Follow-up: Legal owns the regulatory map; platform owns routing; growth owns local campaign rules.
Ownership Question: "Who is accountable for OTP delivery rate in India?" Staff answer: The platform owns the rate and the routing. Legal owns the registration dependency. Both are on the launch checklist with named owners.
Key Takeaway: "Notification delivery is local: carriers, regulations, and norms differ by country. Build a launch checklist, not a one-off."
What clears the Staff bar:
- Residency and consent as design inputs
- Per-country provider routing with measured delivery
- Fallback channel for critical OTP
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Separate transactional, social, and campaign intents and commit to isolating them
- Design a decision layer (preferences, caps, aggregation, quiet hours, suppress-if-active) as a first-class component
- Isolate priority lanes at every shared resource — queue, workers, dependencies, and provider quota
- State an honest guarantee: at-least-once handoff within TTL, with idempotency and collapse making duplicates harmless
- Treat freshness as correctness with per-type TTLs enforced at dequeue and at the provider
- Design multi-provider routing with circuit breakers and warm failover
- Assign ownership: platform owns delivery SLOs; teams own content, targeting, and opt-out rates; legal owns consent
The Bar for This Question#
Mid-level (L4): A queue and workers per channel calling providers, with retries and a preferences table.
Senior (L5): Adds Kafka, templates, device token management, a priority field, and delivery tracking. Competent and scalable, but partitions by channel, promises guaranteed delivery, and leaves fatigue and freshness implicit.
Staff+ (L6): Designs the decision of whether to send, isolates priorities down to provider quota, makes duplicates harmless and stale messages impossible, plans for provider failure, and makes attention a governed resource with named owners. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 The Best Notification System Sends Fewer Notifications#
| Metric | Short-term Effect of More Sends | Long-term Effect |
|---|---|---|
| Sessions | Up | Down after opt-outs |
| Push opt-in rate | Flat | Down — the channel degrades for everyone |
| Uninstalls | Flat | Up |
The Staff position: Suppression is a feature. A platform should be judged partly by what it didn't send.
Why this matters in interviews: It shows you understand the system's real objective, not just its throughput.
10.2 "Guaranteed Delivery" Is a Lie Your Providers Won't Tell You#
The Staff position: Every push provider is best-effort; SMS depends on carriers; email "accepted" says nothing about inboxes. The honest contract is handoff-within-TTL, measured delivery, and a durable inbox.
Why this matters in interviews: Precision about guarantees is a Staff marker.
10.3 Partition by Priority, Not by Channel#
The Staff position: The canonical diagram has one queue per channel. The incident that matters is a campaign and an OTP on the same channel. Priority is the primary partition; channel is secondary.
Why this matters in interviews: It reframes the architecture in one sentence.
10.4 SMS Is a Security and Fraud Surface, Not Just a Channel#
The Staff position: SMS OTP is vulnerable to SIM-swap and is a target for SMS-pumping fraud that runs up costs. It should be the fallback, not the default, for users with the app installed.
Why this matters in interviews: It connects notifications to security, cost, and identity strategy.
10.5 Every Notification Type Needs an Owner Who Can Be Embarrassed#
The Staff position: Types without an owning team and an opt-out metric attached accumulate forever. Registration requires an owner; types with no sends in 90 days are retired.
Why this matters in interviews: Lifecycle governance is what makes a platform sustainable over years.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
At Staff level, the notification system is a platform with lanes and a decision layer. At Principal level, it's the arbiter of a shared, depletable company asset: users' willingness to be interrupted. Every product team draws on it, no team pays for depleting it, and once a user disables notifications, every team — including security and payments — loses the channel. The L7 job is to make that asset visible, budgeted, and governed; to treat provider contracts and SMS spend as financial decisions; and to connect notifications to identity strategy (passkeys vs SMS OTP), compliance (consent across jurisdictions), and brand.
The Org-Level Fault Line#
Central arbitration of attention vs team autonomy over sending.
| Option | What It Buys | What It Costs |
|---|---|---|
| Central arbiter (global caps, ranking, review of new types) | Protects the channel; consistent UX; compliance | Teams feel throttled; ranking disputes escalate to leadership |
| Team autonomy with guardrails (per-type opt-out thresholds, no global cap) | Speed for teams | Sum-of-locally-rational sends still degrades the channel |
| Full autonomy | Maximum team velocity | Tragedy of the commons; opt-out spiral |
The Principal position: Central arbiter for push and SMS (interruptive, scarce, costly), guardrails-only for inbox and email (non-interruptive). Product leadership owns the cap; the platform enforces; teams compete on relevance.
Cost Model#
Assumptions: US-heavy user base, SMS ~$0.008/segment average (international higher), email ~$0.10/1,000, engineer fully loaded ~$250K/year, shared Kafka/Redis allocated by share.
| Scale | Volume | Infra + Provider $/month | Headcount | On-call Load |
|---|---|---|---|---|
| Small | 1M users; 5M pushes, 100K SMS, 2M emails/month | ~$2–4K (SMS ~$800) | 0.5 FTE or vendor | Rare |
| Medium | 50M users; 1.5B pushes, 30M SMS, 300M emails/month | ~$300–350K (SMS ~$240K, infra ~$40K, email ~$30K) | 5–7 engineers | 2–4 pages/month |
| Large | 500M users; 15B pushes, 300M SMS, 3B emails/month | ~$3M+ (SMS dominates) | 15–25 engineers across platform, deliverability, ML ranking | Dedicated 24×7; provider war rooms |
The line item to manage is SMS. Moving 70% of OTPs to push-approval or passkeys at Medium scale saves ~$150K/month — more than the platform team's entire headcount cost.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversibility Cost |
|---|---|---|
| Users disabling push at OS level | One-way (for you) | Re-prompting is limited by the OS; lost opt-ins rarely come back |
| Consent data model | One-way-ish | Retroactively proving consent is impossible |
| Sender identities / domains / IP reputation | One-way-ish | Burned reputation takes weeks to months to rebuild |
| Idempotency and type registry contract | One-way-ish | Every producer integrates against it |
| Queue technology, lane count | Two-way | Internal |
| SMS provider choice | Two-way if multi-provider from the start | Single-provider lock-in makes it one-way |
| Cap value | Two-way | Config with experiment |
The Standard I'd Write#
RFC: User Notification Standard (v1)
Scope: Any push, SMS, email, or in-app notification sent to a user of any product.
MUST:
- Send only through the notification platform; direct provider calls from product services are prohibited.
- Register each type with owner team, priority, TTL, channels, idempotency key source, collapse key, and category for consent.
- Supply an idempotency key per notification.
- Marketing notifications MUST pass central consent and unsubscribe checks at send time.
- New types launch to ≤ 1% of eligible users for 7 days before general availability.
SHOULD:
- Use aggregation for any type expected to fire more than once per user per hour.
- Prefer push or in-app over SMS for anything not security-critical.
- Set TTLs that reflect the notification's value decay.
Exceptions: Approved by the notification platform lead and, for cap bypass, the VP of product.
Success metrics: P0 handoff p95 < 2s during campaigns; push opt-in rate stable or rising quarter over quarter; zero consent violations; SMS spend per MAU down 30% year over year.
What I'd Tell the VP#
Our users' willingness to receive notifications is a shared company asset, and right now every team spends it without a budget. Push opt-in has dropped for three straight quarters, and last month a marketing campaign delayed login codes for 20 minutes. I'm proposing a central notification platform with isolated lanes, so logins and payments are never delayed by campaigns, and a daily attention budget per user that product leadership owns. Separately, SMS is costing us roughly $3M a year and is our weakest login factor; shifting to push approval and passkeys cuts about half of that while improving security. The team cost is 6 engineers, less than the SMS saving alone.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Attention as an asset | "Opt-in rate is a depletable company asset; every team draws on it." |
| Prices the channels | "SMS is 80% of the notification bill; the fix is an identity strategy, not a cheaper provider." |
| Governance design | "Leadership owns the cap, the platform enforces, teams compete on relevance." |
| One-way doors | "OS-level push opt-outs rarely come back. That's the irreversible cost of over-sending." |
| Cross-org connection | "Notifications, identity, and fraud are one conversation about SMS." |
Staff answers that L7 interviewers find insufficient:
- "We'll add a global frequency cap" — with no owner for the value, no dispute process, and no measurement of long-term effect.
- "Multi-provider SMS for resilience" — without addressing cost, fraud, or replacing SMS as the primary factor.
- "Teams own their opt-out rate" — without a mechanism that changes behavior (throttling, review, scorecards).
Appendices
Appendix A: Mechanics in Depth#
A.1 Ingress Dedupe#
key = "dedupe:" + user_id + ":" + idempotency_key
if SET key notification_id NX EX type.dedupe_window_s:
proceed
else:
return 202 { notification_id: GET key, decision: "duplicate" }
A.2 Aggregation Window#
agg_key = user_id + ":" + type.aggregation_key(payload) # e.g. user:post:like
HINCRBY agg:{agg_key} count 1; LPUSH agg:{agg_key}:actors actor (LTRIM 0 4)
if first event and no recent send for agg_key: send now, mark window start
else if no timer: schedule flush at now + window (backoff 10m, 30m, 2h)
flush: render "A, B and N others…", collapse_id = agg_key, send, reset
Aggregation timers use the Distributed Job Scheduler or a delay queue.
A.3 Sender With TTL and Breaker#
on_dequeue(n):
if now >= n.expires_at: record(expired); return
provider = router.pick(n.channel, n.country, breaker_state)
resp = provider.send(n, collapse_id=n.collapse_key, ttl=n.expires_at - now, timeout=3s)
match resp:
ok → record(handed_off)
throttled → requeue with backoff honoring Retry-After (still check TTL)
invalid_token → delete token; record(failed_permanent)
timeout → breaker.record_failure(); retry same collapse_id / same content, ≤ 2 attempts
Appendix B: Data Model#
| Record | Store | Key | Notes |
|---|---|---|---|
| Type registry | Postgres | type_id | Owner, priority, TTL, channels, keys |
| Preferences | Postgres + Redis cache | user_id | Category × channel, quiet hours, timezone |
| Cap counters | Redis | cap:{user}:{yyyy-mm-dd} | TTL 48h |
| Device tokens | Cassandra/DynamoDB | user_id → tokens | Platform, app version, last_refreshed |
| Inbox | Cassandra | (user_id) clustered by created_at DESC | 90-day retention; read state |
| Delivery log | Cassandra → warehouse | notification_id | State transitions, provider IDs |
| Consent log | Append-only store | user_id, category | Provenance, timestamp, source |
Appendix C: Delivery Strategies — Quick Comparison#
| Strategy | Latency | Duplicate Risk | Cost | Use For |
|---|---|---|---|---|
| Push immediate | Seconds | Low with collapse | Free | Most user-visible types |
| Push aggregated | Minutes | Very low | Free | Social activity |
| SMS | Seconds–minutes | Medium (no collapse) | $$$ | OTP fallback, fraud |
| Seconds–minutes | Low | $ | Receipts, digests, campaigns | |
| In-app inbox | Instant on open | None (upsert) | Infra only | Everything with a record |
| Digest (daily/weekly) | Hours | None | $ | Low-urgency updates |
Appendix D: Provider Contract and Client Behavior#
- APNs: HTTP/2;
apns-collapse-id(≤ 64 bytes);apns-expiration;apns-priority10 (immediate) vs 5 (power-considerate); 410 → prune token. - FCM:
collapse_key;ttl;priorityhigh vs normal;UNREGISTERED→ prune token; overuse of high priority for non-user-visible messages can be deprioritized. - SMS: honor provider throttles; per-country sender registration; delivery receipts are partial — treat as signals, not truth.
- Email: separate transactional and marketing domains/IPs; SPF, DKIM, DMARC; List-Unsubscribe header for marketing.
- Client: dedupe by
notification_idon device; inbox read-sync clears badges across devices.
Appendix E: Observability#
notif.handoff_latency_seconds{lane,channel} # platform SLO
notif.expired_total{type}
notif.suppressed_total{type,reason}
notif.duplicate_suspected_total{type}
provider.error_rate{provider,country}, provider.throttled_total{provider}
push.receipt_rate{platform} # client-side sampled
optout_rate{type}, push_disable_rate
sms.cost_per_day{country,type}
| Alert | Threshold | Routes To |
|---|---|---|
| P0 handoff latency | p95 > 5s for 3 min | Platform |
| Provider errors | > 5% for 5 min per provider | Platform |
| Lane lag | P1 > 5 min, P2 > 30 min | Platform |
| Opt-out spike | type opt-out > 2× median | Owning team |
| SMS spend anomaly | daily spend > 150% of 7-day avg | Platform + fraud |
| Bounce rate | > 3% on any campaign | Growth; auto-pause |
Appendix F: Scale Evolution#
| Scale | Design |
|---|---|
| < 1M users | Vendor, or a queue with priority + ESP/SMS APIs |
| 1–50M | Type registry, decision layer, P0/P1/P2 lanes, multi-provider SMS |
| 50–500M | Learned ranking, attention budget, per-country routing, client receipts |
| Multi-region / B2B | User-homed processing, tenant tiers, compliance packages |
What you don't build on day one: learned ranking, send-time optimization, own carrier connections, multi-region failover, tenant tiers.
Appendix G: Multi-Tenancy, Fairness, and Cost#
- Lanes and quotas: P0 reserved; P1/P2 weighted fair across teams; P2 global token bucket.
- Attention budget: per-user daily cap across teams; ranking by type priority × predicted value.
- Chargeback: SMS and email at provider cost per team/type; push by volume share of infrastructure.
- Fairness signal: a team taking > 40% of the P2 lane for > 1 hour needs a scheduled campaign window.
- Lifecycle: types with no sends in 90 days are retired; types without an active owner are disabled.