Hiring BarSupport

Design an Email Delivery Service

Case study88 min read9 diagrams

Technologies referenced in this case study: Apache Kafka · PostgreSQL · Redis · Cassandra · OLAP Databases

Related: Notifications · Webhook Delivery Platform · Rate Limiter · Job Scheduler · Idempotency & Exactly-Once · Multi-Tenancy · Backpressure & Load Shedding · Security Basics

Reading Guide#

Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal View) and the appendices on the MTA scheduler, authentication records and feedback processing
What is an Email Delivery Service? — Why interviewers pick this topic

An email delivery service takes a message from an application — a password reset, a receipt, a weekly newsletter to 4 million subscribers — and gets it into the recipient's mailbox. It exposes an API (and usually SMTP relay) to senders, signs every message so receivers can authenticate it, hands it over SMTP to the recipient's mail servers, and reports back what happened: delivered, deferred, bounced, complained about, opened, clicked, unsubscribed.

The hard part is not speaking SMTP. The hard part is that the receiver decides whether your mail is wanted, and it decides based on your reputation — a score that mailbox providers compute per sending IP and per domain from bounces, spam complaints, spam-trap hits and engagement. Reputation is slow to build, fast to lose, and on shared infrastructure it is pooled across every sender who shares your IPs. One customer who uploads a purchased list can get a whole IP pool throttled or blocklisted, and suddenly password resets from forty unrelated companies land in spam or don't land at all.

Before vs After — the "newsletter ate the password reset" scenario:

Without stream separation:
t=0:        Tenant T9 launches a 6M-recipient campaign on the shared pool, 08:00 local.
t=+2min:    Outbound queue to Gmail: 2.1M messages. Gmail starts returning 421 4.7.28
            (rate limited) on 30% of connections from the pool.
t=+10min:   Password-reset emails for 300 other tenants sit behind the campaign in the
            same per-destination queue. Median send delay: 4s → 23 min.
t=+30min:   Support tickets: "users can't log in". Two tenants' own on-call get paged.
t=+2h:      Campaign drains. Complaint rate for the pool spikes to 0.4% (T9's list is stale).
t=+2 days:  Pool reputation at Gmail drops to "low". 18% of all mail from the pool → spam folder.

With transactional/bulk separation and per-tenant reputation isolation:
t=0:        T9's campaign goes to the bulk stream, on T9's own IP pool, rate-shaped by
            a per-provider throttle that backs off on 421s.
t=+2min:    Transactional stream to Gmail: separate IPs, separate queue, p99 send delay 6s.
t=+10min:   T9's campaign hits 421s; its throttle halves rate. Nobody else notices.
t=+2h:      T9 complaint rate 0.4% → automatic sending pause for T9's bulk stream,
            account review opened, T9 notified with the list segments that drove complaints.
t=+2 days:  Shared transactional pool reputation unchanged. Zero tickets from other tenants.

Why interviewers reach for this question: It looks like "a queue plus workers that send SMTP" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: reputation as a pooled, slowly-moving resource that must be isolated like any other noisy-neighbor resource; per-provider throttling where the receiver sets the rate, not you; feedback loops (bounces, complaints, unsubscribes) that must close in minutes; authentication (SPF, DKIM, DMARC) that is now mandatory for bulk senders; and an SLO — inbox placement — that you cannot measure directly.

Mechanics Refresher: Email Delivery Primitives
PrimitiveHow It WorksProsCons
SMTP handoffLook up the recipient domain's MX records, open TCP (STARTTLS), MAIL FROM / RCPT TO / DATA; a 250 means the receiver accepted responsibilityUniversal; synchronous accept/reject250 ≠ inbox; receivers defer with 4xx at will
4xx deferral vs 5xx rejection4xx = try later (rate limit, greylisting, temp failure); 5xx = permanent (no such user, policy block)Receiver tells you what to doProviders overload codes; enhanced status codes (4.7.x, 5.1.1) carry the real reason
Bounce (DSN)Synchronous 5xx during the SMTP session, or an asynchronous delivery status notification sent to the envelope sender laterTells you the address is badAsync bounces arrive minutes to days later and must be matched to the message (VERP)
Feedback loop (FBL)Provider forwards "this is spam" reports to the sender in ARF formatPer-message complaint signalNot every provider offers one; some give only aggregate spam rates
SPFDNS TXT record listing IPs allowed to send for the envelope domainCheap; ties IPs to a domainBreaks on forwarding; 10-DNS-lookup limit
DKIMSender signs headers + body with a private key; public key published in DNS under a selectorSurvives forwarding; domain-level reputationKey management and rotation per tenant domain
DMARCDomain publishes a policy (none/quarantine/reject) for mail that fails aligned SPF or DKIM, plus where to send reportsStops direct spoofing; required for bulk senders at major providersAlignment mistakes cause legitimate mail to be rejected
Suppression listAddresses that hard-bounced, complained or unsubscribed are never sent to againProtects reputation and complies with lawScope (per tenant? global?) is a policy decision
IP warm-upNew IPs ramp volume over weeks so providers can build a reputation for themAvoids instant throttling of new IPsWeeks of lead time before a dedicated IP is usable at full volume
Open / click tracking1×1 pixel and rewritten links through a redirect domainEngagement data for sendersPixels are prefetched by privacy features; links are clicked by security scanners

For most production systems: Separate transactional and bulk into different streams with different IP pools and queues; sign everything with DKIM on the tenant's own domain and require aligned DMARC; throttle per (IP pool, mailbox provider) adaptively from deferral signals; suppress hard bounces and complaints immediately and per tenant; warm every new IP on a schedule; and treat inbox placement — measured by proxies — as the real SLO. The SMTP client is not the interview. Reputation isolation, receiver-set rate limits and closing the feedback loop are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

An email delivery service is not a queueing question. Everyone can put messages on Kafka and have workers speak SMTP.

It is a shared-reputation and receiver-controlled-throughput question that tests:

  • Whether you notice that the scarce resource is reputation, not CPU or bandwidth — and that it is shared across tenants on the same IPs and domains
  • Whether you let the receiving providers set your send rate, adaptively, instead of pushing as fast as your MTAs can go
  • Whether bounces, complaints and unsubscribes feed back into suppression within minutes, because every repeat send to a complainer costs reputation
  • Whether you define success as inbox placement and time-to-inbox for transactional mail, not "SMTP 250 received"

The key insight: A 250 OK from Gmail means Gmail accepted responsibility for the message — not that a human will see it. The system's real output is reputation, and every design decision either protects it or spends it. Staff candidates design a reputation-isolation system with an SMTP client attached; Senior candidates design an SMTP client with a queue attached.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveAPI → queue → worker pool → SMTPAsks "Transactional, bulk, or both? Our own mail or a multi-tenant platform where customers send to their own lists? What happens to password resets when a customer blasts 6M newsletters?"Asks "Is email a product we sell with a deliverability promise, or plumbing for our own product teams? Who owns sender reputation, and who has authority to stop a tenant from sending?"
IsolationOne queue, scale workers on depth"Separate transactional and bulk streams with separate IP pools and queues; tenants grouped into pools by reputation; high-volume tenants on dedicated IPs"Prices reputation isolation into plans; sets an acceptable-use policy with automated enforcement and a review team
Throughput"Send as fast as the MTAs allow""Per (IP pool, provider) throttle that adapts to 421 deferrals — AIMD, like TCP congestion control. Receivers set the rate"Negotiates and monitors relationships with the top 5 providers (postmaster programs, feedback loops) as a strategic dependency
Feedback"Mark bounced messages failed""Hard bounce and complaint → suppress in under a minute; soft bounces suppress after repeated failures; one-click unsubscribe honored within hours; tenant pause when complaint rate crosses 0.1% for review, 0.3% hard stop"Treats complaint rate as the platform's top-line health metric and ties tenant pricing and onboarding gates to it
Authentication"Add SPF""DKIM on the tenant's domain with 2048-bit keys, aligned DMARC, custom return-path for SPF alignment, scheduled key rotation with DNS pre-publication"Makes "no unauthenticated sending" an org-wide standard; owns the company's own DMARC enforcement roadmap to p=reject
SLO"99.9% of messages sent""Transactional: 99% accepted by the recipient MX within 60s. Deliverability: inbox-placement proxies (seed tests, provider dashboards, engagement) tracked per pool"Defines deliverability as the product SLO and staffs a deliverability team that owns it, separate from infrastructure on-call
Why "isolation" separates levels

L5: "Messages go into one queue; workers pull and send. If we fall behind, autoscale." This works until one tenant's 6M-recipient newsletter fills the queue to Gmail. Gmail starts deferring the pool, every worker is busy retrying deferred campaign mail, and the password resets for everyone else wait behind it. Worse, the campaign's complaints lower the reputation of the IPs that also carry everyone's transactional mail. Autoscaling adds more senders hammering a provider that is already telling you to slow down.

L6: "Two kinds of isolation. Throughput isolation: transactional and bulk are separate streams with separate queues and separate IP pools, so a campaign can never queue ahead of a password reset. Reputation isolation: tenants are grouped into IP pools by their own measured behavior — new and unproven tenants on a quarantine pool, good senders on a trusted pool, large senders on dedicated IPs. A tenant's complaints hurt only the pool they share, and the pool they share is chosen by their track record."

L7: "Reputation is a shared commons; whoever shares it is subsidizing whoever spends it. I'd price dedicated IPs and make the shared trusted pool something tenants earn. And I'd give a trust-and-safety team explicit authority to pause a tenant without engineering approval — the decision has to take minutes, not a meeting."

Why "throughput" separates levels

L5: "Each MTA can do 1,000 messages per second, so 20 MTAs give us 20K/s." The MTAs are not the bottleneck. Gmail, Microsoft and Yahoo decide how fast they accept mail from each IP, and they signal it with 421/451 deferrals and connection refusals. Sending faster than they accept just converts throughput into deferrals, retries and reputation damage.

L6: "The limit is per (sending IP or pool, receiving provider). I keep a throttle per pair — max connections, messages per connection, messages per second — and adapt it: additive increase while acceptance stays above 98%, multiplicative decrease when deferrals cross a few percent. Providers are identified by MX host patterns, not recipient domain, because thousands of domains are hosted by the same few providers."

L7: "Throughput to the top providers is a relationship, not a setting. I'd enroll in their postmaster and feedback programs, track their published sender requirements as external dependencies with owners, and keep a calendar of their policy changes the way we track cloud deprecations."

Why "feedback" separates levels

L5: "Bounced messages are marked failed and reported via webhook." Then the next campaign sends to the same dead addresses again, and the one after that. Repeat sends to invalid addresses and to complainers are exactly what providers measure.

L6: "Feedback closes into suppression automatically. A 5.1.1 user-unknown hard bounce suppresses that address for that tenant in under a minute. A complaint from a feedback loop suppresses it for that tenant across all bulk mail. Soft bounces suppress after, say, 5 consecutive failures over 7 days. Unsubscribes via the one-click header are honored immediately and certainly within 48 hours. The send path checks suppression on every recipient — it's a hot read at 100K lookups per second."

L7: "Suppression is a legal record as well as a reputation tool. I'd define its retention and scope with legal — per tenant, never deleted on request without a re-consent flow — and make it auditable."

Positions to Commit To#

PositionRationale
Transactional and bulk are separate streams: separate queues, IP pools and ideally subdomainsA newsletter must never queue ahead of a password reset or spend the reputation it depends on
Reputation isolation by pool, assigned by measured tenant behaviorReputation is pooled per IP; who shares an IP is a policy decision
Receivers set the rate: adaptive throttles per (pool, provider)Pushing past deferrals buys retries and blocklistings, not throughput
Suppress hard bounces and complaints within a minute, per tenantEvery repeat send to a known-bad address costs reputation
DKIM on the tenant's own domain + aligned DMARC, mandatoryMajor providers require it for bulk senders; domain reputation follows the signing domain
Warm every new IP on a schedule; never cold-start a dedicated IPUnknown IPs sending volume look like spammers
The SLO is time-to-inbox for transactional and inbox placement for bulk — measured by proxies250 OK is an acceptance receipt, not an outcome

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Multi-tenant email platform (customers send to their own users)Untrusted senders and lists; pooled reputation; both transactional and bulkStream separation, reputation pools, per-provider throttles, automated enforcement, per-tenant suppression, event webhooksOne tenant's bad list blocklists a shared pool; campaigns delay resetsTransactional 99% MX-accepted within 60s; tenant complaint rate kept < 0.1%; no repeat send to a suppressed address
First-party email for our own productOne sender, trusted content, one set of domainsTwo streams (transactional, marketing) on our own IPs, strong authentication, suppression shared across productsMarketing burns the reputation the login emails needLogin/receipt mail reaches inbox in < 1 min
Campaign / marketing toolBulk only; latency tolerance in hours; segmentation and scheduling are the productScheduler + list management + send-time throttling; delivery often boughtOver-sending to disengaged segmentsInbox placement and engagement, not latency

🎯 Staff Move: "I'll design a multi-tenant email platform — thousands of customer accounts sending both transactional and bulk mail to their own users. That's where reputation becomes a shared resource and isolation is the whole game. If this were only our own product's mail, I'd still split transactional from marketing, but abuse and tenant enforcement mostly disappear. The interesting problems are reputation isolation, receiver-set throughput and the feedback loop."

Where the Design Splits#

#Fault LineThe Tension
1One Pipeline vs Separate Transactional and Bulk StreamsShared queues and IPs are efficient; separate streams protect latency-critical mail and its reputation at the cost of duplicated infrastructure
2Shared IP Pools vs Dedicated IPsShared pools stay warm and dilute one bad sender; dedicated IPs isolate reputation but need volume and warm-up
3Fixed Send Rates vs Adaptive Per-Provider ThrottlingStatic limits are predictable; receivers change their acceptance rate minute by minute
4Suppression Scope and AuthorityPer-tenant vs global suppression; who may override; how fast feedback closes
5Tracking and the Definition of "Delivered"Open/click tracking gives engagement data but costs privacy, accuracy and sometimes deliverability; what does the SLO actually measure?

How Real Companies Built It#

Why this section belongs here: Mailbox providers publish the rules, email platforms publish their enforcement thresholds, and at least one provider built its whole architecture around stream separation. Naming the numbers shows you know the receiver sets the terms.

Gmail — Authentication and a 0.3% Spam-Rate Line for Bulk Senders#

Google's sender guidelines require every sender to authenticate with SPF or DKIM, send over TLS, have valid forward and reverse DNS, and keep the user-reported spam rate below 0.3%. Senders of more than 5,000 messages a day to Gmail accounts must also have SPF and DKIM, publish a DMARC record (the policy may be none), align the From domain with SPF or DKIM, and support one-click unsubscribe via the List-Unsubscribe-Post and List-Unsubscribe headers for marketing mail. Google recommends keeping spam rate below 0.1% (Gmail sender guidelines). Its FAQ adds that unsubscribe requests should be honored within 48 hours, that transactional messages such as password resets are exempt from the one-click requirement, and that bulk senders above 0.3% are ineligible for mitigation until they stay below it for 7 consecutive days (Gmail sender FAQ).

Staff insight: The receiver writes your requirements. In an interview, quote the thresholds and design the enforcement around them: "I pause a tenant's bulk stream well before it reaches 0.3%, because at 0.3% the provider stops helping us."

Amazon SES — Published Bounce and Complaint Thresholds, Account-Level Suppression#

Amazon SES documents that an account with a bounce rate of 5% or more is placed under review and at 10% or more may have sending paused; for complaints the thresholds are 0.1% for review and 0.5% for a possible pause (SES reputation metrics). Its account-level suppression list automatically adds addresses that hard-bounce or complain, keeps them until removed, counts only hard bounces (not soft), and notes that Gmail does not provide complaint data to SES — a Gmail user pressing "Spam" does not land on the suppression list (SES suppression list).

Staff insight: A platform enforcing thresholds on its own customers is doing exactly what mailbox providers do to the platform — reputation is managed hierarchically. And the Gmail note is a gift in an interview: "The largest provider gives no per-message complaint feed, so I can't rely on suppression alone — I need aggregate spam-rate monitoring per pool and engagement-based list hygiene."

Postmark — Transactional and Broadcast on Separate Infrastructure#

Postmark describes running "parallel but separate sending infrastructures" for transactional and broadcast mail — message streams — so that the two kinds of traffic "do not mix … including IP ranges" (Postmark message streams).

Staff insight: Stream separation is not an optimization; for a provider whose product promise is fast transactional delivery, it is the architecture. Say: "Separate queues aren't enough. If bulk and transactional share IPs, they share reputation, and the campaign's complaints still land on the password reset."

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Workers pull from a queue and send via SMTP""A customer sends 6M newsletters at 8am. What happens to password resets?"Stream separation, head-of-line blocking
"We'll scale MTAs to handle peak""Gmail starts returning 421s on 30% of connections. Now what?"Receiver-set throughput, adaptive throttling
"We put customers on shared IPs""One customer uploads a purchased list and hits spam traps. Who else is hurt?"Reputation pooling, isolation, enforcement
"We give big customers dedicated IPs""A new customer wants a dedicated IP and to send 2M emails tomorrow."Warm-up, saying no with a plan
"We handle bounces""A soft bounce, a hard bounce and a complaint arrive for the same address. What changes?"Suppression semantics, feedback timing
"Our SLO is 99.9% delivered""Delivered means what? Gmail said 250 and the message went to spam."Deliverability as the real SLO
"We track opens""Open rates jumped to 70% overnight. Did your customers get better at email?"Tracking accuracy, privacy prefetch, bot clicks

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Acceptance is the only synchronous part: authenticate the tenant, check quota, drop suppressed recipients, scan content, persist, return a message ID. Everything after is asynchronous. Messages split by stream into transactional and bulk queues keyed by (IP pool, receiving provider) — that key is the unit of throttling, because providers rate-limit per sending IP. The throttle controller adapts each key's rate from the SMTP responses and from the reputation monitor. Feedback — synchronous bounces, asynchronous DSNs, provider complaint reports, one-click unsubscribes — flows into the suppression store within a minute. The metric that tells you the platform is healthy is transactional.time_to_mx_accept_p99 per provider; the metric that tells you it will stay healthy is pool.complaint_rate per pool.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Architecture"Queue + SMTP workers""Acceptance, stream split, per (pool, provider) queues, adaptive throttle, MTAs, feedback into suppression."
Campaign vs reset"Priority field on the queue""Separate streams, separate IP pools; the campaign can't delay the reset or spend its reputation."
Throughput"Scale the MTAs""Receivers set the rate. AIMD throttle per (pool, provider) driven by 421 deferrals."
Bad tenant"Rate-limit them""Quarantine pool for new tenants; automatic pause at complaint and bounce thresholds; review team."
Bounces"Mark failed""Hard bounce and complaint suppress within a minute; soft bounces after repeated failures."
Authentication"SPF""DKIM on the tenant domain, aligned DMARC, custom return-path, scheduled key rotation."
SLO"99.9% sent""Time-to-MX-accept for transactional; inbox-placement proxies per pool for bulk."

Numbers to Bring#

MetricValueWhy It Matters
Gmail spam-rate ceiling< 0.3% required, < 0.1% recommended (Google)Sets your tenant enforcement thresholds — pause well before 0.3%
Gmail bulk-sender threshold5,000+ messages/day to Gmail accountsAbove it: SPF + DKIM + DMARC + alignment + one-click unsubscribe
Unsubscribe processingwithin 48 hours (Google FAQ)Suppression must propagate faster than that — target minutes
SES account review thresholdsbounce ≥ 5%, complaint ≥ 0.1% (AWS)A reference for how platforms police their own tenants
Automated IP warm-up schedule~41 days in one major ESP's documented schedule (Twilio SendGrid)A dedicated IP is weeks, not minutes, from full volume
RFC 5321 retry guidanceretry interval ≥ 30 min; give-up time at least 4–5 daysUpper bound on bulk retry windows; transactional should give up far sooner
Transactional time-to-MX targetp99 < 30–60s (typical)The number login and checkout flows depend on
Share of consumer recipients at top 3–4 providerscommonly 60–80% for consumer lists (illustrative)A handful of throttle keys carry most volume
Concurrent SMTP connections per IP to a large provider~10–50 (typical, adaptive)Throughput per IP is bounded by the receiver, not your hardware
Hard-bounce rate on a clean list< 1–2% (typical)A first send with 8% bounces means a stale or purchased list
Message size~20–100 KB typical HTML; 1–10 MB with attachments1B messages/day × 50 KB ≈ 50 TB/day of bodies
DKIM key size2048-bit RSA (1024 is legacy)DNS TXT record splitting; rotation planning

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on the send API, templates and a Kafka topic, then run out of time before the interviewer asks the only question that matters: "A customer just sent 6 million newsletters from the same IPs your password resets use, and their list is stale. What happens over the next two days?" Compress the plumbing to ~8 minutes and spend the rest on stream separation, reputation isolation, receiver-set throttling, the feedback loop and what "delivered" means.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Tenants send email through an API or SMTP relay — transactional messages like resets and receipts, and bulk campaigns to their subscriber lists. We authenticate and sign each message, deliver it over SMTP to the recipient's provider, process bounces, complaints and unsubscribes, and report events back to the tenant via webhooks and a dashboard."

Then the non-functional requirements, which is where the design lives:

"Three constraints drive everything. One: the receiver controls the rate — Gmail, Microsoft and Yahoo decide how fast they accept mail from each of our IPs, and they decide based on reputation. Two: reputation is shared — every tenant on an IP pool spends from the same account, so isolation is a reputation problem, not just a throughput problem. Three: transactional mail is latency-critical — a password reset that arrives in 20 minutes is a failed login. I'll assume 20,000 tenants, 1 billion messages a day — about 12K per second average and 150K per second at the top of the hour when scheduled campaigns fire — about 15% transactional by count, and roughly 1,500 sending IPs."

Then name the underspecified parts:

"I'd confirm: do tenants bring their own domains — I'll assume yes and require DKIM on them; do we offer dedicated IPs — yes, for tenants above a volume floor; and do we do open and click tracking — yes, opt-out per tenant, with the caveats I'll get to."

🎯 Staff Move: Saying "the receiver controls the rate, and reputation is shared" reframes the problem from "send email" to "manage a shared, receiver-scored resource" — that's the sentence that sets the level.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Message: message_id, tenant_id, stream (transactional / bulk), from_domain, recipients[], body_ref, idempotency_key, campaign_id?, send_at?
  • Recipient delivery: message_id, rcpt_hash, provider (gmail / microsoft / yahoo / other), pool_id, state (queued / deferred / delivered / bounced / suppressed / expired), attempts, next_attempt_at, last_smtp_reply
  • Sending domain: domain, tenant_id, dkim_selectors[] (active + next), return_path_domain, dmarc_status, verified_at
  • IP pool: pool_id, kind (transactional_trusted / bulk_trusted / quarantine / dedicated), ips[], per-IP warmup_day
  • Suppression: (tenant_id, address_hash, scope), reason (hard_bounce / complaint / unsubscribe / manual), created_at, source_message_id

Tenant-facing API:

POST /v1/messages            { from, to[], subject, html, text, stream, idempotency_key, tags[] }
                             → 202 { message_id, accepted: [...], suppressed: [...] }
POST /v1/campaigns           { list_id, template_id, send_at, throttle? }
POST /v1/domains             { domain }  → DNS records to publish (DKIM CNAMEs, return-path, DMARC)
GET  /v1/suppressions?reason=complaint
DELETE /v1/suppressions/{address}   (hard bounce / manual only; complaints need re-consent)
GET  /v1/events?message_id=…        sent | delivered | deferred | bounced | complained | opened | clicked

What the recipient's server sees (headers that matter):

Return-Path: <b+m8f2k1.r77@bounce.mail.tenant.example>     (VERP: encodes message + recipient)
DKIM-Signature: v=1; a=rsa-sha256; d=tenant.example; s=s2026a; h=from:to:subject:date:…
From: Acme <hello@tenant.example>                           (aligned with d= and Return-Path domain)
List-Unsubscribe: <https://u.mail.tenant.example/u/…>, <mailto:unsub+…@mail.tenant.example>
List-Unsubscribe-Post: List-Unsubscribe=One-Click           (bulk stream only)
Feedback-ID: campaign42:tenant123:bulk:esp                  (lets providers aggregate complaints)

🎯 Staff Move: "The return-path encodes the message and recipient, so an asynchronous bounce three hours later maps straight back to one delivery without parsing the bounce body. And the return-path and DKIM d= are both on the tenant's domain, so DMARC alignment passes on SPF and DKIM independently."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk one message in 90 seconds:

  1. A tenant calls POST /messages with a password reset. We authenticate, check the idempotency key, drop the recipient if suppressed, persist the body, return 202 with a message ID — p99 under 50ms.
  2. The stream splitter routes it to the transactional queue for (txn_trusted_pool, gmail) — the provider is resolved from the recipient domain's MX records, cached.
  3. The throttle controller holds a budget per (pool, provider): connections, messages per connection, messages per second. A transactional queue normally has headroom, so the message dispatches in milliseconds.
  4. An MTA signs with the tenant's DKIM key, opens or reuses a TLS SMTP connection from an IP in the pool, and gets 250 2.0.0 OK. That's a delivered event — meaning accepted, not inboxed.
  5. A 4xx → deferred, retried with backoff, and counted by the throttle controller. A 5xx 5.1.1 → bounced and the address is suppressed for this tenant.
  6. Later, if the user hits "Report spam" at a provider with a feedback loop, an ARF report arrives; the address is suppressed for the tenant's bulk mail and a complained event goes to the tenant.

🎯 Staff Move: Say out loud: "The queue key is (IP pool, provider), not tenant or message. That's the grain at which receivers throttle us, so it's the grain at which we schedule." You've now spent ~8 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes this hard is that our throughput is set by receivers, our reputation is shared across tenants, and we can't directly see whether a message reached the inbox. I'd like to go deep on stream separation and reputation pools, adaptive per-provider throttling, the feedback loop into suppression, and authentication and the deliverability SLO. Where would you like to start?"

If no preference: start with stream separation and reputation pools. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Streams and reputation pools (7–8 min)

"Two orthogonal splits. Stream: transactional versus bulk, with separate queues, separate IP pools and a separate tracking and return-path subdomain, because providers score subdomains partly independently. Pool: within each stream, tenants are placed by measured behavior. New tenants start on a quarantine pool with low volume caps for their first 30 days. Tenants with complaint rate under 0.05% and bounce rate under 2% over 30 days graduate to the trusted pool. Tenants sending more than ~1M messages a month to the big providers can buy dedicated IPs, which we warm over 4–6 weeks."

Quantify: "1,500 IPs: maybe 200 for shared transactional, 600 for shared bulk trusted, 100 quarantine, 600 dedicated across ~300 large tenants. A shared bulk IP carrying ~1M messages a day is plenty warm."

Who pays: "A tenant with bad lists pays first — they're paused or moved to quarantine. The quarantine pool's other tenants pay a little, which is why it's for unproven tenants with low caps, not for known-bad ones. Known-bad tenants get suspended, not pooled."


Deep dive 2: Receiver-set throughput (6–7 min)

"Each (pool, provider) has a throttle state: max concurrent connections per IP, max messages per connection, target messages per second. The controller watches SMTP replies in 30-second windows. While acceptance stays above 98%, it increases rate by 5% per window. When deferrals like 421 4.7.28 or 451 4.7.650 exceed 3%, it halves the rate and backs off connections. Providers are keyed by MX host pattern — *.google.com, *.outlook.com, *.yahoodns.net — so a thousand Google-hosted company domains share one throttle."

Quantify: "If Gmail accepts ~30 messages per second per IP from a trusted pool of 600 IPs, the bulk stream can push ~18K/s to Gmail. A 6M-recipient campaign where half are Gmail takes ~3 minutes at full pool rate — but the tenant's own fair share is smaller, so it takes 20–40 minutes, by design."


Deep dive 3: Feedback into suppression (5–6 min)

"Three feedback sources. Synchronous: 5xx during SMTP — classified by enhanced status code; 5.1.1 user unknown is a hard bounce, 5.7.x policy blocks are not hard bounces of the address, they're reputation signals for the pool. Asynchronous: DSNs to the VERP return-path, matched back to the delivery. Complaints: ARF reports from provider feedback loops, matched via our headers. Each lands on a Kafka topic, a suppression writer applies it to a per-tenant suppression table, and the acceptance path reads it — target p99 under 60 seconds from signal to suppression."

"Gmail gives no per-message complaint feed, so for Gmail we watch aggregate spam rate per pool from their postmaster data and per-tenant engagement. Tenants whose Gmail recipients never open or click for 6–12 months get nudged — or forced — to sunset those addresses."


Deep dive 4: Authentication and domains (4–5 min)

"Tenants verify domains by publishing CNAMEs we give them: two DKIM selectors pointing at keys we host, a return-path subdomain pointing at our bounce MX, and we check they publish DMARC. CNAME-delegated DKIM lets us rotate keys without asking the tenant to edit DNS: we publish the next key under the second selector, wait for DNS propagation and a TTL margin — 48 hours — then switch signing. Unverified domains can't send bulk at all."


Deep dive 5: The deliverability SLO and ownership (3–4 min)

"Two SLOs. Transactional: 99% of messages accepted by the recipient MX within 60 seconds, measured per provider — that's ours, and infra on-call pages on it. Deliverability: inbox placement per pool, measured by proxies — seed-list placement tests across providers every hour, provider postmaster reputation, complaint rates and engagement trends. That's owned by a deliverability team, not infra on-call, because the fixes are policy: move a tenant, pause a tenant, slow a pool. Tenants own their lists and content."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: the scarce resource is reputation, scored by receivers and shared by whoever shares our IPs and domains. So transactional and bulk are separate streams; tenants are placed into pools by their own track record; throughput adapts to what each provider accepts; bounces and complaints close into suppression within a minute; every message is DKIM-signed on the tenant's domain with aligned DMARC; and the SLO is time-to-inbox, measured by proxies, not 'sent'."

The evolution closer:

"What I'd build later: per-tenant engagement-based sending — slowing or skipping recipients who haven't engaged in months; BIMI support for verified logos; a dedicated deliverability console for large tenants. What I'd not build: our own content-based spam scoring that tries to out-guess the providers. I'd invest in measuring placement and enforcing list hygiene instead."

🎯 Staff Move: End on who can stop a tenant. "When a tenant's complaint rate crosses 0.1%, their bulk stream pauses automatically and trust-and-safety reviews within a business day. Engineering is never the approval step for protecting the pool."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Template obsession10 min on templating, personalization, list storageOne line; templates are tenant-side or a separate service
SMTP protocol tourExplains HELO, MAIL FROM, RCPT TO in detail"A 250 is an acceptance receipt; 4xx means slow down; 5xx means stop"
Throughput by hardwareSizes MTAs by CPUSizes by receiver acceptance per IP, then counts IPs
One priority queuePriority flag for transactionalSeparate streams with separate IPs — priority doesn't separate reputation
Bounces as a reportEvents onlySuppression closes the loop in under a minute
"Delivered" as the SLO99.9% sentTime-to-MX for transactional, placement proxies for bulk

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Email is the interview where your system's output is graded by a third party using criteria it doesn't fully disclose, and the grade is shared with everyone who sends from your IPs. You cannot make Gmail accept faster. You cannot make a tenant's list cleaner. You cannot see whether a message landed in the inbox or the spam folder. The candidate must design a system that protects a shared, externally scored resource from its own tenants — with automated enforcement, isolation by track record and feedback loops that close in minutes. That is Staff work: blast-radius design where the blast is reputational and the damage lasts for days.

It also has a deceptive cost structure. Sending is cheap — a message costs a few hundredths of a cent in compute and bandwidth. Reputation damage is expensive: a blocklisted pool can take days to recover, delays every tenant on it, and generates churn from customers who never did anything wrong. The design's value is almost entirely in what it prevents.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"It's Tuesday morning. Your shared bulk pool's spam rate at Gmail has climbed from 0.08% to 0.25% over 48 hours, and inbox placement on seed tests dropped from 92% to 71%. No tenant is above your thresholds individually. Walk me through what you do, who you call and what you change."

A candidate who answers with attribution (rank tenants by complaint contribution, not rate — a big tenant at 0.09% may contribute more than a small one at 0.4%), the levers (move the top contributors to their own or quarantine pools, throttle the pool's Gmail rate to slow damage, force engagement-based sunsetting for the worst segments), the owner (deliverability team, not infra on-call), the customer communication (proactive notice to affected tenants with the change) and the prevention (contribution-based alerts, not just per-tenant rate alerts) has run an email platform. A candidate who says "find the bad tenant and block them" has sent email.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Multi-tenant email platform → reputation isolation and enforcement

  • Constraint: untrusted senders, unknown list quality, pooled IP and domain reputation, mixed transactional and bulk
  • Strategy: stream separation, reputation pools by track record, adaptive per-provider throttles, automated thresholds, per-tenant suppression, event webhooks
  • Failure mode: one tenant's list poisons a pool; campaigns delay resets; compromised tenant accounts send phishing
  • Who pays for imperfection: other tenants (reputation and latency), trust-and-safety (abuse), support (tickets about "spam folder")

First-party email for our own product → stream separation and authentication

  • Constraint: one organization, trusted content, our own domains
  • Strategy: transactional and marketing on separate subdomains and IPs; suppression shared across all our products; DMARC at p=reject for the root domain
  • Failure mode: marketing sends to disengaged users and burns the reputation login emails need
  • Who pays: every product team whose critical mail shares the domain

Campaign / marketing tool → scheduling and list hygiene

  • Constraint: bulk only; latency tolerance in hours; segmentation, A/B tests and send-time optimization are the product
  • Strategy: build the campaign layer, buy delivery (or run it on a modest MTA fleet)
  • Failure mode: over-sending; ignoring engagement
  • Who pays: the sender's own domain reputation

2.2 When NOT to Build an Email Delivery Service#

  • You send under ~1M messages a month. A managed sending service is cheaper than one engineer, comes with warm shared pools, and publishes its own enforcement. Your work is authentication on your domain and suppression handling.
  • Email isn't core to your product. Running MTAs means owning provider relationships, blocklist removals, abuse response and an on-call for problems that look like "Gmail is slow." Buy it.
  • You need guaranteed real-time delivery. Email is store-and-forward with receiver-controlled pacing. For time-critical alerts, use push or SMS as primary and email as the record.
  • You want to deliver internal notifications between services. Use a broker. SMTP is for crossing organizational boundaries.

🎯 Staff Insight: "Most companies should run their own email program — domains, authentication, suppression, list hygiene — on someone else's MTAs. You build MTAs when delivery is the product or when volume makes per-message pricing exceed a team's cost."

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Transactional vs bulk mix — most candidates design one queue for both
  • Who owns the lists — tenants' lists are untrusted; your own are not
  • Domain ownership — whether mail is signed with the tenant's domain or yours changes whose reputation accrues
  • Definition of "delivered" — SMTP acceptance vs inbox placement
  • Enforcement authority — who can pause a paying customer, and how fast
  • Peak shape — campaigns scheduled at :00 and :30 produce 10–20× spikes; Black Friday week can be 3–5× a normal week

Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.

2.4 Precise Terminology#

TermWhat It MeansWhy It Matters in the Interview
MTAMail transfer agent — the server that speaks SMTP outboundYour fleet; throughput bounded by receivers
MXDNS record naming a domain's inbound mail serversProvider identification for throttling
Envelope sender / Return-PathMAIL FROM address; where bounces goSPF is checked against its domain; VERP encodes IDs
Header FromThe address the user seesDMARC alignment is checked against its domain
Hard bouncePermanent failure for the address (5.1.1 user unknown)Suppress immediately
Soft bounceTemporary failure (mailbox full, 4xx)Retry; suppress only after repeated failures
Deferral4xx reply meaning "try later," often rate limitingThrottle signal, not an address problem
ComplaintRecipient marked the message as spamThe most damaging signal; suppress and count
Feedback loop (FBL)Provider program forwarding complaints as ARF reportsPer-message complaint data where offered
Spam trapAddress that never opted in; any mail to it signals bad list practicePurchased or scraped lists hit them
IP poolSet of sending IPs used togetherThe unit of shared reputation
Warm-upRamping volume on a new IP over weeksCold IPs at volume look like spammers
AlignmentSPF or DKIM domain matches the header From domainRequired for DMARC to pass
Inbox placementFraction of delivered mail that lands in the inbox, not spamThe real SLO; only measurable by proxies

🎯 Staff Insight: If the interviewer says "guarantee delivery," ask: "Delivered to the MX or to the inbox? I can guarantee we hand it off and record the receiver's answer. Inbox placement is a reputation outcome that I manage, measure by proxies and never fully control."


3. Where the Design Splits#

Every email decision has a technical side (queues, IPs, DNS records) and an organizational side (who may send, who gets paused, who owns reputation). Interviewers grade the second side.

3.1 Fault Line 1: One Pipeline vs Separate Transactional and Bulk Streams#

The tension: A single pipeline is efficient: one set of queues, one IP fleet, every IP kept warm by combined volume. But bulk mail is spiky, tolerant of delay and generates complaints; transactional mail is small, latency-critical and rarely complained about. Sharing infrastructure means sharing both queue position and reputation.

ChoiceWhat WorksWhat BreaksWho Pays
One pipeline, FIFOSimplest; best IP utilizationCampaigns queue ahead of resets; campaign complaints hit reset reputationEvery tenant's end users (late resets)
One pipeline, priority queueResets jump the local queueStill shares IPs → still shares reputation and provider throttlesTenants (reputation coupling)
Separate queues, shared IPsLatency isolationReputation still coupled; provider throttle per IP still sharedTenants (reputation coupling)
Separate streams: queues, IP pools, subdomainsLatency and reputation isolationMore IPs to keep warm; tenants must classify mail honestlyPlatform (IP cost), trust-and-safety (misclassification policing)
Diagram: 3.1 Fault Line 1: One Pipeline vs Separate Transactional and Bulk Streams

Staff default: "Separate streams end to end: separate queues, separate IP pools, separate return-path and tracking subdomains. Tenants declare the stream per message; we audit it — a 'transactional' message sent to 500K recipients with an unsubscribe footer gets reclassified, and repeat abusers lose transactional access."

When to deviate:

  • Very small senders (< 50K/month): a single shared pool is fine; separation can't keep tiny pools warm.
  • Triggered lifecycle mail (welcome series, cart abandonment): it's bulk by behavior — complaint-prone — even though it's triggered one at a time. Route it to bulk.

🧭 Principal Move: "The classification is a policy with teeth. I'd publish what counts as transactional — a message the recipient's own action triggered and expects — and have trust-and-safety audit samples weekly. Misclassification is how bulk reputation leaks into the transactional pool."

❌ Common L5 Trap: "We'll add a priority field so transactional messages jump the queue." Priority fixes local queue order only. The reset still leaves through the same IPs Gmail is deferring because of the campaign, and still inherits the campaign's complaint damage.


3.2 Fault Line 2: Shared IP Pools vs Dedicated IPs#

The tension: Providers score reputation per sending IP (and per domain). Shared pools aggregate many tenants: they stay warm, and one tenant's bad week is diluted. But dilution is also contagion — a bad enough tenant drags the pool down for everyone. Dedicated IPs isolate a tenant's reputation completely, but need consistent volume to stay warm and weeks of warm-up to start.

ChoiceWhat WorksWhat BreaksWho Pays
One big shared poolAlways warm; simpleContagion: worst tenant sets everyone's reputationGood tenants
Tiered shared pools by track recordGood senders share with good senders; new ones quarantinedPool placement policy to build and explainPlatform (policy, tooling)
Dedicated IPs per large tenantFull isolation; tenant owns their reputationNeeds ~volume floor and 4–6 week warm-up; idle IPs go coldTenant (cost, warm-up wait)
Dedicated IPs for everyoneMaximum isolationThousands of cold, low-volume IPs look suspiciousTenants (worse delivery)
Diagram: 3.2 Fault Line 2: Shared IP Pools vs Dedicated IPs

Staff default: "Tiered shared pools: quarantine for new tenants with daily volume caps, trusted for tenants who've earned it, and dedicated IPs for steady high-volume senders. Placement is automatic from 30-day metrics, re-evaluated daily, and a tenant can move down instantly but up only after a sustained clean window. Known-bad tenants are suspended, not dumped into a 'bad' pool that we'd still be paying to keep alive."

When to deviate:

  • A tenant with highly irregular volume (one huge send a quarter) shouldn't get dedicated IPs — they'll be cold every time. Keep them on trusted shared.
  • Regulated or high-trust senders (banks) may require dedicated IPs for audit reasons; warm them with steady transactional volume first.

🧭 Principal Move: "A shared trusted pool is a commons, and good tenants subsidize its protection. I'd make 'trusted pool' an earned status with published criteria and price dedicated IPs so heavy senders fund their own warm-up and isolation."

❌ Common L5 Trap: "Give every customer their own IP so nobody affects anybody." Most tenants send too little to keep an IP warm; providers see thousands of low-volume IPs with no history and treat them with suspicion. Isolation without volume is worse delivery.


3.3 Fault Line 3: Fixed Send Rates vs Adaptive Per-Provider Throttling#

The tension: Static per-provider limits (say 50 msgs/s per IP to Gmail) are easy to reason about. But providers change acceptance continuously based on reputation, time of day and their own load, and they communicate through deferrals. A static limit is either too slow on good days or too fast on bad ones — and too fast means deferrals, retries and a reputation that gets worse because you're ignoring the signal.

ChoiceWhat WorksWhat BreaksWho Pays
No throttle — send as fast as MTAs allowMax throughput when receivers are happyDeferral storms; connection blocks; reputation damageEvery tenant on the pool
Static per-provider limitsPredictable; simpleWrong most of the time; manual tuning by on-callOn-call (tuning), tenants (slow or blocked)
Adaptive AIMD per (pool, provider)Tracks what receivers accept; backs off automaticallyNeeds a central controller and per-key statePlatform team (controller)
Adaptive + per-tenant fair share inside each keyPlus: one tenant's campaign can't take the whole pool's Gmail budgetTwo-level schedulingPlatform team (scheduler)
Diagram: 3.3 Fault Line 3: Fixed Send Rates vs Adaptive Per-Provider Throttling

Staff default: "Adaptive AIMD per (pool, provider), with provider identified by MX pattern. Inside each key, weighted fair share across tenants so a campaign gets its share of the Gmail budget, not all of it. The controller also obeys warm-up caps per IP — a day-10 IP has a hard daily ceiling regardless of how happy Gmail looks. Deferred messages are retried on a schedule — transactional for up to a few hours, bulk for up to 24–72 hours — well inside RFC 5321's 4–5 day give-up guidance, because a newsletter delivered on day four is worse than no newsletter."

When to deviate:

  • Long tail of small providers: a single conservative default throttle per MX host is fine; they're 20–40% of volume spread across thousands of domains.
  • Provider-published limits: where a provider documents connection or rate limits, use them as ceilings for the adaptive controller.

🎯 Staff Insight: "Deferrals are the receiver telling us our rate. Retrying faster is arguing with the referee. The throttle's job is to keep deferrals around 1–2% — enough to know we're near the limit, not so many that we look abusive."

Retrying deferred mail faster than the receiver accepts it multiplies the load you send into an already-throttled provider; capped, backed-off retries let acceptance recover instead of collapsing.

❌ Common L5 Trap: "We'll retry deferred messages every minute so they get through quickly." Every retry is another attempt Gmail counts while it's already telling you to slow down. Aggressive retries turn a rate limit into a block.


3.4 Fault Line 4: Suppression Scope and Authority#

The tension: Suppression protects reputation and is often legally required (unsubscribes). But scope matters: if alice@example.com complains about Tenant A's newsletter, should Tenant B's password reset to her be blocked? And who may remove an address from suppression — the tenant, support, nobody?

ChoiceWhat WorksWhat BreaksWho Pays
Global suppression (all tenants)Maximum reputation protectionBlocks legitimate mail from unrelated tenants; leaks one tenant's signal to othersOther tenants' users (missing resets)
Per-tenant, all streamsRespects tenant boundariesA complaint about marketing blocks the tenant's receipts tooTenant's users (missing receipts)
Per-tenant, per-stream by reasonComplaint/unsubscribe → bulk only; hard bounce → all streamsMore rules to explainDocs / support
Platform-global hard-bounce cache + per-tenant everything elseAvoids every tenant re-discovering a dead addressMust expire — addresses get recycledPlatform (expiry policy)
Diagram: 3.4 Fault Line 4: Suppression Scope and Authority

Staff default: "Per-tenant, by reason and stream. Hard bounce suppresses the address for that tenant on every stream. Complaints and unsubscribes suppress bulk mail only — the user still gets their receipts and resets. A platform-wide hard-bounce hint list, expiring after 6–12 months, warns tenants before they send to a known-dead address. Policy blocks (5.7.x) never suppress the address — they're about us, not the recipient. Tenants can remove hard-bounce entries; complaint entries require re-consent evidence and a support review."

When to deviate:

  • Spam-trap hits suggest the whole list source is bad — suppress the address globally and audit the tenant's list.
  • Legal holds (e.g. regional regulations requiring proof of opt-out): retain suppression records indefinitely with timestamps and source message IDs.

🧭 Principal Move: "Suppression is where reputation policy meets the law. I'd have legal sign off on scope and retention, and make 'never re-send to a complainer without re-consent' a platform guarantee tenants can't override through the API."

❌ Common L5 Trap: "If an address bounces, mark the message failed." Without suppression, the next campaign sends to it again — and providers count repeat sends to invalid addresses as a strong sign of poor list hygiene. The bounce event is a report; the suppression entry is the fix.


3.5 Fault Line 5: Tracking and the Definition of "Delivered"#

The tension: Tenants want opens and clicks. Open tracking uses a remote image; click tracking rewrites every link through your redirect domain. Both have become unreliable: privacy features in some mail clients prefetch images (so "opens" happen without a human), and corporate security scanners click every link (so "clicks" happen without a human — including unsubscribe links). Link rewriting also puts your tracking domain's reputation inside every tenant's message. Meanwhile, "delivered" from SMTP means only accepted.

ChoiceWhat WorksWhat BreaksWho Pays
Open + click tracking on by default, shared domainRich engagement dataInflated opens; bot clicks; shared tracking domain's reputation couples tenantsTenants (bad data), all tenants (domain contagion)
Tracking on, tenant-branded tracking domainsDomain reputation per tenant; links look consistentPer-tenant DNS and TLS certificatesPlatform (cert automation)
Tracking off by default for transactionalFewer moving parts in critical mail; no rewritten reset linksLess dataTenants (no engagement data on resets)
"Delivered" = SMTP 250Simple, observableSays nothing about spam-folder placementTenants (false confidence)
Deliverability = placement proxiesMeasures the real outcomeIndirect: seeds, provider dashboards, engagement trendsDeliverability team (analysis)

Staff default: "Click tracking on for bulk with tenant-branded tracking domains and automated certificates; off by default for transactional so reset links point straight at the tenant. Open tracking available but flagged as approximate — we label machine opens where we can detect prefetch patterns and exclude them from engagement-based sunsetting. Bot clicks are filtered by timing — clicks within a second or two of delivery, or every link in a message at once — and unsubscribe links require one-click POST or a confirmation page, so a scanner's GET doesn't unsubscribe anyone. 'Delivered' in our events means accepted by the MX; we never call it inboxed."

When to deviate:

  • Privacy-sensitive tenants (health, finance): tracking off entirely, at their request.
  • Tenants who need proof of receipt: email can't give it; offer a portal link the recipient must open, and track that.

🎯 Staff Insight: "If engagement data drives suppression and sunsetting, inflated opens keep dead addresses alive. Tracking accuracy is a deliverability concern, not just an analytics one."

❌ Common L5 Trap: "Our SLO is 99.9% delivered, measured by 250 responses." A pool can show 99.9% delivered while 30% goes to spam. The candidate has measured the part of the pipeline they control and declared victory on the part they don't.


4. When It Breaks#

4.1 The Campaign That Ate the Password Resets#

t=0:       08:00 local. 140 tenants have campaigns scheduled for 08:00. Combined: 31M messages.
           Pre-separation design: one queue per provider, shared IPs.
t=+1min:   Gmail queue depth: 11M. Gmail acceptance per IP drops as volume spikes; 421s at 18%.
t=+4min:   Password resets for 2,000 tenants queued behind campaign mail. p99 enqueue→250: 4s → 19 min.
t=+9min:   Tenants' own login-failure alerts fire. Our support queue: 120 tickets.
t=+12min:  Page: transactional.time_to_mx_accept_p99{provider=gmail} > 300s.
t=+20min:  On-call manually moves resets onto a reserve IP range. Cold IPs → more 421s.
t=+70min:  Campaign backlog drains. Resets recover.

Detection: transactional.time_to_mx_accept_p99{provider}; queue.depth{stream, provider}; smtp.deferral_rate{pool, provider}.

Mitigation: separate transactional stream with its own warm IPs; per-tenant fair share inside bulk keys; spread scheduled campaigns across the hour (jitter send times by up to 10–15 minutes unless the tenant opts out).

Prevention: stream separation as architecture, not a flag; load test the 08:00 spike monthly; keep a warm reserve in the transactional pool (≥ 30% headroom).

Owner: email platform on-call (latency); product (campaign send-time jitter as default).

4.2 The Purchased List — Spam Traps on the Shared Pool#

t=0:       Tenant T44 (3 weeks old, graduated early by a manual override) uploads 2.4M addresses.
t=+1h:     First send: 9% hard bounces, 0.6% complaints at FBL providers. Thresholds trip at +40 min,
           but the campaign is 80% sent by then.
t=+6h:     A major blocklist lists 40 IPs in the shared bulk pool (spam-trap hits).
t=+8h:     Microsoft starts rejecting the pool with 5.7.x policy codes. 1,900 tenants affected.
t=+2 days: Delisting requests filed; Microsoft reputation recovers over 5 days.

Detection: tenant.hard_bounce_rate_first_send > 5% (early in the send, not after); blocklist.listed_ips{pool}; smtp.policy_block_rate{pool, provider}.

Mitigation: pause T44 immediately; drain T44's remaining queue to a holding state; move unaffected tenants off listed IPs to clean ones in the same tier (carefully — moving bad traffic spreads the damage).

Prevention: first sends from new tenants are paced — send the first 1–5% of a large list, measure bounces and complaints for 30–60 minutes, then release the rest; list-quality checks at upload (role addresses, syntax, known-dead domain ratio); no manual overrides of quarantine graduation without trust-and-safety sign-off.

Owner: trust-and-safety (tenant enforcement), deliverability team (delisting, pool health).

4.3 The Deferral Storm — Retries Fight the Throttle#

t=0:       Microsoft lowers acceptance for the bulk pool after a complaint spike. 421/451 deferrals: 35%.
t=+5min:   Retry scheduler: deferred messages retried at 1, 2, 4 min. Attempt volume to Microsoft: 2.8×.
t=+15min:  Microsoft starts refusing connections from 120 IPs outright.
t=+30min:  Bulk to Microsoft effectively stopped. Queue: 9M messages. Transactional pool unaffected.
t=+3h:     Retries capped; throttle halved; acceptance recovers to 95% over 4 hours.

Detection: smtp.attempts_per_accepted{pool, provider} (> 1.5 means retries are dominating); smtp.connection_refused_rate.

Mitigation: retries must pass through the same throttle as first attempts — a retry is not free capacity; exponential backoff starting at 15–30 minutes for deferrals.

Prevention: one throttle per (pool, provider) gates all attempts; retry budget per key (retries ≤ 20% of attempts); alert on attempts-per-accepted ratio.

Owner: email platform (scheduler and throttle).

4.4 DKIM Rotation Breaks DMARC#

t=0:       Scheduled DKIM key rotation for 6,000 tenant domains. Signing switches to selector s2026b.
t=+0:      Bug: s2026b public keys were published for only 4,100 domains (DNS provider API rate limit).
t=+10min:  1,900 domains' mail fails DKIM. SPF alignment passes for most — but not for tenants using
           our shared return-path domain. ~600 domains now fail DMARC.
t=+20min:  Tenants with p=reject: their mail is rejected by receivers. Resets fail outright.
t=+45min:  Page from bounce-rate anomaly: 5.7.x DMARC rejects concentrated on 600 domains.

Detection: dkim.selector_published{domain, selector} checked by a resolver probe before switching; bounce.dmarc_reject_rate{domain}.

Mitigation: roll signing back to the old selector (it's still published); re-publish missing keys.

Prevention: rotation is gated per domain on a verified DNS lookup of the new key from multiple public resolvers, 48+ hours after publication; switch in waves (1% → 10% → 100%) watching DMARC failures; never remove the old key until 7 days after the switch.

Owner: email platform (key management).

4.5 Bot Clicks and the Mass Unsubscribe#

A corporate email-security scanner starts following every link in inbound mail to detonate threats — including our tenants' unsubscribe links, which used a simple GET. Over a weekend, 380,000 business recipients across 900 tenants are unsubscribed. Tenants see "unsubscribe rate 14%" on Monday and open urgent tickets.

Detection: unsubscribe.rate{tenant} spike; clicks within 2 seconds of delivery; all links in a message clicked within the same second; user agents from known scanner ranges.

Mitigation: identify scanner-driven unsubscribes by timing and source; restore them with an audit trail, and only where no human signal exists.

Prevention: unsubscribe via RFC 8058 one-click POST (scanners send GETs) or a landing page with a confirm button; bot-click filtering for engagement metrics.

Owner: email platform (tracking edge), with legal sign-off on any restore of unsubscribes.

4.6 The Suppression Migration That Forgot Complaints#

During a move to a new suppression store, the backfill imported hard bounces and unsubscribes but skipped complaint records stored in a legacy table. For six days, tenants sent bulk mail to 2.1M previous complainers. Gmail spam rate on the bulk pool rose from 0.07% to 0.22%, and placement fell from 91% to 78%. Nothing paged, because per-tenant complaint rates stayed under thresholds.

Detection: suppression.hit_rate{reason} dropping sharply after a deploy; per-pool spam rate trend; daily reconciliation count between old and new stores.

Prevention: migrations with dual-read and count reconciliation by reason; alert when suppression hit rate drops > 20% day-over-day.

Owner: email platform (data migration), deliverability team (pool impact).

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Campaign delays resetstransactional.time_to_mx_accept_p99All tenants (without separation)Stream separation, send-time jitterEmail on-call
Bad list on shared poolFirst-send bounce rate, blocklist.listed_ipsEvery tenant on the poolPaced first sends, pause, IP rotationTrust-and-safety + deliverability
Deferral stormsmtp.attempts_per_accepted > 1.5Pool × providerRetries through throttle, backoffEmail platform
DKIM rotation failureDMARC reject spike per domainAffected domainsRollback selector; gated rotationEmail platform
Bot unsubscribesUnsubscribe-rate spike, click timingTenants' listsOne-click POST, bot filterEmail platform + legal
Suppression regressionsuppression.hit_rate{reason} dropPool reputationDual-read migration, reconciliationEmail platform
Compromised tenant API keyContent/URL anomalies, volume spikePool + phishing victimsAuto-pause, key revocationTrust-and-safety
Bounce processor lagfeedback.lag_seconds > 300Repeat sends to bad addressesScale consumers; block bulk sends if lag > 1hEmail on-call

🎯 Staff Insight: The platform's leading indicator is pool complaint contribution, not any one tenant's rate. "I'd alert on the top-10 tenants by share of the pool's complaints, not just tenants over 0.1%. Ten tenants at 0.09% can sink a pool that no single-tenant alert will ever catch."


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framing"Queue and send email reliably"Names reputation as the shared, receiver-scored resource and transactional latency as the critical pathAsks whether deliverability is a product promise and who owns sender reputation across the company
IsolationPriority queueSeparate streams with separate IPs; reputation pools by track record; dedicated IPs with warm-upPrices pools and dedicated IPs; publishes earned-trust criteria; staffs enforcement
ThroughputSizes MTAs by CPUAdaptive AIMD per (pool, provider); retries through the throttle; warm-up capsTreats top providers as strategic dependencies with tracked policy changes
FeedbackBounce eventsHard bounces and complaints suppress within a minute, scoped by tenant and stream; soft-bounce rules; one-click POSTSuppression as a legal record with retention and re-consent policy
AuthenticationSPFDKIM on tenant domains via CNAME, aligned DMARC, custom return-path, gated rotationOrg-wide "no unauthenticated mail" standard; owns DMARC enforcement roadmap
Operations"99.9% sent"Time-to-MX SLO for transactional; placement proxies per pool; contribution-based alertsDeliverability team separate from infra on-call; complaint rate as product KPI

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Names the real resource"The scarce thing is reputation, and it's shared by everyone on the same IPs."
Lets receivers set the rate"Deferrals are the provider telling us our rate — the throttle adapts per pool and provider."
Separates streams properly"Separate queues aren't enough; separate IPs, or the campaign's complaints land on the reset."
Closes the loop"Complaint to suppression in under a minute, scoped to the tenant's bulk mail."
Defines delivered honestly"A 250 is an acceptance receipt. Placement is measured by seeds and provider data."
Knows the thresholds"I pause bulk at 0.1% complaints for review, well before the 0.3% line where providers stop helping."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Scale MTAs to handle peak"Receivers bound throughput; more senders means more deferrals
Priority flag as the reset answerShares IPs, throttles and reputation with campaigns
Bounces without suppressionRepeat sends to bad addresses destroy reputation
No authentication beyond SPFBulk mail without DKIM and DMARC is rejected or spam-foldered at major providers
Per-message dedicated IPs for every tenantCold, low-volume IPs deliver worse
"99.9% delivered" as the SLOMeasures SMTP acceptance, not the outcome

5.4 Common False Positives#

  • Deep SMTP protocol knowledge ≠ email platform design. Knowing every reply code doesn't address reputation pooling or tenant enforcement.
  • A spam-content classifier ≠ deliverability. Providers weight sender reputation and engagement far more than your guess at their filters.
  • "We use a managed service" ≠ no design. Even on bought MTAs, authentication, suppression, list hygiene and stream separation are yours.
  • Detailed retry math ≠ a throttle. A beautiful backoff curve that doesn't share the per-provider budget with first attempts still causes deferral storms.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minMulti-tenant, transactional + bulk; receivers set rate; reputation shared; numbers
Entities & API3–5 minMessage, recipient delivery, domain, pool, suppression; headers that matter
Architecture5–10 min≤ 8 boxes; acceptance → stream split → (pool, provider) queues → throttle → MTAs → feedback
Streams + pools10–18 minSeparate streams with separate IPs; pools by track record; dedicated IPs and warm-up
Throttling18–24 minAIMD per (pool, provider); retries through throttle; fair share inside keys
Feedback + suppression24–31 minBounce classes, VERP, FBLs, scope by tenant and stream, one-click unsubscribe
Auth + SLO31–38 minDKIM via CNAME, DMARC alignment, rotation; time-to-MX and placement proxies
Pivot / wrap38–45 minAbuse, multi-region, cost; close on who can pause a tenant

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"A new customer wants a dedicated IP tomorrow"Warm-up, saying no with a planTrusted shared pool now; dedicated IP warmed over 4–6 weeks with their steady traffic
"Gmail placement dropped 20 points"Diagnosis without direct dataContribution analysis by tenant, seed tests, postmaster data; isolate, throttle, sunset
"A tenant's API key is stolen and used for phishing"Abuse responseAuto-pause on URL and volume anomalies; key revocation; notify; delist
"Make it multi-region"Reputation is per IPIPs are region-bound; keep per-region pools warm; residency for bodies and logs
"Customers want read receipts"Honesty about trackingOpens are approximate; no reliable receipt in email; offer portal-link confirmation
"How do you test deliverability changes?"Measuring the unmeasurableSeed lists, canary pools, staged rollout by pool, placement deltas

6.3 What to Deliberately Skip#

  • Templating and personalization — tenant-side or a separate service; one sentence.
  • SMTP command sequence — "250 accepts, 4xx defers, 5xx rejects."
  • List management UI — mention list hygiene, not the screens.
  • Spam content scoring internals — providers decide; you measure placement.
  • MIME encoding details — irrelevant to the design.

6.4 Follow-Up Questions to Expect#

  1. "A customer sends 6M newsletters at 8am from the IPs your resets use. What happens?"
  2. "Gmail starts deferring 30% of your connections. Walk me through the next hour."
  3. "A new customer uploads a 2M-address list on day one. What do you do before sending it?"
  4. "A hard bounce, a soft bounce and a complaint arrive for the same address from different tenants. What gets suppressed, where?"
  5. "How do you rotate DKIM keys for 6,000 tenant domains without breaking DMARC?"
  6. "What does 'delivered' mean in your SLO, and how do you measure inbox placement?"
  7. "Your pool's spam rate is climbing but no tenant is over threshold. What now?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design an email delivery service."

Staff Answer

"First: is this our own product's mail, or a multi-tenant platform where customers send to their own users? And transactional, bulk, or both? I'll assume multi-tenant with both: 20,000 tenants, a billion messages a day, 12K per second average and around 150K per second when scheduled campaigns fire at the top of the hour, roughly 15% transactional.

Three constraints shape everything. Receivers set our throughput — the big providers decide how fast they accept mail from each IP. Reputation is shared — every tenant on an IP pool spends from the same account. And transactional mail is latency-critical — a 20-minute reset is a failed login. I'll go: entities and the headers that matter → architecture keyed by IP pool and provider → stream separation and reputation pools → adaptive throttling → feedback into suppression → authentication → what 'delivered' means and who owns it."

Why this is L6:

  • Distinguishes first-party from multi-tenant and transactional from bulk before drawing anything
  • Names reputation and receiver-set rate as the constraints, with numbers
  • Previews an outline that ends with the SLO definition and ownership

What L7 adds:

  • Asks whether deliverability is a product promise and who has authority to pause a tenant
  • Asks whether several product teams already run their own sending — the consolidation question
  • Frames pool complaint rate as the platform's top-line health metric
❌ Common L5 Trap

"An API writes messages to Kafka; a pool of workers consumes them, renders templates and sends via SMTP; bounces come back as events; we autoscale workers on lag."

Why this misses: Every piece is reasonable, but it treats email as a throughput problem the sender controls. It has no answer for a campaign delaying resets, a provider deferring the pool, or one tenant's bad list damaging everyone's reputation.


Drill 2: The Reset Behind the Newsletter#

Prompt: "A customer sends 6 million newsletters at 8am. Meanwhile, other customers' users are requesting password resets. What happens in your design?"

Staff Answer

"The newsletter goes to the bulk stream; resets go to the transactional stream. They don't share queues, IPs or return-path subdomains. The campaign is throttled by the bulk pool's per-provider budget, and inside that budget by the tenant's fair share, so it takes 20–40 minutes to drain — intentionally. Resets go out on transactional IPs that carry low, steady, rarely-complained-about volume, with p99 time-to-MX acceptance around 5–10 seconds.

If the newsletter's list is stale and generates complaints, the damage lands on the bulk pool — or on the tenant's dedicated IPs if they have them — not on the IPs the resets use. If complaints cross 0.1% during the send, the remaining campaign pauses automatically."

Why this is L6:

  • Separates on IPs and reputation, not just queue priority
  • Quantifies both drain time and reset latency
  • Shows enforcement acting mid-send

What L7 adds:

  • Makes send-time jitter across the hour the default product behavior to flatten the 08:00 spike
  • Notes that transactional pool headroom (≥ 30%) is a capacity commitment with a cost
❌ Common L5 Trap

"Transactional messages get a high-priority flag, so workers always pick them first."

Why this misses: Priority only reorders the local queue. The reset still leaves through IPs that Gmail is deferring because of the campaign volume, and those IPs inherit the campaign's complaints.


Drill 3: Make It Concrete — IPs, MTAs and Storage#

Prompt: "Size the sending fleet and storage for a billion messages a day."

Staff Answer

"Throughput is bounded by receivers per IP, so I size IPs first. Assume a trusted IP sustains ~20–30 accepted messages per second to a large provider during peaks and ~1–2M messages a day overall. A billion a day at ~1M per IP per day is ~1,000 busy IPs; with headroom for peaks, quarantine and dedicated tenants, ~1,500.

MTAs: an MTA process with async I/O handles a few thousand concurrent SMTP sessions. Peak 150K messages/s at ~300ms per transaction with connection reuse is ~45K concurrent sessions — ~20–30 MTA hosts, plus headroom; each binds many IPs. The MTAs are cheap; the IPs and their reputation are the asset.

Storage: bodies at ~50KB average → 50TB/day; keep 7 days for retries and resends → 350TB in object storage, deduplicated per campaign (one body, many recipients) which cuts it 10–50× for bulk. Events: ~4 per message (accepted, delivered, opened, clicked) → 4B rows/day at ~200 bytes → 800GB/day raw, ~100–200GB compressed in a columnar store. Suppression: maybe 2B entries total at ~40 bytes → 80GB — a sharded key-value store with a bloom filter in front for 150K lookups per second."

Why this is L6:

  • Sizes by receiver acceptance per IP before hardware
  • Notices campaign body deduplication changes storage by an order of magnitude
  • Sizes the suppression lookup as a hot read path

What L7 adds:

  • Prices it: IPs are cheap to rent but expensive to warm; event storage dominates infra spend
  • Tiered event retention by plan as a pricing lever
❌ Common L5 Trap

"A billion a day is 12K per second; each server can send 1,000 per second, so 12 servers plus redundancy."

Why this misses: It ignores the peak shape (10–15× at the top of the hour), and more importantly that receivers cap acceptance per sending IP. Twelve servers on twelve IPs would be deferred into oblivion.


Drill 4: The Dedicated IP Request#

Prompt: "A new enterprise customer signs today. They want a dedicated IP and plan to send 2 million emails tomorrow. What do you tell them?"

Staff Answer

"Yes to the dedicated IP, no to tomorrow on it. A cold IP sending 2M messages on day one looks exactly like a spammer who just rented a server; providers will defer or block it. Tomorrow's send goes out on the trusted shared pool — if their list passes checks — or paced through quarantine if it doesn't.

Meanwhile we start warming two dedicated IPs with their steady traffic — transactional or their most engaged segment first: a few thousand messages a day per provider, roughly doubling while deferral and complaint rates stay healthy, reaching full volume in 4–6 weeks. The controller enforces daily caps per IP per provider and overflows excess to the shared pool so nothing waits. I'd also check their list: when was it last mailed, what's its source? A list not mailed in 12 months will bounce heavily wherever it goes."

Why this is L6:

  • Says no to the unsafe part with a concrete alternative
  • Warm-up schedule with engaged traffic first and overflow to the shared pool
  • Checks list quality before blaming infrastructure

What L7 adds:

  • Builds the warm-up lead time into the enterprise sales process and contract
  • Prices dedicated IPs to include warm-up support and a deliverability review
❌ Common L5 Trap

"Provision a new IP from the pool, assign it to them, and they can send tomorrow."

Why this misses: Provisioning is trivial; reputation isn't. The first day's send defines the IP's reputation with every major provider, and a cold IP at 2M volume starts that reputation in the hole.


Drill 5: Gmail Starts Deferring#

Prompt: "Gmail starts returning 421 on 30% of connections from your bulk pool. Walk me through the next hour."

Staff Answer

"Automatically: the throttle for (bulk_trusted, gmail) sees deferrals above 3% in a 30-second window and halves its rate and connection count; it keeps halving until deferrals drop under ~1–2%, then increases slowly. Deferred messages back off — first retry after 15–30 minutes — and retries pass through the same throttle, so they don't add load. Transactional is a different pool and doesn't notice.

Then I look for cause: is the deferral code a rate limit or a reputation signal? Is it every IP or a subset? Did a tenant start a large send in the last hour, and what's their complaint contribution? If one tenant is driving it, pause or slow that tenant's Gmail traffic specifically. Bulk queue to Gmail grows — that's acceptable; a newsletter 2 hours late is fine. If the deferrals persist for hours, the deliverability team checks postmaster reputation and seed placement."

Why this is L6:

  • The throttle reacts without a human, and retries respect it
  • Distinguishes rate-limit deferrals from reputation problems
  • Accepts bulk delay as the right victim

What L7 adds:

  • Tracks provider-specific deferral patterns over months to set seasonal capacity plans
  • Maintains a relationship channel with major providers for escalations
❌ Common L5 Trap

"Retry deferred messages with short backoff and add more MTAs so throughput recovers."

Why this misses: More MTAs and faster retries send more attempts at a provider that is explicitly asking for fewer. Deferrals become connection refusals, then policy blocks.


Drill 6: The Compromised Tenant#

Prompt: "A tenant's API key leaks. Overnight, it sends 3 million phishing emails from your platform. How do you detect and contain it, and how do you prevent it next time?"

Staff Answer

"Detection should fire within minutes: volume 20× the tenant's baseline, new From domains or display names, URLs not seen before in their mail, URL reputation hits at scan time, a spike in hard bounces from a list we've never seen. Any two of those auto-pause the tenant's sending and revoke the key — pausing a paying customer at 2am is the right call when the alternative is our pool on blocklists.

Containment: purge queued messages for that key; move affected IPs out of rotation if they got listed; file delisting requests; notify the tenant. Prevention: scoped API keys (send-only, domain-restricted, IP allowlists), anomaly detection on content and volume, and a per-tenant daily volume ceiling that requires a request to raise. Abuse is not an edge case for an email platform; it's a constant."

Why this is L6:

  • Multi-signal automated pause, with authority to act without a human
  • Containment covers queued mail, IPs, blocklists and the tenant
  • Prevention via scoped credentials and volume ceilings

What L7 adds:

  • Funds a trust-and-safety function with on-call and clear authority
  • Shares abuse indicators with providers and industry groups
❌ Common L5 Trap

"Rate-limit each tenant to a maximum number of emails per second."

Why this misses: A rate limit spreads the phishing out; it doesn't stop it. Three million messages at a modest rate still fit in one night. Detection must look at what's being sent, not just how fast.


Drill 7: Build vs Buy#

Prompt: "Should we run our own MTAs or use a managed email sending service?"

Staff Answer

"Depends on whether email delivery is our product. If we're a SaaS company sending our own product mail, buy: a managed service gives warm shared pools, provider relationships, feedback-loop enrollment and published enforcement. What stays ours regardless: domains and DKIM/DMARC, stream separation, suppression, list hygiene and the decision of what we send.

Build when delivery is the product — we're selling email to tenants — or when volume makes per-message pricing exceed a team's cost. At roughly $0.10 per thousand, a billion a month is ~$100K a month, which pays for an infra team; at 50M a month it's ~$5K and building is irrational. Even then, the hard part isn't the MTA software — mature open-source and commercial MTAs exist — it's IP reputation management, deliverability operations and abuse response."

Why this is L6:

  • Criteria and a break-even with numbers, not preference
  • Separates what can be bought (MTAs, pools) from what can't (program, authentication, suppression)
  • Locates the build cost in operations, not software

What L7 adds:

  • Prices the deliverability and abuse teams into the build case, not just infra
  • Keeps an exit path: our domains and suppression data are portable between providers
❌ Common L5 Trap

"SMTP is a simple protocol; we can run Postfix on a few servers and save money."

Why this misses: The protocol is simple; reputation, warm-up, provider throttling, feedback loops, blocklist removal and abuse response are not, and the candidate hasn't costed any of them.


Drill 8: Moving to DMARC p=reject Without an Outage#

Prompt: "Security wants our own company domain at DMARC p=reject next month. Twelve teams and four vendors send mail as our domain. How do you ship it?"

Staff Answer

"First, inventory: publish p=none with aggregate reporting (rua) if we haven't, and spend 2–4 weeks reading the reports. They show every source sending as our domain and whether it passes aligned SPF or DKIM. Expect surprises — a survey tool, an old billing system, a recruiting vendor.

Fix each source: give vendors a subdomain with its own DKIM keys, or delegate DKIM via CNAME; move internal senders through the platform. Then ramp: p=quarantine on low-risk subdomains first, then the main domain — using whatever percentage or testing controls the current DMARC specification and our receivers honor — watching reports and support tickets for a week at each step, then p=reject in the same stages. Exceptions get a dated owner. 'Next month' is too fast if the reports show unknown senders — I'd tell security the date the data supports."

Why this is L6:

  • Uses DMARC reports as the inventory before enforcement
  • Subdomains and delegated DKIM for vendors
  • Staged percentage ramp with owners — a policy rollout, not a DNS edit

What L7 adds:

  • Makes "every sender as our domain must be registered and authenticated" a standing policy with procurement review for new vendors
  • Sets the timeline by data, negotiated with security as a shared risk decision
❌ Common L5 Trap

"Update the DMARC TXT record to p=reject."

Why this misses: Every unauthenticated or misaligned sender — vendors, legacy systems — has its mail rejected instantly, including invoices and offer letters. There's no inventory, no ramp and no rollback plan beyond noticing the damage.


Drill 9: The Cost of Email#

Prompt: "Where does the money go in this system, and how would you cut it by 30%?"

Staff Answer

"Infra buckets: event storage and analytics (billions of rows a day), message body storage, MTAs and IPs, the tracking edge. Event storage is usually the largest: cut it by dropping per-recipient 'processed' events tenants never query, compressing in a columnar store and tiering retention by plan — 30 days default, longer as a paid add-on. Body storage: deduplicate campaign bodies and expire after the retry window.

But the biggest cost is people and reputation incidents: deliverability and abuse operations, support tickets about spam folders, churn after a blocklisting. The 30% that matters may be fewer incidents — paced first sends and earlier tenant pauses avoid the multi-day pool recoveries that cost the most."

Why this is L6:

  • Identifies event storage as the dominant infra line with concrete levers
  • Notices that body deduplication is a large, cheap win
  • Recognizes incident cost as larger than infra cost

What L7 adds:

  • Attributes cost per tenant — including the reputation cost of their complaints — and reflects it in pricing
  • Treats retention tiers as product packaging, coordinated with sales
❌ Common L5 Trap

"Use fewer, larger MTA servers and cheaper instances."

Why this misses: MTAs are a small slice of the cost. The candidate optimizes the cheap path and misses storage and the reputation incidents that drive support load and churn.


Drill 10: Multi-Region#

Prompt: "We're expanding to the EU with data residency. How does email delivery work across regions?"

Staff Answer

"IP reputation is per IP, and IPs live in regions, so each region gets its own pools — and they must be warmed before they carry real traffic. I'd start warming EU pools 6–8 weeks before launch with EU tenants' transactional traffic. EU tenants' message bodies, events and suppression lists stay in the EU; the acceptance API, queues, MTAs and tracking edge run there.

Suppression is the tricky part: it's per tenant, so an EU tenant's list lives in the EU. Platform-wide hints like known-dead addresses can be replicated as hashes. On regional failure, EU transactional mail can fail over to a warm standby pool within the EU — failing over to US IPs would breach residency and push volume onto IPs that have never sent that tenant's mail."

Why this is L6:

  • Recognizes IP reputation as regional and requiring warm-up lead time
  • Residency carried through bodies, events and suppression
  • Failover stays within region and uses warm IPs

What L7 adds:

  • Plans the regional launch timeline around warm-up, not deployment
  • Defines residency per tenant contract and audits it with marker tests
❌ Common L5 Trap

"Deploy the same stack in the EU and fail over between regions with DNS."

Why this misses: Cold EU IPs can't take full volume on launch day, and DNS failover to the other region's IPs moves mail onto IPs with no history for those tenants — and out of the residency boundary.


8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Black Friday Morning#

Context: At 06:00 on Black Friday, 1,400 tenants fire campaigns within 20 minutes: 220M messages, 6× a normal morning. Transactional time-to-MX at Microsoft climbs from 8s to 9 minutes; Gmail is fine. The on-call escalates to you.

Questions to Surface First:

  • Is the delay in acceptance, in queueing, or in SMTP to Microsoft specifically?
  • Is the transactional pool itself being deferred, or is something shared between streams saturated?
  • What changed versus last year: new tenants, new IPs, a shared component?
  • Are Microsoft deferral codes rate-limit or reputation codes?

Typical L5 Approach: Adds MTA hosts. Throughput to Microsoft doesn't change, because the bottleneck is Microsoft's acceptance, and the extra hosts open more connections that get deferred.

Staff Approach: Finds that transactional and bulk use separate IPs but share the MTA hosts' DNS resolver cache, which is thrashing under 6× lookup volume; Microsoft MX lookups time out, and transactional sessions stall waiting on DNS. Scales the resolver tier, pins MX records for top providers with longer cache, confirms transactional recovers in 4 minutes. Bulk to Microsoft stays throttled — correctly.

Principal Approach: Treats Black Friday as a scheduled peak with a dedicated readiness review: a shared-component inventory (DNS, TLS session caches, event pipeline) checked for stream isolation, load tests at 8× in October, and tenant guidance to spread sends across the morning with send-time jitter on by default.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Split time-to-MX by stage: acceptance 40ms (fine), queue wait 2s (fine), SMTP session setup 8 min (bad) — only to Microsoft.
TriageSession setup dominated by DNS MX resolution timeouts; resolver cache hit rate 99% → 60% under campaign burst across many recipient domains.
Quick fixScale resolvers ×3; pin MX for top 20 provider patterns with 1h minimum TTL locally.
GuardrailsSeparate resolver pools for transactional MTAs; alert on dns.mx_lookup_p99 > 200ms.
Post-mortemInventory every shared dependency between streams: DNS, TLS, event bus, body store.

Metrics to Watch: transactional.time_to_mx_accept_p99{provider}, smtp.session_setup_p99, dns.mx_lookup_p99, dns.cache_hit_rate

Organizational Follow-up: product ships send-time jitter as default for campaigns scheduled on :00 and :30.

Ownership Question: "Who owns stream isolation?" Staff answer: The email platform team owns it end to end — including shared dependencies like DNS — and proves it with a load test that saturates bulk while measuring transactional latency.

Key Takeaway: "Stream separation that stops at the IP layer isn't separation. Every shared dependency is a path for the campaign to reach the reset."

What clears the Staff bar:

  • Localizes the delay by stage and by provider before adding capacity
  • Finds the shared dependency underneath separated streams
  • Accepts bulk delay as correct while fixing transactional

Deep Dive 2: Silent Failure — 99.8% Delivered, 30% in Spam#

Context: A large tenant says their activation emails "stopped working" two weeks ago: activation rate fell from 62% to 41%. Your dashboard shows 99.8% delivered (250 OK) and no bounce spike.

Questions to Surface First:

  • What do seed-list placement tests show for this tenant's mail, by provider?
  • Did anything change in their mail — content, links, sending domain, tracking domain?
  • Did anything change in ours — pool assignment, IPs, DKIM selectors?
  • What does provider postmaster data say about their domain's reputation?

Typical L5 Approach: Tells the tenant the messages were delivered, with 250 responses as proof.

Staff Approach: Runs seed tests: Gmail inbox 52%, Microsoft 88%. Finds the tenant's activation mail was moved to a different shared transactional pool two weeks ago during an IP rebalancing, and that pool's shared click-tracking domain was recently listed on a domain blocklist because another tenant's links were abused. Disables click tracking for transactional by default, moves the tenant to a branded tracking domain, files delisting.

Principal Approach: Makes "delivered" honest across the product: events say "accepted by provider," dashboards show placement estimates from seeds, and every shared reputation surface — IPs, return-path domains, tracking domains — is inventoried and monitored, with per-tenant branded domains as the default for anything visible in a message.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Seed placement for the tenant's last 3 sends; compare to two weeks ago. Gmail dropped 40 points.
TriageDiff config: pool reassignment two weeks ago; new pool's shared tracking domain listed on a domain blocklist 15 days ago.
Quick fixDisable link rewriting for transactional; tenant on branded tracking domain; request delisting.
GuardrailsDomain-blocklist monitoring for every shared domain; placement checks before and after pool reassignments.
Post-mortemShared tracking domains were a reputation coupling nobody inventoried.

Metrics to Watch: seed.inbox_rate{provider, pool}, blocklist.listed_domains, tenant.engagement_rate trend, pool.reassignments

Organizational Follow-up: deliverability team adds placement checks to the IP rebalancing runbook; tenants with high engagement drops get proactive outreach.

Ownership Question: "Who owns placement for a tenant?" Staff answer: The tenant owns their content and lists; we own every shared reputation surface we put in their mail — IPs, return-path, tracking domain — and must detect when one of ours degrades.

Key Takeaway: "A 250 proves the handoff. Seeds, engagement and provider data are the only evidence of the outcome."

What clears the Staff bar:

  • Doesn't accept SMTP acceptance as proof of success
  • Finds a shared reputation surface beyond IPs
  • Adds placement checks to operational changes that move tenants

Deep Dive 3: Large-Customer Onboarding — Migrating a 50M-Subscriber Sender#

Context: A media company with 50M subscribers is migrating from another provider. They send a daily newsletter (40M recipients) plus 5M transactional messages a day. They want to cut over in two weeks.

Questions to Surface First:

  • What's their current reputation — domain and IPs — at the big providers?
  • Will they bring their sending domains and DKIM, or start fresh?
  • How engaged is the 40M list — when did each segment last open or click?
  • Can they export their suppression list, including complaints and unsubscribes?

Typical L5 Approach: Provisions 40 dedicated IPs, imports the list and switches the newsletter over on day one. Cold IPs at 40M a day are deferred and spam-foldered for weeks.

Staff Approach: Imports their full suppression list first — non-negotiable. Keeps their domain (its reputation travels with them) and sets up DKIM on new selectors. Warms 30–40 dedicated IPs over 5–6 weeks using the most engaged segments first — subscribers who opened in the last 30 days — while the old provider carries the remainder. Ramps by provider, with daily placement checks.

Principal Approach: Turns large migrations into a productized playbook with a deliverability engineer assigned, a contract clause for parallel running with the old provider, and an onboarding gate: no cutover without suppression import and engagement data.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Suppression import: 6.2M entries (bounces, complaints, unsubscribes). Domain verification; DKIM selectors via CNAME.
TriageSegment list by engagement: 14M active in 30 days, 11M in 90 days, 15M older.
Quick fixWeek 1: transactional (steady, engaged) on new IPs. Weeks 2–6: newsletter ramp, most-engaged first.
GuardrailsPer-provider caps per IP per day; auto-overflow to old provider via their API during overlap.
Post-mortem (pre-mortem)What if the 15M stale segment is spam traps? Re-permission campaign before ever sending it from us.

Metrics to Watch: ip.warmup_day, smtp.deferral_rate{ip, provider}, seed.inbox_rate{tenant}, tenant.complaint_rate

Organizational Follow-up: sales learns that dedicated-IP migrations take 4–6 weeks, and quotes accordingly.

Ownership Question: "Who owns the stale 15M?" Staff answer: The tenant owns the list and the decision; we own refusing to put it on our IPs until it's been re-permissioned or sunset.

Key Takeaway: "You can migrate a list in a day. You migrate a reputation over six weeks."

What clears the Staff bar:

  • Suppression import before any send
  • Keeps the tenant's domain, warms new IPs with engaged traffic first
  • Refuses to send stale segments rather than absorbing the damage

Deep Dive 4: Post-Mortem — The Shared Pool Blocklisting#

Context: Forty IPs in the shared bulk pool were listed on a major blocklist for spam-trap hits. For three days, 1,900 tenants saw elevated rejections at providers that consult it. You're leading the post-mortem.

Questions to Surface First:

  • Which tenant(s) sent to the traps? How did their mail reach the shared trusted pool?
  • Were our thresholds checked during the send or only after?
  • Why did the listing affect 1,900 tenants — is the pool too big?
  • How long did detection, delisting and remediation each take?

Typical L5 Approach: Bans the offending tenant and requests delisting.

Staff Approach: Finds a 3-week-old tenant manually graduated to the trusted pool by a sales request, and thresholds evaluated hourly after send rather than during. Ships paced first sends (1–5% sample, then hold), in-send threshold evaluation every 5 minutes, and removes manual graduation without trust-and-safety approval. Splits the trusted pool into smaller sub-pools so one listing affects a fraction of tenants.

Principal Approach: Treats reputation as a governed shared asset: trust-and-safety owns graduation with documented criteria, sales cannot override it, and blast-radius sizing for pools (tenants per pool, volume per IP) is a reviewed standard like cell sizing.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Pause the tenant. Identify listed IPs. Stop routing new mail to them; shift to clean IPs in the same tier.
TriageTenant graduated by override at day 21; list was purchased; first send 9% hard bounces.
Quick fixDelisting requests with remediation evidence; tenant suspended pending review.
GuardrailsPaced first sends; in-send thresholds; sub-pools of ≤ 200 tenants; graduation requires 30 days + T&S approval for exceptions.
Post-mortemThe override path was the root cause; the hourly evaluation was the amplifier.

Metrics to Watch: blocklist.listed_ips{pool}, tenant.first_send_bounce_rate, enforcement.time_to_pause_seconds, pool.tenant_count

Organizational Follow-up: sales enablement on why graduation can't be expedited; affected tenants receive a credit and an explanation.

Ownership Question: "Who decides which tenants share an IP?" Staff answer: Trust-and-safety, by published criteria enforced by the platform. Neither sales nor engineering can override it without a recorded exception.

Key Takeaway: "Who shares an IP is a policy decision. Make it a governed one."

What clears the Staff bar:

  • Finds the policy failure (override) behind the technical one
  • Evaluates thresholds during the send, not after
  • Shrinks the pool's blast radius

Deep Dive 5: Multi-Region Expansion — EU Residency#

Context: The company is launching EU data residency. EU tenants' message bodies, recipient addresses and events must stay in the EU. All sending currently happens from US IPs.

Questions to Surface First:

  • Which data counts: bodies, addresses, events, suppression lists, tracking click logs? (All of them.)
  • How long do new EU IPs need to warm before carrying EU tenants' full volume?
  • Do EU tenants with US recipients still send from EU IPs? (Yes — residency is about the tenant's data, not the recipient's location.)
  • How does suppression work for a tenant with both EU and US accounts?

Typical L5 Approach: Deploys the stack in the EU and moves EU tenants over on launch day. EU IPs are cold; deferrals spike; tenants complain about delivery.

Staff Approach: Starts warming EU pools 8 weeks before launch with EU tenants' transactional mail (bodies already in-region by then). Moves suppression, events and tracking logs for EU tenants into the EU. Launch moves bulk last, by engagement segment, over 3–4 weeks.

Principal Approach: Defines residency as a tenant attribute honored by every platform that touches customer data, audited with marker messages whose content is scanned for outside the region weekly; sets regional launch dates by IP warm-up lead time.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Data inventory: bodies, addresses, events, suppression, click logs, analytics exports.
TriageAnalytics warehouse in the US receives raw events with addresses → residency violation.
Quick fixEU events stay in the EU; warehouse gets aggregated, address-free metrics.
GuardrailsMarker-message test; weekly scan of US stores for EU markers.
Post-mortem (pre-launch)Warm-up starts 8 weeks before launch; EU failover pool also warmed.

Metrics to Watch: residency.marker_found_outside_region (must be 0), eu.smtp.deferral_rate, eu.transactional.time_to_mx_accept_p99

Organizational Follow-up: legal approves the inventory; tenant success schedules migrations.

Ownership Question: "Who proves residency?" Staff answer: The email platform proves it for its stores with automated marker tests; compliance owns company-level attestation.

Key Takeaway: "In email, a new region isn't ready when it's deployed. It's ready when its IPs are warm."

What clears the Staff bar:

  • Plans warm-up lead time into the launch
  • Includes events, suppression and click logs in residency scope
  • Proves residency with tests

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why sender reputation — scored by receivers, pooled per IP and domain — is the scarce resource in email delivery
  • Separate transactional and bulk into streams with separate queues, IP pools and subdomains, and say why priority queues aren't enough
  • Place tenants into reputation pools by measured behavior, and warm dedicated IPs over weeks
  • Design adaptive per-(pool, provider) throttling that respects deferrals, with retries sharing the same budget
  • Close bounces, complaints and unsubscribes into scoped suppression within a minute
  • Configure DKIM on tenant domains, aligned DMARC and safe key rotation
  • Define the SLO honestly: time-to-MX for transactional, placement proxies for bulk
  • Detect and contain abuse from compromised tenant accounts

The Bar for This Question#

Mid-level (L4): Builds an API, a queue and workers that send SMTP with retries. Works for one well-behaved sender. No stream separation, no throttling, no suppression beyond logging bounces.

Senior (L5): Adds a priority flag for transactional mail, SPF and DKIM, bounce webhooks, autoscaling and dedicated IPs for big customers. The gap: priority doesn't separate reputation; MTAs are sized by CPU instead of receiver acceptance; bounces aren't closed into suppression; new IPs are used cold; "delivered" means SMTP 250. The design works in a demo and fails on the first stale list.

Staff+ (L6): Frames the problem as protecting shared, receiver-scored reputation within the first five minutes. Separates streams down to IPs and subdomains, places tenants by track record, throttles adaptively per pool and provider, closes feedback into scoped suppression in under a minute, authenticates on tenant domains with gated rotation, and defines deliverability honestly. Names who pays — tenants with bad lists absorb their own damage; good tenants pay nothing; bulk mail waits so transactional doesn't. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 "Delivered" Is the Most Misleading Word in Email#

ClaimReality
"99.9% delivered"99.9% accepted by an MX; placement unknown
"The provider got it"The provider took responsibility — and may file it as spam
"No bounces, so it worked"Spam-foldered mail doesn't bounce

The Staff position: Report acceptance as acceptance. Measure placement with seeds, provider data and engagement, and never conflate them.

Why this matters in interviews: Defining the SLO honestly is one of the clearest Staff signals in this question.

10.2 Most Tenants Should Not Have a Dedicated IP#

VolumeDedicated IP Outcome
< 100K/monthCold, suspicious, worse delivery
Irregular burstsCold every time it matters
Steady > ~1M/month to big providersGood — tenant owns their reputation

The Staff position: Dedicated IPs are for steady, high-volume senders. Everyone else is better off on a well-policed shared pool.

Why this matters in interviews: "Give every customer their own IP" sounds like isolation and is actually worse delivery.

10.3 Your Rate Limit Lives in Someone Else's Data Center#

LeverWho Controls It
MTA capacityYou
Acceptance per IP per providerThe provider
Reputation that sets acceptanceYour tenants' behavior, scored by the provider

The Staff position: Design the throttle around receiver signals. Hardware scaling is the easy 5%.

Why this matters in interviews: It's the most common wrong answer to "how do you handle peak?"

10.4 Open Tracking Is a Rough Signal, Not a Metric#

Source of "Opens"Human?
Privacy-driven image prefetchNo
Security scannersNo
A person reading the messageYes

The Staff position: Treat opens as approximate, filter machine signals and never let inflated engagement keep dead addresses on a list.

Why this matters in interviews: It shows you understand that analytics feeding suppression is a deliverability concern.

10.5 The Abuse Team Is Part of the Architecture#

The Staff position: On a multi-tenant email platform, a compromised account or a purchased list is a weekly event, not an edge case. Automated pausing, scoped credentials, paced first sends and a trust-and-safety team with authority are system components — as essential as the MTAs.

Why this matters in interviews: Candidates who include enforcement in the architecture diagram are designing the system that actually runs.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer builds an isolated, adaptive, well-authenticated email platform. The Principal engineer notices that the company's root domain is used by marketing's campaign tool, billing's invoice system, the product's transactional mail, a recruiting vendor and a survey tool — and that a complaint spike from one of them lowers the domain reputation the login emails depend on. The L7 problem is governance of a shared reputation asset: which domains and subdomains exist, who may send as them, under what authentication, and who has authority to stop a sender. The MTAs are the easy part.

The Org-Level Fault Line#

One email platform and domain policy vs every team picking its own sender.

OptionWhat WorksWhat BreaksWho Pays
Each team picks a vendor and sends as the root domainFast; teams choose toolsShared domain reputation with no owner; DMARC enforcement impossible; inconsistent suppressionLogin and receipt mail (reputation), security (spoofing)
One central platform for all sendingOne policy, one suppression list, one set of keysPlatform must serve marketing tools, vendors and product; becomes a bottleneckPlatform team (scope)
Domain policy + approved senders on delegated subdomainsTeams keep their tools; each subdomain has its own reputation and keys; root domain at p=rejectRequires a registry of senders and procurement reviewPlatform (registry), teams (onboarding)

🧭 Principal Move: "I don't need every email to flow through one platform. I need every sender registered, authenticated on a delegated subdomain with its own keys, and covered by suppression we can audit. The root domain goes to p=reject, and a new vendor can't send as us until procurement and the email platform sign off."

Cost Model#

Assumptions: managed-service pricing on the order of $0.10 per thousand messages; cloud-hosted MTAs; columnar event store; object storage for bodies; fully loaded engineer ~$250K/year. Illustrative, not vendor quotes.

ScaleVolumeInfra ($/month)HeadcountOn-call LoadNotes
Startup2M messages/month, first-party~$200–500 on a managed service0.25 engShared rotationBuy; own domains, DKIM, suppression
Growth300M/month, first-party + marketing~$30K managed, or ~$15–25K self-run + IPs2–3 eng incl. deliverabilityLight; incidents are reputation, not outagesSelf-running rarely pays here
Platform30B/month, multi-tenant (1B/day)~$300–600K (event store, bodies, MTAs, ~1,500 IPs, tracking edge)25–40 (infra, deliverability, trust-and-safety, support tooling)Dedicated infra rotation + T&S rotationEvent storage and people dominate

The pricing insight: compute is the smallest line. At platform scale, the deliverability and trust-and-safety teams and the event store cost more than the MTAs and IPs combined, and a single multi-day pool blocklisting can cost more in credits and churn than a month of infrastructure.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Which domain tenants' mail is signed with (theirs vs ours)One-wayReputation accrues to the signing domain; switching resets it
Root domain DMARC at p=rejectOne-way-ishEasy to relax in DNS, but relaxing re-opens spoofing and signals weakness
Suppression scope semantics promised to tenantsOne-wayTenants build compliance processes on it
IP ranges and pool layoutTwo-way, slowNew IPs need weeks of warm-up
Throttle algorithmTwo-wayInternal
Event store engineTwo-wayInternal migration
Tracking defaultsTwo-way (with notice)Tenants' analytics change

The Standard I'd Write#

RFC-MAIL-001: Outbound Email Standard
Status: Approved   Owners: Email Platform + Security + Legal

Scope
  Every system, team and vendor sending email as any company-owned domain.

MUST
  1. Register as a sender in the email sender registry before sending.
  2. Send from a delegated subdomain with its own DKIM keys (2048-bit),
     SPF or custom return-path, and DMARC alignment.
  3. Separate transactional and marketing mail by subdomain and stream.
  4. Honor the shared suppression list for complaints and unsubscribes;
     process one-click unsubscribe within 24 hours.
  5. Keep complaint rate below 0.1% per sender; senders above it are paused
     pending review by the email platform.

SHOULD
  1. Sunset recipients with no engagement in 12 months.
  2. Use branded tracking domains; no shared tracking domains in transactional mail.
  3. Pace first sends to new lists (1–5% sample, then hold).

Exceptions
  Filed with Email Platform; Security review for any root-domain sending;
  time-boxed to one quarter.

Success metrics
  - Root domain DMARC policy: p=reject, 100%
  - Unregistered senders seen in DMARC reports: 0
  - Transactional time-to-MX p99: < 60s
  - Pool and sender complaint rates: < 0.1%

What I'd Tell the VP#

"Our login and receipt emails share a reputation with every marketing tool and vendor that sends as our domain, and right now nobody owns that reputation. When marketing has a bad week, customers can't log in. I'm proposing a sender registry, separate subdomains for transactional and marketing mail, and full anti-spoofing enforcement on our main domain within two quarters. It costs about two engineers for six months, mostly coordinating with teams and vendors, and it protects the emails that directly drive revenue. The risk is a vendor we don't know about breaking during enforcement; we'll find them first with reporting before we flip the switch."

Principal Interview Signals#

SignalWhat It Sounds Like
Governs the shared asset"Who may send as our domain is a registry and a policy, not a DNS edit."
Prices people, not servers"Deliverability and abuse teams cost more than the MTAs."
Identifies one-way doors"The signing domain decides whose reputation accrues — that's hard to undo."
Sets enforcement authority"Trust-and-safety can pause any sender without engineering sign-off."
Knows when not to centralize"Teams keep their tools on delegated subdomains; we don't force every email through one pipe."

Staff answers that L7 interviewers find insufficient:

  • "We'll separate transactional and bulk streams" — correct, but ignores the vendors sending as the root domain.
  • "We pause tenants above 0.1% complaints" — good, but no named owner with authority or audit trail.
  • "We'll add an EU region" — no mention of the 6–8 weeks of IP warm-up that set the launch date.

Appendices

Appendix A: Mechanics in Depth#

A.1 The Scheduler and Throttle Loop#

state per key k = (pool, provider):
  rate[k]            : target messages/sec (AIMD)
  conns[k]           : max concurrent SMTP connections per IP
  window[k]          : counts of accepted, deferred, rejected in last 30s
  fair[k][tenant]    : deficit counter for weighted fair share
  ip_cap[ip][provider][day] : warm-up ceiling

every 100ms, for each key k:
  budget = rate[k] * 0.1
  for tenant in round_robin(tenants with ready work in k):
      fair[k][tenant] += weight[tenant] * quantum
      while budget > 0 and fair[k][tenant] > 0 and ready(k, tenant):
          ip = pick_ip(k)  # least-loaded IP under its warm-up cap and conns[k]
          if ip is None: break
          dispatch(next(k, tenant), ip)
          budget -= 1; fair[k][tenant] -= 1

every 30s, for each key k:
  d = window[k].deferred / max(1, window[k].total)
  if window[k].policy_blocks > threshold: state[k] = BLOCKED; rate[k] = 0
  elif d > 0.03:  rate[k] *= 0.5; conns[k] = max(1, conns[k] // 2)
  elif d < 0.01:  rate[k] *= 1.05
  # retries are ordinary ready work: they spend the same budget

A.2 Retry Schedule for Deferrals#

stream          first retry   backoff          give up
transactional   1 min         x2, cap 15 min   4 hours  (a reset older than that is useless)
bulk            15 min        x2, cap 2 h      24-72 hours (tenant-configurable)
RFC 5321 guidance: retry interval >= 30 min; give-up >= 4-5 days (general mail)

A.3 Bounce Classification#

SMTP reply                          class            action
5.1.1 / 5.1.10 user unknown         hard bounce      suppress tenant, all streams; global hint
5.2.2 mailbox full (5xx)            soft-ish         count; suppress after 5 in 7 days
4.x.x temporary                     deferral         retry; throttle signal
5.7.x policy / reputation / DMARC   policy block     no address suppression; pool alert
async DSN via VERP return-path      per DSN code     same table, matched by encoded IDs
ARF complaint via feedback loop     complaint        suppress tenant bulk; count toward rates

Appendix B: Data Model#

CREATE TABLE sending_domains (
  domain             TEXT PRIMARY KEY,
  tenant_id          TEXT NOT NULL,
  dkim_selector_active TEXT NOT NULL,
  dkim_selector_next   TEXT,
  next_published_verified_at TIMESTAMPTZ,      -- gate for rotation
  return_path_domain TEXT NOT NULL,            -- e.g. bounce.mail.tenant.example
  dmarc_policy       TEXT,                     -- observed: none | quarantine | reject
  verified_at        TIMESTAMPTZ
);

CREATE TABLE suppressions (
  tenant_id     TEXT NOT NULL,
  address_hash  BYTEA NOT NULL,                -- normalized, hashed
  scope         TEXT NOT NULL,                 -- all | bulk | list:<id>
  reason        TEXT NOT NULL,                 -- hard_bounce | complaint | unsubscribe | soft_bounce | manual
  source_msg_id TEXT,
  created_at    TIMESTAMPTZ NOT NULL,
  expires_at    TIMESTAMPTZ,                   -- soft bounces only
  PRIMARY KEY (tenant_id, address_hash, scope)
);

CREATE TABLE tenant_reputation_daily (
  tenant_id  TEXT, day DATE, provider TEXT,
  sent BIGINT, hard_bounces BIGINT, complaints BIGINT, unsubscribes BIGINT,
  PRIMARY KEY (tenant_id, day, provider)
);
-- recipient delivery state lives in the queue store, keyed by (pool, provider, due_at);
-- events go to a columnar store partitioned by day, sorted by (tenant_id, message_id).

Appendix C: Coordination Mechanisms#

C.1 Authentication Checks at the Receiver#

Diagram: C.1 Authentication Checks at the Receiver

C.2 Quick Comparison#

MechanismGuaranteesFailure ModeUse For
Stream separation (IPs + queues)Bulk can't delay or taint transactionalShared hidden dependenciesEvery platform
Reputation pools by track recordBad senders share with nobody goodManual overridesMulti-tenant
AIMD throttle per (pool, provider)Rate tracks receiver acceptanceRetries bypassing itAll sending
Warm-up caps per IPNew IPs build reputationSkipped under pressureEvery new IP
Scoped suppressionNo repeat sends to bad addressesMigration regressionsEvery recipient check
VERP return-pathAsync bounces map to one deliveryMissing encodingBounce processing
CNAME-delegated DKIMRotation without tenant DNS editsUnpublished selectorTenant domains
One-click POST unsubscribeScanners can't unsubscribe usersGET-only linksBulk mail

Appendix D: API Contract & Client Behavior#

  • POST /messages is idempotent on idempotency_key for 24 hours; a retry with the same key returns the original message_id.
  • Responses list suppressed recipients explicitly; suppressed recipients are never silently dropped.
  • stream is required; transactional messages with list-like characteristics (many recipients, unsubscribe footers) are reclassified and the tenant is warned.
  • Events are delivered via signed webhooks (see Webhook Delivery Platform) with delivered meaning accepted by the recipient's provider.
  • Complaint suppressions cannot be removed via API; hard-bounce and soft-bounce suppressions can.
  • Tenants must verify sending domains (DKIM CNAMEs, return-path) before sending bulk.

Appendix E: Observability#

Core metrics:

  • transactional.time_to_mx_accept_p99{provider}, queue.depth{stream, pool, provider}
  • smtp.deferral_rate{pool, provider}, smtp.policy_block_rate{pool, provider}, smtp.attempts_per_accepted
  • pool.complaint_rate{provider}, tenant.complaint_contribution{pool}, tenant.hard_bounce_rate
  • feedback.lag_seconds, suppression.hit_rate{reason}
  • seed.inbox_rate{provider, pool}, blocklist.listed_ips, blocklist.listed_domains
  • dkim.selector_published{domain}, bounce.dmarc_reject_rate{domain}

Critical alerts:

AlertThresholdSeverity
Transactional time-to-MX p99 (any top provider)> 120s for 5 minPage (infra)
Feedback processing lag> 5 minPage (infra)
Listed IP or domain in shared poolanyPage (deliverability)
Tenant complaint rate during send> 0.1%Auto-pause + T&S ticket
Pool spam rate trend+50% over 48hDeliverability ticket
Suppression hit rate drop> 20% day-over-dayPage (infra)

Debugging the silent failure: delivered rates, bounce rates and queue depths can all look perfect while placement collapses. Watch seed placement by provider, engagement trends per tenant and pool, and blocklist status of every shared surface — IPs, return-path domains and tracking domains.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 10M/monthManaged service; your own domains, DKIM, suppressionMarketing hurting transactional on the same domain
10M–1B/monthSeparate subdomains and streams; managed or small MTA fleetCost vs building; tenant abuse if multi-tenant
1B–30B/monthOwn MTAs; reputation pools; AIMD throttles; T&S teamPool contagion; regional warm-up; event storage cost
> 30B/monthRegional pools, sub-pool blast-radius sizing, deliverability consoleGovernance across thousands of tenants and providers

What you don't build on day one: your own MTAs, dedicated IPs, multi-region pools, BIMI, engagement-based send-time optimization. Each has a trigger in Section 11.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Reputation pools by track record: quarantine → trusted → dedicated; down instantly, up only after a clean window. See Multi-Tenancy.
  • Weighted fair share inside each (pool, provider) key so one campaign can't take the pool's whole Gmail budget; see Rate Limiter for the token-bucket mechanics.
  • Paced first sends for new tenants and new lists: 1–5% sample, measure, then release.
  • Cost attribution: messages, event rows and dedicated IPs metered per tenant; complaint contribution reported alongside spend.
  • Enforcement as a feature: automatic pauses with clear reasons and a self-serve remediation flow reduce both damage and support load.
  1. Loading the index…