Hiring BarSupport

Design Login and Sessions at Scale

Case study93 min read12 diagrams

Technologies referenced in this case study: Redis · DynamoDB · PostgreSQL · Apache Kafka · Envoy, Kong & NGINX · Kubernetes

Related: Edge Gateway · Rate Limiting · Distributed Cache · Multi-Region Active-Active · Graceful Degradation: Fail Open or Closed · Build vs Buy Framework · Replication

Reading Guide#

Organized for interview use first, reference second. This page covers issuing, storing, revoking and recovering identity. Verifying a token at the edge — JWKS caching, local signature checks, stripping client-supplied identity headers — lives in API Gateway; this page links there rather than repeating it.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on token formats, key rotation and recovery
What is an Identity System? — Why interviewers pick this topic

An identity system answers two questions for every other system in the company: who is this (authentication) and is this still true (session validity). It registers accounts, verifies credentials (passwords, passkeys, one-time codes, federated logins), issues tokens that other services trust, keeps track of which sessions exist, kills them on demand, and lets people who lost every credential get back in without letting an attacker do the same.

The hard part is not hashing a password or signing a JWT. The hard part is that every token you issue is a bearer credential that will be copied, logged, cached and stolen, and the company will one day ask: "How fast can we make every copy of that credential worthless — and what breaks for the other 200 million users while we do it?"

Before vs After — the "stolen refresh token" scenario:

Without revocation design:
t=0:        Malware on a user's laptop copies the browser's 90-day refresh token.
t=+2h:      Attacker replays it from another country. New access tokens issued.
t=+3h:      User notices odd purchases. Clicks "Log out of all devices".
t=+3h:      Logout deletes the user's sessions in one region's cache only.
t=+3h:      Attacker's 1-hour access JWT keeps working — nobody checks it against anything.
t=+4h:      Attacker refreshes again: refresh token is a stateless JWT too. Still valid.
t=+90 days: Refresh token finally expires. Support has 14 tickets. Trust gone.

With server-side sessions + rotation + revocation fan-out:
t=0:        Malware copies refresh token RT1 (family F9, bound to device key).
t=+2h:      Attacker presents RT1 without the device's proof-of-possession → rejected.
            (Weaker variant: no binding — attacker refreshes, gets RT2.)
t=+2h05m:   Real client presents RT1 (already rotated) → reuse detected → family F9 revoked.
t=+2h05m:   session.revoked event published. Gateways' deny list updated in ≤ 30s.
t=+2h06m:   Attacker's access token (10-min TTL) is denied by subject+session deny list.
t=+2h06m:   User forced to re-authenticate with a passkey. Security email sent.

Why interviewers reach for this question: Identity is the purest test of blast-radius thinking. It is the one dependency every request has, so a candidate's choices decide whether a stolen credential is a 10-minute problem or a 90-day problem, and whether an identity outage is a degraded login page or a company-wide outage. The candidate who says "use JWTs, they're stateless" without saying how they would kill one has never been paged for an account takeover.

Mechanics Refresher: The Identity Primitives
PrimitiveHow It WorksProsCons
Password + slow hashStore argon2id(password, salt); verify by recomputingUniversal; no hardware neededPhishable, reused across sites, 50–250ms CPU per verify by design
Passkeys / WebAuthnDevice holds a private key scoped to your domain; server stores the public key and verifies a signed challengePhishing-resistant; nothing reusable on the serverDevice-loss and recovery story is now the weak point
OTP (SMS, email, TOTP)Server sends or shares a short code; user echoes itCheap second factorSMS: SIM swap, cost (~$0.01–0.05 per message), interception; all OTP is phishable
Server-side sessionOpaque random ID → row in a session storeInstant revocation, small cookieLookup on every request (or on every refresh), store must be highly available
Signed token (JWT)Claims + signature; anyone with the public key verifies locallyNo lookup; ~50–500µs verifyCan't be un-issued; revocation lag = remaining lifetime
Access + refresh token pairShort-lived access JWT (5–15 min) + long-lived refresh token checked server-sideLookups only at refresh timeTwo token types to secure; refresh endpoint is the hot path
Refresh token rotationEach refresh returns a new refresh token; reuse of an old one revokes the familyDetects theft of a refresh tokenRaces on flaky mobile networks cause false reuse
Sender-constrained tokensToken bound to a key the client must prove it holds (DPoP, mTLS)Stolen token alone is uselessClient complexity; key storage on device
Federation (OIDC, SAML)Another IdP authenticates; you trust its signed assertionEnterprises keep control of their usersYou inherit their outages and their misconfigurations
Workload identity (SPIFFE, cloud IAM)Platform attests a workload and issues short-lived certs or tokensNo long-lived secrets in configPlatform becomes a tier-0 dependency

For most production systems: passwords hashed with argon2id plus passkeys, short-lived signed access tokens (10 min), server-side refresh sessions with rotation and reuse detection, a revocation event stream that reaches every verifier in under a minute, and federation for enterprise customers. The primitives are not the interview — revocation speed and the identity service's blast radius are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

Authentication is not a crypto question. Nobody is asking you to implement ECDSA.

It is a revocation and blast-radius question that tests:

  • Whether you can say how long a stolen credential stays useful, as a number, and what you'd do to shrink it
  • Whether you know that every token you issue is a cache of a decision, and caches go stale
  • Whether you design the identity service's failure — what still works when it is down — before its happy path
  • Whether you treat account recovery as the real front door, because attackers do

The key insight: Every authentication design is a choice of where the "is this still valid?" check happens and how often. Check on every request and the identity store is in every request's critical path. Check never and a stolen token lives until it expires. Staff candidates pick the check frequency deliberately — per request, per refresh, per risk event — and state the revocation lag it buys.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws login form → auth service → users table → JWTAsks "Consumer, workforce or service-to-service? What's the account-takeover threat, and how fast must 'log out everywhere' take effect?"Asks "How many login systems does the company already have, who owns account recovery policy, and which regulator or customer contract sets our audit requirements?"
Tokens"JWTs, so it's stateless and scales""10-minute access JWTs verified locally, server-side refresh sessions with rotation; revocation lag is bounded at 10 minutes, ≤ 30s for high-risk events via a pushed deny list"Sets token lifetime and binding as an org-wide standard by risk tier; owns the deprecation of the 3 legacy token formats still in circulation
Failure"Run several replicas of the auth service""If the IdP is down, verification keeps working on cached keys; new logins fail; refresh gets a bounded grace extension; step-up actions fail closed"Treats identity as tier 0: cells per region, a written fail-static policy signed by security, quarterly IdP-down game days across every product team
Revocation"Delete the session from the DB""Revocation is an event: session store write, then fan-out to every verifier's deny list, with revocation.propagation_p99 as an SLO"Makes revocation latency a company security KPI and pushes it across the vendor boundary with shared-signals standards
Recovery"Email a reset link""Recovery is the weakest login path, so it gets risk scoring, cooling-off periods, notifications to existing factors, and its own abuse metrics"Decides who owns recovery policy — security sets the floor, product sets friction, support follows a script it can't override — and audits support-assisted recovery
Scale"Shard the users table""Login is ~700/s at peak; refresh is ~25K/s; verification is ~1M/s but local. The hard capacity problem is credential-stuffing waves at 50K/s hitting a 100ms password hash"Prices it: hashing capacity, SMS spend, IdP vendor per-MAU fees, and the headcount of a 24×7 identity on-call versus buying
Why "tokens" separates levels

L5: "We'll issue a JWT on login with a 24-hour expiry. Every service verifies it locally, so there's no central bottleneck." This is a reasonable performance answer. It is also a 24-hour revocation lag that the candidate hasn't noticed. When asked "the user clicks 'log out everywhere'," the answer becomes "we'd add a blacklist," and the blacklist is now a per-request lookup — the bottleneck the JWT was supposed to avoid.

L6: "Access tokens are 10-minute JWTs, verified locally at the gateway as in the API gateway design. The refresh token is opaque and lives in a server-side session store, so refresh — about once every 10 minutes per active client — is where we check 'is this session still alive'. That bounds normal revocation lag at 10 minutes. For high-risk events — password change, takeover detection, admin kill — we push a deny entry keyed by session ID to every verifier within 30 seconds, and it expires after 10 minutes because by then the token has expired anyway. The deny list stays small: it only holds the last 10 minutes of revocations."

L7: "The real problem is that we have four token formats across three generations of products, two of which issue 30-day tokens nobody can revoke. I'd publish a token standard with lifetimes by risk tier, migrate the long-lived ones behind a refresh flow within two quarters, and measure 'maximum credential lifetime in circulation' as a security KPI."

Why "failure" separates levels

L5: "We'll run the auth service in three availability zones with a replicated database." True, and insufficient: the IdP's database is a single logical dependency, a bad deploy or a bad signing-key push hits all three zones at once, and the candidate hasn't said what every other service does when the IdP is unreachable.

L6: "I'll separate what the IdP is needed for. Verifying an existing access token: never needs the IdP — keys are cached. Refreshing a session: needs the session store, which I'll make regional and replicated; if it's unreachable, I extend existing sessions by up to 30 minutes rather than log out 60 million people at once. New logins: fail closed with a clear error, because I can't verify a password I can't read. Sensitive actions — password change, payout, admin — always fail closed."

L7: Recognizes that the IdP is the company's largest correlated-failure domain. "Every product's availability is capped by ours. I'd run identity as cells per region with no cross-region synchronous dependency, publish the fail-static policy so product teams build to it, and run an IdP-down game day every quarter where we turn off login in one region and watch what breaks."

Why "recovery" separates levels

L5: "Forgot password sends a reset link to the email on file, valid for 1 hour." Correct and incomplete. It ignores that the email account may itself be compromised, that support agents can be talked into changing the email, and that the reset link is a bearer credential sitting in an inbox.

L6: "Recovery is the login path with the weakest authentication, so it gets the strongest monitoring. Every recovery notifies all existing factors, high-value accounts get a 24–72 hour cooling-off before a new factor becomes trusted, and support-assisted recovery follows a script with no override. I track recovery.success_then_dispute_rate — the fraction of recoveries later reported as takeovers."

L7: "Recovery policy is a business decision dressed as an engineering one. Security wants a 72-hour delay; growth wants zero friction; support wants a button. I'd make the policy explicit by account tier, give security the veto on the floor, and audit every support-assisted recovery monthly."

Positions to Commit To#

PositionRationale
Short-lived access token + server-side refresh sessionBounds revocation lag at the access-token lifetime (10 min) while keeping the per-request path lookup-free
Revocation is an event with an SLO, not a DELETEA row deleted in one store does nothing to tokens already cached at 400 gateway nodes; propagation must be pushed and measured
Rotate refresh tokens; treat reuse as theftThe only cheap way to detect a copied refresh token; tolerate a 30–60s grace for mobile retries
Verification never depends on the IdP being upCached signing keys mean an IdP outage degrades login, not every API call
Shed credential stuffing before the password hashargon2id at ~100ms means 50K bad attempts/s would need ~5,000 cores; reject by IP, ASN, device and breached-credential signals first
Recovery gets the strongest controls, not the weakestAttackers go through the weakest door; recovery is usually it
Signing keys rotate on a schedule you have rehearsedThe emergency rotation will happen under pressure; if the routine one isn't automated, the emergency one will be an outage

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Consumer login10M–1B accounts, account takeover, low-friction recovery, morning-ramp and attack spikesPasswords + passkeys, short access tokens, server-side refresh sessions, risk engine, layered stuffing defensesCredential stuffing, recovery takeover, IdP outage logs everyone outTakeover rate per million logins, revocation lag ≤ 10 min, login availability 99.95%
Workforce SSO1K–200K employees, very high-value accounts, compliance audits, many SaaS appsOIDC/SAML federation, phishing-resistant MFA mandatory, device posture, SCIM provisioning, short sessions for adminPhished admin session, deprovisioning lag after terminationTerminated employee loses all access ≤ 15 min; every admin action attributable
Service-to-service identity10K–1M workloads, millions of calls/s, no humans in the loopPlatform-attested workload identity, mTLS with certs living hours, no static secretsCA or attestation outage halts deploys; expired certs cause mass failureZero long-lived secrets in config; cert lifetime ≤ 24h; rotation fully automated

🎯 Staff Move: "I'll design consumer login — that's where the scale and the account-takeover pressure are. I'll keep the token and session model general enough that workforce SSO is a federation adapter on top, and I'll treat service-to-service identity as a separate platform problem that shares the signing-key discipline but not the session store. Tell me if you'd rather go deep on workforce or service identity."

Where the Design Splits#

#Fault LineThe Tension
1Stateless Tokens vs Server-Side SessionsRevocation speed vs a lookup on every request; who carries the latency and who carries the stolen-token risk?
2Token Lifetime: Short + Refresh vs Long-LivedEvery minute of lifetime is a minute a stolen token works; every refresh is load on the hottest endpoint you own
3Fail Closed vs Fail Open When the IdP Is UnreachableLock out every user, or let possibly-revoked sessions keep working for a bounded time? Who signs off?
4Central Authorization Service vs Policy in Each ServiceOne consistent answer to "can X do Y?" vs a new tier-0 dependency on every request
5Build vs Buy the Identity ProviderControl over the most security-critical code you own vs a vendor whose outage and breach are now yours

How Real Companies Built It#

Why this section belongs here: Identity is the area where public incident reports teach the most. Each of these shows a fault line from this page playing out in production.

Facebook — "View As" and Access Tokens at Scale (2018)#

In September 2018 Facebook disclosed that three bugs interacting in its "View As" feature let attackers obtain access tokens — which Facebook describes as the digital keys that keep people logged in — for other users; it reset tokens for almost 50 million affected accounts plus another 40 million that had been subject to a "View As" lookup in the previous year, and turned the feature off. Its follow-up said about 30 million people actually had tokens stolen (Facebook security update, Facebook follow-up).

The mitigation was mass token reset — logging ~90 million accounts out — because that was the revocation primitive available at that scale. The bug was in a product feature, not the identity service, yet the identity system carried the response.

Staff insight: When asked "how would you respond to mass token theft?", the answer is a pre-built, rate-limited, targeted bulk revocation tool: revoke sessions by cohort (feature, time window, client version), not "everyone". Facebook reset an extra 40 million accounts as a precaution because the blast radius was hard to bound. Design so you can bound it.

Okta — Support-System Breach and Session Tokens in HAR Files (2023)#

Okta's root-cause report says that between 28 September and 17 October 2023 a threat actor accessed files in its customer support system belonging to 134 customers; some were HAR files containing session tokens, which were used to hijack sessions of five customers. Access came through a service account whose credentials had been saved to an employee's personal Google account, signed into a personal profile on an Okta-managed laptop. Remediations included disabling that account, blocking personal Google profiles on managed devices, and binding admin session tokens to network location (Okta root cause and remediation).

Staff insight: A session token is a bearer credential, and it will leak through channels nobody designed for — debug files, logs, support tickets. Two defenses follow: bind high-value sessions to something the attacker doesn't have (device key, network context), and keep a revocation path that support and security can trigger in seconds. In an interview, say "bearer tokens leak sideways; binding and fast revocation are how we make a leak boring."

Microsoft — Storm-0558 and a Consumer Signing Key (2023)#

Microsoft's investigation reported that Storm-0558 used an acquired Microsoft account (MSA) consumer signing key to forge tokens for Outlook Web Access and Outlook.com. Its leading hypothesis was that a 2021 crash of the consumer signing system produced a crash dump that, through a race condition, contained the key, and that the dump was later reached through a compromised engineering account; enterprise mail accepted the consumer-signed tokens because validation libraries did not automatically enforce key scope (Microsoft Security Response Center).

Staff insight: Two lessons for this page. First, signing keys are the crown jewels — keep them in HSMs or KMS and design emergency rotation as a rehearsed drill. Second, verification must check that a key is allowed to sign this kind of token for this audience, not only that the signature is mathematically valid. Say both in the key-management deep dive.

Google — BeyondCorp and Zanzibar#

Google's BeyondCorp paper describes moving access decisions off the network perimeter: corporate applications are reached over the internet, and access depends on the user and the device rather than on being inside a privileged intranet (Google research). Separately, Google's Zanzibar paper describes a central authorization system storing trillions of access control lists and serving millions of authorization requests per second at a 95th-percentile latency under 10 ms and availability above 99.999% over three years of production use (Google research).

Staff insight: These are the two ends of fault line 4. BeyondCorp says the network is not an identity; every request carries user and device identity. Zanzibar says a central authorization service can be fast and available enough — if you are willing to build what Google built. In an interview, cite Zanzibar to show central authorization is possible, then say what it costs before recommending it.

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"We'll use JWTs""User clicks 'log out of all devices'. How long until a stolen token stops working?"Do you know your revocation lag as a number?
"We store sessions in Redis""Redis is down. What happens to 60 million logged-in users?"Session store as a tier-0 dependency, degraded mode
"We hash passwords with bcrypt""40,000 login attempts per second from 200,000 IPs just started. What happens to your CPU?"Credential stuffing economics, shedding before the hash
"We rotate refresh tokens""A mobile client on a bad network retries the refresh. Did you just log them out?"Reuse-detection races, grace windows
"Forgot password emails a link""The attacker already owns the user's email. Now what?"Recovery as the real front door
"We rotate signing keys""The private key was leaked an hour ago. Walk me through the next 30 minutes."Emergency rotation, verifier key caches, mass re-login

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Three paths with three very different rates. Verification (~1M/s) happens at the gateway with cached keys and never touches the identity plane. Refresh (~25K/s at peak) is the hottest identity endpoint and the place where "is this session still alive?" is answered. Login (~700/s legitimate, up to 50K/s during an attack) is the most CPU-expensive path and the one under attack. Revocation flows the other way — from the session store, through the event bus, out to every verifier — and its latency is the number this whole design is judged by.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Token model"JWT, stateless""10-min access JWT, verified locally. Opaque refresh token checked server-side. Revocation lag ≤ 10 min, ≤ 30s via deny list."
Logout everywhere"Delete sessions""Revoke the refresh family, publish session.revoked, push a session-id deny entry with a 10-min TTL to every verifier."
IdP down"Replicas""Verification continues on cached keys. New logins fail closed. Refresh gets a bounded grace. Step-up fails closed."
Stuffing"Rate limit by IP""Layered: IP/ASN reputation, device signals, breached-password check, per-account limits — all before argon2id."
Recovery"Email a reset link""Risk-scored, notifies all factors, cooling-off for high-value accounts, support follows a no-override script."
Keys"Rotate yearly""Rotate every 30 days, automated, publish-before-sign; emergency rotation drilled quarterly."
Service identity"API keys in env vars""Platform-attested workload identity, mTLS certs living ≤ 24h, no static secrets."

Numbers to Bring#

MetricValueWhy It Matters
argon2id / bcrypt verify cost~50–250ms of CPU per attempt (tuned)1,000 logins/s ≈ 100 cores; a 50K/s stuffing wave ≈ 5,000 cores — shed first
Local JWT verification~50–500µs with a cached keyWhy verification never needs the IdP (see Edge Gateway)
Session store lookup~1–2ms p99 in-region (Redis or DynamoDB)Affordable at refresh rate (25K/s), expensive at request rate (1M/s)
Access-token lifetime5–15 min typical; 10 min hereEquals worst-case revocation lag without a deny list
Refresh session lifetime30-day sliding, 90-day absolute (consumer)How long a device stays signed in without a password
Deny-list propagation targetp99 ≤ 30sShrinks the high-risk revocation window from 10 min to seconds
JWKS cache TTL at verifiers5–15 minKey rotation must publish the new key ≥ 1 TTL before signing with it
NIST SP 800-63B rate limit≤ 100 consecutive failed attempts per account before the authenticator is disabled (NIST SP 800-63B)The ceiling, not the target; per-account limits are one layer of many
NIST SP 800-63B reauthentication at AAL2Overall timeout SHOULD be ≤ 24h, inactivity ≤ 1h (NIST SP 800-63B)The workforce baseline; consumer sessions are usually longer by product choice
Workload certificate lifetime1–24 hoursShort enough that revocation lists are rarely needed
SMS OTP cost~$0.01–0.05 per message, higher in some countriesSMS pumping fraud turns your OTP endpoint into a cost attack
Password reset link lifetime15–60 min, single useA reset link is a bearer credential sitting in an inbox

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the signup form, password hashing and the JWT payload, then run out of time before the interviewer asks the only question that matters: "A token was stolen. How long does it keep working, and what does it cost everyone else to stop it?" Compress the basics to ~10 minutes and spend the rest on revocation, the IdP's failure posture, credential stuffing and recovery.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Users sign up, log in with a password or a passkey, optionally add a second factor, stay logged in on several devices, log out of one or all devices, and recover their account when they lose credentials. Every other service in the company trusts the identity we issue."

Then the non-functional requirements, which is where the design lives:

"Four constraints drive everything. One: a stolen credential must stop working fast — I'll target 10 minutes worst case and 30 seconds for high-risk revocations. Two: identity is in every request's path, so an identity outage must not become a company outage — verification can't depend on the identity service being up. Three: login is under constant attack, so the expensive part — password hashing — must be protected from credential-stuffing volume. Four: recovery is a login path too, and usually the weakest one. Scale: I'll assume 200 million accounts, 60 million daily actives, ~700 logins per second at the morning peak, and around a million authenticated API requests per second."

Then name the underspecified parts:

"A few things I'd confirm: is this consumer, workforce or service-to-service? Do we have enterprise customers who need SSO into our product? Are there regulated accounts — payments, health — that need stronger assurance? I'll assume consumer, with enterprise SSO as a later adapter."

🎯 Staff Move: Putting a number on revocation lag in the first three minutes — "10 minutes worst case, 30 seconds for high-risk" — tells the interviewer you know the token model is a revocation decision, not a performance decision. Everything you draw next can be judged against that number.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Account: account_id, status (active, locked, recovering, deleted), assurance_tier, created_at
  • Credential: credential_id, account_id, type (password, passkey, totp, phone, email, federated), secret_ref (hash or public key), added_at, trusted_after (cooling-off), last_used_at
  • Session (refresh family): session_id, account_id, device_id, family_id, current_refresh_hash, prev_refresh_hash, rotated_at, auth_methods (pwd, passkey, otp), auth_time, expires_at, absolute_expires_at, revoked_at, binding_key_thumbprint
  • SigningKey: kid, alg, status (pending, active, retiring, revoked), published_at, activated_at, retire_after
  • AuditEvent: event_id, account_id, type, actor (user, support agent, system), ip, device, at

Client-facing API:

POST /v1/login                  { identifier, password | passkey_assertion, device_info }
  → 200 { access_token (10 min), refresh_token, expires_in }
  → 401 | 429 | 200 { step_up_required: ["otp","passkey"], challenge_id }

POST /v1/token/refresh          { refresh_token }   DPoP: <proof signed by device key>
  → 200 { access_token, refresh_token (new) }
  → 401 { error: "session_revoked" | "reuse_detected" }

POST /v1/sessions/{id}/revoke                         (log out one device)
POST /v1/sessions/revoke_all                          (log out everywhere, keeps current)
POST /v1/recovery/start         { identifier }
POST /v1/recovery/complete      { recovery_token, new_credential }
GET  /.well-known/jwks.json                           (public keys, cached by verifiers)

Login returns either tokens or a step-up challenge; the risk engine decides which. Refresh is the only endpoint that touches the session store on the hot path.

🎯 Staff Move: "I'm splitting the session from the token. The session is a row I can revoke. The access token is a 10-minute cache of 'this session was valid at issue time'. Once you see the token as a cache entry, its lifetime is just a TTL, and every cache-invalidation tool applies."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the login flow in 90 seconds:

  1. Client submits credentials. The edge drops requests from known-bad IPs, ASNs and device fingerprints and applies per-IP and per-identifier rate limits — before any hashing.
  2. Login service checks the identifier exists and isn't locked, then verifies the password with argon2id (~100ms) or a passkey assertion (~1ms signature verify).
  3. Risk engine scores the attempt: new device, impossible travel, breached-password match, velocity. Outcome: allow, step-up, or block.
  4. Token service creates a session row (refresh family), stores a hash of the refresh token, and signs a 10-minute access JWT with the active key in KMS.
  5. Client calls APIs with the access JWT; the gateway verifies it locally with cached public keys — no identity call.
  6. Every ~10 minutes the client refreshes: the token service looks up the session, checks it isn't revoked, rotates the refresh token, issues a new access token.
  7. Logout or a risk event revokes the session row and publishes session.revoked; the distributor pushes the session ID to every verifier's deny list.

🎯 Staff Move: Say out loud: "Notice the three rates: verification at a million per second never touches identity; refresh at 25,000 per second is where revocation is enforced; login at 700 per second is the expensive one and the one under attack. I'm going to design each path to its own rate." You've now spent ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes identity hard is what happens after a credential leaks and when the identity service fails. I'd like to go deep on four things: how fast we can revoke and how that reaches every verifier, what the rest of the company does when identity is down, how we survive a credential-stuffing wave without melting the password hashers, and how account recovery avoids becoming the takeover path. Where would you like to start?"

If no preference: start with revocation. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Revocation (7–8 min)

"There are three revocation speeds and I'll be explicit about which events get which. Routine logout: revoke the session row — the access token dies at its natural expiry, at most 10 minutes later. High-risk events — password change, 'log out everywhere', takeover detected, admin kill: revoke the row and push the session ID to a deny list held in memory at every gateway node, with a 10-minute TTL, because after 10 minutes the token has expired anyway. Mass events — signing-key compromise: rotate the key, which invalidates every token signed with it."

Quantify: "High-risk revocations are maybe 0.5% of 60 million daily actives — 300,000 per day, 3.5 per second average. With a 10-minute TTL the deny list holds about 2,100 entries in steady state, under 100 KB. During a mass incident that could reach a few million entries — still under 200 MB, which a gateway can hold, but I'd switch to revoking by auth_time cutoff per account instead of listing session IDs."

Who pays: "The user pays up to 10 minutes of exposure for routine logouts. Security signed off on that; for anything they care about, the push path pays 30 seconds. The gateway team pays a small in-memory set and a subscription to the revocation stream."


Deep dive 2: The IdP is down (6–7 min)

"I'll split the dependency by operation. Verification: independent — cached JWKS. Refresh: depends on the session store; if the token service can't reach it, it issues a short extension — a 10-minute access token based on the still-valid refresh token's claims — for up to 30 minutes, but only if the client's refresh token is cryptographically valid and the session isn't on the local deny list. Login: fails closed. Step-up and sensitive actions: fail closed. That's a written policy security signs, not an on-call judgment call."

Quantify: "Without the grace, a 30-minute session-store outage expires every active access token within 10 minutes — 15 to 20 million concurrently active users logged out at once, then all of them retrying login when it comes back: 20 million logins in a few minutes against a 700/s design, or 30,000+ per second of argon2id. The grace prevents the outage from becoming a login stampede."


Deep dive 3: Credential stuffing (6–7 min)

"Stuffing is economics. Attackers try breached email/password pairs at volume. If every attempt reaches argon2id at 100ms, a 50,000/s wave needs 5,000 cores — the attack is a denial of service on login before it's a takeover. So the layers: IP and ASN reputation and per-IP rate limits at the edge; device and client-integrity signals; per-identifier limits — 10 failures per 15 minutes then step-up; a breached-credential check so a correct-but-breached password triggers a reset instead of a login; and the hash budget itself is a bulkhead: login gets a fixed pool, and when it's saturated we serve challenges, not timeouts."

Who pays: "Legitimate users behind shared IPs — mobile carriers, universities — pay extra challenges during waves. Product signs off on a challenge rate ceiling, say 2% of good logins."


Deep dive 4: Recovery (5–6 min)

"Recovery authenticates someone who has, by definition, lost their credentials, so it's the weakest proof we accept. Rules: notify every existing factor on start and completion; for accounts with a passkey or a high-value tier, a cooling-off of 24–72 hours before the new credential can change payout details or remove other factors; reset tokens are single-use, 15 minutes, bound to the browser that started the flow; and support-assisted recovery follows a scripted identity-proofing flow — agents can't skip steps or change contact details directly."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: a token is a cache of an identity decision, so I picked the cache TTL — 10 minutes — and built a push path for the cases where 10 minutes is too long. Verification never depends on the identity service, so its outage degrades login, not the company. Login is protected economically, with shedding before the expensive hash, and recovery is treated as the front door it really is."

The evolution closer:

"What I'd build later: device-bound tokens via DPoP for all first-party clients; passkeys as the default credential, which removes most of the stuffing problem; cross-vendor revocation signals for enterprise customers through the OpenID shared-signals standards; and identity cells per region. What I'd not build: our own HSMs — cloud KMS is the right answer until a regulator says otherwise."

🎯 Staff Move: End on the revocation number and who owns it. Senior candidates end with "and we'd add MFA." Staff candidates end with "and revocation.propagation_p99 is an SLO the identity team carries, reviewed with security every month, because it's the number that decides how bad the next token theft is."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Crypto tourExplains RSA vs ECDSA and JWT header fields for 8 minutesOne sentence: "ES256 or EdDSA, keys in KMS, verified locally"
Signup form designDesigns email verification and username rules in detailNames the entities in 30s and moves to sessions
OAuth flow recitalDraws every redirect of the authorization code flow"Authorization code + PKCE per the OAuth security BCP; I'll skip the redirects unless you want them"
No revocation numberWaits for "how do you log someone out?"States revocation lag as a requirement in Phase 1
No failure postureSays "highly available"Splits verify / refresh / login / step-up and gives each a posture
Recovery as afterthought"Forgot password sends an email" at minute 44Brings recovery into the deep dives as a login path

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Identity is the dependency nobody can route around. A cache outage costs latency; a queue outage costs freshness; an identity outage costs every authenticated request in the company. At the same time it is the system attackers target most, because a single credential unlocks everything that identity guards. That combination — maximal blast radius on failure, maximal value on compromise — forces tradeoffs where both sides are expensive, and the candidate has to say who pays.

It also has an unusually deceptive happy path. A Senior engineer can build a login that works for every honest user on day one. The design is judged entirely by the dishonest ones: the credential stuffer, the malware that copies a refresh token, the caller who convinces a support agent they are the account owner. None of those show up in a demo. Staff candidates design for them before drawing the first box.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"An attacker has a copy of one user's refresh token. Walk me through every place that token, and the access tokens derived from it, are accepted — and how long each keeps accepting it after the user hits 'log out everywhere'."

A candidate who answers with the refresh path (session store, revoked immediately), the access-token path (gateway, up to 10 minutes or ~30s via deny list), the downstream caches (services that cached the identity, bounded by the same TTL), the detection mechanism (rotation reuse, risk signals) and the owner (identity on-call for propagation, trust and safety for the account) has operated an identity system. A candidate who says "we delete the session" has built a login form.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Consumer login. The population is huge, mostly honest, and reuses passwords. The threat model is scale: attackers replay billions of leaked credential pairs, buy stolen session cookies, and socially engineer support. The design centers on cheap verification, aggressive abuse shedding, risk-based step-up, and recovery that is easy for real owners and hard for impostors. Product friction is a real cost here — every extra challenge loses a measurable fraction of logins — so security and growth negotiate every threshold. Correctness bar: account takeovers per million logins, revocation lag, login success rate.

Workforce SSO. The population is small (thousands to a few hundred thousand) but each account is worth far more: one administrator session can reach production, customer data or the identity system itself. The design centers on federation (OIDC and SAML to hundreds of SaaS apps), phishing-resistant MFA as a hard requirement, device posture, short sessions for privileged roles, and deprovisioning: a terminated employee must lose access everywhere within minutes, which means pushing revocation into apps you don't own. Correctness bar: time-to-deprovision, percentage of apps behind SSO, audit completeness.

Service-to-service identity. No humans, no passwords, no recovery flow. Millions of calls per second between workloads that are created and destroyed constantly. The design centers on the platform attesting what a workload is (which cluster, namespace, service account) and issuing short-lived credentials — X.509 certificates for mTLS or signed tokens — that rotate automatically. Standards like SPIFFE define the identity format and how workloads fetch short-lived identity documents (SPIFFE overview). Correctness bar: zero long-lived secrets, rotation success rate, and no deploy blocked by the identity platform. This intent pairs with Service Registry and the workload model in Kubernetes.

🎯 Staff Move: "These share almost nothing operationally. Consumer is about abuse at scale, workforce is about high-value accounts and deprovisioning, service identity is about automation and rotation. If someone asks me to build one system for all three, I'd share the key-management and audit layers and nothing else."

2.2 When NOT to Build Your Own Identity Provider#

SituationWhat to Do InsteadWhy
Fewer than ~1M accounts, no unusual assurance needsBuy a hosted CIAM productYou'd spend 3–5 engineers re-implementing password storage, MFA, recovery and federation that a vendor has hardened
Workforce identity for your own employeesBuy a workforce IdPThe SaaS app catalogue, SCIM connectors and device-posture integrations are the product; you won't out-build them
You need "login with X" onlyUse the social IdP via OIDC; keep a thin local account recordFederation shifts credential risk to a provider with a bigger security team
Service-to-service identity on one cloudUse the cloud's workload identity and IAMThe platform already attests workloads; a parallel CA is a second tier-0 system
Authorization rules that fit in roles and a few attributesPolicy library in each service, not a central authorization serviceA Zanzibar-style service is worth it only when relationships (sharing, folders, orgs) dominate

And within the design, some features you should not build even when you own the IdP:

  • Don't invent a token format. Use JWT/JWS or opaque tokens per the OAuth and OIDC specs. Custom formats lose the libraries, the audits and the interoperability.
  • Don't build the resource owner password credentials grant. The OAuth 2.0 Security Best Current Practice says it MUST NOT be used (RFC 9700).
  • Don't store long-lived secrets for services when the platform can attest them.
  • Don't run your own HSM fleet unless a regulator requires it; cloud KMS gives non-exportable keys with audit logs.

The Staff signal is knowing that identity is the textbook case where buying is the default and building needs a reason — scale economics past ~50–100M MAU, unusual assurance or data-residency requirements, or identity being the product itself. See Buy or Build: The Total-Cost Test and fault line 5.

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Which intentConsumer, workforce and service identity are different systems"I'll assume consumer and treat enterprise SSO as an adapter."
Revocation requirementDetermines token model and lifetime"I'll target 10 min worst case, 30s for high-risk revocations."
Assurance levelsPayments or health data may need step-up and phishing-resistant MFA"I'll design step-up as a first-class session attribute: auth_time and amr."
Number of productsOne login for many apps changes the session model to SSO"If there are several products, I'd centralize the session and issue per-audience tokens."
Regions and residencyEU or other residency rules decide where credential data lives"Home-region per account, with credentials stored only there."
Who owns recovery policySecurity, product and support each want different friction"I'll propose tiers and name security as the owner of the floor."
Existing identity systemsMigration from a legacy hash or token format is often the real project"Is there an existing user base with legacy password hashes I need to migrate on login?"

2.4 Precise Terminology#

TermMeaningCommon Confusion
Authentication (authN)Proving who the caller isConflated with authorization; most "auth service" designs mix both
Authorization (authZ)Deciding whether that caller may do this action on this resourcePushed into the IdP as giant role claims that go stale
IdP / Authorization ServerThe system that authenticates users and issues tokensTreated as the same thing as the gateway that verifies tokens
Access tokenShort-lived credential presented to APIsAssumed to be revocable instantly
Refresh tokenLong-lived credential used only at the token endpoint to get new access tokensSent to APIs, logged, or given the same lifetime as the session
ID tokenOIDC token that tells the client who logged inUsed as an API access token
SessionServer-side record that a user authenticated on a deviceAssumed to exist when tokens are purely stateless
Revocation lagTime from "revoke" to the last moment a copy of the credential is accepted anywhereMeasured at the store write, not at the verifiers
Step-upRe-authenticating with a stronger method before a sensitive actionImplemented as a client-side check
auth_time / amr / acrWhen, how and at what assurance the user last authenticatedIgnored, so a 30-day session can change a password
Sender-constrained tokenToken usable only by the holder of a bound key (DPoP, mTLS)Assumed to be standard; most bearer tokens are not bound
Credential stuffingReplaying breached username/password pairs from other sitesConfused with brute force; the passwords are often correct
ATOAccount takeover — an attacker controlling a real user's accountMeasured only when users complain

3. Where the Design Splits#

Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.

3.1 Fault Line 1: Stateless Tokens vs Server-Side Sessions#

The tension: A server-side session is revocable the instant you delete it, but every request needs a lookup. A signed token needs no lookup, but once issued it is valid until it expires, wherever it has been copied.

StrategyWhat WorksWhat BreaksWho Pays
Opaque session ID, lookup on every requestInstant revocation; tiny cookie; simple mental modelSession store in every request's critical path: 1M lookups/s, +1–2ms each, store outage = company outageEvery product team (latency, availability); the session-store on-call
Long-lived stateless JWT (hours to days)Zero lookups; trivially scalable verificationRevocation lag = remaining lifetime; "log out everywhere" is a lieUsers whose tokens are stolen; support; security
Short JWT + server-side refresh session (hybrid)Lookups only at refresh (~1/40th of request rate); revocation lag bounded by access TTLTwo token types; refresh endpoint is a hot path; lag is still minutes without a push pathIdentity team (more moving parts); users carry ≤ 10 min exposure
Hybrid + pushed deny listHigh-risk revocations in seconds without per-request lookupsDistribution pipeline to every verifier; must be monitoredGateway/platform team holds the deny list; identity owns propagation SLO
Opaque token + introspection with short cacheCentral control; easy revocationIntrospection traffic and latency; cache TTL is the new lagEvery service's latency; the IdP's capacity

The Staff default: hybrid plus deny list. 10-minute access JWTs verified locally (the mechanics are in Edge Gateway), opaque refresh tokens checked against a server-side session store at refresh time, and a pushed deny list for high-risk revocations. The arithmetic: at 1M API requests/s and one refresh per active client per 10 minutes, the session store sees ~25K reads/s instead of 1M — a 40× reduction — and revocation lag is 10 minutes worst case, ~30 seconds for events that matter.

"I want the session store in the refresh path, not the request path. That's a 40× cut in dependency load, and I buy back revocation speed with a push channel for the 0.5% of revocations where minutes matter."

When to deviate:

  • Low-volume, high-assurance apps (admin consoles, banking back-office): opaque sessions with per-request lookup. 200 requests/s doesn't need a token cache, and instant revocation is worth 2ms.
  • Server-rendered web apps on one backend: a classic cookie session in Redis is simpler and fine; JWTs add nothing when the verifier and the session store sit together.
  • Offline-capable clients: longer access tokens are sometimes unavoidable; compensate with device binding.

3.2 Fault Line 2: Token Lifetime — Short + Refresh vs Long-Lived#

The tension: Every minute of access-token lifetime is a minute a stolen copy works. Every minute you remove is refresh load and a chance for a flaky mobile network to log someone out.

Access TTLRefresh load (60M DAU, ~2h active/day)Worst-case revocation lagNotes
1 min~7.2B/day, ~250K/s peak1 minRefresh becomes the biggest service you run; battery and data cost on mobile
5 min~1.4B/day, ~50K/s peak5 minReasonable for high-risk apps
10 min~720M/day, ~25K/s peak10 minDefault; matches common gateway JWKS cache TTLs
60 min~120M/day, ~4K/s peak60 minNeeds the deny list for anything sensitive
24 h~60M/day24 hEffectively unrevocable; avoid

Refresh-token lifetime is a separate decision: how long a device stays logged in without re-authenticating. Consumer default here: 30-day sliding, 90-day absolute, with step-up for sensitive actions based on auth_time, not on whether a session exists.

Rotation and reuse detection. Every refresh returns a new refresh token and invalidates the old one. If an old token is presented again, either the client retried or someone else has a copy — the server can't tell which, so it revokes the whole family. The OAuth security BCP requires refresh tokens for public clients to be either sender-constrained or rotated (RFC 9700).

Diagram: 3.2 Fault Line 2: Token Lifetime — Short + Refresh vs Long-Lived

The race: a mobile client sends RT1, the server rotates to RT2, the response is lost, the client retries with RT1. Without a grace window, that's a false theft signal and a logged-out user. The fix: accept the immediately previous token for 30–60 seconds from the same device binding and return the same RT2 (idempotent rotation). Outside that window, reuse revokes the family.

Who pays: short access TTLs are paid by the identity team (refresh capacity) and mobile users (battery, a few hundred bytes per refresh). Long TTLs are paid by the users whose tokens are stolen. Rotation is paid by users on bad networks if the grace window is wrong — watch refresh.reuse_detected_total split by client version; a spike after a mobile release is a client bug, not an attack.

The Staff default: 10-minute access, rotated refresh with a 60-second idempotent grace, family revocation on reuse, DPoP binding for first-party clients when the client platform supports secure key storage (RFC 9449).

3.3 Fault Line 3: Fail Closed vs Fail Open When the IdP Is Unreachable#

The tension: If the identity plane is down and you fail closed, nobody can do anything — the identity outage becomes a company outage. If you fail open, you might accept a session that was revoked during the outage. See the general framework in Graceful Degradation: Fail Open or Closed.

The mistake is answering this once for the whole system. The Staff answer splits it by operation:

Diagram: 3.3 Fault Line 3: Fail Closed vs Fail Open When the IdP Is Unreachable
PostureWhat WorksWhat BreaksWho Pays
Fail closed everywhereNever accepts a revoked sessionIdP outage = total outage; recovery causes a login stampedeEvery product, every user, revenue
Fail open everywhereProducts stay upAttackers with revoked or expired credentials walk in; sensitive actions unprotectedSecurity, users whose accounts are compromised
Split by operation (Staff default)Existing users keep working ≤ 30 min; new logins and sensitive actions blockedRevocations made during the outage take effect only via deny list or after recoveryUsers who were revoked during the outage get up to 30 min extra exposure; security signs that off

The grace extension, precisely: the token service (or a regional fallback signer) can verify the refresh token's own signature or MAC without the session store if refresh tokens are self-verifiable — for example an opaque random value plus an HMAC over session_id and expiry. It issues a 10-minute access token marked grace=true, repeats at most 3 times (30 minutes), and never for sessions on the local deny list. Services that care (payments, admin) reject grace=true tokens for sensitive operations.

Who signs off: security owns the maximum grace duration; product owns which features accept grace=true. It's a signed table, reviewed twice a year, not a runbook improvisation at 3 a.m.

🎯 Staff Move: "I'd rather let a user whose session was revoked during a 20-minute outage keep browsing for 20 minutes than log 20 million people out and then take 20 million logins in five minutes when we come back. But password changes and payouts fail closed, and security signs the table that says so."

3.4 Fault Line 4: Central Authorization Service vs Policy in Each Service#

The tension: Authentication answers "who". Authorization answers "may they do this to that". A central authorization service gives one consistent answer and one audit trail, and becomes a dependency of every request. Policy inside each service is fast and independent, and drifts.

StrategyWhat WorksWhat BreaksWho Pays
Roles as claims in the tokenZero lookups; simpleStale for the token lifetime; token bloat (a user in 400 groups); can't express "can read doc 9001"Security (stale permissions), users with huge tokens
Policy library in each serviceFast; no new dependency; service owns its dataN implementations drift; audit across services is hardSecurity and compliance (inconsistent rules)
Central policy engine, local evaluationOne policy language and repo; decisions computed in-process from pushed policy and dataData for decisions must be replicated to every service; policy push is a deployPlatform team (distribution); service teams (adoption)
Central relationship-based service (Zanzibar-style)Consistent answers for sharing graphs; one audit trail; handles "folder → doc → user"New tier-0 dependency; needs caching and consistency tokens to be fast and correctPlatform team (huge build); every caller's latency budget

The Staff default: keep authentication central and authorization close to the data. Tokens carry identity and a small, stable set of coarse claims (tenant, account tier, a few scopes) — not permissions. The gateway enforces coarse checks; services enforce resource-level decisions with a shared policy library and a common decision-log format. Move to a central relationship service only when the product's permission model is genuinely a graph (document sharing, nested orgs) and several services need the same answers — that's what Zanzibar was built for, at a cost most companies shouldn't pay first.

When to deviate: collaborative products with sharing (docs, drives, design tools) hit the graph problem early; for them a central relationship service is the right Year-1 investment. See Collaborative Documents for the product side.

3.5 Fault Line 5: Build vs Buy the Identity Provider#

The tension: Identity is the most security-critical code you'll run, which argues for buying it from a specialist. It's also in every request and every product decision about signup friction, which argues for owning it. See Build vs Buy Framework.

OptionWhat WorksWhat BreaksWho Pays
Buy a hosted CIAM / workforce IdPHardened MFA, federation, compliance certifications, fast startPer-MAU pricing at scale; their outage and their breach are yours; limited control of login UX and risk logicFinance (fees); you inherit vendor incidents
Open-source IdP, self-hostedControl, no per-user fees, standard protocolsYou run a tier-0 system: upgrades, CVEs, scaling, on-callYour platform team: 3–6 engineers plus a 24×7 rotation
Build in-houseFull control of UX, risk, data model, cost at 100M+ MAUYears of security work; every protocol edge case is yours10–30 engineers plus security review; the opportunity cost
Hybrid: buy workforce, build or self-host consumerMatches each intent's economicsTwo systems to integrate and auditSecurity (two control planes)

The Staff default: buy workforce identity, always. For consumer identity, buy below ~10M MAU, decide on the numbers between 10M and 100M, and build or self-host above ~100M MAU or when login UX and risk logic are core to the product. Whatever you choose, own the session and revocation layer's interface: products should depend on your token contract, not the vendor's SDK, so a vendor change is a migration, not a rewrite.

The vendor-incident caveat. Buying moves the work, not the risk. The Okta support-system incident above shows a vendor breach reaching customer sessions. If you buy, you still need: your own revocation path, your own monitoring of admin sessions, and a contract that says how fast the vendor notifies you.


4. When It Breaks#

4.1 Signing-Key Leak — The Emergency Rotation#

t=0:       Security finds the active token-signing private key (kid=k42) in a debug
           artifact uploaded to a third-party ticketing tool 6 hours ago.
t=+5min:   Incident declared. Assume forged tokens for any user, any audience.
t=+10min:  Generate k43 in KMS (non-exportable). Publish k43 in JWKS alongside k42.
t=+12min:  Mark k42 "revoked" in the verifier config channel — pushed, not waiting
           for JWKS TTL. Gateways reject kid=k42 within 60s.
t=+13min:  Every access token signed by k42 now fails: ~18M active clients get 401.
t=+13min:  Clients refresh. Refresh tokens are opaque and checked against the session
           store — unaffected by the JWT key — so refresh succeeds with k43 tokens.
t=+15min:  Refresh traffic spikes from 25K/s to ~180K/s. Token service autoscaled from
           pre-warmed capacity, plus jittered client retry (0–60s).
t=+25min:  Refresh traffic back under 40K/s. 0.4% of clients failed and fell back to login.
t=+2h:     Audit: tokens signed with k42 after the leak window carrying unusual audiences
           or sessions that don't exist in the session store = forged. 0 found.

Why it went this well: refresh tokens were not signed by the same key as access tokens, so a JWT-key leak didn't force a password re-login for 60M users. Verifiers honored a pushed key-revocation signal instead of waiting for JWKS cache expiry. Refresh capacity was sized for a "everyone refreshes at once" event.

Detection: tokens.verified_with_unknown_session_total (a valid signature on a session that doesn't exist is a forgery signal), secrets scanning on outbound artifacts, KMS Sign call volume by caller.

Prevention: keys in KMS or HSM, never exportable; signing only through a narrow signer service; per-audience or per-token-type keys so one leak doesn't forge everything; verifiers enforce iss, aud and key scope, not only signature validity — the lesson Microsoft published after Storm-0558.

Owner: identity team (rotation); security (incident command); gateway platform (verifier key config).

4.2 IdP Outage Locks Every Service#

t=0:       Session-store primary in us-east fails over; replica promotion stalls
           (replication lag 40s, failover guard waits). Refresh calls time out.
t=+1min:   refresh.error_rate 0.1% → 92%. Access tokens keep working (local verify).
t=+10min:  Without grace: first wave of 10-min access tokens expires. Users see 401s.
           Apps call refresh in a tight loop: 25K/s → 400K/s retries.
t=+12min:  Token service pods CPU-pinned by retries. Login also shares the pool.
t=+15min:  "Can't log in" trends on social media. Every product reports outage.
t=+25min:  Session store recovers. 15M clients all fall back to full login.
t=+26min:  Login at 40K/s → argon2id pool saturated → login p99 30s → more retries.
t=+70min:  Stable after login admission control and client backoff hotfix.

With the Staff design:
t=+1min:   Refresh fails → token service issues grace tokens (grace=true, 10 min, max 3).
t=+10min:  Users notice nothing. Payouts and password changes show "temporarily unavailable".
t=+25min:  Session store back. Grace tokens refresh normally. No login stampede.

Detection: refresh.error_rate, refresh.grace_issued_total, session_store.replication_lag_seconds, login.queue_depth.

Mitigation: grace extension (fault line 3); separate thread pools for refresh and login so a refresh storm can't starve login; clients with exponential backoff and jitter on refresh, enforced in the shared client SDK; Retry-After honored.

Prevention: session store replicated within the region with automatic failover tested monthly; regional cells so one region's session store never affects another (Multi-Region Active-Active); quarterly IdP-down game day.

Owner: identity on-call (primary); client platform team owns SDK backoff behavior.

4.3 Credential-Stuffing Wave#

t=0:       07:40 local, morning ramp: 600 legit logins/s.
t=+1min:   Attack starts: 45K login attempts/s from ~220K residential proxy IPs,
           each IP sending < 1 attempt/min. Per-IP limits don't trigger.
t=+2min:   argon2id pool (2,000 cores, ~20K hashes/s) saturated. Legit login p99 → 12s.
t=+3min:   Page: login.attempts_rate 75× baseline, login.success_rate 94% → 3%.
t=+5min:   On-call enables wave mode: unknown-device attempts get a challenge before
           hashing; identifiers seen in breach corpus with no prior device → step-up.
t=+8min:   Hash pool load drops to 30%. Legit success rate back to 88% (challenge friction).
t=+30min:  ~0.6% of attempted pairs were valid. Those accounts: forced reset + notification,
           sessions created during the wave revoked.
t=+2h:     Attack stops. Challenge rate decays back to 0.3%.

Detection: login.attempts_rate vs baseline, login.success_rate (a sharp drop is the stuffing signature — most pairs are wrong), login.unique_identifiers_per_min, ratio of failures on non-existent identifiers.

Mitigation: layered shedding before the hash: client-integrity and device signals, ASN and proxy reputation, per-identifier limits, challenges for unknown devices; a fixed hash budget as a bulkhead so overload produces challenges, not timeouts.

Prevention: breached-password screening at signup and at login (NIST SP 800-63B requires checking new passwords against a blocklist of compromised ones); passkey adoption, which removes the password entirely; notify users on new-device logins.

Owner: identity on-call for capacity; trust and safety for the account remediation; product signs off on the challenge-rate ceiling. Rate-limiter mechanics: Rate Limiter.

4.4 Token Cache Serving Revoked Sessions#

t=0:       A team adds a 1-hour in-process cache of "token → user profile + session valid"
           to cut calls to a profile service. Nobody reviews it as an auth change.
t=+3 weeks: User reports takeover. Support runs "log out everywhere" at 14:02.
t=+3 weeks: Gateway deny list blocks the session at 14:02:20. But the internal service
           behind a second ingress path (partner API) uses its own cache: still serves
           the attacker until 15:02.
t=+3 weeks: Attacker exports 1,200 contacts in that hour.

Detection: this is a silent failure — nothing errors. Catch it with a synthetic canary: every 5 minutes, create a session, revoke it, then probe every ingress path with its access token; alert on any 2xx after revocation_time + 60s. Metric: revocation.canary_accept_after_revoke_total.

Mitigation: purge the cache; subscribe that service to session.revoked.

Prevention: a standard: any cache of identity or session validity MUST have TTL ≤ access-token lifetime and MUST subscribe to revocation events; the canary covers every ingress, not just the main gateway. Generic cache-invalidation tradeoffs: Distributed Cache.

Owner: identity team owns the canary and the standard; the service team owns the fix.

4.5 Account Recovery as the Takeover Path#

Diagram: 4.5 Account Recovery as the Takeover Path
t=0:       Attacker controls victim's email via a separate breach.
t=+2min:   Requests password reset. Link arrives in the victim's (now attacker's) inbox.
t=+5min:   Old design: password changed, all sessions revoked — including the real owner's.
           Attacker removes the passkey and adds their own phone as 2FA.
t=+1 day:  Owner locked out of their own account. Support can't verify them because
           every factor on file now belongs to the attacker.

With cooling-off:
t=+5min:   New password works, but account enters limited mode for 24–72h. The existing
           passkey device receives "Someone is recovering your account — this wasn't me".
t=+40min:  Owner taps "this wasn't me". Recovery cancelled, attacker sessions revoked,
           email flagged as compromised.

Detection: recovery.started_rate, recovery.cancelled_by_owner_rate, recovery.success_then_dispute_rate (recoveries later reported as takeovers), support-assisted recoveries per agent per day.

Mitigation: freeze the account, revoke all sessions, restore from the credential history in the audit log.

Prevention: cooling-off and limited mode for high-value tiers; never let recovery remove existing factors immediately; support tooling that can't change contact details without the scripted flow; encourage two independent recovery factors.

Owner: recovery policy — security sets the floor, product sets friction above it; recovery system — identity team; support-assisted recovery — support operations, audited monthly by security.

4.6 Key Rotation Timeline (Routine)#

Routine rotation is the drill that makes 4.1 survivable. The verifier side — caching JWKS and refreshing on an unknown kid — is covered in Edge Gateway; the issuer side is here.

Diagram: 4.6 Key Rotation Timeline (Routine)

Cadence: every 30 days, fully automated, with jwks.verifiers_missing_active_kid (verifiers that fetched JWKS but don't have the active key) as the pre-flight check. An emergency rotation is the same pipeline with the waits removed and a push signal added.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Signing-key leakSecrets scanning hit; tokens.verified_with_unknown_session_total > 0Every token signed by that keyEmergency rotation, pushed key revocation, mass refreshIdentity + security incident command
Session-store outagerefresh.error_rate > 5%All refreshes in the regionGrace extension ≤ 30 min, separate poolsIdentity on-call
Credential stuffinglogin.attempts_rate > 10× baseline, login.success_rate dropLogin latency for everyone; ATO for reused passwordsWave mode, challenges before hashingIdentity on-call + trust and safety
Revoked session still acceptedrevocation.canary_accept_after_revoke_total > 0Every user relying on logoutPurge rogue cache, subscribe to eventsIdentity (canary) + owning service
Revocation pipeline stalledrevocation.propagation_p99 > 60s, consumer lagHigh-risk revocations slow to 10 minRestart distributor; fallback to shorter access TTLIdentity + gateway platform
Recovery takeoverrecovery.success_then_dispute_rate above 0.5%Individual high-value accountsFreeze, restore credentials from audit historySecurity (policy), identity (system)
Refresh-reuse false positivesrefresh.reuse_detected_total by client versionUsers of one app release logged outWiden grace, hotfix clientIdentity + mobile team
SMS pumping fraudotp.sms_sent by destination country vs baselineCost: tens of $K/dayCountry allowlists, per-number limits, prefer TOTP/passkeysIdentity + finance
Workload CA outagesvid.rotation_failures, cert expiry horizon < 2hEvery deploy and service call as certs expireLong enough cert lifetime to ride out CA outages (≥ 2× MTTR)Platform security

🎯 Staff Insight: The most dangerous identity failure is the one that doesn't error: a revoked session that keeps working. Every other row here pages. That one needs a synthetic canary that revokes a session every 5 minutes and checks every ingress path — otherwise you discover it from a user's takeover report.


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists features: signup, login, MFA, logoutNames consumer vs workforce vs service identity; commits; states revocation lag as a requirementAsks how many identity systems exist today, who owns recovery policy, and what audit regime applies
Token modelJWT with an expiryShort access JWT + server-side refresh session + deny list; quantifies lag and refresh loadSets lifetimes and binding by risk tier as a company standard; tracks maximum credential lifetime in circulation
FailureReplicas and failoverSplits verify / refresh / login / step-up; grace extension with a sign-off; separate poolsIdentity as tier-0 cells; fail-static policy published to every product team; IdP-down game days
AbuseRate limit by IPLayered shedding before the hash; hash budget as bulkhead; breached-password checks; challenge ceilingsPrices stuffing defense vs ATO losses; pushes passkey adoption as the structural fix with a target and an owner
RecoveryReset link by emailRisk-scored recovery, cooling-off, notifications to existing factors, support scriptAssigns recovery-policy ownership across security, product and support; audits support-assisted recovery
Operations"Add monitoring"revocation.propagation_p99, revocation canary, refresh.reuse_detected_total, login.success_rateATO rate per million logins and revocation latency as exec-visible security KPIs

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
States revocation lag as a number"Worst case 10 minutes, 30 seconds for high-risk events via the deny list."
Separates the three rates"Verify at a million a second, refresh at 25K, login at 700 — each designed to its own rate."
Knows stuffing is a CPU problem first"At 100ms per hash, 50K attempts a second is 5,000 cores. We shed before hashing."
Splits failure posture by operation"Verification continues, login fails closed, refresh gets a bounded grace, payouts fail closed."
Treats recovery as a login path"Recovery is the weakest authentication we accept, so it gets cooling-off and notifications."
Designs for the silent failure"A revoked session that still works doesn't page. I'd run a revocation canary across every ingress."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"JWTs are stateless, so we don't need a session store"Has no revocation story; "log out everywhere" is impossible
"We'll check a blacklist on every request"Rebuilds the per-request lookup the JWT avoided, without noticing
24-hour or longer access tokensUnrevocable for a day; shows no account-takeover experience
Rolling their own crypto or token formatIgnores decades of library and protocol hardening
Account lockout after N failures as the stuffing defenseLets attackers lock out real users at will; stuffing uses one attempt per account
No mention of recoveryLeaves the weakest door unguarded

5.4 Common False Positives#

  • OAuth flow fluency ≠ identity design. Reciting the authorization code flow with PKCE is table stakes; the Staff question is what happens after the token is issued.
  • Crypto depth ≠ security judgment. Comparing RSA-2048 with Ed25519 impresses briefly; revocation lag and key management decide whether a breach is contained.
  • "Zero trust" vocabulary ≠ a design. It's Staff only if the candidate names what each request carries (user identity, device identity) and where it's verified.
  • MFA everywhere ≠ safety. SMS OTP is phishable and SIM-swappable; a candidate who stops at "add MFA" hasn't thought about phishing-resistant factors or recovery bypassing them.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minPick consumer/workforce/service; state revocation lag and blast-radius constraints
Entities & API3–5 minAccount, credential, session family, signing key, audit event
Architecture5–10 min≤ 8 boxes; three paths and three rates
Revocation10–18 minToken model, refresh rotation, deny list, propagation SLO
Failure posture18–25 minIdP down: verify / refresh / login / step-up split; grace; stampede avoidance
Abuse + recovery25–34 minStuffing shedding, hash bulkhead, recovery cooling-off, support script
Pivot (interviewer's choice)34–42 minKey compromise, enterprise SSO, multi-region, service identity, authorization
Wrap42–45 minToken = cache of a decision; revocation number and owner; evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Our signing key leaked"Key management and blast radiusSeparate keys for access vs refresh; pushed key revocation; mass refresh capacity
"Add enterprise SSO for B2B customers"Federation and deprovisioningPer-tenant OIDC/SAML connections, SCIM, IdP-initiated logout, tenant-level session policy
"Make it multi-region"Where credentials live, what replicatesHome region per account; tokens verifiable anywhere; revocation events replicated globally
"Now services need identity too"Workload identity vs user identityPlatform-attested short-lived certs; no shared secret; user identity propagated as a separate signed context
"A celebrity's account was taken over"Recovery and support controlsTrace via audit log; recovery path; support override; high-value tier
"How do you know logout works?"Silent-failure detectionRevocation canary across every ingress path

6.3 What to Deliberately Skip#

  • Password complexity rules. One sentence: "length minimum, breached-password check, no composition rules, per NIST 800-63B."
  • Every OAuth redirect. Say "authorization code with PKCE" and move on unless asked.
  • Email-verification UX. Mention it exists.
  • Crypto algorithm selection. "ES256 or EdDSA, keys in KMS."
  • CAPTCHA vendor details. "Challenge" is enough; the decision is when to challenge.

6.4 Follow-Up Questions to Expect#

  1. "A user clicks 'log out everywhere'. List every place a stolen token is still accepted, and for how long."
  2. "The session store is down for 20 minutes. What happens to logged-in users, and to new logins?"
  3. "A mobile release causes 2% of users to be logged out every day. What's your first hypothesis?"
  4. "How do you migrate 200 million legacy SHA-1 password hashes to argon2id?"
  5. "An enterprise customer fires an employee. How fast does that person lose access to our product?"
  6. "How do you rotate signing keys without breaking anyone, and how do you do it in an emergency?"
  7. "Who decides how long account recovery takes for a high-value account?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design authentication for our app."

Staff Answer

"Before drawing — is this consumer login, workforce SSO for our employees, or identity between our services? They're three different systems. I'll assume consumer login, around 200 million accounts and 60 million daily actives, with enterprise SSO as a later federation adapter.

The constraints I'll commit to: a stolen credential stops working within 10 minutes in the worst case and within 30 seconds for high-risk revocations; verification of tokens never depends on the identity service being up; login survives credential-stuffing waves of tens of thousands of attempts per second; and recovery is treated as a login path with its own controls. I'll walk through: the session and token model → the three request paths and their rates → revocation and how it reaches every verifier → what happens when identity is down → stuffing → recovery → who owns what."

Why this is L6:

  • Distinguishes three intents and commits to one
  • States revocation lag and failure independence as requirements, before any boxes
  • Previews the outline as decisions ending in ownership

What L7 adds:

  • Asks how many login systems the company already has and whether this is a consolidation
  • Asks who owns recovery policy and which audits (SOC 2, regulators) the design must satisfy
  • Frames identity as tier 0 with an explicit availability target that caps every product's
❌ Common L5 Trap

"Users log in with email and password, we hash with bcrypt, issue a JWT valid for 24 hours, and every service validates it. We'll add MFA and rate limiting."

Why this misses: Every part is reasonable and the design has a 24-hour revocation lag, no failure posture and no recovery story. The interviewer's next question — "the user's token was stolen" — has no good answer.


Drill 2: Revocation Mechanics#

Prompt: "A user reports their account was compromised and clicks 'log out of all devices'. Walk me through exactly what happens."

Staff Answer

"The revoke-all call marks every session family for the account as revoked in the session store, except the caller's current one if it passes step-up, and writes revoked_before = now on the account. It publishes session.revoked events with the session IDs and an account.sessions_revoked_before event. The deny-list distributor pushes those to every gateway node; p99 is under 30 seconds. Gateways reject any access token whose sid is on the list or whose iat is before the account's cutoff, until the entry's 10-minute TTL lapses — after which every pre-revocation access token has expired anyway. Refresh attempts on any revoked family fail immediately at the session store.

Then the account side: force a password change if one is in use, notify all factors, and flag the account for trust-and-safety review. I'd also check revocation.canary hasn't been failing — if any ingress path ignores the deny list, the user's 'log out' is partial."

Why this is L6:

  • Distinguishes refresh-path enforcement (instant) from access-token enforcement (pushed, ≤ 30s)
  • Uses an account-level cutoff so the deny list stays small even for users with many sessions
  • Ties the deny-list TTL to the access-token lifetime with a reason

What L7 adds:

  • Extends revocation across vendor boundaries — partner apps and SaaS receive session-revoked signals via the OpenID shared-signals standards (OpenID CAEP)
  • Makes revocation latency a company security KPI with a monthly review
  • Requires every new ingress path to pass the revocation canary before launch

Drill 3: Make It Concrete — Capacity#

Prompt: "Size the identity plane. How many machines, how much storage?"

Staff Answer

"Three paths. Login: 20 million interactive logins a day, ~230/s average, ~700/s at the morning peak. argon2id tuned to ~100ms of one core: 70 cores at peak. I'd provision a hash pool of 2,000 cores across regions — that's 20K hashes/s — not for legitimate traffic but so a stuffing wave degrades gracefully while shedding kicks in. Passkey logins cost ~1ms and change this math as adoption grows.

Refresh: 60M daily actives, ~2 active hours, one refresh per 10 minutes — 720 million refreshes a day, ~8K/s average, ~25K/s peak. Each is one session-store read, one conditional write and one signature. A remote KMS sign call per token adds milliseconds and a per-request fee at 25K/s, so I'd run a narrow in-region signer service backed by an HSM or KMS-protected key that only it can use — the one place a private key is ever usable.

Session store: ~400M refresh families (≈2 devices per account) × ~500 bytes ≈ 200 GB, ×3 replicas ≈ 600 GB. Fits a sharded Redis cluster with persistence or DynamoDB; I'd pick DynamoDB-style managed storage for durability and TTL-based expiry, with p99 reads ~2ms.

Verification: ~1M/s at gateways, local, ~0.1ms each: ~100 cores across the fleet, which is noise."

Why this is L6:

  • Sizes each path separately and spots that the login pool is sized for attack, not for users
  • Notices that per-call KMS signing doesn't fit the refresh rate
  • Derives storage from sessions per account, not from users

What L7 adds:

  • Prices it: hash pool, session store and SMS spend per month vs vendor per-MAU pricing
  • Sets a passkey-adoption target because it shrinks the most expensive and most attacked path
  • Plans per-region cells so capacity is sized per cell, not globally

Drill 4: The Identity Provider Is Down#

Prompt: "Your session store is unreachable in one region for 25 minutes. What does every other team experience?"

Staff Answer

"API calls with valid access tokens keep working: verification is local. Refresh fails over to grace mode: the token service validates the refresh token's own MAC and expiry, checks the local deny list, and issues a 10-minute access token with grace=true, at most three times. So for 30 minutes, logged-in users notice nothing. New logins fail with a clear 'try again shortly' and Retry-After. Sensitive actions — password change, payout setup, admin — reject grace=true tokens and show 'temporarily unavailable'. Other regions are unaffected because sessions are homed per region and no region calls another synchronously.

The cost: a session revoked during the outage that isn't on the deny list keeps working until grace ends — up to 30 minutes. Security signed off on that maximum. At 26 minutes, if the store is still down, the grace budget runs out and I'd rather extend once more by explicit incident-commander decision than auto-extend."

Why this is L6:

  • Gives each operation its own posture with a named maximum
  • Prevents the login stampede that turns a 25-minute outage into a 70-minute one
  • Names who signed off on the residual risk

What L7 adds:

  • Publishes the fail-static contract so every product knows which of its features accept grace=true
  • Runs this exact scenario as a quarterly game day across product teams
  • Tracks "identity-caused product minutes of downtime" as the identity platform's error budget

Drill 5: Credential Stuffing at the Morning Peak#

Prompt: "45,000 login attempts per second just started, from 200,000 IPs. Most fail. What do you do?"

Staff Answer

"First, protect the hash pool: it does 20K hashes a second, so without shedding legitimate users time out. Switch to wave mode: attempts from unknown devices, or from ASNs dominated by residential proxies, get a challenge before we hash. Identifiers with no prior successful device and an appearance in our breached-credential corpus get step-up. Per-identifier limits drop from 10 failures per 15 minutes to 3. I don't lock accounts — that hands the attacker a way to lock out real users.

Second, find the hits: successful logins during the wave from new devices get revoked and forced through reset with notification. Third, measure friction: challenge rate on known-good devices must stay under 2%, product's ceiling.

After: the structural fix is fewer passwords. Passkey adoption is what makes stuffing irrelevant."

Why this is L6:

  • Treats the wave as a capacity attack first, a takeover attack second
  • Rejects lockout with a reason
  • Names the friction budget and who owns it

What L7 adds:

  • Prices it: ATO losses and support cost per wave vs challenge friction lost logins
  • Sets a company passkey target with a product owner, because it removes the attack surface
  • Shares anonymized attack indicators with peer companies and the gateway team

Drill 6: Multi-Tenant Enterprise SSO#

Prompt: "B2B customers want their employees to log in with their own IdP. One customer has 200,000 employees."

Staff Answer

"Per-tenant federation: each tenant configures an OIDC or SAML connection; we discover the tenant from the email domain or a tenant-specific login URL. Their IdP authenticates; we create our own session with tenant_id and the IdP's amr and auth_time, and issue our normal tokens. Provisioning and deprovisioning via SCIM, so a fired employee's account is deactivated — and its sessions revoked — when the customer's IdP sends the update. Tenant-level policy: session lifetime, required MFA, allowed IP ranges.

Isolation: one tenant's misconfigured IdP or SCIM flood mustn't hurt others — per-tenant rate limits on SCIM and per-connection circuit breakers on metadata fetches. The 200K-employee tenant's Monday-morning login burst — maybe 50K logins in 30 minutes, ~30/s — is easy; their SCIM bulk sync of 200K users is the real load and gets its own queue."

Why this is L6:

  • Keeps our session model; federation is an authentication method, not a new token system
  • Names deprovisioning as the correctness requirement enterprises care about
  • Isolates tenants from each other's misconfigurations

What L7 adds:

  • Prices enterprise SSO as a product feature tier and owns its SLA
  • Accepts continuous-access signals from customer IdPs so their "disable user" reaches our sessions in seconds
  • Decides which tenant-policy knobs to expose and which to keep as company-wide floors

Drill 7: Build vs Buy#

Prompt: "Should we build our own identity provider or buy one?"

Staff Answer

"Workforce: buy, no debate — the SaaS connector catalogue and device-posture integrations are the product. Consumer: depends on three numbers. At 5 million MAU, a hosted CIAM at even a few cents per MAU a month is far cheaper than the 6–10 engineers plus a 24×7 rotation it takes to run identity well. At 150 million MAU, per-MAU fees can reach millions per year, and login UX and risk logic are core to conversion — building or self-hosting an open-source IdP starts to pay. Third number: how many products. If we have five apps, we need central sessions regardless.

Either way I'd own the interface: products consume our token contract and session API, not the vendor's SDK, so changing vendors is a migration of one service."

Why this is L6:

  • Splits by intent and decides with numbers
  • Includes the on-call and security cost of running tier 0, not only build effort
  • Protects reversibility with an owned interface

What L7 adds:

  • Models the 3-year cost curve with MAU growth and the crossover point
  • Negotiates vendor contract terms for incident notification and data export
  • Plans the exit before signing: password hash export format, session migration

Drill 8: Changing Token Lifetime Without an Outage#

Prompt: "Security wants access tokens cut from 60 minutes to 10. How do you roll that out?"

Staff Answer

"The risk isn't the token — it's 6× refresh load and every client that assumed 60 minutes. Steps: (1) capacity: refresh goes from ~4K/s to ~25K/s at peak; scale the token service and session store first and load test at 2×. (2) Shadow: emit the would-be expiry in a claim and log clients that don't refresh before it — that finds hardcoded timers. (3) Canary: 1% of new sessions get 10-minute tokens; watch refresh.error_rate, refresh.reuse_detected_total by client version, and 401 rate at the gateway. (4) Ramp 1% → 10% → 50% → 100% over two weeks, by platform, web first, mobile after the SDK fix ships. (5) Rollback is a config flag; existing tokens keep their original expiry."

Why this is L6:

  • Recognizes the change as a capacity and client-compatibility change
  • Uses shadow → canary → ramp with named metrics
  • Keeps rollback a config change

What L7 adds:

  • Turns token lifetime into a policy by risk tier so the next change is a config edit, not a project
  • Sets a deprecation date for clients older than the SDK that handles short tokens
  • Reports the security gain (revocation lag 60 → 10 min) in the security KPI review

Drill 9: Multi-Region#

Prompt: "We're expanding to the EU and must keep EU users' identity data there."

Staff Answer

"Home region per account. Credentials, sessions and audit events for EU accounts live only in EU cells. The identifier-to-region directory is global but holds only a hashed identifier and a region code. Login hits the nearest edge, which looks up the home region and routes there — +80–120ms for a traveling user, acceptable for login. Access tokens are verifiable anywhere: public keys are published globally, so an EU user calling a US-hosted API needs no cross-region identity call. Refresh goes to the home region.

Revocation events replicate globally: they carry session IDs and account IDs only, which I'd confirm with legal counts as acceptable. A region outage means its users can't log in or refresh — grace mode covers 30 minutes — but other regions are untouched. Replication internals: Replication."

Why this is L6:

  • Separates what must be homed (credentials, sessions) from what can be global (public keys, revocation signals)
  • Keeps verification region-independent
  • Checks the residency status of revocation data instead of assuming

What L7 adds:

  • Makes "identity cell per jurisdiction" a template so the next region is weeks, not quarters
  • Owns the legal interpretation with privacy counsel as a documented decision
  • Plans account migration between regions (user moves countries) as a supported flow

Drill 10: Cost#

Prompt: "Identity costs are growing faster than users. Where's the money going?"

Staff Answer

"Four usual suspects. SMS OTP: at $0.02–0.05 per message and a few percent of logins sending one, it's often the biggest line — and SMS pumping fraud can multiply it; I'd check spend by destination country. Password hashing: attack waves drive hash-pool autoscaling, so cost tracks attacks, not users — shed earlier. Refresh traffic: if a client bug refreshes every minute, refresh load is 10× design. Vendor per-MAU fees if we buy: inactive accounts may still count. Fixes in order: passkeys and TOTP over SMS, country controls on SMS, shedding before hashing, client SDK refresh discipline, purge dormant accounts."

Why this is L6:

  • Knows SMS and attack-driven hashing are the non-obvious cost drivers
  • Connects cost spikes to client bugs and abuse, not growth
  • Orders fixes by savings

What L7 adds:

  • Builds a cost-per-1,000-logins unit metric reported monthly
  • Charges SMS spend back to products that choose SMS as their step-up method
  • Uses passkey adoption as a cost lever in the business case, not only a security one

8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Launch-Day Login Meltdown#

Context: A new product launches with a TV campaign. Signups and logins hit 12× normal in 10 minutes; at the same moment a stuffing botnet piggybacks on the traffic. Login p99 is 25 seconds and the CEO's demo account can't sign in. The on-call escalates to you.

Questions to Surface First:

  • How much of the traffic is legitimate? What's login.success_rate doing — dropping (stuffing) or steady (real users)?
  • Is the hash pool saturated, or is something upstream (risk engine, credential store) the bottleneck?
  • Are signups and logins sharing the same pool?
  • Did the client release for the launch change refresh or retry behavior?

Typical L5 Approach: Scales up the login service and adds per-IP rate limits. The scale-up takes 8 minutes, the botnet's per-IP rate is already below the limit, and new pods saturate on the same hashing cost.

Staff Approach: Separates signal from noise first: success rate fell from 94% to 21%, so most traffic is stuffing. Enables wave mode (challenge unknown devices before hashing), splits signup into its own pool so launch signups aren't starved, and finds the launch client retried failed logins 3 times with no backoff — tripling legitimate load.

Principal Approach: Treats launch-day identity readiness as a launch-checklist item owned by identity: capacity sign-off, wave-mode rehearsal, client SDK version gate. Makes "identity load test at 20× with attack mix" mandatory for any campaign with paid media.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Enable wave mode; confirm hash pool load drops below 70%; protect known-device logins with a priority lane.
Triage78% of attempts are stuffing from residential proxies; launch client retries ×3 with no jitter.
Quick fixServer-side Retry-After honored by the SDK via remote config; separate signup pool.
GuardrailsAlert on login.success_rate < 70% for 3 min; hash-pool saturation auto-enables wave mode.
Post-mortemWhy could a launch client ship retry logic without identity review? Why was wave mode manual?

Metrics to Watch: login.success_rate, login.hash_pool_utilization, login.challenge_rate{device=known}, signup.p99_latency

Organizational Follow-up: identity owns the auth client SDK; apps can't implement their own login retry. Launch checklist adds an identity capacity sign-off.

Ownership Question: "Who decides to enable wave mode, given it adds friction?" Staff answer: Identity on-call, automatically on saturation, within a challenge-rate ceiling product has pre-approved. Above that ceiling, the incident commander decides with product on the call.

Key Takeaway: "Login capacity is sized for the attack, and the shedding decision is pre-approved — not debated during the incident."

What clears the Staff bar:

  • Reads success rate before scaling
  • Protects the hash pool with shedding, not more pods
  • Finds the client-retry multiplier

Deep Dive 2: Silent Failure — Logouts Haven't Propagated for Three Days#

Context: A security researcher reports that after "log out everywhere", their old access token still worked for exactly 10 minutes — not 30 seconds. Investigation shows the deny-list distributor stopped consuming three days ago after a schema change. Nothing paged.

Questions to Surface First:

  • Did revocations still work at the refresh path? (Yes — session store was fine, so lag fell back to 10 minutes, not infinity.)
  • Which high-risk revocations happened in those three days, and what did those sessions do in the 10-minute windows?
  • Why didn't consumer lag or the canary alert?

Typical L5 Approach: Restarts the distributor, fixes the schema mismatch, closes the bug.

Staff Approach: Fixes the consumer, then fixes detection: consumer lag was monitored but routed to a dashboard, and the revocation canary only probed the refresh path, not the gateway. Adds the canary on every ingress with a 60-second threshold, and reviews the 4,100 high-risk revocations from the window for activity after revocation.

Principal Approach: Recognizes that the system degraded gracefully by design — the 10-minute TTL bounded the damage — and that the gap is in owning the SLO. Makes revocation.propagation_p99 a paging SLO, adds schema-compatibility checks to the event bus for every security-relevant topic, and reports the incident in the security KPI review.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Roll the distributor back to the old schema reader; replay the last 10 minutes of revocations.
Triage4,100 high-risk revocations in 3 days; 37 sessions made API calls after revocation, all within 10 min.
Quick fixReview those 37 sessions with trust and safety; notify affected users where data was accessed.
GuardrailsCanary every 5 min per ingress; page if any 2xx after revoke + 60s. Consumer lag > 60s pages.
Post-mortemSchema change on a security topic had no compatibility check; lag alert was not paging.

Metrics to Watch: revocation.propagation_p99, revocation.canary_accept_after_revoke_total, denylist.consumer_lag_seconds, denylist.entries

Ownership Question: "Who owns the deny list — identity or the gateway team?" Staff answer: Identity owns the stream and the propagation SLO; the gateway team owns applying it. The canary measures the end-to-end result, and identity carries the page.

Key Takeaway: "Bounded failure is the reason the 10-minute TTL exists. Measuring the fast path is the reason the canary exists."

What clears the Staff bar:

  • Notices the access TTL bounded the damage, and says so
  • Measures the end-to-end revocation result, not component health
  • Quantifies exposure from the incident window

Deep Dive 3: Large-Customer Onboarding — A 200,000-Employee Bank#

Context: Sales signs a bank with 200,000 employees. Requirements: SAML SSO through their IdP, phishing-resistant MFA, sessions ≤ 8 hours, deprovisioning within 15 minutes, every admin action logged and exportable, and an audit of our identity controls before go-live.

Questions to Surface First:

  • Does their IdP support SCIM and push signals, or do we poll?
  • Do they want their MFA policy or ours to govern? (Theirs — we record their amr.)
  • Where must their audit logs live, and for how long?

Typical L5 Approach: Adds a SAML connection and a tenant flag for 8-hour sessions.

Staff Approach: Treats deprovisioning as the hard requirement: SCIM deactivate → revoke all sessions for those users → push deny entries; plus a 15-minute access-token cap for this tenant and a nightly SCIM full-sync to catch missed events. Enforces tenant policy (session length, IP ranges) in the token service, not in the app. Audit events per tenant, exportable via API.

Principal Approach: Turns the bank's requirements into a standard enterprise tier — session policy, deprovisioning SLA, audit export — priced as a product SKU, and accepts continuous-access signals from customer IdPs so their "user disabled" reaches our sessions in seconds without polling.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignPer-tenant SAML connection; tenant policy object (session ≤ 8h, required amr, IP allowlist).
DeprovisioningSCIM deactivate → revoke sessions → deny list; nightly full sync reconciles drift.
LoadSCIM initial sync of 200K users through a dedicated queue at ≤ 200/s; Monday login burst ~30/s.
AuditPer-tenant audit stream, 7-year retention in their required region, export API.
LaunchPilot with 500 users; measure time-to-deprovision with test accounts weekly.

Metrics to Watch: scim.deprovision_to_revoke_p99{tenant}, saml.assertion_errors{tenant}, tenant.policy_violations

Ownership Question: "Who owns a missed deprovisioning?" Staff answer: If the customer's IdP never sent it, the customer. If we received it and didn't revoke within 15 minutes, identity — which is why we measure deprovision-to-revoke per tenant.

Key Takeaway: "For enterprises, the login is easy. Deprovisioning is the product."

What clears the Staff bar:

  • Puts deprovisioning latency at the center with a measured SLA
  • Keeps tenant policy enforcement in the token service
  • Isolates the big tenant's sync load

Deep Dive 4: Post-Mortem — High-Value Accounts Taken Over Through Support#

Context: Over two weeks, 23 high-follower accounts were taken over. None of the attackers knew a password. All 23 went through support-assisted recovery. You own the post-mortem.

Questions to Surface First:

  • What did the attackers present to support, and which agent actions moved contact details?
  • Did existing factors get notified, and were notifications delivered to the attacker's new contact details?
  • Is there an insider or a single agent pattern?

Typical L5 Approach: Retrains support agents, adds a security question to the support script.

Staff Approach: Finds the control gap: agents could change the email address directly, and the recovery flow then sent everything to the new address. Removes direct contact edits from support tooling, routes all support recovery through the same cooling-off flow with notifications to the previous factors, and adds second-agent approval for high-value tiers.

Principal Approach: Treats support as part of the authentication system. Assigns recovery-policy ownership to security with product and support as consulted parties, audits support-assisted recoveries monthly, and measures recovery.success_then_dispute_rate per channel and per agent as a standing security metric.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateFreeze the 23 accounts; restore credentials from audit history; revoke all sessions.
Triage19 of 23 recoveries handled by agents at one outsourced site; scripts allowed direct email change.
Quick fixRemove direct contact-edit permission; force high-value recoveries through 72h cooling-off.
GuardrailsAlert on support recoveries per agent per day > 3× median; notify previous email on any change.
Post-mortemSupport tooling was outside identity's design review. Recovery policy had no named owner.

Metrics to Watch: recovery.support_assisted_rate, recovery.success_then_dispute_rate{channel}, account.contact_change_without_factor_total

Ownership Question: "Who owns the support tool's account-edit permissions?" Staff answer: Identity owns the permission model for anything that changes a credential or contact detail, even inside support tooling. Support owns the workflow on top.

Key Takeaway: "Every human who can change a credential is part of the authentication system."

What clears the Staff bar:

  • Finds the tooling permission, not just the training gap
  • Notifies previous factors, which the attacker doesn't control
  • Measures recovery abuse per channel

Deep Dive 5: Multi-Region Expansion — Identity Cells#

Context: The company runs identity in one US region. A regional cloud outage took login down globally for 3 hours last quarter. Leadership wants regional independence and EU residency within a year.

Questions to Surface First:

  • Which identity data must be homed (credentials, sessions, audit) and which can be global (public keys, routing directory)?
  • How will a user whose home region is down be treated?
  • Which product services call the identity plane synchronously today?

Typical L5 Approach: Deploys a second active region with a globally replicated user database.

Staff Approach: Cells per region with home-region accounts. Global, read-only directory mapping hashed identifier → home region. Public keys published globally so verification works anywhere. Revocation events replicated to all regions. A home-region outage affects only that region's users and grace mode covers 30 minutes for them.

Principal Approach: Sequences the migration as a multi-quarter program with an explicit order — verification independence first, then refresh, then login — and makes "no synchronous cross-region identity call" a design-review rule for every product service.

Staff Approach — Full Reasoning
PhaseWhat to Do
ScopingClassify identity data: homed vs global; legal confirms revocation events may replicate.
ArchitectureRegion cells: login, token, session store, credential store, audit per region.
MigrationMove accounts to home regions in batches of 1M with dual-read for 7 days per batch.
Failure drillsTurn off a cell's login in a game day; verify other cells unaffected.
LaunchEU cell first for new signups, then migrate existing EU accounts.

Metrics to Watch: login.success_rate{cell}, directory.lookup_latency_p99, revocation.cross_region_replication_lag

Ownership Question: "Who owns the global directory, the one shared component?" Staff answer: Identity platform, with the strictest change process in the stack — it's read-only on the hot path and replicated to every cell so its outage doesn't block login.

Key Takeaway: "Identity becomes regional when its outage stops being global."

What clears the Staff bar:

  • Distinguishes homed from global identity data
  • Keeps the one global component off the critical write path
  • Plans migration in batches with dual-read

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • State revocation lag as a number for any token design and explain how to shrink it
  • Design the hybrid model: short access JWT, server-side refresh sessions with rotation and reuse detection, and a pushed deny list
  • Size login, refresh and verification as three separate paths with their own rates
  • Split IdP failure posture by operation and defend the grace extension with a sign-off
  • Explain credential stuffing as a CPU and economics problem and shed it before the hash
  • Design account recovery with cooling-off, notifications and support controls
  • Run a routine and an emergency signing-key rotation, issuer side and verifier side
  • Decide when to buy an identity provider and how to keep the decision reversible

The Bar for This Question#

Mid-level (L4): Builds a working login: hashed passwords, a session cookie or a JWT, a password reset email, maybe TOTP. Doesn't consider revocation lag, outages or abuse. Would ship something that works for honest users and fails the first takeover report.

Senior (L5): Adds MFA, rate limiting, refresh tokens and replicas. Knows JWTs are hard to revoke. The gap: treats revocation as a database delete, sets token lifetimes without connecting them to exposure, answers "IdP down" with "replicas", and sends a reset link to an email that may already be compromised. The design is plausible and would pass review — and would still turn a stolen refresh token into a 90-day problem.

Staff+ (L6): Frames the problem around revocation and blast radius in the first five minutes. Chooses the hybrid token model with numbers — refresh load, deny-list size, lag. Splits failure posture by operation, sheds stuffing before hashing, treats recovery as a login path, and runs a canary for the silent failure. Names who pays each tradeoff: users carry bounded exposure, product carries challenge friction, security signs the grace table. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 "Stateless JWT" Is a Revocation Decision Disguised as a Performance Decision#

ClaimReality
"JWTs scale because there's no lookup"True for verification; refresh and revocation still need state
"We'll add a blacklist"A per-request blacklist lookup is a session store with extra steps
"Short expiry fixes it"Only if refresh checks server-side state; otherwise refresh is the unrevocable part

The Staff position: Every token design has state somewhere. Choose where — per request, per refresh, per risk event — and state the lag it buys.

Why this matters in interviews: "Stateless" is the word that makes interviewers ask about logout. Have the number ready.

10.2 Account Lockout Is a Gift to Attackers#

DefenseEffect on StuffingEffect on Real Users
Lock after 5 failuresNone — stuffing uses 1 attempt per accountAttackers can lock out any account they name
Per-identifier throttle + step-upSlows targeted guessingFriction only for the targeted account
Pre-hash shedding by device and reputationRemoves most of the waveSmall challenge rate on shared IPs

The Staff position: Throttle and challenge; don't lock. NIST's 100-attempt ceiling is a ceiling, not a design.

Why this matters in interviews: "Lock the account" is the most common stuffing answer and the easiest to dismantle.

10.3 Recovery Is the Real Login Page#

The Staff position: Whatever your strongest factor is, your account is only as strong as the weakest way to replace it. A passkey-protected account with email-only recovery is an email-protected account. Design recovery first, then login.

Why this matters in interviews: Bringing up recovery unprompted is one of the clearest Staff signals on this question.

10.4 Most Companies Should Not Build Central Authorization#

StageRight Answer
Roles and tenantsCoarse claims in tokens + policy library per service
Shared policy, many servicesPolicy-as-code with local evaluation and decision logs
Sharing graphs across many servicesCentral relationship service, Zanzibar-style

The Staff position: Zanzibar proves central authorization can be fast. It doesn't prove you should build it. Wait for the graph.

Why this matters in interviews: Candidates who propose a central authorization service at minute 10 have added a tier-0 dependency without a reason.

10.5 SMS Is a Cost Center and a Liability#

The Staff position: SMS OTP is phishable, SIM-swappable, and an attack target for pumping fraud that bills you per message. Keep it as a fallback, make passkeys and TOTP the default, and watch SMS spend by destination country like a security metric.

Why this matters in interviews: Naming SMS's cost and abuse profile shows you've seen the bill, not just the design doc.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a correct identity system. The Principal engineer notices that the company has five — the original consumer login, the acquired product's login with its own user table, an enterprise SSO bolted on by the B2B team, a partner API using static keys, and service-to-service calls authenticated by a shared secret from 2019. Each has its own token format, its own lifetime and its own idea of "log out". The L7 problem is not one correct login; it is identity as a company-wide trust boundary: one revocation story, one key-management discipline, one recovery policy with a named owner, and a migration path that retires the other four without a flag day.

🧭 Principal Move: "I'm not going to ask whether our login is secure. I'm going to ask what the longest-lived credential in circulation is, which system issued it, and who would notice if it were stolen. That one number tells me where to spend the next year."

The Org-Level Fault Line#

One identity platform vs per-product login.

OptionWhat WorksWhat BreaksWho Pays
Each product owns loginProduct speed; tailored UXN recovery flows, N token formats, N stuffing defenses; no "log out everywhere" across productsSecurity (N attack surfaces), users (N passwords), support
Central identity platform owns everythingOne revocation, one recovery policy, one audit trailPlatform becomes a bottleneck for every signup-flow experimentProduct teams (velocity), growth (experiments)
Platform owns primitives; products own experienceCredentials, sessions, tokens, keys, recovery policy are central; login UI and onboarding are productContract design is hard; needs an SDK and a hosted login with themingIdentity platform (API stewardship and SDK)

🧭 Principal Move: "The platform owns everything that must be right once — credential storage, sessions, token issuance, keys, revocation and the recovery floor. Products own everything that must be fast to change — onboarding screens, copy, which step-up method to offer above the floor. New products may not create a user table; the two existing ones migrate within four quarters."

Cost Model#

Assumptions: consumer-heavy, fully loaded engineer ~$250K/year, cloud list prices, SMS at ~$0.03/message blended, vendor CIAM priced per MAU at negotiated rates that fall with volume (estimates, not quotes).

ScaleMAUInfra ($/month)SMS ($/month)HeadcountOn-call LoadBuy Alternative (approx.)
Startup1M~$1–3K~$1–3K0.5–1 eng integrating a vendorVendor's pager; < 1 page/monthVendor ~$2–10K/month — buy
Growth30M~$25–50K (hash pool, session store, multi-AZ)~$30–60K6–10 eng: identity core (4), abuse (2), recovery/support tooling (1–2), SRE (1–2)Dedicated rotation, 3–6 pages/month, mostly abuse wavesVendor ~$100–300K/month — decide on control needs
Enterprise300M~$300–600K (regional cells, attack headroom, audit storage)~$200–500K unless passkeys displace SMS30–50 eng across identity, abuse, workforce, workload identity, plus security partnersFollow-the-sun, per-cell rotationsVendor fees exceed an in-house team; build or self-host

The pricing insight: at growth scale, the two levers worth more than any infrastructure tuning are SMS volume and attack headroom. Moving 30% of step-ups from SMS to passkeys or TOTP saves ~$10–20K/month and removes a phishing vector; shedding stuffing before the hash avoids sizing the hash pool for the attack. Headcount, not infrastructure, is the dominant cost at every scale.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Password hash algorithm and parametersOne-way-ishCan only migrate on next login; dormant accounts keep the old hash for years
Account identifier model (email as ID vs internal ID)One-wayEvery product's foreign keys; email changes become migrations
Who holds credentials (vendor vs in-house)One-wayPassword hashes may not be exportable in a usable format; forced resets for millions
Token audience and claim contract consumed by servicesOne-way-ishEvery service parses it; versioning takes quarters
Passkey relying-party ID (domain)One-wayPasskeys are bound to it; changing domains strands credentials
Access-token lifetimeTwo-wayConfig, with capacity planning
Grace-extension durationTwo-wayPolicy change with security sign-off
Session store technologyTwo-waySessions expire; dual-write for one refresh lifetime then cut over

🧭 Principal Insight: The passkey relying-party ID and the account identifier model are the decisions I'd slow down on. Everything about tokens can change in a quarter; those can't change in a year.

The Standard I'd Write#

RFC-ID-001: Authentication and Session Standard
Status: Approved   Owners: Identity Platform + Security Engineering

Scope
  Every system that authenticates users, issues or accepts user or service
  credentials, or caches identity or session validity.

MUST
  1. Authenticate users only through the Identity Platform API or hosted login.
  2. Access tokens: lifetime ≤ 15 min; refresh tokens: server-side state,
     rotated, reuse revokes the family.
  3. Any cache of identity or session validity: TTL ≤ access-token lifetime and
     subscribed to session.revoked.
  4. Every ingress path passes the revocation canary before launch.
  5. Signing keys live in KMS/HSM, are non-exportable, rotate ≤ 30 days, and are
     scoped per token type and audience. Verifiers check iss, aud and key scope.
  6. Service-to-service authentication uses platform-issued workload identity;
     no static shared secrets.
  7. Account recovery follows the tiered recovery policy owned by Security.

SHOULD
  1. Offer passkeys as the default credential for new accounts.
  2. Sender-constrain first-party tokens (DPoP or mTLS).
  3. Reject grace-mode tokens for sensitive operations.

Exceptions
  Filed with Identity Platform; reviewed within 10 business days; time-boxed to
  2 quarters; CISO sign-off for any exception to MUST 2, 5 or 7.

Success metrics
  - revocation.propagation_p99 ≤ 30s; canary failures: 0
  - Longest-lived user credential in circulation ≤ 90 days
  - ATO per million logins: tracked monthly, target set yearly by Security
  - Static service secrets in config: 0 by end of year 2
  - Identity-caused product downtime: ≤ 20 min/quarter

What I'd Tell the VP#

"We have five ways to log in and no single way to log someone out. If one of our users' credentials is stolen today, it can stay useful for up to 90 days in two of those systems. I'm proposing one identity platform that owns sessions, keys and recovery policy, while product teams keep their signup and onboarding experience. It's about eight engineers for a year, retires two legacy login systems, cuts the maximum lifetime of a stolen credential from 90 days to 10 minutes, and makes our next identity outage a degraded login page instead of a company-wide outage. The main risk is migration friction for the acquired product's users; I'd move them on their next login, not with a forced reset."

Principal Interview Signals#

SignalWhat It Sounds Like
Measures the worst credential, not the average"What's the longest-lived credential in circulation, and who issued it?"
Prices friction against loss"A 1% challenge rate costs us ~0.3% of logins; stuffing costs us ~$400K a quarter in ATO remediation. Here's the trade."
Redraws ownership"Security owns the recovery floor, product owns friction above it, support follows a script it can't override."
Sets the org's failure posture"Identity's error budget is every product's error budget; we publish the fail-static contract."
Knows when not to standardize"Workload identity shares our key discipline but not our session store; forcing one system serves neither."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a secure login with short tokens and MFA" — correct for one product; ignores the four other login systems in the company.
  • "Security will decide the recovery policy" — names an owner but not the tension with growth and support, nor the audit.
  • "We'll buy an IdP" — no exit plan, no ownership of the token contract, no answer for the vendor's breach.

Appendices

Appendix A: Mechanics in Depth#

A.1 Login and Token Issuance#

Diagram: A.1 Login and Token Issuance

A.2 Refresh With Rotation and Grace#

def refresh(rt, dpop_proof):
    sid, mac_ok, exp = parse_and_verify_mac(rt)          # self-verifiable for grace mode
    if not mac_ok or exp < now(): return 401
    try:
        s = session_store.get(sid, timeout=50ms)
    except Unavailable:
        return grace_issue(sid, dpop_proof)               # ≤ 3 × 10 min, deny list checked
    if s.revoked_at: return 401("session_revoked")
    if not dpop_matches(s.binding_key_thumbprint, dpop_proof): return 401
    h = sha256(rt)
    if h == s.current_refresh_hash:
        new_rt = mint_rt(sid)
        ok = session_store.cas(sid, expect=h, set_current=sha256(new_rt), set_prev=h,
                               rotated_at=now())
        if not ok: return refresh_retry_once()
        return issue_access(s), new_rt
    if h == s.prev_refresh_hash and now() - s.rotated_at < 60s:
        return issue_access(s), last_issued_rt(sid)       # idempotent retry, same RT2
    revoke_family(s.family_id, reason="reuse_detected")   # theft or very late replay
    publish("session.revoked", s.family_id)
    return 401("reuse_detected")

A.3 Legacy Password Hash Migration#

Rehash on successful login: verify with the legacy algorithm, then store argon2id and mark the record migrated. For dormant accounts, wrap the legacy hash — argon2id(legacy_hash) — so the weak hash is no longer stored, and verify by computing the legacy hash first. After ~12 months, accounts still on wrapped hashes get a reset on next login.

Appendix B: Data Model and Session Lifecycle#

CREATE TABLE credentials (
  credential_id   TEXT PRIMARY KEY,
  account_id      TEXT NOT NULL,
  type            TEXT NOT NULL,             -- password | passkey | totp | phone | email | federated
  secret_ref      BYTEA NOT NULL,            -- argon2id hash or passkey public key
  added_at        TIMESTAMPTZ NOT NULL,
  trusted_after   TIMESTAMPTZ NOT NULL,      -- cooling-off end; limited mode before this
  revoked_at      TIMESTAMPTZ
);

-- Session store (key-value, TTL = absolute_expires_at)
session:{sid} → { account_id, device_id, family_id, current_refresh_hash, prev_refresh_hash,
                  rotated_at, auth_time, amr, binding_key_thumbprint, expires_at,
                  absolute_expires_at, revoked_at }
account_cutoff:{account_id} → revoked_before timestamp   -- for revoke-all

Refresh tokens are stored only as hashes; a session-store dump does not yield usable tokens.

Diagram: Appendix B: Data Model and Session Lifecycle

Appendix C: Coordination Mechanisms#

C.1 Revocation Fan-Out#

Diagram: C.1 Revocation Fan-Out

The session-store write and the event are coupled through an outbox or change-data capture, so a revocation can't be committed without being published. Event-bus mechanics: Apache Kafka.

C.2 Quick Comparison#

MechanismRevocation LagPer-Request CostFailure ModeUse For
Access-token expiry≤ access TTL (10 min)NoneLong TTL = long exposureBaseline for all revocations
Refresh-path session checkNext refresh (≤ 10 min)None (per refresh)Session store outageLogout, normal revocation
Pushed deny list~30sIn-memory set lookupDistributor stall (silent)High-risk revocations
Account revoked_before cutoff~30sIn-memory map lookupSame as deny listRevoke-all for one account
Key revocation~60s via pushed key configNoneMass re-auth loadKey compromise
Introspection per request~02–20ms network callIdP in every request pathLow-volume, high-assurance APIs
Client revocation endpoint (OAuth token revocation)Immediate at the issuerNoneOnly affects issuer-side checks; verifiers still need the pushClient-initiated logout (RFC 7009)

Appendix D: API Contract & Client Behavior#

  • Clients refresh at ~80% of access-token lifetime with ±10% jitter; on failure, exponential backoff from 1s capped at 60s, and honor Retry-After.
  • Refresh tokens are stored in platform secure storage (Keychain, Keystore) or HttpOnly, Secure, SameSite cookies on web — never in localStorage.
  • First-party mobile clients generate a non-exportable device key and send DPoP proofs on refresh.
  • 401 session_revoked → clear tokens, go to login. 401 reuse_detected → same, plus show a security notice.
  • Login and recovery responses never reveal whether an identifier exists (uniform responses and timing).
  • Authorization code with PKCE for every OAuth client; no implicit grant, no password grant, per RFC 9700.

Appendix E: Observability#

Core metrics:

  • login.attempts_rate, login.success_rate, login.hash_pool_utilization, login.challenge_rate{device}
  • refresh.rate, refresh.error_rate, refresh.grace_issued_total, refresh.reuse_detected_total{client_version}
  • revocation.propagation_p99, revocation.canary_accept_after_revoke_total (must stay 0), denylist.entries
  • recovery.started_rate, recovery.cancelled_by_owner_rate, recovery.success_then_dispute_rate{channel}
  • tokens.verified_with_unknown_session_total (forgery signal), jwks.verifiers_missing_active_kid
  • otp.sms_sent{country}, ato.confirmed_per_million_logins

Critical alerts:

AlertThresholdSeverity
Revocation canary accepted after revoke> 0Sev-1 page
Valid signature, unknown session> 0 sustained 5 minSev-1 page (possible key compromise)
Revocation propagation p99> 60s for 5 minPage
Login success rate< 70% for 3 minPage + auto wave mode
Refresh error rate> 5% for 2 minPage
Refresh reuse detected> 3× baseline for one client versionPage identity + mobile
SMS sends to one country> 5× 7-day baselinePage (pumping fraud)

Debugging the silent failure: a revoked session that still works produces no errors anywhere. The canary is the only reliable detector; second best is tokens.used_after_revocation computed offline from gateway logs joined to revocation events.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 1M MAUHosted CIAM; vendor sessions; your token contract on topVendor limits on custom risk logic
1–30M MAUHybrid tokens, single-region session store, wave modeStuffing cost, SMS spend, second product wants login
30–300M MAUIdentity platform + SDK, passkeys, deny-list fan-out, regional failoverCross-region outages, residency, legacy formats
> 300M MAURegional cells, shared signals, workload identity everywhereOrganizational: retiring the last legacy login

What you don't build on day one: your own HSMs, a central relationship-based authorization service, DPoP for every client, regional identity cells, cross-vendor shared signals. Each has a trigger in Section 11.

Appendix G: Service-to-Service Identity, Multi-Tenancy and Cost#

Diagram: Appendix G: Service-to-Service Identity, Multi-Tenancy and Cost
  • Workload identity: certificates live ~12 hours and renew at half-life, so a CA outage of up to ~6 hours doesn't break running services; it blocks new workloads. The CA is tier 0 for deploys, not for traffic. User identity rides alongside as a separate signed context, never as a workload certificate. Platform mechanics: Kubernetes; service registration: Service Discovery.
  • Tenant isolation: per-tenant rate limits on SCIM and federation metadata fetches; per-tenant session policy; one tenant's misconfigured IdP produces errors only for that tenant.
  • Cost attribution: tag SMS sends and step-up challenges with the requesting product; products that pick SMS as their step-up method see the bill.
  • Abuse costs are bursty: budget hash-pool headroom and SMS spend for attack weeks, not average weeks; report cost per 1,000 logins monthly.
  1. Loading the index…