Technologies referenced in this case study: Redis · DynamoDB · PostgreSQL · Apache Kafka · Envoy, Kong & NGINX · Kubernetes
Related: Edge Gateway · Rate Limiting · Distributed Cache · Multi-Region Active-Active · Graceful Degradation: Fail Open or Closed · Build vs Buy Framework · Replication
Reading Guide#
Organized for interview use first, reference second. This page covers issuing, storing, revoking and recovering identity. Verifying a token at the edge — JWKS caching, local signature checks, stripping client-supplied identity headers — lives in API Gateway; this page links there rather than repeating it.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal Lens) and the appendices on token formats, key rotation and recovery |
What is an Identity System? — Why interviewers pick this topic
An identity system answers two questions for every other system in the company: who is this (authentication) and is this still true (session validity). It registers accounts, verifies credentials (passwords, passkeys, one-time codes, federated logins), issues tokens that other services trust, keeps track of which sessions exist, kills them on demand, and lets people who lost every credential get back in without letting an attacker do the same.
The hard part is not hashing a password or signing a JWT. The hard part is that every token you issue is a bearer credential that will be copied, logged, cached and stolen, and the company will one day ask: "How fast can we make every copy of that credential worthless — and what breaks for the other 200 million users while we do it?"
Before vs After — the "stolen refresh token" scenario:
Without revocation design:
t=0: Malware on a user's laptop copies the browser's 90-day refresh token.
t=+2h: Attacker replays it from another country. New access tokens issued.
t=+3h: User notices odd purchases. Clicks "Log out of all devices".
t=+3h: Logout deletes the user's sessions in one region's cache only.
t=+3h: Attacker's 1-hour access JWT keeps working — nobody checks it against anything.
t=+4h: Attacker refreshes again: refresh token is a stateless JWT too. Still valid.
t=+90 days: Refresh token finally expires. Support has 14 tickets. Trust gone.
With server-side sessions + rotation + revocation fan-out:
t=0: Malware copies refresh token RT1 (family F9, bound to device key).
t=+2h: Attacker presents RT1 without the device's proof-of-possession → rejected.
(Weaker variant: no binding — attacker refreshes, gets RT2.)
t=+2h05m: Real client presents RT1 (already rotated) → reuse detected → family F9 revoked.
t=+2h05m: session.revoked event published. Gateways' deny list updated in ≤ 30s.
t=+2h06m: Attacker's access token (10-min TTL) is denied by subject+session deny list.
t=+2h06m: User forced to re-authenticate with a passkey. Security email sent.
Why interviewers reach for this question: Identity is the purest test of blast-radius thinking. It is the one dependency every request has, so a candidate's choices decide whether a stolen credential is a 10-minute problem or a 90-day problem, and whether an identity outage is a degraded login page or a company-wide outage. The candidate who says "use JWTs, they're stateless" without saying how they would kill one has never been paged for an account takeover.
Mechanics Refresher: The Identity Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| Password + slow hash | Store argon2id(password, salt); verify by recomputing | Universal; no hardware needed | Phishable, reused across sites, 50–250ms CPU per verify by design |
| Passkeys / WebAuthn | Device holds a private key scoped to your domain; server stores the public key and verifies a signed challenge | Phishing-resistant; nothing reusable on the server | Device-loss and recovery story is now the weak point |
| OTP (SMS, email, TOTP) | Server sends or shares a short code; user echoes it | Cheap second factor | SMS: SIM swap, cost (~$0.01–0.05 per message), interception; all OTP is phishable |
| Server-side session | Opaque random ID → row in a session store | Instant revocation, small cookie | Lookup on every request (or on every refresh), store must be highly available |
| Signed token (JWT) | Claims + signature; anyone with the public key verifies locally | No lookup; ~50–500µs verify | Can't be un-issued; revocation lag = remaining lifetime |
| Access + refresh token pair | Short-lived access JWT (5–15 min) + long-lived refresh token checked server-side | Lookups only at refresh time | Two token types to secure; refresh endpoint is the hot path |
| Refresh token rotation | Each refresh returns a new refresh token; reuse of an old one revokes the family | Detects theft of a refresh token | Races on flaky mobile networks cause false reuse |
| Sender-constrained tokens | Token bound to a key the client must prove it holds (DPoP, mTLS) | Stolen token alone is useless | Client complexity; key storage on device |
| Federation (OIDC, SAML) | Another IdP authenticates; you trust its signed assertion | Enterprises keep control of their users | You inherit their outages and their misconfigurations |
| Workload identity (SPIFFE, cloud IAM) | Platform attests a workload and issues short-lived certs or tokens | No long-lived secrets in config | Platform becomes a tier-0 dependency |
For most production systems: passwords hashed with argon2id plus passkeys, short-lived signed access tokens (10 min), server-side refresh sessions with rotation and reuse detection, a revocation event stream that reaches every verifier in under a minute, and federation for enterprise customers. The primitives are not the interview — revocation speed and the identity service's blast radius are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
Authentication is not a crypto question. Nobody is asking you to implement ECDSA.
It is a revocation and blast-radius question that tests:
- Whether you can say how long a stolen credential stays useful, as a number, and what you'd do to shrink it
- Whether you know that every token you issue is a cache of a decision, and caches go stale
- Whether you design the identity service's failure — what still works when it is down — before its happy path
- Whether you treat account recovery as the real front door, because attackers do
The key insight: Every authentication design is a choice of where the "is this still valid?" check happens and how often. Check on every request and the identity store is in every request's critical path. Check never and a stolen token lives until it expires. Staff candidates pick the check frequency deliberately — per request, per refresh, per risk event — and state the revocation lag it buys.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws login form → auth service → users table → JWT | Asks "Consumer, workforce or service-to-service? What's the account-takeover threat, and how fast must 'log out everywhere' take effect?" | Asks "How many login systems does the company already have, who owns account recovery policy, and which regulator or customer contract sets our audit requirements?" |
| Tokens | "JWTs, so it's stateless and scales" | "10-minute access JWTs verified locally, server-side refresh sessions with rotation; revocation lag is bounded at 10 minutes, ≤ 30s for high-risk events via a pushed deny list" | Sets token lifetime and binding as an org-wide standard by risk tier; owns the deprecation of the 3 legacy token formats still in circulation |
| Failure | "Run several replicas of the auth service" | "If the IdP is down, verification keeps working on cached keys; new logins fail; refresh gets a bounded grace extension; step-up actions fail closed" | Treats identity as tier 0: cells per region, a written fail-static policy signed by security, quarterly IdP-down game days across every product team |
| Revocation | "Delete the session from the DB" | "Revocation is an event: session store write, then fan-out to every verifier's deny list, with revocation.propagation_p99 as an SLO" | Makes revocation latency a company security KPI and pushes it across the vendor boundary with shared-signals standards |
| Recovery | "Email a reset link" | "Recovery is the weakest login path, so it gets risk scoring, cooling-off periods, notifications to existing factors, and its own abuse metrics" | Decides who owns recovery policy — security sets the floor, product sets friction, support follows a script it can't override — and audits support-assisted recovery |
| Scale | "Shard the users table" | "Login is ~700/s at peak; refresh is ~25K/s; verification is ~1M/s but local. The hard capacity problem is credential-stuffing waves at 50K/s hitting a 100ms password hash" | Prices it: hashing capacity, SMS spend, IdP vendor per-MAU fees, and the headcount of a 24×7 identity on-call versus buying |
Why "tokens" separates levels
L5: "We'll issue a JWT on login with a 24-hour expiry. Every service verifies it locally, so there's no central bottleneck." This is a reasonable performance answer. It is also a 24-hour revocation lag that the candidate hasn't noticed. When asked "the user clicks 'log out everywhere'," the answer becomes "we'd add a blacklist," and the blacklist is now a per-request lookup — the bottleneck the JWT was supposed to avoid.
L6: "Access tokens are 10-minute JWTs, verified locally at the gateway as in the API gateway design. The refresh token is opaque and lives in a server-side session store, so refresh — about once every 10 minutes per active client — is where we check 'is this session still alive'. That bounds normal revocation lag at 10 minutes. For high-risk events — password change, takeover detection, admin kill — we push a deny entry keyed by session ID to every verifier within 30 seconds, and it expires after 10 minutes because by then the token has expired anyway. The deny list stays small: it only holds the last 10 minutes of revocations."
L7: "The real problem is that we have four token formats across three generations of products, two of which issue 30-day tokens nobody can revoke. I'd publish a token standard with lifetimes by risk tier, migrate the long-lived ones behind a refresh flow within two quarters, and measure 'maximum credential lifetime in circulation' as a security KPI."
Why "failure" separates levels
L5: "We'll run the auth service in three availability zones with a replicated database." True, and insufficient: the IdP's database is a single logical dependency, a bad deploy or a bad signing-key push hits all three zones at once, and the candidate hasn't said what every other service does when the IdP is unreachable.
L6: "I'll separate what the IdP is needed for. Verifying an existing access token: never needs the IdP — keys are cached. Refreshing a session: needs the session store, which I'll make regional and replicated; if it's unreachable, I extend existing sessions by up to 30 minutes rather than log out 60 million people at once. New logins: fail closed with a clear error, because I can't verify a password I can't read. Sensitive actions — password change, payout, admin — always fail closed."
L7: Recognizes that the IdP is the company's largest correlated-failure domain. "Every product's availability is capped by ours. I'd run identity as cells per region with no cross-region synchronous dependency, publish the fail-static policy so product teams build to it, and run an IdP-down game day every quarter where we turn off login in one region and watch what breaks."
Why "recovery" separates levels
L5: "Forgot password sends a reset link to the email on file, valid for 1 hour." Correct and incomplete. It ignores that the email account may itself be compromised, that support agents can be talked into changing the email, and that the reset link is a bearer credential sitting in an inbox.
L6: "Recovery is the login path with the weakest authentication, so it gets the strongest monitoring. Every recovery notifies all existing factors, high-value accounts get a 24–72 hour cooling-off before a new factor becomes trusted, and support-assisted recovery follows a script with no override. I track recovery.success_then_dispute_rate — the fraction of recoveries later reported as takeovers."
L7: "Recovery policy is a business decision dressed as an engineering one. Security wants a 72-hour delay; growth wants zero friction; support wants a button. I'd make the policy explicit by account tier, give security the veto on the floor, and audit every support-assisted recovery monthly."
Positions to Commit To#
| Position | Rationale |
|---|---|
| Short-lived access token + server-side refresh session | Bounds revocation lag at the access-token lifetime (10 min) while keeping the per-request path lookup-free |
| Revocation is an event with an SLO, not a DELETE | A row deleted in one store does nothing to tokens already cached at 400 gateway nodes; propagation must be pushed and measured |
| Rotate refresh tokens; treat reuse as theft | The only cheap way to detect a copied refresh token; tolerate a 30–60s grace for mobile retries |
| Verification never depends on the IdP being up | Cached signing keys mean an IdP outage degrades login, not every API call |
| Shed credential stuffing before the password hash | argon2id at ~100ms means 50K bad attempts/s would need ~5,000 cores; reject by IP, ASN, device and breached-credential signals first |
| Recovery gets the strongest controls, not the weakest | Attackers go through the weakest door; recovery is usually it |
| Signing keys rotate on a schedule you have rehearsed | The emergency rotation will happen under pressure; if the routine one isn't automated, the emergency one will be an outage |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Consumer login | 10M–1B accounts, account takeover, low-friction recovery, morning-ramp and attack spikes | Passwords + passkeys, short access tokens, server-side refresh sessions, risk engine, layered stuffing defenses | Credential stuffing, recovery takeover, IdP outage logs everyone out | Takeover rate per million logins, revocation lag ≤ 10 min, login availability 99.95% |
| Workforce SSO | 1K–200K employees, very high-value accounts, compliance audits, many SaaS apps | OIDC/SAML federation, phishing-resistant MFA mandatory, device posture, SCIM provisioning, short sessions for admin | Phished admin session, deprovisioning lag after termination | Terminated employee loses all access ≤ 15 min; every admin action attributable |
| Service-to-service identity | 10K–1M workloads, millions of calls/s, no humans in the loop | Platform-attested workload identity, mTLS with certs living hours, no static secrets | CA or attestation outage halts deploys; expired certs cause mass failure | Zero long-lived secrets in config; cert lifetime ≤ 24h; rotation fully automated |
🎯 Staff Move: "I'll design consumer login — that's where the scale and the account-takeover pressure are. I'll keep the token and session model general enough that workforce SSO is a federation adapter on top, and I'll treat service-to-service identity as a separate platform problem that shares the signing-key discipline but not the session store. Tell me if you'd rather go deep on workforce or service identity."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Stateless Tokens vs Server-Side Sessions | Revocation speed vs a lookup on every request; who carries the latency and who carries the stolen-token risk? |
| 2 | Token Lifetime: Short + Refresh vs Long-Lived | Every minute of lifetime is a minute a stolen token works; every refresh is load on the hottest endpoint you own |
| 3 | Fail Closed vs Fail Open When the IdP Is Unreachable | Lock out every user, or let possibly-revoked sessions keep working for a bounded time? Who signs off? |
| 4 | Central Authorization Service vs Policy in Each Service | One consistent answer to "can X do Y?" vs a new tier-0 dependency on every request |
| 5 | Build vs Buy the Identity Provider | Control over the most security-critical code you own vs a vendor whose outage and breach are now yours |
How Real Companies Built It#
Why this section belongs here: Identity is the area where public incident reports teach the most. Each of these shows a fault line from this page playing out in production.
Facebook — "View As" and Access Tokens at Scale (2018)#
In September 2018 Facebook disclosed that three bugs interacting in its "View As" feature let attackers obtain access tokens — which Facebook describes as the digital keys that keep people logged in — for other users; it reset tokens for almost 50 million affected accounts plus another 40 million that had been subject to a "View As" lookup in the previous year, and turned the feature off. Its follow-up said about 30 million people actually had tokens stolen (Facebook security update, Facebook follow-up).
The mitigation was mass token reset — logging ~90 million accounts out — because that was the revocation primitive available at that scale. The bug was in a product feature, not the identity service, yet the identity system carried the response.
Staff insight: When asked "how would you respond to mass token theft?", the answer is a pre-built, rate-limited, targeted bulk revocation tool: revoke sessions by cohort (feature, time window, client version), not "everyone". Facebook reset an extra 40 million accounts as a precaution because the blast radius was hard to bound. Design so you can bound it.
Okta — Support-System Breach and Session Tokens in HAR Files (2023)#
Okta's root-cause report says that between 28 September and 17 October 2023 a threat actor accessed files in its customer support system belonging to 134 customers; some were HAR files containing session tokens, which were used to hijack sessions of five customers. Access came through a service account whose credentials had been saved to an employee's personal Google account, signed into a personal profile on an Okta-managed laptop. Remediations included disabling that account, blocking personal Google profiles on managed devices, and binding admin session tokens to network location (Okta root cause and remediation).
Staff insight: A session token is a bearer credential, and it will leak through channels nobody designed for — debug files, logs, support tickets. Two defenses follow: bind high-value sessions to something the attacker doesn't have (device key, network context), and keep a revocation path that support and security can trigger in seconds. In an interview, say "bearer tokens leak sideways; binding and fast revocation are how we make a leak boring."
Microsoft — Storm-0558 and a Consumer Signing Key (2023)#
Microsoft's investigation reported that Storm-0558 used an acquired Microsoft account (MSA) consumer signing key to forge tokens for Outlook Web Access and Outlook.com. Its leading hypothesis was that a 2021 crash of the consumer signing system produced a crash dump that, through a race condition, contained the key, and that the dump was later reached through a compromised engineering account; enterprise mail accepted the consumer-signed tokens because validation libraries did not automatically enforce key scope (Microsoft Security Response Center).
Staff insight: Two lessons for this page. First, signing keys are the crown jewels — keep them in HSMs or KMS and design emergency rotation as a rehearsed drill. Second, verification must check that a key is allowed to sign this kind of token for this audience, not only that the signature is mathematically valid. Say both in the key-management deep dive.
Google — BeyondCorp and Zanzibar#
Google's BeyondCorp paper describes moving access decisions off the network perimeter: corporate applications are reached over the internet, and access depends on the user and the device rather than on being inside a privileged intranet (Google research). Separately, Google's Zanzibar paper describes a central authorization system storing trillions of access control lists and serving millions of authorization requests per second at a 95th-percentile latency under 10 ms and availability above 99.999% over three years of production use (Google research).
Staff insight: These are the two ends of fault line 4. BeyondCorp says the network is not an identity; every request carries user and device identity. Zanzibar says a central authorization service can be fast and available enough — if you are willing to build what Google built. In an interview, cite Zanzibar to show central authorization is possible, then say what it costs before recommending it.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "We'll use JWTs" | "User clicks 'log out of all devices'. How long until a stolen token stops working?" | Do you know your revocation lag as a number? |
| "We store sessions in Redis" | "Redis is down. What happens to 60 million logged-in users?" | Session store as a tier-0 dependency, degraded mode |
| "We hash passwords with bcrypt" | "40,000 login attempts per second from 200,000 IPs just started. What happens to your CPU?" | Credential stuffing economics, shedding before the hash |
| "We rotate refresh tokens" | "A mobile client on a bad network retries the refresh. Did you just log them out?" | Reuse-detection races, grace windows |
| "Forgot password emails a link" | "The attacker already owns the user's email. Now what?" | Recovery as the real front door |
| "We rotate signing keys" | "The private key was leaked an hour ago. Walk me through the next 30 minutes." | Emergency rotation, verifier key caches, mass re-login |
System Architecture Overview#
Reading the diagram: Three paths with three very different rates. Verification (~1M/s) happens at the gateway with cached keys and never touches the identity plane. Refresh (~25K/s at peak) is the hottest identity endpoint and the place where "is this session still alive?" is answered. Login (~700/s legitimate, up to 50K/s during an attack) is the most CPU-expensive path and the one under attack. Revocation flows the other way — from the session store, through the event bus, out to every verifier — and its latency is the number this whole design is judged by.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Token model | "JWT, stateless" | "10-min access JWT, verified locally. Opaque refresh token checked server-side. Revocation lag ≤ 10 min, ≤ 30s via deny list." |
| Logout everywhere | "Delete sessions" | "Revoke the refresh family, publish session.revoked, push a session-id deny entry with a 10-min TTL to every verifier." |
| IdP down | "Replicas" | "Verification continues on cached keys. New logins fail closed. Refresh gets a bounded grace. Step-up fails closed." |
| Stuffing | "Rate limit by IP" | "Layered: IP/ASN reputation, device signals, breached-password check, per-account limits — all before argon2id." |
| Recovery | "Email a reset link" | "Risk-scored, notifies all factors, cooling-off for high-value accounts, support follows a no-override script." |
| Keys | "Rotate yearly" | "Rotate every 30 days, automated, publish-before-sign; emergency rotation drilled quarterly." |
| Service identity | "API keys in env vars" | "Platform-attested workload identity, mTLS certs living ≤ 24h, no static secrets." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| argon2id / bcrypt verify cost | ~50–250ms of CPU per attempt (tuned) | 1,000 logins/s ≈ 100 cores; a 50K/s stuffing wave ≈ 5,000 cores — shed first |
| Local JWT verification | ~50–500µs with a cached key | Why verification never needs the IdP (see Edge Gateway) |
| Session store lookup | ~1–2ms p99 in-region (Redis or DynamoDB) | Affordable at refresh rate (25K/s), expensive at request rate (1M/s) |
| Access-token lifetime | 5–15 min typical; 10 min here | Equals worst-case revocation lag without a deny list |
| Refresh session lifetime | 30-day sliding, 90-day absolute (consumer) | How long a device stays signed in without a password |
| Deny-list propagation target | p99 ≤ 30s | Shrinks the high-risk revocation window from 10 min to seconds |
| JWKS cache TTL at verifiers | 5–15 min | Key rotation must publish the new key ≥ 1 TTL before signing with it |
| NIST SP 800-63B rate limit | ≤ 100 consecutive failed attempts per account before the authenticator is disabled (NIST SP 800-63B) | The ceiling, not the target; per-account limits are one layer of many |
| NIST SP 800-63B reauthentication at AAL2 | Overall timeout SHOULD be ≤ 24h, inactivity ≤ 1h (NIST SP 800-63B) | The workforce baseline; consumer sessions are usually longer by product choice |
| Workload certificate lifetime | 1–24 hours | Short enough that revocation lists are rarely needed |
| SMS OTP cost | ~$0.01–0.05 per message, higher in some countries | SMS pumping fraud turns your OTP endpoint into a cost attack |
| Password reset link lifetime | 15–60 min, single use | A reset link is a bearer credential sitting in an inbox |
Interview Walkthrough
The most common mistake: Candidates spend 20 minutes on the signup form, password hashing and the JWT payload, then run out of time before the interviewer asks the only question that matters: "A token was stolen. How long does it keep working, and what does it cost everyone else to stop it?" Compress the basics to ~10 minutes and spend the rest on revocation, the IdP's failure posture, credential stuffing and recovery.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"Users sign up, log in with a password or a passkey, optionally add a second factor, stay logged in on several devices, log out of one or all devices, and recover their account when they lose credentials. Every other service in the company trusts the identity we issue."
Then the non-functional requirements, which is where the design lives:
"Four constraints drive everything. One: a stolen credential must stop working fast — I'll target 10 minutes worst case and 30 seconds for high-risk revocations. Two: identity is in every request's path, so an identity outage must not become a company outage — verification can't depend on the identity service being up. Three: login is under constant attack, so the expensive part — password hashing — must be protected from credential-stuffing volume. Four: recovery is a login path too, and usually the weakest one. Scale: I'll assume 200 million accounts, 60 million daily actives, ~700 logins per second at the morning peak, and around a million authenticated API requests per second."
Then name the underspecified parts:
"A few things I'd confirm: is this consumer, workforce or service-to-service? Do we have enterprise customers who need SSO into our product? Are there regulated accounts — payments, health — that need stronger assurance? I'll assume consumer, with enterprise SSO as a later adapter."
🎯 Staff Move: Putting a number on revocation lag in the first three minutes — "10 minutes worst case, 30 seconds for high-risk" — tells the interviewer you know the token model is a revocation decision, not a performance decision. Everything you draw next can be judged against that number.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Account:
account_id,status(active, locked, recovering, deleted),assurance_tier,created_at - Credential:
credential_id,account_id,type(password, passkey, totp, phone, email, federated),secret_ref(hash or public key),added_at,trusted_after(cooling-off),last_used_at - Session (refresh family):
session_id,account_id,device_id,family_id,current_refresh_hash,prev_refresh_hash,rotated_at,auth_methods(pwd,passkey,otp),auth_time,expires_at,absolute_expires_at,revoked_at,binding_key_thumbprint - SigningKey:
kid,alg,status(pending, active, retiring, revoked),published_at,activated_at,retire_after - AuditEvent:
event_id,account_id,type,actor(user, support agent, system),ip,device,at
Client-facing API:
POST /v1/login { identifier, password | passkey_assertion, device_info }
→ 200 { access_token (10 min), refresh_token, expires_in }
→ 401 | 429 | 200 { step_up_required: ["otp","passkey"], challenge_id }
POST /v1/token/refresh { refresh_token } DPoP: <proof signed by device key>
→ 200 { access_token, refresh_token (new) }
→ 401 { error: "session_revoked" | "reuse_detected" }
POST /v1/sessions/{id}/revoke (log out one device)
POST /v1/sessions/revoke_all (log out everywhere, keeps current)
POST /v1/recovery/start { identifier }
POST /v1/recovery/complete { recovery_token, new_credential }
GET /.well-known/jwks.json (public keys, cached by verifiers)
Login returns either tokens or a step-up challenge; the risk engine decides which. Refresh is the only endpoint that touches the session store on the hot path.
🎯 Staff Move: "I'm splitting the session from the token. The session is a row I can revoke. The access token is a 10-minute cache of 'this session was valid at issue time'. Once you see the token as a cache entry, its lifetime is just a TTL, and every cache-invalidation tool applies."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk the login flow in 90 seconds:
- Client submits credentials. The edge drops requests from known-bad IPs, ASNs and device fingerprints and applies per-IP and per-identifier rate limits — before any hashing.
- Login service checks the identifier exists and isn't locked, then verifies the password with argon2id (~100ms) or a passkey assertion (~1ms signature verify).
- Risk engine scores the attempt: new device, impossible travel, breached-password match, velocity. Outcome: allow, step-up, or block.
- Token service creates a session row (refresh family), stores a hash of the refresh token, and signs a 10-minute access JWT with the active key in KMS.
- Client calls APIs with the access JWT; the gateway verifies it locally with cached public keys — no identity call.
- Every ~10 minutes the client refreshes: the token service looks up the session, checks it isn't revoked, rotates the refresh token, issues a new access token.
- Logout or a risk event revokes the session row and publishes
session.revoked; the distributor pushes the session ID to every verifier's deny list.
🎯 Staff Move: Say out loud: "Notice the three rates: verification at a million per second never touches identity; refresh at 25,000 per second is where revocation is enforced; login at 700 per second is the expensive one and the one under attack. I'm going to design each path to its own rate." You've now spent ~9 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes identity hard is what happens after a credential leaks and when the identity service fails. I'd like to go deep on four things: how fast we can revoke and how that reaches every verifier, what the rest of the company does when identity is down, how we survive a credential-stuffing wave without melting the password hashers, and how account recovery avoids becoming the takeover path. Where would you like to start?"
If no preference: start with revocation. It's the question that decides the level.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: Revocation (7–8 min)
"There are three revocation speeds and I'll be explicit about which events get which. Routine logout: revoke the session row — the access token dies at its natural expiry, at most 10 minutes later. High-risk events — password change, 'log out everywhere', takeover detected, admin kill: revoke the row and push the session ID to a deny list held in memory at every gateway node, with a 10-minute TTL, because after 10 minutes the token has expired anyway. Mass events — signing-key compromise: rotate the key, which invalidates every token signed with it."
Quantify: "High-risk revocations are maybe 0.5% of 60 million daily actives — 300,000 per day, 3.5 per second average. With a 10-minute TTL the deny list holds about 2,100 entries in steady state, under 100 KB. During a mass incident that could reach a few million entries — still under 200 MB, which a gateway can hold, but I'd switch to revoking by auth_time cutoff per account instead of listing session IDs."
Who pays: "The user pays up to 10 minutes of exposure for routine logouts. Security signed off on that; for anything they care about, the push path pays 30 seconds. The gateway team pays a small in-memory set and a subscription to the revocation stream."
Deep dive 2: The IdP is down (6–7 min)
"I'll split the dependency by operation. Verification: independent — cached JWKS. Refresh: depends on the session store; if the token service can't reach it, it issues a short extension — a 10-minute access token based on the still-valid refresh token's claims — for up to 30 minutes, but only if the client's refresh token is cryptographically valid and the session isn't on the local deny list. Login: fails closed. Step-up and sensitive actions: fail closed. That's a written policy security signs, not an on-call judgment call."
Quantify: "Without the grace, a 30-minute session-store outage expires every active access token within 10 minutes — 15 to 20 million concurrently active users logged out at once, then all of them retrying login when it comes back: 20 million logins in a few minutes against a 700/s design, or 30,000+ per second of argon2id. The grace prevents the outage from becoming a login stampede."
Deep dive 3: Credential stuffing (6–7 min)
"Stuffing is economics. Attackers try breached email/password pairs at volume. If every attempt reaches argon2id at 100ms, a 50,000/s wave needs 5,000 cores — the attack is a denial of service on login before it's a takeover. So the layers: IP and ASN reputation and per-IP rate limits at the edge; device and client-integrity signals; per-identifier limits — 10 failures per 15 minutes then step-up; a breached-credential check so a correct-but-breached password triggers a reset instead of a login; and the hash budget itself is a bulkhead: login gets a fixed pool, and when it's saturated we serve challenges, not timeouts."
Who pays: "Legitimate users behind shared IPs — mobile carriers, universities — pay extra challenges during waves. Product signs off on a challenge rate ceiling, say 2% of good logins."
Deep dive 4: Recovery (5–6 min)
"Recovery authenticates someone who has, by definition, lost their credentials, so it's the weakest proof we accept. Rules: notify every existing factor on start and completion; for accounts with a passkey or a high-value tier, a cooling-off of 24–72 hours before the new credential can change payout details or remove other factors; reset tokens are single-use, 15 minutes, bound to the browser that started the flow; and support-assisted recovery follows a scripted identity-proofing flow — agents can't skip steps or change contact details directly."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: a token is a cache of an identity decision, so I picked the cache TTL — 10 minutes — and built a push path for the cases where 10 minutes is too long. Verification never depends on the identity service, so its outage degrades login, not the company. Login is protected economically, with shedding before the expensive hash, and recovery is treated as the front door it really is."
The evolution closer:
"What I'd build later: device-bound tokens via DPoP for all first-party clients; passkeys as the default credential, which removes most of the stuffing problem; cross-vendor revocation signals for enterprise customers through the OpenID shared-signals standards; and identity cells per region. What I'd not build: our own HSMs — cloud KMS is the right answer until a regulator says otherwise."
🎯 Staff Move: End on the revocation number and who owns it. Senior candidates end with "and we'd add MFA." Staff candidates end with "and
revocation.propagation_p99is an SLO the identity team carries, reviewed with security every month, because it's the number that decides how bad the next token theft is."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Crypto tour | Explains RSA vs ECDSA and JWT header fields for 8 minutes | One sentence: "ES256 or EdDSA, keys in KMS, verified locally" |
| Signup form design | Designs email verification and username rules in detail | Names the entities in 30s and moves to sessions |
| OAuth flow recital | Draws every redirect of the authorization code flow | "Authorization code + PKCE per the OAuth security BCP; I'll skip the redirects unless you want them" |
| No revocation number | Waits for "how do you log someone out?" | States revocation lag as a requirement in Phase 1 |
| No failure posture | Says "highly available" | Splits verify / refresh / login / step-up and gives each a posture |
| Recovery as afterthought | "Forgot password sends an email" at minute 44 | Brings recovery into the deep dives as a login path |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Identity is the dependency nobody can route around. A cache outage costs latency; a queue outage costs freshness; an identity outage costs every authenticated request in the company. At the same time it is the system attackers target most, because a single credential unlocks everything that identity guards. That combination — maximal blast radius on failure, maximal value on compromise — forces tradeoffs where both sides are expensive, and the candidate has to say who pays.
It also has an unusually deceptive happy path. A Senior engineer can build a login that works for every honest user on day one. The design is judged entirely by the dishonest ones: the credential stuffer, the malware that copies a refresh token, the caller who convinces a support agent they are the account owner. None of those show up in a demo. Staff candidates design for them before drawing the first box.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"An attacker has a copy of one user's refresh token. Walk me through every place that token, and the access tokens derived from it, are accepted — and how long each keeps accepting it after the user hits 'log out everywhere'."
A candidate who answers with the refresh path (session store, revoked immediately), the access-token path (gateway, up to 10 minutes or ~30s via deny list), the downstream caches (services that cached the identity, bounded by the same TTL), the detection mechanism (rotation reuse, risk signals) and the owner (identity on-call for propagation, trust and safety for the account) has operated an identity system. A candidate who says "we delete the session" has built a login form.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Consumer login. The population is huge, mostly honest, and reuses passwords. The threat model is scale: attackers replay billions of leaked credential pairs, buy stolen session cookies, and socially engineer support. The design centers on cheap verification, aggressive abuse shedding, risk-based step-up, and recovery that is easy for real owners and hard for impostors. Product friction is a real cost here — every extra challenge loses a measurable fraction of logins — so security and growth negotiate every threshold. Correctness bar: account takeovers per million logins, revocation lag, login success rate.
Workforce SSO. The population is small (thousands to a few hundred thousand) but each account is worth far more: one administrator session can reach production, customer data or the identity system itself. The design centers on federation (OIDC and SAML to hundreds of SaaS apps), phishing-resistant MFA as a hard requirement, device posture, short sessions for privileged roles, and deprovisioning: a terminated employee must lose access everywhere within minutes, which means pushing revocation into apps you don't own. Correctness bar: time-to-deprovision, percentage of apps behind SSO, audit completeness.
Service-to-service identity. No humans, no passwords, no recovery flow. Millions of calls per second between workloads that are created and destroyed constantly. The design centers on the platform attesting what a workload is (which cluster, namespace, service account) and issuing short-lived credentials — X.509 certificates for mTLS or signed tokens — that rotate automatically. Standards like SPIFFE define the identity format and how workloads fetch short-lived identity documents (SPIFFE overview). Correctness bar: zero long-lived secrets, rotation success rate, and no deploy blocked by the identity platform. This intent pairs with Service Registry and the workload model in Kubernetes.
🎯 Staff Move: "These share almost nothing operationally. Consumer is about abuse at scale, workforce is about high-value accounts and deprovisioning, service identity is about automation and rotation. If someone asks me to build one system for all three, I'd share the key-management and audit layers and nothing else."
2.2 When NOT to Build Your Own Identity Provider#
| Situation | What to Do Instead | Why |
|---|---|---|
| Fewer than ~1M accounts, no unusual assurance needs | Buy a hosted CIAM product | You'd spend 3–5 engineers re-implementing password storage, MFA, recovery and federation that a vendor has hardened |
| Workforce identity for your own employees | Buy a workforce IdP | The SaaS app catalogue, SCIM connectors and device-posture integrations are the product; you won't out-build them |
| You need "login with X" only | Use the social IdP via OIDC; keep a thin local account record | Federation shifts credential risk to a provider with a bigger security team |
| Service-to-service identity on one cloud | Use the cloud's workload identity and IAM | The platform already attests workloads; a parallel CA is a second tier-0 system |
| Authorization rules that fit in roles and a few attributes | Policy library in each service, not a central authorization service | A Zanzibar-style service is worth it only when relationships (sharing, folders, orgs) dominate |
And within the design, some features you should not build even when you own the IdP:
- Don't invent a token format. Use JWT/JWS or opaque tokens per the OAuth and OIDC specs. Custom formats lose the libraries, the audits and the interoperability.
- Don't build the resource owner password credentials grant. The OAuth 2.0 Security Best Current Practice says it MUST NOT be used (RFC 9700).
- Don't store long-lived secrets for services when the platform can attest them.
- Don't run your own HSM fleet unless a regulator requires it; cloud KMS gives non-exportable keys with audit logs.
The Staff signal is knowing that identity is the textbook case where buying is the default and building needs a reason — scale economics past ~50–100M MAU, unusual assurance or data-residency requirements, or identity being the product itself. See Buy or Build: The Total-Cost Test and fault line 5.
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why It Matters | What to Say |
|---|---|---|
| Which intent | Consumer, workforce and service identity are different systems | "I'll assume consumer and treat enterprise SSO as an adapter." |
| Revocation requirement | Determines token model and lifetime | "I'll target 10 min worst case, 30s for high-risk revocations." |
| Assurance levels | Payments or health data may need step-up and phishing-resistant MFA | "I'll design step-up as a first-class session attribute: auth_time and amr." |
| Number of products | One login for many apps changes the session model to SSO | "If there are several products, I'd centralize the session and issue per-audience tokens." |
| Regions and residency | EU or other residency rules decide where credential data lives | "Home-region per account, with credentials stored only there." |
| Who owns recovery policy | Security, product and support each want different friction | "I'll propose tiers and name security as the owner of the floor." |
| Existing identity systems | Migration from a legacy hash or token format is often the real project | "Is there an existing user base with legacy password hashes I need to migrate on login?" |
2.4 Precise Terminology#
| Term | Meaning | Common Confusion |
|---|---|---|
| Authentication (authN) | Proving who the caller is | Conflated with authorization; most "auth service" designs mix both |
| Authorization (authZ) | Deciding whether that caller may do this action on this resource | Pushed into the IdP as giant role claims that go stale |
| IdP / Authorization Server | The system that authenticates users and issues tokens | Treated as the same thing as the gateway that verifies tokens |
| Access token | Short-lived credential presented to APIs | Assumed to be revocable instantly |
| Refresh token | Long-lived credential used only at the token endpoint to get new access tokens | Sent to APIs, logged, or given the same lifetime as the session |
| ID token | OIDC token that tells the client who logged in | Used as an API access token |
| Session | Server-side record that a user authenticated on a device | Assumed to exist when tokens are purely stateless |
| Revocation lag | Time from "revoke" to the last moment a copy of the credential is accepted anywhere | Measured at the store write, not at the verifiers |
| Step-up | Re-authenticating with a stronger method before a sensitive action | Implemented as a client-side check |
auth_time / amr / acr | When, how and at what assurance the user last authenticated | Ignored, so a 30-day session can change a password |
| Sender-constrained token | Token usable only by the holder of a bound key (DPoP, mTLS) | Assumed to be standard; most bearer tokens are not bound |
| Credential stuffing | Replaying breached username/password pairs from other sites | Confused with brute force; the passwords are often correct |
| ATO | Account takeover — an attacker controlling a real user's account | Measured only when users complain |
3. Where the Design Splits#
Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.
3.1 Fault Line 1: Stateless Tokens vs Server-Side Sessions#
The tension: A server-side session is revocable the instant you delete it, but every request needs a lookup. A signed token needs no lookup, but once issued it is valid until it expires, wherever it has been copied.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Opaque session ID, lookup on every request | Instant revocation; tiny cookie; simple mental model | Session store in every request's critical path: 1M lookups/s, +1–2ms each, store outage = company outage | Every product team (latency, availability); the session-store on-call |
| Long-lived stateless JWT (hours to days) | Zero lookups; trivially scalable verification | Revocation lag = remaining lifetime; "log out everywhere" is a lie | Users whose tokens are stolen; support; security |
| Short JWT + server-side refresh session (hybrid) | Lookups only at refresh (~1/40th of request rate); revocation lag bounded by access TTL | Two token types; refresh endpoint is a hot path; lag is still minutes without a push path | Identity team (more moving parts); users carry ≤ 10 min exposure |
| Hybrid + pushed deny list | High-risk revocations in seconds without per-request lookups | Distribution pipeline to every verifier; must be monitored | Gateway/platform team holds the deny list; identity owns propagation SLO |
| Opaque token + introspection with short cache | Central control; easy revocation | Introspection traffic and latency; cache TTL is the new lag | Every service's latency; the IdP's capacity |
The Staff default: hybrid plus deny list. 10-minute access JWTs verified locally (the mechanics are in Edge Gateway), opaque refresh tokens checked against a server-side session store at refresh time, and a pushed deny list for high-risk revocations. The arithmetic: at 1M API requests/s and one refresh per active client per 10 minutes, the session store sees ~25K reads/s instead of 1M — a 40× reduction — and revocation lag is 10 minutes worst case, ~30 seconds for events that matter.
"I want the session store in the refresh path, not the request path. That's a 40× cut in dependency load, and I buy back revocation speed with a push channel for the 0.5% of revocations where minutes matter."
When to deviate:
- Low-volume, high-assurance apps (admin consoles, banking back-office): opaque sessions with per-request lookup. 200 requests/s doesn't need a token cache, and instant revocation is worth 2ms.
- Server-rendered web apps on one backend: a classic cookie session in Redis is simpler and fine; JWTs add nothing when the verifier and the session store sit together.
- Offline-capable clients: longer access tokens are sometimes unavoidable; compensate with device binding.
3.2 Fault Line 2: Token Lifetime — Short + Refresh vs Long-Lived#
The tension: Every minute of access-token lifetime is a minute a stolen copy works. Every minute you remove is refresh load and a chance for a flaky mobile network to log someone out.
| Access TTL | Refresh load (60M DAU, ~2h active/day) | Worst-case revocation lag | Notes |
|---|---|---|---|
| 1 min | ~7.2B/day, ~250K/s peak | 1 min | Refresh becomes the biggest service you run; battery and data cost on mobile |
| 5 min | ~1.4B/day, ~50K/s peak | 5 min | Reasonable for high-risk apps |
| 10 min | ~720M/day, ~25K/s peak | 10 min | Default; matches common gateway JWKS cache TTLs |
| 60 min | ~120M/day, ~4K/s peak | 60 min | Needs the deny list for anything sensitive |
| 24 h | ~60M/day | 24 h | Effectively unrevocable; avoid |
Refresh-token lifetime is a separate decision: how long a device stays logged in without re-authenticating. Consumer default here: 30-day sliding, 90-day absolute, with step-up for sensitive actions based on auth_time, not on whether a session exists.
Rotation and reuse detection. Every refresh returns a new refresh token and invalidates the old one. If an old token is presented again, either the client retried or someone else has a copy — the server can't tell which, so it revokes the whole family. The OAuth security BCP requires refresh tokens for public clients to be either sender-constrained or rotated (RFC 9700).
The race: a mobile client sends RT1, the server rotates to RT2, the response is lost, the client retries with RT1. Without a grace window, that's a false theft signal and a logged-out user. The fix: accept the immediately previous token for 30–60 seconds from the same device binding and return the same RT2 (idempotent rotation). Outside that window, reuse revokes the family.
Who pays: short access TTLs are paid by the identity team (refresh capacity) and mobile users (battery, a few hundred bytes per refresh). Long TTLs are paid by the users whose tokens are stolen. Rotation is paid by users on bad networks if the grace window is wrong — watch refresh.reuse_detected_total split by client version; a spike after a mobile release is a client bug, not an attack.
The Staff default: 10-minute access, rotated refresh with a 60-second idempotent grace, family revocation on reuse, DPoP binding for first-party clients when the client platform supports secure key storage (RFC 9449).
3.3 Fault Line 3: Fail Closed vs Fail Open When the IdP Is Unreachable#
The tension: If the identity plane is down and you fail closed, nobody can do anything — the identity outage becomes a company outage. If you fail open, you might accept a session that was revoked during the outage. See the general framework in Graceful Degradation: Fail Open or Closed.
The mistake is answering this once for the whole system. The Staff answer splits it by operation:
| Posture | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Fail closed everywhere | Never accepts a revoked session | IdP outage = total outage; recovery causes a login stampede | Every product, every user, revenue |
| Fail open everywhere | Products stay up | Attackers with revoked or expired credentials walk in; sensitive actions unprotected | Security, users whose accounts are compromised |
| Split by operation (Staff default) | Existing users keep working ≤ 30 min; new logins and sensitive actions blocked | Revocations made during the outage take effect only via deny list or after recovery | Users who were revoked during the outage get up to 30 min extra exposure; security signs that off |
The grace extension, precisely: the token service (or a regional fallback signer) can verify the refresh token's own signature or MAC without the session store if refresh tokens are self-verifiable — for example an opaque random value plus an HMAC over session_id and expiry. It issues a 10-minute access token marked grace=true, repeats at most 3 times (30 minutes), and never for sessions on the local deny list. Services that care (payments, admin) reject grace=true tokens for sensitive operations.
Who signs off: security owns the maximum grace duration; product owns which features accept grace=true. It's a signed table, reviewed twice a year, not a runbook improvisation at 3 a.m.
🎯 Staff Move: "I'd rather let a user whose session was revoked during a 20-minute outage keep browsing for 20 minutes than log 20 million people out and then take 20 million logins in five minutes when we come back. But password changes and payouts fail closed, and security signs the table that says so."
3.4 Fault Line 4: Central Authorization Service vs Policy in Each Service#
The tension: Authentication answers "who". Authorization answers "may they do this to that". A central authorization service gives one consistent answer and one audit trail, and becomes a dependency of every request. Policy inside each service is fast and independent, and drifts.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Roles as claims in the token | Zero lookups; simple | Stale for the token lifetime; token bloat (a user in 400 groups); can't express "can read doc 9001" | Security (stale permissions), users with huge tokens |
| Policy library in each service | Fast; no new dependency; service owns its data | N implementations drift; audit across services is hard | Security and compliance (inconsistent rules) |
| Central policy engine, local evaluation | One policy language and repo; decisions computed in-process from pushed policy and data | Data for decisions must be replicated to every service; policy push is a deploy | Platform team (distribution); service teams (adoption) |
| Central relationship-based service (Zanzibar-style) | Consistent answers for sharing graphs; one audit trail; handles "folder → doc → user" | New tier-0 dependency; needs caching and consistency tokens to be fast and correct | Platform team (huge build); every caller's latency budget |
The Staff default: keep authentication central and authorization close to the data. Tokens carry identity and a small, stable set of coarse claims (tenant, account tier, a few scopes) — not permissions. The gateway enforces coarse checks; services enforce resource-level decisions with a shared policy library and a common decision-log format. Move to a central relationship service only when the product's permission model is genuinely a graph (document sharing, nested orgs) and several services need the same answers — that's what Zanzibar was built for, at a cost most companies shouldn't pay first.
When to deviate: collaborative products with sharing (docs, drives, design tools) hit the graph problem early; for them a central relationship service is the right Year-1 investment. See Collaborative Documents for the product side.
3.5 Fault Line 5: Build vs Buy the Identity Provider#
The tension: Identity is the most security-critical code you'll run, which argues for buying it from a specialist. It's also in every request and every product decision about signup friction, which argues for owning it. See Build vs Buy Framework.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Buy a hosted CIAM / workforce IdP | Hardened MFA, federation, compliance certifications, fast start | Per-MAU pricing at scale; their outage and their breach are yours; limited control of login UX and risk logic | Finance (fees); you inherit vendor incidents |
| Open-source IdP, self-hosted | Control, no per-user fees, standard protocols | You run a tier-0 system: upgrades, CVEs, scaling, on-call | Your platform team: 3–6 engineers plus a 24×7 rotation |
| Build in-house | Full control of UX, risk, data model, cost at 100M+ MAU | Years of security work; every protocol edge case is yours | 10–30 engineers plus security review; the opportunity cost |
| Hybrid: buy workforce, build or self-host consumer | Matches each intent's economics | Two systems to integrate and audit | Security (two control planes) |
The Staff default: buy workforce identity, always. For consumer identity, buy below ~10M MAU, decide on the numbers between 10M and 100M, and build or self-host above ~100M MAU or when login UX and risk logic are core to the product. Whatever you choose, own the session and revocation layer's interface: products should depend on your token contract, not the vendor's SDK, so a vendor change is a migration, not a rewrite.
The vendor-incident caveat. Buying moves the work, not the risk. The Okta support-system incident above shows a vendor breach reaching customer sessions. If you buy, you still need: your own revocation path, your own monitoring of admin sessions, and a contract that says how fast the vendor notifies you.
4. When It Breaks#
4.1 Signing-Key Leak — The Emergency Rotation#
t=0: Security finds the active token-signing private key (kid=k42) in a debug
artifact uploaded to a third-party ticketing tool 6 hours ago.
t=+5min: Incident declared. Assume forged tokens for any user, any audience.
t=+10min: Generate k43 in KMS (non-exportable). Publish k43 in JWKS alongside k42.
t=+12min: Mark k42 "revoked" in the verifier config channel — pushed, not waiting
for JWKS TTL. Gateways reject kid=k42 within 60s.
t=+13min: Every access token signed by k42 now fails: ~18M active clients get 401.
t=+13min: Clients refresh. Refresh tokens are opaque and checked against the session
store — unaffected by the JWT key — so refresh succeeds with k43 tokens.
t=+15min: Refresh traffic spikes from 25K/s to ~180K/s. Token service autoscaled from
pre-warmed capacity, plus jittered client retry (0–60s).
t=+25min: Refresh traffic back under 40K/s. 0.4% of clients failed and fell back to login.
t=+2h: Audit: tokens signed with k42 after the leak window carrying unusual audiences
or sessions that don't exist in the session store = forged. 0 found.
Why it went this well: refresh tokens were not signed by the same key as access tokens, so a JWT-key leak didn't force a password re-login for 60M users. Verifiers honored a pushed key-revocation signal instead of waiting for JWKS cache expiry. Refresh capacity was sized for a "everyone refreshes at once" event.
Detection: tokens.verified_with_unknown_session_total (a valid signature on a session that doesn't exist is a forgery signal), secrets scanning on outbound artifacts, KMS Sign call volume by caller.
Prevention: keys in KMS or HSM, never exportable; signing only through a narrow signer service; per-audience or per-token-type keys so one leak doesn't forge everything; verifiers enforce iss, aud and key scope, not only signature validity — the lesson Microsoft published after Storm-0558.
Owner: identity team (rotation); security (incident command); gateway platform (verifier key config).
4.2 IdP Outage Locks Every Service#
t=0: Session-store primary in us-east fails over; replica promotion stalls
(replication lag 40s, failover guard waits). Refresh calls time out.
t=+1min: refresh.error_rate 0.1% → 92%. Access tokens keep working (local verify).
t=+10min: Without grace: first wave of 10-min access tokens expires. Users see 401s.
Apps call refresh in a tight loop: 25K/s → 400K/s retries.
t=+12min: Token service pods CPU-pinned by retries. Login also shares the pool.
t=+15min: "Can't log in" trends on social media. Every product reports outage.
t=+25min: Session store recovers. 15M clients all fall back to full login.
t=+26min: Login at 40K/s → argon2id pool saturated → login p99 30s → more retries.
t=+70min: Stable after login admission control and client backoff hotfix.
With the Staff design:
t=+1min: Refresh fails → token service issues grace tokens (grace=true, 10 min, max 3).
t=+10min: Users notice nothing. Payouts and password changes show "temporarily unavailable".
t=+25min: Session store back. Grace tokens refresh normally. No login stampede.
Detection: refresh.error_rate, refresh.grace_issued_total, session_store.replication_lag_seconds, login.queue_depth.
Mitigation: grace extension (fault line 3); separate thread pools for refresh and login so a refresh storm can't starve login; clients with exponential backoff and jitter on refresh, enforced in the shared client SDK; Retry-After honored.
Prevention: session store replicated within the region with automatic failover tested monthly; regional cells so one region's session store never affects another (Multi-Region Active-Active); quarterly IdP-down game day.
Owner: identity on-call (primary); client platform team owns SDK backoff behavior.
4.3 Credential-Stuffing Wave#
t=0: 07:40 local, morning ramp: 600 legit logins/s.
t=+1min: Attack starts: 45K login attempts/s from ~220K residential proxy IPs,
each IP sending < 1 attempt/min. Per-IP limits don't trigger.
t=+2min: argon2id pool (2,000 cores, ~20K hashes/s) saturated. Legit login p99 → 12s.
t=+3min: Page: login.attempts_rate 75× baseline, login.success_rate 94% → 3%.
t=+5min: On-call enables wave mode: unknown-device attempts get a challenge before
hashing; identifiers seen in breach corpus with no prior device → step-up.
t=+8min: Hash pool load drops to 30%. Legit success rate back to 88% (challenge friction).
t=+30min: ~0.6% of attempted pairs were valid. Those accounts: forced reset + notification,
sessions created during the wave revoked.
t=+2h: Attack stops. Challenge rate decays back to 0.3%.
Detection: login.attempts_rate vs baseline, login.success_rate (a sharp drop is the stuffing signature — most pairs are wrong), login.unique_identifiers_per_min, ratio of failures on non-existent identifiers.
Mitigation: layered shedding before the hash: client-integrity and device signals, ASN and proxy reputation, per-identifier limits, challenges for unknown devices; a fixed hash budget as a bulkhead so overload produces challenges, not timeouts.
Prevention: breached-password screening at signup and at login (NIST SP 800-63B requires checking new passwords against a blocklist of compromised ones); passkey adoption, which removes the password entirely; notify users on new-device logins.
Owner: identity on-call for capacity; trust and safety for the account remediation; product signs off on the challenge-rate ceiling. Rate-limiter mechanics: Rate Limiter.
4.4 Token Cache Serving Revoked Sessions#
t=0: A team adds a 1-hour in-process cache of "token → user profile + session valid"
to cut calls to a profile service. Nobody reviews it as an auth change.
t=+3 weeks: User reports takeover. Support runs "log out everywhere" at 14:02.
t=+3 weeks: Gateway deny list blocks the session at 14:02:20. But the internal service
behind a second ingress path (partner API) uses its own cache: still serves
the attacker until 15:02.
t=+3 weeks: Attacker exports 1,200 contacts in that hour.
Detection: this is a silent failure — nothing errors. Catch it with a synthetic canary: every 5 minutes, create a session, revoke it, then probe every ingress path with its access token; alert on any 2xx after revocation_time + 60s. Metric: revocation.canary_accept_after_revoke_total.
Mitigation: purge the cache; subscribe that service to session.revoked.
Prevention: a standard: any cache of identity or session validity MUST have TTL ≤ access-token lifetime and MUST subscribe to revocation events; the canary covers every ingress, not just the main gateway. Generic cache-invalidation tradeoffs: Distributed Cache.
Owner: identity team owns the canary and the standard; the service team owns the fix.
4.5 Account Recovery as the Takeover Path#
t=0: Attacker controls victim's email via a separate breach.
t=+2min: Requests password reset. Link arrives in the victim's (now attacker's) inbox.
t=+5min: Old design: password changed, all sessions revoked — including the real owner's.
Attacker removes the passkey and adds their own phone as 2FA.
t=+1 day: Owner locked out of their own account. Support can't verify them because
every factor on file now belongs to the attacker.
With cooling-off:
t=+5min: New password works, but account enters limited mode for 24–72h. The existing
passkey device receives "Someone is recovering your account — this wasn't me".
t=+40min: Owner taps "this wasn't me". Recovery cancelled, attacker sessions revoked,
email flagged as compromised.
Detection: recovery.started_rate, recovery.cancelled_by_owner_rate, recovery.success_then_dispute_rate (recoveries later reported as takeovers), support-assisted recoveries per agent per day.
Mitigation: freeze the account, revoke all sessions, restore from the credential history in the audit log.
Prevention: cooling-off and limited mode for high-value tiers; never let recovery remove existing factors immediately; support tooling that can't change contact details without the scripted flow; encourage two independent recovery factors.
Owner: recovery policy — security sets the floor, product sets friction above it; recovery system — identity team; support-assisted recovery — support operations, audited monthly by security.
4.6 Key Rotation Timeline (Routine)#
Routine rotation is the drill that makes 4.1 survivable. The verifier side — caching JWKS and refreshing on an unknown kid — is covered in Edge Gateway; the issuer side is here.
Cadence: every 30 days, fully automated, with jwks.verifiers_missing_active_kid (verifiers that fetched JWKS but don't have the active key) as the pre-flight check. An emergency rotation is the same pipeline with the waits removed and a push signal added.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Signing-key leak | Secrets scanning hit; tokens.verified_with_unknown_session_total > 0 | Every token signed by that key | Emergency rotation, pushed key revocation, mass refresh | Identity + security incident command |
| Session-store outage | refresh.error_rate > 5% | All refreshes in the region | Grace extension ≤ 30 min, separate pools | Identity on-call |
| Credential stuffing | login.attempts_rate > 10× baseline, login.success_rate drop | Login latency for everyone; ATO for reused passwords | Wave mode, challenges before hashing | Identity on-call + trust and safety |
| Revoked session still accepted | revocation.canary_accept_after_revoke_total > 0 | Every user relying on logout | Purge rogue cache, subscribe to events | Identity (canary) + owning service |
| Revocation pipeline stalled | revocation.propagation_p99 > 60s, consumer lag | High-risk revocations slow to 10 min | Restart distributor; fallback to shorter access TTL | Identity + gateway platform |
| Recovery takeover | recovery.success_then_dispute_rate above 0.5% | Individual high-value accounts | Freeze, restore credentials from audit history | Security (policy), identity (system) |
| Refresh-reuse false positives | refresh.reuse_detected_total by client version | Users of one app release logged out | Widen grace, hotfix client | Identity + mobile team |
| SMS pumping fraud | otp.sms_sent by destination country vs baseline | Cost: tens of $K/day | Country allowlists, per-number limits, prefer TOTP/passkeys | Identity + finance |
| Workload CA outage | svid.rotation_failures, cert expiry horizon < 2h | Every deploy and service call as certs expire | Long enough cert lifetime to ride out CA outages (≥ 2× MTTR) | Platform security |
🎯 Staff Insight: The most dangerous identity failure is the one that doesn't error: a revoked session that keeps working. Every other row here pages. That one needs a synthetic canary that revokes a session every 5 minutes and checks every ingress path — otherwise you discover it from a user's takeover report.
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | Lists features: signup, login, MFA, logout | Names consumer vs workforce vs service identity; commits; states revocation lag as a requirement | Asks how many identity systems exist today, who owns recovery policy, and what audit regime applies |
| Token model | JWT with an expiry | Short access JWT + server-side refresh session + deny list; quantifies lag and refresh load | Sets lifetimes and binding by risk tier as a company standard; tracks maximum credential lifetime in circulation |
| Failure | Replicas and failover | Splits verify / refresh / login / step-up; grace extension with a sign-off; separate pools | Identity as tier-0 cells; fail-static policy published to every product team; IdP-down game days |
| Abuse | Rate limit by IP | Layered shedding before the hash; hash budget as bulkhead; breached-password checks; challenge ceilings | Prices stuffing defense vs ATO losses; pushes passkey adoption as the structural fix with a target and an owner |
| Recovery | Reset link by email | Risk-scored recovery, cooling-off, notifications to existing factors, support script | Assigns recovery-policy ownership across security, product and support; audits support-assisted recovery |
| Operations | "Add monitoring" | revocation.propagation_p99, revocation canary, refresh.reuse_detected_total, login.success_rate | ATO rate per million logins and revocation latency as exec-visible security KPIs |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| States revocation lag as a number | "Worst case 10 minutes, 30 seconds for high-risk events via the deny list." |
| Separates the three rates | "Verify at a million a second, refresh at 25K, login at 700 — each designed to its own rate." |
| Knows stuffing is a CPU problem first | "At 100ms per hash, 50K attempts a second is 5,000 cores. We shed before hashing." |
| Splits failure posture by operation | "Verification continues, login fails closed, refresh gets a bounded grace, payouts fail closed." |
| Treats recovery as a login path | "Recovery is the weakest authentication we accept, so it gets cooling-off and notifications." |
| Designs for the silent failure | "A revoked session that still works doesn't page. I'd run a revocation canary across every ingress." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "JWTs are stateless, so we don't need a session store" | Has no revocation story; "log out everywhere" is impossible |
| "We'll check a blacklist on every request" | Rebuilds the per-request lookup the JWT avoided, without noticing |
| 24-hour or longer access tokens | Unrevocable for a day; shows no account-takeover experience |
| Rolling their own crypto or token format | Ignores decades of library and protocol hardening |
| Account lockout after N failures as the stuffing defense | Lets attackers lock out real users at will; stuffing uses one attempt per account |
| No mention of recovery | Leaves the weakest door unguarded |
5.4 Common False Positives#
- OAuth flow fluency ≠ identity design. Reciting the authorization code flow with PKCE is table stakes; the Staff question is what happens after the token is issued.
- Crypto depth ≠ security judgment. Comparing RSA-2048 with Ed25519 impresses briefly; revocation lag and key management decide whether a breach is contained.
- "Zero trust" vocabulary ≠ a design. It's Staff only if the candidate names what each request carries (user identity, device identity) and where it's verified.
- MFA everywhere ≠ safety. SMS OTP is phishable and SIM-swappable; a candidate who stops at "add MFA" hasn't thought about phishing-resistant factors or recovery bypassing them.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Pick consumer/workforce/service; state revocation lag and blast-radius constraints |
| Entities & API | 3–5 min | Account, credential, session family, signing key, audit event |
| Architecture | 5–10 min | ≤ 8 boxes; three paths and three rates |
| Revocation | 10–18 min | Token model, refresh rotation, deny list, propagation SLO |
| Failure posture | 18–25 min | IdP down: verify / refresh / login / step-up split; grace; stampede avoidance |
| Abuse + recovery | 25–34 min | Stuffing shedding, hash bulkhead, recovery cooling-off, support script |
| Pivot (interviewer's choice) | 34–42 min | Key compromise, enterprise SSO, multi-region, service identity, authorization |
| Wrap | 42–45 min | Token = cache of a decision; revocation number and owner; evolution |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Our signing key leaked" | Key management and blast radius | Separate keys for access vs refresh; pushed key revocation; mass refresh capacity |
| "Add enterprise SSO for B2B customers" | Federation and deprovisioning | Per-tenant OIDC/SAML connections, SCIM, IdP-initiated logout, tenant-level session policy |
| "Make it multi-region" | Where credentials live, what replicates | Home region per account; tokens verifiable anywhere; revocation events replicated globally |
| "Now services need identity too" | Workload identity vs user identity | Platform-attested short-lived certs; no shared secret; user identity propagated as a separate signed context |
| "A celebrity's account was taken over" | Recovery and support controls | Trace via audit log; recovery path; support override; high-value tier |
| "How do you know logout works?" | Silent-failure detection | Revocation canary across every ingress path |
6.3 What to Deliberately Skip#
- Password complexity rules. One sentence: "length minimum, breached-password check, no composition rules, per NIST 800-63B."
- Every OAuth redirect. Say "authorization code with PKCE" and move on unless asked.
- Email-verification UX. Mention it exists.
- Crypto algorithm selection. "ES256 or EdDSA, keys in KMS."
- CAPTCHA vendor details. "Challenge" is enough; the decision is when to challenge.
6.4 Follow-Up Questions to Expect#
- "A user clicks 'log out everywhere'. List every place a stolen token is still accepted, and for how long."
- "The session store is down for 20 minutes. What happens to logged-in users, and to new logins?"
- "A mobile release causes 2% of users to be logged out every day. What's your first hypothesis?"
- "How do you migrate 200 million legacy SHA-1 password hashes to argon2id?"
- "An enterprise customer fires an employee. How fast does that person lose access to our product?"
- "How do you rotate signing keys without breaking anyone, and how do you do it in an emergency?"
- "Who decides how long account recovery takes for a high-value account?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design authentication for our app."
Staff Answer
"Before drawing — is this consumer login, workforce SSO for our employees, or identity between our services? They're three different systems. I'll assume consumer login, around 200 million accounts and 60 million daily actives, with enterprise SSO as a later federation adapter.
The constraints I'll commit to: a stolen credential stops working within 10 minutes in the worst case and within 30 seconds for high-risk revocations; verification of tokens never depends on the identity service being up; login survives credential-stuffing waves of tens of thousands of attempts per second; and recovery is treated as a login path with its own controls. I'll walk through: the session and token model → the three request paths and their rates → revocation and how it reaches every verifier → what happens when identity is down → stuffing → recovery → who owns what."
Why this is L6:
- Distinguishes three intents and commits to one
- States revocation lag and failure independence as requirements, before any boxes
- Previews the outline as decisions ending in ownership
What L7 adds:
- Asks how many login systems the company already has and whether this is a consolidation
- Asks who owns recovery policy and which audits (SOC 2, regulators) the design must satisfy
- Frames identity as tier 0 with an explicit availability target that caps every product's
❌ Common L5 Trap
"Users log in with email and password, we hash with bcrypt, issue a JWT valid for 24 hours, and every service validates it. We'll add MFA and rate limiting."
Why this misses: Every part is reasonable and the design has a 24-hour revocation lag, no failure posture and no recovery story. The interviewer's next question — "the user's token was stolen" — has no good answer.
Drill 2: Revocation Mechanics#
Prompt: "A user reports their account was compromised and clicks 'log out of all devices'. Walk me through exactly what happens."
Staff Answer
"The revoke-all call marks every session family for the account as revoked in the session store, except the caller's current one if it passes step-up, and writes revoked_before = now on the account. It publishes session.revoked events with the session IDs and an account.sessions_revoked_before event. The deny-list distributor pushes those to every gateway node; p99 is under 30 seconds. Gateways reject any access token whose sid is on the list or whose iat is before the account's cutoff, until the entry's 10-minute TTL lapses — after which every pre-revocation access token has expired anyway. Refresh attempts on any revoked family fail immediately at the session store.
Then the account side: force a password change if one is in use, notify all factors, and flag the account for trust-and-safety review. I'd also check revocation.canary hasn't been failing — if any ingress path ignores the deny list, the user's 'log out' is partial."
Why this is L6:
- Distinguishes refresh-path enforcement (instant) from access-token enforcement (pushed, ≤ 30s)
- Uses an account-level cutoff so the deny list stays small even for users with many sessions
- Ties the deny-list TTL to the access-token lifetime with a reason
What L7 adds:
- Extends revocation across vendor boundaries — partner apps and SaaS receive session-revoked signals via the OpenID shared-signals standards (OpenID CAEP)
- Makes revocation latency a company security KPI with a monthly review
- Requires every new ingress path to pass the revocation canary before launch
Drill 3: Make It Concrete — Capacity#
Prompt: "Size the identity plane. How many machines, how much storage?"
Staff Answer
"Three paths. Login: 20 million interactive logins a day, ~230/s average, ~700/s at the morning peak. argon2id tuned to ~100ms of one core: 70 cores at peak. I'd provision a hash pool of 2,000 cores across regions — that's 20K hashes/s — not for legitimate traffic but so a stuffing wave degrades gracefully while shedding kicks in. Passkey logins cost ~1ms and change this math as adoption grows.
Refresh: 60M daily actives, ~2 active hours, one refresh per 10 minutes — 720 million refreshes a day, ~8K/s average, ~25K/s peak. Each is one session-store read, one conditional write and one signature. A remote KMS sign call per token adds milliseconds and a per-request fee at 25K/s, so I'd run a narrow in-region signer service backed by an HSM or KMS-protected key that only it can use — the one place a private key is ever usable.
Session store: ~400M refresh families (≈2 devices per account) × ~500 bytes ≈ 200 GB, ×3 replicas ≈ 600 GB. Fits a sharded Redis cluster with persistence or DynamoDB; I'd pick DynamoDB-style managed storage for durability and TTL-based expiry, with p99 reads ~2ms.
Verification: ~1M/s at gateways, local, ~0.1ms each: ~100 cores across the fleet, which is noise."
Why this is L6:
- Sizes each path separately and spots that the login pool is sized for attack, not for users
- Notices that per-call KMS signing doesn't fit the refresh rate
- Derives storage from sessions per account, not from users
What L7 adds:
- Prices it: hash pool, session store and SMS spend per month vs vendor per-MAU pricing
- Sets a passkey-adoption target because it shrinks the most expensive and most attacked path
- Plans per-region cells so capacity is sized per cell, not globally
Drill 4: The Identity Provider Is Down#
Prompt: "Your session store is unreachable in one region for 25 minutes. What does every other team experience?"
Staff Answer
"API calls with valid access tokens keep working: verification is local. Refresh fails over to grace mode: the token service validates the refresh token's own MAC and expiry, checks the local deny list, and issues a 10-minute access token with grace=true, at most three times. So for 30 minutes, logged-in users notice nothing. New logins fail with a clear 'try again shortly' and Retry-After. Sensitive actions — password change, payout setup, admin — reject grace=true tokens and show 'temporarily unavailable'. Other regions are unaffected because sessions are homed per region and no region calls another synchronously.
The cost: a session revoked during the outage that isn't on the deny list keeps working until grace ends — up to 30 minutes. Security signed off on that maximum. At 26 minutes, if the store is still down, the grace budget runs out and I'd rather extend once more by explicit incident-commander decision than auto-extend."
Why this is L6:
- Gives each operation its own posture with a named maximum
- Prevents the login stampede that turns a 25-minute outage into a 70-minute one
- Names who signed off on the residual risk
What L7 adds:
- Publishes the fail-static contract so every product knows which of its features accept
grace=true - Runs this exact scenario as a quarterly game day across product teams
- Tracks "identity-caused product minutes of downtime" as the identity platform's error budget
Drill 5: Credential Stuffing at the Morning Peak#
Prompt: "45,000 login attempts per second just started, from 200,000 IPs. Most fail. What do you do?"
Staff Answer
"First, protect the hash pool: it does 20K hashes a second, so without shedding legitimate users time out. Switch to wave mode: attempts from unknown devices, or from ASNs dominated by residential proxies, get a challenge before we hash. Identifiers with no prior successful device and an appearance in our breached-credential corpus get step-up. Per-identifier limits drop from 10 failures per 15 minutes to 3. I don't lock accounts — that hands the attacker a way to lock out real users.
Second, find the hits: successful logins during the wave from new devices get revoked and forced through reset with notification. Third, measure friction: challenge rate on known-good devices must stay under 2%, product's ceiling.
After: the structural fix is fewer passwords. Passkey adoption is what makes stuffing irrelevant."
Why this is L6:
- Treats the wave as a capacity attack first, a takeover attack second
- Rejects lockout with a reason
- Names the friction budget and who owns it
What L7 adds:
- Prices it: ATO losses and support cost per wave vs challenge friction lost logins
- Sets a company passkey target with a product owner, because it removes the attack surface
- Shares anonymized attack indicators with peer companies and the gateway team
Drill 6: Multi-Tenant Enterprise SSO#
Prompt: "B2B customers want their employees to log in with their own IdP. One customer has 200,000 employees."
Staff Answer
"Per-tenant federation: each tenant configures an OIDC or SAML connection; we discover the tenant from the email domain or a tenant-specific login URL. Their IdP authenticates; we create our own session with tenant_id and the IdP's amr and auth_time, and issue our normal tokens. Provisioning and deprovisioning via SCIM, so a fired employee's account is deactivated — and its sessions revoked — when the customer's IdP sends the update. Tenant-level policy: session lifetime, required MFA, allowed IP ranges.
Isolation: one tenant's misconfigured IdP or SCIM flood mustn't hurt others — per-tenant rate limits on SCIM and per-connection circuit breakers on metadata fetches. The 200K-employee tenant's Monday-morning login burst — maybe 50K logins in 30 minutes, ~30/s — is easy; their SCIM bulk sync of 200K users is the real load and gets its own queue."
Why this is L6:
- Keeps our session model; federation is an authentication method, not a new token system
- Names deprovisioning as the correctness requirement enterprises care about
- Isolates tenants from each other's misconfigurations
What L7 adds:
- Prices enterprise SSO as a product feature tier and owns its SLA
- Accepts continuous-access signals from customer IdPs so their "disable user" reaches our sessions in seconds
- Decides which tenant-policy knobs to expose and which to keep as company-wide floors
Drill 7: Build vs Buy#
Prompt: "Should we build our own identity provider or buy one?"
Staff Answer
"Workforce: buy, no debate — the SaaS connector catalogue and device-posture integrations are the product. Consumer: depends on three numbers. At 5 million MAU, a hosted CIAM at even a few cents per MAU a month is far cheaper than the 6–10 engineers plus a 24×7 rotation it takes to run identity well. At 150 million MAU, per-MAU fees can reach millions per year, and login UX and risk logic are core to conversion — building or self-hosting an open-source IdP starts to pay. Third number: how many products. If we have five apps, we need central sessions regardless.
Either way I'd own the interface: products consume our token contract and session API, not the vendor's SDK, so changing vendors is a migration of one service."
Why this is L6:
- Splits by intent and decides with numbers
- Includes the on-call and security cost of running tier 0, not only build effort
- Protects reversibility with an owned interface
What L7 adds:
- Models the 3-year cost curve with MAU growth and the crossover point
- Negotiates vendor contract terms for incident notification and data export
- Plans the exit before signing: password hash export format, session migration
Drill 8: Changing Token Lifetime Without an Outage#
Prompt: "Security wants access tokens cut from 60 minutes to 10. How do you roll that out?"
Staff Answer
"The risk isn't the token — it's 6× refresh load and every client that assumed 60 minutes. Steps: (1) capacity: refresh goes from ~4K/s to ~25K/s at peak; scale the token service and session store first and load test at 2×. (2) Shadow: emit the would-be expiry in a claim and log clients that don't refresh before it — that finds hardcoded timers. (3) Canary: 1% of new sessions get 10-minute tokens; watch refresh.error_rate, refresh.reuse_detected_total by client version, and 401 rate at the gateway. (4) Ramp 1% → 10% → 50% → 100% over two weeks, by platform, web first, mobile after the SDK fix ships. (5) Rollback is a config flag; existing tokens keep their original expiry."
Why this is L6:
- Recognizes the change as a capacity and client-compatibility change
- Uses shadow → canary → ramp with named metrics
- Keeps rollback a config change
What L7 adds:
- Turns token lifetime into a policy by risk tier so the next change is a config edit, not a project
- Sets a deprecation date for clients older than the SDK that handles short tokens
- Reports the security gain (revocation lag 60 → 10 min) in the security KPI review
Drill 9: Multi-Region#
Prompt: "We're expanding to the EU and must keep EU users' identity data there."
Staff Answer
"Home region per account. Credentials, sessions and audit events for EU accounts live only in EU cells. The identifier-to-region directory is global but holds only a hashed identifier and a region code. Login hits the nearest edge, which looks up the home region and routes there — +80–120ms for a traveling user, acceptable for login. Access tokens are verifiable anywhere: public keys are published globally, so an EU user calling a US-hosted API needs no cross-region identity call. Refresh goes to the home region.
Revocation events replicate globally: they carry session IDs and account IDs only, which I'd confirm with legal counts as acceptable. A region outage means its users can't log in or refresh — grace mode covers 30 minutes — but other regions are untouched. Replication internals: Replication."
Why this is L6:
- Separates what must be homed (credentials, sessions) from what can be global (public keys, revocation signals)
- Keeps verification region-independent
- Checks the residency status of revocation data instead of assuming
What L7 adds:
- Makes "identity cell per jurisdiction" a template so the next region is weeks, not quarters
- Owns the legal interpretation with privacy counsel as a documented decision
- Plans account migration between regions (user moves countries) as a supported flow
Drill 10: Cost#
Prompt: "Identity costs are growing faster than users. Where's the money going?"
Staff Answer
"Four usual suspects. SMS OTP: at $0.02–0.05 per message and a few percent of logins sending one, it's often the biggest line — and SMS pumping fraud can multiply it; I'd check spend by destination country. Password hashing: attack waves drive hash-pool autoscaling, so cost tracks attacks, not users — shed earlier. Refresh traffic: if a client bug refreshes every minute, refresh load is 10× design. Vendor per-MAU fees if we buy: inactive accounts may still count. Fixes in order: passkeys and TOTP over SMS, country controls on SMS, shedding before hashing, client SDK refresh discipline, purge dormant accounts."
Why this is L6:
- Knows SMS and attack-driven hashing are the non-obvious cost drivers
- Connects cost spikes to client bugs and abuse, not growth
- Orders fixes by savings
What L7 adds:
- Builds a cost-per-1,000-logins unit metric reported monthly
- Charges SMS spend back to products that choose SMS as their step-up method
- Uses passkey adoption as a cost lever in the business case, not only a security one
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — Launch-Day Login Meltdown#
Context: A new product launches with a TV campaign. Signups and logins hit 12× normal in 10 minutes; at the same moment a stuffing botnet piggybacks on the traffic. Login p99 is 25 seconds and the CEO's demo account can't sign in. The on-call escalates to you.
Questions to Surface First:
- How much of the traffic is legitimate? What's
login.success_ratedoing — dropping (stuffing) or steady (real users)? - Is the hash pool saturated, or is something upstream (risk engine, credential store) the bottleneck?
- Are signups and logins sharing the same pool?
- Did the client release for the launch change refresh or retry behavior?
Typical L5 Approach: Scales up the login service and adds per-IP rate limits. The scale-up takes 8 minutes, the botnet's per-IP rate is already below the limit, and new pods saturate on the same hashing cost.
Staff Approach: Separates signal from noise first: success rate fell from 94% to 21%, so most traffic is stuffing. Enables wave mode (challenge unknown devices before hashing), splits signup into its own pool so launch signups aren't starved, and finds the launch client retried failed logins 3 times with no backoff — tripling legitimate load.
Principal Approach: Treats launch-day identity readiness as a launch-checklist item owned by identity: capacity sign-off, wave-mode rehearsal, client SDK version gate. Makes "identity load test at 20× with attack mix" mandatory for any campaign with paid media.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Enable wave mode; confirm hash pool load drops below 70%; protect known-device logins with a priority lane. |
| Triage | 78% of attempts are stuffing from residential proxies; launch client retries ×3 with no jitter. |
| Quick fix | Server-side Retry-After honored by the SDK via remote config; separate signup pool. |
| Guardrails | Alert on login.success_rate < 70% for 3 min; hash-pool saturation auto-enables wave mode. |
| Post-mortem | Why could a launch client ship retry logic without identity review? Why was wave mode manual? |
Metrics to Watch: login.success_rate, login.hash_pool_utilization, login.challenge_rate{device=known}, signup.p99_latency
Organizational Follow-up: identity owns the auth client SDK; apps can't implement their own login retry. Launch checklist adds an identity capacity sign-off.
Ownership Question: "Who decides to enable wave mode, given it adds friction?" Staff answer: Identity on-call, automatically on saturation, within a challenge-rate ceiling product has pre-approved. Above that ceiling, the incident commander decides with product on the call.
Key Takeaway: "Login capacity is sized for the attack, and the shedding decision is pre-approved — not debated during the incident."
What clears the Staff bar:
- Reads success rate before scaling
- Protects the hash pool with shedding, not more pods
- Finds the client-retry multiplier
Deep Dive 2: Silent Failure — Logouts Haven't Propagated for Three Days#
Context: A security researcher reports that after "log out everywhere", their old access token still worked for exactly 10 minutes — not 30 seconds. Investigation shows the deny-list distributor stopped consuming three days ago after a schema change. Nothing paged.
Questions to Surface First:
- Did revocations still work at the refresh path? (Yes — session store was fine, so lag fell back to 10 minutes, not infinity.)
- Which high-risk revocations happened in those three days, and what did those sessions do in the 10-minute windows?
- Why didn't consumer lag or the canary alert?
Typical L5 Approach: Restarts the distributor, fixes the schema mismatch, closes the bug.
Staff Approach: Fixes the consumer, then fixes detection: consumer lag was monitored but routed to a dashboard, and the revocation canary only probed the refresh path, not the gateway. Adds the canary on every ingress with a 60-second threshold, and reviews the 4,100 high-risk revocations from the window for activity after revocation.
Principal Approach: Recognizes that the system degraded gracefully by design — the 10-minute TTL bounded the damage — and that the gap is in owning the SLO. Makes
revocation.propagation_p99a paging SLO, adds schema-compatibility checks to the event bus for every security-relevant topic, and reports the incident in the security KPI review.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Roll the distributor back to the old schema reader; replay the last 10 minutes of revocations. |
| Triage | 4,100 high-risk revocations in 3 days; 37 sessions made API calls after revocation, all within 10 min. |
| Quick fix | Review those 37 sessions with trust and safety; notify affected users where data was accessed. |
| Guardrails | Canary every 5 min per ingress; page if any 2xx after revoke + 60s. Consumer lag > 60s pages. |
| Post-mortem | Schema change on a security topic had no compatibility check; lag alert was not paging. |
Metrics to Watch: revocation.propagation_p99, revocation.canary_accept_after_revoke_total, denylist.consumer_lag_seconds, denylist.entries
Ownership Question: "Who owns the deny list — identity or the gateway team?" Staff answer: Identity owns the stream and the propagation SLO; the gateway team owns applying it. The canary measures the end-to-end result, and identity carries the page.
Key Takeaway: "Bounded failure is the reason the 10-minute TTL exists. Measuring the fast path is the reason the canary exists."
What clears the Staff bar:
- Notices the access TTL bounded the damage, and says so
- Measures the end-to-end revocation result, not component health
- Quantifies exposure from the incident window
Deep Dive 3: Large-Customer Onboarding — A 200,000-Employee Bank#
Context: Sales signs a bank with 200,000 employees. Requirements: SAML SSO through their IdP, phishing-resistant MFA, sessions ≤ 8 hours, deprovisioning within 15 minutes, every admin action logged and exportable, and an audit of our identity controls before go-live.
Questions to Surface First:
- Does their IdP support SCIM and push signals, or do we poll?
- Do they want their MFA policy or ours to govern? (Theirs — we record their
amr.) - Where must their audit logs live, and for how long?
Typical L5 Approach: Adds a SAML connection and a tenant flag for 8-hour sessions.
Staff Approach: Treats deprovisioning as the hard requirement: SCIM deactivate → revoke all sessions for those users → push deny entries; plus a 15-minute access-token cap for this tenant and a nightly SCIM full-sync to catch missed events. Enforces tenant policy (session length, IP ranges) in the token service, not in the app. Audit events per tenant, exportable via API.
Principal Approach: Turns the bank's requirements into a standard enterprise tier — session policy, deprovisioning SLA, audit export — priced as a product SKU, and accepts continuous-access signals from customer IdPs so their "user disabled" reaches our sessions in seconds without polling.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Per-tenant SAML connection; tenant policy object (session ≤ 8h, required amr, IP allowlist). |
| Deprovisioning | SCIM deactivate → revoke sessions → deny list; nightly full sync reconciles drift. |
| Load | SCIM initial sync of 200K users through a dedicated queue at ≤ 200/s; Monday login burst ~30/s. |
| Audit | Per-tenant audit stream, 7-year retention in their required region, export API. |
| Launch | Pilot with 500 users; measure time-to-deprovision with test accounts weekly. |
Metrics to Watch: scim.deprovision_to_revoke_p99{tenant}, saml.assertion_errors{tenant}, tenant.policy_violations
Ownership Question: "Who owns a missed deprovisioning?" Staff answer: If the customer's IdP never sent it, the customer. If we received it and didn't revoke within 15 minutes, identity — which is why we measure deprovision-to-revoke per tenant.
Key Takeaway: "For enterprises, the login is easy. Deprovisioning is the product."
What clears the Staff bar:
- Puts deprovisioning latency at the center with a measured SLA
- Keeps tenant policy enforcement in the token service
- Isolates the big tenant's sync load
Deep Dive 4: Post-Mortem — High-Value Accounts Taken Over Through Support#
Context: Over two weeks, 23 high-follower accounts were taken over. None of the attackers knew a password. All 23 went through support-assisted recovery. You own the post-mortem.
Questions to Surface First:
- What did the attackers present to support, and which agent actions moved contact details?
- Did existing factors get notified, and were notifications delivered to the attacker's new contact details?
- Is there an insider or a single agent pattern?
Typical L5 Approach: Retrains support agents, adds a security question to the support script.
Staff Approach: Finds the control gap: agents could change the email address directly, and the recovery flow then sent everything to the new address. Removes direct contact edits from support tooling, routes all support recovery through the same cooling-off flow with notifications to the previous factors, and adds second-agent approval for high-value tiers.
Principal Approach: Treats support as part of the authentication system. Assigns recovery-policy ownership to security with product and support as consulted parties, audits support-assisted recoveries monthly, and measures
recovery.success_then_dispute_rateper channel and per agent as a standing security metric.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Freeze the 23 accounts; restore credentials from audit history; revoke all sessions. |
| Triage | 19 of 23 recoveries handled by agents at one outsourced site; scripts allowed direct email change. |
| Quick fix | Remove direct contact-edit permission; force high-value recoveries through 72h cooling-off. |
| Guardrails | Alert on support recoveries per agent per day > 3× median; notify previous email on any change. |
| Post-mortem | Support tooling was outside identity's design review. Recovery policy had no named owner. |
Metrics to Watch: recovery.support_assisted_rate, recovery.success_then_dispute_rate{channel}, account.contact_change_without_factor_total
Ownership Question: "Who owns the support tool's account-edit permissions?" Staff answer: Identity owns the permission model for anything that changes a credential or contact detail, even inside support tooling. Support owns the workflow on top.
Key Takeaway: "Every human who can change a credential is part of the authentication system."
What clears the Staff bar:
- Finds the tooling permission, not just the training gap
- Notifies previous factors, which the attacker doesn't control
- Measures recovery abuse per channel
Deep Dive 5: Multi-Region Expansion — Identity Cells#
Context: The company runs identity in one US region. A regional cloud outage took login down globally for 3 hours last quarter. Leadership wants regional independence and EU residency within a year.
Questions to Surface First:
- Which identity data must be homed (credentials, sessions, audit) and which can be global (public keys, routing directory)?
- How will a user whose home region is down be treated?
- Which product services call the identity plane synchronously today?
Typical L5 Approach: Deploys a second active region with a globally replicated user database.
Staff Approach: Cells per region with home-region accounts. Global, read-only directory mapping hashed identifier → home region. Public keys published globally so verification works anywhere. Revocation events replicated to all regions. A home-region outage affects only that region's users and grace mode covers 30 minutes for them.
Principal Approach: Sequences the migration as a multi-quarter program with an explicit order — verification independence first, then refresh, then login — and makes "no synchronous cross-region identity call" a design-review rule for every product service.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Scoping | Classify identity data: homed vs global; legal confirms revocation events may replicate. |
| Architecture | Region cells: login, token, session store, credential store, audit per region. |
| Migration | Move accounts to home regions in batches of 1M with dual-read for 7 days per batch. |
| Failure drills | Turn off a cell's login in a game day; verify other cells unaffected. |
| Launch | EU cell first for new signups, then migrate existing EU accounts. |
Metrics to Watch: login.success_rate{cell}, directory.lookup_latency_p99, revocation.cross_region_replication_lag
Ownership Question: "Who owns the global directory, the one shared component?" Staff answer: Identity platform, with the strictest change process in the stack — it's read-only on the hot path and replicated to every cell so its outage doesn't block login.
Key Takeaway: "Identity becomes regional when its outage stops being global."
What clears the Staff bar:
- Distinguishes homed from global identity data
- Keeps the one global component off the critical write path
- Plans migration in batches with dual-read
9. Level Expectations Summary#
After studying this case study, you should be able to:
- State revocation lag as a number for any token design and explain how to shrink it
- Design the hybrid model: short access JWT, server-side refresh sessions with rotation and reuse detection, and a pushed deny list
- Size login, refresh and verification as three separate paths with their own rates
- Split IdP failure posture by operation and defend the grace extension with a sign-off
- Explain credential stuffing as a CPU and economics problem and shed it before the hash
- Design account recovery with cooling-off, notifications and support controls
- Run a routine and an emergency signing-key rotation, issuer side and verifier side
- Decide when to buy an identity provider and how to keep the decision reversible
The Bar for This Question#
Mid-level (L4): Builds a working login: hashed passwords, a session cookie or a JWT, a password reset email, maybe TOTP. Doesn't consider revocation lag, outages or abuse. Would ship something that works for honest users and fails the first takeover report.
Senior (L5): Adds MFA, rate limiting, refresh tokens and replicas. Knows JWTs are hard to revoke. The gap: treats revocation as a database delete, sets token lifetimes without connecting them to exposure, answers "IdP down" with "replicas", and sends a reset link to an email that may already be compromised. The design is plausible and would pass review — and would still turn a stolen refresh token into a 90-day problem.
Staff+ (L6): Frames the problem around revocation and blast radius in the first five minutes. Chooses the hybrid token model with numbers — refresh load, deny-list size, lag. Splits failure posture by operation, sheds stuffing before hashing, treats recovery as a login path, and runs a canary for the silent failure. Names who pays each tradeoff: users carry bounded exposure, product carries challenge friction, security signs the grace table. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 "Stateless JWT" Is a Revocation Decision Disguised as a Performance Decision#
| Claim | Reality |
|---|---|
| "JWTs scale because there's no lookup" | True for verification; refresh and revocation still need state |
| "We'll add a blacklist" | A per-request blacklist lookup is a session store with extra steps |
| "Short expiry fixes it" | Only if refresh checks server-side state; otherwise refresh is the unrevocable part |
The Staff position: Every token design has state somewhere. Choose where — per request, per refresh, per risk event — and state the lag it buys.
Why this matters in interviews: "Stateless" is the word that makes interviewers ask about logout. Have the number ready.
10.2 Account Lockout Is a Gift to Attackers#
| Defense | Effect on Stuffing | Effect on Real Users |
|---|---|---|
| Lock after 5 failures | None — stuffing uses 1 attempt per account | Attackers can lock out any account they name |
| Per-identifier throttle + step-up | Slows targeted guessing | Friction only for the targeted account |
| Pre-hash shedding by device and reputation | Removes most of the wave | Small challenge rate on shared IPs |
The Staff position: Throttle and challenge; don't lock. NIST's 100-attempt ceiling is a ceiling, not a design.
Why this matters in interviews: "Lock the account" is the most common stuffing answer and the easiest to dismantle.
10.3 Recovery Is the Real Login Page#
The Staff position: Whatever your strongest factor is, your account is only as strong as the weakest way to replace it. A passkey-protected account with email-only recovery is an email-protected account. Design recovery first, then login.
Why this matters in interviews: Bringing up recovery unprompted is one of the clearest Staff signals on this question.
10.4 Most Companies Should Not Build Central Authorization#
| Stage | Right Answer |
|---|---|
| Roles and tenants | Coarse claims in tokens + policy library per service |
| Shared policy, many services | Policy-as-code with local evaluation and decision logs |
| Sharing graphs across many services | Central relationship service, Zanzibar-style |
The Staff position: Zanzibar proves central authorization can be fast. It doesn't prove you should build it. Wait for the graph.
Why this matters in interviews: Candidates who propose a central authorization service at minute 10 have added a tier-0 dependency without a reason.
10.5 SMS Is a Cost Center and a Liability#
The Staff position: SMS OTP is phishable, SIM-swappable, and an attack target for pumping fraud that bills you per message. Keep it as a fallback, make passkeys and TOTP the default, and watch SMS spend by destination country like a security metric.
Why this matters in interviews: Naming SMS's cost and abuse profile shows you've seen the bill, not just the design doc.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer designs a correct identity system. The Principal engineer notices that the company has five — the original consumer login, the acquired product's login with its own user table, an enterprise SSO bolted on by the B2B team, a partner API using static keys, and service-to-service calls authenticated by a shared secret from 2019. Each has its own token format, its own lifetime and its own idea of "log out". The L7 problem is not one correct login; it is identity as a company-wide trust boundary: one revocation story, one key-management discipline, one recovery policy with a named owner, and a migration path that retires the other four without a flag day.
🧭 Principal Move: "I'm not going to ask whether our login is secure. I'm going to ask what the longest-lived credential in circulation is, which system issued it, and who would notice if it were stolen. That one number tells me where to spend the next year."
The Org-Level Fault Line#
One identity platform vs per-product login.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each product owns login | Product speed; tailored UX | N recovery flows, N token formats, N stuffing defenses; no "log out everywhere" across products | Security (N attack surfaces), users (N passwords), support |
| Central identity platform owns everything | One revocation, one recovery policy, one audit trail | Platform becomes a bottleneck for every signup-flow experiment | Product teams (velocity), growth (experiments) |
| Platform owns primitives; products own experience | Credentials, sessions, tokens, keys, recovery policy are central; login UI and onboarding are product | Contract design is hard; needs an SDK and a hosted login with theming | Identity platform (API stewardship and SDK) |
🧭 Principal Move: "The platform owns everything that must be right once — credential storage, sessions, token issuance, keys, revocation and the recovery floor. Products own everything that must be fast to change — onboarding screens, copy, which step-up method to offer above the floor. New products may not create a user table; the two existing ones migrate within four quarters."
Cost Model#
Assumptions: consumer-heavy, fully loaded engineer ~$250K/year, cloud list prices, SMS at ~$0.03/message blended, vendor CIAM priced per MAU at negotiated rates that fall with volume (estimates, not quotes).
| Scale | MAU | Infra ($/month) | SMS ($/month) | Headcount | On-call Load | Buy Alternative (approx.) |
|---|---|---|---|---|---|---|
| Startup | 1M | ~$1–3K | ~$1–3K | 0.5–1 eng integrating a vendor | Vendor's pager; < 1 page/month | Vendor ~$2–10K/month — buy |
| Growth | 30M | ~$25–50K (hash pool, session store, multi-AZ) | ~$30–60K | 6–10 eng: identity core (4), abuse (2), recovery/support tooling (1–2), SRE (1–2) | Dedicated rotation, 3–6 pages/month, mostly abuse waves | Vendor ~$100–300K/month — decide on control needs |
| Enterprise | 300M | ~$300–600K (regional cells, attack headroom, audit storage) | ~$200–500K unless passkeys displace SMS | 30–50 eng across identity, abuse, workforce, workload identity, plus security partners | Follow-the-sun, per-cell rotations | Vendor fees exceed an in-house team; build or self-host |
The pricing insight: at growth scale, the two levers worth more than any infrastructure tuning are SMS volume and attack headroom. Moving 30% of step-ups from SMS to passkeys or TOTP saves ~$10–20K/month and removes a phishing vector; shedding stuffing before the hash avoids sizing the hash pool for the attack. Headcount, not infrastructure, is the dominant cost at every scale.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Password hash algorithm and parameters | One-way-ish | Can only migrate on next login; dormant accounts keep the old hash for years |
| Account identifier model (email as ID vs internal ID) | One-way | Every product's foreign keys; email changes become migrations |
| Who holds credentials (vendor vs in-house) | One-way | Password hashes may not be exportable in a usable format; forced resets for millions |
| Token audience and claim contract consumed by services | One-way-ish | Every service parses it; versioning takes quarters |
| Passkey relying-party ID (domain) | One-way | Passkeys are bound to it; changing domains strands credentials |
| Access-token lifetime | Two-way | Config, with capacity planning |
| Grace-extension duration | Two-way | Policy change with security sign-off |
| Session store technology | Two-way | Sessions expire; dual-write for one refresh lifetime then cut over |
🧭 Principal Insight: The passkey relying-party ID and the account identifier model are the decisions I'd slow down on. Everything about tokens can change in a quarter; those can't change in a year.
The Standard I'd Write#
RFC-ID-001: Authentication and Session Standard
Status: Approved Owners: Identity Platform + Security Engineering
Scope
Every system that authenticates users, issues or accepts user or service
credentials, or caches identity or session validity.
MUST
1. Authenticate users only through the Identity Platform API or hosted login.
2. Access tokens: lifetime ≤ 15 min; refresh tokens: server-side state,
rotated, reuse revokes the family.
3. Any cache of identity or session validity: TTL ≤ access-token lifetime and
subscribed to session.revoked.
4. Every ingress path passes the revocation canary before launch.
5. Signing keys live in KMS/HSM, are non-exportable, rotate ≤ 30 days, and are
scoped per token type and audience. Verifiers check iss, aud and key scope.
6. Service-to-service authentication uses platform-issued workload identity;
no static shared secrets.
7. Account recovery follows the tiered recovery policy owned by Security.
SHOULD
1. Offer passkeys as the default credential for new accounts.
2. Sender-constrain first-party tokens (DPoP or mTLS).
3. Reject grace-mode tokens for sensitive operations.
Exceptions
Filed with Identity Platform; reviewed within 10 business days; time-boxed to
2 quarters; CISO sign-off for any exception to MUST 2, 5 or 7.
Success metrics
- revocation.propagation_p99 ≤ 30s; canary failures: 0
- Longest-lived user credential in circulation ≤ 90 days
- ATO per million logins: tracked monthly, target set yearly by Security
- Static service secrets in config: 0 by end of year 2
- Identity-caused product downtime: ≤ 20 min/quarter
What I'd Tell the VP#
"We have five ways to log in and no single way to log someone out. If one of our users' credentials is stolen today, it can stay useful for up to 90 days in two of those systems. I'm proposing one identity platform that owns sessions, keys and recovery policy, while product teams keep their signup and onboarding experience. It's about eight engineers for a year, retires two legacy login systems, cuts the maximum lifetime of a stolen credential from 90 days to 10 minutes, and makes our next identity outage a degraded login page instead of a company-wide outage. The main risk is migration friction for the acquired product's users; I'd move them on their next login, not with a forced reset."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Measures the worst credential, not the average | "What's the longest-lived credential in circulation, and who issued it?" |
| Prices friction against loss | "A 1% challenge rate costs us ~0.3% of logins; stuffing costs us ~$400K a quarter in ATO remediation. Here's the trade." |
| Redraws ownership | "Security owns the recovery floor, product owns friction above it, support follows a script it can't override." |
| Sets the org's failure posture | "Identity's error budget is every product's error budget; we publish the fail-static contract." |
| Knows when not to standardize | "Workload identity shares our key discipline but not our session store; forcing one system serves neither." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a secure login with short tokens and MFA" — correct for one product; ignores the four other login systems in the company.
- "Security will decide the recovery policy" — names an owner but not the tension with growth and support, nor the audit.
- "We'll buy an IdP" — no exit plan, no ownership of the token contract, no answer for the vendor's breach.
Appendices
Appendix A: Mechanics in Depth#
A.1 Login and Token Issuance#
A.2 Refresh With Rotation and Grace#
def refresh(rt, dpop_proof):
sid, mac_ok, exp = parse_and_verify_mac(rt) # self-verifiable for grace mode
if not mac_ok or exp < now(): return 401
try:
s = session_store.get(sid, timeout=50ms)
except Unavailable:
return grace_issue(sid, dpop_proof) # ≤ 3 × 10 min, deny list checked
if s.revoked_at: return 401("session_revoked")
if not dpop_matches(s.binding_key_thumbprint, dpop_proof): return 401
h = sha256(rt)
if h == s.current_refresh_hash:
new_rt = mint_rt(sid)
ok = session_store.cas(sid, expect=h, set_current=sha256(new_rt), set_prev=h,
rotated_at=now())
if not ok: return refresh_retry_once()
return issue_access(s), new_rt
if h == s.prev_refresh_hash and now() - s.rotated_at < 60s:
return issue_access(s), last_issued_rt(sid) # idempotent retry, same RT2
revoke_family(s.family_id, reason="reuse_detected") # theft or very late replay
publish("session.revoked", s.family_id)
return 401("reuse_detected")
A.3 Legacy Password Hash Migration#
Rehash on successful login: verify with the legacy algorithm, then store argon2id and mark the record migrated. For dormant accounts, wrap the legacy hash — argon2id(legacy_hash) — so the weak hash is no longer stored, and verify by computing the legacy hash first. After ~12 months, accounts still on wrapped hashes get a reset on next login.
Appendix B: Data Model and Session Lifecycle#
CREATE TABLE credentials (
credential_id TEXT PRIMARY KEY,
account_id TEXT NOT NULL,
type TEXT NOT NULL, -- password | passkey | totp | phone | email | federated
secret_ref BYTEA NOT NULL, -- argon2id hash or passkey public key
added_at TIMESTAMPTZ NOT NULL,
trusted_after TIMESTAMPTZ NOT NULL, -- cooling-off end; limited mode before this
revoked_at TIMESTAMPTZ
);
-- Session store (key-value, TTL = absolute_expires_at)
session:{sid} → { account_id, device_id, family_id, current_refresh_hash, prev_refresh_hash,
rotated_at, auth_time, amr, binding_key_thumbprint, expires_at,
absolute_expires_at, revoked_at }
account_cutoff:{account_id} → revoked_before timestamp -- for revoke-all
Refresh tokens are stored only as hashes; a session-store dump does not yield usable tokens.
Appendix C: Coordination Mechanisms#
C.1 Revocation Fan-Out#
The session-store write and the event are coupled through an outbox or change-data capture, so a revocation can't be committed without being published. Event-bus mechanics: Apache Kafka.
C.2 Quick Comparison#
| Mechanism | Revocation Lag | Per-Request Cost | Failure Mode | Use For |
|---|---|---|---|---|
| Access-token expiry | ≤ access TTL (10 min) | None | Long TTL = long exposure | Baseline for all revocations |
| Refresh-path session check | Next refresh (≤ 10 min) | None (per refresh) | Session store outage | Logout, normal revocation |
| Pushed deny list | ~30s | In-memory set lookup | Distributor stall (silent) | High-risk revocations |
Account revoked_before cutoff | ~30s | In-memory map lookup | Same as deny list | Revoke-all for one account |
| Key revocation | ~60s via pushed key config | None | Mass re-auth load | Key compromise |
| Introspection per request | ~0 | 2–20ms network call | IdP in every request path | Low-volume, high-assurance APIs |
| Client revocation endpoint (OAuth token revocation) | Immediate at the issuer | None | Only affects issuer-side checks; verifiers still need the push | Client-initiated logout (RFC 7009) |
Appendix D: API Contract & Client Behavior#
- Clients refresh at ~80% of access-token lifetime with ±10% jitter; on failure, exponential backoff from 1s capped at 60s, and honor
Retry-After. - Refresh tokens are stored in platform secure storage (Keychain, Keystore) or HttpOnly, Secure, SameSite cookies on web — never in localStorage.
- First-party mobile clients generate a non-exportable device key and send DPoP proofs on refresh.
401 session_revoked→ clear tokens, go to login.401 reuse_detected→ same, plus show a security notice.- Login and recovery responses never reveal whether an identifier exists (uniform responses and timing).
- Authorization code with PKCE for every OAuth client; no implicit grant, no password grant, per RFC 9700.
Appendix E: Observability#
Core metrics:
login.attempts_rate,login.success_rate,login.hash_pool_utilization,login.challenge_rate{device}refresh.rate,refresh.error_rate,refresh.grace_issued_total,refresh.reuse_detected_total{client_version}revocation.propagation_p99,revocation.canary_accept_after_revoke_total(must stay 0),denylist.entriesrecovery.started_rate,recovery.cancelled_by_owner_rate,recovery.success_then_dispute_rate{channel}tokens.verified_with_unknown_session_total(forgery signal),jwks.verifiers_missing_active_kidotp.sms_sent{country},ato.confirmed_per_million_logins
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| Revocation canary accepted after revoke | > 0 | Sev-1 page |
| Valid signature, unknown session | > 0 sustained 5 min | Sev-1 page (possible key compromise) |
| Revocation propagation p99 | > 60s for 5 min | Page |
| Login success rate | < 70% for 3 min | Page + auto wave mode |
| Refresh error rate | > 5% for 2 min | Page |
| Refresh reuse detected | > 3× baseline for one client version | Page identity + mobile |
| SMS sends to one country | > 5× 7-day baseline | Page (pumping fraud) |
Debugging the silent failure: a revoked session that still works produces no errors anywhere. The canary is the only reliable detector; second best is tokens.used_after_revocation computed offline from gateway logs joined to revocation events.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 1M MAU | Hosted CIAM; vendor sessions; your token contract on top | Vendor limits on custom risk logic |
| 1–30M MAU | Hybrid tokens, single-region session store, wave mode | Stuffing cost, SMS spend, second product wants login |
| 30–300M MAU | Identity platform + SDK, passkeys, deny-list fan-out, regional failover | Cross-region outages, residency, legacy formats |
| > 300M MAU | Regional cells, shared signals, workload identity everywhere | Organizational: retiring the last legacy login |
What you don't build on day one: your own HSMs, a central relationship-based authorization service, DPoP for every client, regional identity cells, cross-vendor shared signals. Each has a trigger in Section 11.
Appendix G: Service-to-Service Identity, Multi-Tenancy and Cost#
- Workload identity: certificates live ~12 hours and renew at half-life, so a CA outage of up to ~6 hours doesn't break running services; it blocks new workloads. The CA is tier 0 for deploys, not for traffic. User identity rides alongside as a separate signed context, never as a workload certificate. Platform mechanics: Kubernetes; service registration: Service Discovery.
- Tenant isolation: per-tenant rate limits on SCIM and federation metadata fetches; per-tenant session policy; one tenant's misconfigured IdP produces errors only for that tenant.
- Cost attribution: tag SMS sends and step-up challenges with the requesting product; products that pick SMS as their step-up method see the bill.
- Abuse costs are bursty: budget hash-pool headroom and SMS spend for attack weeks, not average weeks; report cost per 1,000 logins monthly.