Hiring BarSupport

Security Basics for System Design

Foundation37 min read5 diagrams

Go deeper:

  • For the full system design of the service behind envelope encryption, covering DEK caching as a revocation contract, regional isolation and safe key destruction, see Design a Key Management Service.

Why This Matters#

Security in a design interview is not a checklist question. Every candidate can say "HTTPS, OAuth, encrypt the database." The question hiding underneath is: what is the blast radius when one credential leaks, how fast can we make that credential worthless, and who finds out? A design where a single leaked API key reads every tenant's data is broken no matter how strong its ciphers are. A design where a stolen token expires in 10 minutes, is scoped to one tenant and one action, and shows up in an anomaly alert is defensible even if something leaks — and something always leaks.

Interviewers rarely ask "design the security." They probe it in the middle of something else: "How does the mobile app call this API?", "How does the worker get the database password?", "A customer says another customer saw their invoice — what happened?", "We need to rotate the signing key. Walk me through it." The Senior answer names the right technology. The Staff answer names the trust boundary, the credential lifetime, the revocation path, and the owner of the key — and is honest about which threats the design does not handle.

Staff engineers carry a few convictions into every design. Authentication is cheap; authorization is where breaches happen — most real data exposures are a correctly authenticated user asking for an object they shouldn't see. Every secret will leak eventually, so lifetimes and rotation matter more than storage. Encryption at rest mostly protects against stolen disks, not against an attacker with your app's credentials — so access control and key scoping do the real work. And security that depends on every engineer remembering a rule fails; the paved road must make the secure option the default one.

This page covers transport security (TLS and mTLS), the authentication/authorization split, tokens versus sessions, secrets management, encryption at rest and in transit, key rotation, and a 5-minute threat model you can run out loud. A full login and session system lives in Login & Sessions; edge enforcement lives in Edge Gateway.

The 60-Second Version#

  • TLS everywhere, including inside the datacenter. TLS 1.3 costs 1 round trip for a new connection (0 with resumption) and a fraction of a millisecond of CPU; AES-GCM runs at multiple GB/s per core with hardware support. "Internal traffic is trusted" is a perimeter assumption that one compromised pod breaks.
  • mTLS gives services an identity, not just encryption. Each workload gets a short-lived certificate (hours, not years) naming what it is (spiffe://prod/payments-api), and authorization policies use that name instead of IP addresses.
  • AuthN answers "who are you?"; AuthZ answers "may you do this to that object?" Authorization must run on every request, against the specific resource and tenant. Missing object-level checks (IDOR/BOLA) are among the most common real API breaches.
  • Sessions vs tokens is a revocation tradeoff. Server-side sessions revoke instantly at the cost of a lookup (~1ms from Redis). Self-contained JWTs verify locally in tens of microseconds but can't be revoked before expiry — so keep access tokens to 5–15 minutes and put revocation on the refresh token.
  • Secrets belong in a secrets manager, delivered at runtime, scoped per service, and rotated. Not in git, not baked into images, ideally not in long-lived environment variables. Dynamic, per-instance credentials with ~1-hour TTLs make a leaked secret expire on its own.
  • Envelope encryption: data is encrypted with a data key (DEK); the DEK is encrypted with a key-encryption key (KEK) held in a KMS/HSM. Rotating the KEK re-wraps kilobytes of keys, not terabytes of data.
  • Rotation is a capability, not an event. If rotating a key or certificate requires a deploy, a meeting or downtime, it won't happen until an incident forces it. Automate it, overlap old and new, and rotate on a schedule so the path stays exercised.
  • Threat model in 5 minutes: assets, actors, trust boundaries, top 3 threats, and which you're explicitly not defending against.

How Security Works (for System Designers)#

The Four Questions Behind Every Control#

QuestionMechanismFails When
Who is calling? (authentication)Passwords + MFA, OIDC login, API keys, mTLS certificates, signed tokensCredentials are stolen, phished, or long-lived and leaked
May they do this, to this object? (authorization)RBAC, ABAC, relationship-based checks, tenant scoping, policy enginesChecks happen only at the edge or only on the collection, not the object
Can anyone read or alter it on the way? (transport)TLS 1.2+/1.3, mTLS, certificate validationPlaintext internal hops; disabled certificate verification "for testing"
Can anyone read it where it sits? (at rest)Disk/volume encryption, database TDE, application-level envelope encryptionThe attacker uses the application's own credentials — at-rest encryption is transparent to them

🎯 Staff Move: "I'll separate the four questions. Users authenticate at the edge with OIDC; services authenticate to each other with mTLS identities; every service authorizes each request against the tenant and object, not just the route; and the PII columns get application-level envelope encryption so a database snapshot alone is useless."

Trust Boundaries#

A trust boundary is any line where the level of trust changes: the internet to your edge, the edge to internal services, a service to its database, your system to a third party, one tenant's data to another's. Every boundary needs three answers: how the caller proves identity, what is checked, and what is logged.

BoundaryTypical IdentityTypical CheckCommon Gap
Internet → edgeUser session / OAuth access token, API keyToken validation, rate limits, WAF rulesCredential stuffing without per-account rate limits
Edge → serviceForwarded user identity + service mTLSService checks user's permission on the objectServices trust an X-User-Id header anyone inside can set
Service → servicemTLS workload identityAllow-list of caller identities per endpointFlat network: any pod can call any service
Service → datastorePer-service DB credentials (ideally dynamic)DB grants scoped to tables the service ownsOne shared admin password for every service
Tenant → tenantTenant ID bound to the authenticated principalEvery query filtered by tenant, enforced centrallyTenant ID taken from the request body
System → third partyOutbound credentials, webhook signaturesVerify signatures on inbound webhooksUnsigned webhooks accepted from anyone

TLS and mTLS in One Page#

TLS gives three properties: confidentiality (encryption), integrity (tamper detection), and server authentication (the certificate proves the server's name). mTLS adds client authentication: the client also presents a certificate, so both sides know who they're talking to.

PropertyTLS 1.2TLS 1.3
Full handshake round trips2 RTT1 RTT
Resumption1 RTT0-RTT possible (early data is replayable — only for idempotent requests)
Forward secrecyOptional (depends on cipher suite)Mandatory (ephemeral key exchange)
Legacy ciphersMany, some weakRemoved

The expensive part of TLS is not the encryption, it's the handshake and the certificate lifecycle. Connection reuse makes the handshake rare; automated issuance makes the lifecycle safe.

Key Terms#

TermMeaning
PrincipalThe authenticated identity making a request — a user, service, or device
SessionServer-side state keyed by an opaque random ID stored in a cookie
Access tokenShort-lived credential presented on each request (often a signed JWT)
Refresh tokenLonger-lived credential used only to obtain new access tokens; revocable
Scope / claimWhat a token permits (orders:read) and facts about the principal (tenant=acme)
IDOR / BOLAInsecure direct object reference / broken object-level authorization: changing an ID in the URL returns someone else's data
KMS / HSMKey management service / hardware security module — holds root keys and performs crypto without exposing them
DEK / KEKData encryption key (encrypts data) / key encryption key (encrypts DEKs)
Workload identityA cryptographic identity for a running service (e.g., a SPIFFE ID in an X.509 certificate)
Least privilegeEach principal gets only the permissions its job needs, nothing more
Blast radiusWhat an attacker can reach with one compromised credential

Where Security Lives in a Design#

  • Edge — TLS termination, user authentication, coarse rate limits, bot and abuse defences. See Edge Gateway.
  • Service mesh / sidecars — mTLS between services, identity-based allow-lists. See Envoy, Kong & NGINX.
  • Every service — object-level and tenant-level authorization; input validation; audit logging.
  • Data layer — per-service credentials, encryption at rest, field-level encryption for sensitive columns, backups encrypted with separate keys.
  • Platform — secrets manager, KMS, certificate authority, identity provider, policy engine.
  • Pipeline — dependency scanning, secret scanning on commits, signed builds, deploy approvals.

Core Strategies#

Strategy 1: Server-Side Sessions#

login(username, password, mfa_code):
    user = users.get(username)
    if not argon2_verify(user.pw_hash, password) or not totp_ok(user, mfa_code):
        rate_limiter.record_failure(username, client_ip)
        return 401
    sid = random_bytes(32)                         # 256 bits, unguessable, opaque
    sessions.put(sid, {user_id, tenant_id, created, last_seen}, ttl = 30 min idle / 12 h absolute)
    set_cookie("sid", sid, HttpOnly, Secure, SameSite=Lax)

on_request(req):
    s = sessions.get(req.cookie.sid)               # ~0.5–1ms from Redis
    if not s: return 401
    sessions.touch(sid)

When to use: First-party web apps where the same organization runs the frontend and backend. Revocation is instant (delete the key), logout means logout, and the cookie leaks nothing.

Failure mode: Every request depends on the session store. If Redis is down, nobody is logged in — so the session store is tier-0 and needs replication. Cookies need CSRF defences (SameSite plus anti-CSRF tokens for state-changing requests).

Strategy 2: Short-Lived Signed Tokens + Revocable Refresh Tokens#

issue(user):
    access  = jwt_sign(kid = "2026-10", alg = ES256,
                       claims = {sub: user.id, tenant: user.tenant, scope: "orders:read orders:write",
                                 exp: now + 10 min, aud: "orders-api"})
    refresh = random_bytes(32)                     # opaque, stored server-side, rotates on each use
    refresh_store.put(hash(refresh), {user, device, exp: now + 30 days})
    return (access, refresh)

verify(access_token):
    hdr = parse_header(access_token)
    key = jwks_cache.get(hdr.kid)                  # public keys, refreshed every few minutes
    claims = jwt_verify(access_token, key)         # ~20–100µs, no network call
    require(claims.aud == "orders-api" and claims.exp > now - 60s_leeway)
    return claims

When to use: APIs called by mobile apps, SPAs and third parties; many services that must verify identity without a shared session store. The issuer signs with a private key; every service verifies with the public key.

Failure mode: A stolen access token is valid until it expires — there's no "delete." That's why lifetimes stay at 5–15 minutes and refresh tokens are opaque, server-side, rotated on every use, and revoked on logout or password change. Common bugs: accepting alg: none, not checking aud (a token for service A works on service B), putting PII in claims (JWTs are signed, not encrypted), and 24-hour access tokens "to reduce load."

Strategy 3: Workload Identity with mTLS#

# Each pod at startup (via a node agent or sidecar):
csr  = generate_keypair_and_csr(spiffe_id = "spiffe://prod/ns/payments/sa/payments-api")
cert = internal_ca.sign(csr, ttl = 24h)            # attested by the platform, not by a static secret
rotate when 50% of ttl has elapsed                 # renew at ~12h, old cert valid until 24h

# Receiving service policy (enforced in sidecar or library):
allow  spiffe://prod/ns/checkout/sa/checkout-api   -> POST /charges
allow  spiffe://prod/ns/billing/sa/billing-worker  -> POST /refunds
deny   everything else

When to use: Any service-to-service traffic beyond a handful of services. It replaces IP allow-lists (which break with autoscaling) and shared API keys (which leak and never rotate) with identities the platform issues and rotates automatically.

Failure mode: The internal CA becomes tier-0: if it can't issue, new pods can't start, and when certificates expire en masse, everything stops talking. Short lifetimes make the CA's availability critical. Monitor cert.expiry_seconds minimum across the fleet and alert well before the shortest-lived certificates run out.

Strategy 4: Secrets Delivered at Runtime#

# Bad: DB_PASSWORD baked into the image or committed to config
# Better: static secret fetched from a secrets manager at startup, rotated every 30–90 days
# Best: dynamic, per-instance credentials

on_start():
    token = platform_identity_token()              # proves "I am payments-api in prod"
    creds = secrets_manager.issue_db_creds(role = "payments_rw", ttl = 1h)
    db.connect(creds)
    schedule_renewal(at = 0.5 * ttl)               # new creds before old ones expire

secrets_manager.issue_db_creds(role, ttl):
    user = create_db_user(grants = role.grants, expires = now + ttl)
    audit_log(caller, role, user)
    return user

When to use: Always the "better" tier; the "best" tier for databases and cloud APIs that support short-lived credentials. The platform identity bootstraps everything, so there's no secret zero sitting in an environment variable.

Failure mode: The secrets manager is in the startup path of every service. If it's down, nothing new starts — run it highly available, cache leases, and allow existing credentials to keep working until their TTL. Second failure mode: secrets fetched correctly and then logged, put in crash dumps, or exposed in a debug endpoint. Secret scanning on logs and commits catches what reviews miss.

Strategy 5: Envelope Encryption for Data at Rest#

encrypt_record(tenant, plaintext):
    dek = dek_cache.get(tenant) or new_dek(tenant)        # 256-bit AES key, per tenant or per object
    nonce = random_bytes(12)
    ct = aes_gcm_encrypt(dek.key, nonce, plaintext, aad = tenant.id + record.id)
    return {ct, nonce, wrapped_dek: dek.wrapped, kek_id: dek.kek_id}

new_dek(tenant):
    key = random_bytes(32)
    wrapped = kms.encrypt(kek_id = tenant.kek, key)       # one KMS call per new DEK, ~5–20ms
    dek_cache.put(tenant, {key, wrapped, kek_id}, ttl = 5 min, max_uses = 100K)

decrypt_record(rec):
    key = dek_cache.get(rec.wrapped_dek) or kms.decrypt(rec.wrapped_dek)
    return aes_gcm_decrypt(key, rec.nonce, rec.ct, aad = ...)

When to use: Sensitive fields (PII, payment data, health data), multi-tenant stores where customers may require their own keys, and any data where you need crypto-shredding — deleting a tenant's KEK renders all their data unreadable, including in backups.

Failure mode: Calling KMS per record turns a 1ms read into a 15ms read and hits regional KMS request quotas at a few thousand requests per second; cache DEKs with bounded lifetime and usage. AAD that doesn't bind the ciphertext to its row lets an attacker with write access swap encrypted values between rows. And a KMS outage means you cannot decrypt — the KMS is a tier-0 dependency of every read path that uses it.

🎯 Staff Move: "Disk encryption is table stakes and protects against a stolen drive. It does nothing against someone who steals the app's database credentials. For the SSN and bank-account columns, I'll add envelope encryption with per-tenant data keys, so a SQL injection or a leaked read replica returns ciphertext, and deleting a tenant's key shreds their backups too."

Rotation and Revocation: The Hard Sub-Problem#

Issuing a credential is easy. Making a credential stop working — on schedule, or in a hurry after a leak — is where designs fail. Every credential type has a different answer to "how do I kill it?", and the interview signal is knowing all of them.

How Each Credential Dies#

CredentialRevocation MechanismTime to Effective RevocationCost of Faster Revocation
Server-side sessionDelete the session recordImmediateNone beyond the per-request lookup you already pay
JWT access tokenWait for expiry, or add a deny-list check= token lifetime (5–15 min)A deny-list lookup per request — which reintroduces the session store
Refresh tokenDelete server-side recordNext refresh attemptNone
API key (long-lived)Delete from key store; edge cache must expireCache TTL (often 1–5 min)Lower cache TTL → more key-store load
mTLS certificateLet it expire (short-lived) or CRL/OCSP= cert lifetime (hours) for short-lived certsCRL/OCSP infrastructure and its availability
Static DB passwordChange password; every consumer must reloadAs long as it takes to redeploy every consumerCoordination across teams — usually the slowest
Dynamic DB credentialRevoke the lease; DB user droppedSecondsSecrets-manager integration up front
Signing key (JWT issuer)Remove public key from JWKSJWKS cache TTL; all tokens signed by it die at onceMass logout if done abruptly

The pattern: short lifetimes are the revocation mechanism. A credential that lives 10 minutes needs no revocation infrastructure for most leaks. A credential that lives two years needs a fire drill.

Rotating a Signing Key Without Logging Everyone Out#

Day 0    Generate key K2. Publish K2's public key in JWKS alongside K1. Keep signing with K1.
         (Verifiers refresh JWKS every ~5 min; wait at least 2 refresh cycles.)
Day 0+1h Start signing new tokens with K2 (kid = K2). Old K1 tokens still verify.
Day 0+1h+max_token_lifetime
         No valid K1-signed access tokens remain. Remove K1 from JWKS.
Day 0+1d Destroy K1's private key (or archive under break-glass if audit requires).

The same publish → switch → retire overlap works for TLS certificates (serve new cert while old is still valid), KEKs (new DEKs wrap under the new KEK; old DEKs are re-wrapped lazily or by a background job), and database passwords (create a second credential, move consumers, drop the first).

Rotating an Encryption Key: What Actually Gets Re-Encrypted#

Rotated KeyWhat You Re-EncryptVolumeTime
KEK (in KMS)Every wrapped DEK — or nothing, if KMS keeps old key versions for decryptKilobytes to megabytesMinutes, or zero
DEK (per tenant/object)Every record encrypted under that DEKProportional to the dataBackground job: hours to days
Disk/volume keyUsually handled by the storage providerTransparentTransparent

This is why envelope encryption exists: scheduled rotation touches the KEK, which is cheap. DEK re-encryption is reserved for suspected compromise of a DEK itself.

The Emergency Rotation#

When a secret leaks (pushed to a public repo, printed in logs shipped to a vendor, found on a laptop), the clock matters. Automated scanners find public credentials within minutes of a push.

t=0      Secret scanning flags an AWS-style access key in a public commit.
t=+2min  Bot opens incident; key owner identified from the secrets inventory.
t=+5min  Key disabled (not deleted — keep it for forensics). New key issued via secrets manager.
t=+10min Services pick up the new key on their next secret refresh (≤ 5 min cache).
t=+30min Audit logs reviewed for any use of the old key from unknown IPs.
t=+1d    Post-mortem: why was a long-lived static key needed at all?

If step 3 requires a deploy across 12 services owned by 5 teams, the incident lasts hours. That's the real argument for runtime secret delivery: it makes emergency rotation a config change.

🎯 Staff Move: "I'll design rotation before I design storage. Access tokens live 10 minutes, so they need no revocation; refresh tokens are server-side and die on logout; service certificates live 24 hours and rotate at 12; the KEK rotates yearly with no data re-encryption. The only long-lived secret is the KMS root, and that never leaves the HSM."

When NOT to Reach for More Cryptography#

  • The threat is authorization, not secrecy. Encrypting a column doesn't help if the API happily returns it to the wrong tenant.
  • You'd build your own crypto or protocol. Use TLS, standard JWT/JWS libraries, the cloud KMS, and vetted envelope-encryption SDKs. Custom schemes fail in ways reviews don't catch.
  • The data is already public or low-value. Field-level encryption on product descriptions adds latency and key management for nothing.
  • You can't operate the keys. Customer-managed keys where the customer can revoke access at any time are a support and availability commitment, not just a feature checkbox.

Visual Guide#

Request Path Through the Trust Boundaries#

Diagram: Request Path Through the Trust Boundaries

Choosing a User Credential#

Diagram: Choosing a User Credential

Envelope Encryption Write and Read#

Diagram: Envelope Encryption Write and Read

Signing-Key Rotation Lifecycle#

Diagram: Signing-Key Rotation Lifecycle

Implementation Patterns#

Object-Level Authorization Everywhere#

The most common real-world API breach is not broken crypto; it's GET /invoices/8812 returning invoice 8812 to anyone logged in. Fix it structurally:

# Every data access goes through a repository that requires the principal
invoices.get(principal, invoice_id):
    row = db.query("SELECT ... FROM invoices WHERE id = $1 AND tenant_id = $2",
                   invoice_id, principal.tenant_id)          # tenant from the token, never the request
    if not row: return 404                                   # not 403: don't confirm existence
    authz.require(principal, "invoice:read", row)            # role / relationship check
    return row

Pair it with database-level guardrails (row-level security keyed on a session tenant variable) so a forgotten filter fails closed. Use unguessable IDs as defence in depth, never as the control.

Authorization Models#

ModelHow It DecidesGood ForBreaks When
RBACPrincipal has roles; roles have permissionsInternal tools, admin consolesPermissions depend on the specific object ("only my team's docs")
ABACPolicy over attributes of principal, resource, contextCompliance rules, data residency, time-of-day limitsPolicies sprawl and nobody can predict outcomes
Relationship-based (ReBAC)Graph of relations: user → member of → team → owner of → docSharing, nested folders, org hierarchiesGraph lookups add latency; needs a dedicated, consistent service

Centralize the decision engine and the policy language; keep enforcement in each service, at the object.

Passwords and Credential Stuffing#

Store passwords with a memory-hard hash (Argon2id, or bcrypt/scrypt), tuned to ~100–500ms per verification on your hardware. Rate-limit login per account and per IP, check new passwords against known-breached lists, and push MFA — preferably phishing-resistant (WebAuthn/passkeys). Attackers replay billions of leaked username-password pairs; your login endpoint's rate limits are the control that matters.

Threat Modelling in an Interview (5 Minutes, Out Loud)#

  1. Assets: what's worth stealing or breaking? (payment data, PII, account takeover, availability during a launch)
  2. Actors: anonymous internet, authenticated malicious user, compromised partner, insider, compromised dependency.
  3. Trust boundaries: draw them on the diagram — edge, service mesh, data stores, third parties.
  4. Top threats per boundary: use a mnemonic like STRIDE (spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege) to sweep, then pick the three that matter most.
  5. Controls and gaps: name the control for each top threat and say out loud what you are not defending against and why.

🎯 Staff Move: "The assets are card data and account balances. The likeliest attacker is a logged-in user poking at other users' objects, then credential stuffing at login. So my priorities are object-level authorization with tenant scoping, login rate limits with MFA, and keeping card data out of our systems entirely via a tokenizing processor. I'm not designing for a compromised cloud provider — that's accepted risk."

Audit Logs as a Security Control#

Log every authentication, authorization denial, privilege change, secret access, and key operation with principal, resource, decision and request ID. Ship to a store the application can't modify (append-only, separate account). Alert on patterns, not just events: one principal reading 10× its normal object count, many 404s on sequential IDs, secret reads from an unexpected identity.

Failure Scenario: The Missing Tenant Filter#

t=0       Deploy adds GET /v2/reports/{id}. The handler looks up by id only;
          the old v1 handler filtered by tenant in a helper the new code didn't use.
t=+3d     A customer's analyst iterates report IDs in a script to bulk-export their own reports.
          IDs are sequential. The script returns 41,000 reports — 38,000 belong to other tenants.
t=+3d+2h  Customer emails support: "Why can we see other companies' reports?"
t=+3d+3h  Incident declared. Endpoint disabled at the gateway.
t=+3d+6h  Access logs show 2 other principals made similar sequential requests.
t=+5d     Legal determines notification obligations for 140 affected tenants.

Detection: authz.cross_tenant_access (should be impossible — instrument it in the data layer); per-principal objects_read_per_hour anomaly alert; 404-rate spikes on ID-addressed endpoints. Blast radius: every tenant's reports, readable by any authenticated user for three days. Mitigation: disable the route; rotate nothing (no credential leaked) but review all access logs since deploy; notify affected customers. Prevention: tenant scoping enforced by the data-access layer and database row-level security, not by handler code; unguessable IDs; contract tests that call every endpoint as tenant B with tenant A's IDs. Owner: the reports team owns the fix; the platform team owns the data-access layer that should have made the bug impossible; security and legal own disclosure.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Missing object-level authZauthz.cross_tenant_access; sequential-ID scansEvery object of that type, all tenantsDisable route; enforce scoping in data layerOwning service + platform data layer
Leaked static secretSecret scanner hit; use from unknown IP in audit logWhatever the secret can reachDisable, reissue via secrets manager, review usageSecret owner (from inventory) + security
Certificate expirycert.expiry_seconds min < 25% of lifetimeEvery caller of the service, instantlyAutomated renewal; alert on renewal failuresPlatform PKI team
KMS / secrets manager outagekms.error_rate; startup failuresNew pods can't start; reads of encrypted fields failDEK/lease caching; multi-region keysPlatform security infra
Credential stuffingLogin failure rate ↑; many accounts per IPAccount takeoversPer-account and per-IP limits, breached-password check, MFAIdentity team
Over-privileged service credentialAccess-analyzer findings; unused-permission reportsData the service never neededLeast-privilege grants, per-service rolesService team + security review

The Numbers in Context#

NumberValueWhat It Means for Your Design
TLS 1.3 new connection1 RTT (TLS 1.2: 2 RTT)~1ms in-region, ~70–150ms cross-continent. Reuse connections and the cost disappears.
TLS 1.3 resumption0-RTT possibleEarly data is replayable — allow it only for idempotent requests.
AES-GCM with hardware supportMultiple GB/s per coreEncryption in transit is not a throughput concern; handshakes are.
ECDSA/EdDSA token verify~20–100µsLocal JWT verification is effectively free per request.
Session lookup (Redis)~0.5–1msInstant revocation costs about a millisecond.
Access token lifetime5–15 minBounds the damage of a stolen token without revocation infrastructure.
Refresh token lifetimeDays to ~30 days, rotated on useLong-lived, so server-side and revocable.
Session idle / absolute timeout~15–30 min idle / 8–24h absolute (stricter for banking)Product and compliance negotiate this — name who decides.
Workload certificate lifetimeHours to ~1 day, renew at 50%Leaked certs expire on their own; CA availability becomes critical.
Public TLS certificate (ACME-issued)90 days, renew at ~60Short lifetimes force automation, which is the point.
Dynamic DB credential TTL~1 hourLeaked creds die before most attackers use them.
Static secret rotation30–90 days if it can't be dynamicAnything longer is effectively permanent.
KMS call latency~5–20msNever per record. Cache DEKs with bounded lifetime and uses.
KMS regional request quotasThousands to tens of thousands per secondPer-record KMS calls hit quota before they hit latency budgets.
Password hash cost~100–500ms per verify (Argon2id / bcrypt)Slows offline cracking; also a DoS vector — rate-limit login.
AES-GCM nonce96 bits, never reuse with the same keyRandom nonces are safe up to ~2³² messages per key — another reason to rotate DEKs by usage count.

How This Shows Up in Interviews#

Scenario 1: "How does the mobile app authenticate to our API?"#

Don't stop at "OAuth" or "JWT." Say: "OIDC login with authorization code plus PKCE. The app gets a 10-minute access token scoped to our API's audience and a refresh token that's opaque, stored server-side and rotated on every use, with reuse detection — if an old refresh token is replayed, I revoke the whole family. Services verify access tokens locally against a cached JWKS. Logout or password change revokes refresh tokens, so the worst case for a stolen access token is 10 minutes."

Scenario 2: "Design how 200 microservices get database credentials and talk to each other securely." (Full Walkthrough)#

Step 1 — Name the threats. "The two I care about: a compromised pod using a shared credential to reach every database, and lateral movement across a flat network. Encryption in transit matters, but blast radius is the headline."

Step 2 — Give every workload an identity. "The platform issues each pod a 24-hour X.509 certificate with a SPIFFE ID derived from its namespace and service account, rotated at 12 hours. No static secret is needed to get it — the node agent attests the pod."

Step 3 — mTLS with identity-based policy. "All service-to-service calls go over mTLS through sidecars. Each service declares which caller identities may hit which endpoints; default deny. Moving from IP allow-lists to identities means autoscaling and rescheduling don't break policy."

Step 4 — Dynamic database credentials. "Services authenticate to the secrets manager with their workload identity and get a database user scoped to their own schema with a 1-hour TTL. No service can read another service's tables, and a leaked credential expires within the hour."

Step 5 — Plan for the platform failing. "CA and secrets manager are now tier-0. Both run multi-AZ; services cache credentials and keep existing connections when the secrets manager is briefly unavailable; certificate renewal starts at 50% of lifetime so a 6-hour CA outage doesn't cause expiry."

Step 6 — Roll out without an outage. "Permissive mode first — mTLS accepted but not required, with metrics on plaintext callers. Then enforce per namespace, starting with the least critical. Policies go shadow → log-only → enforce."

Step 7 — Owners. "Platform security owns the CA, secrets manager and mesh policy engine. Each service team owns its allow-list — they know their callers. Security reviews any policy that allows *."

Why this is a Staff answer: It leads with blast radius, replaces shared secrets with platform-issued identities, makes the security infrastructure's own failure modes explicit, and rolls out enforcement in stages with clear ownership.

Scenario 3: "A customer says they saw another customer's data. What happened and what do you do?"#

"Most likely an authorization bug, not a breach of crypto: a missing tenant filter, a cache keyed without tenant ID, or a CDN caching a personalized response. First, contain — disable the endpoint or purge the cache. Then scope — access logs since the last deploy that touched that path, per-principal read counts, which tenants were exposed. Then fix structurally — tenant scoping in the data-access layer and row-level security, cache keys that include the principal, Cache-Control: private on personalized responses. Legal and security own disclosure timelines."

Scenario 4: "We need to rotate our JWT signing key. How?"#

"Publish the new public key in the JWKS alongside the old one, wait two verifier refresh cycles, switch signing to the new key, keep the old public key until the longest-lived token signed with it has expired — 15 minutes for access tokens — then remove it. No one gets logged out. If it's an emergency because the old key leaked, I remove it immediately and accept that every access token dies; refresh tokens are opaque and unaffected, so clients silently get new ones."

Advanced Patterns#

PatternHow It WorksWhen to Use
Token exchange / downscopingEdge swaps the user's token for a narrow internal token per downstream callMany internal hops; avoid forwarding broad user tokens
Sender-constrained tokensToken bound to a client key (mTLS-bound or DPoP) so a stolen token alone is uselessHigh-value APIs, open banking, partner integrations
Refresh token rotation with reuse detectionEach refresh returns a new token; replay of an old one revokes the familyMobile and SPA clients
Crypto-shreddingDelete a tenant's KEK to make all their data, including backups, unreadableRight-to-erasure obligations; tenant offboarding
Customer-managed keys (BYOK)Tenant's KEK lives in their KMS account; they can revoke accessEnterprise and regulated customers — an availability commitment
Policy as codeAuthorization policies versioned, reviewed and tested like code; decisions loggedMore than a handful of services with shared authZ rules
Break-glass accessTime-boxed elevated access with approval and full auditProduction debugging without standing admin rights

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

A Staff engineer secures one system well: tight authorization, short-lived credentials, envelope encryption, a sensible threat model. A Principal engineer sees that the company's security posture is the sum of hundreds of local decisions — and that attackers pick the weakest one. Forty services use the paved road; six still have a shared database password from 2019; three teams built their own JWT validation, one without audience checks. No amount of review catches all of it. At L7, security becomes a property of the platform: identity, secrets, keys and authorization primitives that are easier to use correctly than incorrectly, plus inventories that answer "where is every credential, who owns it, and when was it last rotated?" in minutes rather than weeks.

🧭 Principal Move: "We can't review our way to security across 300 services. I'd make the secure path the only easy path: platform-issued workload identity, dynamic secrets, one token-validation library, tenant scoping in the data layer — and an inventory that tells us within an hour who owns any credential that leaks."

The Org-Level Fault Line#

Central security platform vs team-owned security.

OptionWhat WorksWhat BreaksWho Pays
Each team secures its own serviceAutonomy; local contextInconsistent token validation; shared static secrets; no inventorySecurity team during incidents; customers in breaches
Central security review gateExpert eyes on every designBottleneck; reviews catch designs, not drift; teams route around itSecurity team headcount; product velocity
Security platform + paved road (identity, secrets, KMS, authZ library, scanners)Secure by default; inventory for free; rotation automatedPlatform is tier-0; migration of legacy services is slowPlatform security team; legacy teams during migration

The Principal position: fund the paved road and make the review gate light for teams that use it and heavy for teams that don't. Central teams own primitives and inventories; product teams own their authorization rules, because only they know who may see what.

Cost Model#

Assumptions: engineer ~$25K/month fully loaded; managed KMS ~$1/key/month plus ~$3 per 100K requests (order of magnitude); secrets manager ~$0.40/secret/month managed or self-hosted on ~6 nodes; a significant data breach costs millions in response, notification, legal and churn.

ScaleSetupMonthly InfraHeadcountMain Risk Avoided
Small (10 services, 1 region)Managed identity provider, cloud secrets manager, KMS, TLS everywhere~$500–2K0 dedicated; a security-minded engineer part-timeShared admin passwords; IDOR on a launch endpoint
Medium (100 services, 3 regions)Above + service mesh mTLS, dynamic DB creds, secret scanning, authZ library~$10–30K (mesh overhead, scanners, KMS)3–5 platform security engineers (~$100K)Lateral movement; leaked long-lived keys; inconsistent token validation
Large (1,000+ services, global)Internal CA, policy engine, customer-managed keys, credential inventory, red team~$100–300K15–30 across platform security, AppSec, detectionBreach notification events; enterprise deal blockers (BYOK, audits)

The Principal argument is rarely "this prevents a breach" in the abstract; it's "this unblocks enterprise deals that require customer-managed keys and SOC 2 evidence, and it makes emergency rotation a 10-minute task instead of a week-long fire drill."

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal Cost
Token lifetimes, session timeoutsTwo-wayConfig change; product sign-off
mTLS permissive vs strict modeTwo-wayPolicy change per namespace
Session cookies vs self-contained tokens for a public APIMostly one-wayEvery client and SDK in the wild must change
Tenant ID model and where it's enforcedOne-wayEvery query, cache key and index carries it
Encryption key hierarchy (per tenant vs global DEKs)One-wayRe-encrypting all data to change granularity
Identity provider and user ID formatOne-wayEvery system stores user IDs; migrations take years
Collecting sensitive data at all (e.g., storing card numbers)One-way in practiceOnce stored, it's in backups, logs and compliance scope

The Standard I'd Write#

RFC: Service Security Baseline (v1)

Scope: Every production service and data store.

MUST:

  1. All traffic — external and internal — MUST use TLS 1.2+ (1.3 preferred); service-to-service traffic MUST use platform-issued mTLS identities with default-deny policies.
  2. Secrets MUST be delivered at runtime from the secrets manager. No secrets in source, images or long-lived environment variables. Database credentials MUST be dynamic where the engine supports it.
  3. User access tokens MUST expire in ≤ 15 minutes and MUST be validated with the platform library (signature, exp, aud, iss).
  4. Every data access MUST be authorized against the principal's tenant and the specific object, enforced in the data-access layer.
  5. Fields classified as sensitive MUST use envelope encryption with KMS-held KEKs; KEKs rotate at least annually.
  6. Every credential MUST be registered in the inventory with an owner and rotation date.

SHOULD: use sender-constrained tokens for partner APIs; run a threat model for any new external surface.

Exceptions: Time-boxed (≤ 90 days), approved by security, tracked in the inventory.

Success metrics: 100% of services on workload identity; median emergency rotation < 30 minutes; zero static database passwords; zero cross-tenant access findings in quarterly tests.

What I'd Tell the VP#

"Most security breaches don't come from clever hacking; they come from a leaked password or a bug that lets one customer see another's data. Today we have dozens of long-lived passwords spread across teams, and if one leaked, it would take us days to change it everywhere. We want to build shared tools that give every service short-lived credentials automatically and enforce customer data separation in one place. That's three to five engineers for a year. It cuts the time to contain a leak from days to minutes, and it's also what our enterprise prospects keep asking for in security reviews."

Principal Interview Signals#

SignalWhat It Sounds Like
Thinks in blast radius across the org"What can one leaked credential reach today? If the answer is 'every database,' that's the first fix."
Treats rotation speed as a metric"I want median time-to-rotate as a dashboard number, measured in drills, not estimated."
Makes the secure path the easy path"Teams shouldn't choose a JWT library. They should import one that's already right."
Ties security to business outcomes"Customer-managed keys unblock regulated customers; they also make availability depend on the customer's KMS. Sales and SRE both sign off."
Knows what to leave with teams"The platform enforces tenant scoping; teams own who may see what inside a tenant."

Staff answers that L7 interviewers find insufficient:

  • "This service uses mTLS and dynamic secrets." — Correct; silent on the 40 services that don't and how they get there.
  • "We'll do a security review before launch." — Reviews catch designs, not drift, and don't scale to hundreds of services.
  • "We encrypt everything at rest." — True, and irrelevant to the most likely attack: valid credentials or a missing authorization check.

How Real Companies Built It#

These are public, documented examples.

Google: ALTS and Identity-Based Service Authentication#

Google documents ALTS, its mutual authentication and transport encryption system for RPCs inside its infrastructure. ALTS authenticates primarily by identity rather than host — every person, machine and production service has an identity — which lets services be replicated, load-balanced and rescheduled across hosts without changing policy. Credentials are short-lived: Google describes machine handshake certificates rotated every few hours, workload base certificates about every two days, and human certificates valid for 20 hours (Google Cloud documentation).

Staff insight: Encryption is the easy part; who is this? is the valuable part. Saying "workload identity, not IP addresses, with certificates that live hours" is the sentence that separates a mesh design from "we turned on TLS."

Google: BeyondCorp#

Google's BeyondCorp paper describes removing the requirement for a privileged corporate intranet and moving internal applications onto the internet, with access decided by the identity and state of the user and device rather than network location (Google Research).

Staff insight: Perimeter security fails once one machine inside is compromised. The same logic applies inside a production network: "internal traffic is trusted" is a design assumption an interviewer will push on.

AWS KMS: Envelope Encryption#

AWS's KMS documentation defines envelope encryption as encrypting data with a data key and encrypting that data key under another key, with root keys that never leave its FIPS-validated hardware security modules unencrypted. It lists the benefits directly: the encrypted data key can be stored alongside the data, and re-encrypting under a different key means re-wrapping only the data keys rather than the raw data (AWS documentation).

Staff insight: Envelope encryption is what makes key rotation cheap and crypto-shredding possible. When an interviewer asks "how do you rotate the encryption key on 50TB?", the answer is "I don't — I rotate the key that wraps the data keys."

Let's Encrypt: 90-Day Certificates#

Let's Encrypt explains that it issues certificates valid for 90 days for two reasons: shorter lifetimes limit the damage from key compromise and mis-issuance, and they encourage the automation that makes certificate management reliable (Let's Encrypt).

Staff insight: Short lifetimes are a forcing function. A credential that must be renewed every few weeks can't be renewed by hand, so rotation becomes routine — and routine rotation is what makes emergency rotation fast.


Staff Calibration#

What Staff Engineers Say (That Seniors Don't)#

ConceptSenior (L5)Staff (L6)Principal (L7)
Transport"HTTPS at the load balancer""TLS everywhere; mTLS with workload identities so policy follows the service, not the IP""Workload identity is a platform primitive; services can't opt out, and certificate expiry is monitored fleet-wide"
Tokens"Use JWTs""10-minute access tokens verified locally, opaque rotating refresh tokens for revocation, aud checked""One validation library; token lifetimes are policy, not per-team choices"
Authorization"Check the user's role""Check this principal against this object and tenant on every request, enforced in the data layer""Tenant isolation is a platform guarantee tested continuously; teams own the rules within a tenant"
Secrets"Store them in environment variables""Runtime delivery from a secrets manager; dynamic DB creds with 1h TTL; emergency rotation is a config change""Every credential is inventoried with an owner; time-to-rotate is a measured SLO"
Encryption at rest"Enable disk encryption""Disk encryption for stolen hardware; envelope encryption with per-tenant DEKs for sensitive fields""Key hierarchy decided once for the company; BYOK offered where the revenue justifies the availability risk"
Why "Authorization" separates levels

Role checks are correct and common, which is why Seniors stop there. Staff candidates know most real exposures are object-level — a valid user, a valid role, someone else's record — and move enforcement to where the object is loaded. Principal candidates see that relying on each handler to remember is a statistical guarantee of failure across hundreds of endpoints and make isolation a platform property with continuous testing.

Why "Secrets" separates levels

"Use a secrets manager" is the right first step. Staff engineers judge a secrets design by how fast a leaked secret stops working. Principal engineers notice that nobody can answer "who owns this key?" quickly during an incident and build the inventory that makes the answer instant.

Common Interview Traps#

  • Saying "we'll use OAuth" as an authentication answer. OAuth is delegated authorization; say which flow, which token lifetimes, and how revocation works.
  • Long-lived JWTs. A 24-hour access token is a 24-hour breach window with no revocation.
  • Authorization only at the gateway. The gateway knows routes, not objects. Object checks belong where the object is loaded.
  • Trusting identity headers inside the network. Any compromised pod can set X-User-Id. Bind identity to mTLS or a signed token.
  • Treating encryption at rest as access control. It's transparent to anyone with the app's credentials.
  • Calling KMS per record. Latency and quotas; cache DEKs with bounds.
  • No rotation story. If rotating requires a deploy, it won't happen until an incident.
  • Skipping the threat model. Even 5 minutes — assets, actors, boundaries, top 3 threats, accepted risks — shows judgment.

Practice Drill#

Prompt: "We're a B2B SaaS with 2,000 tenants. Our biggest prospect, a bank, requires that their data be encrypted with a key they control and can revoke. Today we use disk encryption and one shared Postgres cluster. What do you propose?"

Staff Answer

First clarify what "control and revoke" must mean: the bank wants the ability to make its data unreadable to us by revoking our access to its key, and evidence that we can't bypass it. Disk encryption with our own keys doesn't meet that, and per-tenant databases for 2,000 tenants are an operational cost we don't want just for one customer. Proposal: application-level envelope encryption with per-tenant key hierarchy. (1) Each tenant gets a KEK; for most tenants it lives in our KMS account, for the bank it lives in their KMS account with a grant allowing our service role to wrap and unwrap. (2) Sensitive fields (documents, PII, financial records — classified with the bank) are encrypted with per-tenant DEKs, AES-256-GCM, with AAD binding ciphertext to tenant and row ID so values can't be swapped across rows. (3) DEKs are cached in memory for 5 minutes or 100K uses, so KMS sees roughly one call per tenant per service instance every few minutes — well inside quotas and adding ~10ms only on cache miss. (4) If the bank revokes the grant, cached DEKs expire within 5 minutes and their data becomes unreadable, including in backups — that's the guarantee they're buying, and it's also an availability risk: a misconfigured policy on their side is an outage for them. We document that in the contract and alert on kms.access_denied per tenant so support can tell them immediately. (5) Search and analytics on encrypted fields are limited — we can't index ciphertext — so we agree which fields stay queryable (e.g., IDs, dates) and which don't. (6) Rotation: their KEK rotates on their schedule with no data re-encryption; DEKs rotate on usage count. Rollout: build it for all tenants with our own KEKs first (two-way door), onboard the bank onto customer-managed keys after a game day where we revoke a test tenant's key and verify the blast radius. Owners: platform security owns the key hierarchy and KMS integration; the data team owns field classification and the search tradeoff; sales and legal own the contract language on availability when the customer revokes.

Why this is L6:

  • Translates "key they control" into a precise guarantee and its cost (availability, searchability).
  • Uses envelope encryption with caching so KMS latency and quotas don't dominate, and binds ciphertext to rows with AAD.
  • Stages the rollout as a two-way door before the one-way customer commitment, with a revocation game day and named owners.

What L7 adds:

  • Prices the feature: a few engineer-quarters to build versus the bank's contract value and the regulated-industry pipeline it unlocks; makes BYOK a priced tier rather than a one-off.
  • Sets the org's key hierarchy standard so every new service inherits per-tenant DEKs instead of retrofitting them.
  • Defines the operational contract across teams: who gets paged when a customer revokes, and how support distinguishes "customer revoked" from "we broke it."

Where This Appears#

  • Login & Sessions — Sessions vs tokens, refresh rotation, MFA, and credential stuffing
  • Edge Gateway — TLS termination, token validation and coarse authorization at the edge
  • Payments — Tokenization, keeping card data out of scope, and audit trails
  • Rate Limiter — Login and API abuse protection
  • Service Registry — Workload identity and who may discover and call whom
  • Object Storage — Encryption at rest, signed URLs, and per-tenant keys
  • Cloud File Sync — Sharing permissions and object-level authorization
  • Feature Flags & Config — Distributing configuration without distributing secrets

Related Foundations & Patterns: API Contracts That Age Well · Latency, Protocols & Tail Amplification · Graceful Degradation · Buy or Build

Related Technologies: Envoy, Kong & NGINX · Kubernetes · PostgreSQL · Redis

  1. Loading the index…