Hiring BarSupport

Design a Key Management Service

Case study89 min read9 diagrams

Technologies referenced in this case study: etcd & ZooKeeper · PostgreSQL · Distributed SQL · Apache Kafka · Kubernetes · Envoy, Kong & NGINX

Related: Security Fundamentals · Auth & Identity · Multi-Region · Consensus Service · Cascading Failures · Graceful Degradation · Multi-Tenancy · Photo Upload and Storage

Reading Guide#

Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal View) and the appendices on the key hierarchy, rotation and crypto-shredding
What is a Key Management Service? — Why interviewers pick this topic

A key management service (KMS) creates, stores, rotates and destroys the cryptographic keys that protect a company's data — and enforces who may use each key, for what, and records every use. Almost no data is encrypted by the KMS directly. Instead, applications encrypt data locally with a data encryption key (DEK) and ask the KMS to wrap that DEK under a key encryption key (KEK) that never leaves it. The wrapped DEK is stored next to the ciphertext. To read, the application asks the KMS to unwrap.

That one indirection — envelope encryption — is what makes the system interesting. The KMS is now on the read path of every encrypted byte in the company. If it's down, nothing decrypts. If it's slow, everything is slow. If a policy is wrong, either the wrong service reads customer data or the right service can't. And if a key is destroyed, every byte under it is gone forever — which is sometimes a disaster and sometimes exactly the point.

Before vs After — the "KMS blip" scenario:

Without DEK caching, static stability or per-caller quotas:
t=0:        KMS regional leader election takes 40s after a host failure.
t=+1s:      Every service that decrypts per request starts failing: 100% of
            unwrap calls time out. Checkout, login, messaging all return 500.
t=+10s:     Clients retry 3× with no jitter. KMS request rate: 60K/s → 400K/s.
t=+40s:     New leader elected; immediately saturated by the retry storm.
t=+6min:    On-call sheds traffic by hand. Company-wide outage: 9 minutes.

With envelope encryption, bounded DEK caches and a statically stable data plane:
t=0:        Same leader election.
t=+1s:      Services keep decrypting with cached DEKs (TTL 10 min, max 1M uses).
t=+1s:      KMS data-plane replicas keep unwrapping from replicated key material;
            only key creation and policy changes pause (control plane).
t=+40s:     Control plane recovers. Customer impact: zero. One ticket filed.

Why interviewers reach for this question: It looks like "store keys in an HSM, expose encrypt and decrypt" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: a tier-0 dependency whose availability bounds the whole company, a cache whose TTL is a revocation promise, a key hierarchy that decides what can be shredded, regional isolation that limits both outages and breaches, and an audit trail that has to be complete for keys used millions of times per second.

Mechanics Refresher: Key Management Primitives
PrimitiveHow It WorksProsCons
Direct encryption by the KMSSend plaintext to the KMS, get ciphertext backKey never leaves the KMS; simpleData size limits (often a few KB); every byte crosses the network; KMS on the hot path
Envelope encryptionLocal DEK encrypts data; KMS wraps the DEK under a KEKBulk crypto is local and fast; KMS sees only small keysPlaintext DEKs exist in app memory; caching decisions matter
Key hierarchyRoot key (HSM) → KEKs → DEKs; each level wraps the one belowRotating a KEK re-wraps kilobytes, not petabytesEvery level is a dependency of every read
HSM (hardware security module)Tamper-resistant device that holds keys and performs crypto insideKeys can't be exported in plaintext; certified boundaryLimited throughput; expensive; operationally heavy
Key versioningA key ID maps to several versions; one encrypts, all decryptRotation without rewriting dataOld versions live as long as any ciphertext
Encryption context / AADNon-secret attributes bound into the authenticated encryptionCiphertext can't be moved between records or tenants; auditableMust be supplied identically on decrypt
Crypto-shreddingDestroy a key so everything encrypted under it is unreadableErases backups and replicas without touching themIrreversible; must be scoped exactly
Audit logEvery key use recorded with caller, key, operation and contextDetection, forensics, complianceBillions of events per day; must be tamper-evident

For most production systems: Envelope encryption with AES-256-GCM DEKs generated per object or per record batch, wrapped by per-tenant or per-service KEKs held in a regional KMS whose root keys never leave HSMs; DEK caches bounded by time and usage; encryption context binding every wrapped key to its owner; a statically stable data plane that keeps unwrapping when the control plane is down; versioned KEKs rotated yearly without re-encrypting data; disable-before-destroy with a waiting period; and an append-only audit trail. The primitives are not the interview — the availability contract, the revocation promise and the blast radius of each key are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

A key management service is not a cryptography question. Everyone can say "AES-256 in an HSM".

It is a tier-0 dependency and blast-radius question that tests:

  • Whether you know the KMS sits on the read path of every encrypted byte, and design its availability above every service that depends on it
  • Whether you treat DEK caching as a tradeoff between availability and revocation latency — and state the revocation promise as a number
  • Whether your key hierarchy is chosen for what you'll need to rotate, revoke and shred, not for elegance
  • Whether you isolate regions and tenants so one compromised key, one bad policy push or one regional outage has a bounded blast radius
  • Whether access policy, separation of duties and audit are part of the design rather than a compliance footnote

The key insight: Every key is a tradeoff between availability (more copies, longer caches, broader scope) and blast radius (fewer copies, shorter caches, narrower scope). A KMS design is the set of places you chose where to sit on that line — per key type, per tenant tier, per region — and the numbers you promised. Senior candidates design the vault; Staff candidates design the line.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Keys in an HSM; services call encrypt and decrypt"Asks "Is this for data at rest across our services, customer-managed keys, or signing? What's the decrypt rate, and what happens to the company if this is down for a minute?"Asks "Which regulatory and customer commitments does this anchor — residency, customer key control, erasure — and which org owns them?"
Hot pathKMS encrypts every record"Envelope encryption: local AES-GCM with DEKs, KMS only wraps and unwraps; DEKs cached 5–10 minutes, bounded by uses"Sets a company-wide envelope SDK so no team calls the KMS per record
Availability"Run the KMS in three zones""Data plane is statically stable: it unwraps from replicated, already-loaded key material when the control plane is down; designed for higher availability than any dependent service"Puts the KMS in the org's tier-0 class with its own error budget, game days and dependency rules — it may depend on nothing but HSMs and its own storage
Rotation & revocation"Rotate keys every year by re-encrypting data""Version the KEK; new wraps use the new version, old versions decrypt; disable is a revocation that takes effect within the DEK cache TTL, which I'd publish"Decides revocation latency per product tier and prices it — shorter caches mean more KMS capacity
Blast radiusOne master key for everything"KEK per tenant per region, DEK per object; encryption context binds ciphertext to its owner; regions share nothing"Draws cell boundaries for keys like for compute; decides which data classes may use multi-region keys
Audit & policy"Log access""Every unwrap logged with caller identity, key, version and context; key admins can't use keys and key users can't change policy"Owns the evidence chain auditors and enterprise customers rely on; makes audit completeness an SLO
Why "hot path" separates levels

L5: "Services call the KMS to encrypt and decrypt each record." At 200,000 record reads a second, that's 200,000 remote calls, each adding 5–20 ms, all hitting HSM-backed capacity. The KMS becomes the most loaded and most critical service in the company, and request quotas — which managed services enforce per account and region — start throttling production.

L6: "Envelope encryption. The service generates or fetches a DEK, encrypts locally with AES-256-GCM — gigabytes a second per core — and stores the wrapped DEK next to the ciphertext. Reads unwrap via the KMS only on a cache miss. The cache is bounded: 10 minutes, a million uses, and keyed by the wrapped-key bytes plus encryption context, so it can't hand one tenant's DEK to another's request. KMS load becomes proportional to distinct keys in use, not to records read."

L7: "Per-record KMS calls are an org-level failure mode, not one team's bug. I'd ship the envelope SDK with sane cache defaults and make direct per-record calls fail review — the same way we don't let teams open raw sockets to the database."

Why "availability" separates levels

L5: "Deploy the KMS across three availability zones with a load balancer." That covers a zone failure. It doesn't cover a bad deploy, a leader election, a policy-store outage or a database migration in the KMS itself — any of which stops every decrypt in the region.

L6: "Split control plane from data plane. The control plane — create key, change policy, schedule deletion — can be strongly consistent and occasionally unavailable. The data plane — wrap, unwrap — runs on replicas that already hold the key material and policies they need, so if the control plane or its database is down, unwraps keep working with the last known state. Deploys roll one cell at a time. The KMS depends on nothing except HSMs and its own replicated storage — no shared service-discovery, no shared config service."

L7: "The KMS's availability is an upper bound on every dependent system's availability. If it's 99.95%, nothing that decrypts per request can be better. I'd set its target an order of magnitude above our best product SLO, fund it like the identity system, and run quarterly game days where we take a KMS region down on purpose."

Why "rotation and revocation" separates levels

L5: "We rotate the master key every year and re-encrypt all the data." Re-encrypting petabytes is months of I/O and a correctness risk, and it isn't what rotation is for. Meanwhile, nobody has said what happens when a customer revokes access.

L6: "Rotation creates a new KEK version; new wraps use it, unwraps pick the version recorded in the wrapped key. No data is rewritten — the KEK protects kilobytes of DEKs. DEK re-encryption is for suspected DEK compromise, not schedules. Revocation is different: disabling a key stops new unwraps immediately at the KMS, but cached DEKs keep working until their TTL — so the revocation promise is 'within 10 minutes', and I'd write that number down."

L7: "The revocation latency is a product commitment for customer-managed keys. A bank will ask 'if I pull my key, how fast does your service stop reading my data?' The answer — and the KMS capacity it implies — differs by tier, and I'd price it."

Positions to Commit To#

PositionRationale
Envelope encryption everywhere; the KMS wraps keys, it doesn't encrypt dataBulk crypto local and fast; KMS load scales with keys, not bytes
Root keys never leave HSMs; KEKs are stored only wrappedA database dump of the KMS is useless without the HSMs
Split control plane and data plane; the data plane is statically stablePolicy and key-creation outages must not stop decrypts
Bounded DEK caches (time and uses) — and the TTL is the published revocation latencyAvailability and revocation are the same knob; say the number
KEK per tenant per region; DEK per object or small batch; encryption context on every wrapBlast radius, shredding granularity and cross-tenant safety
Rotate by versioning KEKs; disable before destroy; destroy only after a waiting periodRotation is cheap; destruction is irreversible
Regions are independent by default; multi-region keys are an explicit, reviewed exceptionRegional isolation bounds outages, breaches and residency violations
Separation of duties and an append-only audit of every key useAdmins can't read data; readers can't change policy; every use is provable

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Internal data-at-rest KMS (every service encrypts its stores)Extreme read rate; tier-0 availability; thousands of callersEnvelope encryption, DEK caching, regional statically stable data plane, HSM-rooted hierarchyKMS outage = company outage; per-record calls throttleUnwrap availability above every dependent SLO; p99 unwrap < 10 ms in-region
Customer-managed keys for a SaaS (tenants control or hold their keys)Tenants can revoke at will; contractual revocation latency; per-tenant auditPer-tenant KEKs, optionally in an external key store the customer runs; short DEK caches; tenant-visible auditRevocation that doesn't revoke; an outage at the customer's key store becomes our outageRevocation effective within the published window; tenant sees every use
Signing key service (tokens, code, artifacts)Private keys must never leak; verifiers everywhereAsymmetric keys in HSMs; sign in the KMS, verify locally with published public keys; overlapping rotationLeaked signing key; verifiers can't fetch new public keysNo private key outside the HSM boundary; rotation without verification failures

🎯 Staff Move: "I'll design the internal data-at-rest KMS — every service in the company encrypting with envelope encryption — because that's where the availability and hot-path problems live. Customer-managed keys are a tier on top of the same hierarchy with shorter caches and tenant-visible audit; I'll cover them as a fault line. Signing is a different system — the hard part there is distributing public keys — and I'd keep it separate."

Where the Design Splits#

#Fault LineThe Tension
1Where Crypto Happens and How Long DEKs LiveRemote crypto (safe, slow, KMS on every read) vs envelope with caching (fast, available, slower revocation)
2Key GranularityOne key per service (few keys, huge blast radius) vs per tenant vs per object (precise shredding, key-count and KMS-load growth)
3Root of Trust: HSM per Operation vs HSM-Rooted HierarchyEvery operation in the certified boundary (throughput-limited) vs intermediate keys in KMS memory (fast, larger exposure)
4Regional Isolation vs Multi-Region KeysIndependent regional keys (contained outages and breaches) vs replicated keys (cross-region DR and reads, wider blast radius)
5Rotation, Revocation and DestructionWhat rotation means, how fast "disable" takes effect, and how to make "destroy" safe and final

How Real Companies Built It#

Why this section belongs here: The largest cloud providers document their key hierarchies and the tradeoffs they made between KMS load, audit granularity and availability. Naming those choices shows you understand the line between availability and blast radius.

AWS KMS — HSM-Bound Keys, Quorum Administration, Independent Regions#

The AWS KMS cryptographic details whitepaper describes a tier of web-facing KMS hosts in front of a fleet of FIPS 140-3 validated HSMs; KMS keys are available only on the HSMs and only in memory for the time needed to process a request, and there is no mechanism to export plaintext KMS keys. Administrative actions on the HSMs require multiple employees under quorum-based access controls, key usage can be isolated within a Region, and use and management of keys is recorded in CloudTrail (AWS KMS cryptographic details: introduction, design goals). Its developer guide states that automatic rotation (default every 365 days, configurable) changes only the current key material, retains all previous material so old ciphertext decrypts transparently, and does not re-encrypt data or rotate data keys; for multi-Region keys, new key material isn't used for encryption until it's available in the primary and every replica (key rotation). Deletion requires a waiting period of 7–30 days (default 30) during which the key can't be used and deletion can be cancelled; after that, data under the key is unrecoverable (deleting keys).

Staff insight: Four design decisions in one system: keys bound to the HSM boundary, humans gated by quorum, regions independent, and destruction slowed down on purpose. In an interview, say: "Rotation versions the KEK and never rewrites data; destroy is disable plus a waiting period, because it's the one operation in the system that can't be undone."

Google Cloud — A Four-Level Key Hierarchy With DEKs Stored Next to Data#

Google's documentation on default encryption at rest describes data split into chunks, each encrypted with its own DEK — two chunks don't share a DEK even for the same customer — using AES-256 by default. DEKs are wrapped by KEKs held centrally in an internal Keystore (one or more KEKs per service); wrapped DEKs are kept with the data chunks for low-latency access, while unwrapping happens centrally. KEKs are in turn wrapped by a Keystore master key, which is protected by a root keystore master key held in a peer-to-peer distributor that runs on dedicated machines in each data center and gossips between instances; the root key is also backed up on secure hardware stored in physical safes in multiple locations. Data is protected by a key set: one key active for encryption and historical keys active for decryption (Google Cloud default encryption at rest).

Staff insight: The hierarchy is designed around the read path: wrapped DEKs live beside the data so a read needs one central unwrap, not a lookup plus an unwrap, and the root of the hierarchy is designed to survive the loss of the services above it. Say: "Fine-grained DEKs, coarse KEKs per service, a tiny root — and the root's own recovery path is part of the design."

Amazon S3 Bucket Keys — Trading Audit Granularity for 99% Fewer KMS Calls#

Amazon S3's documentation explains that without Bucket Keys, SSE-KMS uses a separate data key per object and calls KMS on every request against an encrypted object. With S3 Bucket Keys, S3 obtains a short-lived bucket-level key from KMS, keeps it temporarily, and uses it to create data keys for new objects — reducing KMS request costs by up to 99%. The tradeoffs are documented: subsequent requests that use the bucket-level key don't make KMS API requests or validate access against the KMS key policy; CloudTrail logs the bucket ARN instead of the object ARN, with fewer events; the encryption context becomes the bucket ARN, so policies keyed on object ARNs stop working; and bucket-level keys are fetched at least once per distinct requester so each requester's access still appears in an audit event (S3 Bucket Keys).

Staff insight: This is the availability-versus-blast-radius line drawn in public. An intermediate cached key cuts KMS load by two orders of magnitude and costs per-object policy checks and per-object audit. In an interview, say: "Every cache in front of the KMS is a decision to check policy less often. I'll state what the cache scope is, how long it lives, and what audit granularity we give up."

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Services call the KMS to decrypt""At 200K reads a second, what does that do to latency and the KMS?"Envelope encryption, caching
"We cache DEKs""A customer revokes their key. How long until you stop reading their data?"Revocation latency as a contract
"The KMS runs in three zones""Your KMS database is down for 10 minutes. What still works?"Control vs data plane, static stability
"We rotate keys yearly""Do you re-encrypt the data? How long does that take for 40 PB?"Versioned KEKs vs re-encryption
"Each tenant has its own key""How many keys is that, and what does the KMS do per key?"Key count, granularity cost
"Keys replicate across regions for DR""A key's policy is wrong in one region. How far does that spread?"Regional isolation, blast radius
"We delete the key to erase a tenant""An engineer deletes the wrong key. What happens?"Waiting periods, disable-before-destroy

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Services encrypt and decrypt data locally with DEKs via an envelope SDK; they call the KMS only to generate or unwrap a DEK, and cache results with bounded lifetime and uses. The data plane authenticates the caller over mTLS, applies quotas and policy, and unwraps using KEKs that are themselves stored wrapped and opened with HSM-held keys. It reads keys and policies from a local replica, so a control-plane or database outage doesn't stop unwraps. The control plane owns key creation, rotation, policy changes and destruction, with quorum approval for irreversible operations. Every operation lands in an append-only audit log in a separate account. The metric that tells you the KMS is healthy is in-region kms.unwrap_p99 together with kms.dataplane_staleness — how old the data plane's view of keys and policies is.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Hot path"KMS encrypts each record""Envelope: local AES-GCM; KMS wraps DEKs; bounded DEK cache."
Availability"Three zones""Statically stable data plane; control plane can fail; no dependencies but HSMs and its own store."
Hierarchy"One master key""HSM root → regional domain key → KEK per tenant per region → DEK per object."
Rotation"Re-encrypt yearly""Version the KEK; old versions decrypt; no data rewritten."
Revocation"Disable the key""Disable stops new unwraps now; cached DEKs expire within the published TTL — 10 minutes."
Deletion"Delete the key""Disable, watch usage, quorum-approved destroy after a 7–30 day wait — then crypto-shred."
Audit"Log access""Every use with caller, key version and context; append-only, separate account; completeness is an SLO."

Numbers to Bring#

MetricValueWhy It Matters
AWS KMS symmetric crypto request quota10,000–100,000 req/s shared per account per Region (Region-dependent default)Per-record KMS calls hit this fast
AWS KMS custom key store quota1,800 req/s per key store, not adjustableCustomer-held keys are a throughput ceiling
AWS KMS direct Encrypt limitdata under 4 KBWhy bulk data uses envelope encryption
AWS KMS automatic rotationdefault 365 days; old material retained; no data re-encryptedRotation is cheap when keys are versioned
AWS KMS key deletion waiting period7–30 days, default 30Destruction slowed down on purpose
S3 Bucket Keysup to 99% fewer KMS requestsThe value — and cost — of an intermediate cached key
Remote KMS call~5–20 ms typical in-regionNever per record on a hot path
Local AES-256-GCM~1–5 GB/s per core with hardware AESBulk crypto is effectively free
AES-GCM with random 96-bit nonces≤ 2³² encryptions per key (NIST guidance)Bounds how long one DEK can be used
DEK cache boundstypically 5–15 min and 10⁵–10⁶ usesAvailability vs revocation latency
Audit volume50K unwraps/s ≈ 4.3B events/day; ~0.5–1 KB eachAudit is a big-data system
Illustrative HSM throughputthousands to tens of thousands of symmetric ops/s per HSMWhy HSMs hold roots, not every DEK operation

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on HSMs, AES modes and key ceremonies, then run out of time before the interviewer asks the questions that decide the level: "Your KMS is down for five minutes — what happens to the company?", "A customer revokes their key — how long until you stop reading their data?" and "Someone deletes the wrong key — now what?" Compress the cryptography to ~3 minutes and spend the rest on the hot path, availability, blast radius, rotation and destruction.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"A regional service that creates and manages keys, wraps and unwraps data keys for every service in the company, enforces who may use each key, rotates keys on a schedule, destroys them on request, and records every use. Services encrypt their own data locally; the KMS protects the keys."

Then the non-functional requirements, which is where the design lives:

"Four constraints drive everything. One: the KMS is tier-0 — every encrypted read depends on it, so its availability must exceed every service that uses it. Two: the hot path must not touch the KMS per record. Three: every key has a blast radius — how much data, how many tenants, how many regions — and we choose it deliberately. Four: some operations are irreversible, so destruction is slow and gated. I'll assume 2,000 calling services, 50,000 tenants, about 50,000 unwraps a second at peak after caching, 5,000 new DEKs a second, and three regions that must fail independently."

Then name the underspecified parts:

"I'd confirm: do tenants need to control their own keys? What's the revocation promise? Are there residency rules for keys? Does any data need crypto-shredding for erasure? I'll assume per-tenant keys, an enterprise tier with customer-managed keys, a 10-minute revocation window, keys resident in their region, and crypto-shredding as the erasure mechanism for backups."

🎯 Staff Move: Saying "the KMS's availability is an upper bound on every service that decrypts" in the first two minutes reframes the problem from "store keys safely" to "build the most available service in the company that also has the smallest blast radius". That tension is the interview.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Key (KEK): key_id, owner (tenant or service), region, purpose (wrap / sign / mac), state (ENABLED / DISABLED / PENDING_DESTRUCTION / DESTROYED), primary_version, rotation_period, policy_id
  • KeyVersion: key_id, version, wrapped_material (under the regional domain key), created_at, state
  • Policy: principals allowed per operation, required encryption-context keys, conditions (network, time), admin vs user separation
  • WrappedDEK (stored by callers, not the KMS): key_id, version, ciphertext, nonce, tag — opaque blob
  • AuditEvent: ts, caller, operation, key_id, version, context_hash, decision, request_id

Data-plane API (hot):

POST /v1/keys/{key_id}:generateDataKey   { context }            → plaintext_dek, wrapped_dek
POST /v1/keys/{key_id}:wrap              { plaintext_dek, context } → wrapped_dek
POST /v1/unwrap                          { wrapped_dek, context }   → plaintext_dek    (key_id read from blob)

Control-plane API (cold):

POST   /v1/keys                          { owner, region, purpose, rotation_period }
POST   /v1/keys/{key_id}:rotate          → new primary version
PUT    /v1/keys/{key_id}/policy          { … }                    (versioned, reviewed)
POST   /v1/keys/{key_id}:disable | :enable
POST   /v1/keys/{key_id}:scheduleDestroy { wait_days: 7..30 }     (quorum approval)
POST   /v1/keys/{key_id}:cancelDestroy

🎯 Staff Move: "The wrapped DEK carries the key ID and version, so unwrap doesn't need the caller to know which version encrypted it — rotation is invisible to callers. And unwrap requires the same encryption context used at wrap time — tenant ID and record type — so a wrapped key copied into another tenant's row fails to decrypt."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk one write and one read in 90 seconds:

  1. Write: the orders service needs to store a payment token for tenant t42. Its SDK calls generateDataKey(key=t42-kek, context={tenant:t42, table:payments}). The data plane checks policy, generates a 256-bit DEK, wraps it under the KEK's primary version and returns both. The service encrypts locally with AES-256-GCM and stores {wrapped_dek, nonce, ciphertext} in its row. It may reuse that DEK for more rows for a bounded window.
  2. Read: the service reads the row, looks up wrapped_dek in its DEK cache, misses, and calls unwrap(wrapped_dek, context). The data plane authenticates the caller over mTLS, checks the policy against the caller and context, unwraps the KEK version inside the trust boundary, unwraps the DEK, logs an audit event and returns the plaintext DEK. The SDK caches it for 10 minutes or a million uses and decrypts locally.
  3. KEK unwrap: the KEK version's material is stored wrapped under a regional domain key held in the HSMs. The data plane keeps recently used KEKs unwrapped in protected memory for a short window, so most unwraps don't touch an HSM at all.

🎯 Staff Move: Say out loud: "KMS load is proportional to distinct DEKs in use, not to records read. That's what lets one regional KMS serve the whole company." You've now spent ~8 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes this hard is that every decrypt in the company depends on it, every cache is a delay on revocation, every key has a blast radius, and destroy can't be undone. I'd like to go deep on availability and the cache contract, key granularity and the hierarchy, regional isolation, and rotation, revocation and destruction. Where would you like to start?"

If no preference: start with availability and caching. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Availability and the DEK cache contract (7–8 min)

"Two planes. The control plane — create, rotate, policy, destroy — writes to a strongly consistent store and can be unavailable for minutes without customer impact. The data plane serves wrap and unwrap from a local replica of wrapped KEKs and compiled policies; if replication stops, it keeps serving the last-known-good state up to a staleness bound, say 15 minutes, then fails closed for keys whose state it can't confirm. Deploys roll one cell at a time with automatic rollback on unwrap error rate. Callers cache DEKs for 10 minutes and a million uses, so even a full data-plane outage in a region is invisible for the first 10 minutes to any service with a warm cache."

Quantify: "50,000 unwraps a second at p99 10 ms is ~500 concurrent requests — small. The hard number is availability: if dependent services target 99.99%, the KMS data plane needs better than that, which means no shared dependencies and cell-by-cell deploys."

Who pays: "Revocation latency. Cached DEKs are the reason the KMS can blip without an outage, and they're also why a disabled key keeps working for up to 10 minutes. Security and product sign off on that number."


Deep dive 2: Key granularity and the hierarchy (6–7 min)

"Four levels. Root keys live in HSMs and are never exported. A regional domain key, also HSM-protected, wraps KEK versions stored in the key store. KEKs are per tenant per region — 50,000 tenants × 3 regions is 150,000 keys, trivial to store. DEKs are per object or per small batch, generated freely and stored wrapped beside the data. That gives crypto-shredding per tenant — destroy the KEK and every DEK under it is unrecoverable — without the KMS ever storing a DEK."

Quantify: "5,000 new DEKs a second is 432M wrapped DEKs a day, none stored by the KMS. A per-object DEK costs ~60–100 bytes of wrapped key per object; for 4 KB records, batch several records per DEK to keep overhead under 2%."


Deep dive 3: Regional isolation (5–6 min)

"Each region has its own HSMs, domain keys, key store and audit log; keys are created in one region and usable only there by default. A bad policy push, a corrupted key store or a compromised data-plane host affects one region. Data replicated across regions is re-wrapped under the destination region's KEK by the replication pipeline, which holds a narrow grant in both regions. Multi-region keys — the same key material in several regions — exist only for data classes that must decrypt anywhere during failover, with a review."


Deep dive 4: Rotation, revocation and destruction (5–6 min)

"Rotation adds a KEK version; new wraps use it, old versions stay for decryption; no data is rewritten. If a DEK is suspected compromised, re-encrypt the data under that DEK — that's an incident, not a schedule. Disable is immediate at the data plane and effective everywhere within the 10-minute cache TTL. Destroy is a two-step workflow: schedule with a 7–30 day wait, quorum approval, an alarm on any attempted use during the wait, then destruction of every version and every HSM backup copy of the material."


Deep dive 5: Policy, separation of duties and audit (3–4 min)

"Key administrators can manage keys but not use them; services can use keys but not change policy. Policy changes are versioned, reviewed and rolled out shadow → enforce. Every data-plane decision — allow or deny — emits an audit event with caller identity, key, version and a hash of the context, into an append-only, hash-chained log in a separate account. If the local audit buffer can't be written, the data plane fails closed for that operation: an unaudited decrypt is a decrypt we can't prove happened for a legitimate reason."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: the KMS is the most available service in the company and the one with the smallest blast radius, and those pull in opposite directions. Envelope encryption takes it off the per-record hot path; a statically stable data plane keeps decrypts working through control-plane failures; bounded DEK caches buy availability at the price of a published revocation window; per-tenant, per-region KEKs bound the blast radius and make crypto-shredding precise; versioned rotation is cheap; destruction is slow, gated and final."

The evolution closer:

"What I'd build later: customer-managed keys in an external key store for regulated tenants, a signing service on the same HSM fleet with public-key distribution, automatic detection of per-record KMS callers, and a key-usage inventory that tells us which data each key protects. What I'd not build: our own HSMs or our own cryptography."

🎯 Staff Move: End on who owns a destroyed key. "Destroy requires the key owner and a second approver, waits at least 7 days with alarms on any attempted use, and leaves an audit record that the auditor and the customer can both see. The KMS team owns the mechanism; the data owner owns the decision."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Crypto lecture10 min on AES modes and HSM certificationsOne sentence: AES-256-GCM, FIPS-validated HSMs, standard libraries
KMS on every readEncrypt/decrypt API per recordEnvelope SDK and bounded DEK cache in the first 8 minutes
"Three zones" as availabilityLoad balancer across zonesControl/data plane split, static stability, cell deploys
No revocation number"Disable the key""Effective within the 10-minute cache TTL"
Rotation by re-encryptionBatch job over all dataVersioned KEKs; re-encryption only for DEK compromise
Destroy as a deleteDELETE /keys/{id}Disable, waiting period, quorum, alarm on use, then destroy

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

A KMS is the interview where the most security-critical component is also the most availability-critical one. Every instinct that makes it safer — fewer copies, shorter caches, more checks, HSMs on every operation — makes it less available and slower; every instinct that makes it more available — replication, caching, broad keys — widens the blast radius of a compromise or a mistake. A Senior design picks one side. A Staff design picks a point per key class, states it in numbers, and assigns an owner to each number.

It also hides an irreversible operation inside an ordinary-looking API. In most systems, a bad delete is restored from backup. In a KMS, deleting a key is deleting every backup of every byte under it. The design has to make the one-way door slow, observable and hard to walk through by accident — while still letting it be walked through on purpose, because crypto-shredding is how erasure reaches backups.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"It's 3 a.m. An engineer reports they may have leaked the credentials of a service that can unwrap keys for 8,000 tenants. Walk me through the next hour — what can the attacker do, how do you stop it, and what does it cost the company to stop it?"

A candidate who answers with the scope (which keys that identity's policy covers, per region), the immediate containment (revoke the credential — the policy grants nothing to a revoked identity — rather than disabling 8,000 keys, which would take down every tenant), the exposure window (audit log query: every unwrap by that identity since the suspected leak, with contexts, so we know exactly which DEKs were obtained), the residual risk (DEKs already unwrapped by the attacker stay usable for any ciphertext they can also read, so the question becomes "did they also have data access?"), the remediation (re-encrypt only the data under the exposed DEKs, rotate the KEKs as hygiene) and the cost (hours of targeted re-encryption versus a company-wide key disable) has operated a KMS. A candidate who says "rotate all the keys" has read about one.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Internal data-at-rest KMS → availability and hot path

  • Constraint: thousands of callers, extreme aggregate read rate, tier-0 availability
  • Strategy: envelope SDK, DEK caching, statically stable regional data plane, HSM-rooted hierarchy
  • Failure mode: KMS outage or throttling becomes a company outage; per-record callers
  • Who pays for imperfection: every product team (outages), the KMS on-call (the most consequential pager in the company)

Customer-managed keys → revocation and evidence

  • Constraint: tenants can disable or delete their key at any time and expect us to stop reading their data within a contracted window; some hold the key material in their own infrastructure
  • Strategy: per-tenant KEKs, optionally backed by an external key store; shorter DEK caches for these tenants; tenant-visible audit; clear behavior when the tenant's key store is unreachable
  • Failure mode: revocation that doesn't revoke (long caches, derived plaintext copies); the customer's key store outage becomes an outage of our product for them
  • Who pays: the customer (they own their key's availability), our support team (explaining why their data is unavailable), sales (who promised the window)

Signing key service → private keys and public distribution

  • Constraint: signing keys must never leave the HSM boundary; verifiers are everywhere and must keep working through rotation
  • Strategy: asymmetric keys in HSMs; sign through the KMS; publish public keys (e.g. a JWKS endpoint) with overlap — new key published before use, old key retained until all tokens signed with it expire
  • Failure mode: verifiers that cache public keys too long reject new signatures; a leaked private key forges anything
  • Who pays: every service verifying tokens; security, if a signing key leaks

2.2 When NOT to Build a KMS#

  • You're on a major cloud. Use the provider's KMS and its HSM-backed keys. Build only the layer on top: the envelope SDK, per-tenant key provisioning, caching defaults, audit export and policy guardrails.
  • The data doesn't need application-level encryption. Disk and database encryption with provider-managed keys covers stolen media. Envelope encryption earns its keep when you need per-tenant keys, crypto-shredding, customer key control, or protection against someone holding database credentials.
  • You need a secrets store, not a key service. API tokens and database passwords belong in a secrets manager with leasing and rotation. A KMS protects keys that protect data.
  • You'd write your own cryptography or HSM firmware. Never. Use vetted libraries, standard AEAD modes and certified hardware.

🎯 Staff Insight: "In almost every company, 'design a KMS' should become 'design the key platform around a KMS we buy': the SDK, the hierarchy, the cache contract, the tenant model and the audit pipeline. That's where outages and breaches actually come from."

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Revocation latency — "disable the key" has no meaning until you say how long caches live
  • Who can destroy a key — and whether any single human can do it
  • Key granularity — one key per service, per tenant or per object changes shredding, load and blast radius
  • Regional and residency rules — whether a key may exist outside its region at all
  • Behavior during KMS outages — fail closed everywhere, or serve from cache for a bounded time
  • What happens to plaintext copies — caches, search indexes and logs that crypto-shredding won't reach

Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.

2.4 Precise Terminology#

TermWhat It MeansWhy It Matters in the Interview
DEKData encryption key; encrypts data locallyMany, short-lived, stored wrapped
KEKKey encryption key; wraps DEKs inside the KMSThe unit of policy, rotation and shredding
Root / domain keyTop of the hierarchy, held in HSMsWraps KEK material at rest
Envelope encryptionData under a DEK; DEK wrapped under a KEKTakes the KMS off the per-byte path
Encryption context (AAD)Non-secret attributes bound into authenticated encryptionPrevents moving ciphertext across tenants; auditable
Key versionOne generation of key material under a stable key IDRotation without rewriting data
Static stabilityData plane keeps working with last-known-good state when its dependencies failThe availability property that matters most
Revocation latencyTime from disable to the last successful use anywhereBounded by DEK cache TTL
Crypto-shreddingDestroying a key to make all data under it unreadableErasure that reaches backups
Separation of dutiesDifferent principals administer keys and use keysNo single role can both change and use access

🎯 Staff Insight: If the interviewer says "make it highly available", ask: "Available for what — unwrapping existing data, or creating keys and changing policy? I'd make the first more available than anything in the company and let the second be merely good."


3. Where the Design Splits#

Every key-management decision has a technical side (modes, hierarchies, caches) and an organizational side (who may destroy, what revocation means to a customer, who carries the tier-0 pager). Interviewers grade the second side.

3.1 Fault Line 1: Where Crypto Happens and How Long DEKs Live#

The tension: Sending data to the KMS keeps keys inside the boundary and puts the KMS on every read. Envelope encryption moves bulk crypto local and exposes plaintext DEKs in application memory. Caching unwrapped DEKs makes the KMS's own blips invisible — and makes revocation slower by exactly the cache lifetime.

ChoiceWhat WorksWhat BreaksWho Pays
KMS encrypts/decrypts the dataKeys never leave the KMSSize limits; KMS load = read rate; 5–20 ms per readEvery service (latency), KMS on-call (load)
Envelope, unwrap on every readRevocation is instantKMS load still = read rateKMS capacity
Envelope with bounded DEK cacheKMS load = distinct DEKs; survives blipsRevocation delayed by TTL; plaintext DEKs in memorySecurity (revocation window)
Envelope with long-lived cache (hours)Near-total KMS independenceRevocation takes hours; compromise window growsCustomers (weak revocation)
Diagram: 3.1 Fault Line 1: Where Crypto Happens and How Long DEKs Live

Staff default: "Envelope encryption with an SDK-managed DEK cache bounded at 10 minutes and a million uses, keyed by wrapped-key bytes plus encryption context. For tenants on customer-managed keys, the bound is shorter — 1–5 minutes — and that's what their contract says. Plaintext DEKs never leave process memory, are never logged, and are zeroed on eviction where the runtime allows."

When the KMS blips and every caller retries unwraps without backoff, the recovering service is buried under multiplied load — bounded caches and jittered retries keep a blip from becoming an outage.

When to deviate:

  • Highest-sensitivity keys (root signing, payment PIN keys): no caching; operations inside the HSM; low volume makes it affordable.
  • Batch analytics over millions of records: cache per job with a usage bound, and run in a separate quota class so it can't starve online traffic.

🧭 Principal Move: "The cache TTL is the single number that trades company availability against revocation speed. I'd make it a per-tier policy owned by security, published to customers, and enforced by the SDK — not a constant every team picks."

❌ Common L5 Trap: "Every read calls the KMS to decrypt the record, so keys never leave the KMS." At 200K reads a second that's a remote call on every read, a KMS sized for the company's total read rate, request quotas throttling production, and a company-wide outage the first time the KMS has a bad minute.


3.2 Fault Line 2: Key Granularity#

The tension: Few broad keys are simple and cheap — and one compromise, one policy mistake or one destruction affects everything under them. Fine-grained keys bound blast radius and make crypto-shredding precise — and multiply key counts, policies and KMS cache misses.

ChoiceWhat WorksWhat BreaksWho Pays
One KEK per serviceFew keys; high cache hit rateCan't shred one tenant; one key's compromise or misconfiguration hits all tenantsAll tenants (blast radius)
KEK per tenant per regionPer-tenant shredding, policy and audit; bounded blast radius10⁵–10⁶ keys; more distinct KEKs in hot memoryKMS team (key count)
KEK per tenant per data classShred "all messages" without "all billing records"Key count × classes; policy sprawlTenant admins (complexity)
DEK per recordSmallest exposure per DEKWrapped-key overhead; many unwraps on scansStorage (overhead), KMS (load)
DEK per object or batch, bounded usesGood balance; cache-friendlyBatch boundaries must match deletion unitsService teams (batching logic)
Diagram: 3.2 Fault Line 2: Key Granularity

Staff default: "KEK per tenant per region — the shredding unit matches the erasure unit, which is the tenant. Shared services that hold no tenant data use per-service KEKs. DEKs per object or per batch whose boundary matches what we'd ever want to delete together. Key count — 150,000 KEKs — is trivial for storage; what matters is the working set of unwrapped KEKs in data-plane memory, which I'd size for the active tenants in a 15-minute window."

When to deviate:

  • Per-user crypto-shredding (consumer products with account erasure): KEK per user is too many keys for some KMSs — use per-user wrapping keys, themselves wrapped by a per-shard KEK, stored in a dedicated key table that is excluded from general backups and kept only in its own short-retention backups. Shredding a user means deleting their row and waiting out that retention — state the total as the erasure SLA.
  • Small internal tools: a per-service key is fine; don't build tenant hierarchy for one tenant.

🧭 Principal Move: "Granularity is decided by the erasure unit and the breach unit. I'd ask legal what we must be able to erase independently, and security what one compromised key may expose, then pick the coarsest granularity that satisfies both."

❌ Common L5 Trap: "Each service has one master key that encrypts all its data." Fine until a customer asks you to prove their data is erased from backups, or until that key's policy is misconfigured — and every tenant is affected at once.


3.3 Fault Line 3: Root of Trust — HSM per Operation vs HSM-Rooted Hierarchy#

The tension: Performing every wrap and unwrap inside an HSM keeps all key material inside a certified tamper-resistant boundary — but HSMs are throughput-limited and expensive. Using HSMs only to protect the top of the hierarchy, with KEKs unwrapped into KMS host memory for short periods, scales — and widens the set of machines that hold plaintext key material.

ChoiceWhat WorksWhat BreaksWho Pays
Every operation in an HSMStrongest boundary; simple audit storyThroughput ceiling; HSM fleet cost; HSM outage = KMS outageFinance (HSM fleet), all callers (throughput)
HSM-protected root; KEKs unwrapped in host memory brieflyScales on commodity hosts; HSM load = KEK cache missesHosts hold plaintext KEKs; host hardening becomes criticalKMS team (host security)
Software-only keysCheap, fastNo hardware boundary; compliance often failsSecurity and compliance
Customer-held keys in an external key storeCustomer control and evidenceTheir availability and latency become oursCustomer (availability), us (support)

Staff default: "HSMs hold root and regional domain keys and perform KEK unwraps; data-plane hosts keep recently used KEKs unwrapped in locked memory for a few minutes, so HSM load is proportional to KEK cache misses, not to DEK operations. Data-plane hosts are a hardened, minimal fleet — no shell access, attested boot, no co-tenancy — and their compromise is the scenario the threat model is built around."

When to deviate:

  • Regulated key classes (payment HSM requirements, certain government workloads): operations stay inside the HSM, at lower throughput and higher cost, for that class only.
  • Customer-held keys: route through an external key store with a strict timeout and a clear failure mode — the tenant's data is unavailable when their key store is, and they agreed to that.

🧭 Principal Move: "I'd tier key classes by required boundary — HSM-every-op, HSM-rooted, external — and price each tier. Most data doesn't need the most expensive boundary, and pretending it does just makes the KMS slower for everyone."

❌ Common L5 Trap: "All keys are in the HSM, so every encrypt and decrypt goes to the HSM." At tens of thousands of operations a second, the HSM pool is both the bottleneck and the single point of failure, and the design has no answer when an HSM cluster is down for maintenance.


3.4 Fault Line 4: Regional Isolation vs Multi-Region Keys#

The tension: Independent regional keys mean a region's outage, bad deploy, policy mistake or breach stays in that region — and data replicated elsewhere must be re-wrapped, and a failover region can't decrypt data it never re-wrapped. Multi-region keys let any region decrypt anything — and copy every key's blast radius, and every mistake, to every region.

ChoiceWhat WorksWhat BreaksWho Pays
Strictly regional keysContained outages and breaches; residency by constructionCross-region replication must re-wrap; failover needs pre-wrapped copiesData platform (re-wrap pipeline)
Multi-region replicated keysAny region decrypts; simple DRPolicy and material replicated everywhere; residency harder; one mistake globalSecurity (blast radius)
Regional keys + DR escrow of wrapped KEKsRecovery path without live multi-region useEscrow process and access must be guardedKMS team (escrow process)
Diagram: 3.4 Fault Line 4: Regional Isolation vs Multi-Region Keys

Staff default: "Regional by default: each region has its own HSMs, domain keys, key store, data plane and audit log, sharing nothing. Data replicated across regions is re-wrapped by the replication pipeline under the destination's KEK, so the failover region can decrypt with its own KMS. Multi-region keys are a reviewed exception for data that must be decrypted in any region with no pre-wrapping, and residency-bound tenants can't use them at all."

When to deviate:

  • Global low-latency reads of the same encrypted object (e.g. a global configuration blob): a multi-region key is simpler than re-wrapping per region.
  • Disaster recovery of the KMS itself: escrowed, wrapped root material under quorum control — the recovery path is designed and rehearsed, not improvised.

🧭 Principal Move: "Keys get the same cell boundaries as compute. A regional KMS that shares a policy store, a deploy pipeline stage or an HSM vendor firmware rollout with another region isn't isolated — I'd audit those shared fates explicitly."

❌ Common L5 Trap: "Replicate all keys to all regions so any region can decrypt anything during failover." Now a policy mistake or a compromised key in one region is a compromise in every region, and EU tenants' keys exist in the US.


3.5 Fault Line 5: Rotation, Revocation and Destruction#

The tension: Rotation, revocation and destruction sound alike and are three different operations with different costs. Treating rotation as re-encryption wastes months; treating disable as instant ignores caches; treating destroy as delete risks irreversible data loss.

OperationWhat It Should MeanCommon MistakeWho Pays for the Mistake
RotateNew KEK version for new wraps; old versions decrypt; no data rewrittenRe-encrypting all data on a scheduleStorage and I/O budget; risk of corruption
Re-encryptRewrite data under new DEKs after suspected DEK compromiseNever doing it, or doing it for everythingSecurity (exposure) or ops (cost)
DisableStop new unwraps now; reversible; effective everywhere within cache TTLPromising "instant" revocationCustomers (false assurance)
DestroyPermanent removal of all versions and backups of key material, after a waitImmediate deletion by one personEveryone (unrecoverable data)
Diagram: 3.5 Fault Line 5: Rotation, Revocation and Destruction

Staff default: "Rotate KEKs yearly by adding versions — callers notice nothing. Re-encrypt only data under DEKs we believe exposed. Disable takes effect at the data plane within a second and everywhere within the cache TTL, and that's the number in our docs. Destroy is schedule-only: the key must be disabled first, two approvers, a 7–30 day wait, an alarm that pages the key owner if anything tries to use the key during the wait, and a final step that purges every version — including copies in HSM backups, which are easy to forget."

When to deviate:

  • Confirmed KEK compromise: rotate immediately, then re-wrap all DEKs under the new version (kilobytes per object, not data), then disable the old version after the re-wrap completes.
  • Erasure deadlines: the waiting period is part of the erasure SLA; size the SLA to include it.

🧭 Principal Move: "Destroy is the only operation in the platform that can't be undone, so I'd measure it like a safety system: number of destroys, number cancelled during the wait, number of attempted uses during the wait. A rising cancel rate means our tooling is making it too easy to schedule the wrong key."

❌ Common L5 Trap: "To meet the yearly rotation requirement, we'll re-encrypt all data with a new key every year." That's petabytes of reads and writes for a requirement that's satisfied by versioning the KEK — and it adds a yearly window where half the data is under each key and any bug corrupts records.


4. When It Breaks#

4.1 The Regional Blip That Became an Outage — Retry Storm on Unwrap#

t=0:       KMS data plane in us-east deploys a build with a memory leak in one cell.
t=+4min:   Cell hosts start OOM-restarting. Unwrap error rate: 0.01% → 35%.
t=+4min:   Services without DEK caches (per-record decrypt) fail immediately.
t=+4min:   Their clients retry 5× with no jitter. Unwrap request rate: 50K/s → 310K/s.
t=+6min:   Healthy cells saturate under the retry load. Error rate: 80%.
t=+9min:   Rollback starts. Recovering hosts are hit at full retry rate.
t=+17min:  Load shed by caller class; recovery. Checkout, login, messaging down 13 min.

Detection: kms.unwrap_error_rate{cell}, kms.requests_per_s{caller} vs baseline, sdk.cache_hit_rate by service.

Mitigation: roll back the cell; shed load by caller priority (tier-0 callers first); enforce per-caller token buckets so retry floods are throttled at the edge.

Prevention: cell-by-cell deploys with automatic rollback on unwrap error rate > 0.5%; mandatory envelope SDK with caching and jittered, capped retries; per-caller quotas with priority classes; a dependency rule that services must survive 10 minutes of KMS unavailability on a warm cache.

Owner: KMS team (deploy safety, quotas); service teams (SDK adoption); platform architecture (the dependency rule).

4.2 The Batch Job That Throttled Production#

t=0:       Analytics launches a backfill that decrypts 2B historical records,
           calling unwrap per record (DEK per record, no cache in its code path).
t=+3min:   Unwrap rate from analytics: 0 → 85K/s. Regional quota: 100K/s shared.
t=+4min:   Online services start receiving throttling errors on cache misses.
t=+6min:   Login p99 jumps from 80ms to 2.4s; some logins fail.
t=+15min:  KMS on-call identifies the caller, revokes its grant temporarily.

Detection: kms.throttled_by_caller, kms.requests_per_s{caller} top-N; tier-0 callers' error rates.

Mitigation: per-caller quotas: online tier-0 callers have reserved capacity; batch callers live in a separate, preemptible class.

Prevention: batch jobs use the SDK with per-job DEK caching and a batch quota class; datasets that are re-read in bulk use per-batch DEKs, not per-record; quota requests for new heavy callers go through capacity review.

Owner: KMS team (quota classes); analytics (job design).

4.3 The Wrong Key Scheduled for Destruction#

Day 0:     A tenant offboards. An engineer runs the offboarding script with the wrong
           tenant ID — a live tenant with 40 TB of data. Key scheduled for destruction, 30-day wait.
Day 0:     Key is now PENDING_DESTRUCTION: every unwrap for the live tenant fails
           once caches expire. Their product stops working within 10 minutes.
Day 0+12m: Alarm: "attempted use of key pending destruction" pages the key owner.
Day 0+25m: Destruction cancelled; key re-enabled. Impact: 15 minutes for one tenant.

Detection: alarm on any unwrap attempt against a key that is DISABLED or PENDING_DESTRUCTION; kms.denied_by_state{key}.

Mitigation: cancel destruction, re-enable — possible only because of the waiting period.

Prevention: destroy requires the key to have been DISABLED for at least 7 days with zero attempted uses; two approvers, one from the data owner; the offboarding workflow derives key IDs from the tenant's offboarding record, never from free-text input.

Owner: KMS team (guardrails), tenant lifecycle team (offboarding workflow).

4.4 Revocation That Didn't Revoke#

t=0:       An enterprise tenant disables their customer-managed key during a dispute,
           expecting our service to stop reading their data within 5 minutes (contract).
t=+5min:   KMS denies unwraps. Most services stop.
t=+6h:     Tenant's audit export still shows reads by the search service.
           Search caches DEKs for 24h (custom config), and its index holds plaintext snippets.
t=+1 day:  Contract breach review; legal escalation.

Detection: per-tenant audit of unwraps after disable (should be zero after TTL); periodic test tenant that disables its key and probes every service for successful reads.

Mitigation: force-evict the tenant's DEKs across SDK caches via a revocation broadcast; take the search index for that tenant offline.

Prevention: cache TTL is enforced by the SDK, not configurable per service above the tier maximum; plaintext derived stores (search indexes, caches) must either be encrypted under the tenant's key or registered as revocation consumers that purge on key.disabled.

Owner: KMS team (SDK enforcement, broadcast); search team (derived plaintext); product (the contracted window).

4.5 The Audit Gap#

The audit pipeline's collector falls behind during a traffic spike; the data plane's local buffer fills, and a configuration flag set years ago drops events when the buffer is full rather than failing closed. For 47 minutes, 9% of unwrap events are lost. Nobody notices until an enterprise customer's quarterly audit export shows a gap — and asks whether someone read their data without a record.

Detection: audit.lag_seconds, audit.dropped_events (must be 0), reconciliation of data-plane operation counters against audit event counts per minute.

Mitigation: replay from the data plane's durable local spool where available; disclose the gap with its scope.

Prevention: fail closed when the durable spool is full; spool sized for 1 hour of peak traffic; audit completeness — operations counted versus events stored — measured as an SLO.

Owner: KMS team (data plane), security (audit pipeline and completeness SLO).

4.6 Version Skew After Rotation#

A KEK is rotated in the primary region; the new version starts wrapping new DEKs immediately, but the replica region's key store hasn't received the new version yet. Data written in the primary and replicated within seconds can't be decrypted in the replica for 90 seconds; a failover drill during that window produces decrypt errors for fresh records.

Detection: kms.unwrap_unknown_version{region}, replication lag of key versions.

Prevention: a new version becomes primary only after it's confirmed present in every region that might unwrap it — the same rule documented for multi-region keys in managed KMSs — and regional keys avoid the problem entirely by re-wrapping on replication.

Owner: KMS team.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Data-plane cell failurekms.unwrap_error_rate{cell}Callers routed to that cell, after cache TTLCell rollback; retries to other cellsKMS team
Retry stormkms.requests_per_s > 3× baselineWhole regionPer-caller quotas, jittered SDK retriesKMS team + service teams
Batch caller throttles onlinekms.throttled_by_callerTier-0 services on cache missQuota classes, reserved capacityKMS team
Wrong key scheduled for destroyAttempted use of pending keyOne tenantCancel during waitKMS team + data owner
Revocation not honoredPost-disable unwraps or reads in auditOne tenant's contractSDK-enforced TTL, revocation broadcastKMS team + consumers
Audit lossaudit.dropped_events > 0Evidence for all tenantsFail closed on full spoolKMS team + security
Version skewkms.unwrap_unknown_versionFresh data in replica regionPromote version after full propagationKMS team
Control plane downkms.control_plane_errorsKey creation, policy changesData plane serves last-known-goodKMS team
HSM cluster degradedhsm.op_latency_p99, HSM errorsKEK cache misses in regionKEK memory cache; HSM pool headroomKMS team

🎯 Staff Insight: The KMS's health metric is in-region unwrap success for tier-0 callers, not overall success. "Denials are normal — policy is doing its job. I'd page on tier-0 unwrap errors that aren't policy denials, and on data-plane staleness, because a data plane serving 20-minute-old policy is the quiet failure that turns into a revocation breach."


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framing"Store keys securely and encrypt data"Frames the KMS as tier-0 with a blast-radius tradeoff per key; commits to an intentFrames which regulatory and customer commitments the key platform anchors and who owns each
Hot pathKMS call per recordEnvelope SDK; bounded DEK cache keyed by wrapped key + context; load = distinct DEKsCompany-wide SDK mandate; per-record calls fail review
AvailabilityMulti-AZ deploymentControl/data plane split; static stability; cell deploys; no shared dependencies; quotas by caller classTier-0 class with its own error budget and game days; dependency rules for every service
Hierarchy & blast radiusOne master keyHSM root → regional domain key → KEK per tenant per region → DEK per object; encryption context; regional isolationKey classes tiered by boundary and priced; cells for keys like compute
LifecycleRotate by re-encryptingVersioned rotation; disable with published latency; destroy with wait, quorum, alarms, HSM backup purgeRevocation latency per tier as a priced contract; destroy measured like a safety system
Policy & audit"Log access"Separation of duties; reviewed, staged policy changes; append-only audit; fail closed on audit lossAudit completeness SLO; evidence chain for auditors and customers

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Takes the KMS off the hot path"KMS load scales with distinct DEKs, not with reads."
Names the cache contract"The DEK cache TTL is the revocation latency — 10 minutes, published."
Designs static stability"If the control plane is down, unwraps keep working on last-known-good state."
Chooses granularity by erasure unit"KEK per tenant per region, because the tenant is what we're asked to erase."
Makes destroy slow"Disabled first, two approvers, a 7-day minimum wait, an alarm on any use."
Isolates regions"Replication re-wraps under the destination's KEK; regions share nothing."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
KMS encrypts every recordPuts a remote call and a quota on every read in the company
"Highly available" with no plane splitOne database or policy outage stops every decrypt
Instant revocation claimed with cachesFalse assurance to customers
Rotation means re-encrypting everythingMonths of I/O for a requirement versioning satisfies
Single-step key deletionOne mistake erases data irreversibly
Keys replicated everywhere by defaultGlobal blast radius; residency violations

5.4 Common False Positives#

  • Deep cryptography knowledge ≠ KMS design. Knowing GCM internals doesn't answer what happens when the KMS is down for five minutes.
  • "We use an HSM" ≠ a trust model. If KEKs are unwrapped into host memory, the hosts are part of the boundary; say how they're protected.
  • "We rotate yearly" ≠ a lifecycle. Rotation without revocation latency and a destruction workflow covers the easy third.
  • Compliance vocabulary ≠ audit. A FIPS level doesn't prove every unwrap was logged.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minInternal data-at-rest KMS; tier-0; blast radius; numbers
Entities & API3–5 minKey, version, policy, wrapped DEK, audit event; data vs control API
Architecture5–10 min≤ 8 boxes; envelope SDK, data plane, HSMs, control plane, audit
Availability & caching10–18 minPlane split, static stability, DEK cache bounds, quotas, retry behavior
Hierarchy & granularity18–25 minRoot → domain → KEK per tenant per region → DEK; encryption context
Regions & blast radius25–31 minRegional isolation, re-wrap on replication, multi-region exceptions
Lifecycle & audit31–38 minRotate, disable, destroy; separation of duties; audit completeness
Pivot / wrap38–45 minCustomer-managed keys, build vs buy, cost; close on who can destroy

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"KMS is down for 5 minutes"Static stability, cachingWarm caches cover it; data plane outlives control plane; quotas stop retry storms
"Customer revokes their key"Revocation semanticsImmediate at KMS; within TTL everywhere; derived plaintext stores must comply
"A service credential leaked"Blast radius, auditRevoke identity not keys; audit tells which DEKs; targeted re-encryption
"Rotate everything now"Rotation vs re-encryptionNew versions instantly; re-wrap DEKs if KEK compromised; re-encrypt only exposed data
"Erase tenant t42 including backups"Crypto-shreddingDisable, wait, destroy KEK; enumerate plaintext copies it won't reach
"Go multi-region"Regional isolationRegional keys; re-wrap on replication; multi-region keys only as exceptions

6.3 What to Deliberately Skip#

  • Cipher internals — "AES-256-GCM via a vetted library, random 96-bit nonces, bounded uses per DEK."
  • HSM certification details — "FIPS-validated HSMs" is enough.
  • Key ceremonies — mention quorum for root operations; don't script the ceremony.
  • PKI and certificates — a related but separate system unless asked.
  • Building HSMs or crypto — one sentence on why not.

6.4 Follow-Up Questions to Expect#

  1. "At 200K encrypted reads a second, what's the KMS load and latency impact?"
  2. "Your KMS's database is down for 10 minutes. What still works, and what doesn't?"
  3. "A customer disables their key. Exactly when does every service stop reading their data?"
  4. "How do you rotate keys yearly for 40 PB of data?"
  5. "An engineer schedules the wrong key for deletion. Walk me through it."
  6. "How do keys work when data replicates across three regions?"
  7. "A service credential that can unwrap keys for 8,000 tenants leaked. What now?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design a key management service for our company."

Staff Answer

"Is it for our services encrypting data at rest, for customers managing their own keys, or for signing tokens and artifacts? They're different systems. I'll assume the first, with customer-managed keys as a tier: about 2,000 calling services, 50,000 tenants, three regions, roughly 50,000 unwraps a second after caching and 5,000 new data keys a second.

Four constraints drive it. It's tier-0 — every encrypted read depends on it, so its availability must exceed any service that uses it. The hot path must never call it per record. Every key has a blast radius that we pick deliberately. And destruction is irreversible, so it's slow and gated. I'll go: entities and API → architecture with envelope encryption → availability and caching → hierarchy and granularity → regional isolation → rotation, revocation, destruction and audit."

Why this is L6:

  • Distinguishes three intents and commits, with numbers
  • Frames the KMS as tier-0 with a blast-radius tradeoff
  • Previews an outline that ends with lifecycle and audit, not with HSMs

What L7 adds:

  • Asks which regulatory and customer commitments the platform anchors — residency, customer key control, erasure — and who owns them
  • Asks whether we should build at all on a cloud provider, and what layer we own
❌ Common L5 Trap

"Keys are stored in HSMs. Services call encrypt and decrypt APIs with a key ID. Access is controlled by IAM, and we rotate keys every year."

Why this misses: Every piece is correct and the design still puts a remote call on every read, says nothing about what happens when the KMS is down, has no revocation number, and treats destruction as a footnote. The first follow-up — "200K reads a second" — has no answer.


Drill 2: The KMS Is Down#

Prompt: "Your KMS in one region is completely unavailable for five minutes. What happens?"

Staff Answer

"Depends on which part. If it's the control plane — key creation, policy changes — nothing user-facing happens: the data plane serves unwraps from its replica of keys and policies. If the data plane itself is down in a region, services with warm DEK caches keep decrypting for up to the 10-minute TTL; new writes can keep using their current DEKs within usage bounds. Only cache misses fail: new tenants, cold data, freshly started pods.

What must not happen is a retry storm: the SDK retries with jittered exponential backoff and a cap, and the endpoint enforces per-caller quotas so recovery isn't buried. If it lasts longer than the cache TTL, it's a real outage for affected callers — which is why the data plane deploys cell by cell and depends on nothing but HSMs and its own storage."

Why this is L6:

  • Separates control-plane and data-plane failure
  • Quantifies the grace period the cache buys and what still fails
  • Addresses the retry storm that turns a blip into an outage

What L7 adds:

  • Sets a company rule that every service survives 10 minutes of KMS unavailability on a warm cache, and tests it in game days
  • Notes that pod autoscaling during an incident creates cold caches exactly when the KMS is weakest — and plans for it
❌ Common L5 Trap

"We run the KMS in three availability zones behind a load balancer, so it won't go down."

Why this misses: Zone redundancy doesn't protect against a bad deploy, a leader election, a policy-store outage or a retry storm — the common causes. The design has no answer for what happens when it does go down.


Drill 3: Make It Concrete — Size the Data Plane and Audit#

Prompt: "Size the KMS data plane, the HSM load and the audit log."

Staff Answer

"Unwraps: services do maybe 2 million encrypted reads a second company-wide; with a 97–98% DEK cache hit rate, ~50K unwraps a second reach the KMS. Plus 5K generate-data-key calls. At a p99 of 10 ms, that's ~550 requests in flight per region at peak — a modest fleet; I'd run 3× headroom across cells.

HSM load: data-plane hosts cache unwrapped KEKs for 5 minutes. Active KEKs in a 5-minute window — say 20,000 tenants — means ~70 HSM unwraps a second at steady state, spiking on host restarts. A small HSM pool per region handles that with room for failover.

Audit: 55K events a second × 86,400 ≈ 4.8B events a day at ~600 bytes ≈ 2.9 TB/day raw, a few hundred GB compressed. One year of retention is ~100–150 TB compressed, partitioned by day and indexed by key and caller." (The back-of-envelope calculator helps check the arithmetic.)

Why this is L6:

  • Derives KMS load from cache hit rate, not from read rate
  • Shows HSM load scales with KEK cache misses
  • Sizes audit as a big-data system

What L7 adds:

  • Prices it: audit storage and query likely cost more than the data plane; tiered retention by data class
  • Turns cache hit rate into a per-service metric with targets, because one bad caller dominates load
❌ Common L5 Trap

"Two million reads a second means two million decrypts a second, so we need enough HSMs for two million operations a second."

Why this misses: It assumes per-record KMS calls and HSM-per-operation, sizing the most expensive hardware in the company for a load that envelope encryption and caching reduce by about two orders of magnitude.


Drill 4: Revocation#

Prompt: "An enterprise customer disables their key. When exactly does every one of our services stop reading their data?"

Staff Answer

"The KMS data plane denies new unwraps within about a second of the disable reaching it — the control plane pushes state changes to data-plane replicas with priority. Services already holding the tenant's DEKs keep decrypting until their cache entries expire. For customer-managed-key tenants, the SDK caps that at 5 minutes, enforced in the SDK, not configurable upward by services. So the contract says: 'within 5 minutes'.

The harder part is plaintext we derived: search indexes, analytics extracts, caches of decrypted objects. Those either store data encrypted under the tenant's key or subscribe to key.disabled and purge or lock the tenant's data. A synthetic tenant that disables its key daily and probes every service verifies the promise."

Why this is L6:

  • Gives a precise latency with its cause — the DEK cache bound
  • Enforces the bound centrally, not per service
  • Covers derived plaintext, which caches don't

What L7 adds:

  • Prices shorter windows — a 1-minute window means 5–10× more unwraps for those tenants — and offers it as a tier
  • Makes revocation verification a continuous, customer-visible control
❌ Common L5 Trap

"Disabling the key is instant, so they stop reading immediately."

Why this misses: It's instant at the KMS only. Cached DEKs and plaintext derived stores keep working — and the customer's audit export will show reads after the disable, which reads as a breach of trust.


Drill 5: The Leaked Credential#

Prompt: "A service credential that can unwrap keys for 8,000 tenants was leaked an hour ago. What do you do?"

Staff Answer

"Contain the identity, not the keys. Revoke the credential and rotate the service's identity — policies grant nothing to a revoked principal — rather than disabling 8,000 KEKs, which would take down every one of those tenants. Then query the audit log for every unwrap by that identity since the suspected leak time, with source addresses and contexts. That gives the exact set of DEKs the attacker may hold.

A DEK only matters if the attacker can also read the ciphertext, so the next question is data access: did the same leak include storage credentials? If yes, re-encrypt the data under those specific DEKs — targeted, probably hours of work — and rotate the affected KEKs as hygiene. If the audit shows unwraps from unexpected network locations, it's a confirmed breach with notification obligations per tenant."

Why this is L6:

  • Contains at the identity layer, avoiding self-inflicted outage
  • Uses the audit log to scope exposure precisely
  • Re-encrypts only what's exposed

What L7 adds:

  • Uses encryption context and network conditions in policy so a stolen credential is useless off the service's network
  • Turns per-tenant exposure reports into the customer notification workflow
❌ Common L5 Trap

"Rotate all 8,000 keys immediately and re-encrypt all their data."

Why this misses: Rotation doesn't revoke DEKs already unwrapped, re-encrypting everything takes weeks, and disabling keys to force it causes a self-inflicted outage for 8,000 tenants — while the leaked credential, if not revoked, keeps working.


Drill 6: Crypto-Shredding a Tenant#

Prompt: "A tenant leaves and requires proof that all their data, including backups, is erased. How?"

Staff Answer

"The tenant's data in every region is encrypted under their per-region KEKs. Erasure is: delete live data through normal paths; disable the KEKs; wait the destruction period — 7 to 30 days, with alarms on any attempted use; then destroy every version of the KEK material, including copies in HSM backups. After that, every backup, replica and snapshot holding their ciphertext is unreadable without being touched.

Crypto-shredding doesn't reach plaintext. So before destruction I'd run the inventory: search indexes, analytics extracts, logs, caches and exports that hold their data decrypted, each purged or encrypted under their key. The evidence package is the destruction audit records per KEK version, the inventory sign-offs and the absence of any post-disable unwraps."

Why this is L6:

  • Uses per-tenant KEKs as the shredding unit, across regions
  • Includes the waiting period and HSM backup copies
  • Names the plaintext copies shredding can't reach

What L7 adds:

  • Makes the erasure SLA include the waiting period, contractually
  • Requires every system storing tenant plaintext to register as an erasure consumer
❌ Common L5 Trap

"Delete their rows and their key. Backups expire in 90 days anyway."

Why this misses: Deleting the key immediately risks erasing the wrong tenant with no undo; "backups expire eventually" is not erasure proof; and decrypted copies in indexes and logs aren't covered at all.


Drill 7: Build vs Buy#

Prompt: "Should we build our own KMS or use our cloud provider's?"

Staff Answer

"Buy the KMS; build the key platform around it. A cloud KMS gives HSM-backed keys, regional isolation, rotation, a destruction waiting period and audit — things that take years and a specialized team to build and certify. What we build: the envelope SDK with cache defaults, per-tenant key provisioning, quota and caller-class management in front of provider limits, the revocation broadcast, and audit export into our own evidence store.

I'd build more only for specific reasons: customer-held keys in their own infrastructure, which the provider's external key store features may cover; multi-cloud with one key authority; or provider request quotas that our per-tenant model would exceed even with caching — in which case an intermediate key layer of our own, rooted in the provider's KMS, is the move, with the audit-granularity tradeoff made explicit."

Why this is L6:

  • Separates the undifferentiated vault from the product-specific platform
  • Names concrete triggers for building more
  • Keeps the root of trust in certified hardware either way

What L7 adds:

  • Weighs lock-in: the provider's KMS is a one-way door for ciphertext format and key export — plan the exit at design time
  • Prices provider request costs at 3-year volume, where an intermediate key layer may pay for a team
❌ Common L5 Trap

"We'll build our own with an open-source vault and some HSMs — we want control."

Why this misses: "Control" isn't a requirement. The candidate takes on HSM operations, certification, durability and a tier-0 pager without naming what the provider can't do.


Drill 8: Changing Policy Without an Outage#

Prompt: "Security wants every unwrap to require an encryption context with tenant ID. Today half the callers don't send one. How do you roll it out?"

Staff Answer

"Shadow first. Deploy the policy in report-only mode: every unwrap missing the context is allowed but logged as a would-deny, by caller. That gives a list of non-compliant services with volumes. The SDK's next release sends context automatically, so most callers fix by upgrading.

Then warn: would-deny events page the owning team's channel, with a deadline. Then enforce per caller, starting with the compliant ones and moving to stragglers as they upgrade — canary 1% of each caller's traffic before 100%. Old wrapped DEKs created without context still need to unwrap, so enforcement applies to new wraps first; legacy blobs are re-wrapped with context by a background job before unwrap enforcement for that data. Rollback is a policy version flip, propagated in seconds."

Why this is L6:

  • Shadow → warn → canary → enforce with per-caller data
  • Handles legacy ciphertext that can't satisfy the new rule
  • Keeps rollback fast and cheap

What L7 adds:

  • Uses the rollout as the template for every future policy change — a standard, not a project
  • Tracks policy coverage across the org as a security posture metric
❌ Common L5 Trap

"Update the key policy to require the context and tell teams to update their code."

Why this misses: Every caller that hasn't updated — and every legacy wrapped key without context — starts failing to decrypt the moment the policy lands. It's an outage scheduled by a policy change.


Drill 9: The Cost of Keys#

Prompt: "Our KMS bill and capacity are growing faster than traffic. Where's the cost, and how do you cut it?"

Staff Answer

"Measure unwraps per caller and cache hit rate per caller first — usually a handful of services drive most calls because they don't cache, cache too briefly, or use per-record DEKs on bulk reads. Fixing them is the biggest lever: moving a service from 50% to 98% hit rate cuts its calls 25×.

Then structure: per-batch DEKs for bulk datasets instead of per-record; an intermediate key — a short-lived per-service or per-bucket key derived from the KEK, the way managed object stores do with bucket-level keys — for high-volume, low-sensitivity data, accepting coarser audit and fewer policy checks for that class. Finally retention: audit storage is often the largest line; keep full detail for 90 days and aggregates beyond, except for data classes with contractual audit requirements."

Why this is L6:

  • Starts from per-caller measurement
  • Orders levers by impact
  • Makes the intermediate-key tradeoff explicit — fewer calls, coarser audit and policy

What L7 adds:

  • Charges KMS usage back to calling teams so cache behavior has an owner
  • Defines which data classes may use intermediate keys, so cost savings don't quietly weaken controls
❌ Common L5 Trap

"Cache DEKs for 24 hours everywhere — that cuts KMS calls almost to zero."

Why this misses: It does cut calls — and stretches revocation latency to a day for every tenant, including those with contracts promising minutes. It's a security decision disguised as a cost fix.


Drill 10: Multi-Region#

Prompt: "We're going active-active across three regions. How do keys work?"

Staff Answer

"Regional KMS stacks that share nothing: own HSMs, domain keys, key stores, data planes and audit logs. Each tenant gets a KEK per region it operates in. Data replicated across regions is re-wrapped by the replication pipeline under the destination region's KEK — it unwraps in the source region and wraps in the destination, with a narrow grant on each side — so every region decrypts its own copy with its own KMS and a regional KMS failure affects only that region.

Residency-bound tenants get KEKs only in their allowed regions, and the pipeline refuses to re-wrap their data elsewhere. Multi-region keys, where the same material exists in several regions, are a reviewed exception for data that must decrypt anywhere without pre-wrapping; new versions of those keys are used only after they exist in every replica."

Why this is L6:

  • Regional isolation with re-wrap on replication
  • Residency enforced at the key layer
  • Multi-region keys bounded to explicit exceptions with version-propagation rules

What L7 adds:

  • Audits shared fate — deploy pipelines, policy tooling, HSM firmware rollouts — that could break regional isolation
  • Decides per data class whether cross-region decryptability is worth the blast radius
❌ Common L5 Trap

"Use a single global KMS cluster replicated across all three regions so keys are the same everywhere."

Why this misses: It creates one global blast radius: a bad policy push, a corrupted key store or a compromise affects every region at once, and residency-bound keys now exist outside their region.


8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Cold Caches During a Scale-Out#

Context: A marketing event triples traffic in 10 minutes. Autoscaling adds 6,000 pods across services. Unwrap traffic to the regional KMS jumps from 50K/s to 410K/s, throttling starts, and checkout error rates hit 12%. The on-call escalates to you.

Questions to Surface First:

  • Is this organic cache-miss load from new pods, or a misbehaving caller?
  • Which callers are being throttled, and are tier-0 callers protected?
  • Are new pods unwrapping the same few DEKs thousands of times, or many distinct DEKs?
  • Is the KMS itself healthy, or is it quota enforcement doing its job?

Typical L5 Approach: Requests an emergency quota increase and scales the KMS data plane. Load keeps climbing as more pods start, and the scale-out takes 15 minutes.

Staff Approach: Sees that 6,000 new pods each unwrap the same ~200 hot DEKs on startup — 1.2M unwraps for 200 distinct keys. Prioritizes tier-0 callers in the quota classes, then ships a fix: a node-local DEK cache shared by pods on the same host, and jittered cache warm-up on pod start.

Principal Approach: Treats cold-start amplification as a platform property: every tier-0 dependency gets a capacity model for scale-out events, the KMS publishes its scale-out budget, and the autoscaling policy rate-limits pod creation when downstream tier-0 services report pressure.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Confirm KMS health (latency fine, throttling by quota). Raise priority of checkout and login caller classes; throttle batch classes to zero.
TriageDistinct DEKs requested ≈ 200; requests ≈ 1.2M over 5 min. Per-pod caches start empty.
Quick fixEnable node-local shared DEK cache sidecar (same TTL and bounds); stagger pod warm-up over 60s.
GuardrailsAlert on kms.unwraps_per_distinct_key > 50 per minute; quota headroom target 3× steady state.
Post-mortemLoad test KMS with a 3× scale-out of tier-0 services before every major event.

Metrics to Watch: kms.throttled_by_caller, kms.unwraps_per_distinct_key, sdk.cache_hit_rate{service}, checkout.error_rate

Organizational Follow-up: marketing events go on the capacity calendar; the KMS team gets advance notice and runs a pre-event test.

Ownership Question: "Who owns cold-start load on the KMS?" Staff answer: Calling services own their cache behavior through the SDK; the KMS team owns quota classes that protect tier-0 callers when caches are cold.

Key Takeaway: "A KMS sees your cache misses, not your traffic. Scale-outs turn every pod into a miss at once."

What clears the Staff bar:

  • Distinguishes distinct keys from request volume
  • Protects tier-0 callers with quota classes before scaling
  • Fixes the amplification structurally with shared caches and staggered warm-up

Deep Dive 2: Silent Failure — The Data Plane Serving Stale Policy#

Context: A security engineer removes a decommissioned service's permission to unwrap production keys. Two weeks later, an audit query shows that service still successfully unwrapped keys in one region every day since the change.

Questions to Surface First:

  • Did the policy change reach that region's data plane? When did replication stop?
  • What does the data plane do when it can't refresh policy — serve last-known-good forever, or for a bounded time?
  • Which other changes — disables, destroys, new denies — were also missed?
  • Why didn't any alert fire?

Typical L5 Approach: Pushes the policy change again and confirms the service is denied now.

Staff Approach: Finds that policy replication to one region silently stopped after a certificate rotation on the replication channel; the data plane, designed for static stability, served last-known-good policy with no upper bound. Fixes replication, then adds a staleness bound: beyond 15 minutes without a confirmed policy snapshot, the data plane denies requests for keys whose policy changed after its snapshot, and pages.

Principal Approach: Recognizes that static stability without a staleness bound converts availability into a security hole. Establishes the rule for every statically stable control system in the company — authorization, feature flags, keys: publish the maximum staleness, measure it, and define what happens past it.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Restart policy replication; confirm region's snapshot version matches control plane. Verify the decommissioned service is denied.
TriageReplication failed 15 days ago; data plane served snapshot v8812 since; 37 policy changes missed, including 2 disables.
Quick fixStaleness bound of 15 min; snapshot version exported as a metric; alert on mismatch across regions.
GuardrailsSynthetic policy change every 5 minutes; data plane must reflect it within 60s or page.
Post-mortemAudit every unwrap during the stale window against current policy; notify owners of keys whose disables were missed.

Metrics to Watch: kms.dataplane_staleness_seconds{region} (alert > 900), kms.policy_snapshot_version{region}, synthetic.policy_propagation_seconds

Organizational Follow-up: affected key owners notified; the replication channel's certificate added to the expiry-monitoring inventory.

Ownership Question: "Who owns policy propagation?" Staff answer: The KMS team owns it end to end, including a published propagation bound. Security owns the policy content.

Key Takeaway: "Static stability needs a staleness budget. 'Keep serving the last answer' is safe for minutes and dangerous for weeks."

What clears the Staff bar:

  • Identifies unbounded last-known-good as the root cause, not the cert
  • Adds a staleness bound with defined behavior beyond it
  • Audits the stale window for decisions that should have been denials

Deep Dive 3: Large-Customer Onboarding — A Bank Wants to Hold Its Own Keys#

Context: A large bank will sign only if its data is encrypted under keys held in its own HSMs, on its premises, with the right to cut access at any time. Their HSMs are reachable over a dedicated link with 15 ms latency and 99.9% contracted availability. Sales has signed. You're asked to make it work.

Questions to Surface First:

  • What does the bank expect our service to do when their key store is unreachable — fail, or serve from cache?
  • What revocation window will they accept?
  • What throughput can their HSMs sustain for our unwraps?
  • Which of our derived stores hold their plaintext, and how do those honor revocation?

Typical L5 Approach: Points the bank's KEK at their HSMs and lets every unwrap go there. Their HSMs and link become a hard dependency at our read rate; the first link blip takes down their tenant and floods their HSMs on recovery.

Staff Approach: Uses the bank's key as the top KEK for their tenant only, reached through an external key store proxy with strict timeouts. Their HSM wraps a per-tenant intermediate key that lives in our KMS memory for 5 minutes; DEKs are wrapped by that intermediate key. That bounds unwraps against their HSMs to a few per minute, gives them a 5-minute revocation window, and makes their availability the explicit dependency for their own tenant only.

Principal Approach: Turns "hold your own key" into a product tier with a written contract: availability of the tenant is bounded by their key store, revocation within the stated window, derived stores enumerated, and pricing that covers dedicated proxy capacity and support.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (design review)Load profile: their tenant ~3K unwraps/s after DEK caching → direct to their HSMs would exceed their stated 500 ops/s.
TriageIntermediate tenant key, wrapped by their HSM, cached 5 min in our data plane → ~1 call per data-plane host per 5 min.
Quick fixExternal key store proxy per region with 200 ms timeout and circuit breaker; failure = tenant unavailable, never another tenant.
GuardrailsPer-tenant dashboards of their key store latency and errors, shared with the bank; quarterly revocation drill together.
Post-mortem (pre-mortem)Their link is down for 2 hours: tenant unavailable after 5 minutes, by contract; no impact to others; recovery is gradual with jittered re-fetch.

Metrics to Watch: xks.latency_p99{tenant}, xks.errors{tenant}, tenant.unwrap_denied{reason=external_unavailable}, revocation.drill_seconds

Organizational Follow-up: legal and sales agree the contract language; support gets a runbook for "your key store is down" incidents.

Ownership Question: "Who owns the bank's tenant availability when their HSM is down?" Staff answer: They do, by contract. We own isolating that failure to their tenant and telling them within minutes.

Key Takeaway: "When a customer holds the root, their availability becomes their tenant's availability. Design it so it never becomes anyone else's."

What clears the Staff bar:

  • Shields the customer's HSMs from our read rate with an intermediate key
  • States revocation and availability semantics explicitly
  • Isolates the external dependency to one tenant

Deep Dive 4: Post-Mortem — A Key Destroyed That Shouldn't Have Been#

Context: Nine months ago a team destroyed a KEK they believed unused, after the waiting period passed with no alarms. Today, a restore from a long-retention backup for an unrelated incident fails: 3% of archived records were encrypted under DEKs wrapped by that KEK. They're unrecoverable. You're leading the post-mortem.

Questions to Surface First:

  • Why did the waiting period show no attempted use?
  • How did the team decide the key was unused?
  • What other destroyed keys might protect data in long-retention archives?
  • What does "unused" mean when ciphertext lives in backups for seven years?

Typical L5 Approach: Adds a longer waiting period.

Staff Approach: Finds that the waiting period detects attempted use, and archives are read once a year — so silence proved nothing. Introduces a key-usage inventory: every wrapped DEK creation is recorded with the dataset and its retention, so a key is destroyable only when every dataset it protects has passed its retention or been explicitly erased. Destroy approval now shows that inventory.

Principal Approach: Makes "which data does this key protect, and for how long?" a first-class question for the whole data platform. Destruction tooling consumes the data catalog, and every dataset's retention class names the keys it depends on.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Confirm destruction is final; scope affected records; check whether any plaintext copies exist elsewhere.
TriageWaiting period passed silently because archive reads are yearly; team's "unused" check looked at 30 days of audit.
Quick fixFreeze all scheduled destroys; review each against dataset retention.
GuardrailsKey-usage inventory from generateDataKey contexts (dataset, retention class); destroy blocked if any dataset is within retention.
Post-mortemAttempted-use alarms are necessary, not sufficient. Destruction requires positive evidence that nothing within retention depends on the key.

Metrics to Watch: kms.destroys_blocked_by_inventory, kms.destroy_cancelled_total, inventory.keys_without_dataset (must trend to 0)

Organizational Follow-up: data owners notified; the affected archive's restoration expectations corrected; legal reviews retention obligations for the lost records.

Ownership Question: "Who decides a key can be destroyed?" Staff answer: The data owner decides; the KMS enforces that the decision is backed by the usage inventory, and a second approver signs.

Key Takeaway: "Silence during a waiting period proves only that nothing tried. Destroy needs evidence of what the key protects."

What clears the Staff bar:

  • Spots that attempted-use alarms can't see infrequent readers
  • Builds a key-to-dataset inventory as the gating evidence
  • Freezes pending destroys before more data is lost

Deep Dive 5: Multi-Region Expansion — An EU Region With Resident Keys#

Context: The company opens an EU region and promises EU customers that their keys and data never leave the EU. Today all keys live in a US KMS, including keys for EU tenants' data replicated to the EU for latency.

Questions to Surface First:

  • Which EU tenants' KEKs exist in the US, and which data under them already lives in the EU?
  • Do any multi-region keys span the EU and US?
  • Where do audit logs for EU keys live?
  • How does replication between regions handle EU data?

Typical L5 Approach: Stands up an EU KMS and creates new keys for new EU tenants. Existing EU tenants stay on US keys; audit logs for EU key use still go to a US log account.

Staff Approach: Creates EU KEKs for every EU tenant, re-wraps their DEKs under EU KEKs — a key-level migration, not a data rewrite — and then destroys the US copies of those tenants' KEKs after the waiting period, so any US-resident backups become unreadable. EU audit logs move to an EU log account. Replication refuses to re-wrap EU tenants' data under non-EU keys.

Principal Approach: Defines key residency as a tenant attribute enforced by policy everywhere — KMS, replication, backups, analytics — and proves it with automated checks that the auditor and customers can see.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Inventory: US KEKs for EU tenants, multi-region keys spanning EU/US, US audit of EU key use, US backups of EU data.
Triage2.1B wrapped DEKs for EU tenants under US KEKs → re-wrap job reads and rewrites only the wrapped-DEK fields (~100 bytes each).
Quick fixRe-wrap tenant by tenant: unwrap in US, wrap in EU, verify, flip; then disable US KEKs; destroy after wait and inventory check.
GuardrailsPolicy denies creation of EU-tenant keys outside the EU; residency check scans key stores weekly.
Post-mortem (pre-launch)~200 GB of wrapped-key rewrites, not petabytes of data — schedule over 4 weeks with throttling.

Metrics to Watch: rewrap.tenants_completed, rewrap.verify_failures (must be 0), residency.eu_tenant_keys_outside_eu (must be 0)

Organizational Follow-up: legal approves the plan; EU customers receive a residency attestation once US copies are destroyed.

Ownership Question: "Who proves key residency?" Staff answer: The KMS team proves it for keys and audit logs with automated checks; compliance owns the customer attestation.

Key Takeaway: "Moving keys is cheap; moving data is expensive. Envelope encryption lets residency migrations rewrite kilobytes, not petabytes."

What clears the Staff bar:

  • Re-wraps DEKs instead of re-encrypting data
  • Uses KEK destruction to neutralize copies outside the region
  • Moves audit logs, not just keys

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain envelope encryption and why KMS load should scale with distinct DEKs, not reads
  • Design a control/data plane split with a statically stable data plane and a staleness bound
  • State DEK cache bounds as a published revocation latency, and enforce them centrally
  • Choose key granularity from the erasure and breach units, and build a four-level hierarchy
  • Isolate regions so keys, policies and audit share nothing, with re-wrap on replication
  • Distinguish rotate, re-encrypt, disable and destroy, with costs and safeguards for each
  • Design separation of duties, staged policy rollouts and an audit trail with a completeness SLO
  • Run crypto-shredding end to end, including the plaintext copies it can't reach

The Bar for This Question#

Mid-level (L4): Stores keys in a vault, encrypts data with them, restricts access. Correct cryptography, no system design.

Senior (L5): HSM-backed keys, encrypt/decrypt APIs, IAM policies, yearly rotation, multi-AZ deployment. The gap: the KMS sits on every read; availability is "three zones"; revocation is claimed instant despite caches; rotation means re-encrypting data; destruction is a single call; keys replicate everywhere.

Staff+ (L6): Frames the KMS as tier-0 with a blast-radius tradeoff per key within five minutes. Designs envelope encryption with bounded caches, a statically stable data plane with a staleness bound, a per-tenant per-region hierarchy, regional isolation, versioned rotation, gated destruction and complete audit. Names who pays — security for the revocation window, service teams for cache discipline, the data owner for destroy decisions. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 The KMS Should Be the Most Available Service You Run — and the Least Interesting#

PropertyWhy
Fewest dependenciesAnything it depends on becomes tier-0 too
Slowest feature velocityEvery change risks every decrypt
Deploys one cell at a timeA bad build must not reach a region at once

The Staff position: A KMS earns its reliability by doing very little, very predictably. Feature requests belong in the platform around it.

Why this matters in interviews: Candidates who add features to the KMS core instead of the SDK and platform are adding risk to tier-0.

10.2 Every Cache in Front of a KMS Is a Policy Decision#

CacheWhat You Give Up
DEK cache in servicesRevocation latency = TTL
KEK cache in data planePlaintext KEKs in host memory
Intermediate per-bucket keysPer-object policy checks and audit

The Staff position: Caching is mandatory for availability and cost; state what each cache costs in revocation, exposure and audit.

Why this matters in interviews: "We cache it" is a Senior answer; "we cache it for 10 minutes, and that's our revocation promise" is Staff.

10.3 Rotation Is Overrated; Revocation and Destruction Are Underrated#

OperationHow Often It's DiscussedHow Often It Breaks Things
Scheduled rotationConstantlyRarely, if versioned
RevocationRarelyEvery time caches are forgotten
DestructionAlmost neverCatastrophically, once

The Staff position: Versioned rotation is a solved problem. Spend design time on revocation latency and destruction safeguards.

Why this matters in interviews: Interviewers ask about rotation to see whether you'll talk about the harder two.

10.4 Static Stability Without a Staleness Bound Is a Security Bug#

The Staff position: "Keep serving last-known-good" is the right availability choice for minutes and the wrong security choice for days. Publish the bound and define behavior past it.

Why this matters in interviews: It shows you see availability and security as one tradeoff, not two teams' problems.

10.5 Crypto-Shredding Is Only as Good as Your Plaintext Inventory#

The Staff position: Destroying a key erases ciphertext. Search indexes, caches, logs and analytics extracts hold plaintext the key never touched. Erasure claims require the inventory.

Why this matters in interviews: "Delete the key" is the beginning of the erasure answer, not the end.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer builds an available, isolated, auditable KMS. The Principal engineer notices that encryption decisions are being made by 200 teams independently — some call the KMS per record, some cache keys for a day, some store plaintext in search indexes, some use one key for every tenant — and that the company's commitments to customers (revocation windows, residency, erasure) depend on all of them. The L7 problem is a key platform with enforced defaults: one SDK, one cache contract per tier, one hierarchy, one destruction workflow, one evidence chain — so that customer commitments are properties of the platform, not hopes about each team.

The Org-Level Fault Line#

One key platform with enforced defaults vs per-team encryption choices.

OptionWhat WorksWhat BreaksWho Pays
Each team uses the KMS however it likesFast adoption; flexibilityPer-record callers, unbounded caches, inconsistent granularity; commitments unprovableSecurity (exposure), KMS on-call (load), sales (broken promises)
Key platform with SDK, tiers and guardrailsCommitments enforced centrally; load predictablePlatform team scope; SDK must cover every language and runtimePlatform team (scope), service teams (migration)
Mandated platform plus exceptions processCovers the long tail of special cases safelyExceptions need review capacitySecurity (reviews)

🧭 Principal Move: "The key platform owns the SDK, cache tiers, hierarchy and lifecycle. Teams choose a data class; the class determines key granularity, cache bounds and audit detail. Direct KMS calls outside the SDK are flagged in review and blocked for new services after two quarters."

Cost Model#

Assumptions: managed KMS with per-request and per-key pricing at illustrative list rates; HSM-backed custom key stores where required; audit in a columnar store; fully loaded engineer ~$250K/year.

ScaleVolumeInfra ($/month)HeadcountOn-call LoadNotes
Startup50 services, 1K tenants, 500 unwraps/s~$1–5K (managed KMS requests, keys, audit)1 eng part-timeShared rotationBuy everything; write the SDK wrapper
Growth2K services, 50K tenants, 50K unwraps/s, 3 regions~$40–120K (requests ~40%, audit ~40%, keys and HSM stores ~20%)4–6 engDedicated tier-0 rotation, rare pagesCache hit rate is the biggest lever
Enterprise10K services, 500K tenants, 300K unwraps/s, customer-held keys~$300K–1M12–20 eng (platform, SDKs, external key stores, audit)Per-region tier-0 rotationsIntermediate key layer and own data plane may pay for itself

The pricing insight: request volume is driven by cache behavior of a few callers, and audit is often as expensive as the KMS itself. A 2-point improvement in company-wide cache hit rate (96% → 98%) halves request cost; tiered audit retention by data class is the other large lever.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Destroying any keyOne-wayData under it is gone forever
Wrapped-DEK and ciphertext formatOne-way-ishEvery stored object carries it; changes need dual-read support for years
Key granularity (per service vs per tenant)One-way-ishChanging it means re-wrapping every DEK and re-modeling policy
Provider KMS as root of trustOne-way-ishKeys generally can't be exported; exit means re-wrapping everything
Published revocation windowOne-way-ishCustomers contract against it; shortening costs capacity
DEK cache TTL within the published boundTwo-waySDK config
Rotation periodTwo-wayPolicy config
Audit retention tiersTwo-wayStorage policy, within legal minimums

The Standard I'd Write#

RFC-KEYS-001: Data Encryption and Key Lifecycle Standard
Status: Approved   Owners: Key Platform + Security

Scope
  Every service that encrypts customer or confidential data at rest.

MUST
  1. Encrypt through the Key Platform envelope SDK; no per-record KMS calls.
  2. Declare a data class; the class sets key granularity, DEK cache bounds
     and audit detail. Customer-managed-key tenants use the CMK class.
  3. Bind every wrapped key to an encryption context containing tenant ID
     and dataset.
  4. Register every plaintext-derived store (index, cache, extract) as a
     revocation and erasure consumer.
  5. Destroy keys only through the destroy workflow: disabled ≥ 7 days,
     usage inventory clear, two approvers, 7–30 day wait.

SHOULD
  1. Survive 10 minutes of KMS unavailability on a warm cache.
  2. Use per-batch DEKs for bulk datasets.
  3. Keep keys regional; request multi-region keys through review.

Exceptions
  Filed with Key Platform; security review required; time-boxed to two quarters.

Success metrics
  - Tier-0 unwrap availability: above every dependent service's SLO
  - Revocation drills meeting the published window: 100%
  - Audit completeness (operations vs stored events): 100%
  - Services calling the KMS outside the SDK: 0 by end of year 2

What I'd Tell the VP#

"Every encrypted byte in the company depends on one service, and today 200 teams use it 200 ways — some in ways that throttle production, some in ways that break the promises we make to customers about revoking access and erasing data. I'm proposing a key platform: one SDK, defaults set by data class, and a safe workflow for the one operation that can permanently destroy data. It's about five engineers for a year. It removes a company-wide outage risk, makes our revocation and erasure promises provable to auditors and enterprise customers, and should cut KMS and audit spend by a third through better caching. The main risk is migration effort, so I'd start with the ten services that make 80% of the calls."

Principal Interview Signals#

SignalWhat It Sounds Like
Treats commitments as platform properties"Revocation windows and erasure are enforced by the SDK and workflow, not by each team."
Identifies one-way doors"Destroy, ciphertext format and root of trust are forever; cache TTL isn't."
Prices availability vs blast radius"A 1-minute revocation tier costs 5–10× the unwraps; it's a priced option."
Redraws ownership"Key platform owns mechanism; data owners own destroy decisions; security owns policy."
Knows when not to build"We buy the vault and build the platform; our own data plane only when requests cost more than a team."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a highly available KMS" — correct, but ignores the 200 teams using it inconsistently.
  • "Revocation takes effect within 10 minutes" — good, but no enforcement for services that cache longer.
  • "Destroy has a waiting period" — no usage inventory, so silence during the wait proves nothing.

Appendices

Appendix A: Mechanics in Depth#

A.1 Envelope Encryption in the SDK#

encrypt(record, tenant, dataset):
  ctx = {tenant, dataset}
  dek = active_dek(ctx)                          # reuse within bounds
  if dek is None or dek.uses >= MAX_USES or dek.age >= MAX_AGE:
      plaintext_dek, wrapped = kms.generateDataKey(kek_for(tenant, region), ctx)
      dek = cache.put(wrapped, ctx, plaintext_dek, ttl=class.ttl, max_uses=class.max_uses)
  nonce = random(12)                             # 96-bit random nonce
  ct = aes_gcm_encrypt(dek.key, nonce, record, aad=serialize(ctx))
  dek.uses += 1                                  # stay far below 2^32 per key
  return {wrapped: dek.wrapped, nonce, ct}

decrypt(blob, tenant, dataset):
  ctx = {tenant, dataset}
  dek = cache.get(blob.wrapped, ctx)             # keyed by wrapped bytes + context
  if dek is None:
      dek = cache.put(blob.wrapped, ctx, kms.unwrap(blob.wrapped, ctx), ttl=class.ttl)
  return aes_gcm_decrypt(dek.key, blob.nonce, blob.ct, aad=serialize(ctx))

A.2 The Data-Plane Unwrap#

unwrap(wrapped, ctx, caller):
  key_id, version = parse_header(wrapped)
  snap = local_snapshot()
  if now - snap.confirmed_at > STALENESS_BOUND and snap.changed_after(key_id):
      deny("stale_policy"); page()
  key = snap.keys[key_id]
  require key.state == ENABLED
  require policy_allows(snap.policy[key_id], caller, "unwrap", ctx)
  require quota.take(caller.class)
  kek = kek_cache.get(key_id, version) or hsm.unwrap(domain_key, key.versions[version])
  dek = aead_open(kek, wrapped, aad=ctx)          # fails if ctx differs from wrap time
  audit.spool(caller, "unwrap", key_id, version, hash(ctx), "allow")   # fail closed if spool full
  return dek

A.3 Destroy Workflow#

schedule_destroy(key_id, requester, approver, wait_days in 7..30):
  require key.state == DISABLED for >= 7 days
  require attempted_uses_since_disable(key_id) == 0
  require usage_inventory(key_id).all(dataset -> retention_expired or erased)
  require approver != requester and approver in data_owner_group(key_id)
  key.state = PENDING_DESTRUCTION; key.destroy_at = now + wait_days
  alarm_on_any_use(key_id) → page key owner
on destroy_at:
  purge all versions from key store, data-plane caches, HSM partitions and HSM backups
  audit("destroy", key_id, versions, approvers)

Appendix B: Data Model#

CREATE TABLE keys (
  key_id          TEXT PRIMARY KEY,
  owner           TEXT NOT NULL,            -- tenant or service
  region          TEXT NOT NULL,
  purpose         TEXT NOT NULL,            -- wrap | sign | mac
  data_class      TEXT NOT NULL,            -- sets cache bounds and audit detail
  state           TEXT NOT NULL,            -- ENABLED | DISABLED | PENDING_DESTRUCTION | DESTROYED
  primary_version INT NOT NULL,
  rotation_days   INT NOT NULL DEFAULT 365,
  disabled_at     TIMESTAMPTZ,
  destroy_at      TIMESTAMPTZ,
  policy_version  BIGINT NOT NULL
);

CREATE TABLE key_versions (
  key_id           TEXT NOT NULL REFERENCES keys,
  version          INT NOT NULL,
  wrapped_material BYTEA NOT NULL,          -- under the regional domain key (HSM)
  created_at       TIMESTAMPTZ NOT NULL,
  state            TEXT NOT NULL,           -- ACTIVE | DECRYPT_ONLY | DESTROYED
  PRIMARY KEY (key_id, version)
);

CREATE TABLE key_usage_inventory (
  key_id          TEXT NOT NULL,
  dataset         TEXT NOT NULL,
  retention_until TIMESTAMPTZ NOT NULL,      -- from the data catalog
  first_wrap_at   TIMESTAMPTZ NOT NULL,
  last_wrap_at    TIMESTAMPTZ NOT NULL,
  PRIMARY KEY (key_id, dataset)
);
-- audit events go to an append-only, hash-chained log in a separate account:
-- audit(ts, region, caller, operation, key_id, version, context_hash, decision, request_id)

Appendix C: Coordination Mechanisms#

C.1 Control-to-Data-Plane Propagation#

Diagram: C.1 Control-to-Data-Plane Propagation

C.2 Quick Comparison#

MechanismGuaranteesFailure ModeUse For
Envelope SDKKMS load scales with distinct DEKsTeams bypass the SDKEvery service
Bounded DEK cacheSurvives KMS blips; bounded revocationTTL configured above tier maximumAll data classes
Encryption contextCiphertext bound to tenant and datasetContext omitted at wrap timeEvery wrap
Static data plane + staleness boundUnwraps through control-plane outages, bounded stalenessUnbounded last-known-goodData plane
Per-caller quota classesTier-0 callers protected from batch and retriesClasses misassignedEndpoint
Versioned KEKsRotation without rewriting dataNew version used before propagatedRotation
Destroy workflowNo single-person irreversible lossInventory incompleteDestruction
Hash-chained audit, fail closedComplete, tamper-evident evidenceSpool undersizedEvery operation

Appendix D: API Contract & Client Behavior#

  • Always use the SDK; it sets cache bounds from the data class and refuses configurations above the tier maximum.
  • Retry unwrap failures with exponential backoff and full jitter, at most 3 attempts; never retry policy denials.
  • Treat DISABLED and PENDING_DESTRUCTION denials as terminal for that tenant; surface "tenant key unavailable" rather than a generic 500.
  • Pass the same encryption context on unwrap as on wrap; contexts are case- and order-normalized by the SDK.
  • Never log plaintext DEKs, wrapped DEKs with contexts, or decrypted data; the SDK redacts known fields.
  • On pod start, warm caches with jitter over 30–60 seconds; prefer node-local shared caches for high-fan-out services.

Appendix E: Observability#

Core metrics:

  • kms.unwrap_p99{region, cell}, kms.unwrap_error_rate{reason} (separate policy denials from failures)
  • kms.dataplane_staleness_seconds, kms.policy_snapshot_version{region}
  • kms.requests_per_s{caller, class}, kms.throttled_by_caller, kms.unwraps_per_distinct_key
  • sdk.cache_hit_rate{service}, sdk.cache_ttl_seconds{service} (must be ≤ tier maximum)
  • hsm.op_latency_p99, hsm.errors, kek_cache.hit_rate
  • audit.lag_seconds, audit.dropped_events, audit.completeness_ratio
  • kms.attempted_use_while_disabled{key}, kms.destroy_cancelled_total

Critical alerts:

AlertThresholdSeverity
Tier-0 unwrap errors (non-policy)> 0.5% for 2 minPage
Data-plane staleness> 15 minPage
Audit dropped events> 0Page (security)
Attempted use of a key pending destruction> 0Page key owner
Requests from one caller> 3× its baseline for 5 minTicket → page if tier-0 throttled
SDK cache TTL above tier maximumany serviceTicket

Debugging the silent failure: overall unwrap success hides the dangerous cases. Watch staleness (a data plane happily serving old policy), audit completeness (operations counted versus events stored) and post-disable unwraps per tenant. All three can look perfect on a success-rate dashboard while revocation and evidence are broken.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 100 servicesManaged KMS, per-service keys, thin wrapperPer-record callers; no tenant shredding
100–2K servicesEnvelope SDK, cache tiers, per-tenant KEKs, quotas, destroy workflowRevocation commitments; audit cost
2K–10K servicesKey platform mandate, usage inventory, revocation broadcast, regional isolation auditRequest costs; customer-held keys
> 10K servicesOwn data plane rooted in HSMs, intermediate keys per class, multi-cloud authorityOrg governance of data classes

What you don't build on day one: customer-held keys, your own data plane, multi-region keys, intermediate key layers, a usage inventory. Each has a trigger in Section 11.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Quota classes: tier-0 online, standard online, batch (preemptible); reserved capacity for tier-0; see the rate limiter for the token-bucket mechanics.
  • Per-tenant KEKs isolate policy, audit and shredding; noisy tenants are visible per key in request metrics.
  • Customer-managed-key tenants have shorter cache bounds and pay for the extra unwraps and external key store proxies.
  • Chargeback: KMS requests and audit bytes attributed per calling service, so cache behavior has an owner.
  • Availability math: a service that unwraps on a cold path inherits the KMS's availability; use the availability calculator to see how a 99.95% dependency caps a 99.99% target.
  1. Loading the index…