Hiring BarSupport

Design a Log Aggregation Platform

Case study80 min read8 diagrams

Technologies referenced in this case study: Apache Kafka · Elasticsearch · OLAP Databases · Kubernetes · Apache Flink · Time-Series Databases

Related: Metrics & Alerting Platform · Distributed Tracing · Object Storage · Search Engine · Backpressure & Load Shedding · Multi-Tenancy · Batch and Stream Pipelines · ClickHouse vs Druid vs Pinot · Elasticsearch vs Postgres Full-Text · Security Basics

Reading Guide#

Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal View) and the appendices on agent buffering, storage layout and redaction
What is a Log Aggregation Platform? — Why interviewers pick this topic

A log aggregation platform collects the text that thousands of services write — request logs, errors, debug lines, audit records — from every host and container, moves it to a central place, makes it searchable within seconds, and keeps it for as long as someone is willing to pay for. Engineers use it at 3 a.m. to find the one stack trace that explains an outage; security uses it to reconstruct what an attacker touched; compliance uses it to prove who accessed what.

The hard part is not shipping lines over the network. The hard part is that log volume is unbounded, bursty and correlated with failure: an incident multiplies log output by 5–20× exactly when engineers need search to be fastest. Volume grows faster than headcount, cost scales with bytes nobody reads, one team's debug flag can drown everyone else, and some of those bytes contain passwords, tokens and personal data that must never be stored at all.

Before vs After — the "incident log storm" scenario:

Without per-tenant quotas and tiered storage:
t=0:        Payments service starts failing; retries log a 4KB stack trace per attempt.
t=+1min:    Payments log volume: 40 MB/s → 1.8 GB/s. Total cluster ingest: 2 GB/s → 3.8 GB/s.
t=+3min:    Indexing nodes saturate. Ingest lag for ALL services: 5s → 9 min.
t=+5min:    On-call searches for the payments error: results end 9 minutes ago.
t=+12min:   Index cluster rejects writes; agents buffer locally; some hosts' disks hit 95%.
t=+20min:   Incident resolved without logs. Post-mortem notes "logging was down".

With per-tenant quotas, agent buffering and a durable queue:
t=0:        Same failure. Payments volume spikes to 1.8 GB/s.
t=+10s:     Payments exceeds its burst quota. Pipeline samples repeated stack traces
            (first 100/min per signature kept in full, rest counted) and spills to object storage.
t=+1min:    Other services' ingest lag: unchanged at ~5s.
t=+2min:    On-call searches payments errors: sees the signature, the count, and full examples.
t=+3h:      Spilled raw logs queryable from object storage for the post-mortem, at lower speed.

Why interviewers reach for this question: It looks like "agents → Kafka → Elasticsearch → Kibana" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: backpressure that decides whose logs are lost, the index-everything vs query-on-read cost decision, retention tiers priced per gigabyte, multi-tenant quotas that stop one team's storm from blinding everyone, and redaction that keeps secrets out of storage you can't easily delete from.

Mechanics Refresher: Collection and Storage Options
OptionHow It WorksProsCons
Node agent (daemon per host)Tails files or container stdout, batches, shipsApps stay simple; one agent per nodeAgent CPU/memory competes with the app; must handle rotation
In-process library shippingApp sends logs over the network directlyNo file I/OBackpressure lands inside the app; crashes lose buffered logs
Durable queue (log broker)Agents write to a partitioned, replicated logAbsorbs bursts; decouples ingest from indexing; replayAnother system; retention sizing
Full-text inverted indexEvery token indexed; search by any wordFast needle-in-haystack queriesIndex often comparable to or larger than raw data; expensive at scale
Label index + compressed chunks in object storageIndex only a few labels; scan chunks at query timeVery cheap storage; simple ingestQueries scan bytes; slow for broad searches
Columnar storeFields as columns, compressed; skip indexesFast aggregations; good compressionSchema handling; free-text search weaker than inverted index
Retention tiersHot (fast, expensive) → warm → archivePay for speed only where it's usedTier transitions; queries spanning tiers
RedactionPattern/field-based masking before storageSecrets never land on diskPatterns miss things; CPU cost

For most production systems: a node agent with a bounded disk buffer, a durable queue absorbing bursts, a processing tier that parses, redacts and enforces per-tenant quotas, a short hot tier that is indexed for fast search, and a cheaper object-storage tier queried on read for older or lower-value data — with retention and cost attributed per team. The storage engine is not the interview — whose logs drop under pressure, what you index, and who pays per gigabyte are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

Log aggregation is not a search-engine question. Everyone can put Elasticsearch behind Kafka.

It is a cost and loss-policy question that tests:

  • Whether you decide, in advance, whose logs are dropped, sampled or delayed when volume exceeds capacity
  • Whether you match storage to query patterns — index the little that's searched often, scan the rest
  • Whether you price retention per gigabyte per tier and attribute it to the teams producing the bytes
  • Whether you keep secrets and personal data out of storage that is expensive to rewrite

The key insight: Log volume is driven by the people writing logs, not the people paying for them, and it spikes during incidents — exactly when logs are most valuable. A Senior design scales the cluster. A Staff design sets the policies that bound it: quotas per tenant, a loss order under backpressure, tiered retention with a price tag, and redaction at the edge. The platform's reliability is measured in whether the right logs are searchable during the worst hour, not in average ingest lag.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveAgents → Kafka → Elasticsearch → dashboardAsks "What are the logs for — debugging, security audit or analytics? How much per day, how long kept, and who pays?"Asks "What's the logging budget as a share of infra spend, and which teams' habits drive it?"
Backpressure"Kafka buffers it""Bounded disk buffer on each host, never block the app; durable queue with 24–72h retention; under overload, sample by signature and spill to object storage — a documented loss order"Defines the org-wide loss policy: audit logs never dropped, debug logs first, with sign-off from security
Storage"Index everything in Elasticsearch""Hot tier indexed for 3–7 days; object storage with label or columnar layout for 30–90 days; archive for compliance. Index what's searched, scan what's kept"Prices each tier and makes retention a per-team budget decision with chargeback
Multi-tenancyOne shared cluster"Per-tenant ingest quotas with burst allowance, per-tenant query limits, cardinality limits on labels or fields"Turns quotas into a budget process: teams buy capacity; overages are visible to their leadership
PII"Developers shouldn't log PII""Redact at the agent or pipeline with tested patterns; field-level allowlists for structured logs; deletion path designed for object storage"Makes logging data classification part of the security standard; audits redaction coverage
Cost"Add nodes""Cost per GB ingested and per GB-month retained, per tier, per team; volume reduction at the source"Sets a cost-per-request target for observability and holds teams to it
Why "backpressure" separates levels

L5: "Agents send to Kafka, which buffers when Elasticsearch is slow." That handles a brief slowdown. During a sustained log storm, Kafka fills to its retention, agents back up, and the candidate hasn't said whether the agent then blocks the application's write call (taking down production to save logs), fills the host's disk (same outcome, slower), or drops silently (losing the evidence for the incident).

L6: "There's a loss order, and it's written down. The application never blocks on logging: the agent has a bounded memory buffer and a bounded disk buffer — say 1 GB or two hours. When both are full, it drops oldest-first and emits a dropped-lines counter. Upstream, the queue holds 24–72 hours. If a tenant exceeds its quota, we sample repeated messages by signature — keep the first N per minute in full, count the rest — and spill the overflow to object storage instead of the index. Audit-class logs bypass sampling and have reserved capacity."

L7: "The loss order is a policy decision, not an implementation detail. Security signs off that audit logs are never sampled; product teams agree debug logs are sampled first. I'd publish it, so when a team loses debug logs during their own storm, the answer is the policy they agreed to, not an outage of the logging platform."

Why "storage" separates levels

L5: "Elasticsearch, with daily indices and 30-day retention." Full-text indexing every line costs CPU at ingest and storage often comparable to the raw size, replicated. At 170 TB/day that's petabytes of SSD for 30 days, most of it never queried after the first 48 hours.

L6: "Most log queries hit the last few hours, filter by service and level, and then look for a string. So: a hot tier of 3–7 days, indexed for fast search. Beyond that, compressed chunks in object storage — at roughly a tenth of the bytes and a fraction of the price per GB — queried on read, either by labels with a scan or by a columnar engine with skip indexes. Queries over old data are slower; that's the deal teams accept for keeping 90 days instead of 7."

L7: "Each tier has a price per GB-month and a query latency. I'd let teams choose retention per log stream within a budget, and show them what they're paying for logs nobody has queried in 30 days. The biggest saving is usually not storage engineering — it's deleting debug logs at source."

Why "PII" separates levels

L5: "We'll tell developers not to log sensitive data." They will anyway — a request object dumped on error, a header with a bearer token, an email address in a URL. Once it's in an immutable, compressed chunk in object storage, deleting it means rewriting files.

L6: "Redaction runs before anything is persisted beyond the host: known secret patterns — tokens, keys, card numbers with checksum validation — and field-level allowlists for structured logs. Redaction coverage is tested with seeded canaries. And because something will slip through, the storage layout supports deletion: chunks are addressable by tenant and time so a purge rewrites a bounded set of files."

L7: "Logs are a data store with the same classification rules as databases. I'd put logging into the data classification standard, give security an audit of what fields each service emits, and measure redaction coverage as a compliance metric."

Positions to Commit To#

PositionRationale
Never block the application on logging; bounded agent buffers with counted dropsTaking down production to preserve logs inverts priorities
A durable queue between collection and storage, 24–72h retentionAbsorbs incident bursts and indexing outages; enables replay
A written loss order: debug sampled first, audit neverOverload will happen; the decision must be made before it does
Index only the hot window; query older data on read from object storageMost queries hit the last hours; index cost is wasted on cold data
Per-tenant ingest quotas, query limits and cardinality limitsOne team's storm or label explosion must not blind everyone
Redact at the edge; design storage for deletion anywaySecrets in logs are inevitable; immutable storage makes them expensive
Cost per GB ingested and retained, attributed per teamVolume is driven by producers; they must see the bill

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Operational debugging logsHigh volume, bursty, mostly queried within hours; loss of some debug lines tolerableAgents + queue + hot index + object-storage tier; quotas and sampling under pressureIncident storms blind search; cost explodesIngest-to-searchable p99 < 30s for in-quota tenants; < 0.1% unplanned loss
Security and audit logsCompleteness and integrity matter more than latency; retained for yearsSeparate lane with reserved capacity, no sampling, write-once storage, tamper evidenceSilent gaps; logs altered or deletedZero sampled or dropped records; verifiable gap detection
Analytics from logs (counts, rates, funnels)Aggregations over long ranges; schema mattersEmit metrics or structured events instead; columnar store if logs must be usedUsing text search for aggregations; cardinality blowupsShould usually be a metrics or events pipeline, not logs

🎯 Staff Move: "I'll design the operational debugging platform — high-volume service logs searched mostly during incidents. Audit logs get a separate lane with reserved capacity, no sampling and write-once storage; I'll show where they branch. If the real need is counting things, I'd push teams to emit metrics — counting error lines in a log store is the expensive way to build a dashboard."

Where the Design Splits#

#Fault LineThe Tension
1Index Everything vs Query on ReadFast arbitrary search at high ingest and storage cost, or cheap storage with slower scan-based queries?
2Backpressure: Block, Buffer or DropWhose logs are lost, delayed or sampled when volume exceeds capacity — and who decides?
3Schema on Write vs Schema on ReadParse and type fields at ingest (fast queries, brittle pipelines) or store raw and parse at query time (flexible, slow)?
4Retention Tiers and Who PaysOne retention for all, or per-stream tiers priced and charged back to the producing team?
5Where PII Redaction HappensAt the source, in the pipeline or at query time — and how you delete what slipped through

How Real Companies Built It#

Why this section belongs here: Three well-documented systems made visibly different storage choices for logs. Naming them shows you see indexing as a cost decision, not a default.

Grafana Loki — Index Only Labels, Store Compressed Chunks in Object Storage#

Loki's documentation states that it does not index the contents of logs, only metadata as a set of labels for each log stream; log data is compressed and stored in chunks in an object store such as Amazon S3 or Google Cloud Storage, which keeps the index much smaller than other log tools (Loki overview). Its cardinality guidance warns that high-cardinality labels create many streams, a huge index and thousands of tiny chunks, and notes a default limit of 15 index labels (Loki cardinality).

Staff insight: This is the "query on read" end of the spectrum taken seriously: cheap ingest and storage, paid for with scan-heavy queries and strict label discipline. In an interview, say: "If I index only labels, cardinality limits become a tenant policy — one team putting request IDs in a label breaks the index for everyone."

Uber — From Elasticsearch to a Columnar Store for Schema-Agnostic Logs#

Uber's engineering blog describes moving its logging platform from the ELK stack to ClickHouse. With Elasticsearch, incompatible field types caused type-conflict errors that dropped logs, and the team ran 20+ clusters per region to limit the blast radius of heavy queries and mapping explosions. On ClickHouse, a single node ingested about 300K logs per second — roughly ten times a single Elasticsearch node in their setup — and the platform's hardware cost dropped by more than half while serving more traffic, with Kafka buffering ingestion (Uber blog).

Staff insight: Schema conflicts and mapping explosions are not edge cases at scale; they are the main failure mode of index-everything with dynamic fields. Say: "Dynamic field mapping is a multi-tenant hazard. Either I cap fields per tenant, or I store logs in a layout where a new field can't break someone else's index."

Datadog Husky — Stateless Writers, Object Storage and Separate Query Compute#

Datadog's engineering blog describes Husky, its third-generation event store for logs and other events, as an unbundled, distributed, schemaless, vectorized column store. Writers ingest from Kafka, buffer briefly and upload to blob storage; readers query individual files; compactors merge small files; and durability is pushed to FoundationDB and S3. The post says this separation made it cost-effective to retain data that is rarely queried but must stay immediately queryable, and enabled a new product, Flex Logs (Datadog blog).

Staff insight: Separating ingest, storage and query lets each scale with its own driver — bytes in, bytes kept, queries run — and turns retention into a pricing choice. Say: "Storage and query compute should scale independently, because logs are written a thousand times more than they're read."

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Kafka buffers the spikes""The spike lasts two hours and exceeds Kafka's retention. Whose logs are lost?"Loss policy, backpressure end-to-end
"We index everything""What does 30 days of that cost, and how much of it is ever queried?"Cost awareness, tiering
"Teams send logs to the platform""One team enables debug logging in production at 10× volume. What happens to everyone else?"Multi-tenant quotas
"We parse JSON logs""Two services send status as a string and an integer. What breaks?"Schema conflicts, mapping explosion
"We redact PII""A password was logged for six weeks before anyone noticed. Delete it."Deletion in immutable storage
"Agents tail log files""The host's disk is 95% full and the agent can't ship. What does the agent do?"Never harm the app; bounded buffers
"We keep logs 90 days""Security needs a year, debugging needs a week. One tier or two?"Retention tiers, audit lane

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Agents on every node tail container output, apply a first redaction pass, and batch to an ingest gateway that tags the tenant and checks quotas. A durable queue absorbs bursts. Processors parse and enrich by stream schema, redact again with field-level rules, and apply overload control — sampling repeated messages by signature when a tenant exceeds quota. Every log line lands compressed in object storage; in-quota logs are also indexed in a short hot tier. Audit streams take a separate lane into write-once archive. The query frontend routes by time range: recent queries hit the hot index, older ones scan object storage. The metric that tells you the platform is healthy is ingest.lag_p99_s for in-quota tenants — it should stay under 30 seconds no matter which team is storming.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Collection"Agents ship to Kafka""Bounded memory and disk buffers; never block the app; count every drop."
Overload"Kafka buffers it""Written loss order: sample debug by signature, spill to cold, never sample audit."
Storage"Elasticsearch, 30 days""Index 7 days hot; object storage for 90 days, query on read; archive for audit."
Tenancy"One shared cluster""Per-tenant ingest quotas with burst, query limits, cardinality limits."
Schema"Parse everything""Per-stream schemas; unknown fields kept as raw; type conflicts can't drop logs."
PII"Don't log PII""Redact at agent and pipeline, test with canaries, design storage for deletion."
Cost"Add nodes""Cost per GB ingested and retained per team; reduce at the source."

Numbers to Bring#

MetricValueWhy It Matters
Ingest volume (example scale)~2 GB/s average, 5–10× in incidents (illustrative)~170 TB/day raw; sizes every tier
Average log line~200–1,000 bytes (illustrative)2 GB/s ≈ 2–10M lines/s
Compression for logscommonly ~5–15× (illustrative, data-dependent)170 TB/day raw → ~15–30 TB/day stored
Full-text index overheadoften comparable to raw size before replication (illustrative)Why indexing 90 days is expensive
Object storage price~$0.02/GB-month order of magnitude (check current provider pricing)Cheapest durable tier for warm data
Hot tier window~3–7 days typicalMost queries hit the last hours
Queue retention24–72 hoursCovers an indexing outage plus a weekend
Agent buffermemory ~64–256 MB, disk ~1–5 GB (illustrative)Bounded harm to the host
Fluent Bit buffer limitsmem_buf_limit pauses that input; storage.total_limit_size discards the oldest chunk (docs)A real agent's bounded-buffer loss order
Agent overhead target< ~2–5% of a node's CPU (illustrative)Logging must not tax the product
Ingest-to-searchablep99 < 30 s target for in-quota tenantsThe incident-time promise
Elasticsearch default field limitindex.mapping.total_fields.limit = 1,000Mapping explosion guardrail
Loki default index labels15 per seriesCardinality guardrail
Uber single-node ingest (ClickHouse)~300K logs/s, ~10× their ES nodesColumnar ingest efficiency at scale

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on Elasticsearch shard counts, index templates and Kibana dashboards, then run out of time before the interviewer asks the questions that decide the level: "One team turns on debug logging and triples volume. What happens to everyone else?" and "What does 90 days cost?" Compress the pipeline to ~8 minutes and spend the rest on backpressure and loss policy, storage tiers, multi-tenancy, schema and redaction.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Every service on every host writes logs. We collect them, make them searchable within seconds by service, time, level and text, support live tail during incidents, keep them for a defined retention per stream, and give security an audit lane that's complete and tamper-evident."

Then the non-functional requirements, which is where the design lives:

"Four constraints drive everything. One: volume is bursty and correlated with failure — incidents multiply output 5–20× when search matters most. Two: logging must never take down the application. Three: cost scales with bytes, and producers don't see the bill unless we show it. Four: logs contain secrets and personal data whether we like it or not. I'll assume 3,000 services on 50,000 nodes, 2 GB/s average ingest, 10 GB/s during a bad incident, 7 days of fast search and 90 days of retained, slower search."

Then name the underspecified parts:

"I'd confirm: is this debugging, audit, or both? Who pays? What's the acceptable loss under overload? I'll assume debugging is primary with a separate audit lane, costs charged back per team, and debug-level logs sampled first under pressure."

🎯 Staff Move: Saying "incidents multiply log volume exactly when search matters most" reframes the problem from "store logs" to "stay useful during the worst hour" — that's the sentence that sets the level.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Tenant: a team or service owner — tenant_id, ingest_quota_bytes_s, burst_bytes, retention_policy, cost_center
  • Stream: a labeled sequence — {tenant, service, env, cluster, level}; labels low-cardinality by policy
  • Log record: ts, stream, body (raw), fields (parsed, optional), trace_id?, ingest_ts
  • Chunk: compressed block of records for one stream and time range — chunk_id, tenant, t_min, t_max, bytes, tier
  • Query: tenant, time_range, label_selector, filter (text/field), limit

Ingest API (agent → gateway):

POST /v1/ingest        Content-Encoding: gzip
Authorization: tenant token
{ streams: [ { labels: {service, env, level}, entries: [ [ts, line], … ] } ] }
→ 204 accepted
→ 429 quota exceeded, Retry-After: 2   (agent keeps buffering, backs off)
→ 413 batch too large

Query API:

GET /v1/query?selector={service="checkout",level="error"}&q="timeout"&from=-1h&limit=500
GET /v1/tail?selector={service="checkout"}           (live tail, server-sent stream)
POST /v1/retention   { stream_selector, hot_days, warm_days }

🎯 Staff Move: "429 with Retry-After is the backpressure contract between the platform and agents. The agent never drops on a 429; it buffers within its bounds and retries. Drops only happen when the agent's own bounded buffer is full — and each one is counted."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk one log line in 90 seconds:

  1. Checkout writes a JSON line to stdout. The node agent tails the container log, applies pattern redaction for tokens and card numbers, and adds it to a batch.
  2. Every second or 1 MB, the agent gzips the batch and posts it to the gateway, which authenticates the tenant, checks the token-bucket quota and writes to the queue partitioned by tenant and stream.
  3. A processor parses the line against checkout's stream schema, attaches Kubernetes metadata, runs field-level redaction, and checks the tenant's rate against quota. In quota: indexed into the hot tier and appended to a compressed chunk for object storage. Over quota: sampled by message signature, with full lines still written to object storage.
  4. Within ~10 seconds the line is searchable in the hot tier; within ~5 minutes its chunk is flushed to object storage.
  5. A query for the last hour goes to the hot index; a query for last month goes to object storage, filtered by labels and time, scanned in parallel.

🎯 Staff Move: Say out loud: "Every line lands in object storage, which is cheap and durable. The hot index is an accelerator for the recent window, not the system of record." You've now spent ~8 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes this hard is that volume is unbounded and spikes during incidents, cost scales with bytes nobody reads, and one tenant can hurt all the others. I'd like to go deep on backpressure and the loss order, index vs query-on-read and retention tiers, multi-tenant quotas and cardinality, schema handling, and PII redaction. Where would you like to start?"

If no preference: start with backpressure. It's the question that decides whether logs exist during the incident.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Backpressure and the loss order (7–8 min)

"Four buffers, each bounded. The app's write to stdout never blocks on the network — the container runtime writes to a local file. The agent reads files with a 64 MB memory buffer that spills to a 1 GB disk buffer; when that's full, it stops reading new data and, if the file rotates away, those lines are lost and counted in agent.dropped_lines_total. The gateway returns 429 when a tenant exceeds quota; agents back off. The queue holds 48 hours, so an indexing outage of a day loses nothing. Overload control in the processor samples by signature for over-quota tenants — first 100 per signature per minute in full, the rest counted — but still writes everything compressed to object storage if the storage write path keeps up."

Quantify: "A 1 GB disk buffer at a typical node's 200 KB/s of logs is ~80 minutes of outage tolerance per host. 48 hours of queue at 2 GB/s is ~350 TB raw, ~35–70 TB compressed in the queue — real money, sized deliberately."

Who pays: "The storming tenant loses fidelity on its own repeated messages. Every other tenant pays nothing. Audit streams have reserved capacity and are never sampled."


Deep dive 2: Index vs query on read, and retention tiers (6–7 min)

"Query patterns decide this. Most searches are in the last 24 hours, filtered by service and level, then a text match. So the hot tier — 7 days — is indexed for interactive search in seconds. Everything is also in object storage as compressed, columnar chunks partitioned by tenant, stream and hour, with small per-chunk metadata — min/max time, labels, a bloom filter of tokens. Older queries prune by labels and time, skip chunks via bloom filters, and scan the rest in parallel. A 30-day query over one service is tens of seconds instead of sub-second; that's the explicit trade."

Quantify: "170 TB/day raw. Hot: 7 days indexed and replicated is roughly 170 × 7 × 2 × ~1 ≈ 2.4 PB of SSD-class storage if we indexed everything — which is why the hot tier holds only in-quota, non-sampled data and drops debug after 24 hours. Warm: 90 days × ~17 TB/day compressed ≈ 1.5 PB of object storage at ~$0.02/GB-month ≈ $30K/month."


Deep dive 3: Multi-tenant quotas and cardinality (5–6 min)

"Each tenant has an ingest quota in bytes per second with a burst bucket — say 2× quota for 10 minutes — enforced at the gateway and again in processing. Query limits per tenant: concurrent queries, bytes scanned per query, and a max time range for hot-tier queries. Cardinality limits: max labels per stream (around 10–15), max distinct values per label per day, max parsed fields per stream. Violations are rejected or folded into the raw body, never allowed to grow the shared index without bound."


Deep dive 4: Schema on write vs on read (4–5 min)

"Structured JSON is encouraged and parsed at ingest against a per-stream schema that the owning team registers. Known fields are typed and indexed in the hot tier. Unknown fields stay in the raw body and are searchable as text, never auto-added to a shared mapping. Type conflicts — status as a string in one service and an integer in another — can't collide because fields are namespaced per stream, and a value that fails to parse is kept raw, not dropped."


Deep dive 5: PII redaction and ownership (3–4 min)

"Two passes. The agent runs cheap pattern redaction — bearer tokens, API key formats, card numbers validated with a checksum to avoid false positives. The processor applies per-stream field rules: allowlisted fields kept, known-sensitive fields hashed or dropped. Seeded canary secrets in synthetic traffic verify coverage daily. When something slips through, the purge job rewrites the affected chunks — bounded because chunks are partitioned by tenant and hour. The platform team owns ingest health and quotas; product teams own their volume, schemas and what they log; security owns redaction rules and the audit lane."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: log volume is driven by producers and spikes during incidents, so the platform is defined by its policies — a loss order under backpressure, an indexed hot window over cheap object storage, per-tenant quotas and cardinality limits, schemas that can't break each other, and redaction before persistence — with cost per gigabyte visible to the teams producing it."

The evolution closer:

"What I'd build later: log-to-metric extraction so teams stop querying logs for counts; trace-linked logs; per-team retention self-service with a budget; and federated queries across regions. What I'd not build: indexing every line for 90 days — I'd invest in reducing volume at the source instead."

🎯 Staff Move: End on who owns the bytes. "The platform owns keeping in-quota logs searchable within 30 seconds through any incident. Teams own their volume and see its cost. When a team's own storm gets sampled, that's the policy working, not the platform failing."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Cluster tuning10 min on shards, replicas, index templates"Hot index is an accelerator; object storage is the record"
Kafka as the answer"Kafka buffers it" ends the backpressure discussionFour bounded buffers and a written loss order
One retention30 days for everythingHot, warm and archive tiers with prices
No tenancyOne shared cluster, no quotasIngest quotas, query limits, cardinality limits
PII as a guideline"Developers shouldn't log secrets"Redaction in the pipeline, canaries, deletion path
No costNever states $ per GBCost per GB ingested and retained, per team

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Logging is the interview where the users who generate the load are not the users who pay for it, and where demand spikes at the exact moment the system is most needed. Every engineer adds log lines; nobody removes them. Volume grows with services, traffic and incidents, and the cost line grows with it until a finance review asks why observability costs more than the database fleet.

A Senior design scales storage to match demand. A Staff design recognizes that demand is a policy problem: what gets kept, at what fidelity, for how long, at whose expense — and what gets dropped first when the system is overwhelmed. The architecture follows from those policies, and the policies need owners outside the platform team.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"It's the worst outage of the year. Log volume is 6× normal, one service is producing 80% of it, and the on-call is searching for a specific error in a different service. Walk me through what your platform is doing right now, whose logs are being dropped or sampled, and what that on-call sees."

A candidate who answers with the storming tenant hitting its burst quota, signature sampling applied to that tenant only, full raw lines still landing in object storage, queue depth rising but within its 48-hour bound, other tenants' ingest lag unchanged at seconds, query limits stopping a dozen engineers' 30-day searches from starving the hot tier, and the on-call finding their error in ten seconds has run a logging platform. A candidate who says "Kafka absorbs the spike" has drawn one.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Operational debugging → useful during the worst hour, affordable the rest of the time

  • Constraint: bursty volume, recency-heavy queries, tolerance for sampling repeated debug lines
  • Strategy: bounded agents, durable queue, hot index over object storage, quotas, signature sampling
  • Failure mode: incident storms blind search; cost grows without owners
  • Who pays for imperfection: on-call engineers (blind incidents), finance (unbounded spend), the storming team (sampled fidelity)

Security and audit → complete, intact, retained for years

  • Constraint: no sampling, gap detection, tamper evidence, legal retention
  • Strategy: separate lane with reserved ingest capacity, write-once storage, sequence numbers per source, independent verification
  • Failure mode: silent gaps; deletion by an attacker with access
  • Who pays: security (blind investigations), the company (compliance findings)

Analytics from logs → probably not a logging problem

  • Constraint: long-range aggregations, stable schemas
  • Strategy: emit metrics or structured events; if logs are the only source, extract metrics at ingest
  • Failure mode: dashboards that scan terabytes per refresh
  • Who pays: the platform (query load), everyone else (slow searches during dashboard refreshes)

2.2 When NOT to Build a Log Aggregation Platform#

  • You're small. Under a few hundred GB/day, a managed logging service or a single search cluster is cheaper than a team. The tiering and quota machinery earns its keep at tens of TB/day.
  • You need counts, rates and alerts. Use the metrics platform; a metric is bytes per minute, a log line is bytes per event.
  • You need request-level causality across services. That's distributed tracing; link logs to traces by trace ID rather than reconstructing flows from text.
  • You need business events with a schema. Use an event pipeline into a warehouse; logs are the wrong contract for data other teams depend on.

🎯 Staff Insight: "The cheapest log line is the one never written. Before I build a bigger platform, I'd find the ten streams that make up half the volume and ask whether anyone reads them."

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Debugging vs audit — the completeness bar differs completely
  • Who pays — shared cost hides the producers driving it
  • Acceptable loss — "never lose a log" is unaffordable during a storm; what's the order?
  • Retention — one number for everything vs per-stream tiers
  • Structured vs free text — determines parsing cost and schema risk
  • PII policy — whether logs may contain personal data at all, and how deletion requests apply

Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.

2.4 Precise Terminology#

TermWhat It MeansWhy It Matters in the Interview
StreamRecords sharing one label setUnit of indexing and cardinality
CardinalityDistinct label-value combinationsDrives index size; the multi-tenant hazard
ChunkCompressed block of one stream's records over a time rangeUnit of storage, pruning and deletion
Hot tierRecent data, indexed, on fast storageExpensive; sized by query recency
Query on readScan compressed data at query time, no full indexCheap storage, slower broad queries
Mapping explosionUnbounded growth of indexed field definitionsBreaks shared indices
BackpressureSlowing producers when consumers can't keep upMust stop before it reaches the app
Signature samplingKeep N examples per message template, count the restPreserves information under overload
Ingest lagTime from write to searchableThe incident-time SLO
ChargebackAttributing cost to the producing teamMakes volume visible to its cause
RedactionMasking secrets/PII before persistencePrevents expensive purges

🎯 Staff Insight: If the interviewer says "we can't lose any logs", ask: "Which logs? Audit logs, yes — reserved capacity, no sampling. Debug logs during a 10× storm from one service — I'd rather sample its repeated stack traces than blind every other team. Can we agree a loss order?"


3. Where the Design Splits#

Every logging decision has a technical side (index structures, buffer sizes, compression) and an organizational side (whose logs drop, who pays, who decides retention). Interviewers grade the second side.

3.1 Fault Line 1: Index Everything vs Query on Read#

The tension: A full-text index answers any search in sub-second time but costs ingest CPU and storage comparable to the raw data, replicated, on fast disks. Query-on-read stores compressed chunks cheaply and pays at query time by scanning.

ChoiceWhat WorksWhat BreaksWho Pays
Full-text index, all retentionFast arbitrary search over monthsCost scales with every byte for the whole retention; mapping explosionsFinance (storage), platform (cluster ops)
Labels only + object-storage chunksVery cheap ingest and storageBroad text queries scan lots of bytes; label cardinality must be policedEngineers (slower searches)
Columnar store with skip indexesStrong compression; fast filters and aggregationsFree-text search weaker; schema handling neededPlatform (schema tooling)
Hot index + object-storage warm tierFast where queries are; cheap where data sitsTwo query paths; queries spanning bothPlatform (routing complexity)
Diagram: 3.1 Fault Line 1: Index Everything vs Query on Read

Staff default: "Hot index for 7 days of in-quota data, with debug-level streams kept hot for only 24 hours. Everything written compressed to object storage, partitioned by tenant, stream and hour, with per-chunk metadata and token bloom filters. Queries route by time range; broad queries over cold data require a label selector so they can't scan the whole tenant."

When to deviate:

  • Security investigations that need arbitrary search over months: index the audit lane for longer — it's a fraction of the volume.
  • Small scale: a single indexed cluster with 14–30 days is simpler than two tiers.

🧭 Principal Move: "I'd publish the price of each tier per GB-month and the query latency teams should expect from it. Then retention becomes a choice teams make with their own budget, not a platform default everyone pays for."

❌ Common L5 Trap: "Elasticsearch with 90-day retention and daily indices." At 170 TB/day, that's on the order of tens of petabytes of replicated, indexed storage, most of it never queried after day two. The design isn't wrong — it's unaffordable, and the candidate didn't price it.


3.2 Fault Line 2: Backpressure — Block, Buffer or Drop#

The tension: When ingest exceeds capacity, logs must wait somewhere or be discarded. Waiting in the application blocks production; waiting on the host fills disks; waiting in the queue costs money; dropping loses evidence. The question is whose logs, where, in what order.

ChoiceWhat WorksWhat BreaksWho Pays
Block the app's log callNo lossProduction latency or outage caused by loggingUsers (outage)
Unbounded agent bufferNo loss short-termHost disk or memory exhaustion; agent OOMThe app on that host
Bounded agent buffer, drop oldest, countHost protected; drops visibleLoss during long outagesThe tenant (bounded loss)
Durable queue with long retentionAbsorbs indexing outages and burstsStorage cost; retention limit still existsPlatform (queue cost)
Signature sampling for over-quota tenantsKeeps information, sheds volumeExact counts approximated; some full lines gone from hot tierStorming tenant (fidelity)
Diagram: 3.2 Fault Line 2: Backpressure — Block, Buffer or Drop

Staff default: "The application writes to local stdout files and never blocks on the network. The agent has bounded memory and disk buffers, honors 429s, and drops oldest when full — counting every drop per tenant. The queue retains 48 hours. Over-quota tenants are sampled by message signature in the hot path, while raw lines still go to object storage when its write path has capacity. Audit streams have reserved capacity at every stage and are never sampled."

When to deviate:

  • Audit and security logs: block-or-fail is sometimes legally required — use a separate, synchronous, low-volume path for those specific events, not for all logs.
  • Edge and mobile: client buffers are tiny and lossy; accept it and sample at the source.

🧭 Principal Move: "The loss order is an org policy: audit never, error-level last, debug first, and the storming tenant before anyone else. Security and engineering leadership sign it; the platform enforces it. Then a lost debug line during a team's own storm is policy, not an incident."

❌ Common L5 Trap: "Agents retry until the backend accepts, so nothing is lost." Unbounded retry means unbounded buffering — the agent fills the host's disk or memory and takes the application down with it. See Backpressure & Load Shedding.


3.3 Fault Line 3: Schema on Write vs Schema on Read#

The tension: Parsing at ingest gives typed fields, fast filters and compact columns — and a pipeline that breaks when formats change or two teams disagree on a field's type. Storing raw and parsing at query time is flexible and robust — and slow, with every query re-parsing.

ChoiceWhat WorksWhat BreaksWho Pays
Schema on write, shared dynamic mappingFast field queries; zero setupType conflicts drop or reject logs; mapping explosionEvery tenant (shared breakage)
Schema on write, per-stream registered schemasTyped, isolated per teamTeams must register schemas; drift handlingProduct teams (registration)
Schema on read (raw text)Nothing breaks at ingestEvery query parses; aggregations slowEngineers (query latency)
Hybrid: known fields typed, unknown kept rawRobust and fast for common fieldsTwo representations to explainPlatform (tooling)

Staff default: "Hybrid with per-stream namespaces. Teams register the fields they want typed; those are parsed at ingest and indexed in the hot tier. Everything else stays in the raw body, searchable as text. A field that fails to parse is kept raw with a parse-error tag — never dropped. No shared dynamic mapping, so one team's new field can't push a shared index past its field limit."

When to deviate:

  • Highly regular logs (load balancer access logs): full schema on write, columnar, typed — they're effectively events.
  • Third-party software logs with unstable formats: raw only, with query-time parsing helpers.

🎯 Staff Insight: "Elasticsearch's default limit is 1,000 mapped fields per index for a reason. In a shared index, the 1,001st field from any team breaks ingest for everyone. Per-stream schemas turn a shared failure into a local one."

❌ Common L5 Trap: "We parse JSON and index all fields automatically." It works until one service logs a dynamic map — user IDs as keys — and mints thousands of fields an hour, or two services disagree on whether status is a string. Logs get rejected for tenants who did nothing wrong.


3.4 Fault Line 4: Retention Tiers and Who Pays#

The tension: Retention is the largest cost multiplier: every extra day multiplies every byte. Teams want long retention by default because it's free to them; the platform wants short retention because it pays. Neither has the right incentive.

ChoiceWhat WorksWhat BreaksWho Pays
One retention for all (e.g., 30 days indexed)SimplePays for long retention on chatty debug logsPlatform budget
Tiered, platform-chosenCheaperTeams with real needs (security, compliance) fight exceptionsPlatform (exception handling)
Tiered, team-chosen within a budget + chargebackCost tracks value; teams decideBudget process; teams must understand tiersProduct teams (their own budget)
Keep everything forever in archiveNever miss old dataArchive grows without bound; deletion obligationsFinance, legal (data liability)
Diagram: 3.4 Fault Line 4: Retention Tiers and Who Pays

Staff default: "Three tiers. Hot, indexed: 7 days for info and above, 24 hours for debug. Warm, object storage: 30 days default, up to 90 per stream if the team pays. Archive: audit streams only, write-once, retention set by security and legal. Every stream's bytes ingested and GB-months retained are reported per team monthly, with the top streams by cost and their query counts — 'you paid $14K for this stream; it was queried twice.'"

When to deviate:

  • Regulated environments: retention minimums set by policy override team choice.
  • Early-stage platforms: one tier, short retention; tiering arrives with scale.

🧭 Principal Move: "Chargeback changes behavior more than any compression algorithm. The first month teams see their logging bill next to their query counts, volume drops — I'd plan for that and measure it."

❌ Common L5 Trap: "Keep 30 days for everything; storage is cheap." Object storage is cheap; indexed, replicated SSD storage for 30 days of 170 TB/day is not. And "cheap" multiplied by a volume that doubles yearly is a line item finance will notice.


3.5 Fault Line 5: Where PII Redaction Happens#

The tension: Redacting at the source is most effective but depends on every team's discipline. Redacting in the pipeline is centralized but sees only what patterns can recognize. Redacting at query time leaves secrets in storage. And once data is in compressed, immutable chunks, deletion means rewriting files.

ChoiceWhat WorksWhat BreaksWho Pays
Source only (logging library rules)Never leaves the processDepends on every team; misses ad-hoc dumpsSecurity (gaps)
Agent and pipeline patternsCentral, consistent, no team effortFalse negatives on unknown formats; false positives on IDsPlatform (CPU), engineers (over-masking)
Query-time maskingEasy to change rulesSecrets stored and replicated; breach exposureCompany (liability)
Structured field allowlistsPrecise for structured logsOnly works where logs are structuredProduct teams (schema discipline)

Staff default: "Defense in depth. Logging libraries mask known sensitive fields by type. Agents run fast pattern redaction for credentials and card numbers before anything leaves the host. Processors apply per-stream field allowlists. Daily canaries inject fake secrets through real services and alert if any reach storage unmasked. Storage is partitioned by tenant and hour so a purge rewrites a bounded set of chunks, and the hot index supports delete-by-query for the 7-day window."

When to deviate:

  • Debugging that requires sensitive values (e.g., payment flows): log a reference or a keyed hash, not the value; give a separate, access-controlled store for the rare cases that truly need raw data.
  • Regions with strict data protection: keep logs in-region and apply stricter allowlists.

🧭 Principal Move: "Logs are a database the whole company writes to without review. I'd put logging under the data classification standard, give security a per-service report of emitted fields, and track redaction canary pass rate as a compliance metric."

❌ Common L5 Trap: "We mask PII in the search UI." The secret is still in the index, the replicas, the object store, and any export. A breach or an over-privileged query path exposes it, and a deletion request means finding it everywhere.


4. When It Breaks#

4.1 The Debug Flag — One Team Triples Platform Volume#

t=0:       A team enables DEBUG in production to chase a bug. Their volume: 60 MB/s → 2.4 GB/s.
t=+1min:   Total ingest 2 GB/s → 4.3 GB/s. No per-tenant quotas enforced in processing.
t=+4min:   Hot-tier indexers at 100% CPU. ingest.lag_p99_s for all tenants: 8s → 6 min.
t=+10min:  Queue depth rising 2 TB/min. Live tail lags. Two unrelated incidents begin.
t=+25min:  On-call finds the stream by bytes; asks the team to revert. Lag drains over 40 min.

Detection: tenant.ingest_bytes_s top-N vs quota; ingest.lag_p99_s{in_quota}; indexer.cpu_util.

Mitigation: enforce quota in processing with signature sampling for the over-quota tenant; route its excess directly to object storage, skipping the hot index.

Prevention: quotas enforced by default at gateway and processor; debug level kept hot for 24 hours only; a "debug window" feature that lets a team raise its quota for 30 minutes on a named set of pods, with cost shown up front.

Owner: logging platform on-call; the producing team owns its volume.

4.2 The Full Disk — An Agent Takes Down Its Host#

t=0:       Ingest gateway in one zone rejects connections (cert expired).
t=+5min:   Agents in that zone buffer to disk. Buffer limit misconfigured as "unlimited" on
           an older agent version deployed to 4,000 nodes.
t=+3h:     Node disks at 95%. Kubelet starts evicting pods for disk pressure.
t=+3h10m:  Production services in the zone lose capacity. A logging outage became a
           production outage.

Detection: agent.buffer_bytes{node}, agent.buffer_pct_of_limit, node disk pressure events.

Mitigation: cap agent disk buffers at a fixed size (e.g., 1 GB) and a fraction of free disk; on limit, drop oldest and count; rotate the certificate.

Prevention: agent config enforced by a fleet policy, not per-deployment; certificate expiry alerts 30 days ahead; a chaos test that blackholes the gateway for one zone monthly.

Owner: logging platform (agent fleet and gateway).

4.3 The Mapping Explosion — One Service Breaks the Shared Index#

t=0:       A service starts logging a map keyed by customer ID as a JSON object.
t=+20min:  Shared hot index gains 9,000 new fields; hits its field limit.
t=+21min:  Every write with a new field to that index is rejected; with dynamic mapping,
           12 other services sharing the index start losing logs.
t=+1h:     Teams notice "missing logs"; platform finds the culprit by field growth rate.

Detection: index.field_count growth rate; ingest.rejected_total{reason=mapping}; per-stream new-field rate.

Mitigation: move the stream to its own index or fold unknown fields into the raw body; blocklist the offending field path.

Prevention: no shared dynamic mapping; per-stream registered schemas; unknown fields stored raw; a per-stream limit on parsed fields.

Owner: logging platform (schema policy); producing team (log format).

4.4 Secrets in Logs — Six Weeks of Bearer Tokens#

t=0:       A new HTTP client library logs full request headers at INFO on retry.
t=+6 wks:  Security finds Authorization headers in a log search result.
t=+6 wks:  Tokens for ~2M sessions sit in the hot index, its replicas, 42 days of object
           storage chunks and two analytics exports.

Detection: redaction canaries (a fake token sent through real services daily; alert if found unmasked); periodic scans for credential patterns in stored data.

Mitigation: revoke affected tokens; delete-by-query in the hot tier; rewrite affected chunks in object storage — bounded because chunks are partitioned by tenant and hour; purge exports.

Prevention: agent-level pattern redaction for authorization headers; logging library denies header logging by default; canaries per language runtime.

Owner: security (rules, response), logging platform (purge tooling), producing team (library usage).

4.5 Silent Loss — A Parse Error Path That Drops Lines#

t=0:       A processor deploy changes JSON parsing; lines with invalid UTF-8 now throw.
t=+0:      The exception handler drops the line and increments a counter nobody alerts on.
t=+5 days: An incident: the on-call can't find errors from one service. 3% of its lines
           (those with binary payload fragments) have been dropped for five days.

Detection: end-to-end counts: lines read by agents vs lines stored, per stream, reconciled hourly; processor.dropped_total{reason} with alerts.

Mitigation: fix parser to keep unparseable lines raw with a tag; replay the queue for the last 48 hours (the rest is lost).

Prevention: invariant: the processor never drops a line without a counted, alerted reason; canary streams with malformed content in every deploy's test suite.

Owner: logging platform.

An engineer runs a regex across all services for 90 days during an incident. The query fans out to hundreds of thousands of chunks and saturates the query tier; every other engineer's searches queue behind it during the incident. Mitigation: per-tenant and per-user limits on bytes scanned and concurrent queries; require a label selector for queries beyond the hot window; queue long queries as async jobs with results delivered later.

Detection: query.bytes_scanned{user}, query.queue_wait_p99.

Owner: logging platform (query limits).

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Tenant volume stormtenant.ingest_bytes_s over quotaAll tenants without quotasQuota + signature sampling + cold spillLogging platform
Agent fills host diskagent.buffer_pct_of_limit, disk pressureProduction on affected nodesHard buffer caps, drop oldestLogging platform
Mapping explosionindex.field_count growthTenants sharing the indexPer-stream schemas, raw fallbackLogging platform
Secrets storedCanary found unmaskedSecurity exposureRevoke, purge chunksSecurity + platform
Silent dropsAgent-read vs stored reconciliationOne stream's evidenceNever drop uncounted; replayLogging platform
Query of doomquery.bytes_scanned, queue waitAll searchersScan limits, async jobsLogging platform
Queue retention exceededqueue.oldest_unconsumed_age > 36hLoss of oldest dataScale processors, shed debugLogging platform
Object storage throttlingobjstore.put_throttled_totalWarm-tier writes lagLarger chunks, key prefix spreadingLogging platform

🎯 Staff Insight: The platform's health metric is in-quota ingest lag, not overall lag. "Overall lag will spike whenever a team storms, and it should — they're over quota. I page on lag for tenants within quota, because that number only moves when the platform is the problem."


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framing"Collect and search logs"Names debugging vs audit vs analytics; frames it as staying useful during the worst hour at bounded costFrames observability spend as a budget with owners and a cost-per-request target
Backpressure"Kafka buffers it"Bounded buffers at every stage, never block the app, counted drops, written loss orderLoss order as an org policy signed by security and engineering leadership
StorageIndex everything for 30 daysHot index over object storage, query on read, routing by time, scan limitsPrices tiers per GB-month; retention as a team budget decision
Multi-tenancyShared clusterIngest quotas with burst, query limits, cardinality and field limits, signature samplingQuotas tied to budgets and chargeback; exceptions process
Schema & PIIAuto-parse JSON; "don't log PII"Per-stream schemas, raw fallback; layered redaction with canaries; deletion-ready layoutLogging in the data classification standard; redaction coverage audited
OperationsCluster dashboardsIn-quota lag SLO, end-to-end reconciliation, agent fleet policyVolume reduction at source as a measured program across teams

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Designs for the incident storm"Volume spikes exactly when search matters most; quotas keep everyone else's logs flowing."
Written loss order"Debug sampled first, audit never; drops counted per tenant."
Index only what's queried"The hot index is an accelerator; object storage is the record."
Prices retention"90 days warm is about 1.5 PB of object storage — roughly $30K a month at list-price order of magnitude."
Contains schema failures"Per-stream schemas; a bad field from one team can't reject another team's logs."
Plans for deletion"Something will slip past redaction; chunks are partitioned so a purge is bounded."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Kafka buffers it" as the whole backpressure storyIgnores what happens when buffers fill and who loses
Agent retries without boundsA logging outage becomes a production outage
Indexing everything for months, unpricedUnaffordable at scale; no tiering
No per-tenant quotasOne team's debug flag blinds everyone
Shared dynamic mappingMapping explosions and type conflicts break other tenants
PII handled only at query timeSecrets persist in every replica and export

5.4 Common False Positives#

  • Deep Elasticsearch tuning ≠ log platform design. Shard sizing doesn't address loss policy, quotas or cost attribution.
  • "We use Kafka" ≠ backpressure. A queue moves the buffer; it doesn't decide what happens when the buffer fills.
  • Compression ratios ≠ cost control. Chargeback and volume reduction at source usually save more.
  • Regex redaction ≠ PII safety. Without canaries and a deletion path, it's a hope, not a control.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minDebugging primary, audit lane; incident storms, never block the app, cost, PII; numbers
Entities & API3–5 minTenant, stream, record, chunk; 429 as the backpressure contract
Architecture5–10 min≤ 8 boxes; agent → gateway → queue → processor → hot index + object storage
Backpressure10–18 minBounded buffers, loss order, signature sampling, queue retention
Storage tiers18–25 minHot index vs query on read, routing, retention tiers, prices
Multi-tenancy + schema25–32 minQuotas, query limits, cardinality, per-stream schemas, raw fallback
PII + operations32–38 minLayered redaction, canaries, deletion; in-quota lag SLO, reconciliation
Pivot / wrap38–45 minMulti-region, cost per team, build vs buy; close on who owns the bytes

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"We can't lose any logs"Do you know the cost?Audit lane never sampled; debug has a loss order; price the alternative
"Cut the logging bill by 40%"Cost leversChargeback, debug retention, sampling at source, tier moves — before storage engineering
"Make it multi-region"Data localityRegion-local ingest and storage; federated queries; residency
"Security needs a year of auth logs"Audit laneSeparate stream, write-once archive, gap detection
"A password was logged"DeletionRevoke, purge bounded chunks, fix at source, canary
"How do you know logs aren't missing?"End-to-end verificationAgent-read vs stored reconciliation per stream; canary streams

6.3 What to Deliberately Skip#

  • Index internals — "inverted index for the hot tier; columnar chunks for warm."
  • UI design — say what queries look like, not the screens.
  • Agent implementation details — name the bounds, skip the file-tailing mechanics.
  • Log formats per language — "structured JSON encouraged; raw accepted."
  • Exactly-once ingestion — at-least-once with dedupe by (agent, file, offset) where it matters; say so in one line.

6.4 Follow-Up Questions to Expect#

  1. "One team's volume goes 10× during an incident. What happens to everyone else's logs?"
  2. "The ingest tier is down for an hour. Walk me through every buffer between the app and storage."
  3. "What do you index, for how long, and what does 90 days cost?"
  4. "Two services log status with different types. What breaks?"
  5. "A secret was logged for six weeks. How do you delete it everywhere?"
  6. "How do you know logs aren't silently missing?"
  7. "Security needs a year of retention for auth logs. How does that differ from debug logs?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design a centralized logging system for our company."

Staff Answer

"What are the logs for — debugging services, security audit, or analytics? I'll assume debugging is primary, with a separate audit lane, and that analytics belongs in metrics. Numbers: 3,000 services on 50,000 nodes, about 2 GB/s average and up to 10 GB/s during a bad incident — roughly 170 TB/day raw. Seven days of fast search, 90 days of slower search.

Four constraints shape everything: volume spikes during incidents exactly when search matters most; logging must never take down the application; cost scales with bytes and producers don't see it; and logs will contain secrets. I'll go: entities and the backpressure contract → pipeline → loss order → storage tiers and cost → tenancy and schema → redaction → ownership."

Why this is L6:

  • Separates debugging from audit and analytics before drawing
  • Frames the problem around incident storms and cost, with numbers
  • Previews an outline that ends with ownership

What L7 adds:

  • Asks what share of infra spend observability is today and who owns that budget
  • Asks whether teams already run their own logging stacks — the consolidation question
  • Frames cost per request served as the success metric
❌ Common L5 Trap

"Install an agent on every host that ships logs to Kafka, consume into Elasticsearch with daily indices and 30-day retention, and give engineers Kibana."

Why this misses: Each piece is reasonable; together they don't say what happens when one team storms, what 30 days costs, how schema conflicts are contained or where secrets are removed. The next follow-up — "one team triples volume" — has no answer.


Drill 2: The Ingest Outage#

Prompt: "The ingest tier is unreachable for one hour. Walk me through every buffer between the app and storage."

Staff Answer

"The application writes to stdout; the container runtime writes to a local file with rotation, say 5 files of 100 MB. The app is unaffected — it never talks to the network for logging. The agent can't ship, so its 64 MB memory buffer fills in seconds and spills to a 1 GB disk buffer. A typical node producing 200 KB/s fills that in about 80 minutes — so an hour is covered for most nodes. Chatty nodes at 1 MB/s fill it in ~17 minutes; after that the agent stops reading, falls behind the rotating files, and loses the oldest unread data when files rotate — counted in agent.dropped_lines_total per tenant.

When ingest returns, agents drain with jittered backoff so 50,000 nodes don't stampede the gateway; the gateway's per-tenant quotas allow a recovery burst. Downstream, the queue had nothing to absorb in this scenario — it protects against the indexing tier failing, not ingest. Afterward, I'd report per-tenant loss and consider larger disk buffers for chatty node pools."

Why this is L6:

  • Walks every buffer with sizes and fill times
  • Never blocks the application; drops are bounded and counted
  • Handles recovery stampede and reports loss

What L7 adds:

  • Decides buffer sizing as a cost-vs-loss tradeoff per node pool, agreed with the teams on those pools
  • Adds a zone-level chaos test that blackholes ingest monthly
❌ Common L5 Trap

"The agent retries until the backend comes back, so no logs are lost."

Why this misses: Retrying requires buffering, and unbounded buffering fills the host's disk or memory. The candidate hasn't bounded it — and if the agent's buffer is unbounded, an hour-long logging outage becomes a production outage.


Drill 3: Make It Concrete — Size and Price the Tiers#

Prompt: "Size the storage tiers for 2 GB/s, 7 days hot, 90 days warm. Roughly what does it cost?"

Staff Answer

"2 GB/s × 86,400 ≈ 170 TB/day raw. Warm tier: with ~10× compression, ~17 TB/day; 90 days ≈ 1.5 PB in object storage. At roughly $0.02 per GB-month, that's about $30K a month, plus request costs — which is why chunks should be large, tens to hundreds of MB, not tiny.

Hot tier: if I indexed everything, 7 days of 170 TB with index overhead roughly comparable to raw, compressed maybe 2×, replicated twice, is on the order of 1–2 PB of fast storage — several times the cost of the warm tier. So I'd cut it: debug-level streams hot for only 24 hours, and over-quota data skipped. If debug is half the volume, hot drops to around 0.6–1 PB.

Queue: 48 hours × 2 GB/s ≈ 350 TB raw, ~50 TB compressed, replicated ×3. Compute: processors and indexers sized to peak, 5× average, which argues for elastic processing."

Why this is L6:

  • Sizes each tier with compression, replication and retention
  • Prices the warm tier and shows why the hot tier must be smaller
  • Notices request costs and chunk sizing

What L7 adds:

  • Turns this into cost per GB ingested per tier and publishes it for chargeback
  • Uses the cost estimator to compare self-hosted against vendor pricing at 3-year volume
❌ Common L5 Trap

"170 TB a day times 90 days is 15 PB. We'll need a big Elasticsearch cluster."

Why this misses: Ignores compression, replication and tiering, and doesn't price anything. The answer implies indexing 15 PB on fast storage — an unaffordable design presented as a sizing exercise.


Drill 4: Index or Scan?#

Prompt: "Why not just index everything? Engineers want fast search over all 90 days."

Staff Answer

"Because the cost scales with every byte for the full retention while the queries don't. If 90% of searches hit the last 24 hours, indexing day 60 buys fast answers to a handful of queries at the price of indexing petabytes. I'd measure the query distribution by age — it's usually steep — and size the hot window to cover, say, 95% of queries.

For older data, object storage with good pruning: partition by tenant, stream and hour; per-chunk metadata with time bounds and labels; token bloom filters so a search for a request ID skips most chunks. A 30-day search over one service becomes tens of seconds. For the rare broad search, an async query job. If a team truly needs fast search over 90 days — security investigations — they get a longer hot window for a small, specific stream and pay for it."

Why this is L6:

  • Argues from the query age distribution, with a measurement
  • Describes how cold queries stay usable: pruning, bloom filters, async jobs
  • Offers longer hot retention as a paid exception

What L7 adds:

  • Publishes tier prices so teams make the choice with their own budget
  • Revisits the hot window yearly as query patterns and costs shift
❌ Common L5 Trap

"Fast search matters during incidents, so we should index everything."

Why this misses: Incidents query the last hours, which the hot tier covers. Indexing 90 days to serve post-mortems that could wait 30 seconds is the most expensive way to buy convenience.


Drill 5: The Noisy Tenant#

Prompt: "A team turns on debug logging in production and their volume goes from 60 MB/s to 2.4 GB/s. Everyone else's logs are delayed. Fix it."

Staff Answer

"Immediately: the processor enforces that tenant's quota — say 100 MB/s with a 2× burst for 10 minutes. Above it, signature sampling: group lines by message template, keep the first 100 per signature per minute in full, and emit counts for the rest. Their excess skips the hot index and goes compressed to object storage if write capacity allows, so they can still query it, slower. Other tenants' lag recovers within a minute because indexers are no longer saturated.

Structurally: quotas enforced at the gateway and processor by default; debug streams kept hot for 24 hours; a 'debug window' feature to raise a quota temporarily for named pods with the cost shown up front. And the team sees a 'sampled — over quota' marker in their search results, so they understand what they're looking at."

Why this is L6:

  • Isolates the storming tenant with quotas and signature sampling
  • Preserves information rather than dropping blindly
  • Gives teams a sanctioned way to get more volume temporarily

What L7 adds:

  • Ties quotas to team budgets so raising one is a budget conversation
  • Tracks top volume producers monthly as a reduction program
❌ Common L5 Trap

"Add more indexing nodes so we can handle the extra volume."

Why this misses: Scaling for one team's debug flag means paying peak capacity for a misconfiguration, and it takes longer than the storm. Without isolation, the next team's flag does it again.


Drill 6: Cardinality and Mapping Explosions#

Prompt: "A service starts putting user IDs into a label — or into JSON keys. What happens, and how do you prevent it?"

Staff Answer

"As a label, user IDs create one stream per user: millions of tiny streams, an index that grows with users and chunks too small to compress well — a label-indexed store degrades for everyone sharing it. As JSON keys with dynamic mapping, each user ID becomes a field; a shared index hits its field limit and starts rejecting writes for other tenants too.

Prevention: labels are a short, fixed set per stream — service, env, cluster, level — with a limit on labels per stream and distinct values per label per day, enforced at the gateway with a clear error. Parsed fields are per-stream registered schemas; unknown fields are kept in the raw body, searchable as text, never auto-added to a shared mapping. High-cardinality values like user and request IDs belong in the body or as indexed tokens in bloom filters, not as labels. And the platform reports per-stream cardinality so teams see it before limits hit."

Why this is L6:

  • Explains the failure in both label-indexed and field-indexed stores
  • Enforces limits at ingest with clear errors
  • Gives a correct home for high-cardinality values

What L7 adds:

  • Makes cardinality limits part of the logging standard with an exceptions process
  • Ships linting in the logging libraries that flags dynamic keys before deploy
❌ Common L5 Trap

"Increase the field limit on the index."

Why this misses: It postpones the failure and makes the index slower and heavier for every tenant. The fix is preventing unbounded fields from entering a shared mapping at all.


Drill 7: Build vs Buy#

Prompt: "Should we run our own logging platform or use a vendor?"

Staff Answer

"Below a few TB/day, buy: a vendor gives search, retention tiers and alerting for less than one engineer's cost. Above tens of TB/day, vendor pricing per GB ingested often exceeds the cost of a platform team plus infrastructure, and that's when building — usually on open-source components and object storage — pays off. I'd model 3-year volume, not today's, because log volume tends to grow faster than traffic.

Whichever we pick, three things stay ours: the agent and pipeline layer, so we can redact, sample and route before bytes leave our network — and switch vendors without touching every host; the loss and retention policies; and chargeback to teams. A common hybrid: vendor for hot search, our own object storage for the long tail and audit archive."

Why this is L6:

  • Gives a volume threshold and models future growth
  • Keeps the pipeline under our control for redaction, cost and portability
  • Proposes a hybrid split by tier

What L7 adds:

  • Negotiates vendor contracts on committed volume with an exit plan and data export
  • Treats vendor lock-in risk as a line item: query language, dashboards, alert definitions
❌ Common L5 Trap

"Use a vendor — logging isn't our core business."

Why this misses: True at small scale, but it ignores the volume at which per-GB pricing dominates, and it hands redaction and routing to whoever owns the agent. The decision is a threshold, not a slogan.


Drill 8: Changing a Parsing Rule Without Losing Logs#

Prompt: "You need to change the parser for 400 services' access logs to a new format. How do you ship it?"

Staff Answer

"Parsers are versioned config per stream, not code baked into processors. I'd deploy the new parser in shadow: both versions parse a copy of live traffic; we compare field-level output and parse-failure rates per stream for a day. Streams with differences above a threshold get reviewed before switching.

Rollout by stream: 1% of streams, then 10%, then all, watching processor.parse_errors{stream} and end-to-end line counts. The invariant that makes it safe: a line that fails to parse is stored raw with a parse-error tag, never dropped — so the worst case of a bad parser is less-structured data, not missing data. Rollback is a config change. Dashboards and saved queries that depend on renamed fields get a compatibility alias for 30 days."

Why this is L6:

  • Shadow parsing with field-level comparison before switching
  • Staged rollout per stream with clear metrics
  • The raw-fallback invariant bounds the blast radius

What L7 adds:

  • Gives teams ownership of their stream parsers with platform-provided test harnesses
  • Treats field renames as API changes with deprecation windows
❌ Common L5 Trap

"Update the parser in the processors and redeploy."

Why this misses: A parser bug silently drops or mangles lines for 400 services at once, and nobody notices until an incident needs the missing logs.


Drill 9: Cut the Bill by 40%#

Prompt: "Finance says logging costs too much. Cut it by 40% without hurting incident response."

Staff Answer

"Start with attribution: cost per stream, with query counts. Usually 10–20 streams drive half the volume, and some are rarely queried. Levers in order of ROI:

  1. Volume at source: debug off by default in production, health-check and access-log sampling (keep errors, sample 200s at 1–10%), duplicate stack traces collapsed. Often 30–50% of volume.
  2. Tiering: debug hot for 24 hours, not 7 days; warm default 30 days, longer only if paid for.
  3. Hot-index scope: index only labels and registered fields, keep body text in compressed chunks with bloom filters.
  4. Chargeback: teams see their bill monthly; behavior follows.

Incident response is protected because error-level logs are never sampled, the hot window still covers recent hours, and quotas keep storms isolated. I'd measure time-to-find for a set of standard incident queries before and after."

Why this is L6:

  • Starts with attribution and query counts, not storage tricks
  • Orders levers by ROI with rough impact
  • Defines a metric to prove incident response isn't hurt

What L7 adds:

  • Sets an org-wide observability cost target as a share of infra spend
  • Makes logging cost part of service launch reviews
❌ Common L5 Trap

"Switch to cheaper instances and compress better."

Why this misses: Compression and instance tweaks save a fraction of what reducing volume and retention saves, and they don't change the behavior that keeps growing the bill.


Drill 10: Multi-Region#

Prompt: "We run in four regions. How does logging work across them?"

Staff Answer

"Region-local ingest, processing and storage. Shipping logs across regions doubles egress cost and adds a cross-region dependency to every write, and many logs are subject to residency rules. Each region has its own agents' gateway, queue, hot tier and object storage bucket.

Queries are federated: the query frontend fans out to the regions in scope, each executes locally and returns results or partial aggregates, and the frontend merges. Most incident queries target one region anyway. If a region is unreachable, results are marked partial rather than failing. Audit logs may need a copy in a second region for durability — that's a deliberate, small, replicated stream, not all logs. For data residency, EU logs never leave the EU, including query results cached elsewhere."

Why this is L6:

  • Region-local storage avoids egress and cross-region write dependencies
  • Federated queries with partial-result semantics
  • Replicates only what needs replication; respects residency

What L7 adds:

  • Prices cross-region egress explicitly and makes replication an exception with an owner
  • Defines residency per tenant contract across all observability data, not just logs
❌ Common L5 Trap

"Ship all logs to one central region so engineers can search everything in one place."

Why this misses: Pays cross-region egress on every byte, makes every region's logging depend on one region's availability, and breaks residency. Federation gives the single search box without moving the data.


8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Logs Go Dark During the Outage#

Context: A database failover causes errors across 600 services. Log volume jumps from 2 GB/s to 11 GB/s within two minutes. Ingest lag for all tenants climbs to 14 minutes, and incident responders can't see current errors. The on-call escalates to you.

Questions to Surface First:

  • Where is the bottleneck — gateway, queue, processors or hot-tier indexing?
  • Is the surge broad (600 services) or dominated by a few?
  • Are quotas being enforced, or are all tenants bursting at once?
  • Are responders' queries themselves loading the hot tier?

Typical L5 Approach: Scales the indexing cluster. Shard rebalancing adds load; lag gets worse for 20 minutes before it improves.

Staff Approach: Finds the processors keeping up but hot-tier indexing saturated. Since every tenant is bursting, per-tenant quotas don't help — so applies the platform-wide loss order: debug and info streams skip the hot index (still written to object storage), error and warn keep flowing. Hot-tier lag drops to 20 seconds for error-level logs within 3 minutes.

Principal Approach: Recognizes that broad incidents create correlated bursts that per-tenant quotas can't handle and that log volume during failures is predictable. Builds a platform-level "incident mode" — pre-agreed level-based shedding — and drives a program to collapse repeated error logs at the source (rate-limited loggers per message template).

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Stage metrics: gateway ok, queue ok, processors ok, indexer CPU 100%. Enable level-based shedding: debug/info to object storage only.
Triage70% of surge is repeated stack traces from retry loops; signature sampling would cut it ~20×.
Quick fixEnable signature sampling platform-wide for the incident; error-level examples preserved.
GuardrailsAlert on ingest.lag_p99_s{level=error} > 60 separately from other levels.
Post-mortemIncident mode as a documented, one-click policy; logging libraries get per-template rate limits.

Metrics to Watch: ingest.lag_p99_s{level}, indexer.cpu_util, sampling.suppressed_lines_total{signature}, query.queue_wait_p99

Organizational Follow-up: engineering leadership approves incident-mode shedding as policy; library teams ship template rate limiting.

Ownership Question: "Who decides which logs are shed during a company-wide incident?" Staff answer: The policy is decided in advance by engineering leadership and security; the logging on-call executes it without asking anyone during the incident.

Key Takeaway: "When everyone storms at once, quotas don't help — a pre-agreed level-based loss order does."

What clears the Staff bar:

  • Localizes the bottleneck before scaling
  • Uses a loss order that preserves error-level logs
  • Pushes the fix to the source: rate-limited repeated errors

Deep Dive 2: Silent Failure — Three Days of Missing Logs From One Region#

Context: During a post-mortem, an engineer notices that logs from one region's newest node pool are missing for the last three days. No alerts fired. Ingest lag, error rates and queue depth all look normal.

Questions to Surface First:

  • Are the agents on those nodes running, and what version?
  • What do agent-side counters say about lines read vs shipped?
  • Is there any end-to-end reconciliation between agents and storage?
  • Did anything change for that node pool three days ago?

Typical L5 Approach: Restarts agents on the node pool; logs start flowing. The three days are lost and the cause is unknown.

Staff Approach: Finds the node pool uses a new container runtime that writes logs to a different path; the agent's file discovery doesn't match it, so the agent is healthy and shipping nothing. Fixes discovery; recovers what's still on the nodes' rotated files. Adds end-to-end reconciliation: per node, expected log activity versus received; alerts on nodes with running pods and zero received lines.

Principal Approach: Treats "absence of data" as a first-class signal across observability. Every node and every service has an expected heartbeat in logs, metrics and traces; silence pages. Node pool changes require an observability checklist.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Check agent.lines_shipped{node_pool}: zero for 3 days. Agents healthy.
TriageNew runtime log path not in agent discovery globs; agent reports "no files" as normal.
Quick fixUpdate discovery config; ship surviving rotated files; mark the gap in the platform status.
GuardrailsAlert: node with running pods and zero shipped lines for 10 min; per-service "silent stream" alert.
Post-mortemNode pool rollouts include an observability smoke test: synthetic log line found in search within 60 s.

Metrics to Watch: agent.lines_shipped{node}, stream.silent_minutes{service}, canary.log_found_latency_s

Organizational Follow-up: infrastructure team adds the logging smoke test to node pool rollout automation.

Ownership Question: "Who owns noticing that logs stopped?" Staff answer: The logging platform owns end-to-end completeness signals; the infrastructure team owns running the smoke test when they change the node image.

Key Takeaway: "A healthy agent shipping nothing looks exactly like a quiet service. Alert on silence, not just on errors."

What clears the Staff bar:

  • Recognizes missing data needs its own detection
  • Adds end-to-end reconciliation and canaries
  • Gets the change process for node pools to include observability

Deep Dive 3: Large-Customer Onboarding — A Team With 400 TB/Day#

Context: The CDN team wants to move edge request logs onto the platform: 400 TB/day raw, more than double the platform's current total. They want 30 days of search and dashboards over status codes and latencies.

Questions to Surface First:

  • Are these logs searched for needles, or aggregated for dashboards?
  • What fraction is ever queried? Do they need every 200 response?
  • Is the format regular enough to treat as structured events?
  • Who pays, and what's their budget?

Typical L5 Approach: Triples the hot cluster and onboards them into the shared index. Ingest saturates during their daily peak; every other tenant's lag rises.

Staff Approach: Recognizes these are events, not debug logs: fixed schema, aggregation-heavy queries. Routes them to a dedicated columnar pipeline — typed columns, high compression, fast aggregations — with metrics extracted at ingest for dashboards. Errors (4xx/5xx) and a 1% sample of 2xx are kept searchable for 30 days; full data in object storage for 7 days for investigations. Their quota and cost are separate.

Principal Approach: Uses the onboarding to define a platform tier for high-volume structured events, with its own pricing, and sets the rule that dashboards over logs must be served by extracted metrics.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (design review)Query analysis: 95% of CDN team queries are aggregations by status, POP, path prefix.
TriageFull ingest into the shared hot tier would cost several times their budget and endanger other tenants.
Quick fixDedicated columnar lane; metrics extracted at ingest; errors + 1% of 2xx searchable 30 days.
GuardrailsSeparate quota and processing pool; no shared hot-tier capacity.
Post-mortem (pre-mortem)What happens during a global CDN incident? Their lane sheds 2xx sampling to 0.1% automatically.

Metrics to Watch: lane.cdn.ingest_lag_s, lane.cdn.bytes_stored, extracted_metrics.lag_s, shared-tier lag (must not move)

Organizational Follow-up: CDN team's budget covers the lane; finance sees it as a separate line.

Ownership Question: "Who owns the CDN team's dashboards?" Staff answer: The CDN team owns the queries and sampling choices; the platform owns the lane's isolation and the extraction pipeline.

Key Takeaway: "High-volume regular logs are events. Give them a columnar lane and extracted metrics, not the shared search index."

What clears the Staff bar:

  • Classifies the data by query pattern before choosing storage
  • Isolates the new tenant's capacity and cost
  • Uses sampling and metric extraction instead of indexing everything

Deep Dive 4: Post-Mortem — Card Numbers in Logs for 41 Days#

Context: A compliance scan finds full payment card numbers in log search results. A checkout service started logging raw request bodies on validation errors 41 days ago. The data is in the hot tier, object storage, and a weekly export to the analytics warehouse. You're leading the post-mortem.

Questions to Surface First:

  • Why didn't pipeline redaction catch card numbers?
  • Where exactly does the data exist — tiers, replicas, exports, backups?
  • Who accessed those log streams during the window?
  • Can we purge object storage without rewriting the entire 41 days?

Typical L5 Approach: Adds a regex for 16-digit numbers to the pipeline and deletes the hot index for those days.

Staff Approach: Finds the card numbers were embedded in a URL-encoded body the regex didn't decode. Adds Luhn-validated detection on decoded content, redaction in the logging library for request bodies, and a canary that sends a test card number through checkout daily. Purges hot-tier documents by query, rewrites only the affected tenant-and-hour chunks in object storage, deletes the export partitions, and produces an access report for compliance.

Principal Approach: Treats this as a gap in data governance: logs had no classification and no owner for what fields services emit. Puts logging under the data classification standard, requires that payment services' logs pass through stricter allowlist-only redaction, and makes canary pass rate a compliance control.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Restrict access to affected streams; stop the export job; hotfix checkout to stop logging bodies.
TriageRegex ran on raw text; URL-encoded digits bypassed it. Data in hot tier, 41 days of chunks, 6 export partitions.
Quick fixDecode-then-detect with Luhn validation; purge hot docs; rewrite affected chunks; delete export partitions.
GuardrailsDaily card-number canary through checkout; alert if found unmasked anywhere.
Post-mortemPayment-domain streams switch to allowlist-only fields; exports require classification review.

Metrics to Watch: redaction.canary_pass{runtime,service}, redaction.matches_total{rule}, purge.chunks_rewritten

Organizational Follow-up: compliance notified with access report; security reviews logging libraries across languages.

Ownership Question: "Who owns what a service writes to logs?" Staff answer: The service team owns what they emit; security owns the redaction rules and classification; the platform owns making redaction and purges reliable.

Key Takeaway: "Redaction you don't test with canaries is redaction you hope works. And design storage so you can delete from it."

What clears the Staff bar:

  • Finds why redaction failed, not just adds another rule
  • Purges every copy, including exports, with bounded rewrites
  • Adds continuous verification with canaries

Deep Dive 5: Multi-Region Expansion — Residency for EU Logs#

Context: The company is launching EU data residency. Logs from EU workloads may contain personal data and must stay in the EU. Today all logs flow to a central US platform.

Questions to Surface First:

  • Which services process EU personal data, and do their logs contain it?
  • Can we run a full regional stack, or only regional storage?
  • How do engineers outside the EU search EU logs for incidents?
  • Where do derived artifacts go — extracted metrics, alerts, exports?

Typical L5 Approach: Deploys a second cluster in the EU and routes EU logs there. Leaves the analytics export and alert notifications, which include log excerpts, flowing to the US.

Staff Approach: Full regional stack in the EU: agents, gateway, queue, processors, hot tier and object storage. Federated queries from the global frontend execute in the EU and return results to authorized users under an access policy; results aren't cached outside the EU. Extracted metrics without personal data can flow globally; alert notifications carry links, not excerpts.

Principal Approach: Defines residency as a tenant attribute honored by every observability system — logs, traces, metrics labels, alert payloads — with automated marker tests proving compliance, and a single access-policy layer for cross-region queries.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Inventory: raw logs, hot index, chunks, exports, alert payloads, query caches, extracted metrics.
TriageAlert payloads and analytics exports carry excerpts → residency violations.
Quick fixEU stack; alerts link back to EU search; exports from EU go to an EU warehouse only.
GuardrailsMarker log lines with synthetic personal data in EU services; weekly scan of non-EU stores.
Post-mortem (pre-launch)Cross-region query access audited; on-call outside EU uses federated queries with logging of access.

Metrics to Watch: residency.marker_found_outside_region (must be 0), eu.ingest.lag_p99_s, federated_query.eu.count{user_region}

Organizational Follow-up: legal approves the access policy for non-EU responders; on-call training covers federated search.

Ownership Question: "Who proves EU logs stayed in the EU?" Staff answer: The logging platform proves it for its stores with marker tests; privacy and compliance own the company-level attestation.

Key Takeaway: "Residency includes everything derived from logs — alerts, exports and caches, not just the index."

What clears the Staff bar:

  • Inventories derived data, not just primary storage
  • Uses federated queries instead of moving data
  • Proves residency with tests

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why log volume spikes during incidents and why the platform must stay useful then
  • Design bounded buffering at every stage with a written loss order that never blocks applications
  • Choose what to index (hot window) and what to query on read (object storage), and price both
  • Enforce per-tenant ingest quotas, query limits and cardinality limits, with signature sampling under overload
  • Contain schema conflicts with per-stream schemas and raw fallback
  • Layer PII redaction, verify it with canaries, and design storage so purges are bounded
  • Size and price storage tiers, and attribute cost to producing teams
  • Detect silent loss with end-to-end reconciliation and silence alerts

The Bar for This Question#

Mid-level (L4): Ships logs from hosts to a search cluster with dashboards. Works at small scale. No backpressure story, no quotas, one retention, no redaction.

Senior (L5): Adds agents, Kafka, Elasticsearch with daily indices, retention policies and some parsing. The gap: "Kafka buffers it" ends the backpressure discussion; indexes everything for the full retention without pricing it; no per-tenant isolation; shared dynamic mapping; PII left to developer discipline. The design works on a quiet day and fails during the first incident storm.

Staff+ (L6): Frames the problem as staying useful during the worst hour at a bounded cost within the first five minutes. Bounds every buffer and writes down the loss order. Indexes the hot window and queries the rest on read, with prices. Isolates tenants with quotas, query limits and cardinality limits. Contains schema failures per stream. Layers redaction with canaries and a deletion path. Names who pays — storming tenants lose fidelity on their own repeated logs; everyone else pays nothing. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 Most Logs Are Never Read#

DataTypical Fate
Debug lines older than a dayRarely queried
Successful request logsMostly aggregated, rarely searched
Error logs from the last hoursThe queries that matter

The Staff position: Index what's searched; keep the rest cheaply or not at all. Volume reduction at the source beats storage engineering.

Why this matters in interviews: Pricing retention and admitting most bytes are cold is a Staff signal.

10.2 "Never Lose a Log" Is a Promise You Can't Afford#

PromiseCost
No loss, everBlock apps or buffer unboundedly
No loss for auditSmall reserved lane — affordable
Ordered loss for debugPolicy, sampling — cheap

The Staff position: Promise completeness for audit; publish a loss order for everything else.

Why this matters in interviews: It shows you design the failure behavior instead of pretending there isn't one.

10.3 Logs Are a Bad Metrics System#

NeedWrong ToolRight Tool
Error rate dashboardCount log lines every refreshEmit a counter
Latency percentilesParse durations from logsHistogram metric
Request flowCorrelate log lines by handTrace

The Staff position: Extract metrics at ingest if you must; better, emit them at the source.

Why this matters in interviews: Redirecting analytics to metrics is a strong scoping move.

10.4 Shared Dynamic Mapping Is a Multi-Tenant Bug#

The Staff position: Any shared index where one tenant's new field can reject another tenant's logs is an isolation failure. Per-stream schemas and raw fallback aren't polish; they're tenancy.

Why this matters in interviews: It's the failure Senior designs don't see until it happens.

10.5 Chargeback Saves More Than Compression#

The Staff position: Teams who see their logging bill next to their query counts cut volume. No codec delivers that.

Why this matters in interviews: It shows you treat cost as an organizational problem with a technical assist.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer builds a log platform that stays useful during incidents, isolates tenants and prices its tiers. The Principal engineer notices that logs, metrics and traces are three pipelines with three agents on every node, three bills, three retention policies and three places personal data can leak — and that observability spend is growing faster than revenue. The L7 problem is observability as one budget with one collection layer: a shared agent and pipeline, per-signal storage, consistent tenancy and redaction, and a cost target the organization manages to.

The Org-Level Fault Line#

One observability pipeline vs per-signal stacks.

OptionWhat WorksWhat BreaksWho Pays
Separate stacks for logs, metrics, tracesEach team optimizes its signalThree agents per node, three tenancy models, three redaction gapsNodes (overhead), security (gaps), finance (fragmented spend)
One collection layer, per-signal backendsOne agent, one tenancy and redaction model, one cost reportPipeline team must serve diverse signalsPlatform team (scope)
One vendor for everythingSimplest operationsLock-in; per-GB pricing at scaleFinance (vendor cost)

🧭 Principal Move: "One collection layer — agent, gateway, redaction, tenancy, quotas and cost attribution — shared by logs, metrics and traces. Backends stay specialized. Every team gets one observability bill with three lines."

Cost Model#

Assumptions: object storage at ~$0.02/GB-month order of magnitude; ~10× compression; hot tier on replicated fast storage; fully loaded engineer ~$250K/year. Illustrative ranges.

ScaleVolumeInfra ($/month)HeadcountOn-call LoadNotes
Startup100 GB/day~$1–5K vendor or small clusterPart of one engShared rotationBuy
Growth20 TB/day~$40–120K (hot tier, object storage, queue, compute)4–6 engDedicated rotationChargeback starts paying for itself
Large170+ TB/day, 4 regions~$300K–1M12–20 eng (pipeline, storage, query, agents)Per-region rotationsVolume reduction program is the top ROI

The pricing insight: the hot index usually costs more than the warm tier despite holding a fraction of the days. Shrinking the hot window for debug-level streams and reducing volume at the source typically saves more than any storage-engine migration.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Agent deployed to every nodeOne-way-ishReplacing it touches the whole fleet
Query language exposed to engineersOne-way-ishSaved queries, dashboards, alerts depend on it
Storing unredacted data in immutable archiveOne-wayPurges are expensive; legal exposure persists
Label schema and cardinality policyOne-way-ishChanging labels breaks queries and alerts
Hot window lengthTwo-wayConfig
Storage engine for warm tierTwo-wayMigration, but data is in open formats in object storage
Quotas per tenantTwo-wayConfig

The Standard I'd Write#

RFC-OBS-003: Logging Standard
Status: Approved   Owners: Observability Platform + Security

Scope
  Every service and job that emits logs in any environment.

MUST
  1. Emit logs to stdout or the platform agent; never ship directly from
     application code to storage.
  2. Never block request handling on logging.
  3. Use only the standard label set; high-cardinality values go in the body.
  4. Pass through platform redaction; payment and identity domains use
     allowlist-only fields.
  5. Route security and audit events to the audit lane.

SHOULD
  1. Emit structured JSON with a registered schema per stream.
  2. Rate-limit repeated messages per template in the logging library.
  3. Emit metrics for anything counted on a dashboard.

Exceptions
  Filed with the Observability Platform; security review for redaction
  exceptions; time-boxed to two quarters.

Success metrics
  - In-quota ingest-to-searchable p99: under 30 s
  - Redaction canary pass rate: 100%
  - Logging cost per 1M requests served: tracked per team, trending down
  - Unplanned loss for in-quota tenants: under 0.1%

What I'd Tell the VP#

"Logging is our second-largest infrastructure cost after compute, and it's growing faster than traffic because nobody who writes logs sees the bill. During our last two major outages, the logging platform fell behind exactly when we needed it. I'm proposing three things: per-team quotas and monthly cost reports so teams own their volume, a tiered design where only the last week is fast and expensive, and a pre-agreed policy for what gets dropped during a company-wide incident. I expect a 30–40% cost reduction within two quarters, mostly from teams cutting volume once they see it, and logs that stay searchable during outages."

Principal Interview Signals#

SignalWhat It Sounds Like
Prices observability as a budget"Logging cost per million requests is the number I'd hold teams to."
Identifies one-way doors"Unredacted data in an immutable archive is forever; the hot window isn't."
Redraws ownership"Teams own their bytes; security owns redaction rules; the platform owns isolation."
Unifies collection"One agent per node for all signals, one tenancy model, one bill."
Knows when not to build"Below a few TB a day, buy, and keep the agent layer ours."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a great logging platform" — correct, but ignores the metrics and tracing agents on the same nodes.
  • "Teams can see their usage in a dashboard" — good, but without budgets or targets nothing changes.
  • "We'll add an EU cluster" — no mention of alert payloads, exports and caches derived from logs.

Appendices

Appendix A: Mechanics in Depth#

A.1 Agent Buffering Loop#

config:
  mem_buffer_max = 64MB
  disk_buffer_max = min(1GB, 10% of free disk)
  batch_max = 1MB or 1s

loop:
  read new lines from discovered files (respect rotation, track offsets)
  redact_patterns(line)               # tokens, keys, Luhn-valid card numbers
  append to mem_buffer
  if mem_buffer full: spill oldest batch to disk_buffer
  if disk_buffer full: drop oldest batch; dropped_lines[tenant] += n
  send batches oldest-first:
     204 -> commit offsets
     429 -> backoff(Retry-After, jitter); keep buffered
     5xx/timeout -> exponential backoff with jitter; keep buffered

A.2 Signature Sampling#

signature(line) = hash(template(line))    # numbers, IDs, hex stripped
per tenant, per minute:
  if tenant under quota: keep all
  else:
    count[sig] += 1
    if count[sig] <= 100: keep full line (hot + warm)
    else: warm only (if capacity), emit summary record every 10s:
          { sig, example, suppressed_count }
audit streams: never sampled

A.3 Query Routing#

route(query):
  require label selector with at least one of {service, tenant stream}
  split time range into hot (<= 7d) and warm (> 7d) parts
  hot part  -> index search, limit by tenant concurrency
  warm part -> list chunks by (tenant, stream, hour) -> prune by bloom filter
               -> parallel scan, cap bytes_scanned per query
               -> if over cap: return partial + offer async job
  merge, sort by ts, apply limit

Appendix B: Data Model#

CREATE TABLE tenants (
  tenant_id           TEXT PRIMARY KEY,
  cost_center         TEXT NOT NULL,
  ingest_quota_bps    BIGINT NOT NULL,
  burst_bytes         BIGINT NOT NULL,
  max_labels          INT NOT NULL DEFAULT 12,
  max_query_bytes     BIGINT NOT NULL
);

CREATE TABLE stream_policies (
  tenant_id     TEXT NOT NULL REFERENCES tenants,
  selector      TEXT NOT NULL,          -- e.g. {service="checkout",level="debug"}
  hot_days      INT NOT NULL DEFAULT 7,
  warm_days     INT NOT NULL DEFAULT 30,
  schema_ref    TEXT,                   -- registered parser/schema version
  redaction     TEXT NOT NULL DEFAULT 'standard',  -- standard | allowlist
  PRIMARY KEY (tenant_id, selector)
);

-- chunk metadata (object storage index):
-- chunks(tenant_id, stream_hash, hour, chunk_id, t_min, t_max, bytes,
--        bloom_ref, schema_ref)  -- object key: tenant/stream_hash/yyyy/mm/dd/hh/chunk_id

Appendix C: Coordination Mechanisms#

C.1 Ingest Under Quota#

Diagram: C.1 Ingest Under Quota

C.2 Quick Comparison#

MechanismGuaranteesFailure ModeUse For
Bounded agent bufferHost protectedLong outages lose oldest dataEvery node
429 + Retry-AfterAgents slow down, don't dropAgents ignore itIngest contract
Durable queueIndexing outages absorbedRetention exceededBetween ingest and processing
Tenant quota + burstStorms isolatedQuota too low for real needsEvery tenant
Signature samplingInformation kept under overloadCounts approximateOver-quota tenants
Per-stream schemaConflicts containedUnregistered streams less queryableStructured logs
Redaction canariesCoverage verifiedCanaries not representativeEvery runtime
End-to-end reconciliationSilent loss detectedCounters themselves lostEvery stream

Appendix D: API Contract & Client Behavior#

  • Agents batch up to 1 MB or 1 second, gzip, and retry on 429/5xx with jittered backoff; they never drop on 429.
  • Agents report dropped_lines_total, buffer_bytes and lines_shipped per tenant as metrics.
  • Applications log to stdout; structured JSON preferred, one record per line, timestamps in UTC with milliseconds.
  • Labels: a fixed, small set per stream; request and user IDs go in the body.
  • Queries beyond the hot window require a label selector; large scans run as async jobs.

Appendix E: Observability#

Core metrics:

  • ingest.lag_p99_s{in_quota}, ingest.lag_p99_s{level}
  • agent.dropped_lines_total{tenant}, agent.buffer_pct_of_limit, agent.lines_shipped{node}
  • tenant.ingest_bytes_s, tenant.quota_exceeded, sampling.suppressed_lines_total
  • index.field_count, ingest.rejected_total{reason}, processor.dropped_total{reason}
  • query.bytes_scanned{user}, query.queue_wait_p99
  • redaction.canary_pass, cost.per_gb_by_team

Critical alerts:

AlertThresholdSeverity
In-quota ingest lag p99> 60 s for 5 minPage
Processor drops without reason> 0Page
Redaction canary found unmaskedanySev-1 page (security)
Node with pods and zero shipped lines> 10 minTicket → page if > 5% of nodes
Queue oldest unconsumed age> 36 hPage
Agent buffer > 80% of limit fleet-wide> 5% of nodesTicket

Debugging the silent failure: ingest lag and error rates look fine when data never arrives. Reconcile lines read by agents against lines stored, alert on silent streams and nodes, and run a canary log line through every node pool.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 1 TB/dayVendor, or one search cluster with 14–30 daysFirst incident storm
1–20 TB/dayAgents with bounded buffers, queue, quotas, per-stream schemasHot-tier cost
20–200 TB/dayHot index + object storage, signature sampling, chargebackCross-region egress, multi-signal agents
> 200 TB/day, multi-regionRegional stacks, federated queries, shared observability pipelineOrg cost governance

What you don't build on day one: tiered storage, chargeback, signature sampling, federated queries, an audit lane with tamper evidence. Each has a trigger in Section 11.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Quotas: bytes per second with burst per tenant; raised through a budget request, not a ticket to the platform team.
  • Query fairness: concurrent queries and bytes scanned per tenant and per user; long scans queued as async jobs.
  • Cardinality: labels per stream and distinct values per label per day capped; violations rejected with a clear error.
  • Cost attribution: bytes ingested, GB-months per tier and bytes scanned reported per team monthly, with top streams and their query counts. Rehearse the sizing with the back-of-envelope calculator and the cost estimator.
  1. Loading the index…