Technologies referenced in this case study: Apache Kafka · Elasticsearch · OLAP Databases · Kubernetes · Apache Flink · Time-Series Databases
Related: Metrics & Alerting Platform · Distributed Tracing · Object Storage · Search Engine · Backpressure & Load Shedding · Multi-Tenancy · Batch and Stream Pipelines · ClickHouse vs Druid vs Pinot · Elasticsearch vs Postgres Full-Text · Security Basics
Reading Guide#
Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal View) and the appendices on agent buffering, storage layout and redaction |
What is a Log Aggregation Platform? — Why interviewers pick this topic
A log aggregation platform collects the text that thousands of services write — request logs, errors, debug lines, audit records — from every host and container, moves it to a central place, makes it searchable within seconds, and keeps it for as long as someone is willing to pay for. Engineers use it at 3 a.m. to find the one stack trace that explains an outage; security uses it to reconstruct what an attacker touched; compliance uses it to prove who accessed what.
The hard part is not shipping lines over the network. The hard part is that log volume is unbounded, bursty and correlated with failure: an incident multiplies log output by 5–20× exactly when engineers need search to be fastest. Volume grows faster than headcount, cost scales with bytes nobody reads, one team's debug flag can drown everyone else, and some of those bytes contain passwords, tokens and personal data that must never be stored at all.
Before vs After — the "incident log storm" scenario:
Without per-tenant quotas and tiered storage:
t=0: Payments service starts failing; retries log a 4KB stack trace per attempt.
t=+1min: Payments log volume: 40 MB/s → 1.8 GB/s. Total cluster ingest: 2 GB/s → 3.8 GB/s.
t=+3min: Indexing nodes saturate. Ingest lag for ALL services: 5s → 9 min.
t=+5min: On-call searches for the payments error: results end 9 minutes ago.
t=+12min: Index cluster rejects writes; agents buffer locally; some hosts' disks hit 95%.
t=+20min: Incident resolved without logs. Post-mortem notes "logging was down".
With per-tenant quotas, agent buffering and a durable queue:
t=0: Same failure. Payments volume spikes to 1.8 GB/s.
t=+10s: Payments exceeds its burst quota. Pipeline samples repeated stack traces
(first 100/min per signature kept in full, rest counted) and spills to object storage.
t=+1min: Other services' ingest lag: unchanged at ~5s.
t=+2min: On-call searches payments errors: sees the signature, the count, and full examples.
t=+3h: Spilled raw logs queryable from object storage for the post-mortem, at lower speed.
Why interviewers reach for this question: It looks like "agents → Kafka → Elasticsearch → Kibana" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: backpressure that decides whose logs are lost, the index-everything vs query-on-read cost decision, retention tiers priced per gigabyte, multi-tenant quotas that stop one team's storm from blinding everyone, and redaction that keeps secrets out of storage you can't easily delete from.
Mechanics Refresher: Collection and Storage Options
| Option | How It Works | Pros | Cons |
|---|---|---|---|
| Node agent (daemon per host) | Tails files or container stdout, batches, ships | Apps stay simple; one agent per node | Agent CPU/memory competes with the app; must handle rotation |
| In-process library shipping | App sends logs over the network directly | No file I/O | Backpressure lands inside the app; crashes lose buffered logs |
| Durable queue (log broker) | Agents write to a partitioned, replicated log | Absorbs bursts; decouples ingest from indexing; replay | Another system; retention sizing |
| Full-text inverted index | Every token indexed; search by any word | Fast needle-in-haystack queries | Index often comparable to or larger than raw data; expensive at scale |
| Label index + compressed chunks in object storage | Index only a few labels; scan chunks at query time | Very cheap storage; simple ingest | Queries scan bytes; slow for broad searches |
| Columnar store | Fields as columns, compressed; skip indexes | Fast aggregations; good compression | Schema handling; free-text search weaker than inverted index |
| Retention tiers | Hot (fast, expensive) → warm → archive | Pay for speed only where it's used | Tier transitions; queries spanning tiers |
| Redaction | Pattern/field-based masking before storage | Secrets never land on disk | Patterns miss things; CPU cost |
For most production systems: a node agent with a bounded disk buffer, a durable queue absorbing bursts, a processing tier that parses, redacts and enforces per-tenant quotas, a short hot tier that is indexed for fast search, and a cheaper object-storage tier queried on read for older or lower-value data — with retention and cost attributed per team. The storage engine is not the interview — whose logs drop under pressure, what you index, and who pays per gigabyte are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
Log aggregation is not a search-engine question. Everyone can put Elasticsearch behind Kafka.
It is a cost and loss-policy question that tests:
- Whether you decide, in advance, whose logs are dropped, sampled or delayed when volume exceeds capacity
- Whether you match storage to query patterns — index the little that's searched often, scan the rest
- Whether you price retention per gigabyte per tier and attribute it to the teams producing the bytes
- Whether you keep secrets and personal data out of storage that is expensive to rewrite
The key insight: Log volume is driven by the people writing logs, not the people paying for them, and it spikes during incidents — exactly when logs are most valuable. A Senior design scales the cluster. A Staff design sets the policies that bound it: quotas per tenant, a loss order under backpressure, tiered retention with a price tag, and redaction at the edge. The platform's reliability is measured in whether the right logs are searchable during the worst hour, not in average ingest lag.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Agents → Kafka → Elasticsearch → dashboard | Asks "What are the logs for — debugging, security audit or analytics? How much per day, how long kept, and who pays?" | Asks "What's the logging budget as a share of infra spend, and which teams' habits drive it?" |
| Backpressure | "Kafka buffers it" | "Bounded disk buffer on each host, never block the app; durable queue with 24–72h retention; under overload, sample by signature and spill to object storage — a documented loss order" | Defines the org-wide loss policy: audit logs never dropped, debug logs first, with sign-off from security |
| Storage | "Index everything in Elasticsearch" | "Hot tier indexed for 3–7 days; object storage with label or columnar layout for 30–90 days; archive for compliance. Index what's searched, scan what's kept" | Prices each tier and makes retention a per-team budget decision with chargeback |
| Multi-tenancy | One shared cluster | "Per-tenant ingest quotas with burst allowance, per-tenant query limits, cardinality limits on labels or fields" | Turns quotas into a budget process: teams buy capacity; overages are visible to their leadership |
| PII | "Developers shouldn't log PII" | "Redact at the agent or pipeline with tested patterns; field-level allowlists for structured logs; deletion path designed for object storage" | Makes logging data classification part of the security standard; audits redaction coverage |
| Cost | "Add nodes" | "Cost per GB ingested and per GB-month retained, per tier, per team; volume reduction at the source" | Sets a cost-per-request target for observability and holds teams to it |
Why "backpressure" separates levels
L5: "Agents send to Kafka, which buffers when Elasticsearch is slow." That handles a brief slowdown. During a sustained log storm, Kafka fills to its retention, agents back up, and the candidate hasn't said whether the agent then blocks the application's write call (taking down production to save logs), fills the host's disk (same outcome, slower), or drops silently (losing the evidence for the incident).
L6: "There's a loss order, and it's written down. The application never blocks on logging: the agent has a bounded memory buffer and a bounded disk buffer — say 1 GB or two hours. When both are full, it drops oldest-first and emits a dropped-lines counter. Upstream, the queue holds 24–72 hours. If a tenant exceeds its quota, we sample repeated messages by signature — keep the first N per minute in full, count the rest — and spill the overflow to object storage instead of the index. Audit-class logs bypass sampling and have reserved capacity."
L7: "The loss order is a policy decision, not an implementation detail. Security signs off that audit logs are never sampled; product teams agree debug logs are sampled first. I'd publish it, so when a team loses debug logs during their own storm, the answer is the policy they agreed to, not an outage of the logging platform."
Why "storage" separates levels
L5: "Elasticsearch, with daily indices and 30-day retention." Full-text indexing every line costs CPU at ingest and storage often comparable to the raw size, replicated. At 170 TB/day that's petabytes of SSD for 30 days, most of it never queried after the first 48 hours.
L6: "Most log queries hit the last few hours, filter by service and level, and then look for a string. So: a hot tier of 3–7 days, indexed for fast search. Beyond that, compressed chunks in object storage — at roughly a tenth of the bytes and a fraction of the price per GB — queried on read, either by labels with a scan or by a columnar engine with skip indexes. Queries over old data are slower; that's the deal teams accept for keeping 90 days instead of 7."
L7: "Each tier has a price per GB-month and a query latency. I'd let teams choose retention per log stream within a budget, and show them what they're paying for logs nobody has queried in 30 days. The biggest saving is usually not storage engineering — it's deleting debug logs at source."
Why "PII" separates levels
L5: "We'll tell developers not to log sensitive data." They will anyway — a request object dumped on error, a header with a bearer token, an email address in a URL. Once it's in an immutable, compressed chunk in object storage, deleting it means rewriting files.
L6: "Redaction runs before anything is persisted beyond the host: known secret patterns — tokens, keys, card numbers with checksum validation — and field-level allowlists for structured logs. Redaction coverage is tested with seeded canaries. And because something will slip through, the storage layout supports deletion: chunks are addressable by tenant and time so a purge rewrites a bounded set of files."
L7: "Logs are a data store with the same classification rules as databases. I'd put logging into the data classification standard, give security an audit of what fields each service emits, and measure redaction coverage as a compliance metric."
Positions to Commit To#
| Position | Rationale |
|---|---|
| Never block the application on logging; bounded agent buffers with counted drops | Taking down production to preserve logs inverts priorities |
| A durable queue between collection and storage, 24–72h retention | Absorbs incident bursts and indexing outages; enables replay |
| A written loss order: debug sampled first, audit never | Overload will happen; the decision must be made before it does |
| Index only the hot window; query older data on read from object storage | Most queries hit the last hours; index cost is wasted on cold data |
| Per-tenant ingest quotas, query limits and cardinality limits | One team's storm or label explosion must not blind everyone |
| Redact at the edge; design storage for deletion anyway | Secrets in logs are inevitable; immutable storage makes them expensive |
| Cost per GB ingested and retained, attributed per team | Volume is driven by producers; they must see the bill |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Operational debugging logs | High volume, bursty, mostly queried within hours; loss of some debug lines tolerable | Agents + queue + hot index + object-storage tier; quotas and sampling under pressure | Incident storms blind search; cost explodes | Ingest-to-searchable p99 < 30s for in-quota tenants; < 0.1% unplanned loss |
| Security and audit logs | Completeness and integrity matter more than latency; retained for years | Separate lane with reserved capacity, no sampling, write-once storage, tamper evidence | Silent gaps; logs altered or deleted | Zero sampled or dropped records; verifiable gap detection |
| Analytics from logs (counts, rates, funnels) | Aggregations over long ranges; schema matters | Emit metrics or structured events instead; columnar store if logs must be used | Using text search for aggregations; cardinality blowups | Should usually be a metrics or events pipeline, not logs |
🎯 Staff Move: "I'll design the operational debugging platform — high-volume service logs searched mostly during incidents. Audit logs get a separate lane with reserved capacity, no sampling and write-once storage; I'll show where they branch. If the real need is counting things, I'd push teams to emit metrics — counting error lines in a log store is the expensive way to build a dashboard."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Index Everything vs Query on Read | Fast arbitrary search at high ingest and storage cost, or cheap storage with slower scan-based queries? |
| 2 | Backpressure: Block, Buffer or Drop | Whose logs are lost, delayed or sampled when volume exceeds capacity — and who decides? |
| 3 | Schema on Write vs Schema on Read | Parse and type fields at ingest (fast queries, brittle pipelines) or store raw and parse at query time (flexible, slow)? |
| 4 | Retention Tiers and Who Pays | One retention for all, or per-stream tiers priced and charged back to the producing team? |
| 5 | Where PII Redaction Happens | At the source, in the pipeline or at query time — and how you delete what slipped through |
How Real Companies Built It#
Why this section belongs here: Three well-documented systems made visibly different storage choices for logs. Naming them shows you see indexing as a cost decision, not a default.
Grafana Loki — Index Only Labels, Store Compressed Chunks in Object Storage#
Loki's documentation states that it does not index the contents of logs, only metadata as a set of labels for each log stream; log data is compressed and stored in chunks in an object store such as Amazon S3 or Google Cloud Storage, which keeps the index much smaller than other log tools (Loki overview). Its cardinality guidance warns that high-cardinality labels create many streams, a huge index and thousands of tiny chunks, and notes a default limit of 15 index labels (Loki cardinality).
Staff insight: This is the "query on read" end of the spectrum taken seriously: cheap ingest and storage, paid for with scan-heavy queries and strict label discipline. In an interview, say: "If I index only labels, cardinality limits become a tenant policy — one team putting request IDs in a label breaks the index for everyone."
Uber — From Elasticsearch to a Columnar Store for Schema-Agnostic Logs#
Uber's engineering blog describes moving its logging platform from the ELK stack to ClickHouse. With Elasticsearch, incompatible field types caused type-conflict errors that dropped logs, and the team ran 20+ clusters per region to limit the blast radius of heavy queries and mapping explosions. On ClickHouse, a single node ingested about 300K logs per second — roughly ten times a single Elasticsearch node in their setup — and the platform's hardware cost dropped by more than half while serving more traffic, with Kafka buffering ingestion (Uber blog).
Staff insight: Schema conflicts and mapping explosions are not edge cases at scale; they are the main failure mode of index-everything with dynamic fields. Say: "Dynamic field mapping is a multi-tenant hazard. Either I cap fields per tenant, or I store logs in a layout where a new field can't break someone else's index."
Datadog Husky — Stateless Writers, Object Storage and Separate Query Compute#
Datadog's engineering blog describes Husky, its third-generation event store for logs and other events, as an unbundled, distributed, schemaless, vectorized column store. Writers ingest from Kafka, buffer briefly and upload to blob storage; readers query individual files; compactors merge small files; and durability is pushed to FoundationDB and S3. The post says this separation made it cost-effective to retain data that is rarely queried but must stay immediately queryable, and enabled a new product, Flex Logs (Datadog blog).
Staff insight: Separating ingest, storage and query lets each scale with its own driver — bytes in, bytes kept, queries run — and turns retention into a pricing choice. Say: "Storage and query compute should scale independently, because logs are written a thousand times more than they're read."
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Kafka buffers the spikes" | "The spike lasts two hours and exceeds Kafka's retention. Whose logs are lost?" | Loss policy, backpressure end-to-end |
| "We index everything" | "What does 30 days of that cost, and how much of it is ever queried?" | Cost awareness, tiering |
| "Teams send logs to the platform" | "One team enables debug logging in production at 10× volume. What happens to everyone else?" | Multi-tenant quotas |
| "We parse JSON logs" | "Two services send status as a string and an integer. What breaks?" | Schema conflicts, mapping explosion |
| "We redact PII" | "A password was logged for six weeks before anyone noticed. Delete it." | Deletion in immutable storage |
| "Agents tail log files" | "The host's disk is 95% full and the agent can't ship. What does the agent do?" | Never harm the app; bounded buffers |
| "We keep logs 90 days" | "Security needs a year, debugging needs a week. One tier or two?" | Retention tiers, audit lane |
System Architecture Overview#
Reading the diagram: Agents on every node tail container output, apply a first redaction pass, and batch to an ingest gateway that tags the tenant and checks quotas. A durable queue absorbs bursts. Processors parse and enrich by stream schema, redact again with field-level rules, and apply overload control — sampling repeated messages by signature when a tenant exceeds quota. Every log line lands compressed in object storage; in-quota logs are also indexed in a short hot tier. Audit streams take a separate lane into write-once archive. The query frontend routes by time range: recent queries hit the hot index, older ones scan object storage. The metric that tells you the platform is healthy is
ingest.lag_p99_sfor in-quota tenants — it should stay under 30 seconds no matter which team is storming.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Collection | "Agents ship to Kafka" | "Bounded memory and disk buffers; never block the app; count every drop." |
| Overload | "Kafka buffers it" | "Written loss order: sample debug by signature, spill to cold, never sample audit." |
| Storage | "Elasticsearch, 30 days" | "Index 7 days hot; object storage for 90 days, query on read; archive for audit." |
| Tenancy | "One shared cluster" | "Per-tenant ingest quotas with burst, query limits, cardinality limits." |
| Schema | "Parse everything" | "Per-stream schemas; unknown fields kept as raw; type conflicts can't drop logs." |
| PII | "Don't log PII" | "Redact at agent and pipeline, test with canaries, design storage for deletion." |
| Cost | "Add nodes" | "Cost per GB ingested and retained per team; reduce at the source." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| Ingest volume (example scale) | ~2 GB/s average, 5–10× in incidents (illustrative) | ~170 TB/day raw; sizes every tier |
| Average log line | ~200–1,000 bytes (illustrative) | 2 GB/s ≈ 2–10M lines/s |
| Compression for logs | commonly ~5–15× (illustrative, data-dependent) | 170 TB/day raw → ~15–30 TB/day stored |
| Full-text index overhead | often comparable to raw size before replication (illustrative) | Why indexing 90 days is expensive |
| Object storage price | ~$0.02/GB-month order of magnitude (check current provider pricing) | Cheapest durable tier for warm data |
| Hot tier window | ~3–7 days typical | Most queries hit the last hours |
| Queue retention | 24–72 hours | Covers an indexing outage plus a weekend |
| Agent buffer | memory ~64–256 MB, disk ~1–5 GB (illustrative) | Bounded harm to the host |
| Fluent Bit buffer limits | mem_buf_limit pauses that input; storage.total_limit_size discards the oldest chunk (docs) | A real agent's bounded-buffer loss order |
| Agent overhead target | < ~2–5% of a node's CPU (illustrative) | Logging must not tax the product |
| Ingest-to-searchable | p99 < 30 s target for in-quota tenants | The incident-time promise |
| Elasticsearch default field limit | index.mapping.total_fields.limit = 1,000 | Mapping explosion guardrail |
| Loki default index labels | 15 per series | Cardinality guardrail |
| Uber single-node ingest (ClickHouse) | ~300K logs/s, ~10× their ES nodes | Columnar ingest efficiency at scale |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on Elasticsearch shard counts, index templates and Kibana dashboards, then run out of time before the interviewer asks the questions that decide the level: "One team turns on debug logging and triples volume. What happens to everyone else?" and "What does 90 days cost?" Compress the pipeline to ~8 minutes and spend the rest on backpressure and loss policy, storage tiers, multi-tenancy, schema and redaction.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"Every service on every host writes logs. We collect them, make them searchable within seconds by service, time, level and text, support live tail during incidents, keep them for a defined retention per stream, and give security an audit lane that's complete and tamper-evident."
Then the non-functional requirements, which is where the design lives:
"Four constraints drive everything. One: volume is bursty and correlated with failure — incidents multiply output 5–20× when search matters most. Two: logging must never take down the application. Three: cost scales with bytes, and producers don't see the bill unless we show it. Four: logs contain secrets and personal data whether we like it or not. I'll assume 3,000 services on 50,000 nodes, 2 GB/s average ingest, 10 GB/s during a bad incident, 7 days of fast search and 90 days of retained, slower search."
Then name the underspecified parts:
"I'd confirm: is this debugging, audit, or both? Who pays? What's the acceptable loss under overload? I'll assume debugging is primary with a separate audit lane, costs charged back per team, and debug-level logs sampled first under pressure."
🎯 Staff Move: Saying "incidents multiply log volume exactly when search matters most" reframes the problem from "store logs" to "stay useful during the worst hour" — that's the sentence that sets the level.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Tenant: a team or service owner —
tenant_id,ingest_quota_bytes_s,burst_bytes,retention_policy,cost_center - Stream: a labeled sequence —
{tenant, service, env, cluster, level}; labels low-cardinality by policy - Log record:
ts,stream,body(raw),fields(parsed, optional),trace_id?,ingest_ts - Chunk: compressed block of records for one stream and time range —
chunk_id,tenant,t_min,t_max,bytes,tier - Query:
tenant,time_range,label_selector,filter(text/field),limit
Ingest API (agent → gateway):
POST /v1/ingest Content-Encoding: gzip
Authorization: tenant token
{ streams: [ { labels: {service, env, level}, entries: [ [ts, line], … ] } ] }
→ 204 accepted
→ 429 quota exceeded, Retry-After: 2 (agent keeps buffering, backs off)
→ 413 batch too large
Query API:
GET /v1/query?selector={service="checkout",level="error"}&q="timeout"&from=-1h&limit=500
GET /v1/tail?selector={service="checkout"} (live tail, server-sent stream)
POST /v1/retention { stream_selector, hot_days, warm_days }
🎯 Staff Move: "429 with Retry-After is the backpressure contract between the platform and agents. The agent never drops on a 429; it buffers within its bounds and retries. Drops only happen when the agent's own bounded buffer is full — and each one is counted."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk one log line in 90 seconds:
- Checkout writes a JSON line to stdout. The node agent tails the container log, applies pattern redaction for tokens and card numbers, and adds it to a batch.
- Every second or 1 MB, the agent gzips the batch and posts it to the gateway, which authenticates the tenant, checks the token-bucket quota and writes to the queue partitioned by tenant and stream.
- A processor parses the line against checkout's stream schema, attaches Kubernetes metadata, runs field-level redaction, and checks the tenant's rate against quota. In quota: indexed into the hot tier and appended to a compressed chunk for object storage. Over quota: sampled by message signature, with full lines still written to object storage.
- Within ~10 seconds the line is searchable in the hot tier; within ~5 minutes its chunk is flushed to object storage.
- A query for the last hour goes to the hot index; a query for last month goes to object storage, filtered by labels and time, scanned in parallel.
🎯 Staff Move: Say out loud: "Every line lands in object storage, which is cheap and durable. The hot index is an accelerator for the recent window, not the system of record." You've now spent ~8 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes this hard is that volume is unbounded and spikes during incidents, cost scales with bytes nobody reads, and one tenant can hurt all the others. I'd like to go deep on backpressure and the loss order, index vs query-on-read and retention tiers, multi-tenant quotas and cardinality, schema handling, and PII redaction. Where would you like to start?"
If no preference: start with backpressure. It's the question that decides whether logs exist during the incident.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: Backpressure and the loss order (7–8 min)
"Four buffers, each bounded. The app's write to stdout never blocks on the network — the container runtime writes to a local file. The agent reads files with a 64 MB memory buffer that spills to a 1 GB disk buffer; when that's full, it stops reading new data and, if the file rotates away, those lines are lost and counted in agent.dropped_lines_total. The gateway returns 429 when a tenant exceeds quota; agents back off. The queue holds 48 hours, so an indexing outage of a day loses nothing. Overload control in the processor samples by signature for over-quota tenants — first 100 per signature per minute in full, the rest counted — but still writes everything compressed to object storage if the storage write path keeps up."
Quantify: "A 1 GB disk buffer at a typical node's 200 KB/s of logs is ~80 minutes of outage tolerance per host. 48 hours of queue at 2 GB/s is ~350 TB raw, ~35–70 TB compressed in the queue — real money, sized deliberately."
Who pays: "The storming tenant loses fidelity on its own repeated messages. Every other tenant pays nothing. Audit streams have reserved capacity and are never sampled."
Deep dive 2: Index vs query on read, and retention tiers (6–7 min)
"Query patterns decide this. Most searches are in the last 24 hours, filtered by service and level, then a text match. So the hot tier — 7 days — is indexed for interactive search in seconds. Everything is also in object storage as compressed, columnar chunks partitioned by tenant, stream and hour, with small per-chunk metadata — min/max time, labels, a bloom filter of tokens. Older queries prune by labels and time, skip chunks via bloom filters, and scan the rest in parallel. A 30-day query over one service is tens of seconds instead of sub-second; that's the explicit trade."
Quantify: "170 TB/day raw. Hot: 7 days indexed and replicated is roughly 170 × 7 × 2 × ~1 ≈ 2.4 PB of SSD-class storage if we indexed everything — which is why the hot tier holds only in-quota, non-sampled data and drops debug after 24 hours. Warm: 90 days × ~17 TB/day compressed ≈ 1.5 PB of object storage at ~$0.02/GB-month ≈ $30K/month."
Deep dive 3: Multi-tenant quotas and cardinality (5–6 min)
"Each tenant has an ingest quota in bytes per second with a burst bucket — say 2× quota for 10 minutes — enforced at the gateway and again in processing. Query limits per tenant: concurrent queries, bytes scanned per query, and a max time range for hot-tier queries. Cardinality limits: max labels per stream (around 10–15), max distinct values per label per day, max parsed fields per stream. Violations are rejected or folded into the raw body, never allowed to grow the shared index without bound."
Deep dive 4: Schema on write vs on read (4–5 min)
"Structured JSON is encouraged and parsed at ingest against a per-stream schema that the owning team registers. Known fields are typed and indexed in the hot tier. Unknown fields stay in the raw body and are searchable as text, never auto-added to a shared mapping. Type conflicts — status as a string in one service and an integer in another — can't collide because fields are namespaced per stream, and a value that fails to parse is kept raw, not dropped."
Deep dive 5: PII redaction and ownership (3–4 min)
"Two passes. The agent runs cheap pattern redaction — bearer tokens, API key formats, card numbers validated with a checksum to avoid false positives. The processor applies per-stream field rules: allowlisted fields kept, known-sensitive fields hashed or dropped. Seeded canary secrets in synthetic traffic verify coverage daily. When something slips through, the purge job rewrites the affected chunks — bounded because chunks are partitioned by tenant and hour. The platform team owns ingest health and quotas; product teams own their volume, schemas and what they log; security owns redaction rules and the audit lane."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: log volume is driven by producers and spikes during incidents, so the platform is defined by its policies — a loss order under backpressure, an indexed hot window over cheap object storage, per-tenant quotas and cardinality limits, schemas that can't break each other, and redaction before persistence — with cost per gigabyte visible to the teams producing it."
The evolution closer:
"What I'd build later: log-to-metric extraction so teams stop querying logs for counts; trace-linked logs; per-team retention self-service with a budget; and federated queries across regions. What I'd not build: indexing every line for 90 days — I'd invest in reducing volume at the source instead."
🎯 Staff Move: End on who owns the bytes. "The platform owns keeping in-quota logs searchable within 30 seconds through any incident. Teams own their volume and see its cost. When a team's own storm gets sampled, that's the policy working, not the platform failing."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Cluster tuning | 10 min on shards, replicas, index templates | "Hot index is an accelerator; object storage is the record" |
| Kafka as the answer | "Kafka buffers it" ends the backpressure discussion | Four bounded buffers and a written loss order |
| One retention | 30 days for everything | Hot, warm and archive tiers with prices |
| No tenancy | One shared cluster, no quotas | Ingest quotas, query limits, cardinality limits |
| PII as a guideline | "Developers shouldn't log secrets" | Redaction in the pipeline, canaries, deletion path |
| No cost | Never states $ per GB | Cost per GB ingested and retained, per team |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Logging is the interview where the users who generate the load are not the users who pay for it, and where demand spikes at the exact moment the system is most needed. Every engineer adds log lines; nobody removes them. Volume grows with services, traffic and incidents, and the cost line grows with it until a finance review asks why observability costs more than the database fleet.
A Senior design scales storage to match demand. A Staff design recognizes that demand is a policy problem: what gets kept, at what fidelity, for how long, at whose expense — and what gets dropped first when the system is overwhelmed. The architecture follows from those policies, and the policies need owners outside the platform team.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"It's the worst outage of the year. Log volume is 6× normal, one service is producing 80% of it, and the on-call is searching for a specific error in a different service. Walk me through what your platform is doing right now, whose logs are being dropped or sampled, and what that on-call sees."
A candidate who answers with the storming tenant hitting its burst quota, signature sampling applied to that tenant only, full raw lines still landing in object storage, queue depth rising but within its 48-hour bound, other tenants' ingest lag unchanged at seconds, query limits stopping a dozen engineers' 30-day searches from starving the hot tier, and the on-call finding their error in ten seconds has run a logging platform. A candidate who says "Kafka absorbs the spike" has drawn one.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Operational debugging → useful during the worst hour, affordable the rest of the time
- Constraint: bursty volume, recency-heavy queries, tolerance for sampling repeated debug lines
- Strategy: bounded agents, durable queue, hot index over object storage, quotas, signature sampling
- Failure mode: incident storms blind search; cost grows without owners
- Who pays for imperfection: on-call engineers (blind incidents), finance (unbounded spend), the storming team (sampled fidelity)
Security and audit → complete, intact, retained for years
- Constraint: no sampling, gap detection, tamper evidence, legal retention
- Strategy: separate lane with reserved ingest capacity, write-once storage, sequence numbers per source, independent verification
- Failure mode: silent gaps; deletion by an attacker with access
- Who pays: security (blind investigations), the company (compliance findings)
Analytics from logs → probably not a logging problem
- Constraint: long-range aggregations, stable schemas
- Strategy: emit metrics or structured events; if logs are the only source, extract metrics at ingest
- Failure mode: dashboards that scan terabytes per refresh
- Who pays: the platform (query load), everyone else (slow searches during dashboard refreshes)
2.2 When NOT to Build a Log Aggregation Platform#
- You're small. Under a few hundred GB/day, a managed logging service or a single search cluster is cheaper than a team. The tiering and quota machinery earns its keep at tens of TB/day.
- You need counts, rates and alerts. Use the metrics platform; a metric is bytes per minute, a log line is bytes per event.
- You need request-level causality across services. That's distributed tracing; link logs to traces by trace ID rather than reconstructing flows from text.
- You need business events with a schema. Use an event pipeline into a warehouse; logs are the wrong contract for data other teams depend on.
🎯 Staff Insight: "The cheapest log line is the one never written. Before I build a bigger platform, I'd find the ten streams that make up half the volume and ask whether anyone reads them."
2.3 What the Interviewer Leaves Underspecified#
Interviewers deliberately omit:
- Debugging vs audit — the completeness bar differs completely
- Who pays — shared cost hides the producers driving it
- Acceptable loss — "never lose a log" is unaffordable during a storm; what's the order?
- Retention — one number for everything vs per-stream tiers
- Structured vs free text — determines parsing cost and schema risk
- PII policy — whether logs may contain personal data at all, and how deletion requests apply
Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.
2.4 Precise Terminology#
| Term | What It Means | Why It Matters in the Interview |
|---|---|---|
| Stream | Records sharing one label set | Unit of indexing and cardinality |
| Cardinality | Distinct label-value combinations | Drives index size; the multi-tenant hazard |
| Chunk | Compressed block of one stream's records over a time range | Unit of storage, pruning and deletion |
| Hot tier | Recent data, indexed, on fast storage | Expensive; sized by query recency |
| Query on read | Scan compressed data at query time, no full index | Cheap storage, slower broad queries |
| Mapping explosion | Unbounded growth of indexed field definitions | Breaks shared indices |
| Backpressure | Slowing producers when consumers can't keep up | Must stop before it reaches the app |
| Signature sampling | Keep N examples per message template, count the rest | Preserves information under overload |
| Ingest lag | Time from write to searchable | The incident-time SLO |
| Chargeback | Attributing cost to the producing team | Makes volume visible to its cause |
| Redaction | Masking secrets/PII before persistence | Prevents expensive purges |
🎯 Staff Insight: If the interviewer says "we can't lose any logs", ask: "Which logs? Audit logs, yes — reserved capacity, no sampling. Debug logs during a 10× storm from one service — I'd rather sample its repeated stack traces than blind every other team. Can we agree a loss order?"
3. Where the Design Splits#
Every logging decision has a technical side (index structures, buffer sizes, compression) and an organizational side (whose logs drop, who pays, who decides retention). Interviewers grade the second side.
3.1 Fault Line 1: Index Everything vs Query on Read#
The tension: A full-text index answers any search in sub-second time but costs ingest CPU and storage comparable to the raw data, replicated, on fast disks. Query-on-read stores compressed chunks cheaply and pays at query time by scanning.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Full-text index, all retention | Fast arbitrary search over months | Cost scales with every byte for the whole retention; mapping explosions | Finance (storage), platform (cluster ops) |
| Labels only + object-storage chunks | Very cheap ingest and storage | Broad text queries scan lots of bytes; label cardinality must be policed | Engineers (slower searches) |
| Columnar store with skip indexes | Strong compression; fast filters and aggregations | Free-text search weaker; schema handling needed | Platform (schema tooling) |
| Hot index + object-storage warm tier | Fast where queries are; cheap where data sits | Two query paths; queries spanning both | Platform (routing complexity) |
Staff default: "Hot index for 7 days of in-quota data, with debug-level streams kept hot for only 24 hours. Everything written compressed to object storage, partitioned by tenant, stream and hour, with per-chunk metadata and token bloom filters. Queries route by time range; broad queries over cold data require a label selector so they can't scan the whole tenant."
When to deviate:
- Security investigations that need arbitrary search over months: index the audit lane for longer — it's a fraction of the volume.
- Small scale: a single indexed cluster with 14–30 days is simpler than two tiers.
🧭 Principal Move: "I'd publish the price of each tier per GB-month and the query latency teams should expect from it. Then retention becomes a choice teams make with their own budget, not a platform default everyone pays for."
❌ Common L5 Trap: "Elasticsearch with 90-day retention and daily indices." At 170 TB/day, that's on the order of tens of petabytes of replicated, indexed storage, most of it never queried after day two. The design isn't wrong — it's unaffordable, and the candidate didn't price it.
3.2 Fault Line 2: Backpressure — Block, Buffer or Drop#
The tension: When ingest exceeds capacity, logs must wait somewhere or be discarded. Waiting in the application blocks production; waiting on the host fills disks; waiting in the queue costs money; dropping loses evidence. The question is whose logs, where, in what order.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Block the app's log call | No loss | Production latency or outage caused by logging | Users (outage) |
| Unbounded agent buffer | No loss short-term | Host disk or memory exhaustion; agent OOM | The app on that host |
| Bounded agent buffer, drop oldest, count | Host protected; drops visible | Loss during long outages | The tenant (bounded loss) |
| Durable queue with long retention | Absorbs indexing outages and bursts | Storage cost; retention limit still exists | Platform (queue cost) |
| Signature sampling for over-quota tenants | Keeps information, sheds volume | Exact counts approximated; some full lines gone from hot tier | Storming tenant (fidelity) |
Staff default: "The application writes to local stdout files and never blocks on the network. The agent has bounded memory and disk buffers, honors 429s, and drops oldest when full — counting every drop per tenant. The queue retains 48 hours. Over-quota tenants are sampled by message signature in the hot path, while raw lines still go to object storage when its write path has capacity. Audit streams have reserved capacity at every stage and are never sampled."
When to deviate:
- Audit and security logs: block-or-fail is sometimes legally required — use a separate, synchronous, low-volume path for those specific events, not for all logs.
- Edge and mobile: client buffers are tiny and lossy; accept it and sample at the source.
🧭 Principal Move: "The loss order is an org policy: audit never, error-level last, debug first, and the storming tenant before anyone else. Security and engineering leadership sign it; the platform enforces it. Then a lost debug line during a team's own storm is policy, not an incident."
❌ Common L5 Trap: "Agents retry until the backend accepts, so nothing is lost." Unbounded retry means unbounded buffering — the agent fills the host's disk or memory and takes the application down with it. See Backpressure & Load Shedding.
3.3 Fault Line 3: Schema on Write vs Schema on Read#
The tension: Parsing at ingest gives typed fields, fast filters and compact columns — and a pipeline that breaks when formats change or two teams disagree on a field's type. Storing raw and parsing at query time is flexible and robust — and slow, with every query re-parsing.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Schema on write, shared dynamic mapping | Fast field queries; zero setup | Type conflicts drop or reject logs; mapping explosion | Every tenant (shared breakage) |
| Schema on write, per-stream registered schemas | Typed, isolated per team | Teams must register schemas; drift handling | Product teams (registration) |
| Schema on read (raw text) | Nothing breaks at ingest | Every query parses; aggregations slow | Engineers (query latency) |
| Hybrid: known fields typed, unknown kept raw | Robust and fast for common fields | Two representations to explain | Platform (tooling) |
Staff default: "Hybrid with per-stream namespaces. Teams register the fields they want typed; those are parsed at ingest and indexed in the hot tier. Everything else stays in the raw body, searchable as text. A field that fails to parse is kept raw with a parse-error tag — never dropped. No shared dynamic mapping, so one team's new field can't push a shared index past its field limit."
When to deviate:
- Highly regular logs (load balancer access logs): full schema on write, columnar, typed — they're effectively events.
- Third-party software logs with unstable formats: raw only, with query-time parsing helpers.
🎯 Staff Insight: "Elasticsearch's default limit is 1,000 mapped fields per index for a reason. In a shared index, the 1,001st field from any team breaks ingest for everyone. Per-stream schemas turn a shared failure into a local one."
❌ Common L5 Trap: "We parse JSON and index all fields automatically." It works until one service logs a dynamic map — user IDs as keys — and mints thousands of fields an hour, or two services disagree on whether
statusis a string. Logs get rejected for tenants who did nothing wrong.
3.4 Fault Line 4: Retention Tiers and Who Pays#
The tension: Retention is the largest cost multiplier: every extra day multiplies every byte. Teams want long retention by default because it's free to them; the platform wants short retention because it pays. Neither has the right incentive.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| One retention for all (e.g., 30 days indexed) | Simple | Pays for long retention on chatty debug logs | Platform budget |
| Tiered, platform-chosen | Cheaper | Teams with real needs (security, compliance) fight exceptions | Platform (exception handling) |
| Tiered, team-chosen within a budget + chargeback | Cost tracks value; teams decide | Budget process; teams must understand tiers | Product teams (their own budget) |
| Keep everything forever in archive | Never miss old data | Archive grows without bound; deletion obligations | Finance, legal (data liability) |
Staff default: "Three tiers. Hot, indexed: 7 days for info and above, 24 hours for debug. Warm, object storage: 30 days default, up to 90 per stream if the team pays. Archive: audit streams only, write-once, retention set by security and legal. Every stream's bytes ingested and GB-months retained are reported per team monthly, with the top streams by cost and their query counts — 'you paid $14K for this stream; it was queried twice.'"
When to deviate:
- Regulated environments: retention minimums set by policy override team choice.
- Early-stage platforms: one tier, short retention; tiering arrives with scale.
🧭 Principal Move: "Chargeback changes behavior more than any compression algorithm. The first month teams see their logging bill next to their query counts, volume drops — I'd plan for that and measure it."
❌ Common L5 Trap: "Keep 30 days for everything; storage is cheap." Object storage is cheap; indexed, replicated SSD storage for 30 days of 170 TB/day is not. And "cheap" multiplied by a volume that doubles yearly is a line item finance will notice.
3.5 Fault Line 5: Where PII Redaction Happens#
The tension: Redacting at the source is most effective but depends on every team's discipline. Redacting in the pipeline is centralized but sees only what patterns can recognize. Redacting at query time leaves secrets in storage. And once data is in compressed, immutable chunks, deletion means rewriting files.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Source only (logging library rules) | Never leaves the process | Depends on every team; misses ad-hoc dumps | Security (gaps) |
| Agent and pipeline patterns | Central, consistent, no team effort | False negatives on unknown formats; false positives on IDs | Platform (CPU), engineers (over-masking) |
| Query-time masking | Easy to change rules | Secrets stored and replicated; breach exposure | Company (liability) |
| Structured field allowlists | Precise for structured logs | Only works where logs are structured | Product teams (schema discipline) |
Staff default: "Defense in depth. Logging libraries mask known sensitive fields by type. Agents run fast pattern redaction for credentials and card numbers before anything leaves the host. Processors apply per-stream field allowlists. Daily canaries inject fake secrets through real services and alert if any reach storage unmasked. Storage is partitioned by tenant and hour so a purge rewrites a bounded set of chunks, and the hot index supports delete-by-query for the 7-day window."
When to deviate:
- Debugging that requires sensitive values (e.g., payment flows): log a reference or a keyed hash, not the value; give a separate, access-controlled store for the rare cases that truly need raw data.
- Regions with strict data protection: keep logs in-region and apply stricter allowlists.
🧭 Principal Move: "Logs are a database the whole company writes to without review. I'd put logging under the data classification standard, give security a per-service report of emitted fields, and track redaction canary pass rate as a compliance metric."
❌ Common L5 Trap: "We mask PII in the search UI." The secret is still in the index, the replicas, the object store, and any export. A breach or an over-privileged query path exposes it, and a deletion request means finding it everywhere.
4. When It Breaks#
4.1 The Debug Flag — One Team Triples Platform Volume#
t=0: A team enables DEBUG in production to chase a bug. Their volume: 60 MB/s → 2.4 GB/s.
t=+1min: Total ingest 2 GB/s → 4.3 GB/s. No per-tenant quotas enforced in processing.
t=+4min: Hot-tier indexers at 100% CPU. ingest.lag_p99_s for all tenants: 8s → 6 min.
t=+10min: Queue depth rising 2 TB/min. Live tail lags. Two unrelated incidents begin.
t=+25min: On-call finds the stream by bytes; asks the team to revert. Lag drains over 40 min.
Detection: tenant.ingest_bytes_s top-N vs quota; ingest.lag_p99_s{in_quota}; indexer.cpu_util.
Mitigation: enforce quota in processing with signature sampling for the over-quota tenant; route its excess directly to object storage, skipping the hot index.
Prevention: quotas enforced by default at gateway and processor; debug level kept hot for 24 hours only; a "debug window" feature that lets a team raise its quota for 30 minutes on a named set of pods, with cost shown up front.
Owner: logging platform on-call; the producing team owns its volume.
4.2 The Full Disk — An Agent Takes Down Its Host#
t=0: Ingest gateway in one zone rejects connections (cert expired).
t=+5min: Agents in that zone buffer to disk. Buffer limit misconfigured as "unlimited" on
an older agent version deployed to 4,000 nodes.
t=+3h: Node disks at 95%. Kubelet starts evicting pods for disk pressure.
t=+3h10m: Production services in the zone lose capacity. A logging outage became a
production outage.
Detection: agent.buffer_bytes{node}, agent.buffer_pct_of_limit, node disk pressure events.
Mitigation: cap agent disk buffers at a fixed size (e.g., 1 GB) and a fraction of free disk; on limit, drop oldest and count; rotate the certificate.
Prevention: agent config enforced by a fleet policy, not per-deployment; certificate expiry alerts 30 days ahead; a chaos test that blackholes the gateway for one zone monthly.
Owner: logging platform (agent fleet and gateway).
4.3 The Mapping Explosion — One Service Breaks the Shared Index#
t=0: A service starts logging a map keyed by customer ID as a JSON object.
t=+20min: Shared hot index gains 9,000 new fields; hits its field limit.
t=+21min: Every write with a new field to that index is rejected; with dynamic mapping,
12 other services sharing the index start losing logs.
t=+1h: Teams notice "missing logs"; platform finds the culprit by field growth rate.
Detection: index.field_count growth rate; ingest.rejected_total{reason=mapping}; per-stream new-field rate.
Mitigation: move the stream to its own index or fold unknown fields into the raw body; blocklist the offending field path.
Prevention: no shared dynamic mapping; per-stream registered schemas; unknown fields stored raw; a per-stream limit on parsed fields.
Owner: logging platform (schema policy); producing team (log format).
4.4 Secrets in Logs — Six Weeks of Bearer Tokens#
t=0: A new HTTP client library logs full request headers at INFO on retry.
t=+6 wks: Security finds Authorization headers in a log search result.
t=+6 wks: Tokens for ~2M sessions sit in the hot index, its replicas, 42 days of object
storage chunks and two analytics exports.
Detection: redaction canaries (a fake token sent through real services daily; alert if found unmasked); periodic scans for credential patterns in stored data.
Mitigation: revoke affected tokens; delete-by-query in the hot tier; rewrite affected chunks in object storage — bounded because chunks are partitioned by tenant and hour; purge exports.
Prevention: agent-level pattern redaction for authorization headers; logging library denies header logging by default; canaries per language runtime.
Owner: security (rules, response), logging platform (purge tooling), producing team (library usage).
4.5 Silent Loss — A Parse Error Path That Drops Lines#
t=0: A processor deploy changes JSON parsing; lines with invalid UTF-8 now throw.
t=+0: The exception handler drops the line and increments a counter nobody alerts on.
t=+5 days: An incident: the on-call can't find errors from one service. 3% of its lines
(those with binary payload fragments) have been dropped for five days.
Detection: end-to-end counts: lines read by agents vs lines stored, per stream, reconciled hourly; processor.dropped_total{reason} with alerts.
Mitigation: fix parser to keep unparseable lines raw with a tag; replay the queue for the last 48 hours (the rest is lost).
Prevention: invariant: the processor never drops a line without a counted, alerted reason; canary streams with malformed content in every deploy's test suite.
Owner: logging platform.
4.6 The Query of Doom — A 90-Day Unfiltered Search#
An engineer runs a regex across all services for 90 days during an incident. The query fans out to hundreds of thousands of chunks and saturates the query tier; every other engineer's searches queue behind it during the incident. Mitigation: per-tenant and per-user limits on bytes scanned and concurrent queries; require a label selector for queries beyond the hot window; queue long queries as async jobs with results delivered later.
Detection: query.bytes_scanned{user}, query.queue_wait_p99.
Owner: logging platform (query limits).
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Tenant volume storm | tenant.ingest_bytes_s over quota | All tenants without quotas | Quota + signature sampling + cold spill | Logging platform |
| Agent fills host disk | agent.buffer_pct_of_limit, disk pressure | Production on affected nodes | Hard buffer caps, drop oldest | Logging platform |
| Mapping explosion | index.field_count growth | Tenants sharing the index | Per-stream schemas, raw fallback | Logging platform |
| Secrets stored | Canary found unmasked | Security exposure | Revoke, purge chunks | Security + platform |
| Silent drops | Agent-read vs stored reconciliation | One stream's evidence | Never drop uncounted; replay | Logging platform |
| Query of doom | query.bytes_scanned, queue wait | All searchers | Scan limits, async jobs | Logging platform |
| Queue retention exceeded | queue.oldest_unconsumed_age > 36h | Loss of oldest data | Scale processors, shed debug | Logging platform |
| Object storage throttling | objstore.put_throttled_total | Warm-tier writes lag | Larger chunks, key prefix spreading | Logging platform |
🎯 Staff Insight: The platform's health metric is in-quota ingest lag, not overall lag. "Overall lag will spike whenever a team storms, and it should — they're over quota. I page on lag for tenants within quota, because that number only moves when the platform is the problem."
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | "Collect and search logs" | Names debugging vs audit vs analytics; frames it as staying useful during the worst hour at bounded cost | Frames observability spend as a budget with owners and a cost-per-request target |
| Backpressure | "Kafka buffers it" | Bounded buffers at every stage, never block the app, counted drops, written loss order | Loss order as an org policy signed by security and engineering leadership |
| Storage | Index everything for 30 days | Hot index over object storage, query on read, routing by time, scan limits | Prices tiers per GB-month; retention as a team budget decision |
| Multi-tenancy | Shared cluster | Ingest quotas with burst, query limits, cardinality and field limits, signature sampling | Quotas tied to budgets and chargeback; exceptions process |
| Schema & PII | Auto-parse JSON; "don't log PII" | Per-stream schemas, raw fallback; layered redaction with canaries; deletion-ready layout | Logging in the data classification standard; redaction coverage audited |
| Operations | Cluster dashboards | In-quota lag SLO, end-to-end reconciliation, agent fleet policy | Volume reduction at source as a measured program across teams |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Designs for the incident storm | "Volume spikes exactly when search matters most; quotas keep everyone else's logs flowing." |
| Written loss order | "Debug sampled first, audit never; drops counted per tenant." |
| Index only what's queried | "The hot index is an accelerator; object storage is the record." |
| Prices retention | "90 days warm is about 1.5 PB of object storage — roughly $30K a month at list-price order of magnitude." |
| Contains schema failures | "Per-stream schemas; a bad field from one team can't reject another team's logs." |
| Plans for deletion | "Something will slip past redaction; chunks are partitioned so a purge is bounded." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Kafka buffers it" as the whole backpressure story | Ignores what happens when buffers fill and who loses |
| Agent retries without bounds | A logging outage becomes a production outage |
| Indexing everything for months, unpriced | Unaffordable at scale; no tiering |
| No per-tenant quotas | One team's debug flag blinds everyone |
| Shared dynamic mapping | Mapping explosions and type conflicts break other tenants |
| PII handled only at query time | Secrets persist in every replica and export |
5.4 Common False Positives#
- Deep Elasticsearch tuning ≠ log platform design. Shard sizing doesn't address loss policy, quotas or cost attribution.
- "We use Kafka" ≠ backpressure. A queue moves the buffer; it doesn't decide what happens when the buffer fills.
- Compression ratios ≠ cost control. Chargeback and volume reduction at source usually save more.
- Regex redaction ≠ PII safety. Without canaries and a deletion path, it's a hope, not a control.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Debugging primary, audit lane; incident storms, never block the app, cost, PII; numbers |
| Entities & API | 3–5 min | Tenant, stream, record, chunk; 429 as the backpressure contract |
| Architecture | 5–10 min | ≤ 8 boxes; agent → gateway → queue → processor → hot index + object storage |
| Backpressure | 10–18 min | Bounded buffers, loss order, signature sampling, queue retention |
| Storage tiers | 18–25 min | Hot index vs query on read, routing, retention tiers, prices |
| Multi-tenancy + schema | 25–32 min | Quotas, query limits, cardinality, per-stream schemas, raw fallback |
| PII + operations | 32–38 min | Layered redaction, canaries, deletion; in-quota lag SLO, reconciliation |
| Pivot / wrap | 38–45 min | Multi-region, cost per team, build vs buy; close on who owns the bytes |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "We can't lose any logs" | Do you know the cost? | Audit lane never sampled; debug has a loss order; price the alternative |
| "Cut the logging bill by 40%" | Cost levers | Chargeback, debug retention, sampling at source, tier moves — before storage engineering |
| "Make it multi-region" | Data locality | Region-local ingest and storage; federated queries; residency |
| "Security needs a year of auth logs" | Audit lane | Separate stream, write-once archive, gap detection |
| "A password was logged" | Deletion | Revoke, purge bounded chunks, fix at source, canary |
| "How do you know logs aren't missing?" | End-to-end verification | Agent-read vs stored reconciliation per stream; canary streams |
6.3 What to Deliberately Skip#
- Index internals — "inverted index for the hot tier; columnar chunks for warm."
- UI design — say what queries look like, not the screens.
- Agent implementation details — name the bounds, skip the file-tailing mechanics.
- Log formats per language — "structured JSON encouraged; raw accepted."
- Exactly-once ingestion — at-least-once with dedupe by (agent, file, offset) where it matters; say so in one line.
6.4 Follow-Up Questions to Expect#
- "One team's volume goes 10× during an incident. What happens to everyone else's logs?"
- "The ingest tier is down for an hour. Walk me through every buffer between the app and storage."
- "What do you index, for how long, and what does 90 days cost?"
- "Two services log
statuswith different types. What breaks?" - "A secret was logged for six weeks. How do you delete it everywhere?"
- "How do you know logs aren't silently missing?"
- "Security needs a year of retention for auth logs. How does that differ from debug logs?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design a centralized logging system for our company."
Staff Answer
"What are the logs for — debugging services, security audit, or analytics? I'll assume debugging is primary, with a separate audit lane, and that analytics belongs in metrics. Numbers: 3,000 services on 50,000 nodes, about 2 GB/s average and up to 10 GB/s during a bad incident — roughly 170 TB/day raw. Seven days of fast search, 90 days of slower search.
Four constraints shape everything: volume spikes during incidents exactly when search matters most; logging must never take down the application; cost scales with bytes and producers don't see it; and logs will contain secrets. I'll go: entities and the backpressure contract → pipeline → loss order → storage tiers and cost → tenancy and schema → redaction → ownership."
Why this is L6:
- Separates debugging from audit and analytics before drawing
- Frames the problem around incident storms and cost, with numbers
- Previews an outline that ends with ownership
What L7 adds:
- Asks what share of infra spend observability is today and who owns that budget
- Asks whether teams already run their own logging stacks — the consolidation question
- Frames cost per request served as the success metric
❌ Common L5 Trap
"Install an agent on every host that ships logs to Kafka, consume into Elasticsearch with daily indices and 30-day retention, and give engineers Kibana."
Why this misses: Each piece is reasonable; together they don't say what happens when one team storms, what 30 days costs, how schema conflicts are contained or where secrets are removed. The next follow-up — "one team triples volume" — has no answer.
Drill 2: The Ingest Outage#
Prompt: "The ingest tier is unreachable for one hour. Walk me through every buffer between the app and storage."
Staff Answer
"The application writes to stdout; the container runtime writes to a local file with rotation, say 5 files of 100 MB. The app is unaffected — it never talks to the network for logging. The agent can't ship, so its 64 MB memory buffer fills in seconds and spills to a 1 GB disk buffer. A typical node producing 200 KB/s fills that in about 80 minutes — so an hour is covered for most nodes. Chatty nodes at 1 MB/s fill it in ~17 minutes; after that the agent stops reading, falls behind the rotating files, and loses the oldest unread data when files rotate — counted in agent.dropped_lines_total per tenant.
When ingest returns, agents drain with jittered backoff so 50,000 nodes don't stampede the gateway; the gateway's per-tenant quotas allow a recovery burst. Downstream, the queue had nothing to absorb in this scenario — it protects against the indexing tier failing, not ingest. Afterward, I'd report per-tenant loss and consider larger disk buffers for chatty node pools."
Why this is L6:
- Walks every buffer with sizes and fill times
- Never blocks the application; drops are bounded and counted
- Handles recovery stampede and reports loss
What L7 adds:
- Decides buffer sizing as a cost-vs-loss tradeoff per node pool, agreed with the teams on those pools
- Adds a zone-level chaos test that blackholes ingest monthly
❌ Common L5 Trap
"The agent retries until the backend comes back, so no logs are lost."
Why this misses: Retrying requires buffering, and unbounded buffering fills the host's disk or memory. The candidate hasn't bounded it — and if the agent's buffer is unbounded, an hour-long logging outage becomes a production outage.
Drill 3: Make It Concrete — Size and Price the Tiers#
Prompt: "Size the storage tiers for 2 GB/s, 7 days hot, 90 days warm. Roughly what does it cost?"
Staff Answer
"2 GB/s × 86,400 ≈ 170 TB/day raw. Warm tier: with ~10× compression, ~17 TB/day; 90 days ≈ 1.5 PB in object storage. At roughly $0.02 per GB-month, that's about $30K a month, plus request costs — which is why chunks should be large, tens to hundreds of MB, not tiny.
Hot tier: if I indexed everything, 7 days of 170 TB with index overhead roughly comparable to raw, compressed maybe 2×, replicated twice, is on the order of 1–2 PB of fast storage — several times the cost of the warm tier. So I'd cut it: debug-level streams hot for only 24 hours, and over-quota data skipped. If debug is half the volume, hot drops to around 0.6–1 PB.
Queue: 48 hours × 2 GB/s ≈ 350 TB raw, ~50 TB compressed, replicated ×3. Compute: processors and indexers sized to peak, 5× average, which argues for elastic processing."
Why this is L6:
- Sizes each tier with compression, replication and retention
- Prices the warm tier and shows why the hot tier must be smaller
- Notices request costs and chunk sizing
What L7 adds:
- Turns this into cost per GB ingested per tier and publishes it for chargeback
- Uses the cost estimator to compare self-hosted against vendor pricing at 3-year volume
❌ Common L5 Trap
"170 TB a day times 90 days is 15 PB. We'll need a big Elasticsearch cluster."
Why this misses: Ignores compression, replication and tiering, and doesn't price anything. The answer implies indexing 15 PB on fast storage — an unaffordable design presented as a sizing exercise.
Drill 4: Index or Scan?#
Prompt: "Why not just index everything? Engineers want fast search over all 90 days."
Staff Answer
"Because the cost scales with every byte for the full retention while the queries don't. If 90% of searches hit the last 24 hours, indexing day 60 buys fast answers to a handful of queries at the price of indexing petabytes. I'd measure the query distribution by age — it's usually steep — and size the hot window to cover, say, 95% of queries.
For older data, object storage with good pruning: partition by tenant, stream and hour; per-chunk metadata with time bounds and labels; token bloom filters so a search for a request ID skips most chunks. A 30-day search over one service becomes tens of seconds. For the rare broad search, an async query job. If a team truly needs fast search over 90 days — security investigations — they get a longer hot window for a small, specific stream and pay for it."
Why this is L6:
- Argues from the query age distribution, with a measurement
- Describes how cold queries stay usable: pruning, bloom filters, async jobs
- Offers longer hot retention as a paid exception
What L7 adds:
- Publishes tier prices so teams make the choice with their own budget
- Revisits the hot window yearly as query patterns and costs shift
❌ Common L5 Trap
"Fast search matters during incidents, so we should index everything."
Why this misses: Incidents query the last hours, which the hot tier covers. Indexing 90 days to serve post-mortems that could wait 30 seconds is the most expensive way to buy convenience.
Drill 5: The Noisy Tenant#
Prompt: "A team turns on debug logging in production and their volume goes from 60 MB/s to 2.4 GB/s. Everyone else's logs are delayed. Fix it."
Staff Answer
"Immediately: the processor enforces that tenant's quota — say 100 MB/s with a 2× burst for 10 minutes. Above it, signature sampling: group lines by message template, keep the first 100 per signature per minute in full, and emit counts for the rest. Their excess skips the hot index and goes compressed to object storage if write capacity allows, so they can still query it, slower. Other tenants' lag recovers within a minute because indexers are no longer saturated.
Structurally: quotas enforced at the gateway and processor by default; debug streams kept hot for 24 hours; a 'debug window' feature to raise a quota temporarily for named pods with the cost shown up front. And the team sees a 'sampled — over quota' marker in their search results, so they understand what they're looking at."
Why this is L6:
- Isolates the storming tenant with quotas and signature sampling
- Preserves information rather than dropping blindly
- Gives teams a sanctioned way to get more volume temporarily
What L7 adds:
- Ties quotas to team budgets so raising one is a budget conversation
- Tracks top volume producers monthly as a reduction program
❌ Common L5 Trap
"Add more indexing nodes so we can handle the extra volume."
Why this misses: Scaling for one team's debug flag means paying peak capacity for a misconfiguration, and it takes longer than the storm. Without isolation, the next team's flag does it again.
Drill 6: Cardinality and Mapping Explosions#
Prompt: "A service starts putting user IDs into a label — or into JSON keys. What happens, and how do you prevent it?"
Staff Answer
"As a label, user IDs create one stream per user: millions of tiny streams, an index that grows with users and chunks too small to compress well — a label-indexed store degrades for everyone sharing it. As JSON keys with dynamic mapping, each user ID becomes a field; a shared index hits its field limit and starts rejecting writes for other tenants too.
Prevention: labels are a short, fixed set per stream — service, env, cluster, level — with a limit on labels per stream and distinct values per label per day, enforced at the gateway with a clear error. Parsed fields are per-stream registered schemas; unknown fields are kept in the raw body, searchable as text, never auto-added to a shared mapping. High-cardinality values like user and request IDs belong in the body or as indexed tokens in bloom filters, not as labels. And the platform reports per-stream cardinality so teams see it before limits hit."
Why this is L6:
- Explains the failure in both label-indexed and field-indexed stores
- Enforces limits at ingest with clear errors
- Gives a correct home for high-cardinality values
What L7 adds:
- Makes cardinality limits part of the logging standard with an exceptions process
- Ships linting in the logging libraries that flags dynamic keys before deploy
❌ Common L5 Trap
"Increase the field limit on the index."
Why this misses: It postpones the failure and makes the index slower and heavier for every tenant. The fix is preventing unbounded fields from entering a shared mapping at all.
Drill 7: Build vs Buy#
Prompt: "Should we run our own logging platform or use a vendor?"
Staff Answer
"Below a few TB/day, buy: a vendor gives search, retention tiers and alerting for less than one engineer's cost. Above tens of TB/day, vendor pricing per GB ingested often exceeds the cost of a platform team plus infrastructure, and that's when building — usually on open-source components and object storage — pays off. I'd model 3-year volume, not today's, because log volume tends to grow faster than traffic.
Whichever we pick, three things stay ours: the agent and pipeline layer, so we can redact, sample and route before bytes leave our network — and switch vendors without touching every host; the loss and retention policies; and chargeback to teams. A common hybrid: vendor for hot search, our own object storage for the long tail and audit archive."
Why this is L6:
- Gives a volume threshold and models future growth
- Keeps the pipeline under our control for redaction, cost and portability
- Proposes a hybrid split by tier
What L7 adds:
- Negotiates vendor contracts on committed volume with an exit plan and data export
- Treats vendor lock-in risk as a line item: query language, dashboards, alert definitions
❌ Common L5 Trap
"Use a vendor — logging isn't our core business."
Why this misses: True at small scale, but it ignores the volume at which per-GB pricing dominates, and it hands redaction and routing to whoever owns the agent. The decision is a threshold, not a slogan.
Drill 8: Changing a Parsing Rule Without Losing Logs#
Prompt: "You need to change the parser for 400 services' access logs to a new format. How do you ship it?"
Staff Answer
"Parsers are versioned config per stream, not code baked into processors. I'd deploy the new parser in shadow: both versions parse a copy of live traffic; we compare field-level output and parse-failure rates per stream for a day. Streams with differences above a threshold get reviewed before switching.
Rollout by stream: 1% of streams, then 10%, then all, watching processor.parse_errors{stream} and end-to-end line counts. The invariant that makes it safe: a line that fails to parse is stored raw with a parse-error tag, never dropped — so the worst case of a bad parser is less-structured data, not missing data. Rollback is a config change. Dashboards and saved queries that depend on renamed fields get a compatibility alias for 30 days."
Why this is L6:
- Shadow parsing with field-level comparison before switching
- Staged rollout per stream with clear metrics
- The raw-fallback invariant bounds the blast radius
What L7 adds:
- Gives teams ownership of their stream parsers with platform-provided test harnesses
- Treats field renames as API changes with deprecation windows
❌ Common L5 Trap
"Update the parser in the processors and redeploy."
Why this misses: A parser bug silently drops or mangles lines for 400 services at once, and nobody notices until an incident needs the missing logs.
Drill 9: Cut the Bill by 40%#
Prompt: "Finance says logging costs too much. Cut it by 40% without hurting incident response."
Staff Answer
"Start with attribution: cost per stream, with query counts. Usually 10–20 streams drive half the volume, and some are rarely queried. Levers in order of ROI:
- Volume at source: debug off by default in production, health-check and access-log sampling (keep errors, sample 200s at 1–10%), duplicate stack traces collapsed. Often 30–50% of volume.
- Tiering: debug hot for 24 hours, not 7 days; warm default 30 days, longer only if paid for.
- Hot-index scope: index only labels and registered fields, keep body text in compressed chunks with bloom filters.
- Chargeback: teams see their bill monthly; behavior follows.
Incident response is protected because error-level logs are never sampled, the hot window still covers recent hours, and quotas keep storms isolated. I'd measure time-to-find for a set of standard incident queries before and after."
Why this is L6:
- Starts with attribution and query counts, not storage tricks
- Orders levers by ROI with rough impact
- Defines a metric to prove incident response isn't hurt
What L7 adds:
- Sets an org-wide observability cost target as a share of infra spend
- Makes logging cost part of service launch reviews
❌ Common L5 Trap
"Switch to cheaper instances and compress better."
Why this misses: Compression and instance tweaks save a fraction of what reducing volume and retention saves, and they don't change the behavior that keeps growing the bill.
Drill 10: Multi-Region#
Prompt: "We run in four regions. How does logging work across them?"
Staff Answer
"Region-local ingest, processing and storage. Shipping logs across regions doubles egress cost and adds a cross-region dependency to every write, and many logs are subject to residency rules. Each region has its own agents' gateway, queue, hot tier and object storage bucket.
Queries are federated: the query frontend fans out to the regions in scope, each executes locally and returns results or partial aggregates, and the frontend merges. Most incident queries target one region anyway. If a region is unreachable, results are marked partial rather than failing. Audit logs may need a copy in a second region for durability — that's a deliberate, small, replicated stream, not all logs. For data residency, EU logs never leave the EU, including query results cached elsewhere."
Why this is L6:
- Region-local storage avoids egress and cross-region write dependencies
- Federated queries with partial-result semantics
- Replicates only what needs replication; respects residency
What L7 adds:
- Prices cross-region egress explicitly and makes replication an exception with an owner
- Defines residency per tenant contract across all observability data, not just logs
❌ Common L5 Trap
"Ship all logs to one central region so engineers can search everything in one place."
Why this misses: Pays cross-region egress on every byte, makes every region's logging depend on one region's availability, and breaks residency. Federation gives the single search box without moving the data.
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — Logs Go Dark During the Outage#
Context: A database failover causes errors across 600 services. Log volume jumps from 2 GB/s to 11 GB/s within two minutes. Ingest lag for all tenants climbs to 14 minutes, and incident responders can't see current errors. The on-call escalates to you.
Questions to Surface First:
- Where is the bottleneck — gateway, queue, processors or hot-tier indexing?
- Is the surge broad (600 services) or dominated by a few?
- Are quotas being enforced, or are all tenants bursting at once?
- Are responders' queries themselves loading the hot tier?
Typical L5 Approach: Scales the indexing cluster. Shard rebalancing adds load; lag gets worse for 20 minutes before it improves.
Staff Approach: Finds the processors keeping up but hot-tier indexing saturated. Since every tenant is bursting, per-tenant quotas don't help — so applies the platform-wide loss order: debug and info streams skip the hot index (still written to object storage), error and warn keep flowing. Hot-tier lag drops to 20 seconds for error-level logs within 3 minutes.
Principal Approach: Recognizes that broad incidents create correlated bursts that per-tenant quotas can't handle and that log volume during failures is predictable. Builds a platform-level "incident mode" — pre-agreed level-based shedding — and drives a program to collapse repeated error logs at the source (rate-limited loggers per message template).
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Stage metrics: gateway ok, queue ok, processors ok, indexer CPU 100%. Enable level-based shedding: debug/info to object storage only. |
| Triage | 70% of surge is repeated stack traces from retry loops; signature sampling would cut it ~20×. |
| Quick fix | Enable signature sampling platform-wide for the incident; error-level examples preserved. |
| Guardrails | Alert on ingest.lag_p99_s{level=error} > 60 separately from other levels. |
| Post-mortem | Incident mode as a documented, one-click policy; logging libraries get per-template rate limits. |
Metrics to Watch: ingest.lag_p99_s{level}, indexer.cpu_util, sampling.suppressed_lines_total{signature}, query.queue_wait_p99
Organizational Follow-up: engineering leadership approves incident-mode shedding as policy; library teams ship template rate limiting.
Ownership Question: "Who decides which logs are shed during a company-wide incident?" Staff answer: The policy is decided in advance by engineering leadership and security; the logging on-call executes it without asking anyone during the incident.
Key Takeaway: "When everyone storms at once, quotas don't help — a pre-agreed level-based loss order does."
What clears the Staff bar:
- Localizes the bottleneck before scaling
- Uses a loss order that preserves error-level logs
- Pushes the fix to the source: rate-limited repeated errors
Deep Dive 2: Silent Failure — Three Days of Missing Logs From One Region#
Context: During a post-mortem, an engineer notices that logs from one region's newest node pool are missing for the last three days. No alerts fired. Ingest lag, error rates and queue depth all look normal.
Questions to Surface First:
- Are the agents on those nodes running, and what version?
- What do agent-side counters say about lines read vs shipped?
- Is there any end-to-end reconciliation between agents and storage?
- Did anything change for that node pool three days ago?
Typical L5 Approach: Restarts agents on the node pool; logs start flowing. The three days are lost and the cause is unknown.
Staff Approach: Finds the node pool uses a new container runtime that writes logs to a different path; the agent's file discovery doesn't match it, so the agent is healthy and shipping nothing. Fixes discovery; recovers what's still on the nodes' rotated files. Adds end-to-end reconciliation: per node, expected log activity versus received; alerts on nodes with running pods and zero received lines.
Principal Approach: Treats "absence of data" as a first-class signal across observability. Every node and every service has an expected heartbeat in logs, metrics and traces; silence pages. Node pool changes require an observability checklist.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check agent.lines_shipped{node_pool}: zero for 3 days. Agents healthy. |
| Triage | New runtime log path not in agent discovery globs; agent reports "no files" as normal. |
| Quick fix | Update discovery config; ship surviving rotated files; mark the gap in the platform status. |
| Guardrails | Alert: node with running pods and zero shipped lines for 10 min; per-service "silent stream" alert. |
| Post-mortem | Node pool rollouts include an observability smoke test: synthetic log line found in search within 60 s. |
Metrics to Watch: agent.lines_shipped{node}, stream.silent_minutes{service}, canary.log_found_latency_s
Organizational Follow-up: infrastructure team adds the logging smoke test to node pool rollout automation.
Ownership Question: "Who owns noticing that logs stopped?" Staff answer: The logging platform owns end-to-end completeness signals; the infrastructure team owns running the smoke test when they change the node image.
Key Takeaway: "A healthy agent shipping nothing looks exactly like a quiet service. Alert on silence, not just on errors."
What clears the Staff bar:
- Recognizes missing data needs its own detection
- Adds end-to-end reconciliation and canaries
- Gets the change process for node pools to include observability
Deep Dive 3: Large-Customer Onboarding — A Team With 400 TB/Day#
Context: The CDN team wants to move edge request logs onto the platform: 400 TB/day raw, more than double the platform's current total. They want 30 days of search and dashboards over status codes and latencies.
Questions to Surface First:
- Are these logs searched for needles, or aggregated for dashboards?
- What fraction is ever queried? Do they need every 200 response?
- Is the format regular enough to treat as structured events?
- Who pays, and what's their budget?
Typical L5 Approach: Triples the hot cluster and onboards them into the shared index. Ingest saturates during their daily peak; every other tenant's lag rises.
Staff Approach: Recognizes these are events, not debug logs: fixed schema, aggregation-heavy queries. Routes them to a dedicated columnar pipeline — typed columns, high compression, fast aggregations — with metrics extracted at ingest for dashboards. Errors (4xx/5xx) and a 1% sample of 2xx are kept searchable for 30 days; full data in object storage for 7 days for investigations. Their quota and cost are separate.
Principal Approach: Uses the onboarding to define a platform tier for high-volume structured events, with its own pricing, and sets the rule that dashboards over logs must be served by extracted metrics.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (design review) | Query analysis: 95% of CDN team queries are aggregations by status, POP, path prefix. |
| Triage | Full ingest into the shared hot tier would cost several times their budget and endanger other tenants. |
| Quick fix | Dedicated columnar lane; metrics extracted at ingest; errors + 1% of 2xx searchable 30 days. |
| Guardrails | Separate quota and processing pool; no shared hot-tier capacity. |
| Post-mortem (pre-mortem) | What happens during a global CDN incident? Their lane sheds 2xx sampling to 0.1% automatically. |
Metrics to Watch: lane.cdn.ingest_lag_s, lane.cdn.bytes_stored, extracted_metrics.lag_s, shared-tier lag (must not move)
Organizational Follow-up: CDN team's budget covers the lane; finance sees it as a separate line.
Ownership Question: "Who owns the CDN team's dashboards?" Staff answer: The CDN team owns the queries and sampling choices; the platform owns the lane's isolation and the extraction pipeline.
Key Takeaway: "High-volume regular logs are events. Give them a columnar lane and extracted metrics, not the shared search index."
What clears the Staff bar:
- Classifies the data by query pattern before choosing storage
- Isolates the new tenant's capacity and cost
- Uses sampling and metric extraction instead of indexing everything
Deep Dive 4: Post-Mortem — Card Numbers in Logs for 41 Days#
Context: A compliance scan finds full payment card numbers in log search results. A checkout service started logging raw request bodies on validation errors 41 days ago. The data is in the hot tier, object storage, and a weekly export to the analytics warehouse. You're leading the post-mortem.
Questions to Surface First:
- Why didn't pipeline redaction catch card numbers?
- Where exactly does the data exist — tiers, replicas, exports, backups?
- Who accessed those log streams during the window?
- Can we purge object storage without rewriting the entire 41 days?
Typical L5 Approach: Adds a regex for 16-digit numbers to the pipeline and deletes the hot index for those days.
Staff Approach: Finds the card numbers were embedded in a URL-encoded body the regex didn't decode. Adds Luhn-validated detection on decoded content, redaction in the logging library for request bodies, and a canary that sends a test card number through checkout daily. Purges hot-tier documents by query, rewrites only the affected tenant-and-hour chunks in object storage, deletes the export partitions, and produces an access report for compliance.
Principal Approach: Treats this as a gap in data governance: logs had no classification and no owner for what fields services emit. Puts logging under the data classification standard, requires that payment services' logs pass through stricter allowlist-only redaction, and makes canary pass rate a compliance control.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Restrict access to affected streams; stop the export job; hotfix checkout to stop logging bodies. |
| Triage | Regex ran on raw text; URL-encoded digits bypassed it. Data in hot tier, 41 days of chunks, 6 export partitions. |
| Quick fix | Decode-then-detect with Luhn validation; purge hot docs; rewrite affected chunks; delete export partitions. |
| Guardrails | Daily card-number canary through checkout; alert if found unmasked anywhere. |
| Post-mortem | Payment-domain streams switch to allowlist-only fields; exports require classification review. |
Metrics to Watch: redaction.canary_pass{runtime,service}, redaction.matches_total{rule}, purge.chunks_rewritten
Organizational Follow-up: compliance notified with access report; security reviews logging libraries across languages.
Ownership Question: "Who owns what a service writes to logs?" Staff answer: The service team owns what they emit; security owns the redaction rules and classification; the platform owns making redaction and purges reliable.
Key Takeaway: "Redaction you don't test with canaries is redaction you hope works. And design storage so you can delete from it."
What clears the Staff bar:
- Finds why redaction failed, not just adds another rule
- Purges every copy, including exports, with bounded rewrites
- Adds continuous verification with canaries
Deep Dive 5: Multi-Region Expansion — Residency for EU Logs#
Context: The company is launching EU data residency. Logs from EU workloads may contain personal data and must stay in the EU. Today all logs flow to a central US platform.
Questions to Surface First:
- Which services process EU personal data, and do their logs contain it?
- Can we run a full regional stack, or only regional storage?
- How do engineers outside the EU search EU logs for incidents?
- Where do derived artifacts go — extracted metrics, alerts, exports?
Typical L5 Approach: Deploys a second cluster in the EU and routes EU logs there. Leaves the analytics export and alert notifications, which include log excerpts, flowing to the US.
Staff Approach: Full regional stack in the EU: agents, gateway, queue, processors, hot tier and object storage. Federated queries from the global frontend execute in the EU and return results to authorized users under an access policy; results aren't cached outside the EU. Extracted metrics without personal data can flow globally; alert notifications carry links, not excerpts.
Principal Approach: Defines residency as a tenant attribute honored by every observability system — logs, traces, metrics labels, alert payloads — with automated marker tests proving compliance, and a single access-policy layer for cross-region queries.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Inventory: raw logs, hot index, chunks, exports, alert payloads, query caches, extracted metrics. |
| Triage | Alert payloads and analytics exports carry excerpts → residency violations. |
| Quick fix | EU stack; alerts link back to EU search; exports from EU go to an EU warehouse only. |
| Guardrails | Marker log lines with synthetic personal data in EU services; weekly scan of non-EU stores. |
| Post-mortem (pre-launch) | Cross-region query access audited; on-call outside EU uses federated queries with logging of access. |
Metrics to Watch: residency.marker_found_outside_region (must be 0), eu.ingest.lag_p99_s, federated_query.eu.count{user_region}
Organizational Follow-up: legal approves the access policy for non-EU responders; on-call training covers federated search.
Ownership Question: "Who proves EU logs stayed in the EU?" Staff answer: The logging platform proves it for its stores with marker tests; privacy and compliance own the company-level attestation.
Key Takeaway: "Residency includes everything derived from logs — alerts, exports and caches, not just the index."
What clears the Staff bar:
- Inventories derived data, not just primary storage
- Uses federated queries instead of moving data
- Proves residency with tests
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why log volume spikes during incidents and why the platform must stay useful then
- Design bounded buffering at every stage with a written loss order that never blocks applications
- Choose what to index (hot window) and what to query on read (object storage), and price both
- Enforce per-tenant ingest quotas, query limits and cardinality limits, with signature sampling under overload
- Contain schema conflicts with per-stream schemas and raw fallback
- Layer PII redaction, verify it with canaries, and design storage so purges are bounded
- Size and price storage tiers, and attribute cost to producing teams
- Detect silent loss with end-to-end reconciliation and silence alerts
The Bar for This Question#
Mid-level (L4): Ships logs from hosts to a search cluster with dashboards. Works at small scale. No backpressure story, no quotas, one retention, no redaction.
Senior (L5): Adds agents, Kafka, Elasticsearch with daily indices, retention policies and some parsing. The gap: "Kafka buffers it" ends the backpressure discussion; indexes everything for the full retention without pricing it; no per-tenant isolation; shared dynamic mapping; PII left to developer discipline. The design works on a quiet day and fails during the first incident storm.
Staff+ (L6): Frames the problem as staying useful during the worst hour at a bounded cost within the first five minutes. Bounds every buffer and writes down the loss order. Indexes the hot window and queries the rest on read, with prices. Isolates tenants with quotas, query limits and cardinality limits. Contains schema failures per stream. Layers redaction with canaries and a deletion path. Names who pays — storming tenants lose fidelity on their own repeated logs; everyone else pays nothing. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 Most Logs Are Never Read#
| Data | Typical Fate |
|---|---|
| Debug lines older than a day | Rarely queried |
| Successful request logs | Mostly aggregated, rarely searched |
| Error logs from the last hours | The queries that matter |
The Staff position: Index what's searched; keep the rest cheaply or not at all. Volume reduction at the source beats storage engineering.
Why this matters in interviews: Pricing retention and admitting most bytes are cold is a Staff signal.
10.2 "Never Lose a Log" Is a Promise You Can't Afford#
| Promise | Cost |
|---|---|
| No loss, ever | Block apps or buffer unboundedly |
| No loss for audit | Small reserved lane — affordable |
| Ordered loss for debug | Policy, sampling — cheap |
The Staff position: Promise completeness for audit; publish a loss order for everything else.
Why this matters in interviews: It shows you design the failure behavior instead of pretending there isn't one.
10.3 Logs Are a Bad Metrics System#
| Need | Wrong Tool | Right Tool |
|---|---|---|
| Error rate dashboard | Count log lines every refresh | Emit a counter |
| Latency percentiles | Parse durations from logs | Histogram metric |
| Request flow | Correlate log lines by hand | Trace |
The Staff position: Extract metrics at ingest if you must; better, emit them at the source.
Why this matters in interviews: Redirecting analytics to metrics is a strong scoping move.
10.4 Shared Dynamic Mapping Is a Multi-Tenant Bug#
The Staff position: Any shared index where one tenant's new field can reject another tenant's logs is an isolation failure. Per-stream schemas and raw fallback aren't polish; they're tenancy.
Why this matters in interviews: It's the failure Senior designs don't see until it happens.
10.5 Chargeback Saves More Than Compression#
The Staff position: Teams who see their logging bill next to their query counts cut volume. No codec delivers that.
Why this matters in interviews: It shows you treat cost as an organizational problem with a technical assist.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer builds a log platform that stays useful during incidents, isolates tenants and prices its tiers. The Principal engineer notices that logs, metrics and traces are three pipelines with three agents on every node, three bills, three retention policies and three places personal data can leak — and that observability spend is growing faster than revenue. The L7 problem is observability as one budget with one collection layer: a shared agent and pipeline, per-signal storage, consistent tenancy and redaction, and a cost target the organization manages to.
The Org-Level Fault Line#
One observability pipeline vs per-signal stacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Separate stacks for logs, metrics, traces | Each team optimizes its signal | Three agents per node, three tenancy models, three redaction gaps | Nodes (overhead), security (gaps), finance (fragmented spend) |
| One collection layer, per-signal backends | One agent, one tenancy and redaction model, one cost report | Pipeline team must serve diverse signals | Platform team (scope) |
| One vendor for everything | Simplest operations | Lock-in; per-GB pricing at scale | Finance (vendor cost) |
🧭 Principal Move: "One collection layer — agent, gateway, redaction, tenancy, quotas and cost attribution — shared by logs, metrics and traces. Backends stay specialized. Every team gets one observability bill with three lines."
Cost Model#
Assumptions: object storage at ~$0.02/GB-month order of magnitude; ~10× compression; hot tier on replicated fast storage; fully loaded engineer ~$250K/year. Illustrative ranges.
| Scale | Volume | Infra ($/month) | Headcount | On-call Load | Notes |
|---|---|---|---|---|---|
| Startup | 100 GB/day | ~$1–5K vendor or small cluster | Part of one eng | Shared rotation | Buy |
| Growth | 20 TB/day | ~$40–120K (hot tier, object storage, queue, compute) | 4–6 eng | Dedicated rotation | Chargeback starts paying for itself |
| Large | 170+ TB/day, 4 regions | ~$300K–1M | 12–20 eng (pipeline, storage, query, agents) | Per-region rotations | Volume reduction program is the top ROI |
The pricing insight: the hot index usually costs more than the warm tier despite holding a fraction of the days. Shrinking the hot window for debug-level streams and reducing volume at the source typically saves more than any storage-engine migration.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Agent deployed to every node | One-way-ish | Replacing it touches the whole fleet |
| Query language exposed to engineers | One-way-ish | Saved queries, dashboards, alerts depend on it |
| Storing unredacted data in immutable archive | One-way | Purges are expensive; legal exposure persists |
| Label schema and cardinality policy | One-way-ish | Changing labels breaks queries and alerts |
| Hot window length | Two-way | Config |
| Storage engine for warm tier | Two-way | Migration, but data is in open formats in object storage |
| Quotas per tenant | Two-way | Config |
The Standard I'd Write#
RFC-OBS-003: Logging Standard
Status: Approved Owners: Observability Platform + Security
Scope
Every service and job that emits logs in any environment.
MUST
1. Emit logs to stdout or the platform agent; never ship directly from
application code to storage.
2. Never block request handling on logging.
3. Use only the standard label set; high-cardinality values go in the body.
4. Pass through platform redaction; payment and identity domains use
allowlist-only fields.
5. Route security and audit events to the audit lane.
SHOULD
1. Emit structured JSON with a registered schema per stream.
2. Rate-limit repeated messages per template in the logging library.
3. Emit metrics for anything counted on a dashboard.
Exceptions
Filed with the Observability Platform; security review for redaction
exceptions; time-boxed to two quarters.
Success metrics
- In-quota ingest-to-searchable p99: under 30 s
- Redaction canary pass rate: 100%
- Logging cost per 1M requests served: tracked per team, trending down
- Unplanned loss for in-quota tenants: under 0.1%
What I'd Tell the VP#
"Logging is our second-largest infrastructure cost after compute, and it's growing faster than traffic because nobody who writes logs sees the bill. During our last two major outages, the logging platform fell behind exactly when we needed it. I'm proposing three things: per-team quotas and monthly cost reports so teams own their volume, a tiered design where only the last week is fast and expensive, and a pre-agreed policy for what gets dropped during a company-wide incident. I expect a 30–40% cost reduction within two quarters, mostly from teams cutting volume once they see it, and logs that stay searchable during outages."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices observability as a budget | "Logging cost per million requests is the number I'd hold teams to." |
| Identifies one-way doors | "Unredacted data in an immutable archive is forever; the hot window isn't." |
| Redraws ownership | "Teams own their bytes; security owns redaction rules; the platform owns isolation." |
| Unifies collection | "One agent per node for all signals, one tenancy model, one bill." |
| Knows when not to build | "Below a few TB a day, buy, and keep the agent layer ours." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a great logging platform" — correct, but ignores the metrics and tracing agents on the same nodes.
- "Teams can see their usage in a dashboard" — good, but without budgets or targets nothing changes.
- "We'll add an EU cluster" — no mention of alert payloads, exports and caches derived from logs.
Appendices
Appendix A: Mechanics in Depth#
A.1 Agent Buffering Loop#
config:
mem_buffer_max = 64MB
disk_buffer_max = min(1GB, 10% of free disk)
batch_max = 1MB or 1s
loop:
read new lines from discovered files (respect rotation, track offsets)
redact_patterns(line) # tokens, keys, Luhn-valid card numbers
append to mem_buffer
if mem_buffer full: spill oldest batch to disk_buffer
if disk_buffer full: drop oldest batch; dropped_lines[tenant] += n
send batches oldest-first:
204 -> commit offsets
429 -> backoff(Retry-After, jitter); keep buffered
5xx/timeout -> exponential backoff with jitter; keep buffered
A.2 Signature Sampling#
signature(line) = hash(template(line)) # numbers, IDs, hex stripped
per tenant, per minute:
if tenant under quota: keep all
else:
count[sig] += 1
if count[sig] <= 100: keep full line (hot + warm)
else: warm only (if capacity), emit summary record every 10s:
{ sig, example, suppressed_count }
audit streams: never sampled
A.3 Query Routing#
route(query):
require label selector with at least one of {service, tenant stream}
split time range into hot (<= 7d) and warm (> 7d) parts
hot part -> index search, limit by tenant concurrency
warm part -> list chunks by (tenant, stream, hour) -> prune by bloom filter
-> parallel scan, cap bytes_scanned per query
-> if over cap: return partial + offer async job
merge, sort by ts, apply limit
Appendix B: Data Model#
CREATE TABLE tenants (
tenant_id TEXT PRIMARY KEY,
cost_center TEXT NOT NULL,
ingest_quota_bps BIGINT NOT NULL,
burst_bytes BIGINT NOT NULL,
max_labels INT NOT NULL DEFAULT 12,
max_query_bytes BIGINT NOT NULL
);
CREATE TABLE stream_policies (
tenant_id TEXT NOT NULL REFERENCES tenants,
selector TEXT NOT NULL, -- e.g. {service="checkout",level="debug"}
hot_days INT NOT NULL DEFAULT 7,
warm_days INT NOT NULL DEFAULT 30,
schema_ref TEXT, -- registered parser/schema version
redaction TEXT NOT NULL DEFAULT 'standard', -- standard | allowlist
PRIMARY KEY (tenant_id, selector)
);
-- chunk metadata (object storage index):
-- chunks(tenant_id, stream_hash, hour, chunk_id, t_min, t_max, bytes,
-- bloom_ref, schema_ref) -- object key: tenant/stream_hash/yyyy/mm/dd/hh/chunk_id
Appendix C: Coordination Mechanisms#
C.1 Ingest Under Quota#
C.2 Quick Comparison#
| Mechanism | Guarantees | Failure Mode | Use For |
|---|---|---|---|
| Bounded agent buffer | Host protected | Long outages lose oldest data | Every node |
| 429 + Retry-After | Agents slow down, don't drop | Agents ignore it | Ingest contract |
| Durable queue | Indexing outages absorbed | Retention exceeded | Between ingest and processing |
| Tenant quota + burst | Storms isolated | Quota too low for real needs | Every tenant |
| Signature sampling | Information kept under overload | Counts approximate | Over-quota tenants |
| Per-stream schema | Conflicts contained | Unregistered streams less queryable | Structured logs |
| Redaction canaries | Coverage verified | Canaries not representative | Every runtime |
| End-to-end reconciliation | Silent loss detected | Counters themselves lost | Every stream |
Appendix D: API Contract & Client Behavior#
- Agents batch up to 1 MB or 1 second, gzip, and retry on 429/5xx with jittered backoff; they never drop on 429.
- Agents report
dropped_lines_total,buffer_bytesandlines_shippedper tenant as metrics. - Applications log to stdout; structured JSON preferred, one record per line, timestamps in UTC with milliseconds.
- Labels: a fixed, small set per stream; request and user IDs go in the body.
- Queries beyond the hot window require a label selector; large scans run as async jobs.
Appendix E: Observability#
Core metrics:
ingest.lag_p99_s{in_quota},ingest.lag_p99_s{level}agent.dropped_lines_total{tenant},agent.buffer_pct_of_limit,agent.lines_shipped{node}tenant.ingest_bytes_s,tenant.quota_exceeded,sampling.suppressed_lines_totalindex.field_count,ingest.rejected_total{reason},processor.dropped_total{reason}query.bytes_scanned{user},query.queue_wait_p99redaction.canary_pass,cost.per_gb_by_team
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| In-quota ingest lag p99 | > 60 s for 5 min | Page |
| Processor drops without reason | > 0 | Page |
| Redaction canary found unmasked | any | Sev-1 page (security) |
| Node with pods and zero shipped lines | > 10 min | Ticket → page if > 5% of nodes |
| Queue oldest unconsumed age | > 36 h | Page |
| Agent buffer > 80% of limit fleet-wide | > 5% of nodes | Ticket |
Debugging the silent failure: ingest lag and error rates look fine when data never arrives. Reconcile lines read by agents against lines stored, alert on silent streams and nodes, and run a canary log line through every node pool.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 1 TB/day | Vendor, or one search cluster with 14–30 days | First incident storm |
| 1–20 TB/day | Agents with bounded buffers, queue, quotas, per-stream schemas | Hot-tier cost |
| 20–200 TB/day | Hot index + object storage, signature sampling, chargeback | Cross-region egress, multi-signal agents |
| > 200 TB/day, multi-region | Regional stacks, federated queries, shared observability pipeline | Org cost governance |
What you don't build on day one: tiered storage, chargeback, signature sampling, federated queries, an audit lane with tamper evidence. Each has a trigger in Section 11.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Quotas: bytes per second with burst per tenant; raised through a budget request, not a ticket to the platform team.
- Query fairness: concurrent queries and bytes scanned per tenant and per user; long scans queued as async jobs.
- Cardinality: labels per stream and distinct values per label per day capped; violations rejected with a clear error.
- Cost attribution: bytes ingested, GB-months per tier and bytes scanned reported per team monthly, with top streams and their query counts. Rehearse the sizing with the back-of-envelope calculator and the cost estimator.