Hiring BarSupport

Design Datadog (Metrics & Monitoring) — Staff-Level Case Study

Case study66 min read7 diagrams

Technologies referenced in this case study: Time-Series Databases · Kafka · Flink · Cassandra · Redis · Elasticsearch

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 5
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Appendix A (storage)
Deep Dive3+ hrsEverything, including The Principal Lens and appendices
What is a Metrics & Monitoring Platform? — Why interviewers pick this topic

A metrics platform collects numeric measurements — CPU, request latency, queue depth, orders per minute — tagged with dimensions (host, service, region, endpoint), stores them as time series, serves dashboards and ad-hoc queries, and evaluates alert rules that page humans. Datadog, Prometheus, Google's Monarch, Uber's M3, and Netflix's Atlas are all variations on this shape.

Before vs After — the "one more tag" deploy:

Without cardinality controls:
t=0:       A team adds tag customer_id to http.request.duration (2M customers)
t=+10min:  Active series for one tenant: 1.2M -> 38M
t=+15min:  Ingest shard memory 60% -> 97%; index inserts back up
t=+18min:  Shared shard OOMs; 400 other tenants on it lose ingestion
t=+20min:  Alert evaluator reads empty series; "no data" alerts disabled by default
t=+45min:  A real outage at a different customer goes unpaged
t=+2h:     Postmortem: the monitoring system was the blast radius

With per-tenant cardinality limits:
t=0:       Same deploy, same tag
t=+2min:   New-series rate for the tenant exceeds 50K/min; limiter engages
t=+2min:   Excess series dropped at intake; tenant notified with the offending tag
t=+5min:   Shared shards unaffected; every other tenant's alerts keep evaluating
t=+1h:     Team switches customer_id to a log/trace attribute; metric reverts

Why interviewers reach for this question: It looks like "store numbers with timestamps," so it separates candidates who design a database from candidates who understand three things: cardinality — not data volume — is the unit of cost and failure; the alerting path is the product, and it must be more reliable than anything it monitors; and in a multi-tenant system, one tenant's tag choice can take down everyone's monitoring.

Mechanics Refresher: Storage and Ingestion Options
MechanismHow It WorksProsCons
Pull (Prometheus-style scraping)Server scrapes /metrics endpoints every 10–60 sKnows when a target is down (up == 0); target controls nothingNeeds service discovery and network reachability into every target; awkward for short-lived jobs and SaaS across firewalls
Push (agent-style)Local agent aggregates and ships batches every 10–15 sWorks across NAT/firewalls; natural for SaaS; buffers on network blips"No data" is ambiguous (dead host vs broken pipeline); needs intake-side protection
Gorilla-style compressionDelta-of-delta timestamps, XOR'd float values~1.37 bytes/point vs 16 raw (Facebook, VLDB 2015)Must decode a block to read any point in it
In-memory head + immutable blocksRecent 1–2 h in memory; flushed to compressed, indexed blocksFast writes and recent queriesMemory scales with active series, not bytes
Inverted index on tagstag=value → posting list of series IDsFast tag filteringIndex size and churn scale with cardinality
Rollups (downsampling)Store min/max/sum/count per 1 m, 1 h windows10–100× less data for long rangesPercentiles cannot be rolled up from percentiles
Mergeable sketches (DDSketch, t-digest)Store a quantile sketch per windowPercentiles survive rollup with bounded error (~1–2%)Larger per-point size; tenant must opt into distributions

For most production systems: Agent push for a SaaS product (pull inside a single company's network is fine), Gorilla-style compressed blocks with an in-memory head, an inverted tag index, rollups with sum/count/min/max, and sketches for latency distributions. The storage engine is not the interview — cardinality control, alert-path reliability, and tenant isolation are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

A metrics platform is not a time-series database question. Everyone can describe a TSDB.

It is a cost-of-cardinality and trust-in-alerts question that tests:

  • Whether you identify active series — not bytes — as the unit of memory, cost, and failure
  • Whether you design the alerting path to be more reliable than the systems it watches, with explicit behavior for late and missing data
  • Whether you isolate tenants so one bad tag can't blind everyone
  • Whether you connect resolution, retention, and query flexibility to a price someone pays

The key insight: The monitoring system is the one piece of infrastructure that must keep working during everyone else's outage — when traffic, cardinality, and query load all spike at once. Staff engineers design for that moment: the write path protects the alert path, the alert path protects itself from the query path, and every tenant has a cardinality budget.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveAgents → Kafka → TSDB → dashboardsAsks: SaaS multi-tenant or internal? What's the alert freshness SLO? What's the cardinality per tenant?Asks what monitoring costs as a % of infra spend today, and who decides what gets instrumented
Scale unitData points per secondActive series and series churnSeries per engineer and $ per series — a budget per team
AlertingCron job running queries against the TSDBStreaming or dedicated evaluation path, isolated from dashboards; explicit late-data and no-data semanticsAlert quality as an org SLO: pages per on-call shift, actionable-page ratio, alert ownership
Multi-tenancyTenant ID columnPer-tenant limits on series, ingest rate, query cost; cells for blast radiusPricing and chargeback that make cardinality visible to the engineer adding the tag
RetentionKeep everythingTiered: raw 10 s for ~15 days, 1 m rollups for ~15 months, sketches for percentilesSets the retention contract per data class, priced, with legal and finance
Failure"Replicate the TSDB"Monitors the monitor from an independent failure domain; watchdog alerts; degrade dashboards before alertsDesigns the org's failure posture so the monitoring vendor/platform is not a shared single point of failure
Why "scale unit" separates levels

L5: "10M points per second, 16 bytes each, 14 TB a day." Accurate arithmetic — and the wrong axis. With Gorilla compression the bytes are ~1.4 per point; storage is rarely what breaks.

L6: What breaks is active series. Each series costs memory in the ingest head (~a few KB with its index entries and open chunk), an entry in the inverted index, and a separate object in every query that touches it. A tenant with 1M series at 10 s resolution writes 100K points/s — manageable. The same tenant adding one tag with 1,000 values writes the same number of points per series but now has 1B potential series. "I'll size the system in active series and new-series-per-minute, and I'll put limits on both."

Why "alerting" separates levels

L5: An alert scheduler runs each rule's query against the TSDB every minute and compares to a threshold. It works — and it means dashboards and alerts compete for the same query capacity, exactly when an incident sends 500 engineers to the dashboards.

L6: Alert evaluation gets its own path: rules evaluated on the ingest stream or against a dedicated in-memory recent tier, with its own capacity. Each rule states how it treats late data (wait N seconds before evaluating a window), missing data (no-data alert vs resolve vs hold state), and flapping (hysteresis, for: 5m). "An alert that silently stops evaluating is worse than no alert — people believe it's watching."

Why "failure" separates levels

L5: Replicates storage and ingest for high availability.

L6: Asks: who tells us the monitoring system is down? Not itself. A small, independent watchdog in a different failure domain sends synthetic metrics end-to-end and pages via a separate channel if they don't arrive. And under overload, the system sheds in a defined order: ad-hoc queries first, dashboards second, long-range rollups third — alert evaluation and ingestion of alerted-on series last.

The Staff Positions#

PositionRationale
Size and limit in active series, not bytesMemory, index, and query cost all scale with series; bytes are compressed to ~1.4/point
Per-tenant cardinality limits enforced at intakeOne tenant's tag must not OOM a shared shard
Alert evaluation is isolated from query servingIncidents spike dashboard load exactly when alerts matter most
Explicit late-data and no-data semantics per ruleDefault "no data = OK" silently disables alerts
Downsample with sum/count/min/max + sketchesAverages of averages and percentiles of percentiles are wrong
Monitor the monitor from an independent failure domainA monitoring outage is invisible from inside
Shed load in a defined order: queries → dashboards → alerts lastProtects the one path that pages humans

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Operational alerting (SLO-grade)Freshness: detect in < 1–2 min; never silently missStreaming/in-memory evaluation path, isolated, redundant; explicit no-data handlingMissed page during an outage; alert stormAlert latency p99 < 60–120 s; zero silent evaluator gaps
Exploration & dashboardsFlexible queries over high-cardinality tags, fastInverted index, query fan-out, caching, pre-aggregationQuery of death scanning 10⁷ series; slow dashboards in incidentsp95 dashboard load < 2–3 s for 1 h windows
Long-term analytics & capacity planning13–15 months of history, cheapRollups in object storage, sketches for percentilesWrong percentiles from naive rollups; expensive long scansCorrect aggregates at 1 m/1 h resolution

🎯 Staff Move: "I'll design this as a multi-tenant SaaS like Datadog, and I'll treat the alerting path as the product. Dashboards and long-term analytics matter, but if we're slow on a dashboard, a user is annoyed; if we miss a page, a customer's outage goes unnoticed. So the write path protects the alert path, and cardinality limits protect the write path."

The Five Fault Lines#

#Fault LineThe Tension
1Cardinality: Expressiveness vs CostLet users tag anything (powerful queries) or cap series (predictable cost, frustrated users)?
2Push vs Pull IngestionAgents push (SaaS-friendly, ambiguous absence) or server scrapes (clear liveness, needs reachability)?
3Alert Freshness vs CompletenessEvaluate immediately (fast, false alarms on late data) or wait (correct, slower pages)?
4Shared vs Isolated TenancyPack tenants densely (cheap) or isolate in cells (bounded blast radius, more ops)?
5Resolution vs RetentionKeep raw data long (exact, expensive) or downsample (cheap, lossy — especially percentiles)?

In the Wild: Real Production Systems#

Why this section belongs here: Citing specific, publicly documented monitoring systems shows you've studied the real constraints — cardinality, compression, and global query.

Facebook Gorilla — Compression That Changed the Memory Math#

Facebook's Gorilla paper (VLDB 2015) described an in-memory TSDB for the most recent ~26 hours of data using delta-of-delta timestamp encoding and XOR float compression, getting from 16 bytes to ~1.37 bytes per point on average. That compression let them keep recent data in memory for fast queries, with an on-disk store behind it.

Staff insight: Once points cost ~1.4 bytes, bytes stop being the bottleneck and series overhead dominates. Every modern TSDB inherited this, and it's why cardinality — not volume — is the question.

Google Monarch — Regional Ingestion, Global Query#

Monarch (VLDB 2020) is Google's planet-scale, multi-tenant, in-memory monitoring system. Data is ingested and stored in regional "zones" close to where it's produced, and a global query layer fans out and merges. Monarch pushes aggregation toward the leaves and was explicitly designed so monitoring keeps working regardless of the health of the systems being monitored — for example, it avoids depending on Google's own distributed storage systems for its core path.

Staff insight: "Monitoring must not depend on what it monitors" is a design rule, not a slogan. Say it, and name the dependency you're removing.

Prometheus — Pull, Local Storage, and the Cardinality Warning#

Prometheus (originally from SoundCloud, now a CNCF project) scrapes targets, stores data locally in 2-hour blocks with an in-memory head, and evaluates recording and alerting rules on the same server. Its documentation repeatedly warns against high-cardinality labels like user IDs. Horizontal-scale projects such as Thanos, Cortex/Mimir, and Uber's M3 exist largely to add long-term storage, global query, and multi-tenancy on top of that model.

Staff insight: Prometheus colocates alert evaluation with the scraper — so alerts keep working even if central storage is down. That locality is a reliability feature worth copying in any design.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"We store 10M points/s""How many active series? What happens when a customer adds a user_id tag?"Cardinality as the real unit
"Alert service queries the TSDB every minute""It's an incident; 2,000 engineers open dashboards. Do alerts still fire on time?"Path isolation
"Agents push metrics""A host stops sending. Is it dead or is the pipeline broken?"No-data semantics
"We downsample to 1-minute averages""What's the p99 latency last March?"Rollup correctness, sketches
"Shard by metric name""One metric has 40% of all series. Now what?"Hot shards and sharding keys
"It's highly available""How do you find out that monitoring is down?"Monitoring the monitor

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Intake enforces limits before anything touches shared state. Kafka decouples intake from storage and lets two independent consumers read the same data: the alert evaluators (hot path, streaming, isolated capacity) and the ingest writers (storage path). Queries hit recent data in the ingest heads and older data in blocks. The watchdog lives outside the platform's failure domain and verifies the whole loop — metric in, page out.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Scale"10M points/s, X TB/day""~10⁹ active series across tenants; size memory and index by series; limit new-series rate per tenant."
Ingestion"Agents send to Kafka""Agents aggregate and push every 10–15 s; intake enforces per-tenant rate and cardinality limits before Kafka."
Storage"Cassandra with time buckets""In-memory head with Gorilla chunks, 2 h blocks to object storage, inverted tag index; ~1.4 bytes/point."
Alerting"Cron job queries the TSDB""Streaming evaluators off the log, isolated capacity, evaluation delay for late data, explicit no-data behavior."
Retention"Keep everything a year""Raw 10 s for 15 days; 1 m rollups (sum/count/min/max + sketches) for 15 months."
Failure"Replicate everything""Independent watchdog; shed queries before alerts; cells cap tenant blast radius."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Compressed point size (Gorilla-style)~1.37 bytesBytes are cheap; series are expensive
Raw point16 bytes (8 timestamp + 8 float)Why compression mattered
Memory per active series (head + index)~2–8 KB100M series ≈ 200–800 GB RAM across the fleet
Agent flush interval10–15 sSets the floor for alert latency
Typical alert evaluation interval30–60 sPlus evaluation delay 30–60 s for late data
End-to-end alert latency targetp99 < 60–120 sMetric emitted → page delivered
Series per tenant, typical mid-size10⁵–10⁷Limits are set here, with paid expansion
New-series rate limit per tenant~10K–100K/minCatches tag explosions within minutes
Head block duration~2 hRecent queries served from memory
Raw retention / rollup retention~15 days / ~15 months15 months covers year-over-year comparisons
1 m rollup vs 10 s raw6× fewer points; 1 h rollup ~360×Long-range queries must hit rollups
DDSketch relative error~1–2% configurablePercentiles that survive aggregation

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the storage engine — compression, column layouts, Cassandra schema — and never reach cardinality limits, the alert path, or tenant isolation. Compress storage to 5 minutes. The level is decided by what you protect and in what order.


Phase 1: Requirements & Framing (2–3 minutes)#

Functional scope in one sentence:

"Customers send tagged metrics from agents; we store them, serve dashboards and ad-hoc queries, and evaluate alert rules that notify on-call through pager, chat, or email."

Then the non-functional requirements that shape the design:

QuestionWhy It MattersDefault I'd Assume
Multi-tenant SaaS or internal?Isolation, limits, billingSaaS, ~20K tenants, heavy-tailed sizes
Scale?Memory and index sizing~1B active series total; ~10M points/s ingest
Resolution?Ingest volume, alert latency10 s default; 1 s for opt-in high-res
Alert latency SLO?Evaluation architecturep99 metric-to-page < 2 min
Retention?Storage tiersRaw 15 days; 1 m rollups 15 months
Query patterns?Index and cachingDashboards over 1 h–7 d, ad-hoc tag filtering, group-by

"The numbers I'll design around: a billion active series, and a two-minute metric-to-page SLO that must hold during a customer's worst day — which is exactly when their cardinality and query load spike."


Phase 2: Core Entities & API (1–2 minutes)#

Series      { tenant_id, metric_name, tags: sorted map, series_id = hash(tenant, name, tags) }
Point       { series_id, timestamp (s or ms), value (float64) | sketch }
Monitor     { monitor_id, tenant_id, query, window, threshold(s), eval_delay,
              no_data_policy, for_duration, notify_targets, owner_team }
AlertState  { monitor_id, group_key, state: OK | WARN | ALERT | NO_DATA, since, last_eval }

APIs:

POST /v1/series            batch of {metric, tags[], points[]}   -> 202 (async)
POST /v1/distribution      batch of sketches                     -> 202
GET  /v1/query?q=avg:http.latency{service:checkout} by {region}&from=&to=
POST /v1/monitors          create/update rule
GET  /v1/monitors/{id}/state

"The identity decision that matters: a series is tenant + metric + the full sorted tag set. Every distinct tag combination is a new series with its own memory cost — that's the definition of cardinality."


Phase 3: High-Level Architecture (≤5 minutes)#

Staff candidates spend under 5 minutes here. Draw two consumers off one log.

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Say it in four sentences:

  1. "Agents pre-aggregate locally and push batches every 10–15 s; intake authenticates, stamps the tenant, and enforces rate and cardinality limits before anything reaches shared state."
  2. "Kafka is the durable buffer: partitioned by tenant and series hash, 24–72 h retention, so storage can fall behind or replay without losing data."
  3. "Two independent consumers: alert evaluators consume the stream directly and keep windowed state in memory; ingest writers build compressed in-memory heads and flush 2-hour blocks to object storage with a tag index."
  4. "Queries fan out to heads for recent data and to blocks or rollups for older data, with per-query cost limits."

Phase 4: Transition to Depth (1 minute)#

"Storage engines for time series are well understood; I'll use the Gorilla-style design. The three places I'd like to go deep: cardinality — how we stop one tenant's tag from taking down shared infrastructure; the alerting path — latency, late data, no-data semantics, and isolation from queries; and multi-tenant isolation and cost. Which first?"

Default order: cardinality → alerting → tenancy/cost.


Phase 5: Deep Dives (25–30 minutes)#

Deep dive 1 — Cardinality (8–10 min).

  • Cost per series: head chunk + index postings + series metadata ≈ 2–8 KB of RAM, plus query fan-out cost.
  • Limits at intake, per tenant: active series cap (contracted), new-series-per-minute cap (e.g., 50K/min), per-metric series cap (e.g., 100K), tag value length caps.
  • When a limit trips: drop new series (never existing ones — existing series back alerts), emit a tenant.cardinality_limited event naming the metric and tag, notify the tenant.
  • Detection of the culprit: per-metric HyperLogLog of series and per-tag distinct-value estimates so we can say "tag customer_id on http.request.duration added 3.4M series."
  • Tooling: allow tenants to configure tag allowlists per metric (keep service, endpoint, drop pod_id) — aggregation at intake reduces series before storage.

Deep dive 2 — Alerting path (8–10 min).

  • Evaluators consume Kafka partitions; each owns the monitors whose series hash to its partitions (or evaluates tenant-wide queries via a pre-aggregation stage).
  • Windowed state: for avg(last_5m) > 500ms, keep per-group running sums/counts per 10 s bucket.
  • Evaluation delay: wait 30–60 s past window end before evaluating, because agents flush every 10–15 s and pipelines have jitter.
  • No-data policy per rule: alert, resolve, or keep last state, with a no-data timeout (e.g., 10 min). Default for infrastructure monitors: alert on no data.
  • Hysteresis: for: 5m and separate recovery thresholds to prevent flapping.
  • Notifier: dedup by (monitor, group), group notifications, escalation, and delivery confirmation — a page that fails to send is retried via a second provider.

Deep dive 3 — Tenant isolation and cost (6–8 min).

  • Cells: each cell is a full stack (intake shards, Kafka cluster, writers, evaluators) serving a subset of tenants; the largest tenants get dedicated cells.
  • Query cost limits: max series scanned per query (e.g., 500K), max points returned, timeouts, and per-tenant concurrent query slots.
  • Billing by the cost driver: custom metric series (per month), ingested points or hosts, and retention tier.

Phase 6: Wrap-Up (2–3 minutes)#

"To summarize: cardinality is the unit of cost and failure, so limits live at intake; two independent consumers off a durable log, with alert evaluation isolated from queries; tiered storage with rollups and sketches; cells for blast radius; and a watchdog outside our failure domain. Next I'd build per-tenant cardinality explorers so customers can fix their own tag explosions, SLO-based alerting (burn rates) to replace threshold alerts, and cost attribution by team inside each customer."


Common Timing Mistakes#

MistakeTime LostWhat to Do Instead
Designing the compression algorithm5–8 min"Gorilla-style, ~1.4 bytes/point" and move on
Choosing Cassandra vs a custom TSDB5 minSay the access pattern (append, time-range, tag filter) and one choice
Dashboard UI and query language5 minOne example query; move on
Never discussing late or missing datalevel-decidingEvaluation delay and no-data policy
Treating alerting as a query clientlevel-decidingSeparate path with its own capacity

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Monitoring is infrastructure that people trust implicitly. An engineer who sees no alert assumes nothing is wrong. That trust makes silent failures in the monitoring system uniquely dangerous — and makes this a clean test of whether a candidate designs for the failure mode users can't see.

It's also the purest cardinality problem in system design. Every other system has a natural bound on its keys; a metrics system's keys are whatever tag combinations its users invent. Candidates who don't notice that one line of instrumentation code can multiply storage cost by 1,000× haven't operated one.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"During the customer's worst outage, which of our paths is still guaranteed to work?"

The customer's worst day is our worst day: their error metrics spike, their autoscaler adds 5,000 hosts (new series), their engineers open 300 dashboards, and they create ad-hoc queries over 30 days of data. Every load type spikes at once. If the answer is "everything shares the same TSDB," the design fails when it matters. The Staff answer names the protected path (ingestion of existing series → alert evaluation → notification) and what's shed first (ad-hoc long-range queries).

🎯 Staff Move: "I want an explicit shed order. Under overload we drop ad-hoc queries first, then dashboard refresh rates, then long-range rollup reads. Ingestion of existing series, alert evaluation, and notification delivery are the last things we give up."


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1 — Operational alerting. The job is to tell a human that something is wrong within ~1–2 minutes, and never to silently stop watching. The design priorities are freshness, isolation, redundancy, and precise semantics for late and missing data. Everything else in the platform exists to support this path or is secondary to it.

Intent 2 — Exploration and dashboards. Engineers slice metrics by tags during incidents and reviews. The priorities are flexible filtering and grouping over high-cardinality tags with interactive latency. The risk is the "query of death" — a regex over a tag with 10M values across 30 days.

Intent 3 — Long-term analytics. Capacity planning, year-over-year trends, SLO reporting. The priorities are cheap storage and correct aggregates at coarse resolution. The risk is subtle wrongness: averaging percentiles, or a counter reset turning into a negative rate.

DimensionAlertingDashboardsLong-Term
Latency needSeconds–minutes, alwaysInteractive (1–3 s)Minutes OK
Data windowLast 1–60 minLast 1 h–30 d15 months
Resolution10 s10 s–1 m1 m–1 h
Worst failureSilent missSlow in incidentsWrong aggregates
Who complainsOn-call (after the outage)Engineers mid-incidentCapacity planners, execs

2.2 When NOT to Build This#

SituationUse InsteadWhy
One company, < ~10M seriesManaged Prometheus / a vendorRunning a TSDB fleet is a team; vendors are cheaper below this scale
High-cardinality per-request dataLogs or tracesPer-user or per-request dimensions belong in events, not metrics
Business analytics (revenue by SKU)A data warehouseMetrics are for operations; warehouses handle joins and exactness
Exact billing countsA transactional pipelineMetrics pipelines tolerate loss and approximation

🎯 Staff Move: "The single most useful rule I'd publish: if the tag's values are unbounded — user IDs, request IDs, URLs with parameters — it's a log or trace attribute, not a metric tag."

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Cardinality per tenantSizing and limits"Heavy-tailed: median ~50K series, top 1% ~50M"
Alert latencyEvaluation architecture"p99 < 2 min metric-to-page"
Late data toleranceEvaluation delay"Accept points up to 10 min late for storage; alerts wait 30–60 s"
Metric typesAggregation semantics"Gauges, counters (as rates), distributions (sketches)"
RetentionStorage tiers"15 days raw, 15 months at 1 m"
Isolation promiseCell design"A tenant can't degrade others' alert latency"

2.4 Precise Terminology#

TermPrecise Meaning
SeriesUnique (tenant, metric, sorted tag set); the unit of storage and cost
CardinalityNumber of distinct series (per metric, per tenant, total)
Active seriesSeries that received a point in the recent window (e.g., last 1–2 h)
ChurnRate of series creation/expiry — ephemeral pods and containers drive it
Gauge / Counter / DistributionPoint-in-time value / monotonically increasing total / set of observations summarized by a sketch
RollupAggregation of raw points into coarser windows (sum/count/min/max/sketch)
Evaluation delayTime the evaluator waits after a window closes before evaluating it
No-data policyWhat an alert does when its query returns no points
FlappingAlert toggling repeatedly around a threshold
CellAn independent full stack serving a subset of tenants

3. The Five Fault Lines#

3.1 Fault Line 1: Cardinality — Expressiveness vs Cost#

The tension: Every tag makes queries more powerful and multiplies series. Users want to slice by anything; the platform pays per series in RAM, index, and query fan-out.

StrategyWhat WorksWhat BreaksWho Pays
No limitsMaximum flexibilityOne deploy OOMs shared shards; bill surprisesEvery tenant on the shard; the platform on-call
Hard cap on active series per tenantProtects the platformTenant hits the cap mid-incident and loses new seriesThe tenant — at the worst moment
New-series rate limit + cap, drop new onlyCatches explosions in minutes; existing series (and their alerts) keep workingNew legitimate series (autoscaling) may be delayedTenants with bursty churn
Intake-side tag aggregation (allowlists)Removes unwanted tags before storage; 10–100× series reductionUsers lose dimensions they later wantTenant — by their own choice
Pay per seriesAligns cost with the driverTeams under-instrument to save moneyObservability coverage

Staff default: Rate limit on new series plus a contracted active-series cap, with existing series always accepted; per-metric cardinality caps; tag allowlists configurable per metric; and a cardinality explorer that names the offending tag within minutes. Autoscaling churn gets headroom: the new-series limit is sized to absorb a 2–3× fleet scale-out.

🎯 Staff Move: "When a tenant hits a limit, I drop new series, never existing ones. Existing series are what their alerts are watching."

3.2 Fault Line 2: Push vs Pull Ingestion#

StrategyWhat WorksWhat BreaksWho Pays
Pull (scrape)Target liveness is explicit (up); server controls rateRequires reachability into customer networks; service discovery everywhere; short-lived jobs missedSaaS: infeasible across customer firewalls
Push (agent)Works through NAT/firewalls; agent buffers during outages; batches efficientlyAbsence is ambiguous; intake must defend against floodsPlatform (intake protection); alert authors (no-data semantics)
Hybrid: agent scrapes locally, pushes remotelyLocal pull for liveness, push for transportAgent becomes critical software on every hostAgent team (release safety)

Staff default: Hybrid. The agent scrapes local Prometheus-style endpoints and integrations, aggregates, and pushes to intake every 10–15 s, with a disk buffer (e.g., 15–60 min) during network failures. The agent emits its own heartbeat series, so "host dead" vs "pipeline broken" is distinguishable: a missing heartbeat from one host = host issue; missing heartbeats from 10,000 hosts at once = our issue.

When to deviate: An internal platform inside one network can scrape centrally (Prometheus model) and gains simpler liveness.

3.3 Fault Line 3: Alert Freshness vs Completeness#

The tension: Evaluate a window as soon as it closes and you alert on partial data (false positives on ratios, false negatives on sums). Wait for stragglers and every page is slower.

StrategyWhat WorksWhat BreaksWho Pays
Evaluate at window closeFastestPartial windows: sum(errors) low, error_rate noisyOn-call (false pages), customers (missed ones)
Fixed evaluation delay (30–60 s)Covers normal agent flush + pipeline jitterAdds delay to every alertDetection time (+30–60 s)
Watermark-based (per-partition progress)Evaluates as soon as data is completeNeeds pipeline watermarks; one slow agent can hold backPlatform complexity
Evaluate early, re-evaluate on late dataFast and eventually correctAlerts that fire then un-fire erode trustOn-call trust

Staff default: Fixed evaluation delay (e.g., 60 s for 10 s data) as a per-rule default, configurable down for metrics from low-latency sources; the pipeline tracks per-partition watermarks and exposes alerting.eval_lag so the delay can be tuned from data. Ratios are computed only when both numerator and denominator windows are complete.

Diagram: 3.3 Fault Line 3: Alert Freshness vs Completeness

3.4 Fault Line 4: Shared vs Isolated Tenancy#

StrategyWhat WorksWhat BreaksWho Pays
Fully shared (all tenants on all shards)Best utilization; simple capacity planningAny tenant can degrade all; blast radius = everyoneAll tenants
Shuffle shardingEach tenant on a random subset of shards; two tenants rarely fully overlapHarder capacity reasoning; still shared hardwarePlatform complexity
Cells (independent stacks)Blast radius = one cell (~1–5% of tenants)Lower utilization; cross-cell features (org-wide views) need federationPlatform (ops overhead per cell)
Dedicated cells for largest tenantsIsolates whales, predictableExpensive; needs pricing to matchLargest tenants (priced in)

Staff default: Cells, each serving a few hundred to a few thousand tenants, with shuffle sharding inside a cell and dedicated cells for the top ~1% of tenants by series. Cell count is a lever: more cells = smaller blast radius and more operational overhead. Deploys roll out cell by cell.

3.5 Fault Line 5: Resolution vs Retention#

StrategyWhat WorksWhat BreaksWho Pays
Raw foreverExact historyStorage and scan cost grow linearly; 1-year queries scan 3M points per seriesPlatform (and customer price)
Rollup averages onlyCheapAverages of averages wrong when counts differ; percentiles lostAnyone doing SLO or capacity math
Rollups with sum/count/min/maxCorrect means, extremes, ratesStill no percentilesLatency analysis
Rollups + mergeable sketchesCorrect percentiles within ~1–2% at any resolutionSketch storage per window (~hundreds of bytes–KB)Distribution metrics cost more — priced as such

Staff default: Raw 10 s for 15 days; 1 m rollups (sum/count/min/max; sketches for distributions) for 15 months; query planner picks resolution by range (e.g., > 2 days → 1 m, > 60 days → 1 h derived on read). Counters are stored as rates or deltas with reset handling at ingest.

Diagram: 3.5 Fault Line 5: Resolution vs Retention

4. Failure Modes & Operational Reality#

4.1 Cardinality Explosion#

t=0:       Tenant deploys code adding tag request_path (unnormalized, includes IDs)
t=+2min:   tenant.new_series_rate 3K/min -> 420K/min
t=+2min:   Intake limiter: new series beyond 50K/min dropped; existing accepted
t=+3min:   Event to tenant: metric http.requests, tag request_path, ~3.1M distinct values/10min
t=+3min:   Shard memory +4% (bounded by limit) instead of OOM
t=+40min:  Tenant ships path normalization; new-series rate back to baseline
  • Detection: intake.new_series_rate{tenant}, intake.series_dropped{tenant,metric}, per-metric HLL estimates.
  • Blast radius: Without limits, every tenant on the shard; with limits, only the offending tenant's new series.
  • Mitigation: Drop new series, notify, offer tag allowlist.
  • Prevention: Default allowlists for known-dangerous tag names (user_id, request_id, trace_id); agent-side normalization of URL paths.
  • Owner: Platform owns limits; tenant owns instrumentation.

4.2 Alert Evaluator Lag During a Regional Incident#

t=0:       Cloud region partial outage; 30% of tenants' error metrics spike
t=+1min:   Ingest +40% (retries, error metrics, autoscaled hosts)
t=+2min:   Kafka consumer lag for evaluators: 5s -> 3min on 12 partitions
t=+3min:   alerting.eval_lag_p99 = 190s > 120s SLO -> page platform on-call
t=+5min:   Evaluator autoscale adds 40% capacity; partitions rebalanced
t=+9min:   Lag back under 30s; pages delivered late by up to 4 min
  • Detection: alerting.eval_lag_p99, kafka.consumer_lag{group=evaluators}, notifier.delivery_latency_p99.
  • Blast radius: Delayed pages for tenants on lagging partitions — during their incident.
  • Mitigation: Pre-provisioned evaluator headroom (≥ 2× steady state), prioritize partitions by alert density, shed dashboard queries if they share any resource.
  • Prevention: Evaluators never share compute with query serving; capacity tests at 3× ingest.
  • Owner: Alerting team.

4.3 The Silent No-Data Failure#

The most dangerous failure: a pipeline bug drops one metric family, and alerts on it go quiet because their no-data policy is "resolve."

t=0:       Agent release changes metric name nginx.requests -> nginx.http.requests
t=+10min:  Monitors on nginx.requests return no data; policy = resolve (tenant default)
t=+10min:  Dashboards show a flat line; nobody looks
t=+3days:  Customer's nginx error spike goes unpaged; outage lasts 47 min
  • Detection: Platform-side: alerting.monitors_no_data{tenant} jump after a release; per-metric-name volume drop > 90% after an agent version change.
  • Blast radius: Every monitor on the renamed metric across all tenants on that agent version.
  • Mitigation: Emit both names for 2 releases; notify tenants with affected monitors.
  • Prevention: Metric names are a public API with deprecation windows; agent releases canaried with a "metric-name diff" check; monitor creation defaults to "notify on no data" for infrastructure metrics.
  • Owner: Agent team for naming stability; alerting team for no-data defaults.

🎯 Staff Move: "Metric names are an API. Renaming one silently disables every alert built on it, so they get deprecation windows like any other interface."

4.4 Query of Death#

  • Symptom: A dashboard widget sum:*{*} by {container_id} over 30 days fans out to 40M series; query nodes OOM; dashboards for the cell slow to 30 s.
  • Detection: query.series_scanned_p99, query.memory_peak, query node OOM restarts.
  • Mitigation: Kill the query; per-query series cap (e.g., 500K) and memory cap; cost estimate before execution using the tag index.
  • Prevention: Planner rejects or auto-downsamples queries over cost thresholds; per-tenant query concurrency slots.
  • Owner: Query team.

4.5 Monitoring Is Down and Nobody Knows#

  • Symptom: Intake load balancer misconfiguration drops 100% of traffic for one cell. Dashboards show gaps; alerts go to no-data or resolve.
  • Detection: The independent watchdog (different cloud account, different pager provider) sends synthetic points every 10 s and expects a synthetic alert within 2 min. If not received, it pages the platform team directly.
  • Blast radius: One cell's tenants — and they don't know unless we tell them.
  • Mitigation: Roll back; publish a status page and notify tenants that alerts in the window may be missing.
  • Owner: Platform SRE; the watchdog is owned by a separate team or rotation.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Cardinality explosionintake.new_series_rate{tenant}One tenant (with limits)Drop new series, notifyPlatform + tenant
Evaluator lagalerting.eval_lag_p99Late pages, one cellHeadroom, rebalance, shed queriesAlerting team
Silent no-datamonitors_no_data jump, metric volume dropAll monitors on a metricDual-emit names, notifyAgent + alerting teams
Query of deathquery.series_scanned_p99, OOMsQuery nodes in a cellCost limits, killQuery team
Intake outageWatchdog synthetic alert missingOne cellRollback, status pagePlatform SRE
Kafka partition lossUnder-replicated partitions, consumer errorsData gap for affected seriesReplication factor 3, agent buffer replayStreaming platform
Notifier provider downnotifier.delivery_failures{provider}All pagesFailover to secondary providerAlerting team
Late data backlogWatermark lag per partitionWrong window resultsEvaluation delay, backfill marksIngest team

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Scale modelPoints/s and TB/dayActive series, churn, per-series memory$ per series and per team; instrumentation budget
IngestionAgents → KafkaIntake limits before shared state; agent heartbeatsAgent as fleet-critical software with staged release policy
AlertingScheduled queriesIsolated streaming path; eval delay; no-data; hysteresisAlert quality SLOs (actionable ratio, pages/shift) across the org
StorageTSDB with retentionHead + blocks, rollups with sketches, query-planner resolutionRetention contracts per data class, priced
TenancyTenant IDCells, shuffle sharding, query cost limitsPricing aligned to cost drivers; whale-tenant strategy
FailureReplicationWatchdog, shed order, cell rolloutsMonitoring as an independent failure domain for the whole company

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Cardinality first"Points are 1.4 bytes; series are kilobytes. I'll size and limit by series."
Protect existing series"On limit, drop new series only — existing ones back alerts."
Path isolation"Alert evaluators consume the log directly; dashboards can't starve them."
Late/missing data"Every rule has an evaluation delay and an explicit no-data policy."
Monitor the monitor"A watchdog in a separate account and pager verifies metric-in, page-out every minute."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Sizing only by bytesMisses the dominant cost and failure driver
Alerts as TSDB queries on a cronShares fate with dashboards under incident load
Averaging percentiles in rollupsMathematically wrong long-term data
No tenant limitsOne tenant can blind all
"Monitoring is highly available" with no external checkMonitoring can't detect its own outage

5.4 Common False Positives#

  • Deep compression knowledge ≠ monitoring design. XOR encoding is solved; cardinality isn't.
  • PromQL fluency ≠ platform thinking. Writing queries is different from bounding their cost.
  • "We use Kafka" ≠ decoupling. The point is two independent consumers with different SLOs.
  • Big ingest numbers ≠ scale thinking. Series churn from autoscaling is the scale problem.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minSaaS, 1B series, alert SLO p99 < 2 min
Entities + API3–5 minSeries identity, monitor definition
Architecture5–10 minIntake limits → log → two consumers
Deep dive: cardinality10–20 minCosts, limits, culprit detection
Deep dive: alerting20–30 minStreaming eval, delay, no-data, isolation
Deep dive: tenancy + storage30–40 minCells, query limits, rollups, sketches
Wrap-up40–45 minWatchdog, cost, evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"A customer adds user_id as a tag"Cardinality ownershipIntake limits, drop new only, notify with culprit
"p99 latency over last year?"Rollup correctnessSketches; can't derive from stored percentiles
"Make it 1-second resolution"Cost awareness10× points, same series; price it; limit to opt-in metrics
"How do you know alerts work?"Meta-monitoringWatchdog, synthetic monitors, eval lag SLO
"Add logs and traces"Scope and data modelDifferent stores; shared tags and correlation IDs

6.3 What to Deliberately Skip#

  • Compression bit layouts — name Gorilla, move on.
  • Query language grammar.
  • Dashboard rendering and UI.
  • Log and trace storage (different problem; mention correlation only).

6.4 Follow-Up Questions to Expect#

  1. "How do you handle counters that reset when a process restarts?"
  2. "How would you evaluate 10 million monitors every minute?"
  3. "What does the query planner do for a 90-day range?"
  4. "How do you rebalance a hot Kafka partition that holds one huge tenant?"
  5. "How do you bill fairly — by host, by series, or by points?"
  6. "How would you implement SLO burn-rate alerts?"
  7. "A customer says an alert didn't fire. How do you prove what happened?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design Datadog."

Staff Answer

"Datadog covers metrics, logs, traces, and more — I'll scope to metrics and alerting, which is the core. Three intents compete: operational alerting, interactive exploration, and long-term analytics. They want different storage and different SLOs, and alerting is the one that can't fail silently.

Assumptions: multi-tenant SaaS, ~20K tenants with heavy-tailed sizes, ~1B active series total, 10 s resolution, ~10M points/s, and a metric-to-page p99 under 2 minutes.

The design has three commitments: cardinality limits at intake because series are the unit of cost and failure; two independent consumers off a durable log so alert evaluation never competes with dashboards; and cells so a single tenant can't take down everyone's monitoring. I'll walk the architecture quickly and then go deep on cardinality and the alert path."

Why this is L6:

  • Scopes a huge product to its core and says why
  • Names intents and commits to alerting as the protected path
  • States three architectural commitments up front

What L7 adds:

  • Frames the business constraint: gross margin depends on cost per series, so pricing and architecture are one decision
❌ Common L5 Trap

"Agents send metrics to an API, which writes them into a time-series database like InfluxDB. Dashboards query it, and an alert service runs queries every minute."

Why this misses: A reasonable single-tenant design. The interviewer asks "one customer adds a user_id tag" and "during an incident everyone opens dashboards — do alerts slow down?" — and both break it.


Drill 2: Size It in Series#

Prompt: "How much memory does the ingest tier need?"

Staff Answer

"By series, not points. Assume 1B active series. Each active series in the ingest head costs roughly: an open compressed chunk (~120 points over 20 min at ~1.4 bytes ≈ 170 bytes, plus buffers), series metadata and labels (~200–500 bytes after string interning), and inverted-index postings (~tens of bytes per tag). Call it ~2–4 KB with overhead. 1B × 3 KB ≈ 3 TB of RAM, times replication factor 2 for the head ≈ 6 TB, across maybe 300–600 ingest nodes at 16–32 GB effective each.

Points matter for CPU and network: 1B series / 10 s = 100M points/s if every series reports every interval — in practice ~10–20M/s because many series are sparse.

Churn matters too: containers that live 10 minutes create series that stay in the head for 2 hours. With heavy autoscaling, active series can be 3–5× the 'steady' number."

Why this is L6:

  • Builds memory from per-series cost
  • Separates series (memory) from points (CPU/network)
  • Names churn as a multiplier

What L7 adds:

  • Converts it to $: ~6 TB of RAM plus overhead → cost per million series per month, which becomes the pricing floor
❌ Common L5 Trap

"10M points/s × 16 bytes = 160 MB/s, ~14 TB/day; store 2 hours in memory ≈ 1.1 TB."

Why this misses: Uses raw point size and ignores per-series overhead — which dominates by 10× or more.


Drill 3: The Tag Explosion#

Prompt: "A customer adds customer_id (2M values) as a tag on their busiest metric."

Staff Answer

"Series for that metric go from maybe 5K to up to 10B combinations; in practice millions within minutes.

Intake catches it: the tenant's new-series rate limit (say 50K/min) trips within a minute. We accept points for all existing series — their alerts keep working — and drop new ones beyond the limit. We emit an event and email: 'metric api.latency, tag customer_id, ~2M distinct values; 1.4M series dropped in 10 min.' The per-tag estimate comes from HyperLogLog sketches maintained at intake.

The tenant's options: drop the tag via a per-metric allowlist, move customer_id to traces/logs, or keep a small set (top 100 customers) as a tag and bucket the rest as 'other.' If they need per-customer SLOs for a few thousand enterprise customers, we can price a higher limit.

Shared shards see only a bounded increase — no neighbor impact."

Why this is L6:

  • Protects existing series and alerts
  • Tells the tenant exactly which tag caused it
  • Offers remediations, including the right data type

What L7 adds:

  • Makes cardinality visible before deploy: CI lint on instrumentation diffs that estimates series impact
❌ Common L5 Trap

"Scale out the ingest tier to handle more series."

Why this misses: Unbounded cardinality isn't a capacity problem you can scale out of — it's 1,000× growth from one line of code. It also puts the bill on everyone else.


Drill 4: The Alert Path#

Prompt: "Design the alert evaluation service for 5 million monitors."

Staff Answer

"Two classes of monitors:

  1. Simple per-series or small-group rules (~90%): avg(last_5m):cpu{host:*} > 90. Evaluated in streaming evaluators that consume Kafka partitions. Each evaluator holds rolling window state for the series its monitors reference, updated as points arrive. At the evaluation tick (every 30–60 s, plus a 60 s evaluation delay), it computes per-group results and transitions state.
  2. Complex rules (~10%): cross-metric formulas, large group-bys, anomaly detection. Evaluated by a scheduled evaluator that queries a dedicated recent-data tier (a replica of the in-memory heads), not the dashboard query service.

Monitor assignment: monitors are sharded by tenant and series hash so an evaluator only reads the partitions it needs. State (OK/ALERT/NO_DATA per group) is checkpointed every tick to a replicated store, so a failed evaluator's successor resumes without false recoveries.

Semantics per monitor: evaluation delay, no-data policy and timeout, for duration, recovery threshold. Output to the notifier, which dedups by (monitor, group), batches, and delivers with retry across two providers.

SLO: alerting.eval_lag_p99 < 60s, notifier delivery p99 < 30 s."

Why this is L6:

  • Splits by rule complexity with the right engine for each
  • Isolates evaluation from dashboard queries
  • State checkpointing avoids false recover/re-alert on failover

What L7 adds:

  • Alert quality metrics per tenant team — actionable ratio, flapping monitors — surfaced as product features to reduce noise
❌ Common L5 Trap

"A scheduler runs each monitor's query against the query API every minute."

Why this misses: 5M queries/min on the same query tier as dashboards; during incidents, both spike and alerts lag. Also no state continuity on failover.


Drill 5: Late and Missing Data#

Prompt: "An agent's network is flaky; its points arrive 3 minutes late. What do the alerts do?"

Staff Answer

"With a 60 s evaluation delay, a 3-minute-late point misses its window. Three cases:

  • Threshold on the value (cpu > 90): the window had no points for that host → no-data for that group. If the policy is 'notify on no data after 10 min,' nothing fires yet; if the gap persists, it fires as NO_DATA, not as ALERT — that distinction tells the on-call it's a visibility problem.
  • Sum/count (sum(errors) > 100): the partial window undercounts → potential false negative. For count-type monitors I'd use longer windows (5 m) so a 3-min delay is partially absorbed, or evaluate at watermark.
  • Ratio (errors / requests): only compute when both are present for the window; otherwise skip rather than divide partial by partial.

Storage still accepts the late points (up to a late-arrival limit, e.g., 10 min for raw, then rejected) so dashboards become correct after the fact.

The agent helps: it buffers to disk and flushes in order, and emits a heartbeat so we can tell 'host silent' from 'host late.'"

Why this is L6:

  • Differentiates by aggregation type
  • NO_DATA as a separate state with its own meaning
  • Storage vs alerting late-data policies differ, and says so

What L7 adds:

  • Publishes these semantics in docs as a contract, so tenants choose policies knowingly — and default policies are reviewed like APIs
❌ Common L5 Trap

"We wait until all data arrives before evaluating."

Why this misses: "All data" is unknowable in a push system; waiting unboundedly means alerts never fire for the host that's actually down.


Drill 6: Multi-Tenant Noisy Neighbor#

Prompt: "Your largest customer is 8% of all series and just tripled."

Staff Answer

"8% of series in a shared cell means that customer's growth sets everyone's capacity plan. I'd move them to a dedicated cell — its own intake shards, Kafka cluster, ingest writers, and evaluators. Migration: dual-write from intake to old and new cells, backfill 15 days of raw and 15 months of rollups from blocks (object storage makes this a metadata copy plus rewrite), switch queries per time range, then cut over alerts with state transfer, then stop dual-write.

Inside the shared cells, shuffle sharding keeps remaining large tenants from overlapping fully.

Commercially: their contract should reflect dedicated infrastructure. And their 3× growth — was it intentional? A cardinality review with them often finds 30–50% of series are never queried; we can show them 'series with zero queries in 30 days.'"

Why this is L6:

  • Uses cells to bound blast radius
  • Concrete migration with alert state transfer
  • Brings unused-series data to the customer

What L7 adds:

  • Sets a policy: any tenant above X% of a cell automatically triggers dedicated-cell planning and a pricing review
❌ Common L5 Trap

"Add more nodes to the cluster."

Why this misses: Keeps one tenant's growth as everyone's risk; doesn't change blast radius.


Drill 7: Percentiles Over a Year#

Prompt: "Show me p99 latency per month for the last 12 months."

Staff Answer

"Only possible if the metric was stored as a distribution. If agents sent only a precomputed p99 per host, there's no correct way to get a service-wide or monthly p99 — you can't average or max percentiles into a percentile.

For distributions: agents build a mergeable sketch (DDSketch with ~1% relative error) per flush; we store sketches in raw data and merge them into 1 m rollup sketches. A monthly p99 merges ~43K one-minute sketches per series group, then reads the quantile — error stays ~1%.

Cost: a sketch is hundreds of bytes to a few KB versus ~1.4 bytes per gauge point, so distributions are priced per series separately. I'd default latency metrics from our integrations to distributions and let tenants opt others in."

Why this is L6:

  • Knows percentiles don't aggregate
  • Names a mergeable sketch with its error bound
  • Prices the cost difference

What L7 adds:

  • Makes 'latency is a distribution' the org instrumentation standard so SLO reporting is correct by construction
❌ Common L5 Trap

"Average the daily p99 values for each month."

Why this misses: The average of p99s isn't a p99 of anything; it systematically understates tail latency when traffic varies.


Drill 8: Build vs Buy#

Prompt: "We're a 400-engineer company paying a vendor $3M/year. Should we build our own?"

Staff Answer

"First, why is it $3M? Usually cardinality: custom metrics with container or pod tags, and unused series. A cardinality audit and tag allowlists often cut 30–40% without changing vendors.

Build means running Prometheus-compatible storage (Mimir, Thanos, VictoriaMetrics, or M3) plus Grafana, alerting, agents, and on-call. At our scale that's 4–6 engineers ($1.2–1.8M/yr loaded) plus infrastructure (~$300–600K/yr). Savings exist but are smaller than the headline, and we'd own the monitoring outage risk ourselves.

My recommendation: cut the bill with governance first; revisit build when spend passes ~$6–10M or when we need data residency the vendor can't provide. If we build, use open-source storage — don't write a TSDB."

Why this is L6:

  • Finds the cost driver before switching
  • Realistic TCO including people
  • Names the threshold that changes the answer

What L7 adds:

  • Keeps instrumentation vendor-neutral (OpenTelemetry) so the build/buy decision stays a two-way door
❌ Common L5 Trap

"Prometheus is free; we'd save $3M."

Why this misses: Ignores people, infrastructure, and the risk of owning monitoring reliability.


Drill 9: Rolling Out a New Limit Without Breaking Alerts#

Prompt: "We need to enforce a 1M-series cap on all tenants. 3% of tenants are above it."

Staff Answer

"Never cut existing series on day one — that silently breaks alerts.

  1. Shadow (2 weeks): compute who'd be limited and which metrics; for each over-limit tenant, which monitors reference series that would be dropped.
  2. Communicate: per-tenant report with top metrics by series and 'series never queried in 30 days.' Offer allowlist tooling.
  3. Warn (2–4 weeks): tenants above cap see banners and events; limit applies to new series only above 120% of cap.
  4. Enforce new-series limit at cap. Existing series grandfathered until they go inactive.
  5. Exceptions: contracted higher limits for tenants who pay.

Metrics: tenants_over_cap, series_dropped{tenant}, and — the safety check — monitors_referencing_dropped_series must be zero for enforcement to proceed per tenant."

Why this is L6:

  • Shadow → warn → enforce, never breaking existing series
  • Safety metric tied to alerts
  • Commercial path for exceptions

What L7 adds:

  • Coordinates with sales and support before rollout; limit changes are a product launch, not an infra change
❌ Common L5 Trap

"Enforce the cap; tenants above it will drop data until they reduce series."

Why this misses: Which data drops is arbitrary; alerts break silently for the largest (most valuable) customers.


Drill 10: Multi-Region#

Prompt: "Customers in the EU need data to stay in the EU, and a US region outage shouldn't stop EU alerts."

Staff Answer

"Independent regional stacks — each region is a set of cells with its own intake, log, storage, evaluators, and notifier. A tenant has a home region; agents send there. Nothing on the ingest or alert path crosses regions.

Global pieces are minimal: account/tenant directory and auth (cached regionally, so a global outage doesn't stop ingestion), and the web UI routing to the home region.

Cross-region views for global customers: a federated query layer that fans out to each region and merges — the Monarch pattern — accepting that a region outage makes global views partial, with a banner saying which region is missing.

For the watchdog: each region's watchdog runs in a different region, so a regional outage is detected from outside."

Why this is L6:

  • Keeps alert path region-local
  • Minimal, cached global state
  • Watchdog placed outside the monitored region

What L7 adds:

  • Residency as a contractual product tier with a region-launch playbook reused for each new jurisdiction
❌ Common L5 Trap

"Replicate all data to both regions for availability."

Why this misses: Violates residency and doubles cost; a synchronous dependency between regions couples their failures.


8. Deep Dive Scenarios#

Deep Dive 1: Peak Traffic — A Major Cloud Outage Hits 30% of Tenants#

Context: A large cloud provider has a regional outage. 30% of your tenants are affected at once. Ingest rises 60% (error metrics, retries, autoscaled replacement hosts creating new series), dashboard queries rise 12×, and alert notifications rise 40×. Your status page gets a complaint that alerts arrived 6 minutes late.

Questions to Surface First:

  • Which path is lagging: intake, Kafka consumers for evaluators, evaluation compute, or notification delivery?
  • Are dashboards and evaluators sharing any resource?
  • Is new-series churn from replacement hosts hitting limits and dropping legitimate series?

Typical L5 Approach: Autoscales everything. Autoscaling takes 5–10 minutes to catch up; meanwhile, alerts lag and dashboards time out.

Staff Approach: Applies the shed order: throttles ad-hoc and long-range queries, lowers dashboard auto-refresh to 60 s, prioritizes evaluator consumption over storage writers on the log (storage can catch up later from Kafka), and raises new-series limits temporarily for affected tenants whose churn is from autoscaling. Checks notifier provider rate limits — 40× notifications may hit a paging provider's API limit, so batching and failover matter.

Principal Approach: Treats correlated customer incidents as the design point for capacity, not an anomaly. Sets headroom targets based on the worst historical correlated event (e.g., ≥ 3× ingest and ≥ 10× queries for alerting-critical paths), runs a yearly game day simulating a major cloud-region outage across all cells, and negotiates higher burst limits with paging providers in contract.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Check alerting.eval_lag_p99, notifier.delivery_latency_p99, kafka.consumer_lag{group} per cell
TriageEvaluators lagging on hot partitions; notifier hitting provider 429s
Quick fixShed queries; raise evaluator priority; notifier batching and secondary provider
GuardrailsWatch dropped series for affected tenants; temporary limit raises auto-expire in 24 h
Post-mortemHeadroom policy; provider burst contracts; game day
t=0:       Cloud region degraded; affected tenants' error metrics spike
t=+2min:   Queries 12x; query tier CPU 90%
t=+3min:   Notifications 40x; paging provider returns 429 on 20% of sends
t=+4min:   eval_lag_p99 150s on 18 partitions
t=+6min:   Shed order applied; evaluator lag falls to 40s
t=+8min:   Secondary paging provider engaged; delivery p99 back under 30s

Metrics to Watch: alerting.eval_lag_p99, notifier.delivery_failures{provider}, query.shed_total, intake.series_dropped{tenant}

Organizational Follow-up: Incident comms: a customer-facing note listing which alerts may have been delayed and by how much.

Ownership Question: "Who decides to shed dashboard queries during an incident?" Staff answer: The platform on-call, automatically via load-shedding policy with a manual override. It's pre-approved because the alternative is late pages.

Key Takeaway: "Your peak is your customers' correlated worst day. Protect the page, shed the dashboard."

What clears the Staff bar:

  • Applies a pre-defined shed order
  • Checks the notifier and external providers, not just internal compute
  • Distinguishes autoscaling churn from abusive cardinality

Deep Dive 2: The Silent Failure — Alerts Stopped for One Metric Type#

Context: A customer reports an outage that should have paged. Their monitor on rate(errors) stayed OK for 50 minutes while errors spiked. Your evaluator logs show no errors. The monitor evaluated every minute.

Questions to Surface First:

  • What did the evaluator actually compute each minute? Can we replay it?
  • Did the underlying data arrive and get stored?
  • Was anything deployed recently in the counter-handling path?

Typical L5 Approach: Checks the evaluator is healthy (it is), checks the data is stored (it is), and concludes the customer's threshold was wrong.

Staff Approach: Replays evaluation from the Kafka log for that window with the same code version and finds the rate computation returned 0: a deploy changed counter-reset detection, so every time the customer's process restarted (it was crash-looping), the new counter value was treated as a reset and the delta dropped. The crash loop was the incident — so the metric designed to catch it was suppressed by it. Fix the logic; add golden-test fixtures for counter resets; add an "evaluation audit" that samples monitors and recomputes them independently.

Principal Approach: Establishes that alert correctness is testable and must be tested continuously: synthetic monitors covering every aggregation type (gauge, counter with resets, distribution, ratio, no-data) run in production against synthetic series with known answers, and any mismatch pages the alerting team. Makes "alert evaluation replay" a customer-facing feature so disputes are resolved with evidence.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateReplay the evaluation from the log; compare computed vs expected
TriageIsolate to counter-reset handling; identify the deploy
Quick fixRoll back; re-evaluate affected monitors for the incident window
GuardrailsGolden fixtures in CI; production synthetic monitors per aggregation type
Post-mortemIdentify all tenants whose counter monitors were affected; notify them

Metrics to Watch: alerting.synthetic_mismatch_total{agg_type}, alerting.audit_disagreement_ratio, ingest.counter_resets_detected (a jump after a deploy is a red flag)

Organizational Follow-up: Alerting semantic changes get a separate review class with golden-test sign-off.

Ownership Question: "Who tells affected customers their alerts may not have fired?" Staff answer: We do, proactively, with the window and affected monitors — support owns the communication, alerting team owns the list. Trust in monitoring depends on us disclosing our misses.

Key Takeaway: "An evaluator that computes the wrong answer looks exactly like a healthy one. Test alert semantics continuously with known answers."

What clears the Staff bar:

  • Replays from the log rather than trusting logs of the evaluator
  • Sees the irony — the incident suppressed its own signal — and designs against it
  • Proactive customer disclosure

Deep Dive 3: Large-Customer Onboarding — 200M Series#

Context: An enterprise wants to migrate from a self-hosted Prometheus fleet: 200M active series, 1.5M series/hour churn from Kubernetes, 40K alert rules, and 5 years of data they want imported.

Questions to Surface First:

  • What fraction of series are queried or alerted on?
  • Which labels drive churn (pod, container_id, instance)?
  • Do they really need 5 years of raw data, or rollups?

Typical L5 Approach: Provisions capacity for 200M series in the shared cells and bulk-imports the history.

Staff Approach: Runs a cardinality analysis on their current data first: often ~40–60% of series are never queried in 30 days, and pod-level labels drive most churn. Proposes intake-side aggregation dropping pod for most metrics (keeping deployment and namespace), reducing to ~70–90M series. Places them in a dedicated cell sized for 2× that. Imports history as 1 h rollups for > 15 months, 1 m for the last 15 months. Translates alert rules with automated tooling and runs both systems in parallel for 2–4 weeks comparing alert firings.

Principal Approach: Uses this as the template for enterprise migrations: a migration kit (rule translator, parallel-run comparator, cardinality analyzer), a pricing model that rewards series reduction, and an account plan with a named platform engineer through cutover.

Staff Approach — Full Reasoning
PhaseWhat to Do
AnalysisSeries by metric and label; queried/unqueried; churn sources
ReductionIntake aggregation rules agreed with customer
CapacityDedicated cell at 2× reduced series; churn headroom
HistoryRollup import by age tier
Alert parityTranslated rules; parallel run; diff of firings must be explained before cutover

Metrics to Watch: tenant.active_series, tenant.series_churn_per_hour, migration.alert_parity_mismatches, alerting.eval_lag_p99{cell}

Organizational Follow-up: Solutions engineering and platform co-own enterprise migrations with a checklist.

Ownership Question: "Who signs off on cutover?" Staff answer: The customer's on-call lead, after the parallel-run report shows every alert-firing difference explained. Monitoring migrations fail on trust, not on data.

Key Takeaway: "Onboard the cardinality, not just the capacity. Half of most series are never read."

What clears the Staff bar:

  • Reduces before provisioning
  • Parallel-run alert parity as the cutover gate
  • Rollup-based history import

Deep Dive 4: Post-Mortem — The Monitoring Platform Was Down for 38 Minutes#

Context: A configuration push to the intake fleet contained a malformed tenant-limit file. Intake nodes rejected all traffic for three cells. Customers' dashboards showed gaps; many monitors went to NO_DATA or resolved. Your own internal alerting — which runs on the same platform — also went silent, so the team learned from a customer tweet after 22 minutes.

Questions to Surface First:

  • Why did a config push reach three cells at once?
  • Why didn't intake fail static on bad config?
  • Why was our detection dependent on the thing that failed?

Typical L5 Approach: Adds validation to the config file and rolls back.

Staff Approach: Three fixes. Config pushes roll out cell by cell with automated health gates (ingest rate per cell must stay within 10% of baseline for 10 minutes before the next cell). Intake fails static: a config that fails to parse keeps the last-known-good. And an external watchdog — separate cloud account, separate paging provider — sends synthetic metrics end-to-end and expects a synthetic alert every minute.

Principal Approach: Codifies "monitoring must not depend on what it monitors" as a platform rule: the platform's own alerting runs on an independent minimal stack; every global config channel is audited for blast radius; and customers get a documented status API so their own tooling can detect our outage.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateRoll back config; confirm ingest recovers per cell
TriageIdentify which tenants' monitors resolved or went NO_DATA during the window
Quick fixProactive notice to affected tenants with the outage window
GuardrailsCell-by-cell config rollout with health gates; fail-static on parse errors
Post-mortemIndependent watchdog; independent internal alerting stack
t=0:       Config push to all intake nodes in 3 cells
t=+30s:    Intake rejects all payloads (limit file parse error)
t=+1min:   Agents buffer to disk; ingest rate for 3 cells = 0
t=+2min:   Our internal monitors (on the same cells) go NO_DATA, policy resolve
t=+22min:  Customer tweet; on-call investigates
t=+31min:  Config rolled back
t=+38min:  Agent buffers drained; data gaps partially backfilled

Metrics to Watch: intake.accepted_points{cell} vs baseline, watchdog.synthetic_alert_latency, config.rollout_stage, agent.buffer_bytes

Organizational Follow-up: Config changes are deploys: same staged rollout, same review.

Ownership Question: "Who owns the watchdog?" Staff answer: A different rotation from the platform on-call, on different infrastructure — if the same team owns both on the same stack, it's not independent.

Key Takeaway: "The monitoring system can't detect its own death. Something outside it must."

What clears the Staff bar:

  • Treats config as code with staged rollout
  • Fail-static for control-plane inputs
  • Independent detection path

Deep Dive 5: Multi-Region Expansion — Launching an EU Region#

Context: Legal requires EU customer telemetry to stay in the EU. Some global customers want one dashboard across US and EU. The EU region launches with 5% of current capacity.

Questions to Surface First:

  • Which data counts as customer telemetry — tags can contain personal data (hostnames, user emails in tags)?
  • Is cross-region query acceptable if data doesn't leave the region at rest?
  • How do we handle a tenant with hosts in both regions?

Typical L5 Approach: Deploys a copy of the stack in the EU and routes EU tenants there.

Staff Approach: Independent regional stacks with tenant home regions; agents configured with the regional intake. For tenants spanning regions, each region stores its own hosts' data; a federated query layer merges results at query time without persisting cross-region copies. Alert rules run in the region of their data; cross-region rules evaluate on the federation layer with degraded semantics (partial data flagged) — and the default recommendation is region-local monitors. The EU watchdog runs from outside the EU region.

Principal Approach: Treats residency as a repeatable region-launch capability: a checklist, infrastructure-as-code for a full regional stack, legal review of which fields are telemetry vs account data, and a pricing tier for residency. Plans the capacity ramp so a small new region still has cells, headroom, and on-call coverage.

Staff Approach — Full Reasoning
PhaseWhat to Do
ScopeClassify data: telemetry (regional), account metadata (global, minimized)
BuildFull regional stack from IaC; minimum 2 cells even at low volume
FederationQuery-time merge only; no cross-region persistence
AlertingRegion-local by default; cross-region rules flagged as degraded on partial data
LaunchPilot 20 tenants; compare alert latency and query latency to US
Diagram: Deep Dive 5: Multi-Region Expansion — Launching an EU Region

Metrics to Watch: alerting.eval_lag_p99{region}, federation.partial_results_ratio, residency.cross_region_bytes_at_rest (must be 0)

Organizational Follow-up: Legal owns data classification; platform owns enforcement; each region has local on-call coverage.

Ownership Question: "Who approves a new cross-region feature?" Staff answer: Platform architecture with legal review — any feature that moves telemetry across regions is a residency decision, not a product one.

Key Takeaway: "Regions are independent failure and legal domains. Federate reads; never federate alerts by default."

What clears the Staff bar:

  • Region-local alerting
  • Query-time federation without persisted copies
  • Cross-watchdogs between regions

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why active series — not points or bytes — are the unit of cost, memory, and failure
  • Design intake-side cardinality limits that protect existing series and name the culprit tag
  • Separate alert evaluation from query serving with two consumers off a durable log
  • Specify evaluation delay, no-data policy, and hysteresis for alert rules
  • Design tiered retention with rollups that stay correct (sum/count/min/max, sketches)
  • Isolate tenants with cells, shuffle sharding, and query cost limits
  • Detect a monitoring outage from outside the monitoring system
  • Price the platform per series at three scales and name its one-way doors

The Bar for This Question#

Mid-level (L4): Agents write to a TSDB; dashboards query it; an alert job polls. Sizing by data volume. Single-tenant thinking.

Senior (L5): Adds Kafka, sharded storage, retention, replication, and a separate alert service that queries storage. Solid — but sizes by points, has no cardinality limits, shares query capacity between dashboards and alerts, averages percentiles in rollups, and relies on the platform to detect its own outages.

Staff+ (L6): Sizes and limits by series; isolates the alert path; defines late and no-data semantics; stores sketches for percentiles; uses cells and query cost limits; applies a shed order under load; runs an independent watchdog; names owners for agent naming stability, alert semantics, and tenant instrumentation. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Most Metrics Are Never Read"#

EvidenceImplication
Series-level query logs typically show a large share of series unqueried for 30+ daysThe bill is mostly write-only data
Default integrations emit hundreds of metrics per hostDefaults drive cardinality, not intent
Pod/container labels multiply series with churnMost of that dimensionality is never used in queries

The Staff position: Show "unqueried series" per tenant and per team; default to aggregation that drops ephemeral labels unless requested.

Why this matters in interviews: It turns "scale the TSDB" into "reduce what we store" — a more senior move.

10.2 "Per-User Metrics Are a Category Error"#

EvidenceImplication
User IDs have millions of valuesSeries explode 10⁶×
Per-user questions are per-event questionsLogs and traces answer them with no cardinality penalty
Metrics are for aggregates over timeHigh-cardinality dimensions belong in event stores

The Staff position: Unbounded dimensions go to traces/logs; metrics carry bounded dimensions only.

Why this matters in interviews: It shows you choose the data type by question, not by habit.

10.3 "Threshold Alerts Are Mostly Noise — Alert on SLO Burn"#

EvidenceImplication
CPU/memory thresholds fire without user impactLow actionable-page ratio
Burn-rate alerts (e.g., Google SRE workbook's multi-window approach) tie paging to error budgetPages track user pain
On-call fatigue causes real alerts to be ignoredNoise has a reliability cost

The Staff position: Page on SLO burn rate; route resource thresholds to tickets.

Why this matters in interviews: It connects platform design to on-call health.

10.4 "Monitoring Should Be Less Available Than Alerting"#

EvidenceImplication
Dashboards and alerts share load spikesTreating them equally sacrifices the critical one
A slow dashboard is an annoyance; a missed page is an outageDifferent SLOs are rational

The Staff position: Separate SLOs: alerting 99.99% with p99 < 2 min; dashboards 99.9% with a shed policy.

Why this matters in interviews: It shows you design explicit degradation rather than uniform availability.

10.5 "Build Your Own Only at Planet Scale"#

EvidenceImplication
Google (Monarch), Facebook (Gorilla), Uber (M3), and Netflix (Atlas) built at extreme scaleTheir scale justified dedicated teams
Open-source (Prometheus, Mimir, Thanos, VictoriaMetrics) covers most needsCustom TSDBs are rarely differentiating
Monitoring outages you own are yours to explainBuying transfers operational risk

The Staff position: Buy or run open source; build the governance layer (limits, attribution, standards) yourself.

Why this matters in interviews: Candidates who propose writing a TSDB for a mid-size company signal poor TCO judgment.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer builds a monitoring platform that stays up during incidents. A Principal engineer notices that observability spend is growing faster than infrastructure spend — often 10–30% of the infra bill — because every team instruments independently, cardinality is invisible to the engineer who adds a tag, and alert noise is nobody's metric. The technical platform is solvable; the org's instrumentation economy isn't, until someone makes cost and alert quality visible and owned. L7 decides what to standardize (instrumentation libraries, naming, tag rules, SLO-based paging), what to leave to teams (which metrics matter), and whether to build, buy, or run hybrid.

The Org-Level Fault Line#

Central observability platform with enforced standards vs team-chosen tools and instrumentation.

OptionWhat WorksWhat BreaksWho Pays
Team autonomyTeams pick what fits4 vendors, 3 agents, incompatible naming; cross-team incidents need 4 tools; cost invisibleIncident responders; finance
Central platform, mandatory everythingConsistency, negotiated pricingPlatform becomes a gate; teams wait for integrationsProduct velocity
Paved road: standard SDK + limits + chargeback, opt-out with justificationDefault consistency, visible cost, escape hatchNeeds a platform team and budget attributionPlatform team (~4–8 engineers at mid-size)

The L7 default: Paved road built on OpenTelemetry instrumentation (vendor-neutral), a single agent, naming and tag standards enforced in CI, per-team series budgets with chargeback, and SLO burn-rate paging as the default alert type.

🧭 Principal Move: "The cheapest series is the one never emitted. I'd put the cardinality estimate in the pull request that adds the tag, with the monthly cost next to it."

Cost Model#

Assumptions: fully loaded cost per million active series ~$1.5–3K/month (RAM-heavy ingest heads, index, evaluators, object storage for blocks and rollups, Kafka); 10 s resolution; 15-day raw and 15-month rollup retention; engineer ~$25K/month loaded.

ScaleActive SeriesInfra $/monthHeadcountOn-call Load
Small (one company, self-hosted OSS)~10M~$15–30K3–5 (~$100K)1 rotation; monitoring pages ~weekly
Medium (large company or small SaaS)~500M~$0.75–1.5M30–50Rotations per component (intake, storage, alerting)
Large (multi-tenant SaaS vendor)~10B~$15–30M300+ across platform, agents, query, alerting, SREFollow-the-sun, cell-based rotations

Comparison worth saying: a mid-size company paying a vendor for ~10M custom series at list prices typically pays more than the self-hosted infra above — but the self-hosted column's headcount usually closes the gap until spend reaches several million dollars a year.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

Triggers: Year 1 when a single tenant causes a cross-tenant incident; Year 2 when alert noise or cost reaches leadership; Year 3 when a residency-driven customer or region launch requires it.

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal Cost
Metric naming and tag conventionsOne-way-ishRenames break every dashboard and alert built on them
Alert semantics (no-data defaults, eval delay)One-way-ishChanging defaults silently changes behavior of millions of monitors
Instrumentation SDK / agent protocolOne-way-ishRe-instrumenting hundreds of services; mitigated by OpenTelemetry
Pricing unit (series vs hosts vs points)One-wayCustomers architect around it; changes cause churn
Rollup aggregates stored (no sketches)One-wayPercentile history can't be reconstructed
Storage engineTwo-way (expensive)Dual-write migration with query switching by time range
Cell sizes and countTwo-wayTenant moves between cells
Retention durationsTwo-way (shortening is one-way)Data deleted is gone

The Standard I'd Write#

RFC: Metrics & Alerting Standard v1

Scope: All services emitting operational metrics or defining pages.

MUST:

  • Instrument with the platform OpenTelemetry SDK; metric names follow <domain>.<component>.<measure> with units in the name.
  • Tags have bounded value sets; user_id, request_id, trace_id, raw URLs, and email addresses are prohibited as metric tags (enforced by CI lint and intake).
  • Latency is recorded as a distribution, not as precomputed percentiles.
  • Every paging monitor has an owning team, a runbook link, an explicit no-data policy, and — for user-facing services — is an SLO burn-rate alert.
  • Teams stay within their series budget or file for an increase with cost acknowledgment.

SHOULD:

  • Resource-threshold alerts route to tickets, not pages.
  • Review unqueried series quarterly and remove them.

Exceptions: Filed with the observability platform team; approved by platform lead and the requesting team's director; expire in 6 months.

Success metrics: Actionable-page ratio ≥ 70%; pages per on-call shift ≤ 2 median; observability cost growth ≤ infrastructure cost growth; 100% of paging monitors owned.

What I'd Tell the VP#

Our monitoring bill is growing faster than our infrastructure because engineers can add expensive measurements with one line of code and never see the cost. At the same time, our on-call engineers get paged for things that don't affect customers, which makes them slower to respond when something does. I'm proposing a standard toolkit, a budget per team with visible costs, and paging based on customer-facing reliability targets instead of server thresholds. Expect a 25–40% reduction in observability spend within two quarters and fewer, more meaningful pages. It needs about four engineers for two quarters, and it keeps us vendor-neutral so we can renegotiate or switch providers later.

Principal Interview Signals#

SignalWhat It Sounds Like
Instrumentation economics"Cardinality cost should be visible in the PR that creates it."
Alert quality as an org SLO"I'd track actionable-page ratio and pages per shift for every team."
One-way doors"Metric names and no-data defaults are APIs — changing them silently changes millions of monitors."
Vendor strategy"OpenTelemetry keeps build-versus-buy a two-way door."
Independent failure domain"The company's monitoring must not share fate with the company's production — or with a single vendor's region."

Staff answers that L7 interviewers find insufficient:

  • "Per-tenant limits protect the platform" — correct, but doesn't change the behavior of the engineers generating cardinality.
  • "Alert evaluation is isolated and fast" — but nothing about whether the alerts are worth sending.
  • "We'll use Prometheus and Thanos" — a stack choice, not a strategy for cost, standards, and ownership.

Appendices

Appendix A: Storage Mechanics in Depth#

A.1 Gorilla-Style Compression#

Timestamps: store t0, then delta d1 = t1 - t0, then delta-of-delta dd = (tn - tn-1) - (tn-1 - tn-2)
  dd == 0            -> 1 bit '0'         (regular 10s intervals: almost always)
  dd in [-63, 64]    -> '10' + 7 bits
  larger             -> wider buckets
Values: XOR with previous value
  xor == 0           -> 1 bit '0'         (unchanged gauge)
  else               -> leading/trailing zero counts + meaningful bits
Average ~1.37 bytes/point on production data (Facebook, 2015)

A.2 Head, Blocks, and Compaction#

  • Head: per-series open chunk (~120 points), write-ahead log for crash recovery, in-memory inverted index.
  • Block cut: every 2 h, flush chunks and the index as an immutable block to object storage.
  • Compaction: merge 2 h blocks into larger blocks (e.g., 24 h); produce 1 m rollups; drop expired raw data at 15 days.
  • Query: planner selects head for the last 2 h, raw blocks for ≤ 15 days, rollups beyond.

A.3 Inverted Index#

postings["service=checkout"] = sorted [series_id...]
postings["region=us-east"]   = sorted [series_id...]
query {service:checkout, region:us-east} = intersect(postings...)

Posting lists are compressed (delta + varint or roaring bitmaps). Regex/wildcard matchers expand against the tag-value dictionary first — the expansion size is a query-cost input.

A.4 Mergeable Sketches#

DDSketch buckets values logarithmically so relative error is bounded (e.g., 1%); two sketches merge by adding bucket counts. That makes per-host, per-10 s sketches mergeable into per-service, per-month sketches without error growth — the property percentile rollups need.

Appendix B: Series Identity and Data Model#

ElementConstructionNotes
series_id64/128-bit hash of (tenant, metric, sorted tags)Stable across agents
Tag normalizationlowercase keys, trimmed values, max 200 charsPrevents accidental duplicates
Counter handlingAgent sends deltas or monotonic totals; server detects resets (value decreases)Reset detection is a correctness-critical path (Deep Dive 2)
DistributionSketch per flush intervalPriced separately
Monitor state key(monitor_id, group_key)Checkpointed per evaluation tick

Appendix C: Coordination Mechanisms#

MechanismUsed ForNotes
Kafka partitions by (tenant, series hash)Ordering per series; consumer parallelismSee Kafka
Consumer groups (evaluators, writers)Independent consumption at different speedsStorage lag doesn't delay alerts
Consistent hashing of series → ingest writersHead ownershipReplication factor 2 for the head
Checkpointed evaluator stateFailover without false recoveriesState store with per-tick writes
Cell routerTenant → cell mappingCached at intake; changes via migration protocol

Quick comparison: Evaluating alerts in stream processors (Flink-style) gives the freshest results and natural watermarks; evaluating against a recent-data replica is simpler for complex queries. Most platforms use both, split by rule complexity.

Appendix D: Agent and API Contract#

BehaviorContract
FlushEvery 10–15 s, batched, compressed
BufferingDisk buffer 15–60 min on intake failure; oldest dropped first beyond limit
RetriesExponential backoff with jitter; honor 429/503 Retry-After
HeartbeatAgent emits agent.running every flush — liveness signal
Limits feedbackIntake returns per-batch warnings for dropped series; agent logs and surfaces them
TimestampsAgent clock; intake rejects points > 10 min in the future or older than the late-arrival limit

Appendix E: Observability of the Observability Platform#

E.1 Core Metrics#

Intake:     intake.accepted_points{cell}, intake.series_dropped{tenant}, intake.new_series_rate{tenant}
Log:        kafka.consumer_lag{group}, kafka.under_replicated_partitions
Alerting:   alerting.eval_lag_p99, alerting.monitors_no_data, alerting.synthetic_mismatch_total
Notify:     notifier.delivery_latency_p99, notifier.delivery_failures{provider}
Storage:    ingest.head_series, ingest.memory_used_pct, compaction.backlog
Query:      query.latency_p95, query.series_scanned_p99, query.shed_total
Watchdog:   watchdog.synthetic_alert_latency (from outside)

E.2 Critical Alerts#

AlertThresholdAction
Watchdog synthetic alert missing> 2 minPage via independent provider
Evaluator lagp99 > 120 s for 5 minPage alerting on-call
Notifier deliveryfailures > 1% for 5 minPage; fail over provider
Ingest dropaccepted points < 70% of baseline per cellPage
Head memory> 85% on any writerPage; shed new series
Synthetic mismatchanyPage alerting team

E.3 Control Plane vs Data Plane#

Control plane: monitor CRUD, tenant limits, cell routing, config. Data plane: intake, log, evaluation, notification, storage. The data plane runs on last-known-good control-plane state; control-plane changes roll out cell by cell with health gates (Deep Dive 4).

E.4 Debugging "My Alert Didn't Fire"#

  1. Replay the monitor's evaluation from the log for the window.
  2. Check data arrival: were points accepted, late, or dropped by limits?
  3. Check semantics: evaluation delay, no-data policy, for duration, recovery threshold.
  4. Check notification: state transition recorded? Delivery attempts and provider responses?

Appendix F: Scale Evolution#

F.1 What Works at Each Scale#

ScaleArchitecture
< 1M seriesSingle Prometheus pair + Alertmanager
1M–50MPrometheus with remote write to Mimir/Thanos/M3; central Grafana
50M–1BMulti-tenant cluster, Kafka, separate evaluators, rollups, per-team limits
1B+Cells, shuffle sharding, sketches, federation, multi-region

F.2 Multi-Region Path#

Independent regional stacks; tenant home region; query-time federation; region-local alerts; watchdogs placed outside the region they watch (Drill 10, Deep Dive 5).

F.3 What You Don't Build on Day One#

  • Anomaly-detection alerting (start with thresholds and burn rates)
  • Global federated query
  • 1-second resolution
  • Custom TSDB engine
  • Dedicated cells (until a tenant reaches a meaningful share of a cell)

Appendix G: Multi-Tenancy, Fairness, and Cost#

ConcernMechanism
Ingest fairnessPer-tenant token buckets on points/s; see Rate Limiting
Cardinality fairnessActive-series caps and new-series rate limits per tenant
Query fairnessPer-tenant concurrency slots, series-scanned and memory caps
Alert fairnessEvaluator capacity reserved per cell; per-tenant monitor count limits
Blast radiusCells; dedicated cells for the largest tenants
BillingPer host or per custom-metric series, distributions separately, retention tiers

🧭 Principal Insight: "Charge for what costs us: series and retention. If the pricing unit doesn't match the cost driver, customers will optimize for the invoice and we'll absorb the difference."

Related reading: Stream Processing, Message Queues, Notification Systems, Circuit Breakers, Degraded Mode, Data Pipelines, and Capacity Planning.

  1. Loading the index…