Hiring BarSupport

Design a Distributed Tracing System

Case study79 min read9 diagrams

Technologies referenced in this case study: Apache Kafka · Cassandra · Elasticsearch · OLAP Databases · Time-Series Databases · Kubernetes

Related: Metrics Platform · Stream Processing Engine · Object Storage · Stopping Cascading Failures · Write-Heavy Systems · Backpressure & Load Shedding · Latency, Protocols & Tail Amplification

Go deeper:

  • Traces and logs usually share one collection layer; Log Aggregation covers the agent, tenancy and redaction side of that pipeline.

Reading Guide#

Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal View) and the appendices on sampling math, storage layout and propagation
What is a Distributed Tracing System? — Why interviewers pick this topic

A distributed tracing system follows one request as it crosses dozens of services. Each service records spans — a named, timed operation with a parent — and every span carries the same trace ID, passed from service to service in request headers. Assembled, the spans form a tree that shows where time went: the checkout request took 2.3 seconds, 1.9 of which were a retry loop in the inventory service waiting on a cache that was timing out.

The hard part is not recording timestamps. The hard part is that tracing everything is ruinously expensive — a large company can generate tens of millions of spans per second — so you must decide which requests to keep, often before you know which ones will be interesting. The context must survive every hop, including message queues, thread pools and batch jobs, or traces fall apart into fragments. And the tracing system must never make the application slower or less available than it would be without it.

Before vs After — the "p99 regression nobody can explain" scenario:

Without tracing:
t=0:       Checkout p99 rises from 800ms to 2.4s after a Tuesday deploy train (14 services).
t=+20min:  Each team's dashboard looks "mostly fine" — their p50s haven't moved.
t=+2h:     War room with 9 teams. Everyone's service is fast on its own metrics.
t=+6h:     Someone notices inventory's retry counter rose. Root cause: a client library
           upgrade changed retry backoff from 50ms to 500ms on cache timeouts.
t=+7h:     Rollback. 7 hours of degraded checkout.

With tracing (tail-sampled on latency and errors):
t=0:       Checkout p99 rises to 2.4s.
t=+3min:   On-call opens 20 slowest checkout traces from the last 5 minutes (100% of slow
           traces were kept by the tail sampler).
t=+5min:   Every slow trace shows the same shape: inventory → cache span times out,
           followed by a 500ms gap, then a retry. Span attribute: client_lib=v4.2.0.
t=+12min:  Inventory rolls back the library. p99 recovers.

Why interviewers reach for this question: It's a write-heavy data pipeline (tens of millions of events per second) bolted onto a sampling problem (what to keep), a propagation problem (how context survives every hop), and an organizational problem (every team must instrument, and someone must pay). Candidates who draw "agents → Kafka → Elasticsearch" and stop have designed the cheap part.

Mechanics Refresher: Tracing Primitives
PrimitiveHow It WorksProsCons
Trace / span modelTrace ID shared by all spans; each span has span ID, parent ID, name, start, duration, attributes, statusReconstructs the call tree across servicesSpans arrive from many hosts, out of order, at different times
Context propagationTrace ID, parent span ID and flags travel in headers (W3C traceparent) and in-process contextStandard across languages and vendorsBreaks silently at any hop that doesn't forward it — queues, thread pools, custom RPC
Head samplingDecide at the root, from the trace ID, before anything happens; propagate the decisionCheap; consistent across services; zero bufferingCan't keep "the slow ones" or "the errors" — decided too early
Tail samplingBuffer all spans of a trace, decide after it completes, keep errors/slow/rareKeeps the interesting tracesStateful; all spans of a trace must reach the same decider; memory and complexity
Collection-time (secondary) samplingHash trace ID in the pipeline, keep below a coefficientOne knob to control total write rateUniform; doesn't target interesting traces
Span metricsDerive RED metrics (rate, errors, duration) from 100% of spans before samplingAccurate aggregates even with 1% trace retentionCardinality must be controlled
Trace-ID index onlyStore spans grouped by trace ID in object storage; lookup by IDVery cheap storageSearch by attribute needs another mechanism
ExemplarsMetrics carry trace IDs of sample requestsJump from a latency spike to a traceOnly as good as the sampler's retention

For most production systems: OpenTelemetry instrumentation with W3C Trace Context propagation, a local agent or sidecar, a collector tier that computes span metrics from 100% of spans and then tail-samples (keep all errors, all slow traces, a small probabilistic baseline), a durable buffer in front of storage, and cheap columnar or object storage with a trace-ID lookup plus limited attribute search. The primitives are not the interview — what you keep, what it costs, and making sure tracing never hurts the application are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

Distributed tracing is not a storage question. Everyone can put spans in a database.

It is a sampling-economics and blast-radius question that tests:

  • Whether you decide which traces are worth keeping — and know that the interesting ones can't be identified at the start of the request
  • Whether context propagation is treated as a correctness requirement with gaps you can measure
  • Whether the tracing pipeline is designed to fail without harming the applications it observes
  • Whether someone owns the bill — because tracing cost grows with traffic × services × attributes, and nobody notices until finance does

The key insight: Tracing's value is concentrated in a tiny fraction of requests — the slow, the failed, the rare — while its cost is proportional to all of them. Staff candidates design the system around keeping that fraction (tail sampling plus metrics from 100% of spans) and make the cost visible; Senior candidates design a pipeline that stores a uniform 1% and hope the bad request was in it.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws SDK → agent → Kafka → storage → UIAsks "What questions must traces answer — debugging slow requests, service maps, or auditing business flows? How much are we willing to spend?"Asks "Is tracing a platform every team must adopt, and who funds it — central budget or chargeback?"
Sampling"Sample 1% at the edge""Head-sample a baseline, tail-sample in the collector: keep 100% of errors and traces over the latency SLO, plus ~1% of the rest; compute span metrics before sampling"Sets org-wide sampling policy and budget per team; makes "trace coverage of incidents" the success metric
Propagation"Pass the trace ID in a header""W3C traceparent everywhere; instrumented clients for HTTP, gRPC, queues and thread pools; measure broken-trace rate per service edge"Makes propagation a platform requirement enforced in shared libraries and CI
Pipeline failure"Kafka buffers it""Export is async, bounded and lossy by design — the app drops spans before it ever blocks; collectors shed load by sampling harder"Writes the rule: observability must never be in the availability path; audits it in game days
Storage"Elasticsearch, index everything""Spans grouped by trace ID in object storage for cheap retention; a columnar store for attribute search over a shorter window"Prices retention × indexing per tier; decides build vs buy vs vendor with 3-year cost
OwnershipObservability team owns it"Platform owns pipeline and storage; service teams own instrumentation quality and their sampling overrides; cost shown per team"Redraws: tracing, metrics and logs as one telemetry platform with shared collectors, budgets and governance
Why "sampling" separates levels

L5: "We sample 1% of requests at the entry point and propagate the decision." This is head sampling: cheap and consistent. It also means that when the error rate is 0.1%, you keep 1% of 0.1% — one error trace in 100,000 requests. The p99.9 request the on-call is looking for is almost never there.

L6: "Two decisions, two places. The app head-samples generously — say 10–100% depending on volume — because it costs little to send spans to a local agent. The collector tier tail-samples: it buffers each trace for ~30 seconds after its last span, then keeps every trace with an error, every trace above the endpoint's latency SLO, every trace touching a service flagged for debugging, and 1% of the rest. Before any of that, the collector computes RED metrics from 100% of spans, so dashboards are accurate even though we keep ~2–5% of traces."

L7: "Sampling policy is a budget allocation. I'd give each team a span budget and let them spend it — more baseline for a new service, more retention for a payments flow — with the platform enforcing a ceiling. The success metric isn't 'spans stored'; it's the fraction of incidents where the on-call found a relevant trace in under five minutes."

Why "pipeline failure" separates levels

L5: "The SDK sends spans to an agent, which sends to Kafka. Kafka is durable, so we don't lose data." Then the collector tier falls behind, the agent's buffer fills, and the SDK — configured with a blocking queue — starts holding request threads while it waits to enqueue spans. The observability system has just taken down checkout.

L6: "Every hop is bounded and lossy by design. The SDK exports asynchronously from a fixed-size in-memory queue — 2,048 spans — and drops on overflow, incrementing a counter. The agent does the same. The collector sheds by tightening the sampling rate. We lose traces during a pipeline incident; we never add latency to a request. Tracing data is valuable, but it's never worth an outage."

L7: "I'd make 'telemetry is never in the availability path' a written standard and verify it with game days: black-hole the collectors in one region and confirm application p99 doesn't move. If it does, that's a Sev-2 in the SDK configuration, not in tracing."

Why "storage" separates levels

L5: "Store spans in Elasticsearch so we can search by any attribute." Indexing every attribute of 250,000 spans per second costs more in indexing CPU and storage than everything else in the system combined, and most of those indexes are never queried.

L6: "Two access patterns, two stores. Lookup by trace ID — from an exemplar on a metric, a log line, an error report — is 90% of usage, and it's served by spans grouped by trace ID in object storage with a small trace-ID index: cheap enough for 14–30 days. Search by attribute — 'slow checkout traces for tenant X in the last hour' — goes to a columnar store with a shorter window, say 3–7 days, and only selected attributes."

L7: "Retention and searchability are priced per tier. Payments might buy 30 days of searchable traces; an internal batch service gets 3 days of trace-ID lookup. I'd publish those tiers with their cost so teams choose, rather than defaulting everyone to the expensive one."

Positions to Commit To#

PositionRationale
Tail-sample on errors and latency; head-sample only a baselineThe valuable traces can't be identified at the start of the request
Compute span metrics from 100% of spans before samplingDashboards and SLOs stay accurate while trace retention stays at a few percent
W3C Trace Context everywhere, via shared instrumented clientsOne standard across languages; propagation gaps become measurable
Async, bounded, lossy export at every hopTracing must never add latency or reduce availability
Trace-ID-keyed object storage for retention; columnar store for short-window search90% of lookups are by trace ID; indexing everything is the cost trap
Route all spans of a trace to the same tail-sampling node by trace-ID hashTail sampling requires the whole trace in one place
Show cost per teamTracing spend grows with traffic × services × attributes; invisible costs only grow

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Request-level debugging (latency and errors)Find the slow or failed request within minutes during an incidentTail sampling on errors/latency; trace-ID lookup; short-window search; exemplars from metricsThe bad trace wasn't kept; trace is fragmentedRelevant trace found in < 5 min for ≥ 90% of incidents
Service dependency analyticsAccurate maps, call rates and latency between servicesSpan metrics from 100% of spans; aggregated edges; traces optionalMaps miss rare edges; cardinality explosionEdge rates within a few percent of truth
Business transaction auditEvery order or payment traceable end to end, for complianceNot tracing — durable business event logs with correlation IDs, 100% retentionUsing sampled traces as an audit trail100% completeness, years of retention

🎯 Staff Move: "I'll design for request-level debugging: during an incident, the on-call should find a relevant slow or failed trace in under five minutes. Service maps fall out of span metrics computed from 100% of spans. If someone needs every order traceable for compliance, that's a business event log with correlation IDs — sampled tracing is the wrong tool, and I'd say so."

Where the Design Splits#

#Fault LineThe Tension
1Head vs Tail SamplingDecide cheaply at the start (miss the interesting traces) or expensively at the end (stateful collectors, buffering, routing)?
2Propagation Coverage vs Instrumentation CostAuto-instrument everything (overhead, noise) or hand-instrument (gaps); how to handle queues, batches and thread pools?
3Pipeline Durability vs Application SafetyBuffer durably to never lose spans, or drop freely so the app is never affected?
4Index Everything vs Trace-ID-First StorageRich attribute search at high cost, or cheap retention with limited search?
5Central Budget vs Per-Team OwnershipPlatform pays and controls sampling, or teams own their spend and overrides?

How Real Companies Built It#

Why this section belongs here: Tracing has unusually good public primary sources — a foundational paper, an open-source system's origin story, and a storage design built around cost. Citing them shows you know why the defaults are what they are.

Google — Dapper: Sampling Was the Design, Not an Optimization#

Google's Dapper paper describes a tracing system designed around low overhead, application-level transparency (instrumentation in a small set of common libraries) and ubiquitous deployment. Spans were written to local log files, pulled by per-host daemons and written to regional Bigtable repositories, one row per trace with a column per span; median collection latency was under 15 seconds. The first production version used a uniform sampling rate averaging one trace per 1,024 requests, which the paper reports was still adequate for high-volume services, while lower-traffic workloads motivated an adaptive scheme targeting a rate of sampled traces per unit time, with the probability recorded alongside each trace. Because production generated more than a terabyte of sampled trace data per day, Dapper added a second sampling stage in the collection pipeline: hash the trace ID and keep the span only if the hash falls below a coefficient — so whole traces are kept or dropped together, and the global write rate is one config knob (Google Research).

Staff insight: Two ideas from Dapper still define good designs: instrument the shared libraries, not every service, and make every sampling decision a function of the trace ID so all hosts agree without coordinating. Say: "Any drop decision anywhere in my pipeline hashes the trace ID, so I lose whole traces, never half of one."

Uber — Jaeger: Tracing Grew With the Microservice Count#

Uber has described how its tracing evolved from an early system that couldn't propagate context across services into Jaeger as the company grew from hundreds to more than 2,000 microservices. Early versions used Cassandra for storage; client libraries maintained request context through service chains, with a local agent receiving spans over UDP; and Jaeger introduced sampling strategies that services fetch from the backend via the local agent rather than hard-coding a fixed probability, so the backend can adjust rates to traffic (Uber engineering).

Staff insight: The lesson is that tracing adoption is gated by propagation across an organization's actual RPC and messaging stack, and that sampling must be controllable centrally because you can't redeploy 2,000 services to change a rate. Say: "Sampling configuration is pulled from the control plane, not baked into binaries."

Grafana Labs — Tempo: Object Storage Only, Trace-ID First#

Grafana Tempo is an open-source tracing backend that describes itself as requiring only object storage to operate: it ingests batches in Jaeger, Zipkin, Kafka and OpenTelemetry formats, buffers them, and writes them to S3, GCS, Azure or local disk, positioning itself as cost-efficient and easy to operate (Grafana Tempo).

Staff insight: Designing storage around the cheapest durable medium — and accepting weaker search in exchange — is a legitimate Staff tradeoff once you notice most trace reads start from a known trace ID. Say: "I'd keep weeks of traces in object storage for pennies per gigabyte, and pay for rich search only over the last few days."

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"We sample 1% at the edge""Error rate is 0.1%. During an incident, how many error traces do you have?"Head-sampling blind spot
"We tail-sample in the collector""Spans for one trace arrive at 40 different collectors. How does any one of them decide?"Trace-ID-affinity routing
"We pass the trace ID in headers""The request goes through Kafka and a batch job. Is it still one trace?"Async propagation, span links
"Kafka makes the pipeline durable""The collectors are down for an hour. What happens to the app?"Lossy-by-design export
"Elasticsearch for search""What does indexing 250K spans per second cost?"Storage economics
"Teams add attributes""Someone adds user_id and request_body as attributes. Now what?"Cardinality, PII
"Spans have timestamps""A child span starts before its parent. Why, and what do you show?"Clock skew handling

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The application's only obligation is to hand spans to a local agent without blocking — everything after that is allowed to drop. Load-balancing collectors route every span by a hash of its trace ID, so all spans of one trace meet at the same tail sampler. Span metrics are computed before sampling from 100% of spans, so dashboards and SLOs are accurate. Tail samplers keep errors, slow traces and a small baseline — typically a few percent of traces — and write them through Kafka into two stores: object storage keyed by trace ID for cheap retention, and a columnar store for attribute search over a short window. A control plane pushes sampling policy to SDKs and samplers so rates change without redeploys.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Sampling"1% head sampling""Metrics from 100%; tail-sample errors, slow and rare; small probabilistic baseline."
Tail sampling mechanics"Collector decides""Route by trace-ID hash so one node sees the whole trace; buffer ~30s after the last span."
Propagation"Pass trace ID in headers""W3C traceparent via shared clients; span links across queues and batches; measure broken-trace rate per edge."
App safety"Kafka is durable""Async, bounded, drop-on-full at every hop. Tracing never adds latency."
Storage"Elasticsearch""Object storage by trace ID for weeks; columnar search for days; selected attributes only."
Cost"Scale the cluster""Cost ∝ traffic × spans/request × bytes/span × keep-rate × retention. Show it per team."
Ownership"Observability team""Platform owns pipeline; teams own instrumentation and their budgets."

Numbers to Bring#

MetricValueWhy It Matters
Dapper's first production sampling rate~1 in 1,024 tracesUniform low rates work for high-volume paths, miss low-volume ones
Dapper sampled data volume (as published)> 1TB/day; median collection latency < 15sEven heavily sampled tracing is a big-data pipeline
W3C traceparent IDstrace-id 16 bytes, parent-id 8 bytes, 1 flags byte (sampled bit)Fixed, small header; tracestate limited to 32 list members
OpenTelemetry tail-sampling defaultsdecision_wait 30s, num_traces 50,000 in memoryBuffer size × traces/s determines collector memory
Spans per request (microservice request)~10–100, median ~20–30Multiplies every volume estimate
Span size~200 bytes–1KB serialized; ~5–10× compression in columnar/object storageDrives network and storage cost
Example fleet1M requests/s × 25 spans ≈ 25M spans/s ≈ 12GB/s raw at ~500BWhy storing 100% is not an option at scale
Typical keep rate after tail sampling~1–5% of tracesKeeps 100% of errors and slow traces while storing a few percent
SDK export queue~2,048 spans, batch every ~5s or 512 spansBounded memory; drop on overflow
Object storage price~$0.02/GB-month (standard tier, order of magnitude)Weeks of retention cost pennies per GB
Trace search window3–7 days searchable, 14–30 days by IDMost incident investigation happens within hours

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on the span data model and a storage schema, then run out of time before the interviewer asks the only question that matters: "You keep 1%. The incident is in the 0.1% of requests that fail. Do you have the trace?" Compress the data model to ~5 minutes and spend the rest on sampling, propagation, pipeline safety, storage economics and ownership.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Services emit spans; we propagate trace context across every hop; we collect, sample and store traces; engineers look up a trace by ID, search for slow or failed traces, and see service dependency maps. The primary use is incident debugging."

Then the non-functional requirements, which is where the design lives:

"Three constraints drive everything. One: the tracing system must never hurt the applications — no added latency on the request path, no outages when the pipeline fails. Two: during an incident, the on-call should find a relevant slow or failed trace in under five minutes, which means we must keep the interesting traces, not a random sample. Three: cost. I'll assume 1 million requests per second at the edge, about 25 spans per request — 25 million spans per second, roughly 12GB/s raw. Storing all of it is out of the question, so sampling is the design."

Then name the underspecified parts:

"I'd confirm: is there a budget? Do we need traces for compliance? How long must traces be searchable? I'll assume 7 days searchable, 30 days by trace ID, and that compliance needs are handled by business event logs, not traces."

🎯 Staff Move: Stating "sampling is the design" with the 12GB/s number in the first three minutes tells the interviewer you know where the problem is.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Span: trace_id (16 bytes), span_id (8 bytes), parent_span_id, name, service, kind (server/client/producer/consumer/internal), start_unix_nano, duration_ns, status, attributes{}, events[], links[]
  • Trace: all spans sharing a trace_id; derived, not stored as an entity
  • Sampling decision: trace_id, kept (bool), reason (error / slow / baseline / debug), sample_rate (for re-weighting counts)
  • Span metric series: (service, operation, status, le) → count, duration histogram, exemplar trace IDs

Propagation (what crosses every hop):

traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
             ver-trace_id (16 bytes hex)--------parent_id (8 bytes)-flags (01 = sampled)
tracestate:  platform=p:8;r:3      (vendor-specific, ≤ 32 members)

Query API:

GET /v1/traces/{trace_id}
GET /v1/search?service=checkout&op=POST /cart&min_duration=2s&status=error&since=1h&limit=20
GET /v1/services/{svc}/dependencies?since=1h

🎯 Staff Move: "I'll record the sampling probability on every kept trace. Without it, counts derived from stored traces are wrong — a trace kept at 1% represents 100 requests, an error trace kept at 100% represents one."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk one request in 90 seconds:

  1. Edge receives a request without traceparent, generates a trace ID, makes a head decision (sampled flag on at 10% for this high-volume route — everything is recorded locally; the flag only governs export from this service), and propagates.
  2. Each service's SDK creates server and client spans, forwards traceparent on every outbound call, and exports asynchronously in batches to the agent on localhost.
  3. The agent batches, compresses and forwards to load-balancing collectors, which hash trace_id to pick a tail-sampler node — so all spans of a trace land together.
  4. Before sampling, span metrics are computed: request rate, error rate and duration histograms per service and operation, with exemplar trace IDs.
  5. The tail sampler buffers each trace until 30 seconds after its last span, then applies policy: error → keep; root duration > SLO → keep; debug flag → keep; else keep 1%.
  6. Kept traces go through Kafka to object storage (grouped by trace ID) and to a columnar store with selected attributes.

🎯 Staff Move: Say out loud: "Every drop decision — SDK overflow, agent overflow, collector shedding, sampling — is either blind or a function of the trace ID. Nothing in the pipeline can block the application." You've now spent ~8 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the pipeline, and it's the Senior-level design. What makes tracing hard is four things: keeping the right traces without keeping all of them, keeping context intact across async boundaries, making sure the pipeline's failures never become the application's, and storing weeks of traces for a sane cost. I'd like to go deep on sampling first. Where would you like to start?"

If no preference: start with sampling. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Sampling (8–9 min)

"Head sampling alone decides before we know if the request is interesting. At 1%, with a 0.1% error rate, we keep one error trace per 100,000 requests. So: metrics from 100% of spans, computed in the collector before sampling — that's how dashboards stay accurate. Then tail sampling keeps all errors, everything over the route's latency SLO, everything for services with a debug flag, and 1% baseline. Typical outcome: 2–5% of traces kept, but 100% of the ones the on-call wants."

Quantify: "Tail sampling needs the whole trace in memory. 1M traces/s × 30s decision wait = 30M traces buffered. At ~25 spans × 500 bytes, that's ~375GB of RAM across the sampler tier — say 60 nodes at 8GB each with headroom. That's the real cost of tail sampling, and it's why I'd head-sample very high-volume, low-value routes like health checks to near zero before they reach the sampler."

Who pays: "The platform pays memory for tail samplers; in exchange, the on-call has the trace. For a health-check route, nobody needs the trace — it's head-sampled at 0.01%."


Deep dive 2: Propagation (5–6 min)

"W3C traceparent via instrumented clients in the shared libraries — HTTP, gRPC, database drivers, the queue client. For Kafka, the producer writes traceparent into message headers, and the consumer starts a new span with a link to the producer span — not a child, because one batch consumption may cover messages from 500 traces. Thread pools need context capture at task submission. I measure propagation: for each service edge, the fraction of server spans with no parent where the caller is instrumented — the broken-trace rate. Anything over 1% gets a ticket to the owning team."


Deep dive 3: Pipeline safety (4–5 min)

"SDK: async batch exporter, queue of 2,048 spans, drop-on-full, metrics on drops. Agent: 200MB buffer, drop oldest. Collectors: when CPU or memory crosses 80%, tighten the baseline rate and finally drop to errors-only. The app never waits. We game-day it: black-hole the collectors and confirm app p99 is unchanged."


Deep dive 4: Storage (5–6 min)

"Two stores. Object storage: kept traces grouped into blocks of ~100MB by time, with a trace-ID index and bloom filters per block — lookup by ID touches a few blocks. 3% of 25M spans/s is 750K spans/s; at 500 bytes that's ~32TB/day raw, ~4TB/day compressed 8× — ~120TB for 30 days, roughly $2.5K a month in object storage. Columnar store: kept spans with ~20 promoted attributes (service, operation, status, duration, tenant, region, version) for 7 days — that's what powers 'show me slow checkout traces for tenant X'."


Deep dive 5: Ownership and cost (3–4 min)

"The observability platform owns SDK defaults, the pipeline, storage and the sampling policy service, and pages on pipeline health and trace.lookup_success_rate. Service teams own their instrumentation quality, their custom attributes and their sampling overrides within a budget. Every team sees its span volume and cost monthly."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: tracing's value is in a tiny fraction of requests and its cost is in all of them. So: metrics from 100% of spans, tail sampling to keep the errors and slow traces, propagation enforced through shared libraries and measured per edge, a pipeline that drops rather than blocks, and storage built around trace-ID lookup with a short searchable window."

The evolution closer:

"What I'd build later: correlation with logs and profiles by trace ID, adaptive baseline rates per route to hit a span budget, and trace-derived analytics like critical-path analysis. What I'd not build: 100% retention of all spans, or tracing as an audit log."

🎯 Staff Move: End on the metric that proves it works. "I'd track the fraction of incidents where the on-call found a relevant trace within five minutes. If that's not above 90%, the sampling policy is wrong, no matter how many spans we store."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Data model deep dive10 min on span fields and schemaNames span fields in 30s; spends time on sampling
Storage firstPicks a database before deciding what to keepDecides keep rate, then sizes storage
Uniform sampling"1% everywhere"Tail sampling with policy; metrics from 100%
Ignores async hopsAssumes HTTP everywhereQueues, batches and thread pools in the first 20 minutes
Durable-at-all-costs pipeline"Never lose a span""Drop rather than block; tracing is never in the availability path"
No cost model"Scale the cluster"Bytes/day, $/month, and who pays

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Tracing is the interview where the system you design is judged on whether it helps other people debug other systems, under time pressure, while costing less than the problems it solves. The candidate must reason about a data pipeline whose input grows with every new service and every new attribute, decide which data is valuable before it's needed, and guarantee that the observer never harms the observed. That is Staff work: designing a platform with organization-wide adoption, cost governance and a failure posture.

It also has a subtle failure distribution. A tracing outage is rarely noticed directly — it's noticed during an unrelated incident, when the on-call reaches for a trace and finds nothing. Sampling that keeps the wrong traces fails the same way: silently, until it matters. And a tracing SDK that blocks under backpressure turns an observability problem into a customer-facing outage.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"It's 3am. Checkout error rate just went from 0.05% to 0.4%. The on-call opens your tracing UI. Walk me through exactly what they see, how they got there, and why the trace they need was kept."

A candidate who answers with the alert from span metrics computed on 100% of spans, the exemplar link on the error-rate graph, the tail-sampling rule that kept every error trace, the trace view showing the failing downstream call with its attributes, and the propagation that kept the trace intact through the payment queue has built tracing for on-calls. A candidate who says "they search Elasticsearch for errors" has built a span database.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Request-level debugging → keep the interesting traces

  • Constraint: find a relevant slow or failed trace within minutes during incidents; tracing must not affect applications
  • Strategy: tail sampling on errors/latency; trace-ID lookup; short-window attribute search; exemplars from span metrics
  • Failure mode: interesting trace not kept; trace fragmented; pipeline down when needed
  • Who pays for imperfection: on-calls (longer incidents), customers (longer outages)

Service dependency analytics → aggregates, not traces

  • Constraint: accurate call rates, latencies and error rates per service edge; complete dependency maps
  • Strategy: span metrics computed from 100% of spans in the collector; edges aggregated; traces optional
  • Failure mode: cardinality explosion from unbounded attributes; maps missing rare edges if derived from sampled traces
  • Who pays: the metrics backend (cardinality), architects (wrong maps)

Business transaction audit → not tracing

  • Constraint: every transaction traceable, years of retention, legal defensibility
  • Strategy: durable business events with correlation IDs in an event log or warehouse
  • Failure mode: relying on sampled, short-retention traces as evidence
  • Who pays: compliance, when the trace they need was sampled out

2.2 When NOT to Build a Tracing System#

  • You have a monolith or a handful of services. Structured logs with a request ID and good metrics answer most questions. Tracing earns its keep when a request crosses more services than a human can hold in their head — roughly 5–10+.
  • You need an audit trail. Use business events with correlation IDs and long retention. Sampled tracing will lose exactly the record someone asks for.
  • Your vendor bill is fine and your scale is moderate. A hosted tracing backend with OpenTelemetry instrumentation is usually cheaper than a platform team until span volume makes per-span pricing exceed a few engineers' cost. Keep instrumentation vendor-neutral so the backend is swappable.
  • You want tracing to replace metrics. Traces are for "why is this request slow"; metrics are for "how many are slow". Derive metrics from spans, but don't query stored traces for SLOs.

🎯 Staff Insight: "Instrumentation is the long-lived investment; the backend is replaceable. I'd standardize on OpenTelemetry and W3C Trace Context first — that decision outlives any storage choice."

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Budget — tracing cost scales with traffic, services, spans and attributes; without a budget, every design is either too expensive or too sparse
  • Which traces matter — errors? slow? specific tenants? specific routes?
  • Async boundaries — queues, batch jobs, cron, streaming — where propagation breaks
  • Retention and search — how long, and searchable by what?
  • PII — attributes and span events can capture user data; who scrubs it?
  • Multi-tenancy — do customers see traces (as in a platform product), or only internal engineers?

Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.

2.4 Precise Terminology#

TermWhat It MeansWhy It Matters in the Interview
TraceAll spans sharing a trace IDAssembled after the fact from many hosts
SpanOne timed operation with parent, attributes, statusUnit of volume and cost
Root spanSpan with no parent; the request's entryIts duration is "the request latency"
Context propagationCarrying trace ID and parent span ID across process boundariesBreaks silently; must be measured
Span linkA non-parent reference to another spanFan-in from queues and batches
Head samplingDecision at trace start, propagated via flagCheap, consistent, blind to outcome
Tail samplingDecision after trace completesKeeps interesting traces; stateful
Sample rate / weightProbability a trace was keptNeeded to re-weight counts
ExemplarTrace ID attached to a metric sampleBridge from dashboard to trace
CardinalityNumber of distinct label combinationsExplodes metrics and indexes
Clock skewHosts' clocks disagreeChild spans appear outside parents

🎯 Staff Insight: If the interviewer says "trace every request", ask: "Record every request, or keep every request? Recording locally is cheap; keeping is where the cost is. I'd record broadly and keep selectively."


3. Where the Design Splits#

Every tracing decision has a technical side (sampling algorithm, propagation format, storage layout) and an organizational side (who sets the budget, who fixes broken propagation, who decides what's kept). Interviewers grade the second side.

3.1 Fault Line 1: Head vs Tail Sampling#

The tension: Head sampling is cheap and consistent but decides before the outcome is known. Tail sampling keeps the traces you want but requires buffering whole traces in one place, routing by trace ID, and a decision wait that delays visibility.

ChoiceWhat WorksWhat BreaksWho Pays
Head only (uniform 1%)Cheap; stateless; consistentMisses rare errors and tail latencyOn-call (missing traces)
Head, adaptive per routeLow-traffic routes get higher ratesStill outcome-blindOn-call (fewer misses, still misses)
Tail sampling in collectorsKeeps errors, slow, rareMemory: traces/s × wait × size; routing by trace ID; late spansPlatform (sampler tier)
Head baseline + tail policy + metrics from 100%Accurate metrics; interesting traces kept; bounded costMost components to operatePlatform (complexity)
Diagram: 3.1 Fault Line 1: Head vs Tail Sampling

Staff default: "Span metrics from 100% of spans, computed before sampling. Head-sample in SDKs at a generous, route-dependent rate — 100% for most routes, much lower for health checks and ultra-high-volume internal calls. Tail-sample in collectors routed by trace-ID hash: keep all errors, all traces over the route's SLO, all traces from debug-flagged services or tenants, and 1% of the rest. Record the sample rate on every kept trace."

When to deviate:

  • Moderate volume (< ~50K spans/s): keep 100%. Sampling complexity costs more than the storage.
  • Extreme volume with tight budgets: head-sample harder before the sampler; tail sampling's memory scales with what reaches it.
  • Mobile/browser clients: head decisions only — you can't buffer client spans for tail decisions.

🧭 Principal Move: "I'd define success as incident trace coverage — the fraction of incidents where the on-call found a relevant trace within five minutes — and tune sampling policy against it each quarter. Spans stored is a cost metric, not a value metric."

❌ Common L5 Trap: "Sample 1% at the edge and propagate the decision, so traces are always complete." Complete, yes — and almost never the ones that matter. With a 0.1% error rate, the on-call has one error trace per 100,000 requests, which during a 10-minute incident at 1,000 errors/s means about 6 traces, scattered across unrelated failure modes.


3.2 Fault Line 2: Propagation Coverage vs Instrumentation Cost#

The tension: A trace is only as good as its weakest hop. Auto-instrumentation covers common libraries but adds overhead and noise; manual instrumentation is precise but has gaps. Async boundaries — queues, batches, thread pools, cron — break context unless explicitly handled.

ChoiceWhat WorksWhat BreaksWho Pays
Manual instrumentation per teamPrecise, meaningful spansGaps wherever a team didn't bother; inconsistent namesOn-call (fragmented traces)
Auto-instrumentation agentsBroad coverage quicklyOverhead; noisy spans; version driftService teams (CPU, noise)
Instrumented shared libraries (RPC, HTTP, DB, queue clients)Consistent coverage where it mattersPlatform must maintain clients for every stackPlatform (library ownership)
Span links for async fan-inCorrect semantics for batches and queuesUI must show linked tracesPlatform (UI work)
Diagram: 3.2 Fault Line 2: Propagation Coverage vs Instrumentation Cost

Staff default: "Propagation lives in shared, instrumented clients — HTTP, gRPC, database drivers, queue producers and consumers — owned by the platform. For queues, the producer injects traceparent into message headers; single-message consumers continue the trace as a child; batch consumers start a new trace with span links to each producer span. Thread pool executors are wrapped to capture context at submission. I measure broken-trace rate per service edge and give the owning team a ticket above 1%."

When to deviate:

  • Legacy services you can't change: proxy-level spans from a service mesh or gateway give partial coverage without code changes — timing at the edges, nothing inside.
  • Hot paths with tight CPU budgets: instrument at RPC boundaries only, skip internal spans.

🧭 Principal Move: "Propagation is a platform contract, not a team preference. I'd make the instrumented clients the only supported way to make RPCs and publish messages, and add a CI check that fails builds using raw HTTP clients for internal calls."

❌ Common L5 Trap: "Each service passes the trace ID in a header." Correct for synchronous HTTP — then the order event goes through Kafka, a batch job processes it at 2am, and the trace ends at the producer. The interviewer asks how to see the consumer's failure in the original request's trace, and there's no answer.


3.3 Fault Line 3: Pipeline Durability vs Application Safety#

The tension: Losing spans is bad; slowing or crashing the application is far worse. Durable, never-drop pipelines push backpressure upstream — eventually into the application.

ChoiceWhat WorksWhat BreaksWho Pays
Synchronous exportSimpleEvery request waits on tracing; collector outage = app outageCustomers (latency, outages)
Async export, unbounded queueNo latency addedMemory grows until OOM during collector outageService teams (OOM crashes)
Async, bounded, drop on fullApp unaffectedSpans lost during pipeline incidentsOn-call (gaps in traces during pipeline incidents)
Bounded + durable buffer after the agent (Kafka)Collectors can fail without losing kept tracesKafka is another system; doesn't protect the SDK hopPlatform (Kafka ops)
Diagram: 3.3 Fault Line 3: Pipeline Durability vs Application Safety

Staff default: "Bounded and lossy up to the collectors; durable after sampling. The SDK exports asynchronously from a 2,048-span queue and drops on overflow. The node agent buffers ~200MB and drops oldest. Collectors shed load by tightening sampling — baseline first, then everything except errors. After the tail sampler, kept traces go to Kafka with 24-hour retention, so storage can fall hours behind without loss. Every drop point exports a counter."

When to deviate:

  • Low-volume, high-value spans (e.g., a payments flow at 50 req/s): a larger SDK buffer and retry on export are cheap and safe.
  • Serverless functions: export must flush before the function freezes; use a local extension or accept loss on cold paths.

🎯 Staff Insight: "I'd rather lose a minute of traces than add a millisecond to checkout. That sentence should be in the SDK's default configuration, not just in my head."

❌ Common L5 Trap: "We use Kafka so no spans are ever lost." Kafka protects the hop after the agent. If the SDK uses a blocking queue and the agent can't reach Kafka, request threads block on span export — the observability system becomes the outage.


3.4 Fault Line 4: Index Everything vs Trace-ID-First Storage#

The tension: Engineers want to search traces by any attribute. Indexing every attribute of every kept span is the dominant cost in many tracing backends. Most reads start from a known trace ID.

ChoiceWhat WorksWhat BreaksWho Pays
Full-text/inverted index on all attributesSearch anythingIndex CPU and storage often exceed raw data; short retention to afford itPlatform budget
Wide-column store keyed by trace ID + secondary indexes on a few fieldsFast ID lookup; limited searchSecondary indexes on high-cardinality fields hurtPlatform (index tuning)
Object storage by trace ID + bloom filtersCheapest retention; fast ID lookupNo attribute search without scanningEngineers (must start from an ID)
Object storage + columnar store for promoted attributes (short window)Cheap retention + useful searchTwo stores; attribute promotion governancePlatform (governance)
Diagram: 3.4 Fault Line 4: Index Everything vs Trace-ID-First Storage

Staff default: "Object storage for all kept traces, grouped into time-ordered blocks with a trace-ID index and bloom filters, retained 30 days. A columnar store with ~20 promoted attributes — service, operation, status, duration, region, version, tenant — retained 7 days for search. Aggregates come from span metrics, never from scanning stored traces. Adding a promoted attribute requires a cardinality review."

When to deviate:

  • Low volume: one store with full indexing is simpler and fine below a few TB.
  • Customer-facing tracing product: tenants expect rich search; index per tenant with quotas and price it.

🧭 Principal Move: "Retention and searchability are product tiers with prices. I'd publish them — 3, 7 or 30 days searchable — and charge teams for the tier they pick, instead of defaulting everyone to the most expensive option."

❌ Common L5 Trap: "Put everything in Elasticsearch so we can search any field." At 750K kept spans/s with 30 attributes each, the indexing tier becomes the largest cost center in observability — and the team responds by cutting retention to two days, which means the trace from last week's incident is gone when the post-mortem starts.


3.5 Fault Line 5: Central Budget vs Per-Team Ownership#

The tension: If the platform pays for everything, teams add spans and attributes freely and costs grow without bound. If teams pay, they under-instrument to save money, and traces get worse.

ChoiceWhat WorksWhat BreaksWho Pays
Central budget, no visibilityFrictionless adoptionCosts grow 30–100% a year unnoticedPlatform budget (until finance notices)
Full chargeback per spanCost disciplineTeams cut instrumentation; traces fragmentOn-call (worse traces)
Central baseline + per-team budgets for extrasEveryone gets the default; extras are a choiceBudget governance overheadPlatform (FinOps tooling)

Staff default: "The platform funds a baseline: default instrumentation, errors and slow traces always kept, 1% baseline, 7/30-day retention. Teams get a span budget for extras — custom attributes, higher baseline, longer retention — and see monthly cost per service. Exceeding the budget tightens their baseline automatically, never their error capture."

When to deviate:

  • Early adoption phase: fully central, no budgets — adoption matters more than cost.
  • Customer-facing tracing: cost is passed to customers through pricing.

🧭 Principal Move: "Errors and slow traces are never subject to a team's budget. Budget pressure must never make an incident harder to debug — it only reduces the baseline."

❌ Common L5 Trap: "The observability team just scales the cluster as usage grows." That works until the annual review shows tracing costs tripled while traffic grew 40%, because three teams added request bodies as attributes. Without per-team visibility, nobody can say which.


4. When It Breaks#

4.1 The Observer Takes Down the Observed#

t=0:       Collector tier deploy with a bad config: all collectors reject OTLP with 503.
t=+10s:    Node agents buffer; 200MB fills in ~90s on busy nodes.
t=+2min:   One team's service uses a hand-rolled exporter with a blocking queue.
           Request threads block on span enqueue. Its p99: 80ms → 9s.
t=+3min:   Upstream callers time out and retry. Checkout error rate 0.1% → 6%.
t=+6min:   Page: checkout errors. Nobody suspects tracing — tracing is down, so no traces.
t=+25min:  Thread dump shows threads parked in span exporter. Collector config rolled back.

Detection: sdk.export_queue_full_total, sdk.export_blocked_ms (must be 0), agent.buffer_utilization, collector.accepted_spans_rate drop.

Mitigation: roll back collectors; restart affected service with exporter disabled via flag.

Prevention: shared SDK configuration with async, bounded, drop-on-full export as the only supported mode; CI check rejects custom exporters; quarterly game day black-holing collectors with app p99 as the success criterion.

Owner: observability platform (SDK defaults, game days); service team (custom exporter removal).

4.2 The Fragmented Trace — Propagation Broke at a Queue#

t=0:       Payments team migrates from an instrumented queue client to a new async library.
           traceparent is no longer written to message headers.
t=+3 weeks: Incident: 2% of payments stuck. On-call opens traces for failed checkouts.
t=+3 weeks: Traces end at "publish payment.requested". Consumer-side spans are orphan roots.
t=+3 weeks: 40 minutes spent correlating by order ID across logs instead.

Detection: trace.broken_edge_rate{caller, callee} — fraction of consumer spans with no parent or link when the producer is instrumented; alert when an edge jumps above 1%.

Mitigation: patch the client to inject and extract context; add links on the consumer.

Prevention: propagation tests in the shared client's CI (produce → consume → assert linked); broken-edge dashboard reviewed weekly by the platform.

Owner: payments team (fix), observability platform (detection and client library).

4.3 The Incident That Overflows the Tail Sampler#

t=0:       Dependency slowdown. Traces get longer: 25 spans → 140 spans (retries, fan-out).
           Trace duration: 300ms → 25s, so decision wait extends for many traces.
t=+1min:   Tail sampler memory: 60% → 95%. num_traces limit reached; oldest traces evicted
           before decision — including error traces.
t=+2min:   On-call opens tracing: error traces from the incident window are partial or missing.

Detection: tailsampler.traces_evicted_before_decision, tailsampler.memory_utilization, tailsampler.spans_per_trace_p99.

Mitigation: emergency policy: drop baseline sampling to 0%, keep only errors and slow; scale samplers.

Prevention: memory sized for incident shape, not steady state (3–5× spans per trace); early-decision rule — any span with error status triggers immediate keep for that trace, so errors don't wait for the buffer; cap spans per trace (e.g., 2,000) with a truncation marker.

Owner: observability platform.

4.4 PII and Cardinality — The Attribute Nobody Reviewed#

t=0:       A team adds http.request.body and user.email as span attributes "for debugging".
t=+1 day:  Kept span size: 600B → 9KB. Storage writes ×6. Columnar ingest lags 4 hours.
t=+1 day:  Span metrics accidentally include user.email as a dimension: 30M new series.
t=+2 days: Security review finds customer emails in trace storage retained 30 days,
           accessible to 2,000 engineers.

Detection: spans.bytes_per_span_p99{service} jump; spanmetrics.series_count jump; PII scanner on a sample of attributes.

Mitigation: attribute deny-list applied in collectors (drop or hash); purge affected blocks; restrict access.

Prevention: collector-side attribute allow-list for span metrics dimensions; deny-list and redaction processors for known PII keys and patterns; attribute promotion requires review.

Owner: observability platform (collector policy), security (PII policy), the team (fix).

4.5 Clock Skew — Children Before Parents#

Spans are timestamped by different hosts. A host with a clock 80ms ahead produces a child span that starts before its parent and a trace view that's nonsense. At scale, some fraction of hosts is always skewed by tens of milliseconds.

Mitigation: display-time adjustment — if a child starts before its parent's start or ends after its parent's end on a different host, shift the child's subtree by the minimal offset to fit, and mark it as adjusted. Never rewrite stored timestamps.

Detection: trace.skew_adjusted_spans_pct{host}; persistently skewed hosts get a ticket to infra.

Owner: observability platform (UI), infrastructure (time sync).

4.6 The Silent Cost Explosion#

A new service mesh rollout adds two proxy spans per hop — doubling spans per request from 25 to 75 — and nobody adjusts sampling. Over a quarter, tracing storage and collector costs rise 2.8×. Nothing breaks; the finance review catches it.

Detection: spans_ingested_per_request per service; monthly cost per team.

Mitigation: drop or aggregate proxy spans that duplicate application client/server spans.

Owner: observability platform (policy), mesh team (span configuration).

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
SDK blocks appsdk.export_blocked_ms > 0That service and its callersAsync, bounded, drop-on-full onlyObservability platform
Collector outagecollector.accepted_spans_rate dropTrace gaps fleet-wideAgents buffer then drop; roll backObservability on-call
Broken propagationtrace.broken_edge_rate{edge} > 1%Traces through that edgeFix client; linksOwning service team
Tail sampler overflowtraces_evicted_before_decision > 0Missing traces during incidentsEarly keep on error; emergency policyObservability platform
PII in attributesPII scanner hitsCompliance exposureRedaction processors; purgeSecurity + platform
Cardinality explosionspanmetrics.series_count jumpMetrics backendDimension allow-listObservability platform
Storage write lagkafka.consumer_lag{storage_writer}Recent traces not yet visibleScale writers; Kafka absorbsObservability on-call
Clock skewskew_adjusted_spans_pct{host}Confusing trace viewsDisplay adjustment; fix NTPInfra
Cost growthspans_per_request{service}, monthly $BudgetBudgets, span policyPlatform + teams

🎯 Staff Insight: The most dangerous tracing failure is the one you discover during another incident. "I'd page on tracing pipeline health even though no customer is affected, because the next customer-affecting incident will need it. And I'd verify with a synthetic trace every minute: generate a trace with a known ID, confirm it's queryable within 60 seconds."


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framing"Collect and store spans; build a UI"Names debugging vs analytics vs audit; commits to incident debugging; states the 12GB/s problemAsks whether tracing is a mandated platform and who funds it
SamplingUniform head samplingMetrics from 100%; tail sampling on errors/latency; trace-ID routing; sample rate recordedSampling as budget allocation; incident trace coverage as the success metric
PropagationTrace ID in headersW3C via shared clients; links for queues and batches; broken-edge rate per edgePropagation enforced via libraries and CI as an org contract
Pipeline safety"Kafka for durability"Async, bounded, lossy until sampling; durable after; drop counters; game daysWritten standard: telemetry never in the availability path, verified quarterly
StorageOne database, index everythingObject storage by trace ID + short-window columnar search; promoted attributesRetention/search tiers priced and chosen by teams; build vs buy with 3-year cost
OrganizationObservability team owns itPlatform owns pipeline; teams own instrumentation and budgets; cost per teamOne telemetry platform across traces, metrics and logs with shared governance

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Separates metrics from traces"Metrics from 100% of spans before sampling; traces are kept selectively."
Knows tail sampling's mechanics"All spans of a trace must reach one sampler — route by trace-ID hash."
Protects the application"I'd rather lose a minute of traces than add a millisecond to checkout."
Handles async hops"Batch consumers start a new trace with span links to producers."
Storage follows access pattern"Ninety percent of reads start from a trace ID; I store for that."
Measures value"Incident trace coverage — found a relevant trace in five minutes — is the KPI."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Keep 100% of traces" at large scaleIgnores a 10–100× cost problem
Uniform head sampling as the full answerMisses exactly the traces that matter
Synchronous or blocking exportTracing becomes an availability risk
No mention of async propagationTraces fragment at the first queue
Index every attributeCost trap; leads to short retention
Tracing as an audit logSampled, short-retention data can't be evidence

5.4 Common False Positives#

  • Deep OpenTelemetry API knowledge ≠ tracing system design. Knowing every SDK option is useful; if it doesn't lead to sampling economics and pipeline safety, it's library trivia.
  • "We use a vendor" ≠ no design needed. Vendors still require sampling policy, propagation discipline and cost control — and per-span pricing makes those decisions more expensive to get wrong.
  • Pretty service maps ≠ useful tracing. Maps come from span metrics; the hard part is keeping the trace the on-call needs.
  • Big-data pipeline experience ≠ tracing. Kafka-to-columnar-store skills are necessary; tail-sampling semantics and propagation are what make it tracing.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minDebugging vs analytics vs audit; commit; app safety, find-the-trace, cost; 12GB/s
Entities & propagation format3–5 minSpan fields; traceparent; sample rate on kept traces
Architecture5–10 minSDK → agent → LB collectors → metrics + tail sampler → Kafka → stores
Sampling10–19 minMetrics from 100%; tail policy; trace-ID routing; sampler memory
Propagation19–25 minShared clients; queues, batches, thread pools; broken-edge rate
Pipeline safety + storage25–35 minBounded lossy export; object store + columnar; sizing and $
Pivot (interviewer's choice)35–42 minMulti-region, PII, cost cuts, build vs buy, multi-tenant tracing product
Wrap42–45 minValue in a tiny fraction; keep it; never hurt the app; incident trace coverage

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Cut the tracing bill by 50%"Cost leversDrop duplicate proxy spans, tighten baseline, shorten search window, promote fewer attributes — never cut error capture
"Make it multi-region"Trace localityRegion-local pipelines; cross-region traces joined at query time by trace ID
"We found emails in spans"Data governanceCollector redaction, allow-lists, purge, access controls
"A request spans Kafka and a batch job"Async semanticsHeader propagation, child vs link, new trace for batches
"Customers want to see their traces"Multi-tenancyTenant ID as a first-class field, per-tenant quotas, isolation in storage and query
"How do you know tracing works?"Value measurementSynthetic traces, broken-edge rates, incident trace coverage

6.3 What to Deliberately Skip#

  • UI design — say what the trace view shows (waterfall, critical path, attributes), not how it looks.
  • Every span attribute convention — mention semantic conventions exist; don't list them.
  • Profiling and logs integration — one sentence: "correlated by trace ID."
  • Storage engine internals — unless the interviewer probes block layout or bloom filters.
  • Vendor comparisons — say you'd keep instrumentation vendor-neutral.

6.4 Follow-Up Questions to Expect#

  1. "Error rate is 0.1%. You sample 1%. During an incident, how many error traces do you have?"
  2. "How does a tail sampler decide when spans for one trace arrive at 40 collectors?"
  3. "The collectors are down. What happens to the applications?"
  4. "A request publishes to Kafka and a batch job consumes it at 2am. Show me the trace."
  5. "Size the storage for 30 days at your keep rate."
  6. "Someone added the request body as a span attribute. What breaks?"
  7. "A child span starts 80ms before its parent. Why, and what do you show?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design a distributed tracing system for our microservices."

Staff Answer

"What questions must it answer? Debugging individual slow or failed requests during incidents, service dependency analytics, or auditing business flows end to end? Those are different systems — and the third isn't tracing at all. I'll design for incident debugging, with service maps derived from span metrics.

Constraints: tracing must never add latency or reduce availability; the on-call should find a relevant trace within five minutes during an incident; and cost — at a million requests per second and ~25 spans each, that's 25 million spans per second, ~12GB/s raw. So sampling is the design. I'll go: span model and propagation → pipeline → sampling (metrics from 100%, tail-sample the rest) → propagation across async hops → pipeline safety → storage and cost → ownership."

Why this is L6:

  • Separates three intents and excludes audit explicitly
  • States app safety and a measurable usefulness target before any box
  • Quantifies volume to justify sampling as the core problem

What L7 adds:

  • Asks whether tracing will be mandated and who funds it
  • Asks about existing telemetry systems to converge with metrics and logs
  • Names incident trace coverage as the success metric
❌ Common L5 Trap

"Services use an SDK to send spans to an agent, which forwards to Kafka; consumers write to Elasticsearch; a UI queries it. We sample 1% to control cost."

Why this misses: It's a reasonable pipeline that keeps a uniform 1%, can block the app if the SDK queue is misconfigured, indexes everything at great cost, and breaks at the first queue. The interviewer's next question — "do you have the error trace?" — exposes it.


Drill 2: Tail Sampling Mechanics#

Prompt: "Spans for one trace arrive at 40 different collectors. How does tail sampling work?"

Staff Answer

"Two tiers. The first tier is stateless load-balancing collectors: every span is routed by a consistent hash of its trace ID to one node in the second tier, so all spans of a trace meet at the same tail sampler. When sampler membership changes, consistent hashing moves only ~1/N of traces; traces in flight during a rebalance may be split — I accept a small loss there.

The sampler buffers spans per trace and decides 30 seconds after the last span arrives — or immediately when an error span arrives, so errors don't wait. Policy: error → keep; root duration over route SLO → keep; debug flag → keep; otherwise keep with p=0.01, decided by hashing the trace ID so any late spans get the same decision. Memory: traces/s × wait × trace size — at 1M traces/s, 30s and ~12KB per trace, ~375GB across the tier. Late spans after a decision go to a small cache of recent decisions — kept traces accept them, dropped traces discard them."

Why this is L6:

  • Explains trace-ID affinity routing and consistent hashing behavior on rebalance
  • Early decision on errors and deterministic hash-based probabilistic decisions
  • Sizes sampler memory and handles late spans

What L7 adds:

  • Notes sampler memory is the dominant cost of tail sampling and head-samples low-value routes before they reach it
  • Makes the policy centrally managed and versioned, with changes reviewed like code
❌ Common L5 Trap

"Each collector looks at the spans it has and keeps traces with errors."

Why this misses: A collector seeing 1/40th of a trace can't know whether another part of the trace errored or whether the root was slow. Without trace-ID routing, tail decisions are made on fragments and produce half-kept traces.


Drill 3: Make It Concrete — Size It#

Prompt: "Size the pipeline and storage for 1M requests/s, 25 spans per request, 30 days."

Staff Answer

"Ingest: 25M spans/s × ~500 bytes = ~12.5GB/s raw into collectors; compressed on the wire ~3–4×, so ~3–4GB/s of network into the collector tier. Collectors at ~100–200K spans/s per core for parsing and routing need ~150–250 cores in the first tier.

After tail sampling at ~3%: 750K spans/s, ~375MB/s raw. Object storage, compressed ~8×: ~4TB/day, ~120TB for 30 days, ~$2.5K/month at ~$0.02/GB-month. The columnar search store with ~20 promoted attributes over 7 days: maybe 150 bytes per span compressed → ~10TB, on fast disks. Kafka buffer: 375MB/s × 24h ≈ 32TB raw, maybe 10TB compressed with replication factor 3 → ~30TB. Span metrics: series count is what matters — ~200 services × 50 operations × 3 statuses × 15 buckets ≈ 450K series. The biggest line items are collector compute and the tail-sampler memory, not storage."

Why this is L6:

  • Carries the numbers through every stage with explicit assumptions
  • Identifies compute and memory, not storage, as the cost drivers
  • Sizes span metrics by series count, not volume

What L7 adds:

  • Prices compute and memory per month and compares with vendor per-span pricing at the same volume
  • Identifies head-sampling health checks and duplicate proxy spans as the cheapest 20–30% reduction
❌ Common L5 Trap

"25M spans per second times 500 bytes is about 1PB per day, so we need a petabyte-scale cluster."

Why this misses: It sizes storage for 100% retention without asking whether keeping everything is the goal, and misses that compute and sampler memory dominate once sampling is applied.


Drill 4: The Collectors Are Down#

Prompt: "Your collector tier is completely unavailable for 30 minutes. What happens?"

Staff Answer

"Applications: nothing. The SDK exports asynchronously to the local agent on localhost, which is still up. The agent buffers ~200MB — a minute or two on busy nodes — then drops oldest, incrementing agent.dropped_spans. SDKs never block; if the agent itself is unreachable, the SDK queue of 2,048 spans fills and drops.

Tracing: a 30-minute gap, minus whatever agents buffered. Span metrics — computed in the collector — also gap, which matters more because dashboards and SLO burn alerts depend on them. So I'd keep a minimal fallback: services still emit their own RED metrics directly from the SDK at low cardinality, so alerting doesn't depend on the trace pipeline. Paging: collector unavailability pages the observability on-call even though no customer is affected, because the next incident needs it. Afterward: game day to confirm app p99 didn't move."

Why this is L6:

  • Shows the drop points and that none of them block the app
  • Notices span metrics share fate with collectors and adds a fallback
  • Treats tracing outages as page-worthy despite no customer impact

What L7 adds:

  • Decides whether alerting may depend on the trace pipeline at all — a correlated-failure question across telemetry systems
  • Runs quarterly black-hole drills with app latency as the pass criterion
❌ Common L5 Trap

"Kafka buffers everything, so we don't lose data."

Why this misses: Kafka sits after the collectors in this design; with collectors down, nothing reaches it. And if the SDK is configured to retry or block, the app is affected — the answer doesn't consider the application at all.


Drill 5: Async Propagation#

Prompt: "An API call publishes an event to Kafka. A consumer processes it in a batch of 500 messages. Show me the trace."

Staff Answer

"Producer side: the instrumented producer creates a producer span as a child of the API request span and writes traceparent into the message headers. Consumer side: the batch contains messages from up to 500 different traces, so the consumer's processing span can't be a child of all of them. It starts a new trace with span links to each message's producer span. Per-message processing, if instrumented, can be child spans of the batch span, each with its own link.

In the UI, the API request's trace shows the producer span and a 'linked spans' section pointing to the batch trace; from the batch trace, you can jump back to any of the 500 originating requests. If the consumer handles one message at a time, it continues the producer's trace as a child — simpler, and accurate. The tail sampler needs a rule: if the consumer trace has an error, keep the linked producer traces too — otherwise the context is gone."

Why this is L6:

  • Uses links for fan-in and children for 1:1, with the reason
  • Considers how the UI navigates the relationship
  • Extends tail-sampling policy across links

What L7 adds:

  • Makes queue clients with propagation built-in the only supported clients, with a CI test
  • Measures broken-edge rate on every producer-consumer pair
❌ Common L5 Trap

"Pass the trace ID in the message and have the consumer use it as the parent."

Why this misses: With batching, there's no single parent — choosing the first message's trace misattributes the batch's work to one request and loses the other 499. And the asynchronous delay makes the original trace look like it took hours.


Drill 6: The Attribute Explosion#

Prompt: "A team added user_id and request_body as span attributes. What breaks, and how do you prevent it?"

Staff Answer

"Three things. Storage and ingest: request_body can take span size from ~500 bytes to many kilobytes — 10× ingest and storage for that service. Metrics: if user_id leaks into span-metric dimensions, series count jumps by the number of users — millions of new series, which can take down the metrics backend. Privacy: request bodies and user IDs are personal data now stored for 30 days and readable by every engineer.

Prevention in the collector, not in team discipline: span-metric dimensions come from an allow-list — service, operation, status, region, version — so unknown attributes never become dimensions. A redaction processor drops or hashes known PII keys and patterns, and enforces a per-attribute size cap (e.g., 1KB) and a per-span attribute count cap. New attributes promoted to search require a cardinality and privacy review. And bytes-per-span per service is on a dashboard with an alert on jumps."

Why this is L6:

  • Separates storage, cardinality and privacy impacts
  • Enforces in collectors with allow-lists and caps
  • Adds detection for the next time

What L7 adds:

  • Makes the attribute policy part of the telemetry standard with security co-ownership
  • Ties attribute cost to the team's span budget so the choice is visible
❌ Common L5 Trap

"Ask teams not to add sensitive or high-cardinality attributes."

Why this misses: Guidance without enforcement fails at the first deadline. The collector is the one place every span passes through — that's where policy belongs.


Drill 7: Build vs Buy#

Prompt: "Build our own tracing backend, run an open-source one, or use a vendor?"

Staff Answer

"First decision regardless: instrument with OpenTelemetry and propagate W3C Trace Context, so the backend is swappable. Then it's cost and control. A vendor is fastest and best for small-to-mid volume; per-span or per-GB pricing becomes the problem at scale — at hundreds of thousands of kept spans per second, the bill can exceed a platform team's cost. Running an open-source backend on object storage gives low storage cost with some operational burden — a few engineers. Building from scratch is only justified with unusual requirements at very large scale.

My default: OpenTelemetry collectors that we run — so sampling, redaction and routing are ours regardless — feeding a vendor while volume is moderate, and moving to a self-run backend when the vendor bill crosses roughly the cost of 3–4 engineers. Owning the collector tier means switching backends is a config change."

Why this is L6:

  • Separates instrumentation (long-lived) from backend (swappable)
  • Gives a cost threshold for switching
  • Keeps sampling and redaction under our control in all options

What L7 adds:

  • Projects 3-year volume growth into the vendor-vs-self-run decision
  • Negotiates data export guarantees and avoids vendor-specific instrumentation
❌ Common L5 Trap

"Use the vendor's agent and SDK — it's easiest."

Why this misses: Vendor-specific instrumentation across hundreds of services is the real lock-in; switching later means re-instrumenting everything. And sampling and redaction policy end up in the vendor's control.


Drill 8: Changing Sampling Policy Without an Outage#

Prompt: "You want to change the tail-sampling baseline from 1% to 0.3% and add a rule to keep all traces for a new payments route. How do you ship it?"

Staff Answer

"Sampling policy is versioned config served by the policy service, not baked into binaries. I'd write the change, then shadow-evaluate it: run the new policy in parallel on the samplers, logging its decisions without acting on them, for 24 hours. Compare: kept-trace volume (should drop ~60% of baseline), error and slow trace retention (must be unchanged — that's the invariant), and payments route coverage (should be 100%).

Then roll out to samplers in waves — one zone, then 25%, then 100% — watching tailsampler.kept_rate, storage write rate and a synthetic check that error traces still appear. Rollback is re-activating the previous version. Teams are told a week ahead, because a baseline cut changes what they'll find when browsing non-error traces."

Why this is L6:

  • Policy as versioned config with shadow evaluation
  • States the invariant (error/slow retention unchanged)
  • Staged rollout with rollback and communication

What L7 adds:

  • Ties baseline changes to the budget process, with teams able to buy back baseline for their services
  • Reviews incident trace coverage after the change as the true success test
❌ Common L5 Trap

"Update the sampling rate in each service's config and redeploy."

Why this misses: Redeploying hundreds of services to change a rate is slow and risky, and head-rate changes don't affect tail policy. Without shadow evaluation, nobody knows if error capture changed until an incident.


Drill 9: Cut the Bill in Half#

Prompt: "Finance wants tracing costs cut by 50% this quarter. Where do you look?"

Staff Answer

"Start with where cost actually goes: collector compute, sampler memory, columnar search storage, then object storage. Levers in order of value-per-risk. One: drop spans that add nothing — health checks, readiness probes, duplicate proxy spans that mirror application client/server spans; often 20–30% of volume. Two: head-sample very high-volume, low-value internal routes before they reach collectors — saves collector compute and sampler memory. Three: reduce baseline tail rate from 1% to 0.3% — errors and slow traces unaffected. Four: shorten the searchable window from 7 to 3 days and promote fewer attributes; keep 30-day trace-ID lookup in object storage because it's cheap. Five: cap attribute sizes.

The line I won't cross: error and slow-trace capture. I'd show finance the plan with each lever's saving and confirm incident trace coverage afterward."

Why this is L6:

  • Starts from the cost breakdown, not guesses
  • Orders levers by value and risk, with estimated impact
  • Protects the invariant that makes tracing useful

What L7 adds:

  • Introduces per-team budgets so cost stays controlled after this quarter
  • Reframes the conversation around cost per incident-minute saved
❌ Common L5 Trap

"Cut the sampling rate in half everywhere."

Why this misses: A uniform cut reduces error and slow-trace capture along with everything else — the traces that justify the system — while leaving the waste (health checks, duplicate spans) untouched.


Drill 10: Multi-Region#

Prompt: "We run in three regions. Some requests cross regions. Design tracing for that."

Staff Answer

"Region-local pipelines: agents, collectors, samplers, Kafka and storage in each region, so tracing traffic doesn't cross regions and a regional outage doesn't take down tracing elsewhere. Spans for a cross-region request land in two regions' samplers — so tail decisions can disagree. Two mitigations: errors and slow traces trigger 'keep' in each region independently, which usually agrees because the root's latency is visible in the root region and errors propagate back as error statuses on client spans; and for the probabilistic baseline, the decision is a hash of the trace ID, so every region makes the same choice.

At query time, the query service fans out by trace ID to all regions and merges — cross-region lookups take a bit longer and that's fine. Data residency: if EU traces can't leave the EU, the query service fetches remotely but never replicates EU spans out."

Why this is L6:

  • Region-local pipelines for cost and failure isolation
  • Recognizes split tail decisions and uses hash-based determinism
  • Query-time federation instead of replication; residency respected

What L7 adds:

  • Decides which regions' tracing failures are acceptable during a regional outage — the trace you need is often in the failing region
  • Prices cross-region query egress and sets per-region retention by regulatory needs
❌ Common L5 Trap

"Send all spans to a central region for processing."

Why this misses: Cross-region egress at gigabytes per second is expensive, a central-region outage blinds every region at once, and it may violate data residency.


8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Traces Vanish During Black Friday#

Context: During the biggest sales day of the year, checkout p99 doubles. The on-call opens the tracing UI: traces from the last 20 minutes are sparse, and most error traces are missing. The on-call escalates to you — both for the checkout problem and for tracing.

Questions to Surface First:

  • Is the pipeline dropping (where — SDK, agent, collector, sampler) or lagging (storage writers behind)?
  • Is the tail sampler evicting traces before deciding?
  • Did spans per trace change? (Retries during slowdowns inflate traces.)
  • Are span metrics intact, so the checkout investigation can proceed with exemplars?

Typical L5 Approach: Scales up the storage cluster. Storage wasn't the bottleneck — the tail samplers were evicting traces under memory pressure — so nothing improves.

Staff Approach: Reads drop counters by stage, finds traces_evicted_before_decision climbing, applies the emergency policy (baseline to 0%, errors and slow only), scales samplers, and confirms the early-keep-on-error rule is active. Error traces reappear within minutes; the checkout investigation continues.

Principal Approach: Treats it as a capacity-planning failure for a known peak: sampler memory must be sized for incident-shaped traffic on peak days. Adds peak-day pre-scaling, a pre-approved emergency policy runbook, and a game day before every major sales event.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Stage drop counters: SDK 0, agent 0.1%, collector 0, sampler evictions 40K/s. Apply emergency policy: baseline 0%, keep errors + slow.
TriageSpans per trace p99: 40 → 180 (retry storms in checkout). Decision wait extended. Sampler memory at 98%.
Quick fixScale sampler tier 2×; confirm early-keep on error spans; cap spans per trace at 2,000.
GuardrailsAlert on traces_evicted_before_decision > 0; sampler memory headroom target 3× steady state on peak days.
Post-mortemWhy was the sampler sized for steady state? Why did the emergency policy require a human? Automate: memory > 85% → baseline to 0.

Metrics to Watch: tailsampler.memory_utilization, tailsampler.traces_evicted_before_decision, tailsampler.spans_per_trace_p99, trace.lookup_success_rate

Organizational Follow-up: observability joins peak-event readiness reviews alongside checkout and payments.

Ownership Question: "Who decides to drop baseline sampling during an incident?" Staff answer: The observability on-call, via a pre-approved runbook — and the sampler should do it automatically above a memory threshold.

Key Takeaway: "Traces get bigger exactly when you need them. Size the tail sampler for incidents, not for Tuesdays."

What clears the Staff bar:

  • Localizes the loss to the sampler before scaling anything
  • Protects error capture with an emergency policy
  • Automates the policy so it doesn't depend on a human at 3am

Deep Dive 2: Silent Failure — Six Weeks of Fragmented Traces#

Context: In a post-mortem, an engineer mentions that traces for the order pipeline "always stop at the publish step". Nobody knew how long this had been true. You're asked to investigate.

Questions to Surface First:

  • When did consumer spans stop being linked to producers?
  • Which producer-consumer pairs are affected — one, or many?
  • What changed: client library, framework, message format?
  • Do we measure broken-edge rates at all?

Typical L5 Approach: Fixes propagation in the order consumer. Doesn't check other pipelines.

Staff Approach: Computes broken-edge rate across all producer-consumer pairs from stored spans, finds 11 pairs broken since a messaging library upgrade six weeks ago, patches the library, and adds a propagation test to its CI.

Principal Approach: Makes propagation coverage a tracked platform SLO: broken-edge rate per edge published weekly, with library upgrades gated on propagation tests. Treats "traces silently fragmented for six weeks" as a monitoring gap in the platform itself.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Query consumer spans without parent or links grouped by (producer service, consumer service). 11 pairs > 50% broken.
TriageAll 11 use messaging library v3, which dropped header-based context injection in a refactor.
Quick fixPatch library to inject and extract traceparent; release; teams upgrade.
Guardrailstrace.broken_edge_rate{edge} computed hourly; alert when any edge with > 100 req/s exceeds 5%.
Post-mortemLibrary CI now includes produce → consume → assert-linked test; platform reviews broken edges weekly.

Metrics to Watch: trace.broken_edge_rate{edge}, trace.orphan_root_spans_pct{service}, trace.avg_services_per_trace

Organizational Follow-up: messaging library owners and observability platform agree on a propagation contract with tests.

Ownership Question: "Who owns propagation across a queue?" Staff answer: The messaging library owner implements it; the observability platform detects breakage and owns the test harness.

Key Takeaway: "Propagation fails silently. If you don't measure broken edges, you'll learn about them in a post-mortem."

What clears the Staff bar:

  • Measures the problem fleet-wide instead of fixing one pair
  • Finds the common cause in a shared library
  • Adds detection and a CI contract

Deep Dive 3: Large-Customer Onboarding — A Platform Team Wants 100% Traces#

Context: The payments platform team, newly onboarded to tracing, asks for 100% trace retention for all their services "for compliance and debugging". Their services produce 2M spans/s — 8% of total volume.

Questions to Surface First:

  • What compliance requirement exactly? Is a sampled, 30-day trace the right evidence?
  • What debugging do they do that errors-and-slow plus a baseline doesn't cover?
  • What's the cost of 100% for 30 days at their volume?
  • Could a business event log meet the compliance need instead?

Typical L5 Approach: Sets their sampling rate to 100%. Collector, sampler and storage costs jump ~40%; ingest lags for everyone.

Staff Approach: Separates the two needs. Compliance: business events with correlation IDs in a durable log with years of retention — not traces. Debugging: errors and slow kept at 100% already; raise their baseline to 10% within their budget; enable per-request debug flags so they can force-keep specific flows.

Principal Approach: Uses the request to formalize tiers: a documented "high-retention" tier with a price, and a clear statement in the telemetry standard that traces are not audit records.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (requirements)Clarify: compliance needs every payment's lifecycle for 7 years — that's an event log.
TriageCost of 100% at 2M spans/s: ~86TB/day raw, plus sampler and columnar costs — roughly doubling tracing spend.
Quick fixBaseline 10% for payments; debug flag via header for targeted flows; errors/slow already 100%.
GuardrailsPayments' span budget and monthly cost visible; any increase above budget needs director approval.
Post-mortem (pre-mortem)Confirm with compliance that the event log meets the requirement in writing.

Metrics to Watch: spans.kept_rate{team=payments}, cost.monthly{team=payments}, trace.lookup_success_rate{team=payments}

Organizational Follow-up: compliance documents that business event logs, not traces, are records of evidence.

Ownership Question: "Who pays for higher retention?" Staff answer: The requesting team, from its telemetry budget; errors and slow traces stay platform-funded for everyone.

Key Takeaway: "When someone asks for 100% traces, find out which question they're trying to answer — it's usually an audit question or a debug question, and neither needs 100%."

What clears the Staff bar:

  • Separates audit from debugging and redirects audit to the right system
  • Quantifies the cost of the request
  • Offers targeted alternatives within a budget

Deep Dive 4: Post-Mortem — Customer Emails in Trace Storage#

Context: A security audit finds customer email addresses and partial payment details in span attributes, retained for 30 days and readable by all engineers. You're leading the post-mortem.

Questions to Surface First:

  • Which services emit the PII, and since when?
  • Which stores hold it — object storage, columnar, Kafka, span metrics?
  • Who accessed it?
  • Why didn't any control catch it?

Typical L5 Approach: Asks the offending teams to remove the attributes and deletes the affected indexes.

Staff Approach: Adds collector-side redaction (key deny-list and value patterns for emails, card numbers, tokens), purges affected blocks from object storage and the columnar store, shortens Kafka retention for the affected window, restricts trace access to role-based groups, and adds a PII scanner on a sample of incoming spans.

Principal Approach: Makes telemetry a governed data class: security co-owns the collector policy, access is audited, and new attribute keys from any service are flagged automatically for review.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Deploy collector deny-list for the offending keys. Restrict trace UI access to on-call groups temporarily.
Triage3 services added request payload attributes for debugging; 19 days of data affected.
Quick fixPurge blocks containing affected spans (rewrite blocks without them); purge columnar partitions; age out Kafka.
GuardrailsPattern-based redaction for emails, card numbers, bearer tokens; attribute size caps; sampled PII scanner alerting security.
Post-mortemWhy could any service write any attribute to a store readable by everyone? Add review for new attribute keys.

Metrics to Watch: pii.scanner_hits, collector.redacted_attributes_total, trace.ui_access{role}

Organizational Follow-up: privacy team assesses notification obligations; engineering guidance updated.

Ownership Question: "Who owns PII in telemetry?" Staff answer: Security owns the policy; the observability platform enforces it in collectors; service teams own not emitting it.

Key Takeaway: "Every span attribute is data you're storing. Enforce redaction where all spans pass — the collector."

What clears the Staff bar:

  • Enforces in the collector, not by asking teams
  • Purges every store, including buffers
  • Adds detection and access controls

Deep Dive 5: Multi-Region Expansion — Adding an EU Region With Residency#

Context: The company adds an EU region with strict data residency. Some requests from EU users call shared US services. Tracing is currently centralized in the US.

Questions to Surface First:

  • Do span attributes for EU requests count as personal data?
  • Which services handle EU requests in the US, and what do their spans contain?
  • Can engineers in the US view EU traces?
  • How do we keep cross-region traces debuggable?

Typical L5 Approach: Deploys an EU collector that forwards to the central US pipeline.

Staff Approach: Region-local pipelines and storage in the EU. EU-origin traces stay in the EU; US services handling EU requests strip personal attributes before export (tenant-region tag drives redaction). The query service federates by trace ID across regions with access control on EU data.

Principal Approach: Defines telemetry residency as part of the company's data classification: region tag on every span, redaction rules by region, and audit of where telemetry is stored — applied to traces, logs and metrics together.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Inventory attributes emitted for EU requests; classify personal data.
TriageCentralized pipeline would export EU personal data to the US.
Quick fixEU collectors, samplers and storage; residency tag in tracestate drives region-aware redaction in US collectors.
GuardrailsMarker test: synthetic EU trace with marker attributes; scan US stores for markers daily.
Post-mortem (pre-launch)Confirm query federation respects access controls; test EU-region outage behavior.

Metrics to Watch: residency.marker_found_outside_region (must be 0), query.federated_latency_p99, collector.region_redactions_total

Organizational Follow-up: legal approves the data inventory; access groups for EU telemetry defined.

Ownership Question: "Who proves telemetry residency?" Staff answer: The observability platform proves it with marker tests; compliance owns the attestation.

Key Takeaway: "Traces are data. Residency rules apply to telemetry the same as to databases — and spans travel across regions with the requests."

What clears the Staff bar:

  • Region-local pipelines and storage
  • Redaction driven by a propagated region tag
  • Proves residency with tests

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why uniform head sampling misses the traces that matter, and design metrics-from-100% plus tail sampling
  • Describe tail-sampling mechanics: trace-ID routing, decision wait, early keep on errors, late spans, memory sizing
  • Propagate context across HTTP, gRPC, queues, batches and thread pools — using span links for fan-in — and measure broken edges
  • Design a pipeline that is async, bounded and lossy before sampling and durable after, so tracing never affects applications
  • Size ingest, sampler memory and storage, and identify compute and memory as the cost drivers
  • Choose trace-ID-first object storage plus a short-window columnar search store with promoted attributes
  • Enforce attribute policy (cardinality, PII, size) in collectors
  • Assign ownership: platform pipeline and policy, team instrumentation and budgets, security PII policy

The Bar for This Question#

Mid-level (L4): Instruments services, ships spans to a backend and views traces. Doesn't consider sampling beyond a fixed rate, async propagation or pipeline failure.

Senior (L5): Designs a reasonable pipeline — SDK, agent, Kafka, storage — with head sampling and a searchable store. The gap: uniform sampling that misses errors; blocking or unbounded export; indexing everything; traces that break at queues. The design works in staging and disappoints the first on-call who needs it.

Staff+ (L6): Frames tracing around keeping the valuable fraction and never harming the application. Computes metrics from all spans, tail-samples errors and slow traces with trace-ID routing, makes every hop bounded and lossy, measures propagation, and stores for the trace-ID access pattern. Names who pays — the platform for sampler memory, teams for extra retention, nobody for error capture. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 "Sample 1%" Is a Way of Saying "We'll Miss the Incident"#

ClaimReality
"1% is statistically enough"For averages, yes; for a 0.1% error class, you keep 1 in 100,000
"We'll increase it during incidents"The incident's first 20 minutes are already gone
"Traces are for exploration"Their value is concentrated in incidents

The Staff position: Metrics from 100%, traces for the errors and the slow, and a small baseline for exploration.

Why this matters in interviews: It's the clearest single signal of whether a candidate has used tracing during an incident.

10.2 Most Tracing Cost Is Self-Inflicted Waste#

SourceTypical Share
Health checks and probesOften 10–20% of spans
Duplicate proxy spansUp to 2× spans per hop with a mesh
Unbounded attributesLarge multiplier on bytes per span

The Staff position: Remove waste before reducing sampling of real traffic.

Why this matters in interviews: Cost questions are judgment questions; cutting error capture first is the wrong answer.

10.3 Tracing Should Be Allowed to Fail — Loudly#

ComponentPosture
SDK exportDrop, never block
AgentBuffer briefly, drop oldest
CollectorsShed by sampling harder
StorageLag safely behind Kafka

The Staff position: Tracing is a best-effort system that pages its own team. It never shares fate with the application.

Why this matters in interviews: Availability-minded candidates sometimes over-engineer durability into the wrong hop.

10.4 Index Less, Retain More#

StoreCost Profile
Object storage by trace IDPennies per GB-month
Columnar with promoted attributesModerate
Full inverted index on everythingOften the largest line item

The Staff position: Most investigations start from a trace ID. Optimize for that, and pay for search over a short window.

Why this matters in interviews: It shows storage choices driven by access patterns, not habit.

10.5 Instrumentation Is the Asset; the Backend Is a Commodity#

The Staff position: Standardize on OpenTelemetry and W3C Trace Context and own your collectors. Backends come and go; re-instrumenting 500 services is what you never want to do twice.

Why this matters in interviews: It reframes build-vs-buy around the one-way door.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a tracing system that keeps the right traces and never hurts the application. The Principal engineer notices that the company runs three telemetry pipelines — one for traces, one for metrics, one for logs — each with its own agents on every node, its own collectors, its own vendor contract and its own cost curve growing faster than traffic. The L7 problem is one telemetry platform: shared instrumentation, shared collectors that enforce sampling, redaction and routing for all signals, correlation by trace ID across signals, and a budget model that keeps total observability spend at a deliberate fraction of infrastructure cost.

The Org-Level Fault Line#

One telemetry platform vs per-signal, per-team stacks.

OptionWhat WorksWhat BreaksWho Pays
Per-team tracing choicesTeams pick what they likeTraces break at team boundaries; no shared propagation; N vendor billsOn-call (fragmented traces), finance (N contracts)
Central tracing platform, separate from metrics/logsConsistent tracingThree agents per node; no shared redaction; correlation is manualInfra (agent overhead), security (three PII surfaces)
Unified telemetry platformOne agent, one collector policy, correlated signalsBig migration; platform team becomes critical pathPlatform team (scope), teams (migration)

🧭 Principal Move: "One instrumentation standard and one collector tier for traces, metrics and logs. Collectors are where sampling, redaction, routing and cost attribution happen for every signal. Backends can differ per signal — and change — because the collector tier is ours."

Cost Model#

Assumptions: self-run collectors and samplers; object storage at ~$0.02/GB-month; columnar search on SSD; fully loaded engineer ~$250K/year; ~25 spans per request, ~500 bytes per span.

ScaleVolumeInfra ($/month)HeadcountOn-call LoadNotes
Startup2K req/s, 50K spans/s~$1–3K (vendor or small self-run; keep 100%)< 1 engShared rotationVendor usually cheaper; no tail sampling needed
Growth100K req/s, 2.5M spans/s~$20–50K (collectors, samplers, Kafka, ~10TB object, columnar)3–5 engDedicated rotation, 1–3 pages/monthTail sampling and attribute policy become essential
Enterprise1M+ req/s, 25M+ spans/s, 3 regions~$150–400K10–15 eng (telemetry platform)Per-region rotationsCollector compute and sampler memory dominate; storage is minor

The pricing insight: at enterprise scale, storage is the smallest line; collector CPU and tail-sampler memory are the largest. Dropping waste spans (probes, duplicate proxy spans) at the agent before they reach collectors often saves more than any storage optimization.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Instrumentation standard (OpenTelemetry vs vendor SDK)One-wayRe-instrumenting hundreds of services
Propagation format (W3C Trace Context)One-wayEvery service and every client library
Owning the collector tierOne-way-ishWithout it, sampling and redaction live in a vendor
Trace backendTwo-way (if collectors are ours)Config change + dual-write period
Sampling policyTwo-wayVersioned config
Retention and search windowsTwo-wayCost change
Promoted attributesTwo-waySchema change in columnar store

The Standard I'd Write#

RFC-OBS-003: Telemetry Standard (Traces)
Status: Approved   Owners: Observability Platform + Security

Scope
  Every service that handles requests, messages or scheduled work in production.

MUST
  1. Instrument with the approved OpenTelemetry distribution; propagate W3C Trace
     Context on every inbound and outbound call, message and task.
  2. Export asynchronously with bounded queues; telemetry export MUST NOT block
     request processing or affect availability.
  3. Never emit secrets, credentials, payment data or raw request bodies as attributes.
  4. Use approved shared clients for RPC, HTTP, database and messaging.
  5. Treat traces as diagnostic data, not records of evidence.

SHOULD
  1. Use span links for batch and fan-in consumers.
  2. Keep custom attributes within the team's span budget.
  3. Request promoted search attributes through cardinality review.

Exceptions
  Filed with Observability Platform; security review for any exception to MUST 3.

Success metrics
  - Incident trace coverage (relevant trace found within 5 min): ≥ 90%
  - Broken-edge rate on edges > 100 req/s: < 1%
  - Application latency impact of telemetry pipeline outage: 0 (verified quarterly)
  - Telemetry spend: ≤ agreed % of infrastructure spend

What I'd Tell the VP#

"Our tracing helps when it works, but during the last three major incidents, engineers couldn't find the relevant traces in two of them, and our observability spend is growing twice as fast as traffic. I'm proposing we run our own collection tier that keeps every error and slow request while storing only a few percent of normal traffic, and that we give each team a visible telemetry budget. It costs about four engineers for a year, should cut tracing spend by 30–50%, and makes incident debugging faster. The main risk is migration across 400 services; shared libraries do most of the work."

Principal Interview Signals#

SignalWhat It Sounds Like
Prices compute, not just storage"Storage is the smallest line; collector CPU and sampler memory dominate."
Identifies one-way doors"Instrumentation and propagation are forever; the backend isn't."
Redraws ownership"Platform owns collectors and policy; teams own instrumentation and budgets; security owns PII rules."
Measures value, not volume"Incident trace coverage is the KPI; spans stored is a cost."
Unifies across signals"One collector tier for traces, metrics and logs — one place for redaction and sampling."

Staff answers that L7 interviewers find insufficient:

  • "We'll build excellent tail sampling" — correct, but ignores that metrics and logs pipelines duplicate the same infrastructure.
  • "Teams see their costs" — visibility without a budget process doesn't change behavior.
  • "We'll use OpenTelemetry" — right, but no plan for owning collectors, which is where policy lives.

Appendices

Appendix A: Mechanics in Depth#

A.1 Tail Sampling Policy#

on_span(span):
  t = buffer[span.trace_id] (create if absent, evict LRU if over num_traces)
  t.add(span); t.last_seen = now
  if span.status == ERROR and not t.decided: decide(t, keep=True, reason="error")
  if t.span_count > 2000: t.truncated = True        # stop buffering, keep marker

every 1s:
  for t in buffer where not t.decided and now - t.last_seen > decision_wait (30s):
      root = t.root_span()
      if root and root.duration > slo[root.service, root.name]: decide(t, True, "slow")
      elif debug_flag(t):                                     decide(t, True, "debug")
      elif hash(t.trace_id) < baseline_rate:                  decide(t, True, "baseline")
      else:                                                   decide(t, False, "dropped")

decide(t, keep, reason):
  t.decided = True
  if keep: emit(t.spans, sample_rate = 1.0 if reason != "baseline" else baseline_rate)
  decisions_cache.put(t.trace_id, keep, ttl=5m)       # late spans follow the decision
  free(t.spans)

A.2 Count Re-Weighting#

A trace kept at probability p represents 1/p requests. When computing counts from stored traces (not recommended for SLOs, but useful for exploration), weight each by 1/p: an error trace kept at 1.0 counts once; a baseline trace kept at 0.01 counts 100 times. Span metrics computed before sampling make this unnecessary for dashboards.

A.3 Clock Skew Adjustment (display time)#

for each child c of parent p on a different host:
  if c.start < p.start or c.end > p.end:
      offset = clamp_to_fit(c, p)          # minimal shift so c fits inside p
      shift subtree(c) by offset; mark adjusted
never modify stored timestamps

Appendix B: Data Model#

Span (OTLP-like):
  trace_id        bytes(16)
  span_id         bytes(8)
  parent_span_id  bytes(8) | empty
  name            string
  kind            SERVER | CLIENT | PRODUCER | CONSUMER | INTERNAL
  start_ns, end_ns
  status          UNSET | OK | ERROR, message
  attributes      map<string, scalar>    -- allow-listed for metrics, capped size
  events          [ (time, name, attributes) ]
  links           [ (trace_id, span_id, attributes) ]
  resource        service.name, service.version, deployment.region, host

Object storage layout:
  /tenant/day/block-<ulid>/data.parquet     -- spans sorted by trace_id
                          /index            -- trace_id ranges per row group
                          /bloom            -- bloom filter on trace_id
  block size ~100MB-1GB; compaction merges small blocks hourly

Columnar search table (7d):
  (start_time, service, operation, status, duration_ms, region, version, tenant,
   http.route, error.type, trace_id, + ~10 promoted attributes)
  partitioned by hour, sorted by (service, start_time)

Appendix C: Coordination Mechanisms#

C.1 Trace-ID-Affinity Routing#

Diagram: C.1 Trace-ID-Affinity Routing

C.2 Quick Comparison#

MechanismGuaranteesFailure ModeUse For
Head sampling by trace-ID hashConsistent decision across servicesOutcome-blindBaseline, high-volume routes
Tail sampling with affinity routingKeeps errors/slow tracesMemory overflow; split on rebalanceInteresting traces
Collection-time hash samplingGlobal write-rate knobUniformEmergency volume control
Span metrics before samplingAccurate aggregatesCardinality explosionDashboards, SLOs, exemplars
Bounded drop-on-full queuesApp never blocksData loss during incidentsSDK and agent
Kafka after samplingStorage can lag safelyAnother system to runKept traces
Span linksCorrect fan-in semanticsUI support neededBatch consumers
Collector redactionPII and size controlRules incompleteAll spans

Appendix D: API Contract & Client Behavior#

  • SDKs extract traceparent/tracestate from inbound requests and messages and inject them on every outbound call; if absent, start a new trace.
  • The sampled flag in traceparent governs whether downstream services export spans; downstream services must not override a sampled decision to "not sampled" (they may record locally for metrics).
  • Debug override: a signed request header (internal only) sets a debug flag in tracestate, forcing keep at the tail sampler.
  • SDK defaults: batch export every 5s or 512 spans, queue 2,048, drop on full, export timeout 10s, all asynchronous.
  • Query API: trace by ID (all regions federated); search over the columnar window; dependencies from span metrics.

Appendix E: Observability#

Core metrics:

  • sdk.export_queue_full_total, sdk.export_blocked_ms (must be 0)
  • agent.buffer_utilization, agent.dropped_spans_total
  • collector.accepted_spans_rate, collector.cpu_utilization
  • tailsampler.memory_utilization, tailsampler.traces_evicted_before_decision, tailsampler.kept_rate{reason}
  • trace.broken_edge_rate{edge}, trace.lookup_success_rate, synthetic.trace_visible_seconds
  • spanmetrics.series_count, spans.bytes_per_span_p99{service}, cost.monthly{team}

Critical alerts:

AlertThresholdSeverity
SDK export blocked> 0 ms on any serviceSev-2 page (app impact risk)
Synthetic trace not visible> 120sPage
Traces evicted before decision> 0 sustained 5 minPage
Collector accepted rate drop> 30% vs baselinePage
Span metrics series jump> 2× in 1hPage
PII scanner hit> 0Page security + platform
Broken-edge rate> 5% on edge > 100 req/sTicket to owning team

Debugging the silent failure: tracing failures are discovered during other incidents. Synthetic traces (known trace ID every minute, assert queryable within 60s), broken-edge rates and eviction counters are how you find them first.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 50K spans/sKeep 100%; single backend; vendor or small self-runCost as services multiply
50K–2M spans/sOwn collectors; span metrics; tail sampling; attribute policySampler memory; columnar cost
2M–25M spans/sAffinity-routed sampler ring; object storage + columnar; per-team budgetsMulti-region; residency; agent overhead
> 25M spans/sRegion-local platforms; unified telemetry collectors; waste dropped at agentsOrg governance and budgets

What you don't build on day one: tail sampling, self-run storage, per-team budgets, multi-region federation, unified telemetry. Each has a trigger in Section 11.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Per-team span budgets for extras (baseline above default, custom attributes, longer search); errors and slow traces always platform-funded.
  • Fair sharing in collectors: when shedding, shed the highest over-budget services' baseline spans first, never error spans.
  • Customer-facing tracing (if offered as a product): tenant ID as a first-class field, per-tenant ingest quotas, isolated storage prefixes and query limits; price retention tiers.
  • Cost attribution: spans and bytes metered by service.name → team mapping; monthly report per team.
  • Waste controls at the agent: drop probe and health-check spans; collapse duplicate proxy spans; cap attribute sizes — the cheapest savings are spans that never leave the node.
  1. Loading the index…