Hiring BarSupport

Design LeetCode (Code Execution) — Staff-Level Case Study

Case study70 min read8 diagrams

Technologies referenced in this case study: Apache Kafka · Redis · PostgreSQL · API Gateways

Related case studies and patterns: Distributed Job Scheduler · Message Queues · Rate Limiting · Leaderboard · Auto-Scaling & Capacity · Long-Running Processes · Degraded Mode · Real-time Updates

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to the sections that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 5
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is an Online Judge? — Why interviewers pick this topic

An online judge accepts source code from anonymous or lightly-authenticated users, compiles it, runs it against hidden test cases under strict CPU/memory/time limits, and returns a verdict: Accepted, Wrong Answer, Time Limit Exceeded (TLE), Memory Limit Exceeded (MLE), Runtime Error, Compile Error. LeetCode, Codeforces, HackerRank, AtCoder and every coding-assessment vendor run a version of this.

It is the only common interview prompt where the input to your system is hostile code by design. Every other system treats malicious input as an edge case. Here it is the workload.

Before vs After — Weekly Contest scenario (30K contestants, 90 minutes):

Without a Staff-level design:
t=0:        Contest opens. 30,000 users read problem 1.
t=+4min:    First wave of submissions — 900 submits/s vs a 150/s steady state.
t=+5min:    Shared FIFO queue depth hits 60,000. Practice users are in the same queue.
t=+9min:    p50 verdict latency = 7 minutes. Users resubmit, doubling queue depth.
t=+12min:   Autoscaler adds hosts; cold container images take 90s each to pull.
t=+20min:   Noisy neighbors inflate wall-clock time; 8% of correct solutions get TLE.
t=+95min:   Contest ends. Leaderboard is wrong. Rejudge of 40K submissions takes 3 hours.
t=+1 day:   Contest declared unrated. Forum thread titled "judge is broken again".

With a Staff-level design:
t=-30min:   Contest pool pre-warmed to 3× forecast peak from registration count.
t=0:        Same 30,000 users, same spike.
t=+4min:    Contest lane absorbs 900/s. Practice lane throttled to 50% capacity.
t=+5min:    p95 verdict latency 6s. Penalty time stamped at ingress, not at verdict.
t=+20min:   Timing measured as cgroup CPU time on pinned cores. False-TLE rate < 0.1%.
t=+95min:   Leaderboard final within 2 minutes. Zero rejudges needed.

Why interviewers reach for this question: It compresses four hard problems into one prompt — untrusted-code isolation, bursty queueing with fairness, deterministic measurement on shared hardware, and abuse economics. An L5 candidate designs a queue and a worker. A Staff candidate designs the trust boundary, the capacity plan for a known spike, and the process for when the judge itself is wrong.

Mechanics Refresher: Isolation Options
Isolation MechanismHow It WorksStartupProsCons
Process + rlimitssetrlimit, chroot, run as nobody~5msTrivial, fastShares the kernel and filesystem view; one escape = host
Container (namespaces + cgroups + seccomp)Linux namespaces for pid/net/mount/user, cgroups v2 for CPU/mem/pids, seccomp-bpf syscall allowlist50–300ms warm, 1–3s cold imageMature tooling, dense packingShared host kernel — a kernel CVE is a sandbox escape
User-space kernel (gVisor)Intercepts syscalls in a user-space kernel (Sentry); host sees a tiny syscall surface~150–500msMuch smaller host attack surface, OCI-compatible2–10× overhead on syscall-heavy code; some syscalls unsupported
MicroVM (Firecracker, Cloud Hypervisor)KVM-based minimal VM, own guest kernel~125ms boot, <5 MiB overheadHardware virtualization boundaryNeeds KVM (bare metal or nested virt); more ops complexity
Language VM (V8 isolates, WASM)Runs code inside a language runtime sandbox<5msExtremely dense and fastOnly works for languages that compile to that runtime — not C++/Java/Python native

For most production judges: containers with seccomp + cgroups v2 inside a stronger boundary (gVisor or Firecracker microVMs per worker host or per tenant). Defense in depth — the isolation question is never "which one," it is "how many layers, and what does each one stop."


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

Designing LeetCode is not a queue-and-worker question. Everyone draws a queue and a worker.

It is a trust-boundary and burst-capacity question that tests:

  • Whether you treat submitted code as an attacker, not a user
  • Whether you plan capacity for a spike you can see coming (contests are scheduled)
  • Whether you separate fairness of the verdict from latency of the verdict
  • Whether you own what happens when the judge itself is wrong — bad test data, flaky timing, mass rejudge

The key insight: The judge's product is a trustworthy verdict, not fast execution. A 6-second Accepted that users trust beats a 1-second verdict that is wrong 2% of the time. Every design decision — isolation, timing, queueing — should be evaluated by what it does to verdict trust.

The L5 → L6 → L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws API → queue → Docker worker → DBAsks "is this practice, contest, or assessment? Contest changes everything"Asks "is code execution one product or a platform that practice, contests, assessments and the AI tutor will all consume?"
Isolation"Run it in Docker with no network"Layers seccomp + cgroups + no-network + read-only FS inside a microVM or gVisor boundary; names the kernel-CVE threatOwns the org's sandbox posture: one hardened runtime, patch SLA for kernel CVEs, red-team budget, bug bounty scope
Spikes"Autoscale the workers"Pre-warms to forecast from contest registrations; separate contest and practice lanes; stamps penalty at ingressPrices the idle warm pool vs lost contest credibility; negotiates reserved capacity with the cloud vendor for a scheduled calendar
Timing"Enforce a 2-second timeout"Measures CPU time via cgroups on pinned cores; calibrates per-language multipliers; tracks false-TLE rateTreats timing determinism as an SLO with an owner; funds a dedicated contest hardware class
Failure"Retry failed jobs"Designs rejudge as a first-class workflow; idempotent verdicts keyed by (submission, testset version)Defines the policy for when a contest is unrated, who decides, and how that is communicated
Abuse"Rate limit submissions"Separate quotas for Run vs Submit; CPU-seconds budgets; detects mining, test-case exfiltration, plagiarismFrames abuse as unit economics: cost per free user per month vs conversion, and sets the policy accordingly
Why "first move" separates levels

L5: Starts drawing the pipeline immediately. The pipeline is correct — API, queue, workers, results. But it is designed for the average day, and the average day is not the hard part. The hard part is the scheduled contest where 30K users submit within the same 5-minute window, and every one of them cares about ranking fairness.

L6: "Before I draw anything: practice judging, contest judging, and the free-form Run button have different correctness bars and different traffic shapes. Practice tolerates a 10-second verdict. Contests need fairness under a 6× spike. Run is an abuse magnet. I'll design for contests because that's the hardest shape, and practice falls out of it."

L7: "And I'd note that code execution is going to be requested by at least three other teams within a year — assessments for enterprise hiring, an AI tutor that executes generated code, and notebooks. I'd build the sandbox as a platform with a contract, not as a feature of the judge."

Why "isolation" separates levels

L5: "Docker with network disabled." That stops the obvious attacks. It does not stop a container escape via a kernel vulnerability, because containers share the host kernel. On a public judge, you are running millions of arbitrary programs per day — if a kernel exploit exists, someone will submit it.

L6: Names the layers and what each stops: seccomp stops exotic syscalls, cgroups stop resource exhaustion, network namespace stops exfiltration, read-only rootfs stops persistence, and a VM boundary (Firecracker or gVisor) turns a kernel exploit from "host compromise" into "one throwaway VM compromise." Then states the cost: ~125ms extra boot and ~10–15% density loss.

L7: Recognizes this as an org-level security posture: who patches the guest kernel, what is the CVE response SLA (e.g., 72 hours for critical), and whether the sandbox gets an external audit before the enterprise assessment product launches.

Why "timing" separates levels

L5: Uses wall-clock timeouts. On a shared host, a correct O(n log n) solution that normally takes 900ms can take 1.4s when a neighbor saturates memory bandwidth. The user sees TLE on an accepted solution.

L6: Measures CPU time from cgroup cpu.stat, pins each sandbox to dedicated cores, disables SMT siblings for the contest pool, keeps a wall-clock limit only as a backstop (typically 3× the CPU limit), and continuously runs reference solutions as canaries to detect drift.

L7: Makes verdict determinism a published SLO — "a resubmission of identical code returns an identical verdict 99.9% of the time" — and owns the hardware class decision to hit it.

The Staff Positions#

PositionRationale
Assume every submission is an exploit attemptAt 1M+ submissions/day, the tail of the distribution is actively hostile; design the sandbox for the tail
Defense in depth: container inside a VM boundaryContainers alone share the kernel; the VM boundary converts a kernel CVE from a breach into an incident
Separate lanes for contest, practice, and RunOne queue means the contest spike starves practice or vice versa; lanes let you choose who waits
Stamp fairness at ingress, not at verdictPenalty time and ordering come from the submit timestamp, so queue lag cannot change rankings
CPU time on pinned cores, not wall timeWall time measures your neighbors; CPU time measures the user's code
Pre-warm for scheduled spikes; autoscale for unscheduledContests are on a calendar; reacting with a 90-second cold start is choosing to fail
Rejudge is a product feature, not an incidentTest data will be wrong; the system must re-run tens of thousands of submissions idempotently

The Four Intents#

Four intents drive incompatible designs. Name them and commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Practice JudgingSteady load, cost-sensitiveShared pool, autoscale on queue depth, best-effort latencySlow verdicts during peaksVerdict must be correct; latency p95 ≤ 10s is fine
Contest JudgingScheduled 5–10× spikes, ranking integrityDedicated pre-warmed pool, ingress timestamps, priority laneQueue backlog alters perceived fairness; false TLE changes ranksDeterministic verdicts; p95 ≤ 10s; zero ordering errors
Interactive Run (custom input)High volume, low value per call, abuse magnetTight CPU budget, aggressive rate limits, lowest priority laneCrypto mining, resource burnBest effort; can be shed first
Hiring Assessment (B2B)Contractual SLA, integrity, auditIsolated tenant pool, proctoring signals, full audit trailCandidate disputes, legal exposureEvery verdict reproducible and explainable months later

🎯 Staff Move: "I'll design for contest judging, because it forces every hard decision — a known spike, ranking fairness, and deterministic timing. Practice is a strict subset: the same pipeline with a lower priority lane. I'll call out where the Run button and B2B assessments need different policy, not different architecture."

The Five Fault Lines#

#Fault LineThe Tension
1Isolation Strength vs Startup LatencyStronger boundaries (microVMs) cost boot time and density; weaker ones (containers) share the kernel
2Warm Capacity vs CostA warm pool sized for the contest peak sits idle 95% of the week; a cold pool misses the spike
3Throughput vs Fairness Under SpikeFIFO maximizes simplicity; lanes and priorities protect contests but starve someone else
4Timing Determinism vs Packing DensityPinned cores and disabled SMT make verdicts stable but halve the sandboxes per host
5Judge Platform vs Content OwnershipWho owns a wrong verdict — the platform (runtime, timing) or the problem setter (tests, checker)? Who can trigger a rejudge?

In the Wild: Real Production Systems#

Why this section belongs here: Code execution isolation is a solved-in-public problem. Citing how the hyperscalers isolate untrusted code signals you know the design space is not "Docker vs not Docker."

AWS Lambda — Firecracker MicroVMs#

AWS built Firecracker, an open-source KVM-based virtual machine monitor, to run Lambda and Fargate workloads from mutually untrusting customers on shared hardware. The published design goals are boot times on the order of ~125ms and memory overhead under ~5 MiB per microVM, which makes "one VM per untrusted workload" economically viable where traditional VMs (seconds to boot, hundreds of MB overhead) were not.

Staff insight: Firecracker exists because AWS concluded containers alone were not an adequate boundary between tenants. If AWS will not trust namespaces + seccomp between paying customers, you should not trust it between anonymous internet users submitting C++.

Google gVisor — A User-Space Kernel#

Google open-sourced gVisor, which implements a large part of the Linux syscall interface in a user-space process ("the Sentry") so that sandboxed code never talks directly to the host kernel. It is used for Google's own multi-tenant serverless products. The tradeoff is overhead on syscall-heavy workloads and incomplete syscall coverage.

Staff insight: gVisor is the middle option on the isolation-vs-latency curve — OCI-compatible (so your container tooling works) with a much smaller host attack surface. For a judge where most programs are CPU-bound and do little I/O, the syscall overhead is mostly irrelevant, which makes gVisor unusually well-suited.

IOI / Competitive Programming Judges — isolate and CMS#

The International Olympiad in Informatics' Contest Management System (CMS) uses isolate, a small sandbox built on Linux namespaces and cgroups, designed specifically for measuring CPU time and memory of contestant programs precisely. Large public judges (Codeforces, AtCoder) have long been known among competitors for visible "In queue" backlogs during the busiest rounds — occasionally leading to extended or unrated contests.

Staff insight: Competitive programming judges optimized for measurement precision, not just containment. That's the fault line most candidates miss: a sandbox that is secure but gives a 20% timing variance is still a broken judge.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Run it in Docker""Containers share the kernel. What happens with a kernel exploit?"Threat model depth
"Autoscale on queue depth""Your cold start is 90 seconds and the contest spike lasts 5 minutes. Now what?"Capacity planning for known events
"2-second timeout""Same code gets AC then TLE on resubmit. Why, and how do you fix it?"Measurement vs containment
"Put submissions in a queue""Practice users submit during a contest. Who waits?"Fairness and priority policy
"We store the verdict""The test data was wrong. 40K submissions need rejudging. Walk me through it."Idempotency and operational ownership
"Rate limit per user""Someone is mining crypto via the Run button with 500 accounts."Abuse economics, not just limits

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The request path is deliberately boring — gateway, submission service, queue, dispatcher, judge host. The Staff content is in the annotations: ingress timestamps (fairness decoupled from lag), weighted lanes (who waits during a spike), microVMs on pinned cores (trust boundary + timing determinism), test data cached locally by content hash (no hot object-store reads during a contest), and a control plane that turns the contest calendar into pre-warmed capacity.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Isolation"Docker, no network""Container with seccomp, cgroups v2, no network, read-only root — inside a Firecracker microVM. The VM boundary is what makes a kernel CVE survivable."
Queueing"One queue, many workers""Three lanes — contest, practice, run — with weighted fair dispatch. Run is shed first."
Spikes"Autoscale""Pre-warm from the contest calendar and registration count. Autoscaling is for the unscheduled traffic."
Fairness"Process in order""Stamp ingress_ts at the submission service. Rankings use that, so queue lag can't reorder the leaderboard."
Timing"Timeout after 2s""CPU time from cgroups on pinned cores; wall time is a 3× backstop. Canary reference solutions detect drift."
Wrong tests"Fix the test and rerun""Rejudge is a workflow: new testset version, idempotent verdicts keyed by (submission_id, testset_version), low-priority lane, diff report before publish."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Firecracker microVM boot~125msMakes a VM-per-submission (or per-batch) boundary affordable
Firecracker memory overhead< 5 MiB per VMDensity stays close to containers
Cold container with image pull30–90sWhy autoscaling cannot catch a 5-minute contest spike
Warm sandbox handoff5–20msWhat a pre-warmed pool buys you
Typical per-test time limit1–2s CPUStandard competitive-programming budget
Typical memory limit256 MBSets per-host density: a 64-core, 256 GB host fits ~48 sandboxes with pinning
Test cases per problem50–150A submission is 50–150 executions, not one
Average CPU per judged submission~1.5–3 CPU-secondsIncludes compile (~0.5–1s for C++/Java)
Contest spike vs steady state5–10× in the first 10 min and last 10 minBimodal — plan for both peaks
Weekly contest scale (assumed)30K participants, ~150K submissions in 90 min~28 submits/s average, ~150–300/s peak
Run-to-Submit ratio~3–5 : 1Run is the bigger load and the lower-value one
Verdict latency target (contest)p95 ≤ 10sAbove ~30s, users resubmit and double the load
False-TLE canary rate target< 0.1%Above 1%, contests become contestable
gVisor overhead on syscall-heavy code~2–10×Irrelevant for CPU-bound solutions, painful for I/O-heavy ones

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the CRUD — problems table, submissions table, a REST API — and 3 minutes on the sandbox. The interviewer picked this problem for the sandbox and the spike. Compress the basics to ~10 minutes.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Users browse problems, write code in one of ~20 languages, click Run against custom input or Submit against hidden tests, and get a verdict. We also host weekly contests with a live leaderboard."

Then spend your time on the non-functionals — this is where the design lives:

"Three things make this hard. First, every submission is untrusted code — I'll treat it as an exploit attempt. Second, contests produce a scheduled 5–10× spike in the first and last ten minutes. Third, the verdict has to be trustworthy: the same code should get the same verdict every time, or rankings are meaningless."

Commit to numbers:

"I'll assume 10M submissions/day on practice — about 115/s average, 300/s peak — plus a weekly contest with 30K participants that peaks around 300 submits/s on top. Each submission is ~100 test executions and ~2 CPU-seconds. Contest verdict p95 under 10 seconds. No outbound network from user code, ever."

🎯 Staff Move: Say the threat model out loud in the first three minutes. "I'm going to assume the submitter is an attacker" reframes the entire interview from a queue problem to a security-and-capacity problem — which is what the interviewer wanted to test.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns, don't draw an ER diagram:

  • Problem (id, version, time_limit_ms, memory_limit_mb, checker_type)
  • TestSet (problem_id, version, content_hash, cases[]) — immutable, versioned
  • Submission (id, user_id, problem_id, contest_id?, language, source_hash, ingress_ts, lane)
  • Verdict (submission_id, testset_version, status, cpu_ms, mem_kb, failed_case_idx) — keyed by both

API surface:

POST /submissions            { problem_id, language, source, contest_id? } → 202 { submission_id }
GET  /submissions/{id}       → { status: QUEUED|RUNNING|DONE, verdict?, cpu_ms?, mem_kb? }
POST /runs                   { problem_id, language, source, stdin } → 202 { run_id }
GET  /contests/{id}/standings?page=1

Results are asynchronous. The client polls GET /submissions/{id} with backoff (500ms → 1s → 2s) or subscribes via SSE for the contest page.

🎯 Staff Move: "Verdict is keyed by (submission_id, testset_version), not submission_id alone. That one decision makes rejudging idempotent and auditable — we never overwrite a verdict, we add a new one against a new testset version."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk it in 60 seconds:

  1. Gateway authenticates and applies per-user quotas (Submit 6/min, Run 10/min).
  2. Submission service stamps ingress_ts, stores source in object storage, writes the submission row, enqueues to the right lane.
  3. Dispatcher pulls with weights (contest 70 / practice 25 / run 5 during a contest; 0 / 80 / 20 otherwise).
  4. A judge host takes a warm microVM, injects source, compiles, streams test cases via stdin, compares outputs outside the sandbox, destroys the VM.
  5. Verdict goes to Postgres (durable) and Redis (status cache the client polls).

🎯 Staff Move: "This is the design that works on a Tuesday. It breaks in three places on contest day — the trust boundary, the spike, and timing variance. Let me take those in order of risk."


Phase 4: Transition to Depth (1 minute)#

"Three deep dives matter here: how the sandbox actually contains hostile code, how we absorb a scheduled 10× spike without making rankings unfair, and how we keep verdicts deterministic on shared hardware. I'd start with isolation because a failure there is a security incident, not a slow page. Which would you like?"

If the interviewer has no preference, lead with isolation — it is the most distinctive part of this problem and the one where L5 answers are thinnest.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → pick a position → quantify the cost → name who absorbs it.

Deep dive 1: The sandbox (8–10 min)

"I layer five controls, each stopping a different attack class:"

LayerStopsCost
seccomp-bpf allowlist (~60 syscalls)ptrace, mount, raw sockets, keyctl, kernel attack surface~0; some languages need tuning (JVM, Go runtime)
Namespaces (pid, net, mount, user, ipc, uts)Seeing/killing other processes, network access, host filesystem~0
cgroups v2 (cpu.max, memory.max, pids.max=64)Fork bombs, memory bombs, CPU hogging~0
Read-only rootfs + 64 MB tmpfs + output cap 64 KBPersistence, disk filling, infinite-print DoS~0
Firecracker microVM per sandboxKernel exploits escaping to the host~125ms boot (hidden by warm pool), ~10–15% density

"The first four are containment. The fifth is the blast-radius limiter. If a submission escapes the container via a kernel bug, it lands in a single-use VM with no network and no credentials, which we destroy after the submission. That's an incident report, not a breach."

Diagram: Phase 5: Deep Dives (25–30 minutes)

Key detail most candidates miss — test data never enters the sandbox as files. The harness outside the sandbox streams input via stdin and captures stdout; the expected output and checker stay outside. Otherwise user code can open("/tests/expected_17.txt") and print it.

Deep dive 2: The contest spike (8–10 min)

"Contests are on a calendar. Registration count tells me the peak a week in advance. I pre-warm the contest pool 30 minutes before start to 3× the forecast peak — roughly 300 submits/s × 2 CPU-s × 3 = ~1,800 cores — and keep it until 20 minutes after end for rejudges. Autoscaling handles only the unscheduled practice traffic."

"Fairness is decoupled from latency: penalty time is computed from ingress_ts, stamped at the submission service before the queue. If the queue lags 90 seconds, nobody's rank changes. The only user-visible cost is waiting for the verdict, and we cap that with lane priority."

Deep dive 3: Timing determinism (6–8 min)

"The verdict is CPU time from cgroup cpu.stat, not wall time. Each sandbox gets dedicated physical cores — I disable the SMT sibling for the contest pool. Wall time is a 3× backstop to kill sleep-forever programs. And I run the reference solution for 1% of problems every minute as a canary; if its CPU time drifts more than 5% from baseline, that host class is drained."

"Who pays: the pinning halves sandbox density — a 64-core host goes from ~120 sandboxes to ~48. For the contest pool I pay it. For the Run lane I don't — Run verdicts aren't ranked."


Phase 6: Wrap-Up (2–3 minutes)#

Synthesize, don't restate:

"The pipeline is a standard job queue. What makes it an online judge is three commitments: every submission is treated as an attacker and contained by a VM boundary; fairness is stamped at ingress so queue lag can't change rankings; and verdicts are measured in CPU time on pinned cores so they're reproducible."

The organizational closer:

"The failure I'd worry about most isn't the sandbox — it's wrong test data. That's owned by problem setters, not the platform team, and it causes mass rejudges. I'd make rejudge a first-class workflow with a diff report, and give the contest admin — not an engineer — the button, with the platform team owning the capacity it consumes."

🎯 Staff Move: End on the ownership seam between the platform (runtime, timing, capacity) and content (tests, checkers). That seam is where real incidents come from.


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
10 min on schemaDesigns problems, tags, discussions, user profilesNames 4 entities, keys verdicts by testset version, moves on
"Docker" as the whole security storyMentions no-network and a timeoutLayers controls, names the kernel-CVE threat, adds a VM boundary
Autoscaling as the spike answer"Scale workers on queue depth"Pre-warms from the calendar; autoscaling only for unscheduled load
Ignores fairnessProcesses FIFOStamps ingress_ts; lanes decide who waits
No measurement discussion"2-second timeout"CPU time, pinned cores, canary reference solutions
Stops at happy pathNo rejudge, no wrong-test storyRejudge workflow, idempotent verdicts, content-vs-platform ownership

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

The online judge is a rare prompt where the "obvious" design is correct and also dangerously incomplete. Everyone gets queue + worker right. The level signal is whether you notice that (a) the workload is adversarial, (b) the peak is scheduled and therefore plannable, and (c) the product is a trustworthy measurement, not a computation. Interviewers use it to test threat modeling and capacity planning in one problem.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"A user submits the same C++ file twice, one minute apart, during a contest. The first gets Accepted in 940ms. The second gets TLE. Walk me through why — and what you'd have built so that this doesn't happen."

This question exposes whether the candidate distinguishes wall time from CPU time, understands noisy neighbors (shared L3 cache, memory bandwidth, SMT siblings), and treats verdict determinism as a measurable property with an owner. Candidates who've never operated a judge say "increase the timeout."


2. Problem Framing & Intent#

2.1 The Four Intents — Explained#

Practice judging → cost-first, eventually fast

  • Constraint: steady ~115/s average with a diurnal peak around 300/s; margin matters because this is the free tier
  • Strategy: shared autoscaled pool, weighted lane, can use cheaper isolation (gVisor) and denser packing
  • Failure mode: verdict slows to 20–30s during peaks — annoying, not damaging
  • Who pays for imperfection: free users (wait longer); support volume increases slightly

Contest judging → fairness-first

  • Constraint: scheduled bimodal spike; ranking integrity is the product
  • Strategy: dedicated, pinned, pre-warmed pool; ingress_ts for ordering; rejudge capacity reserved
  • Failure mode: false TLE or backlog that users perceive as unfair → contest unrated
  • Who pays for imperfection: the community's trust in rankings; the contest team's credibility; in the extreme, sponsors

Interactive Run → abuse-first

  • Constraint: 3–5× the volume of Submit, near-zero value per call, easiest free compute on the internet
  • Strategy: lowest priority, tight per-user CPU-seconds budgets (e.g., 60 CPU-s/hour for free, 600 for paid), shed first
  • Failure mode: crypto mining, botnets farming free CPU
  • Who pays for imperfection: the infrastructure bill; ultimately finance

Hiring assessment (B2B) → audit-first

  • Constraint: contractual uptime (e.g., 99.9% during scheduled assessment windows), per-customer isolation, dispute resolution
  • Strategy: separate tenant pool, full provenance per verdict (runtime image digest, host class, testset hash), 1–2 year retention
  • Failure mode: candidate disputes a rejection; customer demands proof
  • Who pays for imperfection: sales/legal; enterprise churn

🎯 Staff Move: "These four share one pipeline but not one policy. The architecture is lanes + pools; the intents differ in priority weight, isolation class, and retention. That's why I'd make lane and pool configurable, not hard-coded."

2.2 When NOT to Build This#

  • You need to execute code for fewer than ~10K runs/day. Use a hosted sandbox service or an open-source judge (e.g., Judge0) on a managed VM. Building microVM orchestration for that volume is resume-driven.
  • The code is trusted (internal CI, your own employees' notebooks). A standard container runner with resource limits is enough — the threat model changes everything.
  • Languages are all JavaScript/WASM-compilable. V8 isolates or WASM runtimes give <5ms startup and dense packing; skip VMs entirely.
  • Latency must be interactive (<100ms), e.g., a REPL-style playground. The judge pipeline (queue + batch tests) is the wrong shape — use a persistent per-session sandbox with a strict idle timeout instead.

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Which languages. 3 languages vs 25 changes image management, warm-pool fragmentation, and seccomp tuning (JVM and Go need more syscalls).
  • Custom checkers. Some problems accept multiple correct answers or floating-point tolerances; the checker is itself code — whose, and is it sandboxed?
  • Interactive problems. Some judges run the user program against an interactor process over pipes — a two-process sandbox.
  • Contest scale and frequency. One 30K-user contest per week vs a 300K-user global event changes whether you own reserved hardware.
  • What users see on failure. Showing the failing test's input leaks hidden tests; hiding it frustrates practice users.

Staff engineers surface these. Senior engineers assume them away.

2.4 Precise Terminology#

TermWhat It MeansWhy the Precision Matters
SandboxThe whole containment stack for one executionSay which layers; "sandbox" alone is not an answer
Trust boundaryThe layer whose failure means host compromiseFor containers it's the kernel; for microVMs it's the hypervisor
CPU timeTime the process spent on CPU (user + sys) from cgroup accountingDeterministic-ish; the basis for TLE
Wall timeElapsed real timeIncludes waiting, sleeping and neighbor interference; only a backstop
Testset versionImmutable, content-hashed bundle of inputs + expected outputs + checkerEnables idempotent rejudge
LaneA logical queue with its own priority weight and quotasEncodes the "who waits" policy
Warm poolPre-booted sandboxes ready for handoffConverts ~125ms–90s startup into ~10ms
RejudgeRe-running existing submissions against a new testset or runtime versionA workflow with capacity, diffing, and approval

3. The Five Fault Lines#

Each fault line below is a place where two reasonable engineers disagree. The Staff job is to pick, quantify, and name who absorbs the cost.

3.1 Fault Line 1: Isolation Strength vs Startup Latency#

The tension: Every layer of isolation costs startup time, runtime overhead, or density. The attacker only needs one gap.

StrategyWhat WorksWhat BreaksWho Pays
Container only (namespaces + cgroups + seccomp)50–300ms warm start, dense, familiar toolingOne kernel CVE (e.g., a privilege-escalation bug in a syscall on the allowlist) = host compromise, including every other user's code and any credentials on the hostSecurity team, and every user whose submission or data was on that host
gVisorSmall host attack surface, OCI-compatible, no KVM needed2–10× slowdown on syscall-heavy programs; some runtimes misbehave on unsupported syscallsUsers of I/O-heavy languages and problems; platform team tuning compatibility
Firecracker microVM + container insideHardware virtualization boundary; ~125ms boot; <5 MiB overheadRequires bare metal or nested virt; guest kernel to patch; more moving partsPlatform team (ops complexity); finance (bare metal cost)
Language isolates (V8/WASM)<5ms start, massive densityCannot run native C++, Java, Python without WASM ports that change performance characteristicsProduct: language support is limited

Staff default: Container controls inside a Firecracker microVM for Submit and contest lanes. The warm pool hides the 125ms boot. For the Run lane, gVisor on shared hosts is acceptable because Run has no ranking impact — but it still gets no network and a strict CPU budget.

Reuse vs fresh sandbox: A microVM reused across submissions must be reset — a leaked background process, a file in /tmp, or a modified environment variable can alter the next user's verdict or leak their code. Staff default: one VM per submission, destroyed after. If density forces reuse, restore from a snapshot (Firecracker supports snapshot/restore) rather than "cleaning up."

When to deviate: Trusted code (internal assessments of employees) → containers alone. Extreme scale with only JS/Python → consider WASM (e.g., Pyodide) and accept the language-compatibility cost.

🎯 Staff Move: "I'm not choosing between container and VM — I'm stacking them. The container controls are for containment; the VM is for blast radius. The question I'd ask the security team is what our kernel CVE patch SLA is, because without a VM boundary that SLA is our sandbox's real security guarantee."


3.2 Fault Line 2: Warm Capacity vs Cost#

The tension: A pool sized for the contest peak is idle ~95% of the week. A pool sized for the average cannot catch a spike that starts in under 60 seconds.

Back-of-envelope:

Contest peak:     300 submits/s × 2 CPU-s = 600 cores busy
Headroom 3×:      1,800 cores  →  ~30 hosts of 64 cores (with SMT off: ~60 hosts)
Duration needed:  T-30min to T+20min after a 90-min contest ≈ 2.3 hours/week
Idle cost if permanent: 60 hosts × ~$3/hr × 168 h ≈ $30K/week
Scheduled cost:         60 hosts × ~$3/hr × 2.3 h ≈ $415/week (+ reservation premium)
StrategyWhat WorksWhat BreaksWho Pays
Always-on peak poolZero spike risk~$1.5M/year mostly idleFinance
Pure reactive autoscalingCheapest30–90s cold starts + image pulls; spike is over before capacity arrivesContestants (lag, unfairness)
Calendar-driven pre-warm + reactive for the restPays for peak only during scheduled windowsNeeds accurate forecasts; unscheduled viral spikes still hurtPlatform team maintains the forecaster
Borrow from practice pool during contestsReuses capacityPractice verdicts slow down 2–5×Practice users (acceptable, if communicated)

Staff default: Calendar-driven pre-warm for contests (forecast = registrations × historical submits/user × peak factor), plus lane weights that borrow practice capacity during the contest window. Reactive autoscaling handles practice diurnal curves with a 20% floor of warm capacity per language.

Warm-pool fragmentation: With 20 languages, a warm pool per language fragments capacity. Mitigation: a language-agnostic base VM snapshot containing all runtimes (~2–4 GB image, cached on local NVMe), so any warm VM can run any language. Cost: bigger image, slower host bootstrap (minutes), but host bootstrap happens once per host, not per submission.

🎯 Staff Move: "Contests are the one spike in distributed systems that's printed on a calendar. Reacting to it with autoscaling is choosing to fail. I'd pre-warm from registrations and let practice absorb the squeeze, because a 20-second practice verdict costs nothing and a 5-minute contest verdict costs the contest."


3.3 Fault Line 3: Throughput vs Fairness Under Spike#

The tension: FIFO is simple and maximizes throughput. It also means a practice user's submission can sit ahead of a contestant's, and a single user with 50 rapid resubmits can occupy 50 slots.

StrategyWhat WorksWhat BreaksWho Pays
Single FIFOSimple, no starvationContest spike blocks everyone; resubmit storms amplifyEveryone equally — which means contests fail
Strict priority (contest > practice > run)Contest always firstPractice and Run starve completely during contestsPractice users (0 throughput for 90 min)
Weighted fair dispatch (70/25/5)Contest dominates, others progressNeeds tuning; weights are policyPractice users wait 2–5× longer, knowingly
Per-user fair share within a laneOne user can't hog the laneMore dispatcher stateHeavy resubmitters wait their turn

Staff default: Weighted fair dispatch across lanes, plus per-user dedupe and fair share within the contest lane: at most 2 in-flight submissions per user per problem; identical source_hash resubmits return the cached verdict instantly.

The decoupling that makes this safe: Fairness in contests is about ordering and penalty, not verdict speed. Stamp ingress_ts at the submission service (before the queue). Penalty = ingress_ts − contest_start + 5–10 min per wrong attempt. Now queue lag affects user experience but not outcome. This is the single most important sentence in the contest deep dive.

Diagram: 3.3 Fault Line 3: Throughput vs Fairness Under Spike

When to deviate: In B2B assessments, each candidate has a private time limit; fairness is per-candidate latency, so a dedicated per-customer pool beats lane weighting.

🎯 Staff Move: "Queue lag is an experience problem, not a fairness problem — as long as the timestamp that decides rankings is taken before the queue. That lets me run the judge hot during the spike without anyone's rank depending on which worker picked them up."


3.4 Fault Line 4: Timing Determinism vs Packing Density#

The tension: Stable timing needs isolated hardware resources (dedicated cores, no SMT siblings, ideally no shared L3 thrash). Density needs sharing.

StrategyWhat WorksWhat BreaksWho Pays
Wall time, shared coresMax density (~120 sandboxes on 64 cores with SMT)20–50% timing variance; false TLEsContestants near the time limit
CPU time, shared coresRemoves waiting time from the measurementCache/memory-bandwidth contention still inflates CPU time 5–20%Contestants with memory-heavy solutions
CPU time, pinned physical cores, SMT offVariance typically < 3–5%Density roughly halves (~48 sandboxes per 64-core host)Finance: contest pool ~2× the hardware
Run-twice-on-borderlineRe-runs verdicts within 10% of the limitExtra CPU (~2–5% of submissions)Platform capacity

Staff default: CPU time on pinned physical cores with SMT disabled for the contest and B2B pools; CPU time on shared cores for practice; plus automatic re-run of any TLE within 10% of the limit before reporting it. Calibrate time limits per language (e.g., Python limits 3–5× C++), and publish them.

Canary detection: Every judge host continuously runs a small set of reference solutions (1% of capacity). Alert on judge.canary_cpu_ms drift > 5% from the host-class baseline; auto-drain the host. This catches thermal throttling, firmware changes, noisy co-tenants and kernel regressions — none of which show up in error rates.

🎯 Staff Move: "I'll pay 2× hardware for the contest pool because timing variance there changes rankings, and I'll refuse to pay it for practice where a borderline TLE just means the user resubmits. Determinism is a per-lane SLO, not a global one."


3.5 Fault Line 5: Judge Platform vs Content Ownership#

The tension: A wrong verdict can come from the platform (runtime version, timing, sandbox bug) or from content (wrong expected output, weak tests, buggy checker). Users can't tell the difference. Someone has to own the fix — and the rejudge.

ModelWhat WorksWhat BreaksWho Pays
Platform owns everythingOne throat to chokePlatform engineers become test-data editors; slow; they lack problem contextPlatform on-call
Content owns everythingDomain experts fix testsContent team triggers 40K-submission rejudges during peakEveryone in the queue
Split: content owns testsets, platform owns rejudge capacity and runtimeClear seams; rejudge rate-limitedNeeds a contract (testset validation, approval, capacity budget)Both, via an agreed process

Staff default: Split ownership with a testset contract:

  • Every new testset version must pass: reference solution AC, at least one known-wrong solution fails, checker runs in the sandbox, max input size within limits.
  • Rejudge is requested by content (contest admin), executed by the platform in a rejudge lane capped at 20% of capacity, with a diff report (how many verdicts flip, in which direction) reviewed before publishing.
  • Verdicts are never overwritten; the "current" verdict points to the latest approved testset version.
Diagram: 3.5 Fault Line 5: Judge Platform vs Content Ownership

🎯 Staff Move: "The rejudge button belongs to the contest admin, but its capacity belongs to me. I'd give them a dry-run diff — '1,240 verdicts flip from AC to WA' — before anything is published, because flipping 1,240 accepted verdicts is a community event, not a bug fix."


4. Failure Modes & Operational Reality#

4.1 Contest-Start Thundering Herd — Full Timeline#

Scenario: Weekly contest, 30K participants. The contest pool was sized from last month's registration, but a popular streamer announced the contest; 52K registered.

t=-30min:  Pre-warm to 3× forecast (forecast based on 30K). Capacity ≈ 900 submits/s burst.
t=0:       Contest opens. Problem pages served from CDN — fine.
t=+3min:   Easy problem submissions begin: 1,100 submits/s. Run clicks: 3,500/s.
t=+4min:   Contest lane lag p95 = 40s. Run lane lag = 4 min.
t=+5min:   Users see "Pending" for 40s and resubmit — 18% duplicate rate.
t=+6min:   Alert: queue_lag_p95{lane=contest} > 30s for 2 min → page.
t=+7min:   On-call shifts weights: contest 90 / practice 10 / run 0 (Run returns 503 + "Run disabled during peak").
t=+8min:   Source-hash dedupe enabled for resubmits (was off for Run lane).
t=+12min:  Reactive capacity from reserved warm hosts arrives. Lag p95 = 8s.
t=+85min:  Second peak (last-minute submissions): absorbed, lag p95 = 12s.
t=+95min:  Leaderboard final at t+97min. Rankings unaffected — ingress_ts ordering.

Detection: queue_lag_p95{lane}, submissions_per_sec{lane}, duplicate_submit_ratio, warm_pool_available{language}.

Mitigation: Shift lane weights (pre-approved playbook), disable Run for non-contest users, enable dedupe, pull reserved capacity.

Prevention: Forecast from live registrations up to T-1h, not last month; auto-add capacity when registrations exceed forecast by 25%; client-side resubmit debounce (disable Submit button until verdict or 30s).

Owner: Platform on-call executes; contest team owns the forecast input.


4.2 Noisy Neighbor → False TLE → Rejudge Storm#

Scenario: A new host class (different CPU generation) was added to the contest pool to handle growth. Its per-core performance is 12% lower for memory-bound code. Nobody recalibrated time limits.

t=0:       New host class takes 30% of contest pool.
t=+20min:  Forum posts: "same code AC then TLE".
t=+25min:  judge.canary_cpu_ms on new class = +14% vs baseline. No alert (threshold was on errors, not drift).
t=+60min:  Contest ends. 2.1% of contest submissions TLE within 15% of limit.
t=+2h:     Contest team requests rejudge of all TLE verdicts. 9K submissions.
t=+2h10m:  Rejudge runs in practice lane, uncapped. Practice lag climbs to 11 min.
t=+4h:     Rejudge complete. 640 verdicts flip TLE → AC. Rankings recomputed. Public apology.

Detection: judge.canary_cpu_ms drift by host class (should have alerted at 5%), verdict_flip_rate_on_rerun, tle_near_limit_ratio (TLEs within 10% of the limit — normally ~0.3%).

Mitigation: Drain the host class from the contest pool; rejudge in a capped rejudge lane.

Prevention: Host classes are admitted to the contest pool only after a calibration run of the full reference suite; per-class time limit multipliers; borderline TLE auto-rerun on a baseline-class host.

Owner: Platform team (host qualification); contest team (communication and rejudge approval).


4.3 Sandbox Escape via Kernel CVE#

Scenario: A public Linux kernel privilege-escalation CVE is disclosed. The exploit uses a syscall on your seccomp allowlist.

t=0:       CVE published with proof-of-concept.
t=+6h:     First submissions containing the PoC appear (security researchers and attackers both read the feed).
t=+6h:     Container-only judge: exploit gains root on host → reads other sandboxes' source, host credentials.
           MicroVM judge: exploit gains root in the guest VM → no network, no creds, VM destroyed after run.
t=+8h:     Detection: sandbox_violation_total spike (seccomp kills, unexpected uid 0 in guest).
t=+24h:    Guest and host kernels patched and rolled via rolling drain.

Detection: sandbox.seccomp_kill_total, sandbox.guest_uid0_events, anomaly on submission_source_signature (known PoC hashes), host integrity monitoring.

Mitigation: Block the syscall in seccomp immediately (config push, minutes) if runtimes tolerate it; emergency kernel patch; quarantine suspicious submissions.

Prevention: VM boundary, no credentials on judge hosts (pull-only, scoped tokens), guest kernel patch SLA ≤ 72h for critical CVEs, source scanning for known PoC signatures, bug bounty including sandbox escapes.

Owner: Security (threat intel, SLA) + platform (patch rollout). The blast radius is determined by an architecture decision made months earlier.


4.4 Wrong Test Data — Silent Mass Misjudgment#

Scenario: A problem setter uploads a testset where case 37's expected output was generated by a buggy reference solution. Correct submissions get Wrong Answer.

Detection signals: problem_ac_rate drops sharply vs difficulty-predicted rate; wa_concentration_by_case (80% of WAs fail on the same case); user reports. The concentration metric is the strongest signal — real WAs spread across cases.

Mitigation: Freeze the problem (stop judging new submissions, return "Pending review"), fix testset, rejudge affected submissions via the rejudge workflow with a diff report.

Prevention: Testset contract (§3.5): multiple independent reference solutions (in two languages) must agree; known-wrong solutions must fail; automated alert when one case accounts for > 50% of WAs within the first 200 submissions.

Owner: Content team owns the testset; platform owns the concentration alert and rejudge capacity.


4.5 Zombie Sandboxes — Capacity Silently Shrinks#

Scenario: A runtime bug means some sandboxes are not destroyed when the harness times out (e.g., the VM process hangs on teardown). Each leaked VM holds 256 MB and 1 pinned core.

Detection: host.allocated_sandboxes − host.active_executions (should be ~0), warm_pool_available trending down without load growth, per-host capacity vs expected.

Mitigation: Reaper process kills any VM older than max_wall × 2; hosts whose leak count exceeds 5 are drained and recycled.

Prevention: Hosts are cattle: recycle every 24h regardless; teardown is part of the canary suite.

Owner: Platform team. This failure never pages on its own — it shows up as "the contest pool was smaller than we thought."


4.6 Free Compute Farming via Run#

Scenario: A botnet of 800 free accounts uses the Run endpoint to mine cryptocurrency, each run using the full 10s wall limit with multi-threaded hashing.

Detection: run_cpu_seconds_per_account distribution (bots hug the limit), run_output_entropy, identical source hashes across many accounts, sign-up velocity by IP/ASN.

Mitigation: Per-account CPU-seconds budget (60 CPU-s/hour free), per-IP/ASN budgets, cap threads (pids.max, 1 CPU per Run sandbox), block known mining signatures.

Prevention: Run is 1-core, 5s, lowest priority, shed first; new accounts get smaller budgets until they build history; CAPTCHA on anomalous sign-up bursts.

Owner: Trust & safety (account policy) + platform (quotas). Finance is the silent victim — the bill arrives a month later.


4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Contest spike exceeds forecastqueue_lag_p95{lane=contest} > 30sAll contestants' experience (not rankings)Shift lane weights, disable Run, pull reserved capacityPlatform on-call
Timing drift on host classjudge.canary_cpu_ms drift > 5%Borderline submissions on that classDrain class, auto-rerun borderline TLEsPlatform
Kernel CVEsandbox.seccomp_kill_total, guest uid0 eventsOne VM (with VM boundary) / whole host (without)seccomp hotfix, emergency patchSecurity + platform
Wrong test datawa_concentration_by_case > 50%Every submission to that problemFreeze problem, fix, rejudge with diffContent team
Zombie sandboxesallocated − active > 0 trendSilent capacity loss, worst during contestsReaper, host recyclePlatform
Compute farmingCPU-s per account near cap, source-hash clusteringCost; Run lane capacityBudgets, account trust tiersTrust & safety
Result store lagverdict_write_lag_ms p99 > 1sUsers see stale "Pending" and resubmitServe status from Redis; backpressure submissionsPlatform
Object store slow for test datatestdata_fetch_ms p99Cold judge hosts can't startLocal NVMe cache by testset hash, pre-fetch before contestPlatform

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Threat model"No network, timeout, Docker"Layered controls + VM boundary; test data never inside the sandboxOrg sandbox posture: patch SLA, red team, bug bounty scope, shared runtime for all code-execution products
CapacityAutoscale workers on queue depthCalendar-driven pre-warm, lanes, borrow-from-practiceReserved capacity contracts, cost per contest, whether to own bare metal
FairnessFIFOingress_ts ordering, per-user fair share, dedupePublishes contest integrity policy: when is a contest unrated, who decides
MeasurementWall-clock timeoutCPU time, pinned cores, canaries, borderline rerunsDeterminism as an SLO with a funded hardware class
Correctness of contentNot mentionedTestset contract, versioned verdicts, rejudge workflowContent/platform org contract; problem-setter tooling as a product
AbuseRate limit per userCPU-second budgets, Run vs Submit policies, farming detectionUnit economics: CPU cost per free MAU vs conversion; policy set with finance

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Treats code as adversarial from the first minute"I'm going to assume every submission is an exploit attempt, because at 10M/day some are."
Separates containment from blast radius"seccomp and cgroups contain; the microVM limits the blast radius when containment fails."
Decouples fairness from latency"Rankings use the ingress timestamp, so the queue can lag without changing outcomes."
Knows measurement is the product"Wall time measures the neighbors. I'll measure CPU time on pinned cores and canary it."
Plans for scheduled spikes"The contest is on a calendar. I'll pre-warm from registrations, not react."
Owns the wrong-verdict path"Rejudge is a workflow with a diff report, not a script someone runs."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Docker with network off" is the entire security answerIgnores the shared-kernel threat — the central risk of the problem
Test cases mounted into the containerUser code can read expected outputs; a correctness hole disguised as a design choice
Autoscaling is the only spike answerIgnores 30–90s cold starts against a 5-minute spike
No distinction between Run and SubmitMisses that Run is 3–5× the volume and the abuse vector
Overwrites verdicts on rejudgeNo audit trail; can't answer "why did my rank change?"
Spends 15 minutes on problem/tag/discussion schemaOptimizing the part of the system that isn't hard

5.4 Common False Positives#

  • Deep knowledge of Linux namespaces ≠ sandbox design. Listing all seven namespaces without stating the trust boundary is trivia.
  • Kubernetes fluency ≠ capacity planning. "HPA on queue depth" is still reactive scaling.
  • Mentioning Firecracker ≠ understanding why. The signal is explaining what the VM boundary changes about blast radius, and what it costs.
  • Elaborate leaderboard design ≠ contest fairness. A perfect sorted set is useless if the timestamp feeding it came from the judge instead of ingress.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing + intent0–4 minThreat model, contest as the hard case, numbers
Entities + API4–6 minVersioned testsets, verdict key, async results
High-level design6–11 minGateway → submission → lanes → judge → results
Deep dive: sandbox11–20 minLayers, VM boundary, test data outside
Deep dive: spike + fairness20–30 minPre-warm, lanes, ingress_ts
Deep dive: timing or rejudge30–40 minCPU time, canaries, rejudge workflow
Wrap-up40–45 minOwnership seam, evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"What if someone submits a fork bomb?"Resource containment basicspids.max=64, cgroups; then pivot to "the harder attack is the kernel"
"Can users read the test cases?"Data-flow securityStdin streaming from outside; expected output never enters the sandbox
"Now support 50 languages."Operability, image managementSingle base snapshot with all runtimes; per-language seccomp profiles; language owners
"Now make Run interactive — a REPL."Recognizing a different shapeSession sandbox with idle timeout, not the batch pipeline
"Enterprise wants to run their own problems with their own checkers."Multi-tenancy + untrusted checkersCheckers are sandboxed too; tenant-dedicated pool
"What if the queue itself goes down?"DurabilitySubmission row written before enqueue; reconciler re-enqueues QUEUED rows older than 60s

6.3 What to Deliberately Skip#

  • Problem browsing, tags, discussions, editorial pages. CDN-cached reads — say so in one sentence.
  • The code editor. Client-side; irrelevant.
  • Plagiarism detection internals. Mention it as an offline post-contest job (token-based similarity, MOSS-style); don't design it.
  • Leaderboard data structure. A sorted set keyed by (solved desc, penalty asc) is enough; link to Leaderboard if pushed.

6.4 Follow-Up Questions to Expect#

  1. "Why not run all test cases in parallel across workers?" — (Fan-out increases VM count 50–100× per submission; latency gain is small because most submissions fail early or finish in 1–3s total. Parallelize only for B2B problems with long tests.)
  2. "How do you handle a submission that compiles for 60 seconds?" — (Separate compile limit, e.g., 10s CPU / 30s wall; template-metaprogramming bombs are real.)
  3. "What does the client do while waiting?" — (Poll with backoff, or SSE for contest pages; show queue position to reduce resubmits.)
  4. "How do you prevent hard-coding answers to known test cases?" — (Hidden tests, randomized large cases, post-contest review; it's a content problem, not a runtime one.)
  5. "How would you judge interactive problems?" — (Two sandboxed processes connected by pipes, interactor outside the user's VM or in a separate one; measure only the user's CPU.)
  6. "What's the cost per submission?" — (~2 CPU-s at ~$0.04/core-hour ≈ $0.00002; the expensive part is idle warm capacity, not execution.)
  7. "How do you roll out a new compiler version?" — (New runtime image = new "runtime version" on verdicts; shadow-judge 1% of submissions on both versions and diff verdicts before switching; never switch during a contest.)

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design LeetCode."

Staff Answer

"Before I draw anything, two framing points. First, the input to this system is arbitrary code from the internet, so I'm going to treat every submission as an exploit attempt — that drives the sandbox design. Second, practice judging, contests, and the Run button have very different shapes: practice is steady and cost-sensitive, contests are a scheduled 5–10× spike where ranking fairness is the product, and Run is high-volume, low-value, and the main abuse vector. I'll design for contests because they force every hard decision, and practice is the same pipeline in a lower-priority lane. Numbers: 10M submissions/day, contest peak ~300/s, ~2 CPU-s per submission, contest verdict p95 under 10s. I'll go: API and entities briefly → pipeline → sandbox → spike and fairness → timing determinism → rejudge."

Why this is L6:

  • Threat model stated before architecture
  • Commits to contest as the design driver and explains why practice falls out of it
  • Gives numbers that later justify capacity decisions

What L7 adds:

  • "I'd also ask whether code execution is a product feature or a platform — assessments and an AI tutor will want the same sandbox."
  • Frames the Run button's cost as a line item against free-tier economics
❌ Common L5 Trap

"Users submit code through an API, we put it in a queue, workers pull it and run it in Docker, and store the result in a database. Let me start with the database schema for problems and submissions..."

Why this misses: Correct, and generic — it's the design for any job queue. The interviewer has to drag the candidate to the sandbox and the contest spike, which are the entire reason this question exists.


Drill 2: The Sandbox#

Prompt: "How do you make sure user code can't harm the system?"

Staff Answer

"Five layers, each stopping a different class of attack. seccomp-bpf with an allowlist of roughly 60 syscalls removes most of the kernel attack surface — no ptrace, mount, raw sockets. Namespaces isolate process, network, filesystem and user views; the network namespace has no interfaces except loopback. cgroups v2 cap CPU, memory at 256 MB, and pids at 64 — that kills fork bombs and memory bombs. Root filesystem is read-only with a 64 MB tmpfs, stdout is capped at 64 KB. And all of that runs inside a single-use Firecracker microVM, because the first four layers share the host kernel — a kernel CVE on an allowed syscall is a host compromise without the VM. With it, it's one throwaway VM with no network and no credentials. Test data never enters the sandbox; the harness streams input over stdin and compares output outside. Cost: ~125ms boot, hidden by a warm pool, and ~10–15% density."

Why this is L6:

  • Maps each layer to the attack it stops, rather than listing technologies
  • Identifies the shared-kernel trust boundary as the real risk
  • Catches the test-data-exfiltration hole

What L7 adds:

  • Sets a guest/host kernel patch SLA (≤ 72h critical) and puts sandbox escapes in bug bounty scope
  • Makes the sandbox a shared org runtime with a security review gate, so the next team doesn't build a weaker one

Drill 3: Make the Spike Concrete#

Prompt: "A contest with 30K participants starts at 8:00. Walk me through capacity."

Staff Answer

"Forecast from registrations: 30K × ~5 submits over 90 min, but bimodal — ~10% of submissions land in minutes 3–10 and another ~15% in the last 10 minutes. Peak ≈ 300 submits/s. Each is ~2 CPU-s, so ~600 busy cores; with 3× headroom and SMT off that's ~60 hosts of 64 cores. The control plane reads the contest calendar and brings the contest pool up at 7:30 — hosts booted, base snapshot on local NVMe, test data for the contest's four problems prefetched by testset hash, and a warm pool of pre-booted VMs. At 8:00 lane weights switch to contest 70 / practice 25 / run 5. Registrations exceeding forecast by 25% auto-add hosts up to T-1h. Pool drains at 9:50 after rejudge buffer. Cost: ~60 hosts × 2.3h ≈ $400 per contest versus ~$30K/week if kept warm permanently."

Why this is L6:

  • Turns a known event into a plan, with a forecast input and a trigger
  • Handles bimodality, not just "peak"
  • Prices the alternative

What L7 adds:

  • Negotiates reserved or scheduled capacity for the weekly calendar; considers whether contest hosts should be owned bare metal
  • Sets policy for mega-events (300K participants): staggered start by region or problem release, decided with the contest team

Drill 4: The Flaky Verdict#

Prompt: "Users say the same code sometimes gets AC and sometimes TLE."

Staff Answer

"That's a measurement problem. First question: are we measuring wall time or CPU time? Wall time includes scheduler delay and neighbors — switch to cgroup CPU time. Second: even CPU time inflates under contention — shared L3, memory bandwidth, and especially SMT siblings. For the contest pool, pin each sandbox to dedicated physical cores and disable SMT. Third: host heterogeneity — check judge.canary_cpu_ms by host class; a new CPU generation or thermal throttling shows up as drift. Fourth: any TLE within 10% of the limit is automatically rerun on a baseline host before we report it. Target: false-TLE rate under 0.1%, measured by rerunning a 1% sample."

Why this is L6:

  • Diagnoses in layers: metric choice, contention, heterogeneity
  • Proposes a detection mechanism (canaries) not just a fix
  • Sets a measurable target

What L7 adds:

  • Publishes determinism as an SLO and ties hardware procurement to it
  • Requires calibration before any new host class enters the contest pool — a standard, not a one-off fix

Drill 5: Wrong Test Data During a Contest#

Prompt: "Twenty minutes into a contest, you realize problem C's test 37 has a wrong expected output."

Staff Answer

"Freeze problem C: new submissions are accepted and timestamped but return 'Pending review' rather than a verdict. Announce it in the contest UI. The content team fixes the testset, which creates testset v2 — immutable, hashed, validated by two reference solutions. Rejudge all problem C submissions against v2 in the contest lane — about 6K submissions, ~12K CPU-s, a few minutes on the contest pool. Because verdicts are keyed by (submission, testset_version), we generate a diff: how many flip WA→AC. The contest admin approves publishing. Rankings recompute from ingress_ts and penalties — nobody is penalized for the wrong WA. If more than ~30% of the contest is affected, the admin decides whether the contest is unrated."

Why this is L6:

  • Freezes rather than keeps producing wrong verdicts
  • Uses versioned testsets and idempotent verdicts
  • Separates who fixes (content), who executes (platform), who decides (admin)

What L7 adds:

  • Writes the unrated-contest policy in advance so it's not decided under pressure
  • Invests in problem-setter tooling (automated cross-validation) because content errors are the #1 source of contest incidents

Drill 6: Multi-Language Support#

Prompt: "We support 5 languages. Product wants 30."

Staff Answer

"Three costs grow with languages: warm-pool fragmentation, seccomp tuning, and time-limit calibration. Fragmentation: move to a single base VM snapshot containing all runtimes so any warm VM serves any language — bigger image (~4 GB), cached once per host. seccomp: JVM, Go and .NET need extra syscalls (threads, madvise, signals); maintain per-language profiles and a test suite of hostile programs per language. Calibration: each language gets a time multiplier derived by running reference solutions; publish them. The long tail — 25 languages with under 1% of traffic each — gets 'best effort' status: allowed in practice, excluded from rated contests until they have a calibration and an owner."

Why this is L6:

  • Identifies the non-obvious scaling costs of languages
  • Uses tiers to contain the operational burden

What L7 adds:

  • Makes language support a contract: each language has a named owner (community or internal) and gets deprecated without one
  • Prices each language's ongoing cost (security review, calibration) before saying yes

Drill 7: The Run Button Is Being Abused#

Prompt: "Our compute bill doubled. Run is 80% of CPU."

Staff Answer

"Run is the free compute endpoint of the internet, so first look at the distribution: CPU-seconds per account. Legit users use a few CPU-s per hour; farms hug the limit. Fix in layers: Run sandboxes get 1 core, 5s wall, pids=16, lowest lane priority; per-account budget of 60 CPU-s/hour for free, 600 for paid; per-IP/ASN budgets to catch account farms; new accounts start with a smaller budget until they have a history. Detect clustering by source hash across accounts. And make the product change: Run on the sample tests only is cheap; custom input with long runtimes is where the burn is."

Why this is L6:

  • Starts from data (per-account distribution), not a blanket limit
  • Uses CPU-seconds, the actual cost unit, as the budget
  • Separates legit and abusive usage with trust tiers

What L7 adds:

  • Frames it as unit economics: cost per free MAU; sets the budget with finance and product against conversion rate
  • Treats abuse budgets as an org-wide policy shared with other free-compute products

Drill 8: Build vs Buy#

Prompt: "Why not just use a hosted code execution API or run Judge0?"

Staff Answer

"Under 10K executions/day or for an internal tool, I would — building microVM orchestration for that volume is waste. At 10M/day with rated contests, three things push toward owning it: timing determinism (we need pinned cores and host qualification, which hosted APIs don't guarantee), contest capacity scheduling (we need a pool that appears at 7:30 on Saturdays), and cost ($0.00002 per submission of CPU vs typical per-execution API pricing that's orders of magnitude higher). What I'd buy: the isolation primitive — Firecracker or gVisor are open source — and the queue. What I'd build: scheduling, lanes, calibration, rejudge."

Why this is L6:

  • Gives a threshold for when building is justified
  • Identifies which requirements a vendor can't meet
  • Buys primitives, builds the differentiating parts

What L7 adds:

  • Plans the exit: abstracts the sandbox interface so switching isolation technology is a two-way door
  • Evaluates whether the sandbox itself could become a product line (e.g., for assessment customers)

Drill 9: Multi-Region#

Prompt: "We're launching in India and Brazil. Do we need judges there?"

Staff Answer

"Judging is latency-tolerant — a 200ms cross-ocean hop on a 5-second verdict is noise. So practice judging can stay centralized. Contests are different in two ways: the fairness timestamp must be taken close to the user so a 250ms RTT doesn't penalize distant users in tie-breaks — stamp ingress_ts at a regional edge submission service — and a regional outage shouldn't kill a global contest, so I'd want a second judge region capable of absorbing the full contest pool. Data residency matters for B2B assessments in some jurisdictions; those tenants get regional pools. Start with regional ingress + central judging; add regional judges only when residency or contest scale demands it."

Why this is L6:

  • Separates where latency matters (timestamping) from where it doesn't (execution)
  • Calls out residency as the real forcing function

What L7 adds:

  • Sets a regional-launch checklist: residency, contest time zones, and a DR test before the first regional contest
  • Prices a second judge region as insurance vs the reputational cost of a cancelled global contest

Drill 10: Changing the Compiler Without an Outage#

Prompt: "We need to upgrade GCC. How do you roll it out?"

Staff Answer

"A compiler upgrade can change verdicts — different optimizations, different UB behavior, different time. So it's a runtime version, recorded on every verdict. Rollout: shadow-judge 1–5% of real submissions on both versions and diff verdicts and CPU time; investigate any flip. Then canary 5% of practice traffic, then 50%, then 100%. Never switch during a contest; contests pin the runtime version at start. Keep the old runtime available for 90 days so disputes can be reproduced. Announce the change and the new time multipliers."

Why this is L6:

  • Treats a runtime change as a verdict-affecting change requiring shadow diffing
  • Pins versions during contests
  • Keeps reproducibility for disputes

What L7 adds:

  • Standardizes runtime lifecycle (support windows, deprecation notices) across all languages
  • Makes verdict reproducibility a contractual guarantee for B2B, with retention tied to it

8. Deep Dive Scenarios#

Deep Dive 1: The Contest That Melted#

Context: Saturday's contest had 48K participants — 60% more than forecast. Verdict latency hit 9 minutes at t+6min, users resubmitted, and the contest team is asking whether to declare it unrated. The on-call escalates to you at t+10min.

Questions to Surface First:

  • Are rankings actually affected, or only experience? (Is penalty time from ingress_ts or verdict time?)
  • Which lane is lagging — contest, or is the Run lane stealing capacity?
  • What fraction of the queue is duplicate resubmits?
  • How much reserved capacity can arrive in the next 5 minutes, and at what cost?

Typical L5 Approach: Scales out the worker deployment, raises the autoscaler max, and watches queue depth fall over 15 minutes. Correct and slow. Doesn't address resubmit amplification or tell the contest team whether rankings are safe.

Staff Approach: Executes the pre-approved spike playbook: shift weights to contest 90 / practice 10 / run 0, enable source-hash dedupe, show queue position in the UI to stop resubmits, pull reserved hosts. Confirms rankings use ingress_ts, and tells the contest team: "Experience is degraded, outcomes are not — no need to unrate."

Principal Approach: Fixes the forecast input, not the incident: capacity is driven by live registrations with an automatic top-up at +25%, and contests above a size threshold require a capacity review 72h ahead. Negotiates a scheduled-capacity reservation for the weekly calendar so the top-up is guaranteed, and writes the unrated-contest criteria so the decision is policy, not mood.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Lane weight shift; disable Run for non-contest users; enable dedupe; UI banner with queue position
TriageCheck duplicate_submit_ratio, queue_lag_p95{lane}, warm pool per language; confirm ranking timestamp source
Quick fixPull reserved hosts; if lag > 10 min, contest admin extends the contest by the lag duration
GuardrailsRejudge lane paused until contest ends; practice lane floor of 10% so it doesn't fully starve
Post-mortemForecast used last month's registrations; no top-up trigger; client allowed unlimited resubmits

Metrics to Watch: queue_lag_p95{lane=contest}, duplicate_submit_ratio, warm_pool_available{language}, submissions_per_sec{lane}, verdict_latency_p95.

Organizational Follow-up: Contest team provides registration feed to the capacity controller; client team adds a Submit debounce; platform adds a +25% auto top-up.

Ownership Question: "Who decides if the contest is unrated?" Staff answer: The contest admin, using written criteria (e.g., verdict lag > 15 min for > 20% of the contest, or ranking-affecting misjudgment). Platform supplies the data within 30 minutes of contest end; it does not make the call.

Key Takeaway: "If fairness is stamped at ingress, a slow judge is an experience incident, not an integrity incident."

What clears the Staff bar:

  • Asks whether rankings are affected before scaling anything
  • Uses a pre-approved playbook rather than improvising
  • Attacks the amplifier (resubmits), not just the capacity

Deep Dive 2: Silent Timing Drift#

Context: A power-user publishes a blog post showing that the same Rust solution runs 18% slower on your judge on weekends. There are no alerts. Your error rates are flat.

Questions to Surface First:

  • What's different about weekend capacity — host class, co-tenancy, spot instances?
  • Are canaries running per host class, and what are their thresholds?
  • How many borderline TLEs were issued on those hosts?

Typical L5 Approach: Increases time limits by 20% across the board. That hides the drift and makes weak solutions pass on fast hosts.

Staff Approach: Finds that weekend autoscaling uses a cheaper, older instance family. Adds per-class canaries with 5% drift alerts, excludes unqualified classes from the contest pool, and reruns borderline TLEs from the affected window.

Principal Approach: Treats "verdict determinism" as an SLO with an owner and a budget. Instance-family choices that affect verdicts now require qualification; procurement cost savings are weighed against the determinism SLO explicitly, not decided by an autoscaler config.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateQuery verdicts by host class; confirm the slower class
TriageCount TLEs within 20% of limit on that class in the window
Quick fixPin contest and B2B pools to qualified classes; rerun affected borderline TLEs
GuardrailsCanary per class; drain at > 5% drift
Post-mortemWhy did an autoscaler config change verdict semantics without review?

Metrics to Watch: judge.canary_cpu_ms{host_class}, tle_near_limit_ratio{host_class}, verdict_flip_rate_on_rerun.

Organizational Follow-up: Host-class admission checklist; autoscaler instance list requires platform approval.

Ownership Question: "Who owns verdict determinism?" Staff answer: The judge platform team, with an explicit SLO (identical code → identical verdict 99.9%) and the authority to veto infra cost changes that break it.

Key Takeaway: "Timing drift never shows up in error rates. If you don't canary the measurement, you don't know your measurement."

What clears the Staff bar:

  • Refuses the blanket limit increase
  • Builds detection for a failure that emits no errors

Deep Dive 3: Enterprise Assessment Customer Onboarding#

Context: A large company wants to use your judge for 200K candidate assessments per year, with their own problems, custom checkers, SSO, 99.9% availability during scheduled windows, and verdict reproducibility for two years.

Questions to Surface First:

  • Do their checkers run in our sandbox? (They're untrusted code too.)
  • What are the concurrency peaks — e.g., university hiring days?
  • Data residency? Retention? Who handles candidate disputes?

Typical L5 Approach: Adds a tenant_id column, lets them upload problems, runs everything on the shared pool.

Staff Approach: Tenant-dedicated pool (or reserved share) so a public contest can't delay their assessment; checkers sandboxed with the same controls; verdict provenance (runtime digest, host class, testset hash) stored for 2 years; old runtimes retained for reproduction.

Principal Approach: Decides whether B2B is a product line with its own SLA, on-call and pricing, or a side feature. If it's a product line, it gets a separate cell (blast-radius isolation from public contests), a contract review for liability on disputed verdicts, and a runtime deprecation policy written into the contract.

Staff Approach — Full Reasoning
DimensionStaff Answer
IsolationReserved pool per large tenant; checkers in sandbox
CapacityScheduled windows from customer calendar → pre-warm like contests
AuditVerdict provenance + 2-year retention; reproducible rejudge
Availability99.9% in windows; separate on-call rotation with customer escalation
DisputeReproduce on archived runtime; report with CPU time trace

Metrics to Watch: tenant_queue_lag_p95, tenant_availability_in_window, checker_sandbox_violations.

Organizational Follow-up: Support playbook for disputes; legal review of verdict liability.

Ownership Question: "Who is paged when an assessment window fails?" Staff answer: Platform on-call for the pool, with a named customer success escalation — and the SLA credits come out of a budget someone owns.

Key Takeaway: "Enterprise customers don't buy execution — they buy a verdict they can defend in a dispute."

What clears the Staff bar:

  • Notices the custom checker is untrusted code
  • Designs for reproducibility, not just correctness

Deep Dive 4: Post-Mortem — Test Data Exfiltration#

Context: Hidden test cases for 40 premium problems appeared on GitHub. Investigation shows users printed them via a side channel: the sandbox gave different error messages depending on input content, and a script submitted thousands of probes.

Questions to Surface First:

  • What does the verdict reveal? (Failing case index, input on failure, stderr?)
  • Was the testset readable from inside the sandbox at any point?
  • How many accounts, and what's the probe rate?

Typical L5 Approach: Rotates the tests and blocks the accounts.

Staff Approach: Classifies it as an information-leak design flaw: verdict detail + unlimited submissions = oracle. Limits what's returned (first failing case index only for hidden tests; no stderr for hidden tests), adds per-problem submission budgets, detects probe patterns (many submits with minimal source diffs, WA on incrementing case index), and regenerates testsets with randomized large cases.

Principal Approach: Recognizes this as a product-policy question with security consequences: how much feedback do we give vs how much we leak. Sets the policy jointly with product (practice problems get rich feedback on sample tests only), and adds leak detection (monitoring public repositories for testset fingerprints) as a standing capability.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateSuspend probing accounts; reduce verdict detail for hidden tests
TriageMap which problems leaked and via which feedback channel
Quick fixRegenerate testsets (v+1), rejudge premium problems
GuardrailsProbe detector on source-diff similarity and submit cadence
Post-mortemFeedback design was never threat-modeled

Metrics to Watch: submits_per_user_per_problem_p99, min_source_edit_distance_between_submits, testset_fingerprint_matches_external.

Organizational Follow-up: Threat-model review for any new verdict detail feature.

Ownership Question: "Who approves adding more detail to verdicts?" Staff answer: Product proposes, security reviews, platform implements — and it goes through the same review as an API change.

Key Takeaway: "A judge with unlimited attempts and rich feedback is an oracle. Feedback is a security surface."

What clears the Staff bar:

  • Identifies the oracle pattern, not just the leak
  • Pairs technical fixes with a feedback policy

Deep Dive 5: Multi-Region Expansion for a Global Contest#

Context: Leadership wants a 250K-participant global contest with users on every continent. Current judge runs in one US region.

Questions to Surface First:

  • Where is ingress_ts taken, and does RTT affect tie-breaks?
  • Can one region absorb ~2,500 submits/s peak?
  • What happens if the judge region fails mid-contest?

Typical L5 Approach: Deploys judges in three regions and routes users to the nearest.

Staff Approach: Regional ingress stamps ingress_ts (clocks disciplined to < 10ms via NTP/PTP), submissions replicate to a central queue; judging stays in two regions (active + warm standby sized for 100%). Test data prefetched to both. Lane weights and capacity from registrations by region.

Principal Approach: Weighs the one-off cost (~2 regions × peak pool × a few hours) against the reputational cost of a failed flagship event; runs a game day with a synthetic 250K-user contest two weeks prior; decides whether a staggered start (problems released by region) is a product answer that removes the infrastructure problem.

Staff Approach — Full Reasoning
DimensionStaff Answer
FairnessTimestamp at regional edge; bound clock skew; tie-break by ingress_ts then submission id
Capacity2,500/s × 2 CPU-s × 2 headroom ≈ 10K cores per region
DRStandby region pre-warmed; queue replicated; failover drill
DataTestsets prefetched and hash-verified in both regions

Metrics to Watch: ingress_clock_skew_ms, cross_region_replication_lag, queue_lag_p95{region}.

Organizational Follow-up: Game day; comms plan for region failure.

Ownership Question: "Who calls the failover mid-contest?" Staff answer: The incident commander from the platform team, within a pre-agreed 5-minute decision window, with the contest admin informed — not consulted.

Key Takeaway: "Latency matters where you take the timestamp, not where you run the code."

What clears the Staff bar:

  • Separates fairness latency from execution latency
  • Plans DR as a drill, not a diagram

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • State the threat model in the first 3 minutes and design a layered sandbox with a VM boundary
  • Explain why test data must never enter the sandbox
  • Size a contest pool from registrations and explain why reactive autoscaling fails
  • Decouple ranking fairness from queue latency with ingress timestamps
  • Explain CPU time vs wall time, noisy neighbors, SMT, and canary-based drift detection
  • Design rejudge as an idempotent, versioned, diff-reviewed workflow
  • Distinguish Run and Submit economically and design abuse budgets in CPU-seconds
  • Articulate the content-vs-platform ownership seam and the org policy for unrated contests

The Bar for This Question#

Mid-level (L4): API, queue, worker in Docker, database. Mentions timeout and no network. Needs prompting to discuss spikes.

Senior (L5): Adds autoscaling, retries, per-user rate limits, maybe multiple languages. Correct system for a normal day. Security story is "containers"; timing is wall-clock; rejudge is not considered.

Staff+ (L6): Treats code as hostile, stacks containment inside a VM boundary, keeps test data outside, pre-warms for scheduled spikes, stamps fairness at ingress, measures CPU time on pinned cores with canaries, and owns the wrong-verdict path through versioned testsets and rejudge. Names who waits, who pays, and who decides. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Containers Are Not a Security Boundary for Public Judges"#

EvidenceImplication
Containers share the host kernelOne kernel CVE on an allowed syscall = host compromise
AWS built Firecracker for multi-tenant serverlessThe hyperscaler's own answer to "is a container enough?"
A public judge runs millions of arbitrary programs dailyExploit attempts are a certainty, not a risk

The Staff position: Containers are a containment mechanism. The security boundary is the VM (or gVisor). Pay the ~125ms and ~10–15% density.

Why this matters in interviews: Saying "Docker is secure enough" is the single fastest way to cap your level on this problem.

10.2 "Queue Latency Doesn't Matter for Contest Fairness"#

EvidenceImplication
Penalty can be computed from ingress timeQueue position becomes irrelevant to ranking
Resubmits are driven by perceived lagShowing queue position is a better fix than more hosts

The Staff position: Stamp at ingress and spend your money on determinism, not on shaving seconds off a verdict nobody is ranked by.

Why this matters in interviews: It shows you can find the invariant that makes a scary problem cheap.

10.3 "Most Judge Incidents Are Content Bugs, Not Platform Bugs"#

EvidenceImplication
Wrong expected outputs and weak tests are the recurring contest complaint in competitive programming communitiesThe highest-leverage investment is testset validation tooling
Platform failures are rarer once sandbox and capacity are matureEngineering attention should shift to the content pipeline

The Staff position: Build the problem-setter validation pipeline (multiple reference solutions, known-wrong solutions, stress generators) before optimizing the runtime further.

Why this matters in interviews: Naming the non-obvious top source of incidents shows operational experience.

10.4 "The Run Button Should Cost Users Something"#

EvidenceImplication
Run is 3–5× Submit volumeIt's the majority of compute
It's the main abuse vectorFree unlimited CPU attracts farms

The Staff position: Budget Run in CPU-seconds per user, visible in the UI, with larger budgets for paid tiers. Free-and-unlimited is a policy choice with a monthly bill.

Why this matters in interviews: It turns a rate-limit answer into an economics answer.

10.5 "Parallelizing Test Cases Is Usually a Mistake"#

EvidenceImplication
Most submissions fail within the first ~10 tests or finish in 1–3s totalParallel fan-out saves little latency
Fan-out multiplies VM count 50–100× per submissionDensity and warm-pool demand explode during spikes

The Staff position: Run tests sequentially with early exit in one sandbox; parallelize only for problems with long individual tests, and only outside contests.

Why this matters in interviews: It shows you evaluate an "obvious optimization" against the spike it would make worse.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer designs a judge. A Principal engineer notices that "execute untrusted code safely" is about to be requested by the assessment product, the AI tutor that runs model-generated code, a notebook feature, and a plugin system — and that four teams building four sandboxes means four security postures, four CVE response times, and four cost curves. The L7 question is not "which sandbox" but "which sandbox platform, with what contract, owned by whom, and what do we refuse to support."

The Org-Level Fault Line#

One execution platform vs per-product sandboxes. Centralizing gives one security posture, one patch SLA, shared warm capacity and shared calendar-driven scheduling. It also creates a bottleneck team and forces the judge's determinism requirements (pinned cores) onto products that don't need them. The L7 answer: centralize the isolation primitive and the security contract (image, seccomp profiles, VM boundary, patch SLA); let products own scheduling policy and pools (lanes, determinism class, retention).

Cost Model#

Assumptions: $0.04 per core-hour (bare metal amortized), ~2 CPU-s per submission, engineers at $250K fully loaded.

ScaleWorkloadCompute / monthHeadcountOn-call load
Small100K submissions/day, monthly contest~$2K (mostly idle warm floor)1–2 engineers part-time; buy/open-source judgeBusiness hours; ~1 page/month
Medium10M/day, weekly 30K contest, Run 3×~$25–40K (practice + Run + ~4 contest windows + warm floor)5–7 engineers (runtime, capacity, security liaison)24/7 rotation; ~4–6 pages/month, clustered on contest days
Large50M/day, daily contests, B2B assessments, AI tutor executions~$150–250K plus reserved-capacity commitments12–18 across platform + B2B cell + securityTwo rotations (public, B2B); SLA credits budget

The dominant cost at medium scale is idle warm capacity and pinned-core density loss, not execution. The dominant cost at small scale is engineers — which is why buying wins there.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
Verdict keyed by (submission, testset_version), append-onlyOne-way (good one)Retrofitting later means migrating billions of rows and losing history
Ingress-timestamp fairness modelOne-way (publicly visible)Changing ranking semantics mid-season breaks trust
Container vs microVM isolationTwo-way if the sandbox is behind an interfaceWeeks, if the harness talks to an abstract "sandbox"; months if Docker APIs leak everywhere
Queue technology (Kafka vs SQS)Two-wayDays to weeks with a dispatcher abstraction
Bare metal vs cloud for contest poolMostly two-wayReserved commitments are 1–3 year contracts
Exposing test-case details in verdictsOne-wayOnce leaked, tests must be regenerated forever

The Standard I'd Write#

RFC: Untrusted Code Execution Standard (v1)

Scope: Any service that executes code not authored by our employees — judge, assessments, AI-generated code, plugins.

MUST:

  • Execute inside the platform sandbox runtime (VM boundary + seccomp + cgroups v2 + no default network).
  • Never mount secrets, credentials or evaluation data inside the sandbox.
  • Record runtime image digest and host class on every result that is shown to a user.
  • Patch guest and host kernels within 72h of a critical CVE.
  • Enforce per-principal CPU-second budgets.

SHOULD:

  • Use calendar-driven capacity for scheduled events.
  • Declare a determinism class (strict / standard / best-effort) per workload.

Exceptions: Filed with the platform team and security; time-boxed to 90 days; reviewed quarterly.

Success metrics: Zero host compromises; ≤ 72h critical patch time; false-TLE rate < 0.1% on strict class; ≤ 1 execution platform per company.

What I'd Tell the VP#

Our code execution is the core of the product and our biggest security exposure: we run millions of programs a day written by strangers. Today it's safe enough, but contests stall when traffic beats our forecast, and timing inconsistencies occasionally make results look unfair — which is what our community notices. I'm proposing we invest roughly five engineers for two quarters to make contest capacity scheduled rather than reactive, make verdicts reproducible, and turn the sandbox into a shared platform that the assessment product and AI tutor will also use. That avoids three teams each building their own sandbox, and it converts our largest reputational risk into a measured SLO.

Principal Interview Signals#

SignalWhat It Sounds Like
Sees the class of systems"Assessments and the AI tutor will need this within a year. I'd build a platform with a contract, not a feature."
Prices the tradeoff"Pinning costs 2× hardware on the contest pool — about $X/month — and I'd pay it only there."
Separates one-way doors"The verdict key and the fairness model are one-way doors; the isolation tech is two-way if we put it behind an interface."
Designs org failure posture"Kernel patch SLA of 72h, sandbox escapes in bug bounty scope, B2B in its own cell."
Knows when not to standardize"Centralize isolation and security; don't force contest determinism on the notebook team."

Staff answers that L7 interviewers find insufficient:

  • "We'll use Firecracker" — correct, but doesn't address who owns patching or how the next team reuses it.
  • "Pre-warm for contests" — correct, but doesn't price the idle pool or propose reserved capacity.
  • "Rejudge with a diff report" — correct, but doesn't set the policy for when a contest is unrated, or who decides.

🧭 Principal Move: "The judge is our first code-execution product, not our last. I'd spend the extra quarter making the sandbox a platform with a written contract, because the alternative is four sandboxes with four different answers to the next kernel CVE."


Appendices

Appendix A: Sandbox Mechanics in Depth#

A.1 The Execution Harness#

The harness runs outside the user's sandbox and owns every trusted artifact.

judge(submission, testset):
    vm = warm_pool.acquire(runtime=submission.language)        # ~10ms
    vm.put_file("/work/main.src", submission.source, ro=False)
    r = vm.exec(compile_cmd, cpu_limit=10s, wall=30s, mem=1GB)
    if r.status != OK: return verdict(COMPILE_ERROR, r.stderr[:4KB])
    for i, case in enumerate(testset.cases):                    # streamed from local NVMe cache
        r = vm.exec(run_cmd, stdin=case.input,
                    cpu_limit=problem.tl, wall=3*problem.tl,
                    mem=problem.ml, pids=64, stdout_cap=64KB)
        if r.cpu_ms > problem.tl:      return verdict(TLE, i, maybe_rerun=near_limit(r))
        if r.oom:                       return verdict(MLE, i)
        if r.exit != 0:                 return verdict(RE, i)
        if not checker(case, r.stdout): return verdict(WA, i)   # checker runs outside, or in its own sandbox
    return verdict(AC, max_cpu, max_mem)
    finally: vm.destroy()                                        # never reuse without snapshot restore

A.2 seccomp Allowlist Strategy#

  • Start from an allowlist (~60 syscalls: read, write, mmap, brk, exit_group, futex, clock_gettime, …).
  • Per-language additions (JVM: clone with thread flags, sched_yield, madvise; Go: signals).
  • Default action KILL_PROCESS and emit sandbox.seccomp_kill_total{syscall} — kills on odd syscalls are attack telemetry.

A.3 Why Destroy, Not Clean#

A "cleaned" sandbox must prove the absence of: background processes, modified files, changed env, kernel state (e.g., keyrings), and page cache side channels. Destroy-and-reacquire from a warm pool (or snapshot restore at ~tens of ms) is cheaper than proving cleanliness.

Appendix B: Data Model#

TableKeyNotes
submissionssubmission_id (ULID)user_id, problem_id, contest_id, language, source_hash, ingress_ts, lane; source in object store
testsets(problem_id, version)content_hash, validated_by[], approved_by, immutable
verdicts(submission_id, testset_version, runtime_version)append-only; status, cpu_ms, mem_kb, failed_case_idx, host_class
current_verdictsubmission_idpointer to the approved verdict row
contest_standingsRedis sorted set per contestscore = solved × 10^9 − penalty_seconds
Diagram: Appendix B: Data Model

Sharding: submissions and verdicts by submission_id hash; user history via a secondary index on (user_id, problem_id, ingress_ts). See Sharding.

Appendix C: Dispatch and Queueing#

C.1 Weighted Fair Dispatch#

weights = contest_active ? {contest:70, practice:25, run:5} : {practice:80, run:20}
on worker_slot_free:
    lane = weighted_choice(nonempty_lanes, weights)       # deficit round robin
    job  = lane.pop_fair_by_user()                        # max 2 in-flight per (user, problem)
    if cache.has(job.source_hash, testset_version): return cached verdict

C.2 Quick Comparison#

MechanismStrengthWeaknessUse For
Kafka partitions per laneDurable, replayablePer-user fairness needs consumer logicLarge scale, rejudge replay
SQS queue per laneManaged, simpleNo ordering guarantee; visibility-timeout tuningMedium scale
Redis streamsLow latencyDurability tuning requiredSmall scale
Postgres SKIP LOCKEDTransactional with submissionsTops out ~few K jobs/sSmall scale, simplicity

Durability rule: write the submission row before enqueue; a reconciler re-enqueues QUEUED rows older than 60s. See Message Queues.

Appendix D: Client Contract#

  • POST /submissions returns 202 with submission_id and queue_position.
  • Poll GET /submissions/{id} with backoff 500ms → 1s → 2s (cap 2s); or SSE on contest pages (Real-time Updates).
  • Submit button disabled until verdict or 30s; identical source within 60s returns the existing submission_id.
  • 429 with Retry-After when user budget exhausted; 503 with a reason for Run during peak shedding.

Appendix E: Observability#

E.1 Core Metrics#

queue_lag_p95{lane}                  verdict_latency_p95{lane}
submissions_per_sec{lane}            duplicate_submit_ratio
warm_pool_available{language}        host.allocated_sandboxes - host.active_executions
judge.canary_cpu_ms{host_class}      tle_near_limit_ratio{host_class}
wa_concentration_by_case{problem}    sandbox.seccomp_kill_total{syscall}
run_cpu_seconds_per_account (p99)    rejudge_lane_utilization

E.2 Critical Alerts#

AlertThresholdAction
Contest lagqueue_lag_p95{lane=contest} > 30s for 2 minPage; spike playbook
Timing driftcanary drift > 5% on a host classAuto-drain class; ticket
Test data suspectone case > 50% of WAs after 200 submitsNotify content on-call
Sandbox anomalyseccomp kills > 10× baselinePage security + platform
Capacity leakallocated − active > 5% of host for 10 minRecycle host

E.3 Debugging the Silent Failure#

Timing drift and zombie sandboxes emit no errors. The only defense is synthetic workload: canary reference solutions and teardown checks that run continuously and compare against baselines.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 10K/dayOpen-source judge on a few VMs, containers + gVisorFirst contest spike
10K–1M/dayQueue + autoscaled workers, lanes, versioned testsetsTiming variance, warm-pool fragmentation
1M–50M/dayFirecracker pools, calendar pre-warm, canaries, rejudge laneMulti-product demand, B2B SLAs
50M+/dayExecution platform with cells per product, reserved capacityOrg coordination, cost attribution

What You Don't Build on Day One#

  • Regional judges (ingress timestamping is enough until residency forces it)
  • Parallel test execution
  • Custom hypervisors or kernels
  • Plagiarism detection in the hot path (it's an offline post-contest job)

Appendix G: Multi-Tenancy and Cost#

  • Chargeback unit: CPU-seconds by lane and product; warm-floor cost allocated by reserved share.
  • Fairness within Run: per-account and per-ASN CPU-second budgets; trust tiers by account age and history.
  • B2B isolation: reserved share or dedicated cell; a public contest must never delay a paid assessment window.
  • Tradeoff summary: strict determinism costs ~2× hardware; VM boundary ~10–15% density; calendar pre-warm costs hours per week instead of a permanently idle peak pool. Spend where verdicts are ranked or contracted, save everywhere else.
  1. Loading the index…