Hiring BarSupport

Design a Multiplayer Game Backend

Case study97 min read10 diagrams

Technologies referenced in this case study: Kubernetes · Redis · Apache Kafka · Cassandra · PostgreSQL

Related: Real-Time Updates with WebSockets · Push vs Poll · Leaderboard · Load Balancing · Autoscaling & Capacity · Latency & Protocols · Auth & Identity · Deployment System · Backpressure · Dating & Proximity Matching

Reading Guide#

Organized for interview use first, reference second. This page designs the backend for a session-based multiplayer game: matchmaking, game-server fleet allocation, the authoritative simulation loop and its network budget, lag compensation, the trust boundary with the client, and the path from "match ended" to durable progression. Persistent sockets for chat and presence live in Real-Time Updates with WebSockets; ranking boards live in Leaderboard; the cluster substrate lives in Kubernetes. This page links to them rather than repeating them.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on the tick loop, lag compensation and allocation
What is a Multiplayer Game Backend? — Why interviewers pick this topic

A multiplayer game backend is everything between "Play" and "Victory" that is not the game client: the login and party services, the matchmaker that groups players into fair matches, the fleet of dedicated game servers that run each match, the simulation inside each server that decides what actually happened, and the post-match pipeline that turns results into ranks, rewards and leaderboards.

The hard part is not storing a player profile. The hard part is that two very different systems share one product. The match itself is a hard-real-time, single-process, in-memory simulation that must finish a frame every 15.6 ms and ship state to ten players over lossy UDP. Everything around the match — queues, allocation, results, inventory — is an ordinary distributed system with durable state, retries and idempotency. Candidates who treat them as one system either put a database on the tick path or try to make a 25-minute match "highly available".

Before vs After — the "patch day" scenario:

Without a designed backend:
t=0:        Patch goes live at 18:00 local in the largest region. 600K players log in
            within 10 minutes, all queueing for the new mode.
t=+2min:    Matchmaker forms 300 matches/s. Fleet had 400 warm servers; cold starts
            take 90 s (image pull + 2 GB map assets). Matches form with no server.
t=+4min:    Tickets time out after 60 s. Clients auto-requeue with no backoff.
            Queue size triples. Matchmaker CPU at 100%, match quality collapses.
t=+9min:    New build crashes at minute 12 of every match on one map. 20% of live
            matches die. Players lose ranked points for "abandoning".
t=+40min:   Support queue at 50K tickets. Social media trending. Rollback impossible:
            clients already patched.

With a designed backend:
t=-24h:     Fleet pre-scaled from forecast: warm buffer 3x normal in the region;
            images and map assets pre-pulled onto every node.
t=0:        Login admission gate meters 2,000 logins/s; queue position shown to players.
t=+2min:    Matchmaker forms 60 matches/s; allocation p99 1.8 s from the warm buffer.
t=+9min:    Canary build (5% of new allocations) shows crash rate 40x baseline on one
            map. Allocation selector stops choosing the new build for that map.
t=+10min:   Crashed matches voided: no rank change, players returned to queue with
            priority. Blast radius: ~1,500 matches, not 40,000.

Why interviewers reach for this question: It is the cleanest test of whether you can separate a real-time hot path from a durable control path and price both. Every candidate can draw "client → server → database". Few notice that the per-tick budget forbids network calls, that the expensive resource is egress bandwidth, that matchmaking is a queueing problem with a quality dial, and that the client is an adversary who holds a copy of your code.

Mechanics Refresher: The Netcode and Session Primitives
PrimitiveHow It WorksProsCons
Dedicated authoritative serverOne server process owns the match state; clients send inputs, server sends stateSingle source of truth; cheat-resistant; fair for all clientsYou pay for every server-minute and every byte of egress
Listen server (player hosts)One player's machine is the serverZero server costHost advantage; host quits = match ends; host can cheat
Peer-to-peer lockstepEvery peer runs the same deterministic simulation from the same inputsTiny bandwidth (inputs only); great for RTS and fighting gamesNeeds bit-exact determinism; slowest peer sets the pace; every peer sees everything
Relay serverA server forwards packets between peers but simulates nothingHides IPs, traverses NAT, cheap CPUNo authority; cheating unchanged
Tick (simulation step)Server advances the world in fixed steps (e.g., 64 per second, 15.6 ms each)Deterministic ordering of inputs; predictable CPUHigher tick = more CPU and bandwidth
Snapshot + delta compressionServer sends world state relative to the last snapshot the client acknowledgedBandwidth drops 5–10× vs full stateNeeds per-client ack tracking on the server
Client-side predictionClient applies its own inputs immediately and reconciles when the server state arrivesYour own movement feels instantMispredictions cause visible corrections
Entity interpolationClient renders other players slightly in the past, between two received snapshotsSmooth movement despite loss and jitterEveryone else is shown ~30–100 ms late
Lag compensationServer rewinds other players to what the shooter saw when validating a hitHits register where you aimedThe target can be hit after reaching cover ("shot behind the wall")
Interest managementServer sends each client only what is relevant (nearby, visible)Bandwidth and cheat exposure both dropCPU per client per tick
Matchmaking ticketA request (player or party, attributes, rating) that waits in a pool until a match function selects itDecouples queueing from match logicPool fragmentation by mode, region and party size
Fleet + allocationPre-started Ready servers; an allocation marks one Allocated and returns its addressAllocation in milliseconds instead of a 30–90 s cold startYou pay for idle Ready servers

For most production systems: a dedicated authoritative server per match for anything competitive, a fixed tick with delta-compressed snapshots and interest management, client prediction plus interpolation plus server-side lag compensation with a bounded rewind window, a rule-based matchmaker whose skill window widens with wait time, and a pre-warmed fleet on Kubernetes (via a game-server operator such as Agones) or a managed hosting service. The primitives are not the interview — the budgets (per-tick CPU, per-player bandwidth, per-ticket wait), where trust ends, and what happens to a match when its server dies are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

A multiplayer game backend is not a WebSocket fan-out question. Anyone can relay messages between ten clients.

It is a budgets-and-trust question that tests:

  • Whether you separate the real-time match (in-memory, single process, disposable) from the durable control path (queues, results, progression) in the first five minutes
  • Whether you can turn "make it feel responsive" into numbers: a 15.6 ms tick budget, ~140 kbps per player down, a 200 ms rewind cap, a 30-second p50 queue
  • Whether you treat matchmaking as a queueing system with a quality dial, and name who pays when the dial moves
  • Whether you draw the trust boundary at the client: the server decides outcomes, and the client only learns what it needs to render

The key insight: The match server is the one place in the system where availability is the wrong goal. A 25-minute match does not need replication, consensus or failover; it needs to run at a fixed tick on a dedicated core and, if its host dies, to end cleanly — voided, players compensated, requeued with priority. The durability budget belongs to the things that outlive the match: the result, the rating change, the items earned. Staff candidates make the match disposable and the result exactly-once, and spend the rest of the interview on the three budgets that actually decide cost and fairness: CPU per tick, bytes per player, seconds per ticket.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws clients → WebSocket gateway → game service → Redis for state → databaseAsks "Competitive shooter, persistent world, or turn-based? How many players per match, what tick rate, and how much latency can the genre hide?"Asks "How many titles will share this platform, what's our egress bill per player-hour, and which parts — identity, matchmaking, fleet, anti-cheat — should every studio stop rebuilding?"
Authority"The server validates moves""One dedicated authoritative process per match. Clients send inputs only; the server simulates, decides hits, and sends each client a filtered, delta-compressed snapshot. Clients predict their own movement and reconcile."Sets the netcode model per genre as a platform decision: authoritative servers for competitive, lockstep or relay for fighting and RTS titles, and owns the cost envelope for each
Matchmaking"Group players with similar MMR from a Redis sorted set""A pool per region × mode × party size. Start at ±100 rating and ≤ 60 ms RTT, widen 50 every 10 s, relax latency after 45 s. Target p50 30 s, p95 90 s. The top 0.1% wait longer by policy."Treats match quality as a product metric with an owner: publishes the wait-vs-fairness curve per mode, and decides which modes to merge when population thins
Failure"Replicate game state so another server can take over""Matches are disposable: if a server dies, the match is voided, ratings untouched, players requeued with priority. Results are durable and idempotent by match_id. The fleet keeps a warm buffer so allocation never waits on a cold start."Designs the org's posture for launch day and patch day: forecasted pre-scaling, login admission control, canary builds by allocation share, and a game day that kills a region's fleet
Trust"We'll add anti-cheat software""The server decides every outcome; the client is untrusted. Lag compensation rewinds at most 200 ms. Fog of war: the server doesn't send enemy positions the player can't see. Results come only from the game server."Treats anti-cheat as a cross-title trust platform: shared detection, shared ban identity, appeal process, and a policy on what the client may run
Scale"Add more game servers""2M players in match = 200K matches. At 5 matches a core, 40K cores. At ~140 kbps per player, ~285 Gbps of peak egress — egress, not compute, is the bill."Prices it: hybrid bare metal for baseline and cloud for bursts, negotiated egress, and a per-title cost-per-player-hour target the studio signs
Why "authority" separates levels

L5: "Clients send their position and actions over a WebSocket to a game service, the service validates them and broadcasts to the other players, and we keep the match state in Redis so any instance can serve it." Reasonable for a turn-based game. For a shooter it puts a network round trip to Redis inside a 15.6 ms frame, uses TCP so one lost packet stalls everything behind it, and trusts the client's position — the input from which every movement cheat is built.

L6: "Each match runs in one dedicated server process that owns the state in memory. Clients send inputs — buttons, aim, a client tick number — over UDP, about 64 times a second. The server applies inputs in tick order, simulates, and decides every hit. Each client gets a snapshot delta-compressed against the last one it acknowledged, filtered to what that player could plausibly see. The client predicts its own movement so it feels instant, interpolates everyone else about two snapshots in the past, and the server rewinds up to 200 ms to check a shot against what the shooter actually saw. No database is ever on the tick path."

L7: "We ship a competitive shooter, a fighting game and a co-op survival game. They should not share a netcode model — the fighting game wants rollback over peer-to-peer with a relay, the shooter wants authoritative servers. They should share identity, matchmaking infrastructure, fleet management, telemetry and anti-cheat. I'd make the netcode model a per-genre choice with a published cost envelope, and everything around the match a platform."

Why "failure" separates levels

L5: "If a game server crashes, we replicate the state to a standby so it can take over the match." That costs a second process per match, a state stream every tick and a failover protocol — for a 25-minute session whose state is worthless once the match ends. It also doesn't address the failures players actually experience: no server available when a match forms, or a bad build crashing thousands of matches at once.

L6: "I design the match to be disposable and the result to be durable. A crashed match is voided: no rating change, any match-only progress discarded, everyone requeued at the front with a 'match ended unexpectedly' message. The rate I track is crashed matches per 10,000, by build and map. The failures I design hardest against are upstream: running out of Ready servers in a region — so the fleet keeps a warm buffer sized to allocation rate × (server start time + autoscaler interval) × 1.5 — and a bad build, so new builds take 5% of new allocations first and the selector stops choosing them when their crash rate rises."

L7: "Our worst days are predictable: launch, patch day, the season reset, a streamer event. I'd make 'pre-scale from forecast, admit logins at a metered rate, canary every build by allocation share' the standard for every title, and run a game day before each season that drains a region's fleet during peak to prove matchmaking fails over without a requeue storm."

Why "matchmaking" separates levels

L5: "Put waiting players in a sorted set by MMR and pair neighbors." It works for one queue at one time of day. It doesn't say what happens to the 2,800-rated player at 4 a.m., to a five-stack of mixed skill, or to a region with 300 people in queue for a niche mode.

L6: "Matchmaking is a queue with a quality dial. Every ticket starts strict — tight rating window, low latency cap, party-size matching — and relaxes with time, so wait is bounded and quality degrades predictably. I'll publish two numbers per mode: p95 wait and p95 rating spread within a match. The dial between them is a product decision, owned by design, and I'd expose it as config, not code."

L7: "Pool fragmentation is the real enemy. Every mode, region, platform and party-size split cuts the pool, and a thin pool forces either long waits or bad matches. I'd give product a population budget: a new mode must show it doesn't push any existing ranked pool below its quality floor, or it ships as a rotating mode."

Positions to Commit To#

PositionRationale
Dedicated authoritative server for anything competitiveThe only model where one process decides outcomes and the client never holds hidden state it doesn't need
Matches are disposable; results are exactly-onceReplicating a 25-minute simulation buys little; losing a rating change or an item costs trust and support tickets
No network call on the tick pathA 64 Hz frame is 15.6 ms; one 2 ms database read per player per tick blows the budget
Egress is the bill; design bytes per player first~140 kbps × millions of players dwarfs compute; interest management and delta compression are cost controls
Matchmaking widens with wait time, by published policyStrict-then-relax bounds wait and makes quality loss predictable and owned
Keep a warm buffer sized to allocation rate × (start time + scaling interval)Cold starts of 30–90 s turn a queue pop into a timeout and a requeue storm
The client is an adversary that holds a copy of your codeInputs only, outcomes on the server, bounded rewind, and no data the player shouldn't see

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Competitive real-time session (team shooter, battle arena, 10–100 players, 10–40 min)Sub-100 ms feel, fairness under latency, cheating paysDedicated authoritative server per match, 30–128 Hz tick, prediction + interpolation + lag compensation, rating-based matchmaking, warm fleetNo server when a match forms; bad build crashes matches; unfair matches; hit-registration disputesAllocation p99 ≤ 2 s; crashed matches < 10 per 10,000; p95 queue ≤ 90 s; p95 in-match RTT ≤ 60 ms for 90% of players
Persistent world (MMO, survival server, 100s–1,000s per shard, months)State outlives every session; density spikes in one placeZone or cell servers with interest management, state persisted on checkpoints and on economic events, shards or layers for populationOne overloaded zone; item duplication on crash; lost progressNo item or currency created or lost across a crash; zone frame time ≤ budget at the design density
Casual / turn-based / async (puzzle, cards, board games, 2–8 players)Seconds of latency are fine; sessions span hours or daysStateless HTTP or a WebSocket gateway, game state in a database, server validates each move, push notifications on your turnStale state on reconnect; double-applied movesEvery move applied exactly once in order; game recoverable from storage at any point

🎯 Staff Move: "I'll design the competitive real-time session — it's where the hard budgets live: tick CPU, per-player bandwidth, lag compensation and fleet allocation. A persistent world reuses the matchmaking, identity and fleet pieces but swaps the disposable match for zone servers with checkpointed state, and I'll point out where. Turn-based is a different, much simpler system — a database and push notifications — so I'll set it aside unless you want it."

Where the Design Splits#

#Fault LineThe Tension
1Authority Model: Dedicated Server vs Peer-to-Peer vs RelayFairness, cheat resistance and a single source of truth vs server cost and operating a global fleet
2Tick Rate and Send Rate vs CPU and BandwidthResponsiveness and hit precision vs cores per match and egress per player
3Lag Compensation: Favor the Shooter vs Favor the TargetHits that land where you aimed vs dying behind cover; how far back the server is willing to rewind
4Matchmaking: Match Quality vs Wait TimeFair, close matches vs bounded queues, with pool fragmentation making both worse
5Fleet Allocation: Warm Buffer vs Cold Start, Packed vs SpreadInstant allocation and isolation vs idle-server cost and node utilization

How Real Companies Built It#

Why this section belongs here: Game studios publish unusually concrete engineering numbers. Each of these shows a fault line from this page in production.

Valve — Source Engine Multiplayer Networking#

Valve's developer documentation describes the Source engine's model in detail: a client-server architecture in which the server simulates the world in discrete ticks — by default a 15 ms timestep, about 66 ticks per second — clients send user commands and receive snapshots (20 per second by default), snapshots are delta-compressed, clients render other entities with a default 100 ms interpolation delay and predict their own movement, and the server performs lag compensation by computing command execution time = current server time − packet latency − client view interpolation and moving other players back to where they were at that time before checking a hit (Valve Developer Community).

Staff insight: This is the reference model for every authoritative shooter: inputs up, snapshots down, prediction for yourself, interpolation for others, rewind on the server for hits. The interview-worthy detail is that the defaults (66 Hz simulation but only 20 Hz snapshots, 100 ms interpolation) were a deliberate bandwidth and smoothness trade, not a limitation — the knobs are tick, send rate and interpolation delay, and each one has a price.

Riot Games — VALORANT's 128-Tick Servers, Netcode and Fog of War#

Riot's engineering posts say VALORANT required 128-tick servers so defenders have time to react to players peeking them, which means a full server frame every 7.8125 ms; to offer 128-tick to every player at an affordable cost the team needed better than three games per core (on its usual 36-core hosts, 108 games or 1,080 players per host), which after reserving 10% for OS and scheduling overhead set a budget of 2.34 ms per frame — and it took the server frame from about 50 ms early in development to sub-2 ms (Riot Games — 128-tick servers). Its netcode post describes a server-authoritative model "to limit the types of cheats possible", client prediction with the server as authority, a server that rewinds to what the shooter saw — with limits on how far it will rewind — and an aim of 35 ms ping for 70% of players (Riot Games — VALORANT netcode). Its Fog of War system withholds enemy positions the player cannot see, using precomputed visibility lookups that take under 2% of server frame time (Riot Games — Fog of War).

Staff insight: Three lessons in one title. Tick rate is a product decision priced in cores per game — Riot did the arithmetic from frame budget to games per core to hosts. Server authority is framed explicitly as an anti-cheat boundary. And information denial (don't send what the player can't see) is the defensive design that makes a whole class of client cheats pointless, at a measured CPU cost.

Agones and Open Match — Open-Source Fleet and Matchmaking Frameworks#

Agones is an open-source platform built on Kubernetes for hosting, scaling and orchestrating dedicated game servers, with GameServer, Fleet, GameServerAllocation and FleetAutoscaler resources (Agones overview). Fleet updates and scale-downs skip Allocated game servers, which are "not deleted until they are specifically shutdown through the game servers SDK, as they are expected to have players on them" (Agones fleet updates); the Buffer autoscaling policy keeps a configured number of Ready servers, evaluated every 30 seconds by default (Agones FleetAutoscaler); and the Packed scheduling strategy bin-packs game servers onto as few nodes as possible for cloud environments while Distributed spreads them for static clusters (Agones scheduling). Open Match, co-founded by Google Cloud and Unity (Google Cloud blog), is a matchmaking framework that handles the player population and concurrent match generation at scale while the game developer writes the match function, director and evaluator (Open Match overview).

Staff insight: Both projects draw the same line this page does: the platform owns the generic, operationally hard part — the Ready/Allocated lifecycle, protection of live matches, buffer autoscaling, ticket storage — and the game team owns the game-specific part: the match function, the server binary. In an interview, "the platform must never kill an Allocated server" is a one-sentence answer to a whole class of Kubernetes-induced outages. See Kubernetes for the substrate.

Amazon GameLift FlexMatch — Rules That Relax Over Time#

AWS's FlexMatch documentation describes expansions that relax match rules when no acceptable match is found: its example requires all players to be within 5 skill levels, widens that to 10 after 15 seconds and to 20 after 10 more seconds, and warns against relaxing player-count requirements before automatic backfill has had time to start (AWS FlexMatch expansions).

Staff insight: "Quality vs wait" is not a philosophical argument; it is a schedule in a config file. Saying "start strict, widen on a published schedule, measure p95 wait and p95 spread" is how you turn the hardest product tradeoff in matchmaking into something an owner can tune.

CCP Games — EVE Online's Time Dilation#

EVE Online runs each solar system on a server node, and a single huge battle can overload one node. CCP's Time Dilation slows the simulation clock for that system when the node falls behind, so every action still happens in order rather than timing out; CCP reported it fully activated on January 18, 2012, and that in fights of more than 1,300 pilots module response time stayed under one second for most of the action, compared with delays of 20, 40 and even 600 seconds before (CCP — Time Dilation).

Staff insight: This is backpressure for a persistent world. When one zone can't keep up, the options are to drop work (lag, rubber-banding, desync), to cap entry (queue at the gate), or to slow everyone down fairly. CCP chose to degrade time instead of correctness — a clean example of choosing which degradation players experience. See Backpressure.

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Clients connect over WebSockets""One packet is lost. What happens to the next 50 ms of updates?"TCP head-of-line blocking; UDP with app-level reliability
"The server is authoritative""Then why does my character move the instant I press a key?"Client prediction and reconciliation
"We use lag compensation""I died after I was already behind the wall. Is that a bug?"Favor-the-shooter tradeoff; rewind cap
"We match by MMR""It's 4 a.m. and a top-50 player is queueing. What happens?"Expansion schedule; quality vs wait; who signs off
"We spin up a server per match""How long does that take, and what does the player see meanwhile?"Warm buffer; allocation latency
"We'll run it on Kubernetes""The cluster upgrades nodes tonight. What happens to live matches?"Allocated protection; drain strategy
"The game server writes results""The server crashes right after writing half the results. Now what?"Idempotent result commit keyed by match_id
"We'll add anti-cheat""What does the client know that it shouldn't?"Information withholding; trust boundary

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Two paths with opposite properties. The match path — the UDP link between clients and one game server — is in memory, single process, hard real-time, and disposable. Everything else is the control and durable path: login, tickets, match formation, allocation, results, progression. It is ordinary distributed-systems work with queues, retries and idempotency keys. The only bridge between them is the allocator handing out an address with a signed connect token, and the game server emitting one result event per match.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Authority"Server validates client moves""Inputs up, state down. One authoritative process per match decides every outcome."
Transport"WebSockets""UDP with sequence numbers and acks for the match; WebSocket for party, queue and chat."
Tick"As fast as possible""64 Hz sim, 15.6 ms frame, ≤ 2.5 ms per match per frame → 5 matches a core."
BandwidthNot mentioned"~140 kbps down, ~55 kbps up per player; interest-filtered, delta-compressed."
Lag compensation"Server checks hits""Rewind other players to what the shooter saw, capped at 200 ms. Above that, the shooter's lag is their problem."
Matchmaking"Sorted set by MMR""Pool per region × mode × party; ±100 rating, widen 50 per 10 s; p50 30 s, p95 90 s."
Fleet"Start a server per match""Warm Ready buffer, allocate in < 2 s, never evict Allocated servers, canary builds by allocation share."
Server crash"Fail over""Void the match, no rating change, requeue with priority. Results are idempotent by match_id."
Cheating"Anti-cheat software""Client is untrusted: no outcomes from the client, bounded rewind, no data the player can't see."

Numbers to Bring#

MetricValueWhy It Matters
Frame budget at 64 / 128 Hz15.6 ms / 7.8 msEverything per tick — inputs, physics, hits, snapshots for every client — fits here
VALORANT server frame~50 ms early → sub-2 ms; needed better than 3 games per core (Riot)Tick rate is priced in games per core
Source engine defaults15 ms timestep (~66 Hz), 20 snapshots/s, 100 ms interpolation (Valve)Simulation rate and send rate are separate knobs
Per-player downstream (design)~250 B × 64 Hz + headers ≈ 140 kbpsTimes 2M players ≈ 285 Gbps of peak egress
Per-player upstream (design)~80 B × 64 Hz + headers ≈ 55 kbpsInputs are small; redundancy (last 3 inputs per packet) covers loss
UDP/IPv4 header overhead28 bytes per packetAt 64 packets/s that's ~14 kbps of pure overhead per direction
Interpolation delay~2 snapshot intervals + jitter buffer (≈ 30–50 ms at 64 Hz)Everyone else is drawn this far in the past
Lag compensation rewind cap150–250 ms (design: 200 ms)Bounds "shot behind cover" and blunts high-latency abuse
Playable RTT for a shooter≤ 60 ms good, 60–100 ms acceptable, > 120 ms degradedDrives region placement and the matchmaker's latency rule
Riot ping goal35 ms for 70% of players (Riot)The scale of edge investment a competitive title makes
Game server cold start30–90 s (image pull + asset load)Why allocation must come from a warm buffer
Allocation from warm buffer< 2 s p99The gap between "match found" and "loading"
Queue targets (design)p50 ≤ 30 s, p95 ≤ 90 s for the middle 90% of ratingsThe other side of the quality dial
FlexMatch example expansion±5 → ±10 at 15 s → ±20 at 25 s (AWS)Relaxation is a schedule, not an algorithm
Agones autoscaler default sync30 s (Agones)Buffer must cover ≥ 30 s of allocations plus start time
Crashed-match target< 10 per 10,000 matches per buildThe number that gates a build's rollout

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on lobbies, friends lists, profile schemas and a WebSocket gateway, then run out of time before the questions that decide the level: "What runs every 15 ms, what does it send, who decides that a shot hit, and what happens when 300 matches form in a second and there are no servers?" Compress the social and profile surface to ~5 minutes and spend the rest on the match loop, matchmaking, allocation and the trust boundary.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Players log in, form parties, queue for a mode, get matched with players of similar skill and good latency to a server, play a 10-player match on a server that decides what happened, and afterwards get rating changes, rewards and leaderboard updates. Spectators and replays are nice-to-have."

Then the non-functional requirements, which is where the design lives:

"Five constraints drive everything. One: the match must feel instant with 30–80 ms of network latency, so the client predicts and the server compensates. Two: the server is authoritative — the client is untrusted and only sends inputs. Three: matchmaking must bound wait — I'll target p50 30 seconds, p95 90 seconds — while keeping matches fair. Four: a match must start within 2 seconds of being formed, so servers come from a warm pool. Five: match results — ratings, items — are durable and applied exactly once, even though the match itself is disposable. Scale: 2.4 million peak concurrent players, about 2 million in matches, 10 players per match — 200,000 concurrent matches across 12 regions, average match 25 minutes."

Then name the underspecified parts:

"A few things I'd confirm: genre and tick rate — a tactical shooter wants 64–128 Hz, a battle arena can live with 30. Is cross-play in scope, which splits or merges pools? Is there a ranked mode where results carry real stakes? I'll assume a competitive team shooter with a ranked queue at 64 Hz."

🎯 Staff Move: Saying "the match is disposable, the result is exactly-once" in the first three minutes tells the interviewer you've already decided where availability and durability budgets go. Everything you draw next can be judged against it.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Player: player_id, account_id, region_home, platform, ratings{mode → (mu, sigma)}, sanctions[]
  • Party: party_id, leader_id, members[] (≤ 5), version
  • Ticket: ticket_id, party_id, mode, created_at, rating, latencies{region → rtt_ms}, state (searching, proposed, assigned, cancelled, expired)
  • Match: match_id, mode, map, teams[[player_id]], region, build_id, server_address, state
  • GameServer: gs_id, fleet, build_id, node, state (Starting, Ready, Allocated, Shutdown), match_id
  • MatchResult: match_id, outcome (completed, voided), per_player{score, placement, abandoned}, server_signature, ended_at

Client-facing control API (TLS, over the session gateway):

POST  /v1/parties/{id}/queue           { mode, latencies: { "eu-west": 24, "eu-central": 31 } }
  → 202 { ticket_id, est_wait_s: 35 }
DELETE /v1/tickets/{id}                (cancel; idempotent)
WS    ticket.updated { state: "assigned", server: "203.0.113.10:7777",
                       connect_token: "<signed, ttl 30s>", match_id }

Match protocol (UDP, game server ↔ client):

C→S  Input   { client_tick, ack_snapshot, inputs[last 3 ticks] }        ~80 B, 64/s
S→C  Snapshot{ server_tick, base_snapshot, delta(entities visible to you) } ~250 B, 64/s

Internal durable API:

GS → results topic   MatchResult keyed by match_id   (exactly-once by consumer dedupe)
POST /internal/allocations  { fleet, region, build_selector, match_id }  → { gs_id, address }

The connect_token is the most important field on this page: it is the only thing that binds the control path (who was matched) to the match path (who may join this server), and it's what keeps a stranger out of a ranked match.

🎯 Staff Move: "I'm keeping two protocols on purpose. Control traffic — party, queue, chat — is low-rate, needs reliability and rides TLS over a WebSocket. Match traffic is 64 packets a second each way where a late packet is worthless, so it rides UDP with my own sequence numbers, acks and redundancy. Mixing them is how a chat burst delays a gunshot."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk one match in 90 seconds:

  1. A party of three queues for ranked. The client attaches measured RTTs to each region's ping beacons. The ticket service writes a ticket into the eu-west × ranked × party3 pool.
  2. Every second, a match function scans the pool, oldest tickets first, and proposes matches within the current rating window and latency cap. A proposal claims its tickets atomically so no ticket lands in two matches.
  3. The director sends the proposal to the allocator in the chosen region, which flips one Ready game server to Allocated and returns its address — about 50–200 ms from the warm buffer.
  4. Each client receives the address and a signed connect token (bound to match_id, player_id, address, 30 s TTL) over its WebSocket and connects over UDP directly to that server — match traffic never passes through a shared Load Balancing, because the whole point of allocation is that exactly one process owns the match.
  5. The server ticks at 64 Hz: apply inputs in tick order, simulate, resolve hits with lag compensation, and send each client a filtered delta snapshot.
  6. At match end the server emits a signed MatchResult keyed by match_id, then shuts itself down; the fleet autoscaler replaces it in the Ready buffer.
  7. The progression service consumes the result, dedupes by match_id, updates ratings and rewards in one transaction per player, and feeds the leaderboard.

🎯 Staff Move: Say out loud: "Notice nothing durable is on the tick path. The server never calls a database during a match — it loads what it needs at allocation and writes one result at the end. If the result pipeline is down, matches keep running and results queue on local disk. If a server dies, we lose one match, not anyone's progress." You've now spent ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path. What makes this hard is four budgets and one boundary: CPU per tick, bytes per player, seconds per ticket, servers per allocation spike — and the trust boundary with a client the player controls. I'd like to go deep on the tick loop and its network budget, lag compensation, matchmaking quality versus wait, and fleet allocation under spikes. Where would you like to start?"

If no preference: start with the tick loop and bandwidth. It decides cost and it's where most candidates have nothing concrete to say.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: The tick loop and the network budget (7–8 min)

"At 64 Hz a frame is 15.6 ms. In that frame the server drains every client's input queue, simulates movement and physics, resolves shots, updates game rules, and builds a snapshot for each of ten players. I budget 2.5 ms of CPU per match per frame, which lets me run five matches on a core at 80% utilization and keep the rest for the OS and jitter. A match that blows its frame budget doesn't crash — it ticks late, and every client sees rubber-banding — so gs.frame_time_p99 per match is the first metric I build."

Quantify bandwidth: "Each snapshot is delta-compressed against the last one that client acknowledged and filtered to entities that player can plausibly see: ~250 bytes typical, ~1 KB on a full-state resync. 64 per second plus 28 bytes of UDP/IP header each is ~17.8 KB/s, about 140 kbps down per player. Upstream is 64 inputs a second, each carrying the last three ticks of input for loss resilience: ~80 bytes, ~55 kbps. 2 million players is ~285 Gbps of egress at peak. That, not CPU, is the line item finance will ask about."

Who pays: "Players on bad connections pay first — at 64 Hz, 3% loss is noticeable, so the client widens its interpolation buffer and players see others slightly later. The platform pays egress. Design pays if we drop to 32 Hz snapshots to save bandwidth: it halves egress and costs hit precision in exactly the moments that matter."


Deep dive 2: Lag compensation and the trust boundary (7–8 min)

"The shooter sees enemies about 40 ms in the past from interpolation plus half their RTT. If the server checked hits against the present, everyone would have to lead their targets by their own ping. So the server keeps a ring buffer of every player's hitbox positions for the last 200 ms — about 13 ticks at 64 Hz — and when a shot arrives it rewinds the others to server_time − one-way latency − interpolation delay, tests the ray, and restores. That's favor-the-shooter: it's why you sometimes die after reaching cover. I cap the rewind at 200 ms so a player at 300 ms ping has to lead, rather than everyone else dying behind walls."

Trust: "The client sends inputs and a claimed view time; the server bounds that claim against the RTT it measures itself, so a client can't ask for an arbitrarily old rewind. Movement is simulated on the server from inputs, so speed or position claims don't exist to forge. And the server only sends an enemy's position when that player could plausibly see it, with a short look-ahead so they don't pop in — what the client never receives, it can't reveal."


Deep dive 3: Matchmaking — quality vs wait (6–7 min)

"Each pool — region × mode × party size — is a queue. At peak in the largest region ranked forms ~33 matches a second, so ~330 players a second enter and with a 30 s average wait about 10,000 tickets sit in the pool. Every second, the match function walks tickets oldest-first and builds candidate teams: rating window starts at ±100, widens by 50 every 10 seconds to a cap of ±400; latency cap starts at 60 ms and relaxes to 90 ms after 45 seconds; party sizes are balanced across teams. Team balance uses rating sums plus uncertainty, so a five-stack doesn't face five solos with the same average."

Quantify: "For the middle 90% of the rating curve, a pool of 10,000 fills ±100 in a few seconds. The top 0.1% — about 10 tickets in that pool at any moment — will never fill ±100. Their window reaches ±400 at 60 seconds, and the match will be lopsided. Product has to choose: longer waits for top players or worse matches for everyone they're placed with. I'd publish both numbers per rating band and let design own the dial."


Deep dive 4: Fleet allocation under spikes (5–6 min)

"A cold game server takes 30–90 s: schedule the pod, pull a 4 GB image if it's not cached, load map assets. Players see a timeout. So the fleet keeps a buffer of Ready servers per region: allocation rate × (server start time + autoscaler interval) × 1.5. In the largest region at 33 matches/s with a 30 s start and a 10 s autoscaler interval, that's ~2,000 Ready servers — about 400 cores of idle capacity, ~4% of the region. The allocator flips one to Allocated in milliseconds. Allocated servers are never evicted by autoscaling or upgrades; they leave when the match ends. Nodes are bin-packed so scale-down can empty whole nodes."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: the match is a disposable, in-memory, authoritative simulation with hard budgets — 15.6 ms of frame time, 140 kbps per player, a 200 ms rewind — and everything around it is an ordinary distributed system whose job is to get ten players onto a warm server in under 2 seconds and to apply the result exactly once. Matchmaking is a queue with a published quality dial. The client is untrusted and only learns what it needs to draw the screen."

The evolution closer:

"What I'd build later: replays and spectating from the snapshot stream, a backfill path for players who leave mid-match, a hybrid fleet with bare-metal baseline and cloud burst, and a shared trust-and-safety service across titles. What I'd not build: our own global fleet scheduler — Agones or a managed hosting service covers it — or state replication for live matches."

🎯 Staff Move: End on the outage you designed out and who owns it. Senior candidates end with "and we'd add anti-cheat." Staff candidates end with "and mm.wait_p95, mm.rating_spread_p95, match.alloc_latency_p99 and crashed matches per 10,000 by build are the four numbers the game platform team reviews weekly, because they tell us if players are waiting, getting bad matches, or losing matches to us."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Social features firstDesigns friends, presence and lobbies for 10 minutesOne sentence: "Party and presence ride the session gateway — same as a chat system"
Database on the hot pathStores player positions in Redis every tick"State lives in the server's memory; one result write per match"
Rating algorithm deep diveDerives Elo update formulas"Bayesian rating with mean and uncertainty; the dial is the expansion schedule"
No bandwidth numbers"UDP is efficient"Bytes per snapshot × rate × players = egress, stated in Phase 5
Failover for matchesDesigns state replication to a standby server"Void, compensate, requeue; make the result exactly-once"
Anti-cheat at minute 44"And we'd add anti-cheat"Names the trust boundary in Phase 1: inputs only from the client

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

A multiplayer game backend is the clearest example of a system with two incompatible halves that a single product hides. The match is hard real-time: a fixed budget every 15.6 ms, no network calls, no garbage-collection pauses longer than a frame, no retries — a late packet is a lost packet. The control path is the opposite: durable, retried, idempotent, eventually consistent. Candidates who apply control-path habits to the match (put state in Redis, use TCP, replicate for availability) build something that lags. Candidates who apply match habits to the control path (keep results in memory, fire-and-forget) build something that loses ranked points.

It also tests whether you can reason about an adversarial client. In most systems the client is a browser you'd prefer to be honest. Here, a meaningful fraction of players are actively trying to make the client lie — and they have the binary. The design has to make lying useless rather than detectable: if the client only sends inputs and only receives what it may see, most cheats have nothing to work with.

And it is quietly a cost question. Competitive titles run tens of thousands of cores and hundreds of gigabits of egress at peak, with peaks that move around the globe with the evening. Tick rate, snapshot size and warm-buffer size are not engineering preferences; they are budget lines.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"A player with 140 ms ping shoots a player with 20 ms ping who has just stepped behind a wall. The game server's host loses power 30 seconds later, two minutes before the end of a ranked match. What does each of the ten players experience — on screen, in their rating and in their inventory — and who decided each of those outcomes?"

A candidate who answers with the hit (the server rewinds the target ~90 ms to what the shooter saw — within the 200 ms cap — so the hit counts and the target dies behind cover; design chose favor-the-shooter and the cap), the crash (clients see the connection time out after ~5 s of no snapshots, the match is voided by the control plane when the server's heartbeat stops, no rating change, match-scoped items discarded, players requeued with priority — owned by the platform team with product's sign-off on the void policy) and the durable state (anything earned before the match — purchases, unlocks — is untouched because the server never held it) has operated a game backend. A candidate who says "we'd fail over to a replica" has designed a database.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Competitive real-time session. Team shooters, battle arenas, sports games: 2–100 players, 10–40 minute matches, everyone in one shared space where milliseconds decide outcomes. The design centers on a dedicated authoritative server per match, a fixed tick between 30 and 128 Hz chosen by genre, client prediction and interpolation, server-side lag compensation, rating-based matchmaking with a latency rule, and a warm fleet in every region where players are. Match state is disposable; results carry stakes. Correctness bar: allocation latency, crashed matches per 10,000, queue p95, rating spread p95 and RTT distribution.

Persistent world. MMOs and survival servers: hundreds to thousands of players sharing a world that outlives every session. The match server becomes a zone or cell server that owns a region of the map, hands entities across boundaries and must persist state — because a crash that rolls back an hour of loot is a support catastrophe, and a crash that duplicates loot is an economic one. The design centers on interest management (each client only hears about entities in its area of interest), checkpointed state plus transactional writes for economic events (trades, purchases, rare drops), and population control: separate shards, layered copies of busy zones, or EVE's approach of slowing time when a node falls behind. Correctness bar: no item or currency created or destroyed across a crash; zone frame time under budget at design density.

Casual, turn-based and async. Card games, board games, word games, puzzle duels: 2–8 players, moves seconds to days apart. There is no tick and no fleet. A move is an HTTP request validated against stored state with an optimistic version check, the new state is persisted, and the opponent gets a push notification or a WebSocket event. This is a plain transactional service plus Push vs Poll — the hardest problems are idempotent moves, reconnect state and push delivery, all covered by Real-Time Updates with WebSockets and ordinary database design. Correctness bar: every move applied exactly once in order; any game recoverable from storage.

🎯 Staff Move: "These share identity, parties, matchmaking and the results pipeline, and almost nothing else. The competitive session needs a disposable real-time server; the persistent world needs a durable one that partitions space; turn-based needs no game server at all. If I'm asked to build one platform for all three, I'd share the control path and let each title choose its match runtime."

2.2 When NOT to Build a Dedicated-Server Game Backend#

SituationWhat to Do InsteadWhy
Turn-based or async playHTTP + database + push notificationsNo tick, no fleet; a server per game would idle 99.9% of the time
2-player fighting or racing games where latency feel dominatesPeer-to-peer rollback netcode with a relay for NAT traversalLowest possible latency between two players; a server in the middle adds a hop; relay hides IPs
RTS with hundreds of unitsDeterministic lockstep, inputs onlySending unit state for 800 units per tick is impractical; inputs are tiny
Co-op PvE with friends, no ranked stakesListen server or relay, optional dedicated for persistenceCheating hurts only the cheater's friends; server cost isn't justified
A studio under ~20 engineers shipping its first multiplayer titleManaged game-server hosting and a managed matchmakerFleet operations across 10+ regions is a full team; buy it until scale proves otherwise
Chat, presence, friends, notificationsThe company's existing real-time platformThese are Chat and WebSocket problems, not game-server problems
Leaderboards and seasonal rankingsA ranking service fed by resultsSee Leaderboard; game servers should emit results, not maintain boards

And within the design, some things you should not build even when you own the backend:

  • Don't replicate live match state for failover. A standby per match doubles cost and adds a per-tick state stream for a session that is worthless after it ends. Void and compensate.
  • Don't put any network call on the tick path. No database, cache or HTTP call inside the simulation loop. Load at allocation; write once at the end; stream telemetry asynchronously.
  • Don't accept results, positions or damage from the client. Every outcome is computed on the server from inputs. "Client-reported kills" is a design that cheaters write for you.
  • Don't send what the player can't see. Every hidden position in a snapshot is information a modified client can display.
  • Don't build your own fleet scheduler on day one. Kubernetes plus a game-server operator, or a managed service, already implements Ready/Allocated lifecycles and protects live matches.

The Staff signal is knowing that a dedicated authoritative server is the right tool for competitive, simultaneous, latency-sensitive play with real stakes, and an expensive mistake for turn-based, two-player-latency-critical, or friends-only co-op. See Build vs Buy Framework and Drill 7.

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Genre and players per matchDecides authority model, tick and bandwidth"Competitive team shooter, 10 players, 64 Hz; I'll note where a battle royale's 100 players changes things."
Ranked stakesDecides result durability and anti-cheat investment"Ranked exists, so results are exactly-once and auditable."
Regions and latency targetDecides fleet footprint and matchmaking latency rule"12 regions; target ≤ 60 ms RTT for 90% of players."
Cross-play and input devicesSplits or merges pools; controller vs mouse fairness"Cross-play on by default, with an opt-out pool for PC-only ranked."
Platform / cloudBare metal vs cloud changes the cost model"Cloud with Kubernetes and a game-server operator; bare metal later for baseline."
Mid-match joinBackfill changes allocation and fairness"Backfill for casual only; ranked never backfills."
Persistence inside a matchDecides whether the server needs checkpoints"Nothing durable is earned mid-match; all rewards come from the result."

2.4 Precise Terminology#

TermMeaningCommon Confusion
Tick / tick rateOne fixed simulation step / steps per secondConfused with snapshot send rate, which can be lower
Send rate / update rateSnapshots per second sent to each clientAssumed equal to tick rate; Source defaults to 66 Hz ticks, 20 Hz snapshots
SnapshotThe server's view of world state at a tick, as sent to one clientAssumed identical for every client; it is filtered per client
Delta compressionEncoding a snapshot relative to one the client acknowledgedEncoding relative to the previous sent snapshot breaks under loss
Client-side predictionClient applies its own inputs immediatelyThought to be authoritative; it is overwritten by the server
ReconciliationClient rewinds to the server's state and replays unacknowledged inputsImplemented as a snap, causing jitter
Interpolation delayHow far in the past the client renders othersSet too low: stutter on loss; too high: everyone looks late
Lag compensationServer rewinds others' hitboxes to the shooter's view timeThought to "remove" latency; it moves the unfairness to the target
Peeker's advantageThe moving player sees the stationary one first, by roughly RTT + interpolationBlamed on tick rate alone
Authoritative serverThe process whose state is the truthConfused with "server that validates client state"
TicketA party's request to be matched, with attributesTreated as a player, not a party
ExpansionA scheduled relaxation of a match rule as tickets ageImplemented ad hoc in code instead of policy
AllocationAssigning a Ready server to a matchConfused with scheduling a new pod
Warm bufferReady servers kept idle to absorb allocationsSized by instinct instead of rate × (start time + scaling interval)
BackfillFilling a slot in a running matchUsed in ranked, where it changes outcomes

3. Where the Design Splits#

Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.

3.1 Fault Line 1: Authority Model — Dedicated Server vs Peer-to-Peer vs Relay#

The tension: A dedicated authoritative server gives one source of truth, equal treatment of every player and a natural trust boundary, but you pay for every server-minute in every region. Peer-to-peer costs nothing to host and minimizes latency between two players, but every peer holds the full state, the slowest peer sets the pace, and any peer can lie.

StrategyWhat WorksWhat BreaksWho Pays
Dedicated authoritative serverServer decides outcomes; filtered snapshots; consistent fairness; easy to observe~40K cores and ~285 Gbps egress at our scale; a global fleet to operateThe company (hosting and egress); platform team (operations)
Listen server (a player hosts)Zero hosting cost; easy for friendsHost advantage (0 ms ping); host leaves = match over; host can modify stateNon-host players (fairness); ranked integrity
P2P deterministic lockstepInputs only: ~1–5 kbps; ideal for RTS with hundreds of unitsBit-exact determinism across platforms; everyone waits for the slowest peer; full information on every machinePlayers on good connections (wait for bad ones); anti-cheat (every client sees everything)
P2P rollback (fighting games)Lowest feel latency for 2 players; local inputs never waitVisual rollbacks on mispredict; scales badly beyond 2–4 playersPlayers (visual corrections); limited to small sessions
Relay serverHides player IPs (stops DDoS of opponents), NAT traversal, cheap CPUNo authority: cheating unchangedSmall hosting cost; players still trust each other
Hybrid: authoritative for outcomes, client-authoritative for cosmeticsSaves server CPU on things that don't matterBoundary creep: "cosmetic" physics becomes gameplayWhoever audits the boundary

The arithmetic: 200,000 concurrent matches at 2.5 ms per 15.6 ms frame is ~32,000 cores busy simulating; at 80% target utilization ~40,000 cores. On 64-core hosts that's ~625 hosts at peak, plus buffer and headroom, ~720. P2P would make that zero — and would make every ranked result something a modified client could dispute.

The Staff default: dedicated authoritative servers for every competitive mode; relay for invite-only co-op and two-player modes where the genre favors rollback. The server decides movement, hits, damage, objectives and the result.

When to deviate:

  • Two-player fighting games: rollback over P2P with a relay; the server mediates matchmaking and results, not frames.
  • RTS with large unit counts: lockstep; reduce information exposure by validating outcomes on a server that replays the input log after the match.
  • Friends-only co-op: listen server with host migration; no ranked results accepted from it.

3.2 Fault Line 2: Tick Rate and Send Rate vs CPU and Bandwidth#

The tension: A higher tick makes inputs land sooner, hits more precise and peeker's advantage smaller. It multiplies CPU per match linearly and bandwidth per player roughly linearly. Tick rate and send rate are separate knobs: simulate at 64 Hz but send at 32 Hz, and you pay half the egress for most of the precision.

StrategyWhat WorksWhat BreaksWho Pays
30 Hz tick, 30 Hz send~33 ms frame; cheap; fine for battle arenas and large-player-count modesHits and peeks feel imprecise in tactical shootersCompetitive players (precision)
64 Hz tick, 32 Hz sendPrecise simulation, half the egressClients see others ~16 ms staler; interpolation buffer growsPlayers at the margin of a peek
64 Hz tick, 64 Hz send (Staff default for a tactical shooter)15.6 ms frame; ~140 kbps per player2× egress vs 32 Hz send; ~5 matches per corePlatform (egress and cores)
128 Hz tick, 128 Hz send7.8 ms frame; smallest peeker's advantage2× CPU and ~2× egress of 64 Hz; needs a frame-time engineering programCompany (cost); engine team (optimization)
Variable tick under loadServer survives overloadSimulation changes feel mid-match; fairness disputesPlayers in the overloaded match
Diagram: 3.2 Fault Line 2: Tick Rate and Send Rate vs CPU and Bandwidth

The arithmetic:

Downstream per player = send_rate × (payload + 28 B UDP/IP header)
  64 Hz × (250 B + 28 B) = 17.8 KB/s ≈ 142 kbps
  32 Hz × (300 B + 28 B) = 10.5 KB/s ≈  84 kbps   (larger deltas: more change between sends)
  128 Hz × (220 B + 28 B) = 31.7 KB/s ≈ 254 kbps

Fleet egress at 2M players in match:
  64 Hz: 2M × 17.8 KB/s = 35.6 GB/s ≈ 285 Gbps
  32 Hz: 2M × 10.5 KB/s = 21.0 GB/s ≈ 168 Gbps

The 28-byte header matters more than people think: at 64 packets a second it's ~14 kbps of pure overhead, which is why games batch everything for a client into one packet per send and keep packets under the path MTU (~1,200 bytes is a safe budget) to avoid fragmentation. See Latency & Protocols.

The Staff default: 64 Hz simulation and send for ranked tactical modes; 30–32 Hz for casual modes and high-player-count modes; snapshot bytes per player as a tracked budget with a regression test in CI. Riot's 128-tick decision shows the price: it required a server frame under ~2.34 ms to stay above three games per core — an engineering program, not a config change.

When to deviate:

  • Battle royale (100 players): 20–30 Hz send with aggressive interest management by distance; full-rate updates only for nearby players.
  • Esports and tournament servers: 128 Hz on dedicated hosts; cost is irrelevant next to fairness.
  • Mobile on cellular: lower send rate, larger interpolation buffer; latency variance dominates.

3.3 Fault Line 3: Lag Compensation — Favor the Shooter vs Favor the Target#

The tension: With interpolation, every client sees other players in the past. If the server checks hits in the present, shooters must lead targets by their own ping — every shot feels wrong. If the server rewinds to what the shooter saw, shots land where aimed, and targets die after reaching cover. Lag compensation does not remove unfairness; it decides who absorbs it.

StrategyWhat WorksWhat BreaksWho Pays
No compensation (check in present)No "behind the wall" deaths; simpleShooters lead by ping; high-ping players can't hit anythingShooters, especially at > 50 ms
Full rewind, uncappedEvery shot lands where aimedA 300 ms player kills people who've been behind cover for a quarter second; high ping becomes an advantageTargets; low-ping players
Rewind with cap (Staff default: 200 ms)Hits feel right for most players; bounded worst casePlayers above the cap must lead; edge cases at the cap feel inconsistentPlayers with > 200 ms latency
Client-side hit detection, server-validatedPerfect feel for the shooterValidation is a heuristic; trust boundary moves to the clientAnti-cheat team; every target
Diagram: 3.3 Fault Line 3: Lag Compensation — Favor the Shooter vs Favor the Target

The trust detail: the client tells the server which tick it was viewing when it fired. The server never trusts that number blindly; it bounds it by the RTT it has measured itself for that client and by the cap, and logs the distribution of requested rewinds per player. A client that consistently claims the maximum allowed rewind is a signal for trust and safety, not a gameplay input.

The Staff default: favor the shooter with a 200 ms cap, rewind only the hitboxes needed for the shot (not the whole world), and keep per-player hitbox history in a fixed ring buffer of ~13 ticks at 64 Hz. Product signs off on the cap: it is a fairness policy, not an implementation detail.

When to deviate:

  • Melee and close-range abilities: smaller cap (~100 ms) because "hit through a door" is more visible.
  • Projectiles with travel time: simulate the projectile on the server from the fire time; rewind only the spawn point.
  • High-latency regions: don't raise the cap; fix placement — a closer region or a better route.

3.4 Fault Line 4: Matchmaking — Match Quality vs Wait Time#

The tension: A match is good when ratings are close, teams balance, latency is low and parties face parties. Every constraint shrinks the set of acceptable partners and increases wait. Every split of the pool — region, mode, platform, party size, input device — shrinks it further. Wait and quality trade against each other per ticket, and the curve is worst exactly where players are most invested: at the top of the ladder and at off-peak hours.

StrategyWhat WorksWhat BreaksWho Pays
Strict windows, no relaxationFair matchesUnbounded waits for the top, bottom and off-peakPlayers at the tails; retention
First-in-first-out, no skillShortest waitStomps; new players quitNew and low-skill players
Strict-then-relax on a schedule (Staff default)Bounded wait; quality degrades predictably; tunableTail players get worse matches; the players matched with them pay tooTail players and their opponents; design owns the schedule
Batch optimization (global assignment every N seconds)Better total match quality per batchAdds N seconds of latency; harder to explainEvery player waits for the batch
More pools (per platform, per input, per party size)Fairer within each poolEach pool is thinner; waits rise everywhereEveryone, through fragmentation
Diagram: 3.4 Fault Line 4: Matchmaking — Match Quality vs Wait Time

The arithmetic: in the largest region at peak, ranked forms ~33 matches/s, so ~330 players/s enter the pool. By Little's law, a 30 s average wait means ~10,000 tickets in the pool. With ratings roughly normal and σ ≈ 300, a ±100 window around the median covers ~26% of the pool — 2,600 candidates, filled in a second. At the 99.9th percentile the density is ~400× lower: ~10 tickets in the whole pool within ±100, most of them already in matches. That player's ticket will relax to ±400 before it fills, and the match will be lopsided by design. At 4 a.m. the arrival rate drops 10×, the pool shrinks 10×, and everyone moves one step down this curve.

The Staff default: a pool per region × mode × party bracket; tickets start strict and relax on a published schedule (rating, then latency, then party mix); match proposals claim tickets atomically (one owner per ticket); team balance by rating sum with uncertainty; two published metrics per mode and rating band — mm.wait_p95 and mm.rating_spread_p95 — and a design owner who signs the schedule. The same shape as Dating & Proximity Matching, with rating distance instead of geographic distance.

When to deviate:

  • Low population (new region, niche mode): merge pools — cross-region with an RTT cap, or cross-platform — before relaxing rating further.
  • Top of the ladder: accept long waits by policy (top players often prefer fair over fast) and show the estimate honestly.
  • Launch week: relax faster; nobody has a stable rating, so tight windows buy little.

3.5 Fault Line 5: Fleet Allocation — Warm Buffer vs Cold Start, Packed vs Spread#

The tension: A match must start within seconds of forming, but a cold game server takes 30–90 s to schedule, pull and load. Keeping Ready servers idle costs money; not keeping them costs players. Separately, packing servers densely onto few nodes lets you scale nodes down cheaply, while spreading them limits how many matches one node failure takes out.

StrategyWhat WorksWhat BreaksWho Pays
Cold start per matchNo idle cost30–90 s from match found to playable; timeouts; requeue stormsPlayers; matchmaking (re-forms matches)
Fixed warm bufferInstant allocation in steady stateSpikes drain it; off-peak it's oversizedPlayers at spikes; finance off-peak
Buffer sized to rate × (start + sync) × 1.5, autoscaled (Staff default)Covers spikes up to ~1.5× expected rateForecast misses (patch day) still drain itPlatform team owns forecast and buffer
Packed scheduling (cloud)Empty nodes can be removed; fewer nodesOne node failure ends more matchesPlayers on the failed node
Distributed scheduling (fixed bare metal)Even load; smaller per-node blast radiusNodes rarely empty; scale-down impossibleFinance (in the cloud)
Diagram: 3.5 Fault Line 5: Fleet Allocation — Warm Buffer vs Cold Start, Packed vs Spread

The arithmetic: in the largest region at peak, 33 allocations/s; a server takes ~30 s from schedule to Ready when its image is pre-pulled. The buffer must cover the start time plus one autoscaler interval (Agones evaluates its buffer every 30 s by default): 33/s × (30 s + 30 s) × 1.5 ≈ 3,000 Ready servers at the default interval, or 33 × (30 + 10) × 1.5 ≈ 2,000 with a 10 s sync interval. That's 400–600 cores idle in a region running ~10,000 busy cores: 4–6%. The cost of the buffer is visible on one line; the cost of not having it shows up as a requeue storm that also burns matchmaking CPU, player patience and support time.

The Staff default: a per-region fleet per build, a Buffer autoscaling policy sized as above with a 10 s sync interval, images and map assets pre-pulled to every node with a node-level cache, Packed scheduling in the cloud, and an absolute rule: nothing evicts an Allocated server — not scale-down, not fleet updates, not node upgrades. Node upgrades cordon nodes and wait for matches to end (max 45 minutes), then drain. The substrate details live in Kubernetes.

When to deviate:

  • Bare metal baseline: Distributed scheduling on owned hosts sized to the daily trough; cloud burst fleets Packed.
  • Predictable spikes (patch day, season start): scheduled buffer increases hours ahead, not reactive autoscaling.
  • Very long sessions (persistent worlds): no warm buffer for zones; capacity is planned, and players queue at the gate.

4. When It Breaks#

4.1 Patch-Day Allocation Exhaustion and the Requeue Storm#

t=0:       Season patch live at 18:00 in eu-west. Normal peak: 33 matches/s.
           Login wave: 450K players in 8 minutes. Matchmaker forms 110 matches/s.
t=+40s:    Ready buffer (2,000) exhausted. Allocator returns "no capacity".
           Fleet autoscaler adds replicas; node autoscaler adds nodes (3-4 min).
           New nodes don't have the new 6 GB image: pull takes 70 s.
t=+2min:   Matchmaker keeps forming matches; proposals fail allocation; tickets
           return to the pool. Clients see "match found" then "match cancelled".
t=+3min:   Clients auto-requeue immediately. Ticket churn 5x. Match function CPU
           at 100%; pool scan time 4 s; quality checks start timing out.
t=+12min:  Capacity arrives. Queue of 180K. Players report "queue broken" and quit.

Why it was bad: the forecast was ignored, allocation failure wasn't fed back to the matchmaker (it kept proposing matches it couldn't host), new nodes paid a cold image pull, and clients retried without backoff — the shape of every Backpressure failure.

Detection: fleet.ready_buffer by region (alert at < 30% of target), match.alloc_failures_per_s, mm.proposals_rejected_ratio, mm.ticket_churn_ratio, node.image_pull_seconds p99.

Prevention: scheduled pre-scale to 3× buffer from 2 hours before patch, with images pre-pulled by a node daemon; matchmaker rate-limits proposals to the allocator's reported capacity and holds tickets in the pool instead of cancelling them; clients show "finding a server" rather than cancelling; login admission gate meters arrivals and shows a position; requeue with jittered backoff.

Owner: game platform team (fleet, allocator, matchmaker); live-ops (patch timing and forecast).

4.2 A Bad Build Crashes Matches at Minute 12#

t=0:       New server build rolled to 100% of the fleet's Ready servers at once.
t=+12min:  A rare interaction on one map dereferences a null after a round reset.
           Every match on that map that reaches round 7 crashes. 20% of matches.
t=+15min:  12,000 matches crashed. Ranked players lose points for "abandoning"
           because the client reported disconnect before the server-crash signal.
t=+30min:  Rollback: fleet switches back to old build for new allocations.
           Running matches on the bad build keep crashing for 25 more minutes.

The Staff design: builds roll out by allocation share, not by replacing servers: the allocator selects the new build for 5% of new matches, then 25%, 50%, 100%, gated on gs.crashes_per_10k and gs.frame_time_p99 by build and map vs the old build. Because matches are disposable and 25 minutes long, a bad build can never be rolled back for matches already running — only stopped from receiving new ones — so the canary share is the blast radius. Crash classification is server-side: a match voided by server crash never costs rating, regardless of what clients reported. The deploy mechanics belong to Deployment System.

Detection: gs.crashes_per_10k by build × map, match.voided_ratio, crash-dump ingestion rate.

Owner: the game team owns the build and the crash; the platform owns the rollout-by-allocation mechanism and the automatic halt.

4.3 Node Upgrade Evicts Live Matches#

A managed Kubernetes cluster runs its scheduled node-pool upgrade. The upgrade drains nodes with a 1-hour timeout but the game-server pods were deployed without the operator's protections — plain Deployments behind a custom allocator — so drain evicts them. 9,000 matches end mid-round across a region in 40 minutes. The fix is structural: game servers run as operator-managed GameServer resources whose Allocated state is respected by scale-down and fleet updates; the platform adds disruption budgets so eviction of an Allocated server is refused; upgrades cordon nodes and wait for natural match end (≤ 45 min) before draining; and the platform team, not the cloud default, owns the upgrade window.

Detection: gs.terminations_while_allocated (should be zero outside crashes), correlated with node events.

Owner: platform team.

4.4 Regional Network Degradation — Rubber-Banding Without an Outage#

t=0:       A transit provider into eu-west starts dropping 4% of packets for players
           on two large ISPs. Servers are healthy. Frame time normal.
t=+5min:   Affected players' clients widen interpolation buffers; prediction
           corrections spike; players rubber-band. Win rate of affected players -6%.
t=+30min:  Social media: "servers are lagging". Status page: all green.
t=+2h:     Support escalates. Network team finds the route issue.

Why it's silent: every server metric is green. The signal is per-client network telemetry, aggregated by ISP and region: client.packet_loss_pct, client.rtt_p95, gs.prediction_corrections_per_min, gs.input_late_ratio. The Staff design: clients report connection quality every 10 s on the control channel; the platform aggregates by (region, ASN) and alerts on deltas versus the trailing week; the matchmaker can steer affected ASNs to a neighboring region within the RTT cap while the network team works the route. A ranked policy decides whether matches with a player above 5% loss for more than 2 minutes are eligible for loss forgiveness.

Owner: platform networking (detection and steering); product (loss-forgiveness policy).

4.5 Results Pipeline Down — Matches Fine, Ranks Frozen#

The results topic's brokers lose quorum in one region for 25 minutes. Servers keep running matches. If servers write results synchronously with retries, they hang at match end, don't return to the fleet, and the Ready buffer drains — turning a pipeline outage into a hosting outage. The Staff design: the game server writes the signed result to local disk and to the topic asynchronously, then exits; a node-level forwarder ships pending results when the topic recovers. The progression service dedupes by match_id, so duplicate deliveries are harmless. Players see "results processing" for a few minutes; nobody loses points. The dedupe design is standard Idempotency.

Detection: results.pending_on_nodes, progression.lag_seconds, match.end_to_rating_p99.

Owner: platform (forwarder and topic); progression team (dedupe and apply).

4.6 Cheat Wave After a Client Update#

A new client build changes a memory layout that an external cheat vendor adapts to within days. Reports spike in high-rated ranked. The Staff design avoids the most damaging cheats structurally — movement and damage are simulated server-side, enemy positions outside a player's plausible view are never sent, rewinds are bounded — so the remaining cheats are aim assistance, which shows up as statistical anomalies: reaction time distributions, accuracy versus rating band, requested-rewind distributions. Detection runs as an offline pipeline over match telemetry, not on the tick path. Enforcement is delayed and batched (a ban wave) so vendors can't easily learn which change triggered detection. Results from flagged matches are reviewed before they hit Leaderboards.

Owner: trust and safety (detection, enforcement, appeals); platform (telemetry, server-side guarantees).

4.7 Silent Matchmaking Degradation#

A config change tightens the rating-uncertainty term, so new accounts get wide uncertainty and are placed as if they were average. Nothing breaks. Over three weeks, mm.rating_spread_p95 in the low-rating band rises from 180 to 420, first-week retention for new players drops 4 points, and nobody connects the two because no alert existed. The Staff design: match quality metrics are first-class SLOs per band with weekly review; matchmaking config changes go through staged rollout by pool with A/B comparison of quality and wait — matchmaking parameters are runtime config and get the same discipline as Feature Flags.

Owner: matchmaking team (metric and rollout); game design (policy).

4.8 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Allocation exhaustionfleet.ready_buffer < 30% target; match.alloc_failures_per_sA region's queue; requeue stormScheduled pre-scale; matchmaker throttles to capacity; hold ticketsPlatform + live-ops
Bad server buildgs.crashes_per_10k by build × mapCanary share of new matchesRollout by allocation share; auto-halt; void without rating lossGame team + platform
Allocated server evictedgs.terminations_while_allocated > 0Every match on drained nodesOperator-managed lifecycle; disruption budgets; wait-for-end upgradesPlatform
Frame-time overrungs.frame_time_p99 > 12 ms at 64 HzMatches on hot nodesLower density per node; profile by map; move noisy neighborsGame team (engine)
Network route degradationclient.packet_loss_pct by ASN × regionPlayers on affected ISPsSteer within RTT cap; loss-forgiveness policyPlatform networking
Results pipeline outageresults.pending_on_nodes; progression.lag_secondsRating updates delayedLocal spool + forwarder; dedupe by match_idPlatform + progression
Cheat waveReport rate by band; anomaly scoresRanked integrity in top bandsServer authority; info withholding; delayed ban wavesTrust and safety
Matchmaking quality driftmm.rating_spread_p95 by bandNew-player retentionQuality SLOs; staged config rolloutMatchmaking + design
Login stormlogin.admission_queue_depth; auth error rateRegion-wide loginAdmission gate with visible position; jittered retryIdentity + platform

🎯 Staff Insight: The failures players remember most rarely page anyone. A lopsided match returns 200s. A 4% packet-loss route shows green servers. A player who lost points to a server crash files a ticket three days later. mm.rating_spread_p95, client.packet_loss_pct by ASN and match.voided_ratio are the three metrics that turn those silent failures into numbers someone owns — build them before launch, not after.


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists features: lobbies, matchmaking, game server, leaderboardNames competitive / persistent / turn-based intents; commits; states "match disposable, result exactly-once"Asks how many titles share the platform and what each costs per player-hour
Netcode"Server validates; WebSockets"Inputs up, filtered delta snapshots down over UDP; prediction, interpolation, capped rewind; tick and send rate as priced knobsChooses the netcode model per genre as a platform policy with a cost envelope per model
MatchmakingSorted set by ratingPools by region × mode × party; strict-then-relax schedule; wait and spread published per band; atomic ticket claimsPopulation budget across modes; merges pools before quality collapses; owns the fairness policy with design
Fleet"Autoscale servers"Warm buffer sized from rate and start time; Allocated never evicted; canary builds by allocation share; pre-pulled imagesHybrid bare-metal baseline + cloud burst; region footprint chosen from latency maps and cost
Trust"Anti-cheat software"Client sends inputs only; bounded rewind; information withholding; results only from serversCross-title trust-and-safety platform: shared detection, ban identity, appeals, client policy
Operations"Add monitoring"gs.frame_time_p99, match.alloc_latency_p99, mm.wait_p95, mm.rating_spread_p95, crashes per 10K by buildLaunch and patch-day readiness as a standard; game days for region loss; cost per player-hour as a KPI

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Separates match path from control path"The match is disposable and in memory; the result is durable and exactly-once."
Prices the tick"64 Hz is 15.6 ms; at 2.5 ms per match per frame that's five matches a core."
Knows egress is the bill"140 kbps a player times 2 million is ~285 Gbps. Snapshot bytes are a cost budget."
Names who absorbs lag"Lag compensation moves unfairness to the target; the 200 ms cap is a product policy."
Treats matchmaking as a queue"10,000 tickets in the pool by Little's law; the top 0.1% will never fill ±100."
Protects live matches"Nothing evicts an Allocated server — not scale-down, not upgrades."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Game state in Redis, updated every tickPuts a network round trip inside a 15.6 ms frame for every player
Client reports position, hits or killsHands the outcome to the party most motivated to lie
TCP for real-time match traffic, unexaminedHead-of-line blocking turns one lost packet into a stall
State replication for match failover as the availability planSpends heavily on the wrong failure; silent on allocation and bad builds
No numbers for bandwidth or CPUCan't reason about cost or density
"Players are matched instantly by skill"No model of wait, pool size or relaxation

5.4 Common False Positives#

  • Netcode vocabulary ≠ backend design. Fluent talk of prediction and rollback with no fleet, allocation or results path is half the system.
  • Rating-algorithm math ≠ matchmaking. Deriving Elo or TrueSkill updates is not the same as designing pools, expansion schedules and fragmentation policy.
  • "We'll use Kubernetes" ≠ fleet design. The question is what happens to an Allocated server during a node upgrade.
  • Kernel-level anti-cheat knowledge ≠ trust design. The structural defense is what the server refuses to accept and refuses to send.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minPick intent; "match disposable, result exactly-once"; tick; scale
Entities & API3–5 minTicket, match, game server lifecycle, result keyed by match_id, connect token
Architecture5–10 min≤ 8 boxes: gateway, tickets, match functions, allocator, game server, results
Tick loop + bandwidth10–17 minFrame budget, density, bytes per player, egress
Lag compensation + trust17–24 minRewind and cap; inputs only; info withholding
Matchmaking24–31 minPools, schedule, Little's law, tails
Fleet31–37 minBuffer sizing, Allocated protection, build canaries
Pivot (interviewer's choice)37–42 minBattle royale, MMO, patch day, multi-region, cost
Wrap42–45 minBudgets, boundary, metrics and owners

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Make it a 100-player battle royale"Interest management and bandwidthDistance-tiered update rates; spatial grid for relevancy; ~20–30 Hz send
"Make it an MMO"Persistent state, spatial partitioningZone servers, handoff, checkpoint + transactional economy writes, population control
"It's patch day"Spikes and backpressureForecast pre-scale, admission gate, matchmaker throttled by capacity
"Players say hit-reg is broken"Lag compensation as policyTelemetry on rewind amounts; cap; region placement; not tick alone
"Cut hosting cost 30%"Cost leversSend rate per mode, density via frame time, bare-metal baseline, egress contracts
"A cheat lets people see through walls"Trust boundaryInformation withholding; what the server sends, not what the client hides

6.3 What to Deliberately Skip#

  • Friends, presence and chat. "Same as a chat and presence system; they ride the session gateway."
  • Store, purchases and inventory UI. "A transactional service; the game server never writes inventory directly."
  • Rating formula derivation. "Bayesian rating with mean and uncertainty; the interesting part is the expansion schedule."
  • Client rendering and animation. "Out of scope beyond interpolation delay."
  • Login internals. "Standard token-based auth; see the identity design."

6.4 Follow-Up Questions to Expect#

  1. "Walk me through one frame on the server. What's in the 15.6 ms?"
  2. "How many bytes per second does each player receive, and how do you reduce it?"
  3. "I shot first and died. Explain what happened on each machine."
  4. "A top-10 player has been queueing for four minutes. What do you do?"
  5. "300 matches form in one second and you have 100 Ready servers. Then what?"
  6. "The game server crashes 20 minutes into a ranked match. What happens to ratings?"
  7. "What can a modified client learn or do in your design, and what can't it?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design the backend for an online multiplayer game."

Staff Answer

"Before drawing — what kind of game? A competitive real-time shooter, a persistent MMO world, and a turn-based card game are three different systems. I'll design a competitive 5v5 shooter with ranked play, and point out what changes for an MMO.

The constraints I'll commit to: one dedicated authoritative server per match — clients send inputs, the server decides outcomes; 64 Hz tick with a 15.6 ms frame and no network calls on the tick path; matchmaking that bounds wait at p95 90 s with a published relaxation schedule; matches start within 2 s from a warm fleet; and the match is disposable while the result is exactly-once. Scale: 2.4M peak concurrent, 200K matches in 12 regions. I'll go: entities → control path and match path → the tick loop and bandwidth → lag compensation and trust → matchmaking → fleet → failure modes."

Why this is L6:

  • Distinguishes intents and commits to one
  • States the authority model and the durability split before boxes
  • Puts numbers on the tick, the queue and allocation up front

What L7 adds:

  • Asks how many titles will share identity, matchmaking, fleet and anti-cheat
  • Frames cost per player-hour as a target the studio signs
  • Asks which launch-day failures the org has already lived through
❌ Common L5 Trap

"Clients connect to a WebSocket gateway, a game service holds match state in Redis, every action is validated and broadcast to other players, a matchmaking service pairs players by MMR from a sorted set, and results are written to Postgres."

Why this misses: Every box works for a turn-based game. For a shooter it puts Redis and TCP inside a 15.6 ms loop, has no answer for prediction or lag compensation, and has no fleet: where does the match run, and how fast can one start?


Drill 2: Core Mechanic — One Server Frame#

Prompt: "Walk me through exactly what the server does in one tick."

Staff Answer

"At tick N, 15.6 ms budget, ~2.5 ms of it ours. One: drain each client's input queue — inputs carry client tick numbers and the last three ticks of input, so a lost packet costs nothing if the next arrives; inputs too far in the future or past are dropped and counted. Two: apply inputs in order and simulate movement and physics with a fixed timestep. Three: resolve shots — for each, rewind the other players' hitboxes to the shooter's bounded view time from the ring buffer, ray-test, restore. Four: game rules — damage, objectives, round state. Five: record this tick's hitboxes in the ring buffer. Six: for each client, compute the visible set, diff against the last snapshot that client acknowledged, encode one packet under ~1,200 bytes and send. Telemetry goes to a lock-free buffer flushed by another thread. No allocation-heavy work and no I/O on this thread."

Why this is L6:

  • Orders the frame and puts lag compensation in it
  • Deltas against the acknowledged snapshot, not the last sent
  • Keeps I/O off the simulation thread

What L7 adds:

  • Makes frame time per match a release gate across every title's engine team
  • Funds a shared netcode library so studios don't each rebuild delta encoding
  • Treats density (matches per core) as the cost KPI engine teams own

Drill 3: Make It Concrete — Capacity and Cost#

Prompt: "Size the fleet and bandwidth for 2 million concurrent players in matches."

Staff Answer

"Matches: 2M / 10 = 200K concurrent.

CPU: 2.5 ms per match per 15.6 ms frame = 16% of a core; 5 matches a core at ~80%. 200K / 5 = 40K cores. On 64-core hosts ~625 at peak; with warm buffer and N+1 per region, ~720.

Bandwidth: 142 kbps down and ~55 kbps up per player. Down: 2M × 17.8 KB/s ≈ 35.6 GB/s ≈ 285 Gbps at peak. If average concurrency is ~45% of peak, that's ~16 GB/s average, ~42 PB a month.

Money, roughly: at negotiated egress of $0.01–0.02/GB, $0.4–0.8M a month; at public list prices several times that. Compute at ~$2K per host-month in the cloud is ~$1.4M a month at peak footprint, less with autoscaling to the trough. So egress is the same order as compute — and send rate, snapshot size and interest filtering are cost levers.

Matchmaking: 2M / 25 min ≈ 1,330 players/s entering queues at peak, ~133 matches/s globally, ~33 in the largest region." The forecasting and headroom method is the one in Autoscaling & Capacity.

Why this is L6:

  • Derives density from frame budget, not a guess
  • Sizes egress and notices it rivals compute
  • Ties queue rate to match length

What L7 adds:

  • Proposes bare-metal for the daily trough (~40% of peak) and cloud for the peak
  • Negotiates egress as a contract term, and places regions where transit is cheap and latency good
  • Sets a cost-per-player-hour target per title

Drill 4: The Dependency Goes Down — Results Pipeline#

Prompt: "The service that records match results is down for 30 minutes. What happens?"

Staff Answer

"Nothing visible in matches — the game server never calls it during play. At match end, the server writes a signed MatchResult to local disk and attempts to publish asynchronously; it then exits so its slot returns to the fleet. A node-level forwarder retries pending results with backoff. When the pipeline recovers, results flow; progression dedupes by match_id and applies each in one transaction per player. Players see 'results processing' and get their rating a few minutes late. The one thing I must not do is block the server's exit on the write — that would drain the Ready buffer and turn a results outage into a hosting outage."

Why this is L6:

  • Decouples the hot path from the durable path
  • Idempotency by match_id
  • Spots the second-order failure: servers that never exit

What L7 adds:

  • Defines the result schema and signature as a cross-title contract
  • Sets an end-to-end SLO from match end to rating applied
  • Makes the forwarder part of the platform node image for every title

Drill 5: Hot Key — The Event Every Player Joins#

Prompt: "A streamer event opens a limited-time mode at 18:00. 800K players queue in the first minute."

Staff Answer

"That's a single pool becoming a hot partition and a fleet spike at once. Pool: shard it by rating band so match functions scan bands in parallel — each band is a separate queue with its own worker, and cross-band matching only happens after expansion. Fleet: this is scheduled, so pre-scale the event mode's fleet to the forecast — say 13K matches in the first ten minutes — two hours ahead with images pre-pulled. Admission: meter queue entry at the rate the fleet can host, show a position and estimate. Matchmaker proposals are throttled to the allocator's reported capacity, so tickets wait in the pool instead of bouncing. Relax rules faster for this mode; ratings are meaningless for a novelty mode."

Why this is L6:

  • Treats the pool as a hot key and shards it by band
  • Pre-scales from forecast, not reactive autoscaling
  • Feeds allocator capacity back to the matchmaker

What L7 adds:

  • Requires live-ops to file a capacity forecast for every event, reviewed by the platform
  • Prices event capacity against expected engagement
  • See Hot Keys for the general pattern

Drill 6: Multi-Tenant — Several Titles, One Platform#

Prompt: "Our studio is launching a second game. Should it share the backend?"

Staff Answer

"Share the control path, not the match runtime. Identity, parties, presence, the ticket service, the allocator, the fleet operator, the results pipeline and telemetry are generic. Match functions, server binaries, tick rates and netcode models are per title. Each title gets its own fleets and pools — a crash loop in title B's build can't consume title A's warm buffer — with per-title quotas on the shared node pools. Results share a schema with a title-specific payload. The risk is noisy neighbors on the nodes: frame time is sensitive to CPU contention, so I'd isolate titles by node pool and pin game-server pods to dedicated cores."

Why this is L6:

  • Draws the platform-vs-title boundary in the right place
  • Isolates fleets and buffers per title
  • Notices CPU contention as a frame-time risk

What L7 adds:

  • Writes the platform contract (APIs, SLOs, quotas) and the onboarding path for title three
  • Decides when a title's needs justify an exception
  • Prices the platform as a chargeback per player-hour

Drill 7: Build vs Buy#

Prompt: "Should we build our own fleet orchestration and matchmaker, or buy?"

Staff Answer

"Fleet orchestration: don't build the scheduler. Agones on Kubernetes gives Ready/Allocated lifecycles, buffer autoscaling and protection for live matches; managed hosting services give the same with less control and higher unit price. I'd start managed or Agones-on-managed-Kubernetes; the decision to own more comes when hosting spend is large enough that bare metal saves real money. Matchmaking: buy or adopt the framework — ticket storage, pools, concurrency — and own the match function, because the match function is the game's fairness policy. Anti-cheat client software: buy. Server-side trust guarantees: always own."

Why this is L6:

  • Separates commodity (scheduling, ticket storage) from differentiating (match logic, trust)
  • Names concrete options and when to switch
  • Keeps the fairness policy in-house

What L7 adds:

  • Keeps an exit plan: match functions and server SDK calls behind internal interfaces
  • Prices managed hosting vs owned fleet at three scales (see Cost Model)
  • See Buy or Build: The Total-Cost Test

Drill 8: Changing Policy Without an Outage — A New Lag Compensation Cap#

Prompt: "Design wants to lower the rewind cap from 200 ms to 120 ms. How do you ship that?"

Staff Answer

"First, measure who it affects: from match telemetry, the distribution of requested rewinds — if 6% of shots in South America request more than 120 ms, those players will suddenly have to lead targets. Then ship it as server config, not a build: per-mode, per-region. Run it in shadow — compute both results, log disagreements — for a week. Then canary in casual modes in one region, comparing hit rate by ping band and complaint rate. Then ranked. Communicate it in patch notes, because it changes feel. And pair it with placement: if high-ping regions are the issue, a closer region is the real fix."

Why this is L6:

  • Quantifies the affected population before shipping
  • Shadow → canary → enforce for a gameplay policy
  • Treats it as runtime config with staged rollout

What L7 adds:

  • Makes fairness parameters a governed set with design, platform and community sign-off
  • Links the decision to region investment: cap changes are cheaper than regions, but not free
  • Publishes the policy so players can see the rule

Drill 9: Multi-Region — Where Does the Match Run?#

Prompt: "Players in one party are in Madrid, London and Stockholm. Where do you host the match?"

Staff Answer

"Each client measures RTT to every candidate region's ping beacon at login and refreshes it every few minutes. A ticket carries those RTTs. For a proposed match, pick the region that minimizes the maximum RTT across all ten players, subject to the latency cap; break ties by available Ready servers. If no region satisfies the cap for everyone, the match isn't proposed until expansion relaxes the cap. The party's worst member decides — that's a product choice; the alternative is that the party's best-connected player decides and the others suffer. The control plane — tickets, pools — can be regional; the allocator is per region; results flow to a global progression store. See Multi-Region."

Why this is L6:

  • Uses measured RTTs, not geo-IP
  • Min-max placement with a cap and a named tradeoff
  • Keeps allocation regional and progression global

What L7 adds:

  • Uses aggregated RTT maps to decide where to open the next region
  • Considers routing investments (peering, private backbone) as an alternative to new regions
  • Sets latency targets by population percentile, not by region

Drill 10: Cost — Cut Hosting 30%#

Prompt: "Finance wants game hosting down 30% next year without hurting players. Where do you look?"

Staff Answer

"Four levers, in order of player impact. One: density — profile frame time by map and mode; moving from 5 to 6 matches a core is 17% of compute with no player-visible change. Two: egress — interest management and quantization reduce snapshot bytes; a 20% cut in bytes per snapshot is 20% of egress. Send rate per mode: casual at 32 Hz halves its egress. Three: placement — bare metal or committed-use for the daily trough, cloud only for peak; troughs are ~40% of peak, and baseline capacity is the cheapest dollar. Four: buffer — tune it per region per hour from allocation forecasts rather than a static peak buffer. I wouldn't touch ranked tick rate; that's the product."

Why this is L6:

  • Ranks levers by player impact
  • Quantifies each
  • Protects the competitive experience explicitly

What L7 adds:

  • Negotiates egress and capacity commitments as a platform contract across titles
  • Builds cost-per-player-hour into each title's business case
  • Sets a policy for which modes may lower tick and send rates

8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Season Launch Queue Collapse#

Context: A new season launches at 18:00 in the two largest regions. Within 5 minutes, queue times climb past 6 minutes, 30% of "match found" notifications end in "match cancelled", and social channels fill with complaints. The live-ops lead escalates to you.

Questions to Surface First:

  • Is the matchmaker forming matches the fleet can't host, or is the fleet idle while the matchmaker stalls?
  • What's fleet.ready_buffer per region and how fast are new servers becoming Ready?
  • Are clients requeueing on cancel, and with what backoff?

Typical L5 Approach: Raises the fleet autoscaler maximum and adds nodes. Nodes arrive in 4 minutes, pull images for 70 seconds, and the cancelled-match loop continues until they do.

Staff Approach: Sees match.alloc_failures_per_s high and Ready buffer at zero. Immediately throttles the matchmaker's proposal rate to the allocator's reported capacity so tickets wait instead of cancelling, enables the login admission gate with visible positions, and switches clients to "finding a server" rather than cancel-and-requeue. Then scales fleets with a pre-pulled image node pool.

Principal Approach: Makes forecast-driven pre-scaling a launch-readiness gate for every title: a capacity plan signed by live-ops and platform 72 hours before any season, a pre-pull job verified per node, and a launch game day that replays last season's arrival curve against staging fleets.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Throttle proposals to allocator capacity; enable admission gate; stop cancel-requeue loop
TriageReady buffer drained at +40 s; new nodes lacked the image; clients retried instantly
Quick fixPre-pull DaemonSet on all nodes; raise buffer to 3× for 6 hours
GuardrailsAlert on Ready buffer < 30%; matchmaker backpressure permanent; client backoff with jitter
Post-mortemWhy was the forecast not applied? Who owns season capacity?

Metrics to Watch: fleet.ready_buffer, match.alloc_failures_per_s, mm.ticket_churn_ratio, node.image_pull_seconds, login.admission_queue_depth

Organizational Follow-up: season launch checklist owned by live-ops with platform sign-off; capacity forecast as a required artifact.

Ownership Question: "Who owns season-launch capacity — the game team or the platform?" Staff answer: Live-ops owns the forecast of players; the platform owns turning it into Ready servers and the backpressure behavior when the forecast is wrong. Both sign the plan.

Key Takeaway: "A matchmaker that ignores fleet capacity converts a capacity shortfall into a retry storm."

What clears the Staff bar:

  • Throttles demand to capacity before adding capacity
  • Finds the cold image pull as the hidden multiplier
  • Turns a one-off fix into backpressure that's always on

Deep Dive 2: Silent Failure — Hit Registration Complaints in One Region#

Context: For three weeks, players in one region post clips of "obvious hits" not counting. Server metrics are green. Community managers ask engineering to "fix hit-reg".

Questions to Surface First:

  • What's the distribution of requested rewind per shot in that region versus others?
  • Is packet loss or jitter elevated for specific ISPs?
  • Did anything change in build, map or server config three weeks ago?

Typical L5 Approach: Proposes raising tick rate to 128 Hz in that region.

Staff Approach: Pulls rewind telemetry: 9% of shots from two ISPs request more than the 200 ms cap since a routing change three weeks ago sent their traffic through a distant exchange. Shots beyond the cap are checked against a later world state and miss. Works with networking to fix the route and temporarily steers those ASNs' matches to a neighboring region within the cap. Adds gs.rewind_clamped_ratio by (region, ASN) as an alerting metric.

Principal Approach: Treats latency to the player population as a platform SLO, with per-ASN monitoring and a peering budget; uses the RTT maps to decide where to add capacity or interconnects, and publishes the lag compensation policy so community teams can explain it.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateConfirm clamped-rewind ratio by ASN; open route ticket with transit provider
TriageRoute change raised RTT by ~80 ms for two ISPs; their shots exceeded the cap
Quick fixSteer affected ASNs to a neighboring region where RTT is within cap
GuardrailsAlert on gs.rewind_clamped_ratio > 2% per ASN; weekly RTT map review
Post-mortemWhy did no server-side metric reflect player-perceived quality?

Ownership Question: "Who owns 'hit-reg feels bad'?" Staff answer: Platform networking owns latency to players and the clamped-rewind signal; the game team owns the cap policy; community gets the dashboard so complaints map to data.

Key Takeaway: "Most hit-reg complaints are latency complaints. Measure the rewind, not the tick."

What clears the Staff bar:

  • Uses lag-compensation telemetry instead of guessing
  • Fixes placement, not tick rate
  • Creates a per-ASN signal so it can't stay silent again

Deep Dive 3: Large-Customer Onboarding — An Esports League Wants Dedicated Servers#

Context: A professional league will run weekly matches on your game. It wants 128 Hz servers, guaranteed hardware, spectator feeds with a 2-minute delay, and match replays for disputes.

Questions to Surface First:

  • Do league servers share fleets, builds and matchmaking with public play?
  • Who can start, pause and restart a match?
  • What evidence do disputes need?

Typical L5 Approach: Allocates normal game servers with a tick-rate override for league matches.

Staff Approach: Creates a separate tournament fleet on dedicated nodes (no co-tenancy, CPU pinned), at 128 Hz with its own build channel frozen during the season. Matches are created by an admin API, not matchmaking, with pause and technical-restart controls. The server records the full input stream and snapshots for replays; a delayed spectator relay fans out the feed. Results go to a separate league store, not public ranked.

Principal Approach: Packages "tournament mode" as a platform product for every title, with pricing, an SLA and a published operations runbook, and sets the policy that tournament builds freeze N days before events.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignDedicated fleet, 128 Hz, pinned cores, frozen build channel
ControlAdmin API to create, pause, restart; audit log of every admin action
EvidenceInput log + snapshots per match; retained per league rules
SpectatingDelayed relay at 2 min to prevent ghosting
IsolationNo shared Ready buffer with public play

Ownership Question: "Who can restart a league match?" Staff answer: League officials through the admin API, with every action audited; engineering on-call only on official request.

Key Takeaway: "Tournament play is a different product with the same binary — isolate its fleet, freeze its build, record everything."

What clears the Staff bar:

  • Isolates hardware and build channel
  • Gives officials controls instead of engineers
  • Records inputs for dispute resolution

Deep Dive 4: Post-Mortem — Duplicated Items After a Zone Server Crash#

Context: In the game's persistent-world mode, a zone server crashed during a busy trading hour. Afterwards, the economy team finds ~4,000 rare items duplicated. Players had traded them; one side's trade was persisted and the other was rolled back to the last checkpoint.

Questions to Surface First:

  • Which writes are checkpointed and which are transactional?
  • Was the trade applied in memory first and persisted later?
  • How do we identify duplicates without punishing honest players?

Typical L5 Approach: Shortens the checkpoint interval from 60 s to 10 s.

Staff Approach: Moves economic events — trades, purchases, rare drops, crafting with rare inputs — out of the checkpoint and into a transactional service: the zone server requests the trade, the service commits both sides atomically with an idempotency key, and only then does the zone server reflect it. Positions and health stay on checkpoints because losing 60 s of movement is harmless. Each item gets a unique instance ID so duplicates are detectable.

Principal Approach: Sets a cross-title rule: anything with real-money or market value is never owned by a game server's memory; game servers request and display, the economy service commits. Adds an economy audit (item creation vs destruction per hour) as a monitored invariant.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateFreeze trading for affected items; snapshot ownership by item instance ID
TriageTrade committed in one checkpoint and lost in the other side's rollback
Quick fixRoute trades through the transactional economy service
GuardrailsItem instance IDs; hourly creation-vs-destruction invariant
Post-mortemWhy was economic state owned by a crash-prone process?

Ownership Question: "Who owns item integrity?" Staff answer: The economy service team owns commits and invariants; zone servers are clients of it. Trust and safety owns remediation for affected players.

Key Takeaway: "Checkpoint what's cheap to lose. Commit what's expensive to duplicate."

What clears the Staff bar:

  • Separates state by cost of loss, not by where it lives
  • Makes economic events atomic and idempotent
  • Adds an invariant that would have caught it

Deep Dive 5: Multi-Region Expansion — Launching in a New Continent#

Context: The game is launching in a region with ~400K expected peak players, currently served from a region 150 ms away.

Questions to Surface First:

  • Where are the players by ASN and city, and what's the RTT map to candidate sites?
  • Is the population large enough for healthy ranked pools?
  • What control-plane pieces must be local?

Typical L5 Approach: Deploys a full copy of the backend in a cloud region in that continent.

Staff Approach: Places game-server fleets (the latency-sensitive part) at one or two sites chosen from client RTT maps; keeps pools regional but allows cross-region matching with an RTT cap when pools are thin; keeps progression global with async result delivery; checks whether ranked pools at 4 a.m. local stay above the quality floor, and if not, merges with the neighboring region at night.

Principal Approach: Decides between cloud regions, bare-metal partners and interconnect investments with a 3-year player forecast; sets the launch criteria (latency percentile, pool health) and the exit criteria if the population doesn't arrive.

Staff Approach — Full Reasoning
PhaseWhat to Do
MeasureClient RTT beacons to candidate sites for 4 weeks pre-launch
PlaceFleets at sites covering 90% of players under 60 ms
PoolsRegional pools; night-time merge rule with RTT cap
Control planeTicket service and allocator local; progression global
LaunchPre-scaled fleet; admission gate; game day before launch

Ownership Question: "Who decides to merge pools at night?" Staff answer: Game design owns the policy and quality floor; matchmaking implements it as config with a schedule.

Key Takeaway: "Put the servers where the latency is; put the pools where the players are; keep progression in one place."

What clears the Staff bar:

  • Uses measured RTT to place fleets
  • Checks pool health off-peak, not only at peak
  • Separates latency-sensitive from latency-tolerant components

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Separate the disposable real-time match from the durable control and results path, and justify where availability and durability budgets go
  • Choose an authority model by genre and defend dedicated authoritative servers for competitive play
  • Derive frame budget, matches per core, bytes per player and fleet egress from tick and send rates
  • Explain prediction, interpolation and lag compensation, and who pays for favor-the-shooter with a rewind cap
  • Design matchmaking pools and expansion schedules, and reason about pool size with Little's law
  • Size a warm buffer from allocation rate and start time, and protect Allocated servers from eviction
  • Draw the client trust boundary: inputs only, outcomes on the server, information withholding
  • Make results exactly-once by match_id and survive results-pipeline outages without draining the fleet

The Bar for This Question#

Mid-level (L4): Builds a lobby, a WebSocket relay and a database. Works for a turn-based game; has no tick, no fleet, no matchmaking model and trusts the client.

Senior (L5): Adds authoritative servers, rating-based matchmaking and autoscaling. Knows UDP is common. The gap: no budgets — no frame time, bytes or queue math; replication proposed for match failover; allocation and cold start unaddressed; anti-cheat treated as software to install rather than a boundary to design. The design would demo well and fall over on patch day.

Staff+ (L6): Separates the match path from the control path in the first five minutes. Prices tick rate in cores and egress, bounds lag compensation as policy, models matchmaking as a queue with a published dial, sizes the warm buffer from rates, protects live matches, and draws the trust boundary at inputs. Names who pays: targets absorb favor-the-shooter, tail players absorb wait or quality, the company absorbs egress, the platform owns allocation and the results pipeline. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 Don't Make Matches Highly Available#

ClaimReality
"Every service must survive host failure"A 25-minute match is worthless once it ends; replication doubles cost for a rare event
"Players hate crashes"Players hate losing points to crashes; voiding fixes that
"Failover is seamless"Migrating a 64 Hz simulation mid-round is rarely seamless and always complex

The Staff position: Make matches disposable and results exactly-once. Spend the reliability budget on allocation, builds and the results path.

Why this matters in interviews: Proposing per-match failover signals generic HA instincts; rejecting it with reasons signals judgment.

10.2 Tick Rate Is Mostly a Cost Decision#

The Staff position: Above ~30 Hz, the players who can feel the difference are the competitive ones, and the cost is linear. Riot's 128-tick required an engineering program to stay above three games per core. Choose tick per mode, and remember that most hit-reg complaints are latency and loss, not tick.

Why this matters in interviews: "Just raise the tick rate" without pricing it is an L5 answer.

10.3 Lag Compensation Doesn't Remove Unfairness; It Assigns It#

The Staff position: Someone always absorbs latency — the shooter who must lead, or the target who dies behind cover. The rewind cap is a fairness policy that product should own and publish.

Why this matters in interviews: Naming the victim of your netcode is the "who pays" move in its purest form.

10.4 The Best Anti-Cheat Is Data You Never Send#

The Staff position: Client-side anti-cheat is an arms race. Server authority over outcomes and information withholding remove whole categories of cheats structurally, with a measurable CPU cost. Buy client anti-cheat; design the server so it matters less.

Why this matters in interviews: "We'll add anti-cheat" is a vendor name. "The client never learns what it can't see" is a design.

10.5 Your Matchmaker's Real Enemy Is Your Product Roadmap#

The Staff position: Every new mode, platform split or party-size queue fragments the pool. Thin pools force long waits or bad matches. Matchmaking quality is protected by saying no to fragmentation — or by rotating modes — more than by any algorithm.

Why this matters in interviews: Bringing population and fragmentation into the conversation is the bridge from Staff to Principal.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs one game's backend well. The Principal engineer notices that the company ships four titles, each of which built its own login, matchmaker, fleet scripts, results pipeline and anti-cheat integration — and that each launch rediscovered the same patch-day failure. Hosting is one of the largest line items after salaries, egress is negotiated title by title, and player identity, bans and purchases are fragmented across studios. The L7 problem is a game platform: which parts every title must share (identity, parties, allocation, fleet operations, results, telemetry, trust and safety), which parts every title must own (the match runtime, the match function, the tick rate), and how to price the platform so studios adopt it instead of forking it.

🧭 Principal Move: "I'm not going to ask whether this title's matchmaker is good. I'm going to ask what we pay per player-hour across titles, how many launches last year had queue incidents in the first hour, and how many separate ban lists we maintain. Those three numbers tell me which parts of the backend should be a platform."

The Org-Level Fault Line#

One game platform vs per-studio backends.

OptionWhat WorksWhat BreaksWho Pays
Each studio owns its backendFits each game; studio autonomy; no central bottleneckEvery launch relearns capacity planning; N ban lists; egress negotiated N timesPlayers at launch; finance; on-call in every studio
One platform owns everything including the match runtimeOne stack, one cost modelForces one netcode model on a shooter and a fighting gameStudios (game feel); platform (special cases)
Platform owns the control path and fleet operations; studios own the match runtime and match functionShared launch readiness, identity, trust, cost leverage; studios keep game feelContract design; platform must stay ahead of studio needsPlatform (stewardship); studios adopt the contract

🧭 Principal Move: "The platform owns everything that must be right once — identity, allocation, Allocated protection, results exactly-once, telemetry, bans. Studios own everything that makes their game feel like their game — the server binary, tick rate, match function. New titles launch on the platform by default; existing titles migrate the control path first, because that's where launch incidents happen."

Cost Model#

Assumptions: fully loaded engineer ~$250K/year; cloud compute ~$2K per 64-core host-month; egress $0.01–0.02/GB negotiated (list prices are several times higher); average concurrency ~45% of peak. Estimates, not quotes.

ScalePeak Players in MatchCompute ($/month)Egress ($/month)HeadcountOn-call Load
Indie / first title20K~$15–25K (managed hosting)~$5–10K2–3 eng, managed servicesShared rotation; launch weeks are the risk
Mid-size300K~$150–250K~$60–120K8–12 eng: fleet, matchmaking, resultsDedicated rotation; 2–5 pages/month
Large competitive title2M~$0.9–1.4M (autoscaled, partly bare metal)~$0.4–0.8M25–40 eng platform + trust and safetyFollow-the-sun; launch war rooms

The pricing insight: compute scales with frame time and egress scales with snapshot bytes — both are engineering choices. A 20% density gain or a 20% snapshot cut at the large scale is worth more per year than the entire matchmaking team. That's why frame time and bytes per player belong on the platform's KPI page next to queue p95.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Authority model (authoritative vs P2P)One-wayChanges the game's code architecture, cheat surface and cost; effectively a rewrite
Rating system semantics shown to playersOne-way-ishPlayers' ranks are a social contract; resets anger the most invested
Wire protocol and client compatibility windowOne-way-ishOld clients in the field; protocol changes need versioned negotiation
Ranked results schema and match_id semanticsOne-way-ishEvery downstream consumer — progression, leaderboards, esports — depends on it
Tick and send rate per modeTwo-wayServer config; announce and canary
Matchmaking expansion scheduleTwo-wayConfig with staged rollout
Hosting provider / fleet operatorTwo-way if servers integrate through a thin SDK interfaceDirect calls to one provider's APIs across the codebase make it one-way

🧭 Principal Insight: The authority model and the meaning of a player's rank are the decisions I'd slow down on. Hosting providers, tick rates and expansion schedules can change in a quarter; the trust boundary and the rank players have ground for a year can't.

The Standard I'd Write#

RFC-GAME-001: Multiplayer Session Platform Standard
Status: Approved   Owners: Game Platform + Trust and Safety

Scope
  Every title with online multiplayer sessions using company infrastructure.

MUST
  1. Competitive modes use a server-authoritative model: clients send inputs;
     outcomes, damage and results are computed on servers.
  2. Game servers make no blocking network calls on the simulation thread.
  3. Match results are emitted once per match, signed by the server, keyed by
     match_id, and applied idempotently.
  4. Allocated game servers are never terminated by scaling, updates or node
     maintenance; maintenance waits for match end.
  5. New server builds receive new allocations by staged share with automatic
     halt on crash-rate regression.
  6. Matchmaking publishes p95 wait and p95 rating spread per mode and band;
     expansion schedules are config owned by game design.
  7. Snapshots send only state the receiving player may perceive.

SHOULD
  1. Use the platform fleet operator and allocator.
  2. Pre-scale from a signed capacity forecast before launches and seasons.
  3. Track cost per player-hour and frame time per match as release metrics.

Exceptions
  Filed with Game Platform; reviewed within 5 business days; MUST 1 exceptions
  (e.g., P2P rollback) require Trust and Safety sign-off and exclude ranked
  rewards with market value.

Success metrics
  - Launch-hour queue incidents: 0 per launch
  - Crashed matches: < 10 per 10,000 per build
  - match.alloc_latency_p99 < 2 s in every region
  - Cost per player-hour: -20% over 2 years at constant tick

What I'd Tell the VP#

"Each of our games built its own online backend, and each launch had the same first-hour problem: more players than servers ready, and a queue that broke instead of waiting. I'm proposing one shared platform for logins, matchmaking infrastructure, servers and results, while each studio keeps control of how its game feels. It's about 30 engineers across the platform and trust and safety, mostly people we already fund in separate studios. The payoff is predictable launches, one ban list for cheaters across all titles, and leverage on our hosting bill — compute and bandwidth are among our largest costs, and engineering choices can cut them by a fifth. I'll report cost per player-hour and launch incidents every quarter."

Principal Interview Signals#

SignalWhat It Sounds Like
Counts launches, not servers"How many launches last year had a first-hour queue incident?"
Prices the experience"Tick rate and snapshot bytes are line items; I'd put cost per player-hour on the dashboard."
Sets the org's failure posture"Pre-scale from a signed forecast, admission gates, and a region-loss game day each season."
Identifies one-way doors"The authority model and rank semantics are forever; the hosting provider isn't."
Knows when not to standardize"Studios keep the match runtime; a fighting game shouldn't run our shooter's netcode."

Staff answers that L7 interviewers find insufficient:

  • "We'll design a great matchmaker for this title" — correct for one title; silent on the four other matchmakers.
  • "We'll add anti-cheat" — a vendor, not a cross-title trust platform with shared ban identity and appeals.
  • "We'll autoscale" — no forecast ownership, no cost target, no launch-readiness standard.

Appendices

Appendix A: Mechanics in Depth#

A.1 The Server Tick Loop#

const TICK_HZ = 64, DT = 1.0 / TICK_HZ           # 15.625 ms
loop:
    t0 = now()
    for client in clients:
        for input in client.queue.drain():       # carries client_tick + last 3 inputs
            if not plausible(input, client): metric("gs.input_rejected"); continue
            apply_input(client.entity, input)
    simulate(world, DT)                          # movement, physics, abilities
    for shot in pending_shots:
        resolve_with_rewind(shot)                # A.2
    run_rules(world)                             # damage, objectives, rounds
    history.record(tick, hitboxes(world))        # ring buffer, ~13 ticks = 200 ms
    for client in clients:
        visible = relevancy(client, world)       # A.4
        pkt = encode_delta(visible, base=client.last_acked_snapshot)
        send_udp(client, pkt)                    # one packet, < 1,200 B
    metric("gs.frame_time_us", now() - t0)
    sleep_until(t0 + DT)

A.2 Lag Compensation#

def resolve_with_rewind(shot):
    shooter = shot.player
    claimed_view = shot.client_view_tick
    measured = server_tick - ticks(shooter.rtt_ms / 2 + shooter.interp_ms)
    view = clamp(claimed_view, measured - 1, measured + 1)     # don't trust the claim
    view = max(view, server_tick - ticks(MAX_REWIND_MS))      # cap: 200 ms
    if view != claimed_view: metric("gs.rewind_clamped")
    targets = candidates_near_ray(shot)                        # only rewind what's needed
    with history.rewound(targets, view):
        hit = raycast(shot.origin, shot.dir, targets)
    if hit: apply_damage(hit.target, shot.weapon)

A.3 Client Prediction and Reconciliation#

on local input i at client tick n:
    pending.append((n, i)); apply_input(me, i); send(Input(n, ack, pending[-3:]))

on snapshot s:
    me = s.my_state                              # server truth at tick s.ack_input
    pending = [p for p in pending if p.tick > s.ack_input]
    for (n, i) in pending: apply_input(me, i)    # replay unacknowledged inputs
    smooth_correction(render_me, me, 100ms)      # blend, don't snap
    buffer_others(s)                             # render at now - interp_delay

A.4 Interest Management#

Spatial grid of 20–50 m cells; each client's relevant set is entities in nearby cells, filtered by a precomputed visibility table between cells (with a look-ahead margin so entities don't pop in). Battle royale adds distance tiers: full rate under 50 m, one in two snapshots under 200 m, one in eight beyond. For a persistent world the same grid drives which zone server owns which cells.

Appendix B: Data Model#

players(player_id PK, account_id, home_region, platform, created_at)
ratings(player_id, mode, mu, sigma, games, updated_at, PK(player_id, mode))
tickets(ticket_id PK, party_id, mode, pool_key, created_at, rating, rtts_json,
        state, match_id)                         -- in Redis; TTL 10 min
matches(match_id PK, mode, map, region, build_id, gs_id, teams_json, state,
        created_at, ended_at)
match_results(match_id PK, outcome, per_player_json, server_sig, received_at)
applied_results(match_id, player_id, applied_at, PK(match_id, player_id))   -- dedupe
items(item_instance_id PK, owner_id, item_def, created_by_event, created_at)

Ticket claims use an atomic compare-and-set on state (searching → proposed) so a ticket can't enter two matches; proposals that fail allocation release their tickets back to searching with their original created_at, so they keep their place in the expansion schedule.

Appendix C: Coordination Mechanisms#

MechanismUsed ForWhy Not Something Else
Single process per matchAll match stateNo coordination at all on the hot path
Atomic CAS on ticket stateOne ticket, one matchCheap, per-ticket; no global lock on the pool
Operator-managed Ready → Allocated transitionOne server, one matchAllocation must be atomic and respect live matches
Signed connect token (TTL 30 s)Binding players to a serverServer verifies offline; no callback to the control plane at connect
Result keyed by match_id + per-player dedupe tableExactly-once rating and rewardsAt-least-once delivery is easy; dedupe makes it effectively once
Transactional economy serviceTrades, purchases, rare dropsCheckpoints lose or duplicate; economic events need atomic commits
Diagram: Appendix C: Coordination Mechanisms

Appendix D: API Contract and Session Lifecycle#

Diagram: Appendix D: API Contract and Session Lifecycle
  • Queue: POST /queue is idempotent per party; re-queue after cancel uses jittered backoff (1–5 s).
  • Connect token: {match_id, player_id, server_addr, exp} signed by the allocator's key; server verifies locally; expired or mismatched tokens are refused.
  • Reconnect: within 60 s a dropped player can rejoin with a fresh token from the gateway; after that the slot is abandoned (casual may backfill; ranked never does).
  • Result delivery: server spools to local disk, publishes asynchronously, exits; forwarder retries; consumers dedupe.

Appendix E: Observability#

MetricAlertWhy
gs.frame_time_p99 by build × map> 12 ms at 64 HzLate ticks = rubber-banding
match.alloc_latency_p99 by region> 2 sPlayers wait between "match found" and loading
fleet.ready_buffer by region< 30% of targetExhaustion imminent
gs.crashes_per_10k by build × map> 2× previous buildHalts the build's allocation share
mm.wait_p95 by mode × bandOver target 15 minQueue health
mm.rating_spread_p95 by mode × bandOver target for 1 daySilent quality drift
client.packet_loss_pct by region × ASN> 2% vs trailing weekRoute problems invisible on servers
gs.rewind_clamped_ratio by region × ASN> 2%Hit-reg complaints incoming
gs.terminations_while_allocated> 0 not from crashesSomething evicted live matches
results.pending_on_nodes> 0 for 5 minResults pipeline down
progression.lag_seconds> 120 sRatings delayed

Control plane vs data plane: a matchmaker or allocator outage stops new matches and pages immediately; a results outage delays ratings and pages in business hours unless pending results grow past a day; running matches have no control-plane dependency.

Appendix F: Scale Evolution#

StageWhat WorksWhat You Add
< 10K concurrentManaged hosting, managed matchmaker, single regionServer authority, results by match_id, warm buffer
10K–300KFleets per region on Kubernetes with a game-server operatorRollout by allocation share, quality SLOs, admission gate
300K–2MMany regions, per-ASN telemetry, pool sharding by bandForecast pre-scaling, tournament fleets, transactional economy
2M+ / multiple titlesShared platform, hybrid bare metal + cloudCost per player-hour KPI, cross-title trust and safety

What you don't build on day one: your own fleet scheduler, match failover, a custom rating algorithm, kernel-level client software, a private network backbone, or batch-optimized matchmaking. Each becomes worth it only when a measured number — cost, latency, wait, cheat rate — says so.

  1. Loading the index…