Technologies referenced in this case study: Redis · Kubernetes · etcd & ZooKeeper · Envoy, Kong & NGINX · Apache Kafka
Related: Chat App · Live Updates · Multiplayer Game · Video Streaming · CDN · Load Balancing · Multi-Region · Latency & Protocols · WebSockets vs SSE vs Long Polling
Reading Guide#
Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal View) and the appendices on bandwidth estimation, layer selection and cascading |
What is a Video Conferencing Service? — Why interviewers pick this topic
A video conferencing service lets a group of people see and hear each other in real time: a 1:1 call, a 12-person standup, a 300-person all-hands, a 20,000-viewer webinar. Clients capture camera and microphone, encode them (typically VP8, VP9, AV1 or H.264 for video, Opus for audio), and send RTP packets over UDP to media servers, which forward them to everyone else. A signalling service — usually WebSocket — handles joining, who is in the meeting, mute state and the session negotiation that sets up the media path.
The hard part is not moving video. The hard part is that every participant has a different network, and the slowest path must not degrade everyone else. A conference has a mouth-to-ear latency budget of roughly 150–300 ms, which leaves no time for retransmission-heavy reliability, no time for buffering, and very little time for routing media through a distant region. Packets will be lost, delayed and reordered; one participant on hotel Wi-Fi will be losing 8% of packets; another behind a corporate firewall can't send UDP at all. The media servers can't decode and re-encode for each viewer without burning a CPU core per stream, so they select which pre-encoded layers to forward to whom — and that selection is the system.
Before vs After — the "one bad connection" scenario:
Without simulcast and per-subscriber adaptation:
t=0: 12-person meeting on one SFU. Each sender uploads one 1.5 Mbps 720p stream.
t=+30s: Participant P7 joins from a train: 1.2 Mbps downlink, 6% loss.
t=+35s: P7 can't receive 11 x 1.5 Mbps. Congestion feedback from P7 reaches senders.
t=+40s: Senders reduce their single encode to 300 kbps to suit P7. Everyone now sees blurry video.
t=+2min: P7's loss triggers keyframe requests every few seconds; every sender emits large
keyframes; other participants' links spike and stutter.
t=+5min: Three people turn off video "because the call is bad". Meeting quality: poor for all 12.
With simulcast, per-subscriber layer selection and receiver-side congestion control:
t=0: Each sender uploads three layers: 180p (150 kbps), 360p (500 kbps), 720p (1.5 Mbps).
t=+30s: P7 joins. The SFU's bandwidth estimate for P7's downlink: ~1 Mbps.
t=+35s: SFU forwards to P7: active speaker at 360p, four thumbnails at 180p. Total ~1.1 Mbps.
t=+40s: Everyone else still receives 720p of the active speaker. No sender changed anything.
t=+2min: P7's loss handled by NACK retransmission from the SFU's cache for P7 only;
keyframe requests from P7 are rate-limited and served from the low layer.
t=+5min: P7 has a usable call. The other 11 never noticed P7 existed.
Why interviewers reach for this question: It looks like "WebRTC plus a media server" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: the topology choice (mesh, SFU, MCU) driven by meeting size and cost; simulcast or SVC so the server adapts per viewer instead of forcing senders to the lowest common denominator; media server placement and cascading across regions when participants are on three continents; packet-loss and jitter handling inside a 150 ms budget; the 10–20% of users who need TURN relays; and the fact that a 20,000-person webinar is a different product from a 20-person meeting.
Mechanics Refresher: Real-Time Media Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| Signalling | WebSocket (or HTTP) channel to exchange session descriptions, ICE candidates, roster and mute state | Flexible; any transport | Not specified by WebRTC — you design it |
| ICE / STUN / TURN | Clients gather candidate addresses; STUN discovers the public address; TURN relays media when direct paths fail | Connects through NATs and firewalls | TURN relays cost bandwidth; TCP/TLS fallback adds latency |
| Mesh (P2P) | Every participant sends its stream to every other participant | No media servers; lowest latency for 2–3 people | Uplink grows as N−1; unusable beyond ~4 |
| SFU | Each participant uploads once; the server forwards selected streams to each subscriber without decoding | Scales to hundreds; cheap CPU; per-viewer selection | Downlink grows with visible tiles; needs simulcast/SVC to adapt |
| MCU | Server decodes all streams, composites one mixed stream per viewer, re-encodes | One downstream per client; good for weak clients and legacy endpoints | ~1 CPU core per output stream; adds 50–200 ms; inflexible layouts |
| Simulcast | Sender encodes 2–3 independent resolutions; SFU picks one per subscriber | Widely supported; simple SFU logic | Uplink cost ~1.5–2× a single stream |
| SVC | One stream with layered spatial/temporal enhancement layers; SFU drops layers | More efficient uplink; finer switching | Codec support (VP9, AV1) and more complex SFU logic |
| NACK / RTX | Receiver asks for retransmission of specific lost packets | Recovers loss with low overhead | Costs a round trip; useless when RTT is near the budget |
| FEC / RED | Sender adds redundant data so receivers rebuild lost packets | No round trip | Constant bandwidth overhead (10–50%) |
| Jitter buffer | Receiver holds packets briefly to reorder and smooth arrival | Smooth playback | Every millisecond of buffer is latency |
| Congestion control | Bandwidth estimation (e.g. transport-wide feedback) drives encoder bitrate and layer choice | Avoids self-inflicted loss | Estimators oscillate on bad links |
For most production systems: SFU topology for anything beyond a 1:1 call, with simulcast (or SVC where codecs allow), per-subscriber layer selection driven by bandwidth estimation, NACK for video and Opus in-band FEC for audio, TURN over UDP with TCP/TLS on 443 as fallback, regional SFUs chosen by participant proximity and cascaded when a meeting spans regions, and a separate broadcast path for audiences beyond a few hundred. The codec isn't the interview. Per-viewer adaptation, server placement and the latency budget are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
Video conferencing is not a "use WebRTC" question. Everyone can draw clients connected to a media server.
It is a per-participant adaptation and latency-budget question that tests:
- Whether you choose the topology (mesh, SFU, MCU) from meeting size, client capability and cost — and know where each breaks
- Whether one participant's bad network stays that participant's problem, through simulcast or SVC and per-subscriber layer selection
- Whether you place media servers by where participants are, and cascade servers when a meeting spans regions, within a ~150–300 ms mouth-to-ear budget
- Whether you separate interactive meetings from broadcast audiences, because a 20,000-viewer webinar has different latency, cost and failure requirements
The key insight: In a conference, the receiver decides what it can take, and the server decides what to send it. Senders encode a few layers once; the SFU picks, per viewer and per second, which layer of which participant to forward. Staff candidates design that selection loop and its placement; Senior candidates design the connection setup.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | WebRTC clients → signalling server → media server | Asks "Meeting sizes? Two people, 12, 300, or a 20,000-viewer webinar? Where are participants? What's the latency target and do we need recording?" | Asks "Is this a product, or a capability embedded across many products? Who owns call quality as a metric, and what's the cost per participant-minute we can afford?" |
| Topology | "SFU" | "Mesh for 1:1 if both peers connect directly; SFU for meetings; MCU only for legacy room systems and phone dial-in; broadcast via SFU cascade or streaming for large audiences" | Prices topologies per participant-minute and sets the default by meeting-size histogram |
| Adaptation | "WebRTC adapts bitrate" | "Simulcast 3 layers (or SVC); SFU selects per subscriber from bandwidth estimate and tile size; only the active speaker gets high resolution; audio always prioritized over video" | Sets quality SLOs (e.g. % of participant-minutes with freeze < 1%) as the product metric, not server uptime |
| Placement | "Deploy in a few regions" | "Each participant connects to the nearest SFU; servers in different regions cascade, forwarding one copy of each stream across the backbone; meeting state in a regional coordinator" | Decides build vs buy of the edge network; negotiates peering and egress, the dominant cost |
| Failure | "Reconnect on failure" | "ICE restart and reattach to a new SFU in < 3 s; audio continues on another path; TURN over TLS 443 for blocked UDP; keyframe and rejoin storms rate-limited" | Designs regional blast radius: meetings pinned to cells; capacity reserved for failover; game days on regional loss |
| Large meetings | "Bigger server" | "Above ~100 active video senders, last-N video (only the N most recent speakers forward video); above ~1,000 viewers, a webinar mode where the audience receives only, via cascaded SFUs or low-latency streaming" | Treats webinars as a separate product with its own pricing, latency tier and capacity planning |
Why "adaptation" separates levels
L5: "WebRTC has built-in congestion control, so video quality adapts." It does — for one sender to one receiver. In a group call through a server forwarding a single stream per sender, the sender's encoder can only adapt to one target. If it adapts to the weakest receiver, everyone gets weak-receiver quality. If it ignores them, the weak receiver drowns.
L6: "Each sender publishes three simulcast layers — roughly 180p at 150 kbps, 360p at 500 kbps, 720p at 1.5 Mbps — or one SVC stream with equivalent layers. The SFU estimates each subscriber's downlink from transport-wide feedback and allocates it: audio first, then the active speaker or pinned tile at the highest layer that fits, then thumbnails at the lowest. A subscriber on 1 Mbps gets a 360p speaker and 180p thumbnails; a subscriber on 20 Mbps gets 720p. Senders never change because of one bad viewer. And if nobody is subscribed to a sender's 720p layer, the SFU tells the sender to stop encoding it — saving their uplink and CPU."
L7: "Adaptation is a quality metric we can measure per participant-minute — freeze rate, resolution delivered, audio concealment rate. I'd make those the SLOs product and infrastructure share, because server uptime can be 100% while a third of users have a bad call."
Why "placement" separates levels
L5: "Put the meeting on a media server in us-east." For a meeting with participants in London, Singapore and São Paulo, every packet now crosses an ocean twice — Singapore to Virginia to London is 250+ ms before any processing, past the budget.
L6: "Each participant connects to the nearest SFU — London, Singapore, São Paulo. The SFUs form a cascade for this meeting: each forwards one copy of each local publisher's stream to the other SFUs over the backbone, and each serves its local subscribers. Last-mile latency is short, and the long-haul hop happens once per stream rather than once per viewer. The meeting's control state lives in one coordinator, chosen in the region of the first joiner or the majority."
L7: "Cascading moves cost from user last-mile to our backbone. At scale, inter-region transfer and egress are the largest line items, so placement is a finance decision as much as a latency one — I'd model cost per participant-minute by region pair and decide where we peer and where we buy transit."
Why "large meetings" separates levels
L5: "The SFU can handle 1,000 participants if we give it a big enough machine." The CPU might hold up; the bandwidth won't. If 1,000 participants each send video and each subscribes to 25 tiles, the server forwards 25,000 streams — and the clients' downlinks and CPUs can't decode 25 streams anyway.
L6: "Interactivity and audience size trade off. Up to ~100 people, everyone can publish; subscribers receive the active speaker plus a page of tiles. Beyond that, last-N: only the N most recent speakers (say 25) forward video; everyone else is audio-only until they speak. Beyond ~1,000, a webinar: a handful of panelists publish; the audience only subscribes, fanned out through a tree of cascaded SFUs, or — if 3–10 s latency is acceptable — through low-latency streaming over a CDN at a fraction of the cost."
L7: "Webinars are a different product: different latency tier, different cost structure, different failure posture — a 20,000-viewer town hall is a scheduled event with a capacity reservation and a dry run, not a meeting that happens to be big."
Positions to Commit To#
| Position | Rationale |
|---|---|
| SFU by default; mesh only for 1:1 with a direct path; MCU only for legacy endpoints and dial-in | SFU scales without decoding; MCU costs a core per output; mesh's uplink explodes past ~4 |
| Simulcast (or SVC) on every video sender; per-subscriber layer selection | One bad network must not lower everyone's quality |
| Audio before video, always | People tolerate frozen video; they leave when audio breaks |
| Connect to the nearest SFU; cascade across regions | Last-mile latency dominates; long-haul once per stream, not per viewer |
| TURN everywhere, with TCP/TLS on 443 as fallback | 10–20% of users can't connect directly; corporate networks block UDP |
| Meetings and webinars are different modes | Interactive latency at 20,000 viewers is expensive and rarely needed |
| Quality SLOs per participant-minute, not server uptime | Servers can be up while calls are bad |
Which Problem Are We Solving?#
Four intents produce four different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| 1:1 calls | 2 participants; lowest latency; cheapest | P2P when ICE finds a direct path; TURN relay otherwise; upgrade to SFU when a third joins | NAT traversal failure; quality on asymmetric links | Call connects in < 2 s; mouth-to-ear < 200 ms |
| Interactive meetings (2–100) | Everyone may speak and show video; heterogeneous networks | SFU, simulcast/SVC, per-subscriber selection, nearest-SFU + cascade | One weak participant degrades all; regional latency | Mouth-to-ear < 300 ms; freeze < 1% of video-minutes |
| Large interactive meetings (100–1,000) | Many participants, few concurrent speakers | SFU cascade, last-N video, audio-level-based forwarding, paged galleries | Server bandwidth and client decode limits | Speakers heard within 300 ms; gallery paging works |
| Webinars / broadcasts (1,000–100,000) | Few presenters; large passive audience; scheduled | Panelists on SFU; audience via SFU fan-out tree or low-latency streaming over CDN | Cost explosion; join storm at start time | Viewers within 1–2 s (SFU tree) or 3–10 s (streaming); join success > 99.9% |
🎯 Staff Move: "I'll design interactive meetings up to a few hundred participants as the core product — SFU-based with simulcast and regional cascading — because that's where per-viewer adaptation and placement are hardest. 1:1 calls are a special case that can go peer-to-peer. Webinars beyond a thousand viewers I'll treat as a separate mode at the end: panelists in a meeting, audience on a broadcast path."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Mesh vs SFU vs MCU | Who bears the cost — client uplink (mesh), server bandwidth (SFU), or server CPU (MCU)? |
| 2 | Simulcast vs SVC, and Who Adapts | Multiple independent encodes vs one layered stream; sender-side vs server-side adaptation |
| 3 | Single Media Server per Meeting vs Regional Cascading | Simple routing vs latency for globally spread participants |
| 4 | Loss and Jitter: Retransmit vs Redundancy vs Degrade | NACK costs a round trip; FEC costs bandwidth; buffering costs latency |
| 5 | Interactive Meeting vs Broadcast for Large Audiences | Sub-second interactivity for everyone vs streaming economics for viewers |
How Real Companies Built It#
Why this section belongs here: Three publicly documented systems show three distinct design points: a custom SFU fleet optimized for cost, cascaded bridges for geography, and an anycast network where every location is part of the media path.
Discord — A Custom SFU, Minimal Signalling, Least-Utilized Server per Region#
Discord's engineering blog describes voice and video built on WebRTC with a homegrown SFU written in C++, chosen for performance and cost and because routing through servers hides users' IP addresses and enables moderation. Because every client connects to a relay server rather than to peers, Discord replaced standard SDP/ICE negotiation with a much smaller signalling exchange (about 1,000 bytes instead of ~10 KB) and used its own encryption instead of DTLS/SRTP for native clients. A guild service watches service discovery and assigns the least-utilized voice server in the chosen region; on server failure, clients reconnect to a newly assigned server. At the time of writing the system served about 2.6 million concurrent voice users on 850+ voice servers in 13 regions and 30+ data centers, with 220 Gbps of egress (Discord engineering blog).
Staff insight: WebRTC is a toolkit, not an architecture. When every call goes through your server anyway, peer-to-peer negotiation machinery is overhead you can strip. Say: "Once I commit to SFU-only, signalling simplifies to 'here is your server and your keys', and server assignment becomes a load-balancing problem I own."
Jitsi — Cascaded Bridges for Geographic Spread#
Jitsi Videobridge, an open-source SFU, supports "relays" (formerly Octo) in which multiple bridges participate in one conference: participants connect to a nearby bridge and the bridges forward media to each other, using ICE and DTLS/SRTP between each pair of bridges (Jitsi Videobridge relay docs). Jitsi's 2022 post on re-enabling cascading on its public service explains the motivation — connect users to a nearby server — and reports a 29% decrease in next-hop latency (from 223 to 158 ms) for endpoints on a different continent from the server, after cascading had been disabled in 2020 while the service scaled to 50+ shards and 2,000+ bridges (Jitsi blog).
Staff insight: Cascading is how you buy latency with backbone bandwidth — and it was operationally hard enough that a large public deployment turned it off during a growth surge and brought it back later. In the interview: "Cascading is the right default for multi-region meetings, but it multiplies inter-server connections; I'd bound which regions cascade with which, rather than full mesh between all servers."
Cloudflare — An Anycast SFU With Per-Track Cascading Trees#
Cloudflare's 2024 post on its real-time media service describes an SFU that runs in every Cloudflare location: each WebRTC peer connection reaches the closest data center through anycast, so there is no region to choose, and media is distributed through "cascading trees that are unique per track" between data centers. It handles NACK retransmission at the edge rather than propagating it back to the publisher, offers TURN over anycast on several ports including 443, and leaves room and signalling to the developer (Cloudflare blog).
Staff insight: Placement can be pushed into the network layer. Per-track trees generalize cascading — each published stream gets its own distribution tree. And edge NACK handling is the point to steal: "Retransmissions should come from the server nearest the receiver, never from the sender across an ocean."
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "We use an SFU" | "One participant is on a 1 Mbps train connection. What does everyone else see?" | Simulcast/SVC, per-subscriber selection |
| "Participants connect to our media server" | "Half the meeting is in Sydney and half in Frankfurt. Where is the server?" | Placement, cascading, latency budget |
| "WebRTC handles NAT traversal" | "A bank's network blocks all UDP. Does the call work?" | TURN, TCP/TLS 443 fallback |
| "The SFU forwards everything" | "It's a 500-person all-hands. How many streams is the SFU sending?" | Last-N, gallery paging, bandwidth math |
| "We record meetings" | "The recorder crashes 50 minutes into a 60-minute board meeting. What's lost?" | Recording architecture, checkpointing |
| "We retransmit lost packets" | "RTT is 250 ms. Does retransmission still help?" | Latency budget, FEC vs NACK |
| "We'll support webinars" | "20,000 people join at 10:00:00 exactly. What happens?" | Join storms, broadcast mode, admission |
System Architecture Overview#
Reading the diagram: Joining is a control-plane operation: geo routing picks the nearest region, signalling authenticates and admits the participant, the meeting coordinator decides which SFU in that region hosts this meeting (or adds a new SFU to the meeting's cascade if the region is new), and the client connects its media directly to that SFU — via TURN only if direct UDP fails. The SFU is the heart of the media plane: it receives each publisher's simulcast layers, estimates each subscriber's bandwidth and forwards the right layer of the right streams. SFUs in different regions hosting the same meeting cascade, sending one copy of each stream across. Recording, broadcast and phone dial-in are subscribers to the SFU, not part of it. The metrics that tell you the service is healthy are quality metrics per participant-minute — freeze ratio and audio concealment ratio — not server CPU.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Topology | "Use an SFU" | "P2P for 1:1 when possible, SFU for meetings, MCU only for legacy room systems and dial-in." |
| Weak participant | "WebRTC adapts" | "Simulcast or SVC; SFU picks layers per subscriber; senders never degrade for one viewer." |
| Global meeting | "Pick a region" | "Nearest SFU per participant, SFUs cascade, one long-haul copy per stream." |
| Firewalls | "STUN" | "TURN on UDP, TCP and TLS 443; 10–20% of sessions relay." |
| Loss | "Retransmit" | "NACK when RTT allows; Opus FEC for audio; drop to lower layer; adaptive jitter buffer." |
| Big meetings | "Bigger server" | "Last-N video, audio-level forwarding; webinars on a broadcast path." |
| SLO | "Server uptime" | "Freeze ratio, audio concealment, join success — per participant-minute." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| Mouth-to-ear latency target | ~150 ms good, ~300 ms tolerable, > 400 ms conversation breaks (typical guidance) | The whole budget: capture, encode, network, jitter buffer, decode, render |
| Opus voice bitrate | ~24–40 kbps typical | Audio is cheap; prioritize it absolutely |
| Simulcast layers (illustrative) | 180p ~150 kbps, 360p ~500 kbps, 720p ~1.5 Mbps | Uplink ~2.2 Mbps per sender; downlink depends on layout |
| Typical downlink per participant | 1–4 Mbps (speaker + gallery) | Sizes SFU egress: egress ≫ ingress |
| Sessions needing TURN relay | commonly 10–20% (illustrative; higher in enterprise) | TURN fleet bandwidth sizing |
| Packet loss where video degrades visibly | ~2–5% without FEC/NACK help | Thresholds for layer downgrade |
| Jitter buffer | ~20–200 ms adaptive | Each ms is latency; adaptive beats fixed |
| SFU per-server forwarding | several Gbps per server, bound by NIC and packet rate (illustrative) | Capacity planning in Gbps and packets/s, not participants |
| Discord voice fleet (2018 post) | ~2.6M concurrent users, 850+ servers, 13 regions (Discord) | A real reference for SFU fleet scale |
| Cascading latency gain | 223 → 158 ms next-hop for cross-continent endpoints (Jitsi) | Why nearest-server plus cascade beats one server |
| Join time target | < 2 s to first audio (typical) | Signalling + ICE + DTLS on the critical path |
| Last-N video senders | ~9–25 (typical) | Bounds per-meeting forwarding regardless of size |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes explaining SDP offers, ICE candidates and STUN, then run out of time before the interviewer asks the only question that matters: "One person in the meeting is on a terrible connection, and the others are on three continents. What does each participant actually receive?" Compress connection setup to ~3 minutes and spend the rest on topology, per-viewer adaptation, placement and cascading, loss handling and large meetings.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"Users create or join meetings by link; participants share audio, video and screen; the service shows a gallery and the active speaker, supports chat, mute and hand-raise, records on request, and scales from 1:1 calls to all-hands meetings, with a separate webinar mode for large audiences."
Then the non-functional requirements, which is where the design lives:
"Three constraints drive everything. One: latency — mouth-to-ear under ~300 ms, ideally near 150, which rules out buffering and limits retransmission. Two: heterogeneity — every participant has a different network, and the worst one must not lower everyone else's quality. Three: egress cost — media servers send far more than they receive, and bandwidth is the dominant cost. I'll assume 3 million concurrent participants at peak, about 500K concurrent meetings with a median of 4 and a long tail to 1,000, global users, and a 2-second join target."
Then name the underspecified parts:
"I'd confirm: is end-to-end encryption required — it constrains recording and transcription; is recording server-side or client-side; and what's the largest interactive meeting versus webinar audience. I'll assume server-side recording on request, optional E2EE that disables server features, interactive up to 1,000 and webinars up to 50,000."
🎯 Staff Move: Saying "the worst participant must not lower everyone else's quality, and egress is the dominant cost" reframes the problem from "connect people" to "adapt per viewer within a latency and cost budget" — that's the sentence that sets the level.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Meeting:
meeting_id,host_id,settings(recording, e2ee, waiting room, max_participants),region_hint,state(scheduled / live / ended) - Participant:
participant_id,meeting_id,user_id?,role(host / panelist / attendee),sfu_id,region,tracks[] - Track:
track_id,kind(audio / video / screen),layers[](rid, resolution, bitrate),muted - SFU assignment:
meeting_id,region,sfu_id,cascade_links[] - Subscription:
subscriber_id,track_id,target_layer,priority(speaker / pinned / thumbnail / off-screen)
Signalling messages (WebSocket, JSON):
→ join { meeting_id, token, capabilities: {codecs, simulcast, svc} }
← joined { participant_id, sfu: {url, ice_servers, turn_creds}, roster[] }
→ publish { tracks: [{kind: video, layers: [q, h, f]}, {kind: audio}] }
→ subscribe { track_id, max_resolution: "360p" } (from client layout)
← speaker_changed { participant_id, audio_level }
← layer_changed { track_id, layer: "h" } (SFU decision, informational)
→ leave {}
Control API:
POST /v1/meetings { settings } → { meeting_id, join_url }
POST /v1/meetings/{id}/recordings → { recording_id }
POST /v1/meetings/{id}/broadcast { mode: "ll-hls" } → { playback_url }
🎯 Staff Move: "The client tells the SFU what it can display — tile sizes from the layout — and the SFU decides what it can afford from the bandwidth estimate. The forwarded layer is the minimum of the two. No subscriber ever receives 720p for a 160-pixel thumbnail."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk one join in 90 seconds:
- The client opens the meeting link; geo routing sends it to the nearest region's signalling service, which validates the token and admits the participant.
- The meeting coordinator checks whether the meeting already has an SFU in this region. If not, the allocator picks the least-loaded SFU by forwarded bandwidth, and the coordinator adds it to the meeting's cascade.
- Signalling returns the SFU address, TURN credentials and the roster. The client runs ICE against the SFU — direct UDP first, TURN on UDP, then TCP/TLS 443 — then DTLS to derive SRTP keys.
- The client publishes audio plus three simulcast video layers. The SFU forwards the participant's audio to everyone and offers the video to subscribers, who request sizes based on their layout.
- For each subscriber, the SFU picks layers from the bandwidth estimate: audio first, active speaker at the best fitting layer, thumbnails at the lowest. If a remote region has subscribers, the SFU forwards one copy of each layer they need to that region's SFU.
- First audio within ~1.5–2 s of clicking join.
🎯 Staff Move: Say out loud: "The SFU never decodes video. It reads RTP headers, picks which packets to forward to whom, and rewrites sequence numbers so layer switches look seamless to the receiver. That's why one server can forward gigabits." You've now spent ~8 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes this hard is that networks are heterogeneous, participants are global, packets get lost inside a 150–300 ms budget, and large meetings break both server bandwidth and client decode. I'd like to go deep on topology choice, per-subscriber adaptation, placement and cascading, loss and jitter, and large meetings versus webinars. Where would you like to start?"
If no preference: start with per-subscriber adaptation. It's the question that decides the level.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: Per-subscriber adaptation (7–8 min)
"Senders publish three simulcast layers, or one SVC stream with three spatial and three temporal layers if the codec supports it. For each subscriber the SFU runs a bandwidth estimator from transport-wide congestion feedback, and every few hundred milliseconds re-allocates: audio for all audible speakers — capped at the top 3 by audio level — then the pinned or active speaker at the highest layer that fits, then gallery tiles from the lowest layer up. Switching up requires a keyframe on the new layer, so upgrades are conservative — hold 2–5 seconds of headroom before switching up, switch down immediately."
Quantify: "A subscriber estimated at 1.2 Mbps gets audio (~100 kbps for three speakers), the speaker at 360p (500 kbps) and four thumbnails at 180p (600 kbps). A subscriber at 8 Mbps gets the speaker at 720p and nine thumbnails at 360p."
Who pays: "The weak participant gets lower resolution — the right victim, since it's their link. Senders pay ~1.5× uplink for simulcast; that's the price of nobody else paying."
Deep dive 2: Placement and cascading (6–7 min)
"Participants connect to the SFU in their nearest region, chosen by a latency probe at join time, not by IP geolocation alone. A meeting with participants in three regions has three SFUs in its cascade. Each SFU forwards each local publisher's needed layers once to each other SFU over our backbone; each remote SFU then fans out to its local subscribers and does its own per-subscriber layer selection. Cascade links are bounded: for a meeting spanning 5+ regions, I'd use a hub topology — regional SFUs connect to one or two hubs — rather than full mesh."
Quantify: "Sydney to Frankfurt is ~250–300 ms RTT on the public internet; via a well-peered backbone maybe 10–20% better. One-way that's ~130–150 ms of the budget — painful but workable. Putting the server in Frankfurt makes Sydney participants pay that on every hop to and from the server; cascading lets everyone's last mile stay short."
Deep dive 3: Loss and jitter (5–6 min)
"Audio: Opus with in-band FEC and packet-loss concealment; at more than ~5% loss, add RED (redundant encoding) — audio is small, so redundancy is cheap. Video: NACK-based retransmission from the SFU's per-subscriber packet cache when RTT allows — the SFU is close to the subscriber, so the round trip is short even if the publisher is far. When loss persists above a few percent, drop the subscriber to a lower layer rather than fighting it. Keyframe requests (PLI) are rate-limited per publisher to about one per second and served where possible by switching the subscriber to a layer that has a recent keyframe."
Deep dive 4: NAT traversal and TURN (4–5 min)
"Every client gets TURN credentials, short-lived and scoped to this session. ICE tries direct UDP to the SFU first — since the SFU has a public IP, direct works for most users — then TURN over UDP, then TURN over TCP and TLS on 443 for networks that block UDP. TURN servers are co-located with SFUs in each region so the relay adds a hop of a few milliseconds, not a detour. I'd expect 10–20% of sessions to relay in consumer traffic and more in enterprise."
Deep dive 5: Large meetings and webinars (3–4 min)
"Up to ~100: everyone may publish; subscribers see a page of tiles. 100–1,000: last-N video — only the N most recent speakers' video is forwarded; others are audio-muted thumbnails until they speak — and a cascade across several SFUs in one region to spread egress. Above that, webinar mode: panelists in a small meeting; the audience subscribes only, through a fan-out tree of SFUs for sub-second latency, or via low-latency HLS over a CDN at 3–10 s latency and a fraction of the cost. Q&A goes through signalling, not media."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: in a conference, the receiver decides what it can take and the server decides what to send. So SFUs with simulcast or SVC and per-subscriber layer selection; nearest-region SFUs cascaded across regions; audio prioritized over video; NACK from the nearest server, FEC for audio, and lower layers under persistent loss; TURN on 443 for locked-down networks; and a separate broadcast path when the audience outgrows interactivity."
The evolution closer:
"What I'd build later: SVC with AV1 to cut uplink, end-to-end encryption via insertable streams for sensitive meetings, server-side noise suppression and transcription as subscribers, and anycast entry to remove region selection. What I'd not build: an MCU for general meetings — the CPU cost and added latency don't pay for themselves outside legacy interop."
🎯 Staff Move: End on what you measure. "I'd page on join success rate and on freeze and audio-concealment ratios per region, not on SFU CPU. A healthy server fleet with a broken peering link still means bad calls."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| SDP and ICE tour | 10 min on offer/answer and candidate types | "ICE finds a path: direct UDP, TURN UDP, TURN TLS 443" in 30 s |
| Codec deep dive | Explains VP8 vs H.264 internals | Names codecs; spends time on layer selection |
| One server per meeting | Puts the meeting in one region | Nearest SFU per participant, cascaded |
| No adaptation story | "WebRTC adapts bitrate" | Simulcast/SVC and per-subscriber allocation volunteered |
| Big meeting = big server | Vertical scaling | Last-N, cascade within a region, webinar mode |
| Server metrics as SLO | CPU and uptime | Freeze ratio, concealment, join success |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Video conferencing is the interview where the system's quality is decided on networks you don't control, in a time budget you can't extend. You can't make a participant's Wi-Fi better, you can't make light cross the Pacific faster, and you can't buffer your way out of loss without breaking conversation. The candidate must design a system where each participant's experience is decoupled from every other participant's network, where servers are placed by where people are, and where degradation is graceful and targeted — lower resolution for one viewer, not frozen video for twelve. That is Staff work: designing adaptation loops and blast-radius boundaries on a real-time path.
It also has a deceptive cost structure. Signalling is trivial in cost. Media egress is everything: a single participant receiving 2 Mbps for an hour is ~900 MB, and a service with 3 million concurrent participants is forwarding terabits per second. Every design choice — how many tiles, which layers, whether a webinar runs on SFUs or a CDN — is a cost choice.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"Your CEO's all-hands has 2,500 people: 400 in the office on corporate Wi-Fi, the rest remote across 14 countries. Three people present; 30 will ask questions. Ten minutes in, the Singapore office says video is frozen but audio is fine. Walk me through the architecture of this meeting and what's happening in Singapore."
A candidate who answers with the meeting's mode (webinar-style: panelists publish, attendees subscribe only, questioners promoted temporarily), the topology (regional SFUs per participant cluster, cascaded through a hub; Singapore participants on a Singapore SFU), the likely cause (the Singapore office's 400 people share one corporate uplink, or the cascade link to Singapore is congested; video freezes while audio survives because audio is prioritized and tiny), the levers (drop Singapore subscribers to a lower layer; check cascade link loss; enable FEC on the cascade link) and the metrics (freeze ratio and bandwidth estimate per region, cascade link loss) has run a conferencing service. A candidate who says "the SFU should scale" has used one.
2. Problem Framing & Intent#
2.1 The Four Intents — Explained#
1:1 calls → peer-to-peer when possible
- Constraint: two participants; lowest latency and cost
- Strategy: ICE for a direct path; TURN relay when it fails; migrate to an SFU when a third person joins
- Failure mode: NAT traversal failures; quality problems that are hard to debug without a server in the path
- Who pays for imperfection: users on restrictive networks (relay latency)
Interactive meetings → SFU with per-subscriber adaptation
- Constraint: everyone may speak; heterogeneous networks; global participants
- Strategy: SFU, simulcast/SVC, bandwidth estimation, nearest-region SFU with cascading
- Failure mode: weakest link degrades all; distant server adds latency; keyframe storms
- Who pays: participants on good networks (if adaptation is missing), the company (egress)
Large interactive meetings → bound the forwarding
- Constraint: hundreds of participants, few concurrent speakers
- Strategy: last-N video, audio-level forwarding of the top 3 speakers, paged galleries, intra-region cascade
- Failure mode: server egress and client decode limits; join storms at the scheduled start
- Who pays: participants who aren't speaking (thumbnails or audio-only)
Webinars → broadcast economics
- Constraint: few presenters, large passive audience, scheduled
- Strategy: panelists in a meeting; audience via SFU fan-out tree or low-latency HLS on a CDN
- Failure mode: 20,000 joins in 60 seconds; interactive-grade cost for passive viewers
- Who pays: the company (if run on SFUs unnecessarily), viewers (latency if on streaming)
2.2 When NOT to Build a Video Conferencing Service#
- Video calling is a feature, not your product. Telehealth visits, support calls, tutoring: use a real-time media platform or an open-source SFU deployment. Your differentiation is the workflow, not RTP forwarding.
- Your audience is passive and latency-tolerant. A product launch to 200,000 viewers is live streaming: ingest, transcode, CDN. See Video Streaming.
- You need sub-50 ms interaction. Real-time games and music collaboration need different tradeoffs — see Multiplayer Game. Conferencing's ~150 ms target is generous by comparison.
- You can't fund the network. Conferencing quality is mostly peering, backbone and egress. Without a budget for regional presence, a managed service with an existing global edge will beat anything you build.
🎯 Staff Insight: "The media server is open-source and commodity. The hard part of conferencing at scale is the network — where servers are, how they're peered, and what egress costs. Build when that network is an asset you'll own anyway."
2.3 What the Interviewer Leaves Underspecified#
Interviewers deliberately omit:
- Meeting-size distribution — the median meeting is small; the tail drives architecture
- Geographic spread — single-region and global meetings are different systems
- Encryption requirements — end-to-end encryption removes server-side recording, transcription and compositing
- Recording — server-side vs client-side, composite vs per-track
- Interop — phone dial-in and room systems often need mixing (MCU-like) components
- Peak shape — meetings start on the hour; 60–70% of joins can land in the first 2 minutes of each hour
Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.
2.4 Precise Terminology#
| Term | What It Means | Why It Matters in the Interview |
|---|---|---|
| SFU | Selective forwarding unit — forwards packets without decoding | The default media server |
| MCU | Multipoint control unit — decodes, mixes, re-encodes | CPU-heavy; legacy and dial-in |
| Simulcast | Multiple independent encodes per sender | Per-subscriber selection |
| SVC | One encode with droppable spatial/temporal layers | More efficient; codec-dependent |
| Layer | One resolution/frame-rate rung | Unit of selection |
| BWE | Bandwidth estimation from congestion feedback | Drives layer choice |
| NACK / RTX | Request and retransmission of lost packets | Costs one round trip |
| FEC / RED | Redundant data to recover loss without round trips | Costs bandwidth |
| PLI / FIR | Picture loss indication — receiver asks for a keyframe | Keyframes are large; storms hurt |
| Jitter buffer | Receiver-side reordering and smoothing | Adds latency |
| TURN | Relay server for media when direct paths fail | Bandwidth cost; TLS 443 fallback |
| Cascade | Multiple SFUs serving one meeting, forwarding to each other | Global meetings |
| Last-N | Forward video only for the N most recent speakers | Bounds large-meeting cost |
| Mouth-to-ear latency | Capture to playback, end to end | The budget everything fits inside |
🎯 Staff Insight: If the interviewer says "make it scale to 10,000 participants," ask: "10,000 who can all speak and show video, or 10,000 watching a few presenters? The first is a last-N meeting with heavy cascading; the second is a broadcast, and I'd design it differently."
3. Where the Design Splits#
Every conferencing decision has a technical side (topology, codecs, packet handling) and an organizational side (what quality we promise, what it costs per minute, who owns the network). Interviewers grade the second side.
3.1 Fault Line 1: Mesh vs SFU vs MCU#
The tension: Someone must do the work of getting N streams to N people. In a mesh, each client uploads N−1 copies. With an SFU, each client uploads once and the server forwards — server bandwidth scales with total subscriptions. With an MCU, the server decodes everyone and composes one stream per viewer — server CPU scales with outputs.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Mesh (P2P) | No servers; lowest latency for 2–3 | Uplink N−1 × bitrate; client CPU for N−1 encodes; IP addresses exposed | Participants (uplink, battery) |
| SFU | One upload; per-subscriber selection; no transcoding | Downlink grows with visible tiles; needs simulcast/SVC | Company (egress bandwidth) |
| MCU | One downstream per client; works for weak/legacy clients | ~1 core per output stream; +50–200 ms; fixed layout | Company (CPU), participants (latency) |
| Hybrid: SFU + mixing for audio-only and dial-in | SFU for clients; mixed audio for phones and room systems | Two media paths to operate | Platform team (complexity) |
Staff default: "SFU for everything with three or more participants. Peer-to-peer for 1:1 when ICE finds a direct path — it's cheaper and lower latency — migrating to an SFU without user-visible interruption when a third person joins. A mixing gateway only at the edges for phone dial-in and legacy room systems. No general-purpose MCU."
When to deviate:
- Very low-end clients (old devices, constrained networks) in a specific market: a server-side composite for those viewers only can beat decoding 9 tiles.
- Privacy-sensitive 1:1 (e.g. telehealth) may prefer always-SFU so IP addresses aren't exposed to the other party.
🧭 Principal Move: "Peer-to-peer 1:1 saves egress — if 40% of our minutes are 1:1 and half connect directly, that's ~20% of the bill. But it also removes server-side quality telemetry for those calls. I'd make that trade explicitly, with client-side telemetry to compensate."
❌ Common L5 Trap: "Use an MCU so each client only receives one stream — that's the most bandwidth-efficient." It is for the client, but every viewer needs their own composite (different layouts, not seeing themselves), so the server burns a core per participant and adds 50–200 ms of decode-mix-encode latency. At 3 million concurrent participants, that's millions of cores.
3.2 Fault Line 2: Simulcast vs SVC, and Who Adapts#
The tension: Per-subscriber adaptation needs multiple qualities of each sender. Simulcast sends independent encodes — simple, universally supported, but costs more uplink and switching needs a keyframe on the target layer. SVC sends one layered stream — more efficient and switchable on any frame at temporal layers, but needs codec support (VP9, AV1) and a smarter SFU.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Single stream, sender adapts | Simplest; minimum uplink | Sender adapts to the weakest receiver → everyone degraded | Participants on good networks |
| Simulcast (2–3 layers) | Broad codec support; SFU just picks a stream | ~1.5–2× uplink; layer switch needs keyframe | Senders (uplink, CPU) |
| SVC (spatial + temporal layers) | ~20–40% less uplink than simulcast (typical); smooth temporal switching | Codec and hardware-encoder support; complex SFU | Platform team (SFU logic) |
| Server transcoding per viewer | Perfect fit for each viewer | CPU per output; added latency | Company (CPU) |
Staff default: "Simulcast with three layers as the baseline because every browser and hardware encoder supports it. SVC with VP9 or AV1 where both client and SFU support it, especially for screen share and large meetings. The SFU adapts per subscriber; the sender adapts only its top layer to its own uplink, and pauses layers nobody is subscribed to."
When to deviate:
- Screen share: prioritize resolution and legibility over frame rate — 1–5 fps at full resolution beats 30 fps at 360p.
- Mobile senders on cellular: two layers, not three, to save uplink and battery.
🎯 Staff Insight: "Pausing unsubscribed layers is a big win nobody mentions. In a 12-person meeting viewing only the speaker in high resolution, eleven senders can stop encoding 720p — less uplink, less CPU, less battery."
❌ Common L5 Trap: "WebRTC's congestion control will adapt the bitrate automatically." Congestion control adapts one sender to one path. In a group call with a single stream per sender, adapting to the weakest subscriber drags everyone down; ignoring them drowns that subscriber. Adaptation in a group needs layers and per-subscriber selection.
3.3 Fault Line 3: Single Media Server per Meeting vs Regional Cascading#
The tension: Putting every participant of a meeting on one SFU is simple: one place for state, no inter-server forwarding. But participants spread across continents pay long-haul latency on every hop, and a single server caps meeting size by its NIC. Cascading puts each participant on a nearby SFU and links the SFUs — better latency and scale, more moving parts.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| One SFU per meeting, region of first joiner | Simple state; no cascade logic | Remote participants pay 100–300 ms extra; size capped by one server | Remote participants (latency) |
| One SFU, region by majority / centroid | Better average | Still bad for the minority; re-homing mid-meeting is disruptive | Minority-region participants |
| Cascade: nearest SFU per participant, linked | Short last mile for all; scales beyond one server | Inter-SFU links; cascade topology management | Platform (backbone cost, complexity) |
| Anycast edge with per-track trees | No region choice; distribution adapts per stream | Requires owning a large edge network | Company (network investment) |
Staff default: "Nearest SFU per participant, chosen by a latency probe to two or three candidate regions at join. SFUs for the same meeting cascade, forwarding only layers that remote subscribers need. Full mesh between cascade members up to ~4 regions; beyond that, a hub-and-spoke cascade through one or two well-connected hubs. Within a region, a meeting larger than one SFU's budget spreads across several SFUs in the same cascade."
When to deviate:
- Single-region meetings (the large majority): one SFU, no cascade.
- Regulated meetings that must keep media in a jurisdiction: pin all SFUs to that region and accept the latency.
🧭 Principal Move: "Cascading shifts cost from users' last mile to our backbone. I'd price inter-region transfer per participant-minute by region pair and decide where to build presence versus where to accept longer last-mile paths — that's a network investment decision, not just a latency one."
❌ Common L5 Trap: "Put the meeting in the region closest to the host." The host's location says nothing about where the other 40 participants are. Half the meeting can end up two oceans away from the server, with every packet crossing twice.
3.4 Fault Line 4: Loss and Jitter — Retransmit vs Redundancy vs Degrade#
The tension: UDP packets get lost and arrive unevenly. You can ask for lost packets again (NACK — costs a round trip), send redundant data so receivers can rebuild them (FEC/RED — costs bandwidth all the time), buffer to smooth arrival (jitter buffer — costs latency), or send less (lower layer — costs quality). All of it must fit inside ~150–300 ms.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| NACK + retransmission | Efficient: only lost packets resent | Useless when RTT approaches the latency budget | Participants on long paths |
| FEC / RED | No round trip; good for bursty loss | 10–50% bandwidth overhead; worsens congestion-caused loss | Everyone's bandwidth |
| Larger jitter buffer | Smooth playback | Adds directly to mouth-to-ear latency | Conversation quality |
| Downgrade layer | Reduces congestion-caused loss at the source | Lower resolution | The affected subscriber |
| Combination by media type and RTT | Each tool used where it's cheap | Tuning complexity | Platform team |
Staff default: "Audio gets Opus in-band FEC and concealment always, and RED above ~5% loss — audio is tiny, so redundancy is cheap insurance. Video gets NACK served from the SFU's packet cache for that subscriber; because the SFU is in the subscriber's region, the round trip is tens of milliseconds even if the publisher is far away. Persistent loss is treated as congestion and answered with a lower layer, not more redundancy. Jitter buffers are adaptive — 20–40 ms on good paths, growing toward 150–200 ms only on bad ones."
When to deviate:
- Cascade links with measurable loss: add FEC on the inter-SFU leg, where bandwidth is ours and a round trip is long.
- Screen share: tolerate more latency for legibility; NACK aggressively.
🎯 Staff Insight: "Retransmissions should come from the server nearest the receiver. A NACK that travels back to a publisher in another continent arrives after the frame was due."
❌ Common L5 Trap: "Use TCP so nothing is lost." TCP's in-order delivery means one lost packet stalls everything behind it until retransmission arrives — a head-of-line freeze of hundreds of milliseconds. Real-time media prefers losing a packet to waiting for it.
3.5 Fault Line 5: Interactive Meeting vs Broadcast for Large Audiences#
The tension: Sub-second latency for everyone is expensive: SFU egress per viewer, cascade fan-out, per-viewer adaptation. Streaming over a CDN (HLS/DASH, including low-latency variants) is far cheaper per viewer and scales effortlessly, but adds seconds of latency — fine for watching, bad for conversation.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Everyone in one interactive meeting | Anyone can speak instantly | Egress and client decode explode; join storms hit SFUs | Company (cost), participants (quality) |
| Last-N interactive (up to ~1,000) | Bounded forwarding; still interactive | Thumbnails for non-speakers; complex allocation | Platform team |
| Webinar via SFU fan-out tree | ~0.5–1.5 s latency; viewers can be promoted | SFU egress per viewer; tree management | Company (egress) |
| Webinar via LL-HLS on CDN | Cheapest per viewer; CDN scale | 3–10 s latency; promotion to speaker needs a mode switch | Viewers (latency) |
Staff default: "Interactive meetings up to ~1,000 with last-N video. Above that, webinar mode: panelists and promoted questioners in a small interactive meeting; the audience on an SFU fan-out tree when the host needs sub-second latency (live Q&A, auctions), or on low-latency HLS through the CDN when 3–10 seconds is fine, which is most town halls. A viewer promoted to ask a question moves from the broadcast path into the panel meeting and back."
When to deviate:
- Interactive large events (classrooms with breakouts, hackathons): stay on SFUs; use breakout rooms rather than one giant room.
- Very large broadcasts (100K+): CDN streaming only; interactivity via chat and polls.
🧭 Principal Move: "Webinars are a separate SKU: a scheduled capacity reservation, a different latency tier and a different price. I'd never let a 50,000-viewer event discover our SFU capacity limits at 10:00:00 on the day."
❌ Common L5 Trap: "Webinars are just big meetings — add more SFUs." Treating 20,000 passive viewers as meeting participants pays interactive-grade egress and per-viewer adaptation for people who will never speak, and puts a 20,000-join storm on the meeting control plane.
4. When It Breaks#
4.1 The Top-of-Hour Join Storm#
t=09:59:30: 180K concurrent meetings in region eu-west.
t=10:00:00: 140K meetings start or restart within 90 seconds. 900K joins.
t=+20s: Signalling p99 join: 400ms → 6s. Meeting coordinator's state store at 95% CPU.
t=+40s: Clients time out at 10s and retry immediately. Join rate doubles.
t=+60s: SFU allocator hands the same least-loaded SFUs to thousands of new meetings
(stale load data, refreshed every 10s). 40 SFUs hit NIC saturation.
t=+2min: Join success rate 99.6% → 81%. Status page: "some users unable to join".
t=+8min: Joins back-pressure; retry storm subsides; success recovers.
Detection: join.success_rate{region}, join.latency_p99, signalling.retry_ratio, sfu.allocations_per_sfu_per_10s.
Mitigation: client retries with jittered exponential backoff (1s, 2s, 4s ± 50%); allocator uses power-of-two-choices on fresh-ish load plus per-SFU allocation rate limits so one SFU can't receive hundreds of meetings in one refresh window.
Prevention: pre-scale signalling and coordinator for :00 and :30 by calendar; capacity reservation for scheduled large meetings; load tests that replay a top-of-hour curve.
Owner: conferencing platform on-call (control plane).
4.2 Robot Voice in Brazil — An ISP Path Goes Bad#
t=0: A peering link between a large Brazilian ISP and our São Paulo region degrades: 7% loss, 80ms jitter.
t=+1min: Audio concealment ratio for that ISP's users: 1% → 18%. Users hear "robot voice".
t=+3min: Video BWE for those subscribers drops; SFU downgrades them to 180p. Video OK-ish.
t=+10min: Support tickets mention one ISP. Server metrics all green.
t=+25min: Network on-call shifts that ISP's traffic to a transit provider. Concealment back to 2%.
Detection: audio.concealment_ratio{region, asn}, rtt_p95{asn}, loss_rate{asn} — quality broken down by network, not just region.
Mitigation: traffic engineering away from the bad path; enable RED for affected subscribers automatically when loss exceeds 5%.
Prevention: per-ASN quality dashboards with anomaly detection; multiple upstreams per region; automatic RED activation.
Owner: network engineering (paths), conferencing platform (adaptive audio).
4.3 The Bank That Blocks UDP#
t=0: A large customer's network team rolls out a new firewall policy: UDP egress blocked, TLS
inspection on 443.
t=+1h: 12,000 employees' meetings: ICE fails on UDP, falls back to TURN-TCP — blocked except 443.
t=+1h: TURN-TLS on 443 works, but TLS inspection terminates and re-originates TLS, adding 60-120ms
and breaking connections every few minutes.
t=+1h: Join success for that customer: 99% → 70%. Ticket: "your service is down for us".
Detection: ice.selected_candidate_type{customer} shift to relay-tls; turn.tls_session_resets; join success by customer.
Mitigation: work with the customer's network team to allowlist media IP ranges and the TURN hostnames from inspection; publish the IP ranges and ports in admin docs.
Prevention: TURN over TLS on 443 always available; published network requirements; an admin connectivity test page; proactive outreach when a large customer's candidate types shift.
Owner: conferencing platform (TURN), customer success (enterprise network requirements).
4.4 The All-Hands Pinned to One SFU#
t=0: A 2,800-person all-hands starts as a normal meeting (not webinar mode). The allocator
places it on one SFU in us-east; cascade added only for other regions.
t=+3min: 1,900 participants on that SFU. Egress: 1,900 × ~2 Mbps = 3.8 Gbps, near NIC limit.
Packet rate: ~600K pps; SFU event loop saturates.
t=+5min: Loss rises for everyone on that SFU; BWE drops; everyone downgraded to 180p.
t=+7min: Audio glitches. CEO's microphone: concealment 12%.
Detection: sfu.egress_gbps / nic_gbps > 0.7, sfu.packet_rate, meeting.participants_per_sfu.
Mitigation: spread the meeting across multiple SFUs in the region (intra-region cascade); switch non-speaking participants to last-N video.
Prevention: per-SFU participant and bandwidth budgets per meeting — when a meeting exceeds ~300 participants on one SFU, the coordinator adds another SFU in the same region; meetings scheduled with more than 500 invitees default to webinar mode or get a capacity reservation.
Owner: conferencing platform (allocator and coordinator).
4.5 The Keyframe Storm#
A popular presenter's stream has 1,200 subscribers across 6 SFUs. A brief uplink glitch causes loss for all subscribers at once; each SFU sees hundreds of keyframe requests (PLIs) and forwards them. The publisher's encoder produces a keyframe every 200 ms — keyframes are 5–10× larger than delta frames — saturating the publisher's uplink, which causes more loss, which causes more PLIs. Video for the presenter freezes for 40 seconds.
Detection: pli.rate{publisher}, keyframe.interval{publisher}, publisher uplink BWE collapse.
Mitigation and prevention: SFUs aggregate PLIs per publisher layer and forward at most one per second; serve recovering subscribers by switching them briefly to a lower layer with a recent keyframe; use long-term reference frames or periodic keyframes on low layers.
Owner: conferencing platform (SFU).
4.6 The Recording That Wasn't#
A board meeting's recorder — a headless subscriber writing a composited file — crashed at minute 52 of 60 due to a memory leak on long meetings. The output file was written only at the end, so the whole recording was lost. Nobody was alerted; the host found out the next day.
Detection: recorder.heartbeat_missing, recording.segments_written{recording_id} stalled; host-visible "recording" indicator tied to segment writes, not process state.
Prevention: record per-track segments (every 5–10 s) to object storage as the meeting runs; composite after the meeting; run a standby recorder for meetings flagged important; notify the host in-meeting if recording stalls for more than 30 s.
Owner: media services team (recording).
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Top-of-hour join storm | join.success_rate, signalling.retry_ratio | Region's new joins | Backoff, allocator rate limits, pre-scale | Platform on-call |
| Bad ISP path | audio.concealment_ratio{asn} | That ISP's users | Traffic engineering, auto-RED | Network eng |
| UDP blocked / TLS inspection | ice.selected_candidate_type shift | One customer | TURN TLS 443, allowlisting | Platform + customer success |
| Meeting outgrows one SFU | sfu.egress / nic > 0.7 | That meeting | Intra-region cascade, last-N | Platform |
| Keyframe storm | pli.rate{publisher} | One publisher's viewers | PLI aggregation, low-layer switch | Platform (SFU) |
| Recorder crash | recording.segments_written stalled | One recording | Segmented recording, standby | Media services |
| SFU host failure | sfu.heartbeat_missing | ~1–2K participants | ICE restart to new SFU < 3 s | Platform on-call |
| Cascade link loss | cascade.loss_rate{region_pair} | Cross-region meetings | FEC on link, reroute via hub | Network eng |
🎯 Staff Insight: Conferencing failures are mostly quality failures, not outages. "Server dashboards will be green while users hear robot voice. I'd break quality metrics down by region, ASN, client version and customer, because that's where bad calls hide."
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | "Connect participants with WebRTC" | Names heterogeneity, the latency budget and egress cost as the design drivers; separates meetings from webinars | Asks whether conferencing is a product or an embedded capability and what cost per participant-minute is affordable |
| Topology | SFU | P2P for 1:1, SFU for meetings, mixing only at edges; reasons from meeting-size distribution | Prices topologies per participant-minute; decides P2P trade of cost vs telemetry |
| Adaptation | "WebRTC adapts" | Simulcast/SVC, per-subscriber layer selection from BWE and layout, audio priority, paused layers | Quality SLOs (freeze, concealment) as shared product and infra metrics |
| Placement | One region per meeting | Nearest SFU via latency probe, cascading with bounded topology, intra-region spread | Network investment: peering, backbone, edge build vs buy |
| Resilience | Reconnect on failure | ICE restart to new SFU < 3 s, TURN TLS 443, PLI aggregation, join backoff | Cell-based regional blast radius, reserved failover capacity, game days |
| Operations | CPU and uptime | Quality by region, ASN, client and customer; join success; segmented recording | Webinars as a scheduled SKU with capacity reservations; quality reviews with large customers |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Decouples participants | "One bad network lowers one subscriber's layer, not every sender's bitrate." |
| Places by participants | "Each participant connects to the nearest SFU; SFUs cascade one copy per stream." |
| Respects the latency budget | "Retransmit from the server nearest the receiver; past ~100 ms RTT, use FEC or a lower layer." |
| Prioritizes audio | "Audio is allocated first, always. Frozen video is annoying; broken audio ends the meeting." |
| Separates broadcast | "Above a thousand passive viewers, it's a webinar on a different path." |
| Measures quality | "I page on concealment and freeze ratios, not SFU CPU." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| MCU for all meetings | CPU per viewer and added latency; doesn't scale economically |
| Single stream per sender | Weakest participant degrades everyone |
| TCP for media | Head-of-line blocking freezes media on any loss |
| One region per meeting, chosen by host | Remote participants blow the latency budget |
| "STUN is enough" | 10–20% of users can't connect without TURN; enterprises block UDP |
| Webinars as big meetings | Interactive-grade cost and join storms for passive viewers |
5.4 Common False Positives#
- Deep SDP and ICE knowledge ≠ conferencing design. Connection setup is 2 seconds of the call; adaptation and placement are the other 59 minutes.
- Codec expertise ≠ system design. Knowing VP9 profiles doesn't answer what a weak subscriber receives.
- "We use WebRTC" ≠ architecture. WebRTC is a toolkit; topology, placement and allocation are your design.
- Big-machine SFUs ≠ large-meeting design. NIC and packet rate saturate first, and clients can't decode 49 tiles anyway.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Meeting sizes, latency budget, heterogeneity, egress cost; numbers |
| Entities & API | 3–5 min | Meeting, participant, track, layers, subscriptions; signalling messages |
| Architecture | 5–10 min | ≤ 8 boxes; signalling, coordinator, allocator, SFU, TURN, cascade |
| Adaptation | 10–18 min | Simulcast/SVC, BWE, per-subscriber allocation, audio priority, paused layers |
| Placement + cascade | 18–25 min | Nearest SFU, latency probes, cascade topology, intra-region spread |
| Loss, jitter, TURN | 25–32 min | NACK from nearest server, FEC for audio, layer downgrade, TURN TLS 443 |
| Large meetings + webinars | 32–38 min | Last-N, fan-out trees, LL-HLS, join storms, recording |
| Pivot / wrap | 38–45 min | Failure, E2EE, cost; close on quality SLOs |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Add end-to-end encryption" | Tradeoff with server features | Insertable-streams E2EE; SFU still forwards; recording, transcription and dial-in disabled or client-side |
| "An SFU host dies mid-meeting" | Recovery path | Clients detect within 1–2 s, signalling assigns new SFU, ICE restart, audio back in < 3 s |
| "Record every meeting" | Cost and reliability | Per-track segments to object storage; composite later; storage tiering |
| "Support 50,000-viewer town halls" | Broadcast separation | Webinar mode, LL-HLS via CDN, capacity reservation, staggered join |
| "Reduce cost by 30%" | Egress awareness | P2P for direct 1:1, paused layers, last-N, lower default gallery resolution |
| "Calls are bad in one country" | Diagnosis | Quality by ASN, path, client version; traffic engineering; local presence |
6.3 What to Deliberately Skip#
- SDP syntax and ICE candidate types — one sentence.
- Codec internals — name them, move on.
- Chat, reactions, hand-raise — signalling messages; one line.
- UI layouts — only as input to subscription sizes.
- DTLS handshake details — "keys are negotiated per connection."
6.4 Follow-Up Questions to Expect#
- "One participant has a 1 Mbps connection with 6% loss. What do the other 11 participants see?"
- "Participants are in Sydney, Frankfurt and Virginia. Where's the media server?"
- "A customer's network blocks UDP. Does the call work, and at what cost?"
- "It's a 2,500-person all-hands. How many streams does each server forward?"
- "An SFU crashes. How long until everyone hears each other again?"
- "RTT is 250 ms and loss is 4%. NACK, FEC or something else?"
- "How do you record a meeting so a crash at minute 52 loses nothing?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design a video conferencing service like the ones you use every day."
Staff Answer
"First, the meeting-size distribution and the largest case: are we designing 1:1 calls, team meetings, all-hands, or 50,000-viewer webinars? And where are users? I'll assume a global service with 3 million concurrent participants at peak, about 500K concurrent meetings, a median meeting of 4, a tail to 1,000 interactive participants, and a separate webinar mode beyond that.
Three constraints shape everything. Latency: mouth-to-ear under 300 ms, ideally near 150, so no buffering and limited retransmission. Heterogeneity: one participant's bad network must not degrade anyone else. Cost: media egress dominates everything. I'll go: entities and signalling → architecture with SFUs → per-subscriber adaptation → placement and cascading → loss, jitter and TURN → large meetings and webinars → failure handling and quality metrics."
Why this is L6:
- Asks for the size distribution and geography before choosing topology
- Names latency, heterogeneity and egress as the drivers, with numbers
- Separates webinars from meetings up front
What L7 adds:
- Asks for the affordable cost per participant-minute and who owns call quality
- Asks whether the network edge is something we own or buy
❌ Common L5 Trap
"Clients use WebRTC. A signalling server exchanges SDP offers and answers over WebSocket, STUN handles NAT traversal, and media goes through an SFU. We scale SFUs horizontally."
Why this misses: Correct components, no design. It doesn't say what a weak participant receives, where the SFU is for a global meeting, what happens when UDP is blocked, or how a 2,000-person meeting differs from a 5-person one.
Drill 2: The Weak Participant#
Prompt: "In a 12-person meeting, one participant joins from a train: 1 Mbps down, 6% packet loss. What happens to their experience, and to everyone else's?"
Staff Answer
"Everyone publishes three simulcast layers. The SFU's bandwidth estimate for the train participant settles around 0.8–1 Mbps. Their allocation: audio for the current speakers first (~100 kbps), the active speaker at 360p (~500 kbps), the rest of the gallery at 180p or paused to audio-only avatars if it doesn't fit. Loss: their audio uses Opus FEC, and with 6% loss the SFU adds RED for them; video losses are recovered by NACK from the SFU's cache — the SFU is near them, so the round trip is short — and if loss persists, the SFU holds them at the lower layer.
Everyone else sees no change. Senders keep encoding all three layers because other subscribers want 720p. The train participant's keyframe requests are aggregated with others' and rate-limited, so they don't force the senders into constant keyframes."
Why this is L6:
- Per-subscriber allocation with audio first and concrete bitrates
- Loss handling split by media type, using the SFU's proximity for NACK
- Explicitly protects other participants, including from keyframe requests
What L7 adds:
- Tracks this as a quality metric (freeze and concealment per participant-minute) by network type
- Considers an audio-only "low bandwidth mode" as a product feature users can choose
❌ Common L5 Trap
"WebRTC's congestion control will lower the bitrate so everyone can keep up."
Why this misses: With one stream per sender, the only way to "keep up" is for every sender to drop to the train participant's level, degrading 11 people for one. And their keyframe requests will hit every sender.
Drill 3: Make It Concrete — Size the SFU Fleet and the Bill#
Prompt: "Size the media fleet for 3 million concurrent participants, and estimate the bandwidth cost."
Staff Answer
"Egress dominates. Assume each participant receives ~2 Mbps on average (speaker plus a few tiles, many on audio-only or small galleries) and sends ~1.5 Mbps (simulcast with paused top layers). Egress: 3M × 2 Mbps = 6 Tbps at peak. Ingress ~4.5 Tbps, plus cascade traffic — say 10% extra.
An SFU host with a 25 Gbps NIC, run at ~50% for headroom and packet-rate limits, forwards ~12 Gbps → ~500 hosts at peak, plus N+1 per cell and failover reserve — ~800 across ~20 regions. TURN: 15% of sessions relay, so ~1 Tbps through TURN — another ~100 hosts.
Cost: peak-to-average of ~3× means ~2 Tbps average egress — ~250 GB/s, roughly 650 PB a month. At an illustrative $0.002–0.01 per GB for a large network with peering, that's ~$1.3–6.5M a month — which is why peering, P2P for 1:1, last-N and paused layers matter more than SFU CPU."
Why this is L6:
- Sizes by egress and NIC, not participants per CPU
- Includes TURN and cascade traffic
- Turns bandwidth into dollars and identifies the levers
What L7 adds:
- Computes cost per participant-minute and compares it to pricing
- Frames peering and edge presence as the main investment decision
❌ Common L5 Trap
"Each server handles about 1,000 participants, so 3,000 servers."
Why this misses: Participants per server is meaningless without bitrate and layout; the real limits are NIC bandwidth and packet rate. And the answer ignores the bill, which is the main constraint.
Drill 4: The Global Meeting#
Prompt: "A meeting has 5 people in Sydney, 6 in Frankfurt and 3 in Virginia. Design the media path."
Staff Answer
"Each participant connects to the SFU in their nearest region, chosen by latency probes at join. The meeting coordinator records three SFUs in the cascade. Each SFU forwards each local publisher's needed layers once to each other SFU — Sydney's 5 publishers send 5 streams to Frankfurt, not 5 × 6. Each SFU then does per-subscriber selection for its local participants.
Latency: Sydney–Frankfurt is the long hop, ~130–150 ms one way over a good backbone. A Frankfurt participant hears a Sydney speaker at roughly 20 ms last mile + 140 ms backbone + 20 ms last mile + ~60 ms encode, jitter buffer and decode ≈ 240 ms — inside the budget. With one SFU in Virginia, Sydney participants would pay the Pacific twice for every conversation with each other — 400+ ms between two people in the same city."
Why this is L6:
- Nearest SFU with cascade, one copy per stream per link
- Adds up the latency budget hop by hop
- Shows why a single-server design hurts even co-located participants
What L7 adds:
- Prices backbone transfer per region pair and decides where to add presence
- Bounds cascade topology (hub-and-spoke past 4 regions) as an org-wide standard
❌ Common L5 Trap
"Put the meeting on a server in the region of whoever created it."
Why this misses: The creator's region is arbitrary. Two Sydney participants talking to each other through Virginia pay a round trip across the Pacific and back for every word.
Drill 5: The SFU Crashes#
Prompt: "An SFU host serving 1,500 participants across 200 meetings dies. What happens?"
Staff Answer
"Clients detect it fast: no RTP or RTCP for ~1–2 seconds and a failed ICE consent check. The signalling connection is separate and survives. The meeting coordinator also sees the SFU's heartbeat lapse in its registry. For each affected meeting, the coordinator assigns a replacement SFU in the same region — spreading the 200 meetings across many SFUs, not one — and signalling tells clients to ICE-restart against the new SFU. Clients republish their tracks; subscriptions are rebuilt from the coordinator's state.
Target: audio back within ~3 seconds, video within ~5 after keyframes. Remote cascade members see the link drop and re-link to the new SFU. The join path is rate-limited and jittered so 1,500 reconnections don't look like a storm. If a whole region fails, clients fall back to the next-nearest region with reserved capacity."
Why this is L6:
- Separate signalling and media failure domains
- Spreads replacements to avoid a secondary hotspot
- Quantified recovery targets with reconnection pacing
What L7 adds:
- Cell-based placement so one host failure is a known, bounded blast radius
- Regular game days that kill SFUs and regions under load
❌ Common L5 Trap
"Run SFUs in active-passive pairs and fail over to the standby."
Why this misses: Media state is per-packet and per-subscriber — there's nothing useful to replicate in real time, and a hot standby doubles cost. Fast reattachment to any healthy SFU is cheaper and recovers as quickly.
Drill 6: The 2,500-Person All-Hands#
Prompt: "The CEO wants a 2,500-person all-hands where anyone can ask a question live. Design it."
Staff Answer
"This is a webinar with promotion, not a 2,500-way meeting. Three presenters plus a moderator are panelists in a small interactive meeting. The audience subscribes only. Because they want live Q&A with low lag, I'd use an SFU fan-out tree rather than streaming: the panel's SFU forwards to regional SFUs, which serve local viewers — ~2,500 × ~1.5 Mbps ≈ 4 Gbps of egress spread across regions.
A participant who raises their hand is promoted: their client starts publishing to the nearest SFU and joins the panel meeting; the moderator demotes them after. Join storm: the meeting is scheduled, so the coordinator reserves SFU capacity 30 minutes ahead and staggers the waiting room. If the company later wants 50,000 viewers, I'd move the audience to low-latency HLS on the CDN and keep Q&A via promotion."
Why this is L6:
- Reframes as panel + audience, with promotion for Q&A
- Fan-out tree sized in bandwidth
- Capacity reservation for a scheduled event
What L7 adds:
- Offers large events as a separate tier with pricing and a dry-run process
- Sets the boundary between SFU tree and CDN by cost per viewer-hour
❌ Common L5 Trap
"Put everyone in one meeting on a large SFU with gallery view."
Why this misses: 2,500 publishers and subscribers on one server saturates its NIC and packet rate; clients can't decode the gallery; and 2,500 people joining at 10:00 hits one allocation decision.
Drill 7: Build vs Buy#
Prompt: "We're a telehealth company. Should we build our own video infrastructure?"
Staff Answer
"Almost certainly not. Our product is the clinical workflow — scheduling, records, prescriptions — and video is a 1:1 or small-group feature. A real-time media platform or a managed SFU service gives us global presence, TURN, SDKs and quality tooling we'd need years to match. What stays ours: identity and access control for rooms, compliance agreements, recording policy, data residency choices, and the UX.
I'd evaluate vendors on regional presence where our patients are, TURN on 443, recording controls and storage location, compliance terms, quality telemetry access, and pricing per participant-minute at our 3-year volume. Building makes sense only if video is the product or volume makes vendor pricing exceed a team plus the network bill — which for telehealth it won't."
Why this is L6:
- Separates differentiating workflow from commodity media infrastructure
- Lists what stays in-house regardless
- Concrete vendor evaluation criteria, including compliance
What L7 adds:
- Requires an exit plan: SDK abstraction so the vendor is swappable
- Prices vendor cost at 3-year volume and negotiates residency commitments
❌ Common L5 Trap
"There are good open-source SFUs; we can deploy one in the cloud and save on vendor costs."
Why this misses: The SFU is the easy part. Regional presence, TURN, quality monitoring, client SDKs across browsers and devices, and 24/7 operations are the real cost, and none of them is a telehealth company's advantage.
Drill 8: Rolling Out a New Congestion Controller Without an Outage#
Prompt: "Your team wrote a new bandwidth estimator that should improve quality on mobile. How do you ship it?"
Staff Answer
"Quality changes can't be validated by unit tests or server metrics, so the rollout is an experiment. The estimator runs in the SFU per subscriber, so I can assign it by subscriber. Step one: shadow mode — compute both estimates, use the old one, log the difference. Step two: 1% of subscribers in two regions, randomized by participant, comparing freeze ratio, delivered resolution, concealment and the rate of layer switches against control. Step three: ramp 5 → 25 → 50 → 100% over a couple of weeks, with automatic rollback if freeze ratio worsens by more than a set margin on any client platform.
I'd slice results by network type, client version and region — an estimator can win on average and lose badly on one carrier. The switch is a per-SFU config flag, so rollback takes seconds."
Why this is L6:
- Shadow mode, randomized experiment, quality metrics as the gate
- Slicing by platform and network to catch regressions hidden by averages
- Fast, config-driven rollback
What L7 adds:
- Establishes a standing quality experimentation platform used for all media changes
- Requires product sign-off on the quality metrics that define "better"
❌ Common L5 Trap
"Test it in staging with network emulation, then deploy to all SFUs."
Why this misses: Lab emulation doesn't capture real carriers, Wi-Fi and devices. Deploying everywhere at once risks degrading millions of calls, with no control group to tell whether quality got better or worse.
Drill 9: Cutting Cost by 30%#
Prompt: "Finance wants the media bill down 30% without users noticing. Where do you look?"
Staff Answer
"Egress is the bill, so I attack bytes forwarded. First, layers nobody watches: make sure senders pause layers with no subscribers — common in meetings where everyone looks at the speaker. Second, gallery defaults: thumbnails at 180p instead of 360p on small screens — users rarely notice. Third, P2P for 1:1 calls with a direct path: if 40% of minutes are 1:1 and half connect directly, that's ~20% of egress gone. Fourth, last-N video for meetings over ~25, and video off for participants who haven't spoken in 10 minutes in very large meetings. Fifth, codec: AV1 or VP9 SVC on capable clients cuts bitrate 20–40% at the same quality.
Then network: shift traffic to settlement-free peering where possible. Each change is A/B tested against freeze ratio and resolution delivered, so 'without users noticing' is measured, not assumed."
Why this is L6:
- Targets egress with specific, quantified levers
- Recognizes P2P and paused layers as large wins
- Gates every change on quality metrics
What L7 adds:
- Builds cost per participant-minute into product decisions (default layouts, features)
- Negotiates peering and transit as a strategic program
❌ Common L5 Trap
"Use cheaper instances and pack more meetings per SFU."
Why this misses: Compute is a small fraction of the bill. Packing harder raises loss and lowers quality without touching the bytes that cost money.
Drill 10: Multi-Region and Residency#
Prompt: "A European customer requires that their meetings' media and recordings never leave the EU. How do you support that?"
Staff Answer
"Residency is a tenant attribute applied at allocation. For that tenant's meetings, the coordinator only places SFUs, TURN servers and recorders in EU regions — even for participants travelling elsewhere, who connect to the nearest EU SFU and accept extra latency. Cascades stay within EU regions. Recordings write to EU object storage; transcription and other media services run in the EU.
Signalling metadata — who joined, when — is also customer data; it stays in an EU control plane. If the EU region fails, meetings fail over to another EU region with reserved capacity, never outside. I'd prove it with marker tests: synthetic EU-tenant meetings whose media and recordings are traced, plus scans of non-EU stores for their identifiers."
Why this is L6:
- Applies residency at allocation, including TURN and recording
- Accepts the latency cost for travelling participants explicitly
- Failover stays in-region; proof by tests
What L7 adds:
- Defines residency as a company-wide tenant attribute honored by every platform
- Prices the reserved EU failover capacity into the enterprise plan
❌ Common L5 Trap
"Route EU users to EU servers using geo-DNS."
Why this misses: Geo-DNS routes by where the user is, not by the tenant's contract — a travelling employee or a non-EU participant would pull the meeting's media outside the EU, and TURN, recordings and failover aren't covered at all.
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — Monday 09:00 in Europe#
Context: At 09:00 CET on the first Monday after a holiday, joins in eu-central are 1.6× the previous peak. Join success drops to 88%, and participants who do join report frozen video. The on-call escalates to you.
Questions to Surface First:
- Is the failure in signalling, allocation, ICE connectivity or media forwarding?
- Are SFUs saturated on NIC or packet rate, or is the allocator concentrating meetings?
- Is TURN carrying more than usual (a network change somewhere)?
- Is the frozen video tied to particular SFUs, ASNs or client versions?
Typical L5 Approach: Scales the SFU fleet. New SFUs take minutes to boot and register; the allocator keeps sending meetings to the same hot hosts in the meantime.
Staff Approach: Finds the allocator using 10-second-old load data and picking the single least-loaded SFU, so each refresh window dumps hundreds of meetings onto a few hosts, saturating their NICs. Switches the allocator to power-of-two-choices with a per-SFU allocation rate cap, drains the hottest SFUs by moving new meetings elsewhere, and enables last-N for meetings over 25 on saturated hosts.
Principal Approach: Treats Monday-after-holiday as a forecastable peak: calendar-driven pre-scaling, a capacity model per region based on the meeting-start curve, and an allocator design reviewed for herd behavior as part of the platform's resilience standards.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Join funnel by stage: signalling OK, allocation OK, ICE OK, first media slow. SFU egress: 30 hosts at > 90% NIC, 400 hosts under 40%. |
| Triage | Allocator picks min-load host from a 10 s snapshot; hundreds of meetings per window land on the same hosts. |
| Quick fix | Power-of-two-choices; cap 20 new meetings per SFU per 10 s; prefer spreading over packing during surges. |
| Guardrails | Alert on load skew (p99/p50 SFU egress > 3); pre-scale by calendar. |
| Post-mortem | Load data freshness vs allocation rate is a design parameter, now documented. |
Metrics to Watch: join.success_rate, sfu.egress_gbps distribution, sfu.allocations_per_10s, participant.freeze_ratio{sfu}
Organizational Follow-up: capacity planning adds holiday-return Mondays to the peak calendar.
Ownership Question: "Who owns allocation fairness?" Staff answer: The conferencing platform team owns the allocator and its load-skew SLO; capacity planning owns the forecast it runs against.
Key Takeaway: "Least-loaded with stale data is most-loaded with fresh data. Allocators need randomness and rate caps."
What clears the Staff bar:
- Localizes the failure in the join funnel before scaling
- Recognizes herd behavior in the allocator
- Uses last-N as a degradation lever while fixing allocation
Deep Dive 2: Silent Failure — Audio Concealment Creeping Up for Two Weeks#
Context: A quality review shows audio concealment ratio rose from 1.5% to 4% over two weeks, concentrated in Android clients. No alerts fired; servers are healthy; ticket volume rose slightly.
Questions to Surface First:
- Which client versions, devices and networks? When did the rise start?
- Did anything ship around then — client release, SFU change, codec setting?
- Is the loss on the network, or are packets arriving late and being discarded by the jitter buffer?
- Is it uplink (their audio sounds bad to others) or downlink (others sound bad to them)?
Typical L5 Approach: Assumes network conditions changed and waits for it to improve.
Staff Approach: Slices by client version: a new Android release changed the audio capture thread priority, causing bursty packet sending under CPU load; receivers' jitter buffers discard late bursts as loss. Rolls back the client setting via remote config, confirms concealment recovers within a day, and adds per-version quality regression gates to the release process.
Principal Approach: Establishes quality SLOs per platform with automated regression detection on every client and server release, owned jointly by client and media teams — and makes "quality regression" a release-blocking class of bug.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Slice concealment by platform, client version, ASN, region. Spike isolated to Android client version N. |
| Triage | Uplink-side problem: Android N publishers' packets arrive in bursts; receivers' jitter buffers discard late packets. |
| Quick fix | Remote-config rollback of capture-thread change; staged re-release with fix. |
| Guardrails | Release gate: concealment and freeze ratio per version vs previous version on 1% rollout. |
| Post-mortem | Server-side metrics couldn't see it; quality telemetry from clients is mandatory. |
Metrics to Watch: audio.concealment_ratio{client_version}, packet.interarrival_jitter{publisher_platform}, jitterbuffer.late_discard_rate
Organizational Follow-up: client teams adopt quality metrics as release criteria alongside crash rate.
Ownership Question: "Who owns call quality?" Staff answer: Shared — client teams own capture and playback quality, the media team owns forwarding and adaptation, and one quality dashboard with per-version gates holds both accountable.
Key Takeaway: "Quality regressions are silent unless clients report quality. Slice by version first."
What clears the Staff bar:
- Distinguishes late packets from lost packets
- Slices by version before blaming the network
- Adds quality gates to client releases
Deep Dive 3: Large-Customer Onboarding — 40,000 Employees, One Network#
Context: A multinational with 40,000 employees moves to your service. Their network routes all office traffic through two regional security gateways that inspect TLS and allow little UDP. Their first all-hands is in three weeks.
Questions to Surface First:
- Which ports and protocols leave their network? Is UDP allowed to any destination?
- Will they allowlist our media IP ranges and exempt TURN from TLS inspection?
- How many concurrent meetings and peak participants per office?
- Where are their offices relative to our regions?
Typical L5 Approach: Tells them to open UDP. They can't before the all-hands; everyone relays through TURN-TLS through inspection proxies, with poor quality.
Staff Approach: Runs the connectivity test page from each office; confirms UDP blocked. Negotiates an allowlist for UDP to published media ranges in two offices immediately and TLS-inspection exemption for TURN hostnames elsewhere. Sizes TURN capacity in their regions for 100% relay in the meantime. Runs the all-hands in webinar mode with a capacity reservation and a dry run a week before.
Principal Approach: Productizes enterprise onboarding: published network requirements, a connectivity diagnostic, a dedicated onboarding engineer for accounts over 10,000 seats, and TURN capacity planned per enterprise customer as part of the sales cycle.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Connectivity tests per office: UDP blocked, 443 inspected. Expected relay rate: ~100% initially. |
| Triage | TURN-TLS through inspection adds 60–120 ms and drops sessions periodically. |
| Quick fix | Allowlist UDP to media ranges in two largest offices; exempt TURN from inspection elsewhere. |
| Guardrails | Customer-specific dashboard: candidate types, join success, freeze ratio by office. |
| Post-mortem (pre-mortem) | All-hands in webinar mode with reserved capacity; dry run at 20% scale one week before. |
Metrics to Watch: ice.selected_candidate_type{customer, office}, turn.egress_gbps{region}, join.success_rate{customer}
Organizational Follow-up: customer success owns the network-requirements checklist before go-live.
Ownership Question: "Who owns their call quality?" Staff answer: We own TURN capacity and a clear published requirement; they own their network policy. A joint dashboard makes the boundary visible.
Key Takeaway: "Enterprise video quality is decided by the customer's firewall. Make the requirements explicit and measure them per office."
What clears the Staff bar:
- Tests connectivity before go-live and sizes TURN for the worst case
- Negotiates specific exemptions rather than "open UDP"
- Runs the big event in webinar mode with a dry run
Deep Dive 4: Post-Mortem — 40 Minutes of Frozen Video on One Continent#
Context: For 40 minutes, cross-region meetings involving South America had frozen video and choppy audio. Local meetings in South America were fine. You're leading the post-mortem.
Questions to Surface First:
- Which cascade links were affected? What did loss and RTT look like on them?
- Did the cascade topology route around the bad link? Why not?
- Did subscribers' layer selection react, or did it keep requesting high layers over a lossy link?
- How was it detected — by metrics or by users?
Typical L5 Approach: Blames the backbone provider and files a ticket.
Staff Approach: Finds the São Paulo–Virginia backbone path had 9% loss; the cascade used it directly because the topology was static full mesh, with no health-based rerouting. Remote SFUs' bandwidth estimates for the cascade link weren't used for layer selection, so Virginia kept forwarding 720p into a lossy link. Ships health-aware cascade routing (reroute via a hub when a link's loss exceeds 2%), cascade-link BWE feeding layer selection, and FEC on cascade links above 1% loss.
Principal Approach: Treats inter-region links as a managed dependency with SLOs, multiple diverse paths per region pair, and quarterly game days that degrade a link under load — and makes cascade topology a control-plane decision, not static configuration.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Quality by region pair: only cross-region meetings via São Paulo–Virginia. Link loss 9%. |
| Triage | Static full-mesh cascade; no reroute; cascade link not treated as a constrained subscriber. |
| Quick fix | Manually re-route cascade via a Miami hub; enable FEC on the link. |
| Guardrails | Health-aware cascade routing; per-link BWE driving layer selection; alert on cascade.loss_rate > 2%. |
| Post-mortem | A cascade link is a subscriber with a bandwidth estimate, not a pipe. |
Metrics to Watch: cascade.loss_rate{region_pair}, cascade.rtt{region_pair}, participant.freeze_ratio{cross_region}
Organizational Follow-up: network engineering adds path diversity for the region pair.
Ownership Question: "Who owns a cascade link?" Staff answer: Network engineering owns the path; the conferencing platform owns how media reacts to the path — rerouting, FEC and layer selection.
Key Takeaway: "Treat each cascade link like a subscriber: estimate its bandwidth, adapt what you send, and route around it when it's sick."
What clears the Staff bar:
- Isolates the failure to a region pair
- Fixes the media system's reaction, not just the network
- Makes cascade routing dynamic
Deep Dive 5: Multi-Region Expansion — Launching in India#
Context: Usage in India is growing fast; meetings currently use Singapore SFUs. Users report latency and frequent quality drops on mobile networks. Leadership approves an India region.
Questions to Surface First:
- What's current RTT from major Indian carriers to Singapore? To a candidate Mumbai region?
- How are users connected: mostly mobile? Which carriers? Is peering possible?
- How much traffic is domestic vs international meetings?
- Are there data-residency requirements for Indian customers?
Typical L5 Approach: Deploys SFUs in Mumbai and points Indian users at them by geo-DNS.
Staff Approach: Measures RTT per carrier to Mumbai vs Singapore with client probes before launch — some carriers route Mumbai traffic via Chennai or even Singapore; negotiates peering with the top carriers; deploys SFUs and TURN in Mumbai with probe-based selection rather than geo-DNS; enables two-layer simulcast defaults for mobile senders on constrained networks; cascades to Singapore for international meetings.
Principal Approach: Treats a new region as a network investment with a business case: projected participant-minutes, egress cost reduction from peering, and quality improvement targets; sets launch criteria on measured quality, not server deployment.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Client RTT probes to candidate sites from top 5 carriers; identify carriers with poor routing. |
| Triage | Two carriers route to Mumbai via distant paths; peering needed. |
| Quick fix | Peering at an Indian internet exchange; SFU + TURN in Mumbai; probe-based selection. |
| Guardrails | Per-carrier quality dashboards; launch only when freeze ratio improves by target margin in a 5% rollout. |
| Post-mortem (pre-launch) | Mobile defaults: two simulcast layers, aggressive audio FEC. |
Metrics to Watch: rtt_p50{asn, region}, participant.freeze_ratio{country}, audio.concealment_ratio{asn}
Organizational Follow-up: network team owns carrier relationships; product owns mobile defaults.
Ownership Question: "When is the region 'launched'?" Staff answer: When the quality metrics for Indian participants hit the target in a randomized rollout — not when the servers are up.
Key Takeaway: "A region is a set of network paths, not a set of servers. Measure paths per carrier before you launch."
What clears the Staff bar:
- Measures carrier paths with client probes before building
- Uses probe-based selection over geo-DNS
- Defines launch by quality, not deployment
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Choose mesh, SFU or MCU from meeting size, endpoints and cost, and explain where each breaks
- Design per-subscriber adaptation with simulcast or SVC, bandwidth estimation, layout-driven caps and audio priority
- Place media servers by participant proximity and cascade them across regions with a bounded topology
- Handle loss and jitter inside a 150–300 ms budget: NACK from the nearest server, FEC for audio, layer downgrade, adaptive jitter buffers
- Provide TURN on UDP, TCP and TLS 443, and size it for 10–20%+ relay
- Separate interactive meetings from webinars, with last-N for large meetings and broadcast paths for audiences
- Size the fleet by egress and NIC, and turn bandwidth into a monthly bill
- Define quality SLOs per participant-minute and slice them by region, ASN and client version
The Bar for This Question#
Mid-level (L4): Describes WebRTC peer connections with a signalling server and STUN. Works for a 1:1 call on friendly networks. No server topology beyond "a media server," no adaptation, no TURN.
Senior (L5): Uses an SFU, TURN, horizontal scaling and reconnection. The gap: a single stream per sender so the weakest participant degrades everyone; one server per meeting in the host's region; TCP or "retransmit everything" for loss; webinars as large meetings; server uptime as the SLO. The design works in a demo and fails on the first global meeting with one participant on a train.
Staff+ (L6): Frames the problem as per-viewer adaptation within a latency and cost budget in the first five minutes. Chooses topology by meeting size, designs simulcast/SVC with per-subscriber allocation, places SFUs by participant with cascading, handles loss by media type and RTT, provides TURN on 443, separates webinars, and measures quality per participant-minute. Names who pays — the weak participant gets lower resolution, senders pay simulcast uplink, the company pays egress — and nobody else pays for one bad network. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 The MCU Is a Legacy Adapter, Not an Architecture#
| Property | SFU | MCU |
|---|---|---|
| Server CPU per viewer | Near zero (forwarding) | ~1 core (decode, mix, encode) |
| Added latency | Few ms | 50–200 ms |
| Per-viewer layout | Client-side, free | Server-side, expensive |
The Staff position: SFU for clients; mixing only at the edge for phones and room systems that need it.
Why this matters in interviews: Choosing an MCU "to save client bandwidth" signals not having done the server math.
10.2 Audio Is the Product; Video Is a Feature#
| Degradation | User Reaction |
|---|---|
| Video freezes, audio fine | Annoyed; meeting continues |
| Audio choppy, video fine | Meeting fails |
| Both degrade gracefully | Acceptable |
The Staff position: Allocate audio first, protect it with FEC and RED, and drop video before audio ever suffers.
Why this matters in interviews: Candidates who prioritize by media type show they know what users actually tolerate.
10.3 Your Latency Budget Is Mostly Spent Before Your Servers See a Packet#
| Segment | Typical Share |
|---|---|
| Capture, encode, packetize | 20–40 ms |
| Last mile + backbone | 20–200 ms |
| Jitter buffer | 20–200 ms |
| Decode, render | 10–30 ms |
| SFU forwarding | 1–5 ms |
The Staff position: Optimize placement and jitter buffers, not SFU code paths. The server's share is the smallest.
Why this matters in interviews: It redirects depth to where latency actually lives.
10.4 "Scale to 10,000 Participants" Is Usually the Wrong Requirement#
| What They Say | What They Mean |
|---|---|
| 10,000 participants | 5 speakers, 9,995 viewers |
| Everyone interactive | Q&A from a handful |
| Low latency for all | Low latency for speakers |
The Staff position: Clarify into panel plus audience, then serve the audience with a broadcast path.
Why this matters in interviews: Reframing a scale requirement into its real shape is a core Staff move.
10.5 Conferencing Is a Networking Business With Some Media Software#
The Staff position: At scale, peering, backbone, regional presence and egress cost decide quality and margin far more than SFU implementation details. The SFU is open-source; the network is the moat.
Why this matters in interviews: It shows you know where the money and the quality are.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer builds an adaptive, well-placed, resilient conferencing service. The Principal engineer notices that the company now has video in four places — the meetings product, a customer-support video feature, a telehealth partner integration and an internal events platform — each with its own media stack, TURN fleet, quality metrics and vendor contracts. And the bill: egress is the largest infrastructure line in the company, negotiated piecemeal. The L7 problem is real-time media as a platform: one media network, one quality standard measured per participant-minute, and a deliberate decision about which parts of the network to own.
The Org-Level Fault Line#
One real-time media platform vs per-product media stacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each product runs its own stack or vendor | Fast; products choose | N TURN fleets, N quality definitions, fragmented egress contracts | Finance (egress), users (inconsistent quality) |
| One media platform, products build on SDKs | Shared network, shared quality tooling, consolidated egress | Platform must serve 1:1, meetings and broadcasts | Platform team (scope), products (migration) |
| Platform for the network, products own experience | Network and SFUs shared; products own UX and workflows | Requires stable APIs and capacity governance | Platform (API stability), products (adoption) |
🧭 Principal Move: "The platform owns the media network — SFUs, TURN, cascading, quality telemetry — and exposes rooms and tracks through SDKs. Products own their workflows and UX. New products don't get to stand up their own TURN fleet; they get a quota on ours."
Cost Model#
Assumptions: average 2 Mbps egress per participant; peak-to-average 3×; blended egress $0.002–0.01/GB depending on peering; SFU hosts with 25 Gbps NICs; fully loaded engineer ~$250K/year. Illustrative.
| Scale | Concurrency (peak) | Infra ($/month) | Headcount | On-call Load | Notes |
|---|---|---|---|---|---|
| Startup | 5K participants | ~$5–20K on a managed media platform | 1–2 eng (product integration) | Shared | Buy; own the UX |
| Growth | 200K participants | ~$150–500K (self-run in cloud regions, cloud egress pricing) | 8–15 eng (media, client SDKs, network) | Dedicated rotation | Egress is 70%+ of infra |
| Global | 3M participants | ~$2–7M (own edge, peering, ~900 media hosts) | 60–120 eng (media, clients, network, quality) | Per-region rotations + network on-call | Peering strategy is the main lever |
The pricing insight: compute is noise; egress and network are the business. Every product decision that changes bytes per participant-minute — default gallery size, thumbnail resolution, P2P for 1:1, last-N thresholds — is a margin decision, and should be reviewed as one.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| E2EE promised to customers | One-way | Withdrawing it is a trust event; it constrains server features forever |
| Client SDK API shape used by products and partners | One-way-ish | Every integration must migrate |
| Building an owned edge network and peering | One-way-ish | Large capital and contract commitments |
| Residency guarantees per tenant | One-way | Contractual |
| Codec choices (adding AV1/SVC) | Two-way | Negotiated per session |
| Cascade topology algorithm | Two-way | Control-plane change |
| Allocator strategy | Two-way | Config |
| Default gallery resolution | Two-way | A/B testable |
The Standard I'd Write#
RFC-MEDIA-001: Real-Time Media Platform Standard
Status: Approved Owners: Media Platform + Network Engineering
Scope
Every product feature that sends or receives real-time audio or video.
MUST
1. Use the media platform SDKs; no product-owned SFU or TURN fleets.
2. Publish simulcast or SVC video; pause unsubscribed layers.
3. Report client quality telemetry: freeze ratio, audio concealment,
resolution delivered, RTT, loss — per participant-minute.
4. Respect tenant residency attributes; never place media outside them.
5. Run events above 1,000 viewers in webinar mode with a capacity reservation.
SHOULD
1. Use peer-to-peer for 1:1 calls when a direct path exists and policy allows.
2. Default thumbnails to the lowest layer on small screens.
3. Ship media-affecting client changes behind remote config with staged rollout.
Exceptions
Filed with Media Platform; network review required for any non-platform media path.
Success metrics
- Participant-minutes with freeze ratio under 1%: >= 98%
- Audio concealment ratio p95: < 3%
- Join success: >= 99.5%
- Egress cost per participant-minute: tracked, trending down
What I'd Tell the VP#
"Video is now in four of our products, each built separately, and bandwidth for it is our largest infrastructure cost. I'm proposing one real-time media platform that every product builds on, with a shared quality standard and consolidated network contracts. It needs about ten engineers for a year to consolidate, and should cut media cost by a quarter to a third through shared capacity and better peering, while making call quality measurable and comparable across products. The main risk is migrating products mid-roadmap; we'd start with the newest feature and move the largest last."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices bytes, not servers | "Every default that changes bytes per participant-minute is a margin decision." |
| Owns the network question | "Peering and edge presence are the real build-vs-buy decision; the SFU is commodity." |
| Identifies one-way doors | "Promising E2EE constrains recording and transcription forever." |
| Sets a platform boundary | "Platform owns media network and telemetry; products own experience." |
| Measures quality as product | "Freeze and concealment per participant-minute are shared SLOs for client and media teams." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a great SFU fleet for the meetings product" — correct, but ignores three other products running their own stacks.
- "We'll add regions to reduce latency" — no business case, no peering strategy, no launch criteria by quality.
- "We'll offer E2EE" — no acknowledgment of the server features it permanently removes.
Appendices
Appendix A: Mechanics in Depth#
A.1 Per-Subscriber Layer Allocation#
every 200-500 ms, for subscriber s:
budget = bwe[s] * 0.9 # headroom for probing and RTCP
allocate audio for top 3 speakers by audio level # ~32 kbps each
candidates = subscriptions[s] sorted by priority:
pinned > active_speaker > visible_gallery > offscreen
for track t in candidates:
want = layer_for(t.requested_resolution) # from client layout
for layer in descending(min(want, t.available_layers)):
if cost(layer) <= budget:
if layer > current[s][t] and not stable_for(s, 3 s): continue
assign(s, t, layer); budget -= cost(layer); break
else: assign(s, t, OFF) # show avatar, audio only
for publisher p: if no subscriber wants p.layer_k: tell p to pause layer_k
A.2 Switching Layers Seamlessly#
simulcast: switching to a higher layer waits for a keyframe on that layer;
SFU requests one (rate-limited), then rewrites RTP sequence numbers,
timestamps and SSRC so the receiver sees one continuous stream
SVC: switching temporal layers can happen at the next frame;
spatial up-switch waits for a switching point / keyframe
down-switch: immediate in both
A.3 Loss Recovery Decision#
audio: Opus in-band FEC always; PLC on receiver; RED when loss > 5%
video: NACK served from SFU's per-subscriber packet cache (last ~1 s)
if rtt_to_subscriber > 100 ms or loss > 10%: lower temporal layer, consider FEC
if loss > 3% sustained 5 s: treat as congestion, lower spatial layer
PLI: aggregate per publisher layer, forward at most 1 per second
jitter: adaptive target from measured inter-arrival variance; 20-200 ms
Appendix B: Data Model#
CREATE TABLE meetings (
meeting_id TEXT PRIMARY KEY,
tenant_id TEXT NOT NULL,
mode TEXT NOT NULL DEFAULT 'meeting', -- meeting | webinar
residency TEXT, -- NULL | 'eu' | ...
e2ee BOOLEAN NOT NULL DEFAULT false,
recording BOOLEAN NOT NULL DEFAULT false,
scheduled_start TIMESTAMPTZ,
expected_size INT
);
-- live state: in a regional store (e.g. Redis) with TTLs refreshed by heartbeats
-- meeting:{id}:sfus -> set of {region, sfu_id}
-- meeting:{id}:participants -> hash participant_id -> {sfu_id, role, tracks}
-- sfu:{id}:load -> {egress_gbps, pps, meetings, updated_at}
-- cascade:{meeting_id} -> list of links {from_sfu, to_sfu, via_hub}
CREATE TABLE recordings (
recording_id TEXT PRIMARY KEY,
meeting_id TEXT NOT NULL,
region TEXT NOT NULL,
segments_uri TEXT NOT NULL, -- per-track segments every 5-10 s
composite_uri TEXT, -- produced after meeting ends
state TEXT NOT NULL -- recording | compositing | ready | failed
);
Appendix C: Coordination Mechanisms#
C.1 Join and SFU Reattachment#
C.2 Quick Comparison#
| Mechanism | Guarantees | Failure Mode | Use For |
|---|---|---|---|
| Simulcast / SVC | Per-subscriber quality | Uplink cost; keyframe dependence | Every video sender |
| Per-subscriber BWE allocation | Weak links don't affect others | Estimator oscillation | Every subscriber |
| Paused unsubscribed layers | Saves sender uplink and CPU | Slow resume if over-paused | Meetings with a dominant speaker |
| Nearest-SFU + cascade | Short last mile globally | Static topology over bad links | Multi-region meetings |
| NACK from SFU cache | Cheap recovery near receiver | Long RTT | Video |
| Opus FEC + RED | Audio survives loss | Bandwidth overhead | Audio |
| PLI aggregation | No keyframe storms | Slower recovery for some | Popular publishers |
| TURN TLS 443 | Connects through strict firewalls | TLS inspection | Enterprise networks |
| Segmented recording | Crash loses ≤ 10 s | Compositing step | Recordings |
Appendix D: API Contract & Client Behavior#
- Clients report layout-derived maximum resolutions per subscription; the SFU never sends more.
- Clients send quality telemetry every 10 seconds: freeze events, concealment, resolution, RTT, loss, candidate type.
- On media loss of 1.5–2 s with signalling alive, clients wait for a reconnect instruction before restarting ICE; on signalling loss, they reconnect with jittered exponential backoff (1, 2, 4, 8 s ± 50%).
- Join retries use jittered backoff; clients never retry joins more than once per second.
- E2EE meetings use insertable-stream encryption; recording, transcription and dial-in are unavailable and the UI says so.
- Webinar attendees cannot publish until promoted; promotion moves them to the panel meeting.
Appendix E: Observability#
Core metrics:
join.success_rate{region},join.time_to_first_audio_p95participant.freeze_ratio{region, asn, client_version},audio.concealment_ratio{...}video.resolution_delivered{...},bwe.estimate_p50{asn}sfu.egress_gbps,sfu.packet_rate,sfu.allocations_per_10scascade.loss_rate{region_pair},cascade.rtt{region_pair}ice.selected_candidate_type{customer},turn.egress_gbpspli.rate{publisher},recording.segments_written
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| Join success rate | < 99% for 5 min (per region) | Page |
| Audio concealment p95 | > 5% for 10 min (per region or top ASN) | Page |
| SFU egress / NIC | > 80% on > 5% of hosts | Page |
| Cascade link loss | > 2% for 5 min | Page (network) |
| Recording segments stalled | > 30 s | Page (media services) + host notice |
| Candidate-type shift for a large customer | relay share +30 points | Ticket (customer success) |
Debugging the silent failure: server health metrics stay green during most quality incidents. Slice client-reported quality by region, ASN, client version, device class and customer; distinguish late packets (jitter buffer discards) from lost packets; and check cascade links separately from last-mile paths.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 10K concurrent | Managed platform or one open-source SFU deployment, 2–3 regions | Cost or quality control |
| 10K–500K | Own SFU fleet, 10+ regions, simulcast, cascading, quality telemetry | Allocator herds; enterprise firewalls; egress bill |
| 500K–5M | Probe-based placement, dynamic cascade topology, webinar mode, SVC | Network cost; peering; regional quality gaps |
| > 5M | Owned edge, anycast entry, per-track distribution trees | Org coordination across products and regions |
What you don't build on day one: your own edge network, SVC, E2EE, anycast, server-side transcription. Each has a trigger in Section 11.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Capacity reservations for scheduled large meetings and webinars; unscheduled meetings share pooled capacity. See Autoscaling for calendar-driven pre-scaling.
- Per-tenant concurrency limits by plan, enforced at join; see Rate Limiter.
- Cell-based placement so one SFU or rack failure affects a bounded set of meetings; see Multi-Region.
- Cost attribution: participant-minutes, egress bytes, TURN bytes and recording storage metered per tenant.
- Defaults are cost policy: thumbnail resolution, last-N thresholds and P2P eligibility are reviewed as margin decisions, not just UX.