Technologies referenced in this case study: Apache Kafka · PostgreSQL · Cassandra · Redis · Kubernetes
Related: CDN & Edge Caching · Object Storage · Handling Large Blobs · Job Scheduler · Managing Long-Running Processes · Metrics & Alerting Platform · Buy or Build · Auth & Identity
Reading Guide#
Organized for interview use first, reference second. The upload mechanics (presigned URLs, multipart, the upload state machine) live in Large File Uploads & Delivery, and the edge cache itself lives in CDN & Edge Caching. This page is about the end-to-end video product and what it costs: which renditions exist, when they are made, where they are served from, and who pays for each choice.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal Lens) and the appendices on the transcode DAG, ABR and storage lifecycle |
What is a Video Streaming Platform? — Why interviewers pick this topic
A video platform accepts uploads (or a live feed), turns each source into a ladder of renditions at different resolutions, bitrates and codecs, cuts every rendition into short segments, publishes a manifest listing them, and serves the segments over HTTP from caches as close to the viewer as possible. The player picks a rendition segment by segment based on measured throughput and buffer — adaptive bitrate (ABR) — so a viewer on a train drops to 360p instead of stalling.
The hard part is not playing an MP4. The hard part is that every rendition you create costs compute once and storage forever, and every byte a viewer watches costs egress every time. At scale the bill is dominated by bytes leaving the network, not by servers.
Before vs After — the "viral upload" scenario:
Without a cost-and-tail design:
t=0: Creator uploads a 12-min 4K video. Pipeline encodes 14 renditions
(H.264 + VP9 + AV1, 144p–2160p) before marking it playable.
t=+38min: First playable. Creator has already tweeted the link; viewers saw "processing".
t=+40min: Link goes viral. 2M viewers in 30 min. Edge caches are cold for every rendition.
t=+42min: Origin egress jumps from 1.5 Tbps to 6 Tbps. Origin p99 TTFB 4s. Rebuffer ratio 3%.
t=+1 month: The same pipeline encoded 14 renditions for 700K other videos that day,
90% of which got fewer than 50 views. Storage bill +18% month over month.
With fast-path publish, popularity-driven encoding and tiered delivery:
t=0: Upload completes. Fast path encodes 360p + 720p H.264 in parallel 4s chunks.
t=+90s: Playable. Remaining H.264 rungs land by t=+8min.
t=+40min: View velocity crosses threshold → advanced-codec ladder queued at high priority.
Mid-tier shield coalesces misses: 1 origin fetch per segment per region.
t=+42min: Origin egress 1.7 Tbps. Edge hit ratio 97%. Rebuffer ratio 0.4%.
t=+1 month: Long-tail videos hold 5 H.264 rungs; sources moved to archive tier after 30 days.
Why interviewers reach for this question: "Design YouTube" looks like a storage question and is really an economics question with a latency budget attached. It separates candidates who draw "upload → S3 → CDN" from candidates who can say how many renditions, how big, served from where, at what cost per GB — and what happens on the night 10 million people press play at 20:00.
Mechanics Refresher: The Video Delivery Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| Progressive download (one MP4) | Player fetches one file with byte ranges | Trivial; any CDN | No adaptation; wastes bytes when viewers quit early; stalls on slow links |
| HLS (RFC 8216) | Media playlist of segments (.ts or fMP4) + multivariant playlist listing renditions | Native on Apple devices; universal | Segment duration sets the latency floor |
| MPEG-DASH | XML manifest (MPD) + fMP4 segments; codec-agnostic | Standard across non-Apple players; flexible | Not native on iOS Safari |
| CMAF | Common fMP4 segment format usable by both HLS and DASH | One set of files for both protocols — halves storage and cache footprint | Needs one encryption scheme (cbcs) for wide DRM coverage |
| Bitrate ladder | Set of (resolution, bitrate, codec) renditions per title | Lets ABR adapt | Each rung = compute once + storage forever |
| Per-title / per-shot encoding | Pick ladder or encoder settings per title or per shot from measured complexity | 20–50% fewer bits for easy content at equal quality (typical industry range) | Many trial encodes; more compute per title |
| ABR (client-side) | Player chooses next segment's rendition from throughput and buffer | Scales; no server state | Oscillation; greedy players fight over shared links |
| Low-latency HLS / DASH | Segments split into ~1s parts, published before the segment completes | 2–5s glass-to-glass over plain HTTP CDNs | More requests, more manifest churn, tighter origin SLOs |
| DRM (Widevine, FairPlay, PlayReady) | Segments encrypted; license server issues keys per session | Required by premium licensors | License server on the startup critical path |
For most production systems: CMAF segments, served as HLS and DASH from one set of files; an H.264 ladder for universal reach plus an advanced codec (VP9, HEVC or AV1) for popular content; 4–6 second VOD segments; client-side ABR with server-provided hints; a commercial CDN until egress spend justifies anything else. The primitives are not the interview — which renditions you are willing to pay for is.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
Video streaming is not a storage question. Anyone can put files in object storage.
It is a cost-and-tail question that tests:
- Whether you see that egress, not compute, dominates the bill — and design to move fewer bytes, from closer, more cheaply
- Whether you treat the long tail differently from the head: most uploads are rarely watched, a few are watched millions of times
- Whether you can turn "fast" into numbers: startup time, rebuffer ratio, publish time, glass-to-glass latency
- Whether you know which decisions are made by people outside engineering — licensing, DRM, regional rights, creator expectations
The key insight: Encode once, store a deliberately chosen set of renditions, and serve them from as close to the viewer as possible. Every rendition is a bet: compute and storage now, against egress saved and quality gained on every future view. Staff candidates price that bet per video based on its popularity; Senior candidates encode everything the same way.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws upload → blob store → transcoder → CDN → player | Asks "UGC, premium catalogue or live? They are three different systems" and commits to one with numbers | Asks "What share of revenue goes to delivery, and which lever — codec, edge, or deletion — moves it most over 3 years?" |
| Encoding | "Transcode every upload into 6 resolutions" | "Fast-path two rungs for publish time; full H.264 ladder for all; advanced codec only past a view-velocity threshold" | Treats encoding efficiency as a company cost lever: funds codec migration (or silicon) against a modeled egress saving with a payback date |
| Delivery | "Put a CDN in front" | "Edge + mid-tier shield, request coalescing, ≥ 95% byte hit ratio target, origin sized for 3× the steady miss rate" | Decides own-edge vs commercial vs hybrid as a multi-year bet; negotiates ISP peering; keeps a commercial CDN as a pressure valve |
| Quality | "Use adaptive bitrate" | "ABR with buffer-based steady state; QoE SLOs: startup p95 ≤ 2.5s, rebuffer ratio ≤ 0.3%, measured per ISP and device" | Makes QoE a business KPI tied to watch time; owns the tradeoff between bitrate caps and egress spend with finance and product |
| Storage | "Store everything in S3" | "Lifecycle tiers by age and views; source to archive at 30 days; renditions re-creatable from source" | Writes the retention and deletion policy for the long tail with legal and creator relations; storage growth must not outrun watch-time growth |
| Ownership | Video team owns everything | Ingest, media pipeline, playback and delivery are separate owners with SLOs at each handoff | Defines the platform contract so live, shorts and premium share one pipeline and one edge without sharing a failure domain |
Why "encoding" separates levels
L5: "Each upload triggers a job that transcodes it into 240p, 360p, 480p, 720p, 1080p and 4K, stores them, and marks the video ready." It works and it is how most people would start. It quietly assumes every video is worth the same spend, and it makes publish time equal to the slowest rung of the longest video.
L6: "I'll split encoding into three tiers by value. The fast path produces 360p and 720p H.264 from parallel 4-second chunks so a 10-minute upload is playable in about 90 seconds. The rest of the H.264 ladder fills in within 10 minutes. The advanced codec — which costs roughly 5× the compute but saves 30–40% of bytes — runs only once a video crosses about 1,000 views in its first day, because below that the compute costs more than the egress it saves."
L7: "The question is what our marginal cost per watch-hour is and which lever bends it. At our volume a 30% bitrate saving on the top 10% of videos is worth more than any other project on the roadmap. I'd fund the codec migration with a payback model — and I'd revisit hardware encoding when software encode becomes the constraint, which is exactly the trade large platforms have made."
Why "delivery" separates levels
L5: "Videos are served from a CDN, so scale is handled." The CDN does handle steady state. The answer is silent on what happens on a cache miss for a 2-hour film with 12 renditions × 1,800 segments each, and on who pays the egress bill.
L6: Separates the head from the tail. The head (popular videos) is pre-warmed or pre-positioned into edge caches; the tail is served from a mid-tier shield that coalesces misses so one origin fetch serves a whole region. "The number I protect is origin egress: if the edge byte hit ratio drops from 97% to 90%, origin load more than triples, and that is the outage."
L7: Frames the edge as a multi-year capital decision. "Below a few hundred Gbps, a commercial CDN wins on every axis. Above tens of Tbps, embedding caches inside ISPs changes our unit cost and our ISPs' transit bill — but it's a 3–5 year program with hardware logistics, and I'd keep a commercial CDN for overflow and for markets where we don't have boxes."
Why "storage" separates levels
L5: Stores the source and every rendition forever in standard-tier object storage. Fine for year one. By year three, storage is the fastest-growing line item and nobody can say which bytes earn their keep.
L6: "Renditions are a cache of the source, not data. I keep the source in archive tier after 30 days, keep only the H.264 ladder for videos with fewer than 10 views in 90 days, and delete advanced-codec renditions that stop being watched — they can be regenerated."
L7: Recognizes that deletion is a policy question with creators, legal and product: what does "we keep your video forever" mean? "I'd write the retention policy with legal: the source is preserved, playback quality for dormant videos may be reduced to a minimal ladder, and the first view after dormancy can trigger re-encoding. That sentence saves more money than any compression project."
Positions to Commit To#
| Position | Rationale |
|---|---|
| Egress is the bill; design to move fewer bytes from closer | At scale, delivery costs roughly 50× or more what transcode compute costs; every architecture choice is judged on bytes and hit ratio |
| Fast-path publish, then fill the ladder | Creators judge the platform by time-to-playable; 2 rungs in ~90s beats 14 rungs in 40 min |
| Encode effort follows popularity | Advanced codecs and per-title tuning pay back only above a view threshold; the long tail gets a cheap universal ladder |
| CMAF segments, one set of files for HLS and DASH | Halves storage and cache footprint versus packaging twice |
| 4–6 s segments for VOD, 1 s parts for low-latency live | Longer segments = fewer requests and better compression; short parts only where latency is the product |
| Client-side ABR, server-side steering hints | The client knows its buffer; the server knows CDN health and cost — combine both |
| Renditions are a regenerable cache; the source is the asset | Lets you tier and delete renditions aggressively without losing anything irreplaceable |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| User-generated uploads at scale (YouTube-like) | Ingest-heavy, long tail, creators want fast publish | Fast-path encode, popularity-driven ladder, tiered storage, pull-through CDN with shield | Transcode backlog on upload spikes; storage outgrows views | Publish p50 ≤ 2 min; no upload lost; rebuffer ≤ 0.3% |
| Premium catalogue (Netflix-like) | Small catalogue (tens of thousands of titles), extreme popularity skew, licensors' DRM rules | Exhaustive per-title/per-shot encoding, pre-positioning to edge before release, DRM on every stream | Cold cache on premiere night; license server on the critical path | Startup p95 ≤ 2s; highest quality per bit; contractual DRM compliance |
| Live streaming (Twitch, sports) | Seconds of glass-to-glass latency, no time to pre-position, sudden audience spikes | Real-time ladder at ingest, LL-HLS/LL-DASH parts, request coalescing at every tier | Ingest drop, origin stampede on a spike, ABR thrash at low buffer | Glass-to-glass 3–5s (LL) or ≤ 30s (standard); no stream dies on one encoder |
🎯 Staff Move: "I'll design for user-generated uploads at YouTube-like scale, because that's where ingest, the long tail and egress all collide. A premium catalogue spends far more compute per title and pre-positions everything; live trades efficiency for seconds of latency. I'll point out where each would change my design, and you can redirect me to either."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Encode Everything Up Front vs Encode on Demand | Pay compute and storage for every rendition of every upload, or accept slower first-view quality for the long tail? |
| 2 | Fixed Ladder vs Per-Title / Per-Shot Encoding | Spend more compute per title to save bandwidth on every view — worth it only where views are high |
| 3 | Own CDN vs Commercial CDN vs Both | Unit cost and control at extreme scale vs capital, logistics and years of build-out |
| 4 | Segment Duration | Short segments cut latency and speed up adaptation; long segments cut request rate and improve compression |
| 5 | Client-Driven ABR vs Server-Guided Steering | The player knows its buffer; the platform knows CDN health, cost and congestion — who decides the rendition and the CDN? |
How Real Companies Built It#
Why this section belongs here: Naming how real video platforms solved these problems shows you've studied the economics, not just the boxes.
YouTube — Custom Transcoding Silicon#
YouTube has said that on average more than 500 hours of video are uploaded every minute, that VP9 takes about 5× more compute to encode than H.264, and that after starting work in 2015 it built a Video (trans)Coding Unit (VCU) — the "Argos" chip — delivering up to 20–33× better compute efficiency than its previous optimized software system on general-purpose servers; the next generation targets AV1 (YouTube blog, ASPLOS 2021 paper).
The paper's extended abstract adds that the VCU fleet spans tens of thousands of servers and that, for offline two-pass encoding, it delivers 8–20× higher throughput than software encoding for H.264 and VP9 respectively (extended abstract).
Staff insight: Better codecs move cost from egress to compute. When upload volume makes software encoding of an advanced codec unaffordable, the choices are: encode it only for popular videos, or change the cost of compute. Saying both out loud — and knowing that only a handful of companies can do the second — is the signal.
Netflix — Per-Title and Per-Shot Encoding#
In December 2015 Netflix described moving from one fixed bitrate ladder for every title to a ladder chosen per title from an analysis of that title's complexity — simple animation needs far fewer bits than grainy action footage to look the same — and later described its Dynamic Optimizer, which chooses encoding parameters per shot to optimize perceptual quality at each bitrate (per-title encoding, Dynamic Optimizer).
Staff insight: Per-title encoding is affordable for Netflix because the catalogue is small and every title is watched many times. For UGC the same idea applies only to the head of the distribution. The interview move is to say where on the popularity curve the extra encode pays back.
Netflix — Open Connect#
Netflix runs its own CDN, Open Connect. Its appliances are provided to qualifying ISP partners at no charge to embed inside their networks, with the same capabilities as the appliances Netflix operates itself, and their content is updated through nightly fills (Open Connect).
Staff insight: Pre-positioning works because a premium catalogue's demand is predictable a day ahead. UGC demand is not, so a UGC platform still needs pull-through caching with a shield. Owning the edge is a decision about decades of traffic, not about this quarter's design. See Content Delivery Network for the cache mechanics.
Apple — HLS Authoring Specification#
Apple's HLS authoring specification for its devices says target segment durations SHOULD be 6 seconds, video segments MUST start with an IDR frame and key frames SHOULD occur every 2 seconds; for Low-Latency HLS the recommended Part Target Duration is 1 second and SHOULD be at least three times the p95 round-trip time; and HEVC variants should use bitrates about 20% below the H.264 values (HLS Authoring Specification, RFC 8216).
Staff insight: Segment duration is not a free parameter — the platform vendor publishes a default, and the low-latency variant ties part duration to network round-trip time. Citing the default and then explaining when you'd deviate is stronger than inventing a number.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "We transcode into 6 resolutions" | "How long until a 1-hour 4K upload is playable? What does that cost per day across all uploads?" | Publish-time design and compute math |
| "We store videos in S3" | "How much storage do you add per day, and what is it in year three?" | Long-tail economics, tiering, deletion |
| "We put it behind a CDN" | "A new episode drops at 20:00 for 10M viewers. What does origin see at 20:00:05?" | Cold cache, coalescing, pre-warming |
| "The player uses adaptive bitrate" | "1M players on one congested ISP all switch down, then up, then down. What happened?" | ABR oscillation, shared bottlenecks |
| "We use HLS with short segments for low latency" | "How short, and what does that do to request rate and encoding efficiency?" | Segment-duration tradeoff with numbers |
| "We'll build our own CDN" | "At what traffic level does that pay back, and what do you do in markets you haven't reached?" | Build-vs-buy as a multi-year bet |
System Architecture Overview#
Reading the diagram: Two almost independent systems share a metadata store. The media pipeline (left) is a batch system judged on publish time and cost per encoded hour. The delivery path (right) is a read system judged on startup time, rebuffer ratio and egress cost. The playback API is the only synchronous control-plane hop on the play path; segments never touch application servers. The number that decides whether the system survives a big night is
cdn.byte_hit_ratio— origin load is proportional to the miss rate, not to the audience.
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| What dominates cost | "Storage and transcoding" | "Egress. At 1B watch-hours a day, delivery is over 1 EB/day; a 1-cent-per-GB difference is about $4B a year." |
| Encoding | "Transcode to all resolutions" | "Fast path 2 rungs in ~90s, full H.264 ladder in ~10 min, advanced codec only above a view threshold." |
| Segments | "Short segments" | "4–6s for VOD; 1s parts for low-latency live. Halving segment length doubles request rate." |
| CDN | "Use a CDN" | "Edge + shield, coalescing, ≥ 95% byte hit ratio; pre-warm the head; protect origin with 3× miss headroom." |
| ABR | "Adaptive bitrate" | "Throughput-based at startup, buffer-based in steady state, hysteresis to stop oscillation; server hints for CDN choice." |
| Storage | "Keep everything" | "Source is the asset; renditions are a cache. Tier by age and views; delete cold advanced-codec rungs." |
| QoE | "Make it fast" | "Startup p95 ≤ 2.5s, rebuffer ratio ≤ 0.3%, start failures ≤ 0.5%, sliced by ISP, device and CDN." |
Numbers to Bring#
All numbers below are this design's working assumptions unless a source is cited. YouTube's own published upload figure is "more than 500 hours every minute" (YouTube blog); the design uses that as its ingest rate.
| Metric | Value | Why It Matters |
|---|---|---|
| Upload ingest rate | 500 h of video/min → 30K h/hour → 720K h/day; peak ~2× | Sizes the encode fleet and the daily storage add |
| Average upload | ~10 min, ~8 Mbps source → ~3.6 GB per hour of video | ~50 uploads/s; ~2.6 PB/day of source |
| H.264 ladder (240p–1080p) | 0.3 + 0.7 + 1.2 + 2.5 + 4.5 = 9.2 Mbps total | ~69 MB per minute of video for the full ladder |
| Storage per minute per rendition | 1080p ≈ 34 MB · 720p ≈ 19 MB · 360p ≈ 5 MB | Multiply by renditions × minutes × uploads/day |
| Renditions per title | 5 H.264 rungs for all; +5–7 advanced-codec rungs for the head; + 2 audio | 7 for the tail, 12–14 for the head |
| Eager rendition storage | ~3 PB/day (H.264 ladder for every upload) | ~1 EB/year before tiering |
| Encode compute (software) | ~3 core-hours per video-hour for the H.264 ladder; advanced codec ~5× | ~90K cores busy for H.264 alone at 720K h/day |
| Segment size (4s at 2.5 Mbps) | ~1.25 MB | One request per 4s per viewer = 0.25 req/s |
| Peak concurrent viewers | ~75M (1B watch-hours/day ÷ 24 × 1.8 peak factor) | ~19M segment req/s at 4s segments; ~38M at 2s |
| Peak egress | ~190 Tbps (75M × 2.5 Mbps) | The number the CDN contract is written around |
| Edge byte hit ratio target | ≥ 95% (≥ 99% after mid-tier) | Origin sees ≤ 1% of bytes ≈ 1.9 Tbps at peak |
| Egress cost | owned edge + peering ~$0.002–0.005/GB; commercial committed ~$0.005–0.02/GB | 1.1 EB/day × $0.01 = ~$11M/day |
| Startup time | p50 ≤ 1s, p95 ≤ 2.5s | Every extra second before first frame loses viewers who never return to the video |
| Rebuffer ratio | ≤ 0.3% of watch time | The stall a viewer feels most; set the SLO per ISP and device, not just globally |
| Publish time | first playable p50 ≤ 2 min; full H.264 ladder ≤ 15 min | Creator-facing SLO |
| Live glass-to-glass | Standard HLS (6s segments) 15–30s · LL-HLS (1s parts) 3–5s · WebRTC < 1s | Each step down multiplies cost per viewer |
Interview Walkthrough
A 45-minute plan for "Design YouTube." The goal is to spend under 10 minutes on the boxes and the rest on the three places the system actually breaks: the encode pipeline under load, the edge under a spike, and the cost curve over time.
Phase 1: Requirements & Framing (2–3 minutes)#
Say this, nearly word for word:
"Before I draw anything — there are three different products hiding in 'design YouTube': user uploads at scale, a premium catalogue like Netflix, and live streaming. They share a CDN and a player, but the pipeline and the economics are different. I'll design for user uploads, and I'll call out what changes for the other two.
Scale assumptions: 500 hours uploaded per minute, about 1 billion watch-hours a day, viewers worldwide, mostly on mobile. Functional scope: upload, process, publish, play with adaptive bitrate, basic view counts. Out of scope unless you want them: recommendations, comments, monetization, search.
The constraints I'll design to: a creator sees their video playable within about 2 minutes; viewers start playback in under 2.5 seconds at p95 with under 0.3% of watch time spent rebuffering; and the cost per watch-hour goes down over time, because egress is the line item that decides whether this business works."
Why this works: It names three intents, commits, attaches numbers to "fast", and announces the cost thesis in the first two minutes. The interviewer now expects a conversation about bytes and tails, not about CRUD.
Phase 2: Core Entities & API (1–2 minutes)#
Video { video_id, owner_id, state, duration_s, source_uri, created_at,
visibility, rights{regions[], drm_required}, popularity_tier }
Rendition { video_id, codec, height, bitrate_kbps, state, segment_prefix,
storage_tier, bytes, created_at, last_served_at }
EncodeTask { task_id, video_id, chunk_idx, ladder_rung, priority, attempt,
lease_owner, lease_expires_at }
PlaySession { session_id, video_id, user_id?, device, cdn, started_at } # analytics
POST /videos → { video_id, upload_id } # starts multipart
PUT <presigned part URL> → 200 # client → storage
POST /videos/{id}/complete → 202 { state: PROCESSING }
GET /videos/{id}/play?device=... → { manifest_url (signed, 6h), license_url?,
cdn_hints[], start_rung }
GET <cdn>/v/{id}/{rendition}/seg_{n}.m4s → segment (immutable, cache 30d)
POST /qoe/beacons → 204 # startup, stalls, bitrate switches
Say: "Segments are immutable and content-addressed by video, rendition and index, so they cache forever. Manifests for VOD are cacheable too once the ladder is final; I'll give them a short TTL while rungs are still being added. The play endpoint is the only dynamic call on the play path — it signs URLs, checks rights and region, and hands the player CDN hints."
Phase 3: High-Level Architecture (≤5 minutes)#
Staff candidates spend under 5 minutes here. Draw two lanes: the media pipeline and the delivery path. Everything else is detail you add on demand.
Narrate: "Upload goes straight to storage with presigned multipart — the API never sees the bytes; that pattern is covered in large-blob handling. Completion emits an event. The scheduler splits the source into GOP-aligned 4-second chunks and fans out tasks by priority: fast-path rungs first, the rest of the H.264 ladder next, advanced codec only when popularity says so. Workers write CMAF segments to the rendition origin. On the read side: player calls the play API, gets a signed manifest, fetches segments through edge and shield caches."
Phase 4: Transition to Depth (1 minute)#
"The boxes are standard. The interesting decisions are three: how much encoding we do and when, because that sets publish time and the compute-plus-storage bill; how we protect origin when a video or a premiere goes viral, because that's where outages come from; and how ABR and segment length trade latency against cost. I'd like to start with encoding economics — is that where you want to go, or would you rather start with delivery?"
Phase 5: Deep Dives (25–30 minutes)#
Deep dive 1 — Encoding economics (8–10 min). Walk the arithmetic. 720K hours/day × ~3 core-hours per hour for the H.264 ladder ≈ 2.2M core-hours/day ≈ 90K cores. Advanced codec at ~5× would add ~450K cores if applied to everything. Then the popularity curve: "If 10% of videos earn about 90% of watch time — a typical UGC skew I'd verify in our own data — applying the advanced codec to that 10% costs ~45K cores and saves ~30% of bytes on ~90% of egress. Applying it to the other 90% costs ~400K cores to save bytes on 10% of egress. So the threshold is a number, not a philosophy." Mention chunked parallel encoding for publish time and the priority classes.
Deep dive 2 — Delivery under a spike (8–10 min). Edge + mid-tier shield; request coalescing (one upstream fetch per object per tier while others wait); pre-warm the first 3 segments of the top rungs for anything the recommendation system is about to push; origin sized for 3× normal miss traffic. "If edge byte hit ratio drops from 97% to 91%, origin load triples. That's the incident, so cdn.byte_hit_ratio per region is a paging metric." Defer cache mechanics to the CDN case.
Deep dive 3 — ABR, segments and QoE (6–8 min). 4s segments for VOD (Apple's spec defaults to 6s; I'll go slightly shorter for faster adaptation on mobile). Throughput-based start at a conservative rung, buffer-based steady state with hysteresis. QoE beacons → startup time, rebuffer ratio, bitrate, sliced by ISP and CDN. Server-side hints choose the CDN; the client chooses the rung.
If time remains — live: real-time ladder at ingest, 1s LL-HLS parts, glass-to-glass 3–5s, and why a 10M-viewer spike is a coalescing problem, not a transcoding problem.
Phase 6: Wrap-Up (2–3 minutes)#
"To summarize: the media pipeline optimizes publish time and cost per encoded hour by spending effort according to popularity. Delivery optimizes origin protection and egress cost with tiered caches and coalescing. QoE is measured on the client and sliced by ISP and CDN. What I'd build later: per-title encoding for the head, a hybrid edge once egress crosses tens of Tbps, and a retention policy for the dormant tail written with legal. The one-way doors are the segment format and the DRM scheme; everything else can change behind the manifest."
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| Explaining multipart upload in detail | 5–8 min | One sentence: "presigned multipart, API never sees bytes" — link to large-blob handling |
| Designing the recommendation system or comments | 10+ min | Scope it out in Phase 1 |
| Drawing 15 boxes before any numbers | 8 min | Two lanes, then numbers |
| Debating HLS vs DASH | 3–5 min | "CMAF, served as both" |
| Never getting to cost | Whole signal | Put the egress number in Phase 1 |
| Skipping live because "it's similar" | Pivot fails | One sentence on what changes: no pre-positioning, 1s parts, coalescing |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
"Design YouTube" is the most common media prompt because it contains every cost tension of a large-scale read system in one product. The ingest side is a batch pipeline with a backlog problem. The read side is a cache with a 1,000:1 popularity skew. The bill is dominated by a resource — egress bandwidth — that most backend engineers never see on a dashboard. A Senior engineer can build every box. The Staff signal is noticing that the boxes are not where the risk is: the risk is a viral upload that the encode queue can't absorb, a premiere night with cold caches, and a storage curve that grows faster than watch time. Interviewers use it to see whether you can put a dollar sign and an owner on each of those.
1.2 The L5 vs L6 Contrast — Visual#
The Senior path is correct and ends at scaling the metadata database — which is a few thousand QPS and never the bottleneck. The Staff path spends its time where the money and the outages are.
1.3 The Staff Question That Cuts Through Everything#
"For a video that will be watched N times, what is the cheapest set of renditions — and the cheapest place to serve them from — that still meets our startup and rebuffer targets?"
Every fault line is a version of this question. Encode up front or on demand: depends on N. Per-title encoding: pays back above some N. Own the edge: depends on the sum of N across the catalogue. Segment duration: trades request cost against latency per view. The candidate who asks it out loud has reframed the interview from "what boxes" to "what is each byte worth."
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1: User-generated uploads at scale. Ingest is enormous and unpredictable: 500 hours a minute, spiky around events and time zones. Most uploads are watched rarely; a small fraction are watched enormously and unpredictably — a video can go from 0 to 5M views in an hour. Creators measure the platform by time-to-playable and by quality on their own phone. The design centre is a priority-aware encode pipeline, popularity-driven encoding effort and pull-through caching with a shield because you cannot predict which video to pre-position.
Intent 2: Premium catalogue. Tens of thousands of titles, each expensive to license and watched by millions. Releases are scheduled, so demand is predictable to the hour. Spending hours of compute per title on per-shot optimization pays back across millions of views. Licensors impose DRM, regional windows and output-protection rules; the license server is on the startup critical path. The design centre is encode quality per bit and pre-positioning — fill edge caches overnight before the release.
Intent 3: Live streaming. No time to optimize encodes — the ladder is produced in real time at ingest. No time to pre-position — the content didn't exist a second ago. Audiences spike from 10K to 10M in minutes when a match goes to penalties. The design centre is latency budget (glass-to-glass), ingest redundancy (dual encoders, dual ingest points), and request coalescing at every tier, because every viewer requests the same newest part within the same second.
🎯 Staff Move: "These three share a player and a CDN but not a cost model. UGC is a tail problem, premium is a quality-per-bit problem, live is a latency problem. I'll design UGC, and for each fault line I'll say in one sentence what the premium and live answers would be."
2.2 When NOT to Build Your Own Video Pipeline#
Most companies that need video should not build any of this.
| Situation | Better Choice | Why |
|---|---|---|
| Video is a feature, not the product (course platform, product demos, support clips) | Managed video platform or cloud media services | Encoding, packaging, DRM and player SDKs for a per-minute price; your engineers work on the product |
| Under ~10 Gbps of peak egress | Commercial CDN + cloud transcoding service | Egress is a few thousand dollars a month; an owned stack costs more in salaries than it saves |
| Short clips under ~60s, rarely re-watched | Progressive MP4 with 2–3 renditions | ABR and segmenting add complexity that a 15-second clip doesn't need |
| Video calls and real-time interaction (< 500 ms) | WebRTC / SFU architecture | HTTP segment streaming cannot reach interactive latency; different system — see Push vs Poll |
| Internal recordings, compliance archives | Object storage + on-demand transcode | Watched almost never; optimize storage, not delivery |
| Premium licensed content without a DRM team | Buy DRM-as-a-service and a packager | Licensor audits and multi-DRM key management are a specialty; mistakes cost licenses |
🎯 Staff Insight: The build-vs-buy line is set by egress volume and by whether video quality is a competitive differentiator. Until both are true, buying is cheaper and better. See Buy or Build for the general framework.
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It Matters | What to Ask or Assume |
|---|---|---|
| Upload volume and video length | Drives encode fleet and storage add | "500 h/min, average 10 min; peak 2×" |
| Watch volume and device mix | Drives egress and the ladder | "~1B watch-hours/day, ~70% mobile; average delivered bitrate ~2.5 Mbps" |
| Publish-time expectation | Fast path vs batch | "Playable p50 ≤ 2 min, full quality ≤ 15 min" |
| Retention promise | Long-tail storage forever? | "Source kept; rendition set may shrink for dormant videos — need legal sign-off" |
| DRM / rights | License server on the play path; regional blocks | "UGC: no DRM, but region and visibility checks; premium: multi-DRM" |
| Live in scope? | Different latency architecture | "VOD first; live as an extension" |
| Max resolution | 4K rungs are 3–4× the bytes of 1080p | "Up to 2160p for sources that have it, head only" |
| Who pays for egress | Own edge vs commercial | "Assume we're large enough to negotiate and peer" |
2.4 Precise Terminology#
| Term | Precise Meaning | Common Confusion |
|---|---|---|
| Rendition / rung | One encoded version: codec + resolution + bitrate | Not the same as a "format" or container |
| Bitrate ladder | The full set of renditions for a title | Fixed ladder vs per-title ladder |
| Segment | A few seconds of one rendition, independently decodable (starts on a keyframe) | Segment ≠ chunk; "chunk" usually means an encode work unit |
| GOP / keyframe interval | Distance between IDR frames; segments must align to them | Longer GOP compresses better; must divide the segment length |
| Part (LL-HLS) | Sub-segment (~1s) published before the full segment completes | Not a separate rendition |
| Manifest / playlist | Lists renditions (multivariant) or segments (media playlist) | Live manifests change every part; VOD manifests are static once complete |
| CMAF | Common fragmented-MP4 format usable by both HLS and DASH | A container format, not a protocol |
| Byte hit ratio | Bytes served from cache ÷ bytes served | Request hit ratio can be high while byte hit ratio is low (manifests hit, segments miss) |
| Startup time | Play press → first frame rendered | Excludes ad time; includes manifest, license and first segments |
| Rebuffer ratio | Stall time ÷ (stall + play time) | Not "number of rebuffers"; one 10s stall can matter more than five 0.2s ones |
| Glass-to-glass latency | Camera capture → viewer screen | Not segment latency; includes encode, packaging, CDN and player buffer |
| Egress | Bytes leaving your network (or your CDN's) to viewers | The dominant cost; priced per GB or per committed Gbps (95th percentile billing) |
3. Where the Design Splits#
Each fault line below has options, who pays, a Staff default and when to deviate. The question underneath all five is the one from 1.3: for a video watched N times, what is the cheapest way to meet the QoE targets?
3.1 Fault Line 1: Encode Everything Up Front vs Encode on Demand#
The tension: Encoding every rendition at upload makes every first view perfect and wastes compute and storage on videos nobody watches. Encoding on demand saves both and makes the first viewer of a cold rendition wait — or watch a lower rung.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Full ladder, all codecs, at upload | Every view gets the best rendition from day one; simple pipeline | Compute ~6× the H.264-only cost; storage ~2× forever; publish time = slowest rung | Finance (compute + storage), creators (slow publish) |
| Universal ladder up front, advanced codec on popularity | Cheap baseline for everyone; bytes saved where views concentrate | First hours of a viral video served in H.264 (~30% more bytes) | Delivery budget during the ramp |
| Minimal ladder up front, everything else just-in-time | Lowest storage; good for archives | First viewer of a rung waits for an encode or gets a lower rung; JIT encode capacity must handle spikes | Viewers of cold content (quality), on-call (JIT fleet spikes) |
| Just-in-time packaging only | Store one mezzanine per rung, package HLS/DASH/encryption at the edge or origin on request | Saves duplicate packaging; CPU at origin on every miss | Origin team (CPU per miss) |
The arithmetic that sets the threshold. An advanced-codec ladder costs about 5× the H.264 encode: roughly 15 core-hours per video-hour, about $0.30 at $0.02 per core-hour, so ~$0.05 for a 10-minute video. It saves about 30% of bytes. A full view of that video at 2.5 Mbps delivers ~190 MB; 30% saved is ~56 MB, worth ~$0.0003 at a blended $0.005/GB. Compute alone breaks even at ~180 full views. Two corrections push the number up: the average view watches maybe 40% of the video (→ ~450 views), and the extra rungs (~0.45 GB for 10 minutes) cost ~$0.05 a year in warm storage, roughly doubling the cost to recover. That lands near 1,000 views — hence the threshold of 1,000 views in 24 hours, lowered when the platform predicts the video is trending.
Staff default: universal H.264 ladder for every upload, fast path first; advanced codec driven by view velocity; dormant videos shrink to two rungs.
When to deviate:
- Premium catalogue: encode everything exhaustively up front — every title is in the head.
- Archive or compliance video: store the source only; encode on first request and accept a 30–60s wait.
- Encode hardware makes advanced codecs cheap: if per-hour encode cost drops 10–20×, the threshold drops by the same factor and "encode everything" becomes reasonable again.
🎯 Staff Move: "I won't pick 'encode everything' or 'encode on demand' on principle. Advanced-codec encoding costs about 30 cents per video-hour and saves about 0.17 cents per watch-hour; with partial viewing and a year of storage for the extra rungs, it pays back at roughly 1,000 views. That's my threshold, and I'd re-derive it whenever compute or egress pricing changes."
3.2 Fault Line 2: Fixed Ladder vs Per-Title or Per-Shot Encoding#
The tension: A fixed ladder (same bitrates for every title) is predictable and cheap to compute. It overspends bits on simple content — slides, animation, a talking head — and underspends on grainy, high-motion content. Per-title and per-shot encoding fix both by running analysis or trial encodes, which costs compute per title.
| Strategy | Extra Compute | Bandwidth Saved | Fits |
|---|---|---|---|
| Fixed ladder | 0 | 0 (baseline) | The tail; live |
| Complexity-classified ladder (3–5 ladder templates picked by a fast analysis pass) | +5–10% | Large on easy content (screen recordings, animation) | UGC baseline |
| Per-title convex hull (trial encodes at several resolutions/bitrates, pick the efficient frontier) | +2–5× | Typically 20–40% at equal perceptual quality for content that differs from the "average" title | The head of UGC; premium |
| Per-shot optimization | +5–10× | Further gains over per-title by allocating bits shot by shot | Premium catalogue; the very top of UGC |
Who pays: the encode budget pays up front; the delivery budget and the viewer's data plan collect the savings on every view. Live can't use either — there's no time for trial encodes — so live ladders are fixed and slightly over-provisioned.
Measuring quality: you need a perceptual metric (VMAF, SSIM or an internal model) to say "equal quality". Without one, per-title encoding is guesswork and every ladder change is an argument. Treat the metric pipeline as a prerequisite, owned by the media pipeline team.
Staff default: complexity-classified ladders for every upload (cheap, catches screen recordings and animation), full per-title encoding for the head, per-shot never for UGC unless hardware changes the cost.
🎯 Staff Insight: Per-title encoding is a bandwidth project funded by the encode budget. Present it with both numbers — extra core-hours and saved petabytes — or it will lose every prioritization meeting to features.
3.3 Fault Line 3: Own CDN vs Commercial CDN vs Both#
The tension: A commercial CDN is instant, global and someone else's pager. At very large egress volumes, owning cache servers — in your own points of presence and, further, embedded inside ISP networks — lowers unit cost and improves QoE, but it is a multi-year capital program with hardware, logistics and peering negotiations.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Single commercial CDN | Day-one global reach; no capex | Unit cost at scale; correlated outage; limited control of steering | Finance (per-GB), everyone during a provider outage |
| Multi-CDN with steering | Price leverage, failover, per-region best performer | Steering logic, config drift, cache warmed N times | Delivery team (complexity) |
| Owned edge + commercial overflow | Lowest unit cost on the bulk of traffic; control of cache policy | Capex, hardware refresh every 3–5 years, ops team, peering | Company (capital), infra org (headcount) |
| ISP-embedded appliances | Bytes never cross transit; best QoE; ISPs save transit too | Only viable with traffic big enough that ISPs want the boxes; logistics in hundreds of networks | Company (multi-year program) |
Rough break-even. At ~$0.01/GB commercial, 1 Tbps sustained average is 10.8 PB/day ≈ $108K/day ≈ $40M/year. An owned edge that serves the same traffic at $0.002–0.004/GB all-in (hardware amortized, colocation, transit/peering, staff) saves roughly $25M/year per Tbps — once the fixed cost of a 40–80 person edge organization ($15–25M/year) is covered. Below a few Tbps average, buy. Above tens of Tbps, owning wins if you can execute. The details of cache design are in Content Delivery Network; the decision framework is in Buy or Build.
Staff default for our design (190 Tbps peak): hybrid. Owned edge and ISP peering carry the steady head; one or two commercial CDNs carry overflow, new markets and the long tail where our caches would thrash.
When to deviate: a startup or mid-size platform should use one commercial CDN with committed pricing and add a second only when spend or a provider outage justifies it.
3.4 Fault Line 4: Segment Duration — Latency vs Efficiency and Request Rate#
The tension: Shorter segments let live streams run closer to real time and let ABR react faster. Longer segments mean fewer HTTP requests, fewer keyframes (better compression), smaller manifests and better cache efficiency.
| Segment Length | Requests per Viewer-Hour | At 75M Concurrent | Live Latency (≈ 3 segments buffered) | Compression Cost |
|---|---|---|---|---|
| 2 s | 1,800 | ~38M req/s | ~6–10 s | Keyframe every 2s: a few % more bits |
| 4 s | 900 | ~19M req/s | ~12–15 s | Baseline for this design |
| 6 s (Apple's default target) | 600 | ~12.5M req/s | ~18–30 s | Slightly better than 4s |
| 10 s | 360 | ~7.5M req/s | 30 s+ | Best; slow ABR reaction |
| LL parts 1 s inside 4–6 s segments | 3,600 part requests (or fewer with blocking preload) | Edge must hold requests open | ~3–5 s | Same encode; packaging and CDN cost rise |
The request rate matters because CDNs charge per request at some tiers, because each request has a fixed overhead in TLS and HTTP processing at the edge, and because manifests for live are re-fetched every part.
Who pays: short segments are paid for by the edge (request rate) and by bandwidth (keyframe overhead); long segments are paid for by viewers on bad networks (slower adaptation → more stalls) and by live audiences (latency).
Staff default: 4 s segments with 2 s keyframes for VOD (two GOPs per segment gives the packager flexibility); 6 s segments with 1 s parts for low-latency live where latency is in the product promise; standard 6 s segments for live where 20–30 s is acceptable — most live content does not need LL.
🎯 Staff Move: "Segment length is a cost knob, not a latency knob. Halving it doubles request rate across 75 million viewers. I'll use 4-second segments for VOD and reserve 1-second parts for the live events where latency is part of the product, like sports betting or interactive streams."
3.5 Fault Line 5: Client-Driven ABR vs Server-Guided Steering#
The tension: The player knows its buffer level and recent throughput; the platform knows which CDN is healthy, which ISP is congested, what each byte costs and what a million other players are seeing. Pure client ABR is scalable and robust but myopic. Pure server control is informed but adds a control plane to the play path and cannot see the device's buffer.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Throughput-based client ABR | Fast startup decision | Noisy estimates → oscillation; over-estimates on bursty mobile links | Viewers (stalls, quality flapping) |
| Buffer-based client ABR | Stable in steady state; few stalls | Slow start; under-uses capacity at startup | Viewers (low quality early) |
| Hybrid client ABR (throughput at start, buffer in steady state) | Industry default | Needs tuning per device class | Player team |
| Server-side steering hints (CDN choice, rung cap per ISP during congestion, start rung) | Uses fleet-wide knowledge; can shed load during incidents | A control-plane dependency; must degrade to client-only | Playback platform team |
| Server-side ABR decision per segment | Total control | Adds latency to every segment; doesn't see device buffer | Everyone |
Staff default: hybrid client ABR with hysteresis (up-switch needs headroom and a hold time; down-switch is immediate), plus server hints delivered with the play response and refreshed every few minutes: preferred CDN order, a starting rung, and an optional bitrate cap per ISP or region during congestion. If the hint service is down, the player uses cached hints or none — never blocks playback. This is a degraded-mode design: the server can improve playback but cannot stop it.
🎯 Staff Insight: The server-side bitrate cap is the most powerful incident lever you have. During an ISP congestion event, capping that ISP's viewers at 720p cuts their egress by ~45% and usually removes the congestion that caused the stalls. Decide in advance who may pull it — the delivery on-call — and who must be told — partnerships, because the ISP will ask.
4. When It Breaks#
4.1 Transcoding Backlog After a Viral Upload Spike#
A global event (an election result, a celebrity incident, a natural disaster) produces a 3× upload surge in 40 minutes, concentrated in long phone videos.
t=0: Upload rate 500 → 1,400 h/min. Encode demand ~2.8× fleet capacity.
t=+5min: encode.queue_age_p99 1min → 9min. Fast-path and ladder tasks share one queue.
t=+12min: Fast-path p50 publish time 90s → 14min. Creators retry uploads (duplicates +20%).
t=+15min: Page: encode.fastpath_queue_age_p99 > 5min for 5m.
t=+18min: On-call: pause advanced-codec and re-encode classes (frees ~30% of fleet).
t=+20min: Scheduler moves full-ladder tasks behind fast path. Autoscaler requests
preemptible capacity: +40% in 25 min.
t=+45min: Fast path back under 2 min. Ladder backlog 6 hours, drains overnight.
t=+1 day: Post-mortem: duplicate uploads (same content hash) were encoded twice.
Detection: encode.queue_age_p99{class}, publish.time_to_playable_p50, encode.tasks_pending{class}, upload.duplicate_content_ratio.
Mitigation: strict priority classes — fast path, ladder, advanced codec, re-encode/backfill — with the lower classes preemptible; dedup by content hash before encoding; shed the advanced-codec class first because it is pure optimization.
Prevention: capacity model with 2× headroom for the fast path only (cheap: it's 2 rungs); preemptible spot capacity for everything else; quarterly load test replaying a 3× surge. Scheduling mechanics are in Job Scheduler and Managing Long-Running Processes.
Owner: media pipeline on-call; capacity planning owns the headroom model.
4.2 Cold Cache on a Premiere Night#
A scheduled episode releases at 20:00 local. 8M viewers press play in the first five minutes.
t=-24h: Nobody pre-warmed: the release workflow publishes metadata at 20:00 only.
t=0: 8M play requests in 300s. Every edge cache is cold for every rung.
t=+5s: Edge miss rate 100% for segment 1. Mid-tier coalesces per region, but 600 edge
sites × 6 rungs × first 5 segments each go upstream.
t=+20s: Origin egress 0.4 → 5 Tbps. Origin TTFB p99 300ms → 4s.
t=+30s: Startup p95 2s → 9s. Players time out and retry → miss traffic doubles.
t=+2min: Page: qoe.startup_time_p95 > 5s and origin.egress_gbps > 3x baseline.
t=+4min: Edge caches warm. Hit ratio 97%. Startup recovers. 11% of viewers had given up.
Detection: qoe.startup_time_p95{title}, cdn.byte_hit_ratio{region}, origin.egress_gbps, origin.ttfb_p99, player.start_failure_rate.
Mitigation: request coalescing at edge and shield; origin rate limiting with Retry-After honored by players (with jitter); start viewers on a lower rung for the first 10 seconds during a known spike.
Prevention: release workflow publishes segments to origin hours early and pre-warms the first 60 seconds of the top 4 rungs into every edge site; metadata flips visibility at 20:00. For premium releases, pre-position the whole title overnight.
Owner: delivery team for coalescing and pre-warm; content operations owns the release checklist.
4.3 ABR Oscillation Storm#
A player release changes the throughput estimator to a shorter window. On a congested ISP, a million players see the same dip at the same time.
t=0: Player v8.2 ships to 30% of Android. Estimator window 10s → 3s.
t=+2h: Evening peak on ISP-X. Link utilization 92%.
t=+2h01: Players on ISP-X see throughput dip, switch 1080p → 480p together.
t=+2h02: Link frees up. Players estimate high throughput, switch back to 1080p together.
t=+2h03: Link saturates again. Repeat every ~40s.
t=+2h20: qoe.bitrate_switches_per_min{isp=X} 0.3 → 4.1. Rebuffer ratio 0.2% → 1.8%.
t=+2h25: Page on rebuffer ratio by ISP. Delivery on-call caps ISP-X at 720p via hints.
t=+2h30: Oscillation stops. Rebuffer 0.3%. Player rollback started.
Detection: qoe.bitrate_switches_per_min{isp,player_version}, qoe.rebuffer_ratio{isp}, cdn.egress_gbps{isp} showing a sawtooth.
Mitigation: server-side rung cap for the affected ISP; halt the player rollout.
Prevention: up-switch hysteresis (headroom + 20s hold) and randomized hold times so players desynchronize; player releases staged by ISP and device with QoE guardrails per stage — a player is a fleet of 75M distributed control loops.
Owner: player team for the algorithm and rollout; delivery on-call for the cap.
4.4 CDN Region Failure Causes an Origin Stampede#
A commercial CDN loses a large region; traffic shifts to our owned edge and a second CDN, both cold for that region's long tail.
t=0: CDN A loses region EU-West. DNS/steering moves 3 Tbps to CDN B + owned edge.
t=+30s: New caches miss on the tail. Origin egress 0.9 → 9 Tbps (10×). Capacity: 3×.
t=+1min: Origin TTFB p99 8s. Startup failures 6%. Players retry → load climbs.
t=+3min: Origin protection: priority shedding — lowest 3 rungs always served; 1080p+
requests get 429 Retry-After 10s with jitter. Steering caps region at 720p.
t=+15min: Caches warm. Origin 3.5 Tbps. Caps lifted in 25% steps over 20 min.
Detection: cdn.availability{provider,region}, origin.egress_gbps, origin.shed_requests_total, player.start_failure_rate{region}.
Mitigation: priority shedding at origin by rung; steering-based rung cap; stagger traffic moves (25% per minute) rather than flipping a region at once.
Prevention: keep the secondary CDN partially warm by sending it a steady 10–20% share; size origin for the largest single-region failover, not for average miss rate; game day twice a year.
Owner: delivery team; vendor management for the CDN incident report.
4.5 Storage Costs Growing Faster Than Views#
The slow failure: nothing pages, but the storage line grows 40% a year while watch time grows 15%.
What's happening: every upload adds ~3 PB/day of renditions and ~2.6 PB/day of source. Without tiering, year 3 holds ~6 EB, and the bytes are overwhelmingly in videos that earn almost no views. Watch time is concentrated in recent and popular videos; storage is concentrated in old and unpopular ones.
Detection: storage.bytes{tier}, storage.bytes_per_watch_hour (the ratio that should be flat or falling), storage.cost_growth_vs_watchtime_growth reviewed monthly.
Mitigation: lifecycle policy by age and views: source to archive at 30 days (it's only needed for re-encodes); rendition sets shrink as videos cool; advanced-codec rungs deleted after 180 days unserved; dormant videos keep two rungs. Re-encoding from archive on a spike costs minutes of compute — cheap next to years of storage.
Prevention: storage cost per watch-hour as a quarterly KPI with a named owner; retention policy agreed with legal and creator relations. Storage mechanics are in Object Storage.
Owner: media platform for the policy engine; finance partner for the KPI; legal for the retention promise.
4.6 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Encode backlog | encode.queue_age_p99{class=fastpath} > 5min | New uploads (creators) | Priority classes, shed advanced codec, preemptible burst | Media pipeline |
| Cold premiere | qoe.startup_time_p95{title} > 5s | One title's audience | Coalescing, pre-warm, lower start rung | Delivery + content ops |
| ABR oscillation | qoe.bitrate_switches_per_min{isp} > 2 | One ISP × player version | Rung cap hint, rollout halt | Player team / delivery on-call |
| CDN region loss | cdn.availability{provider,region} < 99% | A region's viewers | Staggered steering, origin shedding by rung | Delivery |
| Storage drift | storage.bytes_per_watch_hour rising 3 months | Margin | Lifecycle tiers, deletion policy | Media platform + finance |
| License server slow | drm.license_latency_p99 > 500ms | All DRM titles' startup | Cache licenses per session, regional license replicas | Playback platform |
| Manifest error | player.manifest_parse_errors spike | Every viewer of affected titles | Roll back packager; immutable segments let you republish manifests | Media pipeline |
| Live ingest drop | live.ingest_gap_seconds > 2 | One live channel | Dual ingest with automatic failover; slate | Live team |
🎯 Staff Insight: Two classes of failures dominate: the ones you see in minutes (cold cache, CDN loss, ABR storms), which are all versions of "origin load is proportional to the miss rate"; and the one you see in quarters (storage drift), which is a policy failure. The first class needs coalescing and shedding; the second needs an owner and a KPI.
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | Lists features: upload, transcode, play | Names UGC vs premium vs live; commits; states cost thesis with an egress number | Asks what share of revenue delivery consumes and which lever moves it over 3 years |
| Encoding | Fixed ladder for every upload | Fast path, priority classes, popularity threshold derived from break-even math | Funds codec or hardware programs with a payback model; owns the efficiency roadmap |
| Delivery | "CDN in front" | Shield, coalescing, byte hit ratio as SLO, origin headroom for failover | Own vs buy vs hybrid edge as a multi-year bet; ISP relationships; CDN contract structure |
| QoE | "Adaptive bitrate" | Startup, rebuffer, start failures — SLOs sliced by ISP/device/CDN; hysteresis | QoE tied to watch time and revenue; decides bitrate caps vs egress with product and finance |
| Failure | Retries and replicas | Spike, cold cache, ABR storm, CDN loss — each with detection and shedding | Game days, staggered failover policy, correlated risk across shared edge for live and VOD |
| Organization | One video team | Pipeline, player, delivery, content ops with handoff SLOs | Platform contract shared by VOD, shorts and live; retention policy with legal |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Leads with cost | "Egress dominates. At 1 EB a day, a cent per GB is about $4 billion a year." |
| Encodes by value | "Advanced codec pays back after about 1,000 views; below that it's wasted compute." |
| Protects origin by design | "Origin load is proportional to the miss rate, so byte hit ratio per region is a paging metric." |
| Quantifies segment choices | "2-second segments double request rate versus 4. I'll pay that only for live." |
| Separates publish from quality | "Playable in 90 seconds with two rungs; full quality fills in." |
| Owns the long tail | "Renditions are a cache of the source. Dormant videos shrink to two rungs." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Spends 10 minutes on the metadata database | Metadata is a few thousand QPS; the hard parts are bytes |
| "Stream the file from our servers" | Ignores CDN, egress economics and range/segment delivery |
| Encodes every upload in every codec at every resolution, no numbers | No sense of compute or storage cost at 720K hours a day |
| "Short segments are better" without a request-rate number | Treats a cost knob as free |
| No QoE metrics | Can't tell if the system works for viewers |
| WebRTC for 75M viewers of VOD | Wrong tool: per-viewer server state, no HTTP caching |
5.4 Common False Positives#
- Codec trivia ≠ system design. Explaining B-frames, CABAC and motion vectors impresses briefly; if it doesn't lead to a cost or publish-time decision, it's trivia.
- Naming HLS and DASH ≠ understanding delivery. The signal is what segment length does to request rate and latency.
- "Netflix built Open Connect" ≠ a build decision. Without traffic volume and break-even reasoning, citing it suggests copying a company 1,000× your size.
- Kubernetes for transcoding ≠ scheduling design. Running workers on Kubernetes is fine; the signal is priority classes, preemption and chunk-level retry.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Pick UGC/premium/live; numbers; cost thesis |
| Entities & API | 3–5 min | Video, rendition, encode task; immutable segments; play API |
| Architecture | 5–10 min | Two lanes: pipeline and delivery; ≤ 10 boxes |
| Encoding economics | 10–20 min | Fast path, ladder, popularity threshold, priority classes |
| Delivery | 20–30 min | Shield, coalescing, hit ratio, spikes, CDN strategy |
| ABR & QoE | 30–37 min | Segment length, hybrid ABR, SLOs, server hints |
| Pivot | 37–42 min | Live, premium/DRM, multi-region, cost cut |
| Wrap | 42–45 min | Cost levers, one-way doors, what's next |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Now make it live" | Can you rebuild for latency? | Real-time ladder at ingest, 1s parts, dual ingest, coalescing at every tier |
| "We're Netflix now" | Does intent change the design? | Exhaustive per-title encoding, pre-positioning, multi-DRM license server on the play path |
| "Cut delivery cost 30%" | Cost levers in order | Advanced codec for the head, per-title for easy content, hit ratio, bitrate caps on small screens, own/peer edge |
| "A video gets 5M views in an hour" | Spike handling | Coalescing, priority re-encode, advanced codec queued, no origin stampede |
| "Viewers in one country rebuffer" | Operational forensics | Slice QoE by ISP/CDN/player version; steering; ISP congestion vs CDN fault |
| "Shorts: 30-second vertical videos" | Workload shift | Prefetch next videos, progressive or 2s segments, startup time dominates, fewer rungs |
6.3 What to Deliberately Skip#
- Upload resumability details — one sentence; link to Handling Large Blobs.
- Codec internals — name codecs and their compute/byte tradeoff only.
- Recommendations, comments, search — scoped out; mention that view counts feed popularity.
- View-count accuracy — a separate aggregation problem; see Ad Click Aggregation for the pattern.
- Metadata sharding — the metadata DB is small; a sharded PostgreSQL or Cassandra is fine, and saying why you're skipping it is the signal.
6.4 Follow-Up Questions to Expect#
- "How long does a 1-hour 4K upload take to become playable, and how would you make it faster?"
- "How much storage do you add per day, and what's the bill in year three?"
- "What does origin see when a premiere starts at 20:00 for 8 million viewers?"
- "Why 4-second segments? What changes at 2 seconds? At 10?"
- "How does the player decide which rendition to fetch, and how do you stop it oscillating?"
- "At what scale would you build your own CDN?"
- "How do you get live latency under 5 seconds, and what does it cost?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design YouTube."
Staff Answer
"There are three products hiding in that prompt — user uploads at scale, a premium catalogue, and live — and they differ in where the cost and risk sit. I'll design user uploads and call out where premium and live diverge.
Numbers I'll use: 500 hours uploaded a minute, about a billion watch-hours a day, 70% mobile, average delivered bitrate ~2.5 Mbps. That's over 1 EB of egress a day and roughly 190 Tbps at peak — so delivery cost is the line item that matters, and the metadata database is not the hard part. Targets: playable within 2 minutes of upload, startup p95 under 2.5 seconds, rebuffer ratio under 0.3%. I'll walk through: the encode pipeline and how much encoding each video deserves, delivery and origin protection, ABR and segment length, then failure modes and owners."
Why this is L6:
- Commits to one intent and says what the other two would change
- Turns "scale" into egress bytes and dollars in the first minute
- Deprioritizes the metadata DB with a reason
What L7 adds:
- Asks what share of revenue goes to delivery and what the 3-year target cost per watch-hour is
- Asks whether live, shorts and VOD share one pipeline and edge today — the platform question
- Notes that retention promises to creators are a policy decision, not an engineering default
❌ Common L5 Trap
"Users upload to an upload service which stores the file in S3, a transcoding service converts it to multiple resolutions, metadata goes in a sharded MySQL, and a CDN serves the videos. We'll use Kafka between services and autoscale everything."
Why this misses: Every box is right and there's not a single number. It can't answer "how long until playable", "what does it cost", or "what does origin see at 20:00". The interviewer has to drag every tradeoff out.
Drill 2: The Core Mechanic — Publish Time#
Prompt: "A creator uploads a 1-hour 4K video. How long until viewers can watch it, and how do you make that fast?"
Staff Answer
"Upload time is the creator's network; I'll talk about what happens after complete. Encoding a 1-hour 4K source sequentially into a full ladder would take hours. Instead, the scheduler splits the source at keyframes into ~4-second chunks — about 900 chunks — and fans out tasks per chunk per rung. The fast path is 360p and 720p H.264: 1,800 small tasks, each a few seconds of work, run in parallel across the fleet, then a stitch step writes the manifest. That's playable in a few minutes even for a 1-hour video. The rest of the H.264 ladder, including 1080p, fills in over the next 10–20 minutes at normal priority. 1440p and 2160p in an advanced codec run only if the video gets traction, because those rungs are the most expensive to encode and store.
Two details make this work: chunks must split on closed GOPs so they encode independently, and each chunk task is idempotent — keyed by (video_id, rung, chunk_idx) — so a preempted worker just means a retry of one 4-second chunk."
Why this is L6:
- Separates time-to-playable from time-to-full-quality
- Quantifies chunk count and parallelism
- Makes preemption cheap by design (chunk-level idempotent retry)
What L7 adds:
- Sets a creator-facing SLO (p50 ≤ 2 min for 10-minute uploads) and reports it per region
- Asks whether hardware encoding changes the fast path economics enough to include 1080p
Drill 3: Make It Concrete — Storage per Day#
Prompt: "How much storage do you add per day? What about year three?"
Staff Answer
"720K hours uploaded a day. Source at ~8 Mbps is ~3.6 GB per hour → ~2.6 PB/day. The H.264 ladder from 240p to 1080p totals ~9.2 Mbps, about 4.1 GB per video-hour → ~3 PB/day. Advanced-codec rungs for the ~10% of videos that cross the popularity threshold add ~2.7 GB per hour for those → ~0.2 PB/day. Call it ~6 PB/day, ~2 EB/year, ~6 EB in year three if we keep everything at full fidelity.
That's why the storage design is a lifecycle, not a bucket: source moves to an archive tier after 30 days (it's only needed for re-encodes), cooling videos shrink their rendition set, and dormant videos keep two rungs. With that, most of the bytes sit in tiers that cost a fraction of standard storage, and the metric I'd watch is storage bytes per watch-hour — it should be flat or falling."
Why this is L6:
- Derives the number from stated bitrates rather than guessing
- Separates source (asset) from renditions (cache)
- Names a ratio metric that catches drift
What L7 adds:
- Turns retention into a written policy with legal and creator relations
- Models the year-3 bill at each tiering policy and picks one with finance
Drill 4: Dependency Down — The Origin Is Struggling#
Prompt: "Origin TTFB p99 just went from 200ms to 6 seconds during peak. What do you do?"
Staff Answer
"First question: is origin slow because of more misses, or slow on its own? I check cdn.byte_hit_ratio by region and provider. If hit ratio dropped — a CDN region failed over, a cache purge, a premiere — it's a load problem. If hit ratio is normal, it's an origin fault: storage latency, a bad deploy, a network path.
For a load problem: confirm shields are coalescing; cap the affected region at 720p via steering hints, which cuts bytes per view ~45%; shed by rung at origin — always serve the bottom three rungs, return 429 with Retry-After and jitter for the top rungs. Players fall back a rung instead of stalling. Then lift caps in 25% steps.
For an origin fault: fail origin reads over to the replica region — renditions are replicated because they're immutable and cheap to copy — and roll back whatever changed."
Why this is L6:
- Diagnoses miss-driven load vs origin fault before acting
- Uses quality degradation (rung caps) as a load-shedding tool
- Recovers in steps to avoid a second stampede
What L7 adds:
- Makes "largest single-region failover" the origin sizing rule, written into the capacity standard
- Runs twice-yearly game days with each CDN provider
Drill 5: Hot Key — One Video, 5M Viewers in an Hour#
Prompt: "A video uploaded 20 minutes ago goes viral: 5 million viewers in an hour. What breaks?"
Staff Answer
"Three things, in order. Delivery: every edge site misses on every segment the first time. Coalescing at the edge and the regional shield means one origin fetch per segment per region — about 50 regional shields × 5 rungs × 150 segments for a 10-minute video — trivial for origin. The danger is if coalescing is off or misconfigured: then 5M viewers × first segment is a stampede.
Encoding: the video only has the H.264 ladder. View velocity crosses the threshold in minutes; the scheduler queues the advanced-codec ladder at high priority. That saves ~30% of egress for the rest of its life — at 5M views that's tens of TB. Metadata: the play API reads video metadata 5M times in an hour — ~1,400 req/s for one key — cache it in the play service with a 30-second TTL.
What doesn't break: transcoding the viral video itself — it's already playable."
Why this is L6:
- Identifies coalescing as the mechanism that turns a hot object into a non-event
- Connects popularity to encoding effort in real time
- Names the metadata hot key with a number and a cheap fix
What L7 adds:
- Asks whether the recommendation system can signal "about to promote" so delivery pre-warms before the spike
- Treats hot-object handling as a shared edge capability used by live as well
Drill 6: Multi-Tenant — Live and VOD on One Edge#
Prompt: "We're adding live sports on the same CDN and edge fleet as VOD. What do you worry about?"
Staff Answer
"Live and VOD want opposite things from the same caches. Live is a tiny working set — the newest few parts of each rung — requested by millions within the same second; VOD is a huge working set with a long tail. Without isolation, a big match can evict VOD's warm tail, and VOD's tail churn can slow live's hot parts.
So: separate cache namespaces or capacity reservations per class; live parts in memory with very short TTLs, VOD segments on disk; separate shields for live so a live surge doesn't saturate VOD's origin path. Live manifests and parts need request collapsing that holds requests until the part exists — blocking playlist reload in LL-HLS — or the origin sees a request per viewer per second.
Ownership: live has its own on-call and SLOs (glass-to-glass, ingest gaps), but both run on a shared edge platform with per-class quotas."
Why this is L6:
- Explains why the workloads conflict (working set and timing)
- Isolates by capacity and namespace rather than separate fleets
- Names the live-specific coalescing mechanism
What L7 adds:
- Defines the edge as a platform with per-product quotas and chargeback
- Decides when a marquee event justifies dedicated capacity or a commercial CDN overflow contract
Drill 7: Build vs Buy — Our Own CDN#
Prompt: "Finance says the CDN bill is our second-largest expense. Should we build our own?"
Staff Answer
"Depends on volume and how predictable it is. At ~$0.01/GB, every 1 Tbps of sustained average traffic is about $40M a year. An owned edge — servers in colocation and peering at exchanges, later caches inside ISPs — might deliver at $0.002–0.004/GB all-in, but it carries a fixed cost: a 40–80 person organization, hardware refresh every 3–5 years, and years to reach coverage. Below a few Tbps average I'd negotiate harder, add a second CDN for leverage, and turn on the advanced codec for the head — that's a 30% byte cut with no capex. Above tens of Tbps, a hybrid wins: own the bulk in our top markets, keep commercial CDNs for overflow, new markets and the tail. I'd start with the five metros that carry the most traffic, measure unit cost and QoE against the CDN, and expand only if the numbers hold."
Why this is L6:
- Gives the break-even arithmetic, including fixed cost
- Lists cheaper levers to pull first
- Proposes an incremental, measurable path instead of a big bang
What L7 adds:
- Treats it as a 5-year capital program with board-level visibility
- Plans ISP relationships and peering policy as a partnerships function, not just engineering
- Keeps a commercial contract as a strategic hedge
Drill 8: Policy Change Without an Outage — New Codec Rollout#
Prompt: "We want to move the head of the catalogue from VP9 to AV1. How do you roll it out?"
Staff Answer
"Two separate rollouts: encoding and playback. Encoding: start re-encoding the top 1% of videos by watch time in AV1, in the backfill priority class so it never competes with uploads. Playback: the manifest advertises AV1 only to devices that report hardware decode support — software AV1 decode drains mobile batteries — so the device capability list is the gate. Stage by device family: 1% → 10% → 50% → 100%, with QoE guardrails per stage: startup time, rebuffer ratio, start failures, and battery-related abandonment where we have it. Keep the VP9 rungs until AV1 has run at 100% for a quarter, then let the lifecycle policy delete them as they go unserved.
Success metric: egress per watch-hour on AV1-capable devices, compared against a VP9 holdout."
Why this is L6:
- Separates encode capacity from playback exposure
- Gates on hardware decode capability
- Stages with QoE guardrails and a holdout to prove the saving
What L7 adds:
- Builds the business case: AV1-capable device share × byte saving × egress price vs re-encode compute
- Coordinates with device partners and OEM decode support roadmaps
Drill 9: Cost — Cut Delivery 30%#
Prompt: "The CFO wants delivery cost down 30% next year without hurting QoE. Where do you look?"
Staff Answer
"Attribute first: bytes by codec, rung, device class, region, CDN. Then the levers in order of cost-to-implement:
- Advanced codec on more of the head — lower the view threshold; ~30% fewer bytes on whatever share of watch time moves.
- Complexity-classified ladders — screen recordings and animation are often over-encoded by 2×.
- Device-aware caps — a phone in portrait doesn't need 1080p; capping small screens at 720p cuts their bytes ~45% with little perceptible loss. Product signs off.
- Hit ratio — every point of edge hit ratio moves bytes from expensive paths (origin, transit) to cheap ones.
- Contract mix — shift volume to the cheapest CDN per region; renegotiate commits.
- Owned edge — only if volume justifies it; it's a multi-year lever, not next year's.
I'd model each one's saving, cost and QoE risk and pick a portfolio that reaches 30% with a holdout for every change."
Why this is L6:
- Attributes before cutting
- Orders levers by cost and risk
- Names who signs off on quality-affecting changes
What L7 adds:
- Sets a multi-year target for cost per watch-hour and makes it a standing KPI
- Recognizes that bitrate caps are a product decision with brand risk
Drill 10: Multi-Region — Rights and Residency#
Prompt: "We're launching in a country that requires certain content to be blocked and some user data to stay in-country."
Staff Answer
"Two different requirements. Content rights are enforced at the play API: the video's rights.regions is checked against the viewer's region before a signed manifest is issued, and signed URLs are short-lived so they can't be shared across borders easily. That's a control-plane check; segments themselves are immutable and can be cached anywhere unless the license forbids it — premium licensors sometimes do, in which case those titles' renditions are only placed on in-country edge.
Data residency usually applies to user data — accounts, watch history, QoE beacons with identifiers — not to public video bytes. So: in-country storage for user-linked data, aggregated non-identifying QoE metrics exported to the global pipeline, and an in-country origin replica only if required for licensed content."
Why this is L6:
- Separates content rights from data residency
- Enforces rights in the control plane, not in caches
- Exports aggregates instead of moving user data
What L7 adds:
- Builds a rights and residency policy engine shared by every product line
- Puts legal review into the market-launch checklist with a standard template
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — The Championship Final#
Context: A live final peaks at 11M concurrent viewers, 4× the previous record. At the 80th minute, startup failures hit 9% and existing viewers report stalls. The on-call escalates to you.
Questions to Surface First:
- Is it ingest (the stream itself), packaging, origin, or edge?
live.ingest_gap_secondsvsorigin.ttfb_p99vscdn.byte_hit_ratio. - Are failures concentrated on one CDN, ISP or device class?
- Are players requesting parts that don't exist yet (clock skew, too-aggressive LL settings)?
- Did anything change: player version, packager config, steering weights?
Typical L5 Approach: Scales origin servers horizontally and asks the CDN to add capacity. Origin recovers partially; stalls continue because the edge is the bottleneck.
Staff Approach: Finds that one CDN's edge is returning 404s for parts requested ~1s before they exist, and the players retry immediately — a retry storm into origin. Shifts 30% of traffic to the second CDN in steps, caps the top rung, and turns on blocking playlist reload so requests wait at the edge instead of retrying.
Principal Approach: Treats it as a readiness failure: the event was 4× the record with no load test at that scale. Establishes an event-readiness review for anything forecast above 2× the previous peak — capacity reservations with each CDN, a rehearsal with synthetic load, and a pre-agreed degradation ladder signed off by the sports partnership owner.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Slice start failures by CDN: 85% on CDN A. Edge logs show 404 on not-yet-existing parts, then immediate retries. |
| Triage | Player version 9.1 lowered the live-edge offset to 2 parts; clock skew on some devices makes them request ahead. |
| Quick fix | Steering: CDN A 70% → 40% in 10% steps. Server hint raises live-edge offset to 3 parts. Rung cap at 720p for 15 min. |
| Guardrails | Players must back off with jitter on 404 for future parts; edge configured to hold requests for parts within 2× part duration. |
| Post-mortem | Why did a latency-tightening player change ship the week of the final? Why was there no 4× load test? |
Metrics to Watch: live.start_failure_rate{cdn}, cdn.status_404_rate{cdn,content=live}, origin.requests_per_sec{live}, live.glass_to_glass_p50
Organizational Follow-up: player release freeze 7 days before marquee events; event-readiness review with capacity reservations.
Ownership Question: "Who decides to trade latency for stability mid-event?" Staff answer: The live on-call, using a pre-approved degradation ladder (raise live-edge offset → cap top rung → shift CDN weights). Product signed off on the ladder in advance, so nobody negotiates during the incident.
Key Takeaway: "In live, a request for something that doesn't exist yet is the most expensive request you'll serve. Hold it at the edge; never let it become a retry."
What clears the Staff bar:
- Slices by CDN and player version before scaling anything
- Identifies the retry storm as the amplifier
- Uses a pre-approved degradation ladder
Deep Dive 2: Silent Failure — The Ladder That Stopped at 720p#
Context: A creator support ticket says "my videos are blurry." Investigation shows that for 9 days, about 18% of uploads never got 1080p. No alert fired: every video was PLAYABLE.
Questions to Surface First:
- Which uploads: by source codec, resolution, device, region?
- Did 1080p tasks fail, never get created, or get created and starve?
- Why didn't an alert fire — what does "success" mean in our metrics?
Typical L5 Approach: Finds a worker crash on a new phone's HEVC source variant, fixes the decoder flag, re-runs the failed tasks.
Staff Approach: Fixes the crash, then fixes the definition of success: the pipeline measured "video reached PLAYABLE" but nothing measured "video reached its expected ladder." Adds
encode.ladder_completeness— the share of videos with every expected rung within 1 hour — as a paging SLO, and a reconciler that compares expected vs actual rungs hourly.
Principal Approach: Generalizes to every derived asset: thumbnails, captions, audio tracks. Establishes "expected vs actual derived assets" as a platform-level reconciliation with a dashboard per asset type, and adds a canary corpus of source files from new devices to the pipeline's release checks.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Query videos older than 1h with missing 1080p rung: 410K. All sources from two new phone models. |
| Triage | The 1080p task fails on a 10-bit HEVC source; after 3 attempts it goes to a dead-letter queue nobody watches. |
| Quick fix | Decoder fix; replay the DLQ at backfill priority; prioritize videos by views. |
| Guardrails | encode.ladder_completeness < 99.5% pages; DLQ depth per task type alerts at > 1,000. |
| Post-mortem | PLAYABLE was the only success signal; DLQ had no owner. |
Metrics to Watch: encode.ladder_completeness, encode.dlq_depth{task_type}, encode.task_failure_rate{source_codec}
Organizational Follow-up: DLQs get an owner and an SLA on creation; new-device source samples are added to the pipeline test corpus each quarter.
Ownership Question: "Who owns a task in the dead-letter queue?" Staff answer: The media pipeline team, with a 24-hour triage SLA. A DLQ without an owner is a silent data-loss queue.
Key Takeaway: "PLAYABLE is not done. Measure the ladder you promised, not the first rung you shipped."
What clears the Staff bar:
- Distinguishes partial success from success in the metrics
- Finds the ownerless DLQ behind the bug
- Adds reconciliation of expected vs actual outputs
Deep Dive 3: Large-Customer Onboarding — A Broadcaster Brings 400,000 Hours of Archive#
Context: A broadcaster partnership will upload 400K hours of archive in 6 weeks, plus 50 live channels. Their contract requires DRM on all content and geo-restriction to three countries.
Questions to Surface First:
- How does 400K hours compare to daily capacity? (~55% of a normal day's upload volume — but spread over 6 weeks, ~1.3% extra per day.)
- What DRM systems and security levels does the contract require? Who holds keys?
- Which content needs advanced codecs on day one vs on popularity?
Typical L5 Approach: Lets the archive flow through the normal upload pipeline and adds DRM encryption to the packaging step.
Staff Approach: Routes the archive through a bulk-ingest class below normal uploads, rate-limited to keep the fast-path SLO intact. Adds multi-DRM packaging (CMAF
cbcsso one encrypted set serves Widevine, FairPlay and PlayReady), a license service with per-title policy (regions, security level), and region checks in the play API. Live channels get dedicated ingest pairs.
Principal Approach: Sees that DRM and rights are now platform capabilities, not a partner integration. Funds a rights-and-licensing service owned jointly by the playback platform and business affairs, so the next partner is configuration, and negotiates the contract's security-level requirements with partnerships before engineering commits.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Capacity | Bulk class gets a quota of ~10% of the fleet off-peak, 0% when fast-path queue age > 2 min. 400K h × 3 core-h ≈ 1.2M core-hours ≈ 13 days of a 4K-core quota. |
| DRM | Encrypt once with cbcs; license service p99 < 150ms in each region; license caching per session to keep it off the segment path. |
| Rights | rights.regions and rights.window on the video; play API enforces; signed URLs TTL 6h. |
| Live | Dual encoders and dual ingest points per channel; automatic failover with slate on loss. |
| Launch | Shadow the license service with synthetic sessions for a week; 3 pilot titles; then full catalogue. |
Metrics to Watch: encode.queue_age_p99{class=bulk}, drm.license_latency_p99, drm.license_denials{reason}, play.rights_denials{region}
Ownership Question: "Who decides whether a title can play in a country?" Staff answer: Business affairs owns the rights data; the playback platform owns enforcement. Engineering never edits rights by hand.
Key Takeaway: "Bulk work gets its own class with its own quota. A partner's archive must never be the reason a creator waits."
What clears the Staff bar:
- Sizes the archive against daily capacity before designing
- Isolates bulk ingest from the creator SLO
- Encrypts once for all DRM systems
Deep Dive 4: Post-Mortem — The Month Storage Grew 22%#
Context: The monthly cloud bill shows storage up 22% while watch time grew 1%. Finance asks for an explanation and a plan by Friday. You own the post-mortem.
Questions to Surface First:
- Which tier and which object classes grew: source, renditions, multipart leftovers, thumbnails?
- Did a lifecycle rule stop running or change?
- Did upload behaviour change (more 4K, longer videos)?
Typical L5 Approach: Finds that a lifecycle rule was disabled during a migration, re-enables it, and reports the fix.
Staff Approach: Re-enables the rule, then finds three causes: the disabled rule (60% of growth), a new 4K-default camera app doubling average source bitrate (25%), and incomplete multipart uploads never aborted (15%). Adds an abort rule for incomplete multipart uploads older than 7 days, a lifecycle-rule drift check, and
storage.bytes_per_watch_houras a weekly alert.
Principal Approach: Makes storage efficiency a standing KPI owned by the media platform with a finance partner, adds lifecycle rules to infrastructure-as-code with policy checks so they can't be disabled silently, and opens the retention policy discussion with legal — the 4K-source trend will continue regardless of hygiene.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Break down bytes added by prefix and storage class for the month. |
| Triage | Rule disabled 34 days ago during a bucket migration; 4K sources up from 9% to 21% of uploads; 1.1 PB of incomplete multipart parts. |
| Quick fix | Re-enable rule; abort incomplete uploads > 7 days; downscale source copies above 4K to a mezzanine after the ladder completes (with legal sign-off on fidelity). |
| Guardrails | IaC policy check on lifecycle rules; weekly storage.bytes_per_watch_hour review; alert at +5% month over month. |
| Post-mortem | Migration runbook lacked a lifecycle checklist; storage growth was reviewed quarterly, not weekly. |
Metrics to Watch: storage.bytes{class,tier}, storage.incomplete_multipart_bytes, upload.source_bitrate_p50, storage.bytes_per_watch_hour
Ownership Question: "Who owns the storage bill?" Staff answer: The media platform team owns the bytes and the lifecycle engine; finance owns the budget; neither can fix it alone.
Key Takeaway: "Storage doesn't page. Give it a ratio metric and an owner, or it grows until finance finds it."
What clears the Staff bar:
- Decomposes growth into causes with percentages
- Catches the multipart-leftover class (see Large File Uploads & Delivery)
- Turns a one-time fix into a guardrail
Deep Dive 5: Multi-Region Expansion — Launching in a Market With Weak Transit#
Context: The platform launches in a large market where international transit is expensive and congested at evening peak. Early QoE: startup p95 6s, rebuffer ratio 2.4%.
Questions to Surface First:
- Where is the bottleneck: CDN presence in-country, ISP interconnects, or last-mile?
- What device mix and data-plan constraints apply? (Many viewers on prepaid mobile data.)
- Which content is watched — is the head local or global?
Typical L5 Approach: Adds a cloud region in-country as an origin and points the CDN at it.
Staff Approach: Measures per ISP: the bottleneck is international transit into the two largest ISPs. Moves the head in-country — a commercial CDN with local presence plus, for the two largest ISPs, a direct peering arrangement. Adjusts the ladder for the market: an extra low rung (~150 kbps) and data-saver defaults on mobile. Results: startup p95 2.4s, rebuffer 0.5%.
Principal Approach: Treats market entry as a delivery-strategy decision: when a market's traffic exceeds a threshold, start the owned-edge or ISP-embedded program there; until then, commercial CDN with local presence. Writes the market-entry playbook so QoE targets, ladder adaptations and partnership steps are standard.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Diagnose | QoE by ISP and hour: degradation tracks evening peak on two ISPs; others fine. |
| Delivery | CDN with in-country PoPs; peering with the two largest ISPs; pre-warm local head nightly. |
| Ladder | Add 144p/150 kbps rung; data-saver mode caps at 480p by default on cellular. |
| Measure | QoE holdouts per change; qoe.rebuffer_ratio{isp} weekly with the partnerships team. |
| Expand | Traffic forecast triggers owned-edge evaluation at a set Gbps threshold. |
Metrics to Watch: qoe.startup_time_p95{isp}, qoe.rebuffer_ratio{isp,hour}, cdn.byte_hit_ratio{country}, egress.cost_per_gb{country}
Ownership Question: "Who owns ISP relationships?" Staff answer: A partnerships or network-strategy team with delivery engineering as the technical owner. Engineering can't sign peering agreements alone.
Key Takeaway: "In a new market, the bottleneck is usually an interconnect, not a server. Measure per ISP before building anything."
What clears the Staff bar:
- Diagnoses per ISP and hour
- Adapts the ladder to the market's devices and data costs
- Brings in the partnerships owner
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why video streaming is a cost-and-tail problem and put an egress number on it within the first two minutes
- Derive storage per day and encode compute from upload rate, ladder bitrates and core-hours per video-hour
- Design a priority-aware encode pipeline: fast path, ladder fill, popularity-driven advanced codec, bulk and backfill classes
- Derive a popularity threshold for extra encoding from compute cost, byte savings and egress price
- Choose segment length with request-rate and latency numbers, and say when low-latency parts are worth it
- Describe hybrid ABR with hysteresis and the server-side hints that can cap quality during incidents
- Protect origin with shields, coalescing, rung-based shedding and failover sizing
- Price own vs commercial vs hybrid edge and name the traffic level where each wins
- Write a lifecycle and retention policy for the long tail and name who signs off
The Bar for This Question#
Mid-level (L4): Draws upload → storage → transcoder → CDN → player and explains HLS at a high level. Uses a fixed ladder and a single CDN. No numbers beyond "millions of users". Would build something that works for a few thousand videos and has no plan for spikes, cost or the tail.
Senior (L5): Adds chunked upload, a transcoding queue, multiple resolutions, a CDN, adaptive bitrate and a sharded metadata store. Knows HLS and DASH. The gap: treats every video the same, never prices egress, says "short segments" without a request-rate number, and has no origin-protection story for a premiere. Plausible and competent — and would produce a storage bill that grows faster than the business and a 20:00 outage on the first big release.
Staff+ (L6): Frames the problem around egress and the long tail in the first minutes. Splits publish time from full quality. Derives the popularity threshold for advanced encoding. Protects origin with coalescing and rung-based shedding and sizes it for the largest regional failover. Sets QoE SLOs sliced by ISP and device, and names owners for the pipeline, player, delivery and content-ops handoffs. Knows when not to build any of it. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 "Video Is a Networking Business That Happens to Use Storage"#
| Cost Line (this design, rough) | Annual |
|---|---|
| Egress at a blended $0.003/GB, ~1.1 EB/day | ~$1.2B |
| Storage, year 1 corpus with tiering | ~$100–200M |
| Encode compute (H.264 for all, advanced for the head) | ~$25–35M |
The Staff position: Optimize for bytes delivered and where they're delivered from. A 10% bitrate saving is worth more than halving the encode fleet.
Why this matters in interviews: Candidates who spend their time on storage layout or the metadata DB are optimizing the smallest line items.
10.2 "Most Videos Should Never Get the Good Codec"#
| Video Class | Share of Videos (assumed) | Share of Watch Time (assumed) | Right Ladder |
|---|---|---|---|
| Head | ~10% | ~90% | Full H.264 + advanced codec, per-title for the very top |
| Torso | ~30% | ~9% | Full H.264 |
| Tail | ~60% | ~1% | H.264, shrinking to 2 rungs when dormant |
The Staff position: Encoding effort is an investment that pays back per view. Spend it where views are.
Why this matters in interviews: "Encode everything in AV1" sounds modern; it's a compute bill for videos nobody watches.
10.3 "Low Latency Is a Product Feature You Pay For, Not a Default"#
Standard HLS at 20–30 seconds is fine for most live content. LL-HLS at 3–5 seconds roughly triples to quadruples the request rate per viewer and tightens every origin and edge SLO. WebRTC under 1 second gives up HTTP caching altogether.
The Staff position: Default to standard latency; enable low latency per event where interaction, betting or spoilers make seconds matter — and have product say so.
Why this matters in interviews: "We'll use low-latency everywhere" signals that you haven't priced it.
10.4 "Owning a CDN Is a Ten-Year Decision Disguised as a Cost Project"#
The savings are real at extreme scale, but the commitment includes hardware generations, ISP relationships in hundreds of networks, a logistics operation, and an organization that will exist for a decade. The cheaper levers — codecs, hit ratio, contracts — should be exhausted first.
The Staff position: Build the edge only when traffic is large, predictable and growing, and keep a commercial CDN as both overflow and leverage.
Why this matters in interviews: Citing Open Connect without the volume math is copying; citing it with the math is judgment.
10.5 "Rebuffer Ratio Is the Only Metric the CEO Needs"#
Startup time, bitrate, resolution and start failures all matter to engineers. Stalls are what viewers remember and what most directly ends sessions.
The Staff position: One headline QoE metric — rebuffer ratio, sliced by ISP and device — with the rest as diagnostics. Every cost-cutting change ships with a rebuffer-ratio guardrail.
Why this matters in interviews: Naming one metric that leadership can track shows you can communicate a technical system upward.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer designs a pipeline and an edge that meet QoE targets at a defensible cost. The Principal engineer notices that the company's margin is set by cost per watch-hour, that it is driven by three levers with very different time horizons — codec efficiency (quarters), edge ownership (years) and retention policy (a legal and creator-relations decision) — and that the company currently runs three video stacks: VOD, shorts and live, each with its own encoder settings, its own CDN contract and its own idea of quality. The L7 problem is setting the video platform's economic and organizational shape for the next three to five years: one pipeline with workload classes, one edge with per-product quotas, one QoE definition, and a cost-per-watch-hour target that every team can see.
🧭 Principal Move: "I'd make cost per watch-hour and rebuffer ratio the two numbers every video team reports, and fund the levers in order of payback: advanced codec for the head this year, a hybrid edge in our top markets over three years, and a retention policy for the dormant tail that legal and creator relations sign. Each lever has a different owner, so the plan has to be written as a portfolio, not as one project."
The Org-Level Fault Line#
One video platform vs per-product video stacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each product runs its own stack (VOD, shorts, live, ads) | Speed; each team tunes for its workload | 3–4 encoders, 3–4 CDN contracts, inconsistent QoE metrics, no volume leverage | Finance (weaker contracts), viewers (inconsistent quality) |
| One central video team owns everything | One edge, one contract, one QoE definition | Bottleneck; live's latency needs fight VOD's efficiency needs in one roadmap | Product teams (velocity) |
| Platform owns primitives; products own experiences | Pipeline with workload classes, packaging, edge, QoE telemetry are platform; ladders, latency targets and player UX are product-configurable | Contract design and quota governance are hard | Platform team (API stewardship, chargeback) |
🧭 Principal Insight: The platform should own everything whose cost scales with bytes — encode infrastructure, packaging, edge, contracts — and expose knobs for everything whose value differs by product: ladder templates, latency class, startup strategy.
Cost Model#
Assumptions: average delivered bitrate 2.5 Mbps; source 8 Mbps; H.264 ladder 9.2 Mbps total; software encode ~3 core-hours per video-hour at $0.02 per core-hour; storage blended across tiers; fully loaded engineer ~$250K/year.
| Scale | Uploads / Watch | Egress ($/month) | Storage ($/month) | Encode ($/month) | Headcount | On-call |
|---|---|---|---|---|---|---|
| Startup | 500 h/day uploaded · 1M watch-h/day (~34 PB/month delivered) | ~$170–350K (commercial CDN, $0.005–0.01/GB) | ~$5–15K | ~$1K (or a managed service) | 3–5 eng; buy encoding and player | Shared rotation |
| Growth | 50K h/day · 100M watch-h/day (~3.4 EB/month) | ~$10–25M (multi-CDN, committed) | ~$1–3M | ~$100K | 30–50 eng: pipeline, player, delivery, QoE data | Pipeline and delivery rotations |
| Hyperscale | 720K h/day · 1B watch-h/day (~34 EB/month) | ~$70–150M (hybrid owned edge + commercial overflow) | ~$10–25M and growing | ~$1.5–3M (more with advanced codecs) | 300+ eng incl. edge hardware, peering, codec research | Follow-the-sun per layer |
The pricing insight: at every scale in this model, egress costs roughly 50× or more what encode compute costs. A team of 5 engineers (~$1.25M/year) that lowers average bitrate 5% saves ~$0.5–1.2M/month at growth scale and ~$3.5–7.5M/month at hyperscale. At startup scale the same team saves ~$10K/month — buy the managed service and spend the engineers on product.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
Segment container (CMAF fMP4 vs TS) and encryption scheme (cbcs vs cenc) | One-way | Repackaging the whole catalogue; player and DRM changes on every device family |
| Owning an edge network / ISP-embedded caches | One-way | Multi-year capital and partner commitments; unwinding strands hardware and relationships |
| Retention promise to creators ("we keep your video forever at full quality") | One-way | Public commitment; changing it is a trust and legal event |
| Deleting sources | One-way | Lost forever; renditions can be regenerated, sources cannot |
| Ladder rungs and bitrates | Two-way | Re-encode the head; tail follows lifecycle |
| Segment duration for VOD | Two-way | Repackage (not re-encode) if GOP structure allows |
| CDN vendor mix and steering weights | Two-way | Config change, contract cycle |
| Popularity threshold for advanced codec | Two-way | Config change; re-derive when prices change |
The Standard I'd Write#
RFC-VID-001: Video Encoding, Packaging and Delivery Standard
Status: Approved Owners: Video Platform + Delivery Engineering
Scope
Every product that stores or serves video: VOD, shorts, live, ads, previews.
MUST
1. Encode through the shared pipeline using a workload class
(fastpath, ladder, advanced, bulk, backfill); no product runs its own encoders.
2. Package as CMAF; encrypt with cbcs where DRM is required.
3. Serve segments only through the platform edge or contracted CDNs via steering;
no direct origin URLs to clients.
4. Emit standard QoE beacons (startup, stalls, bitrate switches, start failures)
with player version, device, ISP and CDN dimensions.
5. Keep the source for every video; treat renditions as regenerable.
6. Attach a lifecycle policy to every rendition set at creation.
SHOULD
1. Use 4–6 s VOD segments with 2 s keyframes; 1 s parts only for low-latency live classes.
2. Gate advanced codecs on device hardware decode capability.
3. Ship player changes in stages per device family with rebuffer-ratio guardrails.
Exceptions
Filed with Video Platform; decided within 10 business days; time-boxed to 2 quarters.
Exceptions to MUST 5 need legal sign-off.
Success metrics
- Cost per watch-hour: −15% year over year
- Rebuffer ratio ≤ 0.3% globally, ≤ 0.6% in every market with > 1M daily viewers
- Publish time p50 ≤ 2 min for uploads under 15 min
- Storage bytes per watch-hour: flat or falling quarter over quarter
- Products running their own encoders or CDN contracts: 0 by end of year 2
What I'd Tell the VP#
"Delivering video is our largest cost after content, and it scales with every hour watched. We can bend that curve with three levers on three timelines: better compression for our most-watched videos this year, a lower-cost delivery network in our biggest markets over three years, and a clear retention policy for videos nobody watches. Together they target a 15% annual reduction in cost per hour watched while holding stall rates under 0.3%. The first lever needs about five engineers; the second is a multi-year capital program I'd bring back as a separate proposal with market-by-market numbers. The retention policy needs legal and creator relations, not engineering."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices the business, not the boxes | "Egress is fifty times our encode bill. Our cost per watch-hour is the number I'd manage." |
| Sequences levers by horizon | "Codec this year, edge over three, retention as policy. Different owners, one plan." |
| Identifies one-way doors | "Container and encryption format, the edge, the retention promise, deleting sources. Everything else is config." |
| Redraws ownership | "Platform owns what scales with bytes; products own what differs in value." |
| Brings in non-engineering owners | "Licensors set DRM, legal sets retention, partnerships owns ISPs. I design around their decisions." |
Staff answers that L7 interviewers find insufficient:
- "We'll use AV1 to save bandwidth" — correct lever, no payback model, no device-capability plan, no owner.
- "We'll build our own CDN like Netflix" — no traffic threshold, no fixed-cost estimate, no plan for markets without boxes.
- "We'll keep everything in cold storage" — ignores that the retention promise is a policy decision with creators and legal.
Appendices
Appendix A: Mechanics in Depth#
A.1 The Chunked Transcode DAG#
def plan(video):
if dedup.exists(video.content_hash): # re-upload of identical bytes
return link_renditions(video, dedup.get(video.content_hash))
chunks = split_points(video.source, target_s=4) # closed GOP boundaries only
for rung in FASTPATH: # 360p, 720p H.264
for i in range(len(chunks)):
enqueue(task(video.id, rung, i), cls="fastpath")
for rung in LADDER - FASTPATH:
for i in range(len(chunks)):
enqueue(task(video.id, rung, i), cls="ladder")
def run(task): # idempotent per (video, rung, chunk)
key = f"{task.video_id}/{task.rung}/{task.chunk}"
if output_exists(key): # replay after preemption → no-op
return
with lease(key, ttl="2m", heartbeat="20s"):
out = encode(read_chunk(task), task.rung)
put_if_absent(key, out) # write-once segment objects
if all_chunks_done(task.video_id, task.rung):
package_and_publish(task.video_id, task.rung) # single writer per rung
The idempotent-task and lease mechanics are general; see Workflows, Sagas & Compensation and Job Scheduler.
A.2 Hybrid ABR With Hysteresis#
def next_rung(state, rungs):
tput = ewma_throughput(window_s=10) # bytes/s, both fast and slow EWMA, take min
if state.phase == "startup":
return highest(rungs, bitrate <= 0.7 * tput) # conservative start
if state.buffer_s < 3:
return rungs.lowest # panic
if state.buffer_s < 8:
return highest(rungs, bitrate <= 0.8 * tput) # immediate down-switch
up = rungs.above(state.current)
if (up and state.buffer_s > 25 and tput > 1.4 * up.bitrate
and now() - state.last_switch > hold_s()): # hold 20s + random 0–10s
return up
return min(state.current, hints.rung_cap or rungs.highest)
The random component in the hold time desynchronizes players that share a bottleneck — the fix for the oscillation storm in 4.3.
A.3 Live Pipeline#
Budget for 4 s glass-to-glass with LL-HLS: capture + contribution ~500 ms, real-time encode ~500 ms, packaging one 1 s part ~1 s, CDN ~200 ms, player buffer ~2 parts ≈ 2 s.
Appendix B: Data Model#
videos (video_id PK, owner_id, state, duration_s, source_uri, source_tier,
content_hash, visibility, rights_regions[], drm_required,
popularity_tier, created_at, last_viewed_at)
renditions (video_id, codec, height, bitrate_kbps, state, segment_prefix,
storage_tier, bytes, created_at, last_served_at,
PRIMARY KEY (video_id, codec, height))
encode_tasks (task_id PK, video_id, rung, chunk_idx, class, attempt,
lease_owner, lease_expires_at, status)
view_counters (video_id, hour_bucket, views) -- feeds popularity tiers
qoe_beacons → stream to Kafka → aggregates by ISP, device, CDN, player version
Metadata is small: ~100M+ videos × ~2 KB ≈ hundreds of GB, read-heavy and cacheable — PostgreSQL sharded by video_id or Cassandra both work. View counts and QoE beacons are high-volume streams through Kafka into aggregation; the pattern is the same as in Ad Click Aggregation.
Appendix C: Coordination Mechanisms#
| Mechanism | Used For | Why |
|---|---|---|
Kafka topic video.uploaded | Trigger planning | Durable, replayable; partition by video_id |
| Priority task queues per class | Encode scheduling | Fast path can't starve; lower classes preemptible |
| Leases with heartbeat | Worker ownership of a chunk | Preempted worker's chunk is retried by another after 2 min |
| Write-once segment objects | Output | Duplicate execution is harmless |
| Single packager per rung | Manifest writes | Avoids concurrent manifest edits |
Conditional state updates on videos.state | PROCESSING → PLAYABLE → LADDER_COMPLETE | Out-of-order completion events are no-ops |
| Hourly reconciler | Expected vs actual rungs | Catches silent ladder gaps (Deep Dive 2) |
Workers run well on Kubernetes with preemptible node pools for the ladder, advanced and backfill classes and on-demand nodes for the fast path.
Appendix D: API Contract & Client Behavior#
GET /videos/{id}/play?device=android&hdr=0&codecs=avc1,vp09,av01
200 {
"manifest_url": "https://cdn-a.example/v/abc123/master.m3u8?sig=...&exp=6h",
"alt_manifests": ["https://cdn-b.example/v/abc123/master.m3u8?sig=..."],
"license_url": null,
"start_rung": "720p",
"rung_cap": null,
"hints_ttl_s": 300
}
Client rules: start at start_rung or lower; on segment failure try the same segment once on the next CDN in alt_manifests before switching down; on 429 honor Retry-After with ±30% jitter; never retry a future live part faster than half the part duration; refresh hints every hints_ttl_s and keep playing with stale hints if the hint call fails.
Caching rules: segments Cache-Control: public, max-age=2592000, immutable; VOD manifests short TTL (~2 s) until LADDER_COMPLETE, then long; live manifests ≤ part duration with blocking reload.
Appendix E: Observability#
E.1 Core Metrics#
# Viewer experience (client beacons)
qoe.startup_time_p50 / p95{device,isp,cdn}
qoe.rebuffer_ratio{device,isp,cdn,player_version}
player.start_failure_rate{reason}
qoe.bitrate_avg_kbps / qoe.bitrate_switches_per_min
# Delivery
cdn.byte_hit_ratio{provider,region}
origin.egress_gbps / origin.ttfb_p99 / origin.shed_requests_total
egress.cost_per_gb{provider,country}
# Pipeline
encode.queue_age_p99{class}
publish.time_to_playable_p50
encode.ladder_completeness
encode.dlq_depth{task_type}
# Economics
storage.bytes{tier} / storage.bytes_per_watch_hour
cost.per_watch_hour (monthly)
E.2 Critical Alerts#
| Alert | Threshold | Page |
|---|---|---|
| Rebuffer ratio by ISP | > 1% for 10 min on an ISP with > 100K viewers | Delivery on-call |
| Startup p95 | > 5 s for 5 min globally or per title | Delivery on-call |
| Byte hit ratio | Drop > 3 points in 10 min in a region | Delivery on-call |
| Fast-path queue age | p99 > 5 min for 5 min | Media pipeline |
| Ladder completeness | < 99.5% over 1 h | Media pipeline |
| Live ingest gap | > 2 s on a channel | Live on-call |
E.3 Debugging "Users Say It Buffers"#
Slice qoe.rebuffer_ratio by ISP → CDN → device → player version → title. A single ISP at evening peak means interconnect congestion (steer, cap); a single CDN means a provider issue (shift weights); a single player version means a regression (halt rollout); a single title means a cold cache or a bad encode. Client-side telemetry is the only place the viewer's experience exists; see Metrics & Alerting Platform.
Appendix F: Scale Evolution#
| Scale | What Works |
|---|---|
| < 1K uploads/day, < 10 Gbps | Managed encoding service, one CDN, fixed ladder, progressive MP4 for short clips |
| 1K–100K uploads/day | Own pipeline with priority classes, CMAF, multi-CDN, QoE beacons |
| 100K+ uploads/day, Tbps egress | Popularity-driven codecs, lifecycle tiers, shields with coalescing, steering hints |
| Tens of Tbps+ | Hybrid owned edge, ISP peering and embedded caches, hardware encode evaluation |
What you don't build on day one: per-shot encoding, owned edge, hardware encoders, server-side ABR, low-latency live, your own DRM license server. Each one has a trigger in the 3-year path.
Multi-region: metadata and the play API run active-active per region; renditions are replicated to two or more origin regions because they're immutable (replication is simple — see Replication); sources live in one home region plus archive copy. See Multi-Region.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Encode fairness: per-creator quotas on the bulk class so one partner's archive can't starve others; fast path is never quota-limited for ordinary uploads, but upload rate limits apply per account (Rate Limiting).
- Edge fairness: per-product cache quotas (VOD, shorts, live, ads) so one class can't evict another; live gets reserved memory capacity during marquee events.
- Cost allocation: chargeback by bytes delivered and core-hours consumed per product; storage by bytes-month per tier. Teams that can't see their bytes can't reduce them.
- Tradeoff summary:
| Lever | Saves | Costs | Who Signs Off |
|---|---|---|---|
| Advanced codec for the head | ~30% bytes on most watch time | Compute; device gating | Video platform |
| Per-title ladders | 20–40% on easy content | Compute per title; quality metric pipeline | Video platform |
| Small-screen bitrate caps | ~45% bytes on capped sessions | Perceived quality | Product |
| Lifecycle tiers and dormant ladders | Most storage growth | Re-encode on revival | Media platform + legal |
| Owned edge | 50–80% unit cost on owned traffic | Multi-year capex and org | Executive leadership |