Hiring BarSupport

Design a Video Streaming Platform

Case study86 min read10 diagrams

Technologies referenced in this case study: Apache Kafka · PostgreSQL · Cassandra · Redis · Kubernetes

Related: CDN & Edge Caching · Object Storage · Handling Large Blobs · Job Scheduler · Managing Long-Running Processes · Metrics & Alerting Platform · Buy or Build · Auth & Identity

Reading Guide#

Organized for interview use first, reference second. The upload mechanics (presigned URLs, multipart, the upload state machine) live in Large File Uploads & Delivery, and the edge cache itself lives in CDN & Edge Caching. This page is about the end-to-end video product and what it costs: which renditions exist, when they are made, where they are served from, and who pays for each choice.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on the transcode DAG, ABR and storage lifecycle
What is a Video Streaming Platform? — Why interviewers pick this topic

A video platform accepts uploads (or a live feed), turns each source into a ladder of renditions at different resolutions, bitrates and codecs, cuts every rendition into short segments, publishes a manifest listing them, and serves the segments over HTTP from caches as close to the viewer as possible. The player picks a rendition segment by segment based on measured throughput and buffer — adaptive bitrate (ABR) — so a viewer on a train drops to 360p instead of stalling.

The hard part is not playing an MP4. The hard part is that every rendition you create costs compute once and storage forever, and every byte a viewer watches costs egress every time. At scale the bill is dominated by bytes leaving the network, not by servers.

Before vs After — the "viral upload" scenario:

Without a cost-and-tail design:
t=0:        Creator uploads a 12-min 4K video. Pipeline encodes 14 renditions
            (H.264 + VP9 + AV1, 144p–2160p) before marking it playable.
t=+38min:   First playable. Creator has already tweeted the link; viewers saw "processing".
t=+40min:   Link goes viral. 2M viewers in 30 min. Edge caches are cold for every rendition.
t=+42min:   Origin egress jumps from 1.5 Tbps to 6 Tbps. Origin p99 TTFB 4s. Rebuffer ratio 3%.
t=+1 month: The same pipeline encoded 14 renditions for 700K other videos that day,
            90% of which got fewer than 50 views. Storage bill +18% month over month.

With fast-path publish, popularity-driven encoding and tiered delivery:
t=0:        Upload completes. Fast path encodes 360p + 720p H.264 in parallel 4s chunks.
t=+90s:     Playable. Remaining H.264 rungs land by t=+8min.
t=+40min:   View velocity crosses threshold → advanced-codec ladder queued at high priority.
            Mid-tier shield coalesces misses: 1 origin fetch per segment per region.
t=+42min:   Origin egress 1.7 Tbps. Edge hit ratio 97%. Rebuffer ratio 0.4%.
t=+1 month: Long-tail videos hold 5 H.264 rungs; sources moved to archive tier after 30 days.

Why interviewers reach for this question: "Design YouTube" looks like a storage question and is really an economics question with a latency budget attached. It separates candidates who draw "upload → S3 → CDN" from candidates who can say how many renditions, how big, served from where, at what cost per GB — and what happens on the night 10 million people press play at 20:00.

Mechanics Refresher: The Video Delivery Primitives
PrimitiveHow It WorksProsCons
Progressive download (one MP4)Player fetches one file with byte rangesTrivial; any CDNNo adaptation; wastes bytes when viewers quit early; stalls on slow links
HLS (RFC 8216)Media playlist of segments (.ts or fMP4) + multivariant playlist listing renditionsNative on Apple devices; universalSegment duration sets the latency floor
MPEG-DASHXML manifest (MPD) + fMP4 segments; codec-agnosticStandard across non-Apple players; flexibleNot native on iOS Safari
CMAFCommon fMP4 segment format usable by both HLS and DASHOne set of files for both protocols — halves storage and cache footprintNeeds one encryption scheme (cbcs) for wide DRM coverage
Bitrate ladderSet of (resolution, bitrate, codec) renditions per titleLets ABR adaptEach rung = compute once + storage forever
Per-title / per-shot encodingPick ladder or encoder settings per title or per shot from measured complexity20–50% fewer bits for easy content at equal quality (typical industry range)Many trial encodes; more compute per title
ABR (client-side)Player chooses next segment's rendition from throughput and bufferScales; no server stateOscillation; greedy players fight over shared links
Low-latency HLS / DASHSegments split into ~1s parts, published before the segment completes2–5s glass-to-glass over plain HTTP CDNsMore requests, more manifest churn, tighter origin SLOs
DRM (Widevine, FairPlay, PlayReady)Segments encrypted; license server issues keys per sessionRequired by premium licensorsLicense server on the startup critical path

For most production systems: CMAF segments, served as HLS and DASH from one set of files; an H.264 ladder for universal reach plus an advanced codec (VP9, HEVC or AV1) for popular content; 4–6 second VOD segments; client-side ABR with server-provided hints; a commercial CDN until egress spend justifies anything else. The primitives are not the interview — which renditions you are willing to pay for is.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

Video streaming is not a storage question. Anyone can put files in object storage.

It is a cost-and-tail question that tests:

  • Whether you see that egress, not compute, dominates the bill — and design to move fewer bytes, from closer, more cheaply
  • Whether you treat the long tail differently from the head: most uploads are rarely watched, a few are watched millions of times
  • Whether you can turn "fast" into numbers: startup time, rebuffer ratio, publish time, glass-to-glass latency
  • Whether you know which decisions are made by people outside engineering — licensing, DRM, regional rights, creator expectations

The key insight: Encode once, store a deliberately chosen set of renditions, and serve them from as close to the viewer as possible. Every rendition is a bet: compute and storage now, against egress saved and quality gained on every future view. Staff candidates price that bet per video based on its popularity; Senior candidates encode everything the same way.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws upload → blob store → transcoder → CDN → playerAsks "UGC, premium catalogue or live? They are three different systems" and commits to one with numbersAsks "What share of revenue goes to delivery, and which lever — codec, edge, or deletion — moves it most over 3 years?"
Encoding"Transcode every upload into 6 resolutions""Fast-path two rungs for publish time; full H.264 ladder for all; advanced codec only past a view-velocity threshold"Treats encoding efficiency as a company cost lever: funds codec migration (or silicon) against a modeled egress saving with a payback date
Delivery"Put a CDN in front""Edge + mid-tier shield, request coalescing, ≥ 95% byte hit ratio target, origin sized for 3× the steady miss rate"Decides own-edge vs commercial vs hybrid as a multi-year bet; negotiates ISP peering; keeps a commercial CDN as a pressure valve
Quality"Use adaptive bitrate""ABR with buffer-based steady state; QoE SLOs: startup p95 ≤ 2.5s, rebuffer ratio ≤ 0.3%, measured per ISP and device"Makes QoE a business KPI tied to watch time; owns the tradeoff between bitrate caps and egress spend with finance and product
Storage"Store everything in S3""Lifecycle tiers by age and views; source to archive at 30 days; renditions re-creatable from source"Writes the retention and deletion policy for the long tail with legal and creator relations; storage growth must not outrun watch-time growth
OwnershipVideo team owns everythingIngest, media pipeline, playback and delivery are separate owners with SLOs at each handoffDefines the platform contract so live, shorts and premium share one pipeline and one edge without sharing a failure domain
Why "encoding" separates levels

L5: "Each upload triggers a job that transcodes it into 240p, 360p, 480p, 720p, 1080p and 4K, stores them, and marks the video ready." It works and it is how most people would start. It quietly assumes every video is worth the same spend, and it makes publish time equal to the slowest rung of the longest video.

L6: "I'll split encoding into three tiers by value. The fast path produces 360p and 720p H.264 from parallel 4-second chunks so a 10-minute upload is playable in about 90 seconds. The rest of the H.264 ladder fills in within 10 minutes. The advanced codec — which costs roughly 5× the compute but saves 30–40% of bytes — runs only once a video crosses about 1,000 views in its first day, because below that the compute costs more than the egress it saves."

L7: "The question is what our marginal cost per watch-hour is and which lever bends it. At our volume a 30% bitrate saving on the top 10% of videos is worth more than any other project on the roadmap. I'd fund the codec migration with a payback model — and I'd revisit hardware encoding when software encode becomes the constraint, which is exactly the trade large platforms have made."

Why "delivery" separates levels

L5: "Videos are served from a CDN, so scale is handled." The CDN does handle steady state. The answer is silent on what happens on a cache miss for a 2-hour film with 12 renditions × 1,800 segments each, and on who pays the egress bill.

L6: Separates the head from the tail. The head (popular videos) is pre-warmed or pre-positioned into edge caches; the tail is served from a mid-tier shield that coalesces misses so one origin fetch serves a whole region. "The number I protect is origin egress: if the edge byte hit ratio drops from 97% to 90%, origin load more than triples, and that is the outage."

L7: Frames the edge as a multi-year capital decision. "Below a few hundred Gbps, a commercial CDN wins on every axis. Above tens of Tbps, embedding caches inside ISPs changes our unit cost and our ISPs' transit bill — but it's a 3–5 year program with hardware logistics, and I'd keep a commercial CDN for overflow and for markets where we don't have boxes."

Why "storage" separates levels

L5: Stores the source and every rendition forever in standard-tier object storage. Fine for year one. By year three, storage is the fastest-growing line item and nobody can say which bytes earn their keep.

L6: "Renditions are a cache of the source, not data. I keep the source in archive tier after 30 days, keep only the H.264 ladder for videos with fewer than 10 views in 90 days, and delete advanced-codec renditions that stop being watched — they can be regenerated."

L7: Recognizes that deletion is a policy question with creators, legal and product: what does "we keep your video forever" mean? "I'd write the retention policy with legal: the source is preserved, playback quality for dormant videos may be reduced to a minimal ladder, and the first view after dormancy can trigger re-encoding. That sentence saves more money than any compression project."

Positions to Commit To#

PositionRationale
Egress is the bill; design to move fewer bytes from closerAt scale, delivery costs roughly 50× or more what transcode compute costs; every architecture choice is judged on bytes and hit ratio
Fast-path publish, then fill the ladderCreators judge the platform by time-to-playable; 2 rungs in ~90s beats 14 rungs in 40 min
Encode effort follows popularityAdvanced codecs and per-title tuning pay back only above a view threshold; the long tail gets a cheap universal ladder
CMAF segments, one set of files for HLS and DASHHalves storage and cache footprint versus packaging twice
4–6 s segments for VOD, 1 s parts for low-latency liveLonger segments = fewer requests and better compression; short parts only where latency is the product
Client-side ABR, server-side steering hintsThe client knows its buffer; the server knows CDN health and cost — combine both
Renditions are a regenerable cache; the source is the assetLets you tier and delete renditions aggressively without losing anything irreplaceable

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
User-generated uploads at scale (YouTube-like)Ingest-heavy, long tail, creators want fast publishFast-path encode, popularity-driven ladder, tiered storage, pull-through CDN with shieldTranscode backlog on upload spikes; storage outgrows viewsPublish p50 ≤ 2 min; no upload lost; rebuffer ≤ 0.3%
Premium catalogue (Netflix-like)Small catalogue (tens of thousands of titles), extreme popularity skew, licensors' DRM rulesExhaustive per-title/per-shot encoding, pre-positioning to edge before release, DRM on every streamCold cache on premiere night; license server on the critical pathStartup p95 ≤ 2s; highest quality per bit; contractual DRM compliance
Live streaming (Twitch, sports)Seconds of glass-to-glass latency, no time to pre-position, sudden audience spikesReal-time ladder at ingest, LL-HLS/LL-DASH parts, request coalescing at every tierIngest drop, origin stampede on a spike, ABR thrash at low bufferGlass-to-glass 3–5s (LL) or ≤ 30s (standard); no stream dies on one encoder

🎯 Staff Move: "I'll design for user-generated uploads at YouTube-like scale, because that's where ingest, the long tail and egress all collide. A premium catalogue spends far more compute per title and pre-positions everything; live trades efficiency for seconds of latency. I'll point out where each would change my design, and you can redirect me to either."

Where the Design Splits#

#Fault LineThe Tension
1Encode Everything Up Front vs Encode on DemandPay compute and storage for every rendition of every upload, or accept slower first-view quality for the long tail?
2Fixed Ladder vs Per-Title / Per-Shot EncodingSpend more compute per title to save bandwidth on every view — worth it only where views are high
3Own CDN vs Commercial CDN vs BothUnit cost and control at extreme scale vs capital, logistics and years of build-out
4Segment DurationShort segments cut latency and speed up adaptation; long segments cut request rate and improve compression
5Client-Driven ABR vs Server-Guided SteeringThe player knows its buffer; the platform knows CDN health, cost and congestion — who decides the rendition and the CDN?

How Real Companies Built It#

Why this section belongs here: Naming how real video platforms solved these problems shows you've studied the economics, not just the boxes.

YouTube — Custom Transcoding Silicon#

YouTube has said that on average more than 500 hours of video are uploaded every minute, that VP9 takes about 5× more compute to encode than H.264, and that after starting work in 2015 it built a Video (trans)Coding Unit (VCU) — the "Argos" chip — delivering up to 20–33× better compute efficiency than its previous optimized software system on general-purpose servers; the next generation targets AV1 (YouTube blog, ASPLOS 2021 paper).

The paper's extended abstract adds that the VCU fleet spans tens of thousands of servers and that, for offline two-pass encoding, it delivers 8–20× higher throughput than software encoding for H.264 and VP9 respectively (extended abstract).

Staff insight: Better codecs move cost from egress to compute. When upload volume makes software encoding of an advanced codec unaffordable, the choices are: encode it only for popular videos, or change the cost of compute. Saying both out loud — and knowing that only a handful of companies can do the second — is the signal.

Netflix — Per-Title and Per-Shot Encoding#

In December 2015 Netflix described moving from one fixed bitrate ladder for every title to a ladder chosen per title from an analysis of that title's complexity — simple animation needs far fewer bits than grainy action footage to look the same — and later described its Dynamic Optimizer, which chooses encoding parameters per shot to optimize perceptual quality at each bitrate (per-title encoding, Dynamic Optimizer).

Staff insight: Per-title encoding is affordable for Netflix because the catalogue is small and every title is watched many times. For UGC the same idea applies only to the head of the distribution. The interview move is to say where on the popularity curve the extra encode pays back.

Netflix — Open Connect#

Netflix runs its own CDN, Open Connect. Its appliances are provided to qualifying ISP partners at no charge to embed inside their networks, with the same capabilities as the appliances Netflix operates itself, and their content is updated through nightly fills (Open Connect).

Staff insight: Pre-positioning works because a premium catalogue's demand is predictable a day ahead. UGC demand is not, so a UGC platform still needs pull-through caching with a shield. Owning the edge is a decision about decades of traffic, not about this quarter's design. See Content Delivery Network for the cache mechanics.

Apple — HLS Authoring Specification#

Apple's HLS authoring specification for its devices says target segment durations SHOULD be 6 seconds, video segments MUST start with an IDR frame and key frames SHOULD occur every 2 seconds; for Low-Latency HLS the recommended Part Target Duration is 1 second and SHOULD be at least three times the p95 round-trip time; and HEVC variants should use bitrates about 20% below the H.264 values (HLS Authoring Specification, RFC 8216).

Staff insight: Segment duration is not a free parameter — the platform vendor publishes a default, and the low-latency variant ties part duration to network round-trip time. Citing the default and then explaining when you'd deviate is stronger than inventing a number.

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"We transcode into 6 resolutions""How long until a 1-hour 4K upload is playable? What does that cost per day across all uploads?"Publish-time design and compute math
"We store videos in S3""How much storage do you add per day, and what is it in year three?"Long-tail economics, tiering, deletion
"We put it behind a CDN""A new episode drops at 20:00 for 10M viewers. What does origin see at 20:00:05?"Cold cache, coalescing, pre-warming
"The player uses adaptive bitrate""1M players on one congested ISP all switch down, then up, then down. What happened?"ABR oscillation, shared bottlenecks
"We use HLS with short segments for low latency""How short, and what does that do to request rate and encoding efficiency?"Segment-duration tradeoff with numbers
"We'll build our own CDN""At what traffic level does that pay back, and what do you do in markets you haven't reached?"Build-vs-buy as a multi-year bet

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Two almost independent systems share a metadata store. The media pipeline (left) is a batch system judged on publish time and cost per encoded hour. The delivery path (right) is a read system judged on startup time, rebuffer ratio and egress cost. The playback API is the only synchronous control-plane hop on the play path; segments never touch application servers. The number that decides whether the system survives a big night is cdn.byte_hit_ratio — origin load is proportional to the miss rate, not to the audience.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
What dominates cost"Storage and transcoding""Egress. At 1B watch-hours a day, delivery is over 1 EB/day; a 1-cent-per-GB difference is about $4B a year."
Encoding"Transcode to all resolutions""Fast path 2 rungs in ~90s, full H.264 ladder in ~10 min, advanced codec only above a view threshold."
Segments"Short segments""4–6s for VOD; 1s parts for low-latency live. Halving segment length doubles request rate."
CDN"Use a CDN""Edge + shield, coalescing, ≥ 95% byte hit ratio; pre-warm the head; protect origin with 3× miss headroom."
ABR"Adaptive bitrate""Throughput-based at startup, buffer-based in steady state, hysteresis to stop oscillation; server hints for CDN choice."
Storage"Keep everything""Source is the asset; renditions are a cache. Tier by age and views; delete cold advanced-codec rungs."
QoE"Make it fast""Startup p95 ≤ 2.5s, rebuffer ratio ≤ 0.3%, start failures ≤ 0.5%, sliced by ISP, device and CDN."

Numbers to Bring#

All numbers below are this design's working assumptions unless a source is cited. YouTube's own published upload figure is "more than 500 hours every minute" (YouTube blog); the design uses that as its ingest rate.

MetricValueWhy It Matters
Upload ingest rate500 h of video/min → 30K h/hour → 720K h/day; peak ~2×Sizes the encode fleet and the daily storage add
Average upload~10 min, ~8 Mbps source → ~3.6 GB per hour of video~50 uploads/s; ~2.6 PB/day of source
H.264 ladder (240p–1080p)0.3 + 0.7 + 1.2 + 2.5 + 4.5 = 9.2 Mbps total~69 MB per minute of video for the full ladder
Storage per minute per rendition1080p ≈ 34 MB · 720p ≈ 19 MB · 360p ≈ 5 MBMultiply by renditions × minutes × uploads/day
Renditions per title5 H.264 rungs for all; +5–7 advanced-codec rungs for the head; + 2 audio7 for the tail, 12–14 for the head
Eager rendition storage~3 PB/day (H.264 ladder for every upload)~1 EB/year before tiering
Encode compute (software)~3 core-hours per video-hour for the H.264 ladder; advanced codec ~5×~90K cores busy for H.264 alone at 720K h/day
Segment size (4s at 2.5 Mbps)~1.25 MBOne request per 4s per viewer = 0.25 req/s
Peak concurrent viewers~75M (1B watch-hours/day ÷ 24 × 1.8 peak factor)~19M segment req/s at 4s segments; ~38M at 2s
Peak egress~190 Tbps (75M × 2.5 Mbps)The number the CDN contract is written around
Edge byte hit ratio target≥ 95% (≥ 99% after mid-tier)Origin sees ≤ 1% of bytes ≈ 1.9 Tbps at peak
Egress costowned edge + peering ~$0.002–0.005/GB; commercial committed ~$0.005–0.02/GB1.1 EB/day × $0.01 = ~$11M/day
Startup timep50 ≤ 1s, p95 ≤ 2.5sEvery extra second before first frame loses viewers who never return to the video
Rebuffer ratio≤ 0.3% of watch timeThe stall a viewer feels most; set the SLO per ISP and device, not just globally
Publish timefirst playable p50 ≤ 2 min; full H.264 ladder ≤ 15 minCreator-facing SLO
Live glass-to-glassStandard HLS (6s segments) 15–30s · LL-HLS (1s parts) 3–5s · WebRTC < 1sEach step down multiplies cost per viewer

Interview Walkthrough

A 45-minute plan for "Design YouTube." The goal is to spend under 10 minutes on the boxes and the rest on the three places the system actually breaks: the encode pipeline under load, the edge under a spike, and the cost curve over time.

Phase 1: Requirements & Framing (2–3 minutes)#

Say this, nearly word for word:

"Before I draw anything — there are three different products hiding in 'design YouTube': user uploads at scale, a premium catalogue like Netflix, and live streaming. They share a CDN and a player, but the pipeline and the economics are different. I'll design for user uploads, and I'll call out what changes for the other two.

Scale assumptions: 500 hours uploaded per minute, about 1 billion watch-hours a day, viewers worldwide, mostly on mobile. Functional scope: upload, process, publish, play with adaptive bitrate, basic view counts. Out of scope unless you want them: recommendations, comments, monetization, search.

The constraints I'll design to: a creator sees their video playable within about 2 minutes; viewers start playback in under 2.5 seconds at p95 with under 0.3% of watch time spent rebuffering; and the cost per watch-hour goes down over time, because egress is the line item that decides whether this business works."

Why this works: It names three intents, commits, attaches numbers to "fast", and announces the cost thesis in the first two minutes. The interviewer now expects a conversation about bytes and tails, not about CRUD.

Phase 2: Core Entities & API (1–2 minutes)#

Video        { video_id, owner_id, state, duration_s, source_uri, created_at,
               visibility, rights{regions[], drm_required}, popularity_tier }
Rendition    { video_id, codec, height, bitrate_kbps, state, segment_prefix,
               storage_tier, bytes, created_at, last_served_at }
EncodeTask   { task_id, video_id, chunk_idx, ladder_rung, priority, attempt,
               lease_owner, lease_expires_at }
PlaySession  { session_id, video_id, user_id?, device, cdn, started_at }   # analytics

POST /videos                         → { video_id, upload_id }            # starts multipart
PUT  <presigned part URL>            → 200                                # client → storage
POST /videos/{id}/complete           → 202 { state: PROCESSING }
GET  /videos/{id}/play?device=...    → { manifest_url (signed, 6h), license_url?,
                                         cdn_hints[], start_rung }
GET  <cdn>/v/{id}/{rendition}/seg_{n}.m4s   → segment (immutable, cache 30d)
POST /qoe/beacons                    → 204   # startup, stalls, bitrate switches

Say: "Segments are immutable and content-addressed by video, rendition and index, so they cache forever. Manifests for VOD are cacheable too once the ladder is final; I'll give them a short TTL while rungs are still being added. The play endpoint is the only dynamic call on the play path — it signs URLs, checks rights and region, and hands the player CDN hints."

Phase 3: High-Level Architecture (≤5 minutes)#

Staff candidates spend under 5 minutes here. Draw two lanes: the media pipeline and the delivery path. Everything else is detail you add on demand.

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Narrate: "Upload goes straight to storage with presigned multipart — the API never sees the bytes; that pattern is covered in large-blob handling. Completion emits an event. The scheduler splits the source into GOP-aligned 4-second chunks and fans out tasks by priority: fast-path rungs first, the rest of the H.264 ladder next, advanced codec only when popularity says so. Workers write CMAF segments to the rendition origin. On the read side: player calls the play API, gets a signed manifest, fetches segments through edge and shield caches."

Phase 4: Transition to Depth (1 minute)#

"The boxes are standard. The interesting decisions are three: how much encoding we do and when, because that sets publish time and the compute-plus-storage bill; how we protect origin when a video or a premiere goes viral, because that's where outages come from; and how ABR and segment length trade latency against cost. I'd like to start with encoding economics — is that where you want to go, or would you rather start with delivery?"

Phase 5: Deep Dives (25–30 minutes)#

Deep dive 1 — Encoding economics (8–10 min). Walk the arithmetic. 720K hours/day × ~3 core-hours per hour for the H.264 ladder ≈ 2.2M core-hours/day ≈ 90K cores. Advanced codec at ~5× would add ~450K cores if applied to everything. Then the popularity curve: "If 10% of videos earn about 90% of watch time — a typical UGC skew I'd verify in our own data — applying the advanced codec to that 10% costs ~45K cores and saves ~30% of bytes on ~90% of egress. Applying it to the other 90% costs ~400K cores to save bytes on 10% of egress. So the threshold is a number, not a philosophy." Mention chunked parallel encoding for publish time and the priority classes.

Deep dive 2 — Delivery under a spike (8–10 min). Edge + mid-tier shield; request coalescing (one upstream fetch per object per tier while others wait); pre-warm the first 3 segments of the top rungs for anything the recommendation system is about to push; origin sized for 3× normal miss traffic. "If edge byte hit ratio drops from 97% to 91%, origin load triples. That's the incident, so cdn.byte_hit_ratio per region is a paging metric." Defer cache mechanics to the CDN case.

Deep dive 3 — ABR, segments and QoE (6–8 min). 4s segments for VOD (Apple's spec defaults to 6s; I'll go slightly shorter for faster adaptation on mobile). Throughput-based start at a conservative rung, buffer-based steady state with hysteresis. QoE beacons → startup time, rebuffer ratio, bitrate, sliced by ISP and CDN. Server-side hints choose the CDN; the client chooses the rung.

If time remains — live: real-time ladder at ingest, 1s LL-HLS parts, glass-to-glass 3–5s, and why a 10M-viewer spike is a coalescing problem, not a transcoding problem.

Phase 6: Wrap-Up (2–3 minutes)#

"To summarize: the media pipeline optimizes publish time and cost per encoded hour by spending effort according to popularity. Delivery optimizes origin protection and egress cost with tiered caches and coalescing. QoE is measured on the client and sliced by ISP and CDN. What I'd build later: per-title encoding for the head, a hybrid edge once egress crosses tens of Tbps, and a retention policy for the dormant tail written with legal. The one-way doors are the segment format and the DRM scheme; everything else can change behind the manifest."

Common Timing Mistakes#

MistakeTime LostFix
Explaining multipart upload in detail5–8 minOne sentence: "presigned multipart, API never sees bytes" — link to large-blob handling
Designing the recommendation system or comments10+ minScope it out in Phase 1
Drawing 15 boxes before any numbers8 minTwo lanes, then numbers
Debating HLS vs DASH3–5 min"CMAF, served as both"
Never getting to costWhole signalPut the egress number in Phase 1
Skipping live because "it's similar"Pivot failsOne sentence on what changes: no pre-positioning, 1s parts, coalescing

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

"Design YouTube" is the most common media prompt because it contains every cost tension of a large-scale read system in one product. The ingest side is a batch pipeline with a backlog problem. The read side is a cache with a 1,000:1 popularity skew. The bill is dominated by a resource — egress bandwidth — that most backend engineers never see on a dashboard. A Senior engineer can build every box. The Staff signal is noticing that the boxes are not where the risk is: the risk is a viral upload that the encode queue can't absorb, a premiere night with cold caches, and a storage curve that grows faster than watch time. Interviewers use it to see whether you can put a dollar sign and an owner on each of those.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

The Senior path is correct and ends at scaling the metadata database — which is a few thousand QPS and never the bottleneck. The Staff path spends its time where the money and the outages are.

1.3 The Staff Question That Cuts Through Everything#

"For a video that will be watched N times, what is the cheapest set of renditions — and the cheapest place to serve them from — that still meets our startup and rebuffer targets?"

Every fault line is a version of this question. Encode up front or on demand: depends on N. Per-title encoding: pays back above some N. Own the edge: depends on the sum of N across the catalogue. Segment duration: trades request cost against latency per view. The candidate who asks it out loud has reframed the interview from "what boxes" to "what is each byte worth."


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1: User-generated uploads at scale. Ingest is enormous and unpredictable: 500 hours a minute, spiky around events and time zones. Most uploads are watched rarely; a small fraction are watched enormously and unpredictably — a video can go from 0 to 5M views in an hour. Creators measure the platform by time-to-playable and by quality on their own phone. The design centre is a priority-aware encode pipeline, popularity-driven encoding effort and pull-through caching with a shield because you cannot predict which video to pre-position.

Intent 2: Premium catalogue. Tens of thousands of titles, each expensive to license and watched by millions. Releases are scheduled, so demand is predictable to the hour. Spending hours of compute per title on per-shot optimization pays back across millions of views. Licensors impose DRM, regional windows and output-protection rules; the license server is on the startup critical path. The design centre is encode quality per bit and pre-positioning — fill edge caches overnight before the release.

Intent 3: Live streaming. No time to optimize encodes — the ladder is produced in real time at ingest. No time to pre-position — the content didn't exist a second ago. Audiences spike from 10K to 10M in minutes when a match goes to penalties. The design centre is latency budget (glass-to-glass), ingest redundancy (dual encoders, dual ingest points), and request coalescing at every tier, because every viewer requests the same newest part within the same second.

🎯 Staff Move: "These three share a player and a CDN but not a cost model. UGC is a tail problem, premium is a quality-per-bit problem, live is a latency problem. I'll design UGC, and for each fault line I'll say in one sentence what the premium and live answers would be."

2.2 When NOT to Build Your Own Video Pipeline#

Most companies that need video should not build any of this.

SituationBetter ChoiceWhy
Video is a feature, not the product (course platform, product demos, support clips)Managed video platform or cloud media servicesEncoding, packaging, DRM and player SDKs for a per-minute price; your engineers work on the product
Under ~10 Gbps of peak egressCommercial CDN + cloud transcoding serviceEgress is a few thousand dollars a month; an owned stack costs more in salaries than it saves
Short clips under ~60s, rarely re-watchedProgressive MP4 with 2–3 renditionsABR and segmenting add complexity that a 15-second clip doesn't need
Video calls and real-time interaction (< 500 ms)WebRTC / SFU architectureHTTP segment streaming cannot reach interactive latency; different system — see Push vs Poll
Internal recordings, compliance archivesObject storage + on-demand transcodeWatched almost never; optimize storage, not delivery
Premium licensed content without a DRM teamBuy DRM-as-a-service and a packagerLicensor audits and multi-DRM key management are a specialty; mistakes cost licenses

🎯 Staff Insight: The build-vs-buy line is set by egress volume and by whether video quality is a competitive differentiator. Until both are true, buying is cheaper and better. See Buy or Build for the general framework.

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It MattersWhat to Ask or Assume
Upload volume and video lengthDrives encode fleet and storage add"500 h/min, average 10 min; peak 2×"
Watch volume and device mixDrives egress and the ladder"~1B watch-hours/day, ~70% mobile; average delivered bitrate ~2.5 Mbps"
Publish-time expectationFast path vs batch"Playable p50 ≤ 2 min, full quality ≤ 15 min"
Retention promiseLong-tail storage forever?"Source kept; rendition set may shrink for dormant videos — need legal sign-off"
DRM / rightsLicense server on the play path; regional blocks"UGC: no DRM, but region and visibility checks; premium: multi-DRM"
Live in scope?Different latency architecture"VOD first; live as an extension"
Max resolution4K rungs are 3–4× the bytes of 1080p"Up to 2160p for sources that have it, head only"
Who pays for egressOwn edge vs commercial"Assume we're large enough to negotiate and peer"

2.4 Precise Terminology#

TermPrecise MeaningCommon Confusion
Rendition / rungOne encoded version: codec + resolution + bitrateNot the same as a "format" or container
Bitrate ladderThe full set of renditions for a titleFixed ladder vs per-title ladder
SegmentA few seconds of one rendition, independently decodable (starts on a keyframe)Segment ≠ chunk; "chunk" usually means an encode work unit
GOP / keyframe intervalDistance between IDR frames; segments must align to themLonger GOP compresses better; must divide the segment length
Part (LL-HLS)Sub-segment (~1s) published before the full segment completesNot a separate rendition
Manifest / playlistLists renditions (multivariant) or segments (media playlist)Live manifests change every part; VOD manifests are static once complete
CMAFCommon fragmented-MP4 format usable by both HLS and DASHA container format, not a protocol
Byte hit ratioBytes served from cache ÷ bytes servedRequest hit ratio can be high while byte hit ratio is low (manifests hit, segments miss)
Startup timePlay press → first frame renderedExcludes ad time; includes manifest, license and first segments
Rebuffer ratioStall time ÷ (stall + play time)Not "number of rebuffers"; one 10s stall can matter more than five 0.2s ones
Glass-to-glass latencyCamera capture → viewer screenNot segment latency; includes encode, packaging, CDN and player buffer
EgressBytes leaving your network (or your CDN's) to viewersThe dominant cost; priced per GB or per committed Gbps (95th percentile billing)

3. Where the Design Splits#

Each fault line below has options, who pays, a Staff default and when to deviate. The question underneath all five is the one from 1.3: for a video watched N times, what is the cheapest way to meet the QoE targets?

3.1 Fault Line 1: Encode Everything Up Front vs Encode on Demand#

The tension: Encoding every rendition at upload makes every first view perfect and wastes compute and storage on videos nobody watches. Encoding on demand saves both and makes the first viewer of a cold rendition wait — or watch a lower rung.

StrategyWhat WorksWhat BreaksWho Pays
Full ladder, all codecs, at uploadEvery view gets the best rendition from day one; simple pipelineCompute ~6× the H.264-only cost; storage ~2× forever; publish time = slowest rungFinance (compute + storage), creators (slow publish)
Universal ladder up front, advanced codec on popularityCheap baseline for everyone; bytes saved where views concentrateFirst hours of a viral video served in H.264 (~30% more bytes)Delivery budget during the ramp
Minimal ladder up front, everything else just-in-timeLowest storage; good for archivesFirst viewer of a rung waits for an encode or gets a lower rung; JIT encode capacity must handle spikesViewers of cold content (quality), on-call (JIT fleet spikes)
Just-in-time packaging onlyStore one mezzanine per rung, package HLS/DASH/encryption at the edge or origin on requestSaves duplicate packaging; CPU at origin on every missOrigin team (CPU per miss)
Diagram: 3.1 Fault Line 1: Encode Everything Up Front vs Encode on Demand

The arithmetic that sets the threshold. An advanced-codec ladder costs about 5× the H.264 encode: roughly 15 core-hours per video-hour, about $0.30 at $0.02 per core-hour, so ~$0.05 for a 10-minute video. It saves about 30% of bytes. A full view of that video at 2.5 Mbps delivers ~190 MB; 30% saved is ~56 MB, worth ~$0.0003 at a blended $0.005/GB. Compute alone breaks even at ~180 full views. Two corrections push the number up: the average view watches maybe 40% of the video (→ ~450 views), and the extra rungs (~0.45 GB for 10 minutes) cost ~$0.05 a year in warm storage, roughly doubling the cost to recover. That lands near 1,000 views — hence the threshold of 1,000 views in 24 hours, lowered when the platform predicts the video is trending.

Staff default: universal H.264 ladder for every upload, fast path first; advanced codec driven by view velocity; dormant videos shrink to two rungs.

When to deviate:

  • Premium catalogue: encode everything exhaustively up front — every title is in the head.
  • Archive or compliance video: store the source only; encode on first request and accept a 30–60s wait.
  • Encode hardware makes advanced codecs cheap: if per-hour encode cost drops 10–20×, the threshold drops by the same factor and "encode everything" becomes reasonable again.

🎯 Staff Move: "I won't pick 'encode everything' or 'encode on demand' on principle. Advanced-codec encoding costs about 30 cents per video-hour and saves about 0.17 cents per watch-hour; with partial viewing and a year of storage for the extra rungs, it pays back at roughly 1,000 views. That's my threshold, and I'd re-derive it whenever compute or egress pricing changes."

3.2 Fault Line 2: Fixed Ladder vs Per-Title or Per-Shot Encoding#

The tension: A fixed ladder (same bitrates for every title) is predictable and cheap to compute. It overspends bits on simple content — slides, animation, a talking head — and underspends on grainy, high-motion content. Per-title and per-shot encoding fix both by running analysis or trial encodes, which costs compute per title.

StrategyExtra ComputeBandwidth SavedFits
Fixed ladder00 (baseline)The tail; live
Complexity-classified ladder (3–5 ladder templates picked by a fast analysis pass)+5–10%Large on easy content (screen recordings, animation)UGC baseline
Per-title convex hull (trial encodes at several resolutions/bitrates, pick the efficient frontier)+2–5×Typically 20–40% at equal perceptual quality for content that differs from the "average" titleThe head of UGC; premium
Per-shot optimization+5–10×Further gains over per-title by allocating bits shot by shotPremium catalogue; the very top of UGC

Who pays: the encode budget pays up front; the delivery budget and the viewer's data plan collect the savings on every view. Live can't use either — there's no time for trial encodes — so live ladders are fixed and slightly over-provisioned.

Measuring quality: you need a perceptual metric (VMAF, SSIM or an internal model) to say "equal quality". Without one, per-title encoding is guesswork and every ladder change is an argument. Treat the metric pipeline as a prerequisite, owned by the media pipeline team.

Staff default: complexity-classified ladders for every upload (cheap, catches screen recordings and animation), full per-title encoding for the head, per-shot never for UGC unless hardware changes the cost.

🎯 Staff Insight: Per-title encoding is a bandwidth project funded by the encode budget. Present it with both numbers — extra core-hours and saved petabytes — or it will lose every prioritization meeting to features.

3.3 Fault Line 3: Own CDN vs Commercial CDN vs Both#

The tension: A commercial CDN is instant, global and someone else's pager. At very large egress volumes, owning cache servers — in your own points of presence and, further, embedded inside ISP networks — lowers unit cost and improves QoE, but it is a multi-year capital program with hardware, logistics and peering negotiations.

StrategyWhat WorksWhat BreaksWho Pays
Single commercial CDNDay-one global reach; no capexUnit cost at scale; correlated outage; limited control of steeringFinance (per-GB), everyone during a provider outage
Multi-CDN with steeringPrice leverage, failover, per-region best performerSteering logic, config drift, cache warmed N timesDelivery team (complexity)
Owned edge + commercial overflowLowest unit cost on the bulk of traffic; control of cache policyCapex, hardware refresh every 3–5 years, ops team, peeringCompany (capital), infra org (headcount)
ISP-embedded appliancesBytes never cross transit; best QoE; ISPs save transit tooOnly viable with traffic big enough that ISPs want the boxes; logistics in hundreds of networksCompany (multi-year program)

Rough break-even. At ~$0.01/GB commercial, 1 Tbps sustained average is 10.8 PB/day ≈ $108K/day ≈ $40M/year. An owned edge that serves the same traffic at $0.002–0.004/GB all-in (hardware amortized, colocation, transit/peering, staff) saves roughly $25M/year per Tbps — once the fixed cost of a 40–80 person edge organization ($15–25M/year) is covered. Below a few Tbps average, buy. Above tens of Tbps, owning wins if you can execute. The details of cache design are in Content Delivery Network; the decision framework is in Buy or Build.

Staff default for our design (190 Tbps peak): hybrid. Owned edge and ISP peering carry the steady head; one or two commercial CDNs carry overflow, new markets and the long tail where our caches would thrash.

When to deviate: a startup or mid-size platform should use one commercial CDN with committed pricing and add a second only when spend or a provider outage justifies it.

3.4 Fault Line 4: Segment Duration — Latency vs Efficiency and Request Rate#

The tension: Shorter segments let live streams run closer to real time and let ABR react faster. Longer segments mean fewer HTTP requests, fewer keyframes (better compression), smaller manifests and better cache efficiency.

Segment LengthRequests per Viewer-HourAt 75M ConcurrentLive Latency (≈ 3 segments buffered)Compression Cost
2 s1,800~38M req/s~6–10 sKeyframe every 2s: a few % more bits
4 s900~19M req/s~12–15 sBaseline for this design
6 s (Apple's default target)600~12.5M req/s~18–30 sSlightly better than 4s
10 s360~7.5M req/s30 s+Best; slow ABR reaction
LL parts 1 s inside 4–6 s segments3,600 part requests (or fewer with blocking preload)Edge must hold requests open~3–5 sSame encode; packaging and CDN cost rise

The request rate matters because CDNs charge per request at some tiers, because each request has a fixed overhead in TLS and HTTP processing at the edge, and because manifests for live are re-fetched every part.

Who pays: short segments are paid for by the edge (request rate) and by bandwidth (keyframe overhead); long segments are paid for by viewers on bad networks (slower adaptation → more stalls) and by live audiences (latency).

Staff default: 4 s segments with 2 s keyframes for VOD (two GOPs per segment gives the packager flexibility); 6 s segments with 1 s parts for low-latency live where latency is in the product promise; standard 6 s segments for live where 20–30 s is acceptable — most live content does not need LL.

🎯 Staff Move: "Segment length is a cost knob, not a latency knob. Halving it doubles request rate across 75 million viewers. I'll use 4-second segments for VOD and reserve 1-second parts for the live events where latency is part of the product, like sports betting or interactive streams."

3.5 Fault Line 5: Client-Driven ABR vs Server-Guided Steering#

The tension: The player knows its buffer level and recent throughput; the platform knows which CDN is healthy, which ISP is congested, what each byte costs and what a million other players are seeing. Pure client ABR is scalable and robust but myopic. Pure server control is informed but adds a control plane to the play path and cannot see the device's buffer.

Diagram: 3.5 Fault Line 5: Client-Driven ABR vs Server-Guided Steering
StrategyWhat WorksWhat BreaksWho Pays
Throughput-based client ABRFast startup decisionNoisy estimates → oscillation; over-estimates on bursty mobile linksViewers (stalls, quality flapping)
Buffer-based client ABRStable in steady state; few stallsSlow start; under-uses capacity at startupViewers (low quality early)
Hybrid client ABR (throughput at start, buffer in steady state)Industry defaultNeeds tuning per device classPlayer team
Server-side steering hints (CDN choice, rung cap per ISP during congestion, start rung)Uses fleet-wide knowledge; can shed load during incidentsA control-plane dependency; must degrade to client-onlyPlayback platform team
Server-side ABR decision per segmentTotal controlAdds latency to every segment; doesn't see device bufferEveryone

Staff default: hybrid client ABR with hysteresis (up-switch needs headroom and a hold time; down-switch is immediate), plus server hints delivered with the play response and refreshed every few minutes: preferred CDN order, a starting rung, and an optional bitrate cap per ISP or region during congestion. If the hint service is down, the player uses cached hints or none — never blocks playback. This is a degraded-mode design: the server can improve playback but cannot stop it.

🎯 Staff Insight: The server-side bitrate cap is the most powerful incident lever you have. During an ISP congestion event, capping that ISP's viewers at 720p cuts their egress by ~45% and usually removes the congestion that caused the stalls. Decide in advance who may pull it — the delivery on-call — and who must be told — partnerships, because the ISP will ask.


4. When It Breaks#

4.1 Transcoding Backlog After a Viral Upload Spike#

A global event (an election result, a celebrity incident, a natural disaster) produces a 3× upload surge in 40 minutes, concentrated in long phone videos.

t=0:       Upload rate 500 → 1,400 h/min. Encode demand ~2.8× fleet capacity.
t=+5min:   encode.queue_age_p99 1min → 9min. Fast-path and ladder tasks share one queue.
t=+12min:  Fast-path p50 publish time 90s → 14min. Creators retry uploads (duplicates +20%).
t=+15min:  Page: encode.fastpath_queue_age_p99 > 5min for 5m.
t=+18min:  On-call: pause advanced-codec and re-encode classes (frees ~30% of fleet).
t=+20min:  Scheduler moves full-ladder tasks behind fast path. Autoscaler requests
           preemptible capacity: +40% in 25 min.
t=+45min:  Fast path back under 2 min. Ladder backlog 6 hours, drains overnight.
t=+1 day:  Post-mortem: duplicate uploads (same content hash) were encoded twice.

Detection: encode.queue_age_p99{class}, publish.time_to_playable_p50, encode.tasks_pending{class}, upload.duplicate_content_ratio.

Mitigation: strict priority classes — fast path, ladder, advanced codec, re-encode/backfill — with the lower classes preemptible; dedup by content hash before encoding; shed the advanced-codec class first because it is pure optimization.

Prevention: capacity model with 2× headroom for the fast path only (cheap: it's 2 rungs); preemptible spot capacity for everything else; quarterly load test replaying a 3× surge. Scheduling mechanics are in Job Scheduler and Managing Long-Running Processes.

Owner: media pipeline on-call; capacity planning owns the headroom model.

4.2 Cold Cache on a Premiere Night#

A scheduled episode releases at 20:00 local. 8M viewers press play in the first five minutes.

t=-24h:    Nobody pre-warmed: the release workflow publishes metadata at 20:00 only.
t=0:       8M play requests in 300s. Every edge cache is cold for every rung.
t=+5s:     Edge miss rate 100% for segment 1. Mid-tier coalesces per region, but 600 edge
           sites × 6 rungs × first 5 segments each go upstream.
t=+20s:    Origin egress 0.4 → 5 Tbps. Origin TTFB p99 300ms → 4s.
t=+30s:    Startup p95 2s → 9s. Players time out and retry → miss traffic doubles.
t=+2min:   Page: qoe.startup_time_p95 > 5s and origin.egress_gbps > 3x baseline.
t=+4min:   Edge caches warm. Hit ratio 97%. Startup recovers. 11% of viewers had given up.

Detection: qoe.startup_time_p95{title}, cdn.byte_hit_ratio{region}, origin.egress_gbps, origin.ttfb_p99, player.start_failure_rate.

Mitigation: request coalescing at edge and shield; origin rate limiting with Retry-After honored by players (with jitter); start viewers on a lower rung for the first 10 seconds during a known spike.

Prevention: release workflow publishes segments to origin hours early and pre-warms the first 60 seconds of the top 4 rungs into every edge site; metadata flips visibility at 20:00. For premium releases, pre-position the whole title overnight.

Owner: delivery team for coalescing and pre-warm; content operations owns the release checklist.

4.3 ABR Oscillation Storm#

A player release changes the throughput estimator to a shorter window. On a congested ISP, a million players see the same dip at the same time.

t=0:       Player v8.2 ships to 30% of Android. Estimator window 10s → 3s.
t=+2h:     Evening peak on ISP-X. Link utilization 92%.
t=+2h01:   Players on ISP-X see throughput dip, switch 1080p → 480p together.
t=+2h02:   Link frees up. Players estimate high throughput, switch back to 1080p together.
t=+2h03:   Link saturates again. Repeat every ~40s.
t=+2h20:   qoe.bitrate_switches_per_min{isp=X} 0.3 → 4.1. Rebuffer ratio 0.2% → 1.8%.
t=+2h25:   Page on rebuffer ratio by ISP. Delivery on-call caps ISP-X at 720p via hints.
t=+2h30:   Oscillation stops. Rebuffer 0.3%. Player rollback started.

Detection: qoe.bitrate_switches_per_min{isp,player_version}, qoe.rebuffer_ratio{isp}, cdn.egress_gbps{isp} showing a sawtooth.

Mitigation: server-side rung cap for the affected ISP; halt the player rollout.

Prevention: up-switch hysteresis (headroom + 20s hold) and randomized hold times so players desynchronize; player releases staged by ISP and device with QoE guardrails per stage — a player is a fleet of 75M distributed control loops.

Owner: player team for the algorithm and rollout; delivery on-call for the cap.

4.4 CDN Region Failure Causes an Origin Stampede#

A commercial CDN loses a large region; traffic shifts to our owned edge and a second CDN, both cold for that region's long tail.

Diagram: 4.4 CDN Region Failure Causes an Origin Stampede
t=0:       CDN A loses region EU-West. DNS/steering moves 3 Tbps to CDN B + owned edge.
t=+30s:    New caches miss on the tail. Origin egress 0.9 → 9 Tbps (10×). Capacity: 3×.
t=+1min:   Origin TTFB p99 8s. Startup failures 6%. Players retry → load climbs.
t=+3min:   Origin protection: priority shedding — lowest 3 rungs always served; 1080p+
           requests get 429 Retry-After 10s with jitter. Steering caps region at 720p.
t=+15min:  Caches warm. Origin 3.5 Tbps. Caps lifted in 25% steps over 20 min.

Detection: cdn.availability{provider,region}, origin.egress_gbps, origin.shed_requests_total, player.start_failure_rate{region}.

Mitigation: priority shedding at origin by rung; steering-based rung cap; stagger traffic moves (25% per minute) rather than flipping a region at once.

Prevention: keep the secondary CDN partially warm by sending it a steady 10–20% share; size origin for the largest single-region failover, not for average miss rate; game day twice a year.

Owner: delivery team; vendor management for the CDN incident report.

4.5 Storage Costs Growing Faster Than Views#

The slow failure: nothing pages, but the storage line grows 40% a year while watch time grows 15%.

Diagram: 4.5 Storage Costs Growing Faster Than Views

What's happening: every upload adds ~3 PB/day of renditions and ~2.6 PB/day of source. Without tiering, year 3 holds ~6 EB, and the bytes are overwhelmingly in videos that earn almost no views. Watch time is concentrated in recent and popular videos; storage is concentrated in old and unpopular ones.

Detection: storage.bytes{tier}, storage.bytes_per_watch_hour (the ratio that should be flat or falling), storage.cost_growth_vs_watchtime_growth reviewed monthly.

Mitigation: lifecycle policy by age and views: source to archive at 30 days (it's only needed for re-encodes); rendition sets shrink as videos cool; advanced-codec rungs deleted after 180 days unserved; dormant videos keep two rungs. Re-encoding from archive on a spike costs minutes of compute — cheap next to years of storage.

Prevention: storage cost per watch-hour as a quarterly KPI with a named owner; retention policy agreed with legal and creator relations. Storage mechanics are in Object Storage.

Owner: media platform for the policy engine; finance partner for the KPI; legal for the retention promise.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Encode backlogencode.queue_age_p99{class=fastpath} > 5minNew uploads (creators)Priority classes, shed advanced codec, preemptible burstMedia pipeline
Cold premiereqoe.startup_time_p95{title} > 5sOne title's audienceCoalescing, pre-warm, lower start rungDelivery + content ops
ABR oscillationqoe.bitrate_switches_per_min{isp} > 2One ISP × player versionRung cap hint, rollout haltPlayer team / delivery on-call
CDN region losscdn.availability{provider,region} < 99%A region's viewersStaggered steering, origin shedding by rungDelivery
Storage driftstorage.bytes_per_watch_hour rising 3 monthsMarginLifecycle tiers, deletion policyMedia platform + finance
License server slowdrm.license_latency_p99 > 500msAll DRM titles' startupCache licenses per session, regional license replicasPlayback platform
Manifest errorplayer.manifest_parse_errors spikeEvery viewer of affected titlesRoll back packager; immutable segments let you republish manifestsMedia pipeline
Live ingest droplive.ingest_gap_seconds > 2One live channelDual ingest with automatic failover; slateLive team

🎯 Staff Insight: Two classes of failures dominate: the ones you see in minutes (cold cache, CDN loss, ABR storms), which are all versions of "origin load is proportional to the miss rate"; and the one you see in quarters (storage drift), which is a policy failure. The first class needs coalescing and shedding; the second needs an owner and a KPI.


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists features: upload, transcode, playNames UGC vs premium vs live; commits; states cost thesis with an egress numberAsks what share of revenue delivery consumes and which lever moves it over 3 years
EncodingFixed ladder for every uploadFast path, priority classes, popularity threshold derived from break-even mathFunds codec or hardware programs with a payback model; owns the efficiency roadmap
Delivery"CDN in front"Shield, coalescing, byte hit ratio as SLO, origin headroom for failoverOwn vs buy vs hybrid edge as a multi-year bet; ISP relationships; CDN contract structure
QoE"Adaptive bitrate"Startup, rebuffer, start failures — SLOs sliced by ISP/device/CDN; hysteresisQoE tied to watch time and revenue; decides bitrate caps vs egress with product and finance
FailureRetries and replicasSpike, cold cache, ABR storm, CDN loss — each with detection and sheddingGame days, staggered failover policy, correlated risk across shared edge for live and VOD
OrganizationOne video teamPipeline, player, delivery, content ops with handoff SLOsPlatform contract shared by VOD, shorts and live; retention policy with legal

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Leads with cost"Egress dominates. At 1 EB a day, a cent per GB is about $4 billion a year."
Encodes by value"Advanced codec pays back after about 1,000 views; below that it's wasted compute."
Protects origin by design"Origin load is proportional to the miss rate, so byte hit ratio per region is a paging metric."
Quantifies segment choices"2-second segments double request rate versus 4. I'll pay that only for live."
Separates publish from quality"Playable in 90 seconds with two rungs; full quality fills in."
Owns the long tail"Renditions are a cache of the source. Dormant videos shrink to two rungs."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Spends 10 minutes on the metadata databaseMetadata is a few thousand QPS; the hard parts are bytes
"Stream the file from our servers"Ignores CDN, egress economics and range/segment delivery
Encodes every upload in every codec at every resolution, no numbersNo sense of compute or storage cost at 720K hours a day
"Short segments are better" without a request-rate numberTreats a cost knob as free
No QoE metricsCan't tell if the system works for viewers
WebRTC for 75M viewers of VODWrong tool: per-viewer server state, no HTTP caching

5.4 Common False Positives#

  • Codec trivia ≠ system design. Explaining B-frames, CABAC and motion vectors impresses briefly; if it doesn't lead to a cost or publish-time decision, it's trivia.
  • Naming HLS and DASH ≠ understanding delivery. The signal is what segment length does to request rate and latency.
  • "Netflix built Open Connect" ≠ a build decision. Without traffic volume and break-even reasoning, citing it suggests copying a company 1,000× your size.
  • Kubernetes for transcoding ≠ scheduling design. Running workers on Kubernetes is fine; the signal is priority classes, preemption and chunk-level retry.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minPick UGC/premium/live; numbers; cost thesis
Entities & API3–5 minVideo, rendition, encode task; immutable segments; play API
Architecture5–10 minTwo lanes: pipeline and delivery; ≤ 10 boxes
Encoding economics10–20 minFast path, ladder, popularity threshold, priority classes
Delivery20–30 minShield, coalescing, hit ratio, spikes, CDN strategy
ABR & QoE30–37 minSegment length, hybrid ABR, SLOs, server hints
Pivot37–42 minLive, premium/DRM, multi-region, cost cut
Wrap42–45 minCost levers, one-way doors, what's next

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Now make it live"Can you rebuild for latency?Real-time ladder at ingest, 1s parts, dual ingest, coalescing at every tier
"We're Netflix now"Does intent change the design?Exhaustive per-title encoding, pre-positioning, multi-DRM license server on the play path
"Cut delivery cost 30%"Cost levers in orderAdvanced codec for the head, per-title for easy content, hit ratio, bitrate caps on small screens, own/peer edge
"A video gets 5M views in an hour"Spike handlingCoalescing, priority re-encode, advanced codec queued, no origin stampede
"Viewers in one country rebuffer"Operational forensicsSlice QoE by ISP/CDN/player version; steering; ISP congestion vs CDN fault
"Shorts: 30-second vertical videos"Workload shiftPrefetch next videos, progressive or 2s segments, startup time dominates, fewer rungs

6.3 What to Deliberately Skip#

  • Upload resumability details — one sentence; link to Handling Large Blobs.
  • Codec internals — name codecs and their compute/byte tradeoff only.
  • Recommendations, comments, search — scoped out; mention that view counts feed popularity.
  • View-count accuracy — a separate aggregation problem; see Ad Click Aggregation for the pattern.
  • Metadata sharding — the metadata DB is small; a sharded PostgreSQL or Cassandra is fine, and saying why you're skipping it is the signal.

6.4 Follow-Up Questions to Expect#

  1. "How long does a 1-hour 4K upload take to become playable, and how would you make it faster?"
  2. "How much storage do you add per day, and what's the bill in year three?"
  3. "What does origin see when a premiere starts at 20:00 for 8 million viewers?"
  4. "Why 4-second segments? What changes at 2 seconds? At 10?"
  5. "How does the player decide which rendition to fetch, and how do you stop it oscillating?"
  6. "At what scale would you build your own CDN?"
  7. "How do you get live latency under 5 seconds, and what does it cost?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design YouTube."

Staff Answer

"There are three products hiding in that prompt — user uploads at scale, a premium catalogue, and live — and they differ in where the cost and risk sit. I'll design user uploads and call out where premium and live diverge.

Numbers I'll use: 500 hours uploaded a minute, about a billion watch-hours a day, 70% mobile, average delivered bitrate ~2.5 Mbps. That's over 1 EB of egress a day and roughly 190 Tbps at peak — so delivery cost is the line item that matters, and the metadata database is not the hard part. Targets: playable within 2 minutes of upload, startup p95 under 2.5 seconds, rebuffer ratio under 0.3%. I'll walk through: the encode pipeline and how much encoding each video deserves, delivery and origin protection, ABR and segment length, then failure modes and owners."

Why this is L6:

  • Commits to one intent and says what the other two would change
  • Turns "scale" into egress bytes and dollars in the first minute
  • Deprioritizes the metadata DB with a reason

What L7 adds:

  • Asks what share of revenue goes to delivery and what the 3-year target cost per watch-hour is
  • Asks whether live, shorts and VOD share one pipeline and edge today — the platform question
  • Notes that retention promises to creators are a policy decision, not an engineering default
❌ Common L5 Trap

"Users upload to an upload service which stores the file in S3, a transcoding service converts it to multiple resolutions, metadata goes in a sharded MySQL, and a CDN serves the videos. We'll use Kafka between services and autoscale everything."

Why this misses: Every box is right and there's not a single number. It can't answer "how long until playable", "what does it cost", or "what does origin see at 20:00". The interviewer has to drag every tradeoff out.


Drill 2: The Core Mechanic — Publish Time#

Prompt: "A creator uploads a 1-hour 4K video. How long until viewers can watch it, and how do you make that fast?"

Staff Answer

"Upload time is the creator's network; I'll talk about what happens after complete. Encoding a 1-hour 4K source sequentially into a full ladder would take hours. Instead, the scheduler splits the source at keyframes into ~4-second chunks — about 900 chunks — and fans out tasks per chunk per rung. The fast path is 360p and 720p H.264: 1,800 small tasks, each a few seconds of work, run in parallel across the fleet, then a stitch step writes the manifest. That's playable in a few minutes even for a 1-hour video. The rest of the H.264 ladder, including 1080p, fills in over the next 10–20 minutes at normal priority. 1440p and 2160p in an advanced codec run only if the video gets traction, because those rungs are the most expensive to encode and store.

Two details make this work: chunks must split on closed GOPs so they encode independently, and each chunk task is idempotent — keyed by (video_id, rung, chunk_idx) — so a preempted worker just means a retry of one 4-second chunk."

Why this is L6:

  • Separates time-to-playable from time-to-full-quality
  • Quantifies chunk count and parallelism
  • Makes preemption cheap by design (chunk-level idempotent retry)

What L7 adds:

  • Sets a creator-facing SLO (p50 ≤ 2 min for 10-minute uploads) and reports it per region
  • Asks whether hardware encoding changes the fast path economics enough to include 1080p

Drill 3: Make It Concrete — Storage per Day#

Prompt: "How much storage do you add per day? What about year three?"

Staff Answer

"720K hours uploaded a day. Source at ~8 Mbps is ~3.6 GB per hour → ~2.6 PB/day. The H.264 ladder from 240p to 1080p totals ~9.2 Mbps, about 4.1 GB per video-hour → ~3 PB/day. Advanced-codec rungs for the ~10% of videos that cross the popularity threshold add ~2.7 GB per hour for those → ~0.2 PB/day. Call it ~6 PB/day, ~2 EB/year, ~6 EB in year three if we keep everything at full fidelity.

That's why the storage design is a lifecycle, not a bucket: source moves to an archive tier after 30 days (it's only needed for re-encodes), cooling videos shrink their rendition set, and dormant videos keep two rungs. With that, most of the bytes sit in tiers that cost a fraction of standard storage, and the metric I'd watch is storage bytes per watch-hour — it should be flat or falling."

Why this is L6:

  • Derives the number from stated bitrates rather than guessing
  • Separates source (asset) from renditions (cache)
  • Names a ratio metric that catches drift

What L7 adds:

  • Turns retention into a written policy with legal and creator relations
  • Models the year-3 bill at each tiering policy and picks one with finance

Drill 4: Dependency Down — The Origin Is Struggling#

Prompt: "Origin TTFB p99 just went from 200ms to 6 seconds during peak. What do you do?"

Staff Answer

"First question: is origin slow because of more misses, or slow on its own? I check cdn.byte_hit_ratio by region and provider. If hit ratio dropped — a CDN region failed over, a cache purge, a premiere — it's a load problem. If hit ratio is normal, it's an origin fault: storage latency, a bad deploy, a network path.

For a load problem: confirm shields are coalescing; cap the affected region at 720p via steering hints, which cuts bytes per view ~45%; shed by rung at origin — always serve the bottom three rungs, return 429 with Retry-After and jitter for the top rungs. Players fall back a rung instead of stalling. Then lift caps in 25% steps.

For an origin fault: fail origin reads over to the replica region — renditions are replicated because they're immutable and cheap to copy — and roll back whatever changed."

Why this is L6:

  • Diagnoses miss-driven load vs origin fault before acting
  • Uses quality degradation (rung caps) as a load-shedding tool
  • Recovers in steps to avoid a second stampede

What L7 adds:

  • Makes "largest single-region failover" the origin sizing rule, written into the capacity standard
  • Runs twice-yearly game days with each CDN provider

Drill 5: Hot Key — One Video, 5M Viewers in an Hour#

Prompt: "A video uploaded 20 minutes ago goes viral: 5 million viewers in an hour. What breaks?"

Staff Answer

"Three things, in order. Delivery: every edge site misses on every segment the first time. Coalescing at the edge and the regional shield means one origin fetch per segment per region — about 50 regional shields × 5 rungs × 150 segments for a 10-minute video — trivial for origin. The danger is if coalescing is off or misconfigured: then 5M viewers × first segment is a stampede.

Encoding: the video only has the H.264 ladder. View velocity crosses the threshold in minutes; the scheduler queues the advanced-codec ladder at high priority. That saves ~30% of egress for the rest of its life — at 5M views that's tens of TB. Metadata: the play API reads video metadata 5M times in an hour — ~1,400 req/s for one key — cache it in the play service with a 30-second TTL.

What doesn't break: transcoding the viral video itself — it's already playable."

Why this is L6:

  • Identifies coalescing as the mechanism that turns a hot object into a non-event
  • Connects popularity to encoding effort in real time
  • Names the metadata hot key with a number and a cheap fix

What L7 adds:

  • Asks whether the recommendation system can signal "about to promote" so delivery pre-warms before the spike
  • Treats hot-object handling as a shared edge capability used by live as well

Drill 6: Multi-Tenant — Live and VOD on One Edge#

Prompt: "We're adding live sports on the same CDN and edge fleet as VOD. What do you worry about?"

Staff Answer

"Live and VOD want opposite things from the same caches. Live is a tiny working set — the newest few parts of each rung — requested by millions within the same second; VOD is a huge working set with a long tail. Without isolation, a big match can evict VOD's warm tail, and VOD's tail churn can slow live's hot parts.

So: separate cache namespaces or capacity reservations per class; live parts in memory with very short TTLs, VOD segments on disk; separate shields for live so a live surge doesn't saturate VOD's origin path. Live manifests and parts need request collapsing that holds requests until the part exists — blocking playlist reload in LL-HLS — or the origin sees a request per viewer per second.

Ownership: live has its own on-call and SLOs (glass-to-glass, ingest gaps), but both run on a shared edge platform with per-class quotas."

Why this is L6:

  • Explains why the workloads conflict (working set and timing)
  • Isolates by capacity and namespace rather than separate fleets
  • Names the live-specific coalescing mechanism

What L7 adds:

  • Defines the edge as a platform with per-product quotas and chargeback
  • Decides when a marquee event justifies dedicated capacity or a commercial CDN overflow contract

Drill 7: Build vs Buy — Our Own CDN#

Prompt: "Finance says the CDN bill is our second-largest expense. Should we build our own?"

Staff Answer

"Depends on volume and how predictable it is. At ~$0.01/GB, every 1 Tbps of sustained average traffic is about $40M a year. An owned edge — servers in colocation and peering at exchanges, later caches inside ISPs — might deliver at $0.002–0.004/GB all-in, but it carries a fixed cost: a 40–80 person organization, hardware refresh every 3–5 years, and years to reach coverage. Below a few Tbps average I'd negotiate harder, add a second CDN for leverage, and turn on the advanced codec for the head — that's a 30% byte cut with no capex. Above tens of Tbps, a hybrid wins: own the bulk in our top markets, keep commercial CDNs for overflow, new markets and the tail. I'd start with the five metros that carry the most traffic, measure unit cost and QoE against the CDN, and expand only if the numbers hold."

Why this is L6:

  • Gives the break-even arithmetic, including fixed cost
  • Lists cheaper levers to pull first
  • Proposes an incremental, measurable path instead of a big bang

What L7 adds:

  • Treats it as a 5-year capital program with board-level visibility
  • Plans ISP relationships and peering policy as a partnerships function, not just engineering
  • Keeps a commercial contract as a strategic hedge

Drill 8: Policy Change Without an Outage — New Codec Rollout#

Prompt: "We want to move the head of the catalogue from VP9 to AV1. How do you roll it out?"

Staff Answer

"Two separate rollouts: encoding and playback. Encoding: start re-encoding the top 1% of videos by watch time in AV1, in the backfill priority class so it never competes with uploads. Playback: the manifest advertises AV1 only to devices that report hardware decode support — software AV1 decode drains mobile batteries — so the device capability list is the gate. Stage by device family: 1% → 10% → 50% → 100%, with QoE guardrails per stage: startup time, rebuffer ratio, start failures, and battery-related abandonment where we have it. Keep the VP9 rungs until AV1 has run at 100% for a quarter, then let the lifecycle policy delete them as they go unserved.

Success metric: egress per watch-hour on AV1-capable devices, compared against a VP9 holdout."

Why this is L6:

  • Separates encode capacity from playback exposure
  • Gates on hardware decode capability
  • Stages with QoE guardrails and a holdout to prove the saving

What L7 adds:

  • Builds the business case: AV1-capable device share × byte saving × egress price vs re-encode compute
  • Coordinates with device partners and OEM decode support roadmaps

Drill 9: Cost — Cut Delivery 30%#

Prompt: "The CFO wants delivery cost down 30% next year without hurting QoE. Where do you look?"

Staff Answer

"Attribute first: bytes by codec, rung, device class, region, CDN. Then the levers in order of cost-to-implement:

  1. Advanced codec on more of the head — lower the view threshold; ~30% fewer bytes on whatever share of watch time moves.
  2. Complexity-classified ladders — screen recordings and animation are often over-encoded by 2×.
  3. Device-aware caps — a phone in portrait doesn't need 1080p; capping small screens at 720p cuts their bytes ~45% with little perceptible loss. Product signs off.
  4. Hit ratio — every point of edge hit ratio moves bytes from expensive paths (origin, transit) to cheap ones.
  5. Contract mix — shift volume to the cheapest CDN per region; renegotiate commits.
  6. Owned edge — only if volume justifies it; it's a multi-year lever, not next year's.

I'd model each one's saving, cost and QoE risk and pick a portfolio that reaches 30% with a holdout for every change."

Why this is L6:

  • Attributes before cutting
  • Orders levers by cost and risk
  • Names who signs off on quality-affecting changes

What L7 adds:

  • Sets a multi-year target for cost per watch-hour and makes it a standing KPI
  • Recognizes that bitrate caps are a product decision with brand risk

Drill 10: Multi-Region — Rights and Residency#

Prompt: "We're launching in a country that requires certain content to be blocked and some user data to stay in-country."

Staff Answer

"Two different requirements. Content rights are enforced at the play API: the video's rights.regions is checked against the viewer's region before a signed manifest is issued, and signed URLs are short-lived so they can't be shared across borders easily. That's a control-plane check; segments themselves are immutable and can be cached anywhere unless the license forbids it — premium licensors sometimes do, in which case those titles' renditions are only placed on in-country edge.

Data residency usually applies to user data — accounts, watch history, QoE beacons with identifiers — not to public video bytes. So: in-country storage for user-linked data, aggregated non-identifying QoE metrics exported to the global pipeline, and an in-country origin replica only if required for licensed content."

Why this is L6:

  • Separates content rights from data residency
  • Enforces rights in the control plane, not in caches
  • Exports aggregates instead of moving user data

What L7 adds:

  • Builds a rights and residency policy engine shared by every product line
  • Puts legal review into the market-launch checklist with a standard template

8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — The Championship Final#

Context: A live final peaks at 11M concurrent viewers, 4× the previous record. At the 80th minute, startup failures hit 9% and existing viewers report stalls. The on-call escalates to you.

Questions to Surface First:

  • Is it ingest (the stream itself), packaging, origin, or edge? live.ingest_gap_seconds vs origin.ttfb_p99 vs cdn.byte_hit_ratio.
  • Are failures concentrated on one CDN, ISP or device class?
  • Are players requesting parts that don't exist yet (clock skew, too-aggressive LL settings)?
  • Did anything change: player version, packager config, steering weights?

Typical L5 Approach: Scales origin servers horizontally and asks the CDN to add capacity. Origin recovers partially; stalls continue because the edge is the bottleneck.

Staff Approach: Finds that one CDN's edge is returning 404s for parts requested ~1s before they exist, and the players retry immediately — a retry storm into origin. Shifts 30% of traffic to the second CDN in steps, caps the top rung, and turns on blocking playlist reload so requests wait at the edge instead of retrying.

Principal Approach: Treats it as a readiness failure: the event was 4× the record with no load test at that scale. Establishes an event-readiness review for anything forecast above 2× the previous peak — capacity reservations with each CDN, a rehearsal with synthetic load, and a pre-agreed degradation ladder signed off by the sports partnership owner.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Slice start failures by CDN: 85% on CDN A. Edge logs show 404 on not-yet-existing parts, then immediate retries.
TriagePlayer version 9.1 lowered the live-edge offset to 2 parts; clock skew on some devices makes them request ahead.
Quick fixSteering: CDN A 70% → 40% in 10% steps. Server hint raises live-edge offset to 3 parts. Rung cap at 720p for 15 min.
GuardrailsPlayers must back off with jitter on 404 for future parts; edge configured to hold requests for parts within 2× part duration.
Post-mortemWhy did a latency-tightening player change ship the week of the final? Why was there no 4× load test?

Metrics to Watch: live.start_failure_rate{cdn}, cdn.status_404_rate{cdn,content=live}, origin.requests_per_sec{live}, live.glass_to_glass_p50

Organizational Follow-up: player release freeze 7 days before marquee events; event-readiness review with capacity reservations.

Ownership Question: "Who decides to trade latency for stability mid-event?" Staff answer: The live on-call, using a pre-approved degradation ladder (raise live-edge offset → cap top rung → shift CDN weights). Product signed off on the ladder in advance, so nobody negotiates during the incident.

Key Takeaway: "In live, a request for something that doesn't exist yet is the most expensive request you'll serve. Hold it at the edge; never let it become a retry."

What clears the Staff bar:

  • Slices by CDN and player version before scaling anything
  • Identifies the retry storm as the amplifier
  • Uses a pre-approved degradation ladder

Deep Dive 2: Silent Failure — The Ladder That Stopped at 720p#

Context: A creator support ticket says "my videos are blurry." Investigation shows that for 9 days, about 18% of uploads never got 1080p. No alert fired: every video was PLAYABLE.

Questions to Surface First:

  • Which uploads: by source codec, resolution, device, region?
  • Did 1080p tasks fail, never get created, or get created and starve?
  • Why didn't an alert fire — what does "success" mean in our metrics?

Typical L5 Approach: Finds a worker crash on a new phone's HEVC source variant, fixes the decoder flag, re-runs the failed tasks.

Staff Approach: Fixes the crash, then fixes the definition of success: the pipeline measured "video reached PLAYABLE" but nothing measured "video reached its expected ladder." Adds encode.ladder_completeness — the share of videos with every expected rung within 1 hour — as a paging SLO, and a reconciler that compares expected vs actual rungs hourly.

Principal Approach: Generalizes to every derived asset: thumbnails, captions, audio tracks. Establishes "expected vs actual derived assets" as a platform-level reconciliation with a dashboard per asset type, and adds a canary corpus of source files from new devices to the pipeline's release checks.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Query videos older than 1h with missing 1080p rung: 410K. All sources from two new phone models.
TriageThe 1080p task fails on a 10-bit HEVC source; after 3 attempts it goes to a dead-letter queue nobody watches.
Quick fixDecoder fix; replay the DLQ at backfill priority; prioritize videos by views.
Guardrailsencode.ladder_completeness < 99.5% pages; DLQ depth per task type alerts at > 1,000.
Post-mortemPLAYABLE was the only success signal; DLQ had no owner.

Metrics to Watch: encode.ladder_completeness, encode.dlq_depth{task_type}, encode.task_failure_rate{source_codec}

Organizational Follow-up: DLQs get an owner and an SLA on creation; new-device source samples are added to the pipeline test corpus each quarter.

Ownership Question: "Who owns a task in the dead-letter queue?" Staff answer: The media pipeline team, with a 24-hour triage SLA. A DLQ without an owner is a silent data-loss queue.

Key Takeaway: "PLAYABLE is not done. Measure the ladder you promised, not the first rung you shipped."

What clears the Staff bar:

  • Distinguishes partial success from success in the metrics
  • Finds the ownerless DLQ behind the bug
  • Adds reconciliation of expected vs actual outputs

Deep Dive 3: Large-Customer Onboarding — A Broadcaster Brings 400,000 Hours of Archive#

Context: A broadcaster partnership will upload 400K hours of archive in 6 weeks, plus 50 live channels. Their contract requires DRM on all content and geo-restriction to three countries.

Questions to Surface First:

  • How does 400K hours compare to daily capacity? (~55% of a normal day's upload volume — but spread over 6 weeks, ~1.3% extra per day.)
  • What DRM systems and security levels does the contract require? Who holds keys?
  • Which content needs advanced codecs on day one vs on popularity?

Typical L5 Approach: Lets the archive flow through the normal upload pipeline and adds DRM encryption to the packaging step.

Staff Approach: Routes the archive through a bulk-ingest class below normal uploads, rate-limited to keep the fast-path SLO intact. Adds multi-DRM packaging (CMAF cbcs so one encrypted set serves Widevine, FairPlay and PlayReady), a license service with per-title policy (regions, security level), and region checks in the play API. Live channels get dedicated ingest pairs.

Principal Approach: Sees that DRM and rights are now platform capabilities, not a partner integration. Funds a rights-and-licensing service owned jointly by the playback platform and business affairs, so the next partner is configuration, and negotiates the contract's security-level requirements with partnerships before engineering commits.

Staff Approach — Full Reasoning
PhaseWhat to Do
CapacityBulk class gets a quota of ~10% of the fleet off-peak, 0% when fast-path queue age > 2 min. 400K h × 3 core-h ≈ 1.2M core-hours ≈ 13 days of a 4K-core quota.
DRMEncrypt once with cbcs; license service p99 < 150ms in each region; license caching per session to keep it off the segment path.
Rightsrights.regions and rights.window on the video; play API enforces; signed URLs TTL 6h.
LiveDual encoders and dual ingest points per channel; automatic failover with slate on loss.
LaunchShadow the license service with synthetic sessions for a week; 3 pilot titles; then full catalogue.

Metrics to Watch: encode.queue_age_p99{class=bulk}, drm.license_latency_p99, drm.license_denials{reason}, play.rights_denials{region}

Ownership Question: "Who decides whether a title can play in a country?" Staff answer: Business affairs owns the rights data; the playback platform owns enforcement. Engineering never edits rights by hand.

Key Takeaway: "Bulk work gets its own class with its own quota. A partner's archive must never be the reason a creator waits."

What clears the Staff bar:

  • Sizes the archive against daily capacity before designing
  • Isolates bulk ingest from the creator SLO
  • Encrypts once for all DRM systems

Deep Dive 4: Post-Mortem — The Month Storage Grew 22%#

Context: The monthly cloud bill shows storage up 22% while watch time grew 1%. Finance asks for an explanation and a plan by Friday. You own the post-mortem.

Questions to Surface First:

  • Which tier and which object classes grew: source, renditions, multipart leftovers, thumbnails?
  • Did a lifecycle rule stop running or change?
  • Did upload behaviour change (more 4K, longer videos)?

Typical L5 Approach: Finds that a lifecycle rule was disabled during a migration, re-enables it, and reports the fix.

Staff Approach: Re-enables the rule, then finds three causes: the disabled rule (60% of growth), a new 4K-default camera app doubling average source bitrate (25%), and incomplete multipart uploads never aborted (15%). Adds an abort rule for incomplete multipart uploads older than 7 days, a lifecycle-rule drift check, and storage.bytes_per_watch_hour as a weekly alert.

Principal Approach: Makes storage efficiency a standing KPI owned by the media platform with a finance partner, adds lifecycle rules to infrastructure-as-code with policy checks so they can't be disabled silently, and opens the retention policy discussion with legal — the 4K-source trend will continue regardless of hygiene.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateBreak down bytes added by prefix and storage class for the month.
TriageRule disabled 34 days ago during a bucket migration; 4K sources up from 9% to 21% of uploads; 1.1 PB of incomplete multipart parts.
Quick fixRe-enable rule; abort incomplete uploads > 7 days; downscale source copies above 4K to a mezzanine after the ladder completes (with legal sign-off on fidelity).
GuardrailsIaC policy check on lifecycle rules; weekly storage.bytes_per_watch_hour review; alert at +5% month over month.
Post-mortemMigration runbook lacked a lifecycle checklist; storage growth was reviewed quarterly, not weekly.

Metrics to Watch: storage.bytes{class,tier}, storage.incomplete_multipart_bytes, upload.source_bitrate_p50, storage.bytes_per_watch_hour

Ownership Question: "Who owns the storage bill?" Staff answer: The media platform team owns the bytes and the lifecycle engine; finance owns the budget; neither can fix it alone.

Key Takeaway: "Storage doesn't page. Give it a ratio metric and an owner, or it grows until finance finds it."

What clears the Staff bar:

  • Decomposes growth into causes with percentages
  • Catches the multipart-leftover class (see Large File Uploads & Delivery)
  • Turns a one-time fix into a guardrail

Deep Dive 5: Multi-Region Expansion — Launching in a Market With Weak Transit#

Context: The platform launches in a large market where international transit is expensive and congested at evening peak. Early QoE: startup p95 6s, rebuffer ratio 2.4%.

Questions to Surface First:

  • Where is the bottleneck: CDN presence in-country, ISP interconnects, or last-mile?
  • What device mix and data-plan constraints apply? (Many viewers on prepaid mobile data.)
  • Which content is watched — is the head local or global?

Typical L5 Approach: Adds a cloud region in-country as an origin and points the CDN at it.

Staff Approach: Measures per ISP: the bottleneck is international transit into the two largest ISPs. Moves the head in-country — a commercial CDN with local presence plus, for the two largest ISPs, a direct peering arrangement. Adjusts the ladder for the market: an extra low rung (~150 kbps) and data-saver defaults on mobile. Results: startup p95 2.4s, rebuffer 0.5%.

Principal Approach: Treats market entry as a delivery-strategy decision: when a market's traffic exceeds a threshold, start the owned-edge or ISP-embedded program there; until then, commercial CDN with local presence. Writes the market-entry playbook so QoE targets, ladder adaptations and partnership steps are standard.

Staff Approach — Full Reasoning
PhaseWhat to Do
DiagnoseQoE by ISP and hour: degradation tracks evening peak on two ISPs; others fine.
DeliveryCDN with in-country PoPs; peering with the two largest ISPs; pre-warm local head nightly.
LadderAdd 144p/150 kbps rung; data-saver mode caps at 480p by default on cellular.
MeasureQoE holdouts per change; qoe.rebuffer_ratio{isp} weekly with the partnerships team.
ExpandTraffic forecast triggers owned-edge evaluation at a set Gbps threshold.

Metrics to Watch: qoe.startup_time_p95{isp}, qoe.rebuffer_ratio{isp,hour}, cdn.byte_hit_ratio{country}, egress.cost_per_gb{country}

Ownership Question: "Who owns ISP relationships?" Staff answer: A partnerships or network-strategy team with delivery engineering as the technical owner. Engineering can't sign peering agreements alone.

Key Takeaway: "In a new market, the bottleneck is usually an interconnect, not a server. Measure per ISP before building anything."

What clears the Staff bar:

  • Diagnoses per ISP and hour
  • Adapts the ladder to the market's devices and data costs
  • Brings in the partnerships owner

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why video streaming is a cost-and-tail problem and put an egress number on it within the first two minutes
  • Derive storage per day and encode compute from upload rate, ladder bitrates and core-hours per video-hour
  • Design a priority-aware encode pipeline: fast path, ladder fill, popularity-driven advanced codec, bulk and backfill classes
  • Derive a popularity threshold for extra encoding from compute cost, byte savings and egress price
  • Choose segment length with request-rate and latency numbers, and say when low-latency parts are worth it
  • Describe hybrid ABR with hysteresis and the server-side hints that can cap quality during incidents
  • Protect origin with shields, coalescing, rung-based shedding and failover sizing
  • Price own vs commercial vs hybrid edge and name the traffic level where each wins
  • Write a lifecycle and retention policy for the long tail and name who signs off

The Bar for This Question#

Mid-level (L4): Draws upload → storage → transcoder → CDN → player and explains HLS at a high level. Uses a fixed ladder and a single CDN. No numbers beyond "millions of users". Would build something that works for a few thousand videos and has no plan for spikes, cost or the tail.

Senior (L5): Adds chunked upload, a transcoding queue, multiple resolutions, a CDN, adaptive bitrate and a sharded metadata store. Knows HLS and DASH. The gap: treats every video the same, never prices egress, says "short segments" without a request-rate number, and has no origin-protection story for a premiere. Plausible and competent — and would produce a storage bill that grows faster than the business and a 20:00 outage on the first big release.

Staff+ (L6): Frames the problem around egress and the long tail in the first minutes. Splits publish time from full quality. Derives the popularity threshold for advanced encoding. Protects origin with coalescing and rung-based shedding and sizes it for the largest regional failover. Sets QoE SLOs sliced by ISP and device, and names owners for the pipeline, player, delivery and content-ops handoffs. Knows when not to build any of it. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 "Video Is a Networking Business That Happens to Use Storage"#

Cost Line (this design, rough)Annual
Egress at a blended $0.003/GB, ~1.1 EB/day~$1.2B
Storage, year 1 corpus with tiering~$100–200M
Encode compute (H.264 for all, advanced for the head)~$25–35M

The Staff position: Optimize for bytes delivered and where they're delivered from. A 10% bitrate saving is worth more than halving the encode fleet.

Why this matters in interviews: Candidates who spend their time on storage layout or the metadata DB are optimizing the smallest line items.

10.2 "Most Videos Should Never Get the Good Codec"#

Video ClassShare of Videos (assumed)Share of Watch Time (assumed)Right Ladder
Head~10%~90%Full H.264 + advanced codec, per-title for the very top
Torso~30%~9%Full H.264
Tail~60%~1%H.264, shrinking to 2 rungs when dormant

The Staff position: Encoding effort is an investment that pays back per view. Spend it where views are.

Why this matters in interviews: "Encode everything in AV1" sounds modern; it's a compute bill for videos nobody watches.

10.3 "Low Latency Is a Product Feature You Pay For, Not a Default"#

Standard HLS at 20–30 seconds is fine for most live content. LL-HLS at 3–5 seconds roughly triples to quadruples the request rate per viewer and tightens every origin and edge SLO. WebRTC under 1 second gives up HTTP caching altogether.

The Staff position: Default to standard latency; enable low latency per event where interaction, betting or spoilers make seconds matter — and have product say so.

Why this matters in interviews: "We'll use low-latency everywhere" signals that you haven't priced it.

10.4 "Owning a CDN Is a Ten-Year Decision Disguised as a Cost Project"#

The savings are real at extreme scale, but the commitment includes hardware generations, ISP relationships in hundreds of networks, a logistics operation, and an organization that will exist for a decade. The cheaper levers — codecs, hit ratio, contracts — should be exhausted first.

The Staff position: Build the edge only when traffic is large, predictable and growing, and keep a commercial CDN as both overflow and leverage.

Why this matters in interviews: Citing Open Connect without the volume math is copying; citing it with the math is judgment.

10.5 "Rebuffer Ratio Is the Only Metric the CEO Needs"#

Startup time, bitrate, resolution and start failures all matter to engineers. Stalls are what viewers remember and what most directly ends sessions.

The Staff position: One headline QoE metric — rebuffer ratio, sliced by ISP and device — with the rest as diagnostics. Every cost-cutting change ships with a rebuffer-ratio guardrail.

Why this matters in interviews: Naming one metric that leadership can track shows you can communicate a technical system upward.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a pipeline and an edge that meet QoE targets at a defensible cost. The Principal engineer notices that the company's margin is set by cost per watch-hour, that it is driven by three levers with very different time horizons — codec efficiency (quarters), edge ownership (years) and retention policy (a legal and creator-relations decision) — and that the company currently runs three video stacks: VOD, shorts and live, each with its own encoder settings, its own CDN contract and its own idea of quality. The L7 problem is setting the video platform's economic and organizational shape for the next three to five years: one pipeline with workload classes, one edge with per-product quotas, one QoE definition, and a cost-per-watch-hour target that every team can see.

🧭 Principal Move: "I'd make cost per watch-hour and rebuffer ratio the two numbers every video team reports, and fund the levers in order of payback: advanced codec for the head this year, a hybrid edge in our top markets over three years, and a retention policy for the dormant tail that legal and creator relations sign. Each lever has a different owner, so the plan has to be written as a portfolio, not as one project."

The Org-Level Fault Line#

One video platform vs per-product video stacks.

OptionWhat WorksWhat BreaksWho Pays
Each product runs its own stack (VOD, shorts, live, ads)Speed; each team tunes for its workload3–4 encoders, 3–4 CDN contracts, inconsistent QoE metrics, no volume leverageFinance (weaker contracts), viewers (inconsistent quality)
One central video team owns everythingOne edge, one contract, one QoE definitionBottleneck; live's latency needs fight VOD's efficiency needs in one roadmapProduct teams (velocity)
Platform owns primitives; products own experiencesPipeline with workload classes, packaging, edge, QoE telemetry are platform; ladders, latency targets and player UX are product-configurableContract design and quota governance are hardPlatform team (API stewardship, chargeback)

🧭 Principal Insight: The platform should own everything whose cost scales with bytes — encode infrastructure, packaging, edge, contracts — and expose knobs for everything whose value differs by product: ladder templates, latency class, startup strategy.

Cost Model#

Assumptions: average delivered bitrate 2.5 Mbps; source 8 Mbps; H.264 ladder 9.2 Mbps total; software encode ~3 core-hours per video-hour at $0.02 per core-hour; storage blended across tiers; fully loaded engineer ~$250K/year.

ScaleUploads / WatchEgress ($/month)Storage ($/month)Encode ($/month)HeadcountOn-call
Startup500 h/day uploaded · 1M watch-h/day (~34 PB/month delivered)~$170–350K (commercial CDN, $0.005–0.01/GB)~$5–15K~$1K (or a managed service)3–5 eng; buy encoding and playerShared rotation
Growth50K h/day · 100M watch-h/day (~3.4 EB/month)~$10–25M (multi-CDN, committed)~$1–3M~$100K30–50 eng: pipeline, player, delivery, QoE dataPipeline and delivery rotations
Hyperscale720K h/day · 1B watch-h/day (~34 EB/month)~$70–150M (hybrid owned edge + commercial overflow)~$10–25M and growing~$1.5–3M (more with advanced codecs)300+ eng incl. edge hardware, peering, codec researchFollow-the-sun per layer

The pricing insight: at every scale in this model, egress costs roughly 50× or more what encode compute costs. A team of 5 engineers (~$1.25M/year) that lowers average bitrate 5% saves ~$0.5–1.2M/month at growth scale and ~$3.5–7.5M/month at hyperscale. At startup scale the same team saves ~$10K/month — buy the managed service and spend the engineers on product.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Segment container (CMAF fMP4 vs TS) and encryption scheme (cbcs vs cenc)One-wayRepackaging the whole catalogue; player and DRM changes on every device family
Owning an edge network / ISP-embedded cachesOne-wayMulti-year capital and partner commitments; unwinding strands hardware and relationships
Retention promise to creators ("we keep your video forever at full quality")One-wayPublic commitment; changing it is a trust and legal event
Deleting sourcesOne-wayLost forever; renditions can be regenerated, sources cannot
Ladder rungs and bitratesTwo-wayRe-encode the head; tail follows lifecycle
Segment duration for VODTwo-wayRepackage (not re-encode) if GOP structure allows
CDN vendor mix and steering weightsTwo-wayConfig change, contract cycle
Popularity threshold for advanced codecTwo-wayConfig change; re-derive when prices change

The Standard I'd Write#

RFC-VID-001: Video Encoding, Packaging and Delivery Standard
Status: Approved   Owners: Video Platform + Delivery Engineering

Scope
  Every product that stores or serves video: VOD, shorts, live, ads, previews.

MUST
  1. Encode through the shared pipeline using a workload class
     (fastpath, ladder, advanced, bulk, backfill); no product runs its own encoders.
  2. Package as CMAF; encrypt with cbcs where DRM is required.
  3. Serve segments only through the platform edge or contracted CDNs via steering;
     no direct origin URLs to clients.
  4. Emit standard QoE beacons (startup, stalls, bitrate switches, start failures)
     with player version, device, ISP and CDN dimensions.
  5. Keep the source for every video; treat renditions as regenerable.
  6. Attach a lifecycle policy to every rendition set at creation.

SHOULD
  1. Use 4–6 s VOD segments with 2 s keyframes; 1 s parts only for low-latency live classes.
  2. Gate advanced codecs on device hardware decode capability.
  3. Ship player changes in stages per device family with rebuffer-ratio guardrails.

Exceptions
  Filed with Video Platform; decided within 10 business days; time-boxed to 2 quarters.
  Exceptions to MUST 5 need legal sign-off.

Success metrics
  - Cost per watch-hour: −15% year over year
  - Rebuffer ratio ≤ 0.3% globally, ≤ 0.6% in every market with > 1M daily viewers
  - Publish time p50 ≤ 2 min for uploads under 15 min
  - Storage bytes per watch-hour: flat or falling quarter over quarter
  - Products running their own encoders or CDN contracts: 0 by end of year 2

What I'd Tell the VP#

"Delivering video is our largest cost after content, and it scales with every hour watched. We can bend that curve with three levers on three timelines: better compression for our most-watched videos this year, a lower-cost delivery network in our biggest markets over three years, and a clear retention policy for videos nobody watches. Together they target a 15% annual reduction in cost per hour watched while holding stall rates under 0.3%. The first lever needs about five engineers; the second is a multi-year capital program I'd bring back as a separate proposal with market-by-market numbers. The retention policy needs legal and creator relations, not engineering."

Principal Interview Signals#

SignalWhat It Sounds Like
Prices the business, not the boxes"Egress is fifty times our encode bill. Our cost per watch-hour is the number I'd manage."
Sequences levers by horizon"Codec this year, edge over three, retention as policy. Different owners, one plan."
Identifies one-way doors"Container and encryption format, the edge, the retention promise, deleting sources. Everything else is config."
Redraws ownership"Platform owns what scales with bytes; products own what differs in value."
Brings in non-engineering owners"Licensors set DRM, legal sets retention, partnerships owns ISPs. I design around their decisions."

Staff answers that L7 interviewers find insufficient:

  • "We'll use AV1 to save bandwidth" — correct lever, no payback model, no device-capability plan, no owner.
  • "We'll build our own CDN like Netflix" — no traffic threshold, no fixed-cost estimate, no plan for markets without boxes.
  • "We'll keep everything in cold storage" — ignores that the retention promise is a policy decision with creators and legal.

Appendices

Appendix A: Mechanics in Depth#

A.1 The Chunked Transcode DAG#

Diagram: A.1 The Chunked Transcode DAG
def plan(video):
    if dedup.exists(video.content_hash):                 # re-upload of identical bytes
        return link_renditions(video, dedup.get(video.content_hash))
    chunks = split_points(video.source, target_s=4)      # closed GOP boundaries only
    for rung in FASTPATH:                                 # 360p, 720p H.264
        for i in range(len(chunks)):
            enqueue(task(video.id, rung, i), cls="fastpath")
    for rung in LADDER - FASTPATH:
        for i in range(len(chunks)):
            enqueue(task(video.id, rung, i), cls="ladder")

def run(task):                                            # idempotent per (video, rung, chunk)
    key = f"{task.video_id}/{task.rung}/{task.chunk}"
    if output_exists(key):                                # replay after preemption → no-op
        return
    with lease(key, ttl="2m", heartbeat="20s"):
        out = encode(read_chunk(task), task.rung)
        put_if_absent(key, out)                           # write-once segment objects
    if all_chunks_done(task.video_id, task.rung):
        package_and_publish(task.video_id, task.rung)     # single writer per rung

The idempotent-task and lease mechanics are general; see Workflows, Sagas & Compensation and Job Scheduler.

A.2 Hybrid ABR With Hysteresis#

def next_rung(state, rungs):
    tput = ewma_throughput(window_s=10)                   # bytes/s, both fast and slow EWMA, take min
    if state.phase == "startup":
        return highest(rungs, bitrate <= 0.7 * tput)      # conservative start
    if state.buffer_s < 3:
        return rungs.lowest                               # panic
    if state.buffer_s < 8:
        return highest(rungs, bitrate <= 0.8 * tput)      # immediate down-switch
    up = rungs.above(state.current)
    if (up and state.buffer_s > 25 and tput > 1.4 * up.bitrate
            and now() - state.last_switch > hold_s()):     # hold 20s + random 0–10s
        return up
    return min(state.current, hints.rung_cap or rungs.highest)

The random component in the hold time desynchronizes players that share a bottleneck — the fix for the oscillation storm in 4.3.

A.3 Live Pipeline#

Diagram: A.3 Live Pipeline

Budget for 4 s glass-to-glass with LL-HLS: capture + contribution ~500 ms, real-time encode ~500 ms, packaging one 1 s part ~1 s, CDN ~200 ms, player buffer ~2 parts ≈ 2 s.

Appendix B: Data Model#

videos           (video_id PK, owner_id, state, duration_s, source_uri, source_tier,
                  content_hash, visibility, rights_regions[], drm_required,
                  popularity_tier, created_at, last_viewed_at)
renditions       (video_id, codec, height, bitrate_kbps, state, segment_prefix,
                  storage_tier, bytes, created_at, last_served_at,
                  PRIMARY KEY (video_id, codec, height))
encode_tasks     (task_id PK, video_id, rung, chunk_idx, class, attempt,
                  lease_owner, lease_expires_at, status)
view_counters    (video_id, hour_bucket, views)            -- feeds popularity tiers
qoe_beacons      → stream to Kafka → aggregates by ISP, device, CDN, player version

Metadata is small: ~100M+ videos × ~2 KB ≈ hundreds of GB, read-heavy and cacheable — PostgreSQL sharded by video_id or Cassandra both work. View counts and QoE beacons are high-volume streams through Kafka into aggregation; the pattern is the same as in Ad Click Aggregation.

Appendix C: Coordination Mechanisms#

MechanismUsed ForWhy
Kafka topic video.uploadedTrigger planningDurable, replayable; partition by video_id
Priority task queues per classEncode schedulingFast path can't starve; lower classes preemptible
Leases with heartbeatWorker ownership of a chunkPreempted worker's chunk is retried by another after 2 min
Write-once segment objectsOutputDuplicate execution is harmless
Single packager per rungManifest writesAvoids concurrent manifest edits
Conditional state updates on videos.statePROCESSING → PLAYABLE → LADDER_COMPLETEOut-of-order completion events are no-ops
Hourly reconcilerExpected vs actual rungsCatches silent ladder gaps (Deep Dive 2)

Workers run well on Kubernetes with preemptible node pools for the ladder, advanced and backfill classes and on-demand nodes for the fast path.

Appendix D: API Contract & Client Behavior#

GET /videos/{id}/play?device=android&hdr=0&codecs=avc1,vp09,av01
200 {
  "manifest_url": "https://cdn-a.example/v/abc123/master.m3u8?sig=...&exp=6h",
  "alt_manifests": ["https://cdn-b.example/v/abc123/master.m3u8?sig=..."],
  "license_url": null,
  "start_rung": "720p",
  "rung_cap": null,
  "hints_ttl_s": 300
}

Client rules: start at start_rung or lower; on segment failure try the same segment once on the next CDN in alt_manifests before switching down; on 429 honor Retry-After with ±30% jitter; never retry a future live part faster than half the part duration; refresh hints every hints_ttl_s and keep playing with stale hints if the hint call fails.

Caching rules: segments Cache-Control: public, max-age=2592000, immutable; VOD manifests short TTL (~2 s) until LADDER_COMPLETE, then long; live manifests ≤ part duration with blocking reload.

Appendix E: Observability#

E.1 Core Metrics#

# Viewer experience (client beacons)
qoe.startup_time_p50 / p95{device,isp,cdn}
qoe.rebuffer_ratio{device,isp,cdn,player_version}
player.start_failure_rate{reason}
qoe.bitrate_avg_kbps / qoe.bitrate_switches_per_min

# Delivery
cdn.byte_hit_ratio{provider,region}
origin.egress_gbps / origin.ttfb_p99 / origin.shed_requests_total
egress.cost_per_gb{provider,country}

# Pipeline
encode.queue_age_p99{class}
publish.time_to_playable_p50
encode.ladder_completeness
encode.dlq_depth{task_type}

# Economics
storage.bytes{tier} / storage.bytes_per_watch_hour
cost.per_watch_hour (monthly)

E.2 Critical Alerts#

AlertThresholdPage
Rebuffer ratio by ISP> 1% for 10 min on an ISP with > 100K viewersDelivery on-call
Startup p95> 5 s for 5 min globally or per titleDelivery on-call
Byte hit ratioDrop > 3 points in 10 min in a regionDelivery on-call
Fast-path queue agep99 > 5 min for 5 minMedia pipeline
Ladder completeness< 99.5% over 1 hMedia pipeline
Live ingest gap> 2 s on a channelLive on-call

E.3 Debugging "Users Say It Buffers"#

Slice qoe.rebuffer_ratio by ISP → CDN → device → player version → title. A single ISP at evening peak means interconnect congestion (steer, cap); a single CDN means a provider issue (shift weights); a single player version means a regression (halt rollout); a single title means a cold cache or a bad encode. Client-side telemetry is the only place the viewer's experience exists; see Metrics & Alerting Platform.

Appendix F: Scale Evolution#

ScaleWhat Works
< 1K uploads/day, < 10 GbpsManaged encoding service, one CDN, fixed ladder, progressive MP4 for short clips
1K–100K uploads/dayOwn pipeline with priority classes, CMAF, multi-CDN, QoE beacons
100K+ uploads/day, Tbps egressPopularity-driven codecs, lifecycle tiers, shields with coalescing, steering hints
Tens of Tbps+Hybrid owned edge, ISP peering and embedded caches, hardware encode evaluation

What you don't build on day one: per-shot encoding, owned edge, hardware encoders, server-side ABR, low-latency live, your own DRM license server. Each one has a trigger in the 3-year path.

Multi-region: metadata and the play API run active-active per region; renditions are replicated to two or more origin regions because they're immutable (replication is simple — see Replication); sources live in one home region plus archive copy. See Multi-Region.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Encode fairness: per-creator quotas on the bulk class so one partner's archive can't starve others; fast path is never quota-limited for ordinary uploads, but upload rate limits apply per account (Rate Limiting).
  • Edge fairness: per-product cache quotas (VOD, shorts, live, ads) so one class can't evict another; live gets reserved memory capacity during marquee events.
  • Cost allocation: chargeback by bytes delivered and core-hours consumed per product; storage by bytes-month per tier. Teams that can't see their bytes can't reduce them.
  • Tradeoff summary:
LeverSavesCostsWho Signs Off
Advanced codec for the head~30% bytes on most watch timeCompute; device gatingVideo platform
Per-title ladders20–40% on easy contentCompute per title; quality metric pipelineVideo platform
Small-screen bitrate caps~45% bytes on capped sessionsPerceived qualityProduct
Lifecycle tiers and dormant laddersMost storage growthRe-encode on revivalMedia platform + legal
Owned edge50–80% unit cost on owned trafficMulti-year capex and orgExecutive leadership
  1. Loading the index…