Problem & requirements
Focus on two capabilities: uploading videos and watching them. Clarify the clients (mobile, web, smart TV), DAU (assume 5 million), average time spent, international reach, supported resolutions, whether encryption is needed, and the maximum upload size (assume 1 GB). The service should support fast uploads, smooth streaming, changeable quality, low infrastructure cost, high availability, and reliability.
Leveraging existing cloud services such as blob storage and a CDN is the right call; no interviewer expects you to build a global CDN in an hour, and real companies, Netflix included, rely heavily on cloud infrastructure.
Back-of-the-envelope estimation
Assume 5M DAU each watching 5 videos a day, 10% of users uploading one video a day, and an average video size of 300 MB. Daily new storage is 5M × 0.1 × 300 MB = 500,000 × 300 MB = 150 TB per day, before transcoding multiplies it.
CDN cost dominates. Daily views are 5M × 5 = 25M; at 300 MB each that is 7.5M GB served. At about $0.02 per GB, that is 25,000,000 × 0.3 GB × $0.02 = $150,000 per day, or roughly $55 million a year. Check: 7.5 × 10^6 GB × 0.02 = 1.5 × 10^5 dollars. That number is why later optimisations focus on serving only popular content from the CDN.
High-level design
Three pieces form the system. The client runs on phones, browsers, and TVs. The CDN stores and streams video bytes from edges close to viewers. API servers handle everything else: feed recommendations, upload URLs, metadata updates, user sign-up.
Uploading touches more components. Original storage is blob storage for raw uploads. Transcoding servers convert videos into multiple formats and bitrates. Transcoded storage holds the outputs, which are then distributed to the CDN. A completion queue records finished transcoding events, and a completion handler consumes them to update the metadata DB and metadata cache, which hold video URLs, titles, sizes, and owner info.
Uploading and streaming flows
Upload runs as two parallel processes. The video bytes go to original storage, then through transcoding, then to transcoded storage and the CDN, with a completion event triggering metadata updates. In parallel, the client sends metadata such as title and description to API servers, which write it to the metadata cache and database.
For streaming, the player does not download the whole file first; it fetches small chunks continuously. Streaming protocols such as MPEG-DASH, Apple HLS, Microsoft Smooth Streaming, and Adobe HDS define how video is segmented and requested. With adaptive bitrate streaming, the video is encoded at several quality levels and the player switches between them based on measured bandwidth, so a viewer on a weak network sees lower resolution instead of a frozen screen. Videos stream directly from the nearest CDN edge.
Video transcoding and the DAG model
Raw video is huge and comes in many formats, while devices support different codecs and resolutions and networks vary. Transcoding compresses video, converts it to compatible formats, and produces multiple bitrates for adaptive streaming. Encoding formats have two parts: a container such as MP4 or WebM, and codecs such as H.264, VP9, or HEVC.
Different creators need different processing: watermarks, custom thumbnails, high-definition variants. To keep the pipeline flexible and parallel, tasks are described as a directed acyclic graph (DAG), similar to Facebook's streaming video engine. The original video is split into video, audio, and metadata; video tasks such as inspection, encoding at several resolutions, thumbnail generation, and watermarking can run concurrently where dependencies allow, and outputs are merged at the end.
Transcoding architecture components
The preprocessor splits the video into small Group of Pictures (GOP) aligned segments that can be played independently, generates the DAG from configuration files written by engineers, and caches segments and metadata in temporary storage so failed tasks can be retried without re-uploading.
The DAG scheduler splits the graph into stages of tasks and sends them to the resource manager. The resource manager keeps a task queue (priority queue of tasks), a worker queue (worker utilisation), and a running queue (what is currently executing); a task scheduler picks the best task and worker pair, dispatches the job, and records it as running. Task workers execute the tasks, such as encoding or watermarking, and the result becomes the encoded video, for example funny_720p.mp4.
Optimisations: speed, safety and cost
For speed, the client splits a video into GOP-aligned chunks and uploads them in parallel, so a failed chunk can be resumed rather than restarting. Upload centers close to users, often CDN edges, shorten network paths. Message queues between stages decouple them, so encoding can start on chunks while later chunks are still uploading.
For safety, presigned URLs let clients upload directly to object storage for a short time without exposing credentials, and only authorised users get them. Copyright protection uses DRM systems such as FairPlay, Widevine, and PlayReady, AES encryption with keys released only to authorised players, and visual watermarks.
For cost, remember that video popularity follows a long tail. Serve only the most popular videos from the CDN and less-watched videos from your own high-capacity storage servers. Encode fewer versions of rarely watched content, distribute some videos only regionally, and consider building your own CDN or partnering with ISPs, as large platforms do.
Error handling & wrap-up
Errors are either recoverable, such as a segment failing to transcode, which is retried a few times before returning an error, or non-recoverable, such as a malformed video file, which stops the job and returns an error to the client. Component-specific handling follows the same logic: retry failed uploads, split videos server-side if an old client cannot, re-fetch from temporary storage after a task worker dies, let replicas take over when an API server or metadata DB node fails, and use standby nodes for the resource manager.
Further discussion points include scaling the API tier horizontally, database replication and sharding, live streaming (lower latency requirements, smaller processing chunks, different protocols), and takedowns for copyright or illegal content detected at upload or via user reports.
Key numbers
Key terms
- Transcoding
- Converting a video into other formats, codecs, and bitrates for compatibility and efficient streaming.
- Adaptive bitrate streaming
- Switching between pre-encoded quality levels during playback based on the viewer's bandwidth.
- MPEG-DASH / HLS
- HTTP-based streaming protocols that serve video as a sequence of small segments with a manifest.
- DAG
- A directed acyclic graph of tasks whose edges express dependencies, enabling parallel processing.
- GOP (Group of Pictures)
- A self-contained run of frames that can be decoded independently, used as the unit for chunking.
- Presigned URL
- A time-limited URL granting permission to upload or download a specific object in storage.
- DRM
- Digital rights management systems that restrict playback of protected content to authorised devices.
- Completion handler
- A worker that consumes transcoding-finished events and updates metadata stores.
Common mistakes
- Routing video bytes through API servers instead of uploading directly to blob storage.
- Forgetting that CDN egress, not storage, is the dominant cost.
- Describing transcoding as one monolithic step rather than a parallel DAG of tasks.
- Serving every video, including the long tail, from the CDN.
- Not distinguishing recoverable from non-recoverable errors.
- Ignoring adaptive bitrate and assuming one encoded file fits every device and network.
Further study
- SVE: Distributed Video Processing at Facebook Scale (SOSP 2017)
- Netflix Open Connect
- Netflix TechBlog on per-title encode optimization
- Apple HTTP Live Streaming specification (RFC 8216)
- Warehouse-Scale Video Acceleration: Co-design and Deployment in the Wild (Google, ASPLOS 2021)