AI System Design

Theory

Study notes for every chapter: the requirements, the numbers, the design and the trade-offs. Read a chapter, then practise it in the simulator, the design challenge and the mock interview.

Volume 1

1

Scale from Zero to Millions of Users

Growing a web product from one box to millions of users is not one big leap but a sequence of small, well-understood moves: split the database out, add redundancy, cache aggressively, push static bytes to the edge, make web servers stateless, and decouple work with queues. Each step removes the current bottleneck or single point of failure and exposes the next one. This chapter is the vocabulary every later design builds on.

7 min read
2

Back-of-the-Envelope Estimation

Back-of-the-envelope estimation is the habit of turning vague requirements into rough numbers (requests per second, storage, bandwidth, machines) before drawing boxes. The goal is the right order of magnitude, not precision, so that design choices are grounded in reality. It rests on three toolkits: powers of two, latency intuitions, and availability arithmetic.

6 min read
3

A Framework for System Design Interviews

A system design interview is a simulated collaboration on an ambiguous problem, and the interviewer is judging how you think as much as what you draw. A four-step structure (clarify scope, propose a high-level design and get buy-in, dive deep, wrap up) keeps the conversation focused and makes sure you spend time where it earns signal. This chapter turns that structure into a time budget and a set of habits.

5 min read
4

Design a Rate Limiter

A rate limiter caps how many requests a client may make in a time window, protecting services from abuse, runaway clients and cost blowouts. The design centers on where the limiter sits, which counting algorithm it uses, and how counters stay correct when many limiter instances share them. Redis-backed counters with atomic scripts and clear 429 responses are the standard answer.

6 min read
5

Design Consistent Hashing

Consistent hashing maps both servers and keys onto the same circular hash space so that adding or removing a server moves only a small fraction of keys. It replaces the fragile hash(key) % N scheme, under which nearly every key relocates when N changes. Virtual nodes make the distribution even and are what production systems actually deploy.

5 min read
6

Design a Key-Value Store

A distributed key-value store exposes just put(key, value) and get(key), yet building one that is scalable, highly available and tunably consistent touches nearly every core distributed-systems idea. The classic Dynamo-style answer combines consistent hashing, replication with quorums, vector clocks for conflicts, gossip for failure detection, hinted handoff and Merkle trees for repair, and an LSM-tree storage engine. Every piece exists to keep serving requests while machines and networks fail.

7 min read
7

Design a Unique ID Generator in Distributed Systems

Database auto-increment works on one server but fails once writes are spread across many machines and data centers. This chapter compares multi-master increments, UUIDs, a central ticket server and Twitter's Snowflake layout, and shows why a 64-bit, time-ordered ID assembled from timestamp, machine and sequence bits wins. The interesting details are bit budgeting and clock behavior.

5 min read
8

Design a URL Shortener

A URL shortener maps a long URL to a short, unique alias and redirects anyone who visits the alias. The interesting decisions are how to generate collision-free short codes, which HTTP redirect status to return, and how to serve a read-heavy workload cheaply with caching.

6 min read
9

Design a Web Crawler

A web crawler starts from seed URLs, downloads pages, extracts new links, and repeats, eventually covering billions of documents. The hard parts are being polite to websites, prioritising what matters, avoiding duplicate work, and surviving the endless strangeness of the real web.

6 min read
10

Design a Notification System

A notification system delivers push, SMS, and email messages on behalf of many internal services through third-party providers like APNs, FCM, Twilio, and SendGrid. The design moves from a fragile single server to a decoupled pipeline of per-channel queues and workers, then adds the reliability, deduplication, rate limiting, and monitoring that make it trustworthy.

6 min read
11

Design a News Feed System

A news feed shows each user a constantly updating list of posts from the people they follow. The central decision is when to do the work of assembling feeds, at write time (fan-out on write), at read time (fan-out on read), or a hybrid that treats celebrities differently, backed by several layers of caching.

5 min read
12

Design a Chat System

A chat system delivers messages between users in real time, supports small group chats, shows who is online, and keeps several devices per user in sync. It relies on long-lived WebSocket connections to stateful chat servers, a key-value store for message history, per-channel ordered message IDs, and heartbeat-based presence.

6 min read
13

Design a Search Autocomplete System

Search autocomplete suggests the most popular completions for whatever a user has typed so far, within about a hundred milliseconds of each keystroke. The design separates an offline data-gathering pipeline that builds a trie with cached top-k results from a fast query service that only reads it.

5 min read
14

Design YouTube

A video platform must accept large uploads, transcode each video into many formats and resolutions, and stream them smoothly to devices all over the world. The design leans on blob storage, a DAG-based transcoding pipeline, CDNs for delivery, and cost-aware choices about which videos live where.

6 min read
15

Design Google Drive

A cloud file store must keep every user's files durable, private and identical across phones, laptops and the web, while spending as little bandwidth and disk as possible. The core ideas are splitting files into content-addressed blocks so only changed pieces move, keeping a strongly consistent metadata database as the source of truth, and pushing change notifications to every device so they can pull what is new.

7 min read
16

Design a Proximity Service

A proximity service answers 'what businesses are near me?' for a point and a radius, the backbone of local search in apps like Yelp. The whole problem reduces to indexing two-dimensional points so that range queries are cheap; geohash, quadtree and Google S2 each solve it with different trade-offs, and the read-heavy, rarely changing data lets caching and replicas do the rest.

7 min read

Volume 2

17

Design Nearby Friends

Nearby friends shows a user which of their opted-in friends are currently within a few miles, updating as everyone moves. Unlike a proximity service, every data point moves constantly, so the design is a real-time fan-out system: WebSocket connections, a TTL location cache, and a Redis pub/sub channel per user that pushes each location update to that user's friends.

6 min read
18

Design Google Maps

A maps product combines three systems: a high-volume location ingestion pipeline, a map rendering path that serves pre-computed tiles from a CDN, and a navigation service that runs shortest-path searches over hierarchical routing tiles and predicts arrival times from live traffic. Most of the difficulty is in pre-processing massive geographic data so that each online request is cheap.

6 min read
19

Design a Distributed Message Queue

A distributed message queue decouples producers from consumers and absorbs bursts, but modern designs in the Kafka mould go further: they persist messages in a replicated, append-only log so consumers can replay history. The design rests on partitioned topics, sequential disk IO with batching, consumer groups with rebalancing, and leader-follower replication whose acknowledgement settings trade latency for durability.

6 min read
20

Design a Metrics Monitoring and Alerting System

A metrics system collects numeric time series such as CPU, request latency and queue depth from a large fleet, stores them efficiently for a year, and turns them into dashboards and alerts. The design is shaped by a constant, heavy write load, bursty read load, and the observation that time-series data compresses extremely well when stored in a purpose-built database.

6 min read
21

Design an Ad Click Event Aggregation System

Online advertising bills by clicks, so a click aggregation system must count billions of events per day correctly, quickly and auditably. The design is a streaming MapReduce pipeline between two Kafka queues, with careful handling of event time, late data, exactly-once processing and reconciliation against a batch recomputation.

6 min read
22

Hotel Reservation System

A hotel booking platform looks like ordinary CRUD until two guests try to grab the last room at the same moment. The interesting design work is in modelling inventory per room type per night, making reservations idempotent, and choosing a concurrency-control strategy that stays correct without sacrificing too much throughput.

7 min read
23

Distributed Email Service

Designing a Gmail-scale email service means replacing the classic one-server-per-mailbox model with stateless web tiers, queues, a partitioned metadata store and an object store for attachments. The deep problems are choosing a metadata database that serves per-user reads at enormous scale, keeping mail flowing reliably through sending and receiving pipelines, and protecting deliverability against spam.

7 min read
24

S3-like Object Storage

Object storage trades the rich semantics of file systems for massive scale, durability and low cost, exposing a flat namespace of immutable objects over HTTP. Building one means separating metadata from data, packing small objects into large files, and choosing replication or erasure coding to hit durability targets across failure domains.

7 min read
25

Real-time Gaming Leaderboard

A live leaderboard must rank millions of players and answer top-N and my-rank queries instantly as scores change. Relational ORDER BY queries collapse at this scale, so the design centres on Redis sorted sets, with careful thought about server-authoritative scoring, persistence, and how to shard when one node is no longer enough.

6 min read
26

Payment System

A payment backend for an e-commerce site has low throughput but extremely high stakes: money must never be lost, charged twice or left unaccounted for. The design rests on a payment service that orchestrates external providers, a double-entry ledger, idempotency everywhere, careful retry and failure tracking, and daily reconciliation as the final safety net.

6 min read
27

Digital Wallet

A digital wallet must move money between accounts at a million transfers per second while never losing or creating a cent and being able to prove every balance. The design journey runs from naive sharded caches, through distributed transaction protocols like TC/C and Saga, to event sourcing replicated with Raft, which delivers both throughput and auditability.

7 min read
28

Stock Exchange

An electronic exchange matches buy and sell orders with microsecond-level, highly predictable latency while guaranteeing fairness and a perfectly deterministic record. The design keeps the critical path on a single server, sequences every event, runs a matching engine over in-memory order books, and rebuilds state from an event log for high availability.

7 min read