Challenges
Apply what you've learned with no hand-holding.
Boss battles
Each win is checked against the live simulator — no honor system.
Survive 10K requests/sec
Traffic is ramping to 10,000 req/s. Your starting server does 1,000 req/s.
Fight itRight-size a 150M-user service
Set daily users to 100M+ and provision servers and storage for your workload.
Fight itStop API abuse
10,000 req/s of bot traffic is hitting a 3,000 req/s API alongside real users.
Fight itSurvive a hot-key stampede
10,000 req/s on a 2,000 req/s DB, and a hot key is about to expire.
Fight itRebalance with minimal movement
Scale your cache tier from 3 to 6 nodes.
Fight itRe-balance the bits
30 datacenters, 50 machines each, 2,000 IDs per millisecond per machine — and a clock is about to jump backwards.
Fight itLose a server under load
6K req/s of mixed fast and slow requests. Then API 1 crashes.
Fight itDrain the backlog
1M jobs are queued and every consumer just crashed.
Fight itThe primary just died
9K reads/s and 1K writes/s. The primary crashes.
Fight itDesign a shard key
12K writes/s, and a celebrity account is driving 30% of them.
Fight itPayments ledger under partition
You run a bank's balance store across two regions. The link between them just failed.
Fight itSurvive a 2-node outage
5 replicas per key, 2 of them down, and two clients writing the same key.
Fight itDesign real systems
One challenge per system in both volumes. Make the calls, avoid the traps, get a design review.
🗝️ Key-Value Store · Chapter 6
Build your own Dynamo. Nodes will die.
Starting point: Client → Single server with an in-memory hash map. Pick the components you'd add. Some of these are traps — each choice has a real trade-off.
- ▸10 TB of data, growing 2x per year
- ▸1M reads and 500K writes / second
- ▸Get/put p99 under 10ms
- ▸Stay writable when nodes or a whole rack fail