AI System Design

Missions

Every mission is a real engineering problem. Solve it in the simulator.

L1

Scaling: One server isn't enough. What would you do?

Your app just got popular. A single server handled the early users fine — now traffic is climbing past what one machine can serve. Keep it healthy.

L1

Back-of-the-Envelope Estimation: How many servers do you need? Guess — then prove it.

Your manager asks how much hardware a new service needs before anything is built. Turn users, actions and payload sizes into QPS, servers and storage — then provision it and see if you were right.

L2

Rate Limiter: Your API is under attack. Good luck.

A botnet just pointed 5,000 req/s at your API on top of your real users. Everything is falling over. Protect the system without locking out real customers.

L2

Caching: Your database is dying. Add a cache.

Reads are hammering your database far beyond what it can serve. Most requests ask for the same popular data over and over. Save the database.

L2

Consistent Hashing: You have 4 servers and 10,000 keys. Now add server #5.

You're distributing cache keys across servers. With plain modulo hashing, adding one server reshuffles almost everything. Find a better way.

L2

Unique ID Generator: Mint 2 million IDs a second. No duplicates. Ever.

Every order, message and user needs a unique 64-bit ID that sorts by time — across dozens of datacenters, with no central bottleneck. Pick a scheme, then break it.

L2

Load Balancing: Spread the work so no single server drowns.

Traffic is uneven and one server keeps melting while others idle. Distribute the load.

L3

Message Queues: Your queue has a million pending jobs. Fix the backlog.

A traffic spike is producing work faster than you can process it. Keep the system responsive.

L2

Database Replication: Your database just died. Good luck.

Your single database is a single point of failure and can't keep up with reads. Make it survivable.

L3

Database Sharding: One database can't hold it all. Split it.

Your dataset and write volume outgrew a single database. Partition the data so it scales.

L3

CAP & Consistency: Pick two. Actually, pick your poison.

A network partition splits your cluster. Do you keep serving (risking stale data) or stop (staying consistent)?

L3

Key-Value Store: Two replicas just died. Keep reads and writes flowing.

You're building a Dynamo-style key-value store: data partitioned on a hash ring, N replicas per key. Machines fail all the time. Tune quorums, conflict resolution and repair so the store stays available without silently losing data.

Full real-system missions (URL shortener, chat, news feed, video, payments) live in Challenges.