AI System Design
← Learn
Level 1beginner

Scaling

One server isn't enough. What would you do?

Depth:
1

Mission

Your app just got popular. A single server handled the early users fine — now traffic is climbing past what one machine can serve. Keep it healthy.
2

Interactive Simulation

Before we explain anything — play. Push it until it breaks, then fix it.

Requests
500/s
Latency
114ms
p95 343ms
Error rate
0.0%
CPU
50%
DB load
13%
Est. cost
$560/mo
illustrative
Accepted
500/s
Latency (ms)
Error rate (%)

Controls

Experiment

Try breaking it — crank traffic past capacity.

System Score91
3

What just happened?

As you pushed traffic past the server's capacity, CPU pinned at 100%, latency shot up, and requests started failing. A single server has a hard ceiling.

4

The concept

Scaling means adding capacity to handle more load. Vertical scaling makes one machine bigger (more CPU/RAM) — simple, but there's a physical ceiling and it's a single point of failure. Horizontal scaling adds more machines behind a load balancer, so work is spread out and you can grow almost without limit.

5

Trade-offs

Nothing is free. Here's what this solution costs you.

Cost
More servers = more money. Horizontal scaling trades dollars for headroom.
Complexity
Load balancers, health checks, and stateless design add moving parts.
Diminishing returns
Adding API servers stops helping once the database saturates.
6

In the real world

Conceptually similar to running an app across an autoscaling group behind an L7 load balancer in any major cloud.

7

Mini quiz

Question 1 of 30 correct

You 10× your traffic and add more API servers, but latency is still terrible. What's the most likely bottleneck?

8

Interview me

The app becomes your interviewer. One question, in your own words.

9

Boss challenge

Survive 10K requests/sec

Traffic is ramping to 10,000 req/s. Your starting server does 1,000 req/s.

Goal: Keep error rate under 2% and latency under 200ms using scaling and caching.

Use the simulator above with no hints. These checks update live as you play.

10

Interview question

“Explain vertical vs horizontal scaling, why stateless services matter for horizontal scaling, and how you'd identify the next bottleneck after adding servers.”

Next: Back-of-the-Envelope Estimation