Mission
Interactive Simulation
Before we explain anything — play. Push it until it breaks, then fix it.
Controls
Experiment
Try breaking it — crank traffic past capacity.
What just happened?
As you pushed traffic past the server's capacity, CPU pinned at 100%, latency shot up, and requests started failing. A single server has a hard ceiling.
The concept
Scaling means adding capacity to handle more load. Vertical scaling makes one machine bigger (more CPU/RAM) — simple, but there's a physical ceiling and it's a single point of failure. Horizontal scaling adds more machines behind a load balancer, so work is spread out and you can grow almost without limit.
Trade-offs
Nothing is free. Here's what this solution costs you.
In the real world
Conceptually similar to running an app across an autoscaling group behind an L7 load balancer in any major cloud.
Mini quiz
You 10× your traffic and add more API servers, but latency is still terrible. What's the most likely bottleneck?
Interview me
The app becomes your interviewer. One question, in your own words.
Boss challenge
Survive 10K requests/sec
Traffic is ramping to 10,000 req/s. Your starting server does 1,000 req/s.
Goal: Keep error rate under 2% and latency under 200ms using scaling and caching.
Use the simulator above with no hints. These checks update live as you play.
Interview question
“Explain vertical vs horizontal scaling, why stateless services matter for horizontal scaling, and how you'd identify the next bottleneck after adding servers.”