AI System Design
← Learn
Level 2beginner

Rate Limiter

Your API is under attack. Good luck.

Depth:
1

Mission

A botnet just pointed 5,000 req/s at your API on top of your real users. Everything is falling over. Protect the system without locking out real customers.
2

Interactive Simulation

Before we explain anything — play. Push it until it breaks, then fix it.

🚨 API overwhelmed — 5.5K/s hitting a 3.0K/s API. Add a rate limiter.
Real users rejected
45.5%
the true cost of a limiter
Requests
5.5K/s
Latency
7.00s
p95 21.00s
Error rate
45.5%
CPU
100%
Est. cost
$600/mo
illustrative
Accepted
3.0K/s
Rejected
2.5K/s
Latency (ms)
Error rate (%)

Traffic

Rate Limiter

Break it

System Score38
3

What just happened?

The flood of bot traffic saturated your API, so real users' requests failed too. A single global limit protected the API but couldn't tell bots from users — it rejected both. Keying the limit per client (API key or IP) throttled the handful of loud bots while quiet real users sailed through.

4

The concept

A rate limiter caps how many requests a client can make in a time window. It protects backends from abuse and accidental overload. Three classic algorithms: Fixed Window (simple, but bursty at window edges), Sliding Window (smoother, more accurate), and Token Bucket (allows controlled bursts while enforcing an average rate).

5

Trade-offs

Nothing is free. Here's what this solution costs you.

Global vs per-key
A global bucket is cheap but punishes everyone equally; per-key buckets need shared state per client.
Too tight
Aggressive limits reject legitimate users and hurt experience.
Too loose
Generous limits let abuse through and defeat the point.
State cost
Accurate limiting needs shared counters — extra infra and latency.
6

In the real world

Conceptually similar to API gateway rate limiting backed by a Redis-like store.

7

Mini quiz

Question 1 of 40 correct

Which algorithm best allows short bursts while keeping a steady average rate?

8

Interview me

The app becomes your interviewer. One question, in your own words.

9

Boss challenge

Stop API abuse

10,000 req/s of bot traffic is hitting a 3,000 req/s API alongside real users.

Goal: Keep the API under capacity (3K req/s) while rejecting under 5% of legitimate traffic.

Use the simulator above with no hints. These checks update live as you play.

10

Interview question

“Compare token bucket, fixed window, and sliding window. Where would you enforce limits, what key would you use, and how do you handle distributed counters?”

Next: Caching