AI System Design
← Learn
Level 2intermediate

Database Replication

Your database just died. Good luck.

Depth:
1

Mission

Your single database is a single point of failure and can't keep up with reads. Make it survivable.
2

Interactive Simulation

Before we explain anything — play. Push it until it breaks, then fix it.

The leader is drowning in reads. Route reads to replicas.
Replication lag
—
Stale reads
0.0%
Write latency
1.60s
Lost on failover
everything
acked writes
Requests
6.8K/s
Latency
1.60s
p95 4.81s
Error rate
58.5%
CPU
100%
DB load
100%
Est. cost
$660/mo
illustrative
Accepted
2.8K/s
Rejected
4.0K/s
Latency (ms)
Error rate (%)

Workload

Replication

Mode

Leader acks immediately. Fastest writes, but followers lag and a crash can lose recent writes.

Break it

System Score34
3

What just happened?

Read replicas absorbed read traffic and gave you a failover target, so a primary failure no longer meant total outage.

4

The concept

Replication keeps copies of your data on multiple machines. A common setup is leader/follower: writes go to the leader and replicate to followers, which serve reads. This scales reads and provides redundancy for failover.

5

Trade-offs

Nothing is free. Here's what this solution costs you.

Replication lag
Async followers can serve stale reads.
Write scaling
Replication scales reads, not writes — that needs sharding.
Durability vs latency
Sync replication never loses writes but every write waits on the slowest follower.
6

In the real world

Conceptually similar to managed Postgres/MySQL with read replicas and automated failover.

7

Mini quiz

Question 1 of 30 correct

Read replicas primarily scale…

8

Interview me

The app becomes your interviewer. One question, in your own words.

9

Boss challenge

The primary just died

9K reads/s and 1K writes/s. The primary crashes.

Goal: Keep errors under 1%, lose zero acknowledged writes, and keep stale reads under 5%.

Use the simulator above with no hints. These checks update live as you play.

10

Interview question

“Explain leader/follower, sync vs async replication, replication lag, and failover.”

Next: Database Sharding