Problem & requirements
We need IDs that are unique across the whole system, numeric, fit in 64 bits, are ordered by creation time (roughly, so newer items sort after older ones), and can be generated at more than 10,000 per second. Ordering matters because many queries, like a timeline or a pagination cursor, want to sort by recency without an extra timestamp index.
A single relational database with AUTO_INCREMENT satisfies everything except scale and availability: it becomes a bottleneck and a single point of failure, and it cannot be shared across data centers without a network hop on every insert. The challenge is to generate IDs independently on many machines while preserving uniqueness and approximate order.
Back-of-the-envelope
A 41-bit millisecond timestamp covers 2^41 ms ≈ 2.2 x 10^12 ms. Dividing by 1,000 x 86,400 x 365 (about 3.15 x 10^10 ms per year) gives roughly 69.7 years before overflow, measured from a custom epoch. That is why a custom epoch close to launch day matters; counting from 1970 would waste decades of range.
A 12-bit sequence allows 2^12 = 4,096 IDs per millisecond per machine, or about 4.1 million per second per machine, far above the 10,000-per-second requirement. Ten bits of machine identity (5 for datacenter, 5 for machine) allow 2^5 = 32 datacenters x 32 machines = 1,024 generators. Total: 1 + 41 + 5 + 5 + 12 = 64 bits.
Options considered
Multi-master replication uses auto-increment on each of k databases with a step of k, so server 1 issues 1, k+1, 2k+1 and so on. It scales with the number of databases, but IDs do not increase with time across servers and adding or removing a server breaks the step scheme. UUIDs are 128-bit values any machine can generate with negligible collision risk and no coordination, but they are too long for the 64-bit requirement, are non-numeric strings in many systems, and random versions are not time-ordered, which also hurts B-tree index locality.
A ticket server is a single database that hands out auto-increment values to everyone, as Flickr popularized. It is simple and numeric and works for medium scale, but it is a single point of failure, and running several ticket servers reintroduces synchronization problems. Each option fails at least one requirement, which motivates a layout that embeds time and machine identity directly into the ID.
The Snowflake bit layout
Snowflake divides a 64-bit integer into sections, most significant first. The sign bit (1 bit) is always 0 so the ID stays positive in signed integer types and is reserved for future use. The timestamp (41 bits) is milliseconds since a custom epoch. The datacenter ID (5 bits) and machine ID (5 bits) identify the generator. The sequence number (12 bits) counts IDs generated within the current millisecond on that machine and resets to 0 each millisecond.
Because the timestamp occupies the high bits, sorting IDs numerically sorts them by time, and you can recover the creation time from any ID: shift right by 22 bits and add the epoch. Datacenter and machine IDs are assigned at deployment and fixed while running, since changing them carelessly could create duplicates.
- 1 bit: sign, always 0
- 41 bits: milliseconds since custom epoch
- 5 bits: datacenter ID (32 values)
- 5 bits: machine ID (32 per datacenter)
- 12 bits: per-millisecond sequence (4,096 values)
id >> 22 recovers the milliseconds since the epoch. The low 22 bits make IDs unique among generators and within one millisecond.Generating an ID step by step
On each request the generator reads the current time in milliseconds and subtracts the epoch. If the time equals the last-seen millisecond, it increments the sequence; if the sequence would overflow past 4,095, it spins until the next millisecond. If the time is newer, it resets the sequence to 0. It then assembles the ID with bit shifts: (ts << 22) | (dc << 17) | (machine << 12) | seq.
All of this happens in local memory with a short critical section, so generation takes well under a microsecond and requires no network coordination. That is the key architectural win: each application server can embed the generator, removing the central bottleneck entirely while still producing globally unique values.
Deep dive: clocks and tuning
The design assumes each machine's clock moves forward. In reality clocks drift, and NTP (Network Time Protocol) corrections can step a clock backwards. If the time goes backwards, the generator could reissue IDs it has already produced. A careful implementation detects this (current time less than last timestamp) and either waits until the clock catches up or refuses to issue IDs and raises an alert. Across machines, small skew means IDs are only approximately time-ordered, which is acceptable for most feeds.
Section lengths are tunable to the workload. A system with low concurrency but a longer lifespan could shrink the sequence and give more bits to the timestamp. A system with few datacenters but many machines per site could shift bits from datacenter to machine ID. Do the arithmetic for each section explicitly rather than copying the default split.
Trade-offs, high availability and wrap-up
Since generation is local, an ID generator is only as available as the service it is embedded in, and there is no shared component to fail. The operational risks move elsewhere: assigning machine IDs uniquely (often via a coordination service like ZooKeeper or from deployment configuration) and keeping clocks healthy. Run the generator in every instance and monitor for clock regressions and sequence exhaustion.
Compared with alternatives, Snowflake trades the perfect global ordering of a single counter for scalability and availability, and trades the zero-configuration of UUIDs for compactness and sortability. Newer formats such as ULID and UUIDv7 adopt the same timestamp-first idea in 128 bits, which is worth mentioning when 64 bits is not a hard constraint.
Key numbers
Key terms
- Snowflake ID
- A 64-bit ID composed of a timestamp, datacenter ID, machine ID and per-millisecond sequence.
- Custom epoch
- A chosen starting instant from which timestamps are counted to maximize the usable range.
- Sequence number
- A counter that distinguishes IDs generated on the same machine within the same millisecond.
- UUID
- A 128-bit identifier generated without coordination, typically random and not time-ordered.
- Ticket server
- A centralized database whose auto-increment counter issues IDs to all clients.
- NTP
- Network Time Protocol, used to synchronize machine clocks, which can step clocks backward.
- Clock skew
- The difference between clocks on different machines, which makes cross-machine ordering approximate.
Common mistakes
- Proposing UUIDs when the requirement says numeric, 64-bit and time-sortable.
- Ignoring what happens when the clock moves backwards after an NTP correction.
- Using the Unix epoch and losing most of the 41-bit timestamp range.
- Forgetting the sign bit, producing negative IDs in signed 64-bit languages.
- Not explaining how machine IDs are assigned uniquely across deployments.
- Claiming IDs are strictly globally ordered across machines.
Further study
- Twitter Snowflake (2010 announcement and open-source release)
- Flickr ticket servers engineering blog post (2010)
- RFC 9562: UUIDs including UUIDv7
- Instagram engineering: Sharding and IDs at Instagram
- ULID specification