Every machine in a cluster has its own clock, and no two agree. The difference is usually milliseconds, sometimes seconds, and occasionally much worse after a virtual machine pause, a bad time source or a manual change. Code that compares timestamps from different machines inherits that disagreement silently, and the resulting bugs, lost writes, expired locks that are still in use, and out-of-order events, rarely show up in tests.

This article is a map of time in distributed systems. It starts with physical clocks, how they drift and how synchronization corrects them, then separates the three questions programs ask of time: how long something took, what time it is, and which of two events came first. Each question needs a different clock. Detailed treatments of hybrid logical clocks, vector clocks, NTP and TrueTime exist on this site and are linked here; this page shows where each fits and how to choose.

Advertisement

Physical clocks drift

A computer keeps time by counting the oscillations of a quartz crystal. The crystal's frequency is not exactly its nominal value and changes with temperature and age, so the clock gains or loses time at a rate measured in parts per million. A clock that is 10 ppm fast gains 10 microseconds every second, which is about 0.86 seconds per day. Left alone, two servers drift apart without bound.

Synchronization protocols correct this. NTP estimates the offset from reference servers by timing request and response messages and assuming the network delay is symmetric; any asymmetry becomes error. Over the public internet the result is typically accurate to milliseconds, on a good local network to much less, and PTP with hardware timestamping in network cards can reach the sub-microsecond range. The daemon applies corrections in two ways. It slews the clock, running it slightly faster or slower until the offset is gone, or, when the offset is large, it steps the clock, jumping it to the right value, possibly backwards.

Leap seconds add one more discontinuity. Some systems repeat a second, others smear it by running the clock slightly slow over many hours, as large cloud providers do. Mixing smeared and non-smeared time sources in one fleet creates offsets of up to half a second during the smear. Leap seconds are due to be abolished by 2035, but existing systems still have to handle them until then.

Three questions, three kinds of clock: never answer one with the clock built for anotherQuartz oscillatordrifts with temperatureKernel clocksREALTIME, MONOTONICNTP / PTP daemonslews or steps timeHow long did it take?monotonic durationWhat time is it?wall clock, +/- offsetWhat happened first?logical clocktimeouts, leaseslatency metricslogs, TTLs, UIhuman-facing timeLamport / HLCtotal order, causalVector clocksdetect concurrencyTrueTime-stylebounded uncertaintyticksadjustsPhysical time answers durations and human-facing times; ordering across machines needs logical or bounded-uncertainty clocks.
The oscillator ticks, the kernel exposes wall and monotonic clocks, and the sync daemon adjusts them. Each application question maps to a different clock.

Wall-clock time versus monotonic time

Operating systems expose several clocks. On Linux, CLOCK_REALTIME is wall-clock time: seconds since the Unix epoch, affected by steps, manual changes and leap-second handling. CLOCK_MONOTONIC counts from an arbitrary starting point, never goes backwards and is not stepped, though NTP may still slew its rate. CLOCK_MONOTONIC_RAW is not slewed at all, and CLOCK_BOOTTIME is monotonic but keeps counting during suspend.

The rule is short. Use the monotonic clock for durations: timeouts, retries, rate limits, latency measurements and lease expiry. Use wall-clock time for values humans read or that must mean the same thing on another machine, such as log timestamps and absolute expiry dates, and accept that it can be wrong by the current offset. A monotonic reading has no meaning on another machine, so never send it over the network.

import time

# Wrong: wall-clock time can jump backwards or forwards when NTP steps the clock.
start = time.time()
do_request()
elapsed = time.time() - start          # can be negative, or include a 1 s leap

# Right: the monotonic clock never goes backwards (it has no meaning across machines).
start = time.monotonic()
do_request()
elapsed = time.monotonic() - start

# Java: System.nanoTime() for durations, Instant.now() for timestamps.
# Go:   time.Now() carries a monotonic reading since Go 1.9, so time.Since(start)
#       is safe; t.Round(0) strips it, and so does serializing the time.
Advertisement

Why timestamps cannot order events

It is tempting to order events across machines by wall-clock timestamp. Many databases do exactly that for conflict resolution: last write wins keeps the value with the highest timestamp. Consider two application servers, A and B, where B's clock is 300 ms behind A's.

Real timeEventStamped withStored winner
t = 0 msUser sets email to old@x.com through server A10:00:00.500 (A's clock)old@x.com
t = 100 msUser corrects email to new@x.com through server B10:00:00.300 (B's clock)old@x.com: the correction has the lower stamp
laterRead from any replicaold@x.com: the newer write is silently discarded

Nothing failed, no error was logged, and the user's correction disappeared. Any two writes closer together than the clock offset can be ordered the wrong way. Smaller offsets shrink the window but never close it, and a single step or VM pause can open it wide. If you must use last-write-wins, stamp writes from one place per key, for example the leader for that key, or use a clock that guarantees a later event gets a larger stamp when the two are causally related.

Lamport clocks: ordering without physical time

Lamport's 1978 insight was that distributed systems usually need causal order, not real-time order. Event a happened before b if they are on the same process and a came first, if a is a send and b the matching receive, or by transitivity. A Lamport clock is a counter per process that guarantees: if a happened before b, then the timestamp of a is smaller than that of b.

import threading

class LamportClock:
    def __init__(self, node_id):
        self.node_id = node_id
        self.counter = 0
        self.lock = threading.Lock()

    def tick(self):                      # local event or send
        with self.lock:
            self.counter += 1
            return (self.counter, self.node_id)

    def on_receive(self, remote_counter):
        with self.lock:
            self.counter = max(self.counter, remote_counter) + 1
            return (self.counter, self.node_id)

# Tuples compare counter first, node id second: a total order that respects causality.
# If event a happened before b, then stamp(a) < stamp(b). The converse does NOT hold.

Rerun the email example with Lamport clocks. Server A's write gets counter 7. The user's client receives the response carrying 7, and sends the correction to B with that value; B sets its counter to at least 8 before stamping. The correction now wins because it causally follows the first write, whatever the physical clocks say.

The guarantee is one-way. A smaller Lamport timestamp does not mean the event happened first; two concurrent events get arbitrary relative order, broken by node ID. That is fine for a deterministic total order, as in replicated state machines or mutual exclusion, but it cannot tell you whether two writes conflicted. Lamport stamps also have no relation to wall time, so you cannot ask for 'all events before 10:00'.

The rest of the toolbox

ClockGuaranteeSizeUse it for
Monotonic clocknever goes backwards on one machineone numberdurations, timeouts, lease expiry
Wall clock (synchronized)close to real time, within the sync offsetone numberlogs, TTLs, human-facing times
Lamport clockhappened-before implies smaller stampcounter + node idtotal order, replicated logs
Hybrid logical clockLamport guarantee, stays close to wall time64-bit value typicallyMVCC timestamps, snapshot reads by time
Vector clocksmaller vector if and only if happened-beforeone entry per nodedetecting concurrent writes, causal delivery
Bounded-uncertainty timeinterval guaranteed to contain real timetwo numbersexternal consistency with commit wait

Hybrid logical clocks combine the Lamport rule with physical time, so stamps are causally consistent and still usable for time-based queries; several distributed SQL databases use them for transaction timestamps. Vector clocks give the full converse: comparing two vectors tells you whether one event happened before the other or whether they were concurrent, at the cost of one entry per participant. Google's TrueTime returns an interval rather than an instant; Spanner waits out the uncertainty before making a commit visible, so timestamp order matches real-time order. Its cost is latency proportional to the uncertainty, which is why it depends on GPS and atomic clock references.

Leases: when you must trust physical time

Some decisions cannot avoid physical time. A lease grants exclusive rights, such as leadership or a lock, for a duration, and its holder must stop acting before the grantor considers it expired. Safety depends on the clocks of both sides running at nearly the same rate, not on their agreeing on the time of day, which is why leases use monotonic durations rather than absolute timestamps.

import time

MAX_DRIFT = 200e-6      # assumed worst-case rate error: 200 microseconds per second
SAFETY = 0.050          # extra margin for scheduling pauses, seconds

class Lease:
    def acquire(self, grant_rpc, duration_s):
        sent = time.monotonic()          # measured BEFORE the request leaves
        token = grant_rpc(duration_s)    # server starts its clock when it grants
        # Count from `sent`, not from the reply: the grant may have been issued
        # at any point between send and receive.
        # Holder and grantor may drift in opposite directions, hence 2 * MAX_DRIFT.
        self.expires = sent + duration_s * (1 - 2 * MAX_DRIFT) - SAFETY
        self.token = token               # fencing token, sent with every write
        return token

    def valid(self):
        return time.monotonic() < self.expires

Three details carry the safety argument. The holder starts its countdown when it sends the request, the earliest moment the grant could have been issued. It shortens the duration by twice an assumed maximum drift rate, because both clocks can err in opposite directions. The 200 microseconds per second used here is the conservative bound the Spanner paper assumes; use your own measured bound. And it still sends a fencing token with every write, because a process can pause for garbage collection or VM migration right after checking valid() and act after expiry anyway. The storage system rejects any write whose token is older than the newest it has seen.

Operating clocks in production

  • Monitor offset, not just daemon health. Export the offset reported by chronyc tracking or your NTP daemon and alert well below the tolerance of any system that relies on timestamps.
  • Use several sources. One bad upstream server can drag a fleet. Configure several sources, prefer the provider's local time service in the cloud, and do not mix smeared and non-smeared sources.
  • Know your tolerance. Some databases configure a maximum clock offset and remove a node that exceeds it; read that setting and alert before reaching it.
  • Watch for pauses. VM live migration, suspended containers and long garbage-collection pauses look like time jumping forwards for the process. Lease holders must re-check validity after every blocking call.
  • Stamp logs in UTC with offsets recorded. When debugging across machines, order events by trace and causal IDs first and timestamps second.

Failure modes and trade-offs

The common failures follow from the three questions. Using wall-clock time for durations produces negative latencies and timeouts that fire early or never after a step. Using wall-clock time for ordering loses writes under last-write-wins. Using logical clocks for human-facing time gives meaningless values. Trusting a lease without drift margin and fencing produces two leaders.

The trade-off is cost against guarantee. Wall clocks are free and approximate. Lamport and hybrid clocks add a few bytes per message and give causal order. Vector clocks detect concurrency but grow with the number of participants. Bounded-uncertainty clocks give real-time order but need hardware time references and add latency to every commit.

Related reading

The detailed pages: hybrid logical clocks, vector clocks for causal delivery and consistent cuts and causal consistency. For synchronization itself, NTP; for bounded uncertainty, Spanner TrueTime; and for leases with fencing tokens, distributed lock architecture.

What to do next

  1. Search your code for wall-clock reads used to compute durations, such as time.time() differences or System.currentTimeMillis() subtraction, and switch them to monotonic clocks.
  2. List every place that orders or resolves conflicts by timestamp, and write down what happens when two writes land within the clock offset.
  3. Export clock offset from every host and alert on it, with thresholds set below the tolerance of your databases.
  4. Audit leases and locks: countdown from send time, a drift margin, and fencing tokens checked by the storage layer.
  5. Where you need causal order, add a Lamport or hybrid logical clock to messages rather than trusting physical time.
  6. Where you need to detect concurrent updates, evaluate vector clocks or version vectors instead of last-write-wins.
  7. Run a game day that steps one host's clock by a second, and watch what breaks.
Key takeaway: Clocks answer three different questions. Use the monotonic clock for durations, timeouts and leases; use synchronized wall-clock time for human-facing times and accept that it is off by the current offset; use logical clocks, Lamport, hybrid or vector, for ordering events across machines. Timestamps from different machines cannot order writes closer together than their offset, so last-write-wins with wall clocks silently loses data. Leases need a drift margin and fencing tokens, and every fleet needs offset monitoring.