Chronyc: Auditing Kernel Clock Drift, Triaging NTP Stratum Synchronization, and Enforcing Precision Time Tracking in Production
The invisible culprit behind this multi-million-pound outage is a silent drift of 350 milliseconds across the server fleet. In everyday human life, a third of a second is less than the blink of an eye. In high-speed distributed computing, however, physical time is a strict law: when machines disagree on the exact sequence of events, distributed databases assume disaster has struck and halt to prevent data corruption.
To audit, discipline, and repair this chaotic drift at the Linux kernel level, systems engineers rely on chronyc, the command-line control tool for the Chrony time synchronisation suite.
When you are troubleshooting a live crisis, you do not have time to wade through pages of theory. You need to know immediately whether your machine is locked onto a reliable global clock or floating dangerously adrift.
The single most valuable command to run right now is:
chronyc -n tracking
Reference ID : CB007106 (195.0.113.6)
Stratum : 2
Ref time (UTC) : Tue Aug 18 06:01:40 2026
System time : 0.000012411 seconds slow of NTP time
Last offset : -0.000003120 seconds
RMS offset : 0.000008451 seconds
Frequency : 14.218 ppm slow
Residual freq : -0.002 ppm
Skew : 0.045 ppm
Root delay : 0.012451820 seconds
Root dispersion : 0.001045120 seconds
Update interval : 64.5 seconds
Leap status : Normal
In a dozen lines of clear output, this command confirms that the host is synchronised to a primary upstream server (195.0.113.6) and is lagging nominal UTC by merely 12.4 microseconds (0.000012411 seconds slow)βwell inside the sub-millisecond envelope required for distributed databases to operate safely.
What It Does in Plain English
Every computer contains a tiny quartz crystal oscillator that vibrates at a specific frequency to tick off fractions of a second, much like a microscopic digital wristwatch. However, these physical crystals are inherently imperfect: changes in ambient server room temperature, power fluctuations, and component ageing cause them to tick slightly too fast or too slow. In virtualised cloud environments, where dozens of guest machines share a single physical CPU, virtual clocks can suddenly stall for hundreds of milliseconds during hypervisor migrations.
chronyc is the administrative interface for chronyd, the background daemon engineered to keep Linux clocks accurate across unpredictable networks and virtualised servers. While standard clock tools merely tell you what time it is right now, chronyc inspects the internal mathematical models of the operating system to reveal how your hardware is drifting relative to global atomic standards. It allows engineers to inspect upstream atomic clocks, calculate network packet delay, measure jitter, and smoothly nudge the Linux kernel clock without jarring running software.
Core Flags & Quick Start
The chronyc utility interacts with the background daemon through a local UNIX domain socket (typically at /run/chrony/chronyd.sock), providing direct telemetry and live configuration changes without requiring daemon restarts.
Essential Flags and Commands
| Flag / Command | Purpose | Production Context |
|---|---|---|
-n |
Suppresses reverse DNS lookups | Displays raw IP addresses instantly, preventing shell lockups during DNS outages. |
-c |
Outputs comma-separated values (CSV) | Ideal for piping metrics into monitoring agents like Prometheus Node Exporter or Datadog. |
-m |
Multi-command mode | Runs multiple commands in a single execution (e.g., chronyc -m tracking sources sourcestats). |
tracking |
System clock overview | Shows stratum level, phase offset, root delay, dispersion, and crystal frequency drift. |
sources -v |
Upstream server breakdown | Displays health, poll intervals, and reachability status for all configured time sources. |
sourcestats -v |
Statistical regression metrics | Evaluates error bounds, sample stability, and Allan deviation for each upstream source. |
ntpdata |
Deep packet-level telemetry | Diagnoses network jitter, transit delays, and timestamping layers for a specific peer. |
makestep |
Immediate phase step | Manually forces the clock to jump directly to true time, bypassing gradual slewing. |
Theoretical & Architectural Foundations
To wield chronyc effectively, it helps to understand how the Linux kernel manages the bridge between physical quartz crystals and abstract digital time.
Monotonic vs Real-Time Clocks
The Linux kernel maintains distinct representations of time through POSIX system calls, documented thoroughly in the Linux Kernel Timekeeping Documentation:
CLOCK_REALTIME: The conventional wall-clock time (UTC) that human beings read. Because it represents real-world date and time, it can be adjusted, stepped backwards, or modified by leap seconds. Applications that measure elapsed durations usingCLOCK_REALTIMErisk calculating negative intervals if an administrator or service steps the clock backwards.CLOCK_MONOTONICandCLOCK_MONOTONIC_RAW: Clocks that represent elapsed time since an arbitrary starting point (usually system boot).CLOCK_MONOTONICis guaranteed never to jump backwards, although its tick frequency is gently disciplined.CLOCK_MONOTONIC_RAWprovides direct access to un-slewed hardware ticks from the CPU Time Stamp Counter (TSC).
Mathematical Slewing versus Stepping
When a server's clock drifts away from true atomic time, the operating system can correct the error using one of two strategies:
- Stepping: The clock is instantaneously jumped to the correct time. While fast, this creates a temporal discontinuity. If the clock jumps backwards, logs record events in reverse order, database write leases expire unexpectedly, and security tokens fail validation.
- Slewing: The daemon subtly speeds up or slows down the kernel tick rate (typically by up to 83,333 parts per million, or a 12:1 ratio) until the difference shrinks to zero. Slewing preserves strict monotonicity, guaranteeing that time flows uniformly forward without disrupting running software.
Why Chrony Outperforms Legacy NTPD in Virtualisation
The classic ntpd software was designed in an era of dedicated physical servers and stable networks, as defined in RFC 5905: Network Time Protocol Version 4. When run inside modern public cloud virtual machines (such as AWS EC2, Google Cloud Engine, or Azure VMs), legacy daemons often struggle:
- Hypervisor Steal Time: When a physical host is busy or migrates a virtual machine across physical hardware, the guest operating system experiences momentary pauses lasting anywhere from milliseconds to seconds. Legacy daemons often mistake this pause for a broken crystal oscillator and panic or detach from time sources.
- Dynamic Convergence:
chronyduses linear regression over historical sample points to continuously calculate frequency drift. When resuming from a hypervisor pause, Chrony calculates the accumulated drift within a handful of network packets and rapidly slews the clock back to accuracy in minutes rather than hours.
The Stratum Hierarchy and Hardware PTP
Network Time Protocol relies on a tiered pyramid of trust: * Stratum 0: Atomic clocks, Rubidium oscillators, and GNSS/GPS satellite constellations. * Stratum 1: Dedicated servers directly attached to Stratum 0 devices via hardware interfaces. * Stratum 2: Networked servers synchronising over local or wide-area networks with Stratum 1 systems.
To eliminate the network jitter inherent in standard network cards, the IEEE 1588 Precision Time Protocol Specification uses hardware timestamping directly at the physical network interface layer (PHY). Chrony natively bridges standard NTP over the internet and sub-microsecond PTP hardware clocks (/dev/ptpX) inside modern data centres.
5 Genuine Real-World Production Use-Cases
Use-Case 1: Auditing Upstream NTP Sources & Isolating Falsetickers
Scenario
A financial transaction routing node begins experiencing sporadic timestamp rejections. As the systems engineer on duty, you must inspect the pool of upstream time servers, verify stratum integrity, and identify any erratic servers broadcasting faulty time measurements.
Command Execution
chronyc -n sources -v
Realistic Terminal Output
.-- Source mode '^' = server, '=' = peer, '#' = local clock.
/ .- Source state '*' = current best, '+' = combined, '-' = not combined,
| / 'x' = may be in error, '~' = too variable, '?' = unusable.
|| .- xxxx [ * upstairs ]
|| / [ - downstairs ]
|| / [ = equal ]
|| / [ ? unknown ]
|| /
MS Name/IP address Stratum Poll Reach LastRx Last sample
===============================================================================
^* 195.0.113.6 1 6 377 22 -1204ns[ -1510ns] +/- 1240us
^+ 198.51.100.14 2 6 377 21 +3120ns[ +2814ns] +/- 4110us
^- 203.0.113.88 2 6 377 24 +12.4ms[ +12.4ms] +/- 18ms
^x 192.0.2.199 3 6 377 23 +245.1ms[+245.1ms] +/- 120ms
To check statistical stability and sample weighting for each server, execute:
chronyc -n sourcestats -v
.- Number of sample points in measurement.
/ .- Number of sample points bearing weight.
/ / .- Frequency of sample collection (hz).
/ / / .- Estimated frequency error (ppm).
/ / / / .- Estimated error bounds.
/ / / / / .- Standard deviation of offset.
/ / / / / /
Name/IP Address NP NR Tr FReq-ppm Error-ppm Std-Dev
===============================================================================
195.0.113.6 32 18 0 -0.002 0.012 1.2us
198.51.100.14 32 20 0 +0.014 0.035 3.4us
203.0.113.88 32 16 0 +1.120 0.450 850.0us
192.0.2.199 8 4 0 +48.100 12.300 45.2ms
Line-by-Line Explanation
^* 195.0.113.6: The^indicates a standard network server; the*signifies that this node is currently selected as the active primary clock source.^+ 198.51.100.14: The+identifies an acceptable backup candidate whose measurements are blended into the synchronisation algorithm.^- 203.0.113.88: The-indicates a server discarded from the main calculation because its offset exceeds precision thresholds.^x 192.0.2.199: Thexexplicitly flags a falsetickerβa broken or drifting server whose time contradicts the majority consensus of the pool.Reach: 377: An octal representation of an 8-bit register (11111111in binary), confirming that the last eight consecutive poll attempts succeeded without a single lost packet.Std-Devinsourcestats: While195.0.113.6shows exceptional stability with a standard deviation of just $1.2\,\mu\text{s}$, the rogue server192.0.2.199exhibits an unacceptably erratic variance of $45.2\,\text{ms}$.
What the Admin Does Next
- Immediately remove the falseticker (
192.0.2.199) from your configuration file in/etc/chrony/chrony.conf(or/etc/chrony.confdepending on your Linux distribution, as detailed in the ArchWiki Chrony Reference). - Replace the high-jitter host
203.0.113.88with a closer Stratum 1 or Stratum 2 server to maintain a resilient pool of at least four healthy sources.
Use-Case 2: Diagnosing Distributed Database Clock Skew & Split-Brain Risks
Scenario
A globally distributed CockroachDB cluster reports rising transaction commit latency. CockroachDB relies on Hybrid Logical Clocks (HLC) and deliberately aborts database transactions whenever the physical clock difference across cluster nodes exceeds 500 milliseconds. You need to calculate the node's maximum possible time uncertainty.
Command Execution
chronyc tracking
Realistic Terminal Output
Reference ID : 0A640001 (10.100.0.1)
Stratum : 2
Ref time (UTC) : Tue Aug 18 06:03:12 2026
System time : 0.000045102 seconds slow of NTP time
Last offset : -0.000008120 seconds
RMS offset : 0.000021400 seconds
Frequency : 22.415 ppm fast
Residual freq : +0.004 ppm
Skew : 0.082 ppm
Root delay : 0.004210000 seconds
Root dispersion : 0.000840100 seconds
Update interval : 128.4 seconds
Leap status : Normal
Line-by-Line Explanation
System time : 0.000045102 seconds slow: The local operating system clock is lagging nominal NTP time by $45.1\,\mu\text{s}$.RMS offset : 0.000021400 seconds: The long-term Root-Mean-Square offset is $21.4\,\mu\text{s}$, confirming strong mathematical consistency.Frequency : 22.415 ppm fast: The physical crystal on this motherboard naturally runs 22.415 parts per million faster than normal; Chrony continuously commands the Linux kernel to slew the tick rate to compensate.Root delay($4.21\,\text{ms}$) andRoot dispersion($0.84\,\text{ms}$): Total round-trip network transit time to the primary atomic clock and the cumulative error bound across the network path.
To calculate maximum possible time error on this node: $$\text{Max Uncertainty} = |\text{System Time Offset}| + \text{Root Dispersion} + \frac{\text{Root Delay}}{2}$$ $$\text{Max Uncertainty} = 0.0451\,\text{ms} + 0.8401\,\text{ms} + \frac{4.2100\,\text{ms}}{2} = 2.9902\,\text{ms}$$
What the Admin Does Next
Because the calculated worst-case uncertainty ($\approx 2.99\,\text{ms}$) is well below the database's 500 ms safety threshold, you can confidently rule out local clock drift and refocus your investigation on Raft consensus queues or database lock contention.
Use-Case 3: Enforcing Safe Clock Slewing in Regulated Financial Environments
Scenario
A server cold-boots after routine hardware maintenance with an onboard hardware clock offset of $+3.8\,\text{seconds}$. The server hosts an audited financial ledger where stepping the clock backwards would violate ISO-27001 compliance by creating duplicate or non-sequential timestamps. You must confirm that the system safely slews rather than jumps.
Command Execution
Check whether a manual step can be forced:
chronyc makestep
200 OK
Verify that the system has settled and is tracking continuously:
chronyc -n tracking
Realistic Terminal Output
Reference ID : 0A640001 (10.100.0.1)
Stratum : 2
Ref time (UTC) : Tue Aug 18 06:05:01 2026
System time : 0.000001012 seconds fast of NTP time
Last offset : +0.000000412 seconds
RMS offset : 0.000002100 seconds
Frequency : 18.110 ppm slow
Residual freq : +0.001 ppm
Skew : 0.012 ppm
Root delay : 0.002104000 seconds
Root dispersion : 0.000412000 seconds
Update interval : 64.1 seconds
Leap status : Normal
Line-by-Line Explanation
makestep: If issued manually,chronydsteps the clock immediately only if permitted by the configuration rules.System time : 0.000001012 seconds fast: Confirms that following adjustment, the clock is running within $1.012\,\mu\text{s}$ of atomic time.Leap status : Normal: Indicates that no unhandled leap seconds or clock step operations are pending.
Directive in chrony.conf |
Cold Boot Behaviour | Runtime Behaviour | Operational Safety |
|---|---|---|---|
makestep 1.0 3 |
Steps if offset > 1.0s (first 3 updates only) | Strictly slews clock gradually | Safe: Preserves timestamp monotonicity during live transactions. |
makestep 0 -1 |
Steps on every update | Steps on every update | Dangerous: Causes time to jump backwards or forwards unpredictably. |
What the Admin Does Next
- Verify that
/etc/chrony/chrony.confcontainsmakestep 1.0 3rather than unrestricted step policies. - In production deployment pipelines, run
chronyc makestepduring the pre-boot orchestration scripts before transactional database daemons are launched.
Use-Case 4: Triaging Network Asymmetry and Packet Loss on UDP Port 123
Scenario
An edge node deployed across a software-defined WAN (SD-WAN) exhibits erratic clock drift. You suspect that network routing asymmetry between inbound and outbound paths, or an upstream firewall dropping UDP packets, is distorting time synchronisation.
Command Execution
chronyc -n ntpdata 195.0.113.6
Realistic Terminal Output
Remote address : 195.0.113.6 (195.0.113.6)
Remote port : 123
Local address : 198.51.100.22 (198.51.100.22)
Leap status : Normal
Stratum : 1
Ref time (UTC) : Tue Aug 18 05:59:12 2026
Offset : -0.000014210 seconds
Delay : 0.018420100 seconds
Dispersion : 0.000120100 seconds
Jitter : 0.000412000 seconds
Total tx count : 4120
Total rx count : 4118
Total valid rx : 4118
Total rx errors : 0
Interleaved : No
Authenticated : No
TX timestamping : Kernel
RX timestamping : Kernel
Total filter samples: 8
To check whether the local daemon is dropping packets under heavy load:
chronyc serverstats
NTP packets received : 142011
NTP packets dropped : 14
Command packets received : 320
Command packets dropped : 0
Client log records dropped : 0
NTP timestamp hits : 141997
NTP timestamp misses : 14
Line-by-Line Explanation
Total tx count : 4120vsTotal rx count : 4118: Shows that exactly two packets were lost across the WAN link out of more than 4,000 sent.Jitter : 0.000412000 seconds: Measured packet jitter is approximately $412\,\mu\text{s}$, well within normal operational limits ($< 1\,\text{ms}$).TX/RX timestamping : Kernel: Confirms that socket timestamping (SO_TIMESTAMPING) is handled inside the Linux kernel network stack, avoiding delay caused by user-space software scheduling.NTP packets dropped : 14: Confirms that upstream firewalls and local traffic shaping policies are not aggressively throttling time packets.
What the Admin Does Next
- If packet drops spike during peak hours, configure edge routers to apply Differentiated Services Code Point (DSCP) classification
CS6orEF(Expedited Forwarding) to UDP port 123. - If asymmetric BGP routing paths exist between data centres, apply a calibrated
offsetparameter inchrony.confto compensate for known transmission path differentials.
Use-Case 5: Synchronising Air-Gapped Networks with Local GPS and PTP Hardware
Scenario
An air-gapped industrial infrastructure has no connection to the public internet. High-precision synchronisation must be maintained using an onsite GPS receiver connected to a Pulse-Per-Second (PPS) hardware line (/dev/pps0) and an IEEE 1588 Precision Time Protocol master clock.
Command Execution
chronyc -n refclocks -v
Realistic Terminal Output
.- Number of sample points in measurement.
/ .- Number of sample points bearing weight.
/ / .- Frequency of sample collection (hz).
/ / / .- Estimated frequency error (ppm).
/ / / / .- Estimated error bounds.
/ / / / / .- Standard deviation of offset.
/ / / / / /
Refclock name NP NR Tr FReq-ppm Error-ppm Std-Dev
==============================================================================
PPS0 64 38 0 -0.001 0.002 24ns
NMEA(ttyS0) 16 8 0 +0.410 0.120 450us
PTP0 64 42 0 +0.000 0.001 12ns
To see how the daemon combines these hardware reference sources:
chronyc -n sources
MS Name/IP address Stratum Poll Reach LastRx Last sample
===============================================================================
#* PPS0 0 4 377 12 -14ns[ -18ns] +/- 32ns
#+ PTP0 0 4 377 11 +8ns[ +5ns] +/- 18ns
#- NMEA(ttyS0) 0 4 377 12 +124ms[ +124ms] +/- 250ms
Line-by-Line Explanation
#* PPS0: The#indicates a directly connected hardware clock; the*confirms thatPPS0is the selected master source, maintaining time within an incredible 14 nanoseconds (-14ns).#+ PTP0: The IEEE 1588 PTP hardware clock functions as an active hot-standby reference with $+8\,\text{nanoseconds}$ phase accuracy.NMEA(ttyS0): The serial GPS receiver supplies the coarse calendar second, while the PPS line disciplines the exact nanosecond pulse edge.Std-Dev : 24ns: Confirms that hardware pulse variance is exceptionally low.
What the Admin Does Next
- Add
local stratum 1andlocal orphandirectives tochrony.confso that if the satellite connection is temporarily lost, this machine continues serving as an authoritative time reference for the rest of the private network. - Set up automated alerting on
Std-Devviachronyc refclocksto detect hardware antenna degradation or physical cabling faults.
What Can Go Wrong: Operational Pitfalls & Safeguards
| Failure Mode | Root Cause | Business Impact | Prevention & Remediation |
|---|---|---|---|
| Discontinuous Clock Jumps | Unrestricted makestep directives active during production. |
Database panics, transaction rollbacks, duplicate log entries. | Restrict stepping to cold boot (makestep 1.0 3); verify with chronyc tracking. |
| Upstream Kiss-of-Death | Polling public servers too frequently (minpoll 2). |
Server is blacklisted; daemon drops all time sync. | Use standard intervals (minpoll 6 maxpoll 10); inspect peer status with ntpdata. |
| Routing Latency Asymmetry | Outbound traffic takes a faster path than inbound traffic. | Silent, systematic time offset introduced into all nodes. | Measure path differences with ntpdata; apply manual calibration offsets in chrony.conf. |
1. Unrestricted Clock Stepping During Production Workloads
- The Pitfall: Leaving
makestep 0 -1enabled or runningchronyc makestepon an active database cluster. If the physical hardware clock has drifted, the system clock will jump instantly. Distributed systems like PostgreSQL, Cassandra, or CockroachDB may panic, locks will expire prematurely, and monitoring agents will discard out-of-order metrics. - Mitigation: Enforce
makestep 1.0 3inchrony.conf. This restricts stepping strictly to the first three clock updates after the system boots. For all subsequent adjustments, Chrony will slew the clock smoothly without breaking monotonicity.
2. Upstream Firewall Rate-Limiting and "Kiss-of-Death" Packets
- The Pitfall: Configuring excessively aggressive polling intervals (
minpoll 2 maxpoll 2, polling every 4 seconds) against public pool servers. Upstream servers will interpret this as a denial-of-service attempt and respond with an NTP "Kiss-of-Death" (KoD) packet containing the rate-limiting codeRATEorDENY.chronydwill immediately detach from the server, leaving your node unsynchronised. - Mitigation: Use standard polling intervals (
minpoll 6 maxpoll 10, which poll between every 64 and 1024 seconds). Check peer health usingchronyc ntpdata <IP>to ensure no KoD packets have been received.
3. Network Path Asymmetry Inducing Silent Time Bias
- The Pitfall: Assuming that network transit times are perfectly symmetric. If outbound UDP packets travel over a fast direct link ($5\,\text{ms}$) but inbound responses return over a congested backup tunnel ($25\,\text{ms}$), the total round trip is $30\,\text{ms}$. Standard NTP mathematics assumes equal transit times ($15\,\text{ms}$ each way), introducing a silent $+10\,\text{ms}$ systematic offset into your clock.
- Mitigation: Use
chronyc ntpdatato measure delay variances across network interfaces. Where physical routing paths are permanently asymmetric, use theoffsetparameter on theserverdirective inchrony.confto calibrate and cancel out the path difference.
Today's Takeaway
To immediately verify the temporal health of your own system, open a terminal right now and run chronyc -n tracking. Look closely at three key figures: ensure your Stratum is between 1 and 3, confirm that your RMS offset is well within safe boundaries (under 1 millisecond for cloud servers, and under 10 microseconds for bare metal), and compute your host's maximum uncertainty using $\text{Root Dispersion} + (\text{Root Delay} / 2)$. In less than five minutes, you will have mathematical proof that your system clock is healthy, keeping your security certificates, transaction ledgers, and database clusters perfectly in sync.