Sar: Auditing Historical System Metrics, Triaging Ephemeral Resource Spikes, and Diagnosing Production Outages
Yet by 3:22 AM, when your secure shell terminal finally connects to the offending machine, the storm has passed into eerie, absolute silence. Processor utilisation hovers at an innocent four percent, memory allocation is calm, and network traffic flows without the slightest friction. Running interactive live monitors like top, htop, or vmstat shows an operating environment in pristine health. The catastrophe happened, service-level agreements were shattered, and bewildered engineers are left staring at a completely placid system with zero active clues remaining in memory.
This "phantom outage" is one of the most frustrating realities of systems administration. Modern operating systems do not pause to wait for human witnesses when resource contention peaks. Ephemeral spikesβa thirty-second burst of un-indexed database queries, lock contention storms, hypervisor CPU stealing, or network buffer overflowsβstrike with devastating speed, wreak havoc, and vanish before an engineer can type a command.
To solve these crimes after the perpetrator has left the scene, Linux provides a built-in flight data recorder: the System Activity Reporter, or sar(1), the cornerstone of the venerable sysstat telemetry suite. Rather than offering fleeting point-in-time glances that evaporate the moment you exit your terminal, sar continuously captures, structures, and archives detailed kernel performance statistics into lightweight local archives.
If you suspect a server is misbehaving right now, the single most useful command to get an immediate, multi-second health check across user processes, kernel routines, and I/O wait times is:
sar -u 1 5
Linux 6.8.0-48-generic (edge-prod-compute-04) 08/17/2026 _x86_64_ (16 CPU)
03:15:01 PM CPU %user %nice %system %iowait %steal %idle
03:15:02 PM all 8.25 0.00 2.10 0.05 0.10 89.50
03:15:03 PM all 45.30 0.00 14.20 0.15 0.20 40.15
03:15:04 PM all 82.10 0.00 16.40 0.00 0.30 1.20
03:15:05 PM all 84.00 0.00 15.80 0.00 0.10 0.10
03:15:06 PM all 12.40 0.00 4.20 0.00 0.00 83.40
Average: all 46.41 0.00 10.54 0.04 0.14 42.87
In five seconds, this command reveals what live tools often obscure: a sudden three-second compute burst peaking at 84% user execution and 16% kernel overhead, before instantly returning to 83% idle capacity.
What It Does and How It Works
At its heart, sar operates like a digital security camera system for your server's core subsystems. Unlike heavy third-party monitoring agents that consume significant memory and stream megabytes of JSON over the network, the sysstat architecture was designed for near-zero overhead, operational resilience, and deterministic local storage.
The architecture decouples the collection of raw telemetry from its analysis:
CPU & Tasks"] PMem["/proc/meminfo
Memory & Swap"] PDisk["/proc/diskstats
Block Storage I/O"] PNet["/proc/net/dev
Network Interfaces"] end sadc["sadc Engine
System Activity Data Collector"] saDD[("/var/log/sysstat/saDD
Compact Binary Logs")] sarQuery["sar CLI
Human-Readable Forensic Reports"] sadfTrans["sadf CLI
JSON / CSV / SVG Structured Export"] Kernel -->|Periodic Polling via sa1 / systemd| sadc sadc -->|Raw Binary Structs| saDD saDD -->|Parse & Compute Deltas| sarQuery saDD -->|Transpile Data| sadfTrans
The Engine Components
- The System Activity Data Collector (
sadc(8)): Written in highly optimized C,sadcis the backend collector. Triggered periodically by systemd timers or cron, it reads kernel accounting structures directly from the/procfilesystemβincluding/proc/stat,/proc/meminfo,/proc/diskstats, and/proc/net/dev. Instead of burning CPU cycles converting these metrics into text strings or schemas on the fly, it writes the raw binary memory structures straight to/var/log/sysstat/saDD(whereDDrepresents the day of the month). - The Automation Wrappers (
sa1andsa2): Lightweight shell scripts manage the collection lifecycle. Thesa1script invokessadcto gather regular interval samples, whilesa2runs daily housekeeping, producing consolidated summary reports and pruning logs according to your retention policies. - The Forensic Reader (
sar): When you query historical performance,saracts as an analytical reader. It does not poll the kernel again; instead, it reads the historical binary archives, calculates mathematical deltas between consecutive sample points, adjusts for time elapsed, converts raw tick counts into per-second rates or percentages, and formats the output into clean terminal tables. - The Modern Transpiler (
sadf(1)): Designed to integrate historical data with modern observability pipelines,sadfextracts the proprietary binary logs into industry-standard JSON, CSV, or XML for ingestion into spreadsheets, Elasticsearch, or dashboard engines.
The Investigator's Toolkit: Core Flags
Navigating system metrics during a critical outage post-mortem requires moving across hardware dimensions quickly. The following flags form the core vocabulary of historical investigation:
| Flag | Subsystem Profiled | Primary Kernel Source | Primary Diagnostic Utility |
|---|---|---|---|
-u |
CPU Utilization (%usr, %sys, %iowait, %steal) |
/proc/stat |
General compute exhaustion & hypervisor contention |
-q |
Task Run Queue & Load Averages | /proc/loadavg |
Thread starvation, queue latency, & CPU backlog |
-r |
Memory Allocation & Page Buffers | /proc/meminfo |
Application working set growth vs. kernel cache eviction |
-S |
Swap Space Allocation & Exhaustion | /proc/meminfo |
Memory pressure spillover & swap thrashing |
-B |
Kernel Paging Statistics | /proc/vmstat |
Virtual memory page-ins, page-outs, and page faults |
-d |
Block Device I/O & Queue Depth | /proc/diskstats |
Storage bottlenecks, latency spikes, and write stalls |
-n |
Network Activity (DEV, EDEV, TCP, ETCP) |
/proc/net/* |
Interface packet drops, bandwidth limits, TCP resets |
-w |
Task Creation & Context Switching | /proc/stat |
Context-switch storms and fork-bomb behavior |
-f |
Historical Log File Target | Binary Archives | Queries historical daily archives (/var/log/sysstat/saDD) |
-s / -e |
Start / End Time (hh:mm:ss) |
N/A | Isolates forensic queries to the exact alert window |
Five Real-World Production Investigations
Case 1: Diagnosing Invisible Hypervisor CPU Steal
The Scenario
An e-commerce payment microservice running inside a public cloud virtual machine suffered a flurry of HTTP 504 gateway timeouts between 03:00 UTC and 03:45 UTC on the 17th of the month. Database queries showed no lock contention, yet payment worker threads stopped processing requests. You need to verify whether the bottleneck was caused by runaway application code, system call overhead, or noisy neighbors on the underlying cloud host.
Forensic Command
sar -u -q -s 03:00:00 -e 03:45:00 -f /var/log/sysstat/sa17
Terminal Output
Linux 6.8.0-48-generic (pay-prod-app-02) 08/17/2026 _x86_64_ (8 CPU)
03:00:00 AM CPU %user %nice %system %iowait %steal %idle
03:10:00 AM all 14.20 0.00 3.10 0.12 0.15 82.43
03:20:00 AM all 18.50 0.00 4.20 0.10 42.80 34.40
03:30:00 AM all 16.10 0.00 3.80 0.08 55.40 24.62
03:40:00 AM all 15.80 0.00 3.50 0.11 38.10 42.49
03:50:00 AM all 12.10 0.00 2.90 0.05 0.10 84.85
Average: all 15.34 0.00 3.50 0.09 27.31 53.76
03:00:00 AM runq-sz plist-sz ldavg-1 ldavg-5 ldavg-15 blocked
03:10:00 AM 2 1024 0.85 0.92 0.88 0
03:20:00 AM 48 1180 14.20 8.50 3.20 0
03:30:00 AM 86 1250 28.60 18.40 9.10 0
03:40:00 AM 35 1190 11.30 14.10 10.20 0
03:50:00 AM 1 1018 0.75 2.10 4.80 0
Average: 34 1132 11.14 8.80 5.64 0
Line-by-Line Breakdown
%steal(Surging to 55.40% at 03:30:00 AM): Shows that your virtual machine was ready to execute instructions, but the physical host hypervisor refused to allocate CPU cycles, diverting physical processor time to other virtual machines.%user(16.10%) and%system(3.80%): Confirms that application code and kernel system calls were nowhere near exhausting the serverβs 8 virtual cores.%iowait(0.08%): Proves that local disk operations and storage stalls were not responsible for the slowdown.runq-sz(86 waiting tasks): On an 8-vCPU system, having 86 tasks queued means more than 10 runnable threads were lined up per core, starved of execution time purely because the hypervisor had paused the virtual CPU.blocked(0): Confirms processes were not stuck in uninterruptible disk sleep.
What the Administrator Does Next
Do not waste time refactoring application code or tuning runtime garbage collection. The problem is infrastructure virtualization contention. Immediately migrate the virtual machine away from the oversubscribed host, move to dedicated instance classes with guaranteed compute allocations (such as switching from burstable t4g instances to compute-optimized c6i instances), or raise a priority ticket with your cloud platform provider regarding hypervisor overcommitment.
Case 2: Tracking Memory Leaks and Out-Of-Memory Crashes
The Scenario
A high-throughput Redis caching server suddenly stopped answering requests at 02:40 UTC. The Redis daemon vanished without logging a shutdown reason or generating an application core dump. You suspect the Linux OOM Killer executed a forceful termination, and you need historical proof to determine whether physical RAM or kernel caches were depleted.
Forensic Command
sar -r -S -f /var/log/sysstat/sa17 -s 02:00:00 -e 03:00:00
Terminal Output
Linux 6.8.0-48-generic (cache-prod-redis-01) 08/17/2026 _x86_64_ (32 CPU)
02:00:00 AM kbmemfree kbavail kbmemused %memused kbbuffers kbcached kbcommit %commit
02:10:00 AM 16543200 28450100 48956800 74.74 450120 12540000 42100000 64.27
02:20:00 AM 8234100 18120400 57265900 87.43 310200 9850000 58400000 89.16
02:30:00 AM 1240500 3100200 64259500 98.11 85400 1840000 72100000 110.07
02:40:00 AM 410200 850100 65089800 99.37 12000 420000 84500000 129.00
02:50:00 AM 48250100 56100200 17249900 26.34 210400 8120000 18200000 27.78
Average: 14935620 21324200 50564380 77.20 213624 6554000 55060000 84.06
02:00:00 AM kbswpfree kbswpused %swpused kbswpcad %swpcad
02:10:00 AM 8388604 0 0.00 0 0.00
02:20:00 AM 6144000 2244604 26.76 450100 20.05
02:30:00 AM 1200400 7188204 85.69 1850400 25.74
02:40:00 AM 0 8388604 100.00 2450100 29.21
02:50:00 AM 8388604 0 0.00 0 0.00
Average: 4824362 3564242 42.49 950120 15.00
Line-by-Line Breakdown
%memused(Climbing to 99.37%) vskbcached(Plummeting from 12.5 GB to 420 MB): Demonstrates that the kernel aggressively evicted almost all filesystem page caches (kbcached) and buffers (kbbuffers) to satisfy growing anonymous memory requests from Redis.%commit(Reaching 129.00%): Indicates that application processes requested 29% more memory than the physical RAM and swap space could collectively provide under the system's memory overcommit policy.%swpused(Hitting 100.00% at 02:40:00 AM): Confirms total swap exhaustion. With physical memory at near zero and swap fully consumed, the kernel had no remaining options to relieve pressure.- The Recovery at 02:50:00 AM (
%memuseddropping to 26.34%): This sudden drop in used memory is the signature of an abrupt process death caused by the kernel Out-of-Memory Killer.
What the Administrator Does Next
Verify the OOM kill event by checking the kernel ring buffer (dmesg -T | grep -i oom-killer). Enforce strict memory ceilings in Redis configuration files (maxmemory 50gb coupled with an eviction policy such as volatile-lru), configure Linux virtual memory overcommit settings via sysctl (vm.overcommit_memory = 2 with vm.overcommit_ratio), or add physical memory to the host.
Case 3: Pinpointing Storage Bottlenecks and Controller Saturation
The Scenario
A primary PostgreSQL database server experienced severe query commit stalls and connection pool exhaustion at 08:30 UTC. You must determine whether the underlying high-performance NVMe storage array suffered from write throughput saturation, read starvation, or hardware controller queue lockups.
Forensic Command
sar -d -p -s 08:15:00 -e 09:00:00 -f /var/log/sysstat/sa17
Terminal Output
Linux 6.8.0-48-generic (db-prod-pg-01) 08/17/2026 _x86_64_ (64 CPU)
08:15:00 AM DEV tps rkB/s wkB/s areq-sz aqu-sz await r_await w_await %util
08:20:00 AM nvme0n1 850.40 12400.00 45100.00 67.62 0.82 0.96 0.45 1.10 22.40
08:30:00 AM nvme0n1 4200.10 18200.00 489000.00 120.76 64.50 15.35 2.10 16.50 99.80
08:40:00 AM nvme0n1 4850.80 15100.00 512000.00 108.66 82.10 16.92 2.30 18.20 100.00
08:45:00 AM nvme0n1 1100.20 9500.00 62000.00 64.99 1.10 1.05 0.50 1.20 28.10
Average: nvme0n1 2750.38 13800.00 277025.00 90.51 37.13 10.57 1.34 11.75 62.58
Line-by-Line Breakdown
%util(99.80% to 100.00%): The storage device was actively busy servicing I/O requests for 100% of the measurement interval, indicating the storage subsystem was completely saturated.aqu-sz(Average Queue Size: 82.10 requests): A significant backlog of unserviced I/O operations accumulated in the operating system device queue waiting for physical disk access.w_await(18.20 ms) vsr_await(2.30 ms): Isolates the latency issue specifically to writes. While reads completed promptly (2.3 ms), write latency degraded by more than 1600% compared to the baseline at 08:20 AM (1.10 ms).wkB/s(Over 512 MB/s of sustained sequential writes): Proves that write volume overwhelmed the storage controller cache, forcing database transactions into synchronous disk wait stalls.
What the Administrator Does Next
The PostgreSQL Write-Ahead Log (WAL) flushing mechanism is hitting physical device write limits. Separate the database workloads across dedicated physical disks: place pg_wal on a dedicated NVMe drive isolated from the main base/ data directory, adjust PostgreSQL checkpoint settings (checkpoint_timeout = 15min, max_wal_size = 16GB) to smooth write spikes, and ensure battery-backed write-caching is enabled on hardware RAID controllers.
Case 4: Uncovering Network Interface Saturation and Packet Drops
The Scenario
An ingress API gateway routing traffic to Kubernetes services experienced client connection timeouts and dropped requests between 14:00 UTC and 14:30 UTC. You need to find out whether the dropouts were caused by physical link-layer saturation or kernel TCP connection queue limits.
Forensic Command
sar -n DEV,EDEV,ETCP -s 14:00:00 -e 14:35:00 -f /var/log/sysstat/sa17
Terminal Output
Linux 6.8.0-48-generic (gw-prod-k8s-ingress) 08/17/2026 _x86_64_ (16 CPU)
14:00:00 PM IFACE rxpck/s txpck/s rxkB/s txkB/s rxcmp/s txcmp/s rxmcst/s %ifutil
14:10:00 PM eth0 85400.10 92100.40 85200.00 94100.00 0.00 0.00 0.00 75.28
14:20:00 PM eth0 125000.50 138000.80 119500.00 124200.00 0.00 0.00 0.00 99.36
14:30:00 PM eth0 128500.20 141200.10 121000.00 124800.00 0.00 0.00 0.00 99.84
Average: eth0 112966.93 123767.10 108566.67 114366.67 0.00 0.00 0.00 91.49
14:00:00 PM IFACE rxerr/s txerr/s coll/s rxdrop/s txdrop/s txcarr/s rxfram/s rxfifo/s
14:10:00 PM eth0 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
14:20:00 PM eth0 0.00 0.00 0.00 1450.80 0.00 0.00 0.00 0.00
14:30:00 PM eth0 0.00 0.00 0.00 2890.40 0.00 0.00 0.00 0.00
Average: eth0 0.00 0.00 0.00 1447.07 0.00 0.00 0.00 0.00
14:00:00 PM atmptf/s estres/s retrans/s isegerr/s ores_s
14:10:00 PM 0.12 0.45 2.10 0.00 0.00
14:20:00 PM 18.40 42.10 1850.50 0.00 0.00
14:30:00 PM 24.80 58.30 3120.00 0.00 0.00
Average: 14.44 33.62 1657.53 0.00 0.00
Line-by-Line Breakdown
%ifutil(99.84% oneth0): The virtual network interface reached 100% capacity on its 1 Gbps (approximately 125 MB/s) link.rxdrop/s(2,890.40 dropped incoming frames/sec): The kernel network stack and hardware ring buffers dropped nearly three thousand incoming packets every second because buffer descriptor rings were completely full.retrans/s(3,120 TCP retransmissions/sec): The direct transport-layer fallout of network packet loss. Retransmission rates skyrocketed from 2.1 per second to over 3,100 per second, triggering exponential backoff delays and cascading client timeouts.atmptf/s(24.80 failed connection attempts/sec): Incoming TCP three-way handshakes (SYNpackets) failed or timed out during initial connection establishment.
What the Administrator Does Next
The network interface has hit its maximum throughput. To resolve:
1. Increase network interface ring buffers: sudo ethtool -G eth0 rx 4096 tx 4096.
2. Expand kernel connection backlog limits: sudo sysctl -w net.core.netdev_max_backlog=10000 and sudo sysctl -w net.ipv4.tcp_max_syn_backlog=8192.
3. Upgrade the virtual machine instance to a tier featuring enhanced networking (such as AWS ENA or SR-IOV supporting 10 to 25 Gbps bandwidth) and distribute ingress traffic across multiple gateway instances using a load balancer.
Case 5: Exporting Structured Incident Data for Post-Mortems
The Scenario
Following a major outage, leadership requires full telemetry data to be imported into a centralized Jupyter analysis notebook and Elasticsearch cluster to establish an accurate incident timeline. You need to convert the proprietary binary sa17 file into clean, structured CSV and RFC-compliant JSON without losing metric precision.
Command for CSV Export
sadf -d -s 03:00:00 -e 04:00:00 -f /var/log/sysstat/sa17 -- -u -r > incident_metrics.csv
Command for JSON Export
sadf -j -s 03:00:00 -e 04:00:00 -f /var/log/sysstat/sa17 -- -u -q -d > incident_telemetry.json
Structured JSON Output
{
"sysstat": {
"hosts": [
{
"nodename": "edge-prod-compute-04",
"sysname": "Linux",
"release": "6.8.0-48-generic",
"machine": "x86_64",
"number-of-cpus": 16,
"date": "2026-08-17",
"statistics": [
{
"timestamp": {
"time": "03:15:00",
"utc": 1,
"interval": 600
},
"cpu-load": [
{
"cpu": "all",
"usr": 82.10,
"nice": 0.00,
"sys": 16.40,
"iowait": 0.00,
"steal": 0.30,
"idle": 1.20
}
],
"queue": {
"runq-sz": 64,
"plist-sz": 1420,
"ldavg-1": 18.25,
"ldavg-5": 12.10,
"ldavg-15": 5.40,
"blocked": 0
}
}
]
}
]
}
}
Technical Utility and Pipeline Integration
- The JSON generated by
sadf -jincludes rich host metadata (host architecture, CPU count, sampling interval, and UTC timestamps). - This format allows immediate parsing in Python pandas pipelines (
pd.read_json()), command-line filtering viajq, or direct forwarding via Logstash and Fluentbit into centralized dashboards. - The double dash (
--) syntax cleanly separatessadfexport options from standardsarmetric selection flags (-u,-q,-d).
Production Pitfalls and How to Avoid Them
Working with historical system telemetry presents subtle traps that can mislead even seasoned engineers during an incident review.
(Incident Completely Hidden!)"] end Actual -.-> Sampling Sampling --> Recorded
1. The Metric Smoothing (Aliasing) Trap
By default, older Linux distributions configure /etc/cron.d/sysstat to sample system metrics every 10 minutes (*/10 * * * *).
- The Danger: If an application suffers a catastrophic CPU lockup or thread storm that lasts for 45 seconds before crashing, a 10-minute sampling cadence averages that load across 600 seconds. A 45-second 100% spike averaged across 10 minutes registers as a harmless 7.5% increase in CPU usage. The outage becomes invisible in the historical logs.
- The Solution: On critical production nodes, increase the collection resolution to 1 minute or 10 seconds. Under modern systemd setups, customize the collection timer:
sudo systemctl edit sysstat-collect.timer
Add an aggressive, lightweight sampling interval:
[Unit]
Description=High-Resolution 1-Minute Metric Collection Timer
[Timer]
OnCalendar=
OnCalendar=*:0/1
AccuracySec=1s
Reload the systemd daemon to activate:
sudo systemctl daemon-reload
sudo systemctl restart sysstat-collect.timer
2. Timezone Mismatches During Incidents
When cloud alerts trigger in UTC while an engineer's shell session uses local time (such as BST, EST, or JST), passing -s 03:00:00 -e 03:30:00 queries local log entries rather than the actual alert window.
- The Solution: Always enforce explicit UTC evaluation when running queries across distributed infrastructure by prefixing your command with the
TZenvironment variable:
TZ=UTC sar -u -f /var/log/sysstat/sa17
3. Binary Incompatibility Across Operating System Upgrades
The files stored in /var/log/sysstat/saDD contain raw, unpadded C data structures written directly from memory. They are not architecture-portable and are often incompatible across different major releases of sysstat. Attempting to read a file created under version 11 on a server running version 12 will result in Invalid system activity file errors.
- The Solution: If you need to store performance logs for compliance or long-term auditing, export them to structured JSON or CSV at generation time using
sadf:
sadf -j /var/log/sysstat/sa17 -- -A > /archived/telemetry/$(hostname)_sa17.json
Today's Takeaway
The difference between a frantic, speculative outage call and a calm, evidence-based post-mortem lies in whether your servers record their own operational history. Right now, open a terminal on your Linux server and verify that sysstat telemetry is actively running by executing sar -q. If you receive an error indicating that data collection is not enabled, install the suite through your package manager, open /etc/default/sysstat, set ENABLED="true", and start the background recorder with sudo systemctl enable --now sysstat. Ensuring this lightweight service runs in the background costs negligible CPU cycles, but provides the indispensable evidence you will need when the next 3:00 AM production emergency strikes.