Powernews Thursday, 20 August 2026 at 07:01 CEST
UNIX COMMAND OF THE DAY

Atop: Auditing Historical System Workloads, Replaying Ephemeral Microburst Telemetry, and Inspecting Kernel Process Accounting in Production

It is 3:14 AM on a damp Tuesday when the piercing vibration of an on-call phone tears through the quiet of a dark bedroom. Your heart rate spikes before your eyes can fully focus on the harsh glare of the screen: a core database replica has abruptly dropped out of service, customer transactions are stalling, and alert notifications are cascading across the screen. You stumble out of bed, flip open your laptop, and authenticate over your secure connection with half-awake urgency. But by the time your terminal cursor blinks back at you, the storm has passed. The graphs have settled into a flat, deceptive calm, the servers hum along as if nothing happened, and you are left wide awake in the dark, wondering what on earth just brought your infrastructure to its knees.
Key Takeaway
Essential takeaway summary for Atop: Auditing Historical System Workloads, Replaying Ephemeral Microburst Telemetry, and Inspecting Kernel Process Accounting in Production.

This is the quintessential ghost in the machine that every system administrator dreads. When you turn to standard diagnostics like top or htop, or glance at monitoring dashboards that calculate averages on thirty-second intervals, you see only a tranquil snapshot of the present: processor usage sits at a placid 12%, memory pressure is normal, and storage throughput is well within safe thresholds. Yet somewhere during those critical missing minutes, an aggressive burst of resource contention brought the system to a standstill before vanishing completely.

Standard process monitors fail during these outages because they are instantaneous observersβ€”they only report the state of the machine at the exact second they take a sample. If a rogue batch job or unthrottled background script spawns thousands of sub-tasks that exhaust CPU pipelines or storage queues and terminate within a two-second window, standard tools never even register that those tasks existed. To solve these transient incidents, you need more than a live snapshot; you need a continuous, high-fidelity flight data recorder. That tool is atop.

The single most useful command you can run to cut through the confusion and get an immediate, full-spectrum assessment of your system is:

atop 10

Running this command launches an interactive, live console that aggregates system metrics across a ten-second sampling window. In the upper half of the display, atop presents a unified view of your processor cores, physical memory, swap space, block devices, and network throughput. In the lower pane, active processes are dynamically ranked by resource consumption, with any bottlenecked subsystem highlighted in warning colors so you can diagnose issues at a glance.

ATOP - ip-10-0-4-128          2026/08/20  05:15:22          -----------          10s elapsed
PRC | sys    0.84s | user   3.12s | #proc    284 | #zombie    0 | #exit     14 |
CPU | sys       8% | user     31% | irq       1% | idle    360% | wait      0% |
CPL | avg1    0.42 | avg5    0.68 | avg15   0.55 | csw   14280 | intr    8920 |
MEM | tot    31.1G | free   18.4G | slab    1.8G | dirty  14.2M | buff  412.0M |
SWP | tot     8.0G | free    8.0G |              |              | vmcom   6.2G |
DSK |          nvme0n1 | busy      3% | read      12 | write    184 | avq    0.12 |
NET | transport    | tcpi    1240 | tcpo    1180 | udpi       4 | udpo       4 |

PID  TID S  CPU   SYSCPU   USRCPU   VGROW   RGROW  RDDSK  WRDSK  RNET  WNET  CMD          1/5
 3842    - S  34%    0.62s    2.78s      0K      0K     0K   1.2M    0K    0K  postgres
 1290    - S   4%    0.12s    0.28s      0K      4K     0K   128K    0K    0K  node_exporter
 4102    - S   1%    0.08s    0.02s      0K      0K     0K     0K    0K    0K  containerd

1. What It Does in Plain English

At its core, atop is a comprehensive performance monitor and forensic retrospective tool for Linux environments. While tools like top show you what is happening right now, atop records aggregate hardware utilisation alongside fine-grained, per-process telemetry over time.

Crucially, it captures ephemeral processes that spawn and exit between sampling intervals by listening directly to kernel-level accounting events. Furthermore, atop runs as a lightweight background daemon that writes this telemetry into compressed daily log files. When a production incident occurs at 3:00 AM, an engineer logging on at 9:00 AM does not need to guess what happened: they can rewind atop to the exact minute of the failure and replay the incident second by second.


2. Core Flags & Quick Start

Mastering atop involves learning its key interactive modifiers and command-line flags:

  • -r [file | YYYYMMDD]: Opens a historical daily log file for retrospective post-mortem playback.
  • -b HH:MM: Specifies the start time boundary when opening a historical log.
  • -e HH:MM: Specifies the end time boundary for targeted investigations.
  • -m: Switches both system and process views to memory-centric metrics (virtual size, resident memory, cache growth, and page faults).
  • -d: Toggles the disk storage view, showing per-process read/write throughput alongside hardware queue depths.
  • -n: Activates network subsystem profiling (requires the optional netatop kernel module).
  • -y: Expands the view to show individual thread-level scheduling details (TID).
  • -w file [interval [samples]]: Writes real-time raw binary telemetry to a specified log file.

3. The Architectural Engine: Kernel Accounting, /proc Virtual Sampling, and Ephemeral Process Tracking

Traditional monitoring agents collect data exclusively by polling the /proc virtual filesystem at scheduled intervals:

  1. /proc/stat: Aggregate CPU tick allocation across user, system, idle, and I/O wait states.
  2. /proc/meminfo: Real-time allocation of physical RAM, buffers, and reclaimable cache.
  3. /proc/diskstats: Monotonic counters for storage operations, sector transfers, and queue times.
  4. /proc/net/dev: Network packet and byte transmission counters per interface.
  5. /proc/[pid]/stat and /proc/[pid]/io: Per-process CPU cycles, resident memory, and storage I/O.
flowchart TD subgraph Kernel["Linux Kernel Subsystems"] ProcStat["/proc/stat, /proc/meminfo, /proc/diskstats
(Periodic System Metrics)"] ProcPid["/proc/[pid]/stat, /proc/[pid]/io
(Active Process Sampling)"] PACCT["Kernel Process Accounting: acct(2)
(Triggered on Process Exit)"] NetAtop["netatop.ko Module
(Per-Socket Interception)"] end Engine["atop Collection Engine"] ProcStat --> Engine ProcPid --> Engine PACCT --> Engine NetAtop --> Engine TUI["Interactive Console
(Live & Historical Playback)"] DailyLog["Compressed Binary Daily Log
(/var/log/atop/atop_YYYYMMDD)"] Engine --> TUI Engine --> DailyLog

The Ephemeral Blind Spot and Process Accounting (PACCT)

The fundamental limitation of polling /proc lies in short-lived processes. If a build pipeline or shell script spawns ten thousand sub-processes that run for fifty milliseconds each, a monitor taking snapshots every ten seconds will miss nearly all of them. The system CPU graph will show a massive spike, but the process list will appear empty.

atop overcomes this blind spot by connecting to Linux Kernel Process Accounting (PACCT). When PACCT is enabled, the kernel writes an accounting record to a spool file the moment any task terminates. This record contains: * The exact command string and exit status. * Total CPU execution time (user and system). * Start time, total elapsed runtime, and user ID. * Cumulative memory footprint and disk I/O transferred.

During each sampling interval, atop reads /proc for active processes and drains the PACCT spool for all processes that completed during that window. It presents these finished tasks with an E (exit) state marker, ensuring no resource consumer escapes auditability.

Network Attribution via Netatop

Standard Linux kernels do not expose per-process network bandwidth through /proc. To bridge this gap, atop can pair with netatop, an optional kernel module that registers lightweight hooks in the network stack to attribute incoming and outgoing traffic directly to individual process sockets.

Compressed Binary Persistence

Rather than generating sprawling text or unstructured JSON logs, the atop.service background daemon stores raw binary data structs in /var/log/atop/atop_YYYYMMDD. Compressed with zlib, an entire host's granular, second-by-second historical telemetry typically consumes less than 50 megabytes per day.


4. Five Real-World Enterprise-Grade Use Cases

Use Case 1: Retrospective Post-Mortem Incident Triage

Scenario

A Kubernetes worker node experienced intense CPU throttling at 03:02 AM, causing several application containers to time out. Standard metrics recorded the CPU saturation, but the offending processes were gone by the time alerts fired. We must rewind the daily log /var/log/atop/atop_20260820 to inspect the interval between 02:55 AM and 03:10 AM.

Command Invocation

atop -r /var/log/atop/atop_20260820 -b 02:55 -e 03:10

Key interactive controls during playback: * Press t: Step forward one sampling interval. * Press T: Step backward one interval. * Press b: Jump directly to a timestamp (e.g. 03:02). * Press P: Sort processes by CPU usage. * Press z: Pause auto-advance on a specific frame.

Realistic Terminal Output

ATOP - k8s-node-worker-09    2026/08/20  03:02:10          -----------          10m00s elapsed
PRC | sys   42.10s | user  348.12s | #proc    412 | #zombie    0 | #exit   14820 |
CPU | sys      12% | user      87% | irq       1% | idle      0% | wait      0% |
CPL | avg1   18.42 | avg5    8.12 | avg15   3.45 | csw  489120 | intr  210940 |
MEM | tot    62.8G | free    4.1G | slab    4.2G | dirty  82.1M | buff  820.0M |
SWP | tot     0.0G | free    0.0G |              |              | vmcom  48.2G |
DSK |          nvme1n1 | busy     42% | read    4812 | write  18490 | avq    1.84 |

PID   TID S  CPU   SYSCPU   USRCPU   VGROW   RGROW  RDDSK  WRDSK  ST  EXC  CMD          1/8
18942     - E  78%   31.20s  284.10s      0K      0K   1.2G     0K  _E    0  rotate_logs.sh
19844     - E  14%    8.40s   48.20s      0K      0K     0K   412M  _E    0  gzip
 2104     - S   4%    1.12s    8.14s      0K      0K    12K   4.1M  --    -  kubelet
 3912     - S   2%    0.84s    4.21s      4K      8K   180K   1.2M  --    -  containerd
Triage Parameter Observation & Finding
Incident Classification Ephemeral Microburst Workload
Process Turnover 14,820 short-lived processes spawned and exited (#exit 14820)
Primary Root Cause rotate_logs.sh (PID 18942) invoking unthrottled gzip (PID 19844)
Subsystem Impact 99% combined CPU saturation and run-queue load spike to 18.42

Line-by-Line Telemetry Breakdown

  • PRC | #exit 14820: During this ten-minute window, 14,820 processes terminated. This rapid turnover explains why point-in-time monitors saw nothing.
  • CPU | user 87% | sys 12% | idle 0%: Aggregate compute capacity was 99% saturated, causing container throttling.
  • PID 18942 | ST _E: The _E flag confirms that rotate_logs.sh was an exited ephemeral process captured via kernel accounting.
  • PID 19844 | CMD gzip: Shows that the rotation script spawned an unthrottled compression worker, consuming 412 MB of write throughput and dominating CPU time.

What the Administrator Does Next

  1. Update the log rotation configuration to constrain I/O and processor scheduling priority: bash nice -n 19 ionice -c 3 gzip
  2. Place the maintenance task in a dedicated systemd slice with strict limits (CPUQuota=50%) so it cannot overwhelm production workloads.

Use Case 2: Deep Memory Leak & Virtual Allocation Auditing

Scenario

A backend routing service displays progressive resident memory growth over three days. The engineering team needs to distinguish between active heap allocation, shared library mappings, and unmanaged growth using atop -m.

Command Invocation

atop -m 5

Realistic Terminal Output

ATOP - srv-payment-prod-02   2026/08/20  05:22:40          -----------           5s elapsed
MEM | tot    31.1G | free    1.2G | slab    2.8G | dirty  44.1M | buff  112.0M |
SWP | tot    16.0G | free    8.2G |              |              | vmcom  42.1G |
PAG | scan  184200 | steal 142100 | stall      0 |              | swin       0 |

PID  TID MINFLT MAJFLT  VSTEXT  VSIZE  RSIZE  PSIZE  VGROW  RGROW   MEM  CMD          1/4
 8492    -    840      0    1.8M  34.2G  24.1G  24.1G   128M    94M   77%  payment_router
 1102    -     12      0   14.2M   1.2G 412.0M 380.0M     0K     0K    1%  dockerd
 1048    -      0      0    4.1M 820.0M 142.0M 110.0M     0K     0K    0%  systemd-journal

Line-by-Line Telemetry Breakdown

  • MEM | vmcom 42.1G | tot 31.1G: Committed virtual memory (vmcom) stands at 42.1 GB, significantly exceeding physical RAM (31.1 GB).
  • PAG | scan 184200 | steal 142100: The kernel page scanner is active, scanning 184,200 pages and reclaiming 142,100 per interval to stave off out-of-memory panics.
  • PID 8492 | payment_router:
    • VSIZE (34.2G): Total virtual memory reserved by the process.
    • RSIZE (24.1G): Actual physical RAM currently mapped to this application.
    • VGROW (128M): Virtual address space expanded by 128 MB within this single five-second interval.
    • RGROW (94M): Physical resident memory increased by 94 MB in the same five seconds.

What the Administrator Does Next

  1. Trigger a heap allocation report directly from the live process without restarting it: bash gdb --batch --pid 8492 --eval-command "call (void)malloc_stats()"
  2. Enable native heap profiling via memory allocator environment flags: bash MALLOC_CONF="prof:true,prof_prefix:/var/log/jeprof/router" ./payment_router
  3. Establish a hard ceiling using systemd memory controls: ini # /etc/systemd/system/payment_router.service.d/override.conf [Service] MemoryMax=26G MemoryHigh=22G

Use Case 3: Pinpointing Storage Queue Depth & Disk I/O Saturation

Scenario

A PostgreSQL cluster reports sudden database query latency spikes exceeding 450ms. We need to verify whether the bottleneck is caused by client read queries, background checkpoint flushes, or Write-Ahead Log (WAL) synchronisation using atop -d.

Command Invocation

atop -d 5

Realistic Terminal Output

ATOP - db-pg-primary-01      2026/08/20  05:31:15          -----------           5s elapsed
DSK |          nvme0n1 | busy     98% | read     140 | write   8420 | avq   12.40 |
DSK |          nvme1n1 | busy      2% | read       0 | write    110 | avq    0.04 |

PID  TID   RDDSK   WRDSK  WCANCL  DSK  CMD                        1/6
 4120    -      0K   48.2M      0K  68%  postgres: checkpointer
 4124    -      0K   18.4M      0K  22%  postgres: walwriter
 4890    -   12.4M    1.2M      0K   6%  postgres: client_worker
 1280    -      0K    4.0K      0K   0%  systemd-journald
Storage Metric Observed Value Diagnostic Context
Device Saturation nvme0n1 at 98% busy Storage layer is operating at continuous maximum capacity
Average Queue (avq) 12.40 requests pending Block scheduler queue is heavily backlogged
Dominant Writer Checkpointer (48.2 MB/s) Writing bulk dirty pages without adequate pacing
Contending Writer WAL Writer (18.4 MB/s) Forced to compete for bandwidth on the same block device

Line-by-Line Telemetry Breakdown

  • DSK | nvme0n1 | busy 98%: The primary NVMe drive is handling operations 98% of the time, causing read/write queues to stall.
  • DSK | nvme0n1 | avq 12.40: Average Queue Depth. Over twelve concurrent I/O operations are waiting in the block layer scheduler.
  • PID 4120 | WRDSK 48.2M | DSK 68%: The Postgres checkpointer is pushing 48.2 MB/s of dirty buffer writeouts, monopolising storage bandwidth.
  • PID 4124 | WRDSK 18.4M | DSK 22%: The walwriter is contending on the same disk, delaying transaction commits.

What the Administrator Does Next

  1. Spread the checkpoint workload across a wider time interval in postgresql.conf: ini checkpoint_completion_target = 0.9 max_wal_size = 32GB checkpoint_timeout = 30min
  2. Move the Write-Ahead Log to the underutilised secondary device (nvme1n1): bash systemctl stop postgresql mv /var/lib/postgresql/16/main/pg_wal /mnt/nvme1n1/pg_wal ln -s /mnt/nvme1n1/pg_wal /var/lib/postgresql/16/main/pg_wal systemctl start postgresql

Use Case 4: Thread-Level Scheduling & Lock Contention

Scenario

A Go-based API gateway displays irregular latency spikes under load despite overall CPU utilisation remaining below 40%. We suspect thread lock contention and scheduling delays. We inspect thread-level details with atop -y and atop -g.

Command Invocation

atop -y -g 5

Realistic Terminal Output

ATOP - edge-gw-prod-04       2026/08/20  05:40:02          -----------           5s elapsed
PRC | sys    4.10s | user   18.20s | #proc    140 | #trun      8 | #tslpi   1840 |
CPU | sys      18% | user      24% | irq       2% | idle    356% | wait      0% |
CPL | csw    84920 | intr    48100 |              |              |              |

PID   TID  RUID    EUID    ST  CPU   SYSCPU   USRCPU  AVGSLI  WCHAN   CMD          1/8
 5912     -  root    root    --  38%    3.80s   15.20s   0.1ms  -       api_gateway
    -  5913  root    root    --  12%    1.20s    4.80s   0.0ms  futex   gw-worker-01
    -  5914  root    root    --  11%    1.10s    4.40s   0.0ms  futex   gw-worker-02
    -  5915  root    root    --  10%    1.00s    4.00s   0.0ms  futex   gw-worker-03
    -  5916  root    root    --   5%    0.50s    2.00s   0.0ms  futex   gw-metrics
 1042     -  systemd systemd --   1%    0.08s    0.12s   1.2ms  epoll_  systemd-resolved

Line-by-Line Telemetry Breakdown

  • CPL | csw 84920: Context Switches. Nearly 85,000 thread context switches occurred in five seconds, pointing to severe scheduling churn.
  • TID 5913-5915 | WCHAN futex:
    • The TID column identifies distinct execution threads inside the parent process (PID 5912).
    • WCHAN shows futex (Fast Userspace Mutex), indicating that threads are spending CPU cycles waiting for shared locks rather than processing network requests.
  • AVGSLI (0.0ms): The average scheduling latency slice is near zero, confirming that worker threads wake, immediately collide on mutex locks, and yield back to the kernel.

What the Administrator Does Next

  1. Profile lock contention using perf: bash perf record -s -F 99 -p 5912 -g -- sleep 10 perf report -n --stdio
  2. Refactor shared mutex structures into partitioned caches or lock-free data channels.
  3. Pin worker processes to specific CPU cores using systemd affinity settings: ini # /etc/systemd/system/api_gateway.service.d/override.conf [Service] CPUAffinity=0-3

Use Case 5: Automated Telemetry Ingestion & Extraction with Atopsar

Scenario

As part of an automated staging benchmark, we need to collect high-frequency system telemetry during a load test and programmatically verify that CPU and disk utilisation stay within target limits using atopsar.

Command Invocation (Data Collection)

# Record 12 samples at 5-second intervals directly to a binary capture file
atop 5 12 -w /tmp/benchmark_run.raw

Command Invocations (Metric Extraction)

1. Extract CPU Subsystem Activity (atopsar -c)
atopsar -c -r /tmp/benchmark_run.raw
srv-perf-bench-01  6.8.0-45-generic  2026/08/20

-------------------------- CPU ACTIVITY --------------------------
05:50:05  cpu  %usr  %nice  %sys  %irq  %softirq  %steal  %idle
05:50:10  all    24      0     6     0         1       0     69
05:50:15  all    48      0    12     0         2       0     38
05:50:20  all    82      0    14     0         3       0      1
05:50:25  all    85      0    13     0         2       0      0
05:50:30  all    79      0    11     0         2       0      8
05:50:35  all    30      0     8     0         1       0     61
05:50:40  all    12      0     4     0         0       0     84
------------------------------------------------------------------
2. Extract Block Storage Utilization (atopsar -d)
atopsar -d -r /tmp/benchmark_run.raw
srv-perf-bench-01  6.8.0-45-generic  2026/08/20

-------------------------- DISK ACTIVITY -------------------------
05:50:05  device     busy  read/s  KB/read  writ/s  KB/writ  avque
05:50:10  nvme0n1      4%       0        0     210       32   0.08
05:50:15  nvme0n1     18%      12      128     840       64   0.45
05:50:20  nvme0n1     88%     110      512    4120      128   8.40
05:50:25  nvme0n1     96%     142      512    5490      128  14.20
05:50:30  nvme0n1     91%      98      256    4910      128  11.10
05:50:35  nvme0n1     22%       4       64     920       32   0.50
05:50:40  nvme0n1      3%       0        0     180       16   0.05
------------------------------------------------------------------
3. Automated CI/CD Validation Script

To evaluate performance benchmarks in continuous integration pipelines, use awk to parse peak disk saturation:

#!/usr/bin/env bash
set -euo pipefail

RAW_LOG="/tmp/benchmark_run.raw"
MAX_ALLOWED_BUSY=85

# Extract maximum disk utilisation percentage from the test run
PEAK_BUSY=$(atopsar -d -r "${RAW_LOG}" | awk '$2 ~ /nvme0n1/ { gsub(/%/, "", $3); if($3>max) max=$3 } END { print max }')

echo "Calculated Peak Disk Saturation: ${PEAK_BUSY}%"

if [ "${PEAK_BUSY}" -ge "${MAX_ALLOWED_BUSY}" ]; then
    echo "CRITICAL: Benchmark failed. Storage saturation (${PEAK_BUSY}%) exceeded threshold (${MAX_ALLOWED_BUSY}%)." >&2
    exit 1
fi

echo "SUCCESS: Hardware subsystem performance within permissible bounds."
exit 0

What the Administrator Does Next

Integrate this verification script into automated deployment pipelines. If a pull request introduces an unindexed database query or unbuffered file writes, the test suite automatically fails and preserves benchmark_run.raw for developer review.


5. Subsystem Saturation Rules and Alerting Thresholds

atop evaluates hardware utilisation against defined thresholds, dynamically altering terminal colours to highlight bottlenecks:

Subsystem Metric Monitored Warning Threshold (Cyan) Critical Threshold (Red) Mathematical Condition
CPU Core Utilisation Utilization >= 70% Utilization >= 90% $\sum (\text{user} + \text{sys} + \text{irq}) \div \text{ncpus} \ge 0.90$
Memory Available RAM Available <= 10% Available <= 5% $(\text{MemFree} + \text{Cached} + \text{SReclaim}) \div \text{MemTotal} \le 0.05$
Storage Device Busy DSK %busy >= 70% DSK %busy >= 90% $(\Delta \text{io_ticks} \div \Delta \text{interval}) \ge 0.90$
Storage Queue Depth avq >= 3.0 avq >= 8.0 $\text{I/O Requests in Flight} \ge 8.0$
Paging Frequency Pages Scanned > 0 Pages Stolen > 100/s Continuous active page reclamation under memory pressure

6. What Can Go Wrong: Operational Pitfalls & Mitigations

Pitfall 1: Unrestricted Log Retention and Disk Exhaustion

By default, atop records daily binary logs to /var/log/atop/. On servers running high process volumes (such as CI/CD runners or busy container hosts), these files can grow quickly. If /var/log resides on your root filesystem, unmonitored logs can exhaust available disk space.

Mitigation Strategy: Enforce strict retention limits in /etc/default/atop or configure a dedicated logrotate rule:

# /etc/default/atop or /etc/sysconfig/atop
# Increase interval from 600s to 1200s if storage is limited
LOGINTERVAL=600
# Retain raw daily logs for precisely 7 days
LOGGENERATIONS=7

Pitfall 2: High Fork Rates and Process Accounting Spool Overhead

When kernel process accounting (PACCT) is active on servers experiencing thousands of process forks per second (such as unoptimised shell loops), writing exit records to /var/cache/atop.acct can generate disk I/O overhead.

Mitigation Strategy: If accounting overhead becomes noticeable, disable PACCT while keeping standard /proc sampling active:

# Deactivate process accounting system-wide if necessary
sysctl -w kernel.acct="0 0 0"

Pitfall 3: Binary Version Incompatibilities

atop binary log files record raw internal structs that can vary between major software releases (e.g. attempting to read logs generated by atop 2.8 using atop 2.11). Reading an incompatible file results in an error:

raw file /var/log/atop/atop_20260820 has incompatible format

Mitigation Strategy: Ensure your analysis host uses the same atop version as the server where the telemetry was captured, or use a lightweight container matching the host version.


7. Automated Maintenance: Production Log Retention Architecture

To keep log storage clean and predictable across servers, deploy this maintenance script to prune and manage historical log files:

#!/usr/bin/env bash
# ==============================================================================
# /usr/local/sbin/atop-prune.sh
# Automated maintenance script for atop historical log directories.
# ==============================================================================
set -euo pipefail

LOG_DIR="/var/log/atop"
RETENTION_DAYS=14
DISK_MAX_USAGE=85

if [ ! -d "${LOG_DIR}" ]; then
    echo "Log directory ${LOG_DIR} does not exist. Skipping."
    exit 0
fi

# Check disk utilisation for the log mount point
CURRENT_USAGE=$(df -P "${LOG_DIR}" | awk 'NR==2 {gsub(/%/, "", $5); print $5}')

if [ "${CURRENT_USAGE}" -gt "${DISK_MAX_USAGE}" ]; then
    echo "WARNING: Disk usage at ${CURRENT_USAGE}% exceeds maximum (${DISK_MAX_USAGE}%). Accelerating prune."
    # Aggressively remove files older than 3 days if disk is constrained
    find "${LOG_DIR}" -type f -name "atop_*" -mtime +3 -delete
fi

# Standard retention purge
echo "Executing standard purge for logs older than ${RETENTION_DAYS} days."
find "${LOG_DIR}" -type f -name "atop_*" -mtime +"${RETENTION_DAYS}" -delete

# Compress uncompressed raw daily dumps if present
find "${LOG_DIR}" -type f -name "atop_20*" ! -name "*.gz" -mtime +1 -exec gzip -9 {} +

echo "atop log management tasks successfully completed."
exit 0

Deploy the script using a scheduled systemd timer:

# /etc/systemd/system/atop-prune.timer
[Unit]
Description=Daily cleanup of atop historical log files
RefuseManualStart=no
RefuseManualStop=no

[Timer]
OnCalendar=*-*-* 01:00:00
Persistent=true

[Install]
WantedBy=timers.target

8. Authoritative References & Further Reading

To learn more about Linux system profiling and atop instrumentation, consult these technical resources:


9. Today's Takeaway

Standard live monitors only show you what is happening right now; atop gives you an exact, second-by-second record of what happened in the past. In the next five minutes, check whether the recording daemon is active on your machine by running systemctl status atop. Then, launch atop -r in your terminal and press t to step through your system's recent history. You will immediately see the short-lived processes, cron tasks, and background spikes that standard dashboards miss.

πŸ›‘οΈ Schede di Revisione Redazionale & Statistiche AI β–Ύ
πŸ“° Verifiche Redazionali (100% SOTA)
FactCheckerAgent (Web & Technical Verification) APPROVED
Verified technical flags, physics formulas, and working external links.
GuardianStyleReviewer (Brand & Typography) APPROVED
Enforces Guardian brand color tokens (#052962, #c70000), uppercase kickers, and callout boxes.
EditorialQualityReviewer (Academic Rigor & Depth) APPROVED
Verified >1,500 word academic length, working links, and didactic goal satisfaction.
πŸ“Š Statistiche AI & Token Telemetry
Engine: gemini-3.6-pro
Auth: Google Gemini Ultra OAuth Session (~/.config/antigravity)
Prompt Tokens: 1,353
Completion Tokens: 8,524
Token Totali: 9,877
Costo API: $0.00 (Google Ultra Plan)
← Back to UNIX Command of the Day Archive
MAPPA STORICA πŸ“ Bologna