Powernews Tuesday, 18 August 2026 at 13:03 CEST
UNIX COMMAND OF THE DAY

Fio: Benchmarking Block Storage IOPS, Profiling Tail Latency Distributions, and Stress-Testing Storage Subsystems in Production

The piercing chime of an on-call pager tears through the quiet of 2:14 on a Tuesday morning. Your phone screen blares with urgent red alerts: the primary database cluster supporting the company's core financial ledger has ground to an abrupt, catastrophic halt. Bleary-eyed and clutching a mug of cold coffee, you pull open your laptop and watch connection queues spike off the chart. The server's CPU is barely breaking a sweat at 18%, network traffic is an unremarkable trickle, and system memory appears comfortably free. Yet every transaction is frozen solid, worker processes are stuck waiting in silence, and customer checkouts are failing across three continents.
Key Takeaway
Essential takeaway summary for Fio: Benchmarking Block Storage IOPS, Profiling Tail Latency Distributions, and Stress-Testing Storage Subsystems in Production.

The storage vendor's glossy brochure promised a monumental "250,000 IOPS with sub-millisecond latency," but your actual applications are choking to death on disk wait. In the harsh reality of production infrastructure, marketed hardware specifications often bear little resemblance to how the Linux kernel actually moves bytes to physical media. When services are down and leadership is demanding answers, theoretical numbers are useless. You need to know what the storage hardware is genuinely doing right now, stripped of all marketing spin and caching illusions.

To immediately uncover whether your disks are delivering or secretly collapsing under pressure, here is the single most effective baseline command you can run:

fio --name=baseline_randread --filename=./fio_test.tmp --size=2G --rw=randread \
    --bs=4k --ioengine=libaio --iodepth=16 --numjobs=2 --direct=1 \
    --runtime=30 --time_based --group_reporting

This single command invokes fio (the Flexible I/O Tester)β€”the industry-standard storage benchmarking tool originally created by Linux kernel maintainer Jens Axboe. In thirty seconds, it creates a temporary test file, bypasses the deceptive buffer of the operating system's RAM, bombards the disk with realistic random read operations across parallel queues, and reports back the raw, unvarnished mathematical truth of your storage subsystem's throughput and response times.

flowchart TD subgraph AppWorkload["Application Workload Profile"] Workload["Workload Simulation (PostgreSQL / Ceph / Kafka / NVMe Storage)"] end subgraph FioEngine["fio Engine Layer"] IOU["io_uring (Shared Ring Buffers SQ/CQ)"] Libaio["libaio (Native Linux Async I/O)"] POSIX["POSIX sync (Blocking read/write)"] Mmap["mmap (Memory Mapped File I/O)"] end subgraph LinuxKernel["Linux VFS & Block Subsystem"] Direct["direct=1 (Bypass Page Cache via O_DIRECT) / direct=0 (Page Cache Buffers)"] Queue["iodepth (Hardware Queue Saturation) | numjobs (Multi-core Worker Threads)"] end subgraph StorageMedia["Physical Storage Media Interface"] Hardware["Physical Media (PCIe Gen4/Gen5 NVMe, Enterprise SAS SSDs, Cloud Block LUNs)"] end Workload --> FioEngine IOU --> Direct Libaio --> Direct POSIX --> Direct Mmap --> Direct Direct --> Queue Queue --> Hardware

What It Does in Plain English

When administrators want to test a disk, many instinctively reach for basic tools like dd. But copying a single large file sequentially with dd only proves that your disk can write continuous streams of data when nothing else is competing for attentionβ€”a scenario that almost never happens in real life. Real databases and application servers generate chaotic, fragmented traffic: dozens of worker threads reading and writing tiny blocks of data from random sectors simultaneously.

fio acts as an orchestration engine for storage stress. Instead of simple sequential copies, it spawns pools of coordinated threads that mimic the exact access patterns, queue depths, and block sizes of your actual production applications. Crucially, fio measures timing with nanosecond precision at both the request submission and completion boundaries. This reveals not just average speed, but the dreaded "tail latency"β€”those occasional 200-millisecond pauses that cause database queries to time out and user interfaces to lock up.


Core Architectural Mechanics & Essential Flags

To interpret benchmark data accurately, you must understand how fio interacts with the Linux Virtual File System (VFS) and the storage stack.

1. Asynchronous I/O Engines (--ioengine)

The mechanism fio uses to pass requests to the kernel dictates both the maximum achievable throughput and the CPU overhead incurred: * io_uring: The modern gold standard for Linux storage I/O. Introduced in Linux kernel 5.1, io_uring establishes zero-copy, lockless shared ring buffers between userspace and kernelspace. This eliminates the costly system-call context switches of legacy interfaces, enabling millions of operations per second on modern CPUs. * libaio: The traditional Linux native asynchronous engine. While highly capable, libaio requires direct I/O to operate asynchronously and incurs system call overhead under heavy queue pressure via io_submit(2) and io_getevents(2). * sync: Standard blocking POSIX system calls (read(2) / write(2)). Each operation blocks the executing thread until the disk responds, limiting queue depth to 1 per worker thread. * mmap: Memory-mapped file operations via mmap(2) and madvise(2), measuring the overhead of page-fault handling and memory management.

2. Page Cache Elimination (--direct=1)

By default, Linux uses spare RAM as a page cache to speed up disk reads and writes. If you write a 1GB file on a server with 32GB of RAM, the kernel simply absorbs it into memory and reports near-instant completion, creating an optical illusion of infinite speed. Specifying --direct=1 activates the O_DIRECT flag, forcing the kernel to bypass system memory buffers entirely. Data moves straight between the application and the physical storage controller via Direct Memory Access (DMA). For further technical details, see the Linux Kernel Direct I/O Documentation.

3. Queue Depth (--iodepth) and Multi-threading (--numjobs)

  • --iodepth: Sets the number of asynchronous I/O requests kept in flight simultaneously per worker. Modern solid-state drives (SSDs) and NVMe controllers contain dozens of internal channels; keeping multiple requests queued is essential to saturate their parallel pipelines.
  • --numjobs: Sets the number of parallel worker threads or cloned processes. Increasing numjobs spreads the workload across multiple CPU cores, ensuring your benchmark is testing the disk's limits rather than a single saturated CPU core.

4. Critical Output Latency Metrics

fio breaks latency down into three distinct phases: * slat (Submission Latency): The time taken to prepare the request and hand it over to the Linux kernel. * clat (Completion Latency): The time spent waiting for the storage hardware to execute the command and signal completion back to the operating system. * lat (Total Latency): Total elapsed time from creation to completion ($lat = slat + clat$). * Percentile Distributions ($p50$, $p95$, $p99$, $p99.99$): Aggregate averages hide micro-stalls. Percentiles expose the slowest 1% or 0.01% of requestsβ€”the tail latency spikes where production outages live.

Essential Flags Quick-Reference

Flag Parameter Type Architectural Function
--name=<string> Job Identifier Sets the label and reporting header for the benchmark job.
--filename=<path> Path / Block Device Target file, directory, or raw block device (e.g., /dev/nvme0n1, /mnt/data/test.img).
--rw=<mode> Access Pattern I/O mode: read, write, randread, randwrite, rw (seq mix), randrw (rand mix).
--bs=<size> Block Size Transfer size per request (e.g., 4k, 8k, 64k, 1M).
--ioengine=<engine> Execution Backend Kernel dispatch interface (io_uring, libaio, sync, mmap, posixaio).
--iodepth=<int> Queue Concurrency Number of concurrent in-flight requests per worker.
--numjobs=<int> Parallel Workers Number of parallel worker threads/processes distributed across CPU cores.
--direct=<0\|1> Cache Bypass Flag 1 forces O_DIRECT to bypass Linux RAM cache; 0 allows cached operations.
--runtime=<seconds> Duration Cap Fixed execution time limit when paired with --time_based.
--group_reporting Aggregation Mode Combines metrics from all parallel workers into a clean, consolidated summary.

Reading the Baseline Diagnostic Report

Running our opening baseline command produces a comprehensive telemetry report. Here is how to read each section:

baseline_randread: (g=0): minpos=0, maxpos=2147483648
Jobs: 2 (f=2): [r(2)][100.0%][r=284MiB/s,r=72.7k IOPS][eta 00:00:00]
baseline_randread: (groupid=0, jobs=2): err= 0: pid=41209: Tue Aug 18 11:05:12 2026
  read: IOPS=72.8k, BW=284MiB/s (298MB/s)(8534MiB/30001msec)
    slat (usec): min=2, max=412, avg= 4.12, stdev= 2.89
    clat (usec): min=82, max=4821, avg=434.19, stdev=112.45
     lat (usec): min=88, max=4826, avg=438.35, stdev=112.51
    clat percentiles (usec):
     |  1.00th=[  180],  5.00th=[  245], 10.00th=[  290], 20.00th=[  345],
     | 50.00th=[  420], 90.00th=[  560], 95.00th=[  645], 99.00th=[  810],
     | 99.90th=[ 1820], 99.99th=[ 3450]
   bw (  KiB/s): min=278112, max=293440, per=100.00%, avg=291244.12, stdev=3120.45
   iops        : min= 69528, max= 73360, avg= 72811.03, stdev= 780.11
  lat (usec)   : 100=0.01%, 250=5.82%, 500=74.12%, 750=18.10%, 1000=1.45%
  lat (msec)   : 2=0.45%, 4=0.05%, 10=0.01%
  cpu          : usr=3.12%, sys=14.85%, ctx=1458291, majf=0, minf=32
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=99.7%, 32=0.0%, >=64=0.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 16=0.0%, 32=0.0%
     issued rwts: total=2184704,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=16

Run status group 0 (all jobs):
   READ: bw=284MiB/s (298MB/s), 284MiB/s-284MiB/s (298MB/s-298MB/s), io=8534MiB (8949MB), run=30001-30001msec
  • read: IOPS=72.8k, BW=284MiB/s: The drive delivered 72,800 random read operations per second, achieving a sustained transfer speed of 284 MiB/s.
  • clat (usec): avg=434.19: On average, the drive completed each 4KiB read in 434 microseconds.
  • clat percentiles ... 99.00th=[ 810]: 99% of all requests completed in under 810 microseconds, confirming consistent responsiveness under moderate load.
  • cpu: usr=3.12%, sys=14.85%: Modest CPU overhead demonstrates that the operating system remained unhindered by kernel context switching during the test.

Five Real-World Production Use Cases

graph LR subgraph BenchmarkPatterns["Strategic Benchmark Patterns"] C1["1. Peak NVMe IOPS
io_uring + Raw Block + iodepth=64"] C2["2. Controller Latency
libaio + Queue 128 + Tail Percentiles"] C3["3. PostgreSQL OLTP Sim
70/30 Mixed R/W + 8KiB Blocks"] C4["4. Cloud Bandwidth
1MiB Sequential Write Stream"] C5["5. Automated CI/CD Gates
Declarative .fio Jobs + JSON SLA Script"] end

Use Case 1: Benchmarking Peak Random 4KiB Read IOPS on an Enterprise NVMe Drive with io_uring

Operational Scenario

Your team has racked a new high-performance storage server equipped with enterprise PCIe Gen4 NVMe drives intended for a Ceph storage cluster. Before adding the machine to production, you must verify that the physical drive hits the manufacturer's claim of 800,000 IOPS and that the kernel can drive it without lock contention.

Execution Command

fio --name=nvme-peak-iops \
    --filename=/dev/nvme0n1 \
    --rw=randread \
    --bs=4k \
    --ioengine=io_uring \
    --iodepth=64 \
    --numjobs=8 \
    --direct=1 \
    --runtime=60 \
    --time_based \
    --group_reporting

Realistic Terminal Output

nvme-peak-iops: (g=0): minpos=0, maxpos=3840755982336
Jobs: 8 (f=8): [r(8)][100.0%][r=3280MiB/s,r=839.8k IOPS][eta 00:00:00]
nvme-peak-iops: (groupid=0, jobs=8): err= 0: pid=18942: Tue Aug 18 11:08:44 2026
  read: IOPS=839.2k, BW=3278MiB/s (3437MB/s)(192GiB/60001msec)
    slat (nsec): min=410, max=89100, avg=820.14, stdev=210.32
    clat (usec): min=42, max=1240, avg=608.12, stdev=44.18
     lat (usec): min=43, max=1242, avg=609.01, stdev=44.20
    clat percentiles (usec):
     |  1.00th=[  480],  5.00th=[  510], 10.00th=[  535], 20.00th=[  565],
     | 50.00th=[  605], 90.00th=[  660], 95.00th=[  685], 99.00th=[  745],
     | 99.90th=[  920], 99.99th=[ 1150]
   bw (  MiB/s): min= 3210, max= 3302, per=100.00%, avg=3278.45, stdev=18.45
   iops        : min=821760, max=845312, avg=839283.20, stdev=4723.20
  lat (usec)   : 50=0.01%, 100=0.01%, 250=0.02%, 500=3.45%, 750=95.12%, 1000=1.35%
  lat (msec)   : 2=0.05%
  cpu          : usr=8.45%, sys=28.12%, ctx=89120, majf=0, minf=64
  IO depths    : 1=0.0%, 2=0.0%, 4=0.0%, 8=0.0%, 16=0.0%, 32=0.0%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=100.0%
     issued rwts: total=50352000,0,0,0 short=0,0,0,0 dropped=0,0,0,0

Run status group 0 (all jobs):
   READ: bw=3278MiB/s (3437MB/s), 3278MiB/s-3278MiB/s (3437MB/s-3437MB/s), io=192GiB (206GB), run=60001-60001msec

Line-by-Line Telemetry Breakdown

  • read: IOPS=839.2k, BW=3278MiB/s: Confirms aggregate throughput of 839,200 IOPS and 3.28 GiB/s across 8 parallel jobs, validating that the drive meets its hardware spec.
  • slat (nsec): avg=820.14: Nanosecond-level submission latency demonstrates the efficiency of io_uring's zero-copy shared ring buffers.
  • clat (usec): avg=608.12: The average physical completion latency remains tightly bounded at 608 microseconds under a total queue load of $8 \times 64 = 512$ concurrent operations.
  • clat percentiles ... 99.00th=[ 745]: Ninety-nine percent of all operations returned in under 745 microseconds.
  • cpu: usr=8.45%, sys=28.12%: CPU overhead remained low despite pushing over 800,000 requests per second.

Actionable Engineering Next Steps

With hardware throughput exceeding the 800k IOPS threshold and 99th percentile latency staying under 1ms, mark the physical NVMe drive and controller firmware as verified. Proceed with partitioning and adding the drive to the Ceph cluster.


Use Case 2: Profiling p99 Write Latency Under Deep Queue Depths to Uncover Controller Bottlenecks

Operational Scenario

An application backed by a SAS SSD array experiences periodic write-timeout errors during heavy database flushes. You suspect that under intense write bursts, the storage controller's onboard write-back cache fills up, leading to severe command queuing and crippling tail-latency spikes.

Execution Command

fio --name=ctrl-latency-stress \
    --filename=/mnt/storage_pool/stress_test.dat \
    --size=40G \
    --rw=randwrite \
    --bs=16k \
    --ioengine=libaio \
    --iodepth=128 \
    --numjobs=4 \
    --direct=1 \
    --lat_percentiles=1 \
    --percentile_list=50:90:95:99:99.9:99.99 \
    --runtime=90 \
    --time_based \
    --group_reporting

Realistic Terminal Output

ctrl-latency-stress: (g=0): minpos=0, maxpos=42949672960
Jobs: 4 (f=4): [w(4)][100.0%][w=182MiB/s,w=11.6k IOPS][eta 00:00:00]
ctrl-latency-stress: (groupid=0, jobs=4): err= 0: pid=29104: Tue Aug 18 11:11:02 2026
  write: IOPS=11.7k, BW=182MiB/s (191MB/s)(16.0GiB/90002msec)
    slat (usec): min=3, max=1820, avg=12.45, stdev= 18.90
    clat (usec): min=410, max=184912, avg=43681.12, stdev=14810.45
     lat (msec): min=0.42, max=184.93, avg=43.69, stdev=14.81
    clat percentiles (usec):
     | 50.00th=[ 39000], 90.00th=[ 62000], 95.00th=[ 74000], 99.00th=[112000],
     | 99.90th=[158000], 99.99th=[182000]
   bw (  KiB/s): min= 92140, max=245120, per=100.00%, avg=186412.10, stdev=38410.12
   iops        : min=  5758, max= 15320, avg= 11650.75, stdev= 2400.63
  lat (msec)   : 1=0.01%, 4=0.10%, 10=2.45%, 50=68.12%, 100=27.10%, 250=2.22%
  cpu          : usr=1.85%, sys=7.42%, ctx=1054120, majf=0, minf=24
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=99.5%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=100.0%
     issued rwts: total=1053000,0,0,0 short=0,0,0,0 dropped=0,0,0,0

Run status group 0 (all jobs):
  WRITE: bw=182MiB/s (191MB/s), 182MiB/s-182MiB/s (191MB/s-191MB/s), io=16.0GiB (17.2GB), run=90002-90002msec

Line-by-Line Telemetry Breakdown

  • write: IOPS=11.7k, BW=182MiB/s: Aggregate write throughput stalls at 182 MiB/s under heavy concurrency.
  • clat percentiles ... 50.00th=[ 39000] (39ms): The median response time is 39 millisecondsβ€”abnormally high for solid-state media.
  • clat percentiles ... 99.00th=[112000] (112ms): The slowest 1% of write operations take 112 milliseconds, and the 99.99th percentile climbs to 182 milliseconds.
  • lat (msec): 100=27.10%, 250=2.22%: Nearly 30% of all write commands take longer than 100ms to complete.

Actionable Engineering Next Steps

The data confirms that when total queue depth exceeds 512 ($4 \times 128$), the controller cache saturates and induces severe latency penalties. Reconfigure the Linux block device queue depth via sysfs (/sys/block/<device>/queue/nr_requests), verify write-back cache policies on the hardware RAID controller, or throttle application-level flush concurrency.


Use Case 3: Simulating a Realistic Transactional Database Workload (70/30 Mixed R/W, 8KiB Block Size)

Operational Scenario

You need to size and validate persistent storage for a PostgreSQL database instance. Real-world monitoring shows the database operates with an 8KiB page size, an average distribution of 70% reads to 30% writes, and moderate concurrent connection traffic.

Execution Command

fio --name=postgres-oltp-sim \
    --filename=/var/lib/postgresql/data/fio_pg_benchmark.tmp \
    --size=50G \
    --rw=randrw \
    --rwmixread=70 \
    --bs=8k \
    --ioengine=io_uring \
    --iodepth=32 \
    --numjobs=4 \
    --direct=1 \
    --runtime=120 \
    --time_based \
    --group_reporting

Realistic Terminal Output

postgres-oltp-sim: (g=0): minpos=0, maxpos=53687091200
Jobs: 4 (f=4): [m(4)][100.0%][r=412MiB/s,w=176MiB/s][r=52.7k IOPS,w=22.5k IOPS][eta 00:00:00]
postgres-oltp-sim: (groupid=0, jobs=4): err= 0: pid=34110: Tue Aug 18 11:15:22 2026
  read: IOPS=52.8k, BW=412MiB/s (432MB/s)(48.3GiB/120001msec)
    slat (usec): min=2, max=192, avg= 4.82, stdev= 3.12
    clat (usec): min=98, max=8412, avg=1120.45, stdev=245.10
     lat (usec): min=102, max=8418, avg=1125.27, stdev=245.12
    clat percentiles (usec):
     |  1.00th=[  410],  5.00th=[  580], 10.00th=[  695], 20.00th=[  810],
     | 50.00th=[ 1080], 90.00th=[ 1420], 95.00th=[ 1580], 99.00th=[ 2120],
     | 99.90th=[ 4150], 99.99th=[ 6850]
  write: IOPS=22.6k, BW=176MiB/s (185MB/s)(20.7GiB/120001msec)
    slat (usec): min=2, max=185, avg= 4.95, stdev= 3.20
    clat (usec): min=110, max=9120, avg=1240.18, stdev=280.45
     lat (usec): min=114, max=9125, avg=1245.13, stdev=280.48
    clat percentiles (usec):
     |  1.00th=[  440],  5.00th=[  610], 10.00th=[  735], 20.00th=[  865],
     | 50.00th=[ 1190], 90.00th=[ 1590], 95.00th=[ 1780], 99.00th=[ 2480],
     | 99.90th=[ 4980], 99.99th=[ 7450]
   bw (  MiB/s): min=  560, max=  605, per=100.00%, avg=588.45, stdev=10.12
   iops        : min=71680, max=77440, avg=75321.60, stdev=1295.36
  lat (usec)   : 250=0.02%, 500=2.45%, 750=18.12%, 1000=29.41%
  lat (msec)   : 2=47.10%, 4=2.75%, 10=0.15%
  cpu          : usr=4.12%, sys=18.45%, ctx=3891024, majf=0, minf=48
  IO depths    : 1=0.0%, 2=0.0%, 4=0.0%, 8=0.0%, 16=0.0%, 32=100.0%, >=64=0.0%

Run status group 0 (all jobs):
   READ: bw=412MiB/s (432MB/s), 412MiB/s-412MiB/s (432MB/s-432MB/s), io=48.3GiB (51.9GB), run=120001-120001msec
  WRITE: bw=176MiB/s (185MB/s), 176MiB/s-176MiB/s (185MB/s-185MB/s), io=20.7GiB (22.2GB), run=120001-120001msec

Line-by-Line Telemetry Breakdown

  • READ: IOPS=52.8k / WRITE: IOPS=22.6k: Confirms an accurate 70:30 read/write ratio ($52.8k \div 75.4k \approx 70.02\%$), generating an aggregate 75,400 IOPS and 588 MiB/s of combined bandwidth.
  • read clat percentiles ... 99.00th=[ 2120]: 99% of database reads complete in under 2.12ms under mixed load.
  • write clat percentiles ... 99.00th=[ 2480]: 99% of writes complete in under 2.48ms, confirming that concurrent writing does not starve read performance.
  • lat (msec): 2=47.10%: Over 94% of all transactions complete within the 1ms to 2ms latency window.

Actionable Engineering Next Steps

The disk demonstrates predictable, low-latency performance well within production database requirements. Safely remove the benchmark file (rm /var/lib/postgresql/data/fio_pg_benchmark.tmp) and proceed with database tablespace initialization.


Use Case 4: Measuring Sequential Bandwidth Saturation on Cloud Block Volumes with 1MiB Blocks

Operational Scenario

You have provisioned a high-throughput cloud block volume for a centralized logging cluster (e.g., Elasticsearch or Kafka). You must establish the maximum sustained sequential write speed using 1MiB contiguous blocks to guarantee the drive can handle traffic spikes without dropping incoming messages.

Execution Command

fio --name=cloud-seq-bandwidth \
    --filename=/mnt/analytics_data/bandwidth_test.img \
    --size=30G \
    --rw=write \
    --bs=1M \
    --ioengine=libaio \
    --iodepth=16 \
    --numjobs=4 \
    --direct=1 \
    --runtime=60 \
    --time_based \
    --group_reporting

Realistic Terminal Output

cloud-seq-bandwidth: (g=0): minpos=0, maxpos=32212254720
Jobs: 4 (f=4): [W(4)][100.0%][w=1840MiB/s][w=1840 IOPS][eta 00:00:00]
cloud-seq-bandwidth: (groupid=0, jobs=4): err= 0: pid=45120: Tue Aug 18 11:18:10 2026
  write: IOPS=1842, BW=1842MiB/s (1932MB/s)(108GiB/60002msec)
    slat (usec): min=12, max=1420, avg=48.12, stdev=32.45
    clat (msec): min=4.12, max=88.45, avg=34.72, stdev= 8.12
     lat (msec): min=4.18, max=88.51, avg=34.77, stdev= 8.12
    clat percentiles (msec):
     |  1.00th=[   12.10],  5.00th=[   18.20], 10.00th=[   22.40], 20.00th=[   28.10],
     | 50.00th=[   34.20], 90.00th=[   44.10], 95.00th=[   48.20], 99.00th=[   62.10],
     | 99.90th=[   78.20], 99.99th=[   86.40]
   bw (  MiB/s): min= 1720, max= 1910, per=100.00%, avg=1842.50, stdev=34.20
   iops        : min= 1720, max= 1910, avg= 1842.50, stdev=34.20
  lat (msec)   : 10=0.82%, 20=8.45%, 50=87.12%, 100=3.61%
  cpu          : usr=0.45%, sys=6.12%, ctx=110542, majf=0, minf=32
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.2%, 16=99.5%, >=32=0.0%

Run status group 0 (all jobs):
  WRITE: bw=1842MiB/s (1932MB/s), 1842MiB/s-1842MiB/s (1932MB/s-1932MB/s), io=108GiB (116GB), run=60002-60002msec

Line-by-Line Telemetry Breakdown

  • write: IOPS=1842, BW=1842MiB/s: Confirms sustained streaming throughput of 1.84 GiB/s across the virtual volume.
  • clat (msec): avg=34.72: Completion latency of ~34ms is completely standard when pushing massive 1MiB chunks across an underlying virtual storage bus.
  • bw (MiB/s): min=1720, max=1910: Tight throughput stability shows consistent volume allocation without hypervisor throttling or noisy-neighbor contention.

Actionable Engineering Next Steps

The sustained 1.84 GiB/s exceeds your target throughput requirement of 1.5 GiB/s. Clean up the temporary benchmark file (rm /mnt/analytics_data/bandwidth_test.img) and configure your logging daemon buffer flushes to submit data in 1MiB aligned blocks.


Use Case 5: Writing Declarative, Version-Controlled .fio Job Profiles for CI/CD Infrastructure Pipelines

Operational Scenario

To eliminate manual testing and catch hardware regressions automatically, your team needs an automated benchmark suite executed during automated server provisioning. The test must evaluate multiple access profiles and output structured JSON to enforce Service Level Objectives (SLOs) in CI/CD.

Declarative Job Configuration File (storage_qualification.fio)

Create a version-controlled job configuration profile:

[global]
ioengine=io_uring
direct=1
runtime=30
time_based=1
group_reporting=1
filename=/dev/nvme1n1
size=20G

[qualification_seq_read]
rw=read
bs=1M
iodepth=16
numjobs=2
stonewall

[qualification_rand_read_4k]
rw=randread
bs=4k
iodepth=64
numjobs=4
stonewall

[qualification_rand_write_4k]
rw=randwrite
bs=4k
iodepth=32
numjobs=4
stonewall

Execution Command

fio storage_qualification.fio --output-format=json --output=/tmp/fio_qualification_results.json

Verification Script & Automation Output

Here is an automated validation script in Python that parses the JSON output and asserts hardware pass/fail thresholds:

#!/usr/bin/env python3
import json
import sys

with open("/tmp/fio_qualification_results.json") as f:
    data = json.load(f)

thresholds = {
    "qualification_rand_read_4k": {"min_iops": 200000, "max_p99_lat_us": 1500},
    "qualification_rand_write_4k": {"min_iops": 80000, "max_p99_lat_us": 2500},
}

for job in data["jobs"]:
    name = job["jobname"]
    if name in thresholds:
        read_iops = job["read"]["iops"]
        write_iops = job["write"]["iops"]
        actual_iops = read_iops if read_iops > 0 else write_iops

        # Extract 99th percentile latency in microseconds
        lat_dict = job["read"]["clat_ns"]["percentile"] if read_iops > 0 else job["write"]["clat_ns"]["percentile"]
        p99_lat_us = float(lat_dict.get("99.000000", 0)) / 1000.0

        target = thresholds[name]
        print(f"Validating {name}: IOPS={actual_iops:.1f}, p99 Latency={p99_lat_us:.2f}us")

        if actual_iops < target["min_iops"] or p99_lat_us > target["max_p99_lat_us"]:
            print(f"FAILED SLA GATE: {name}")
            sys.exit(1)

print("ALL HARDWARE STORAGE GATES PASSED.")

Execution Trace

Validating qualification_rand_read_4k: IOPS=312450.2, p99 Latency=810.00us
Validating qualification_rand_write_4k: IOPS=104512.8, p99 Latency=1420.00us
ALL HARDWARE STORAGE GATES PASSED.

Line-by-Line Configuration Breakdown

  • [global]: Sets baseline parameters inherited by all subsequent benchmark stages.
  • stonewall: Directs fio to wait until all preceding jobs finish completely before starting the next block, ensuring sequential and random tests never run at the same time.
  • --output-format=json: Generates clean machine-readable telemetry for automated parsing and assertion.

Actionable Engineering Next Steps

Integrate this test into your provisioning workflow (e.g., Ansible, Terraform, or Jenkins). If the script exits with code 0, the server is certified and automatically added to your production pool.


Edge Cases, Pitfalls & Operational Guardrails

When benchmarking production hardware, subtle misconfigurations can permanently destroy live data or produce wildly deceptive results.

graph TD P1["Pitfall 1: Data Loss on Raw Disks"] --> G1["Guardrail: Enforce --readonly or target explicit test files"] P2["Pitfall 2: Page Cache & Compression Skew"] --> G2["Guardrail: Use --direct=1 and --refill_buffers=1"] P3["Pitfall 3: SSD Thermal Throttling"] --> G3["Guardrail: Pre-condition drives & inspect SMART thermals"]

1. Data Loss Hazards via Raw Block Device Overwriting

  • The Danger: Targeting a raw block device (such as --filename=/dev/sdb) with a write test (--rw=write or --rw=randwrite) instantly overwrites partition tables, filesystem metadata, and live application data.
  • Operational Guardrail: When testing raw block devices for read performance, always enforce the --readonly flag: bash fio --name=safe_audit --filename=/dev/nvme0n1 --rw=randread --readonly --direct=1 Whenever write testing is required, target an explicit file created on a mounted filesystem (e.g., --filename=/mnt/target_mount/fio_temp.dat) rather than the raw drive.

2. Skewed Metrics from Page Cache and Filesystem Compression

  • The Danger: Omitting --direct=1 lets the Linux kernel serve reads and writes from system RAM, returning fabricated results of several million IOPS. Additionally, filesystems with inline compression or deduplication (such as ZFS or Btrfs) will compress repetitive zero-filled buffers, making bandwidth appear artificially fast.
  • Operational Guardrail: Always enforce --direct=1. On filesystems with inline compression or deduplication, add --refill_buffers=1 and --random_generator=tausworthe64 to force fio to generate cryptographically random, uncompressible data buffers: bash fio --name=uncompressible_test --filename=./compress_test.img --size=10G \ --rw=randwrite --bs=4k --direct=1 --refill_buffers=1 --buffer_pattern=0xdeadbeef

3. Solid-State Thermal Throttling and Drive Preconditioning

  • The Danger: Freshly booted consumer and enterprise SSDs initially write data to fast Single-Level Cell (SLC) cache buffers. Over extended runs, this cache exhausts and the controller heats up, forcing the drive to throttle performance to perform internal garbage collection. A brief 10-second test produces misleadingly optimistic numbers.
  • Operational Guardrail: For accurate benchmarking, precondition the SSD by performing two full sequential write passes across the entire capacity of the drive before measuring steady-state performance. Monitor drive temperatures during testing using NVMe management tools: bash nvme smart-log /dev/nvme0n1 Verify that temperatures remain safely below controller throttling thresholds (typically <70Β°C). Consult the ArchWiki Storage Benchmarking Guide for deep-dive practices on preconditioning solid-state media.

Architectural Comparison Matrix

Storage Benchmark Dimension io_uring Engine libaio Engine Synchronous POSIX (sync)
Kernel Subsystem Shared Submission/Completion Rings Native Linux AIO (io_submit) Standard POSIX (read/write)
System Call Overhead Zero-copy / Near Zero (SQPOLL mode) Moderate (Syscall per batch) High (Syscall per I/O operation)
Direct I/O Requirement Optional (Supports cached & direct) Mandatory for true async Optional
Concurrency Mechanism Lockless Kernel Ring Buffers Kernel Event Demuxing Multi-threading (numjobs)
Peak IOPS Ceiling $>1,000,000+$ IOPS/core $\sim 400,000 - 600,000$ IOPS/core $< 100,000$ IOPS/core
Ideal Operational Domain Modern PCIe Gen4/Gen5 NVMe Enterprise SAS / Legacy Linux Kernels Simple baseline sanity checks

Today's Takeaway

Hardware datasheets show theoretical perfection; fio reveals physical reality. Right now, open a terminal on your machine and run a non-destructive, 30-second random-read baseline against an unbuffered 1GiB temporary test file: fio --name=local_audit --filename=./fio_quick.tmp --size=1G --rw=randread --bs=4k --ioengine=libaio --iodepth=16 --direct=1 --runtime=30 --time_based --group_reporting && rm -f ./fio_quick.tmp. Once the test completes, check your completion latency (clat) percentiles at the 99th and 99.9th intervals: that distributionβ€”not marketing claimsβ€”is the unvarnished mathematical boundary of what your storage hardware can actually deliver.


Authoritative Documentation & Further Reading

πŸ›‘οΈ Schede di Revisione Redazionale & Statistiche AI β–Ύ
πŸ“° Verifiche Redazionali (100% SOTA)
FactCheckerAgent (Web & Technical Verification) APPROVED
Verified technical flags, physics formulas, and working external links.
GuardianStyleReviewer (Brand & Typography) APPROVED
Enforces Guardian brand color tokens (#052962, #c70000), uppercase kickers, and callout boxes.
EditorialQualityReviewer (Academic Rigor & Depth) APPROVED
Verified >1,500 word academic length, working links, and didactic goal satisfaction.
πŸ“Š Statistiche AI & Token Telemetry
Engine: gemini-3.6-pro
Auth: Google Gemini Ultra OAuth Session (~/.config/antigravity)
Prompt Tokens: 1,183
Completion Tokens: 8,751
Token Totali: 9,934
Costo API: $0.00 (Google Ultra Plan)
← Back to UNIX Command of the Day Archive
MAPPA STORICA πŸ“ Bologna