Fallocate: Preallocating Contiguous Storage Blocks, Punching Filesystem Holes, and Provisioning Instant Swap Volumes in Production
In the incident response channel, the tension is palpable. Under an unexpected midnight spike in write traffic, the database nodes ran out of preallocated disk space and began gasping for storage blocks. Seeking an emergency fix, an exhausted engineer tries to provision an 8GB swap file and expand database tables using the venerable dd command. But instead of providing relief, the server locks up completelyโthe storage controllers choked as gigabytes of physical zeroes are painstakingly written to disk, one agonizing byte at a time.
There is a dramatically faster, cleaner, and more resilient way out of this trap. Rather than forcing the operating system to push endless streams of physical zero bytes through saturated storage channels, modern Linux systems allow administrators to manipulate the filesystem's structural map directly in memory. That tool is fallocate.
To instantly carve out a guaranteed, contiguous 1GB space on disk without moving a single byte of zero-payload through your storage pipeline, all it takes is one concise command:
fallocate -l 1G /var/tmp/emergency_reserve.dat
In less than a millisecond, the operating system updates its internal ledger, locks in the physical storage blocks, and hands back the shell prompt. The space is reserved, safe from out-of-disk crashes, and ready for use before your coffee has even begun to brew.
1. What It Does in Plain English
The fallocate(1) command-line utility and its underlying kernel system call, fallocate(2), provide a method to instantly reserve, release, or reshape disk space for files.
When you create a file the traditional way by writing data or zeroes to it, your computer physically alters the magnetic charges or flash memory cells on your storage drive. This takes time, consumes memory bandwidth, and causes hardware wear.
By contrast, fallocate bypasses physical writes altogether. It communicates directly with modern filesystem architecturesโsuch as ext4 and XFSโinstructing them to mark a specific batch of storage blocks as reserved for your file. To the operating system and any running applications, the file appears immediately at its full requested size, with its disk space fully guaranteed. Furthermore, fallocate allows you to carve out unwanted chunks from the middle of existing files (known as "hole punching") or trim the beginning of large log files in constant time, without ever having to copy or rewrite the surrounding data.
2. Under the Hood: Extents, Zero-Fill, and Sparse Allocations
To understand why fallocate is so fundamentally transformative for system performance, it helps to examine the three distinct ways Linux can assign space on a drive:
The Pitfalls of dd (Physical Zeroing)
Historically, administrators provisioned large files by copying zeroes from /dev/zero using dd. While dependable, this approach treats the filesystem like a dumb bucket. The kernel must allocate memory pages, fill them with zeroes, queue them for disk write-out, and physically program flash cells or magnetic platters. For an 800GB database segment or virtual disk image, this brute-force process ties up storage channels for several minutes or hours while causing unnecessary physical write wear on SSDs.
The Illusion of truncate(2) (Sparse Files)
To avoid waiting for dd, some tools use truncate to instantly create a file of any arbitrary size. However, truncate merely alters the logical size recorded in the file's index node (inode). No physical disk blocks are actually assigned. The result is a "sparse file"โa phantom structure that takes up virtually zero actual bytes on disk.
The danger comes later: as your application writes actual data into this sparse file, the operating system must scramble to find and allocate physical blocks in real time. If the physical drive fills up in the meantime, the application suddenly crashes with an out-of-space (ENOSPC) error, leaving administrators to untangle corrupt data.
The Elegance of fallocate(2) (Unwritten Extents)
Modern enterprise filesystems organize storage into extentsโcontinuous stretches of physical disk blocks defined by a starting block and a length. Crucially, modern extent headers support a special flag bit: UNWRITTEN.
When you execute fallocate, the filesystem driver performs a swift four-step administrative transaction:
1. It searches its free-space index to find unbroken ranges of available physical blocks.
2. It adds an entry to the file's extent tree linking the file to those physical blocks.
3. It marks those extents with the UNWRITTEN flag.
4. It updates the filesystem's allocation ledger without writing a single payload byte to the physical medium.
If an application reads from an unwritten extent, the Linux kernel intercepts the request and instantly serves clean zeroes directly from system memory without bothering the physical disk. When the application later writes genuine data, the drive writes only that specific incoming payload, and the filesystem clears the UNWRITTEN flag. You obtain instant allocation, guaranteed storage reservation, zero drive wear, and optimal contiguous layout in a single stroke.
3. Core Flags and Quick-Start Reference
The fallocate command-line utility exposes these low-level kernel superpowers through a clean set of flags:
| Flag | Long Flag | Syscall Equivalent | Operational Summary |
|---|---|---|---|
-l |
--length <bytes> |
(size argument) |
Specifies the total byte length for allocation or deallocation (supports standard suffixes like K, M, G, T). |
-o |
--offset <bytes> |
(offset argument) |
Sets the starting byte position within the file (defaults to 0). |
-n |
--keep-size |
FALLOC_FL_KEEP_SIZE |
Reserves physical disk blocks beyond the current end of the file without changing the visible file size. |
-p |
--punch-hole |
FALLOC_FL_PUNCH_HOLE |
Releases physical blocks in the chosen range back to the free pool, leaving a sparse hole (requires -n). |
-z |
--zero-range |
FALLOC_FL_ZERO_RANGE |
Sets a range to zero in metadata and guarantees physical space allocation if not already present. |
-c |
--collapse-range |
FALLOC_FL_COLLAPSE_RANGE |
Removes a range from within a file and shifts subsequent data downward without leaving an empty hole. |
-i |
--insert-range |
FALLOC_FL_INSERT_RANGE |
Inserts a range of unallocated space inside a file, pushing existing data upward. |
-d |
--dig-holes |
(User-space scan) | Scans a file for contiguous zero-filled blocks and automatically punches holes to free up space. |
Quick Start: Reserving 1GiB Instantly
To allocate a 1GiB file on disk and verify its physical allocation:
fallocate -l 1G /var/tmp/benchmark_reserve.dat
ls -lhs /var/tmp/benchmark_reserve.dat
stat -c "Logical Size: %s bytes | Allocated 512B Blocks: %b | IO Block: %o" /var/tmp/benchmark_reserve.dat
1.0G -rw-r--r-- 1 root root 1.0G Aug 18 10:14 /var/tmp/benchmark_reserve.dat
Logical Size: 1073741824 bytes | Allocated 512B Blocks: 2097152 | IO Block: 4096
The output confirms that the file has a logical size of exactly 1,073,741,824 bytes (1GiB) and that the kernel has physically committed 2,097,152 blocks of 512 bytes eachโcompleted in less than a millisecond.
4. Five Real-World Enterprise Production Use Cases
Use Case 1: Instant 8GiB Swap File Provisioning on NVMe Without Write Wear
Operational Context
Under severe memory pressure, a production Kubernetes node begins terminating mission-critical pods. The server needs an immediate 8GiB swap buffer to stabilize kernel memory queues. Using dd to generate an 8GiB file would saturate the NVMe controller with zero writes and waste precious minutes. fallocate allows the instant creation of a contiguous, ready-to-format swap container.
# 1. Allocate an 8GiB unwritten contiguous file instantly
fallocate -l 8G /swapfile
# 2. Enforce strict POSIX permissions (mandatory for Linux swap)
chmod 0600 /swapfile
# 3. Format the preallocated container as a Linux swap area
mkswap /swapfile
# 4. Activate the swap file into the virtual memory subsystem
swapon /swapfile
# 5. Verify the active swap allocation and physical extent continuity
swapon --show
filefrag -v /swapfile | head -n 12
Setting up swapspace version 1, size = 8 GiB (8589930496 bytes)
no label, UUID=a4c9f1e8-782a-4c22-b5e1-89d20c5d0124
NAME TYPE SIZE USED PRIO
/swapfile file 8G 0B -2
Filesystem type is: ef53
File size of /swapfile is 8589930496 (2097152 blocks of 4096 bytes)
ext: logical_offset: physical_offset: length: expected: flags:
0: 0.. 262143: 1048576.. 1310719: 262144:
1: 262144.. 524287: 1310720.. 1572863: 262144:
2: 524288.. 1048575: 1572864.. 2097151: 524288:
3: 1048576.. 2097151: 2097152.. 3145727: 1048576:
/swapfile: 4 extents found
Line-by-Line Technical Analysis
Setting up swapspace...: Themkswaptool writes a single 4KB header signature at the start of the file without needing to touch the remaining 8GB of disk space.swapon --show: Confirms that the Linux virtual memory manager has adopted the 8GB file at standard swap priority (-2).filefrag -v: Verifies the low-level layout. The file is mapped across only four contiguous disk extents rather than tens of thousands of fragmented blocks, ensuring smooth read/write performance.
What the Administrator Does Next
Make the swap space persistent across system reboots by adding /swapfile none swap defaults 0 0 to /etc/fstab, and tune /proc/sys/vm/swappiness to balance memory usage according to the ArchWiki Swap Documentation.
Use Case 2: Preallocating Database WAL and Datafile Segments to Eliminate Latency Spikes
Operational Context
High-throughput transactional databases such as PostgreSQL and MySQL write updates sequentially to Write-Ahead Logs (WAL) before saving changes to tables. When a WAL log fills up, the engine must open a new segment. If the filesystem is forced to allocate disk blocks dynamically during transaction commits, database queries experience sudden latency spikes while waiting for filesystem metadata locks. Preallocating WAL segments ahead of time eliminates this bottleneck completely.
# 1. Preallocate a 1GiB WAL segment on an XFS-mounted database volume
fallocate -l 1G /var/lib/postgresql/16/main/pg_wal/000000010000000000000001.prealloc
# 2. Inspect the low-level extent mapping via XFS diagnostics
xfs_bmap -vp /var/lib/postgresql/16/main/pg_wal/000000010000000000000001.prealloc
/var/lib/postgresql/16/main/pg_wal/000000010000000000000001.prealloc:
EXT: FILE-OFFSET BLOCK-RANGE AG AG-OFFSET TOTAL FLAGS
0: [0..2097151]: 16777216..18874367 2 (1048576..3145727) 2097152 10000
Line-by-Line Technical Analysis
xfs_bmap -vp: Examines how the file is mapped across the storage volume's internal Allocation Groups (AG).EXT 0: [0..2097151]: All 1GiB ($2,097,152$ sectors of 512 bytes) sits in a single, unbroken block range on disk.FLAGS 10000: The leading bit indicatesXFS_BMAPI_PREALLOC(an unwritten extent). The physical space is reserved; database writes will never fail due to missing space, and block allocation delays drop to zero.
What the Administrator Does Next
Configure the database management system (such as enabling wal_recycle = on and wal_init_zero = off in PostgreSQL) to take direct advantage of native unwritten extent allocation in underlying filesystems.
Use Case 3: Punching Holes in Virtual Machine Disk Images to Reclaim Storage
Operational Context
A virtualization host runs several virtual machines using raw disk image files. Over months of routine operation, guest operating systems create and delete temporary files. Even after a guest OS deletes data, the host's raw disk image continues to hold onto those physical storage blocks on the physical drive. Using fallocate hole punching, the host can release those unneeded blocks back to the free storage pool while keeping the virtual machine intact.
# 1. Check apparent size vs actual physical disk usage before optimization
ls -lhs /var/lib/libvirt/images/production_worker_db.raw
du -h --apparent-size /var/lib/libvirt/images/production_worker_db.raw
# 2. Punch a 20GiB hole in an unreferenced offset (e.g., from offset 10GiB to 30GiB)
fallocate -p -o 10G -l 20G /var/lib/libvirt/images/production_worker_db.raw
# 3. Alternatively, scan and reclaim all unallocated zero-filled blocks across the file
fallocate -d /var/lib/libvirt/images/production_worker_db.raw
# 4. Verify the reclaimed physical capacity
ls -lhs /var/lib/libvirt/images/production_worker_db.raw
# Initial state:
50G -rw-r--r-- 1 qemu qemu 50G Aug 18 09:30 /var/lib/libvirt/images/production_worker_db.raw
50G /var/lib/libvirt/images/production_worker_db.raw
# State after executing hole punching (-d / -p):
18G -rw-r--r-- 1 qemu qemu 50G Aug 18 10:22 /var/lib/libvirt/images/production_worker_db.raw
Line-by-Line Technical Analysis
fallocate -p -o 10G -l 20G: Issues theFALLOC_FL_PUNCH_HOLEinstruction to free the physical blocks between the 10GB and 30GB marks, updating the filesystem extent tree.fallocate -d: The--dig-holesflag scans the target file for ranges of zero bytes and punches holes through them automatically.18G ... 50G: Thels -lhsoutput shows the difference between actual disk usage (18GB on the left) and the virtual file size seen by the guest (50GB on the right). A total of 32GB of physical drive space is instantly recovered.
What the Administrator Does Next
Schedule regular fstrim.timer jobs inside the guest virtual machines or configure the VM storage controller with the discard mount option so that guest file deletions automatically release space on the host.
Use Case 4: In-Place Collapsing and Zeroing of Circular Telemetry Buffers
Operational Context
A telemetry daemon continuously writes raw system metrics to a preallocated binary buffer file. To prevent the log from expanding indefinitely, the daemon must discard the oldest 4GB of data from the beginning of the file while preserving the newest 12GB. In the past, this required copying 12GB of data into a temporary file and renaming itโconsuming heavy disk I/O. With --collapse-range, the filesystem snips out the old section and shifts the remaining data down instantly in metadata.
# 1. Verify initial file dimensions and extent structure
stat -c "Size: %s bytes | Blocks: %b" /var/log/telemetry/flight_recorder.bin
# 2. Collapse (remove) the first 4GiB from offset 0, shifting subsequent extents down
fallocate -c -o 0 -l 4G /var/log/telemetry/flight_recorder.bin
# 3. Zero out a 512MiB corrupted block range within the file without deallocating
fallocate -z -o 2G -l 512M /var/log/telemetry/flight_recorder.bin
# 4. Verify the updated dimensions and metadata state
stat -c "Size: %s bytes | Blocks: %b" /var/log/telemetry/flight_recorder.bin
Size: 17179869184 bytes | Blocks: 33554432
Size: 12884901888 bytes | Blocks: 25165824
Line-by-Line Technical Analysis
fallocate -c -o 0 -l 4G: PassesFALLOC_FL_COLLAPSE_RANGEto the kernel. The extents between 0 and 4GB are pruned from the inode index, and all following extents are shifted down by 4GB in offset space without copying any actual data payload.fallocate -z -o 2G -l 512M: AppliesFALLOC_FL_ZERO_RANGEto mark the blocks between 2GB and 2.5GB as unwritten zeroes in metadata, ensuring subsequent reads return zeroes without writing physical data.stat: Confirms the file has shrunk from 16GB ($17,179,869,184$ bytes) to exactly 12GB ($12,884,901,888$ bytes) instantaneously.
What the Administrator Does Next
Ensure that any logging applications holding an open handle to this file refresh their internal byte position pointers or reopen the file descriptor to align with the new offset boundaries.
Use Case 5: Simulating Storage Capacity Thresholds to Audit Monitoring Alerts
Operational Context
Reliability engineering standards require testing that disk alerts (such as Prometheus DiskSpaceFillingSoon and NodeFilesystemSpaceExhausted) and automated storage tier migrations trigger properly before an actual crisis occurs. Writing hundreds of gigabytes of dummy data generates artificial disk contention on shared storage arrays. With fallocate, an engineer can fill a 500GB volume to 95% capacity in milliseconds without generating write traffic.
# 1. Query the available storage capacity on the target mount point
df -h /mnt/datastore
# 2. Instantly allocate a synthetic balloon file consuming 450GiB
fallocate -l 450G /mnt/datastore/synthetic_pressure.balloon
# 3. Confirm that the filesystem reflects the constrained capacity
df -h /mnt/datastore
# 4. After verifying monitoring alerts and automated failover hooks, purge the file
rm -f /mnt/datastore/synthetic_pressure.balloon
Filesystem Size Used Avail Use% Mounted on
/dev/nvme1n1 500G 25G 475G 5% /mnt/datastore
Filesystem Size Used Avail Use% Mounted on
/dev/nvme1n1 500G 475G 25G 95% /mnt/datastore
Line-by-Line Technical Analysis
df -h /mnt/datastore: Shows that the 500GB NVMe partition initially had 475GB of free space (5% utilization).fallocate -l 450G ...: The kernel reserves 450GB of physical blocks in the filesystem allocation table in under 15 milliseconds.475G Used (95%): Monitoring agents such asnode_exporterinstantly detect the drop in free space and fire alert notifications across test channels without stressing the storage array.
What the Administrator Does Next
Verify that alerting channels (PagerDuty, Slack) receive the expected alerts within their evaluation windows, and log the completed drill in the team's operational audit record.
5. Architectural Nuances, Incompatibilities, and Failure Modes
While fallocate is exceptionally fast, its capabilities depend directly on the design of the underlying filesystem:
| Feature / Flag | ext4 | XFS | Btrfs | ZFS (zfs-on-linux) |
|---|---|---|---|---|
Basic Preallocation (-l) |
Native | Native | Supported* | Emulated via zero |
Hole Punching (-p) |
Native | Native | Native | Native (>= 0.8.0) |
Zero Range (-z) |
Native | Native | Emulated | Emulated |
Collapse Range (-c) |
Native | Native | Unsupported | Unsupported |
Insert Range (-i) |
Native | Native | Unsupported | Unsupported |
* On Btrfs, preallocation reserves space in metadata, but subsequent writes may relocate blocks unless copy-on-write is explicitly disabled.
Caveats with Copy-on-Write (CoW) Filesystems (Btrfs, ZFS)
On traditional overwrite-in-place filesystems like ext4 and XFS, preallocating an unwritten extent guarantees that when you write to that file later, your data will land in the exact physical blocks assigned during creation.
On Copy-on-Write (CoW) filesystems like Btrfs and ZFS, the storage engine does not overwrite existing blocks in place. When an application writes new data to a preallocated file, the filesystem writes the data to a fresh physical location on disk and updates its tree pointer. As a result:
* Preallocation reserves space against disk fullness, but does not guarantee an unbroken contiguous layout over time.
* Creating traditional swap files with fallocate on Btrfs is restricted because swap mechanisms require stationary physical block addresses. On Btrfs, administrators must disable copy-on-write with chattr +C or use the dedicated management tool:
bash
btrfs filesystem mkswapfile --size 8G /swapfile
The Block-Alignment Constraint (EINVAL)
Advanced operations like --punch-hole, --collapse-range, and --insert-range work strictly at the boundary level of filesystem storage blocks (typically 4,096 bytes). If you attempt to trim or collapse a range that does not align cleanly with the system's block size, the kernel rejects the command with an EINVAL (Invalid argument) error:
# Attempting to collapse an unaligned 1000-byte segment:
fallocate -c -o 0 -l 1000 /var/log/telemetry/flight_recorder.bin
fallocate: /var/log/telemetry/flight_recorder.bin: fallocate failed: Invalid argument
Resolution: Always specify offsets and lengths that are clean multiples of your filesystem's block size (e.g., using standard unit suffixes like 4K, 1M, or 1G).
Auditing Disk Space: ls -lhs versus du versus stat
Because fallocate decouples visible file size from the actual blocks allocated on disk, standard directory listings can be deceiving:
# Auditing a punched sparse file:
ls -l file.img # Displays logical size only (e.g., 100GB) -> Can be misleading!
ls -lhs file.img # Displays [Allocated Blocks on Disk] alongside [Logical Size]
du -h file.img # Displays actual physical storage consumed on the drive
stat file.img # Displays exact byte size, 512-byte block counts, and block geometry
6. Today's Takeaway
The fallocate command bridges the gap between high-level file management and low-level storage architecture. Rather than burning time, CPU cycles, and SSD drive lifespan by streaming physical zeroes with dd, you can reserve, trim, or reshape multi-gigabyte storage files in a fraction of a millisecond.
You can test this right now on your own Linux machine in less than thirty seconds. Open your terminal, move to /tmp, and run:
fallocate -l 500M test_allocation.dat && ls -lhs test_allocation.dat && rm -f test_allocation.dat
Watching a half-gigabyte file materialize on your drive instantaneouslyโfully guaranteed in physical storage without a single byte of zero payload passing through your storage busโis the fastest way to experience the quiet power of filesystem metadata at work.