Free: Auditing Kernel Memory Allocations, Demystifying Buffer Cache Dynamics, and Triaging Swap Saturation in Production
Before anyone hits the panic button or authorises costly infrastructure changes, you open an SSH terminal and type five keystrokes that cut straight through the confusion:
free -h -w
In less than two seconds, the terminal returns a clean grid of figures that reveals the truth. The operating system has not run out of memory; instead, the Linux kernel has opportunistically borrowed nearly 80 gigabytes of otherwise idle RAM to cache frequently read files from disk. The system is in peak health. The command that prevented an unnecessary production outage is the classic Unix utility free.
For decades, this modest command has stood guard on almost every Linux system in the world, translating the kernel's arcane internal accounting into numbers engineers can reason about. But reading those numbers correctly requires unlearning one of the oldest instincts in personal computing: the belief that empty RAM is good RAM.
1. What It Does in Plain English
To understand free, one must understand how modern operating systems manage memory. In an everyday desktop operating system, a user might assume that memory works like a water tank: empty space is safe, and full space means danger. But in high-performance computing, unused memory is wasted memory.
The free command provides an instantaneous snapshot of system-wide physical memory (RAM) and virtual swap memory. Rather than treating memory as a monolithic bucket, it queries the kernel to separate memory into distinct operational categories:
- Locked Memory: Memory actively reserved by running applications for their heaps, execution stacks, and private runtime data.
- Shared Memory: Memory blocks allocated for inter-process communication or temporary in-memory filesystems (
tmpfs). - Dynamic Cache: Memory temporarily harnessed by the kernel to mirror disk reads and buffer pending writes. This memory can be instantly reclaimed whenever an application asks for more space.
By presenting these distinctions clearly, free lets engineers determine whether a system is genuinely starving for resources or simply operating at maximum caching efficiency.
2. Core Flags & Quick Start
Maintained under the upstream procps-ng project, free provides several essential switches to tailor its output for human inspection or automated monitoring:
| Flag | Long Flag | Purpose & Operational Value |
|---|---|---|
-h |
--human |
Formats byte counts automatically into human-readable binary units (KiB, MiB, GiB). |
-w |
--wide |
Prevents combining buffers and cache, displaying them as separate columns. |
-b, -k, -m, -g |
--bytes, --kibi, etc. |
Forces fixed output units in bytes, kibibytes, mebibytes, or gibibytes for automated scripts. |
-s [N] |
--seconds [N] |
Continuously polls and refreshes memory statistics every N seconds. |
-c [N] |
--count [N] |
Limits polling cycles to N iterations when used alongside -s. |
-t |
--total |
Adds a summary row calculating the mathematical sum of physical RAM and swap space. |
-l |
--lohi |
Displays detailed low and high memory subsystem breakdowns (relevant on 32-bit systems). |
The Essential Quick Start
For standard incident triage, the wide, human-readable combination is the industry benchmark:
free -h -w
total used free shared buffers cache available
Mem: 62Gi 14Gi 2.1Gi 1.2Gi 850Mi 45Gi 46Gi
Swap: 8.0Gi 256Mi 7.7Gi
In this output, total reports 62 GiB of usable physical RAM. Although free shows that raw unallocated memory (free) is only 2.1 GiB, the available column confirms that the operating system can immediately provide 46 GiB of memory to incoming workloads without having to evict active processes or touch swap space.
3. Kernel Memory Architecture & /proc/meminfo Deconstruction
Under the hood, free is a lightweight user-space wrapper around /proc/meminfo, a dynamic pseudo-file generated on the fly by the Linux kernel's virtual memory manager.
(AnonPages)"] end subgraph Reclaimable["Dynamic Kernel Caching (Reclaimable)"] Buffers["I/O Buffers
(Raw Block Device Metadata)"] Cache["Page Cache & Slab
(File Pages & Inodes)"] end subgraph FreeMem["Raw Free Memory"] RawFree["MemFree
(Unallocated Buddy Pages)"] end end Reclaimable -.-> Avail["MemAvailable
(Immediate Headroom for New Applications)"] RawFree -.-> Avail
The Virtual Pipeline: From Kernel Counters to User Space
When free executes, it reads /proc/meminfo. The kernel computes these figures by aggregating memory zone watermarks, page tables, slab descriptors, and per-CPU virtual memory statistics (vm_stat).
The key metrics reported by free map to specific kernel calculations:
$$\text{MemTotal} = \text{Total Physical RAM} - \text{Kernel Binary Image} - \text{Reserved Architecture Pages}$$
$$\text{MemFree} = \sum \text{Free Pages in Kernel Buddy Allocator} \times \text{Page Size}$$
$$\text{Buffers} = \text{In-Memory Raw Block Device I/O Descriptors (Filesystem Metadata)}$$
$$\text{Cached} = \text{Page Cache} - \text{Shmem} + \text{Tmpfs Allocation}$$
$$\text{SReclaimable} = \text{Slab Allocator Inode and Dentry Caches Marked for Shrinking}$$
$$\text{SwapTotal} - \text{SwapFree} = \text{Used Swap Pages}$$
Demystifying "Linux Ate My RAM"
Because reading data from storage drives (even high-speed NVMe SSDs) is hundreds of times slower than reading directly from physical RAM, the Linux kernel automatically allocates unused memory to mirror recently accessed files. This mechanism is known as the Page Cache.
Kernel memory falls into two fundamental categories:
- Anonymous Memory (
AnonPages): Memory allocated by user applications for heaps, variables, and execution stacks. It has no backing file on disk. Anonymous memory cannot be dropped unless it is paged out to swap storage or the owning process is terminated. - File-Backed Memory (
Active(file)/Inactive(file)): Memory pages that mirror data stored on disk. When an application reads a file, the kernel keeps those 4 KiB memory pages in RAM. If another process needs that physical RAM, the kernel can immediately discard clean file-backed pages with zero disk write penalty, because the permanent copy already resides on disk.
When basic monitoring scripts calculate used = total - free, they mistakenly treat file-backed page cache as consumed RAM, triggering false alarms.
The Historic Shift: MemFree vs. MemAvailable
Before Linux kernel version 3.14, sysadmins estimated available headroom using an approximate formula:
$$\text{Free}_{\text{Estimated}} \approx \text{MemFree} + \text{Buffers} + \text{Cached}$$
This heuristic was dangerous because it assumed that 100% of buffers and cached memory could be reclaimed on demand. In practice, significant portions cannot be freed immediately:
- Dirty Pages: Memory containing writes that have not yet been flushed to disk cannot be dropped without waiting for disk I/O.
- Shared Memory &
tmpfs: Stored within the page cache, but cannot be discarded without swapping. - Zone Watermarks (
min_free_kbytes): The kernel must preserve a minimum safety margin of memory per zone to handle network packets and hardware interrupts without deadlocking.
To provide an authoritative figure, Linux 3.14 introduced MemAvailable directly in the kernel source, calculated using a battle-tested reclamation formula:
$$\text{MemAvailable} = \text{MemFree} - W_{\text{low}} + \left( \text{Active(file)} + \text{Inactive(file)} - \min\left(\frac{\text{Active(file)} + \text{Inactive(file)}}{2}, W_{\text{low}}\right) \right) + \left( \text{SReclaimable} - \min\left(\frac{\text{SReclaimable}}{2}, W_{\text{low}}\right) \right)$$
where $W_{\text{low}}$ represents the sum of low memory watermarks across all active memory zones. Today, MemAvailable is the single most reliable metric for evaluating true application memory headroom.
4. Five Real-World Production Use Cases
Use Case 1: Wide-Mode Memory Auditing Under Database Saturation
Scenario: During a nightly batch indexing job on a dedicated PostgreSQL database server, a monitoring alert warns that memory utilisation has crossed 95%. The sysadmin must verify whether memory is being consumed by runaway database connection processes or harmless read caching.
free -h -w
total used free shared buffers cache available
Mem: 125Gi 38Gi 5.2Gi 4.1Gi 3.8Gi 78Gi 82Gi
Swap: 16Gi 0B 16Gi
PostgreSQL Processes & Heaps"] CacheMem["Buffers & Cache (81.8 GiB)
Shared Buffers & Page Cache"] FreeMem["Free (5.2 GiB)
Raw Unallocated"] end CacheMem -.-> AvailMem["Real Usable Headroom: 82 GiB"] FreeMem -.-> AvailMem
Line-by-Line Telemetry Dissection:
* total (125Gi): Total physical RAM recognized by the kernel and hypervisor.
* used (38Gi): Memory actively claimed by PostgreSQL backend processes, connection pools, and operating system services.
* free (5.2Gi): Pages currently sitting completely unused in the kernel buddy allocator.
* shared (4.1Gi): Memory dedicated to PostgreSQL shared buffers (shared_buffers) via POSIX shared memory.
* buffers (3.8Gi): Raw block I/O metadata holding disk synchronization structures.
* cache (78Gi): File-backed page cache holding read tables and unwritten pages.
* available (82Gi): Real application allocation capacity before any swapping is required.
What the Administrator Does Next: The sysadmin observes that available memory stands at a comfortable 82 GiB and Swap is completely untouched (0B used). No failover, restart, or node resize is necessary. The sysadmin updates the monitoring rule to track available rather than raw free memory.
Use Case 2: Continuous Real-Time Sampling During Deployment Spikes
Scenario: A newly compiled microservice container is rolled out to a staging environment. Engineers suspect a memory leak in a memory allocator that continuously claims heap space under synthetic user load.
free -s 1 -c 5 -m
total used free shared buff/cache available
Mem: 31892 4210 22100 120 5582 27120
Swap: 4096 0 4096
Mem: 31892 4980 21330 120 5582 26350
Swap: 4096 0 4096
Mem: 31892 5750 20560 120 5582 25580
Swap: 4096 0 4096
Mem: 31892 6520 19790 120 5582 24810
Swap: 4096 0 4096
Mem: 31892 7290 19020 120 5582 24040
Swap: 4096 0 4096
Line-by-Line Telemetry Dissection:
* Iteration 1: used starts at 4,210 MiB, and available is 27,120 MiB.
* Iterations 2 through 5: Exactly 770 MiB of memory shifts from free to used every second.
* buff/cache stays completely flat at 5,582 MiB, confirming the growth is not legitimate disk logging or caching.
* available drops in lockstep with free, falling from 27,120 MiB to 24,040 MiB in just four seconds.
What the Administrator Does Next: The sysadmin notes an alarming allocation rate of ~770 MiB per second. The engineer immediately cancels the canary rollout, preventing the microservice from exhausting container cgroup memory limits and triggering an abrupt Out-Of-Memory (OOM) crash.
Use Case 3: Isolating Shared Memory (tmpfs) and IPC Leaks
Scenario: A Kubernetes worker node exhibits a persistent drop in available memory. However, standard process inspection tools like top and ps show no individual process consuming abnormal amounts of RAM.
free -m
total used free shared buff/cache available
Mem: 64120 12400 1850 38200 49870 12800
Swap: 8192 512 7680
Standard Process Heaps"] subgraph CacheZone["Buff/Cache (49,870 MiB)"] SharedLeak["Shared / tmpfs Leak (38,200 MiB)
Non-Reclaimable Caching"] RealCache["Clean Page Cache (11,670 MiB)
Reclaimable"] end RawFreeMem["Raw Free (1,850 MiB)"] end RealCache --> RealAvail["Real Usable Headroom: 12,800 MiB"] RawFreeMem --> RealAvail
Line-by-Line Telemetry Dissection:
* shared (38200): An unusually high 38.2 GiB is tied up in shared memory segments or tmpfs mounts.
* buff/cache (49870): The aggregate cache shows nearly 50 GiB, but over 76% of it is consumed by the shared allocation.
* available (12800): The kernel reports only 12.8 GiB available, correctly recognizing that shared memory cannot be reclaimed through standard cache eviction.
What the Administrator Does Next: The sysadmin investigates POSIX shared memory allocations and ramdisk mounts:
df -h /dev/shm && ipcs -m
The output reveals abandoned shared memory segments left behind by crashed container processes. The sysadmin clears the orphaned segments, instantly releasing 38 GiB of physical RAM back to the operating system.
Use Case 4: Automating Threshold Calculations in Production Healthcheck Scripts
Scenario: A site reliability engineer needs to build an automated node health check for a cluster watchdog daemon. The script must evaluate exact, byte-level available memory without rounding errors or unit conversion mistakes.
free -b
total used free shared buffers cache available
Mem: 67425718272 18253611008 2147483648 1073741824 536870912 46487752704 47561498624
Swap: 8589934592 0 8589934592
The engineer writes an automated verification script:
#!/usr/bin/env bash
set -euo pipefail
# Extract total and available memory in exact bytes
read -r MEM_TOTAL MEM_AVAIL < <(free -b | awk '/^Mem:/ {print $2, $7}')
# Calculate integer percentage of available memory
AVAIL_PERCENT=$(( (MEM_AVAIL * 100) / MEM_TOTAL ))
echo "Deterministic Healthcheck: ${AVAIL_PERCENT}% memory headroom available (${MEM_AVAIL} bytes)."
# Threshold check: Alert if available memory drops below 10%
if [ "${AVAIL_PERCENT}" -lt 10 ]; then
echo "CRITICAL: Available memory below threshold. Initiating pod eviction cordon." >&2
# Execution hook: /usr/local/bin/cordon-node.sh
exit 1
fi
Line-by-Line Telemetry Dissection:
* free -b: Emits figures directly in raw bytes, preventing floating-point rounding inaccuracies.
* $2, $7: Parses MemTotal ($67,425,718,272$ bytes) and MemAvailable ($47,561,498,624$ bytes) with precision.
* AVAIL_PERCENT: Accurately returns $70\%$ headroom.
What the Administrator Does Next: The engineer deploys the script to the cluster monitoring agent, ensuring reliable automated remediation without false eviction storms.
Use Case 5: Evaluating Kernel Memory Subsystem Tuning
Scenario: A systems architect is tuning an I/O-intensive file server. Under heavy load, the server swaps out active application memory to disk even though ample page cache is available. The architect adjusts vm.swappiness and vm.vfs_cache_pressure to prioritise reclaiming directory and inode caches before swapping processes.
Pre-Optimization Telemetry:
free -h
total used free shared buff/cache available
Mem: 62Gi 48Gi 1.1Gi 2.0Gi 13Gi 11Gi
Swap: 8.0Gi 4.2Gi 3.8Gi
The architect applies new virtual memory parameters using sysctl:
sysctl -w vm.swappiness=10
sysctl -w vm.vfs_cache_pressure=50
Post-Optimization Telemetry (Under Identical Workload):
free -h
total used free shared buff/cache available
Mem: 62Gi 36Gi 2.4Gi 2.0Gi 24Gi 23Gi
Swap: 8.0Gi 120Mi 7.9Gi
Line-by-Line Telemetry Dissection:
* Pre-tuning: Swap usage is elevated at 4.2Gi, while buff/cache is constrained to 13Gi. The kernel was prematurely writing application memory to swap to maintain cache space.
* Post-tuning: Swap usage drops to 120Mi, buff/cache expands to 24Gi, and available headroom climbs from 11Gi to 23Gi.
What the Administrator Does Next: The architect confirms that tuning vm.swappiness=10 prevents premature swapping of idle processes while vm.vfs_cache_pressure=50 preserves directory caches. The settings are made permanent in /etc/sysctl.d/99-memory-performance.conf.
5. Production Pitfalls & Architectural Blindspots
Interpreting free accurately requires an understanding of edge cases where numbers on the screen can be deceptive:
| Blindspot | Mechanism | Diagnostic Strategy |
|---|---|---|
| The Legacy Fallacy | Obsolete -+/ buffers/cache math assumed 100% cache reclaimability. |
Rely exclusively on the modern available column. |
| Unreclaimable Slab Leaks | Kernel memory allocations (SUnreclaim) consume RAM without showing in process tables. |
Inspect /proc/meminfo and run slabtop. |
| Static HugePages | Dedicated Hugetlb allocations are pinned in RAM and removed from available. |
Check grep -i huge /proc/meminfo. |
| Hypervisor Ballooning | Virtual machine balloon drivers dynamically inflate, shrinking guest memory. | Check vmware-toolbox-cmd or lsmod \| grep balloon. |
1. The Legacy Fallacy: -/+ buffers/cache
Engineers accustomed to legacy Linux systems (such as CentOS 6 or Debian 7) often look for the historic second output line:
# OBSOLETE OUTPUT FORMAT (Pre-procps-ng 3.3.10)
-/+ buffers/cache: 12450 51670
This legacy line computed:
$$\text{Used}_{\text{Legacy}} = \text{used} - (\text{buffers} + \text{cached})$$
$$\text{Free}_{\text{Legacy}} = \text{free} + (\text{buffers} + \text{cached})$$
This formulation is unreliable because it assumes that every single cached byte can be discarded immediately, ignoring locked dirty pages, shared memory allocations, and non-reclaimable kernel metadata. Modern tooling should always reference MemAvailable.
2. The Unreclaimable Slab Trap (SUnreclaim)
The Linux kernel allocates internal objects (such as directory entries, file inodes, and socket buffers) using the Slab Allocator. The slab is divided into two parts:
SReclaimable: Caches that can be shrunk on demand under memory pressure.SUnreclaim: Structures pinned permanently in memory (such as tracking data or out-of-tree kernel driver allocations).
If a kernel module leaks slab memory, used will rise and available will drop, yet tools like ps and top will show zero user-space processes consuming RAM. To diagnose this, inspect /proc/meminfo directly:
grep -E 'Slab|SReclaimable|SUnreclaim' /proc/meminfo
If SUnreclaim accounts for gigabytes of memory, launch slabtop to pinpoint the specific kernel object causing the leak.
3. Dedicated HugePages Allocation (Hugetlb)
High-performance applications like large relational databases and virtualization hosts often pre-allocate Static HugePages (typically 2 MiB or 1 GiB page sizes) to prevent translation lookaside buffer (TLB) bottlenecks:
# Pre-allocate 16,000 HugePages of 2 MiB (~31.25 GiB)
sysctl -w vm.nr_hugepages=16000
Because HugePages are permanently reserved for dedicated workloads, standard Linux applications cannot use them. Consequently, free reports this memory as subtracted from free and excluded from available. When diagnosing mysterious memory consumption, always verify HugePage reservations:
grep -i huge /proc/meminfo
4. Hypervisor Ballooning Distortions
In virtualized environments (such as VMware ESXi, KVM, or public cloud instances), hypervisors dynamically balance host memory across guests using memory balloon drivers (e.g., virtio_balloon, vmw_balloon).
When the physical host experiences memory pressure, it instructs the guest balloon driver to inflate, locking guest pages so the host can reclaim them. Depending on the hypervisor configuration:
MemTotalreported byfreemay shrink dynamically without warning.usedmemory inside the guest may increase artificially.
If memory metrics change unexpectedly inside a virtual instance, check for active balloon drivers:
vmware-toolbox-cmd stat balloon 2>/dev/null || lsmod | grep balloon
6. Comprehensive Troubleshooting Workflow
When investigating memory alerts, avoid relying on free in isolation. Follow this structured triage workflow:
(Memory utilized by Page Cache; ignore low 'free')"] CheckAvail -- No --> Step2["Run: vmstat 1 5
(Inspect swap paging: si / so)"] CheckThrash{"Are si / so > 0?"} Step2 --> CheckThrash CheckThrash -- Yes (Active Thrashing) --> Step3["Identify Top RSS Processes
ps -eo pid,user,pmem,rss,vsz,comm --sort=-rss | head"] CheckThrash -- No (Memory Pinned) --> StepIPC["Audit Slab & Shared Memory
grep -E 'Slab|Shmem' /proc/meminfo"] Step3 --> CheckOOM["Audit Kernel OOM Killer Logs
dmesg -T | grep -i oom-killer"] StepIPC --> CheckOOM
Step 1: Rapid Headroom Assessment
Run free -h -w. If available is greater than 20% of total, the host is under no immediate memory pressure, regardless of how small the raw free column appears.
Step 2: Correlate with I/O and Paging Activity
If available is low, determine whether the host is actively thrashing swap space by inspecting /proc/vmstat with vmstat:
vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
2 1 524288 104856 40960 524288 128 256 1024 2048 4500 8200 45 25 10 20 0
Focus on the si (swap-in) and so (swap-out) columns. If so is consistently above zero, the kernel is actively evicting memory pages to disk, causing severe latency spikes.
Step 3: Identify Top Memory-Consuming Processes
Pinpoint which processes are consuming physical resident memory:
ps -eo pid,user,pmem,rss,vsz,comm --sort=-rss | head -n 10
RSS(Resident Set Size): The actual physical RAM currently occupied by the process.VSZ(Virtual Size): The total virtual address space mapped by the process, including shared libraries and unallocated buffers.
Step 4: Audit Kernel Out-Of-Memory (OOM) Events
If a service has vanished or crashed unexpectedly, check whether the kernel's Out-Of-Memory Killer terminated it to protect system stability:
dmesg -T | grep -E -i 'oom[-_]killer|out of memory|killed process'
[Tue Aug 18 02:14:22 2026] Out of memory: Kill process 18421 (node) score 852 or sacrifice child
[Tue Aug 18 02:14:22 2026] Killed process 18421 (node) total-vm:4214820kB, anon-rss:3145728kB, file-rss:0kB, shmem-rss:0kB
The log entry provides the exact virtual size and physical footprint (anon-rss) of the process at the moment of termination, giving developers the empirical data needed to fix memory leaks.
7. Today's Takeaway
Open a terminal on your computer right now and run free -h -w. Look past the free column and look straight at available. If your available memory is high and swap usage is at or near zero, your operating system is running smoothly, using its memory to accelerate file access and keep everyday operations fast. Whenever you hear alarms about memory exhaustion, remember the golden rule of modern Linux memory triage: never judge system health by how much RAM is empty, but by how much is available.