Slabtop: Auditing Kernel Memory Allocations, Triaging Dentry and Inode Cache Leaks, and Diagnosing Slab Exhaustion in Production
Instead, you find a ghost.
When you open top or htop and tally up every running program on the server—the database engine, the reverse proxies, the background queues, the monitoring daemons—they account for barely a quarter of the machine's sixty-four gigabytes of physical memory. The process list looks calm, peaceful, and lightly loaded. Yet the operating system is gasping for air, the kernel is threatening to kill running services, and tens of gigabytes of physical RAM have seemingly vanished into thin air.
The memory hasn’t been consumed by any user application. It has been absorbed directly by the Linux kernel itself, hidden inside internal memory caches that standard process monitors simply cannot see. When conventional diagnostic tools leave you in the dark, the tool you need is slabtop.
To cut through the fog immediately and see exactly which kernel caches are claiming your memory, run a sorted snapshot in your terminal:
sudo slabtop -s c -o | head -n 20
Active / Total Objects (% used) : 4829140 / 5012480 (96.3%)
Active / Total Slabs (% used) : 156640 / 156640 (100.0%)
Active / Total Caches (% used) : 118 / 174 (67.8%)
Active / Total Size (% used) : 1421042.88K / 1475120.50K (96.3%)
Minimum / Average / Maximum Object : 0.02K / 0.29K / 18.25K
OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME
1842912 1840102 99% 0.10K 47254 39 189016K buffer_head
1420512 1389410 97% 0.20K 74764 19 299056K dentry
894100 872110 97% 0.61K 34388 26 550208K ext4_inode_cache
241020 235100 97% 1.06K 16068 15 257088K xfs_inode
84120 80112 95% 2.00K 10515 8 168240K kmalloc-2k
In less than a second, the missing gigabytes are accounted for. The kernel's internal caches—from file metadata descriptors (dentry and inode) to raw network packet buffers—are laid bare on your screen.
What It Does in Plain English
The slabtop utility is a real-time monitor for the Linux kernel’s internal memory management engine.
While standard diagnostic tools like free, ps, and top track what user-space applications are doing, they treat the kernel largely as a black box. The Linux kernel, however, constantly allocates and deallocates millions of tiny, specialized data structures behind the scenes: file paths, network packet headers, process control blocks, and filesystem pointers.
Rather than constantly requesting raw chunks of memory and building these structures from scratch every time a file is opened or a network packet arrives, the kernel maintains pre-allocated pools of memory called "slabs." The slabtop command reads directly from /proc/slabinfo and presents a live, sortable breakdown of these pools. It tells you exactly how many objects exist, how much RAM they occupy, and whether that memory can be reclaimed safely or represents a stubborn, unyielding leak.
Core Flags and Quick Reference
You can run slabtop interactively as a live dashboard (similar to top) or in single-shot batch mode for scripts and terminal pipelines.
| Flag | Purpose | Operational Context |
|---|---|---|
-s c |
Sort by Cache Size | Orders caches by total memory footprint. This is your primary diagnostic view. |
-s a |
Sort by Active Objects | Identifies which kernel structures have the highest raw count of active items. |
-s n |
Sort by Name | Arranges caches alphabetically for cross-referencing against kernel symbols. |
-s l |
Sort by Loss (Fragmentation) | Surfaces caches with the highest internal fragmentation and wasted space. |
-o |
Once (Batch Mode) | Emits a single clean snapshot without terminal control characters, ideal for pipelines. |
-d <n> |
Delay Interval | Configures the display refresh rate in seconds (default is 3 seconds). |
The Architectural Anatomy of Linux Kernel Memory
To make sense of what slabtop reports during an outage, it helps to understand how Linux allocates memory under the hood.
• Anonymous Memory (Heap, Stack)
• Page Cache (File Buffers, mmap)"] Buddy --> KernelAlloc["Kernel Object Allocator
(SLAB / SLUB / SLOB)"] KernelAlloc --> SReclaim["SReclaimable (Reclaimable Under Pressure)
• dentry (Directory Cache)
• ext4_inode_cache / xfs_inode
• buffer_head"] KernelAlloc --> SUnreclaim["SUnreclaim (Pinned / Critical Allocations)
• task_struct & mm_struct
• skbuff_head_cache (Network)
• kmalloc-X General Caches"]
The Buddy Allocator vs. The SLAB Engine
At the lowest layer of Linux memory management sits the Buddy Allocator. It manages memory in chunks called page frames—typically 4 KiB in size. The Buddy Allocator is brilliant at preventing memory fragmentation when allocating large blocks, but terribly inefficient for tiny kernel structures.
If the kernel needs a 64-byte directory record (dentry), asking the Buddy Allocator for a full 4 KiB page frame would waste over 98% of that memory.
To solve this, Linux uses an object-caching layer called the SLAB Allocator (and its modern, streamlined successor, the SLUB Allocator). The SLAB engine requests full pages from the Buddy Allocator and slices them into neat, pre-sized slots designed specifically for particular kernel objects. When an object is no longer needed, its memory is not returned to the wider system; instead, it sits ready in the slab cache, allowing the kernel to reuse it instantly without CPU overhead.
The Golden Distinction: SReclaimable vs. SUnreclaim
When you inspect system memory via /proc/meminfo, total slab memory is split into two critical categories:
$$\text{Slab} = \text{SReclaimable} + \text{SUnreclaim}$$
-
SReclaimable(Reclaimable Slab Memory): These are opportunistic caches that improve performance, such as directory entries (dentry) and file metadata (inode_cache). When applications demand more memory, the kernel's background memory manager automatically shrinks these caches and hands the memory back to user programs. A highSReclaimablefigure is generally healthy behavior. -
SUnreclaim(Unreclaimable Slab Memory): These are pinned, active operational structures that the kernel cannot discard under memory pressure. They include active process tables (task_struct), network buffers (sk_buff), and general kernel allocations (kmalloc). IfSUnreclaimgrows steadily over time, the system will eventually run out of memory and crash, regardless of how much swap space or free disk you have.
Five Real-World Troubleshooting Scenarios
1. Diagnosing 'Hidden' Memory Saturation: Dentry and Inode Cache Runaways
The Scenario
A high-traffic web server hosting an image delivery platform triggers urgent low-memory alerts. Monitoring graphs show 60 GiB of the server's 64 GiB physical RAM in use. However, running ps aux --sort=-%mem shows that all running web server workers and background processes combined are using barely 8 GiB. Free memory is near zero, and disk operations are slowing to a crawl.
The Diagnostic Command
slabtop -s c -o | head -n 15
Active / Total Objects (% used) : 42194812 / 42201000 (100.0%)
Active / Total Slabs (% used) : 2221105 / 2221105 (100.0%)
Active / Total Caches (% used) : 98 / 142 (69.0%)
Active / Total Size (% used) : 38869337.50K / 38875000.00K (100.0%)
Minimum / Average / Maximum Object : 0.02K / 0.92K / 18.25K
OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME
28140224 28140224 100% 0.20K 1481064 19 23697024K dentry
12540192 12540192 100% 0.61K 482315 26 15434080K ext4_inode_cache
841200 839100 99% 0.10K 21569 39 86276K buffer_head
120400 118200 98% 0.21K 6688 18 26752K vm_area_struct
45200 44100 97% 1.00K 2825 16 45200K kmalloc-1k
Line-by-Line Explanation
- Lines 1–5 (Header): Total active slab memory has ballooned to roughly 37 GiB ($38,869,337.50\text{ KiB}$) across more than 42 million objects, sitting at 100% saturation.
- Line 7 (
dentry): Directory entry caches are holding 28.14 million objects and consuming nearly 24 GiB of RAM. Every time a program navigates a directory path, the kernel caches the path lookup here. - Line 8 (
ext4_inode_cache): File metadata caches are holding 12.54 million file records, occupying over 15 GiB of RAM. - Lines 9–11: Other caches (
buffer_head,vm_area_struct) are small and operating normally within baseline parameters.
What the Admin Does Next
Investigation reveals that a daily backup script was searching through millions of temporary files on disk, flooding the kernel with path lookups. Although dentry and inode caches are technically reclaimable, holding tens of millions of them can cause severe latency spikes when user processes suddenly demand RAM and force the kernel to clean up millions of records at once.
To make the kernel reclaim directory and inode caches more aggressively before memory gets tight, tune the virtual filesystem cache pressure dynamically with sysctl:
# Check current VFS cache pressure (default is 100)
sysctl vm.vfs_cache_pressure
# Increase aggressiveness of kernel cache reclamation
sudo sysctl -w vm.vfs_cache_pressure=150
2. Triaging Network Socket Buffer Leaks Under High Connection Churn
The Scenario
An API edge gateway starts dropping incoming requests during a morning traffic surge. CPU utilization is modest (under 40%), and user-space memory looks normal, but the kernel logs begin reporting network starvation: TCP: out of memory -- consider tuning tcp_mem. We need to inspect whether network socket buffers are getting pinned and exhausting kernel memory.
The Diagnostic Command
slabtop -s a -o | head -n 16
Active / Total Objects (% used) : 8941200 / 9102400 (98.2%)
Active / Total Slabs (% used) : 455120 / 455120 (100.0%)
Active / Total Caches (% used) : 112 / 168 (66.7%)
Active / Total Size (% used) : 12450810.00K / 12672400.00K (98.3%)
Minimum / Average / Maximum Object : 0.02K / 1.39K / 18.25K
OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME
4120500 4120000 99% 0.25K 257531 16 1030124K skbuff_head_cache
2410200 2408900 99% 2.00K 301275 8 4820400K kmalloc-2k
984000 982100 99% 0.50K 61500 16 492000K tcp_bind_bucket
512000 510200 99% 0.38K 24380 21 195040K tw_sock_TCP
312000 309100 99% 1.00K 19500 16 312000K sock_inode_cache
Line-by-Line Explanation
- Line 7 (
skbuff_head_cache): Shows 4.12 million active network packet header structures (struct sk_buff), consuming over 1 GiB of RAM. - Line 8 (
kmalloc-2k): Consumes 4.82 GiB across 2.41 million objects. Network interface drivers frequently usekmalloc-2kto store raw incoming packet payloads. - Line 9 (
tcp_bind_bucket): Nearly 1 million network binding records are active in kernel memory. - Line 10 (
tw_sock_TCP): Over 510,000 TCP sockets are lingering in theTIME_WAITteardown state.
What the Admin Does Next
The gateway's upstream clients are closing HTTP connections after every single request instead of reusing persistent connections (Keep-Alive). This churn leaves hundreds of thousands of sockets trapped in TIME_WAIT, pinning network slab structures and exhausting kernel socket space.
The administrator immediately enables TCP socket reuse and tightens timeout limits to drain lingering sockets:
# Allow immediate reuse of TIME_WAIT sockets for outgoing traffic
sudo sysctl -w net.ipv4.tcp_tw_reuse=1
# Cap the maximum lingering TIME_WAIT buckets and lower the FIN timeout
sudo sysctl -w net.ipv4.tcp_max_tw_buckets=262144
sudo sysctl -w net.ipv4.tcp_fin_timeout=15
3. Auditing Container Cgroup kmem Allocations Under Ephemeral Process Spawning
The Scenario
On a multi-tenant Kubernetes worker node running CI/CD build pipelines, the kernel abruptly terminates container pods with Out-Of-Memory errors. When developers look at their application logs, memory usage appears well below container limits. The suspicion is that rapid, short-lived subprocess invocations (fork/exec loops inside build scripts) are generating kernel descriptors charged to the container’s Kernel Memory Cgroup.
The Diagnostic Command
slabtop -s a -o | head -n 16
Active / Total Objects (% used) : 3410280 / 3491200 (97.7%)
Active / Total Slabs (% used) : 198400 / 198400 (100.0%)
Active / Total Caches (% used) : 104 / 152 (68.4%)
Active / Total Size (% used) : 7120400.00K / 7291200.00K (97.7%)
Minimum / Average / Maximum Object : 0.02K / 2.08K / 18.25K
OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME
842100 841000 99% 4.00K 105262 8 3368400K task_struct
842100 840900 99% 1.06K 56140 15 898240K mm_struct
842100 841100 99% 0.69K 36613 23 585808K sighand_cache
842100 840800 99% 0.75K 39598 21 631425K files_cache
842100 841200 99% 0.06K 3302 64 52832K pid
Line-by-Line Explanation
- Line 7 (
task_struct): Over 841,000 active process descriptor records are occupying 3.36 GiB of non-reclaimable RAM. Every single running thread or process in Linux requires one of these 4 KiB structures. - Line 8 (
mm_struct): 840,900 virtual memory map descriptors are consuming nearly 900 MiB. - Lines 9–11 (
sighand_cache,files_cache,pid): Exactly match thetask_structcount, representing signal handlers, file tables, and process ID allocations for each spawned process. - The Root Problem: A broken test runner script inside a container was spawning thousands of un-reaped subprocesses in a loop. Because the parent process never collected their exit codes, the kernel had to hold their unreclaimable process structures in memory, directly charging the container's cgroup memory quota until it tripped the OOM killer.
What the Admin Does Next
- Identify which container cgroup is generating the runaway process structures:
bash cat /sys/fs/cgroup/system.slice/docker-*.scope/memory.stat | grep -E "kernel|slab" - Place a strict process limit (
pids.max) on the container runtime to prevent fork bombs from consuming kernel memory:bash echo 2048 > /sys/fs/cgroup/kubepods.slice/kubepods-burstable.slice/pids.max
4. Telemetry Pipeline: Automated Early-Warning System for SUnreclaim Growth
The Scenario
In large cloud deployments, slow kernel memory leaks—often introduced by third-party drivers, filesystem filters, or custom eBPF hooks—can take weeks to exhaust physical memory. Standard monitoring systems that only look at free RAM miss the gradual shift from reclaimable memory to dangerous, unreclaimable kernel memory. We need an automated monitoring script that extracts slab metrics and exports them to Prometheus before servers hit critical levels.
The Automation Script
#!/usr/bin/env bash
# /usr/local/bin/telemetry_slab_audit.sh
# Automated kernel slab telemetry probe for Prometheus Node Exporter
set -euo pipefail
METRICS_FILE="/var/lib/node_exporter/textfile_collector/kernel_slab.prom"
TEMP_FILE="${METRICS_FILE}.tmp"
# Read /proc/meminfo to capture aggregate slab metrics (in KiB)
S_RECLAIMABLE=$(awk '/SReclaimable:/ {print $2}' /proc/meminfo)
S_UNRECLAIM=$(awk '/SUnreclaim:/ {print $2}' /proc/meminfo)
TOTAL_SLAB=$((S_RECLAIMABLE + S_UNRECLAIM))
# Extract top 3 consuming individual caches via slabtop batch mode
TOP_CACHES_JSON=$(slabtop -s c -o | awk '
BEGIN { print "[" }
NR > 7 && count < 3 {
if (count > 0) print ","
printf " {\"name\": \"%s\", \"size_kb\": %d, \"active_objs\": %d}", $8, substr($7, 1, length($7)-1), $2
count++
}
END { print "\n]" }
')
# Generate Prometheus metrics format
cat <<EOF > "${TEMP_FILE}"
# HELP node_kernel_slab_reclaimable_bytes Reclaimable kernel slab memory in bytes
# TYPE node_kernel_slab_reclaimable_bytes gauge
node_kernel_slab_reclaimable_bytes $((S_RECLAIMABLE * 1024))
# HELP node_kernel_slab_unreclaim_bytes Unreclaimable kernel slab memory in bytes
# TYPE node_kernel_slab_unreclaim_bytes gauge
node_kernel_slab_unreclaim_bytes $((S_UNRECLAIM * 1024))
# HELP node_kernel_slab_total_bytes Total allocated kernel slab memory in bytes
# TYPE node_kernel_slab_total_bytes gauge
node_kernel_slab_total_bytes $((TOTAL_SLAB * 1024))
EOF
mv "${TEMP_FILE}" "${METRICS_FILE}"
# Alerting Logic: Trigger warning if SUnreclaim exceeds 15% of total physical RAM
MEM_TOTAL=$(awk '/MemTotal:/ {print $2}' /proc/meminfo)
UNRECLAIM_RATIO=$(awk "BEGIN {print ($S_UNRECLAIM / $MEM_TOTAL) * 100}")
if (( $(echo "$UNRECLAIM_RATIO > 15.0" | bc -l) )); then
logger -p user.crit "CRITICAL: Kernel SUnreclaim memory has reached ${UNRECLAIM_RATIO}% of total RAM. Immediate kernel leak audit required."
fi
Line-by-Line Explanation
- Lines 10–13: Parses
/proc/meminfoto calculate aggregate reclaimable and unreclaimable slab sizes in kilobytes. - Lines 16–25: Invokes
slabtop -s c -oin batch mode, safely parsing the top three memory consumers without crashing on terminal formatting codes. - Lines 28–39: Writes clean Prometheus metric gauges to the node-exporter directory for dashboard visualization.
- Lines 41–47: Computes the ratio of unreclaimable slab memory against total system memory. If it crosses 15%, it sends a critical alert to syslog.
What the Admin Does Next
The sysadmin installs this script as a cron job running every 5 minutes. When an alert fires, engineering teams have days of advance notice to triage leaking drivers or schedule rolling node restarts during normal business hours rather than suffering an unprompted midnight crash.
5. Auditing Memory Compaction and Reclamation Post-Drop Signals
The Scenario
Before migrating a large database instance or spinning up a huge virtual machine that demands contiguous physical memory blocks, you need to verify how much kernel memory can actually be freed. We run a controlled cache drop and use slabtop before and after to verify what was released and what remains pinned.
Command Execution Sequence
# Step 1: Capture baseline slab cache allocation
slabtop -s c -o | head -n 12 > /tmp/slab_baseline.txt
# Step 2: Flush dirty file buffers to disk before dropping caches
sync
# Step 3: Trigger full cache purge (free pagecache, dentries, and inodes)
echo 3 | sudo tee /proc/sys/vm/drop_caches
# Step 4: Capture immediate post-drop slab state
slabtop -s c -o | head -n 12 > /tmp/slab_post_drop.txt
# Step 5: Compare baseline against post-drop state
diff -u /tmp/slab_baseline.txt /tmp/slab_post_drop.txt || true
Output Comparison and Diagnostic Diff
--- /tmp/slab_baseline.txt
+++ /tmp/slab_post_drop.txt
@@ -1,11 +1,11 @@
- Active / Total Objects (% used) : 18450120 / 18501200 (99.7%)
- Active / Total Size (% used) : 14205012.00K / 14250000.00K (99.7%)
+ Active / Total Objects (% used) : 412010 / 512000 (80.5%)
+ Active / Total Size (% used) : 1120400.00K / 1210500.00K (92.6%)
OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME
-10450120 10450000 99% 0.20K 550006 19 8800096K dentry
- 5210400 5210000 99% 0.61K 200400 26 3206400K ext4_inode_cache
+ 152400 151000 99% 0.20K 8021 19 128336K dentry
+ 84100 82000 97% 0.61K 3234 26 51744K ext4_inode_cache
310200 305000 98% 2.00K 38775 8 620400K kmalloc-2k
140100 138000 98% 1.00K 8756 16 140096K kmalloc-1k
Line-by-Line Explanation
- Header Diff: Total slab memory dropped dramatically from 14.2 GiB to 1.12 GiB, freeing over 13 GiB of physical memory back to the general pool.
dentryDiff: Directory cache objects dropped from 10.45 million (8.8 GiB) down to 152,400 (128 MiB)—confirming that 8.6 GiB of directory metadata was cleanly reclaimed.ext4_inode_cacheDiff: File metadata structures shrank from 5.21 million (3.2 GiB) down to 84,100 (51.7 MiB).- Retained Allocations: The remaining 152,400
dentryobjects represent active mount points, open file descriptors, and current working directories held by running applications. These are in active use and cannot be evicted until those processes close.
What the Admin Does Next
With the memory successfully released and contiguous page frames restored, the admin can proceed with initializing the memory-intensive workload without risk of immediate allocation throttling.
Operational Hazards: What Can Go Wrong
1. Lock Contention and Monitoring Overhead on Large Multi-Socket Servers
On large multi-socket NUMA (Non-Uniform Memory Access) servers with hundreds of CPU cores and terabytes of RAM, polling /proc/slabinfo or running slabtop with rapid refresh intervals (e.g. -d 0.5) can create serious CPU bottlenecks.
(Lock Acquisition Overhead)"] Cache1 --> Lock
To aggregate statistics across the entire system, the kernel must iterate over every per-CPU and per-node cache, acquiring internal locks along the way. On heavily loaded systems, polling slab metrics too aggressively can stall active workloads.
slabtop or /proc/slabinfo faster than every 30 to 60 seconds. For interactive terminal inspection, stick to the default 3-second refresh delay.2. Misinterpreting Normal Caching as a Memory Leak
The most common mistake sysadmins make with kernel memory is panicking when they see large dentry or inode_cache values.
Linux is intentionally designed to put unused RAM to work. If you have 64 GiB of memory and your applications only need 16 GiB, the kernel will use the remainder to cache directory trees and file metadata so that future disk access is lightning-fast.
Running drop_caches repeatedly in a cron job to keep memory "free" hurts performance, forcing the system to read metadata from slow physical storage over and over again.
SReclaimable (dentry, inode), the server is doing its job; the kernel will free that space automatically when apps need it. But if memory is dominated by SUnreclaim or generic kmalloc pools that refuse to shrink, you are dealing with a genuine kernel leak.3. Hardened Permissions on /proc/slabinfo
On security-hardened Linux distributions (such as systems with Grsecurity patches, enforcing SELinux policies, or kernel.kptr_restrict configured), unprivileged users cannot view slab details:
fopen /proc/slabinfo: Permission denied
To allow monitoring tools or junior engineers to triage kernel memory without granting full, unrestricted root access, grant specific permissions via Linux capabilities or targeted sudo rules:
# Check slabinfo access permissions
ls -la /proc/slabinfo
# Run with elevated privileges or dedicated sudo rules
sudo slabtop -s c -o
Today's Takeaway
Open a terminal on your machine right now and run:
sudo slabtop -s c -o | head -n 20
Look at the top two lines under the column headings. You will almost certainly see dentry, buffer_head, or your filesystem's inode cache sitting at the top.
Take note of the total size in the header and compare it with the output of free -h. In five minutes, you have verified exactly how much of your operating system's memory is quietly accelerating your file access behind the scenes—and you now have the exact command you need when memory seemingly vanishes into thin air.