Systemd-cgtop: Profiling Real-Time Control Group Resource Saturation, Auditing Container Slices, and Triaging Multi-Tenant Throttling in Production
For decades, systems administrators have relied on diagnostic tools that view the operating system as an undifferentiated, flat soup of individual processes. When a single program went rogue in the 1990s, top or htop would point an unambiguous finger at its process identifier (PID). But modern Linux deployments no longer operate as single, isolated applications. Today, a production server is a bustling metropolis of container runtimes, system services, virtual machines, and background maintenance workers, often sharing generic names like python3, node, java, or envoy.
When a microservice misbehaves by spawning hundreds of lightweight workersβeach quietly nibbling an unremarkable 0.4% of total processor capacityβtraditional process monitors report that everything is fine. In reality, that swarm of workers is collectively starving your database of disk bandwidth and overwhelming the kernel's scheduler. Traditional process-level tools suffer from structural blindness: they cannot see the forest for the trees.
To regain control during an outage, you need an instrument that steps back from individual threads and aggregates compute, memory, and storage metrics at the architectural boundary of services, containers, and user sessions. That tool is systemd-cgtop.
For an immediate, system-wide health check that cuts through the noise and ranks every service by real-time processor consumption, run:
systemd-cgtop -c -d 1
Control Group Tasks %CPU Memory Input/s Output/s
/ 1842 74.2 14.2G 4.1M 12.8M
system.slice 812 52.1 9.1G 3.8M 11.2M
system.slice/docker.service 410 38.4 4.2G 2.1M 8.4M
machine.slice 620 18.0 3.8G 256.0K 1.2M
machine.slice/libvirt-qemu@storage01.service 128 14.2 3.1G 180.0K 1.1M
user.slice 410 4.1 1.3G 4.0K 40.0K
Within two seconds, the fog clears. Across 1,842 individual tasks, the machine's resources are laid bare: the Docker engine inside system.slice is responsible for 38.4% of CPU consumption and generating 8.4 MB/s of disk writes, while virtual machines under machine.slice account for another 18% of CPU load.
What It Does in Plain English
Think of a modern Linux server as an apartment block with a shared water, gas, and electricity supply. If one tenant runs a commercial laundry business in their flat, the water pressure drops for everyone else. If your utility company only monitors the water meters on individual taps, finding out why the fourth floor has no water requires checking hundreds of faucets one by one. Control Groups (cgroups) put a master utility meter on each apartment's main supply pipe.
systemd-cgtop is the dashboard that reads those master meters in real time. Instead of tracking thousands of individual processes, it groups resources according to systemdβs service hierarchy. It tallies up processor load, memory consumption, task counts, and disk input/output rates across high-level system services (system.slice), virtual machines and container runtimes (machine.slice), and logged-in user sessions (user.slice).
Core Flags & Navigation
The utility provides intuitive keyboard shortcuts and command-line switches for adjusting refresh rates, sorting metrics, and filtering the hierarchy:
| Flag | Long Option | Operational Description |
|---|---|---|
-c |
--order=cpu |
Sort the control group hierarchy by aggregate CPU load (default). |
-m |
--order=memory |
Sort the control group hierarchy by current memory footprint. |
-i |
--order=io |
Sort the control group hierarchy by combined disk read and write throughput. |
-t |
--order=tasks |
Sort the control group hierarchy by active process and thread count. |
-p |
--order=path |
Sort the control group hierarchy alphabetically by canonical cgroup path. |
-d SEC |
--delay=SEC |
Specify the sampling and screen refresh interval in seconds (e.g. 0.5). |
-n INT |
--iterations=INT |
Exit automatically after a set number of sampling iterations. |
-b |
--batch |
Run in non-interactive batch mode without terminal escape codes. |
-r |
--raw |
Output numeric metrics as raw, unscaled byte and counter values. |
--depth=INT |
--depth=INT |
Restrict hierarchical tree traversal depth to suppress noise. |
While systemd-cgtop is running interactively, you can switch sorting modes on the fly by pressing C (CPU), M (Memory), I (I/O), T (Tasks), or P (Path) on your keyboard. Pressing + or - adjusts the refresh rate dynamically, while q exits the interface.
Architectural Foundations: Slices, Trees, and Unified Accounting
To interpret systemd-cgtop effectively, it helps to understand how the modern Linux kernel organises resources under the Linux Kernel Control Group v2 Architecture.
In legacy cgroups v1, every resource typeβCPU, memory, block I/O, and process IDsβlived in its own isolated filesystem under /sys/fs/cgroup/<controller>/. This uncoordinated design meant memory writebacks could not be accurately linked to the I/O controller, and process classification required fragile manual synchronisation across multiple directories.
The unified cgroups v2 hierarchy resolves this by placing every running process into a single tree rooted at /sys/fs/cgroup/. Internal nodes enforce resource distribution policies, while leaf nodes house the active processes.
(-.slice)"] --> SystemSlice["system.slice
(Core Daemons & Background Services)"] Root --> MachineSlice["machine.slice
(Virtual Machines & Container Pods)"] Root --> UserSlice["user.slice
(Interactive User Sessions)"] SystemSlice --> Nginx["system.slice/nginx.service"] SystemSlice --> Postgres["system.slice/postgresql-16.service"] MachineSlice --> PodA["machine.slice/libpod-a8f1b...service"] MachineSlice --> PodB["machine.slice/qemu-101-analytics.scope"] UserSlice --> UserSession["user.slice/user-1000.slice"]
The systemd init system acts as the single manager of this unified hierarchy, grouping services into distinct resource partitions termed slices via systemd.slice(5):
-.slice: The global root slice containing all managed resources across the entire operating system.system.slice: The default partition for background system daemons, OS-level microservices, and systemd service units.user.slice: The partition assigned to interactive user sessions, desktop environments, and per-user daemon instances.machine.slice: The dedicated operational partition automatically instantiated for virtual machines (KVM/QEMU via libvirt) and container runtimes (Docker, Podman, containerd, andsystemd-nspawn).
systemd-cgtop gathers its data by sampling the kernel's pseudo-filesystem interfaces inside these cgroup directories. CPU usage is calculated from differential deltas in cpu.stat (usage_usec). Memory usage is parsed from memory.current (which includes anonymous memory, swap, and active page caches). Block I/O read and write rates are computed by differential sampling of byte counters in io.stat. Furthermore, the unified hierarchy supports Linux Kernel Pressure Stall Information (PSI), tracking the exact percentage of wall-clock time that tasks within a slice spend stalled waiting for CPU cycles, memory pages, or disk I/O.
Five Real-World Production Use Cases
1. Real-Time CPU Profiling with Sub-Second Sampling to Isolate Bursting Microservices
The Scenario
An unpredictable spike in CPU load is degrading API responsiveness. A microservice inside system.slice is executing high-frequency, short-lived compute bursts that finish too quickly to register on standard five-second interval monitoring tools.
The Command
systemd-cgtop -c -d 0.2 --depth=3
Realistic Terminal Output
Control Group Tasks %CPU Memory Input/s Output/s
system.slice 940 184.2 8.2G 1.1M 2.4M
system.slice/api-gateway.service 64 142.6 2.1G 890.0K 1.8M
system.slice/api-gateway.service/worker.slice 48 138.4 1.8G 850.0K 1.7M
system.slice/postgresql-16.service 32 28.1 4.2G 120.0K 540.0K
system.slice/systemd-journald.service 1 8.2 140.2M 0.0B 48.0K
system.slice/prometheus-node-exporter.service 4 2.1 42.0M 0.0B 0.0B
Line-by-Line Breakdown
- The
-d 0.2flag sets a 200-millisecond sampling frequency, capturing transient CPU spikes before they get smoothed out by time averaging. - The
--depth=3parameter truncates tree traversal at the third hierarchical level, keeping deep sub-cgroup noise off the screen. system.slice/api-gateway.serviceis drawing 142.6% CPU (representing roughly 1.4 fully saturated CPU cores on a multi-core system).- The sub-slice
worker.sliceaccounts for 138.4% of that consumption, confirming that the load stems from child worker threads rather than the parent supervisor process.
What the Admin Does Next
Apply a dynamic resource constraint on the offending unit without restarting the service, using systemd Resource Management Directives:
sudo systemctl set-property api-gateway.service CPUQuota=100% CPUWeight=50
This enforces a hard cap restricting the service to a maximum of one CPU core equivalent per scheduling period while reducing its CPU scheduling weight from the default of 100 down to 50 during contentious periods.
2. Triaging Multi-Tenant Container Memory Saturation across machine.slice
The Scenario
A multi-tenant host running multiple containerised workloads experiences unexpected Out-Of-Memory (OOM) killer terminations. The administrator must pinpoint which specific container is consuming excessive memory before critical production pods are terminated.
The Command
systemd-cgtop -m --depth=2
Realistic Terminal Output
Control Group Tasks %CPU Memory Input/s Output/s
machine.slice 842 42.1 28.6G 3.4M 8.9M
machine.slice/libpod-a8f1b2c3d4e5f678...service 128 12.4 14.2G 1.2M 4.1M
machine.slice/libpod-3b9c0d1e2f3a4b5c...service 64 8.1 8.4G 800.0K 2.2M
machine.slice/libpod-9e8d7c6b5a4f3e2d...service 32 2.0 4.1G 240.0K 910.0K
system.slice 412 18.2 2.4G 410.0K 1.1M
user.slice 18 0.1 410.2M 0.0B 0.0B
Line-by-Line Breakdown
- Sorting by memory (
-m/--order=memory) exposes absolute memory occupancy (combining heap, anonymous allocations, and mapped page buffers). - The container identified by
machine.slice/libpod-a8f1b2c3d4e5...dominates host memory, holding 14.2 GB of the 28.6 GB allocated tomachine.slice. - This container is dangerously close to triggering a host-wide memory exhaustion event for neighbouring workloads.
What the Admin Does Next
Inspect the container's active memory limits and memory pressure metrics directly via the cgroup filesystem:
cat /sys/fs/cgroup/machine.slice/libpod-a8f1b2c3d4e5*.service/memory.pressure
cat /sys/fs/cgroup/machine.slice/libpod-a8f1b2c3d4e5*.service/memory.max
To prevent host-wide memory depletion while avoiding an immediate hard kill, apply a high-watermark throttling limit:
sudo systemctl set-property libpod-a8f1b2c3d4e5f678.service MemoryHigh=12G MemoryMax=15G
Under cgroups v2, exceeding MemoryHigh does not trigger the OOM killer; instead, the kernel throttles the offending cgroup's running processes and aggressively reclaims its page caches.
3. Auditing Continuous Block I/O Saturation on High-Density Virtualised Storage Nodes
The Scenario
Storage volumes on a shared virtualisation hypervisor are experiencing severe I/O queue wait states. While overall CPU load is normal, virtual machines hosted under machine.slice suffer latency degradation due to an unconstrained workload performing continuous sequential writes.
The Command
systemd-cgtop -i
Realistic Terminal Output
Control Group Tasks %CPU Memory Input/s Output/s
machine.slice/qemu-101-vm-analytics.scope 16 14.0 8.0G 4.2K 184.2M
machine.slice/qemu-102-vm-webcore.scope 32 22.1 16.0G 142.0K 1.4M
system.slice/systemd-journald.service 1 1.2 180.4M 0.0B 820.0K
system.slice/prometheus.service 8 3.4 1.8G 412.0K 210.0K
user.slice/user-1000.slice 12 0.0 210.0M 0.0B 0.0B
Line-by-Line Breakdown
- Sorting by I/O throughput (
-i/--order=io) aggregates real-time block storage read and write rates. machine.slice/qemu-101-vm-analytics.scopeis writing data at a rate of 184.2 MB/s (Output/s), saturating the shared drive controller.- In contrast, mission-critical infrastructure (
qemu-102-vm-webcore.scope) is restricted to minimal I/O bandwidth, causing latency spikes for connected clients.
What the Admin Does Next
Limit the analytics VM's aggregate block I/O throughput at the hypervisor level via systemctl using block device major/minor paths:
# Identify the backing block device
ls -l /dev/disk/by-id/nvme-eui.*
# Enforce a 50 MB/s write limit on NVMe block device /dev/nvme0n1
sudo systemctl set-property qemu-101-vm-analytics.scope IOWriteBandwidthMax="/dev/nvme0n1 50M"
This immediately updates the cgroup v2 controller interface io.max, enforcing kernel-level throttling on the analytics VM and restoring disk bandwidth for the web cluster.
4. Non-Interactive Headless Telemetry Ingestion for Automated Cron Audits
The Scenario
An automated auditing script executed via cron must periodically capture raw, unformatted cgroup metrics to evaluate whether batch maintenance tasks running inside system.slice have exceeded compliance thresholds.
The Command
systemd-cgtop -b -n 1 --raw > /var/log/audit/cgroup-snapshot-$(date +%s).raw
Realistic Terminal Output (Raw Batch Format)
Path Tasks CPU Memory Input Output
/ 1842 742000000 15247187968 4300120 13421772
system.slice 812 521000000 9768249344 3984588 11744051
system.slice/docker.service 410 384000000 4509715660 2202009 8808038
machine.slice 620 180000000 4080218931 262144 1258291
user.slice 410 41000000 1395864371 4096 40960
Line-by-Line Breakdown
- The
-b(--batch) flag disables interactive ANSI control characters and curses rendering, producing clean tabular ASCII streams suitable for stream editors likeawk,sed, or custom shell scripts. - The
-n 1flag instructssystemd-cgtopto perform a single measurement cycle and exit immediately with status code 0. - The
--rawflag outputs exact, unscaled numeric counters: - CPU: Nanoseconds or microseconds of cumulative runtime.
- Memory: Exact byte totals (e.g.
15247187968bytes $\approx 14.2$ GiB). - Input / Output: Raw data transfer counts in bytes per second.
What the Admin Does Next
Integrate the batch snapshot into a lightweight alerting script located at /usr/local/bin/cgroup-memory-check.sh:
#!/usr/bin/env bash
set -euo pipefail
# Ingest current batch metrics
SNAPSHOT=$(systemd-cgtop -b -n 1 --raw)
# Extract memory consumed by system.slice in bytes
SYSTEM_MEM_BYTES=$(echo "$SNAPSHOT" | awk '$1 == "system.slice" {print $4}')
THRESHOLD_BYTES=$(( 16 * 1024 * 1024 * 1024 )) # 16 GiB Limit
if [ "$SYSTEM_MEM_BYTES" -gt "$THRESHOLD_BYTES" ]; then
logger -p user.crit "CRITICAL: system.slice memory footprint ($SYSTEM_MEM_BYTES bytes) exceeds threshold ($THRESHOLD_BYTES bytes)."
/usr/local/bin/trigger-pagerduty-alert.sh "Memory threshold breach in system.slice"
fi
5. Auditing Thread Pool Proliferation and Fork Bomb Prevention via Task Limits
The Scenario
A microservice running an internal thread pool contains a bug that spawns orphan worker threads upon every unhandled HTTP timeout. Over several hours, total thread consumption rises, threatening to exhaust the kernel process table limit (/proc/sys/kernel/pid_max) and causing system-wide fork failures.
The Command
systemd-cgtop -t --depth=3
Realistic Terminal Output
Control Group Tasks %CPU Memory Input/s Output/s
system.slice/payment-processor.service 14820 12.4 2.8G 4.0K 12.0K
system.slice/docker.service 410 18.1 3.2G 120.0K 410.0K
machine.slice 320 14.2 4.1G 210.0K 820.0K
system.slice/postgresql-16.service 32 2.1 2.1G 12.0K 84.0K
user.slice 18 0.0 140.0M 0.0B 0.0B
Line-by-Line Breakdown
- Sorting by task count (
-t/--order=tasks) lists cgroups by total instantiated execution contexts (both primary processes and POSIX threads managed by the kernel scheduler). system.slice/payment-processor.servicehas accumulated 14,820 active tasks, consuming a massive proportion of available system thread handles despite maintaining modest CPU utilisation (12.4%).- Without immediate intervention, this single unit will trigger host-wide
EAGAIN: Resource temporarily unavailableerrors on all subsequentfork()andclone()syscalls across unrelated services.
What the Admin Does Next
Mitigate the runaway thread allocation instantly via the TasksMax control parameter:
# Impose an immediate ceiling of 512 total concurrent tasks on the unit
sudo systemctl set-property payment-processor.service TasksMax=512
The systemd controller immediately writes 512 to /sys/fs/cgroup/system.slice/payment-processor.service/pids.max. Any subsequent thread allocation attempts beyond this limit fail gracefully within the boundary of the payment-processor cgroup, preserving host stability.
What Can Go Wrong: Architectural Pitfalls & Measurement Quirks
Operating systemd-cgtop in high-throughput environments requires awareness of several operational nuances:
Pitfall 1: Ghost Slices and Missing Accounting Controllers
On older Linux installations or systems operating in hybrid cgroups v1 mode, running systemd-cgtop may output empty dashes across %CPU, Memory, and Input/Output columns:
Control Group Tasks %CPU Memory Input/s Output/s
system.slice/nginx.service 12 - - - -
This occurs when accounting controllers are disabled in the unit configuration to minimise kernel scheduling overhead.
Recovery and Prevention
Verify unified cgroup v2 status and ensure accounting directives are enabled globally in /etc/systemd/system.conf or locally within the target service unit file:
[Service]
CPUAccounting=yes
MemoryAccounting=yes
IOAccounting=yes
TasksAccounting=yes
Apply configuration changes immediately:
sudo systemctl daemon-reload
Pitfall 2: Sampling Frequency Jitter and Kernel Overhead
Invoking systemd-cgtop with sub-millisecond refresh rates (e.g. -d 0.01) on servers hosting tens of thousands of active cgroups forces the utility to continuously traverse and parse thousands of virtual files in /sys/fs/cgroup/. This introduces measurable kernel CPU overhead and lock contention within the virtual file system.
Mitigation
In production environments with deep cgroup hierarchies, restrict the sampling interval to $\ge 1.0$ second (-d 1) and constrain hierarchical traversal depth using --depth=2 or --depth=3.
Pitfall 3: Metric Incoherence from Ephemeral Cgroups
Short-lived containers or batch jobs that launch, execute, and terminate within a window shorter than the sampling interval (such as serverless functions finishing in 50 milliseconds) may never register on the live display of systemd-cgtop. Their resource consumption is aggregated into the parent slice upon termination, causing sudden, unexplained spikes in slice-level metrics without corresponding active child units.
Mitigation
For microsecond-level ephemeral execution tracing, complement systemd-cgtop with eBPF-based instrumentation (such as bpftrace or execsnoop).
Reference Manuals & Authoritative Resources
For deeper exploration of the Linux kernel control group architecture and systemd slice management, consult the following documentation:
- systemd-cgtop(1) β Linux Manual Pages
- Control Group v2 Official Linux Kernel Documentation
- systemd.resource-control(5) β Resource Control Settings
- Linux Kernel Documentation: Pressure Stall Information (PSI)
- Arch Linux Control Groups Architecture & Configuration Guide
- systemd.slice(5) β Slice Unit Configuration
Today's Takeaway
The single most valuable operational habit you can establish today is to replace the reflex of running flat process monitors with hierarchical cgroup inspection during your initial triage phase. Open an active terminal session on your primary workstation or staging server and execute systemd-cgtop -m --depth=2. Within five seconds, you will observe the precise memory distribution of your machine broken down by system services, user workspaces, and container enginesβrevealing the structural architecture of your Linux operating system rather than an unorganised list of processes.