Systemd-analyze: Profiling Boot Latency Bottlenecks, Tracing Dependency Critical Chains, and Auditing Unit Security Exposure in Production
You SSH into one of the freshly launched virtual machines expecting to find runaway processes consuming memory or CPUs pegged at 100%. Instead, you are met with an eerie, baffling calm. Processor utilization sits near zero, memory is practically untouched, yet the essential reverse proxies and payment services have not even begun listening on their network ports. Traditional diagnostic commands like top show an idle system, while scrolling through thousands of lines in kernel logs reveals no obvious crash. The machine is alive, but its startup process is trapped in an invisible traffic jam.
When an operating system stalls during boot, searching blindly through log files is like trying to find a blown fuse in a skyscraper without a floor plan. To restore your systems before the outage breaches customer service agreements, you need immediate, second-by-second visibility into every phase of the initialization sequence. You need to know exactly which process started, which dependencies it waited on, and why your services were prevented from launching. This is where systemd-analyze becomes an indispensable tool.
To instantly cut through the confusion and pinpoint the exact bottleneck holding up your system, run the single most powerful diagnostic command in the systemd suite:
$ systemd-analyze critical-chain
In less than a second, systemd-analyze critical-chain reconstructs the entire sequence of startup events, tracing the critical path backwards from your application down to the core operating system. Instead of guessing whether a slow disk mount, a lagging cloud metadata query, or an unconfigured network interface delayed your service, you receive a clear, hierarchical timeline that highlights precisely which background task stalled the boot pipeline and how many seconds it consumed.
What It Does in Plain English
At its core, systemd-analyze is a native diagnostic and profiling utility built directly into systemd, the initialization system that manages services and resources on modern Linux distributions. When a Linux machine powers on, it does not simply execute startup scripts one after another in a straight line. Instead, it launches dozens of tasks in parallelβconfiguring network adapters, mounting filesystems, setting up security boundaries, and starting background daemons.
systemd-analyze acts as an x-ray machine for this complex startup process. It translates the internal, highly concurrent startup lifecycle of a Linux host into human-readable performance metrics, structured execution trees, and visual timelines.
Beyond measuring raw boot speed, the utility serves two other vital functions in production environments: 1. Static Validation: It acts as an offline validator for service configuration files, catching syntax errors, missing binaries, and broken dependencies before they are deployed to live servers. 2. Security Auditing: It evaluates the security posture of running system services against modern Linux kernel sandboxing features, assigning each service a clear risk rating based on how well it is isolated from the rest of the operating system.
Theoretical Foundations and System Architecture
Modern Linux systems orchestrated by systemd abandon the sequential, imperative shell-script model historically used by SysV init in favor of an aggressively parallelized, event-driven directed acyclic graph (DAG). To understand how systemd-analyze inspects this environment, we must examine the state machinery inside PID 1, the D-Bus serialization protocols, and the kernel primitives underpinning service sandboxing.
(Storage, Crypto, Local Mounts)"] --> S2["basic.target
(Sockets, Timers, IPC Channels)"] S2 --> S3["default.target / multi-user.target
(Network, Web Proxies, Applications)"] end S1 -.->|"Monotonic Timestamps"| DBUS["org.freedesktop.systemd1.Manager"] S2 -.->|"Monotonic Timestamps"| DBUS S3 -.->|"Monotonic Timestamps"| DBUS subgraph ToolSpace["systemd-analyze Diagnostic Engine"] DBUS -->|"D-Bus IPC Socket"| SA["systemd-analyze"] SA --> SA1["time & blame
(Latency Profiling)"] SA --> SA2["critical-chain
(Dependency Path Tracing)"] SA --> SA3["plot
(SVG Chronograph)"] SA --> SA4["security
(Kernel Sandbox Auditing)"] SA --> SA5["verify & condition
(Static Validation)"] end
1. Directed Acyclic Graph (DAG) Orchestration and Target Milestones
During system initialization, systemd treats units (services, mounts, sockets, and targets) as vertices within a directed graph, with edges defined by relational and ordering constraints. The relational directivesβsuch as Requires=, Wants=, BindsTo=, and PartOf=βdetermine dependency sets, whereas ordering directivesβpredominantly Before= and After=βdictate scheduling order.
PID 1 structures boot progression around landmark synchronization milestones known as targets:
1. sysinit.target: Mounts early low-level filesystems, activates cryptographic storage volumes, initializes keyrings, enables swap space, and synchronizes the system clock.
2. basic.target: Establishes fundamental operating environments, activating listening sockets, IPC message buses, and system timer events.
3. default.target (typically linked to multi-user.target on servers or graphical.target on workstations): Coordinates application daemons, reverse proxies, container runtimes, and user workloads.
Because systemd launches services concurrently whenever their prerequisites are met, total boot duration is governed by the longest uninterrupted sequential path through the DAG: the critical path.
2. In-Memory State Serialization and D-Bus IPC Interrogation
The init system tracks state transitions internally via high-resolution monotonic timestamps (CLOCK_MONOTONIC). Every unit moves through a deterministic lifecycle state machine:
$$\text{inactive} \longrightarrow \text{activating} \longrightarrow \text{active} \longrightarrow \text{deactivating} \longrightarrow \text{failed}$$
During these transitions, PID 1 records nanosecond-accurate timestamps corresponding to:
- InactiveExitTimestamp: When service activation was scheduled.
- ActiveEnterTimestamp: When the unit successfully signaled readiness.
- ActiveExitTimestamp: When deactivation commenced.
- InactiveEnterTimestamp: When cleanup finalized.
systemd-analyze does not parse text logs to infer these metrics. Instead, it opens a UNIX domain socket connection to /run/systemd/private or communicates over the system D-Bus broker to query the org.freedesktop.systemd1.Manager interface. Through methods such as GetUnitByPID and property queries against the org.freedesktop.systemd1.Unit interface, systemd-analyze computes the precise delta of process forks, I/O waits, and synchronous state handshakes.
3. Sandboxing Evaluation Against Kernel Primitives
The systemd-analyze security subcommand audits unit configurations against Linux kernel isolation boundaries. It evaluates declarative directives against a weighted risk matrix spanning several kernel primitives:
- Namespaces (
CLONE_NEW*): Validates filesystem and isolation boundaries through directives such asProtectSystem=strict(mounts/usr,/boot, and/etcread-only via mount namespaces),ProtectHome=yes(renders/root,/home, and/run/userinaccessible), andPrivateTmp=yes(allocates an isolated file system namespace for/tmpand/var/tmp). - Control Groups (cgroups v2): Evaluates resource constraint boundaries and memory/PID limits that prevent fork-bomb attacks and memory exhaustion.
- Seccomp Filters (Secure Computing Mode): Examines
SystemCallFilter=directives to determine whether a service is restricted to an approved list of syscalls compiled into a Berkeley Packet Filter (BPF) program, preventing execution of high-risk interfaces likeptrace,kexec_load, or raw socket creation. - POSIX Capabilities: Assesses
CapabilityBoundingSet=andAmbientCapabilities=to ensure processes execute without ambient root privileges, stripping dangerous capabilities likeCAP_SYS_ADMIN,CAP_NET_ADMIN, andCAP_DAC_OVERRIDE. - Privilege Escalation Controls: Audits
NoNewPrivileges=yes, which enforces thePR_SET_NO_NEW_PRIVSflag viaprctl(2)to ensure child processes cannot acquire elevated privileges via setuid/setgid binaries.
systemd-analyze security compiles these checks into an Exposure Score from $0.0$ (hyper-secure and isolated) to $10.0$ (unrestricted root exposure).
Core Subcommand Matrix & Quick Start
The table below outlines the primary subcommands available within systemd-analyze:
| Subcommand | Primary Operational Objective | How It Works |
|---|---|---|
time |
Print overall boot and initialization duration. | Queries kernel, initramfs, and userspace monotonic clock deltas from PID 1. |
blame |
List individual unit startup delays in descending order. | Calculates the time delta ($\Delta t = \text{ActiveEnter} - \text{InactiveExit}$) for each unit. |
critical-chain |
Trace the sequential chain of dependencies holding up a target. | Walks backwards through the directed acyclic graph to find blocking services. |
plot |
Generate a visual SVG timeline of system orchestration. | Serializes start, active, and exit timestamps into a graphical vector chronograph. |
security |
Audit unit security and sandboxing settings. | Evaluates directives against weighted kernel namespace, seccomp, and capability matrices. |
verify |
Perform offline syntax and dependency validation. | Statically parses unit configuration files to detect typos, syntax errors, and missing binaries. |
condition |
Test conditional startup rules in real time. | Evaluates conditional expressions against the current host environment. |
Quick Start: Macroscopic Host Telemetry
To establish the overarching boot latency baseline on any Linux host, invoke systemd-analyze time:
$ systemd-analyze time
Output:
Startup finished in 1.452s (kernel) + 2.180s (initrd) + 8.643s (userspace) = 12.275s
graphical.target reached after 8.612s in userspace
This output separates initialization into three distinct phases:
- kernel: Time elapsed from bootloader handover until kernel initialization completes and /init executes.
- initrd: Time spent within the initial ramdisk (loading storage drivers, decrypting LUKS encrypted volumes, and executing pivot_root(2)).
- userspace: Time required from PID 1 execution until the primary target milestone (graphical.target or multi-user.target) completes its transactional activation.
5 Everyday Real-World Production Use Cases
Use Case 1: Triage Cloud VM Cold-Start Latency in Auto-Scaling Fleets
The Scenario
An auto-scaling fleet of compute instances in an AWS or GCP infrastructure pool is exhibiting severe cold-start delays. New instances require over 90 seconds to join the cluster load balancer during traffic spikes. We need to isolate which service initialization sequences are consuming execution time during cloud-init provisioning.
The Command
$ systemd-analyze blame | head -n 12
Terminal Output
38.412s cloud-init.service
18.231s cloud-config.service
12.104s cloud-init-local.service
4.512s dkms.service
3.120s systemd-networkd-wait-online.service
1.845s lvm2-monitor.service
1.102s systemd-udev-settle.service
0.890s dev-nvme0n1p1.device
0.421s systemd-journal-flush.service
0.312s polkit.service
0.180s systemd-logind.service
0.150s ssh.service
Line-by-Line Technical Analysis
38.412s cloud-init.service: The primarycloud-initmodule stalled for over 38 seconds while running userdata scripts and pulling external instance metadata.18.231s cloud-config.service: The second-stage cloud configuration engine spent 18 seconds executing blocking package repository updates and SSH key distributions.12.104s cloud-init-local.service: Early disk discovery and local networking verification consumed 12 seconds.4.512s dkms.service: Dynamic Kernel Module Support attempted to verify and rebuild kernel module headers on every single boot.3.120s systemd-networkd-wait-online.service: The network waiter blocked downstream services until all detected network interfaces achieved an operational carrier state.1.102s systemd-udev-settle.service: An obsolete synchronization service forced the system to wait until all hardware device discovery events were dispatched.
Remediation and Action
The systems engineer recognizes that dynamic kernel compilation (dkms.service) and udev-settle are completely redundant in pre-baked, immutable cloud virtual machine images. Furthermore, cloud-init was configured to run unneeded repository index updates on every startup.
The engineer masks the unnecessary services and updates /etc/cloud/cloud.cfg to eliminate redundant modules:
$ sudo systemctl mask systemd-udev-settle.service
$ sudo systemctl mask dkms.service
Disabling these legacy tasks reduces VM boot duration by more than 40 seconds across the autoscaling pool.
Use Case 2: Resolve Critical-Chain Dependency Bottlenecks Delaying Web Proxies
The Scenario
In a high-availability infrastructure environment, an edge NGINX reverse proxy (nginx.service) fails to start for nearly 45 seconds after a server reboot. The operations team suspects network issues, but standard log viewers fail to show which upstream dependency in the graph is holding up the web server.
The Command
$ systemd-analyze critical-chain nginx.service
Terminal Output
The time when unit became active or started is printed after the "@" character.
The time the unit took to start is printed after the "+" character.
nginx.service @44.812s +120ms
network-online.target @44.680s
systemd-networkd-wait-online.service @14.210s +30.450s
systemd-networkd.service @13.890s +310ms
systemd-udevd.service @4.120s +1.230s
systemd-sysusers.service @3.850s +250ms
sysinit.target @3.780s
systemd-journal-flush.service @1.210s +2.550s
var-log.mount @1.120s +80ms
local-fs.target @1.100s
Line-by-Line Technical Analysis
nginx.service @44.812s +120ms: NGINX took only 120 milliseconds to initialize its worker pools, but activation was blocked until 44.8 seconds after boot.network-online.target @44.680s: NGINX configuredAfter=network-online.target, deferring its execution until the operating system signaled complete network availability.systemd-networkd-wait-online.service @14.210s +30.450s: This was the critical bottleneck. The service stalled the boot process for 30.45 seconds before timing out.systemd-networkd.service @13.890s +310ms: The core network daemon started cleanly in just 310ms.- The remaining chain confirms that foundational storage (
local-fs.target,sysinit.target) initialized rapidly in under 4 seconds.
Remediation and Action
The system has multiple physical network interfaces configured via systemd-networkd, but a secondary interface (such as eth1) is disconnected and lacks a physical carrier link. systemd-networkd-wait-online.service was waiting for all configured interfaces to establish a connection before releasing network-online.target.
The engineer overrides the wait service configuration using a drop-in configuration file:
$ sudo systemctl edit systemd-networkd-wait-online.service
The engineer specifies that the waiter should only monitor the primary active interface:
[Service]
ExecStart=
ExecStart=/usr/lib/systemd/systemd-networkd-wait-online -i eth0 --operational-state=routable
After applying this change, systemd-networkd-wait-online.service completes in under 400ms, allowing nginx.service to come online within 5 seconds of host boot.
Use Case 3: Zero-Trust Unit Security Audit and Remediation
The Scenario
A security compliance mandate requires that all custom backend microservices running on production bare-metal infrastructure adhere to a strict Zero-Trust containment profile. We must audit backend-worker.service to identify privilege escalation risks and apply kernel-level sandboxing.
The Command
$ systemd-analyze security backend-worker.service
Terminal Output
NAME DESCRIPTION EXPOSURE
[!] PrivateNetwork= Service has full access to host's network 0.5
[!] User=/DynamicUser= Service runs as root user 0.8
[!] CapabilityBoundingSet= Service has all POSIX capabilities assigned 0.7
[!] ProtectHome= Service has full access to home directories 0.4
[!] ProtectSystem= Service has full access to the OS file hierarchy 0.6
[!] SystemCallFilter= Service may execute arbitrary system calls 0.8
[!] PrivateTmp= Service shares /tmp with other processes 0.4
[!] ProtectKernelTunables= Service may alter kernel sysctls 0.2
[!] ProtectControlGroups= Service may alter cgroup hierarchies 0.2
[!] RestrictAddressFamilies= Service may allocate any socket address family 0.3
[!] NoNewPrivileges= Service processes may acquire new privileges 0.5
-> Overall exposure level for backend-worker.service: 9.2 UNSAFE
Line-by-Line Technical Analysis
User=/DynamicUser= (0.8): The service runs unconfined as UID 0 (root).SystemCallFilter= (0.8): The process has unrestricted access to all 400+ Linux system calls, including dangerous memory manipulation and debugging hooks.CapabilityBoundingSet= (0.7): The process retains broad Linux capabilities, including administrative privileges and raw filesystem access.ProtectSystem= (0.6)&ProtectHome= (0.4): Root filesystems (/usr,/etc,/home) are mounted read-write to the process.NoNewPrivileges= (0.5): The process can leverage SUID binaries to escalate privileges.Overall exposure level: 9.2 UNSAFE: The unit represents an unconfined attack surface vulnerable to system takeover if compromised.
Remediation and Action
The engineer drafts a security hardening drop-in configuration via systemctl edit backend-worker.service, implementing directives governed by systemd.exec(5):
[Service]
# Enforce execution under dynamic unprivileged UID/GID
DynamicUser=yes
# Filesystem Sandboxing via Mount Namespaces
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
ReadWritePaths=/var/log/backend-worker
# Kernel Sandboxing & Seccomp BPF Filters
NoNewPrivileges=yes
CapabilityBoundingSet=
AmbientCapabilities=
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
RestrictNamespaces=yes
RestrictRealtime=yes
MemoryDenyWriteExecute=yes
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources @debug
The engineer reloads the unit and audits the security posture once more:
$ sudo systemctl daemon-reload
$ sudo systemctl restart backend-worker.service
$ systemd-analyze security backend-worker.service
-> Overall exposure level for backend-worker.service: 1.4 OK
The microservice is now safely sandboxed inside an unprivileged namespace and protected by seccomp filters.
Use Case 4: Automating Visual Boot Waterfall Telemetry in Golden Image Pipelines
The Scenario
A platform engineering team builds automated machine images (such as AMIs or Packer artifacts) for a global Kubernetes cluster. To prevent boot performance regressions before images reach production, the CI pipeline must generate and validate visual boot chronographs.
The Command
Generate the visual SVG chronograph:
$ systemd-analyze plot > /tmp/boot-waterfall.svg
Parse the generated vector file to extract programmatic timing boundaries for CI assertions:
$ python3 -c '
import xml.etree.ElementTree as ET
tree = ET.parse("/tmp/boot-waterfall.svg")
root = tree.getroot()
text_nodes = [elem.text for elem in root.iter() if elem.text and "Startup finished" in elem.text]
print(text_nodes[0])
'
Terminal Output
Startup finished in 980ms (kernel) + 1.120s (initrd) + 4.230s (userspace) = 6.330s
Line-by-Line Technical Analysis
systemd-analyze plot: Connects to D-Bus and generates an interactive SVG chart illustrating the precise start, active, and completion times for every service on the system.- Distinct color-coded bars delineate phases: - Solid Red: Time spent actively executing in userspace CPU initialization. - Solid Black: Time spent waiting on asynchronous I/O, device readiness, or upstream target dependencies. - Light Gray: Background active running duration following initialization.
- The Python script parses the generated XML structure to validate total boot duration against the continuous integration service level agreement (SLA) threshold of less than 8.0 seconds.
Remediation and Action
The platform engineer integrates this script into the Packer post-processor stage. If any newly introduced package adds an unexpected blocking dependency (such as an unoptimized monitoring agent or logging daemon), the SVG artifact is published to an observability dashboard, and the pull request build fails automatically before the image reaches staging.
Use Case 5: CI/CD Pre-Flight Unit File Validation and Condition Testing
The Scenario
During a continuous deployment run, an engineer modifies a mission-critical systemd unit file (payment-gateway.service) with new conditional gates and sandbox directives. Deploying a broken unit file with invalid syntax or missing dependencies could cause initialization loops or service failures across thousands of production nodes.
The Command
First, perform an offline syntax and dependency verification of the unit file directly within the git repository:
$ systemd-analyze verify ./services/payment-gateway.service
Second, evaluate whether the deployment environment satisfies the runtime virtualization and hardware conditions required by the service:
$ systemd-analyze condition 'ConditionVirtualization=kvm' 'ConditionPathExists=/etc/ssl/certs/payment-vault.crt'
Terminal Output
[/home/ci-runner/repo/services/payment-gateway.service:24] Unknown lvalue 'ProtctSystem' in section 'Service'
[/home/ci-runner/repo/services/payment-gateway.service:31] Executable path is not absolute: bin/payment-gateway
payment-gateway.service: Unit configured to use NetworkManager.service, but unit does not exist.
ConditionVirtualization=kvm: yes
ConditionPathExists=/etc/ssl/certs/payment-vault.crt: no
ConditionPathExists=/etc/ssl/certs/payment-vault.crt was not met
Line-by-Line Technical Analysis
verifycatches a syntax typo on line 24:'ProtctSystem'(misspelling ofProtectSystem), which would otherwise fail silently or be ignored at runtime.verifydetects an invalid relative executable path on line 31:bin/payment-gatewayviolates systemd's requirement for absolute binary paths (such as/usr/bin/payment-gateway).verifydetects a missing dependency: the unit referencesNetworkManager.service, which is not installed in this minimal server environment.conditionevaluates runtime constraints against the host:ConditionVirtualization=kvmsucceeds (the host runs on KVM), butConditionPathExistsfails because the required TLS certificate is missing.
Remediation and Action
The engineer corrects the unit configuration prior to merging the pull request:
- Fixes the directive name to ProtectSystem=strict.
- Resolves the binary path to /usr/local/bin/payment-gateway.
- Replaces the NetworkManager.service dependency with systemd-networkd.service.
- Adds a pre-flight provisioning step to ensure the certificate exists at /etc/ssl/certs/payment-vault.crt.
These validation checks are added to the repository's git pre-commit hooks and CI linting pipelines to prevent configuration errors from ever reaching production servers.
Operational Caveats, Nuances & What Can Go Wrong
While systemd-analyze is an invaluable tool for systems engineering, misinterpreting its telemetry can lead to misguided optimizations or broken service environments.
1. blame Elapsed Time vs. CPU Saturation
A common misconception when interpreting systemd-analyze blame is treating initialization duration as equivalent to CPU resource consumption.
$$\Delta t_{\text{elapsed}} \neq \text{CPU Time}$$
systemd-analyze blame measures wall-clock time from InactiveExitTimestamp to ActiveEnterTimestamp. A unit that consumes negligible CPU but waits on an asynchronous network event, hardware discovery, or an external timer will register a high duration:
15.210s systemd-networkd-wait-online.service # 0.01s CPU time, 15.20s waiting on link carrier
Masking or removing services based solely on blame metrics can break downstream services that rely on these synchronization milestones. Always cross-reference blame against systemd-analyze critical-chain to determine if a slow unit is actually on the critical path to service availability.
2. The Scope of --user Session Instances
By default, systemd-analyze inspects the system-level daemon (PID 1). However, in modern Linux deployments running rootless containers, user-level multiplexers, or developer environments, services are often orchestrated by user-level systemd managers.
Invoking systemd-analyze on a rootless Podman orchestrator or user session will return only system-level metrics, missing user-space container initialization delays. Engineers must explicitly pass the --user flag to audit user-level session managers:
$ systemd-analyze --user blame
$ systemd-analyze --user security user-database.service
3. Over-Hardening and Privilege Revocation Hazards
When utilizing systemd-analyze security to harden units, aggressively applying directives without understanding their operational requirements can cause subtle runtime failures:
PrivateTmp=yes: Isolates the unit's/tmpdirectory into a private namespace. This breaks services that rely on shared/tmpUNIX domain sockets for local IPC (such as legacy X11 connections or local database sockets).ProtectSystem=strict: Mounts the entire OS filesystem tree read-only. If a daemon needs to write logs, create lockfiles, or rotate cache directories outside of/var/logor/run, it will encounterEPERM(Operation not permitted) orEROFS(Read-only file system) errors unless explicitReadWritePaths=exemptions are defined.ProtectHome=yes: Renders/home,/root, and/run/usercompletely inaccessible. Services requiring access to user data will fail immediately upon execution.SystemCallFilter=@system-service: Restricts system calls using a Seccomp BPF filter. If a library or runtime (such as JVM, Go, or Node.js) attempts an unlisted syscall (such as memory management or threading operations likeclone3ormprotect), the kernel will terminate the process with aSIGSYSsignal.
To prevent outages, test security overrides in a staging environment and monitor the system journal for seccomp audit violations:
$ journalctl -xe _COMM=backend-worker | grep -i "seccomp"
Authoritative Documentation & Further Reading
For deeper exploration of systemd architecture, initialization mechanics, and kernel interfaces, consult the following documentation:
- systemd-analyze(1) β Linux Manual Page
- systemd.exec(5) β Execution Environment Configuration
- systemd.unit(5) β Unit Configuration Syntax
- ArchWiki: Improving System Boot Performance
- Linux Kernel Documentation: Seccomp BPF
- Freedesktop.org: systemd System and Service Manager Architecture
Today's Takeaway
To understand and optimize your system's boot profile right now, open a terminal on your Linux machine and execute systemd-analyze critical-chain. Trace the highlighted branches to see which sequential dependencies dictate your system's startup latency. Then, pick one custom service on your machine and run systemd-analyze security <unit-name>.service to evaluate its sandboxing posture against modern Linux kernel isolation primitives.