Coredumpctl: Retrieving Kernel Crash Dumps, Inspecting Native Stack Traces, and Triaging Daemon Segfaults in Production
When a compiled service crashes without leaving behind so much as a dying log entry, it usually means the operating system terminated it immediately—a casualty of an illegal memory access or an unrecoverable hardware exception. In traditional Linux environments, diagnosing what happened meant digging through obscure filesystem paths, hoping that crash snapshots were neither disabled by default resource limits nor overwritten by subsequent background tasks.
Under modern Linux systems powered by systemd, you do not need to guess. You can immediately see every recent crash across the entire machine with a single command:
coredumpctl list --since "2 hours ago" --no-pager
TIME PID UID GID SIG COREFILE EXE SIZE
Tue 2026-08-17 01:45:12 UTC 142100 1001 1001 SIGABRT present /usr/bin/node 18.4M
Tue 2026-08-17 02:14:08 UTC 148291 33 33 SIGSEGV present /usr/sbin/envoy-proxy 142.1M
In less than a second, you have your culprit: at 02:14:08 UTC, the process /usr/sbin/envoy-proxy operating under process ID 148291 terminated due to a fatal segmentation fault (SIGSEGV), leaving behind a 142.1 MB compressed forensic snapshot ready for immediate triage.
What It Does in Plain English
When a running program encounters a fatal condition it cannot handle—such as trying to read memory it does not own or encountering a failed assertion—the Linux kernel takes an instantaneous snapshot of the program’s entire state. This snapshot captures the exact contents of its memory, CPU registers, active threads, and function call stacks at the microsecond of failure. This artifact is known as a core dump.
Historically, managing core dumps was messy: uncompressed multi-gigabyte files would scatter across arbitrary directories, fill disks to capacity, or disappear entirely if limits were misconfigured.
coredumpctl is the dedicated management and diagnostic tool for systemd-coredump. It intercepts crashes directly from the kernel, compresses them on the fly, indexes them alongside system logs, and stores them under strict disk quotas. With coredumpctl, systems engineers can search crash histories, view stack traces directly in the terminal without installing heavyweight debuggers, extract raw memory images, or immediately launch interactive debugging sessions against the exact state that failed.
The Architecture of Linux Core Dump Ingestion
To effectively triage crashes under operational pressure, it helps to understand how the operating system turns a low-level hardware fault into an indexed, searchable crash record.
From Kernel Fault to User-Space Pipe
When a program violates memory access permissions—such as dereferencing a NULL pointer—the CPU’s Memory Management Unit (MMU) raises a hardware page fault. The Linux kernel identifies the access as invalid and dispatches a fatal SIGSEGV signal to the process. If the application has no custom handler for this signal, the kernel initiates a core dump via its internal do_coredump() function.
Traditionally, the destination for core dumps was configured via /proc/sys/kernel/core_pattern as documented in core(5). In older systems, this file pointed to a static file path pattern (such as /tmp/core.%e.%p), which often flooded the root filesystem with massive, uncompressed binary files.
Modern Linux distributions configure /proc/sys/kernel/core_pattern to point directly to a user-space pipe:
|/lib/systemd/systemd-coredump %P %u %g %s %t %c %h
When a process crashes:
1. Piping the Memory: The kernel pauses the dying process and executes /lib/systemd/systemd-coredump, streaming the process memory down an anonymous UNIX pipe.
2. Passing Context: The kernel passes key parameters via command-line specifiers: Process ID (%P), User ID (%u), Group ID (%g), Signal number (%s), Crash timestamp (%t), and Soft resource limits (%c).
3. Metadata Harvesting: The systemd-coredump worker queries /proc/$PID/cgroup to identify the originating systemd unit and captures environmental attributes including the executable path, working directory, and command-line arguments.
4. Journal Logging: The worker logs structured metadata and an automated stack trace preview directly into systemd-journald.
5. Compressed Storage: If the crash payload fits within the storage limits defined in /etc/systemd/coredump.conf, it is compressed in real time using Zstandard (zstd) or LZ4 and saved to /var/lib/systemd/coredump/.
Anatomical Dissection of the ELF Core Dump
A Linux core dump is formatted as an Executable and Linkable Format (ELF) file of type ET_CORE. It contains no executable code sections, but instead records the memory layout and register state of the application at the exact instant it crashed.
An ET_CORE file consists of two primary segments:
-
PT_NOTEHeaders (Execution State): -NT_PRSTATUS: Records thread status, terminating signal number, faulting thread ID, and all CPU registers (such as instruction pointerRIP, stack pointerRSP, and general-purpose registers on x86_64). -NT_PRPSINFO: Captures high-level execution parameters, including process name and argument list (argv). -NT_AUXV: Preserves the ELF Auxiliary Vector, enabling debuggers to calculate memory offsets when Address Space Layout Randomization (ASLR) is enabled. -NT_FILE: Maps memory ranges to shared libraries on disk (.sofiles), preventing redundant duplication of static library code inside the dump. -NT_SIGINFO: Contains the faulting memory address and kernel signal code. -
PT_LOADHeaders (Memory Projections): - Defines the actual memory pages dumped from the process address space. To conserve disk space, the kernel's dump filter (/proc/$PID/coredump_filter) records dynamic heap and stack data while omitting read-only executable files that can be read directly from disk during analysis.
Core Flags & Command Reference
The table below outlines the primary subcommands and operational flags used with coredumpctl:
| Command / Flag | Purpose & Operational Scope |
|---|---|
coredumpctl list |
Summarises all indexed crash events recorded across the system journal. |
coredumpctl info [MATCH] |
Displays detailed metadata, register states, journal context, and a generated stack backtrace. |
coredumpctl debug [MATCH] |
Decompresses the core dump and launches an interactive debugger (GDB) targeting the matching binary. |
coredumpctl dump [MATCH] -o PATH |
Decompresses and extracts the raw ELF ET_CORE file to a specified path on disk. |
--since="WINDOW" / --until="WINDOW" |
Filters crash searches to a specific time window (e.g. "1 hour ago", "yesterday"). |
-u UNIT / --unit=UNIT |
Restricts queries exclusively to crashes originating from a specific systemd unit. |
-S SIGNAL / --signal=SIGNAL |
Filters crash events by signal name or number (e.g. SIGSEGV, SIGABRT, 11). |
--no-pager |
Disables terminal pagination, outputting raw text suitable for scripts and pipelines. |
5 Real-World Production Use Cases
The following scenarios illustrate how Site Reliability Engineers and Systems Administrators triage incidents in real-world Linux environments.
Use Case 1: Querying and Filtering Crash Events Across High-Throughput Production Hosts
The Scenario
A microservices gateway running envoy-proxy experiences intermittent connection dropouts across containerized instances. The service generates millions of access log lines per hour, making manual log scanning impractical. You must determine whether the daemon is crashing repeatedly under specific signals, correlate the events with its systemd service unit, and establish an exact timeline.
The Exact Command
coredumpctl list /usr/sbin/envoy-proxy \
--unit=envoy.service \
--signal=SIGSEGV \
--since="2026-08-17 00:00:00" \
--reverse \
--no-pager
Realistic Terminal Output
TIME PID UID GID SIG COREFILE EXE SIZE
Tue 2026-08-17 02:14:08 UTC 148291 33 33 SIGSEGV present /usr/sbin/envoy-proxy 142.1M
Tue 2026-08-17 01:12:44 UTC 139102 33 33 SIGSEGV present /usr/sbin/envoy-proxy 138.6M
Mon 2026-08-16 23:45:01 UTC 128450 33 33 SIGSEGV present /usr/sbin/envoy-proxy 141.0M
Line-by-Line Breakdown of the Output
TIME: The exact timestamp when the kernel caught the fatal signal and passed the dump tosystemd-coredump.PID: The process ID assigned by the host kernel in its root namespace (148291,139102,128450).UID / GID: The numeric user and group IDs under which the process executed (33:33, corresponding towww-data).SIG: The fatal terminating signal (SIGSEGV, signal 11), confirming memory access violations rather than standard exits (SIGTERM) or software assertion aborts (SIGABRT).COREFILE: The status of the core dump on disk.presentindicates the full compressed dump is available; other values includemissing(pruned by retention policies) ortruncated(exceeded size limits).EXE: The absolute filesystem path to the executable binary mapped in memory.SIZE: The compressed size of the core dump file in/var/lib/systemd/coredump/.
Actionable Next Steps
The crash history reveals a repeating failure every 60 to 75 minutes. You note the most recent PID (148291) and proceed to inspect its stack backtrace using coredumpctl info.
Use Case 2: Inspecting Crash Metadata, Registers, and Backtraces Without Interactive Debuggers
The Scenario
Security policies on a hardened production server prohibit running interactive debuggers like GDB. You need to inspect PID 148291 directly from the shell: identify the faulting instruction pointer, inspect CPU registers, and view a complete stack trace without installing external debugging tools.
The Exact Command
coredumpctl info 148291
Realistic Terminal Output
PID: 148291 (envoy-proxy)
UID: 33 (www-data)
GID: 33 (www-data)
Signal: 11 (SEGV)
Timestamp: Tue 2026-08-17 02:14:08 UTC (42min ago)
Command Line: /usr/sbin/envoy-proxy -c /etc/envoy/envoy.yaml --concurrency 4
Executable: /usr/sbin/envoy-proxy
Control Group: /system.slice/envoy.service
Unit: envoy.service
Slice: system.slice
Boot ID: 4f1a28cb947b420f913d8d64197e93ab
Machine ID: 9a38f71295db4679872e45da84cf1299
Hostname: edge-gw-prod-01
Storage: /var/lib/systemd/coredump/core.envoy-proxy.33.4f1a28cb947b420f913d8d64197e93ab.148291.1786932848000000.zst (present)
Disk Size: 142.1M
Message: Process 148291 (envoy-proxy) of user 33 dumped core.
Module /usr/sbin/envoy-proxy with build-id 8a4c2f0d912e73a14b986100234a9ef12781cb92
Stack trace of thread 148294:
#0 0x000055d13a98ef42 _ZN5Envoy4Http15HeaderMapImpl8addCopyERKNS0_9LowerCaseStringENSt7__cxx1112basic_stringIcSt11char_traitsIcESaIcEEE (envoy-proxy + 0x1a8ef42)
#1 0x000055d13a890114 _ZN5Envoy6Router12FilterConfig16processHeaderMapERNS_4Http13RequestHeaderMapE (envoy-proxy + 0x1990114)
#2 0x000055d13a771b90 _ZN5Envoy6Router6Filter13decodeHeadersERNS_4Http16RequestHeaderMapEb (envoy-proxy + 0x1871b90)
#3 0x000055d13a6540c1 _ZN5Envoy4Http18ConnectionManager15decodeHeadersInternalEv (envoy-proxy + 0x17540c1)
#4 0x00007f8a3ec2b609 start_thread (libpthread.so.0 + 0x9609)
#5 0x00007f8a3eb50353 clone (libc.so.6 + 0x122353)
Line-by-Line Breakdown of the Output
PID / UID / GID: Validates process identity and security boundaries, confirming execution under the unprivilegedwww-dataaccount.Signal: 11 (SEGV): Verifies that the CPU raised a memory segmentation fault.Command Line: Shows the exact runtime arguments supplied at launch (-c /etc/envoy/envoy.yaml --concurrency 4).Control Group / Unit: Confirms the crash occurred withinsystem.slice/envoy.service, proving the service was not terminated by an external orchestrator.Storage: Provides the exact path to the Zstandard-compressed core dump file on disk.Module ... build-id: Displays the unique GNUbuild-idSHA-1 hash (8a4c2f...) embedded during compilation, ensuring precise matching with debugging symbols.Stack trace of thread 148294: Identifies worker thread148294as the specific thread that encountered the fault.- Frame
#0 (0x000055d13a98ef42): Pinpoints the failure insideEnvoy::Http::HeaderMapImpl::addCopy(), indicating an issue during string copying into an HTTP header map. - Frame
#1 (0x000055d13a890114): Traces the caller toEnvoy::Router::FilterConfig::processHeaderMap(), confirming the crash happened during routing filter evaluation.
Actionable Next Steps
The crash is isolated to header processing in the routing filter. You attach this stack trace directly to the incident ticket and proceed to run a detailed inspection or export the dump for deeper analysis.
Use Case 3: Launching Interactive GDB Sessions Seamlessly with Debug Symbol Integration
The Scenario
The stack backtrace confirms the crash occurred in HeaderMapImpl::addCopy(). However, you need to inspect local variable values and CPU registers to determine whether the fault was caused by a NULL pointer dereference, an unaligned memory access, or memory corruption. You want an automated way to launch GDB (The GNU Project Debugger) against the exact core payload, matching binaries, and source code symbols.
The Exact Command
coredumpctl debug 148291 \
--debugger-arguments="-ex 'set pagination off' -ex 'bt full' -ex 'info registers' -ex 'quit'"
(For interactive step-through debugging, run coredumpctl debug 148291 without the batch --debugger-arguments flags).
Realistic Terminal Output
PID: 148291 (envoy-proxy)
Executable: /usr/sbin/envoy-proxy
GNU gdb (Ubuntu 12.1-0ubuntu1~22.04.2) 7.6.1
Copyright (C) 2022 Free Software Foundation, Inc.
Reading symbols from /usr/sbin/envoy-proxy...
Downloading separate debug info for /usr/sbin/envoy-proxy via debuginfod...
Reading symbols from /root/.cache/debuginfod_client/8a4c2f0d912e73a14b986100234a9ef12781cb92/debuginfo...done.
[New LWP 148294]
[New LWP 148291]
[New LWP 148292]
[New LWP 148293]
Core was generated by `/usr/sbin/envoy-proxy -c /etc/envoy/envoy.yaml --concurrency 4'.
Program terminated with signal SIGSEGV, Segmentation fault.
#0 0x000055d13a98ef42 in Envoy::Http::HeaderMapImpl::addCopy (this=0x0,
key=..., value="Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...")
at source/common/http/header_map_impl.cc:184
184 HeaderString& new_value = header->value();
#0 0x000055d13a98ef42 in Envoy::Http::HeaderMapImpl::addCopy (this=0x0,
key=..., value="Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...")
at source/common/http/header_map_impl.cc:184
new_value = <error reading variable>
header = 0x0
#1 0x000055d13a890114 in Envoy::Router::FilterConfig::processHeaderMap (
this=0x55d13cc40810, headers=...) at source/common/router/config_impl.cc:342
entry = 0x55d13cc48920
rax 0x0 0
rbx 0x7f8a3c829000 140231649980416
rcx 0x55d13cc48920 94357326498080
rdx 0x0 0
rsi 0x7f8a3c829120 140231649980704
rdi 0x0 0
rbp 0x7f8a3c8290e0 0x7f8a3c8290e0
rsp 0x7f8a3c829090 0x7f8a3c829090
rip 0x000055d13a98ef42 0x55d13a98ef42 <Envoy::Http::HeaderMapImpl::addCopy+34>
eflags 0x10206 [ PF IF RF ]
Line-by-Line Breakdown of the Output
Downloading separate debug info ... via debuginfod: GDB contactsdebuginfodto automatically fetch matching debug symbols based on the binary'sbuild-id, rendering exact source filenames and line numbers (source/common/http/header_map_impl.cc:184).[New LWP ...]: GDB enumerates all kernel threads present when the crash happened.this=0x0: The root cause is revealed. Thethispointer passed toHeaderMapImpl::addCopyis0x0(nullptr).key=..., value="Bearer eyJhbGci...": Confirms the thread crashed while handling an authorization header containing a JWT token.rax 0x0 / rdi 0x0: Register inspection confirms registerrdi(used to pass the first argument andthispointer in x86_64 ABI) holds0x0. The crash occurred when instruction0x55d13a98ef42attempted to dereference[rdi + offset].
Actionable Next Steps
You have conclusive evidence: the routing filter called addCopy() on an uninitialised header map pointer when processing bearer tokens. You file a bug report with the development team citing source/common/http/header_map_impl.cc:184 and provide the input token context.
Use Case 4: Extracting and Compressing Raw Core Payloads for Off-Host Analysis
The Scenario
Production servers run in a secure, isolated network zone without developer tooling or internet connectivity. Company policy forbids running debuggers on production hosts. You must export the uncompressed ELF core dump from systemd-coredump, compress it using multi-threaded Zstandard compression, and transfer it to an off-host analysis system.
The Exact Command
coredumpctl dump 148291 \
--output=/var/crash/envoy_crash_pid148291_$(date +%Y%m%d_%H%M%S).core
Verify the exported ELF file and compress it:
file /var/crash/envoy_crash_pid148291_*.core && \
zstd -19 --rm -T0 /var/crash/envoy_crash_pid148291_*.core
Realistic Terminal Output
PID: 148291 (envoy-proxy)
Executable: /usr/sbin/envoy-proxy
Storage: /var/lib/systemd/coredump/core.envoy-proxy.33.4f1a28cb947b420f913d8d64197e93ab.148291.1786932848000000.zst (present)
Disk Size: 142.1M
Dumped core to /var/crash/envoy_crash_pid148291_20260817_024510.core.
/var/crash/envoy_crash_pid148291_20260817_024510.core: ELF 64-bit LSB core file, x86-64, version 1 (SYSV), SVR4-style, from '/usr/sbin/envoy-proxy -c /etc/envoy/envoy.yaml --concurrency 4', real uid: 33, effective uid: 33, real gid: 33, effective gid: 33, execfn: '/usr/sbin/envoy-proxy', platform: 'x86_64'
/var/crash/envoy_crash_pid148291_20260817_024510.core : 78.42% (1.42 GiB => 312 MiB, /var/crash/envoy_crash_pid148291_20260817_024510.core.zst)
Line-by-Line Breakdown of the Output
Dumped core to ...:coredumpctl dumpextracts and decompresses the payload to the target file path.file ...: Thefileutility validates that the exported file is a valid 64-bit ELF core dump (e_type=ET_CORE).ELF 64-bit LSB core file ... platform: 'x86_64': Verifies that register layouts and thread context conform to the standard System V ABI specification.zstd -19 --rm -T0: Compresses the 1.42 GB raw memory image down to 312 MB using high-ratio compression across all available CPU threads (-T0), removing (--rm) the uncompressed source file once complete.
Actionable Next Steps
Transfer the compressed artifact (/var/crash/envoy_crash_pid148291_20260817_024510.core.zst) via scp or rsync to a dedicated analysis environment equipped with full source repositories and debugging symbols.
Use Case 5: Hardening Storage Limits, Disk Quotas, and Retention Policies in /etc/systemd/coredump.conf
The Scenario
A backend worker enters a rapid crash-restart loop under systemd supervision. Each crash attempts to write a 4 GB memory dump to /var/lib/systemd/coredump/. Without storage limits, this crash storm will rapidly exhaust disk space on the root filesystem. You must configure production storage thresholds and verify that active policies prevent disk starvation.
The Exact Configuration & Command
Create a configuration drop-in file at /etc/systemd/coredump.conf.d/50-production-limits.conf:
[Coredump]
# Store core images directly on disk (external to the systemd journal)
Storage=external
# Compress core dumps using real-time Zstandard compression
Compress=yes
# Maximum total memory size of a process to capture (discard if larger)
ProcessSizeMax=2G
# Maximum compressed disk space an individual core dump may consume
ExternalSizeMax=1G
# Enforce strict disk bounds across all aggregated coredumps
MaxUse=10%
KeepFree=20%
Inspect the active configuration using systemd-analyze:
systemd-analyze cat-config systemd/coredump.conf
Realistic Terminal Output
# /etc/systemd/coredump.conf
# See coredump.conf(5) for details.
[Coredump]
#Storage=external
#Compress=yes
#ProcessSizeMax=2G
#ExternalSizeMax=2G
#JournalSizeMax=768M
#MaxUse=
#KeepFree=
# /etc/systemd/coredump.conf.d/50-production-limits.conf
[Coredump]
Storage=external
Compress=yes
ProcessSizeMax=2G
ExternalSizeMax=1G
MaxUse=10%
KeepFree=20%
Line-by-Line Breakdown of Configuration Parameters
Storage=external: Saves raw memory dumps to/var/lib/systemd/coredump/while recording metadata in the journal. SettingStorage=nonelogs metadata while completely discarding raw memory—ideal for strict privacy environments.Compress=yes: Applies Zstandard compression before writing to disk, reducing storage consumption by 60% to 85%.ProcessSizeMax=2G: Directs the handler to discard dumps from processes whose memory footprint exceeds 2 GB, preventing memory-heavy caches from exhausting disk capacity.ExternalSizeMax=1G: Enforces a maximum on-disk compressed file size for any single core dump.MaxUse=10%: Caps cumulative disk usage for all stored core dumps to 10% of the underlying filesystem capacity.KeepFree=20%: Pauses core dump saving if free filesystem space falls below 20%, safeguarding the root volume against disk exhaustion.
Actionable Next Steps
Reload the systemd manager to apply the changes:
systemctl daemon-reload
Check the systemd-tmpfiles-clean.timer unit to confirm that old core files are purged automatically according to retention schedules (defaulting to 3 days).
Operational Edge Cases: Symbols, Namespaces, and Data Privacy
Managing core dumps in production infrastructure requires addressing three common challenges:
1. Stripped Binaries and Symbol Table Resolution
Production binaries are almost always compiled with debugging symbols stripped (strip -s) to reduce executable size. When analyzing stripped binaries, coredumpctl info shows numeric memory offsets (e.g. app + 0x14a2b) instead of human-readable function names.
To resolve stripped backtraces:
- Check the Module ... build-id line in coredumpctl info or run readelf -n <binary>.
- Use debuginfod. When running coredumpctl debug, GDB automatically contacts configured debuginfod servers using the build-id and downloads the matching DWARF symbols without modifying the binary on disk.
2. Container Namespace Boundaries (Docker, Podman, Kubernetes)
When a process crashes inside a container, the host kernel intercepts the fault because /proc/sys/kernel/core_pattern is a host-wide setting.
The systemd-coredump worker runs in the host's root namespace, inspecting the container process via /proc/$PID/root/. Running coredumpctl debug directly on the host may fail to resolve shared libraries that exist only inside the container image.
To debug container crashes accurately:
- Export the core dump: coredumpctl dump <PID> -o /tmp/container.core
- Run GDB inside a container instance created from the identical image:
bash
docker run --rm -v /tmp/container.core:/tmp/core:ro -v /usr/src/app:/src:ro my-app:latest gdb /app/bin /tmp/core
3. Mitigating Memory Exposure and Data Leakage
Because a core dump captures memory pages from the dying process, it can contain sensitive information such as decrypted TLS keys, database passwords, and personal data.
To protect sensitive data:
- Filesystem Permissions: Ensure /var/lib/systemd/coredump/ is restricted to mode 0750 owned by root:root or root:systemd-coredump. Core dump files are created with mode 0640 by default.
- Selective Memory Exclusion: Applications handling secret keys can call madvise(..., MADV_DONTDUMP) on sensitive buffers, directing the kernel to exclude those memory pages from core dumps.
- Selective Disabling via prctl: In PCI-DSS or HIPAA environments, sensitive daemons can call prctl(PR_SET_DUMPABLE, 0) during initialization, instructing the kernel to discard core dumps completely upon termination.
What Can Go Wrong: Diagnostic Pitfalls and How to Recover
Pitfall 1: Missing Core Dumps Due to Zeroed Resource Limits (RLIMIT_CORE)
An application crashes with SIGSEGV, but coredumpctl list shows no records, or lists the status as none or missing.
The Cause
The process was started with a core file limit of zero (ulimit -c 0). Although systemd-coredump intercepts crashes via a kernel pipe, the kernel checks the process’s soft RLIMIT_CORE limit before streaming data. If the limit is zero, the kernel aborts dump generation.
The Fix & Recovery
Check the active limits of the running service using prlimit:
prlimit --pid $(pgrep envoy-proxy) --core
If SOFT is 0, create a drop-in override for the systemd service:
systemctl edit envoy.service
Add the following configuration:
[Service]
LimitCORE=infinity
Apply the changes and restart the service:
systemctl daemon-reload && systemctl restart envoy.service
Pitfall 2: Setuid and Privileged Process Dump Ingestion Failures
Services that change user privileges (such as a master process running as root spawning worker processes under www-data) fail to generate core dumps when a worker crashes.
The Cause
Linux kernel security protections prevent processes that have executed setuid() from dumping core by default. This prevents unprivileged users from accessing memory segments populated while the process had root privileges.
The Fix & Recovery
Check the current kernel setting:
sysctl fs.suid_dumpable
If set to 0 (disabled), set it to mode 2 ("safe" mode), which allows core dumps through a user-space pipe (such as systemd-coredump) while preventing unprivileged file writes:
sysctl -w fs.suid_dumpable=2
echo "fs.suid_dumpable = 2" >> /etc/sysctl.d/99-security-coredump.conf
Pitfall 3: Cascading Disk I/O Saturation During High-Volume Crash Loops
When a multi-threaded service crashes in a tight loop across hundreds of threads, the kernel launches concurrent systemd-coredump workers. This can cause severe disk contention and memory pressure, affecting adjacent healthy services.
The Cause
Unthrottled kernel core dump dispatching combined with a lack of rate limiting on the ingestion socket.
The Fix & Recovery
Configure rate limiting on the systemd-coredump socket unit:
systemctl edit systemd-coredump.socket
Add rate limits to throttle crash ingestion storms:
[Socket]
TriggerLimitIntervalSec=30s
TriggerLimitBurst=5
Apply the configuration:
systemctl daemon-reload
If a crash loop exceeds 5 crashes in 30 seconds, systemd temporarily pauses socket activation, preventing disk I/O saturation.
Today's Takeaway
The difference between an unresolved outage and a rapid root-cause diagnosis often comes down to post-mortem visibility. Open your terminal right now and run coredumpctl list --no-pager. If the system returns a history of application crashes, pick a recent PID and run coredumpctl info <PID> to see how systemd maps memory registers, signals, and stack traces without requiring any third-party tools. If coredumpctl reports no records or missing files, inspect your critical service unit files to ensure LimitCORE=infinity is set and verify that /etc/systemd/coredump.conf is configured with appropriate storage limits.
Authoritative Documentation & Technical References
- systemd-coredump(8) Manual — freedesktop.org
- coredumpctl(1) Reference Manual — freedesktop.org
- coredump.conf(5) Configuration Guide — freedesktop.org
- core(5) Linux Programmer's Manual — man7.org
- GDB Documentation: Debugging Core Dumps — GNU Project
- ArchWiki: Core Dump Management and Troubleshooting
- Debuginfod Architecture and Client Setup — Sourceware / ELFutils