Capsh: Auditing Process Capability Boundaries, Dropping Privileged Syscalls, and Hardening Container Runtimes in Production
Yet, as the forensic timeline unfolds, the anticipated catastrophe fails to materialize. The attackerβs post-exploitation payloads collapse in a cascade of permission errors. Attempts to load kernel modules, manipulate network routing tables, trace adjacent memory spaces, or inspect raw disk blocks are abruptly terminated by the kernel with EPERM. The server survived intact not because of luck, but because the daemon had been stripped of universal administrative power before execution. It operated under a surgically truncated envelope of POSIX capabilities managed by capsh.
Under the classic Unix security model dating back to the 1970s, permissions were an all-or-nothing proposition: a process was either an unprivileged account or it possessed sovereign authority over every file, socket, memory segment, and kernel subsystem on the machine. Modern Linux dismantled this blunt binary by splitting the monolithic power of UID 0 into dozens of discrete, modular privileges known as POSIX capabilities. The capsh (Capability Shell) utility is the definitive diagnostic lens and execution harness for this system, allowing engineers to inspect, trim, and enforce granular security boundaries on running services and background jobs.
Before configuring complex server hardening rules, the single most useful practical command you can execute on any Linux machine is an immediate capability audit of your current shell environment:
capsh --print
Current: =
Bounding set =cap_chown,cap_dac_override,cap_dac_read_search,cap_fowner,cap_fsetid,
cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_linux_immutable,cap_net_bind_service,
cap_net_broadcast,cap_net_admin,cap_net_raw,cap_ipc_lock,cap_ipc_owner,cap_sys_module,
cap_sys_rawio,cap_sys_chroot,cap_sys_ptrace,cap_sys_pacct,cap_sys_admin,cap_sys_boot,
cap_sys_nice,cap_sys_resource,cap_sys_time,cap_sys_tty_config,cap_mknod,cap_lease,
cap_audit_write,cap_audit_control,cap_setfcap,cap_mac_override,cap_mac_admin,
cap_syslog,cap_wake_alarm,cap_block_suspend,cap_audit_read,cap_perfmon,cap_bpf,
cap_checkpoint_restore
Ambient set =
Current IAB:
Securebits: 00/0x0/1'b0 (no-revert-caps)
secure-noroot: no (unlocked)
secure-no-suid-fixup: no (unlocked)
secure-keep-caps: no (unlocked)
secure-no-ambient-raise: no (unlocked)
uid=0(root) gid=0(root) groups=0(root)
guarded: unsupported
This diagnostic snapshot queries the kernel credential structure directly, exposing active user credentials, the full 41-capability bounding set, ambient privileges, and security bits.
What It Does in Plain English
Rather than granting a program blanket administrative privileges through the superuser account, capsh lets systems engineers grant applications the exact, minimal subset of administrative abilities required for their functionβsuch as opening low-numbered network ports (CAP_NET_BIND_SERVICE), setting system clocks (CAP_SYS_TIME), or locking memory pages (CAP_IPC_LOCK).
By wrapping process execution with capsh, administrators permanently prevent any application compromise from escalating into a full host takeover. If an attacker achieves arbitrary code execution within a service constrained by capsh, their payloads are trapped within a sandbox lacking the system calls needed to alter hardware, read unauthorized files, or pivot across the operating system.
Core Flags and Quick Start Reference
The capsh utility provides a flexible grammar for manipulating process credentials, Linux security flags, and capability sets defined within the Linux capabilities(7) manual.
| Flag Parameter | Operational Mechanism & Kernel Effect |
|---|---|
--decode=<hex> |
Decodes a raw hexadecimal bitmask (such as those found in /proc/<PID>/status) into human-readable capability names. |
--drop=<caps> |
Irreversibly removes the specified comma-delimited capabilities from the processβs Bounding capability set. |
--caps=<caps> |
Explicitly configures the threadβs Permitted, Inheritable, and Effective capability sets using libcap string syntax. |
--inh=<caps> |
Configures the processβs Inheritable capability set prior to executing the target binary. |
--ambient=<caps> |
Raises capabilities into the processβs Ambient set (Linux 4.3+), ensuring non-root privilege inheritance across execve(2). |
--user=<user> |
Drops UID and GID credentials to an unprivileged account while orchestrating capability preservation. |
--secbits=<mask\|names> |
Applies Linux Securebits flags (such as SECBIT_NOROOT) to prevent UID transformations from regaining root powers. |
--no-new-privs |
Invokes the prctl(PR_SET_NO_NEW_PRIVS, 1) syscall, forbidding the process and its children from gaining privileges via SUID/SGID binaries. |
--print |
Outputs an exhaustive diagnostic snapshot of the current process credentials, capability sets, and securebits status. |
Architectural Deep Dive: Kernel Credential Lifecycle Across execve
To utilize capsh effectively in production, engineers must understand how the Linux kernel manages process privileges under the hood. Within the Linux kernel, process credentials are encapsulated inside the struct cred data structure defined in include/linux/cred.h. As documented in the Kernel Credential Architecture Documentation, each Linux thread maintains five distinct capability sets:
System-wide ceiling for capabilities across execve"] PRM["Permitted Set (P)
Limiting ceiling for thread's active powers"] INH["Inheritable Set (I)
Preserved across execve when file caps match"] EFF["Effective Set (E)
Actively tested by kernel for syscall authorization"] AMB["Ambient Set (A)
Preserved across non-root execve without file caps"] BND --> PRM PRM --> EFF PRM --> INH INH --> AMB AMB --> PRM AMB --> EFF end
- Permitted ($P$): The absolute upper bound of capabilities the thread is authorized to assume. Capabilities in this set can be moved into the Effective set at runtime by an unprivileged thread using
cap_set_proc(3). - Inheritable ($I$): A mask of capabilities preserved across an
execve(2)system call if and only if the target binary has matching inheritable capabilities configured via extended filesystem attributes (security.capability). - Effective ($E$): The capability bitmask actively evaluated by kernel subsystems during privilege checks (such as whether
capable(CAP_NET_BIND_SERVICE)returns true). - Bounding Set ($X$): A process-level filter and monotonic ceiling. A capability cannot be gained through file-capability-based
execve(2)transformations unless it exists within the process's Bounding set. - Ambient Set ($A$): Introduced in Linux 4.3 (via
PR_CAP_AMBIENT), this set allows non-root processes to preserve and inherit capabilities across anexecve(2)call to an unprivileged binary without requiring extended file attributes on the binary itself.
The Transformation Formula of execve(2)
When an executable binary is invoked through the execve(2) system call, the kernel transitions the process from its current credentials ($P, I, E, X, A$) to new credentials ($P', I', E', X', A'$) according to strict kernel logic:
$$P' = (P_{\text{inh}} \cap F_{\text{inh}}) \cup (F_{\text{prm}} \cap X) \cup A$$
$$A' = \begin{cases} A & \text{if UID } \neq 0 \text{ and binary is not setuid} \ 0 & \text{if UID is altered or ambient inheritance is cleared} \end{cases}$$
$$E' = \begin{cases} P' & \text{if } F_{\text{eff}} \text{ bit is set on the binary} \ A' & \text{if file has no capabilities, but ambient capabilities exist} \ 0 & \text{otherwise} \end{cases}$$
Understanding this mathematical transition explains why traditional daemon privilege dropping often failed: calling setuid(non_zero) historically wiped the Permitted and Effective capability sets unless complex prctl configurations were applied. The capsh utility abstracts this underlying kernel complexity into repeatable, deterministic CLI invocations.
5 Production Use Cases for Senior SREs and Sysadmins
Scenario 1: Forensic Privilege Auditing via /proc/<PID>/status Bitmask Decoding
The Operational Scenario
During an infrastructure-wide security compliance audit, an auditor alerts you that a critical proprietary payment ingestion service (payment_gw_srv, PID 41022) is executing with an unknown, non-standard privilege profile. The auditor suspects that the process was launched with unmetered administrative authority. You need to inspect the live process status without interrupting traffic, extract its capability masks, and decode them into human-readable capability identifiers.
Execution Command
Inspect the process status via procfs, isolate the capability bitmasks, and pipe the hexadecimal representations directly into capsh --decode:
grep -E '^Cap(Inh|Prm|Eff|Bnd|Amb):' /proc/41022/status
capsh --decode=$(awk '/CapEff:/ {print $2}' /proc/41022/status)
Representative Terminal Output
CapInh: 0000000000000000
CapPrm: 0000000000000420
CapEff: 0000000000000420
CapBnd: 000001ffffffffff
CapAmb: 0000000000000000
0x0000000000000420=cap_net_bind_service,cap_sys_nice
Line-by-Line Technical Analysis
CapInh: 0000000000000000: The process does not pass capabilities to child binaries via standard inheritable file attributes.CapPrm: 0000000000000420: The process holds a 64-bit integer bitmask where bit 5 (1 << 5 = 0x20, representingCAP_NET_BIND_SERVICE) and bit 10 (1 << 10 = 0x400, representingCAP_SYS_NICE) are set ($0x20 + 0x400 = 0x420$).CapEff: 0000000000000420: The kernel actively enforces only these two capabilities during execution.CapBnd: 000001ffffffffff: The bounding set remains fully open (bits 0 through 40 set), which allows the process to assume other capabilities if it executes a binary configured with file capabilities.CapAmb: 0000000000000000: No ambient capabilities are active.- Decoded Output:
capshconfirms that the process holds only two specific administrative privileges: binding ports below 1024 (cap_net_bind_service) and adjusting process scheduling priorities (cap_sys_nice).
SRE Next Steps
- Document the decoded capabilities in the serviceβs architectural specification.
- Verify that the application genuinely requires
cap_sys_nice(for example, to set low-latency thread scheduling). - If
cap_sys_niceis unused, eliminate it from the systemd unit file or container specification to enforce absolute least privilege.
Scenario 2: Sandboxing Ingress Daemons by Excising Dangerous Bounding Capabilities
The Operational Scenario
You are deploying an edge-facing reverse proxy (envoy or haproxy) that terminates public TLS traffic. The proxy runs as an unprivileged user, but you must guarantee that even if an attacker achieves arbitrary code execution and executes a local binary with Set-UID (suid-root) permissions, the attacker cannot mount filesystems, load kernel modules, inject into adjacent processes, or sniff raw network interfaces. You must restrict the Bounding capability ceiling before handing execution over to the application.
Execution Command
Use capsh to permanently drop CAP_SYS_ADMIN, CAP_NET_RAW, CAP_SYS_PTRACE, and CAP_SYS_MODULE from the bounding set, enforce no-new-privs, drop privileges to the proxyuser account (UID 1002), and execute the proxy daemon:
capsh --drop=cap_sys_admin,cap_net_raw,cap_sys_ptrace,cap_sys_module \
--no-new-privs \
--user=proxyuser \
-- -c "exec /opt/proxy/bin/reverse_proxy --config /etc/proxy/proxy.yaml"
To verify the bounding set reduction on the spawned process:
capsh --decode=$(awk '/CapBnd:/ {print $2}' /proc/$(pgrep -u proxyuser reverse_proxy)/status)
Representative Terminal Output
0x000001ffffde7ddf=cap_chown,cap_dac_override,cap_dac_read_search,cap_fowner,cap_fsetid,
cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_linux_immutable,cap_net_bind_service,
cap_net_broadcast,cap_net_admin,cap_ipc_lock,cap_ipc_owner,cap_sys_rawio,cap_sys_chroot,
cap_sys_pacct,cap_sys_boot,cap_sys_nice,cap_sys_resource,cap_sys_time,cap_sys_tty_config,
cap_mknod,cap_lease,cap_audit_write,cap_audit_control,cap_setfcap,cap_mac_override,
cap_mac_admin,cap_syslog,cap_wake_alarm,cap_block_suspend,cap_audit_read,cap_perfmon,
cap_bpf,cap_checkpoint_restore
Line-by-Line Technical Analysis
--drop=cap_sys_admin,...: Instructscapshto callprctl(PR_CAPBSET_DROP, ...)for each named capability. This is a one-way operation; once dropped from the Bounding set, neither this process nor any of its child processes can ever recover these capabilities.--no-new-privs: Callsprctl(PR_SET_NO_NEW_PRIVS, 1). This prevents SUID-root binaries (such as/usr/bin/sudoor/usr/bin/passwd) from granting the process elevated capabilities onexecve(2).--user=proxyuser: Changes the real, effective, and saved UIDs and GIDs to those ofproxyuser.-- -c "exec ...": Terminatescapshoptions and passes the command string to be executed under the newly constrained security context.- Decoded Output: In the resulting bounding set mask (
0x000001ffffde7ddf), bits 16 (CAP_NET_RAW), 19 (CAP_SYS_PTRACE), 21 (CAP_SYS_ADMIN), and 27 (CAP_SYS_MODULE) are explicitly stripped.
SRE Next Steps
- Attempt a diagnostic attachment to the running proxy using
gdborstraceas an unprivileged user; confirm that the kernel rejects the attachment due to the missingCAP_SYS_PTRACE. - Codify this bounding restriction directly into the systemd service file:
ini CapabilityBoundingSet=~CAP_SYS_ADMIN CAP_NET_RAW CAP_SYS_PTRACE CAP_SYS_MODULE NoNewPrivileges=true
Scenario 3: Non-Root Low-Port Binding via Ambient Capabilities (Linux 4.3+)
The Operational Scenario
A development team has packaged a Go-based microservice that must listen on standard HTTP (80) and HTTPS (443) ports. Company security policy strictly forbids running application processes as UID 0. Previously, engineers assigned file capabilities to the binary using setcap 'cap_net_bind_service=+ep' /opt/app/server. However, this binary is frequently updated via CI/CD pipelines deployed to ephemeral NVMe mountpoints mounted with the nosuid mount option (which ignores file capabilities) or immutable containers where extended file attributes are wiped during packaging. You must use Ambient capabilities to pass CAP_NET_BIND_SERVICE across execve(2) to a standard unprivileged user account (appuser, UID 1001) without relying on filesystem extended attributes.
Execution Command
Execute the service using capsh, raising CAP_NET_BIND_SERVICE into both the Inheritable and Ambient sets while changing to the unprivileged UID:
capsh --keep=1 \
--user=appuser \
--inh=cap_net_bind_service \
--ambient=cap_net_bind_service \
-- -c "exec /opt/app/server --listen-addr :80"
To verify the socket binding and capability profile while the daemon runs:
ss -tulpn | grep :80
grep -E '^(Uid|Cap.*):' /proc/$(pgrep -u appuser server)/status
Representative Terminal Output
tcp LISTEN 0 4096 0.0.0.0:80 0.0.0.0:* users:(("server",pid=48192,fd=3))
Uid: 1001 1001 1001 1001
CapInh: 0000000000000400
CapPrm: 0000000000000400
CapEff: 0000000000000400
CapBnd: 000001ffffffffff
CapAmb: 0000000000000400
Line-by-Line Technical Analysis
--keep=1: Configures the thread'sSECBIT_KEEP_CAPSsecurebit (prctl(PR_SET_KEEPCAPS, 1)). This prevents the kernel from clearing the Permitted set when the process transitions from UID 0 to UID 1001.--user=appuser: Drops the process credentials to UID/GID 1001.--inh=cap_net_bind_service: Adds capability 10 (CAP_NET_BIND_SERVICE) to the Inheritable set. A capability must exist in both the Permitted and Inheritable sets before it can be added to the Ambient set.--ambient=cap_net_bind_service: Callsprctl(PR_CAP_AMBIENT, PR_CAP_AMBIENT_RAISE, CAP_NET_BIND_SERVICE, 0, 0). When/opt/app/serveris executed, the kernel copies this ambient capability into the new binary's Permitted and Effective sets.- Status Output Verification:
Uid: 1001: The process runs completely unprivileged.CapEff: ...0400: The process effectively holdsCAP_NET_BIND_SERVICE.CapAmb: ...0400: The ambient capability was preserved across theexecveboundary without touching extended file attributes.
SRE Next Steps
- Transition the application launch configuration to standard systemd parameters in
/etc/systemd/system/microservice.service:ini [Service] User=appuser Group=appuser AmbientCapabilities=CAP_NET_BIND_SERVICE CapabilityBoundingSet=CAP_NET_BIND_SERVICE ExecStart=/opt/app/server --listen-addr :80 - Remove any leftover file capability assignments (
setcap -r /opt/app/server) to eliminate conflicting security states.
Scenario 4: Locking Credential Transitions using Securebits Flags in Multi-Tenant Runners
The Operational Scenario
You maintain a multi-tenant continuous integration (CI/CD) execution worker. Untrusted user jobs run inside isolated runtime shells. While the runner drops the worker process to an unprivileged account (ci-worker, UID 1000), you must ensure that malicious build scripts cannot exploit legacy Set-UID behaviors or attempt to regain capabilities if an attacker executes code paths that invoke setuid(0). You must lock down kernel credential transformations permanently using Securebits.
Execution Command
Launch the runner process while setting Securebits to disable root capability elevation (SECBIT_NOROOT), disable Set-UID capability fixups (SECBIT_NO_SETUID_FIXUP), and apply their immutable locks (SECBIT_NOROOT_LOCKED, SECBIT_NO_SETUID_FIXUP_LOCKED):
# Securebit bitmask values:
# SECBIT_NOROOT (1 << 0) = 0x01
# SECBIT_NOROOT_LOCKED (1 << 1) = 0x02
# SECBIT_NO_SETUID_FIXUP (1 << 2) = 0x04
# SECBIT_NO_SETUID_FIXUP_LOCKED (1 << 3) = 0x08
# Cumulative Mask = 0x0F (or 15)
capsh --secbits=0x0f \
--user=ci-worker \
-- -c "id; capsh --print"
Inside the sandboxed runner shell, verify what happens when a process attempts to invoke a standard SUID binary like su:
su -c "whoami"
Representative Terminal Output
uid=1000(ci-worker) gid=1000(ci-worker) groups=1000(ci-worker)
Current: =
Bounding set =cap_chown,cap_dac_override,cap_dac_read_search,cap_fowner,cap_fsetid,
cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_linux_immutable,cap_net_bind_service,
cap_net_broadcast,cap_net_admin,cap_net_raw,cap_ipc_lock,cap_ipc_owner,cap_sys_module,
cap_sys_rawio,cap_sys_chroot,cap_sys_ptrace,cap_sys_pacct,cap_sys_admin,cap_sys_boot,
cap_sys_nice,cap_sys_resource,cap_sys_time,cap_sys_tty_config,cap_mknod,cap_lease,
cap_audit_write,cap_audit_control,cap_setfcap,cap_mac_override,cap_mac_admin,
cap_syslog,cap_wake_alarm,cap_block_suspend,cap_audit_read,cap_perfmon,cap_bpf,
cap_checkpoint_restore
Ambient set =
Current IAB:
Securebits: 0f/0xf/4'b1111 (revert-caps)
secure-noroot: yes (locked)
secure-no-suid-fixup: yes (locked)
secure-keep-caps: no (unlocked)
secure-no-ambient-raise: no (unlocked)
uid=1000(ci-worker) gid=1000(ci-worker) groups=1000(ci-worker)
su: Authentication failure
(Under the hood, su fails because the kernel refuses to assign root capabilities to the process even though the binary has the SUID bit set, breaking the privilege elevation mechanism completely).
Line-by-Line Technical Analysis
--secbits=0x0f: Applies the combined bitmask:SECBIT_NOROOT($1$): Modifies the kernel behavior such that when a thread changes its UID to 0, it does not automatically receive full capabilities in its Permitted set.SECBIT_NOROOT_LOCKED($2$): MakesSECBIT_NOROOTpermanent; child processes cannot clear this bit even if they possessCAP_SETPCAP.SECBIT_NO_SETUID_FIXUP($4$): Disables the kernel's automatic capability adjustments during UID transitions.SECBIT_NO_SETUID_FIXUP_LOCKED($8$): Locks theNO_SETUID_FIXUPconfiguration permanently.Securebits: 0f/0xf/4'b1111: Confirms that bits 0 through 3 are active and locked.secure-noroot: yes (locked): Even if an exploit triggers an arbitrarysetuid(0)call, the resulting thread remains devoid of all capabilities.
SRE Next Steps
- Incorporate these securebit definitions into your container runtime configurations (e.g., CRI-O, containerd) or execution wrappers.
- Review the Linux prctl(2) reference for fine-grained capability isolation controls.
Scenario 5: Constructing Minimalist Capability Envelopes for Ephemeral Namespaces and Chroots
The Operational Scenario
You are developing an incident response sandbox to safely unpack and inspect untrusted tarballs and binaries recovered from compromised hosts. You need to create an isolated filesystem sandbox using chroot, but standard chroot environments are vulnerable to breakout techniques if the sandboxed process retains CAP_SYS_CHROOT, CAP_MKNOD, or CAP_SYS_ADMIN. You must construct a capability envelope that enters the chroot, drops UID, and strips the Bounding and Permitted sets down to only the basic capabilities needed for analysis.
Execution Command
Use capsh to enter the isolated directory tree /var/jail/sandbox, drop all capabilities from the bounding set except CAP_DAC_READ_SEARCH and CAP_SYS_PTRACE (for debugging), switch to an unprivileged account, and start a shell:
capsh --chroot=/var/jail/sandbox \
--drop=all \
--inh=cap_dac_read_search,cap_sys_ptrace \
--caps="cap_dac_read_search,cap_sys_ptrace=ep" \
--user=analyst \
-- -c "/bin/sh"
Inside the sandbox, test whether administrative operations (such as creating device nodes) are blocked:
mknod /dev/sdb1 b 8 17
Representative Terminal Output
mknod: /dev/sdb1: Operation not permitted
Checking the capabilities of the running sandboxed shell from the host:
capsh --decode=$(awk '/CapEff:/ {print $2}' /proc/$(pgrep -u analyst sh)/status)
0x0000000000080004=cap_dac_read_search,cap_sys_ptrace
Line-by-Line Technical Analysis
--chroot=/var/jail/sandbox: Instructscapshto callchroot("/var/jail/sandbox")while the process still retains the requiredCAP_SYS_CHROOTcapability.--drop=all: Empties the Bounding capability set entirely. No future binary executed inside this sandbox can ever assume any capability outside what is explicitly passed via ambient or inheritable flags.--inh=cap_dac_read_search,cap_sys_ptrace: Establishes an explicit Inheritable mask for forensic inspection tools.--caps="cap_dac_read_search,cap_sys_ptrace=ep": Sets the Permitted ($P$) and Effective ($E$) sets directly.--user=analyst: Drops process execution to an unprivileged forensic user.Operation not permitted: The kernel blocks themknodsystem call becauseCAP_MKNODis completely absent from the process's Effective and Permitted sets.
SRE Next Steps
- Pair this capability restriction with network namespace isolation to ensure zero network egress:
unshare -n capsh .... - Consult the ArchWiki Linux Capabilities Guide to fine-tune specific capability lists for forensic debugging tools.
What Can Go Wrong: Operational Hazards and Failure Modes
Working with Linux capabilities requires precision. Misconfigured capability sequences or missing flags can result in hard-to-diagnose startup errors or silent security vulnerabilities.
Monotonic Security Ceiling"] P["Permitted Set (P)
Absolute Process Limit"] I["Inheritable Set (I)
Matches File Capabilities"] E["Effective Set (E)
Active Syscall Enforcement"] A["Ambient Set (A)
Inherited Across Non-Root Execve"] X --> P P --> I P --> E I -->|Must be present in I and P| A A --> E
Pitfall 1: The Ambient Inversion Trap (EINVAL on Ambient Raise)
A frequent error when working with ambient capabilities is attempting to raise an Ambient capability without first ensuring it exists in both the Permitted and Inheritable sets:
# INCORRECT: Will immediately fail with EINVAL
capsh --user=appuser --ambient=cap_net_bind_service -- -c "/opt/app/server"
capsh: setting ambient capabilities: Invalid argument
Root Cause
Under the kernel rules enforced in kernel/capability.c, a capability cannot be added to the Ambient set unless it is already present in both the Permitted and Inheritable sets of the current thread.
The Fix
Always declare --inh alongside --ambient, and specify --keep=1 when dropping from root to an unprivileged user:
# CORRECT:
capsh --keep=1 --user=appuser \
--inh=cap_net_bind_service \
--ambient=cap_net_bind_service \
-- -c "/opt/app/server"
Pitfall 2: Privilege Leakage via Omitted --no-new-privs
If you drop capabilities and switch to an unprivileged user, but omit the --no-new-privs flag, the process remains vulnerable to privilege elevation via installed Set-UID binaries:
# VULNERABLE: Lacks --no-new-privs
capsh --drop=cap_sys_admin --user=sandboxed_user -- -c "/bin/bash"
If the unprivileged session executes a Set-UID root binary (such as /usr/bin/pkexec or a misconfigured administrative utility), the kernel executes the target with the full privileges of root, bypassing your intended capability restrictions.
The Fix
Always include --no-new-privs whenever constructing sandboxed environments:
# SECURE:
capsh --drop=cap_sys_admin --no-new-privs --user=sandboxed_user -- -c "/bin/bash"
Pitfall 3: Irreversible Bounding Set Truncation
Capabilities dropped from the Bounding set cannot be recovered by the current process or any of its child processes. If you drop a capability inside an interactive administrative shell, that privilege is permanently lost for the remainder of that terminal session:
capsh --drop=cap_sys_ptrace --
# From this shell, attempting to debug any process will fail permanently:
strace -p 1201
strace: attach: ptrace(PTRACE_SEIZE, 1201): Operation not permitted
The Fix
Use capsh as an execution wrapper for specific services rather than running commands inside an interactive shell, and launch target binaries using exec to cleanly replace the shell process. For full flag specifications, consult the official capsh(1) manual.
Today's Takeaway
Right now, open a terminal on one of your staging or development servers and inspect the effective capabilities of your running web server:
capsh --decode=$(awk '/CapEff:/ {print $2}' /proc/$(pgrep -u root -f nginx | head -n1)/status)
If the command returns a saturated mask (0x000001ffffffffff), your process is running with the legacy all-or-nothing root model, exposing the entire host to compromise if a daemon vulnerability is exploited. In just five minutes, you can begin isolating these daemons: test dropping unnecessary bounding capabilities, assign ambient capabilities for low-port binding, and harden your systemd units using capsh.