Setpriv: Dropping Process Privileges, Enforcing No-New-Privs Boundaries, and Hardening Unprivileged Services in Production
The post-incident post-mortem reveals a sobering truth familiar to seasoned system administrators: merely giving a background service an ordinary user account is not the ironclad vault we wish it were. On a typical operating system, a supposedly harmless program can still rummage through the file system, stumble upon an old administrative utility left behind by a forgotten maintenance script, and trick the system into elevating its access. Traditional administrative tools such as su, sudo, or runuser were built for an older era of interactive computing. Because they are themselves large, privileged programs with vast capabilities, using them to launch tightly restricted background jobs is like trying to build a secure bank vault with a master key taped to the front door.
This is where setpriv(1) comes to the rescue. Think of setpriv as a meticulous airport security checkpoint for your computerβs programs. Right before an application is allowed to start, setpriv strips away dangerous group memberships, permanently forbids the program from ever asking for higher powers, and hands over only the exact, microscopic abilities it genuinely needsβsuch as listening on a secure web port. Once the program launches through setpriv, the operating system locks those boundaries down so firmly that even if an attacker completely hijacks the running application, they remain trapped in an inescapable digital sandbox.
To inspect how your operating system currently views your process permissionsβor to see what privileges an application will have before you lock it downβthe most fundamental diagnostic tool at your disposal is the --dump flag:
# Interrogate the current process credential topology
setpriv --dump
uid: 0
euid: 0
gid: 0
egid: 0
supplementary groups: 0
no_new_privs: 0
Inheritable capabilities: (none)
Ambient capabilities: (none)
Capability bounding set: cap_chown,cap_dac_override,cap_dac_read_search,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_linux_immutable,cap_net_bind_service,cap_net_broadcast,cap_net_admin,cap_net_raw,cap_ipc_lock,cap_ipc_owner,cap_sys_module,cap_sys_rawio,cap_sys_chroot,cap_sys_ptrace,cap_sys_pacct,cap_sys_admin,cap_sys_boot,cap_sys_nice,cap_sys_resource,cap_sys_time,cap_sys_tty_config,cap_mknod,cap_lease,cap_audit_write,cap_audit_control,cap_setfcap,cap_mac_override,cap_mac_admin,cap_syslog,cap_wake_alarm,cap_block_suspend,cap_audit_read,cap_perfmon,cap_bpf,cap_checkpoint_restore
Securebits: none
When you need to run an untrusted batch script or automated task with ironclad isolation, the single most practical command you will ever run strips supplementary groups, enforces exact user IDs, and permanently disables privilege escalation with --no-new-privs:
# Execute the untrusted batch script within a strictly enforced no-new-privs envelope
setpriv --re-exec --no-new-privs \
--ruid 1001 --euid 1001 \
--rgid 1001 --egid 1001 \
--clear-groups \
/usr/bin/bash /opt/ci/build-task.sh
What It Does in Plain English
In simple terms, setpriv gives you direct, user-space control over the Linux kernelβs internal process credentials (struct cred) right before calling a program. Instead of launching a service with full administrator rights or blindly relying on standard user accounts, setpriv configures the precise execution envelope: dropping auxiliary groups, setting user IDs, configuring fine-grained Linux capabilities, and activating the kernel's one-way security switches.
| Flag | Category | Operational Mechanism |
|---|---|---|
--no-new-privs |
Execution Guard | Invokes the prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0) system call, permanently disallowing privilege escalation across execve() through SUID/SGID bits or filesystem capability attributes. |
--inh-caps <caps> |
Capability Set | Configures the inheritable capability set ($P_{\text{inheritable}}$), defining the set of capabilities preserved across an execve() boundary. |
--ambient-caps <caps> |
Capability Set | Configures the ambient capability set ($P_{\text{ambient}}$), enabling non-root processes to preserve and utilize specific POSIX capabilities across execve() without disk-based file capabilities. |
--bounding-set <caps> |
Capability Set | Manipulates the capability bounding set ($P_{\text{bounding}}$), establishing a strict mathematical upper bound on the capabilities a process tree can ever acquire. |
--securebits <bits> |
Security Flags | Modifies kernel securebits flags controlling UID transition semantics, root privilege disablement (SECBIT_NOROOT), and modification locks. |
--ruid <uid>, --euid <uid> |
Credentials | Sets the real and effective user identity (UID) for the target process. |
--rgid <gid>, --egid <gid> |
Credentials | Sets the real and effective group identity (GID) for the target process. |
--clear-groups |
Credentials | Empties the process supplementary group list via setgroups(0, NULL), neutralizing access to auxiliary host filesystems and administrative sockets. |
--re-exec |
Execution Order | Forces setpriv to re-execute itself internally to ensure that security options, capabilities, and user IDs are applied in the exact sequence required by the kernel. |
--dump |
Diagnostic | Interrogates the calling processβs current credential state, capability sets, securebits, and no_new_privs status, emitting an unadulterated state dump. |
Theoretical Foundations & Kernel Mechanics
To master setpriv, one must discard the simplistic Unix assumption that permissions are solely governed by the triad of rwx file permissions and UID 0. Modern Linux security is governed by five discrete capability vectors, fine-grained securebits bitmasks, and the irreversible execution boundary known as PR_SET_NO_NEW_PRIVS.
β’ UID / EUID / SUID / FSUID
β’ GID / EGID / SGID / FSGID
β’ Supplementary Groups List"] Flags["Security Bitmasks
β’ Securebits Flag Mask
β’ no_new_privs Bit Flag (0 or 1)"] end subgraph Caps["POSIX Capability Vectors"] P_perm["P_permitted: Upper bound on active capabilities"] P_eff["P_effective: Actively asserted operational privileges"] P_inh["P_inheritable: Capabilities preserved via execve()"] P_bnd["P_bounding: System-wide ceiling for task tree"] P_amb["P_ambient: Capabilities preserved across non-root execve()"] end StructCred --> Caps end
The Five Capability Vectors
Under the capabilities(7) subsystem, the monolithic privilege of root is partitioned into distinct units (e.g., CAP_NET_BIND_SERVICE, CAP_SYS_ADMIN, CAP_DAC_READ_SEARCH). Every thread in the kernel maintains five distinct capability bitmasks within its struct cred:
- Permitted ($P_{\text{permitted}}$): The definitive limiting superset of capabilities that the thread may actively assume.
- Effective ($P_{\text{effective}}$): The exact capability bitmask currently asserted by the kernel when performing operational authorization checks (e.g., authorizing a
bind()on port 443). - Inheritable ($P_{\text{inheritable}}$): A set of capabilities passed down through an
execve()call when the target binary on disk possesses matching file capability masks ($F_{\text{inheritable}}$). - Bounding Set ($P_{\text{bounding}}$): A kernel-level ceiling. Even if a binary on disk possesses elevated filesystem capabilities, the executing process cannot attain any capability excluded from its bounding set.
- Ambient Set ($P_{\text{ambient}}$): Introduced in Linux kernel 4.3 to solve the architectural challenge of running unprivileged daemons needing isolated root capabilities. Ambient capabilities survive an
execve()execution across non-root user boundaries without requiring extended filesystem attributes (setcap) on the target binary.
Transformation Equations Across execve()
When a process invokes the execve() system call, the kernel recalculates the new process credential sets ($P'$) using deterministic Boolean logic:
$$P'{\text{ambient}} = (\text{no_new_privs} \lor \text{SECBIT_NOROOT}) \ ? \ 0 : P{\text{ambient}}$$ $$P'{\text{permitted}} = (P{\text{inheritable}} \cap F_{\text{inheritable}}) \cup (F_{\text{permitted}} \cap P_{\text{bounding}}) \cup P'{\text{ambient}}$$ $$P'{\text{effective}} = F_{\text{effective}} \ ? \ P'{\text{permitted}} : P'{\text{ambient}}$$ $$P'{\text{bounding}} = P{\text{bounding}}$$
Prior to the introduction of $P_{\text{ambient}}$, passing a capability down to an unprivileged daemon required either setting extended attributes on the binary (setcap cap_net_bind_service=+ep /usr/bin/binary) or invoking the binary as root and requiring the binary itself to implement libcap to drop privileges internally. The former introduces a persistent disk vulnerability (any local user running that binary inherits the capability), while the latter requires the process to execute initialization code as UID 0. setpriv solves this cleanly by injecting the capability directly into $P_{\text{ambient}}$ and $P_{\text{inheritable}}$ before invoking execve().
The PR_SET_NO_NEW_PRIVS Invariant
Documented under the prctl(2) manual and the Linux Kernel Documentation on No New Privs, the no_new_privs bit is an immutable execution barrier. When no_new_privs is asserted (1):
- The
execve()system call guarantees never to grant privileges that exceed the privileges of the parent process. - Set-user-ID (
SUID) and set-group-ID (SGID) bits on binaries are completely ignored by the loader; they execute as the calling user. - Filesystem capability attributes (
setfcap) on binaries are completely ignored. - Linux Security Modules (LSMs) like AppArmor and SELinux are prevented from transitioning to a domain with greater permissions than the caller.
- Critical Invariant: Once enabled,
no_new_privscan never be cleared by the process, its child threads, or any subprocesses throughout its entire lifecycle.
The Securebits State Machine
The Linux kernel maintains a bitmask of security flags known as securebits (see capabilities(7)). These flags control how the kernel transitions capabilities when a process changes its UID:
SECBIT_NOROOT: Disables the traditional POSIX behavior where settingUID 0automatically grants all capabilities in $P_{\text{permitted}}$ and $P_{\text{effective}}$.SECBIT_NO_SETUID_FIXUP: Disables the automatic capability adjustments that occur when transitioning betweenUID 0and non-zero UIDs.SECBIT_KEEP_CAPS: Allows a thread to preserve capabilities in its permitted set when switching fromUID 0to a non-zero UID.- Locked Variants (
*_LOCKED): Each securebit has a corresponding lock bit (e.g.,SECBIT_NOROOT_LOCKED). Once locked, that bit cannot be toggled even by a process retainingCAP_SETPCAP.
Architectural Comparison: setpriv vs. Legacy Delegation Mechanisms
| Vector / Metric | sudo (8) | su (1) | runuser (1) | setpriv (1) |
|---|---|---|---|---|
| Execution Path | SUID root binary | SUID root binary | Direct syscall wrapper | Direct syscall wrapper |
| SUID Attack Plane | Severe | Severe | Moderate | Zero (Immune) |
| Enforces NNP | Optional / No | No | No | Deterministic (--no-new-privs) |
| Ambient Caps | Complex / No | Unsupported | Unsupported | Native (--ambient-caps) |
| Bounding Set Trim | Unsupported | Unsupported | Unsupported | Native (--bounding-set) |
| Securebits Locks | Unsupported | Unsupported | Unsupported | Native (--securebits) |
| Supplementary Groups | PAM Managed | Shell Managed | System Managed | Zeroed Explicitly (--clear-groups) |
Legacy tools such as sudo and su operate by leveraging the setuid binary bit to transition the execution context into root, execute PAM validation modules, and subsequently drop down into the designated user. If an exploit exists within PAM, dynamic linkers, or the binary itself, complete system compromise occurs. Conversely, setpriv is an unprivileged utility that executes the exact syscall sequence required to strip privileges downward before calling execve(), eliminating privilege escalation vectors entirely.
Five Enterprise Production Use Cases
The following five end-to-end production scenarios illustrate the defensive utilization of setpriv across systems engineering, container sandboxing, and telemetry collection workflows.
Use Case 1: Hardening Ad-Hoc Worker Execution with --no-new-privs
The Production Scenario
Your production CI/CD infrastructure executes dynamic, user-submitted shell scripts and third-party build containers. The build agent runs under a non-root service account (UID 1001, GID 1001). However, the host filesystem contains legacy system utilities with SUID-root permissions (e.g., /usr/bin/chfn, /usr/bin/gpasswd, or custom mounted utilities). A malicious build script attempts to exploit known local privilege escalation vulnerabilities in these SUID binaries to obtain root access to the node.
The Command
# Execute the untrusted batch script within a strictly enforced no-new-privs envelope
setpriv --re-exec --no-new-privs \
--ruid 1001 --euid 1001 \
--rgid 1001 --egid 1001 \
--clear-groups \
/usr/bin/bash /opt/ci/build-task.sh
Mock Terminal Execution & Verification
To demonstrate the deterministic protection of this command, consider a scenario where the worker script attempts to execute /usr/bin/passwd (a classic SUID-root binary):
# Simulating the untrusted worker attempting to invoke SUID binaries
setpriv --no-new-privs --ruid 1001 --euid 1001 --rgid 1001 --egid 1001 --clear-groups /bin/sh -c '
echo "=== Context State ==="
whoami
id
echo "=== Attempting SUID Binary Execution ==="
passwd --status
'
=== Context State ===
ci-worker
uid=1001(ci-worker) gid=1001(ci-worker) groups=1001(ci-worker)
=== Attempting SUID Binary Execution ===
passwd: Authentication token manipulation error
passwd: permission denied
Line-by-Line Technical Analysis
setpriv --no-new-privs: Invokesprctl(PR_SET_NO_NEW_PRIVS, 1). The kernel sets theno_new_privsbit intask_struct. From this nanosecond onward, any invocation of anexecvesystem call guarantees that SUID and SGID bits are rendered entirely inert.--ruid 1001 --euid 1001: Explicitly assigns the Real User ID and Effective User ID to1001viasetresuid().--rgid 1001 --egid 1001: Enforces strict Real and Effective Group ID alignment viasetresgid().--clear-groups: Invokessetgroups(0, NULL), purging all inherited supplementary group IDs (such asdocker,wheel, oradm), neutralizing auxiliary file permission bypasses.passwd: permission denied: The binary/usr/bin/passwdcontainsrwsr-xr-xpermissions on disk. Normally, execution transitionsEUIDto0. Becauseno_new_privsis active, the kernel ignores the SUID attribute;passwdexecutes strictly asUID 1001, fails to read/etc/shadow, and immediately aborts.
Next Operational Step
Integrate this invocation schema into systemd service definitions (ExecStart=...) or the low-level container runtime runner scripts, ensuring that all dynamic batch invocations inherit this immutable execution boundary.
Use Case 2: Granting Ephemeral Low-Port Binding via Ambient Capabilities
The Production Scenario
A mission-critical edge microservice written in Go (an HTTP/3 reverse proxy) must bind directly to privileged network ports (80/tcp and 443/tcp). Platform engineering policy strictly forbids running the application as root. Furthermore, security compliance disallows setting permanent extended file capabilities on disk (setcap 'cap_net_bind_service=+ep' /usr/local/bin/proxy) because local malicious actors could execute the binary to hijack privileged sockets.
The Command
# Launch the proxy as unprivileged user proxy-svc with ambient low-port binding capabilities
setpriv --re-exec \
--ruid 10002 --euid 10002 \
--rgid 10002 --egid 10002 \
--clear-groups \
--inh-caps +net_bind_service \
--ambient-caps +net_bind_service \
/usr/local/bin/edge-proxy --listen-addr 0.0.0.0:443
Mock Terminal Execution & Diagnostic Output
# Executing diagnostic probe to verify ambient capability inheritance
setpriv --ruid 10002 --euid 10002 --rgid 10002 --egid 10002 --clear-groups \
--inh-caps +net_bind_service \
--ambient-caps +net_bind_service \
-- /bin/sh -c '
echo "--- Process Credentials ---"
id
echo "--- Effective Capabilities in Runtime ---"
cat /proc/self/status | grep -E "Cap(Inh|Prm|Eff|Bnd|Amb)"
echo "--- Binding Privileged Port Test ---"
python3 -c "import socket; s = socket.socket(); s.bind((\"0.0.0.0\", 443)); print(\"SUCCESS: Bound to port 443\")"
'
--- Process Credentials ---
uid=10002(proxy-svc) gid=10002(proxy-svc) groups=10002(proxy-svc)
--- Effective Capabilities in Runtime ---
CapInh: 0000000000000400
CapPrm: 0000000000000400
CapEff: 0000000000000400
CapBnd: 000001ffffffffff
CapAmb: 0000000000000400
--- Binding Privileged Port Test ---
SUCCESS: Bound to port 443
Line-by-Line Technical Analysis
--inh-caps +net_bind_service: Sets the 10th bit (0x0000000000000400) within the process's inheritable capability mask ($P_{\text{inheritable}}$).--ambient-caps +net_bind_service: Crucially raisesCAP_NET_BIND_SERVICEin the ambient capability mask ($P_{\text{ambient}}$). Under Linux kernel rules, a capability must be present in both $P_{\text{inheritable}}$ and $P_{\text{permitted}}$ before it can enter $P_{\text{ambient}}$.CapEff: 0000000000000400: The bitmask calculation confirms that upon callingexecve()forpython3, the kernel automatically populated $P_{\text{effective}}$ withCAP_NET_BIND_SERVICE, despite the user beingUID 10002and the binary possessing no extended disk attributes.s.bind(("0.0.0.0", 443)): The kernel's socket subsystem evaluatessecurity_socket_bind(), queries $P_{\text{effective}}$, finds bit 10 enabled, and permits binding to low-port numbers ($< 1024$) without root credentials.
Next Operational Step
Verify with ss -tulpn | grep 443 that the process is bound and operating under UID 10002. Remove any lingering setcap file capabilities from binaries on disk across all deployment nodes.
Use Case 3: Restricting Process Capability Bounding Sets for Telemetry Collectors
The Production Scenario
A third-party observability daemon (node-telemetry-agent) requires administrative visibility to trace system calls (CAP_SYS_PTRACE) and bypass filesystem directory read restrictions (CAP_DAC_READ_SEARCH) to calculate disk utilization metrics. However, if the agent is compromised, you must guarantee that it cannot manipulate network routing tables (CAP_NET_ADMIN), alter kernel parameters (CAP_SYS_ADMIN), load kernel modules (CAP_SYS_MODULE), or alter file ownerships (CAP_CHOWN).
The Command
# Strip all bounding capabilities except DAC read and process trace
setpriv --re-exec \
--bounding-set -all,+dac_read_search,+sys_ptrace \
--ruid 10050 --euid 10050 \
--rgid 10050 --egid 10050 \
--clear-groups \
--inh-caps +dac_read_search,+sys_ptrace \
--ambient-caps +dac_read_search,+sys_ptrace \
/opt/telemetry/node-telemetry-agent
Mock Terminal Execution & Diagnostic Output
# Testing capability bounding set pruning
setpriv --bounding-set -all,+dac_read_search,+sys_ptrace \
--ruid 10050 --euid 10050 --rgid 10050 --egid 10050 --clear-groups \
--inh-caps +dac_read_search,+sys_ptrace \
--ambient-caps +dac_read_search,+sys_ptrace \
-- /bin/sh -c '
echo "=== Current Process Bounding Mask ==="
cat /proc/self/status | grep "CapBnd"
echo "=== Decoding Bitmask ==="
capsh --decode=$(cat /proc/self/status | grep "CapBnd" | awk "{print \$2}")
'
=== Current Process Bounding Mask ===
CapBnd: 0000000000080004
=== Decoding Bitmask ===
0x0000000000080004=cap_dac_read_search,cap_sys_ptrace
Line-by-Line Technical Analysis
--bounding-set -all,+dac_read_search,+sys_ptrace: First zeroes all 41 bits of the capability bounding set mask, then selectively sets Bit 2 (CAP_DAC_READ_SEARCH,0x4) and Bit 19 (CAP_SYS_PTRACE,0x80000). The resulting bitmask is0x0000000000080004.- $P_{\text{bounding}}$ Enforcement: This bounding set serves as an irreversible process ceiling. Even if an adversary exploits a local buffer overflow in
node-telemetry-agentand executes an embedded binary that is marked SUID-root or has full filesystem capabilities, the kernel limits the maximum attainable permitted capability set to exactly0x0000000000080004. - The process retains the absolute minimum capability profile necessary to perform filesystem walks across restricted paths and trace thread status via
/proc/$PID/, eliminating lateral kernel exploitation pathways.
Next Operational Step
Embed this capability bounding restriction into the container orchestrator security context or systemd service file (CapabilityBoundingSet=CAP_DAC_READ_SEARCH CAP_SYS_PTRACE).
Use Case 4: Setting Securebits to Permanently Prevent Root Re-Acquisition
The Production Scenario
In a hardened multi-tenant container runtime, dynamically scheduled jobs run under specialized isolation profiles. Standard Linux security mechanisms permit a process with CAP_SETUID or an execution path that reaches UID 0 to automatically regain full POSIX capability sets via standard kernel "setuid fixup" logic. You need to permanently break the association between UID 0 and kernel capabilities, ensuring that even if a thread transitions its identity to UID 0, it acquires zero kernel privileges.
The Command
# Launch a task with locked noroot and locked no-setuid-fixup securebits
setpriv --re-exec \
--securebits +noroot,+noroot_locked,+no_setuid_fixup,+no_setuid_fixup_locked \
--no-new-privs \
/usr/libexec/isolated-tenant-executor
Mock Terminal Execution & Diagnostic Output
# Inspecting securebits flags and capability retention behavior
setpriv --securebits +noroot,+noroot_locked,+no_setuid_fixup,+no_setuid_fixup_locked \
--no-new-privs --dump
uid: 0
euid: 0
gid: 0
egid: 0
supplementary groups: 0
no_new_privs: 1
Inheritable capabilities: (none)
Ambient capabilities: (none)
Capability bounding set: (none)
Securebits: noroot noroot_locked no_setuid_fixup no_setuid_fixup_locked (0x2b)
Line-by-Line Technical Analysis
--securebits +noroot: SetsSECBIT_NOROOT(Bit 0). When this bit is active, the kernel'scap_emulate_setxuid()function will not modify the processβs capability sets when transitioningEUIDto0.UID 0ceases to possess administrative superiority.+noroot_locked: SetsSECBIT_NOROOT_LOCKED(Bit 1). Freezes theSECBIT_NOROOTstate. No system call, regardless of privilege, can clear this flag for the lifetime of the process tree.+no_setuid_fixup: SetsSECBIT_NO_SETUID_FIXUP(Bit 2). Disables capability recalculation when changing between zero and non-zero UIDs.+no_setuid_fixup_locked: SetsSECBIT_NO_SETUID_FIXUP_LOCKED(Bit 3). Freezes the setuid fixup policy permanently.Securebits: 0x2b: Validates the combined bitmask ($1 | 2 | 8 | 32 = 43 = \text{0x2B}$). The execution thread is permanently decoupled from the historical superuser security model.
Next Operational Step
Deploy this securebit configuration wrapper around all dynamic runtime container entrypoints to ensure that multi-tenant processes cannot leverage legacy POSIX setuid semantics to escape unprivileged confinement.
Use Case 5: Establishing Hermetic Multi-Tenant Enclaves with Group Stripping and Explicit Discretionary Access Controls
The Production Scenario
You are designing an automated CI/CD runner host where isolated tasks must read from specific project directories but must be strictly blocked from accessing local UNIX domain sockets (such as /var/run/docker.sock, /run/containerd/containerd.sock, or /var/run/libvirt/libvirt-sock). These sockets are protected by supplementary group permissions (e.g., docker, containerd, or libvirt). When spawning execution threads, standard user transitions often inherit the parent's supplementary group memberships, allowing runners to send commands directly to administrative container engines.
The Command
# Execute the untrusted job within a completely isolated DAC enclave with explicit credentials
setpriv --re-exec \
--clear-groups \
--ruid 20001 --euid 20001 \
--rgid 20001 --egid 20001 \
--no-new-privs \
--bounding-set -all \
/usr/bin/runner-executor --task-dir /data/projects/tenant-alpha
Mock Terminal Execution & Diagnostic Output
# Spawning the sandbox and evaluating socket access and group membership
setpriv --clear-groups \
--ruid 20001 --euid 20001 \
--rgid 20001 --egid 20001 \
--no-new-privs \
--bounding-set -all \
-- /bin/sh -c '
echo "=== Process Identities ==="
id
echo "=== Auxiliary Socket Access Check ==="
if [ -S /var/run/docker.sock ]; then
cat /var/run/docker.sock 2>&1
else
echo "Socket check simulated: Access Denied via Discretionary Access Control"
fi
'
=== Process Identities ===
uid=20001(tenant-runner) gid=20001(tenant-runner) groups=20001(tenant-runner)
=== Auxiliary Socket Access Check ===
cat: /var/run/docker.sock: Permission denied
Line-by-Line Technical Analysis
--clear-groups: Executessetgroups(0, NULL). In standard Linux user management, if the parent process belongs to supplementary groups (GID 998(docker),GID 999(wheel)), simplesetuid()calls do not clear supplementary groups unless explicitly instructed.setprivstrips all secondary groups, reducing the process's group context strictly to its primaryEGID.--ruid 20001 --euid 20001 --rgid 20001 --egid 20001: Forces complete identity separation. Discretionary Access Control (DAC) verifies file access strictly againstUID 20001andGID 20001.--bounding-set -all: Eradicates all 41 capabilities from $P_{\text{bounding}}$. The runner cannot invokechroot, cannot intercept network packets, cannot manipulate mounts, and cannot attach debuggers (ptrace) to other processes.cat: /var/run/docker.sock: Permission denied: Because supplementary groupdockerwas expunged, the kernel enforces file modesrw-rw---- root dockeron the socket, blocking the runner from executing Docker daemon API commands.
Next Operational Step
Audit all execution pipelines using ps -eo pid,user,group,supgrp,args to confirm that runner subprocesses present empty supplementary group masks (supgrp -).
Operational Pitfalls, Diagnostics, and Failure Recovery
While setpriv provides deterministic security boundaries, improper parameter ordering or flag misconfigurations can lead to subtle failure modes.
--ambient-caps +net_bind_service ..."] --> S1["Success: Ambient matches Inheritable boundary"] B2["setpriv --re-exec
--bounding-set -all
--ruid 1000 ..."] --> S2["Success: Drops bounding set before identity change"] end
Pitfall 1: Ambient Capabilities Without Matching Inheritable Set
The Error
Attempting to raise ambient capabilities without configuring the inheritable set results in immediate kernel rejection:
setpriv --ambient-caps +net_bind_service /usr/bin/whoami
setpriv: libcap-ng is unable to update the ambient capability set: Invalid argument
Kernel Mechanics
The Linux kernel enforces a strict operational rule: a capability cannot be added to $P_{\text{ambient}}$ unless it is already present in both $P_{\text{permitted}}$ and $P_{\text{inheritable}}$. If you specify --ambient-caps alone, $P_{\text{inheritable}}$ does not contain the capability, causing the prctl(PR_CAP_AMBIENT, PR_CAP_AMBIENT_RAISE, ...) system call to return -EINVAL.
The Fix
Always define --inh-caps before --ambient-caps:
setpriv --inh-caps +net_bind_service --ambient-caps +net_bind_service /usr/bin/whoami
Pitfall 2: Ordering of Identity Drops and Privilege Modification
The Error
When invoking setpriv without --re-exec, developers often mistakenly sequence arguments such that user identity drops occur before capability bounding set or securebit configurations are evaluated.
# Potential race or permission failure in manual sequence
setpriv --ruid 1000 --bounding-set -all /usr/bin/id
setpriv: setresuid failed: Operation not permitted
Kernel Mechanics
Modifying the capability bounding set ($P_{\text{bounding}}$) requires the caller to possess CAP_SETPCAP. If setpriv switches the user identity to an unprivileged UID 1000 first, the process loses CAP_SETPCAP and cannot subsequently modify the bounding set or securebits.
The Fix
Use the --re-exec flag. This instructs setpriv to validate and parse all options first, configure capabilities, bounding sets, securebits, and groups in the exact, kernel-mandated sequence, and finally execute the target payload.
Pitfall 3: Troubleshooting with Diagnostic Interrogation
When debugging permission failures within an application launched by setpriv, do not guess the active state. Use /proc/$PID/status and capsh to decode runtime capability masks.
# Step 1: Query the hexadecimal capability status of the process
grep -E "Cap(Inh|Prm|Eff|Bnd|Amb)" /proc/$(pgrep edge-proxy)/status
CapInh: 0000000000000400
CapPrm: 0000000000000400
CapEff: 0000000000000400
CapBnd: 0000000000000400
CapAmb: 0000000000000400
# Step 2: Decode the hexadecimal vector into human-readable capabilities
capsh --decode=0000000000000400
0x0000000000000400=cap_net_bind_service
If CapEff does not match the requirements of your application, inspect whether no-new-privs or missing inheritable caps blocked the credential transfer across the execve() boundary.
Today's Takeaway
The era of delegating process execution through monolithic, SUID-laden utilities like sudo or relying solely on user IDs for process isolation has ended. In the next five minutes, open a terminal on your workstation and run setpriv --no-new-privs --dump to inspect your shell session's security posture; then, identify a single background daemon currently running as root solely to bind a privileged port or monitor system metrics, and rewrite its launcher script to utilize setpriv --re-exec --inh-caps --ambient-caps --no-new-privs. By shifting your infrastructure toward deterministic, kernel-enforced capability boundaries, you eliminate entire classes of privilege escalation vulnerabilities before malicious payloads can even attempt execution.