Skip to content

feat(extruntime): raise capabilities at startup and report missing ones - #181

Draft
achoimet wants to merge 5 commits into
mainfrom
feat/raise-capabilities
Draft

achoimet wants to merge 5 commits into
mainfrom
feat/raise-capabilities

Conversation

@achoimet

Copy link
Copy Markdown
Member

Why

Extensions such as extension-container and extension-host run as a non-root user and get their capabilities as file capabilities of their binary (setcap …+eip). With the effective bit, the kernel refuses to exec the binary as soon as one listed capability is not granted to the container: a cluster whose policy forbids a single capability (e.g. NET_ADMIN) gets a crash-looping extension instead of one where only the network attacks are unavailable.

Dropping the effective bit (+ip) lets the binary start with whatever the container grants, but the capabilities then stay in the permitted set. This PR adds what the extensions need to work that way.

What

In extruntime:

  • RaiseCapabilities() moves the permitted set into the effective set on all threads (syscall.AllThreadsSyscall, which needs a binary built without cgo — the extensions build with CGO_ENABLED=0). Capabilities are per thread in Linux, so raising them on the calling thread only is not enough for Go.
  • MissingCapabilities(...) / RequireCapabilities(what, ...) check the bounding set: it is what the container's securityContext grants, and what the root helpers the extensions spawn (runc, nsenter, tc, …) inherit. Actions can fail fast in prepare with the missing capabilities named.
  • LogMissingCapabilities(...) warns at startup.
  • No-ops outside Linux.

Verified on Linux

A test binary with file capabilities cap_setuid,cap_setgid,cap_net_admin,cap_sys_admin, run as uid 10000 in Docker with several OS threads:

file capabilities NET_ADMIN granted result
+eip (today) yes starts
+eip (today) no exec: operation not permitted (the crash loop)
+ip yes starts, NET_ADMIN effective on all 8 threads
+ip no starts, MissingCapabilities reports NET_ADMIN, the granted ones effective on all 8 threads

Used by the follow-up PRs in extension-container and extension-host (setcap +ip, raise at startup, fail fast per action).

Extensions run as a non-root user and get their capabilities as file
capabilities of their binary. With the effective bit (setcap ...+eip), the
kernel refuses to exec the binary as soon as one of them is not granted to
the container, so an extension missing a single capability crash-loops.
Without it (+ip) the binary starts with what the container grants, but the
capabilities stay in the permitted set.

- RaiseCapabilities moves the permitted set into the effective set on all
  threads (AllThreadsSyscall; needs a binary built without cgo)
- MissingCapabilities and RequireCapabilities check the bounding set, which
  the container's securityContext decides and the root helpers spawned by the
  extension inherit, so actions can fail fast with the missing ones named
- LogMissingCapabilities warns at startup
@claude

claude Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Claude finished @achoimet's task in 2m 5s —— View job


Review complete

  • Check unresolved review threads from prior Claude reviews and resolve if fixed
  • Review new/changed code in this PR (correctness, security, concurrency, bugs)
  • Post findings / summary

Prior review thread: The earlier comment about the cgo fallback path applying Capset to an unpinned OS thread (capabilities_linux.go) was already addressed in 9b6215b and the thread is resolved — under cgo, RaiseCapabilities now raises nothing and returns an explicit error instead of silently raising on a single, unpinned thread. No action needed.

New changes since that review (commits 4eb4515, d183e38, 05bfde7): I re-read capabilities.go, capabilities_linux.go, capabilities_other.go, and both test files end-to-end.

  • The capability numbering table in capabilities.go matches the kernel's capability.h values (spot-checked NET_ADMIN=12, BPF=39, CHECKPOINT_RESTORE=40).
  • missingCapabilities's "usable when in bounding set AND (held OR root-helper-obtainable)" logic, and the no_new_privs gating of root helpers, is consistent with the documented kernel semantics and is exercised by the added table-driven tests (including the legacy-kernel BPF/SYS_ADMIN fallback).
  • missingHeldCapabilities / resolve correctly checks the effective set (including bits ≥32 via the two-word CapUserData combination) and shares the same legacy-capability fallback.
  • RaiseCapabilities preserves the inheritable set read via Capget before calling Capset, and its cgo-vs-no-cgo error handling matches the fix from the earlier round.

No new correctness, security, or concurrency issues found.

Comment thread extruntime/capabilities_linux.go Outdated
…e root helpers

MissingCapabilities only checked the bounding set, which is an upper bound:
it reported capabilities as present that neither the extension nor its root
helpers can use. It now requires the capability in the bounding set and
either held by the extension (permitted set) or obtainable by the root
helpers it spawns: SETUID and SETGID held and no_new_privs off. Under
no_new_privs (allowPrivilegeEscalation: false) an exec gains nothing beyond
what its parent held, so the helpers only get what the extension holds.

Capabilities newer than the kernel (BPF, PERFMON before 5.8,
CHECKPOINT_RESTORE before 5.9) were part of SYS_ADMIN, which now stands in
for them instead of reporting them as missing, e.g. BPF on kernel 5.4.
… itself

Some actions use a capability in the extension's own process (a syscall,
writing a file), which a root helper cannot do for them. They must check the
effective set, which also reports what RaiseCapabilities could not raise.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant