Skip to content

feat(preflight): check capabilities and kernel before loading BPF - #191

Merged
maxgio92 merged 10 commits into
mainfrom
feat/preflight
Sep 19, 2026
Merged

maxgio92 merged 10 commits into
mainfrom
feat/preflight

Conversation

@maxgio92

@maxgio92 maxgio92 commented Sep 12, 2026 •

Copy link
Copy Markdown
Owner

xcover run checks its environment before loading any BPF object so unsupported hosts get an actionable error instead of libbpf's bare "invalid argument". Missing CAP_BPF and CAP_PERFMON (or CAP_SYS_ADMIN) in the effective set is a hard failure that names the missing capabilities. A kernel release older than 6.6 is an advisory only, because distribution kernels backport uprobe_multi (RHEL 9.4 ships it on 5.14). --skip-preflight bypasses both checks and --userspace-bpf implies it.

Table tests cover the release parser, the capability decision and the skip paths.

@dodoazzurro dodoazzurro left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed #191 (adds internal/preflight — kernel version advisory + capability check before BPF load).

Logic checks out:

  • ParseRelease handles distro suffixes correctly (6.12.0-1-amd64 → 6.12, 6.18.44-r0-gcp-6.18 → 6.18), and the test table covers the tricky cases (6.6-rc1, non-numeric major/minor, empty string, bare "6").
  • Kernel check is advisory-only (logs a warning, never blocks) — correct, since backports (RHEL 9.4/5.14) can't be detected from the version string alone.
  • Capability check is the actual gate: CAP_SYS_ADMIN alone short-circuits as sufficient, otherwise both CAP_BPF and CAP_PERFMON are required — matches how the perfmon/bpf cap split works (pre-5.8 kernels only had CAP_SYS_ADMIN for bpf()).
  • --userspace-bpf correctly implies skipping preflight (no kernel BPF path involved there).

One thing worth a second look: daemonize() calls o.preflight() once in the parent (to fail fast before forking) and then the re-exec'd daemon calls it again via Run() → setup() → preflight(). Not wrong, just means the check — and any kernel advisory warning — fires twice when detaching. Harmless (cheap syscalls) but slightly redundant logging in the daemon's log file.

Nothing blocking.

… BPF

On kernels older than 6.6 (no BPF_TRACE_UPROBE_MULTI) or without the
needed privileges, xcover failed with libbpf's bare "invalid argument"
or an EPERM deep in the attach path.

Add internal/preflight with a kernel release parser tolerant of distro
suffixes, a comparison against the 6.6 minimum, and an effective
capability check that accepts CAP_SYS_ADMIN or CAP_BPF plus
CAP_PERFMON. Both checks return sentinel errors with actionable
messages. The version check is a heuristic, so the error names the
skip flag. The capability check reads the caller's effective set and
cannot see user-namespace confinement.
Run the kernel and capability checks after setup in the foreground path
and before forking in --detach mode, so the error reaches the user
instead of only the daemon log. --userspace-bpf implies skipping, since
bpftime needs neither kernel uprobe_multi nor privileges.
Distribution kernels backport uprobe_multi: RHEL 9.4 ships it on 5.14.
A hard uname-based gate refused hosts where xcover works. Warn when
the release looks older than 6.6, naming the backport case and the skip
flag, and keep the capability check as the only hard failure.
Kernel row: 6.6 upstream or a distribution backport, checked with a
warning at start. Privileges row: the effective capability set is
checked and cannot see user-namespace confinement. Regenerate the run
command page for --skip-preflight.
Kernel row: 6.6 upstream or a distribution backport, checked with a
warning at start. Privileges row: the effective capability set is
checked and cannot see user-namespace confinement.
The capability check fails before any BPF load, so its error matched
none of the runtime-environment markers and an unprivileged run failed
instead of skipping. Match the sentinel and cover the matcher.
The parent ran preflight to fail fast and the re-exec'd daemon ran it
again through setup, so the kernel advisory landed twice in the log.
Forward --skip-preflight to the child unless the user set it.
… CLI

The setcap hint named no file and did not mention that file
capabilities do not apply under go run or on nosuid mounts. The kernel
advisory recommended --skip-preflight, which also disables the hard
capability check, so the checker no longer mentions the flag.
Attach failures are still only warned about on this branch, so the
limitation still holds until the lifecycle fix lands.
@maxgio92

Copy link
Copy Markdown
Owner Author

Fixed: preflight now runs once when detaching. The child gets --skip-preflight because the parent has just run the same checks with the same credentials.

Also on this branch:

  1. The e2e harness skips on the capability error when unprivileged instead of failing.
  2. The setcap hint names the binary and says file capabilities do not apply under go run or nosuid mounts; the kernel advisory no longer recommends the skip flag, which would also disable the capability check.
  3. README keeps the "reports ready with 0% coverage" caveat until fix: fail loudly instead of writing a wrong report #185 lands.

The capability check reads the namespace-local effective set, so inside a
rootless container it can pass and the BPF load then fails with libbpf's
bare EPERM. xcover now warns when /proc/self/uid_map is not the identity
map. The troubleshooting page and the README lead with the preflight
message as the missing-privilege symptom and keep the libbpf error for
--skip-preflight, user namespaces and seccomp or LSM denials.
@maxgio92
maxgio92 merged commit 62352eb into main Sep 19, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants