Skip to content

Wayfinder: fleet remediation — enforce the seams, kill silent failure (2026-07 audit) #247

Description

@rubenhensen

This issue is a wayfinder map. It is an index, not a store: each decision lives in exactly
one place — its own child ticket — and is only gisted here. The original epic body (the audit's
evidence record) is preserved verbatim in the first comment.
The session-by-session narrative that used to fill the Notes section is archived, verbatim, in
part 1,
part 2 and
part 3; newer
session notes are separate comments on this issue. Notes below is standing rules only.

Destination

Every open question in the 2026-07 fleet-audit remediation programme is decided and implemented — on main, CI green, and where the change is a contract, carrying a gate that goes red if it regresses. The map is finished when no issue carrying wayfinder:247 remains open.

Notes

Domain. Nine repos across two orgs implementing PostGuard: IBE/IBS crypto (pg-core), a PKG service, C ABI + .NET bindings, a Rust file-transfer service (cryptify), a JS SDK monorepo with three email clients, an e2e harness, and Terraform ops. The organising insight from the July 2026 audit (11 repos, code review + a year of failure archaeology) is that the dominant failure class is cross-repo contract drift — contracts existed only as convention, so they drifted silently ~40 times a year. The strategy is: remove seams where possible, and make the surviving ones executable. Prose specs are explicitly out — fixtures and CI gates are the spec.

This map carries execution, not just decisions. Wayfinder normally stops at the decision; this effort deliberately overrides that. A ticket is done when the thing is built, not when the approach is chosen.

Done bar. Merged to main with CI green, and a gate that fails on regression where the ticket touches a contract. Publishing a release and deploying to prod are explicitly not required to close a ticket — with one deliberate exception, #348, whose deliverable is a published advisory. For the ops repo, "merged" means merged and applied, since Terraform's merge is its deploy.

Membership is the wayfinder:247 label, and so is the frontier. The tree hit GitHub's hard cap of 100 sub-issues on 2026-09-02, which silently stopped the map accepting new work and made four charted tickets invisible (#396). The label is now the membership record — on every member, open and closed, so a reopened ticket returns to the frontier — and the parent/child links are UI convenience. Unwire a child when you close it, in the same step as the resolution comment and the index entry; the cap counts closed children too, and this map filled 100 slots in five weeks. Run the query; do not read a list, and do not count leaves from a number written here:

# every open member, across all seven repos
for r in encryption4all/postguard encryption4all/postguard-js \
         encryption4all/postguard-e2e encryption4all/postguard-business \
         encryption4all/postguard-docs encryption4all/renovate-config \
         privacybydesign/postguard-ops; do
  gh issue list -R "$r" --label wayfinder:247 --state open --limit 200 \
    --json number,title,assignees --jq ".[] | \"$r#\(.number) \(.assignees|length) \(.title)\""
done

The frontier is the open, unblocked, unassigned members. This replaced a subIssues walk that only ever returned depth-1 children, so takeable leaves sitting inside umbrellas never appeared in it. Umbrellas are wired blocked by their own open children and drop off the frontier the moment those close — check them for the open-but-done shape rather than waiting to trip over it. A closed wired blocker is not the only blocker: no dependency edge models a release pipeline, an unpublished artifact, or "this PR before that one", so a ticket can read takeable and not be. Where that is true it is said on the ticket.

Division of labour. dobby (the dobby-coder GitHub App) writes the code and tests but cannot push .github/workflows/*.yml — it lacks workflows: write, so it leaves the YAML in a comment for a maintainer to apply. Both halves of that handover fail, in two distinct ways: a patch never posted (#324 — check the URL it is claimed to be at), and a patch posted correctly and never applied (#272, and again on postguard-e2e — checking the URL passes, only reading the file catches it). postguard and postguard-js machine-read their own workflows (pg-core/tests/ci_wiring.rs, packages/pg-js/tests/ci-wiring.test.ts), so an unapplied patch there reds a suite; delivery.yml and all of postguard-e2e are unguarded. A workflow handover is verified by reading the file on main and then by reading the step's output in a run on main — a green PR is evidence about the PR, and a green wiring test is evidence about jobs and contexts. Neither is evidence a step landed. privacybydesign/postguard-ops has no coding agent at all — zero dobby-coder PRs in its history — so its tickets can never be dispatched however well written.

Verify before you build — and the thing to re-read is rarely the source file. Read the ticket's target at origin/main (git show origin/main:<path>), never the working tree. Then extend the same suspicion to every other kind of premise, because each of these has cost this map at least one cycle: a ticket's line numbers (stale, and once inverted); its citations (closing an issue is exactly what makes a reference to it wrong); its dates (a ticket can instruct an agent to write a stale date onto the one line recording currency); its closed sets (an allowlist has to be checked against every emitter, not against the reader — there are five, and checking two and inferring three is how pg-cli and dotnet were nearly dropped); its cost estimates (one was inverted by four minutes of running the formatter); and its premises against its own siblings, which keep moving while it waits, and not monotonically — one merge made #370 more dispatchable and less, at once. A PR is authoritative about its own closes, so ask it rather than inferring from timing: gh api graphql -f query='{repository(owner:"O",name:"R"){pullRequest(number:N){closingIssuesReferences(first:10){nodes{number}}}}}'.

An agent's self-report is unreliable in both directions, and silence carries no information. Four shapes, all observed: a handover reported that never happened (#324); "CI is green" three seconds after the push that started the runs, true of the previous sha (#343); "no changes were made" over a complete, green, 40-check branch, with an invitation to re-dispatch that would have discarded it (#362, then #364); and a /dobby on a pull request, which routes to the review pipeline and blocks nothing, so the PR merges as written (#384). Read the branch and the checks of the sha the PR points at now — never the narration. Dispatch acknowledgement latency is bimodal (seconds when dobby is free, unbounded when rate-limited) and its absence means nothing at all: it cannot distinguish not started from never arrived. Never re-fire on silence — ask the queue; a second /dobby buys a duplicate, not a retry. And verify a verification before trusting it: a watcher written to poll for acks reported clean negatives because its gh calls were failing into || echo 0.

Harvest owns more than reading a label, and it runs at both ends of a session. Harvest by label, not by state — a closed ticket still carrying wayfinder:in-flight is an unfinished harvest, and that has been the rule rather than the anomaly. Clearing the mark is not the harvest; a cleared mark is indistinguishable from a finished one. A closed dispatch ticket is not a finished one — diff what the PR touched against what the ticket asked for, because a trailer closes an issue on the strength of a branch nobody read. Three free commands stand between a finished branch and a review queue and none is visible to any query: un-draft (gh pr ready — four consecutive dobby dispatches delivered complete green work and lost only the finalize step), re-run a flake, and check the release PRs. Re-measure the mid-flight set every time rather than carrying the previous session's table forward — this map has twice described a PR as awaiting review that had merged days earlier, and once counted a PR that had been superseded six weeks before. An open-PR count cannot tell a stalled review from a dead PR. An in-flight mark older than a day with no branch is failed, not pending; an in-flight mark outside every frontier query is a claim nobody will collect — wire the ticket in or drop the mark.

The dispatch gate, and the five questions it does not ask. The gate is in the skill; what this map has added is what slips past it. Gate 1 is settled by reading the code the ticket points at, not by reading the ticket — a body can be specific, confident and well-written and still describe a move the tree does not offer. Ask "what artifact does this ticket write, and can the agent write that kind of artifact" before reading the body: a workflow file, an org-wide App installation, and edits to someone else's PR branch are all gate-5 rejects knowable from the deliverable's path or verb. Check a batch for file collisions, which no gate question asks — two dispatches against one file is a bad merge waiting to happen, and wiring blocked by at charting time for that reason alone is good practice. Check that a ticket's own tests are reachable under its own scope fence. Answer gate 3 by reading the target repo's test config, never by guessing which directory gates. And expect a pre-flight amendment: a dispatch body is checked against the code it names and not against the code that names it, so grep the identifier rather than re-reading the ticket's line numbers. Amendments demonstrably work — the tickets that need none are the ones a previous wayfinder session wrote to the gate, and whose premises were measurements rather than readings. Roughly seven in eight swept candidates fail, almost always on gate 1: the frontier's remaining wayfinder:task tickets are mostly grilling tickets wearing task labels. Relabel when you take one.

The fail-open family: for every new gate or signal, ask what it reports when its subject is absent rather than wrong. Instances, all real: a required check with a paths filter and nothing to report (#299); a metric that drops what it cannot parse, so an empty panel reads as a confident zero (#371); a wayfinder:in-flight query against a repo where the label was never created; an in-flight query against a repo with no coding agent; pnpm --filter skipping a package with no lint script, silently, ten lines below the guard that fixes exactly that (examples.yml); alloy validate exiting 0 on an empty endpoint URL; and the inversion — a check with something real to say and no permission to say it, where four days of awaiting approval is indistinguishable from four days of nothing wrong (#394). Two habits: test the checker against a known-bad input before believing its pass, and a rationale written into a code comment is a claim like any other.

Reaching main is roughly half the distance to a user. The other half is a release PR that arrives pre-gated — check-runs: 0, action_required, BLOCKED — and waits on a human who is never notified. The discriminator is the actor: app/github-actions (what secrets.GITHUB_TOKEN makes release-plz) gets zero check-runs where rubenhensen and app/dobby-coder get dozens. It is one approval per sha, and a release PR re-shas whenever main moves under it, so the busier main is the further the release falls behind — a treadmill, not a checkpoint. It has stranded a security fix on main and on no registry three times. #334 owns this. Not every zero-runs PR is that gate — CONFLICTING looks identical through gh pr checks. Check the release PRs at every harvest, because an empty in-flight sweep now reads as done while releases sit still. Downstream, a caret resolves a fix; it does not require one, and a pre-1.0 caret cannot cross a minor boundary — consumers pinned ^0.5.x can never resolve a 0.6.x fix, which is the one case cargo audit/npm audit cannot resolve for a consumer. And check that a fix reaches the code path before calling something either a gap or a vulnerability — it has pointed both ways here.

Measure registries and pipelines together, and believe the pipeline first. A registry read lags a publish and fails in the exact shape of one that never happened: for 2m08s after @e4a/pg-js@2.6.0 published, dist-tags.latest was stale, GET /<pkg>/<version> 404'd, time.modified read three weeks old, and cache-busting changed none of it — the JSON endpoint is no better than npm view, both read one lagging document, and no registry field distinguishes not yet from never. The pipeline log is the authority; the registry only confirms. One gh api .../jobs call separates the publish failed from it has not run yet from it landed and the registry lags, and only the first is a finding. crates.io's API needs a User-Agent or its 404 looks like a crate never published.

Repo posture: admin merges are the normal path here, deliberately. dobby-coder cannot self-approve and neither can a maintainer authoring their own PR. Classic protection keeps enforce_admins: false, so --admin still walks past the review requirement — but a separate ruleset main: required checks with bypass_actors: [] means it cannot walk past a red or absent gate. postguard-js has the same posture via a ruleset split. bypass_actors is ruleset-level, not per-rule — that is why differential enforcement needs two objects, and why re-adding one actor to the wrong ruleset silently makes 20 contexts advisory again. This repo's required-context registry is asserted by the ruleset-drift job against the pins in ci_wiring.rs; postguard-js's 20 are still unguarded (#331). Both repos squash-merge, so git merge-base --is-ancestor reports not on main for work that plainly is — read the squash commit on main, not the PR's head sha.

History-preserving imports silently close issues here. ba380a14 closed live issue #146 via an imported commit's Closes encryption4all/postguard-website#146, resolved against this repo's numbering. Nothing warned. Audit closing keywords before and after any such merge; the guard is postguard-js#139. Measure the hazard rather than fearing it — postguard-dotnet's nine keywords all target already-closed items, which made a history-preserving import the cheap option there. And when retiring a repo, turn its release automation off first: release-plz-pr re-created cryptify's superseded release PR one minute after it was closed. Automation off and README finalized in one PR, then close the PRs, then transfer, then archive.

Repo consolidation is finished. All five source repos — the four app repos plus cryptify — are transferred and archived as of 2026-08-09 (#294), and the 62 issues stranded in archived read-only repos are moved. Nothing in the fleet is left half-moved.

Skills. /grilling and /domain-modeling for the wayfinder:grilling tickets; /prototype for wayfinder:prototype. Repo-durable knowledge lives in root CLAUDE.md — read it before touching any gate, and update it in the same PR when a ticket teaches something lasting. It carries its own 4,000-byte budget with a test behind it; do not relieve pressure here by spending it there.

Findings not yet in Notes

Decisions so far

Not yet specified

  • Which routes, fields, and headers actually get removed under #257, and when. Collection is ops#70/ops#71. #371 narrowed this patch rather than clearing it, and inverted what it is waiting on. The earlier note here said the data was unreadable as a deprecation signal until #371 added an unknown bucket; that is wrong, and the bucket does not make it readable. The split to carry: the path label is bounded by the server's own route table, so route-level removal is answerable once scraping runs — but no measurement distinguishes one client version from another when the client does not send the header, so that half waits on nothing and is an announcement plus a window, not a metric. Split it that way when it graduates rather than filing one ticket that half-can-be-answered.
  • What reaches Cockpit Loki, and what log_format prod should run instead. Opened by ops#72 decision 1: Cockpit log shipping is wanted in prod, and ops#73 leaves the Alloy blocks written and switched off behind two absent credentials, so turning it on is now a two-secret operation with no code change — which is exactly why the question has to be answered before someone does it. The procolix vhost sets no access_log, so nginx's default combined format applies and logs $remote_addr plus the full $request; cryptify's routes are GET /usage?<email>, /fileupload/<uuid>, /fileupload/<uuid>/status and /filedownload/<filename>. So the default arrangement ships client IPs, recipient email addresses and download references to a third-party log store — the same identifier set apps/website/src/lib/reportScrub.ts exists to strip out of Sentry. Not yet a ticket because the sharp question is which log_format (or which redacting map) prod should run, not whether the pipeline should exist. And nginx is not the only sink. cryptify's own application log writes the accounting key — since #402 a namespaced unproven:/proven:/api-key: key, so still the canonicalized claimed sender email on the default tier — together with that address's 14-day used_bytes, at info, on every finalize: the log::info! above the rolling-limit check in upload_finalize. Found while resolving #387 and deliberately left there, because changing what an operational log line says is a judgement about diagnostic value and belongs to this decision rather than to a security fix. So the question is which format and which application log lines, with two code owners rather than one — settling the nginx half alone would ship the same identifier set from a different file and read as done.
  • Whether we backport security fixes to the previous pre-1.0 minor line. A consumer pinned ^0.5.x can never resolve a 0.6.x fix, and no lockfile audit can fix that for them. The fix is either a patched 0.5.x or an advisory that says plainly no such patch exists. decide the fleet dependency-audit gate: tool per ecosystem, required or scheduled, and its voice #451 ruled it out of the audit gate: it is release policy, and it probably meets task: publish GHSA and RUSTSEC advisories for the sender-spoofing defect once the fix ships #348's advisory wording first.
  • Whether postguard-business and postguard-docs get required-check protection. business has no branch protection and no rulesets. docs has review-only protection. So decide the fleet dependency-audit gate: tool per ecosystem, required or scheduled, and its voice #451's PR check can only be advisory there (business#143, docs#132), and any other gate added to those repos is advisory too.

Out of scope

  • An RFC, a protocol registry, or capability negotiation. The audit's central negative finding: it would have prevented zero of the year's ~60 incidents. Recorded here so it is not re-litigated — executable contracts are the answer instead.
  • Product bugs the audit surfaced but deliberately tracks in place, not as children of this map: #197, the mobile Yivi cluster (postguard-website#264/#270/#271/#272), cryptify#47, postguard-website#280.
  • The superseded programme: EPIC #201 and its phases #202/#208/#213/#214/#215, closed when this epic replaced them. Its phases 0–3 delivered the harness, the pinned contracts, and the drift detector that this map builds on.
  • Zeroizing USK/MSK on drop in pg-core. A known-open security gap (no zeroize dependency yet), but never part of the audit's remediation set and never a child here. Needs its own effort.
  • Moving map-budget onto the fleet filing action. It comments on the map instead of filing an issue, a different kind of output (#460).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    epicLarge cross-cutting initiative tracked via sub-issueswayfinder:mapThe wayfinder map for an effort; its child issues are its tickets

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions