Skip to content

fix(bin): sync fork with upstream through 54663948 - #22

Merged
sanis merged 10 commits into
mainfrom
fm/fm-upstream-sync-12
Sep 2, 2026
Merged

sanis merged 10 commits into
mainfrom
fm/fm-upstream-sync-12

Conversation

@sanis

@sanis sanis commented Sep 2, 2026

Copy link
Copy Markdown
Owner

Intent

Sync this fork (origin = sanis/firstmate) with upstream = kunchenguid/firstmate.

Scope of this sync is pinned to the seven upstream commits that were the fork's
backlog at intake, reached through upstream commit 5466394:

5466394 fix(pi): deliver captain outcomes as deterministic transcript entries (kunchenguid#3312)
7d4b517 fix(bin): support process events under symlinked homes (kunchenguid#3484)
ee58e39 fix(bin): retire public follow-ups in remote homes (kunchenguid#3479)
f42a629 fix(bin): bound wake drain presentation lock waits (kunchenguid#3475)
f2ee922 fix(bin): defer inactive reconciliation during startup (kunchenguid#3480)
41d0ab3 fix: surface inbound Relay media to responding agents (kunchenguid#3442)
355f46f fix: isolate new Herdr server environments (kunchenguid#2792)

Requirements:

  1. Ship branch fm/fm-upstream-sync-12, created from the fork's main.
  2. Integrate upstream by MERGE, never by rebase. The fork's own 72 commits are the
    fork's history and every one of them must survive the operation.
  3. Resolve the two expected conflicts on the merits:
  • tests/lib.sh, which carries the fork-local fixture-root fix that resolves
    fixture roots physically so the claim gate passes where the system temp
    directory is a symlink. If upstream has changed the same region for its own
    reason, keep BOTH intents rather than dropping the fork's. If upstream has
    genuinely implemented the same physical-resolution behaviour itself, prefer
    upstream's version and say so explicitly.
  • tests/fm-public-followup.test.sh.
  1. Verify nothing was silently lost: confirm the fork-local divergences are still
    present, and confirm each upstream commit's substance actually arrived, rather
    than only confirming that the merge completed.
  2. Run lint and the full portable test suite. Where a test fails, first prove
    whether it also fails on the pristine pre-merge baseline before treating it as
    caused by this merge. A failure red on both sides is pre-existing and is
    reported, not fixed here. A failure red only on the merged head is a genuine
    merge regression and must be fixed before shipping.
  3. Deliver through the pipeline to a pull request with green CI.

Scope discipline:

This task renews the fork. It does not fix upstream's bugs, nor the fork's own
pre-existing ones. A defect found in incoming upstream code, or a pre-existing
problem this merge exposes, is REPORTED in the pull request description and not
fixed here, unless the merge itself cannot pass without the fix, in which case
that must be stated explicitly.

Pull request description constraints:

  • The description must contain no local filesystem path of any kind. Redact any
    quoted terminal output to a placeholder before it goes in the description.
  • Report the upstream advance beyond the pinned seven commits as a flag for a
    following sync, explicitly not as scope taken on by this one.

What Changed

  • Merged seven upstream commits into the fork by merge (never rebase), so all 72 fork commits survive: deterministic captain-outcome transcript entries for the pi branch extension, process-event support under symlinked homes, retirement of public follow-ups in remote homes, a bounded presentation-lock wait in the wake drain, deferred inactive reconciliation during startup, inbound Relay media surfaced to responding agents, and isolation of new Herdr server environments — across bin/, .pi/extensions/fm-branch-supervision.ts, docs, and the portable test suite.
  • Resolved both expected conflicts by keeping both intents rather than dropping the fork's. In tests/lib.sh, upstream has genuinely converged on resolving the fixture root physically, so upstream's form of that half is used (with its TMPDIR trailing-slash normalization); the fork's rollback on a resolution failure is retained because upstream returns without removing the directory mktemp just created, leaking a fixture root no cleanup registry tracks. In tests/fm-public-followup.test.sh, upstream's remote-secondmate fixture and bounded-lock traps sit alongside the fork's de-time-bombed clock (iso_utc_at helper plus PF_TEST_NOW-derived seed).
  • Fixed one genuine merge regression: upstream's bounded wake-drain lock wait exits zero while printing WAKE DRAIN SKIPPED, so the afk-return catch-up gate could read a skipped drain as a presented queue. bin/fm-afk-return.sh now treats that marker as a lifecycle failure that holds the gate with no open blocker, with the contract documented in the script header, docs/architecture.md, and covered by new cases in tests/fm-afk-return.test.sh.

Reported, Not Fixed

  • Upstream's fixture-root resolution leaks the freshly created directory when physical resolution fails. The fork's rollback keeps this suite correct here, but the defect remains in upstream's own code and is not fixed by this sync.
  • Upstream has advanced past the pinned commit 54663948 since this backlog was taken at intake. Those newer commits are deliberately outside this sync's scope and are flagged here as the starting point for the next one.

Risk Assessment

✅ Low: The fix is a four-line, correctly-streamed guard in the fork-only bin/fm-afk-return.sh that closes the only path where a skipped drain cleared the catch-up gate, leaves every upstream-merged file byte-identical to the pinned commit, and is covered by a deterministic both-directions regression test whose only weakness is one redundant assertion.

Testing

Exercised the away-mode catch-up gate end-to-end against the real wake drain with a contended queue lock: the fixed build refuses to clear catch-up (exit 3, explanatory lifecycle evidence, gate file retained, durable wake preserved, guard still blocking) and only clears after an uncontended retry that actually presents the wake and its acknowledgement command, whereas the pre-fix build on the identical scenario reports catch-up clear with the wake undrained. The new regression case fails before the fix and passes after, the six pre-existing fm-afk-return cases stay green on both sides, the neighbouring wake-queue and wake-drain suites pass, and bin/fm-wake-drain.sh was verified byte-identical to the upstream pin. No failures, no setup problems; worktree left clean and scratch homes removed.

Evidence: Captain CLI transcript — fixed build gates catch-up on a skipped wake drain, then clears after an uncontended retry

Source: Captain CLI transcript — fixed build gates catch-up on a skipped wake drain, then clears after an uncontended retry

\### FIXED build (target commit 3b3d183) — captain returns while the wake queue lock is held

$ bin/fm-afk-return.sh begin
fm-afk-return: catch-up must finish before the captain request
catch-up lifecycle: durable wake drain was skipped while the queue lock was held; retry catch-up before ordinary work
catch-up wake: WAKE DRAIN SKIPPED: queue lock remains held by live pid 29042 after 2s; retry on the next drain.
fm-afk-return: handle each blocker now, or close it with resolved [key=...] and append a durable reclassification reason, then run bin/fm-afk-return.sh check
[exit 3]

$ bin/fm-afk-return.sh guard   # can ordinary captain work proceed?
fm-afk-return: return catch-up is pending; remediate or durably reclassify every listed blocker, then run bin/fm-afk-return.sh check
[exit 3]

$ ls state/.afk-return-catchup ; wc -l < state/.wake-queue
state/.afk-return-catchup  (catch-up gate still open)
undrained durable wake rows: 1

--- lock holder exits; captain retries catch-up ---

$ bin/fm-afk-return.sh check
catch-up lifecycle: durable wake drain was skipped while the queue lock was held; retry catch-up before ordinary work
catch-up wake: WAKE DRAIN SKIPPED: queue lock remains held by live pid 29042 after 2s; retry on the next drain.
catch-up wake: 1788357704 1 signal task.status signal: needs-decision [key=demo]: deploy window needs a call
catch-up wake: wake annotation: latest wake-EVENT observed at drain, not current state: task.status: needs-decision [key=demo]: deploy window needs a call
catch-up wake: OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
catch-up wake: task [key=demo] needs-decision: deploy window needs a call
catch-up wake: OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 28999.1788357704.UqVh0W
fm-afk-return: catch-up clear; ordinary captain work may proceed
[exit 0]

$ bin/fm-afk-return.sh guard
[exit 0]

state/.afk-return-catchup  (gate cleared)
Evidence: Same scenario on the pre-fix build — catch-up reported clear while the durable wake was never drained

Source: Same scenario on the pre-fix build — catch-up reported clear while the durable wake was never drained

\### PRE-FIX build (merge commit, before 3b3d183) — same scenario

$ bin/fm-afk-return.sh begin
catch-up wake: WAKE DRAIN SKIPPED: queue lock remains held by live pid 40359 after 2s; retry on the next drain.
fm-afk-return: catch-up clear; ordinary captain work may proceed
[exit 0]

$ bin/fm-afk-return.sh guard   # can ordinary captain work proceed?
[exit 0]

state/.afk-return-catchup absent  (catch-up gate CLEARED)
undrained durable wake rows still in state/.wake-queue: 1
Evidence: Regression test fails before the fix, passes after

Source: Regression test fails before the fix, passes after

\### tests/fm-afk-return.test.sh :: test_skipped_wake_drain_keeps_catchup_gated

--- WITHOUT the fix (bin/fm-afk-return.sh from the merge commit) ---
ok - return catch-up precedes Bearings, owns live blocker remediation, preserves evidence once, and clears idempotently
ok - tmux and Herdr blockers require the same explicit durable reclassification before ordinary work
ok - needs-decision remains reportable without masquerading as a firstmate-actionable blocker
ok - AFK return re-drains published wakes until handling acknowledges
ok - away-mode re-entry fails closed while the prior return catch-up is pending
ok - check retries recorded terminal teardown and keeps catch-up gated until success
not ok - skipped wake drain cleared the catch-up gate (rc=0): catch-up wake: WAKE DRAIN SKIPPED: queue lock remains held by live pid 80670 after 1s; retry on the next drain.
fm-afk-return: catch-up clear; ordinary captain work may proceed
[exit 1]

--- WITH the fix (target commit) ---
ok - return catch-up precedes Bearings, owns live blocker remediation, preserves evidence once, and clears idempotently
ok - tmux and Herdr blockers require the same explicit durable reclassification before ordinary work
ok - needs-decision remains reportable without masquerading as a firstmate-actionable blocker
ok - AFK return re-drains published wakes until handling acknowledges
ok - away-mode re-entry fails closed while the prior return catch-up is pending
ok - check retries recorded terminal teardown and keeps catch-up gated until success
ok - skipped wake drain keeps catch-up gated until an uncontended drain presents the queue
[exit 0]
Evidence: Key contrast (fixed vs pre-fix, same contended-lock scenario)
FIXED:
$ bin/fm-afk-return.sh begin
fm-afk-return: catch-up must finish before the captain request
catch-up lifecycle: durable wake drain was skipped while the queue lock was held; retry catch-up before ordinary work
catch-up wake: WAKE DRAIN SKIPPED: queue lock remains held by live pid ... after 2s; retry on the next drain.
[exit 3] gate retained, 1 undrained durable wake row preserved

PRE-FIX:
$ bin/fm-afk-return.sh begin
catch-up wake: WAKE DRAIN SKIPPED: queue lock remains held by live pid ... after 2s; retry on the next drain.
fm-afk-return: catch-up clear; ordinary captain work may proceed
[exit 0] gate CLEARED, 1 undrained durable wake row still queued

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 infos
  • ⚠️ bin/fm-wake-drain.sh:421 - Upstream commit f42a629 (bounded wake-drain lock waits), arriving byte-identical in this merge, adds a new non-fatal outcome: when the queue lock is held by a live pid for longer than FM_STATUS_PRESENTATION_LOCK_TIMEOUT (default 10s), fm-wake-drain.sh prints 'WAKE DRAIN SKIPPED: ...' on stdout and exits 0 (bin/fm-wake-drain.sh:414-425). The fork-only consumer bin/fm-afk-return.sh:162 predates that contract and reads exit 0 as 'drained'. Concrete sequence: captain runs bin/fm-afk-return.sh begin while the watcher (or a concurrent drain) holds $STATE/.wake-queue lock >10s -> drain exits 0 with only the SKIPPED line on stdout -> lifecycle_ok stays 1; no WAKE_ACK_REQUIRED on stderr so the invalid-ack branch never fires; scan_open_blockers (bin/fm-afk-return.sh) only inspects .meta/.status blocked decisions and so cannot see an undrained queue -> the function reaches rm -f &#34;$GATE&#34;; clear_delivery_artifacts and prints 'catch-up clear; ordinary captain work may proceed' (bin/fm-afk-return.sh:212-215) while durable wakes were never drained or presented. Pre-merge, fm_lock_acquire_wait blocked until acquisition, so the drain could not return 0 without draining; this is a merge-introduced weakening of the fork's catch-up gate invariant. The skip line is still appended as wake evidence, so it is visible rather than silent, and the daemon consumer (bin/fm-supervise-daemon.sh:1696-1730) is fail-safe (handled=0 -> fallback wake, missing ack -> 'retaining durable wakes', return 1). Per the intent's scope discipline ('a defect found in incoming upstream code, or a pre-existing problem this merge exposes, is REPORTED in the pull request description and not fixed here'), this should be reported in the PR description rather than fixed in this sync; the remedy (teaching fm-afk-return.sh the new skip contract so it gates instead of clearing) extends the change beyond its stated intent, so it needs the author's authorization.

🔧 Fix: gate afk-return catch-up on skipped wake drain
2 infos still open:

  • ℹ️ tests/fm-afk-return.test.sh:344 - The retry direction of the new regression test is proven by lines 342-343 (rc=0 and gate removed), but the final assertion assert_contains &#34;$out&#34; &#39;catch-up wake:&#39; cannot distinguish a real re-presentation from leftover evidence. On the skipped run, fm-afk-return.sh:179 appends the drain stdout as wake evidence, so the gate durably holds the record evidence\twake\tWAKE DRAIN SKIPPED: queue lock remains held.... The check run calls preserve_evidence (fm-afk-return.sh:153), which copies that line into the retry's evidence file, and print_evidence (fm-afk-return.sh:106-112) emits it as catch-up wake: WAKE DRAIN SKIPPED: .... The assertion therefore passes even if the uncontended drain presented nothing at all. Strengthen it by asserting on something only the second drain can produce - the seeded row's payload (signal: .../task.status) or the retry's --ack-through &lt;n&gt; with n>0, which run_return captures via 2>&1. Test-only, non-user-visible.
  • ℹ️ bin/fm-wake-drain.sh:376 - Informational, no change requested: I traced the other outcomes of the upstream bounded-wait change to confirm the fix is correctly scoped to the one path that cleared the catch-up gate. STATUS PRESENTATION SKIPPED (line 376) returns 1 before status_commit_presentation_snapshot, so no cursor/snapshot is advanced and the caller at line 620 still emitted the raw wake rows and WAKE_ACK_REQUIRED - a deferred annotation, not a lost wake. STATUS PRESENTATION INCOMPLETE (line 384) sets rc=1 and likewise skips the commit. fm-session-start.sh:738 prints the drain output verbatim so the skip line reaches the captain, and fm-supervise-daemon.sh:1696-1730 retains durable wakes when handling does not complete. The non-124 lock failure exits 1 and is already caught by the pre-existing drain-failure branch. No additional gating is required, and no upstream file was modified.
✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-afk-return.test.sh — 7/7 ok on the target commit, including the new test_skipped_wake_drain_keeps_catchup_gated
  • Regression direction: temporarily restored bin/fm-afk-return.sh from HEAD~1 (the merge commit) and re-ran bash tests/fm-afk-return.test.sh — the new case fails (not ok - skipped wake drain cleared the catch-up gate (rc=0)), then passes again with the fix restored; worktree returned clean
  • Manual end-to-end captain transcript with the real bin/ scripts: seeded a durable wake via fm_wake_append, held $STATE/.wake-queue.lock with a live pid, ran bin/fm-afk-return.sh begin, bin/fm-afk-return.sh guard, then released the lock and ran bin/fm-afk-return.sh check (FM_STATUS_PRESENTATION_LOCK_TIMEOUT=2)
  • Same manual scenario replayed against the pre-fix bin/fm-afk-return.sh (HEAD~1) next to the unmodified upstream drain, showing the gate clearing while the wake stayed undrained
  • git diff --quiet 5466394 HEAD -- bin/fm-wake-drain.sh — confirms the drain is byte-identical to the upstream pin (no fork patch)
  • Exercised the fork-local fixture-root divergence: sourced tests/lib.sh and asserted fm_test_tmproot returns a physically resolved path (/private/tmp/... equal to its own pwd -P)
  • bash tests/fm-wake-queue.test.sh — exit 0
  • bash tests/fm-wake-drain-open-decisions.test.sh — exit 0
  • bash tests/fm-wake-drain-unread-status.test.sh — exit 0
⚠️ **Document** - 1 info
  • ℹ️ docs/watcher-continuity.md:63 - Judgment call on placement, not an open gap. docs/watcher-continuity.md owns the wake-drain presentation-deadline contract and already states that a live initial queue-lock holder makes the drain emit one PID-naming advisory and skip the whole drain. I did NOT add the consumer-side invariant ('a skipped drain is not a drained queue; a gating consumer must not read exit 0 as presented') there, because that would create a second prose copy of a contract whose decision procedure now lives in the bin/fm-afk-return.sh header - exactly the synchronization the one-owner rule forbids. The gate owner carries it instead, and the daemon consumer (bin/fm-supervise-daemon.sh) was already fail-safe by construction rather than by prose. Flagging only so the placement is a visible choice; no action needed unless a third exit-code-only consumer of the drain appears, at which point the invariant would be worth hoisting to the drain's own script header.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

RooseveltAdvisors and others added 10 commits September 1, 2026 08:05
* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose
)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract
)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots
…kunchenguid#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>
Merge upstream/main (7 commits) into the fork, preserving all 72 of our
own commits. Merged rather than rebased so the fork's history survives
intact.

Incoming upstream work:
- fix(pi): deliver captain outcomes as deterministic transcript entries (kunchenguid#3312)
- fix(bin): support process events under symlinked homes (kunchenguid#3484)
- fix(bin): retire public follow-ups in remote homes (kunchenguid#3479)
- fix(bin): bound wake drain presentation lock waits (kunchenguid#3475)
- fix(bin): defer inactive reconciliation during startup (kunchenguid#3480)
- fix: surface inbound Relay media to responding agents (kunchenguid#3442)
- fix(bin): isolate new Herdr server environments (kunchenguid#2792)

Two conflicts, both in test files:

tests/lib.sh - fm_test_tmproot carried two fork-local intents and upstream
took only one of them. Upstream now resolves the fixture root physically
itself, so that half has genuinely converged and upstream's form is used.
Upstream did not take the rollback half: its resolution failure path
returns without removing the directory mktemp had just created, leaking a
fixture root no cleanup registry knows about, and our own
tests/fm-test-fixture-cleanup.test.sh asserts that rollback. Kept both
intents - upstream's TMPDIR trailing-slash normalization and physical
resolution, plus our rollback that restores traversal before removing the
root, because a locked directory nested inside a locked parent is never
reached by a walk that could not descend into that parent.

tests/fm-public-followup.test.sh - both sides changed the same region for
unrelated reasons. Upstream added a remote-secondmate fixture with a
cleanup trap and a bounded-lock test; our fork had de-time-bombed the
suite by pinning the clock and deriving thread windows from it instead of
literal calendar dates. Kept both intents: upstream's fixture and traps
alongside our exported FMX_NOW_OVERRIDE, iso_utc_at helper and
PF_TEST_NOW-derived seed windows.
@coderabbitai

coderabbitai Bot commented Sep 2, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Team

Run ID: aa55a61e-9d37-4d2e-b51f-83a2c6ab71f6


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sanis
sanis merged commit 79f7f4f into main Sep 2, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants