Skip to content

feat: align firstmate fork with upstream supervision updates - #50

Closed
knowttl wants to merge 51 commits into
upstream-basefrom
fm/fm-fork-realign-upstream
Closed

knowttl wants to merge 51 commits into
upstream-basefrom
fm/fm-fork-realign-upstream

Conversation

@knowttl

@knowttl knowttl commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Intent

Bring the firstmate fork (knowttl/firstmate) fully up to date with upstream (kunchenguid/firstmate), keeping the fork fixes that are still needed and dropping the ones that are not.
The captain's decision, following the fork drift review report: keep #44 (authorized follow-on phases: AGENTS.md section 10 tells firstmate to file each item as soon as its work is authorized, including later phases gated on another item or a date, and the secondmate charter treats later phases authorized in a routed message as routed work to file and dispatch), #45 (pending-reply and Herdr-lab process identity uses the Linux /proc start ticks instead of ps lstart so a host clock step does not make a live process read as dead) and #46 (Linux remote job worker process identity uses /proc start ticks, still comparing pre-upgrade lstart records); discard #48 (ready gated backlog work) and discard #49 with its revert (already a no-op).
The captain approved the review's recommended route: reset onto upstream and re-apply only the kept fixes, then fix up the fork's main with a force-push, which is a separate step that needs his explicit word.

What Changed

  • Bring in upstream supervision and worker updates, including attended /quiet behavior, Claude Code supervision notes, and a memory pressure guard.
  • Keep Linux remote job workers from being mistaken for stale processes after a host clock change, while accepting older process records.
  • Remove the discarded ready-gated backlog wake mechanism and its tests.

Risk Assessment

⚠️ Medium: The reviewed fork fixes are bounded, but this upstream realignment changes many additional files that remain unreviewed in this pass.

Testing

No baseline test commands were supplied. Four focused test scripts passed; live checks captured a Herdr viewer, process identities, generated charter text, and the removed command. The live firstmate filing flow stopped at Claude workspace trust, so the result is inconclusive.

  • Live validation: ⚠️ inconclusive - 5 of 7 scenarios driven live against the product
Scenario Result Live Evidence
A firstmate receives authorization for a later gated phase and files it immediately ⏸️ untested no Claude required interactive workspace trust before processing the request. Preapprove the disposable worktree for Claude, then rerun this live scenario.
A secondmate charter tells the agent to file authorized later phases on arrival ✅ pass live Generated secondmate charter command output and bash tests/fm-brief.test.sh
A pending-reply sender records Linux start ticks, and its clock-step and legacy-record checks pass ✅ pass live Live process identity command output; bash tests/fm-pending-reply.test.sh
A named Herdr lab viewer records start ticks and detaches without affecting the default session ✅ pass live Live Herdr lab viewer command output; bash tests/fm-herdr-lab.test.sh
A remote job keeps one worker across a drifted legacy lock and replaces it when its code changes ✅ pass live bash tests/fm-remote-job.test.sh exercised real worker processes, drifted legacy locks, replacement, and job execution.
The discarded ready-work command is unavailable ✅ pass live bash -c 'bin/fm-ready-work.sh' exited 127: No such file or directory.
The discarded #49 change adds no behavior to exercise ⏸️ untested no There is no live-validatable surface for a change already canceled by its revert.
Evidence: Live Herdr lab viewer
viewer attached to fm-lab-identity-2353028-7822; launcher_start=proc-starttime=17832825; viewer_start=proc-starttime=17832828; viewer_detached
Evidence: Live process identity
live_pid=2353043 pending=proc-starttime=17832799 worker=starttime=17832799 legacy_match=yes
Evidence: Generated secondmate charter
Later phases the main firstmate authorizes in a routed message are routed work: file each one in your backlog when it arrives, with its dependencies, and dispatch it when it becomes ready without waiting to be asked again.
- Outcome: ⚠️ 1 warning across 2 runs (12m25s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped
  • ⚠️ docs/turnend-guard.md - merge conflict rebasing onto origin/main
🔧 **Review** - 1 issue found ✅
  • 🚨 bin/fm-remote-job-lib.sh:1062 - A legacy Linux lock can outlive its worker. If that PID is reused by another worker launched with the same command, the starttime=*:*) branch accepts it without comparing process start time. fm_remote_job_worker_owned_alive (bin/fm-remote-job-lib.sh:1084) can then treat the new process as the lock owner, and fm_remote_job_start_linux_worker (bin/fm-remote-job-lib.sh:1206) can stop its process tree during an upgrade. Preserving ownership across clock steps is intentional, but safely distinguishing a drifted legacy owner from PID reuse needs an explicit compatibility decision; the legacy record lacks start ticks.

✅ No issues found.

⚠️ **Test** - 1 warning

✅ No issues found.

  • Live validation: ✅ go - 4 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Generate a secondmate charter and see instructions to file authorized later phases on arrival and dispatch them when ready ✅ pass live bash tests/fm-brief.test.sh generated and checked the secondmate charter
Read a live pending-reply sender's Linux start ticks and its pre-upgrade ps identity ✅ pass live Manual real-process identity check; bash tests/fm-pending-reply.test.sh
Attach a real Herdr lab viewer and detach it while keeping the lab session isolated ✅ pass live Named Herdr lab provision, viewer start, status, viewer stop, and teardown; bash tests/fm-herdr-lab.test.sh
Ensure a live remote worker with a drifted legacy record stays single and is replaced in place when its code changes ✅ pass live bash tests/fm-remote-job.test.sh ran a real worker, exercised a drifted legacy lock, and replaced the worker in place
Step the host clock and observe pending replies, Herdr viewer ownership, and remote workers retain live process identity while rejecting PID reuse ⏸️ untested no This worktree has no authority to change the host clock. A disposable VM with clock-step authority would permit a live check; the targeted tests simulated the clock change.
Inspect the realigned branch and find the three retained fork fixes above the upstream history ⏸️ untested no Git ancestry has no running product surface to drive live. The repository history was checked directly.
  • bash tests/fm-brief.test.sh

  • bash tests/fm-pending-reply.test.sh

  • bash tests/fm-remote-job.test.sh

  • bash tests/fm-herdr-lab.test.sh

  • Real-process fm_pending_reply_pid_identity and fm_pending_reply_ps_identity check

  • Named Herdr lab provision, viewer start, status, viewer stop, and teardown through bin/fm-herdr-lab.sh

  • git merge-base HEAD e9a6675ed188f3d77639cfe753451e07d68ab6a6, branch log, and git status --short

  • ⚠️ live validation verdict: inconclusive (5 of 7 scenarios were driven live against the product); untested: A firstmate receives authorization for a later gated phase and files it immediately, The discarded fix: prevent away-mode escalation injection wedges #49 change adds no behavior to exercise

  • Live validation: ⚠️ inconclusive - 5 of 7 scenarios driven live against the product

Scenario Result Live Evidence
A firstmate receives authorization for a later gated phase and files it immediately ⏸️ untested no Claude required interactive workspace trust before processing the request. Preapprove the disposable worktree for Claude, then rerun this live scenario.
A secondmate charter tells the agent to file authorized later phases on arrival ✅ pass live Generated secondmate charter command output and bash tests/fm-brief.test.sh
A pending-reply sender records Linux start ticks, and its clock-step and legacy-record checks pass ✅ pass live Live process identity command output; bash tests/fm-pending-reply.test.sh
A named Herdr lab viewer records start ticks and detaches without affecting the default session ✅ pass live Live Herdr lab viewer command output; bash tests/fm-herdr-lab.test.sh
A remote job keeps one worker across a drifted legacy lock and replaces it when its code changes ✅ pass live bash tests/fm-remote-job.test.sh exercised real worker processes, drifted legacy locks, replacement, and job execution.
The discarded ready-work command is unavailable ✅ pass live bash -c 'bin/fm-ready-work.sh' exited 127: No such file or directory.
The discarded #49 change adds no behavior to exercise ⏸️ untested no There is no live-validatable surface for a change already canceled by its revert.
  • bash tests/fm-pending-reply.test.sh
  • bash tests/fm-herdr-lab.test.sh
  • bash tests/fm-remote-job.test.sh
  • bash tests/fm-brief.test.sh
  • Provisioned, attached, inspected, detached, and tore down a named Herdr lab through bin/fm-herdr-lab.sh
  • Read pending-reply and remote-job identities from a live process
  • Generated a secondmate charter with bin/fm-brief.sh
  • Invoked the discarded bin/fm-ready-work.sh path
  • Launched a disposable Claude primary, inspected its trust dialog, and removed the lab
  • git status --short
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

knowttl and others added 30 commits September 23, 2026 08:19
* fix: file authorized gated work when it is authorized, not at dispatch

AGENTS.md section 10 told supervisors to file a backlog item before
dispatch, so a later phase authorized behind another item or a date was
never filed and stayed invisible to the teardown and session-start
re-evaluation. The secondmate charter's "act only on routed tasks" line
also read as needing a fresh route for each already-authorized phase.

File each item as soon as its work is authorized, including every later
phase gated on another item (blocked-by) or a date, and state in the
charter that authorized later phases are routed work to file on arrival
and dispatch when ready. The generated-charter test asserts the new rule.

* no-mistakes(document): Align secondmate routing guidance with authorized phases

* no-mistakes(ci): Restored AGENTS.md section 7 to its exact pre-b7fe1d5e routing sentence. No other files changed. git diff --check and fm-doc-audience-check passed. The two CI failures are unrelated pre-existing flakes being handled separately
* fix(bin): keep pending-reply sender and lab viewer identity stable across clock steps

The pending-reply recovery sender check and the Herdr lab viewer ownership
check identified processes by ps lstart text, which on Linux is the
wall-clock-derived boot time plus start ticks. A host clock step (WSL2 steps
about every 30 seconds) re-renders it, so a live recovery sender read as dead
and a running lab viewer pair read as not owned.

Both now use /proc/<pid>/stat start ticks where readable, like fm_pid_identity
and task_process_identity, and keep the ps form elsewhere. Records already
written in the legacy lstart form are still compared through ps, so an
upgrade does not strand an in-flight recovery or a running viewer.

* no-mistakes(document): Document clock-stable process identity and legacy records
…#46)

* fix(bin): keep Linux remote job worker identity stable across clock steps

The remote job worker identified its own processes (lock owner, staging
owner, job claims, lanes, command groups) by `ps -o lstart=` text. On Linux,
procps renders lstart from the current boot time, which moves whenever the
wall clock is stepped (NTP, VM or WSL2 time sync, resume). After a step a
healthy worker no longer matched its own lock record, so every remote call
started another detached supervisor beside it, the losers restarted for
minutes, the serving loop blocked on live lanes it thought had exited, a
competing worker reclaimed the live lock, and running jobs were published as
"remote job worker stopped before this job completed".

- Record Linux process identity as starttime=<stat field 22>, which no clock
  step moves; Darwin keeps ps lstart, unchanged.
- Keep records written by earlier workers comparable: an lstart record is
  compared as lstart, and a Linux lock owner still recorded as lstart is
  identified by pid and exact command, so an update replaces it in place and
  drains supervisors already piled beside it instead of stranding it.
- A serving worker that has lost its ownership lock now stops its own active
  execution and exits on a stop signal instead of re-arming, and never writes
  quarantine into a lock it does not own.

* no-mistakes(document): Document Linux remote worker identity and shutdown behavior
* fix(bin): wake a home when queued work becomes ready on its own

Queued backlog work gated on a hold date or on blockers could become
ready without any turn in the home - a date passing, or a blocker closed
by a captain answer, a hand-run tasks-axi done, or work elsewhere - and
nothing noticed until the next teardown or session start.

bin/fm-ready-work.sh owns the backstop: the watcher runs a ready-work
scan on the base heartbeat cadence and wakes once per readiness
transition, teardown names the work its own close unblocked and records
it as surfaced, and live-gated queued work (a future hold date, or
blockers that are in flight or live-gated themselves) now counts as
supervision need while undated holds never keep a watcher alive.

* no-mistakes(review): Surface gated work after durable wake delivery

* no-mistakes(review): Remove unused surface mode and document bare mutation limit

* no-mistakes(review): Use fixed ready-work timeout

* no-mistakes(document): Correct ready-work documentation and secondmate supervision limit

* no-mistakes(lint): Fix ShellCheck warnings in ready-work tests

* no-mistakes(ci): Fixed both lint failures by removing a redundant wake-library load and stopping ShellCheck from recursively analyzing the new ready-work import through watcher and supervision callers. The ready-work script remains linted as its own root. Local lint, source-aware checks, and ready-work behavior tests pass
* fix(bin): bound each away-mode digest so oversized escalations still deliver

escalate_flush joined the whole escalation buffer into one digest and
typed it as a single backend argument. A catch-all replay of a long
status span made that digest hundreds of kilobytes, which exceeds the
kernel's 128 KiB single-argument limit (herdr pane send-text never
execs) and tmux's ~16 KB command limit. The send failed before reaching
the pane, the buffer was kept, and every retry resent the same growing
digest until the captain returned.

Each flush now sends only the oldest items that fit FM_INJECT_MAX_BYTES
(default 1000, below tmux, the kernel, and the Claude-on-Herdr
head-truncation window), truncates an item too long to fit alone with a
marker naming its status log, and removes only the delivered lines so
the rest go in later batches. A backend send refusal is now logged as
such with the digest size instead of blaming the composer.

* no-mistakes(review): Bound escalation buffers and preserve every status-log pointer

* no-mistakes(review): Preserve status pointers and simplify escalation digests

* no-mistakes(review): Validate status pointers and report buffer update failures

* no-mistakes(review): Preserve escalation sources without parsing status text

* no-mistakes(document): Document bounded away-mode escalation batches

* no-mistakes(ci): Fixed the Bash 3.2 parse failure in bin/fm-afk-return.sh by using the metadata check already used by the daemon. The return test suite, repository lint, and diff check pass locally. Stock macOS Bash 3.2 is unavailable here, so CI must confirm that parser check
This reverts commit fad8567.

Superseded by upstream 683b3eb (bound the away digest and log why a
delivery failed) and 1d3ac67 (refuse a Herdr Claude submit that would
send only a message tail), which fix the same oversized-digest wedge.
Sync with upstream ef595d8. Conflict resolutions:
- AGENTS.md: keep both the fork's .ready-work* and upstream's
  secondmate-liveness state entries.
- bin/fm-watch.sh: source both fm-ready-work.sh and
  fm-secondmate-liveness-lib.sh.
- bin/fm-remote-job-worker.sh: take upstream's fd-bound lost-lock
  shutdown (a8572f6) over the fork's worker_owns_lock; keep the fork's
  start-tick identity call sites.
- docs/remote-secondmates.md: upstream structure plus the fork's Linux
  start-tick identity sentences.
- tests/fm-remote-job.test.sh: union of both sides' fixture variables
  and cleanup.
…h their launch config (kunchenguid#5799)

* fix(bin): pass the profile effort to OpenCode workers through their launch config

The dispatch profile's effort axis was recorded in task metadata but never
reached an OpenCode worker: the launch wrote only a permission grant into
the config it constructs.

OpenCode 1.18.32's config schema carries per-model reasoning effort as
agent.<name>.variant, so the chosen effort is now merged into the same
OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to
the resolved model. With no effort chosen the launch stays byte-identical.

Fixes kunchenguid#1373

* no-mistakes(review): gate OpenCode effort variant by model provider family

* no-mistakes(document): docs(opencode): note provider-family gating for effort variant
…nchenguid#5815)

* fix: preserve cancellation as no verdict in crew state

Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state.

Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation.

Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass.

* fix(review): Verify PR disposition before reclassifying terminal validation runs

* fix(test): Add captured cancellation replay coverage for resolver and fleet

* fix(document): Clarify cancellation and terminal delivery documentation
…henguid#5812)

* fix: declare worker background and pipeline waits

Require ship and scout workers to declare owned-work waits with the existing
paused verb before ending a turn or waiting on a pipeline or long command.
Keep the first-sight alert and existing liveness classification unchanged;
subsequent inspection follows the existing long pause cadence.

Validation: emitted brief regression failed before the instruction change
and passes afterward. Public watcher/drain regressions cover the first
alert, repeated wedge suppression, bounded rechecks, and undeclared idle
alarms using isolated backend fixtures. Brief suite, pinned lint, Bash
syntax, documentation inventory, and whitespace checks pass.
No real worker harness was exercised for wait behavior.

* fix(document): Clarify declared worker waits and documentation ownership

* fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
…5770)

* fix(bin): bound each lint root in its own ShellCheck process

CI job "Lint 1" died twice at about ten minutes because the two shard
workers each packed about 110 canonical roots into one unbounded ShellCheck
process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh
into a partition with other heavy roots, so the pair outgrew the 16 GiB
runner before anything could name a culprit.

Run one canonical root per ShellCheck process under an enforced envelope:
a wall deadline plus terminate-then-kill grace via the shared
fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the
child before exec (default a 4 GiB address-space cap, so two workers stay
inside a 16 GiB job with headroom). A root that exceeds the envelope fails
by name with a recorded reason - timeout, memory, signal, or
limit-unavailable - instead of taking the runner down. The per-root
watchdog runs in its own process group so the owner's group sweep cannot
orphan the bounded subtree, and fm_exec_timed now starts the same
escalation when its parent dies before it can be signalled.
FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a
configured bound cannot be enforced on the host rather than lint uncapped.
Each root's begin/end, reason, duration, and peak RSS stream to stderr in
partition mode and append to a retained <telemetry>.roots.tsv sidecar
uploaded beside the partition telemetry.

Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources
full analysis, complete and disjoint partition inventory, workflow lint,
and the backend-purity check, with byte-identical diagnostics across
jobs=1/2 proven by tests/fm-lint.test.sh.

* fix(bin): fail closed on unenforceable lint bounds and size the cap

Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the
run with named errors before any root starts: a missing fm-timeout-lib.sh,
a watchdog that cannot actually bound a probe command, or a host that
rejects the address-space limit all stop the run rather than lint uncapped.
The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a
single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record
the run's final exit status after backend-purity and workflow checks
instead of the pre-check lint status.

The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v
bounds virtual address space rather than resident memory, and ShellCheck's
GHC runtime keeps roughly a third of that space as reservation, so 6 GiB
yields about a 4 GiB working heap budget. A Linux measurement during this
change showed eleven real canonical roots running out of memory under the
earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB
resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job.
Roots that still exceed the cap keep failing by name, and the sidecar's
per-root peak RSS keeps roots approaching the budget visible.

tests/fm-lint.test.sh now proves the memory primitive where it can be
proven: on hosts that accept ulimit -v a perl allocator is refused under a
256 MiB limit and reported by name as a memory death, the pinned ShellCheck
lints a small file under the configured cap and is named when a far smaller
cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the
bounded cases skip on macOS, which cannot enforce the address-space limit.

* no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes

* no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller

* no-mistakes(document): Clarify bounded lint documentation and telemetry

* no-mistakes(document): Correct bounded lint documentation and sidecar path

* docs(bin): restore the per-root memory cap sizing rationale

The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and
dropped the sizing reasoning the change is required to record: address
space vs resident memory, the GHC reservation share, the measured 4 GiB
failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity
arithmetic. Restore it beside the default while keeping the corrected
"not a resident-memory ceiling" framing.

* no-mistakes(review): Document memory cap RSS reduction threshold and first candidate

* no-mistakes(review): Scope owner-death escalation docs to the perl watchdog

* no-mistakes(document): Clarify bounded lint and timeout documentation

* no-mistakes(review): Install perl watchdog signal handlers before forking the command

* no-mistakes(document): Correct bounded lint documentation and stale watcher comments

* no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear

* no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified

* no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM

* no-mistakes(review): Classify memory deaths from root stderr, not source excerpts

* no-mistakes(review): Match only whole runtime memory-error lines for memory reason

* no-mistakes(document): Clarify lint memory classification in script documentation

* no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots
…uid#5732)

* fix(bin): bound the watcher cleanup marker-lock wait

tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1
with "watcher pid <pid> did not exit within 10s of TERM". The watcher
had processed the TERM and was inside watcher_cleanup, where the
recovery-marker publish waits on state/.watcher-down.lock through an
unbounded fm_lock_acquire_wait. A live foreign holder of that lock
leaves the TERM'd watcher spinning in its own EXIT trap until the lock
frees or a second signal short-circuits the trap.

fm_recovery_transition now takes an optional bound and both
release-lock paths plus publish honour it through a new in-process
fm_lock_acquire_wait_max. watcher_cleanup passes
FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is
skipped, the singleton stays behind as ordinary dead-pid evidence, and
the next arm's clear-stale-lock still republishes it.

Regression test drives a real watcher with .watcher-down.lock held by
a live foreign process and asserts a single TERM still stops it.

* no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults

* no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention

* no-mistakes(document): Document bounded watcher cleanup and recovery

* no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation

* no-mistakes: apply agent fixes

* no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE

* no-mistakes(review): Hold marker lock through a failed cleanup acquire
…kunchenguid#5845)

* test: stop the leaked unreachable watcher before remote e2e cleanup

The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the
remote_env shell function, so $! named the function's subshell rather than
the watcher. Killing that subshell left the unreachable-leg watcher running,
and its one-second liveness probe kept invoking the fake ssh, which rewrites
ssh.count in the temp root. When a probe landed while the EXIT trap was
removing the root, rm failed with "Directory not empty" after every
assertion had passed.

Exec the watcher from the backgrounded function so the recorded pid is the
watcher itself, and assert the stopped watcher stops probing and writing its
state. Cleanup also stops a watcher left running by a failed assertion and
removes the root through fm_test_remove_tree, so a run that fails before
retirement does not strand the read-only spawn hooks directory.

Closes kunchenguid#5836

* no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal

* no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline
…nchenguid#4806)

* fix: stop quarantining ordinary shared-captain source updates

* no-mistakes(document): Rewrap remote inherit header so usage prints fully
* docs: make calm easier to read

Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved.

* docs: restore reload case in calm override lead-in

The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879)

* fix: route second-mate signal wakes by their new status span

A second mate's status log is a shared channel carrying many independently
keyed decisions, so judging its signal rows by every decision still open in
the whole log pinned each routine update to main behind any unrelated
parked hold. scopeForUnreadWake (the one owner for Pi and the attended
supervision host) now judges a second-mate signal row by the lines presented
since the last drain, bounded by the existing status-presentation cursor:
a decision, blocked, resolution, or captain-held line, or a line declaring
the key of a still-open decision, keeps the whole row on main, and any
cursor problem falls back to the whole log. Keys are read only at the
status parser's declared positions, with readable time stamps stripped as
bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are
unchanged, and stale and signal rows for one mate keep independent verdicts.

The supervision branch now treats a second mate's done and merged lines as
relayed child outcomes, and fm-teardown refuses the branch actor second-mate
retirement through the existing role-partition helper in both postures.

* no-mistakes(review): Route second-mate resolutions to main only when closing open decision

* no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally

* no-mistakes(document): Clarify second-mate wake routing and retirement documentation

* no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882)

register-extension took the extension lifecycle lock and then the source
lock, while reconcile republishing an unhandled extension result holds the
source lock and reaches the lifecycle lock through the extension host's
process-event path. Both waits are unbounded and both owners stay alive, so
the two could wait on each other forever and freeze the home's monitoring
cycle.

register-extension now takes the source lock first, matching every other
path that holds both. The lifecycle lock still spans binding resolution
through registration publication, so binding retirement stays serialized.

A new lifecycle-order section in the extension-binding suite, run in the
default aggregate, holds a re-registration inside binding resolution while
reconcile republishes that source's unhandled result and requires both to
finish within a bound.

Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath

The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before
looking up the repository's own hooks directory. When core.hooksPath reached git
through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child
process, the lookup found the wrapper directory again and exited 0, so the
repository's real hook - such as a pre-push publish guard - never ran and the
push succeeded.

The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's
config files decide its hooks directory, and a failed lookup exits nonzero
instead of skipping the hook. AI-trailer stripping is unchanged.

Fixes kunchenguid#5871

* no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
…hrashing guard (kunchenguid#5903)

* feat(jev): add the guard framework and the memory RSS/swap thrashing guard

A Jev guard is a bounded read-only host diagnostic that turns one class of
resource pressure into a machine-readable audit record and a one-line verdict.
This lands the framework contract (docs/jev-guards.md) with one representative
family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash
wrapper, stdlib-only python engine, and a behavioral test through the CLI.

* no-mistakes(review): fix jev mem guard fail-open unknown and contract

* no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs

* no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test

* no-mistakes(review): register docs/jev-guards.md in audience inventory

* no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds

* no-mistakes(review): assert exact exit code in fail-forcing test leg

* no-mistakes(review): tolerate any stdout encoding in text output

* no-mistakes(document): Fix guard contract dash style and output wording
…enguid#5900)

* fix(bin): stop slow GitHub reads from starving and waking the contributions poll

The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending.

* fix(bin): report the bound when a signal death leaks through fm_run_timed

fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallback so the suite runs on the macOS system bash.

* no-mistakes(review): Rotate contribution polling and verify generated budget behavior

* no-mistakes(review): Stabilize contribution rotation across successful observation refreshes

* no-mistakes(review): Exclude settled contributions from live observation rotation

* no-mistakes(document): Document contribution poll rotation and observation reserves

* no-mistakes(ci): Captain, fixed timeout handling with a one-line change preserving natural exit 137. Full session-start and contributions suites, targeted regressions, and lint passed. The timeout suite still fails on the known Bash 3.2 BASHPID issue, left unchanged as instructed. Logs retained in .no-mistakes/ci-evidence/. Remote CI was not rerun

* no-mistakes(ci): Fixed the fixture’s lock-acquisition race with a one-line bounded wait. Forced contention reproduced the CI error before the fix and passed afterward; the ordinary held-lock case, lint, syntax, and diff checks passed. Local Bash 3.2 failures remain: “cleanup lock bound 08 gave up before the marker lock freed” (also reproduced without the fix) and “TERM did not stop a watcher blocked inside a poll”. Remote CI was not rerun
…railers (kunchenguid#5859)

* feat: add keep AI trailers setting

* no-mistakes(document): Drop duplicate keep-ai-trailers line, update Cursor attribution note

* no-mistakes(document): Honor keep-ai-trailers for Devin worker attribution

* no-mistakes(document): Qualify Devin attribution note with keep-ai-trailers flag

* no-mistakes(review): Inherit keep-ai-trailers into secondmate homes
…sh-command popup cannot hide the composer (kunchenguid#5876)

* Fix Herdr composer reads blinded by the slash-command popup

Capture the full visible viewport for every Herdr composer state and content read instead of a bounded tail.
Claude Code renders its slash-command popup between the composer and the pane bottom, which pushes the composer outside a tail window.
The pre-Enter payload proof then read an empty composer, judged a typed command unsent, and cleared it without pressing Enter.
The proof-lines value now bounds only the clear cost, not the capture size.
Growing the window adds rows above the composer only, so bottom-most shape selection and prior verdicts are unchanged.
The suffix refusal is kept, and the shared inbox pending-line read stays a bounded tail.
Portable regressions cover the popup-below-composer layout, and the live submit-confirmation guard gains a third exit scenario.
The runtime-backends verification record documents the Herdr 0.9.0 and Claude Code 2.1.283 run.

Closes kunchenguid#5533

* no-mistakes(document): Clarify composer capture bound ownership

* no-mistakes(review): Drop unrequired bracketed-paste Enter fallback from live guard

* no-mistakes(review): Latch trust prompt, check idle composer arm first
…configurable bound (kunchenguid#5917)

Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s).

A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886)

* Keep lab tmux sockets on short private paths

* no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown

* no-mistakes(document): Update lab helper documentation for isolated tmux sockets
jcpoyser and others added 21 commits September 28, 2026 20:31
* Allow silent task-level no-change outcomes

* no-mistakes(review): Exclude silent outcomes from captain-return handoffs

* fix: look up supervision receipts by exact sequence

* no-mistakes(review): Suppress silent notes in away-return brief

* no-mistakes(review): Clarify visible notes; remove unused mode

* no-mistakes(review): Clarify silent outcomes and avoid false drain promises

* no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928)

* feat: make /quiet a statement where the attended supervision host runs

On a home that opted into the supervision host, quiet mode is what the
attended host already does, so /quiet now enters nothing there instead of
launching the quiet daemon and writing a record that would park a present
captain's main.

- bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the
  attended host runs, or that the session is paused while its
  broken-session latch holds; a quiet enter refuses there before writing
  anything.
- Where the home opted in but the attended host lacks a part (engine,
  tools, verified mirror writer, identifiable main session, valid mirror),
  quiet-check names it and quiet mode falls back to the daemon.
- Under a live away record on that home, quiet-check and a quiet enter
  refuse and name the record, so the return runs first, whatever
  state/.afk says.
- A quiet enter records mode: quiet in the posture record, so start and
  start-native launch the quiet daemon without FM_AFK_MODE, and the away
  refusal wording fires only for away.
- bin/fm-host-mirror.sh check validates the dialog mirror read-only and
  exits 1 on a missing, unreadable, or invalid mirror.
- The quiet and afk skills and the supervision-host docs describe the new
  behavior; homes without the opt-in and Pi homes keep the daemon path.

* no-mistakes(review): Archive the quiet record when a quiet daemon start fails

* no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper
…id#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
…ardown (kunchenguid#5997)

* WIP: retire task-keyed watcher markers and orphan journals at teardown

Re-applies old PR kunchenguid#5584 on current main: teardown retires the
turn-ended .seen-* signature and an orphaned Herdr presentation
journal whose workspace is already gone, and the wake-drain rotates
its own dead scratch files. Not yet validated through no-mistakes.

* no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
…unchenguid#6002)

* fix(tests): disable Claude Code's auto-updater during live harness runs

fm_live_gate let a live run proceed without ever setting
DISABLE_AUTOUPDATER, so a live Claude test could let the real updater
repoint ~/.local/bin/claude into a temporary directory and stop every
Claude process on the machine from starting. Export
DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run
proceed, and assert the export in tests/fm-live-gate.test.sh, including
that it reaches a child process the same way a real harness pane would
inherit it.

* no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope

* no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
…cevent record (kunchenguid#6010)

* fix(bin): ring the inbox doorbell only for a newly published procevent result

publish_result rewrote a worker's captured Lavish round idempotently on
every reconcile, unconditionally moved an already-acknowledged inbox
record back out of handled/, and rang the doorbell every time - so an
already-processed round rang the owning worker on every cycle. Snapshot
the existing active and handled records before the idempotent write and
ring, or move anything, only when the write actually created a fresh
record; re-delivery of a still-open round is left to the inbox's own
re-ring ladder.

* no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record

* no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
…unchenguid#6032)

* fix(bin): make the Claude Stop auto-arm refuse arguments before arming

A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real
supervision-host park owned by its short-lived tool process, leaving
supervision down once that process exited. The Stop hook passes no
arguments, so -h/--help now prints usage and any other argument is refused
before anything is sourced, read, or armed.

* no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations

* no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean

* docs: restore the bin/ toolbelt intro's manual-use clause

The document step dropped "interactive entrypoints work by hand too" from
docs/scripts.md, which still holds for most bin/ scripts.

* no-mistakes(document): Clarify Claude auto-arm manual-use guidance
…uid#6039)

* feat(calm): show supervision sailboat and anchor notes on Claude Code

The Calm mod follows a bounded display tail copy of the outcome store,
which bin/fm-branch-outcome.sh append now refreshes, and the supervision
host's latch, and appends one dim transcript line per visible routine
outcome, captain outcome, and latch change, replaying unread and
unprocessed outcomes at session start. It shows them whenever the mod is
active, regardless of config/calm, and never marks anything read.

* fix(calm): show each supervision note once per session on Claude Code

Claude Code 2.1.283 stores ui.log lines in the session and restores them
on --continue, so the mod records how far each session has followed the
outcome store and a resume replays only newer outcomes. It also checks
file existence before reads so absent files do not log debug errors.
The live guard gains the supervision-notes scenario and the dated
2.1.283 record documents the observed behavior.

* docs: name the Claude supervision note row as the engine draws it

* no-mistakes(review): Seed outcome tail on present and anchor first tail on markers

* no-mistakes(review): Seed outcome tail at session start; replay against start markers

* no-mistakes(review): Bound outcome tail by bytes; reread recently changed files

* no-mistakes(review): Skip store validation when outcome tail already exists

* no-mistakes(document): Clarify bounded Claude supervision note replay

* no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
…enguid#6033)

* fix(bin): read a quiet-mode record as a present captain, never hold-for-return

Daemon-backed quiet mode writes the away-posture record marked mode: quiet,
but the entry announcement, read-back, and session-start digest rendered it
as "hold-for-return only", and the spend cap and PR merge gate treated it as
away. A present captain's requested actions could then be held for a return
that was not coming.

bin/fm-afk-contract.sh now owns which posture a record is (the mode
subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet
record announces, reads back, and appears in the digest as a present captain
holding nothing; merges under it stay attended and it binds no spend cap. An
away record is unchanged, an /afk entry over quiet mode rewrites the record
as away, and a quiet entry never turns a standing away record quiet.

* no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance

* no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure

* no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed

* no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
…merge (kunchenguid#6053)

* fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged

require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for
the recorded pr= before refusing a different URL, so a task's later PR is
accepted once its earlier PR's merge is confirmed, while it keeps refusing
while the bound PR is still unmerged.

* no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…#46)

* fix(bin): keep Linux remote job worker identity stable across clock steps

The remote job worker identified its own processes (lock owner, staging
owner, job claims, lanes, command groups) by `ps -o lstart=` text. On Linux,
procps renders lstart from the current boot time, which moves whenever the
wall clock is stepped (NTP, VM or WSL2 time sync, resume). After a step a
healthy worker no longer matched its own lock record, so every remote call
started another detached supervisor beside it, the losers restarted for
minutes, the serving loop blocked on live lanes it thought had exited, a
competing worker reclaimed the live lock, and running jobs were published as
"remote job worker stopped before this job completed".

- Record Linux process identity as starttime=<stat field 22>, which no clock
  step moves; Darwin keeps ps lstart, unchanged.
- Keep records written by earlier workers comparable: an lstart record is
  compared as lstart, and a Linux lock owner still recorded as lstart is
  identified by pid and exact command, so an update replaces it in place and
  drains supervisors already piled beside it instead of stranding it.
- A serving worker that has lost its ownership lock now stops its own active
  execution and exits on a stop signal instead of re-arming, and never writes
  quarantine into a lock it does not own.

* no-mistakes(document): Document Linux remote worker identity and shutdown behavior
@knowttl
knowttl force-pushed the fm/fm-fork-realign-upstream branch from 9d83eda to f9d1688 Compare September 29, 2026 03:47
@knowttl
knowttl changed the base branch from main to upstream-base September 29, 2026 04:10
@knowttl

knowttl commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Superseded by #51 (clean branch off upstream-base; this PR's head was rebased onto the diverged fork main by a CI repair).

@knowttl knowttl closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.