Skip to content

fix: bound watcher wake consumption for successor continuity - #99

Merged
dnth merged 1 commit into
mainfrom
fm/infoconnect-secondmate-watcher-recurring-death
Sep 4, 2026
Merged

dnth merged 1 commit into
mainfrom
fm/infoconnect-secondmate-watcher-recurring-death

Conversation

@dnth

@dnth dnth commented Sep 4, 2026

Copy link
Copy Markdown
Owner

Intent

Fix the recurring death of the OMP primary watcher successor chain in the shared Pi/OMP watcher core, bin/fm-primary-watch-core.ts. Root cause, proven in the scout report data/infoconnect-secondmate-watcher-recurring-death/report.md: since #97 (f0ec61a) sendWake awaited a before_agent_start acknowledgement that OMP only emits when a wake starts a new turn; a wake steered into an already-running turn never produces it, so owner.restoring stayed true forever and every later actionable close was enqueued with no successor arm and no follow-up. The InfoConnect secondmate home lost its watcher within a minute of every arm because its crew ends turns every 30-50s while the secondmate is still inside its own recovery turn; the sibling homes on pre-#97 code never died.

Required change, exactly and minimally: bound the sendWake consumption wait with a new positive-integer knob FM_WATCH_WAKE_CONSUME_TIMEOUT_MS (default 15000ms); on timeout delete the wake token and return generationIsLive(owner) - true, never false - so processPendingActionables does not re-deliver the same steer and the successor arm chain continues. Do NOT remove the restoring gate and do NOT re-arm from the close handler; those are deliberate. Idle-session wakes, where before_agent_start fires, must still return the real consumption result unchanged. The durable .wake-queue row is never touched by this path.

Acceptance criteria: AC1 the bounded wait as described; AC2 acknowledged path unchanged, timeout path returns true, durable queue row preserved; AC3 regression tests in tests/fm-omp-primary.test.sh and tests/fm-pi-watch-extension.test.sh prove that two consecutive actionable closes with no before_agent_start yield a third arm invocation within the bound, exactly one firstmate-watcher-wake per close, and no duplicate delivery (both tests fail against the unpatched core with a timeout waiting for the third arm and pass with the fix); AC4 tests/fm-omp-primary.test.sh, tests/fm-pi-watch-extension.test.sh, and tests/fm-watch-arm.test.sh pass and bin/fm-lint.sh is clean.

Documentation decisions: the knob is listed once in docs/configuration.md next to the other watcher continuity knobs; the one-sentence contract lives in docs/watcher-continuity.md, the continuity owner, plus one regression-coverage sentence there; a dated regression entry with exact commands and output is in docs/verification/supervision.md. AGENTS.md, skills, and the supervision protocol docs are intentionally unchanged because the extension-owned continuity contract they describe still holds.

Deliberate trade-off accepted: after the bound elapses, a wake that sat unconsumed inside a running turn is treated as consumed, so a /new or /resume within that window no longer carries it in the replacement handoff; the durable queue row still survives for the next drain, which is the guarantee the fleet ran on before #97. An adapter-level message_end acknowledgement is a possible follow-up and is intentionally not part of this change. This touches firstmate shared tracked material; the firstmate-coding-guidelines skill was read first.
Firstmate-Validation-Generation: 2c1f027f7e4c9e9b50bf66f4f9821638

What Changed

  • Added a bounded FM_WATCH_WAKE_CONSUME_TIMEOUT_MS wait (default 15000 ms) so unacknowledged wakes cannot stall the successor arm chain, while preserving durable queue rows.
  • Documented the timeout contract and added dated supervision verification coverage.
  • Added OMP and Pi regression tests for successor re-arming, one wake per close, and duplicate prevention.

Risk Assessment

✅ Low: The bounded acknowledgement wait, timeout consumption semantics, durable queue preservation, successor-chain behavior, and documentation are implemented consistently with the stated intent; no material source defect was substantiated.

Testing

Ran the OMP, Pi, and real arm-layer targeted suites; the new unacknowledged-wake scenarios demonstrate successor arm continuity, exactly-once delivery per close, and durable queue preservation. No source or transient worktree artifacts remain.

Evidence: OMP continuity regression evidence
ok - OMP unacknowledged wake delivery keeps the successor chain and delivers once per close
omp-unacknowledged-wake-ok
Evidence: Pi continuity regression evidence
ok - Pi unacknowledged wake delivery keeps the successor chain and delivers once per close

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-omp-primary.test.sh
  • bash tests/fm-pi-watch-extension.test.sh
  • bash tests/fm-watch-arm.test.sh
  • Inspected the timeout path in bin/fm-primary-watch-core.ts and confirmed token cleanup, true timeout result, and no durable queue mutation assertion in the regression test.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…eer cannot park the successor chain

The shared Pi/OMP watcher core (bin/fm-primary-watch-core.ts) waited forever
for a delivered wake to be acknowledged through before_agent_start. OMP only
emits that event when the wake starts a new turn; a wake queued as a steer
into an already-running turn never does, so owner.restoring stayed true and
every later actionable close was enqueued without a successor arm or a
follow-up. A secondmate whose crew ends turns every 30-50s lost its watcher
within a minute of every arm.

sendWake now races the acknowledgement against a new
FM_WATCH_WAKE_CONSUME_TIMEOUT_MS bound (default 15000ms); on timeout it drops
the wake token and returns generationIsLive(owner), so the pending loop never
re-sends the same steer and the successor chain continues. The restoring gate
and the close handler are unchanged.

Regression tests in tests/fm-omp-primary.test.sh and
tests/fm-pi-watch-extension.test.sh drive two consecutive actionable closes
with no before_agent_start and prove a third arm starts, exactly one wake is
delivered per close, and the durable wake-queue row is untouched; both fail
against the unbounded core.

Claude-Session: https://claude.ai/code/session_018sbgQjjCZMmpvcfvxyan1Q
@dnth
dnth merged commit 304e870 into main Sep 4, 2026
15 checks passed
@dnth
dnth deleted the fm/infoconnect-secondmate-watcher-recurring-death branch September 4, 2026 14:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant