Skip to content

fix: retain authorized follow-on work and stable process identity - #51

Closed
knowttl wants to merge 3 commits into
upstream-basefrom
fm/fm-fork-realign-upstream-clean
Closed

knowttl wants to merge 3 commits into
upstream-basefrom
fm/fm-fork-realign-upstream-clean

Conversation

@knowttl

@knowttl knowttl commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Intent

Bring the firstmate fork (knowttl/firstmate) fully up to date with upstream (kunchenguid/firstmate), keeping the fork fixes that are still needed and dropping the ones that are not.
The captain's decision, following the fork drift review report: keep #44 (authorized follow-on phases: AGENTS.md section 10 tells firstmate to file each item as soon as its work is authorized, including later phases gated on another item or a date, and the secondmate charter treats later phases authorized in a routed message as routed work to file and dispatch), #45 (pending-reply and Herdr-lab process identity uses the Linux /proc start ticks instead of ps lstart so a host clock step does not make a live process read as dead) and #46 (Linux remote job worker process identity uses /proc start ticks, still comparing pre-upgrade lstart records); discard #48 (ready gated backlog work) and discard #49 with its revert (already a no-op).
The captain approved the review's recommended route: reset onto upstream and re-apply only the kept fixes, then fix up the fork's main with a force-push, which is a separate step that needs his explicit word.

What Changed

  • Require firstmate and secondmates to file authorized later phases when routed, including work gated by dependencies or dates.
  • Use Linux /proc start ticks for pending-reply senders and Herdr-lab viewers, while continuing to recognize existing ps lstart records.
  • Use Linux /proc start ticks for remote job worker identity, preserve comparisons with older records, and recognize older workers for in-place replacement.

Risk Assessment

⚠️ Medium: The retained fixes match the stated intent, but the large upstream sync includes changed files outside this pass’s verified coverage.

Testing

All four targeted test scripts passed. A real named Herdr lab attached and detached a viewer with start-tick ownership, and the remote-job test exercised live worker recovery from a drifted legacy record. Live host clock adjustment was outside the workspace boundary; repository history established the upstream relationship. No UI layout changed, so no screenshot was captured.

  • Live validation: ✅ go - 3 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Generating a secondmate charter includes instructions to file authorized later phases on arrival and dispatch them when ready ✅ pass live bash tests/fm-brief.test.sh generated and checked the secondmate charter
A pending-reply recovery sender remains recognized after a host clock step ⏸️ untested no Stepping the host clock would change system state outside the worktree. A disposable Linux VM with clock-setting permission would allow a live clock-step check.
A real Herdr lab viewer attaches with start-tick ownership and detaches cleanly ✅ pass live Named lab provision, viewer start, viewer stop, and teardown completed; the live ownership record contained /proc start ticks
A remote job worker with a drifted legacy identity avoids duplicate supervisors and continues serving jobs ✅ pass live bash tests/fm-remote-job.test.sh started real workers, injected a drifted legacy lock record, verified one supervisor remained, replaced the worker, and ran a job
The fork contains the current upstream base and only the three retained fork fixes ⏸️ untested no Commit ancestry and the absence of discarded commits have no runtime product surface to drive; the available evidence is the repository history comparison.
Evidence: Live Herdr viewer ownership
Real named Herdr lab: viewer attached; launcher_start=proc-starttime=18046608; viewer_start=proc-starttime=18046611. Viewer stopped and lab teardown completed.

Base and landing

This PR targets upstream-base (a fork branch equal to upstream 2d833ff147cd26a5c461e914e06854e0eb2707ce) so it shows exactly the three retained fork commits on top of upstream.
The fork's main (eda9e3f2e72bcf9f0e2228e25853c176680cbb16) is untouched.
Landing is a separate step: a --force-with-lease push of the validated head to fork main, naming that SHA, after explicit approval.
Nothing in this PR merges or rewrites main.
The fork's CI workflows trigger only for pull requests to main, so no CI checks register on this base.
The pipeline's CI step was therefore skipped by ending its wait (run aborted after push and PR), and the green local validation (rebase, review, test, document, lint, push, PR, all with no findings) is the accepted evidence.
This was decided by the supervising firstmate under the captain's away instructions, because waiting out the four-hour CI timeout could never produce a check.

Kept and dropped

Known limitation

A legacy Linux lock record written before this change carries no start ticks.
During the one-time upgrade, such a lock that outlives its worker could be mistaken for a PID-reused worker launched with the same command, and that other worker's process tree could be stopped.
This is rare and applies only across the upgrade.
It is #46's existing behavior and was deliberately not redesigned here; it is tracked as a separate follow-up.

How running homes converge after the force-push

The guarded paths never force, so each home needs one manual reset after the force-push.

  • bin/fm-ff-lib.sh advances only by git merge --ff-only.
    A home on the old fork history reports skipped: diverged.
  • The redundancy proof (reset --keep) applies only to secondmate homes and only when their whole local result is already in the target.
    It fails here because fix: surface newly ready gated backlog work #48 is dropped, so those homes stay put and keep a state/.secondmate-update-reconcile/<id>.pending record.
  • The primary checkout gets no proof, so bin/fm-update.sh prints firstmate: skipped: diverged from origin/main.
  • bin/fm-fleet-sync.sh covers only projects/ clones, and /updatefirstmate inherits all of the above.
  • Primary checkout: git fetch origin && git reset --keep origin/main, first, because secondmate local sync follows the primary's main.
  • Local secondmate homes (nextrade-mate-d7, flex-mate-x1): git -C <home> checkout --detach origin/main, clear the stale reconcile marker, then bin/fm-secondmate-restart.sh.
  • Remote secondmate nx-remote-b1 on Brytton-Desktop: reset the code root first (bin/fm-remote-secondmate-control.sh update refuses on the divergence), then the home, then restart; confirm both are clean and at the old head before resetting.
  • Crew worktrees are pool slots that spawn from the default branch, so nothing needs resetting; an in-flight branch cut from fork main after fix: surface newly ready gated backlog work #48 must be rebased so it does not bring fix: surface newly ready gated backlog work #48 back.
  • After the resets, restart secondmates and re-arm the primary's supervision, because fix: surface newly ready gated backlog work #48's fm-ready-work.sh disappears; leftover state/.ready-work* files are inert.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - medium risk

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Generating a secondmate charter includes instructions to file authorized later phases on arrival and dispatch them when ready ✅ pass live bash tests/fm-brief.test.sh generated and checked the secondmate charter
A pending-reply recovery sender remains recognized after a host clock step ⏸️ untested no Stepping the host clock would change system state outside the worktree. A disposable Linux VM with clock-setting permission would allow a live clock-step check.
A real Herdr lab viewer attaches with start-tick ownership and detaches cleanly ✅ pass live Named lab provision, viewer start, viewer stop, and teardown completed; the live ownership record contained /proc start ticks
A remote job worker with a drifted legacy identity avoids duplicate supervisors and continues serving jobs ✅ pass live bash tests/fm-remote-job.test.sh started real workers, injected a drifted legacy lock record, verified one supervisor remained, replaced the worker, and ran a job
The fork contains the current upstream base and only the three retained fork fixes ⏸️ untested no Commit ancestry and the absence of discarded commits have no runtime product surface to drive; the available evidence is the repository history comparison.
  • bash tests/fm-brief.test.sh
  • bash tests/fm-pending-reply.test.sh
  • bash tests/fm-herdr-lab.test.sh
  • bash tests/fm-remote-job.test.sh
  • Real named Herdr lab: provision, viewer start, inspect ownership record, viewer stop, teardown
  • Live /proc and legacy identity reads through fm_pending_reply_pid_identity and fm_pending_reply_ps_identity
  • git merge-base HEAD origin/upstream-base, git diff --name-only origin/upstream-base..HEAD, and git log origin/upstream-base..HEAD
  • git status --short
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

* fix: file authorized gated work when it is authorized, not at dispatch

AGENTS.md section 10 told supervisors to file a backlog item before
dispatch, so a later phase authorized behind another item or a date was
never filed and stayed invisible to the teardown and session-start
re-evaluation. The secondmate charter's "act only on routed tasks" line
also read as needing a fresh route for each already-authorized phase.

File each item as soon as its work is authorized, including every later
phase gated on another item (blocked-by) or a date, and state in the
charter that authorized later phases are routed work to file on arrival
and dispatch when ready. The generated-charter test asserts the new rule.

* no-mistakes(document): Align secondmate routing guidance with authorized phases

* no-mistakes(ci): Restored AGENTS.md section 7 to its exact pre-b7fe1d5e routing sentence. No other files changed. git diff --check and fm-doc-audience-check passed. The two CI failures are unrelated pre-existing flakes being handled separately
* fix(bin): keep pending-reply sender and lab viewer identity stable across clock steps

The pending-reply recovery sender check and the Herdr lab viewer ownership
check identified processes by ps lstart text, which on Linux is the
wall-clock-derived boot time plus start ticks. A host clock step (WSL2 steps
about every 30 seconds) re-renders it, so a live recovery sender read as dead
and a running lab viewer pair read as not owned.

Both now use /proc/<pid>/stat start ticks where readable, like fm_pid_identity
and task_process_identity, and keep the ps form elsewhere. Records already
written in the legacy lstart form are still compared through ps, so an
upgrade does not strand an in-flight recovery or a running viewer.

* no-mistakes(document): Document clock-stable process identity and legacy records
…#46)

* fix(bin): keep Linux remote job worker identity stable across clock steps

The remote job worker identified its own processes (lock owner, staging
owner, job claims, lanes, command groups) by `ps -o lstart=` text. On Linux,
procps renders lstart from the current boot time, which moves whenever the
wall clock is stepped (NTP, VM or WSL2 time sync, resume). After a step a
healthy worker no longer matched its own lock record, so every remote call
started another detached supervisor beside it, the losers restarted for
minutes, the serving loop blocked on live lanes it thought had exited, a
competing worker reclaimed the live lock, and running jobs were published as
"remote job worker stopped before this job completed".

- Record Linux process identity as starttime=<stat field 22>, which no clock
  step moves; Darwin keeps ps lstart, unchanged.
- Keep records written by earlier workers comparable: an lstart record is
  compared as lstart, and a Linux lock owner still recorded as lstart is
  identified by pid and exact command, so an update replaces it in place and
  drains supervisors already piled beside it instead of stranding it.
- A serving worker that has lost its ownership lock now stops its own active
  execution and exits on a stop signal instead of re-arming, and never writes
  quarantine into a lock it does not own.

* no-mistakes(document): Document Linux remote worker identity and shutdown behavior
@knowttl

knowttl commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Landed: fork main was updated to upstream 00679ae plus the three kept fixes (a97b7cd), the same content as this branch rebased onto the latest upstream commit. Closing this validation vehicle.

@knowttl knowttl closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant