Skip to content

fix(bin): preserve process identity across host clock steps - #52

Merged
knowttl merged 2 commits into
mainfrom
fm/fm-reland-clock-fixes
Sep 30, 2026
Merged

knowttl merged 2 commits into
mainfrom
fm/fm-reland-clock-fixes

Conversation

@knowttl

@knowttl knowttl commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Intent

Re-land two fixes that were dropped when the fork's local main was hard-reset to match upstream: "fix(bin): keep process identity stable across host clock steps" (#45, commit 59c9f6d) and "fix(bin): prevent Linux remote job worker pileups after clock changes" (#46, commit a97b7cd). Context: after the reset, the remote second mate host's remote job worker no longer starts on current code. fm-remote-doctor.sh --fix fails with "remote job worker did not report ready after startup", and each attempt leaves another stuck fm-remote-job-worker.sh process behind (23 on the host at last count). The captain approved dispatching this fix to get that host working again.

What Changed

  • Use Linux kernel start ticks for lab viewer, recovery sender, and remote worker process identities so host clock steps preserve ownership checks.
  • Preserve legacy ps lstart records and recognize remote worker lock owners by PID and exact command, preventing supervisor pileups and allowing replacement when worker code changes.
  • Add regression coverage for clock steps, PID reuse, legacy records, and worker upgrades; document the remote worker identity contract.

Risk Assessment

✅ Low: The changes are bounded to clock-stable process identity and legacy-record compatibility, with no substantiated correctness or intent-conformance defects.

Testing

Three targeted test files and live worker/watcher checks passed, including a pre-fix ownership counterexample. Clock effects were simulated because namespace isolation was denied. Herdr preparation was blocked, preventing real viewer and visual evidence. Disposable state was cleaned up.

  • Live validation: ⚠️ inconclusive - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Repeated worker startup recognizes a live owner with a drifted legacy record and retains exactly one supervisor ✅ pass live live-clock-check.log
The recovered worker executes a queued remote log read successfully ✅ pass live live-clock-check.log: job exit=0 and returned remote log payload
Redundant Linux supervisors drain, and replacing stale worker code stops the old process group before executing another job ✅ pass live tests/fm-remote-job.test.sh; remote-job-tests.log
Worker ownership rejects incorrect start ticks and an unrelated command ✅ pass live live-clock-check.log
The watcher preserves an active recovery sender across clock-related ps drift and escalates after that sender exits ✅ pass live live-clock-check.log: real fm-watch.sh runs with simulated ps clock drift
A real Herdr viewer remains owned across clock drift, accepts legacy records, and detaches safely ⏸️ untested no Tried the required lab prepare command with all home, configuration, and state paths isolated inside the worktree. Preparation refused because its safety tripwire requires a running default session. T…
Evidence: Live worker and watcher transcript

Source: Live worker and watcher transcript

ready worker pid=1443127 identity=starttime=31478212
wrong start ticks rejected for the live worker
unrelated command rejected even with a legacy start record
baseline rejects the still-running owner with a drifted legacy record
current code recognizes the same live legacy owner
ensure 1 kept worker pid=1443127
ensure 2 kept worker pid=1443127
ensure 3 kept worker pid=1443127
ensure 4 kept worker pid=1443127
supervisor count after repeated ensure=1
job exit=0
schema=fm-remote-delta.v1
status=delta
path=state/replies.log
from_offset=0
to_offset=29
from_prefix_sha256=e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855
to_prefix_sha256=e2662b5f4ee32c73b2af4accddeff00ec0481b974d2a0352ef94f7048c752ca5
payload_sha256=e2662b5f4ee32c73b2af4accddeff00ec0481b974d2a0352ef94f7048c752ca5
payload_bytes=29
reason=

clock recovery job completed
legacy ps sender identity changes when ps renders a clock step
live sender identity=proc-starttime=31478620 cmdline-hex=736c6565700031323000 phase=recovery_sending
legacy ps sender record still recognizes the real live process
mismatched sender identity rejected
check: rearm-resurface
dead sender outcome=unknown phase=escalated
worker tree stopped through the product cleanup interface
Evidence: Herdr live preparation blocker

Source: Herdr live preparation blocker

fm-herdr-lab: fleet-state tripwire requires exactly one running default session
- Outcome: ⏭️ skipped across 1 run (10m41s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⏭️ **Test** - skipped
  • ⚠️ live validation verdict: inconclusive (5 of 6 scenarios were driven live against the product); untested: A real Herdr viewer remains owned across clock drift, accepts legacy records, and detaches safely
  • Live validation: ⚠️ inconclusive - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Repeated worker startup recognizes a live owner with a drifted legacy record and retains exactly one supervisor ✅ pass live live-clock-check.log
The recovered worker executes a queued remote log read successfully ✅ pass live live-clock-check.log: job exit=0 and returned remote log payload
Redundant Linux supervisors drain, and replacing stale worker code stops the old process group before executing another job ✅ pass live tests/fm-remote-job.test.sh; remote-job-tests.log
Worker ownership rejects incorrect start ticks and an unrelated command ✅ pass live live-clock-check.log
The watcher preserves an active recovery sender across clock-related ps drift and escalates after that sender exits ✅ pass live live-clock-check.log: real fm-watch.sh runs with simulated ps clock drift
A real Herdr viewer remains owned across clock drift, accepts legacy records, and detaches safely ⏸️ untested no Tried the required lab prepare command with all home, configuration, and state paths isolated inside the worktree. Preparation refused because its safety tripwire requires a running default session. T…
  • TMPDIR="$PWD/.test-clock-tmp" bash tests/fm-remote-job.test.sh
  • TMPDIR="$PWD/.test-clock-tmp" bash tests/fm-pending-reply.test.sh
  • TMPDIR="$PWD/.test-clock-tmp" bash tests/fm-herdr-lab.test.sh
  • bash ~/.no-mistakes/evidence/01M3SNGYYS8919CG59S3FMWEW3/live-clock-check.sh
  • Executed the base commit's worker ownership interface against the same live worker: it rejected the drifted record; current code accepted it.
  • unshare --user --map-root-user --time true refused with Operation not permitted; clock effects were simulated without changing host time.
  • Ran isolated bin/fm-herdr-lab.sh prepare fm-lab-clock-gate; the required default-session tripwire was unavailable.
  • Stopped disposable worker trees, removed transient worktree files, and confirmed clean git status.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

* fix(bin): keep pending-reply sender and lab viewer identity stable across clock steps

The pending-reply recovery sender check and the Herdr lab viewer ownership
check identified processes by ps lstart text, which on Linux is the
wall-clock-derived boot time plus start ticks. A host clock step (WSL2 steps
about every 30 seconds) re-renders it, so a live recovery sender read as dead
and a running lab viewer pair read as not owned.

Both now use /proc/<pid>/stat start ticks where readable, like fm_pid_identity
and task_process_identity, and keep the ps form elsewhere. Records already
written in the legacy lstart form are still compared through ps, so an
upgrade does not strand an in-flight recovery or a running viewer.

* no-mistakes(document): Document clock-stable process identity and legacy records
…#46)

* fix(bin): keep Linux remote job worker identity stable across clock steps

The remote job worker identified its own processes (lock owner, staging
owner, job claims, lanes, command groups) by `ps -o lstart=` text. On Linux,
procps renders lstart from the current boot time, which moves whenever the
wall clock is stepped (NTP, VM or WSL2 time sync, resume). After a step a
healthy worker no longer matched its own lock record, so every remote call
started another detached supervisor beside it, the losers restarted for
minutes, the serving loop blocked on live lanes it thought had exited, a
competing worker reclaimed the live lock, and running jobs were published as
"remote job worker stopped before this job completed".

- Record Linux process identity as starttime=<stat field 22>, which no clock
  step moves; Darwin keeps ps lstart, unchanged.
- Keep records written by earlier workers comparable: an lstart record is
  compared as lstart, and a Linux lock owner still recorded as lstart is
  identified by pid and exact command, so an update replaces it in place and
  drains supervisors already piled beside it instead of stranding it.
- A serving worker that has lost its ownership lock now stops its own active
  execution and exits on a stop signal instead of re-arming, and never writes
  quarantine into a lock it does not own.

* no-mistakes(document): Document Linux remote worker identity and shutdown behavior
@knowttl
knowttl force-pushed the fm/fm-reland-clock-fixes branch from 4b6eb9f to 9223540 Compare September 30, 2026 18:15
@knowttl
knowttl merged commit 04ccd9f into main Sep 30, 2026
18 of 19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant