Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
d8e2bb3
fix(pi): invoke Bash helpers correctly on native Windows (#3843)
cr101 Sep 6, 2026
71ec401
test(bin): pin teardown outcomes for squash-merged rebased branches (…
MortenGad Sep 6, 2026
29015a2
fix(bin): keep supervision armed for registered custom checks (#3860)
gyute Sep 6, 2026
64d3905
fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)
kunchenguid Sep 7, 2026
cf7e2fa
fix(bearings): repair board listening and decision reconciliation (#3…
kunchenguid Sep 7, 2026
1533f47
fix(bin): allow pooled spawns without a git origin (#3885)
kunchenguid Sep 7, 2026
5592cb6
feat(tests): run live harness guards by default when available (#3889)
kunchenguid Sep 7, 2026
6d396da
fix(bin): refuse test runs in the primary checkout when a task marker…
3264studios Sep 7, 2026
d4eb228
fix(bin): bound stale alarms for backlog captain holds (#3842)
mremond Sep 7, 2026
0b9f518
feat(pi): resolve extension-registered providers in the supervision b…
0x7067 Sep 7, 2026
ffd2c89
fix(bin): gate secondmate wake-loop stall alerts on real queue no-pro…
cisrd Sep 7, 2026
36fd955
test(calm): harden the export-DOM render step and record Pi 0.85.1 ev…
cisrd Sep 7, 2026
3af74fe
fix(bin): make every counted wake queue row presentable or retired (#…
cisrd Sep 7, 2026
98b37d4
fix(bin): use system stat for Darwin BSD formats (#3305)
0x7067 Sep 7, 2026
72bfdd0
fix(spawn): carry attribution-off policy in every claude launch (#3945)
NewAiCoder-bot Sep 7, 2026
891dc51
fix(bin): derive watcher beacon staleness grace from poll cadence (#3…
NewAiCoder-bot Sep 8, 2026
b84e0e3
fix(procevent): reap orphaned runners and prevent launch storms (#3904)
kunchenguid Sep 8, 2026
38ea36a
Merge upstream round 4 through b84e0e362fac
Sep 12, 2026
dc76f2e
Merge fork monitoring fix into upstream sync round 4
Sep 12, 2026
335a112
test: preserve monitoring guard gates in portable coverage
Sep 12, 2026
db91abe
test: balance CI lanes and split watcher triage within existing bounds
Sep 14, 2026
5722375
Merge fork planning and design relaunch update into upstream sync rou…
Sep 14, 2026
0b838c8
test: isolate backend lock fixtures and align split-suite ordering
Sep 14, 2026
d5fdeb3
test: refresh serial shard hints from measured CI durations
Sep 14, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 12 additions & 4 deletions .agents/skills/bearings/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,11 +88,13 @@ Board answers are acted on later under the normal authority rules; this skill's
## Lavish board mode

`/bearings lavish` adds one deliverable beside the unchanged chat digest: the interactive fleet board, a myfirstmate-styled Lavish page where the captain answers Captain's Call items directly instead of replying in chat.
`bin/fm-bearings-board.sh` owns every board mechanic - the stable board path, fm-bearings-board.v1 payload validation, template injection, Lavish session establishment, the any-origin answer binding, and arm-if-absent registration - so the per-invocation work is composing the payload and running its `build`.
`bin/fm-bearings-board.sh` owns every board mechanic - the stable board path, fm-bearings-board.v1 payload validation, template injection, live Lavish session verification and ended-session reopening, the any-origin answer binding, and listener registration - so the per-invocation work is composing the payload and running its `build`.

Compose the payload from the same snapshot with the same ranking judgment as the chat digest, plus these board rules:

- A Captain's Call decision key is the captain-held TASK ID from `decisions_open` (legacy `<origin>-decision-<key>` rows are already task ids); a merge card's key is `merge.<task-id>`; the Charted Next dispatch picker's key is `dispatch.charted`.
- Before carding a hold, check that its SUBJECT has not already landed, and omit it when it has. `build` drops a card whose task or PR appears in the payload's own landed rows, and one whose task is no longer an open captain call. When a hold waits on one specific PR, put that PR in the card's `pr_url`. When it concerns a published version, put the artifact and numeric three-part version in the card's structured `subject`; landed rows for releases carry the same identity, and a matching or newer version drops the card. Identity matching is structured only, so verify any subject without one of these identities against current reality before carding it.
- Never author a `reconcile` option on any card. `build` gives every decision card the standard reconcile choice itself, and the payload validator reserves that value across all card types; recommendations must name an authored option.
- Compose exactly one decision card per captain-held task id. When one task carries multiple questions, consolidate all of them and their options into that card; never emit duplicate cards with the same task-id key.
- Decision cards carry agent-authored copy: a short noun-phrase title, one-line `about` and `decide` context rows, and option labels with hints, with the recommended option marked.
- Card `type` (decision, merge, credential) is your composing judgment from the row's content; no backlog field types a card for you.
Expand All @@ -102,14 +104,20 @@ Compose the payload from the same snapshot with the same ranking judgment as the
- Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words.

Run `build` once after composing the payload.
Its serve-first sequence publishes the board, establishes or resumes its Lavish session with `lavish-axi`, and only then binds and arms the polling source; use the session URL it prints in the chat digest.
Never bind or arm the board before that session exists.
Never run `lavish-axi poll` for the board yourself: the armed source's supervised runner owns the blocking poll, and the watcher's ordinary reconcile restarts it, so no conversational turn ever blocks on the board.
Its serve-first sequence publishes the board, establishes and verifies its Lavish session with `lavish-axi`, reopens an ended session when necessary, and only then binds the answer source and proves a live polling listener; use the session URL it prints in the chat digest.
Never bind or arm the board before its session is listed open.
Never run `lavish-axi poll` for the board yourself: the armed source's supervised runner owns the blocking poll, and both the build and the watcher's ordinary reconcile repair a missing listener, so no conversational turn ever blocks on the board.

### Handling a board wake

A board answer arrives as an ordinary `procevent lavish <source-id> <sequence>` check wake. Identify it by comparing the wake source id with `bin/fm-procevent-lavish.sh source-id "$(bin/fm-bearings-board.sh path)"`, regardless of which answer kinds the result contains; then load `process-event-sources` and follow its contract for the result read, adapter classification, and the handled acknowledgement.
Decision answers need no routing from you: the runner feeds the board's binding into `bin/fm-captain-hold.sh`'s one keyed-answer intake, which closes or releases each answered captain-held task at answer time; reconcile any `skipped:` key yourself with a direct `answer`, and when the captain's answer is "later", record it as a deferral with `bin/fm-captain-hold.sh hold <id> --reason "<reason>" --until <date>` instead of a closure.
A current structured Reconcile selection closes nothing: the versioned board context carries its exact selected option separately from any typed note, and the adapter routes that selection only into a durable re-check request while preserving the note as provenance.
The rollout-compatible old context still feeds ordinary non-reconcile answers, but its bare or separator-annotated reconcile values and every structurally uncertain choice feed neither intake and remain announced for deliberate handling.
Verify the call's latest state, then retire the request through `bin/fm-captain-hold.sh reconcile close <id> --evidence-file <path>` when it turns out to be moot, or `reconcile note <id> --note-file <path>` when it is genuinely still open.
Both outcomes refuse without that pending board-created request, and `bin/fm-captain-hold.sh reconcile list` names every request still outstanding.
A remote-secondmate card whose task is absent from the main backlog remains on the board unchanged, but its reconcile request is refused in the main home until the separately tracked owner-aware routing follow-up can query and mutate the authoritative secondmate home; handle the announced capture without claiming that a request or reconciliation succeeded.
`captain-hold-lifecycle` owns why a reconcile may never be recorded as the captain's answer.
Route the non-decision keys yourself:

- `merge.<task-id>` is the captain's explicit merge order; follow the merge ruling below.
Expand Down
22 changes: 13 additions & 9 deletions .agents/skills/bearings/assets/board-template.html
Original file line number Diff line number Diff line change
Expand Up @@ -548,22 +548,26 @@
var fd = new FormData(form);
var value = fd.get("answer");
var note = (fd.get("note") || "").trim();
/* picked option, optionally annotated; a bare note is itself the answer */
var answer = value ? (note ? value + " - " + note : value) : note;
if (!answer) return;
if (utf8ByteLength(answer) > 512) {
var displayAnswer = value ? (note ? value + " - " + note : value) : note;
if (!displayAnswer) return;
if (utf8ByteLength(displayAnswer) > 512) {
answerLimit.textContent = "Answer is too long to queue (512 bytes maximum).";
answerLimit.classList.add("is-visible");
return;
}
if (window.lavish && window.lavish.queuePrompt) {
/* close carries the composer-declared close mode: "release" frees a
captain-gated work item instead of completing a question task */
var ctxData = { question: item.key, answer: answer };
/* The versioned context keeps the selected option separate from its
note, while close carries the composer-declared completion mode. */
var ctxData = {
schema: "fm-bearings-answer.v1",
question: item.key,
selection: value || "",
note: note
};
if (item.close) ctxData.close = item.close;
window.lavish.queuePrompt(
"Captain's Call answer - " + item.title + ": " + answer,
{ tag: "choice", text: item.title + " -> " + answer, element: form,
"Captain's Call answer - " + item.title + ": " + displayAnswer,
{ tag: "choice", text: item.title + " -> " + displayAnswer, element: form,
data: ctxData }
);
}
Expand Down
16 changes: 13 additions & 3 deletions .agents/skills/captain-hold-lifecycle/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,14 +24,24 @@ After inventorying the whole report and review surface, run `bin/fm-captain-hold
A completed investigation, a completed ADR design, and an ended visual review use this same owner and completion command; a design profile or visual tool, including Lavish, never owns a parallel completion policy.
Run the command in the originating work's authoritative `FM_HOME`; secondmate-owned work registers in that secondmate home's backlog, and a question already held anywhere is never re-registered as a second row.
Do not close a captain-held task merely because the originating investigation completed, its report was archived, its visual review ended, or its task was torn down.
Holding the work item the question gates is safe for exactly that reason: cleanup keeps such a row open with the finished work's deliverable recorded and returns it to the queue, so it still reads as the captain's own call and only `answer` closes it.
Holding the work item the question gates is safe for exactly that reason: cleanup keeps such a row open with the finished work's deliverable recorded and returns it to the queue, so it still reads as the captain's own call.
Only `answer` with the captain's words or an evidence-backed `reconcile close` may close it.

Never close anything the captain owns without recording what he actually said: `bin/fm-captain-hold.sh answer` writes his exact words into the task and closes it in the same act, with `--release` when the answer frees a captain-gated work item to proceed instead of completing a question.
When the answer changes what a task must build, follow `AGENTS.md` section 7's Validate contract to preserve the captain's words in the brief and steer the worker.
When the captain says "later", that is an answer too: re-hold with `bin/fm-captain-hold.sh hold <id> --reason "<reason>" --until <date>` so the item leaves the live Captain's Call and resurfaces on its date, instead of leaving a live-looking card or fabricating a closure.
"A keyed answer closes its matching captain-held task" is one capability with one owner, `bin/fm-captain-hold.sh answers`, and every channel that carries a captain answer feeds it the same task id and answer; a channel never maps keys to tasks, records a decision, or closes anything itself.
Chat already feeds it through `bin/fm-send.sh --resolve-key`, and a captured-answer source feeds it once bound with `bin/fm-captain-hold.sh bind <source-id>`; bind before arming the source, and key each structured question by the held task's id.
An unbound source and a key that names no captain-held task both simply feed nothing: the answer is still captured and firstmate is still woken, and closing falls back to the direct command above.
One answer value is reserved and closes nothing: `reconcile` means "go re-check reality", never "the captain answered", so the shared intake refuses it from every channel and creates nothing.
A bound captured source uses a separate seam: its adapter omits reconcile from keyed answers and emits the selected task id through `reconciles`, the generic runner feeds that into `reconcile-requests`, and the intake verifies the source binding and the local captain-held task before filing the durable board request.
A remote-secondmate card whose task is absent from the main backlog therefore remains announced but cannot create a main-home request; owner-aware request and mutation routing to the authoritative secondmate home is a separate follow-up.
That board-created request is yours to work off in the turn that receives it: `bin/fm-captain-hold.sh reconcile close <id> --evidence-file <path>` records the EVIDENCE and closes a moot call, while `reconcile note <id> --note-file <path>` annotates a genuinely active call and leaves it held.
Both outcomes refuse unless that task still has the pending request created by the captain's board selection, so neither is a standalone way to mutate a captain call.
A normal captain answer also retires any pending request because the call is settled, including close, release, and idempotent replay paths.
A retirement failure makes the command fail without reversing the already-durable answer, close, or note, and `reconcile list` keeps the surviving request visible for retry.
`reconcile list` names every request still outstanding.
Never use `answer` for an evidence-only moot call: `answer` records what the captain said, while `reconcile close` records verified evidence.
A captain-held task closed outside this owner leaves no durable answer, so the completion gate keeps failing until `answer` records the decision the captain actually gave.
Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create held tasks.
Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose.
Expand All @@ -49,8 +59,8 @@ The absence of a routed work item is not a divergence and the guard never requir
3. Hold that task - or create one captain-held task for the review's open questions - with a concise reason carrying the question and options.
4. Run `complete` with the full captain-held inventory for that review pass.
5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat.
6. Close each call only through `answer` (or a channel that feeds `answers`), through `--until` when the captain defers it, or confirm a channel already closed it.
7. Confirm Bearings reflects the outcome: answered calls leave Captain's Call, released work resumes, and deferred calls sit in Charted Next with their date.
6. Close each call only through `answer` (or a channel that feeds `answers`), close a board-requested moot call through evidence-backed `reconcile close`, record a still-active reconciliation through `reconcile note`, use `--until` when the captain defers it, or confirm a channel already closed it.
7. Confirm Bearings reflects the outcome: answered or reconciled-moot calls leave Captain's Call, released work resumes, active reconciliations remain held, and deferred calls sit in Charted Next with their date.

`bin/fm-captain-hold.sh --help` owns command syntax, close modes, legacy-identity compatibility, completion attestation, retry behavior, and close ordering.
`docs/captain-hold-lifecycle.md` records the mechanism and regression evidence without restating this policy.
5 changes: 3 additions & 2 deletions .agents/skills/firstmate-coding-guidelines/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,9 +102,10 @@ Every such check needs two tests, because they fail for different reasons:
- A portable regression in `tests/` that pins the logic with real processes and no harness, so CI enforces the classifier everywhere it runs tmux.
Drive the signals apart deliberately and assert the verdict survives losing one; assert the divergence itself so the case cannot go quietly vacuous.
Confirm which signal a given construction actually blinds on each supported platform rather than assuming, because the same trick can break different sources on macOS and Linux.
- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`), env-gated and self-skipping, that exercises every INSTALLED harness for real and fails naming the harness and version.
- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`) that exercises every INSTALLED harness for real and fails naming the harness and version.
Report an absent harness explicitly rather than passing silently over it, and refuse a pass that checked nothing.
This guard is opt-in and on-demand because standard CI has neither harness binaries nor credentials; run it after every harness upgrade and before trusting refreshed per-harness evidence.
Open it with `fm_live_gate` from `tests/lib.sh`, which is the single owner of that decision: a guard that spends no model tokens runs by default wherever its tools are installed, a guard that submits prompts stays opt-in, and its own variable or `FM_LIVE` forces it on (an absent tool then fails rather than skips) or off.
The portable serial CI lane has no credentials and installs the public Pi package, so token-free guards exercise the available Pi surfaces while unavailable tools capability-skip; run a prompt-submitting guard after every harness upgrade and before trusting refreshed per-harness evidence.

Record the dated per-harness result in `docs/verification/runtime-backends.md`, and point at the live guard as the command that refreshes it, rather than leaving a version-scoped observation to rot into a false claim.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ The primary integration was verified on 2026-07-08 with OpenCode 1.17.6.
Throwing from `session.idle` does not block `opencode run`, so the primary adapter treats the event as passive and uses `client.session.promptAsync` to force one follow-up turn when `../../../bin/fm-turnend-guard.sh` returns 2.
The follow-up was verified in the interactive TUI.
`opencode run` can exit before displaying a queued follow-up, so the adapter steps aside in headless mode.
On native Windows, the operational-input adapter runs its Bash helper through `bash`; macOS and Linux invoke it directly.

The companion `.opencode/plugins/fm-primary-watch-arm.js` owns normal TUI watcher supervision, wakes it with `client.session.promptAsync`, and coordinates with the guard before a blind-turn follow-up.
The PreToolUse-equivalent watcher-arm seatbelt blocks by throwing from `tool.execute.before`.
1 change: 1 addition & 0 deletions .agents/skills/harness-adapters/references/harness/pi.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,7 @@ Pi sets `PI_CODING_AGENT=true` for its children as its harness-detection marker.
The primary turn-end behavior was verified on 2026-07-09 with Pi 0.80.5.
`.pi/extensions/fm-primary-turnend-guard.ts` listens for logical-run `agent_settled`, not per-tool-loop `turn_end`, and uses `pi.sendUserMessage(..., { deliverAs: "followUp" })` to force one guarded follow-up when `../../../bin/fm-turnend-guard.sh` returns 2.
Without `deliverAs: "followUp"`, Pi rejects the send while the agent is still processing.
On native Windows, the extension runs its session-start, both PreToolUse, turn-end, and operational-input Bash helpers through `bash`; macOS and Linux invoke those helpers directly.

The primary watcher protocol also requires `.pi/extensions/fm-primary-pi-watch.ts`.
The Pi engine auto-discovers both tracked project-local extensions once the project is trusted.
Expand Down
Loading
Loading