From 69d660ad6167271daf09e8c5521581c03cb9a4f8 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Wed, 16 Sep 2026 22:32:26 -0700 Subject: [PATCH 01/37] feat(bin): add opt-in typed dispatch resolution (#4692) * feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior --- .../references/common/dispatch.md | 1 + .agents/skills/quota-array-dispatch/SKILL.md | 1 + AGENTS.md | 3 +- bin/fm-bootstrap.sh | 53 +- bin/fm-control-lib.sh | 11 +- bin/fm-dispatch-resolve.sh | 404 +++++++++++ bin/fm-env-lib.sh | 31 + bin/fm-quota-axi-lib.sh | 48 +- bin/fm-quota-choose.sh | 31 +- bin/fm-test-run.sh | 12 + bin/fm-x-lib.sh | 22 +- docs/configuration.md | 59 +- docs/documentation-audiences.json | 4 + docs/examples/crew-dispatch.json | 2 +- docs/verification/dispatch-resolve.md | 73 ++ tests/fm-bootstrap.test.sh | 96 ++- tests/fm-dispatch-resolve.test.sh | 638 ++++++++++++++++++ tests/fm-gotmp.test.sh | 18 +- tests/fm-quota-choose.test.sh | 15 +- 19 files changed, 1435 insertions(+), 87 deletions(-) create mode 100755 bin/fm-dispatch-resolve.sh create mode 100644 bin/fm-env-lib.sh create mode 100644 docs/verification/dispatch-resolve.md create mode 100755 tests/fm-dispatch-resolve.test.sh diff --git a/.agents/skills/harness-adapters/references/common/dispatch.md b/.agents/skills/harness-adapters/references/common/dispatch.md index 96db331b557..eda57198865 100644 --- a/.agents/skills/harness-adapters/references/common/dispatch.md +++ b/.agents/skills/harness-adapters/references/common/dispatch.md @@ -7,6 +7,7 @@ Load this with the selected tool reference for dispatch, start, or adapter verif Use the router's detection and safety sections for static crew and secondmate harness resolution and all explicit overrides. `config/crew-dispatch.json` can override that static default for one crewmate or scout with concrete harness, model, and effort axes. For a profile array, load `quota-array-dispatch` after establishing harness and provider facts here. +When the opt-in `bin/fm-dispatch-resolve.sh` is on, its `clear` answer already names the concrete axes; `docs/configuration.md` "Typed dispatch resolution" owns that contract. `../secondmate-provisioning/SKILL.md` owns inherited local material. Its harness consequence is that a secondmate's workers receive literal `config/crew-harness` and `config/crew-dispatch.json`, while the primary-only `config/secondmate-harness` is never inherited because secondmates do not spawn secondmates. diff --git a/.agents/skills/quota-array-dispatch/SKILL.md b/.agents/skills/quota-array-dispatch/SKILL.md index c2b9f05ece5..4b988f1baab 100644 --- a/.agents/skills/quota-array-dispatch/SKILL.md +++ b/.agents/skills/quota-array-dispatch/SKILL.md @@ -33,6 +33,7 @@ Authoritative multi-provider routing - including provider discovery from the har Use it only when the brief already fixed the candidate order and every candidate's provider is the harness's primary family. It does not replace the reasoning-class, runway-feasibility, or authentication gates above. Firstmate can optionally arm `bin/fm-procevent-quota.sh` for a recurring mid-task check that wakes when the tracked provider drops below its configured threshold or its runway becomes `exhausted_now`. +The opt-in `bin/fm-dispatch-resolve.sh` (`docs/configuration.md` "Typed dispatch resolution") applies the same eligibility gates and `spendPriority` argmax in code after a typed rule match; it never removes this skill's authority, and its `ambiguous`, `escalate`, and `error` outcomes return here. ## Read the default TOON diff --git a/AGENTS.md b/AGENTS.md index 12bd53b73af..65a3197944d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -68,7 +68,7 @@ README.md public overview and development notes .claude/mods/ Claude Code mods (function-hooks plugins), committed; Calm's module may load through CLAUDE_CODE_ENABLE_FUNCTION_HOOKS or tengu_plugin_hooks_modules, but activates only when CLAUDE_CODE_ENABLE_FUNCTION_HOOKS is exactly "1" and is otherwise a complete no-op (docs/calm.md) skills/ standalone public installer-facing skills, committed; not loaded by firstmate bin/ helper scripts, committed; read each script's header before first use -.env optional Relay pairing token (presence-gates section 14) and mail-plane credentials (schema: docs/configuration.md "Mail plane"); LOCAL, gitignored +.env optional Relay pairing token (presence-gates section 14), mail-plane credentials (schema: docs/configuration.md "Mail plane"), and typed dispatch resolution key TYPESAFE_API_KEY (presence-gates bin/fm-dispatch-resolve.sh; docs/configuration.md "Typed dispatch resolution"); LOCAL, gitignored config/crew-harness crewmate harness override; LOCAL, gitignored; absent or "default" = same as firstmate. Inherited as the literal file: a concrete primary adapter value also controls a secondmate home's own crewmates (section 4) config/claude-permission-mode optional one-token permission posture for every Claude worker launch: absent or "bypass" keeps --dangerously-skip-permissions, "auto" launches with --permission-mode auto; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Claude permission mode" config/crew-dispatch.json optional crewmate dispatch profiles; LOCAL, gitignored; firstmate-maintained but human-editable natural-language rules that choose a per-task harness/model/effort profile (section 4). Inherited by secondmate homes @@ -227,6 +227,7 @@ When every candidate is tight, preserve the captain's strongest-reasoning class Break genuine evidence ties without array-order or harness bias. `quota-axi` owns how model or product windows relate to bounding account windows and remains data-only. Load `quota-array-dispatch` before choosing among a matched profile array; that skill is the single owner of the TOON-first spendPriority selection procedure. +Run `bin/fm-dispatch-resolve.sh` directly on the written brief in the same turn, with no preflight, and on `clear` pass its `profile:` line to `fm-spawn` unless you state a reason to override; `ambiguous`, `escalate`, `error`, and off all mean the intake above, unchanged (contract: `docs/configuration.md` "Typed dispatch resolution"). The generic effort fallback and its precedence are owned by `harness-adapters`: explicit captain and standing configured effort win; otherwise use low for well-understood explicit work, xhigh for ambiguous investigation or design, intermediate levels proportionally, and never max without explicit captain preference. Do not add model-specific versions of that policy. diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 1c550c71f10..31792fa37ba 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -156,6 +156,10 @@ # nothing; bin/fm-brief.sh uses it to gate scout Lavish hosting. set -u +TYPESAFE_API_KEY_PRIVATE=${TYPESAFE_API_KEY:-} +export -n TYPESAFE_API_KEY_PRIVATE 2>/dev/null || true +unset TYPESAFE_API_KEY + SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" @@ -169,6 +173,10 @@ DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" . "$SCRIPT_DIR/fm-backlog-transition-lib.sh" # shellcheck source=bin/fm-quota-axi-lib.sh disable=SC1091 . "$SCRIPT_DIR/fm-quota-axi-lib.sh" +# shellcheck source=bin/fm-control-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-control-lib.sh" +# shellcheck source=bin/fm-env-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-env-lib.sh" # shellcheck source=bin/fm-tangle-lib.sh disable=SC1091 . "$SCRIPT_DIR/fm-tangle-lib.sh" # shellcheck source=bin/fm-ff-lib.sh disable=SC1091 @@ -1102,7 +1110,7 @@ EOF } crew_dispatch_validate() { - local file err + local file err verified_harnesses typed_key typed_active=false file="$CONFIG/crew-dispatch.json" [ -f "$file" ] || return 0 if ! command -v jq >/dev/null 2>&1; then @@ -1113,8 +1121,17 @@ crew_dispatch_validate() { echo "CREW_DISPATCH: invalid config/crew-dispatch.json - malformed JSON" return 0 fi - err=$(jq -r ' - def verified($h): ["claude","codex","opencode","pi","pi-signed","grok","kimi","cursor","agy","muse","rovo","omp"] | index($h); + typed_key=$TYPESAFE_API_KEY_PRIVATE + [ -n "$typed_key" ] || typed_key=$(fmx_env_get TYPESAFE_API_KEY "$FM_HOME/.env") + [ -z "$typed_key" ] || typed_active=true + if $typed_active; then + verified_harnesses=$(fm_control_harnesses | jq -Rsc 'split("\n") | map(select(length > 0))') + else + verified_harnesses='["claude","codex","opencode","pi","pi-signed","grok","kimi","cursor","agy","muse","rovo","omp"]' + fi + err=$(jq -r --argjson typed "$typed_active" --argjson verified_harnesses "$verified_harnesses" --arg provider_re "$FM_QUOTA_PROVIDER_ID_RE" ' + def verified($h): $verified_harnesses | index($h); + def provider_id($p): ($p | type) == "string" and ($p | test($provider_re)); def effort_ok($h; $m; $e): if $e == null then true elif ($e | type) != "string" then false @@ -1139,7 +1156,21 @@ crew_dispatch_validate() { + (if has("default") then [profiles(.default)[]?] else [] end)); def malformed_optional_fields($items): ($items | any(has("model") and (((.model | type) != "string") or (.model | length) == 0))) - or ($items | any(has("effort") and (((.effort | type) != "string") or (.effort | length) == 0))); + or ($items | any(has("effort") and (((.effort | type) != "string") or (.effort | length) == 0))) + or ($typed and ($items | any(has("provider") and (provider_id(.provider) | not)))); + # A quota floor, on a rule or a profile: bin/fm-dispatch-resolve.sh applies + # it in code against one quota-axi row, so scope and min_percent must be + # concrete; a rule floor also names the provider whose row it reads. + def floor_bad($f; $need_provider): + ($f | type) != "object" + or (($f.scope | type) != "string") or (($f.scope | length) == 0) + or (($f.min_percent | type) != "number") or ($f.min_percent < 0) or ($f.min_percent > 100) + or (if $need_provider + then (provider_id($f.provider) | not) + else ($f | has("provider")) + end); + def malformed_profile_floors($items): + ($items | any(has("floor") and floor_bad(.floor; false))); def bad_efforts: configured_profiles | map({h: .harness, m: .model, e: .effort}) @@ -1156,7 +1187,13 @@ crew_dispatch_validate() { elif [(.rules // [])[]? | select((.use? | type) == "array" and (.use | length) == 0)] | length > 0 then "each rule needs at least one use profile" elif [(.rules // [])[]? | profiles(.use?)[]? | select(type != "object")] | length > 0 then "each use profile must be an object" elif [(.rules // [])[]? | profiles(.use?)[]? | select((.harness? | type) != "string" or (.harness | length) == 0)] | length > 0 then "each use profile needs harness" - elif malformed_optional_fields([(.rules // [])[]? | profiles(.use?)[]?]) then "use profile model and effort must be non-empty strings when present" + elif malformed_optional_fields([(.rules // [])[]? | profiles(.use?)[]?]) then + if $typed then "use profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\\z when present" + else "use profile model and effort must be non-empty strings when present" + end + elif $typed and malformed_profile_floors([(.rules // [])[]? | profiles(.use?)[]?]) then "use profile floor needs scope and min_percent 0..100" + elif $typed and ([(.rules // [])[]? | select(has("approval") and .approval != "captain")] | length > 0) then "approval must be \"captain\" when present" + elif $typed and ([(.rules // [])[]? | select(has("floor") and floor_bad(.floor; true))] | length > 0) then "rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\\z" elif [(.rules // [])[]? | select(has("select") and ((.select? | type) != "string" or (.select | length) == 0))] | length > 0 then "select must be a non-empty string" elif [(.rules // [])[]? | .select? // empty | select(. != "quota-balanced")] | length > 0 then "unknown select: " + ([ (.rules // [])[]? | .select? // empty | select(. != "quota-balanced") ] | unique | join(", ")) @@ -1164,7 +1201,11 @@ crew_dispatch_validate() { elif has("default") and ((.default | type) == "array" and (.default | length) == 0) then "default needs at least one profile" elif has("default") and ([profiles(.default)[]? | select(type != "object")] | length) > 0 then "each default profile must be an object" elif has("default") and ([profiles(.default)[]? | select((.harness? | type) != "string" or (.harness | length) == 0)] | length) > 0 then "each default profile needs harness" - elif has("default") and malformed_optional_fields([profiles(.default)[]?]) then "default profile model and effort must be non-empty strings when present" + elif has("default") and malformed_optional_fields([profiles(.default)[]?]) then + if $typed then "default profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\\z when present" + else "default profile model and effort must be non-empty strings when present" + end + elif $typed and has("default") and malformed_profile_floors([profiles(.default)[]?]) then "default profile floor needs scope and min_percent 0..100" else (configured_profiles | map(.harness) diff --git a/bin/fm-control-lib.sh b/bin/fm-control-lib.sh index 516a00b4364..7bb4d580ec6 100644 --- a/bin/fm-control-lib.sh +++ b/bin/fm-control-lib.sh @@ -61,10 +61,15 @@ fm_control_verb_allowed() { # # The harnesses whose control mechanics are verified. Mirrors AGENTS.md # section 4's verified-adapter list; an unverified adapter is refused rather # than guessed at, exactly as a spawn on it would be. +fm_control_harnesses() { + printf '%s\n' claude codex opencode pi pi-signed grok kimi cursor gemini muse rovo omp agy +} + fm_control_harness_supported() { # - case "${1-}" in - claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy) return 0 ;; - esac + local harness + while read -r harness; do + [ "$harness" = "${1-}" ] && return 0 + done < <(fm_control_harnesses) return 1 } diff --git a/bin/fm-dispatch-resolve.sh b/bin/fm-dispatch-resolve.sh new file mode 100755 index 00000000000..12f67dbcb00 --- /dev/null +++ b/bin/fm-dispatch-resolve.sh @@ -0,0 +1,404 @@ +#!/usr/bin/env bash +# fm-dispatch-resolve.sh - resolve one concrete crewmate or scout dispatch +# profile from a task brief with typesafe.ai's System One model (Jev), opt-in. +# +# Usage: +# fm-dispatch-resolve.sh [--project ] +# +# Opt-in gate: TYPESAFE_API_KEY non-empty in this process environment, else a +# TYPESAFE_API_KEY= line in $FM_HOME/.env read with fmx_env_get, the same +# accessor as FMX_PAIRING_TOKEN (bin/fm-env-lib.sh). The environment wins. +# Absent in both: one "dispatch-resolve: off" line on stderr, nothing on +# stdout, exit 0, no network call, so firstmate dispatches exactly as today. +# The key lives in one shell variable and reaches curl as a header read from +# a file descriptor, never on argv; nothing logs or writes it. +# +# What it does when on with at least one rule: one POST to +# https://api.typesafe.ai/v1/systemone with the project name and the whole brief as +# state and ONE Choice question whose +# options are every rule's `when` from config/crew-dispatch.json plus one +# fixed generic none option. Jev returns the matched rule, a probability per +# option, and a confidence. Everything after that is jq: the confidence +# floor, the rule's declared `approval` and `floor`, each profile's declared +# `provider` and `floor`, the quota rows from ONE quota-axi --json snapshot, +# and the spendPriority argmax over the eligible candidates. The model never +# sees quota, catalogs, approvals, `why`, or `use`. With no rules, it returns +# a non-clear result so firstmate keeps using the existing intake. +# docs/configuration.md "Crew dispatch profiles" owns the declared fields and +# "Typed dispatch resolution" owns this tool's operator contract. +# +# Output (stdout, TOON-style block): +# dispatch-resolve: +# status: clear | ambiguous | escalate | error +# model/latency_ms/tokens, rule (when excerpt) and confidence, probabilities +# reason: +# candidate: : provider=.. scope=.. remaining=..% spendPriority=.. runway=.. -> eligible | eligible, unranked: | not eligible: +# profile: --harness [--model ] [--effort ] (status clear only) +# clear -> pass the profile line to fm-spawn.sh unless you state a reason to override +# ambiguous -> confidence below the floor; decide as today from the probabilities +# escalate -> the rule requires captain approval, no candidate is rankable, or a genuine tie +# error -> API, network, response, or quota-axi failure; decide as today +# Every outcome exits 0 so an intake is never blocked by this tool. +# Exit 2 only for a usage or configuration error (unreadable brief, an +# existing unreadable rules file, malformed rules, or missing jq), which is +# actionable, never selected around. +# +# Environment: +# TYPESAFE_API_KEY is the only resolver-specific environment setting. +# +# Authority: this tool never replaces firstmate's judgment, quota-array-dispatch, +# the captain-approval gate, or fm-spawn.sh validation; it publishes one +# inspectable answer plus every candidate's evidence, in code. +set -u + +TYPESAFE_API_KEY_PRIVATE=${TYPESAFE_API_KEY:-} +export -n TYPESAFE_API_KEY_PRIVATE 2>/dev/null || true +unset TYPESAFE_API_KEY + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-$FM_ROOT}" +CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" + +# shellcheck source=bin/fm-quota-axi-lib.sh +. "$SCRIPT_DIR/fm-quota-axi-lib.sh" +# shellcheck source=bin/fm-control-lib.sh +. "$SCRIPT_DIR/fm-control-lib.sh" +# shellcheck source=bin/fm-env-lib.sh +. "$SCRIPT_DIR/fm-env-lib.sh" +# shellcheck source=bin/fm-timing-lib.sh +. "$SCRIPT_DIR/fm-timing-lib.sh" + +CONFIDENCE_FLOOR=0.6 +TS_MODEL=jev-latest +TS_BASE=https://api.typesafe.ai +TS_TIMEOUT=5 +DEFAULT_WHEN="No listed rule applies to this task." + +die() { printf 'error: %s\n' "$1" >&2; exit 2; } +no_rules() { + printf 'dispatch-resolve:\n status: escalate\n reason: no rules to match\n' + exit 0 +} +usage() { + awk ' + NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit } + ' "$0" +} + +BRIEF='' PROJECT='' RULES_PATH="$CONFIG/crew-dispatch.json" RULES='' +while [ $# -gt 0 ]; do + case "$1" in + --project) [ $# -ge 2 ] || die "--project needs a value"; PROJECT=$2; shift 2 ;; + -h|--help) usage; exit 0 ;; + -*) die "unknown flag $1" ;; + *) [ -z "$BRIEF" ] || die "one brief file only"; BRIEF=$1; shift ;; + esac +done + +# ---- opt-in gate --------------------------------------------------------------- +if [ -z "$TYPESAFE_API_KEY_PRIVATE" ]; then + TYPESAFE_API_KEY_PRIVATE=$(fmx_env_get TYPESAFE_API_KEY "$FM_HOME/.env") +fi +if [ -z "$TYPESAFE_API_KEY_PRIVATE" ]; then + echo "dispatch-resolve: off (TYPESAFE_API_KEY absent from the environment and $FM_HOME/.env)" >&2 + exit 0 +fi + +# ---- inputs -------------------------------------------------------------------- +[ -n "$BRIEF" ] || die "brief file required (see --help)" +[ -r "$BRIEF" ] || die "brief file not readable: $BRIEF" +[ -e "$RULES_PATH" ] || [ -L "$RULES_PATH" ] || no_rules +[ -r "$RULES_PATH" ] || die "rules file not readable: $RULES_PATH" +command -v jq >/dev/null 2>&1 || die "jq required" +RULES=$(mktemp) || die "mktemp failed" +trap 'rm -f "$RULES"' EXIT +cp "$RULES_PATH" "$RULES" || die "could not snapshot rules file: $RULES_PATH" +chmod 400 "$RULES" || die "could not protect rules snapshot" +VERIFIED_HARNESSES=$(fm_control_harnesses | jq -Rsc 'split("\n") | map(select(length > 0))') + +# The fields this tool consumes must be well formed; bootstrap owns the wider +# schema diagnostic, but an intake never selects around a malformed file. +rules_err=$(jq -r --argjson verified_harnesses "$VERIFIED_HARNESSES" --arg provider_re "$FM_QUOTA_PROVIDER_ID_RE" ' + def verified($h): $verified_harnesses | index($h); + def provider_id($p): ($p | type) == "string" and ($p | test($provider_re)); + def effort_ok($h; $m; $e): + if $e == null then true + elif ($e | type) != "string" then false + elif $e == "ultra" then (($h == "pi" or $h == "pi-signed") and (($m | type) == "string") and ($m | startswith("codex-native/")) and ($m | length) > 13) + elif $h == "claude" then (["low","medium","high","xhigh","max"] | index($e)) != null + elif $h == "codex" then ((["low","medium","high","xhigh"] | index($e)) != null or ($e == "max" and $m == "gpt-5.6-luna")) + elif $h == "grok" or $h == "agy" then (["low","medium","high"] | index($e)) != null + elif $h == "pi" or $h == "pi-signed" or $h == "omp" or $h == "muse" then (["low","medium","high","xhigh","max"] | index($e)) != null + elif $h == "rovo" then (["low","medium","high","max"] | index($e)) != null + elif $h == "opencode" or $h == "kimi" or $h == "cursor" then false + else true end; + def profiles($v): if ($v | type) == "array" then $v elif ($v | type) == "object" then [$v] else [] end; + def floor_bad($f; $need_provider): + ($f | type) != "object" + or (($f.scope | type) != "string") or (($f.scope | length) == 0) + or (($f.min_percent | type) != "number") or ($f.min_percent < 0) or ($f.min_percent > 100) + or (if $need_provider + then (provider_id($f.provider) | not) + else ($f | has("provider")) + end); + def profile_bad($p): + ($p | type) != "object" + or (($p.harness | type) != "string") or (($p.harness | length) == 0) + or ($p | has("model") and ((.model | type) != "string" or (.model | length) == 0)) + or ($p | has("effort") and ((.effort | type) != "string" or (.effort | length) == 0)) + or ($p | has("provider") and (provider_id(.provider) | not)) + or ($p | has("floor") and floor_bad(.floor; false)); + def duplicate_profiles($items): + ($items | map([.harness, (.model // null), (.effort // null)] | @json)) as $keys + | ($keys | length) != ($keys | unique | length); + if type != "object" then "top-level value must be an object" + elif has("rules") and (.rules | type) != "array" then "rules must be an array" + elif any((.rules // [])[]; type != "object") then "each rule must be an object" + elif any((.rules // [])[]; (.when | type) != "string" or (.when | length) == 0) then "each rule needs non-empty when" + elif any((.rules // [])[]; (profiles(.use) | length) == 0) then "each rule needs at least one use profile" + elif any((.rules // [])[]; has("approval") and .approval != "captain") then "approval must be \"captain\" when present" + elif any((.rules // [])[]; has("select") and ((.select | type) != "string" or (.select | length) == 0)) then "select must be a non-empty string" + elif any((.rules // [])[]; has("select") and .select != "quota-balanced") then + "unknown select: " + ([.rules[] | select(has("select") and .select != "quota-balanced") | .select] | unique | join(", ")) + elif any((.rules // [])[]; has("floor") and floor_bad(.floor; true)) then "rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\\z" + elif any((.rules // [])[] | profiles(.use)[]; profile_bad(.)) then "each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\\z when present" + elif any((.rules // [])[]; duplicate_profiles(profiles(.use))) then "each rule use must not contain duplicate harness, model, and effort profiles" + elif any((.rules // [])[] | profiles(.use)[]; (verified(.harness) | not)) then "each use profile must name a verified harness" + elif any((.rules // [])[] | profiles(.use)[]; (effort_ok(.harness; .model; .effort) | not)) then "each use profile effort must be supported by its harness and model" + elif has("default") and (profiles(.default) | length) == 0 then "default must be a profile object or non-empty profile array" + elif has("default") and any(profiles(.default)[]; profile_bad(.)) then "each default profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\\z when present" + elif has("default") and duplicate_profiles(profiles(.default)) then "default must not contain duplicate harness, model, and effort profiles" + elif has("default") and any(profiles(.default)[]; (verified(.harness) | not)) then "each default profile must name a verified harness" + elif has("default") and any(profiles(.default)[]; (effort_ok(.harness; .model; .effort) | not)) then "each default profile effort must be supported by its harness and model" + else empty end +' "$RULES" 2>/dev/null) || die "malformed rules file: $RULES_PATH (not JSON)" +[ -z "$rules_err" ] || die "malformed rules file: $RULES_PATH - $rules_err" + +missing_provider=$(jq -r ' + def profiles($v): if ($v | type) == "array" then $v elif ($v | type) == "object" then [$v] else [] end; + ((.rules // [])[] | profiles(.use)[] | select(has("provider") | not) | "use\t\(.harness)"), + (profiles(.default // null)[] | select(has("provider") | not) | "default\t\(.harness)") +' "$RULES" | while IFS=$'\t' read -r location harness; do + if ! fm_quota_single_provider_for_harness "$harness" >/dev/null; then + printf '%s\t%s\n' "$location" "$harness" + break + fi +done) +if [ -n "$missing_provider" ]; then + IFS=$'\t' read -r location harness <<< "$missing_provider" + die "malformed rules file: $RULES_PATH - $location profiles whose harness lacks one authoritative provider family require provider: $harness" +fi + +# ---- harness -> provider map, from the single owner in fm-quota-axi-lib.sh ----- +PMAP='{}' +while IFS= read -r h; do + [ -n "$h" ] || continue + p=$(fm_quota_single_provider_for_harness "$h" 2>/dev/null) || p='' + PMAP=$(jq -c --arg h "$h" --arg p "$p" '. + {($h): (if $p == "" then null else $p end)}' <<<"$PMAP") +done < <(jq -r ' + def profiles($v): if ($v | type) == "array" then $v elif ($v | type) == "object" then [$v] else [] end; + ([((.rules // [])[]) | profiles(.use)[]] + profiles(.default // null)) + | map(.harness) | unique | .[]' "$RULES") + +RULE_COUNT=$(jq -r '(.rules // []) | length' "$RULES") + +emit_error() { + local reason=$1 + echo "dispatch-resolve: error ($reason)" >&2 + printf 'dispatch-resolve:\n status: error\n reason: %s\n' "$reason" + exit 0 +} + +if [ "$RULE_COUNT" -eq 0 ]; then + no_rules +fi + +RESP_FILE=$(mktemp) || die "mktemp failed" +QUOTA=$(mktemp) || { rm -f "$RESP_FILE"; die "mktemp failed"; } +trap 'rm -f "$RULES" "$RESP_FILE" "$QUOTA"' EXIT +LAT_MS=null +command -v curl >/dev/null 2>&1 || emit_error "curl not installed" + REQUEST=$(jq -n --rawfile brief "$BRIEF" --arg project "$PROJECT" --arg model "$TS_MODEL" \ + --arg none_criterion "$DEFAULT_WHEN" --slurpfile rules "$RULES" ' + ($rules[0]) as $cfg | + ($cfg.rules | to_entries | map({key: ("rule_" + ((.key + 1) | tostring)), value: .value.when}) | from_entries) as $criteria | + { + model: $model, + state: {task: {project: $project, brief: $brief}}, + questions: { + rule: { + type: "choice", + instructions: "Which ONE dispatch rule best fits `task` (read `task.brief` and `task.project`)? Each option is the rule'"'"'s own matching condition; pick `default` when no rule'"'"'s condition is met, including when a rule'"'"'s own exemption text excludes this task.", + criteria: ($criteria + {default: $none_criterion}) + } + } + }') + T0=$(fm_timing_now_ms) + HTTP=$(printf '%s' "$REQUEST" | curl -sS --max-time "$TS_TIMEOUT" -o "$RESP_FILE" -w '%{http_code}' \ + -X POST "$TS_BASE/v1/systemone" -H 'Content-Type: application/json' \ + -H @/dev/fd/3 3< <(printf 'Authorization: Bearer %s\n' "$TYPESAFE_API_KEY_PRIVATE") \ + --data-binary @- 2>/dev/null) || HTTP=000 + T1=$(fm_timing_now_ms) + LAT_MS=$(( T1 - T0 )) + [ "$HTTP" = 200 ] || emit_error "http $HTTP after ${LAT_MS} ms: $(head -c 200 "$RESP_FILE" 2>/dev/null | tr '\n' ' ')" +jq -e --slurpfile rules "$RULES" ' + (($rules[0].rules | to_entries | map("rule_" + ((.key + 1) | tostring))) + ["default"] | sort) as $choices | + (.answers.rule.choice | type) == "string" and + (.answers.rule.confidence | type) == "number" and + .answers.rule.confidence >= 0 and .answers.rule.confidence <= 1 and + (.answers.rule.probabilities | type) == "object" and + ((.answers.rule.probabilities | keys | sort) == $choices) and + all(.answers.rule.probabilities[]; type == "number" and . >= 0 and . <= 1) and + ((.answers.rule.probabilities | [.[]] | add) as $total | $total >= 0.99 and $total <= 1.01) and + ((has("usage") | not) or + ((.usage | type) == "object" and + (.usage.input_tokens | type) == "number" and + (.usage.output_tokens | type) == "number"))' \ + "$RESP_FILE" >/dev/null 2>&1 || emit_error "response is not a rule Choice answer" + +# ---- quota evidence: one quota-axi --json snapshot ----------------------------- +command -v quota-axi >/dev/null 2>&1 || emit_error "quota-axi not installed" +quota-axi --json > "$QUOTA" 2>/dev/null || emit_error "quota-axi --json failed" +fm_quota_json_valid < "$QUOTA" || emit_error "quota-axi --json returned an invalid snapshot" + +# ---- resolution: declared gates + quota evidence + argmax, all in jq ------------ +RESULT=$(jq -n --arg floor "$CONFIDENCE_FLOOR" --argjson lat "$LAT_MS" --arg none_criterion "$DEFAULT_WHEN" --argjson pmap "$PMAP" \ + --slurpfile resp "$RESP_FILE" --slurpfile rules "$RULES" --slurpfile quota "$QUOTA" ' + ($resp[0]) as $r | ($rules[0]) as $cfg | ($quota[0]) as $q | ($r.answers.rule) as $a | + def profiles($v): if ($v | type) == "array" then $v elif ($v | type) == "object" then [$v] else [] end; + def prov($p): ([$q.providers[] | select(.provider == $p)] | first) // null; + def rows($p): (prov($p) | .quotaSemantics.effectiveAvailability // []); + def bare($m): ($m | split("/") | last); + def provider_of($c): ($c.provider // $pmap[$c.harness] // null); + def measured($p): + (prov($p) != null and (["known", "partial"] | index(prov($p).quotaSemantics.status)) != null); + def applicable($p; $m): + (bare($m)) as $bare | + [rows($p)[] | select( + .scope == "all_models" or .scope == "all_products" or + ($m != "" and (.scope == ("model:" + $bare) or .scope == ("product:" + $bare))) + )]; + def floor_state($f; $p): + if $f == null then "none" + elif prov($p) == null or (measured($p) | not) then "unknown" + else [rows($p)[] | select(.scope == $f.scope)] as $matches + | if ($matches | length) == 0 or any($matches[]; .status != "known") then "unknown" + elif any($matches[]; .effectivePercentRemaining < $f.min_percent) then "below" + else "ok" + end + end; + def evidence($rows): + $rows | map({scope, status, pct: (.effectivePercentRemaining // null), runway: (.runway.status // null), spendPriority: (.selection.spendPriority // null)}); + def evaluate($c): + (provider_of($c)) as $p | + if $p == null then {profile: $c, eligible: false, reason: "no provider family for harness \($c.harness); declare provider on the profile"} + elif prov($p) == null then {profile: $c, provider: $p, eligible: true, unranked: true, reason: "provider \($p) not in the quota snapshot"} + else + (applicable($p; ($c.model // ""))) as $rows | + (evidence($rows)) as $bounds | + (floor_state($c.floor; $p)) as $profile_floor_state | + if any($rows[]; (.runway.status // "") == "exhausted_now") then + ($rows | map(select((.runway.status // "") == "exhausted_now")) | first) as $bad | + {profile: $c, provider: $p, bounds: $bounds, scope: $bad.scope, pct: ($bad.effectivePercentRemaining // null), runway: $bad.runway.status, eligible: false, reason: "runway exhausted_now at \($bad.scope)"} + elif any($rows[]; .status == "known" and (.effectivePercentRemaining | type) == "number" and .effectivePercentRemaining <= 0) then + ($rows | map(select(.status == "known" and (.effectivePercentRemaining | type) == "number" and .effectivePercentRemaining <= 0)) | first) as $bad | + {profile: $c, provider: $p, bounds: $bounds, scope: $bad.scope, pct: $bad.effectivePercentRemaining, runway: $bad.runway.status, eligible: false, reason: "0% remaining at \($bad.scope)"} + elif $profile_floor_state == "below" then + ([rows($p)[] | select( + .scope == $c.floor.scope and + .effectivePercentRemaining < $c.floor.min_percent + )] | first) as $floor_row | + {profile: $c, provider: $p, bounds: $bounds, scope: ($floor_row.scope // $c.floor.scope), pct: ($floor_row.effectivePercentRemaining // null), runway: ($floor_row.runway.status // null), eligible: false, reason: "profile floor \($c.floor.scope) below \($c.floor.min_percent)%"} + elif (measured($p) | not) then + ($rows | first) as $row | + {profile: $c, provider: $p, bounds: $bounds, scope: ($row.scope // null), pct: ($row.effectivePercentRemaining // null), runway: ($row.runway.status // null), eligible: true, unranked: true, unknown: true, reason: "provider \($p) unmeasured (\(prov($p).quotaSemantics.status))"} + elif ($rows | length) == 0 then + {profile: $c, provider: $p, bounds: $bounds, eligible: true, unranked: true, unknown: true, reason: "no applicable quota row for provider \($p)"} + elif $profile_floor_state == "unknown" then + ([rows($p)[] | select(.scope == $c.floor.scope)] | first) as $floor_row | + {profile: $c, provider: $p, bounds: $bounds, scope: $c.floor.scope, pct: ($floor_row.effectivePercentRemaining // null), runway: ($floor_row.runway.status // null), eligible: true, unranked: true, unknown: true, reason: "profile floor \($c.floor.scope) is unverifiable: not rankable"} + elif any($rows[]; .status != "known") then + ($rows | map(select(.status != "known")) | first) as $bad | + {profile: $c, provider: $p, bounds: $bounds, scope: $bad.scope, eligible: true, unranked: true, unknown: true, reason: "quota row \($bad.scope) unknown: not rankable"} + elif any($rows[]; (.selection.spendPriority | type) != "number") then + ($rows | map(select((.selection.spendPriority | type) != "number")) | first) as $bad | + {profile: $c, provider: $p, bounds: $bounds, scope: $bad.scope, pct: $bad.effectivePercentRemaining, runway: $bad.runway.status, eligible: true, unranked: true, reason: "spendPriority missing or non-numeric at \($bad.scope): not rankable"} + else + ($rows | min_by(.selection.spendPriority)) as $limiting | + {profile: $c, provider: $p, bounds: $bounds, scope: $limiting.scope, pct: $limiting.effectivePercentRemaining, + spendPriority: $limiting.selection.spendPriority, runway: $limiting.runway.status, eligible: true, reason: "ok"} + end + end; + ($a.choice) as $choice | + (if ($choice | test("^rule_[1-9][0-9]*$")) + then ($choice | ltrimstr("rule_") | tonumber) + else null end) as $rule_number | + (if $choice == "default" then null + elif $rule_number != null and $rule_number <= (($cfg.rules // []) | length) then $cfg.rules[$rule_number - 1] + else null end) as $rule | + (if $rule == null then "none" else floor_state($rule.floor; $rule.floor.provider) end) as $rule_floor_state | + (if $choice != "default" and $rule == null then [] + elif $rule == null then profiles($cfg.default // null) + else profiles($rule.use) + end) as $answer_use | + (if $choice != "default" and $rule == null then {invalid: "rule \($choice) is not in the rules file"} + elif $rule == null then {source: "default", use: profiles($cfg.default // null), note: "no rule matched"} + elif ($rule.approval // "") == "captain" then {source: $choice, escalate: "rule requires the captain'"'"'s explicit approval before dispatch"} + elif $rule_floor_state == "unknown" then {source: $choice, escalate: "rule \($choice) floor \($rule.floor.provider)/\($rule.floor.scope) is unverifiable"} + elif $rule_floor_state == "below" + then {source: "default", use: profiles($cfg.default // null), note: "rule \($choice) floor \($rule.floor.scope) below \($rule.floor.min_percent)%: fall through to default"} + else {source: $choice, use: profiles($rule.use), note: "rule matched"} end) as $sel | + { + model: $r.model, latency_ms: $lat, tokens: ($r.usage // null), + rule: $choice, + rule_when: (if $rule == null then $none_criterion else $rule.when end | .[0:60]), + confidence: $a.confidence, probabilities: $a.probabilities + } as $ev | + if $sel.invalid then $ev + {status: "error", reason: $sel.invalid} + elif $a.confidence < ($floor | tonumber) then + $ev + {status: "ambiguous", reason: "confidence \($a.confidence) below floor \($floor)", candidates: ($answer_use | map(evaluate(.)))} + elif $sel.escalate then + $ev + {status: "escalate", reason: $sel.escalate, candidates: ($answer_use | map(evaluate(.)))} + elif ($sel.use | length) == 0 then $ev + {status: "escalate", reason: "no profiles configured for \($sel.source)", note: $sel.note, candidates: []} + else + ($sel.use | map(evaluate(.))) as $cands | + ([$cands[] | select(.eligible and ((.unranked // false) | not))]) as $elig | + ([$cands[] | select(.unranked)]) as $unranked | + if ($elig | length) == 0 then $ev + {status: "escalate", reason: "no rankable eligible candidate", note: $sel.note, candidates: $cands} + else + ($elig | max_by(.spendPriority)) as $best | + ([$elig[] | select(.spendPriority == $best.spendPriority)] | length) as $ties | + if $ties > 1 then $ev + {status: "escalate", reason: "genuine spendPriority tie", note: $sel.note, candidates: $cands} + else $ev + {status: "clear", note: $sel.note, candidates: $cands, chosen: $best} + + (if ($unranked | length) > 0 then + {unranked_note: "\($unranked | length) eligible candidate(s) unranked (\([$unranked[].provider] | unique | join(", ")))"} + else {} end) + end + end + end') || emit_error "resolution failed" + +TEXT=$(jq -r ' + def flat: tostring | gsub("[\t\r\n]"; " "); + def show($value): ($value // "-") | flat; + def shell_arg: flat | @sh; + "dispatch-resolve:", + " status: \(.status | flat)", + " model: \(show(.model)) latency_ms: \(show(.latency_ms)) tokens: \(show(.tokens.input_tokens))/\(show(.tokens.output_tokens))", + " rule: \(.rule | flat) (\(.rule_when | flat)) confidence: \(.confidence | flat)", + " probabilities: \([.probabilities | to_entries[] | "\(.key | flat)=\(.value | flat)"] | join(" "))", + (if .reason then " reason: \(.reason | flat)" else empty end), + (if .note then " note: \(.note | flat)" else empty end), + (if .unranked_note then " note: \(.unranked_note | flat)" else empty end), + (.candidates[]? | " candidate: \(.profile.harness | flat):\(show(.profile.model))" + + (if .provider then " provider=\(.provider | flat)" else "" end) + + (if .scope then " scope=\(.scope | flat) remaining=\(show(.pct))% spendPriority=\(show(.spendPriority)) runway=\(show(.runway))" else "" end) + + (if (.bounds // [] | length) > 1 then " bounds=" + ([.bounds[] | "\(.scope | flat):\(show(.pct))%/\((.runway // .status) | flat)"] | join(",")) else "" end) + + " -> " + (if .unranked then "eligible, unranked: \(.reason | flat): disclosed uncertainty" elif .eligible then "eligible" else "not eligible: \(.reason | flat)" end)), + (if .chosen then " profile: --harness \(.chosen.profile.harness | shell_arg)" + + (if .chosen.profile.model then " --model \(.chosen.profile.model | shell_arg)" else "" end) + + (if .chosen.profile.effort then " --effort \(.chosen.profile.effort | shell_arg)" else "" end) else empty end)' <<<"$RESULT") || emit_error "output rendering failed" +printf '%s\n' "$TEXT" +exit 0 diff --git a/bin/fm-env-lib.sh b/bin/fm-env-lib.sh new file mode 100644 index 00000000000..fd27ead1c1c --- /dev/null +++ b/bin/fm-env-lib.sh @@ -0,0 +1,31 @@ +# shellcheck shell=bash +# Shared .env-style file accessor. +# Usage: . bin/fm-env-lib.sh +# +# This file is the single owner of the one-key .env read: the Relay pairing +# token (bin/fm-x-lib.sh and its callers) and the optional typesafe.ai +# dispatch key (bin/fm-dispatch-resolve.sh) both resolve their value through +# fmx_env_get, so those opt-in secrets in $FM_HOME/.env are parsed by one rule. +# (bin/fm-mail.sh loads its whole .env block itself under the same env-wins +# contract.) The value is printed to the caller's command substitution only; +# nothing is logged. + +# fmx_env_get +# Read the value of KEY from a .env-style file: last assignment wins; tolerates a +# leading "export ", surrounding whitespace, and one layer of matching single or +# double quotes. Prints nothing (and succeeds) when the file or key is absent, so +# callers can treat empty output as "unset". +fmx_env_get() { + local key=$1 file=$2 line val + [ -f "$file" ] || return 0 + line=$(grep -E "^[[:space:]]*(export[[:space:]]+)?${key}=" "$file" 2>/dev/null | tail -n1) || return 0 + [ -n "$line" ] || return 0 + val=${line#*=} + val=${val#"${val%%[![:space:]]*}"} # strip leading whitespace + val=${val%"${val##*[![:space:]]}"} # strip trailing whitespace (incl. CR) + case "$val" in + \"*\") val=${val#\"}; val=${val%\"} ;; + \'*\') val=${val#\'}; val=${val%\'} ;; + esac + printf '%s' "$val" +} diff --git a/bin/fm-quota-axi-lib.sh b/bin/fm-quota-axi-lib.sh index 0ade3fb7db9..7a2df68a440 100644 --- a/bin/fm-quota-axi-lib.sh +++ b/bin/fm-quota-axi-lib.sh @@ -10,6 +10,7 @@ # what keeps an older build from reaching a dispatch intake at all. FM_QUOTA_AXI_MIN=0.1.29 +FM_QUOTA_PROVIDER_ID_RE='^[a-z0-9]+(-[a-z0-9]+)*\z' fm_quota_axi_compatible() { local timeout=${1:-} output parts major minor patch extra @@ -42,7 +43,7 @@ fm_quota_axi_compatible() { } fm_quota_json_valid() { - jq -se ' + jq -se --arg provider_re "$FM_QUOTA_PROVIDER_ID_RE" ' length == 1 and (.[0] | type) == "object" and (.[0] | @@ -51,7 +52,7 @@ fm_quota_json_valid() { (([.providers[].provider] | length) == ([.providers[].provider] | unique | length)) and all(.providers[]; (.provider | type) == "string" and - (.provider | test("^[a-z0-9]+(-[a-z0-9]+)*$")) and + (.provider | test($provider_re)) and (.quotaSemantics | type) == "object" and (.quotaSemantics.status as $semantics_status | (["known", "partial", "unknown"] | index($semantics_status)) != null and @@ -91,3 +92,46 @@ fm_quota_json_valid() { ) ' >/dev/null 2>&1 } + +fm_quota_single_provider_table() { + printf '%s\n' \ + 'claude claude' \ + 'codex codex' \ + 'grok grok' \ + 'kimi kimi' \ + 'cursor cursor' \ + 'agy agy' \ + 'muse meta' +} + +fm_quota_single_provider_for_harness() { + local harness provider + while read -r harness provider; do + if [ "$harness" = "$1" ]; then + printf '%s\n' "$provider" + return 0 + fi + done < <(fm_quota_single_provider_table) + return 1 +} + +fm_quota_provider_for_harness() { + case "$1" in + omp) + case "${2:-}" in + openai-codex/*) printf 'codex\n' ;; + claude-bridge/*) printf 'claude\n' ;; + *) return 1 ;; + esac + ;; + claude) printf 'claude\n' ;; + codex) printf 'codex\n' ;; + opencode) printf 'codex\n' ;; + pi|pi-signed) printf 'pi\n' ;; + grok) printf 'grok\n' ;; + kimi) printf 'kimi\n' ;; + cursor) printf 'cursor\n' ;; + muse) printf 'meta\n' ;; + *) return 1 ;; + esac +} diff --git a/bin/fm-quota-choose.sh b/bin/fm-quota-choose.sh index 3c7fa891c56..4bfe89247bf 100755 --- a/bin/fm-quota-choose.sh +++ b/bin/fm-quota-choose.sh @@ -24,7 +24,8 @@ # candidate remains eligible under the captured quota evidence. # # Multi-provider limitation: this helper maps each harness to ONE primary -# provider family (see provider_for_harness below) and checks quota for that +# provider family (fm_quota_provider_for_harness in bin/fm-quota-axi-lib.sh) +# and checks quota for that # family only. Some harnesses can run models from several providers - for # example, Pi and OpenCode may dispatch xAI, Anthropic, or other models - so a # candidate whose established provider differs from the harness's primary family @@ -309,31 +310,11 @@ fi printf '%s\n' "$QUOTA_JSON" | fm_quota_json_valid || die "invalid quota-axi provider data" # provider_for_harness [] -# Map a firstmate harness name to its primary quota-axi provider family. -# Multi-provider harnesses (Pi, OpenCode) map to their primary family only; see -# the header limitation note. omp is keyed on the candidate model prefix instead -# and has no family for any other prefix (see the header). Authoritative -# multi-provider routing is owned by AGENTS.md section 4 and the -# quota-array-dispatch skill, not this helper. +# The harness -> primary provider family table is owned by +# fm_quota_provider_for_harness in bin/fm-quota-axi-lib.sh; see the header +# limitation note for why one family per harness is all this helper checks. provider_for_harness() { - case "$1" in - omp) - case "${2:-}" in - openai-codex/*) printf 'codex\n' ;; - claude-bridge/*) printf 'claude\n' ;; - *) return 1 ;; - esac - ;; - claude) printf 'claude\n' ;; - codex) printf 'codex\n' ;; - opencode) printf 'codex\n' ;; - pi|pi-signed) printf 'pi\n' ;; - grok) printf 'grok\n' ;; - kimi) printf 'kimi\n' ;; - cursor) printf 'cursor\n' ;; - muse) printf 'meta\n' ;; - *) return 1 ;; - esac + fm_quota_provider_for_harness "$@" } # effective_for_provider_model diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 1f1bda6b8b9..0bfc3e941ec 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -396,6 +396,7 @@ family_for_basename() { fm-branch-supervision.test.sh|fm-busy-adapter-wiring.test.sh|\ fm-busy-state.test.sh|fm-classify-corr-token.test.sh|\ fm-claude-stop-autoarm.test.sh|fm-cursor-harness.test.sh|\ + fm-dispatch-resolve.test.sh|\ fm-extension-binding.test.sh|fm-gitignore-config.test.sh|\ fm-no-mistakes-required.test.sh|fm-peek-remote.test.sh|\ fm-pending-reply.test.sh|fm-pi-branch-extension.test.sh|\ @@ -698,6 +699,7 @@ tests/fm-control.test.sh 54301 tests/fm-cursor-harness.test.sh 30103 tests/fm-cursor-primary-live-e2e.test.sh 21 tests/fm-cursor-primary.test.sh 54947 +tests/fm-dispatch-resolve.test.sh 1800 tests/fm-daemon.test.sh 26870 tests/fm-documentation-audiences.test.sh 732 tests/fm-extension-binding.test.sh 7398 @@ -1408,6 +1410,7 @@ families_for_changed_path() { printf '%s\n' session-bootstrap printf '%s\n' "__script__:fm-procevent-quota.test.sh" printf '%s\n' "__script__:fm-quota-choose.test.sh" + printf '%s\n' "__script__:fm-dispatch-resolve.test.sh" ;; bin/fm-procevent-quota.sh) printf '%s\n' "__script__:fm-procevent-quota.test.sh" @@ -1415,6 +1418,15 @@ families_for_changed_path() { bin/fm-quota-choose.sh) printf '%s\n' "__script__:fm-quota-choose.test.sh" ;; + bin/fm-dispatch-resolve.sh) + printf '%s\n' "__script__:fm-dispatch-resolve.test.sh" + ;; + bin/fm-env-lib.sh) + # The one .env accessor, sourced by bin/fm-x-lib.sh (Relay token) and + # bin/fm-dispatch-resolve.sh (TYPESAFE_API_KEY). + printf '%s\n' pr-forge + printf '%s\n' "__script__:fm-dispatch-resolve.test.sh" + ;; .pi/extensions/fm-branch-supervision.ts|.pi/extensions/lib/fm-async-exec.ts|\ .pi/extensions/lib/fm-branch-dispatch.ts|.pi/extensions/lib/fm-native-contract.ts) # The portable suites that actually load these files, named one by one. diff --git a/bin/fm-x-lib.sh b/bin/fm-x-lib.sh index aae910db8cb..fb1f1a61e6e 100644 --- a/bin/fm-x-lib.sh +++ b/bin/fm-x-lib.sh @@ -8,6 +8,7 @@ # # This file is sourced, never executed. It defines: # fmx_env_get - read one KEY=VALUE from a .env-style file +# (defined by bin/fm-env-lib.sh, sourced here) # fmx_load_config - resolve FMX_TOKEN, FMX_RELAY, FMX_DRY, FMX_MAX, # and FMX_THREAD_MAX (env wins over .env) # fmx_auth_header_file - write the bearer header to a 0600 temp file @@ -56,24 +57,9 @@ if ! command -v fm_backlog_atomic_transition >/dev/null 2>&1; then . "$_FM_X_LIB_DIR/fm-backlog-transition-lib.sh" fi -# Read the value of KEY from a .env-style file: last assignment wins; tolerates a -# leading "export ", surrounding whitespace, and one layer of matching single or -# double quotes. Prints nothing (and succeeds) when the file or key is absent, so -# callers can treat empty output as "unset". -fmx_env_get() { - local key=$1 file=$2 line val - [ -f "$file" ] || return 0 - line=$(grep -E "^[[:space:]]*(export[[:space:]]+)?${key}=" "$file" 2>/dev/null | tail -n1) || return 0 - [ -n "$line" ] || return 0 - val=${line#*=} - val=${val#"${val%%[![:space:]]*}"} # strip leading whitespace - val=${val%"${val##*[![:space:]]}"} # strip trailing whitespace (incl. CR) - case "$val" in - \"*\") val=${val#\"}; val=${val%\"} ;; - \'*\') val=${val#\'}; val=${val%\'} ;; - esac - printf '%s' "$val" -} +# fmx_env_get lives in bin/fm-env-lib.sh, the single owner of .env parsing. +# shellcheck source=bin/fm-env-lib.sh +. "$_FM_X_LIB_DIR/fm-env-lib.sh" fmx_poll_shim_content() { local home=$1 root=$2 diff --git a/docs/configuration.md b/docs/configuration.md index cb6ead4c3c1..46949796ef2 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -426,8 +426,10 @@ This section is the single owner of the canonical schema and its per-field seman "rules": [ { "when": "", + "approval": "captain", + "floor": { "scope": "", "min_percent": 20, "provider": "" }, "use": [ - { "harness": "", "model": "", "effort": "" } + { "harness": "", "model": "", "effort": "", "provider": "", "floor": { "scope": "", "min_percent": 50 } } ], "why": "" } @@ -438,10 +440,23 @@ This section is the single owner of the canonical schema and its per-field seman } ``` -Per rule, `when` and `use` are required. +Per rule, `when` and `use` are required; the top-level `rules` array itself may be absent or empty for a default-only configuration. Both `use` and the optional top-level `default` accept either one profile object or a non-empty array of profile objects. The single-object form stays fully backward-compatible, and every profile needs `harness`. Profile `model` and `effort` fields and rule `why` are optional. +Rule `approval` and `floor`, and profile `provider` and `floor` are optional declarations that only [typed dispatch resolution](#typed-dispatch-resolution-env-typesafe_api_key) applies in code; without that opt-in they are inert, and firstmate's own intake reads them as ordinary hints. +The resolver supplies the fixed neutral Choice option `No listed rule applies to this task.` for work that matches no listed rule. +`approval` accepts only `"captain"` and means a task the rule matches is never dispatched from the tool's answer alone. +A rule `floor` names the quota-axi `provider` and `scope` whose `effectivePercentRemaining` must be at least `min_percent` for the rule's profiles to apply. +A known percentage below it makes the tool resolve among `default` instead; an absent or unknown row or unmeasured provider makes the floor unverifiable and escalates without authorizing default routing. +A profile `provider` optionally names the quota-axi provider family whose rows apply to that profile; when present, profile and rule-floor provider IDs must match the strict whole-string pattern `^[a-z0-9]+(-[a-z0-9]+)*\z`. +Bootstrap validates resolver-only `approval`, `floor`, and present `provider` values only while typed resolution is active; without the key those inert fields and the pre-existing verified-harness baseline preserve bootstrap behavior. +Typed resolution additively recognizes `gemini` because AGENTS.md section 4 verifies it for crewmate and scout dispatch. +The opted-in resolver has authoritative single-provider mappings for `claude`, `codex`, `grok`, `kimi`, `cursor`, `agy`, and `muse`; every other verified harness must declare `provider` explicitly, including multi-provider `pi`, `pi-signed`, `omp`, and `opencode` and unmapped `gemini` and `rovo`. +Its single-provider table is separate from the frozen legacy mapping used by `fm-quota-choose.sh`, so additions cannot alter no-key routing. +The resolver returns an actionable configuration error before any request when such a profile omits it. +A profile `floor` contains only `scope` and `min_percent`, always uses that profile's provider, and makes that one candidate ineligible below `min_percent` on the named scope. +An absent or unknown named row also makes the candidate unrankable and is reported as an unverifiable floor, not as a known shortfall. `ultra` is native-only: the model-aware validation contract and launch mapping are owned by `bin/fm-harness.sh validate-native-effort` and `bin/fm-spawn.sh` respectively. Codex `max` is valid when the profile selects `gpt-5.6-luna`, whose installed catalog entry supports that reasoning level. An omitted model or effort means the selected harness uses its own default for that axis. @@ -449,13 +464,48 @@ Every profile array is an implicit quota-aware choice resolved through `quota-ar If no dispatch rule fits, firstmate resolves `default` through the same object-or-array path before falling back to `config/crew-harness`. Except for `ultra`, which refuses unsupported profiles under the native-effort contract above, an effort value the chosen harness does not accept is recorded as `effort=` in task meta for traceability but omitted from the launch flags. Bootstrap reports unsupported harness/model/effort combinations as a `CREW_DISPATCH` diagnostic when they are visible in the file. -See [`docs/examples/crew-dispatch.json`](examples/crew-dispatch.json) for a starting point to copy into local `config/crew-dispatch.json`. +See [`docs/examples/crew-dispatch.json`](examples/crew-dispatch.json) for a starting point to copy into local `config/crew-dispatch.json`; its Pi default declares the `claude` provider required for typed resolution of that Anthropic model. When the file exists, bootstrap validates it with `jq`. Valid files stay silent by default; with `FM_BOOTSTRAP_VERBOSE_FACTS=1`, bootstrap emits `BOOTSTRAP_INFO: crew dispatch active config/crew-dispatch.json`, one `BOOTSTRAP_INFO:` fact per rule, and one fact for the optional default profile set. -Malformed JSON, an empty or malformed rule/default array, an unverified harness, or an effort value unsupported by that harness is reported as `CREW_DISPATCH: invalid config/crew-dispatch.json - ...`; missing `jq` is reported through the normal `MISSING: jq` install-consent flow. +Malformed JSON, malformed rules, an empty or malformed profile array, an unverified harness, or an effort value unsupported by that harness is reported as `CREW_DISPATCH: invalid config/crew-dispatch.json - ...`. +While typed resolution is active, malformed `approval`, `floor`, and present `provider` declarations receive the same diagnostic; without the key those inert declarations preserve the pre-existing bootstrap behavior. +Missing `jq` is reported through the normal `MISSING: jq` install-consent flow. While the file remains present, no crewmate or scout spawn may proceed without an explicit resolved harness; malformed configuration must be reported and corrected rather than selected around. Secondmate homes inherit this file from the primary, so a secondmate's own crewmates apply the same dispatch profile behavior. +## Typed dispatch resolution (.env TYPESAFE_API_KEY) + +`bin/fm-dispatch-resolve.sh` resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model (Jev), so the rule match that firstmate otherwise reasons out in its own context becomes one short tool turn. +It is off unless `TYPESAFE_API_KEY` is non-empty in the calling environment or the home's gitignored `.env` holds a `TYPESAFE_API_KEY=` line; the environment wins, matching the Relay and mail-plane contracts, and the Relay accessor in `bin/fm-env-lib.sh` reads the line. +Off means one `dispatch-resolve: off` line on stderr, nothing on stdout, exit 0, and no network call, so firstmate dispatches exactly as it does without the tool. +This section is the single owner of the tool's operator contract; the script header owns its exact flags and output lines, and "Crew dispatch profiles" above owns the declared rule and profile fields it applies. +Rules come only from the effective home's `config/crew-dispatch.json`; `FM_CONFIG_OVERRIDE` selects the config directory for tests and specialized setup like the other scripts. + +```sh +bin/fm-dispatch-resolve.sh data//brief.md --project # TOON block on stdout +``` + +Firstmate invokes the resolve path directly after writing the brief, without a preflight; the absent-key off line is handled exactly like every other non-clear outcome. +When on and at least one rule exists, the tool sends the project name and the whole brief as state and asks one Choice question whose options are every rule's `when` plus the fixed neutral option for no matching rule; the model never sees quota, catalogs, `why`, `use`, or approvals. +An absent rules file, a default-only file, or `rules: []` returns the non-clear reason `no rules to match` without a model or quota request, leaving firstmate's existing routing in control; an existing but unreadable or malformed rules file, including a broken symlink, remains an actionable exit 2 configuration error. +Everything after the answer runs in code: the confidence floor, the matched rule's `approval` and `floor`, each candidate's `provider` and `floor`, every applicable account-wide and model/product row from one `quota-axi --json` snapshot, and the numeric `spendPriority` argmax over candidates using each candidate's limiting row. +Known applicable rows from a provider with partial quota semantics remain rankable; rows whose own status is not known remain unrankable. +Any applicable `exhausted_now` row or known zero bound makes that candidate ineligible, and a known profile-floor shortfall does the same before unrelated quota uncertainty is considered. +Missing or nonnumeric `spendPriority` evidence is never ranked, and every candidate is printed beside its evidence or the reason it was not rankable, including on ambiguous and approval-gated outcomes that emit no profile. +On the opted-in path, duplicate concrete profiles with the same harness, model, and effort inside one rule or the default array are configuration errors rather than ties. +The result is one of `clear` (a `profile:` line ready for `fm-spawn.sh`), `ambiguous` (confidence below the floor), `escalate` (an approval-gated rule, unverifiable rule floor, nothing rankable, or a genuine tie), or `error` (API, network, malformed response metadata, rendering, or quota-axi failure), and every one of them exits 0. +Response probabilities must contain exactly every offered choice, use numeric values from 0 through 1, and sum to approximately 1 within 0.01. +Only a usage or configuration error exits 2: an unreadable brief, an existing but unreadable or malformed canonical rules file, or missing `jq`, each reported and never selected around. +Missing `curl` is a normal structured `error` outcome with exit 0 so firstmate uses today's routing. +The tool never replaces firstmate's judgment, `quota-array-dispatch`, the captain-approval gate, or `fm-spawn.sh` validation; `AGENTS.md` section 4 owns what firstmate does with each outcome. +By accepted design, a `clear` result does not enforce catalog/authentication, reasoning-class, or completion-runway gates. +Firstmate passes its profile line unless it states a reason to override, such as the brief's reasoning class or an eligible-unranked-candidate note; every non-clear result returns to the full existing intake. + +The resolver and bootstrap copy an environment-provided key into a non-exported private variable and unset `TYPESAFE_API_KEY` before launching child processes, so the secret is absent from child environments. +The resolver sends the key to `curl` only as a header read from a file descriptor, never on argv, and nothing prints, logs, or writes it. +The resolver fixes the endpoint at `https://api.typesafe.ai`, model at `jev-latest`, confidence floor at 0.6, and request timeout at 5 seconds; `TYPESAFE_API_KEY` is its only resolver-specific environment setting. +The live rule-match evidence is recorded in [`verification/dispatch-resolve.md`](verification/dispatch-resolve.md). + ## Toolchain On session start the first mate detects what its required toolchain is missing or too old and lists each problem with either an exact install command or manual instructions. @@ -1040,6 +1090,7 @@ FMX_RELAY_URL=https://myfirstmate.io # optional Relay endpoint override, mainl FMX_ENV_FILE= # optional alternate .env file for direct Relay client invocations; bootstrap still checks $FM_HOME/.env FMX_DRY_RUN= # truthy previews Relay replies and dismissals to state/x-outbox/ without posting or requiring a token FMX_X_REPLY_MAX_CHARS=280 # X reply per-message split budget; values below 50 clamp to 50 +TYPESAFE_API_KEY= # typed dispatch resolution opt-in, from the environment or .env; absent means bin/fm-dispatch-resolve.sh is off (docs/configuration.md "Typed dispatch resolution") FMX_DISCORD_REPLY_MAX_CHARS=1900 # Discord reply per-message split budget; values below 50 clamp to 50, values above 2000 reset to 1900 FMX_X_THREAD_MAX=25 # maximum messages in one auto-split reply thread FMX_FOLLOWUP_MAX_AGE_SECS=604800 # local window for posting Relay completion follow-ups (7 days) diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index 618caa941ec..f639220b551 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -448,6 +448,10 @@ "path": "docs/verification/dispatch-auth.md", "audience": "maintainer-verification" }, + { + "path": "docs/verification/dispatch-resolve.md", + "audience": "maintainer-verification" + }, { "path": "docs/verification/lint-option-a.md", "audience": "maintainer-verification" diff --git a/docs/examples/crew-dispatch.json b/docs/examples/crew-dispatch.json index b404e95e777..97c5ad38db1 100644 --- a/docs/examples/crew-dispatch.json +++ b/docs/examples/crew-dispatch.json @@ -21,6 +21,6 @@ ], "default": [ { "harness": "codex", "model": "gpt-5.5", "effort": "medium" }, - { "harness": "pi", "model": "anthropic/claude-sonnet-5", "effort": "medium" } + { "harness": "pi", "model": "anthropic/claude-sonnet-5", "effort": "medium", "provider": "claude" } ] } diff --git a/docs/verification/dispatch-resolve.md b/docs/verification/dispatch-resolve.md new file mode 100644 index 00000000000..58152196181 --- /dev/null +++ b/docs/verification/dispatch-resolve.md @@ -0,0 +1,73 @@ +# Typed dispatch resolution verification + +Audience: maintainer verification. + +This record supports the opt-in `bin/fm-dispatch-resolve.sh` contract owned by [`../configuration.md`](../configuration.md) ("Typed dispatch resolution") and the declared rule and profile fields owned there under "Crew dispatch profiles". +It records only facts that must be re-established when the typesafe.ai model, its API, or firstmate's dispatch rules change. +Task chronology, the captain's rules, and the briefs themselves stay in the private scout report. + +## The API the tool depends on + +Verified 2026-09-16 against `https://api.typesafe.ai`. +`GET /v1/models` listed `jev-latest` and `jev-preview`, both released 2026-09-10; a `jev-latest` request answered as `jev-1.13.0`. +`POST /v1/systemone` takes `{model, state, questions}`; a `choice` question returns `{choice, probabilities, confidence}` with the probabilities summing to 1. +Observed error shapes: 401 `authentication_error` for a bad key, 403 when the header is missing, 422 with a `detail[].loc` naming the offending field, 400 `api_usage_error` for an unknown model, 405 on GET. +No rate-limit headers were present on any response; every response carried `x-typesafe-request-id`. +Observed end-to-end latency from a Mac was 123 to 348 ms per request, with the server's own upstream time at 4 to 60 ms. + +## Live rule match against real briefs + +Run 2026-09-16 with the key injected for the one command through the vault (`av inject +TYPESAFE_API_KEY -- ...`), model `jev-latest`, confidence floor 0.6, timeout 5 s, one `quota-axi --json` snapshot for the whole run. +Rules: the captain's five-rule file with a captain-authored none option, one `approval: captain` rule, two rule floors on `model:fable`, and declared `provider` on the Pi profiles. +Briefs: 15 real briefs from this home's recent work plus 10 synthetic ones written to hit each rule. + +| Measure | Result | +| --- | --- | +| Rule matched the hand label | 20 of 25 | +| Resolved to the hand-labeled profile | 20 of 25 | +| Outcomes: clear / ambiguous / escalate / error | 18 / 1 / 6 / 0 | +| Clear results with a wrong profile | 0 | +| API latency (min / median / max) | 152 / 214 / 348 ms | +| Wall time per call including jq (min / median / max) | 198 / 261 / 396 ms | +| Input tokens per brief (min / median / max) | 1,279 / 3,114 / 4,538 | +| Output tokens | 150 to 152 | +| API errors | 0 | + +Of the five disagreements, one was a wrong hand label (the brief quoted the bug-fix rule's wording verbatim), three were real briefs the model read as the approval-gated design rule at 0.66 to 0.86 confidence and escalated by design, each of which the captain had in fact dispatched at the strongest-reasoning class, and one was a synthetic tweak that came back ambiguous at 0.41 confidence and was handed back to firstmate. +A lean request that asks only the rule Choice matched the full request (rule, profile, and status) on all 25 briefs, which is why the shipped tool asks one question and keeps every gate in code. +That table records the 2026-09-16 run with the captain-authored none option. +A second live run on 2026-09-17 used the same 25 briefs, held one quota snapshot constant through a fake `quota-axi`, and exercised a copy of this branch with the shipped neutral `No listed rule applies to this task.` option and option-free interface. + +| Measure | Result | +| --- | --- | +| Rule matched the hand label | 20 of 25 | +| Resolved to the hand-labeled profile | 18 of 25 | +| Outcomes: clear / ambiguous / escalate / error | 17 / 2 / 6 / 0 | +| Clear results with a profile other than the hand label | 1 | +| API latency (min / median / max) | 137 / 220 / 1,795 ms | +| Input tokens per brief (min / median / max) | 754 / 2,589 / 4,013 | +| Output tokens | 60 to 62 | +| API errors | 0 | + +The maximum latency was one outlier; the next slowest request was 309 ms. +The differing clear result was a synthetic small tweak that matched the simple-bug-fix rule at 0.90 and selected `cursor-grok-4.6-medium` instead of the hand-labeled `cursor-grok-4.6-high`: the tweak exemption removed from the none-option text belongs in that rule's own `when` text. +Two default-labeled briefs became ambiguous. + +## Offline behavior + +`tests/fm-dispatch-resolve.test.sh` drives the public interface with a fake `curl` that records argv, the request body, the header read from file descriptor 3, and whether the secret reached its environment, plus a fake `quota-axi` that performs the same environment check. +It proves firstmate can invoke the resolve path without a preflight, rules are snapshotted once from the isolated home's canonical `config/crew-dispatch.json`, and dynamic output fields are flattened to one line. +It proves the absent key (environment and `.env`) prints one stderr line, nothing on stdout, exits 0, and never invokes `curl` or `quota-axi`. +It proves absent, default-only, and empty-rules files return `no rules to match` without a model or quota request, while a broken rules-file symlink exits 2 as unreadable. +It proves the documented starter configuration resolves its Pi default through the declared Claude provider, a `.env` key turns the tool on, and the environment wins over it. +It proves the key is absent from child environments, never appears on `curl` argv, and arrives only as the bearer header on the descriptor. +It proves the request uses the fixed endpoint and model, carries only the project, brief, and rule Choice with one option per rule plus the fixed neutral none option, and never carries `why`, `use`, or quota. +It proves the clear, fixed-floor ambiguous with candidate evidence, escalate (approval with candidate evidence, unverifiable rule floor, tie, nothing rankable), known rule-floor fall-through, known and unverifiable profile-floor evidence, explicit-provider and provider-ID enforcement, authoritative Agy and explicit-provider Gemini routing, partial providers, eligible unranked candidates and their clear-result note, concrete quota vetoes and profile-floor shortfalls taking precedence over uncertainty, account-wide quota veto, limiting-bound ranking, missing-curl and quota-axi failures, HTTP 429 and 500, transport failure, malformed usage, zero-mass or malformed probabilities or confidence, malformed or duplicate profile, invalid selector, removed-option rejection, and out-of-range rule ID paths behave as the contract states, with configuration errors exiting 2 before any network call. +`tests/fm-bootstrap.test.sh` proves bootstrap ignores resolver-only fields without the typed key, validates each malformed shape when the environment or home `.env` activates typed resolution, and prevents an environment-provided key from reaching child processes. + +```console +$ bash tests/fm-dispatch-resolve.test.sh | tail -1 +# all fm-dispatch-resolve tests passed +``` + +A live run needs a key and is not part of the suite; rerun the table above by pointing the tool at a brief with the key injected for that one command. diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index 561f8100aa1..d8cc824f0dd 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -135,6 +135,13 @@ add_real_jq() { real_jq=$(command -v jq 2>/dev/null) || fail "jq is required for dispatch profile validation tests" cat > "$fakebin/jq" <> "\$FM_TEST_CHILD_ENV_LOG" + else + printf 'clean\n' >> "\$FM_TEST_CHILD_ENV_LOG" + fi +fi exec '$real_jq' "\$@" SH chmod +x "$fakebin/jq" @@ -1098,7 +1105,7 @@ test_crew_dispatch_active_rules_are_verbose_bootstrap_info() { } test_crew_dispatch_validation() { - local label body expect mode case_dir fakebin out n + local label body expect mode case_dir fakebin out child_env n n=0 while IFS='^' read -r label body mode expect; do [ -n "$label" ] || continue @@ -1110,7 +1117,7 @@ test_crew_dispatch_validation() { fakebin=$(make_fake_toolchain "$case_dir") add_real_jq "$fakebin" out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ - FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + TYPESAFE_API_KEY=test-key FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") case "$mode" in empty) [ -z "$out" ] || fail "$label: expected silence, got: $out" ;; @@ -1126,21 +1133,22 @@ codex Luna max effort is accepted^{"rules":[{"when":"big feature","use":{"harnes codex unsupported model max effort is flagged^{"rules":[{"when":"big feature","use":{"harness":"codex","model":"gpt-5","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: codex:max unsupported grok max effort is flagged^{"rules":[{"when":"deep current work","use":{"harness":"grok","model":"grok-4","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: grok:max unsupported grok xhigh effort is flagged^{"rules":[{"when":"deep current work","use":{"harness":"grok","model":"grok-4","effort":"xhigh"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: grok:xhigh -native pi ultra is accepted^{"rules":[],"default":{"harness":"pi","model":"codex-native/gpt-6-astra","effort":"ultra"}}^empty^ -native signed pi ultra is accepted^{"rules":[{"when":"native reasoning","use":{"harness":"pi-signed","model":"codex-native/gpt-6-astra","effort":"ultra"}}]}^empty^ -ordinary pi ultra is refused^{"default":{"harness":"pi","model":"openai-codex/gpt-6-astra","effort":"ultra"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: pi:ultra -missing native model ultra is refused^{"default":{"harness":"pi","effort":"ultra"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: pi:ultra -empty native model ultra is refused^{"default":{"harness":"pi","model":"codex-native/","effort":"ultra"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: pi:ultra +native pi ultra is accepted^{"rules":[],"default":{"harness":"pi","model":"codex-native/gpt-6-astra","effort":"ultra","provider":"codex"}}^empty^ +native signed pi ultra is accepted^{"rules":[{"when":"native reasoning","use":{"harness":"pi-signed","model":"codex-native/gpt-6-astra","effort":"ultra","provider":"codex"}}]}^empty^ +ordinary pi ultra is refused^{"default":{"harness":"pi","model":"openai-codex/gpt-6-astra","effort":"ultra","provider":"codex"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: pi:ultra +missing native model ultra is refused^{"default":{"harness":"pi","effort":"ultra","provider":"codex"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: pi:ultra +empty native model ultra is refused^{"default":{"harness":"pi","model":"codex-native/","effort":"ultra","provider":"codex"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: pi:ultra codex harness ultra is refused^{"default":{"harness":"codex","model":"codex-native/gpt-6-astra","effort":"ultra"}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: codex:ultra -pi max effort is accepted^{"rules":[{"when":"deep coding","use":{"harness":"pi","model":"openai-codex/gpt-5.6-sol","effort":"max"}}]}^empty^ -pi-signed max effort is accepted^{"rules":[{"when":"signed coding","use":{"harness":"pi-signed","model":"openai-codex/gpt-5.6-sol","effort":"max"}}]}^empty^ +pi max effort is accepted^{"rules":[{"when":"deep coding","use":{"harness":"pi","model":"openai-codex/gpt-5.6-sol","effort":"max","provider":"codex"}}]}^empty^ +pi-signed max effort is accepted^{"rules":[{"when":"signed coding","use":{"harness":"pi-signed","model":"openai-codex/gpt-5.6-sol","effort":"max","provider":"codex"}}]}^empty^ muse shared efforts are accepted^{"rules":[{"when":"muse low","use":{"harness":"muse","effort":"low"}},{"when":"muse medium","use":{"harness":"muse","effort":"medium"}},{"when":"muse high","use":{"harness":"muse","effort":"high"}},{"when":"muse xhigh","use":{"harness":"muse","effort":"xhigh"}},{"when":"muse max","use":{"harness":"muse","effort":"max"}}]}^empty^ unsupported muse ultra effort is flagged^{"rules":[{"when":"muse ultra","use":{"harness":"muse","effort":"ultra"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: muse:ultra agy model profile is accepted^{"rules":[{"when":"agy work","use":{"harness":"agy","model":"gemini-3.8-flash-high"}}]}^empty^ +gemini profile with explicit provider is accepted^{"rules":[{"when":"gemini work","use":{"harness":"gemini","model":"gemini-3.8-flash-high","provider":"google"}}]}^empty^ agy low medium high efforts are accepted^{"rules":[{"when":"agy low","use":{"harness":"agy","effort":"low"}},{"when":"agy medium","use":{"harness":"agy","effort":"medium"}},{"when":"agy high","use":{"harness":"agy","effort":"high"}}]}^empty^ unsupported agy xhigh effort is flagged^{"rules":[{"when":"agy xhigh","use":{"harness":"agy","effort":"xhigh"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:xhigh unsupported agy max effort is flagged^{"rules":[{"when":"agy max","use":{"harness":"agy","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:max -unsupported opencode effort is flagged^{"rules":[{"when":"opencode work","use":{"harness":"opencode","model":"anthropic/claude-sonnet-4-5","effort":"high"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: opencode:high +unsupported opencode effort is flagged^{"rules":[{"when":"opencode work","use":{"harness":"opencode","model":"anthropic/claude-sonnet-4-5","effort":"high","provider":"claude"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: opencode:high kimi model profile is accepted^{"rules":[{"when":"kimi work","use":{"harness":"kimi","model":"kimi-code/k3"}}]}^empty^ unsupported kimi effort is flagged^{"rules":[{"when":"kimi work","use":{"harness":"kimi","model":"kimi-code/k3","effort":"high"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: kimi:high cursor model profile is accepted^{"rules":[{"when":"cursor work","use":{"harness":"cursor","model":"cursor-grok-4.5-high"}}]}^empty^ @@ -1149,18 +1157,80 @@ array use with quota-balanced is accepted^{"rules":[{"when":"big feature","use": array use without select is accepted^{"rules":[{"when":"big feature","use":[{"harness":"claude"},{"harness":"codex"}]}]}^empty^ one-element array use is accepted^{"rules":[{"when":"focused feature","use":[{"harness":"claude"}]}]}^empty^ default array is accepted^{"default":[{"harness":"pi","model":"anthropic/claude-sonnet-5"},{"harness":"grok"}]}^empty^ +provider-less multi-provider profile remains accepted without opt-in^{"rules":[{"when":"cross-provider work","use":{"harness":"opencode","model":"anthropic/claude-sonnet-4-5"}}],"default":{"harness":"pi","model":"anthropic/claude-sonnet-5"}}^empty^ one-element default array is accepted^{"default":[{"harness":"codex"}]}^empty^ empty array use is flagged^{"rules":[{"when":"big feature","use":[]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - each rule needs at least one use profile array profile without harness is flagged^{"rules":[{"when":"big feature","use":[{"model":"gpt-5.5"}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - each use profile needs harness -array profile with malformed model is flagged^{"rules":[{"when":"big feature","use":[{"harness":"codex","model":5}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings when present +array profile with malformed model is flagged^{"rules":[{"when":"big feature","use":[{"harness":"codex","model":5}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present +resolve fields are accepted^{"rules":[{"when":"hard design","approval":"captain","floor":{"scope":"model:fable","min_percent":20,"provider":"claude"},"use":[{"harness":"pi","model":"openai-codex/gpt-5.6-sol","provider":"codex"},{"harness":"codex","model":"gpt-5.6-sol","floor":{"scope":"all_models","min_percent":50}}]}],"default":[{"harness":"pi","model":"kimi-code/k3","provider":"kimi","floor":{"scope":"all_models","min_percent":10}}]}^empty^ +non-captain approval is flagged^{"rules":[{"when":"hard design","approval":"firstmate","use":{"harness":"claude"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - approval must be "captain" when present +rule floor without provider is flagged^{"rules":[{"when":"hard design","floor":{"scope":"model:fable","min_percent":20},"use":{"harness":"claude"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\z +rule floor uppercase provider is flagged^{"rules":[{"when":"hard design","floor":{"scope":"model:fable","min_percent":20,"provider":"CLAUDE"},"use":{"harness":"claude"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\z +rule floor out of range is flagged^{"rules":[{"when":"hard design","floor":{"scope":"model:fable","min_percent":120,"provider":"claude"},"use":{"harness":"claude"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\z +empty profile provider is flagged^{"rules":[{"when":"images","use":[{"harness":"pi","model":"openai-codex/gpt-5.6-sol","provider":""}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present +whitespace profile provider is flagged^{"rules":[{"when":"images","use":[{"harness":"pi","model":"openai-codex/gpt-5.6-sol","provider":" claude"}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present +newline profile provider is flagged^{"rules":[{"when":"images","use":[{"harness":"pi","model":"openai-codex/gpt-5.6-sol","provider":"claude\n"}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present +profile floor without scope is flagged^{"rules":[{"when":"images","use":[{"harness":"codex","floor":{"min_percent":50}}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile floor needs scope and min_percent 0..100 +profile floor provider override is flagged^{"rules":[{"when":"images","use":{"harness":"codex","floor":{"scope":"all_models","min_percent":50,"provider":"claude"}}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - use profile floor needs scope and min_percent 0..100 unknown select is flagged^{"rules":[{"when":"big feature","use":[{"harness":"claude"},{"harness":"codex"}],"select":"mystery"}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - unknown select: mystery array profile codex max without Luna model is flagged^{"rules":[{"when":"big feature","use":[{"harness":"codex","effort":"max"}]}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: codex:max empty default array is flagged^{"default":[]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - default needs at least one profile non-object default array entry is flagged^{"default":["codex"]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - each default profile must be an object default array profile without harness is flagged^{"default":[{"model":"gpt-5.5"}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - each default profile needs harness -default array malformed effort is flagged^{"default":[{"harness":"codex","effort":3}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - default profile model and effort must be non-empty strings when present +default array malformed effort is flagged^{"default":[{"harness":"codex","effort":3}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - default profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present +default profile floor without min_percent is flagged^{"default":[{"harness":"codex","floor":{"scope":"all_models"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - default profile floor needs scope and min_percent 0..100 +default profile floor provider override is flagged^{"default":{"harness":"codex","floor":{"scope":"all_models","min_percent":50,"provider":"claude"}}}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - default profile floor needs scope and min_percent 0..100 ROWS - pass "bootstrap validates crew-dispatch.json and reports malformed or unverified configs" + + case_dir="$TMP_ROOT/dispatch-opt-in-gate" + mkdir -p "$case_dir/home/config" + printf '%s\n' manual > "$case_dir/home/config/backlog-backend" + fakebin=$(make_fake_toolchain "$case_dir") + add_real_jq "$fakebin" + + printf '%s\n' '{"rules":[{"when":"legacy malformed model","use":{"harness":"codex","model":5}}]}' > "$case_dir/home/config/crew-dispatch.json" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ "$out" = 'CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings when present' ] \ + || fail "no-key use-profile diagnostic changed from main, got: $out" + + printf '%s\n' '{"default":{"harness":"codex","effort":3}}' > "$case_dir/home/config/crew-dispatch.json" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ "$out" = 'CREW_DISPATCH: invalid config/crew-dispatch.json - default profile model and effort must be non-empty strings when present' ] \ + || fail "no-key default-profile diagnostic changed from main, got: $out" + + printf '%s\n' '{"rules":[{"when":"legacy metadata","approval":"firstmate","floor":{"scope":"all_models","min_percent":200,"provider":"CLAUDE"},"use":{"harness":"claude","provider":"Anthropic","floor":{"scope":"all_models"}}}]}' > "$case_dir/home/config/crew-dispatch.json" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ -z "$out" ] || fail "resolver-only fields must be ignored without the typed key, got: $out" + printf '%s\n' 'TYPESAFE_API_KEY=test-key' > "$case_dir/home/.env" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ "$out" = 'CREW_DISPATCH: invalid config/crew-dispatch.json - use profile model and effort must be non-empty strings, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present' ] \ + || fail "typed .env key must activate resolver-field validation, got: $out" + + rm -f "$case_dir/home/.env" + printf '%s\n' '{"rules":[{"when":"gemini work","use":{"harness":"gemini","model":"gemini-3.8-flash-high","provider":"google"}}]}' > "$case_dir/home/config/crew-dispatch.json" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ "$out" = 'CREW_DISPATCH: invalid config/crew-dispatch.json - unverified harness: gemini' ] \ + || fail "no-key bootstrap must preserve its former verified-harness baseline, got: $out" + printf '%s\n' 'TYPESAFE_API_KEY=test-key' > "$case_dir/home/.env" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ -z "$out" ] || fail "typed resolution should add verified Gemini crewmate routing, got: $out" + + rm -f "$case_dir/home/.env" + : > "$case_dir/child-env.log" + out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + TYPESAFE_API_KEY=test-key FM_TEST_CHILD_ENV_LOG="$case_dir/child-env.log" \ + FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") + [ -z "$out" ] || fail "environment-key validation should remain silent, got: $out" + child_env=$(cat "$case_dir/child-env.log") + [ -n "$child_env" ] || fail "bootstrap child environment probe did not run" + assert_not_contains "$child_env" 'secret-present' "bootstrap children never inherit the typesafe key" + pass "bootstrap gates resolver fields and additive harnesses on the typed key" } test_bootstrap_reporting diff --git a/tests/fm-dispatch-resolve.test.sh b/tests/fm-dispatch-resolve.test.sh new file mode 100755 index 00000000000..0c4c28c71ad --- /dev/null +++ b/tests/fm-dispatch-resolve.test.sh @@ -0,0 +1,638 @@ +#!/usr/bin/env bash +# Behavior tests for bin/fm-dispatch-resolve.sh. +# +# Drives the public argv and environment interface with a fake curl on PATH +# that records argv, the request body it read from stdin, and the header it +# read from file descriptor 3, and answers with a canned typesafe.ai response. +# A fake quota-axi serves the selected schema-5 fixture. No case touches the +# network, and the absent-key case proves the tool makes no call +# at all. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +TOOL="$ROOT/bin/fm-dispatch-resolve.sh" +TMP_ROOT=$(fm_test_tmproot fm-dispatch-resolve) +HOME_DIR="$TMP_ROOT/home" +FAKEBIN=$(fm_fakebin "$TMP_ROOT") +NO_CURL_BIN="$TMP_ROOT/no-curl-bin" +LOG="$TMP_ROOT/log" +BRIEF="$TMP_ROOT/brief.md" +BASE_RULES="$TMP_ROOT/rules.json" +RULES="$HOME_DIR/config/crew-dispatch.json" +QUOTA="$TMP_ROOT/quota.json" +BASE_PATH=$PATH +mkdir -p "$HOME_DIR/config" "$LOG" "$NO_CURL_BIN" +for command_name in bash chmod cp dirname jq mktemp rm; do + ln -s "$(command -v "$command_name")" "$NO_CURL_BIN/$command_name" +done + +cat > "$BRIEF" <<'MD' +# Task +Fix the off-by-one in the pager: root cause is the `<=` on line 40 of pager.sh, expected behavior is one page per call. +MD + +cat > "$BASE_RULES" <<'JSON' +{ + "rules": [ + { + "when": "New feature work on the app.", + "floor": { "scope": "model:fable", "min_percent": 20, "provider": "claude" }, + "use": { "harness": "claude", "model": "fable", "effort": "xhigh" }, + "why": "SECRET-WHY-TEXT feature work wants the strongest model" + }, + { + "when": "The task generates images.", + "use": [ + { "harness": "pi", "model": "openai-codex/gpt-5.6-sol", "provider": "codex" }, + { "harness": "codex", "model": "gpt-5.6-sol", "floor": { "scope": "all_models", "min_percent": 50 } } + ] + }, + { + "when": "Genuinely very difficult design or planning work.", + "approval": "captain", + "use": { "harness": "claude", "model": "fable", "effort": "xhigh" } + }, + { + "when": "A simple bug fix with a stated root cause.", + "use": [ + { "harness": "claude", "model": "sonnet", "effort": "high" }, + { "harness": "cursor", "model": "cursor-grok-4.6-medium" }, + { "harness": "kimi", "model": "kimi-code/k3" } + ] + } + ], + "default": [ + { "harness": "claude", "model": "opus" }, + { "harness": "cursor", "model": "cursor-grok-4.6-high" } + ] +} +JSON +cp "$BASE_RULES" "$RULES" + +write_quota() { # [] + local path=$1 cursor=$2 claude=${3:--0.4627} + cat > "$path" < + cat > "$1" < "$FAKEBIN/curl" <<'SH' +#!/usr/bin/env bash +# Fake curl: records argv (minus the -o target), the stdin body, and the header +# read from fd 3, then answers with FAKE_CURL_RESPONSE and FAKE_CURL_HTTP. +set -u +if [ -n "${TYPESAFE_API_KEY+x}" ] || [ -n "${TYPESAFE_API_KEY_PRIVATE+x}" ]; then + printf 'curl:secret-present\n' >> "${CHILD_ENV_LOG:?}" +else + printf 'curl:clean\n' >> "${CHILD_ENV_LOG:?}" +fi +out='' +while [ $# -gt 0 ]; do + case "$1" in + -o) out=$2; shift 2 ;; + *) printf '%s\n' "$1" >> "${FAKE_CURL_LOG:?}/argv"; shift ;; + esac +done +cat > "$FAKE_CURL_LOG/body" +cat /dev/fd/3 > "$FAKE_CURL_LOG/header" 2>/dev/null || printf 'fd3 unreadable\n' > "$FAKE_CURL_LOG/header" +if [ -n "${FAKE_CURL_MUTATE_SOURCE:-}" ]; then + cp "$FAKE_CURL_MUTATE_SOURCE" "${FAKE_CURL_MUTATE_TARGET:?}" +fi +if [ "${FAKE_CURL_FAIL:-0}" = 1 ]; then + exit 7 +fi +cp "${FAKE_CURL_RESPONSE:?}" "$out" +printf '%s' "${FAKE_CURL_HTTP:-200}" +SH +chmod +x "$FAKEBIN/curl" + +cat > "$FAKEBIN/quota-axi" <<'SH' +#!/usr/bin/env bash +set -u +if [ -n "${TYPESAFE_API_KEY+x}" ] || [ -n "${TYPESAFE_API_KEY_PRIVATE+x}" ]; then + printf 'quota-axi:secret-present\n' >> "${CHILD_ENV_LOG:?}" +else + printf 'quota-axi:clean\n' >> "${CHILD_ENV_LOG:?}" +fi +printf '%s\n' "$*" >> "${QUOTA_AXI_CALLS:?}" +[ "${FAKE_QUOTA_FAIL:-0}" = 1 ] && exit 1 +[ "${1:-}" = --json ] || exit 2 +cat "${QUOTA_AXI_FIXTURE:?}" +SH +chmod +x "$FAKEBIN/quota-axi" + +RESPONSE="$TMP_ROOT/response.json" +export FAKE_CURL_LOG="$LOG" FAKE_CURL_RESPONSE="$RESPONSE" QUOTA_AXI_CALLS="$LOG/quota-axi.calls" QUOTA_AXI_FIXTURE="$QUOTA" CHILD_ENV_LOG="$LOG/child-env" + +reset_log() { + rm -rf "$LOG" + mkdir -p "$LOG" +} + +# run [args...]: the tool with fakebin first on +# PATH and an isolated FM_HOME; TYPESAFE_API_KEY comes from the caller's env. +run() { + local __exit=$1 __out=$2 __err=$3 _out _code + shift 3 + _out=$(PATH="$FAKEBIN:$BASE_PATH" FM_HOME="$HOME_DIR" "$TOOL" "$@" 2> "$TMP_ROOT/stderr") + _code=$? + printf -v "$__exit" '%s' "$_code" + printf -v "$__out" '%s' "$_out" + printf -v "$__err" '%s' "$(cat "$TMP_ROOT/stderr")" +} + +run_without_curl() { + local __exit=$1 __out=$2 __err=$3 _out _code + shift 3 + _out=$(PATH="$NO_CURL_BIN" FM_HOME="$HOME_DIR" TYPESAFE_API_KEY="$KEY" "$TOOL" "$@" 2> "$TMP_ROOT/stderr") + _code=$? + printf -v "$__exit" '%s' "$_code" + printf -v "$__out" '%s' "$_out" + printf -v "$__err" '%s' "$(cat "$TMP_ROOT/stderr")" +} + +KEY='test-key-9f1c2d3e-never-on-argv' +code='' out='' err='' + +# --- absent key: off, silent on stdout, no network, no quota read ----------- +reset_log +write_response "$RESPONSE" rule_4 0.9 +run code out err "$BRIEF" --project pager +expect_code 0 "$code" "absent key exits 0" +assert_equals '' "$out" "absent key prints nothing on stdout" +assert_contains "$err" 'dispatch-resolve: off (TYPESAFE_API_KEY absent from the environment and' "absent key explains itself on stderr" +assert_absent "$LOG/argv" "absent key never calls curl" +assert_absent "$LOG/quota-axi.calls" "absent key never reads quota-axi" +pass "absent key is off: one stderr line, exit 0, no network call" + +# --- .env key, and the environment wins over it ------------------------------ +printf '%s\n' '# local secrets' 'FMX_PAIRING_TOKEN=abc' "export TYPESAFE_API_KEY=\"$KEY\"" > "$HOME_DIR/.env" +reset_log +run code out err "$BRIEF" --project pager +expect_code 0 "$code" ".env key resolves" +assert_contains "$out" ' status: clear' ".env key produces a clear result" +assert_contains "$(cat "$LOG/header")" "Authorization: Bearer $KEY" ".env key reaches curl on the fd header" +reset_log +TYPESAFE_API_KEY=env-wins run code out err "$BRIEF" --project pager +assert_equals 'Authorization: Bearer env-wins' "$(cat "$LOG/header")" "environment key wins over .env" +rm -f "$HOME_DIR/.env" +OVERRIDE_CONFIG="$TMP_ROOT/override-config" +mkdir -p "$OVERRIDE_CONFIG" +cp "$BASE_RULES" "$OVERRIDE_CONFIG/crew-dispatch.json" +reset_log +TYPESAFE_API_KEY=$KEY FM_CONFIG_OVERRIDE="$OVERRIDE_CONFIG" run code out err "$BRIEF" --project pager +assert_contains "$out" ' status: clear' "FM_CONFIG_OVERRIDE selects the canonical rules directory" +pass "TYPESAFE_API_KEY= in .env activates the tool; environment and config overrides work" + +# --- clear: request shape, secret handling, argmax -------------------------- +reset_log +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" --project pager +expect_code 0 "$code" "clear exits 0" +assert_contains "$out" 'dispatch-resolve:' "TOON block header" +assert_contains "$out" ' status: clear' "clear status" +assert_contains "$out" ' rule: rule_4 (A simple bug fix with a stated root cause.) confidence: 0.9' "rule and confidence line" +assert_contains "$out" " profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'" "argmax picks the highest spendPriority" +assert_contains "$out" 'candidate: claude:sonnet provider=claude scope=all_models remaining=79% spendPriority=-0.4627 runway=projected_exhaustion -> eligible' "every candidate is accounted for" +assert_contains "$out" 'candidate: kimi:kimi-code/k3 provider=kimi -> eligible, unranked: provider kimi unmeasured (unknown): disclosed uncertainty' "unmeasured provider stays listed as eligible and unranked" +assert_contains "$out" ' note: 1 eligible candidate(s) unranked (kimi)' "clear results flag eligible unranked candidates once" +assert_not_contains "$out" '--effort' "cursor profile without effort emits no --effort" +argv=$(cat "$LOG/argv") +assert_not_contains "$argv" "$KEY" "the key never appears on curl argv" +assert_contains "$argv" 'https://api.typesafe.ai/v1/systemone' "the request uses the fixed typesafe.ai endpoint" +assert_contains "$argv" $'--max-time\n5' "the request uses the fixed five-second timeout" +assert_contains "$argv" '@/dev/fd/3' "the header is read from a file descriptor" +assert_equals "Authorization: Bearer $KEY" "$(cat "$LOG/header")" "curl receives the bearer header on fd 3" +assert_equals $'curl:clean\nquota-axi:clean' "$(cat "$LOG/child-env")" "the API key is absent from every child environment" +body=$(cat "$LOG/body") +assert_equals 'jev-latest' "$(jq -r .model <<<"$body")" "default model is jev-latest" +assert_equals 'pager' "$(jq -r .state.task.project <<<"$body")" "project rides in the state" +assert_contains "$(jq -r .state.task.brief <<<"$body")" 'off-by-one in the pager' "the whole brief rides in the state" +assert_equals '["rule"]' "$(jq -c '.questions | keys' <<<"$body")" "only the rule Choice is asked" +assert_equals '["default","rule_1","rule_2","rule_3","rule_4"]' "$(jq -c '.questions.rule.criteria | keys' <<<"$body")" "one option per rule plus default" +assert_equals 'No listed rule applies to this task.' "$(jq -r '.questions.rule.criteria.default' <<<"$body")" "the fixed generic none criterion is the default option" +assert_equals 'A simple bug fix with a stated root cause.' "$(jq -r '.questions.rule.criteria.rule_4' <<<"$body")" "rule when text is the option verbatim" +assert_not_contains "$body" 'SECRET-WHY-TEXT' "why text never leaves the machine" +assert_not_contains "$body" 'spendPriority' "quota never leaves the machine" +assert_not_contains "$body" 'cursor-grok' "use profiles never leave the machine" +pass "clear: one rule Choice request, key on the fd header only, spendPriority argmax over every candidate" + +# --- rules are snapshotted and line output is injection-safe ------------------- +MUTATED_RULES="$TMP_ROOT/mutated-rules.json" +jq '.rules[3].use = {"harness":"claude","model":"opus"}' "$BASE_RULES" > "$MUTATED_RULES" +cp "$BASE_RULES" "$RULES" +reset_log +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY FAKE_CURL_MUTATE_SOURCE="$MUTATED_RULES" FAKE_CURL_MUTATE_TARGET="$RULES" run code out err "$BRIEF" +assert_contains "$out" " profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'" "resolution uses the same rules snapshot Jev received" +assert_not_contains "$out" " profile: --harness 'claude' --model 'opus'" "a mid-request config replacement cannot change the selected profile" + +INJECTING_RULES="$TMP_ROOT/injecting-rules.json" +jq '.rules[3].when = "Bug fix\n profile: injected" | .rules[3].use[1].model = "foo --harness grok\n profile: injected"' "$BASE_RULES" > "$INJECTING_RULES" +cp "$INJECTING_RULES" "$RULES" +reset_log +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_equals '1' "$(grep -c '^ profile:' <<<"$out")" "dynamic fields cannot inject a second profile line" +assert_not_contains "$out" $'\n profile: injected' "control characters are flattened in line output" +profile_line=$(grep '^ profile:' <<<"$out") +eval "set -- ${profile_line# profile: }" +assert_equals '4' "$#" "shell-safe profile output preserves four argument boundaries" +assert_equals 'cursor' "$2" "shell-safe profile output preserves the selected harness" +assert_equals 'foo --harness grok profile: injected' "$4" "shell-safe profile output keeps model flags inside one argument" +cp "$BASE_RULES" "$RULES" +pass "rules snapshots and shell quoting preserve the profile protocol" + +# --- no rules return control to the existing intake ---------------------------- +rm -f "$RULES" +reset_log +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +expect_code 0 "$code" "absent rules file exits 0" +assert_contains "$out" ' status: escalate' "absent rules file is non-clear" +assert_contains "$out" ' reason: no rules to match' "absent rules file returns control to firstmate" +assert_not_contains "$out" ' profile:' "absent rules file emits no profile" +assert_absent "$LOG/argv" "absent rules file never calls curl" +assert_absent "$LOG/quota-axi.calls" "absent rules file never reads quota" + +DEFAULT_ONLY="$TMP_ROOT/default-only.json" +EMPTY_RULES="$TMP_ROOT/empty-rules.json" +printf '%s\n' '{"default":[{"harness":"claude","model":"opus"},{"harness":"cursor","model":"cursor-grok-4.6-high"}]}' > "$DEFAULT_ONLY" +printf '%s\n' '{"rules":[],"default":[{"harness":"claude","model":"opus"},{"harness":"cursor","model":"cursor-grok-4.6-high"}]}' > "$EMPTY_RULES" +for direct_rules in "$DEFAULT_ONLY" "$EMPTY_RULES"; do + cp "$direct_rules" "$RULES" + reset_log + TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" + expect_code 0 "$code" "no-rule resolution exits 0: $direct_rules" + assert_contains "$out" ' status: escalate' "no-rule resolution is non-clear: $direct_rules" + assert_contains "$out" ' reason: no rules to match' "no-rule resolution returns control to firstmate: $direct_rules" + assert_not_contains "$out" ' profile:' "no-rule resolution emits no profile: $direct_rules" + assert_absent "$LOG/argv" "no-rule resolution never calls curl: $direct_rules" + assert_absent "$LOG/quota-axi.calls" "no-rule resolution never reads quota: $direct_rules" +done + +AGY_RULE="$TMP_ROOT/agy-rule.json" +printf '%s\n' '{"rules":[{"when":"Agy work.","use":{"harness":"agy"}}]}' > "$AGY_RULE" +cp "$AGY_RULE" "$RULES" +cat > "$RESPONSE" <<'JSON' +{"model":"jev-1.13.0","answers":{"rule":{"type":"choice","choice":"rule_1","confidence":0.99,"probabilities":{"rule_1":0.99,"default":0.01}}},"usage":{"input_tokens":100,"output_tokens":60}} +JSON +reset_log +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" 'candidate: agy:- provider=agy scope=all_models remaining=64% spendPriority=0.4 runway=through_reset -> eligible' "agy uses its resolver-only authoritative quota provider" +assert_contains "$out" " profile: --harness 'agy'" "provider-less agy rule resolves" + +GEMINI_RULE="$TMP_ROOT/gemini-rule.json" +printf '%s\n' '{"rules":[{"when":"Gemini work.","use":{"harness":"gemini","model":"gemini-3.8-flash-high","provider":"google"}}]}' > "$GEMINI_RULE" +cp "$GEMINI_RULE" "$RULES" +reset_log +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" 'candidate: gemini:gemini-3.8-flash-high provider=google scope=all_models remaining=72% spendPriority=0.3 runway=through_reset -> eligible' "Gemini resolves through its explicit provider" +assert_contains "$out" " profile: --harness 'gemini' --model 'gemini-3.8-flash-high'" "Gemini is a typed verified dispatch harness" + +cp "$ROOT/docs/examples/crew-dispatch.json" "$RULES" +cat > "$RESPONSE" <<'JSON' +{"model":"jev-1.13.0","answers":{"rule":{"type":"choice","choice":"default","confidence":0.9,"probabilities":{"rule_1":0.02,"rule_2":0.02,"rule_3":0.02,"default":0.94}}},"usage":{"input_tokens":812,"output_tokens":60}} +JSON +reset_log +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: clear' "the documented example passes opted-in resolution" +assert_contains "$out" 'candidate: pi:anthropic/claude-sonnet-5 provider=claude' "the documented Pi default uses its declared Claude provider" +assert_not_contains "$err" 'malformed rules file' "the documented example reaches resolution" +cp "$BASE_RULES" "$RULES" +pass "no-rule fallback, Agy, Gemini, and documented configurations resolve" + +# --- ambiguous: fixed confidence floor ----------------------------------------- +reset_log +write_response "$RESPONSE" rule_4 0.41 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +expect_code 0 "$code" "ambiguous exits 0" +assert_contains "$out" ' status: ambiguous' "below the floor is ambiguous" +assert_contains "$out" ' reason: confidence 0.41 below floor 0.6' "ambiguous names the floor" +assert_contains "$out" 'candidate: claude:sonnet provider=claude scope=all_models remaining=79% spendPriority=-0.4627 runway=projected_exhaustion -> eligible' "ambiguous preserves matched candidate evidence" +assert_contains "$out" 'candidate: kimi:kimi-code/k3 provider=kimi -> eligible, unranked: provider kimi unmeasured (unknown): disclosed uncertainty' "ambiguous preserves eligible unranked candidate evidence" +assert_not_contains "$out" ' profile:' "ambiguous emits no profile line" +pass "ambiguous: confidence below the fixed floor hands the decision back" + +# --- escalate: captain approval ------------------------------------------------ +reset_log +write_response "$RESPONSE" rule_3 0.95 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +expect_code 0 "$code" "escalate exits 0" +assert_contains "$out" ' status: escalate' "approval-gated rule escalates" +assert_contains "$out" " reason: rule requires the captain's explicit approval before dispatch" "escalate names the approval gate" +assert_contains "$out" 'candidate: claude:fable provider=claude scope=model:fable remaining=15% spendPriority=-0.79 runway=projected_exhaustion bounds=all_models:79%/projected_exhaustion,model:fable:15%/projected_exhaustion -> eligible' "approval escalation preserves matched candidate evidence" +assert_not_contains "$out" ' profile:' "escalate emits no profile line" +pass "escalate: a rule declared approval: captain never yields a profile" + +# --- rule floor fails: fall through to default ------------------------------- +reset_log +write_response "$RESPONSE" rule_1 0.97 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: clear' "rule floor fall-through still resolves" +assert_contains "$out" ' note: rule rule_1 floor model:fable below 20%: fall through to default' "rule floor fall-through is explained" +assert_contains "$out" " profile: --harness 'cursor' --model 'cursor-grok-4.6-high'" "fall-through resolves among the default profiles" +assert_not_contains "$out" 'candidate: claude:fable' "the floored rule's own profile is not a candidate" + +MISSING_RULE_FLOOR="$TMP_ROOT/missing-rule-floor.json" +jq '(.providers[] | select(.provider == "claude") | .quotaSemantics.effectiveAvailability) |= map(select(.scope != "model:fable"))' "$QUOTA" > "$MISSING_RULE_FLOOR" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$MISSING_RULE_FLOOR" run code out err "$BRIEF" +assert_contains "$out" ' status: escalate' "an unverifiable rule floor escalates" +assert_contains "$out" ' reason: rule rule_1 floor claude/model:fable is unverifiable' "the unverifiable rule floor names its provider and scope" +assert_not_contains "$out" ' profile:' "an unverifiable rule floor never authorizes default routing" +pass "rule floor: known shortfall falls through while unavailable evidence escalates" + +# --- declared provider and profile floor -------------------------------------- +reset_log +write_response "$RESPONSE" rule_2 0.99 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" 'candidate: pi:openai-codex/gpt-5.6-sol provider=codex scope=all_models remaining=31%' "declared provider routes a Pi profile to the codex row" +assert_contains "$out" 'candidate: codex:gpt-5.6-sol provider=codex scope=all_models remaining=31% spendPriority=- runway=projected_exhaustion -> not eligible: profile floor all_models below 50%' "profile floor makes a candidate ineligible with its reason" +assert_contains "$out" " profile: --harness 'pi' --model 'openai-codex/gpt-5.6-sol'" "the remaining eligible candidate wins" + +FLOOR_BOUNDS="$TMP_ROOT/floor-bounds.json" +jq '(.providers[] | select(.provider == "codex") | .quotaSemantics.effectiveAvailability) += [ + {"scope":"model:gpt-5.6-sol","status":"known","effectivePercentRemaining":10,"runway":{"status":"projected_exhaustion"},"selection":{"spendPriority":-0.9}} +]' "$QUOTA" > "$FLOOR_BOUNDS" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$FLOOR_BOUNDS" run code out err "$BRIEF" +assert_contains "$out" 'candidate: codex:gpt-5.6-sol provider=codex scope=all_models remaining=31% spendPriority=- runway=projected_exhaustion bounds=all_models:31%/projected_exhaustion,model:gpt-5.6-sol:10%/projected_exhaustion -> not eligible: profile floor all_models below 50%' "a failed profile floor reports its named row while retaining all bounds" + +FLOOR_WITH_UNKNOWN="$TMP_ROOT/floor-with-unknown.json" +jq '(.providers[] | select(.provider == "codex") | .quotaSemantics) |= (.status = "partial" | .effectiveAvailability += [ + {"scope":"model:gpt-5.6-sol","status":"unknown","runway":{"status":"unknown"}} +])' "$QUOTA" > "$FLOOR_WITH_UNKNOWN" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$FLOOR_WITH_UNKNOWN" run code out err "$BRIEF" +assert_contains "$out" 'candidate: codex:gpt-5.6-sol provider=codex scope=all_models remaining=31% spendPriority=- runway=projected_exhaustion bounds=all_models:31%/projected_exhaustion,model:gpt-5.6-sol:-%/unknown -> not eligible: profile floor all_models below 50%' "a known profile-floor shortfall wins over unrelated unknown model evidence" + +MISSING_PROFILE_FLOOR_RULES="$TMP_ROOT/missing-profile-floor-rules.json" +jq '.rules[1].use[1].floor.scope = "model:missing"' "$BASE_RULES" > "$MISSING_PROFILE_FLOOR_RULES" +cp "$MISSING_PROFILE_FLOOR_RULES" "$RULES" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" 'candidate: codex:gpt-5.6-sol provider=codex scope=model:missing remaining=-% spendPriority=- runway=- -> eligible, unranked: profile floor model:missing is unverifiable: not rankable: disclosed uncertainty' "a missing profile floor remains eligible but unranked" +assert_not_contains "$out" 'profile floor model:missing below' "missing profile evidence is not described as a shortfall" +assert_contains "$out" " profile: --harness 'pi' --model 'openai-codex/gpt-5.6-sol'" "another candidate may clear without misrepresenting missing floor evidence" +cp "$BASE_RULES" "$RULES" +pass "declared provider and profile floor evidence are applied in code" + +# --- malformed ranking evidence is never ordered ------------------------------- +reset_log +NONNUMERIC="$TMP_ROOT/nonnumeric-spend-priority.json" +jq '(.providers[] | select(.provider == "cursor") | .quotaSemantics.effectiveAvailability[] | select(.scope == "all_models") | .selection.spendPriority) = "high"' "$QUOTA" > "$NONNUMERIC" +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$NONNUMERIC" run code out err "$BRIEF" +assert_contains "$out" 'candidate: cursor:cursor-grok-4.6-medium provider=cursor scope=all_models remaining=91% spendPriority=- runway=through_reset -> eligible, unranked: spendPriority missing or non-numeric at all_models: not rankable: disclosed uncertainty' "a nonnumeric spendPriority remains eligible but unranked" +assert_contains "$out" " profile: --harness 'claude' --model 'sonnet' --effort 'high'" "numeric evidence wins without mixed-type ordering" +pass "nonnumeric spendPriority evidence is never ranked" + +# --- partial providers retain their known row evidence -------------------------- +reset_log +PARTIAL="$TMP_ROOT/partial.json" +jq '(.providers[] | select(.provider == "cursor") | .quotaSemantics.status) = "partial"' "$QUOTA" > "$PARTIAL" +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$PARTIAL" run code out err "$BRIEF" +assert_contains "$out" 'candidate: cursor:cursor-grok-4.6-medium provider=cursor scope=all_models remaining=91% spendPriority=0.7597 runway=through_reset -> eligible' "a known row from a partial provider remains rankable" +assert_contains "$out" " profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'" "partial provider evidence can win the argmax" + +PARTIAL_UNKNOWN="$TMP_ROOT/partial-unknown.json" +jq '(.providers[] | select(.provider == "cursor") | .quotaSemantics) |= (.status = "partial" | .effectiveAvailability += [ + {"scope":"model:cursor-grok-4.6-medium","status":"unknown","runway":{"status":"unknown"}} +])' "$QUOTA" > "$PARTIAL_UNKNOWN" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$PARTIAL_UNKNOWN" run code out err "$BRIEF" +assert_contains "$out" 'candidate: cursor:cursor-grok-4.6-medium provider=cursor scope=model:cursor-grok-4.6-medium remaining=-% spendPriority=- runway=- bounds=all_models:91%/through_reset,model:cursor-grok-4.6-medium:-%/unknown -> eligible, unranked: quota row model:cursor-grok-4.6-medium unknown: not rankable: disclosed uncertainty' "an unknown exact-model row preserves partial known evidence without ranking" +assert_contains "$out" ' note: 2 eligible candidate(s) unranked (cursor, kimi)' "clear result lists every provider with unranked uncertainty" +assert_contains "$out" " profile: --harness 'claude' --model 'sonnet' --effort 'high'" "another measured candidate can clear" + +PARTIAL_EXHAUSTED="$TMP_ROOT/partial-exhausted.json" +jq '(.providers[] | select(.provider == "cursor") | .quotaSemantics) |= (.status = "partial" | .effectiveAvailability += [ + {"scope":"model:cursor-grok-4.6-medium","status":"unknown","runway":{"status":"unknown"}} +] | .effectiveAvailability[] |= if .scope == "all_models" then .effectivePercentRemaining = 0 | .runway.status = "exhausted_now" else . end)' "$QUOTA" > "$PARTIAL_EXHAUSTED" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$PARTIAL_EXHAUSTED" run code out err "$BRIEF" +assert_contains "$out" 'candidate: cursor:cursor-grok-4.6-medium provider=cursor scope=all_models remaining=0% spendPriority=- runway=exhausted_now bounds=all_models:0%/exhausted_now,model:cursor-grok-4.6-medium:-%/unknown -> not eligible: runway exhausted_now at all_models' "known exhaustion vetoes a candidate despite unknown exact-model evidence" +assert_contains "$out" ' note: 1 eligible candidate(s) unranked (kimi)' "an exhausted candidate is excluded from the unranked uncertainty note" + +UNKNOWN_EXHAUSTED="$TMP_ROOT/unknown-exhausted.json" +jq '(.providers[] | select(.provider == "cursor") | .quotaSemantics) = { + "status":"unknown","effectiveAvailability":[ + {"scope":"all_models","status":"unknown","runway":{"status":"exhausted_now"}} + ] +}' "$QUOTA" > "$UNKNOWN_EXHAUSTED" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$UNKNOWN_EXHAUSTED" run code out err "$BRIEF" +assert_contains "$out" 'candidate: cursor:cursor-grok-4.6-medium provider=cursor scope=all_models remaining=-% spendPriority=- runway=exhausted_now -> not eligible: runway exhausted_now at all_models' "unknown provider semantics cannot mask concrete exhaustion" + +NO_APPLICABLE="$TMP_ROOT/no-applicable.json" +jq '(.providers[] | select(.provider == "cursor") | .quotaSemantics.effectiveAvailability) = [ + {"scope":"model:other","status":"known","effectivePercentRemaining":91,"runway":{"status":"through_reset"},"selection":{"spendPriority":0.8}} +]' "$QUOTA" > "$NO_APPLICABLE" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$NO_APPLICABLE" run code out err "$BRIEF" +assert_contains "$out" 'candidate: cursor:cursor-grok-4.6-medium provider=cursor -> eligible, unranked: no applicable quota row for provider cursor: disclosed uncertainty' "a candidate without an applicable row remains eligible but unranked" +assert_contains "$out" ' note: 2 eligible candidate(s) unranked (cursor, kimi)' "no-applicable-row uncertainty appears in the clear-result note" +pass "partial and missing quota evidence remain eligible but unranked" + +# --- provider-wide rows remain bounds beside exact model rows ------------------ +reset_log +BOUNDED="$TMP_ROOT/bounded.json" +jq '(.providers[] | select(.provider == "claude") | .quotaSemantics.effectiveAvailability) += [ + {"scope":"model:sonnet","status":"known","effectivePercentRemaining":99,"runway":{"status":"through_reset"},"selection":{"spendPriority":0.9}} +]' "$QUOTA" > "$BOUNDED" +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$BOUNDED" run code out err "$BRIEF" +assert_contains "$out" 'candidate: claude:sonnet provider=claude scope=all_models remaining=79% spendPriority=-0.4627' "the limiting provider-wide row drives ranking" +assert_contains "$out" 'bounds=all_models:79%/projected_exhaustion,model:sonnet:99%/through_reset' "all applicable quota bounds are disclosed" + +EXHAUSTED_WIDE="$TMP_ROOT/exhausted-wide.json" +jq '(.providers[] | select(.provider == "claude") | .quotaSemantics.effectiveAvailability[] | select(.scope == "all_models")) |= (.effectivePercentRemaining = 0 | .runway.status = "exhausted_now")' "$BOUNDED" > "$EXHAUSTED_WIDE" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$EXHAUSTED_WIDE" run code out err "$BRIEF" +assert_contains "$out" 'candidate: claude:sonnet provider=claude scope=all_models remaining=0%' "the exhausted account-wide bound is the candidate evidence" +assert_contains "$out" '-> not eligible: runway exhausted_now at all_models' "a healthy exact row cannot bypass an exhausted account-wide bound" +pass "provider-wide and exact quota rows combine into one limiting candidate" + +# --- default choice ------------------------------------------------------------ +reset_log +write_response "$RESPONSE" default 0.88 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' rule: default (No listed rule applies to this task.)' "default names the fixed neutral none option" +assert_contains "$out" ' note: no rule matched' "default is explained" +assert_contains "$out" " profile: --harness 'cursor' --model 'cursor-grok-4.6-high'" "default resolves by argmax" +pass "default: no rule matched resolves among the default profiles" + +# --- genuine tie escalates --------------------------------------------------------- +reset_log +TIE="$TMP_ROOT/tie.json" +write_quota "$TIE" 0.5 0.5 +write_response "$RESPONSE" default 0.88 +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$TIE" run code out err "$BRIEF" +assert_contains "$out" ' status: escalate' "tie escalates" +assert_contains "$out" ' reason: genuine spendPriority tie' "tie is named" +pass "tie: equal spendPriority never breaks by array order" + +# --- nothing rankable escalates ------------------------------------------------- +reset_log +NONE="$TMP_ROOT/none.json" +jq '.providers |= map(if .provider == "cursor" or .provider == "claude" then .quotaSemantics.effectiveAvailability |= map(.runway.status = "exhausted_now") else . end)' "$QUOTA" > "$NONE" +TYPESAFE_API_KEY=$KEY QUOTA_AXI_FIXTURE="$NONE" run code out err "$BRIEF" +assert_contains "$out" ' status: escalate' "no rankable candidate escalates" +assert_contains "$out" ' reason: no rankable eligible candidate' "no-candidate reason" +assert_contains "$out" '-> not eligible: runway exhausted_now' "exhausted candidates keep their reason" +pass "no rankable candidate: the tool escalates instead of guessing" + +# --- quota-axi is read exactly once -------------------------------------------- +reset_log +write_response "$RESPONSE" rule_4 0.9 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +expect_code 0 "$code" "quota-axi path exits 0" +assert_equals '--json' "$(cat "$LOG/quota-axi.calls")" "quota-axi --json is called exactly once" +assert_contains "$out" " profile: --harness 'cursor' --model 'cursor-grok-4.6-medium'" "quota-axi snapshot drives the argmax" +reset_log +TYPESAFE_API_KEY=$KEY FAKE_QUOTA_FAIL=1 run code out err "$BRIEF" +expect_code 0 "$code" "quota-axi failure exits 0" +assert_contains "$out" ' status: error' "quota-axi failure is an error outcome" +assert_contains "$out" ' reason: quota-axi --json failed' "quota-axi failure is named" +pass "quota evidence comes from one quota-axi --json read, and its failure is an error outcome" + +# --- API and response failures are error outcomes, exit 0 ---------------------- +reset_log +run_without_curl code out err "$BRIEF" +expect_code 0 "$code" "missing curl exits 0" +assert_contains "$out" ' status: error' "missing curl is a structured error outcome" +assert_contains "$out" ' reason: curl not installed' "missing curl is named in the TOON block" +assert_contains "$err" 'dispatch-resolve: error (curl not installed)' "missing curl is also reported on stderr" +reset_log +TYPESAFE_API_KEY=$KEY FAKE_CURL_HTTP=429 run code out err "$BRIEF" +expect_code 0 "$code" "http 429 exits 0" +assert_contains "$out" ' status: error' "http 429 is an error outcome" +assert_contains "$out" ' reason: http 429 after' "http status is reported" +assert_contains "$err" 'dispatch-resolve: error (http 429' "error also goes to stderr" +reset_log +TYPESAFE_API_KEY=$KEY FAKE_CURL_FAIL=1 run code out err "$BRIEF" +expect_code 0 "$code" "curl failure exits 0" +assert_contains "$out" ' reason: http 000 after' "transport failure reads as http 000" +reset_log +printf '%s\n' '{"model":"jev","answers":{}}' > "$RESPONSE" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' reason: response is not a rule Choice answer' "a malformed answer is an error outcome" +reset_log +write_response "$RESPONSE" rule_4 0.9 +jq '.usage = "bad"' "$RESPONSE" > "$TMP_ROOT/malformed-usage.json" +mv "$TMP_ROOT/malformed-usage.json" "$RESPONSE" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "malformed usage is an error outcome" +assert_contains "$out" ' reason: response is not a rule Choice answer' "malformed usage cannot break text rendering silently" +reset_log +write_response "$RESPONSE" rule_4 0.9 +jq 'del(.answers.rule.probabilities.default)' "$RESPONSE" > "$TMP_ROOT/malformed-probabilities.json" +mv "$TMP_ROOT/malformed-probabilities.json" "$RESPONSE" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "missing probability choice is an error outcome" +assert_contains "$out" ' reason: response is not a rule Choice answer' "probabilities must name every offered choice" +reset_log +write_response "$RESPONSE" rule_4 0.9 +jq '.answers.rule.probabilities.rule_4 = "high"' "$RESPONSE" > "$TMP_ROOT/malformed-probabilities.json" +mv "$TMP_ROOT/malformed-probabilities.json" "$RESPONSE" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "nonnumeric probability is an error outcome" +assert_contains "$out" ' reason: response is not a rule Choice answer' "probabilities must be numeric and bounded" +reset_log +write_response "$RESPONSE" rule_4 0.9 +jq '.answers.rule.probabilities[] = 0' "$RESPONSE" > "$TMP_ROOT/malformed-probabilities.json" +mv "$TMP_ROOT/malformed-probabilities.json" "$RESPONSE" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "a zero-mass probability distribution is an error outcome" +assert_contains "$out" ' reason: response is not a rule Choice answer' "probabilities must sum to approximately one" +reset_log +write_response "$RESPONSE" rule_4 2 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "out-of-range confidence is an error outcome" +assert_contains "$out" ' reason: response is not a rule Choice answer' "out-of-range confidence is a malformed answer" +reset_log +write_response "$RESPONSE" rule_9 0.9 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "an unknown rule id is an error outcome" +assert_contains "$out" ' reason: rule rule_9 is not in the rules file' "unknown rule id is named" +write_response "$RESPONSE" rule_0 0.9 +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +assert_contains "$out" ' status: error' "rule zero is an error outcome" +assert_contains "$out" ' reason: rule rule_0 is not in the rules file' "rule zero cannot alias the final rule" +reset_log +TYPESAFE_API_KEY=$KEY FAKE_CURL_HTTP=500 run code out err "$BRIEF" +assert_contains "$out" ' status: error' "http 500 is a TOON error outcome" +pass "API, transport, and response failures are error outcomes with exit 0" + +# --- configuration errors exit 2 and select nothing ---------------------------------- +reset_log +TYPESAFE_API_KEY=$KEY run code out err +expect_code 2 "$code" "missing brief exits 2" +assert_contains "$err" 'brief file required' "missing brief is named" +rm -f "$RULES" +ln -s "$TMP_ROOT/missing-rules-target.json" "$RULES" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +expect_code 2 "$code" "broken canonical rules symlink exits 2" +assert_contains "$err" "rules file not readable: $RULES" "broken rules symlink is actionable" +rm -f "$RULES" +printf '%s\n' '{"rules":[' > "$RULES" +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" +expect_code 2 "$code" "non-JSON rules exits 2" +assert_contains "$err" 'not JSON' "non-JSON rules is named" +for bad in \ + '{"rules":[{"when":"x","use":{"harness":"claude"},"approval":"firstmate"}]}|approval must be "captain" when present' \ + '{"rules":[{"when":"x","use":{"harness":"claude"},"select":"mystery"}]}|unknown select: mystery' \ + '{"rules":[{"when":"x","use":{"harness":"claude"},"floor":{"scope":"model:fable","min_percent":20}}]}|rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\z' \ + '{"rules":[{"when":"x","use":{"harness":"claude"},"floor":{"scope":"model:fable","min_percent":20,"provider":"CLAUDE"}}]}|rule floor needs scope, min_percent 0..100, and provider matching ^[a-z0-9]+(-[a-z0-9]+)*\z' \ + '{"rules":[{"when":"x","use":{"harness":"claude","provider":""}}]}|each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present' \ + '{"rules":[{"when":"x","use":{"harness":"claude","provider":" claude"}}]}|each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present' \ + '{"rules":[{"when":"x","use":{"harness":"claude","provider":"claude\n"}}]}|each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present' \ + '{"rules":[{"when":"x","use":{"harness":"codex","floor":{"scope":"all_models","min_percent":20,"provider":"claude"}}}]}|each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present' \ + '{"rules":[{"when":"x","use":[{"harness":"codex","model":"gpt-5.5","effort":"high"},{"harness":"codex","model":"gpt-5.5","effort":"high"}]}]}|each rule use must not contain duplicate harness, model, and effort profiles' \ + '{"rules":[{"when":"x","use":{"harness":"codex"}}],"default":[{"harness":"claude","model":"opus"},{"harness":"claude","model":"opus"}]}|default must not contain duplicate harness, model, and effort profiles' \ + '{"rules":[{"when":"x","use":{"harness":"spaceship"}}]}|each use profile must name a verified harness' \ + '{"rules":[{"when":"x","use":{"harness":"grok","effort":"max"}}]}|each use profile effort must be supported by its harness and model' \ + '{"rules":[{"when":"x","use":{"harness":"opencode","model":"anthropic/claude-sonnet-4-5"}}]}|use profiles whose harness lacks one authoritative provider family require provider: opencode' \ + '{"rules":[{"when":"x","use":{"harness":"rovo"}}]}|use profiles whose harness lacks one authoritative provider family require provider: rovo' \ + '{"rules":[{"when":"x","use":{"harness":"codex"}}],"default":{"harness":"pi","model":"anthropic/claude-sonnet-5"}}|default profiles whose harness lacks one authoritative provider family require provider: pi'; do + printf '%s\n' "${bad%%|*}" > "$RULES" + TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" + expect_code 2 "$code" "malformed rules exit 2: ${bad#*|}" + assert_contains "$err" "malformed rules file: $RULES - ${bad#*|}" "malformed rules are named: ${bad#*|}" +done +assert_absent "$LOG/argv" "configuration errors never reach the network" +cp "$BASE_RULES" "$RULES" +for removed in --json --rules --quota; do + TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" "$removed" + expect_code 2 "$code" "removed option is rejected: $removed" + assert_contains "$err" "unknown flag $removed" "removed option has no public path: $removed" +done +TYPESAFE_API_KEY=$KEY run code out err "$BRIEF" --bogus +expect_code 2 "$code" "unknown flag exits 2" +run code out err --help +expect_code 0 "$code" "--help exits 0" +assert_contains "$out" 'Usage:' "--help prints usage" +pass "configuration errors exit 2 before any network call" + +printf '# all fm-dispatch-resolve tests passed\n' diff --git a/tests/fm-gotmp.test.sh b/tests/fm-gotmp.test.sh index d29c540e648..2555e85f1b5 100755 --- a/tests/fm-gotmp.test.sh +++ b/tests/fm-gotmp.test.sh @@ -76,12 +76,13 @@ SH ln -s "$ROOT/bin/fm-gate-refuse-lib.sh" "$fake/bin/fm-gate-refuse-lib.sh" # fm-pr-lib.sh: teardown uses its canonical task-ID validator for poll cleanup. ln -s "$ROOT/bin/fm-pr-lib.sh" "$fake/bin/fm-pr-lib.sh" - # fm-public-followup-lib.sh (and the fm-x-lib.sh it sources): teardown sources - # it for the relay-activation gate on the promised-public-reply check. Neither - # does anything in this fixture, which has no .env, but both are real siblings - # teardown now requires. + # fm-public-followup-lib.sh (and the fm-x-lib.sh and fm-env-lib.sh it + # sources): teardown sources it for the relay-activation gate on the + # promised-public-reply check. None does anything in this fixture, which has + # no .env, but all three are real siblings teardown now requires. ln -s "$ROOT/bin/fm-public-followup-lib.sh" "$fake/bin/fm-public-followup-lib.sh" ln -s "$ROOT/bin/fm-x-lib.sh" "$fake/bin/fm-x-lib.sh" + ln -s "$ROOT/bin/fm-env-lib.sh" "$fake/bin/fm-env-lib.sh" ln -s "$ROOT/bin/fm-secondmate-registry-lib.sh" "$fake/bin/fm-secondmate-registry-lib.sh" ln -s "$ROOT/bin/fm-secondmate-parent-lib.sh" "$fake/bin/fm-secondmate-parent-lib.sh" # Receiver-wake retirement sources the pending-reply library, which in turn @@ -176,12 +177,13 @@ SH ln -s "$ROOT/bin/fm-gate-refuse-lib.sh" "$fake/bin/fm-gate-refuse-lib.sh" # fm-pr-lib.sh: teardown uses its canonical task-ID validator for poll cleanup. ln -s "$ROOT/bin/fm-pr-lib.sh" "$fake/bin/fm-pr-lib.sh" - # fm-public-followup-lib.sh (and the fm-x-lib.sh it sources): teardown sources - # it for the relay-activation gate on the promised-public-reply check. Neither - # does anything in this fixture, which has no .env, but both are real siblings - # teardown now requires. + # fm-public-followup-lib.sh (and the fm-x-lib.sh and fm-env-lib.sh it + # sources): teardown sources it for the relay-activation gate on the + # promised-public-reply check. None does anything in this fixture, which has + # no .env, but all three are real siblings teardown now requires. ln -s "$ROOT/bin/fm-public-followup-lib.sh" "$fake/bin/fm-public-followup-lib.sh" ln -s "$ROOT/bin/fm-x-lib.sh" "$fake/bin/fm-x-lib.sh" + ln -s "$ROOT/bin/fm-env-lib.sh" "$fake/bin/fm-env-lib.sh" ln -s "$ROOT/bin/fm-secondmate-registry-lib.sh" "$fake/bin/fm-secondmate-registry-lib.sh" ln -s "$ROOT/bin/fm-secondmate-parent-lib.sh" "$fake/bin/fm-secondmate-parent-lib.sh" ln -s "$ROOT/bin/fm-pending-reply-lib.sh" "$fake/bin/fm-pending-reply-lib.sh" diff --git a/tests/fm-quota-choose.test.sh b/tests/fm-quota-choose.test.sh index 88ae72fbdbb..095e292365f 100755 --- a/tests/fm-quota-choose.test.sh +++ b/tests/fm-quota-choose.test.sh @@ -27,6 +27,7 @@ NO_APPLICABLE="$LAB/no-applicable.json" APPLICABLE_VETO="$LAB/applicable-veto.json" MUSE_EXHAUSTED="$LAB/muse-exhausted.json" MUSE_POSITIVE="$LAB/muse-positive.json" +AGY_POSITIVE="$LAB/agy-positive.json" TOON="$LAB/quota.toon" RENDERER_TOON="$LAB/renderer-quota.toon" EMPTY_TOON="$LAB/empty-quota.toon" @@ -247,10 +248,10 @@ fi [ "$err" = "error: unknown harness: bogus" ] || fail "unknown harness returned: $err" ok "unknown harness fails closed" -if err=$(call_choose --snapshot "$LAB/captured.json" --candidate claude:default --candidate agy:default 2>&1); then +if err=$(call_choose --snapshot "$LAB/captured.json" --candidate claude:default --candidate rovo:default 2>&1); then fail "trailing unsupported harness was hidden by an earlier selection" fi -[ "$err" = "error: unknown harness: agy" ] || fail "trailing unsupported harness returned: $err" +[ "$err" = "error: unknown harness: rovo" ] || fail "trailing unsupported harness returned: $err" if err=$(call_choose --snapshot "$LAB/captured.json" --candidate claude:default --candidate 'claude:' 2>&1); then fail "trailing empty model was hidden by an earlier selection" @@ -551,11 +552,13 @@ fi [ "$out" = "none" ] || fail "exhausted Meta quota returned: $out" ok "Muse uses Meta quota" -if err=$(call_choose --snapshot "$LAB/captured.json" --candidate agy:default 2>&1); then - fail "unsupported harness unexpectedly dispatched" +jq '.providers += [{"provider":"agy","windows":[],"quotaSemantics":{"status":"known","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":25,"runway":{"status":"through_reset"}}]}}]' \ + "$LAB/captured.json" > "$AGY_POSITIVE" +if err=$(call_choose --snapshot "$AGY_POSITIVE" --candidate agy:default 2>&1); then + fail "legacy quota chooser unexpectedly accepted Agy" fi -[ "$err" = "error: unknown harness: agy" ] || fail "unsupported harness returned: $err" -ok "unsupported harness is rejected" +printf '%s\n' "$err" | grep -F 'unknown harness: agy' >/dev/null || fail "legacy Agy rejection changed: $err" +ok "Agy remains resolver-only" jq '.providers += [.providers[] | select(.provider == "claude")]' "$LAB/captured.json" > "$DUPLICATE" if err=$(call_choose --snapshot "$DUPLICATE" --candidate claude:default 2>&1); then From 334fa1226d4efb9bda832be09017b8df300488f2 Mon Sep 17 00:00:00 2001 From: Tiago Date: Thu, 17 Sep 2026 03:25:51 -0300 Subject: [PATCH 02/37] fix(bin): read the latest status event so buried declarations and open decisions aren't lost (#3753) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test: reproduce buried status declarations in shared readers * fix: share status event reads and preserve open blockers * fix: retain terminal scout and ship status declarations * no-mistakes(review): Fix status chronology, legacy completions, and reader performance * no-mistakes(review): Share terminal decision reconciliation across fleet snapshots * no-mistakes(review): Unify terminal supersession across cached folds and consumers * no-mistakes(review): Filter per-key status history while preserving terminal chronology * no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells * no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses * no-mistakes(document): Document latest-event status read and kind-scoped fold cursor * no-mistakes(lint): Quote literal done in test for-lists for SC1010 * ci: expect 19 snapshot/fleet-view tests This branch adds a fleet-snapshot regression, so the stock macOS Bash lane's hardcoded guard of 18 'ok - ' lines fails on the new count. Bump the guard and its message to 19. * no-mistakes(review): Restore multiline child outcome reporting * no-mistakes(review): Select ledger terminal events through bounded shared reader * no-mistakes(review): Report newest open decision instead of preferring blocked * no-mistakes(review): Require colon before ship/scout terminal supersession in fold * no-mistakes(review): Gate socket-down override on latest event; drop lock matrix * no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions * no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold * no-mistakes(test): Update fleet-view expectations to newest-open-decision rule * no-mistakes(document): Align status-read docs with fold-resolved crew state * no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers * no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree * test: fold terminal-cleanup snapshot coverage into the completed-scout case Keep the ship/scout/secondmate supersession assertions without adding a nineteenth top-level fleet-view test, so CI can stay at the upstream suite count. * no-mistakes(document): Clarify socket-down override expiry in architecture doc * ci: retrigger flaky contribution check --- .agents/skills/fmx-respond/SKILL.md | 2 +- bin/fm-captain-hold.sh | 26 +- bin/fm-classify-lib.sh | 293 +++++++++++++----- bin/fm-crew-state.sh | 38 ++- bin/fm-fleet-snapshot.sh | 2 +- bin/fm-inactive-reconcile.sh | 34 +- bin/fm-watch.sh | 4 +- docs/architecture.md | 9 +- tests/fm-captain-hold-lifecycle.test.sh | 16 +- tests/fm-classify-decision-key.test.sh | 146 +++++++++ tests/fm-crew-state.test.sh | 193 ++++++++++++ tests/fm-fleet-snapshot-view.test.sh | 51 ++- tests/fm-inactive-reconcile.test.sh | 67 +++- tests/fm-send-resolve-key.test.sh | 13 +- ...m-wake-drain-open-decisions-cursor.test.sh | 63 ++++ tests/fm-watch-triage.test.sh | 33 +- 16 files changed, 838 insertions(+), 152 deletions(-) diff --git a/.agents/skills/fmx-respond/SKILL.md b/.agents/skills/fmx-respond/SKILL.md index 39cafc2961f..9ad57af9b04 100644 --- a/.agents/skills/fmx-respond/SKILL.md +++ b/.agents/skills/fmx-respond/SKILL.md @@ -151,7 +151,7 @@ Treat `state/x-inbox/` as the source of truth and process **every** file you fin 1. **Gather live fleet state once.** Compose answers from what this instance genuinely knows right now: - `data/backlog.md` "## In flight" - the work currently moving. - - `state/*.status` - the latest line of each in-flight job, for fresh phase detail. + - `state/*.status` - the latest status event of each in-flight job, for fresh phase detail. - `data/projects.md` - the active projects, for naming what you work on in plain terms. Translate every internal item into an outcome. Example: a backlog line `fix-login-k3 - repair OAuth redirect (repo: yourapp)` becomes "patching a sign-in redirect bug on one of the apps" - no id, no repo name unless it is already public. 2. **Drain every pending mention.** For each `state/x-inbox/*.json` file: diff --git a/bin/fm-captain-hold.sh b/bin/fm-captain-hold.sh index 8a30c89c546..facc86505cb 100755 --- a/bin/fm-captain-hold.sh +++ b/bin/fm-captain-hold.sh @@ -442,23 +442,6 @@ meta_value() { # grep "^$2=" "$1" 2>/dev/null | tail -1 | cut -d= -f2- || true } -origin_open_decisions() { # - local origin=$1 meta="$STATE/$1.meta" status_file="$STATE/$1.status" open kind last verb - open=$(status_open_decisions "$status_file") - [ -n "$open" ] || return 0 - [ -f "$meta" ] || { printf '%s' "$open"; return 0; } - kind=$(meta_value "$meta" kind) - [ -n "$kind" ] || kind=ship - if [ "$kind" != secondmate ]; then - last=$(last_status_line "$status_file") - verb=$(status_line_verb "$last") - case "$verb" in - done|failed) return 0 ;; - esac - fi - printf '%s' "$open" -} - # A resolution record written by this script or by the retired # fm-decision-hold.sh. Both carry the same leader-then-captain-decision shape. body_has_resolution_record() { # @@ -1629,7 +1612,7 @@ reconcile_note() { } command_complete() { - local origin=${1:-} meta previous='' supplied='' keys='' entry key status_file open raw_open has_meta=0 transfer_rc resolved + local origin=${1:-} meta previous='' supplied='' keys='' entry key status_file open has_meta=0 transfer_rc resolved local resolved_how attested_by_prefix='' [ "$#" -ge 2 ] || { usage >&2; exit 2; } validate_slug origin-id "$origin" @@ -1673,8 +1656,7 @@ EOF fi status_file="$STATE/$origin.status" - raw_open=$(status_open_decisions "$status_file") - open=$(origin_open_decisions "$origin") + open=$(status_open_decisions "$status_file") if [ -n "$open" ] && [ -z "$keys" ]; then fail "origin $origin still has open captain decisions in its status stream; hold a captain task for what remains, or answer them, before attesting --none" fi @@ -1700,7 +1682,7 @@ EOF "captain-held [key=$key]: tracked by $keys" || transfer_rc=$? [ "$transfer_rc" -ne 2 ] || fail "cannot append the captain-held transfer for $origin/$key" done < FM_CLASSIFY_RESOLVE_VERB_DEFAULT='resolved' FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT='captain-held' -# Return the last non-blank line of a status file (empty if missing/blank). -last_status_line() { - local f=$1 - [ -e "$f" ] || return 0 - grep -v '^[[:space:]]*$' "$f" 2>/dev/null | tail -1 +# How many trailing lines the latest-event read parses before it widens to the +# whole file. A status record and its continuation prose sit within a few lines +# of the log's end, so this bounds the watcher's per-poll read on a long-lived +# log while a log whose tail holds no event still gets a full pass. +FM_CLASSIFY_EVENT_WINDOW_LINES=200 + +# Return the last recognized status event, ignoring continuation prose and blanks +# (empty if missing/blank), and with the event before it. +# The optional previous event is what this reader returned before the latest one +# was appended, so a consumer can name the head it is superseding; asking for it +# always reads the whole file, since a bounded window cannot bound two events. +# This is an event read; status_current_line below reconciles open decisions. +last_status_line() { # [] + local f=$1 scan='' + [ -f "$f" ] && [ -r "$f" ] || return 0 + if [ "$#" -gt 1 ]; then + scan=$(_fm_status_event_scan < "$f") || : + elif ! scan=$(tail -n "$FM_CLASSIFY_EVENT_WINDOW_LINES" "$f" 2>/dev/null | _fm_status_event_scan); then + scan=$(_fm_status_event_scan < "$f") || : + fi + [ "$#" -lt 2 ] || printf -v "$2" '%s' "${scan%%$'\n'*}" + printf '%s\n' "${scan##*$'\n'}" +} + +# Print "\n" for the status lines on stdin, and +# return 1 when the stream holds no recognized event at all, so a caller reading +# a bounded window knows to widen it. A stream without events keeps its last +# nonblank line as the latest, matching the read this replaced. +# Keep decision-closing events: skipping a resolved line would revive its opener. +# A bare legacy free-text line counts as an event only when a captain token leads +# it, so continuation prose that merely mentions one cannot hide a declaration. +_fm_status_event_scan() { + local line last='' prev='' fallback='' verb legacy_re + legacy_re="^[[:space:]]*(${FM_CAPTAIN_RE:-$FM_CLASSIFY_CAPTAIN_RE_DEFAULT})" + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in *[![:space:]]*) fallback=$line ;; *) continue ;; esac + case "$line" in *:*) status_line_verb "$line" verb ;; *) verb='' ;; esac + case "$verb" in + working|needs-decision|blocked|done|failed|note|\ + "${FM_CLASSIFY_PAUSED_VERB:-$FM_CLASSIFY_PAUSED_VERB_DEFAULT}"|\ + "${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT}"|\ + "${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT}") prev=$last; last=$line ;; + *) _fm_classify_matches "$line" "$legacy_re" && { prev=$last; last=$line; } ;; + esac + done + printf '%s\n%s\n' "$prev" "${last:-$fallback}" + [ -n "$last" ] +} + +# 0 when matches the extended regex case-insensitively, leaving +# the caller's nocasematch setting untouched. +_fm_classify_matches() { # + local matched=1 restore_case=0 + shopt -q nocasematch || { shopt -s nocasematch; restore_case=1; } + [[ "$1" =~ $2 ]] && matched=0 + [ "$restore_case" -eq 0 ] || shopt -u nocasematch + return "$matched" } # 0 if the given (last) status line's leading verb is a real terminal captain verb @@ -156,8 +208,7 @@ status_is_terminal_verb() { status_is_captain_relevant() { local line=$1 verb [ -n "$line" ] || return 1 - status_is_paused "$line" && return 1 - verb=$(status_line_verb "$line") + status_line_verb "$line" verb case "$verb" in working|resolved|captain-held|"${FM_CLASSIFY_PAUSED_VERB:-$FM_CLASSIFY_PAUSED_VERB_DEFAULT}") return 1 @@ -168,7 +219,7 @@ status_is_captain_relevant() { done|needs-decision|blocked|failed) return 0 ;; esac fi - printf '%s' "$line" | grep -qiE "${FM_CAPTAIN_RE:-$FM_CLASSIFY_CAPTAIN_RE_DEFAULT}" + _fm_classify_matches "$line" "${FM_CAPTAIN_RE:-$FM_CLASSIFY_CAPTAIN_RE_DEFAULT}" } # 0 if a status line's leading verb is the pause verb (paused: ). A pure @@ -231,9 +282,10 @@ status_paused_until() { # -> epoch on stdout # after a later, unrelated event": a subsequent done/paused/working line silently # masks a still-open needs-decision. status_open_decisions is the ONE authoritative # statement of the status-fold contract that fixes this - a needs-decision/blocked -# line OPENS a keyed decision, and only an explicit resolution or a verified -# captain-held backlog transfer referencing that key CLOSES it; a later unrelated -# terminal line never clears an open captain decision. +# line OPENS a keyed decision, and an explicit resolution or a verified +# captain-held backlog transfer referencing that key CLOSES it. +# Ship/scout terminal declarations supersede stale log decisions; a secondmate's +# terminal event may describe other work and cannot close an unrelated decision. # Who WRITES the closing line is owned elsewhere: the answering firstmate closes # at answer time through fm-send's --resolve-key (bin/fm-send.sh header), and a # worker self-closes only a blocker that cleared without an answer (bin/fm-brief.sh @@ -312,7 +364,11 @@ _fm_classify_is_corr_token() { # return 1 } -status_line_verb() { # -> leading verb word +# Printed, or assigned to when one is given, so a per-line caller on a +# hot path can take the verb without forking a command substitution. Under bash's +# dynamic scope an named like one of this function's own locals (v, out, +# word) would be assigned here and lost, so callers pass a distinct name. +status_line_verb() { # [] -> leading verb word local v=${1%%:*} out='' word v=${v%%\[*} v=${v#"${v%%[![:space:]]*}"} @@ -321,23 +377,24 @@ status_line_verb() { # -> leading verb word # contain a correlation token is returned byte-for-byte as before, so every # line without one keeps its exact historical verb, spacing included. case "$v" in - *corr=*) ;; - *) printf '%s' "$v"; return 0 ;; + *corr=*) + # Retain the first word, then drop only recognised tokens from the remaining + # whole words. Anything unrecognised stays, so prose still matches no verb. + word=${v%%[[:space:]]*} + out=$word + v=${v#"$word"} + v=${v#"${v%%[![:space:]]*}"} + while [ -n "$v" ]; do + word=${v%%[[:space:]]*} + v=${v#"$word"} + v=${v#"${v%%[![:space:]]*}"} + _fm_classify_is_corr_token "$word" && continue + out="$out $word" + done + ;; + *) out=$v ;; esac - # Retain the first word, then drop only recognised tokens from the remaining - # whole words. Anything unrecognised stays, so prose still matches no verb. - word=${v%%[[:space:]]*} - out=$word - v=${v#"$word"} - v=${v#"${v%%[![:space:]]*}"} - while [ -n "$v" ]; do - word=${v%%[[:space:]]*} - v=${v#"$word"} - v=${v#"${v%%[![:space:]]*}"} - _fm_classify_is_corr_token "$word" && continue - out="$out $word" - done - printf '%s' "$out" + if [ "$#" -gt 1 ]; then printf -v "$2" '%s' "$out"; else printf '%s' "$out"; fi } # 0 when a complete "[key=...]" token sits in the documented position before # the line's first colon (or anywhere on a line that has no colon at all). @@ -465,18 +522,39 @@ _fm_is_pending_reply_escalation() { # esac } -_fm_decision_fold_line() { # - local open=$1 line=$2 resolve=$3 held=$4 verb key note - # Blank-line guard. A `case` glob answers "does this line hold any non-space - # character" in one pattern match; the equivalent ${line//[[:space:]]/} costs - # tens of milliseconds per line under bash 3.2's global bracket-class - # substitution, which is the whole per-line cost of both folds on a status log - # of ordinary width. Same verdict, bounded cost. +_fm_status_kind() { + local meta=${1%.status}.meta kind=${2:-} line + if [ -z "$kind" ]; then + [ -f "$meta" ] && [ -r "$meta" ] && [ ! -L "$meta" ] || { printf unknown; return 0; } + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in kind=*) kind=${line#kind=} ;; esac + done < "$meta" + kind=${kind:-ship} + fi + case "$kind" in ship|scout|secondmate) printf '%s' "$kind" ;; *) printf unknown ;; esac +} + +_fm_decision_fold_line() { # + local open=$1 line=$2 resolve=$3 held=$4 kind=$5 verb key note + # Declaration guard. A transition's verb ends at a colon, or - in the colonless + # form _fm_decision_key still accepts below - at a complete "[key=...]" token. + # A line holding neither is continuation prose, a bare word, or blank, and can + # never move the set. A `case` glob answers that in one pattern match; the + # equivalent parameter expansion costs tens of milliseconds per line under bash + # 3.2's global bracket-class substitution, which is the whole per-line cost of + # both folds on a status log of ordinary width. Same verdict, bounded cost. case "$line" in - *[![:space:]]*) ;; + *:*|*\[key=*\]*) ;; + *) printf '%s' "$open"; return 0 ;; + esac + status_line_verb "$line" verb + case "$line" in + *:*) case "$verb:$kind" in done:ship|done:scout|failed:ship|failed:scout) return 0 ;; esac ;; + esac + case "$verb" in + needs-decision|blocked|"$resolve"|"$held") ;; *) printf '%s' "$open"; return 0 ;; esac - verb=$(status_line_verb "$line") key=$(_fm_decision_key "$line") || { printf '%s' "$open"; return 0; } _fm_decision_key_transition_allowed "$key" "$(status_line_note "$line")" \ || { printf '%s' "$open"; return 0; } @@ -497,27 +575,52 @@ _fm_decision_fold_line() { # \t\t" line per still-open decision, in -# most-recently-opened-last order; prints nothing when none are open. Pure read of -# the file, no globals beyond the optional FM_CLASSIFY_RESOLVE_VERB override. This -# is the durable open-set the fleet snapshot and any point-in-time consumer must use -# instead of trusting the last status line. +# most-recently-opened-last order; prints nothing when none are open. Reads the +# status file, plus its sibling `.meta` for the task kind the terminal rule needs +# when the caller passes no ; no globals beyond the optional +# FM_CLASSIFY_RESOLVE_VERB override. This is the durable open-set the fleet +# snapshot and any point-in-time consumer must use instead of trusting the last +# status line. # The scan_open_decisions wrapper below enumerates a whole directory rather than # a single caller-chosen path, so a status file that is itself a symlink (e.g. # escaping the state directory) is rejected outright with a plain [ -L ] check # before any read - a cheap builtin, unlike fm_wake_latest_event's O_NOFOLLOW # subprocess read, which exists for that function's much narrower payload-driven # path resolution rather than this directory-local glob. -status_open_decisions() { # - local f=$1 line resolve held open='' +status_open_decisions() { # [] + local f=$1 kind=${2:-} line resolve held open='' verb [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 0 + kind=$(_fm_status_kind "$f" "$kind") resolve=${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT} held=${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT} while IFS= read -r line || [ -n "$line" ]; do - open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held") + status_line_verb "$line" verb + case "$verb" in + needs-decision|blocked|done|failed|"$resolve"|"$held") + open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held" "$kind") + ;; + esac done < "$f" printf '%s' "$open" } +# Resolve the log's current declaration at one boundary for crew-state consumers. +# Any decision the fold still holds open wins over unrelated events, and the +# fold's most recently opened record supplies it; the latest recognized event +# stands when nothing is open. +# Actual run/pane evidence is still reconciled by fm-crew-state.sh. +status_current_line() { # + local open key verb note current='' + open=$(status_open_decisions "$1" "$2") + while IFS=$'\t' read -r key verb note; do + case "$verb" in ?*) current="$verb [key=$key]: $note" ;; esac + done < has a record in a folded "\t\t" open set. _fm_open_set_has() { # case "$1" in @@ -552,33 +655,50 @@ EOF # the question is settled outright, so a structured row still open behind it is a # contradiction between the two records - see fm-captain-hold.sh's `diverged`. # -# Semantics are not re-derived here: every line goes through the same +# Semantics are not re-derived here: every candidate line goes through the same # _fm_decision_fold_line rule the two folds use, and the reported verb is read -# off the transitions that rule produces. Only lines whose parsed key equals the -# requested one can move that key, so a caller-supplied key other than "default" -# lets the scan pre-filter the stream to lines carrying its token and stay cheap -# on a long log. +# off the transitions that rule produces. +# +# One `grep` pre-selects those candidates so the bash fold below costs the log's +# TRANSITIONS rather than its whole lifetime length - status files are only ever +# appended to, and this runs per open task on every supervision presentation. +# The pre-select deliberately over-includes: it takes any line whose leading word +# could be a fold verb (including the ship/scout terminals, which carry no key +# token), and the fold alone decides which of them really moves the set. A line +# whose leading word is followed by neither whitespace, a colon, nor a bracket +# tag cannot be a transition, because the fold's own declaration guard rejects it. status_key_closing_verb() { # - local f=$1 want=$2 line resolve held open='' was verb='' stream + local f=$1 want=$2 line resolve held open='' was verb='' kind event candidates [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 0 [ -n "$want" ] || return 0 + kind=$(_fm_status_kind "$f") resolve=${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT} held=${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT} - if [ "$want" = default ]; then - stream=$(cat "$f") || return 0 - else - stream=$(grep -F "[key=$want]" "$f") || stream='' - fi - [ -n "$stream" ] || return 0 + candidates=$(grep -E \ + "^[[:space:]]*(needs-decision|blocked|done|failed|$resolve|$held)[[:space:]:[]" \ + "$f") || [ "$?" -eq 1 ] || candidates=$(cat "$f") while IFS= read -r line || [ -n "$line" ]; do + status_line_verb "$line" event + case "$event:$kind" in + done:ship|done:scout|failed:ship|failed:scout) ;; + *) + case "$event" in + needs-decision|blocked|"$resolve"|"$held") ;; + *) continue ;; + esac + if [ "$want" != default ]; then + case "$line" in *"[key=$want]"*) ;; *) continue ;; esac + fi + ;; + esac was=0 _fm_open_set_has "$open" "$want" && was=1 - open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held") + open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held" "$kind") if [ "$was" = 1 ] && ! _fm_open_set_has "$open" "$want"; then - verb=$(status_line_verb "$line") + verb=$event fi done <:`), `offset`, `ident`, then the folded open set. # FM_OPEN_DECISIONS_FOLD_VERSION must be bumped whenever # _fm_decision_fold_line semantics change, so persisted state from an older -# interpretation is discarded and rebuilt from byte 0. +# interpretation is discarded and rebuilt from byte 0; the kind suffix does the +# same when a task kind changes, because kind changes the fold below. # # Cursor invalidation is deliberately minimal, matching how status files are # ACTUALLY used in this repo: every one is created once (`>`) and only ever @@ -678,10 +800,18 @@ _fm_open_decisions_cursor_path() { # # and closes. # 5: status_line_verb now also reads through an UNBRACKETED correlation token, # so lines that previously folded as ordinary status become opens and closes. +# 6: a done/failed line on a ship or scout closes every open decision, and the +# persisted version now carries the task kind, so cursors folded without that +# terminal rule are discarded. +# 7: that terminal rule now fires only for a line carrying a colon, so a cursor +# folded when bare prose could close every open decision is discarded. +# 8: a colonless line without a complete "[key=...]" token is no longer a +# transition at all, so a cursor holding a phantom decision that bare prose +# opened - which no later line could close - is discarded. # Version 4 was already spent on the bracketed-tag parser change above, and a # cursor persisted under that reading predates this one, so it must still be # discarded and rebuilt from byte 0 under the new reading. -FM_OPEN_DECISIONS_FOLD_VERSION=5 +FM_OPEN_DECISIONS_FOLD_VERSION=8 # Portable device:inode identity for the rotation/recreation check below. _fm_open_decisions_file_ident() { # -> strongest available identity @@ -756,8 +886,10 @@ _fm_status_read_span() { # status_open_decisions_incremental() { # [] local f=$1 captured_end=${2:-} cf offset ident open='' trusted_open='' cursor_data first rest offset_line ident_line local version='' size actual_size cur_ident resolve held chunk_file chunk_size line cursor_dirty=0 - local target_cursor + local target_cursor kind fold_version [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 0 + kind=$(_fm_status_kind "$f") + fold_version="$FM_OPEN_DECISIONS_FOLD_VERSION:$kind" cf=$(_fm_open_decisions_cursor_path "$f") offset=0 ident='' @@ -769,7 +901,7 @@ status_open_decisions_incremental() { # [] case "$first" in version=*) version=${first#version=} - [ "$version" = "$FM_OPEN_DECISIONS_FOLD_VERSION" ] || version='' + [ "$version" = "$fold_version" ] || version='' rest=${cursor_data#*$'\n'} offset_line=${rest%%$'\n'*} case "$offset_line" in @@ -847,7 +979,7 @@ status_open_decisions_incremental() { # [] resolve=${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT} held=${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT} while IFS= read -r line || [ -n "$line" ]; do - open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held") + open=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held" "$kind") done < "$chunk_file" rm -f "$chunk_file" offset=$size @@ -856,7 +988,7 @@ status_open_decisions_incremental() { # [] if [ "$cursor_dirty" -eq 1 ]; then target_cursor="$cf.tmp.$$" { - printf 'version=%s\n' "$FM_OPEN_DECISIONS_FOLD_VERSION" + printf 'version=%s\n' "$fold_version" printf 'offset=%s\n' "$offset" printf 'ident=%s\n' "$cur_ident" if [ -n "$open" ]; then printf '%s' "$open"; fi @@ -1369,8 +1501,9 @@ EOF # a caller explicitly requests a migration snapshot. status_open_decisions_cursor_offset() { # local f=$1 cf offset=0 ident='' version='' cursor_data first rest open='' - local offset_line ident_line cur_ident size + local offset_line ident_line cur_ident size fold_version [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 1 + fold_version="$FM_OPEN_DECISIONS_FOLD_VERSION:$(_fm_status_kind "$f")" cf=$(_fm_open_decisions_cursor_path "$f") if [ -e "$cf" ] || [ -L "$cf" ]; then [ -f "$cf" ] && [ -r "$cf" ] && [ ! -L "$cf" ] || return 1 @@ -1379,7 +1512,7 @@ status_open_decisions_cursor_offset() { # case "$first" in version=*) version=${first#version=} - [ "$version" = "$FM_OPEN_DECISIONS_FOLD_VERSION" ] || version='' + [ "$version" = "$fold_version" ] || version='' rest=${cursor_data#*$'\n'} offset_line=${rest%%$'\n'*} case "$offset_line" in @@ -1422,7 +1555,7 @@ status_open_decisions_cursor_offset() { # fi if [ -n "${FM_STATUS_CURSOR_SNAPSHOT_FILE:-}" ]; then { - printf 'version=%s\n' "$FM_OPEN_DECISIONS_FOLD_VERSION" + printf 'version=%s\n' "$fold_version" printf 'offset=%s\n' "$offset" printf 'ident=%s\n' "$cur_ident" if [ -n "$open" ]; then printf '%s' "$open"; fi @@ -1633,14 +1766,16 @@ $1 EOF } -_fm_status_open_decision_origins() { # +_fm_status_open_decision_origins() { # [] local f=$1 line open='' after key verb note number=0 origins='' - local resolve held + local resolve held kind + kind=$(_fm_status_kind "$f" "${2:-}") resolve=${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT} held=${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT} while IFS= read -r line || [ -n "$line" ]; do number=$((number + 1)) - after=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held") + after=$(_fm_decision_fold_line "$open" "$line" "$resolve" "$held" "$kind") + [ -n "$after" ] || origins='' key=$(_fm_decision_key "$line") || { open=$after; continue; } verb=$(status_line_verb "$line") note=$(status_line_note "$line") @@ -1731,7 +1866,7 @@ status_span_first_actionable_record() { # [record- || { failed=1; break; } while IFS= read -r _line || [ -n "$_line" ]; do prefix_lines=$((prefix_lines + 1)); done < "$prefix_file" fi - origins=$(_fm_status_open_decision_origins "$full_file") || { failed=1; break; } + origins=$(_fm_status_open_decision_origins "$full_file" "$(_fm_status_kind "$f")") || { failed=1; break; } folded=1 fi live_line=$(while IFS=$(printf '\t') read -r _key _line; do @@ -1962,7 +2097,7 @@ signal_crew_provably_working() { # ... return 0 } -# 0 (terminal/actionable) if a stale window's last status line is +# 0 (terminal/actionable) if a stale window's latest recognized status event is # captain-relevant; 1 otherwise, including the no-status case. A 1 only means # "non-terminal"; the always-on watcher then applies crew_is_provably_working, # while the away-mode daemon applies its persistence recheck. diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 49ab696156f..8ef77cf25dd 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -72,18 +72,22 @@ # FAILED record whose daemon an explicit probe proves down reads unknown, # never failed: an instrument failure must not read as work failure # (nm_daemon_probe_down). -# 3. Reconcile the status log: if its last line says needs-decision/blocked but +# 3. Reconcile the status log through fm-classify-lib.sh's status_current_line: +# open decisions survive unrelated events and continuation prose cannot +# hide a declaration. Ship/scout terminal declarations supersede stale log +# decisions. If it says needs-decision/blocked but # the run-step shows the run moved on, the log is deterministically stale and # is flagged superseded. A genuinely parked run plus a needs-decision log # agree, and are reported as parked. A `blocked:` line that reports a # refused or missing daemon socket remains blocked even if an attributed -# run record is stale or terminal. Other daemon, timeout, or unreachability +# run record is stale or terminal, for as long as that blocker is still the +# log's latest event. Other daemon, timeout, or unreachability # claims are superseded BECAUSE THE RUN IS ALIVE when the run is # running/fixing with recent reported activity: a killed or timed-out drive # call is not daemon death, so that claim is answered by steering the crew # to reattach, not by escalating. # 4. No run for this crew (pre-validation, or kind=scout): fall back to the -# recorded backend's pane busy state, then the status log's last line only +# recorded backend's pane busy state, then the resolved status declaration # when its verb maps to a recognized run-state. Decision-only events such as # `resolved` never become current state or detail. # 5. Missing meta or torn-down worktree: report unknown · none. If no run is @@ -170,11 +174,6 @@ fi # --- status log ------------------------------------------------------------ -# Last non-empty status line; fm-classify-lib.sh owns leading-verb normalization. -log_last_line() { - [ -f "$LOG" ] || return 1 - grep -v '^[[:space:]]*$' "$LOG" 2>/dev/null | tail -1 -} # Map a status-log verb onto a canonical state for the fallback path. `paused` is # the deliberate-external-wait verb (fm-classify-lib.sh's FM_CLASSIFY_PAUSED_VERB): # a crew with no active run and an idle pane that declared a known external wait @@ -195,7 +194,7 @@ map_log_state() { # esac } -LOG_LINE=$(log_last_line || true) +LOG_LINE=$(status_current_line "$LOG" "$KIND") LOG_VERB=$(status_line_verb "$LOG_LINE") # --- remote secondmate: the true source is the remote endpoint --------------- @@ -860,15 +859,22 @@ if [ "$HAVE_RUN" = 1 ]; then # # A refused or missing daemon socket is positive daemon-down evidence and # outranks any attributed run record, including a terminal one left behind - # after the daemon stopped. Other blocked claims caused by a timed-out drive - # call are contradicted only when the run reports recent - # activity; the answer is then to steer the crew to reattach without touching - # the shared daemon. + # after the daemon stopped, but only while that blocker is itself the log's + # LATEST recognized event: a later event of any kind means the crew has moved + # on, and the attributed run is the better witness again. The evidence is + # therefore read off that latest event, not off the reconciled declaration - + # the two are the same line while the blocker is current, and when they differ + # the open blocker is by definition no longer the log's tip. Other blocked + # claims caused by a timed-out drive call are contradicted only when the run + # reports recent activity; the answer is then to steer the crew to reattach + # without touching the shared daemon. case "$LOG_VERB" in needs-decision|blocked) + LOG_LATEST=$(last_status_line "$LOG") if [ "$LOG_VERB" = blocked ] \ - && log_reports_daemon_socket_down "$LOG_LINE"; then - emit blocked status-log "$(status_line_note "$LOG_LINE")${SEP}daemon socket down despite attributed run record" + && [ "$(status_line_verb "$LOG_LATEST")" = blocked ] \ + && log_reports_daemon_socket_down "$LOG_LATEST"; then + emit blocked status-log "$(status_line_note "$LOG_LATEST")${SEP}daemon socket down despite attributed run record" fi if [ "$RUN_STATE" != parked ]; then if [ "$RUN_STATE" = working ]; then @@ -962,7 +968,7 @@ if [ "$KIND" != secondmate ]; then esac fi -# Fall back to the status log's last line, but ONLY when its verb maps to a real +# Fall back to the resolved status declaration, but ONLY when its verb maps to a real # run-state. A decision-closing event - resolved: (fm-classify-lib.sh's # FM_CLASSIFY_RESOLVE_VERB), and any future decision-only sibling - is NOT a state: # it exists solely to CLOSE a keyed decision in the durable fold, so a trailing diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 94f632e98c3..296159ce04e 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -800,7 +800,7 @@ task_json_lines() { # never clear another concern's keyed decision. A parked/blocked state, or a # non-authoritative status-log/none read on a still-live task, keeps the fold's # open decision surfacing. - open_decisions_tsv=$(status_open_decisions "$status_log") + open_decisions_tsv=$(status_open_decisions "$status_log" "$kind") if [ "$kind" != secondmate ] && \ { { { [ "$current_source" = run-step ] || [ "$current_source" = pane ]; } \ && [ "$current_state" != parked ] && [ "$current_state" != blocked ]; } \ diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 8b2457376bf..5cbaf9e63d2 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -351,20 +351,23 @@ notice_parent_report_failed() { # queue_notice_once "$record" "inactive-reconcile:$fingerprint" "$payload" || true } -# The whole terminal line a child's ledger ends in, or non-zero when the ledger -# is absent, unusable, still being appended (no trailing newline yet), or does -# not end in a done or failed line. +# The whole terminal event a child's ledger states, or non-zero when the ledger +# is absent, unusable, or states no done or failed event (1), or when that event +# is the line still being appended (2, no trailing newline yet). The event is +# selected through the shared latest-event reader, so the ledger path owns a +# terminal record whose continuation prose trails it, and an unfinished line of +# ordinary prose withholds nothing. child_terminal_ledger_line() { # local status=$1 snapshot last marker='__FM_LEDGER_SNAPSHOT_END__' [ -f "$status" ] && [ ! -L "$status" ] && [ -s "$status" ] || return 1 + last=$(last_status_line "$status") + case "$(status_line_verb "$last")" in done|failed) ;; *) return 1 ;; esac snapshot=$(cat "$status"; printf '%s' "$marker") || return 1 - case "$snapshot" in *$'\n'"$marker") ;; *) return 1 ;; esac - snapshot=${snapshot%"$marker"} - last=$(printf '%s' "$snapshot" | grep -v '^[[:space:]]*$' | tail -1) - case "$(status_line_verb "$last")" in - done|failed) printf '%s\n' "$last" ;; - *) return 1 ;; + case "$snapshot" in + *$'\n'"$marker") ;; + "$last$marker"|*$'\n'"$last$marker") return 2 ;; esac + printf '%s\n' "$last" } # Claim one already-delivered inactive fallback as the delivery of this ledger @@ -405,12 +408,11 @@ report_child_ledger_locked() { # pr=$(pr_for_task "$meta" "$last") incarnation=$(meta_incarnation "$meta") fingerprint=$(sha256_text "$incarnation|$id|$state|ledger|$last") - previous=$(grep -v '^[[:space:]]*$' "$status" 2>/dev/null \ - | tail -2 | awk 'NR == 1 { first = $0 } NR == 2 { print first }' || true) - predecessor_head=$(sha256_text "$previous") outcome_key="child-outcome-$id-$state-${fingerprint:0:8}" ensure_record "$fingerprint" "$id" "$incarnation" "$state" "$outcome_key" direct upstream "$pr" || return 1 [ -n "$RECORD_PENDING" ] || return 0 + last_status_line "$status" previous >/dev/null + predecessor_head=$(sha256_text "$previous") if claim_inactive_report_for_ledger "$id" "$incarnation" "$state" "$fingerprint" "$predecessor_head"; then # The fallback line is already on the parent channel. This reported ledger # receipt records that its richer rendering owes no second publication. @@ -483,8 +485,9 @@ reconcile_direct_child_locked() { # /dev/null 2>&1 && return 0 # A ledger that states its own outcome is the ledger-first path's to deliver. - if [ -n "$self" ] && child_terminal_ledger_line "$status" >/dev/null; then - return 0 + if [ -n "$self" ]; then + child_terminal_ledger_line "$status" >/dev/null + case "$?" in 0|2) return 0 ;; esac fi age=$(last_activity_age "$meta" "$status" "$turn") [ "$age" -ge "$FM_INACTIVE_RECONCILE_SECS" ] || return 0 @@ -493,7 +496,8 @@ reconcile_direct_child_locked() { # /dev/null + case "$?" in 0|2) return 0 ;; esac fi case "$state_line" in 'state: done '*) state='done' ;; diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 3b1966bded6..0f85d3b7c7b 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -30,7 +30,7 @@ # absorbed instead with its own long re-surface cadence, # never as a wedge, and that recheck reason names which # human the wait is on. Only when neither absorb class -# applies does the log's last line decide: +# applies does the log's latest recognized status event decide: # terminal (captain-relevant) or non-terminal (no verb), # both surfaced at once. A provably-working stale past the # wedge threshold also surfaces, with an "escalation N" @@ -2457,7 +2457,7 @@ EOF wake "stale: $w" fi elif stale_is_terminal "$w" "$STATE"; then - # The log's last line is captain-relevant - but that alone is not + # The log's latest status event is captain-relevant - but that alone is not # proof the crew is actually done: a crew's own status log gets no # new entry once firstmate hands it to a no-mistakes validation # (AGENTS.md's sparse status-reporting contract), so the log can diff --git a/docs/architecture.md b/docs/architecture.md index 8079e672c21..8026d38e6b9 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -76,7 +76,7 @@ Each `fm-wake-drain.sh` presentation runs the same liveness guard as the supervi Routine watcher polling, supervision no-ops, elapsed waiting time, and absorbed benign wakes stay silent. A declared external wait or an attended verified captain-held transfer trades that silence for one bounded recheck per pause window, naming which human the wait is on; while the away-posture record exists, captain-held work waits without rechecks and remains visible in the return brief. Crew status files are append-only wake-event logs, not current-state fields. -Because of that, a per-wake read of only the latest line can bury an earlier still-open `needs-decision`/`blocked` under later unrelated appends; `fm-wake-drain.sh` prints a separate, fleet-wide OPEN DECISIONS section on every presentation (including the empty-queue path session-start relies on), built through `fm-classify-lib.sh`'s cursor-backed incremental scan using the authoritative `status_open_decisions` fold semantics so the buried decision keeps surfacing until it is explicitly resolved while each presentation folds only new status-log appends. +Because of that, a per-wake read of only the latest line can bury an earlier still-open `needs-decision`/`blocked` under later unrelated appends; `fm-wake-drain.sh` prints a separate, fleet-wide OPEN DECISIONS section on every presentation (including the empty-queue path session-start relies on), built through `fm-classify-lib.sh`'s cursor-backed incremental scan using the authoritative `status_open_decisions` fold semantics so the buried decision keeps surfacing until that fold closes it while each presentation folds only new status-log appends. The drain coordinates that fold and its annotations through a locked fleet-wide snapshot whose `.status-presentation-cursor` manifest records each status file's identity plus independent annotation and outcome-backstop byte offsets. [`pi-supervision-branch.md`](pi-supervision-branch.md#lost-wake-outcome-backstop) owns the bounded lost-wake backstop that uses the latter offset. A queued signal annotation prints every status line still unread at that cursor, while the fleet-wide UNREAD STATUS section prints `note:` lines and reserved-key pending-reply resolutions once even on an empty-queue drain because those verbs never enter the OPEN DECISIONS fold. @@ -86,7 +86,7 @@ The explicit resolution is written by the actor that answers, not the busy worke This home's answerer close, pending-reply escalation close, and captain-held transfer use the provenance-guarded append owned by `bin/fm-wake-lib.sh`, so they advance the watcher marker only across their own bytes when all earlier bytes were already announced; pending or interleaved foreign bytes fail toward an ordinary wake. A turn-ended-only queue row omits its historical status annotation when that status file exactly matches the same seen marker. Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line. -`bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed, except that a `blocked:` event reporting a refused or missing daemon socket outranks a potentially stale active run record. +`bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed, except that a `blocked:` event reporting a refused or missing daemon socket outranks a potentially stale active run record only while that socket-down declaration is itself the log's latest recognized event, since any later event, including another `blocked:` one, means the crew moved on. For other daemon, timeout, or unreachability claims, a running or fixing run with recent pipeline-reported activity supersedes the event and names reattachment as the recovery instead of surfacing a false block. [`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules. It also owns which binding run wins when more than one recorded run binds to the same worktree: a live run outranks a terminal one, so a crashed run sitting at the worktree's own commit never reports a healthy task as failed while its live successor is still validating. @@ -95,7 +95,7 @@ During no-mistakes' `ci` monitor phase, it also reads the ci step log tail becau The most recent recognized ci log marker wins, so checks-green monitoring reports done while a later re-arm, failed-check, or issue marker returns the crew to working. `bin/fm-crew-state.sh` owns the evidence guard that recognizes ended CI monitors after green checks, including cancelled runs and skipped rebase steps; a passed run alone never proves a forge merge. In the coarse runs-ledger fallback, which has no steps table and no ci log, a terminal failed record whose daemon an explicit `daemon status` probe proves down reports unknown as unverified instead: an instrument failure must never read as work failure. -Only when no matching run exists does it consult semantic busy state; exact busy reports working, exact idle permits fallback to a status-log event whose verb maps to a recognized run-state, and unknown or a dead pane stays unknown instead of trusting a stale log. +Only when no matching run exists does it consult semantic busy state; exact busy reports working, exact idle permits fallback to the log's resolved current declaration - the newest decision the fold still holds open, otherwise the latest recognized event - when its verb maps to a recognized run-state, and unknown or a dead pane stays unknown instead of trusting a stale log. Decision-only events such as `resolved` never become current state or leak their prose into the current-state detail. In that status-log fallback, a declared external wait reports the distinct `paused` state with its reason. The semantic branch reports working only on an exact busy verdict and names the source that produced it; an unknown verdict never becomes working, never permits the status-log fallback, and never becomes a silent idle. @@ -157,9 +157,10 @@ On Pi and pi-signed the away daemon is no longer launched: the ordinary supervis A presence-gated sub-supervisor (`bin/fm-supervise-daemon.sh`) still extends this for walk-away supervision on the other harnesses: the `/afk` skill starts it through the tracked foreground helper `bin/fm-afk-start.sh` once the record exists, after which the watcher reverts to daemon-managed one-shot mode and the daemon self-handles routine wakes in bash. The watcher and daemon share `bin/fm-classify-lib.sh` for captain-relevant status verbs, declared-wait vocabulary (a `paused:` external wait and a verified `captain-held` transfer alike, through one combined predicate), and status-scan primitives. Terminal verbs remain captain-relevant, while a nonterminal progress verb cannot become terminal merely because its prose contains a legacy free-text token such as `merged`; bare legacy free-text lines remain compatible. +The shared latest-event read takes the most recent line that leads with a recognized verb or legacy token, so continuation prose and trailing blank lines after a multi-line record cannot hide a declared wait. Both supervisors classify the status bytes appended since they last classified that log, never its last line alone, and report every actionable event through the captured endpoint before committing that position. The watcher's `.seen-*` and `.hb-surfaced-` markers and the daemon's `.subsuper-seen-status-` marker independently track reported file state and successfully classified position, so an unchanged unreadable state reports once without advancing past unread content, while a changed state retries and an unusable position re-reads the whole log. -A keyed `needs-decision` or `blocked` transition accepted by the whole-file decision fold is retired only when that fold proves the exact opening closed, while a reserved-key transition the fold rejects surfaces as a reconciliation signal without becoming an open decision. +A keyed `needs-decision` or `blocked` transition accepted by the whole-file decision fold is retired only when that fold retires it - an explicit close for its exact key, or a terminal declaration by the ship or scout that owns the log - while a reserved-key transition the fold rejects surfaces as a reconciliation signal without becoming an open decision. The fold remains the sole owner of open/closed semantics, including same-key reopening and reserved-key handling, shared with the durable OPEN DECISIONS surface. The always-on watcher also uses that library's absorb classification on no-verb signals and first-sighting stale panes before status-log terminality is trusted, while the daemon maintains distinct wedge and declared-wait recheck cadences. The daemon's declared-wait window ages against the crew's own latest status line rather than against pane busy state, because a declared wait can legitimately hold a pane busy, and only a status append that stops declaring the wait ends that routing and restores wedge detection. diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index 97dc3fc0663..fd496a82cd4 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -1245,22 +1245,32 @@ test_terminal_single_owner_status_decision_does_not_block_empty_inventory() { mkdir -p "$home/data/$id" tasks_in "$home" add "$id" "Review a terminal sample finding" --kind scout --repo sample --start >/dev/null write_origin_meta "$home" "$id" - printf 'needs-decision [key=default]: choose route A or route B\ndone: report complete\n' \ + printf 'blocked [key=access]: waiting\ndone: report complete\nnote: cleanup complete\n' \ > "$home/state/$id.status" printf '# Terminal sample review\n\nNo unresolved captain choice remains.\n' > "$home/data/$id/report.md" open=$(bash -c '. "$1"; status_open_decisions "$2"' _ \ "$ROOT/bin/fm-classify-lib.sh" "$home/state/$id.status") - assert_contains "$open" "default" "fixture must retain the raw stale status decision" + [ -z "$open" ] || fail "the shared fold retained a pre-terminal blocker" run_captain "$home" complete "$id" --none >/dev/null \ || fail "terminal single-owner stale status decision blocked empty inventory completion" run_captain "$home" verify "$id" >/dev/null \ || fail "terminal single-owner stale status decision blocked inventory verification" + printf 'blocked [key=access]: reopened\nnote: more cleanup\n' >> "$home/state/$id.status" + if run_captain "$home" complete "$id" --none > "$home/reopened.out" 2> "$home/reopened.err"; then + fail "completion accepted a genuinely reopened post-terminal decision" + fi + if run_captain "$home" verify "$id" > "$home/reopened-verify.out" 2> "$home/reopened-verify.err"; then + fail "verification accepted a genuinely reopened post-terminal decision" + fi + printf 'resolved [key=access]: answered\nfailed: investigation ended\nnote: final cleanup\n' >> "$home/state/$id.status" + run_captain "$home" complete "$id" --none >/dev/null || fail "resolved reopening blocked completion" + run_captain "$home" verify "$id" >/dev/null || fail "resolved reopening blocked verification" run_teardown "$home" "$id" >/dev/null 2> "$home/terminal-teardown.err" \ || fail "terminal single-owner stale status decision blocked teardown: $(cat "$home/terminal-teardown.err")" secondmate=sample-secondmate write_origin_meta "$home" "$secondmate" secondmate - printf 'needs-decision [key=route]: choose route A or route B\ndone: heartbeat complete\n' \ + printf 'blocked [key=route]: waiting\ndone: heartbeat complete\nnote: cleanup complete\n' \ > "$home/state/$secondmate.status" if run_captain "$home" complete "$secondmate" --none \ > "$home/secondmate-terminal.out" 2> "$home/secondmate-terminal.err"; then diff --git a/tests/fm-classify-decision-key.test.sh b/tests/fm-classify-decision-key.test.sh index 8c4196a8a26..e6ede61d1e6 100755 --- a/tests/fm-classify-decision-key.test.sh +++ b/tests/fm-classify-decision-key.test.sh @@ -338,3 +338,149 @@ EOF test_closing_verb_separates_resolution_from_durable_transfer test_closing_verb_tracks_the_last_transition_in_both_positions + +# The per-key read pre-selects candidate lines by their leading verb before the +# bash fold sees them, and the resolve/durable-transfer verbs are overridable, so +# an overridden verb buried behind unrelated history must still close its key. +test_closing_verb_honors_overridden_transition_verbs() { + local dir f i + dir=$(case_dir closing-verb-overrides) + f="$dir/task.status" + printf 'kind=ship\n' > "$dir/task.meta" + printf 'blocked [key=route]: waiting\n' > "$f" + for ((i = 0; i < 200; i++)); do + printf 'note: routine reply\nworking: still going\nContinuation prose here.\n' >> "$f" + done + printf 'answered [key=route]: settled\n' >> "$f" + [ "$(FM_CLASSIFY_RESOLVE_VERB=answered status_key_closing_verb "$f" route)" = answered ] \ + || fail "an overridden resolve verb stopped closing its key" + [ "$(status_key_closing_verb "$f" route)" = blocked ] \ + || fail "without the override the same line must leave the key open" + printf 'blocked [key=access]: waiting\nawaiting-captain [key=access]: handed off\n' >> "$f" + [ "$(FM_CLASSIFY_CAPTAIN_HELD_VERB=awaiting-captain status_key_closing_verb "$f" access)" = awaiting-captain ] \ + || fail "an overridden durable-transfer verb stopped closing its key" + pass "overridden resolve and durable-transfer verbs still close keys behind unrelated history" +} + +test_closing_verb_filters_unrelated_history_without_subshell_growth() { + local dir f want tag size i level small large + dir=$(case_dir closing-verb-processes) + f="$dir/task.status" + printf 'kind=secondmate\n' > "$dir/task.meta" + for want in route default; do + tag="[key=$want]" + [ "$want" != default ] || tag='' + for size in 1 1000; do + printf 'blocked corr=0123456789abcdef %s: waiting\n' "$tag" > "$f" + for ((i = 0; i < size; i++)); do + printf 'note: routine reply\nworking: mentions [key=%s] in prose\ndone: another task finished\nfailed: unrelated work\nPR ready https://example.com/pull/1\n\n' "$want" >> "$f" + if [ "$want" != default ]; then + printf 'blocked [key=other]: another question\nresolved [key=other]: answered\n' >> "$f" + fi + done + printf 'resolved corr=0123456789abcdef: %s answered\nnote: cleanup complete\n' "$tag" >> "$f" + : > "$dir/children-$size" + ( + level=$BASH_SUBSHELL + set -T + trap 'if [ "$BASH_SUBSHELL" -gt "$level" ]; then printf x >> "$dir/children-$size"; fi' DEBUG + status_key_closing_verb "$f" "$want" > "$dir/output" + ) + [ "$(cat "$dir/output")" = resolved ] || fail "$want lost its resolution behind unrelated history" + done + small=$(wc -c < "$dir/children-1") + large=$(wc -c < "$dir/children-1000") + [ "$large" -le "$((small + 20))" ] || fail "$want launches subprocess work for unrelated history ($small -> $large)" + done + pass "per-key reads retain resolutions without subprocess work growing with unrelated history" +} + +test_closing_verb_filter_preserves_terminal_chronology() { + local dir f kind want tag terminal expected + dir=$(case_dir closing-verb-terminals) + f="$dir/task.status" + for kind in ship scout secondmate; do + printf 'kind=%s\n' "$kind" > "$dir/task.meta" + for want in access default; do + tag="[key=$want]" + [ "$want" != default ] || tag='' + for terminal in 'done' failed; do + printf 'blocked %s: waiting\n' "$tag" > "$f" + case "$terminal" in + done) printf 'done: report saved\n' >> "$f" ;; + failed) printf 'failed corr=0123456789abcdef [key=other]: task failed\n' >> "$f" ;; + esac + printf 'note: cleanup complete\n' >> "$f" + expected=$terminal + [ "$kind" != secondmate ] || expected=blocked + [ "$(status_key_closing_verb "$f" "$want")" = "$expected" ] || fail "$kind/$want lost $terminal chronology" + printf 'needs-decision: [key=%s] reopened\nnote: more cleanup\n' "$want" >> "$f" + [ "$(status_key_closing_verb "$f" "$want")" = needs-decision ] || fail "$kind/$want lost a post-terminal reopening" + done + done + done + pass "per-key filtering retains ship/scout terminals, reopenings, and secondmate blockers" +} + +test_closing_verb_filters_unrelated_history_without_subshell_growth +test_closing_verb_honors_overridden_transition_verbs +test_closing_verb_filter_preserves_terminal_chronology + +test_bare_prose_cannot_impersonate_a_terminal_declaration() { + local dir f kind word open + dir=$(case_dir prose-terminal) + open=$(printf 'route\tneeds-decision\tA or B?\n') + for kind in ship scout; do + for word in 'done' failed; do + f="$dir/$kind-$word.status" + printf 'kind=%s\n' "$kind" > "$dir/$kind-$word.meta" + printf 'needs-decision [key=route]: A or B?\npaused: waiting on the vendor\nSteps remaining:\n %s\n' \ + "$word" > "$f" + assert_fold "$f" "$open" "$kind: bare '$word' prose" + [ "$(status_key_closing_verb "$f" route)" = needs-decision ] \ + || fail "$kind: bare '$word' prose closed a still-open key" + f="$dir/$kind-$word-real.status" + printf 'kind=%s\n' "$kind" > "$dir/$kind-$word-real.meta" + printf 'needs-decision [key=route]: A or B?\n%s: real outcome\n' "$word" > "$f" + assert_fold "$f" '' "$kind: genuine $word supersedes" + [ "$(status_key_closing_verb "$f" route)" = "$word" ] \ + || fail "$kind: genuine $word no longer supersedes the open key" + done + done + pass "prose without a colon cannot impersonate a ship or scout terminal declaration" +} + +test_bare_prose_cannot_impersonate_a_terminal_declaration + +test_bare_prose_cannot_open_or_close_a_decision() { + local dir f word blocked + dir=$(case_dir prose-decision) + blocked=$(printf 'default\tblocked\tneed release access\n') + for word in blocked needs-decision resolved; do + f="$dir/open-$word.status" + printf 'kind=ship\n' > "$dir/open-$word.meta" + printf 'working: investigating the deploy\nOptions considered:\n %s\n' "$word" > "$f" + assert_fold "$f" '' "bare '$word' prose opened a decision" + + f="$dir/close-$word.status" + printf 'kind=ship\n' > "$dir/close-$word.meta" + printf 'blocked: need release access\nSteps remaining:\n %s\n' "$word" > "$f" + assert_fold "$f" "$blocked" "bare '$word' prose moved an open decision" + done + + f="$dir/keyed-colonless.status" + printf 'kind=ship\n' > "$dir/keyed-colonless.meta" + printf 'blocked [key=access]\n' > "$f" + assert_fold "$f" "$(printf 'access\tblocked\tblocked [key=access]\n')" \ + "a keyed colonless line stopped opening its key" + printf 'resolved [key=access]\n' >> "$f" + assert_fold "$f" '' "a keyed colonless line stopped closing its key" + + f="$dir/real-resolution.status" + printf 'kind=ship\n' > "$dir/real-resolution.meta" + printf 'blocked: need release access\nresolved: access granted\n' > "$f" + assert_fold "$f" '' "a genuine resolution stopped closing its decision" + pass "only a colon-bearing or keyed line is a decision transition in the fold" +} + +test_bare_prose_cannot_open_or_close_a_decision diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index a2ccd001ac8..2f3faf3ca3f 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -740,6 +740,44 @@ test_socket_refusal_over_terminal_run_reports_blocked() { pass "socket refusal over a terminal attributed run reports blocked" } +# The socket-down override is evidence about the log's CURRENT tip, not a latch: +# once the crew appends any later event the attributed run is the better witness. +test_socket_refusal_override_expires_when_the_crew_moves_on() { + reset_fakes + local d out + d=$(new_case daemon-socket-refused-superseded) + make_repo_on_branch "$d/wt" fm/feat-ds + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ds.meta" "window=fm:fm-feat-ds" "worktree=$d/wt" "kind=ship" + printf 'blocked: no-mistakes daemon socket is missing\n' > "$d/state/feat-ds.status" + FM_FAKE_AXI_STATUS="$(run_fixing_active_recent fm/feat-ds)" + out=$(run_crew_state "$d" feat-ds) + assert_contains "$out" "state: blocked" "socket-down as the latest event still outranks a live run" + assert_contains "$out" "source: status-log" "the override remains status-log evidence" + assert_contains "$out" "daemon socket down despite attributed run record" "the override names its reason" + + printf 'working: reattached and continuing\n' >> "$d/state/feat-ds.status" + out=$(run_crew_state "$d" feat-ds) + assert_contains "$out" "state: working" "a later working event hands the reading back to the live run" + assert_contains "$out" "source: run-step" "the superseded override no longer emits status-log state" + assert_not_contains "$out" "daemon socket down despite attributed run record" \ + "a stale socket-down blocker cannot override a live run forever" + + # The later event does not have to be one the decision fold accepts. A blocked + # line on a reserved key whose note does not speak that namespace is folded as + # ordinary status, so the socket-down blocker stays the reconciled declaration + # while the tip of the log has moved on; the override reads the tip, not the + # declaration, so the stale daemon evidence stays retired. + printf 'blocked: no-mistakes daemon socket is missing\nblocked [key=pending-reply-t3]: still waiting on the answer\n' \ + > "$d/state/feat-ds.status" + out=$(run_crew_state "$d" feat-ds) + assert_contains "$out" "state: working" "a later unfolded blocked event also hands the reading back to the run" + assert_contains "$out" "source: run-step" "the retired override emits no status-log state" + assert_not_contains "$out" "daemon socket down despite attributed run record" \ + "an unrelated later blocker cannot republish stale socket-down evidence" + pass "socket-down evidence outranks a live run only while it is the log's latest event" +} + # And the claim half: an ordinary blocked line over the same live run keeps the # generic reading, so the sharper one cannot fire on every superseded block. test_ordinary_blocked_over_live_run_keeps_plain_superseded() { @@ -1959,9 +1997,158 @@ test_no_run_idle_pane_paused() { assert_contains "$out" "state: paused" "paused log -> paused" assert_contains "$out" "source: status-log" "idle pause -> status-log source" assert_contains "$out" "holding for the upstream tool release" "the pause reason is carried in the detail" + printf 'The release window opens tomorrow.\n\n' >> "$d/state/feat-pause.status" + out=$(run_crew_state "$d" feat-pause) + assert_contains "$out" "state: paused" "continuation prose and trailing blanks preserve the pause" + assert_contains "$out" "holding for the upstream tool release" "multiline pause preserves its declared reason" pass "no run + idle pane on a paused: status reports state: paused with its reason" } +test_secondmate_open_block_survives_unrelated_append() { + reset_fakes + local d out suffix gen + d=$(new_case buried-block) + mkdir -p "$d/wt" + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/mate.meta" "window=fm:fm-mate" "worktree=$d/wt" "kind=secondmate" "harness=claude" + gen=$("$ROOT/bin/fm-busy-event.sh" arm "$d/state" mate) + "$ROOT/bin/fm-busy-event.sh" apply "$d/state" mate busy --gen "$gen" --source claude-hook --event user-prompt-submit + for suffix in '' 'note: unrelated progress' 'resolved [key=other]: unrelated answer' 'working: continuing another task' 'done: another task completed' 'failed: another task failed' $'done: another task completed\nnote: cleanup complete' $'failed: another task failed\nnote: cleanup complete'; do + printf 'blocked [key=access]: need release access\n%s\n' "$suffix" > "$d/state/mate.status" + out=$(run_crew_state "$d" mate) + assert_contains "$out" "state: blocked" "open blocker survives '$suffix' with a busy endpoint" + assert_contains "$out" "need release access" "the open blocker's reason remains visible" + done + printf 'resolved [key=access]: access granted\n' >> "$d/state/mate.status" + out=$(run_crew_state "$d" mate) + assert_contains "$out" "state: unknown" "matching resolution clears the blocker" + assert_not_contains "$out" "need release access" "closed blocker is not resurrected" + pass "a busy secondmate keeps its open blocker until that exact key closes" +} + +test_newest_open_decision_supplies_the_reported_detail() { + reset_fakes + local d out gen + d=$(new_case newest-open-decision) + mkdir -p "$d/wt" + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/mate.meta" "window=fm:fm-mate" "worktree=$d/wt" "kind=secondmate" "harness=claude" + gen=$("$ROOT/bin/fm-busy-event.sh" arm "$d/state" mate) + "$ROOT/bin/fm-busy-event.sh" apply "$d/state" mate busy --gen "$gen" --source claude-hook --event user-prompt-submit + printf 'blocked [key=a]: staging is down\nneeds-decision [key=b]: pick a rollout order\n' > "$d/state/mate.status" + out=$(run_crew_state "$d" mate) + assert_contains "$out" "state: parked" "the newer open decision is the reported state" + assert_contains "$out" "pick a rollout order" "the newer open decision supplies the detail" + printf 'blocked [key=c]: the deploy host went away\n' >> "$d/state/mate.status" + out=$(run_crew_state "$d" mate) + assert_contains "$out" "state: blocked" "a newer blocker takes the report back" + assert_contains "$out" "the deploy host went away" "the newest blocker supplies the detail" + printf 'resolved [key=c]: host restored\n' >> "$d/state/mate.status" + out=$(run_crew_state "$d" mate) + assert_contains "$out" "state: parked" "closing the newest decision falls back to the next open one" + assert_contains "$out" "pick a rollout order" "the still-open older decision is not lost" + pass "the most recently opened decision supplies the reported state and detail" +} + +test_single_owner_terminal_declaration_supersedes_stale_decision() { + reset_fakes + local d kind opener terminal out key expected + d=$(new_case terminal-stale-decision) + mkdir -p "$d/wt" + make_fakebin "$d" >/dev/null + arm_idle_record "$d/state" task + for kind in scout ship; do + fm_write_meta "$d/state/task.meta" "window=fm:fm-task" "worktree=$d/wt" "kind=$kind" "harness=claude" + for opener in needs-decision blocked; do + for terminal in 'done' failed; do + printf '%s [key=choice]: an earlier decision\n%s: final outcome\nContinuation prose.\n\n' \ + "$opener" "$terminal" > "$d/state/task.status" + out=$(run_crew_state "$d" task) + assert_contains "$out" "state: $terminal" "$kind terminal declaration supersedes stale $opener" + assert_contains "$out" "final outcome" "the terminal declaration supplies the detail" + printf 'note: cleanup complete\n' >> "$d/state/task.status" + out=$(run_crew_state "$d" task) + assert_contains "$out" "state: unknown" "$kind cleanup note does not revive a pre-terminal $opener" + assert_not_contains "$out" "an earlier decision" "superseded decision detail stays absent after cleanup" + expected=parked + [ "$opener" != blocked ] || expected=blocked + for key in choice new-choice; do + printf '%s [key=%s]: reopened after completion\nnote: more cleanup\n' "$opener" "$key" >> "$d/state/task.status" + out=$(run_crew_state "$d" task) + assert_contains "$out" "state: $expected" "$kind retains a post-terminal $opener for $key" + assert_contains "$out" "reopened after completion" "the reopened decision supplies the detail" + printf 'resolved [key=%s]: answered\n' "$key" >> "$d/state/task.status" + out=$(run_crew_state "$d" task) + assert_contains "$out" "state: unknown" "matching resolution clears the reopened decision" + assert_not_contains "$out" "an earlier decision" "closing a reopened decision cannot revive pre-terminal decisions" + done + done + done + done + pass "ship and scout terminal declarations supersede stale decisions" +} + +test_latest_status_preserves_legacy_completions() { + local d event line + d=$(new_case latest-legacy) + for event in 'PR ready https://example.com/pull/1' 'checks green' 'ready in branch fm/topic' merged 'PR READY https://example.com/pull/1'; do + printf 'paused: awaiting release\n%s\nMore detail: cleanup complete.\n\n' "$event" > "$d/state/task.status" + line=$(last_status_line "$d/state/task.status") + [ "$line" = "$event" ] || fail "legacy completion '$event' was hidden by an earlier pause" + status_is_captain_relevant "$line" || fail "legacy completion is no longer captain-relevant" + status_is_paused "$line" && fail "legacy completion retained pause handling" + printf 'working: following up on merged work\n' >> "$d/state/task.status" + line=$(last_status_line "$d/state/task.status") + [ "$line" = 'working: following up on merged work' ] || fail "later working event did not supersede legacy completion" + status_is_captain_relevant "$line" && fail "legacy prose made a working event captain-relevant" + printf 'paused: waiting on upstream PR #123 to land\nOnce it is %s I will rebase and continue.\n\n' "$event" > "$d/state/task.status" + line=$(last_status_line "$d/state/task.status") + [ "$line" = 'paused: waiting on upstream PR #123 to land' ] || fail "continuation prose mentioning '$event' hid a multi-line pause: $line" + status_is_paused "$line" || fail "a multi-line pause lost pause handling behind prose mentioning '$event'" + done + ( + shopt -u nocasematch + FM_CAPTAIN_RE='custom-event:' status_is_captain_relevant 'CUSTOM-EVENT: ready' || fail "custom captain regex lost case-insensitive matching" + shopt -q nocasematch && fail "captain matching changed caller shell options" + FM_CAPTAIN_RE='custom-event:' status_is_captain_relevant 'done: ready' && fail "custom captain regex did not replace defaults" + shopt -s nocasematch + status_is_captain_relevant 'unrelated prose' && fail "ordinary prose became captain-relevant" + shopt -q nocasematch || fail "captain matching cleared caller shell options" + ) || fail "captain matching changed regex or shell-option behavior" + pass "latest status retains legacy completion events and shared captain matching" +} + +test_latest_status_subshell_work_does_not_grow_with_history() { + local d size i level small large window + d=$(new_case latest-processes) + window=${FM_CLASSIFY_EVENT_WINDOW_LINES:-200} + for size in "$window" "$((window * 10))"; do + { + for ((i = 0; i < size; i++)); do + printf 'working corr=0123456789abcdef [key=phase]: progress\nMore detail: still working.\n' + done + printf 'PR ready https://example.com/pull/1\npaused corr=0123456789abcdef [key=release]: awaiting release\n\n' + } > "$d/state/task.status" + : > "$d/children-$size" + ( + level=$BASH_SUBSHELL + set -T + trap 'if [ "$BASH_SUBSHELL" -gt "$level" ]; then printf x >> "$d/children-$size"; fi' DEBUG + last_status_line "$d/state/task.status" > "$d/output" + ) + [ "$(cat "$d/output")" = 'paused corr=0123456789abcdef [key=release]: awaiting release' ] \ + || fail "latest status lost correlation-token parsing on a long log" + done + small=$(wc -c < "$d/children-$window") + large=$(wc -c < "$d/children-$((window * 10))") + [ "$large" -le "$((small + 20))" ] || fail "latest status shell work grows with history ($small -> $large)" + printf 'paused: awaiting a long quiet tail\n' > "$d/state/task.status" + for ((i = 0; i < 500; i++)); do printf 'continuation prose %s\n' "$i" >> "$d/state/task.status"; done + [ "$(last_status_line "$d/state/task.status")" = 'paused: awaiting a long quiet tail' ] \ + || fail "a declared pause buried under a long prose tail was hidden" + pass "latest status subprocess work stays bounded and still reads past a long prose tail" +} + test_no_run_idle_pane_custom_paused_verb() { reset_fakes local d; d=$(new_case custom-paused) @@ -2746,8 +2933,14 @@ test_stale_blocked_superseded test_daemon_claim_over_live_run_reads_run_alive test_socket_refusal_over_stale_fixing_run_reports_blocked test_socket_refusal_over_terminal_run_reports_blocked +test_socket_refusal_override_expires_when_the_crew_moves_on test_ordinary_blocked_over_live_run_keeps_plain_superseded test_genuine_daemon_down_reports_blocked +test_secondmate_open_block_survives_unrelated_append +test_newest_open_decision_supplies_the_reported_detail +test_single_owner_terminal_declaration_supersedes_stale_decision +test_latest_status_preserves_legacy_completions +test_latest_status_subshell_work_does_not_grow_with_history test_genuine_parked_not_superseded test_scalar_gate_parked_not_superseded test_gate_block_parked_not_superseded diff --git a/tests/fm-fleet-snapshot-view.test.sh b/tests/fm-fleet-snapshot-view.test.sh index 71b2acbdc18..4ee0e9bf219 100755 --- a/tests/fm-fleet-snapshot-view.test.sh +++ b/tests/fm-fleet-snapshot-view.test.sh @@ -893,7 +893,7 @@ test_open_decision_clears_on_keyed_resolution() { # must not linger as pending. Decisions come purely from the keyed fold reconciled # against the crew lifecycle; report prose never opens or reopens a decision. test_completed_scout_report_is_pointer_not_pending() { - local home fakebin out + local home fakebin out kind terminal id phase single mate single_state mate_state home=$(make_home completed-scout) mkdir -p "$home/projects/scout-wt" "$home/data/lavish-103" fm_write_meta "$home/state/lavish-103.meta" \ @@ -918,6 +918,55 @@ test_completed_scout_report_is_pointer_not_pending() { and (.hints.open_decisions | length) == 0 and .hints.scout_report_present == true ' >/dev/null || fail "a completed scout report must be a pointer, not a pending decision: $out" + + # Same terminal-supersession contract across ship/scout/secondmate, both snapshot + # modes, and reopen/resolve after cleanup. + home=$(make_home terminal-cleanup) + mkdir -p "$home/projects/task" + fakebin=$(make_fakebin "$home") + for kind in ship scout secondmate; do + for terminal in 'done' failed; do + id="$kind-$terminal" + fm_write_meta "$home/state/$id.meta" \ + "window=firstmate:fm-$id" "worktree=$home/projects/task" \ + "kind=$kind" "harness=claude" + record_claude_idle "$home/state" "$id" + printf 'blocked [key=access]: waiting\nneeds-decision [key=choice]: choose a route\n%s: final outcome\nnote: cleanup complete\n' \ + "$terminal" > "$home/state/$id.status" + done + done + for phase in terminal reopened resolved; do + case "$phase" in + terminal) single='[]'; mate='["access","choice"]'; single_state=unknown; mate_state=parked ;; + reopened) single='["access","new-choice"]'; mate='["access","choice","new-choice"]'; single_state=parked; mate_state=parked ;; + resolved) single='[]'; mate='["choice"]'; single_state=unknown; mate_state=parked ;; + esac + for kind in ship scout secondmate; do + for terminal in 'done' failed; do + id="$kind-$terminal" + case "$phase" in + reopened) printf 'blocked [key=access]: reopened access\nneeds-decision [key=new-choice]: a new choice\nnote: more cleanup\n' >> "$home/state/$id.status" ;; + resolved) printf 'resolved [key=access]: access granted\nresolved [key=new-choice]: answered\nnote: final cleanup\n' >> "$home/state/$id.status" ;; + esac + done + done + out=$(PATH="$fakebin:$PATH" FM_HOME="$home" "$SNAPSHOT" --json) + printf '%s' "$out" | jq -e --argjson single "$single" --argjson mate "$mate" \ + --arg single_state "$single_state" --arg mate_state "$mate_state" ' + .tasks | length == 6 and all(.[]; + (.kind == "secondmate") as $persistent + | (.hints.open_decisions | map(.key) | sort) == (if $persistent then $mate else $single end) + and .current_state.state == (if $persistent then $mate_state else $single_state end) + and .hints.blocked_event == (if $persistent then $mate else $single end | index("access") != null) + and .hints.pending_decision == (if $persistent then $mate else $single end | any(. != "access"))) + ' >/dev/null || fail "$phase snapshot revived a completed decision or lost a current one: $out" + out=$(PATH="$fakebin:$PATH" FM_HOME="$home" "$SNAPSHOT" --secondmate-home-summary) + printf '%s' "$out" | jq -e --argjson single "$single" --argjson mate "$mate" ' + (.decisions_open | map({id,key}) | sort_by(.id,.key)) == + (([ ("ship-done","ship-failed","scout-done","scout-failed") as $id | $single[] | {id:$id,key:.} ] + + [ ("secondmate-done","secondmate-failed") as $id | $mate[] | {id:$id,key:.} ]) | sort_by(.id,.key)) + ' >/dev/null || fail "$phase home summary revived a completed decision or lost a current one: $out" + done pass "a completed scout's stale decision surfaces as a report pointer, not pending" } diff --git a/tests/fm-inactive-reconcile.test.sh b/tests/fm-inactive-reconcile.test.sh index 9726fb6a1df..0c57b55691e 100755 --- a/tests/fm-inactive-reconcile.test.sh +++ b/tests/fm-inactive-reconcile.test.sh @@ -183,6 +183,8 @@ test_local_secondmate_delivers_terminal_ledger_line() { FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" [ "$(grep -c 'child-outcome-child-done' "$MAIN/state/mate.status")" = 1 ] \ || fail "a second poll delivered the same ledger line again" + printf 'Report at /tmp/report.md\n' >> "$MATE/state/child.status" + age "$MATE/state/child.status" FM_FAKE_CREW_STATE='done' run_reconcile "$MATE" --startup ! grep -q 'inactive-outcome-' "$MAIN/state/mate.status" \ || fail "the inactive path reported a child the ledger delivery already owned" @@ -190,6 +192,61 @@ test_local_secondmate_delivers_terminal_ledger_line() { pass "secondmate delivers a child's terminal ledger line once, on the next poll, from the ledger alone" } +# A terminal record written as a multi-line block belongs to the ledger path +# whether the block lands before or during the state read: it is delivered once, +# under the ledger's own outcome key, and the inactive fallback stays out of it. +test_secondmate_multiline_terminal_outcome_is_delivered_once() { + local terminal timing key + for terminal in 'done' failed; do + for timing in before during; do + make_world "multiline-$terminal-$timing"; bind_secondmate local + write_child "$MATE" child 'working: finishing validation' + if [ "$timing" = before ]; then + printf '%s: validation finished\nSee the report for details.\n\n' "$terminal" >> "$MATE/state/child.status" + age "$MATE/state/child.status" + else + cat > "$WORLD/fakebin/fm-crew-state.sh" <<'SH' +#!/usr/bin/env bash +printf '%s: validation finished\nSee the report for details.\n\n' "$FM_FAKE_CREW_STATE" >> "$FM_STATE_OVERRIDE/$1.status" +printf 'state: %s · source: fake\n' "$FM_FAKE_CREW_STATE" +SH + fi + FM_FAKE_CREW_STATE="$terminal" run_reconcile "$MATE" --startup + age "$MATE/state/child.status" + FM_FAKE_CREW_STATE="$terminal" run_reconcile "$MATE" --startup + run_report "$MATE" child + key=$(reported_outcome_key "$MATE" child "$terminal") \ + || fail "$terminal with trailing prose arriving $timing state read was not owned by the ledger" + grep -Fq "$terminal [key=$key]: child child $terminal: validation finished" "$MAIN/state/mate.status" \ + || fail "$terminal with trailing prose arriving $timing state read was lost: $(cat "$MAIN/state/mate.status" 2>/dev/null)" + [ "$(wc -l < "$MAIN/state/mate.status" | tr -d ' ')" = 1 ] \ + || fail "$terminal with trailing prose arriving $timing state read was delivered twice" + [ "$(outcome_count "$MATE" reported)" = 1 ] \ + || fail "multiline $terminal outcome did not retain exactly one receipt" + done + done + pass "multiline terminal outcomes are reported once before or during a state read" +} + +# A child that dies mid-prose cannot hide an outcome its run already proves: an +# unterminated continuation line states no terminal event, so the inactive +# fallback still reports the attributed failure upward. +test_secondmate_unterminated_prose_reports_run_outcome() { + make_world unterminated-prose; bind_secondmate local + write_child "$MATE" child 'working: compiling' + printf 'Still going' >> "$MATE/state/child.status" + age "$MATE/state/child.status" + FM_FAKE_CREW_STATE='failed' run_reconcile "$MATE" --startup + grep -Fq "failed [key=inactive-outcome-mate-child-failed]: inactive terminal child=child" "$MAIN/state/mate.status" \ + || fail "an unterminated prose line withheld a proven failure: $(cat "$MAIN/state/mate.status" 2>/dev/null)" + [ "$(outcome_count "$MATE" reported)" = 1 ] || fail "the fallback report did not retain its receipt" + age "$MATE/state/child.status" + FM_FAKE_CREW_STATE='failed' run_reconcile "$MATE" --startup + [ "$(wc -l < "$MAIN/state/mate.status" | tr -d ' ')" = 1 ] \ + || fail "the proven failure was reported twice" + pass "an unterminated continuation line does not withhold a proven child outcome" +} + # A busy child cannot keep later ledger outcomes from being visited, and is # retried on the next poll after its lifecycle lock becomes available. test_busy_child_does_not_starve_later_ledger_outcomes() { @@ -376,14 +433,18 @@ test_secondmate_partial_ledger_line_waits_for_newline() { make_world partial; bind_secondmate local write_child "$MATE" child 'working: nearly there' printf 'done: half writ' >> "$MATE/state/child.status" - FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" - [ ! -e "$MAIN/state/mate.status" ] || ! grep -q 'child-outcome-' "$MAIN/state/mate.status" \ + age "$MATE/state/child.status" + FM_FAKE_CREW_STATE='done' run_reconcile "$MATE" --startup + [ ! -s "$MAIN/state/mate.status" ] \ || fail "an unterminated ledger line was delivered: $(cat "$MAIN/state/mate.status")" printf 'ten\n' >> "$MATE/state/child.status" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" key=$(reported_outcome_key "$MATE" child 'done') || fail "completed ledger receipt key missing" grep -Fq "done [key=$key]: child child done: half written" "$MAIN/state/mate.status" \ || fail "the completed line was not delivered once its newline landed" + FM_FAKE_CREW_STATE='done' run_reconcile "$MATE" --startup + [ "$(wc -l < "$MAIN/state/mate.status" | tr -d ' ')" = 1 ] \ + || fail "completing the partial line delivered the outcome twice" pass "a ledger line still being appended waits for its newline" } @@ -835,6 +896,8 @@ SH test_main_direct_terminal_presentation_receipt test_local_secondmate_delivers_terminal_ledger_line +test_secondmate_multiline_terminal_outcome_is_delivered_once +test_secondmate_unterminated_prose_reports_run_outcome test_busy_child_does_not_starve_later_ledger_outcomes test_secondmate_ledger_delivery_carries_report_and_failure test_pr_field_requires_recorded_pr_or_ready_signal_line diff --git a/tests/fm-send-resolve-key.test.sh b/tests/fm-send-resolve-key.test.sh index fc51a123a55..62a8f05d2d5 100755 --- a/tests/fm-send-resolve-key.test.sh +++ b/tests/fm-send-resolve-key.test.sh @@ -13,8 +13,8 @@ # text: # 1. An answer send closes the open decision, including the answer-starts-work # scenario where the worker never writes a matching resolved line. -# 2. A routine steer without the flag never closes anything, and a working:/ -# done: line still cannot clear a captain decision. +# 2. A routine steer without the flag never closes anything, and a working: +# line still cannot clear a captain decision. # 3. A key that is not open refuses BEFORE anything is sent (mistype safety). # 4. The close happens at enqueue: a failed doorbell ring still closes the # answered key (the record is durably sent), while a failed ENQUEUE - the @@ -243,15 +243,18 @@ test_routine_steer_never_closes() { run_send "$fb" "$home" "$log" t3 "unrelated nudge, keep going"; rc=$? expect_code 0 "$rc" "a routine steer should still succeed" printf 'working: resumed\n' >> "$home/state/t3.status" - printf 'done: unrelated milestone\n' >> "$home/state/t3.status" if grep -F 'resolved' "$home/state/t3.status" >/dev/null; then fail "a routine steer wrote a resolved line: $(cat "$home/state/t3.status")" fi out=$(drain_out "$home") printf '%s' "$out" | grep -F '[key=schema]' >/dev/null \ - || fail "a routine steer (or later working/done lines) cleared an unanswered captain decision: $out" - pass "fm-send: a send without --resolve-key never closes a decision, and working/done still cannot" + || fail "a routine steer or later working line cleared an unanswered captain decision: $out" + printf 'done: task complete\nnote: cleanup complete\n' >> "$home/state/t3.status" + run_send "$fb" "$home" "$log" t3 --resolve-key schema "answer to a stale decision" > "$dir/terminal.out" 2> "$dir/terminal.err"; rc=$? + expect_code 1 "$rc" "an answer to a terminally superseded decision must refuse" + [ ! -e "$home/state/t3.inbox/002.msg" ] || fail "a stale decision answer was delivered" + pass "fm-send preserves decisions through routine work and refuses superseded terminal decisions" } test_not_open_key_refuses_before_send() { diff --git a/tests/fm-wake-drain-open-decisions-cursor.test.sh b/tests/fm-wake-drain-open-decisions-cursor.test.sh index c0f8c5fe2f6..d206cea72e4 100755 --- a/tests/fm-wake-drain-open-decisions-cursor.test.sh +++ b/tests/fm-wake-drain-open-decisions-cursor.test.sh @@ -347,6 +347,69 @@ test_previous_fold_cache_is_refolded_under_current_semantics() { pass "an old fold cache is rebuilt once before same-version incremental reads resume" } +test_terminal_supersession_reaches_cached_drains() { + local dir state status cursor out kind terminal expected closing ident size span pass_number + for kind in scout ship secondmate; do + for terminal in 'done' failed; do + dir=$(make_case "terminal-$kind-$terminal") + state="$dir/state"; status="$state/task.status"; cursor="$state/.task.open-decisions-cursor"; out="$dir/drain.out" + printf 'kind=%s\n' "$kind" > "$state/task.meta" + printf 'blocked [key=access]: waiting\n' > "$status" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$dir/drain.err" || fail "initial blocked drain failed" + assert_contains "$(cat "$out")" 'task [key=access] blocked: waiting' "initial blocker must surface" + printf '%s: report saved\nnote: cleanup complete\n' "$terminal" >> "$status" + expected=''; closing=$terminal + if [ "$kind" = secondmate ]; then expected=$'access\tblocked\twaiting'; closing=blocked; fi + for pass_number in 1 2; do + if [ "$pass_number" = 2 ]; then + ident=$(sed -n 's/^ident=//p' "$cursor") + size=$(LC_ALL=C wc -c < "$status" | tr -d '[:space:]') + printf 'version=5\noffset=%s\nident=%s\naccess\tblocked\twaiting' "$size" "$ident" > "$cursor" + fi + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$dir/drain.err" || fail "$kind terminal drain failed" + if [ "$kind" = secondmate ]; then + assert_contains "$(cat "$out")" 'task [key=access] blocked: waiting' "secondmate blocker must survive $terminal and cache migration" + else + assert_not_contains "$(cat "$out")" 'OPEN DECISIONS' "$kind pre-terminal blocker resurfaced after $terminal or cache migration" + fi + bash -c '. "$1"; [ "$(status_open_decisions "$2")" = "$3" ] && [ "$(status_open_decisions_incremental "$2")" = "$3" ] && [ "$(status_key_closing_verb "$2" access)" = "$4" ]' \ + _ "$ROOT/bin/fm-classify-lib.sh" "$status" "$expected" "$closing" \ + || fail "$kind whole-file, incremental, and key-history reads disagree with terminal supersession" + done + span=$(bash -c '. "$1"; status_span_first_actionable "$2" 0' _ "$ROOT/bin/fm-classify-lib.sh" "$status") + if [ "$kind" = secondmate ]; then + assert_contains "$span" 'blocked [key=access]: waiting' "secondmate opening must remain actionable" + else + assert_not_contains "$span" 'waiting' "$kind superseded opening remained actionable in a captured span" + fi + printf 'blocked [key=access]: reopened\nneeds-decision [key=new]: a new decision\nnote: more cleanup\n' >> "$status" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$dir/drain.err" || fail "reopened drain failed" + assert_contains "$(cat "$out")" 'task [key=access] blocked: reopened' "post-terminal reopening must surface" + assert_contains "$(cat "$out")" 'task [key=new] needs-decision: a new decision' "post-terminal new key must surface" + printf 'resolved [key=access]: answered\nresolved [key=new]: answered\nnote: final cleanup\n' >> "$status" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$dir/drain.err" || fail "resolved drain failed" + assert_not_contains "$(cat "$out")" 'OPEN DECISIONS' "matching resolutions must close reopened decisions" + done + done + pass "terminal supersession reaches whole-file reads, incremental drains, old caches, and captured spans" +} + +test_kind_changes_invalidate_folded_decisions() { + local dir state status kind expected + dir=$(make_case cursor-kind-change); state="$dir/state"; status="$state/task.status" + printf 'blocked [key=access]: waiting\ndone: report saved\nnote: cleanup complete\n' > "$status" + for kind in unknown ship secondmate scout; do + [ "$kind" = unknown ] || printf 'kind=%s\n' "$kind" >> "$state/task.meta" + case "$kind" in unknown|secondmate) expected=$'access\tblocked\twaiting' ;; *) expected='' ;; esac + bash -c '. "$1"; [ "$(status_open_decisions_incremental "$2")" = "$3" ] && [ "$(status_open_decisions "$2")" = "$3" ]' \ + _ "$ROOT/bin/fm-classify-lib.sh" "$status" "$expected" \ + || fail "cached decisions did not follow the current $kind metadata without a status append" + done + pass "folded decisions are rebuilt when task-kind evidence changes" +} + +test_terminal_supersession_reaches_cached_drains +test_kind_changes_invalidate_folded_decisions test_truncated_log_falls_back_to_a_full_refold_not_a_dropped_decision test_same_size_rewrite_is_detected_via_inode_identity test_read_failure_preserves_state_for_retry diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 65ae867770b..e50aedd2f78 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -302,6 +302,10 @@ test_stale_is_terminal_classifier() { stale_is_terminal "default:w1:p2" "$state" || fail "terminal herdr stale status not resolved through metadata" printf 'working: compiling\n' > "$state/nonterm.status" stale_is_terminal "sess:fm-nonterm" "$state" && fail "non-terminal stale classified terminal" + printf 'paused: waiting on upstream PR #123 to land\nOnce it is merged I will rebase and continue.\n' > "$state/prose-pause.status" + stale_is_terminal "sess:fm-prose-pause" "$state" && fail "prose mentioning a legacy token escalated a multi-line pause as terminal" + status_is_paused_or_captain_held "$(last_status_line "$state/prose-pause.status")" \ + || fail "prose mentioning a legacy token hid a multi-line pause from the wait cadence" stale_is_terminal "sess:fm-missing" "$state" && fail "stale with no status classified terminal" pass "stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign" } @@ -311,6 +315,11 @@ test_classifier_primitives() { dir=$(make_case classify-primitives); state="$dir/state" printf 'working: a\n\ndone: b\n\n' > "$state/x.status" [ "$(last_status_line "$state/x.status")" = "done: b" ] || fail "last_status_line did not return the last non-blank line" + printf 'paused [corr=aaaa1111bbbb2222]: waiting for release\nMore detail: still waiting.\n\n' > "$state/x.status" + [ "$(last_status_line "$state/x.status")" = 'paused [corr=aaaa1111bbbb2222]: waiting for release' ] \ + || fail "continuation prose hid the last declared status verb" + printf 'merged\n\n' > "$state/x.status" + [ "$(last_status_line "$state/x.status")" = merged ] || fail "legacy free-text status was lost" status_is_captain_relevant "done: b" || fail "done: not recognized as captain-relevant" status_is_captain_relevant "needs-decision [key=q1]: b" || fail "keyed needs-decision not recognized as captain-relevant" status_is_captain_relevant "working: b" && fail "working: wrongly recognized as captain-relevant" @@ -1473,6 +1482,27 @@ test_secondmate_status_note_surfaced_despite_busy_agent() { pass "a secondmate's status note surfaces even while its own agent is busy" } +test_secondmate_buried_block_wakes_despite_busy_agent() { + local dir state fakebin out suffix pid + for suffix in '' 'note: unrelated progress' 'resolved [key=other]: unrelated answer'; do + dir=$(make_case "secondmate-buried-block-${#suffix}"); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + printf 'kind=secondmate\n' > "$state/mate.meta" + printf 'blocked [key=access]: need release access\n%s\n' "$suffix" > "$state/mate.status" + [ "$(status_line_verb "$(status_current_line "$state/mate.status" secondmate)")" = blocked ] \ + || fail "unrelated '$suffix' hid an open blocker from current-state resolution" + export FM_FAKE_CREW_STATE='state: working · source: pane · harness busy' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "busy secondmate's blocker did not wake after '$suffix'" + grep -F "signal: $state/mate.status" "$out" >/dev/null \ + || fail "busy secondmate's blocker was not surfaced" + grep -F "$state/mate.status" "$state/.wake-queue" >/dev/null \ + || fail "busy secondmate's blocker was not durably queued" + done + pass "a secondmate blocker wakes despite busy evidence and later unrelated appends" +} + test_self_announced_close_does_not_rewake_but_next_note_does() { local dir state fakebin out status_file pid rc dir=$(make_case self-close-quiet); state="$dir/state"; fakebin="$dir/fakebin"; out="$dir/watch.out" @@ -2917,7 +2947,7 @@ test_secondmate_paused_resurfaces_in_normal_mode() { window="test:fm-secondmate-held" printf 'idle awaiting external\n' > "$capture_file" printf 'window=%s\nkind=secondmate\n' "$window" > "$state/secondmate-held.meta" - printf 'paused: awaiting the upstream release\n' > "$statusf" + printf 'paused: awaiting the upstream release\nThe release window opens tomorrow.\n\n' > "$statusf" back=$(( $(date +%s) - 500 )) if [ "$(uname)" = Darwin ]; then touch -mt "$(date -r "$back" '+%Y%m%d%H%M.%S')" "$statusf" else touch -m -d "@$back" "$statusf"; fi @@ -5101,6 +5131,7 @@ test_turn_ended_invalid_churn_deadline_surfaced test_turn_ended_surfaced_batch_opens_no_partial_deadline test_working_note_not_working_surfaced test_secondmate_status_note_surfaced_despite_busy_agent +test_secondmate_buried_block_wakes_despite_busy_agent test_self_announced_close_does_not_rewake_but_next_note_does test_actionable_signal_surfaced test_needs_decision_signal_payload_marked_for_branch_exclusion From fa93097162d16f70a070044b8ccece037a38e3e6 Mon Sep 17 00:00:00 2001 From: Cody <72239807+codyjohnsontx@users.noreply.github.com> Date: Thu, 17 Sep 2026 01:53:16 -0500 Subject: [PATCH 03/37] fix(bin): launch codex crewmates with codex's hook layer disabled (#4689) * fix(spawn): launch codex crewmates with codex's hook layer disabled A freshly launched Codex worker never reached its instructions. Codex stopped it on an interactive "Hooks need review" modal whose selection sits on "Review hooks", which is neither trusting nor declining. Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow navigation, so the selection cannot be moved, and pre-accepting the prompt by writing Codex's own trust store would record an operator consent that was never given. The hooks are the machine's own ~/.codex/hooks.json plus any project's .codex/hooks.json. A crewmate needs neither: its turn-end signal is the -c notify= program on the same launch, and Firstmate's project hooks are primary-session infrastructure that stands down in a child worktree. Crewmate and scout launches now pass --disable hooks. That is the opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted hooks; disabling the feature runs none of them and leaves the operator's ~/.codex untouched. An unknown feature name is a hard Codex error, so a release that drops the flag fails the launch loudly instead of silently restoring the modal. A secondmate is a primary in its own home and keeps the project hooks its turn-end guard and session-start digest ride on. Verified on codex-cli 0.151.0: the modal is gone and the turn-end notification still lands. This unblocks the second review that every finished pull request is supposed to get. Fixes kunchenguid/firstmate#4673 * no-mistakes(review): Fix contradictory hook count in Codex verification record --- .../references/harness/codex.md | 9 ++ bin/fm-spawn.sh | 24 ++++- bin/fm-test-run.sh | 3 +- docs/verification/runtime-backends.md | 57 +++++++++++ tests/fm-codex-hook-layer-live-e2e.test.sh | 97 +++++++++++++++++++ tests/fm-spawn-dispatch-profile.test.sh | 46 +++++++++ 6 files changed, 234 insertions(+), 2 deletions(-) create mode 100755 tests/fm-codex-hook-layer-live-e2e.test.sh diff --git a/.agents/skills/harness-adapters/references/harness/codex.md b/.agents/skills/harness-adapters/references/harness/codex.md index 7ae33b57bf5..d68486f12e2 100644 --- a/.agents/skills/harness-adapters/references/harness/codex.md +++ b/.agents/skills/harness-adapters/references/harness/codex.md @@ -20,6 +20,15 @@ A directory trust dialog appears on the first run for a repository root: "Do you Accept it with Enter and verify the instructions begin processing. The decision persists for the repository, so later worktrees of the same project skip it. +## Hook trust + +A second dialog, "Hooks need review - N hooks are new or changed", appears whenever the machine's `~/.codex/hooks.json` or a project's own `.codex/hooks.json` carries a hook Codex has not persisted trust for. +It is unanswerable rather than merely inconvenient: its selection starts on "Review hooks", which is neither trusting nor declining, and Firstmate's key plane carries Enter, Escape and Ctrl-C with no arrow navigation. +Writing Codex's own trust store to pre-accept it would manufacture an operator consent that was never given. +So crewmate and scout launches disable Codex's hook layer outright (`bin/fm-spawn.sh`'s launch template owns the flag), which is the opposite of `--dangerously-bypass-hook-trust` - that flag RUNS the untrusted hooks. +A crewmate loses nothing: its turn-end signal is the `-c notify=` program on the same launch, and the Firstmate hooks in a project's `.codex/hooks.json` are primary-session infrastructure that stands down in a child worktree. +A secondmate is a primary in its own home and keeps its hooks, so an unanswerable modal there is still possible and is the operator's own hook review to settle. + ## Skill popup A `$` invocation opens a `$` autocomplete popup. diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 88fa2fcfabe..d9867cca406 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -1685,11 +1685,33 @@ launch_template() { fi printf '%s' '__MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # --disable hooks (equivalent to -c features.hooks=false) turns codex's whole + # lifecycle-hook layer off for CREWMATE and SCOUT launches only. + # Without it a crewmate launch parks forever on codex's hook-trust modal + # ("N hooks are new or changed"), whose selection sits on "Review hooks" - + # neither trusting nor declining. Firstmate's key plane carries Enter, Escape + # and Ctrl-C with no arrow navigation, so the selection cannot be moved, and + # pre-accepting the prompt by writing codex's own trust store would manufacture + # an operator consent that was never given. The hooks it asks about are the + # OPERATOR's machine-level ~/.codex/hooks.json plus any project-local + # .codex/hooks.json, and a crewmate needs none of them: its turn-end signal is + # the -c notify= program on this same launch (verified still firing with hooks + # disabled, codex-cli 0.151.0), and firstmate's own .codex/hooks.json registers + # PRIMARY-session infrastructure that already stands down in a child worktree. + # This is the opposite of --dangerously-bypass-hook-trust, which RUNS untrusted + # hooks; disabling the feature runs none of them and leaves the operator's + # ~/.codex untouched. An unknown feature name is a hard codex error, so a future + # release that drops this flag fails the launch loudly instead of silently + # restoring the modal. + # A secondmate is a firstmate PRIMARY in its own home, and its turn-end guard, + # session-start digest, and cd/arm seatbelts are exactly those project hooks + # (docs/turnend-guard.md, docs/sessionstart-nudge.md, docs/cd-guard.md), so the + # secondmate launch deliberately keeps hooks on. codex) if [ "$kind" = secondmate ]; then printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' else - printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox -c "notify=[\"bash\",\"-c\",\"touch __TURNEND__\"]" "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox --disable hooks -c "notify=[\"bash\",\"-c\",\"touch __TURNEND__\"]" "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' fi ;; opencode) printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' opencode __MODELFLAG__--prompt "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 0bfc3e941ec..958e96740c4 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -344,7 +344,8 @@ family_for_basename() { fm-cmux-claude-composer-live-e2e.test.sh|\ fm-composer-matrix-live-e2e.test.sh|\ fm-composer-codex-idle-live-e2e.test.sh|\ - fm-codex-continuity-live-e2e.test.sh|fm-grok-continuity-live-e2e.test.sh|\ + fm-codex-continuity-live-e2e.test.sh|fm-codex-hook-layer-live-e2e.test.sh|\ + fm-grok-continuity-live-e2e.test.sh|\ fm-cursor-primary-live-e2e.test.sh|\ fm-grok-stop-live-e2e.test.sh|fm-harness-adapter-instructions-live-e2e.test.sh|\ fm-harness-liveness-drift-live-e2e.test.sh|\ diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index c7c5f183a52..8cb1627c1dc 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -508,6 +508,63 @@ The lab home was deleted and the test entry was removed from the store and verif That automated spawn case runs against a fake claude, so it asserts the store entry and the launch command and nothing more; the live arms above are what establish that the entry actually suppresses the dialog. The composer-classification record below observes the same gate from the other side, where an untrusted worktree left Claude, Grok, and Muse unverified because the guard reads a first-launch trust dialog as an unreadable composer. +## Codex hook trust + +Verified 2026-09-16 on codex-cli 0.151.0, macOS arm64, in a fresh linked worktree of this repository. + +Codex gates hooks it has no persisted trust for behind an interactive modal. +A crewmate launch built by `bin/fm-spawn.sh` was driven under a real PTY and stopped there before the brief was ever submitted: + +```text +Hooks need review +12 hooks are new or changed. +Hooks can run outside the sandbox after you trust them. +> 1. Review hooks + 2. Trust all and continue + 3. Continue without trusting (hooks won't run) +Press enter to confirm or esc to go back +``` + +The selection starts on "Review hooks", which is neither trusting nor declining, and Firstmate's key plane carries only Enter, Escape, and C-c with no arrow navigation, so the selection cannot be moved. +That count covers every hook Codex had no persisted trust for, drawn from both the machine's own `~/.codex/hooks.json` and this repository's tracked `.codex/hooks.json`. +Writing Codex's own trust store to pre-accept the modal would record an operator consent that was never given, so it is not an option either. + +`codex --help` documents `--dangerously-bypass-hook-trust` as "Run enabled hooks without requiring persisted hook trust for this invocation", which RUNS the untrusted hooks. +That is the opposite of what an unattended worker needs, so the control used is the hook feature flag: + +```sh +codex features list | grep '^hooks' +codex --disable hooks features list | grep '^hooks' +codex --disable no_such_feature features list +``` + +```text +hooks stable true +hooks stable false +Error: Unknown feature flag: no_such_feature +``` + +The last arm is what makes the control safe to depend on: an unknown feature name is a hard error, so a release that renames or drops the flag fails the launch loudly instead of silently restoring the modal. + +The same launch with the hook layer disabled reached the composer with no modal, answered the prompt, and fired the turn-end program that rides the launch rather than any hook: + +```sh +codex --dangerously-bypass-approvals-and-sandbox --disable hooks \ + -c "notify=[\"bash\",\"-c\",\"touch $TURNEND\"]" "Say ACK and stop." +``` + +```text +> Say ACK and stop. +- ACK, captain. +$ ls "$TURNEND" + +``` + +`tests/fm-codex-hook-layer-live-e2e.test.sh` is the command that refreshes this record. +It captures the launch `bin/fm-spawn.sh` actually builds, replays those exact flags against the installed Codex, and fails naming the harness and version if the hook layer comes back on. +It spends no model tokens, so it runs by default wherever Codex is installed. +The portable half, `tests/fm-spawn-dispatch-profile.test.sh`, pins the split the launch template makes: a crewmate launches hook-free while a secondmate, which runs a primary session on this repository's own project hooks, keeps them. + ## Composer classification matrix The shared composer classifier (`bin/fm-composer-lib.sh`, `fm_composer_classify_screen`) owns every composer shape fleet-wide; each backend contributes only a capture and a capability descriptor. diff --git a/tests/fm-codex-hook-layer-live-e2e.test.sh b/tests/fm-codex-hook-layer-live-e2e.test.sh new file mode 100755 index 00000000000..ff47efe1642 --- /dev/null +++ b/tests/fm-codex-hook-layer-live-e2e.test.sh @@ -0,0 +1,97 @@ +#!/usr/bin/env bash +# Live guard for the codex crewmate launch's hook posture. +# +# The verdict here comes from the installed codex, not from a stub: a stub can +# only confirm the assumption already written into it, and what this guard +# protects is exactly a vendor-owned surface. Codex blocks a fresh crewmate +# launch on an unanswerable "Hooks need review" modal whenever the machine's +# ~/.codex/hooks.json or a project's .codex/hooks.json carries a hook it has no +# persisted trust for, so the crewmate launch disables codex's hook layer +# outright (bin/fm-spawn.sh's launch template owns the flag). +# +# The guard replays the REAL launch flags fm-spawn builds - captured from a +# spawn driven through a fake pane - against the installed codex and asks codex +# itself whether hooks ended up disabled. If a codex release renames or drops +# the feature, the flag becomes a hard "Unknown feature flag" error and this +# guard fails naming the harness and version instead of letting the modal +# silently come back. +# +# It spends no model tokens (`codex features list` resolves configuration only), +# so it runs by default wherever codex is installed. +set -u + +# shellcheck source=tests/fixtures.sh +. "$(dirname "${BASH_SOURCE[0]}")/fixtures.sh" + +fm_live_gate default-on FM_CODEX_HOOK_LAYER_LIVE codex + +CODEX_VERSION=$(codex --version 2>&1) +TMP_ROOT=$(fm_test_tmproot fm-codex-hook-layer-live) + +# capture_codex_launch : spawns a codex crewmate +# against a fake pane and echoes the literal launch command firstmate sent. +capture_codex_launch() { + local name=$1 + shift + local case_dir home proj wt fakebin launchlog id + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + proj="$case_dir/project" + wt="$case_dir/wt" + launchlog="$case_dir/launch.log" + id="codex-hook-layer-$name" + fakebin=$(fm_test_make_spawn_fakebin "$case_dir/fake") + fm_test_spawn_home "$home" codex + fm_test_spawn_brief "$home" "$id" + fm_git_worktree "$proj" "$wt" "wt-$name" + : > "$launchlog" + FM_FAKE_LAUNCH_LOG="$launchlog" \ + fm_test_run_spawn "$home" "$wt" "$fakebin" "$id" "$proj" "$@" >/dev/null 2>&1 || + fail "codex $CODEX_VERSION: fm-spawn could not build a crewmate launch" + cat "$launchlog" +} + +# codex_global_flags : the flags between the codex executable +# and the positional brief, which is everything codex itself is configured by. +codex_global_flags() { + local launch=$1 flags + flags=${launch#*codex } + flags=${flags%%\"\$(*} + printf '%s' "$flags" +} + +test_installed_codex_disables_hooks_for_the_captured_crewmate_launch() { + local launch flags state + launch=$(capture_codex_launch ship --mode no-mistakes --yolo off) + flags=$(codex_global_flags "$launch") + + # The whole point: every flag firstmate will launch with, handed to the real + # codex, must leave the hook layer off. `features list` reports the effective + # state after those flags are applied and contacts no model. + state=$(eval "codex $flags features list" 2>&1) || + fail "codex $CODEX_VERSION rejected firstmate's crewmate launch flags: $state" + case "$state" in + *"Unknown feature flag"*) + fail "codex $CODEX_VERSION no longer knows the hook feature firstmate disables: $state" + ;; + esac + printf '%s\n' "$state" | awk '$1 == "hooks" { print $NF }' | grep -qx false || + fail "codex $CODEX_VERSION left hooks enabled for firstmate's crewmate launch flags, so a fresh launch can park on the hook-trust modal" + + printf 'ok - codex %s runs a firstmate crewmate launch with its hook layer disabled\n' "$CODEX_VERSION" +} + +test_installed_codex_still_reports_the_hook_feature() { + local listing + listing=$(codex features list 2>&1) || + fail "codex $CODEX_VERSION could not list its feature flags: $listing" + printf '%s\n' "$listing" | awk '{ print $1 }' | grep -qx hooks || + fail "codex $CODEX_VERSION no longer publishes a hook feature flag; firstmate's crewmate launch needs a new control" + + printf 'ok - codex %s still publishes the hook feature flag firstmate disables\n' "$CODEX_VERSION" +} + +test_installed_codex_still_reports_the_hook_feature +test_installed_codex_disables_hooks_for_the_captured_crewmate_launch + +echo "# all fm-codex-hook-layer-live-e2e tests passed" diff --git a/tests/fm-spawn-dispatch-profile.test.sh b/tests/fm-spawn-dispatch-profile.test.sh index 745a1381552..f7274693ae7 100755 --- a/tests/fm-spawn-dispatch-profile.test.sh +++ b/tests/fm-spawn-dispatch-profile.test.sh @@ -451,6 +451,50 @@ test_codex_omits_max_effort_for_unsupported_model() { pass "codex omits max for models without the catalog capability" } +# Codex parks a crewmate launch forever on its unanswerable hook-trust modal +# unless the launch turns the hook layer off. These two cases pin the split: +# a crewmate runs hook-free, a secondmate keeps the project hooks that carry its +# own primary-session turn-end guard and session-start digest. +test_codex_crewmate_launch_disables_the_hook_layer() { + local rec id out status launch + id=profile-codex-hooks-z4c + rec=$(make_spawn_case profile-codex-hooks codex "$id") + read_case_record "$rec" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "codex crewmate spawn should succeed"$'\n'"$out" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "--disable hooks" \ + "codex crewmate launch did not disable the hook layer that blocks it on a trust modal" + # The opposite posture: this flag RUNS the untrusted hooks instead of + # disabling them, so a launch must never reach for it. + assert_not_contains "$launch" "--dangerously-bypass-hook-trust" \ + "codex crewmate launch ran the operator's untrusted hooks instead of disabling them" + # Firstmate goes blind without the turn-end signal, which rides this same + # launch rather than any hook. + assert_contains "$launch" "notify=" \ + "codex crewmate launch lost the turn-end notify program" + pass "a codex crewmate launches with no hook layer and keeps its turn-end signal" +} + +test_codex_secondmate_launch_keeps_the_hook_layer() { + local rec id sm out status launch + id=profile-codex-secondmate-hooks-z4d + rec=$(make_spawn_case profile-codex-secondmate-hooks codex "$id") + read_case_record "$rec" + sm="$CASE_DIR/secondmate-home" + make_seeded_secondmate_home "$sm" "$id" + + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$sm" --secondmate) + status=$? + expect_code 0 "$status" "codex secondmate spawn should succeed"$'\n'"$out" + launch=$(cat "$LAUNCH_LOG") + assert_not_contains "$launch" "--disable hooks" \ + "codex secondmate launch disabled the project hooks its own primary supervision depends on" + pass "a codex secondmate keeps the project hook layer its primary session runs on" +} + test_grok_threads_model_and_reasoning_effort() { local rec id out status launch id=profile-grok-z5 @@ -1386,6 +1430,8 @@ test_claude_threads_model_and_effort test_codex_threads_model_and_effort test_codex_threads_model_and_max_effort test_codex_omits_max_effort_for_unsupported_model +test_codex_crewmate_launch_disables_the_hook_layer +test_codex_secondmate_launch_keeps_the_hook_layer test_grok_threads_model_and_reasoning_effort test_grok_omits_invalid_max_reasoning_effort test_grok_omits_invalid_xhigh_reasoning_effort From 3eb5b6334a80e06083e3837f0032a5cec39b8e52 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Micka=C3=ABl=20R=C3=A9mond?= Date: Thu, 17 Sep 2026 12:21:19 +0200 Subject: [PATCH 04/37] fix(bin): settle terminal contribution observations (Fixes #4669, Fixes #4670) (#4710) * fix(bin): settle terminal contributions and wake once per read-failure episode A contribution whose last good observation is merged or closed is final: poll no longer re-reads it, projection keeps it fresh, and a stale error recorded beside it is cleared once. A genuine forge-read failure on an open contribution still records its error on every cycle but prints the unavailable wake only when it starts a failure episode; a successful read ends the episode. Open PRs linked from done tasks keep being observed. The false unavailable beside a complete observation was budget exhaustion mid-observation, already fixed by #4661. * fix(review): Settle terminal contribution owners * fix(review): Deduplicate shared contribution failure episodes * fix(test): Preserve settled terminal contribution records --- bin/fm-contributions.jq | 8 +- bin/fm-contributions.sh | 46 +++++++++- tests/fm-contributions.test.sh | 152 ++++++++++++++++++++++++++++++++- 3 files changed, 198 insertions(+), 8 deletions(-) diff --git a/bin/fm-contributions.jq b/bin/fm-contributions.jq index 3b3f16fcf4b..3b1748802b9 100644 --- a/bin/fm-contributions.jq +++ b/bin/fm-contributions.jq @@ -47,7 +47,9 @@ def projected($input; $saved; $now; $max_age): | ($record.observation // {}) as $o | (if $record.error == null and $record.observation != null and ($o.head | sha) then $o.head else null end) as $observed_head | (($record.checked_at // "") | try fromdateiso8601 catch null) as $checked - | ($checked != null and ($now - $checked) >= 0 and ($now - $checked) <= $max_age + # A merged or closed observation is final; poll never re-reads it, so it never expires. + | ($record.error == null and ($o.state | IN("merged","closed"))) as $final + | (($final or ($checked != null and ($now - $checked) >= 0 and ($now - $checked) <= $max_age)) and (if $record.kind == "pr" then $observed_head != null else $record.error == null and $record.observation != null end) and ($k.url | startswith("https://github.com/"))) as $fresh @@ -91,7 +93,7 @@ def projected($input; $saved; $now; $max_age): elif $o.can_merge == true then {actor:"captain",reason:"checks green; merge approval needed"} else {actor:"maintainer",reason:"delivery awaits the maintainer"} end) as $action | $k + {kind:($record.kind // (if ($k.url | contains("/issues/")) then "issue" else "pr" end)), - checked_at:$record.checked_at,checked:$fresh,head:($observed_head // $recorded_head // $o.head),verdict:$verdict,reviews:$reviews, + checked_at:$record.checked_at,checked:$fresh,final:$final,head:($observed_head // $recorded_head // $o.head),verdict:$verdict,reviews:$reviews, distinct_checks:($checks | length),missing_verdicts:(($no_verdict | length) + (($o.absent_checks // []) | length)), pending_checks:($pending | length),failed_checks:($failed | length), stale_verdicts:((if $stale then 1 else 0 end) + ([$reviews[] | select(.freshness == "STALE")] | length)), @@ -113,6 +115,6 @@ def summary($rows; $errors): stale_verdicts:([$rows[].stale_verdicts] | add // 0), missing_verdicts:([$rows[].missing_verdicts] | add // 0), unreadable_records:$errors, - valid_until:([$rows[].checked_at | try (fromdateiso8601) catch 0] | min // 0), + valid_until:([$rows[] | select(.final | not) | .checked_at | try (fromdateiso8601) catch 0] | min // 0), captain:[$rows[] | select(.actor == "captain") | {task,url,kind,head,reason:(.reason[:240]),hold, verdict_freshness:.verdict.freshness,verdict_head:.verdict.head,verdict_source:.verdict.source,checked_at}]}; diff --git a/bin/fm-contributions.sh b/bin/fm-contributions.sh index 4481c913db6..f0c9949ffbd 100755 --- a/bin/fm-contributions.sh +++ b/bin/fm-contributions.sh @@ -34,11 +34,16 @@ # and spends at most FM_CONTRIBUTIONS_BUDGET seconds on forge reads (default 20, # 1..25). Each gh call is bounded by the remaining budget and five seconds. # Oldest observations go first, so a large corpus progresses across polls. -# Each distinct URL is observed once per poll and applied to every owner. When +# Each distinct URL is observed once per poll and applied to every owner. A +# final observation applies to every owner without another forge read. When # the budget runs out mid-observation, the poll ends with that URL's records # untouched; only a genuine forge failure or head change records an error. # API failure leaves error evidence; an expired or absent observation is not # silence. FM_CONTRIBUTIONS_MAX_AGE (default 900 seconds) bounds freshness. +# A URL whose last good observation is merged or closed is final: it is +# never re-read, stays fresh, and a stale error beside it is cleared once. +# A genuine failure prints its unavailable line only when it starts an episode +# (no prior owner has an error); a successful read ends the episode. # FM_CONTRIBUTIONS_NOW supplies an ISO UTC clock for tests, otherwise UTC now. # FM_CONTRIBUTIONS_READY_LABEL selects the equivalent triage label, default # ready-for-pr. Labels are matched case-insensitively and exactly. @@ -129,7 +134,8 @@ project() { --arg all "${2:-}" ' projected($input[0];$saved[0];$now;$max_age) as $rows | summary($rows;($errors + (if $input[0].backlog.present == true then 0 else 1 end))) - | .valid_until += $max_age + # Final rows never expire; a home holding only final rows is valid from now. + | .valid_until = (if ($rows | length) > 0 and all($rows[]; .final) then $now else .valid_until end) + $max_age | .captain_omitted = ([0, (.captain | length) - 20] | max) | .captain |= .[:20] | . + (if $all == "--all" then {rows:$rows} else {} end)' @@ -262,6 +268,28 @@ publish_pending() { # task canonical-url record-file done < <(jq -r '. as $r | .pending[] | .token | select(. as $t | ($r.notified // [] | index($t)) == null)' "$record") } +settle_final() { # canonical-url task... : copy the URL's final observation to every owner + local url=$1 task + shift + jq -n --slurpfile saved "$TMP/saved.json" --arg url "$url" ' + [$saved[0][] | .records[] | select(.url == $url + and (.observation.state | IN("merged","closed")))] as $final + | ([$final[] | select(.error == null)] | first) // ($final | first)' > "$TMP/final.json" + for task in "$@"; do + fm_pr_task_id_valid "$task" || { printf 'contributions: invalid durable task id\n'; continue; } + jq -n --slurpfile saved "$TMP/saved.json" --arg task "$task" --arg url "$url" ' + [$saved[0][] | select(.task == $task) | .records[] | select(.url == $url)] | first' > "$TMP/old.json" + if jq -e '. == null' "$TMP/old.json" >/dev/null; then + jq -n --slurpfile final "$TMP/final.json" ' + $final[0] + {error:null,pending:[],notified:[]}' > "$TMP/row.json" + write_record "$task" "$TMP/row.json" + elif jq -e '.error != null' "$TMP/old.json" >/dev/null; then + jq '.error = null' "$TMP/old.json" > "$TMP/row.json" + write_record "$task" "$TMP/row.json" + fi + done +} + poll() { local task url old kind error observed local -a row @@ -280,12 +308,24 @@ poll() { [ "${#row[@]}" -ge 2 ] || continue [ "$(date +%s)" -lt "$DEADLINE" ] || break url=${row[0]} + # A contribution with a final observation is not re-read for any owner. + if jq -ne --slurpfile saved "$TMP/saved.json" --arg url "$url" --args \ + 'any($ARGS.positional[] as $task | [$saved[0][] | select(.task == $task) | .records[] | select(.url == $url)] | first; + . != null and (.observation.state | IN("merged","closed")))' "${row[@]:1}" >/dev/null; then + settle_final "$url" "${row[@]:1}" + continue + fi observed=0 observe "$url" || observed=$? # An observation the budget cut short is unmeasured, not unavailable: keep # every owner's prior record so the URL is observed first next poll. [ "$BUDGET_EXHAUSTED" -eq 0 ] || break - [ "$observed" -eq 0 ] || printf 'contributions: observation unavailable for %s\n' "$url" + # Wake once per failure episode: only when no owner has a prior error. + if [ "$observed" -ne 0 ] && jq -ne --slurpfile saved "$TMP/saved.json" --arg url "$url" --args \ + 'all($ARGS.positional[] as $task | [$saved[0][] | select(.task == $task) | .records[] | select(.url == $url)] | first; + .error == null)' "${row[@]:1}" >/dev/null; then + printf 'contributions: observation unavailable for %s\n' "$url" + fi case "$url" in */issues/*) kind=issue ;; *) kind="pr" ;; esac for task in "${row[@]:1}"; do fm_pr_task_id_valid "$task" || { printf 'contributions: invalid durable task id\n'; continue; } diff --git a/tests/fm-contributions.test.sh b/tests/fm-contributions.test.sh index c17ecebc08a..24e1d3b9909 100755 --- a/tests/fm-contributions.test.sh +++ b/tests/fm-contributions.test.sh @@ -124,7 +124,10 @@ case "$*" in 'pr view '*headRefOid*) cat "$FORGE/head" ;; 'pr view '*state*) printf 'OPEN\n' ;; 'api repos/o/r/pulls/8') - jq -n --arg head "$(cat "$FORGE/head")" '{state:"open",user:{login:"author"},head:{sha:$head},draft:false,mergeable:true,merged_at:null}' ;; + jq -n --arg head "$(cat "$FORGE/head")" --arg state "$(cat "$FORGE/state" 2>/dev/null || printf open)" ' + {state:(if $state == "open" then "open" else "closed" end),user:{login:"author"},head:{sha:$head},draft:false, + mergeable:(if $state == "open" then true else null end), + merged_at:(if $state == "merged" then "2026-09-16T07:00:00Z" else null end)}' ;; 'api repos/o/r/issues/9') jq -n --slurpfile labels "$FORGE/labels.json" '{state:"open",user:{login:"author"},labels:$labels[0]}' ;; 'api repos/o/r/issues/'*'/events?'*) jq -s . "$FORGE/events.json" ;; @@ -560,6 +563,7 @@ case "$fault:$*" in printf '%s\n' "$(( $(cat "$FORGE/clock") + 100 ))" > "$FORGE/clock" printf 'HTTP 502\n' >&2; exit 1 ;; fail:'api repos/o/r/pulls/8/reviews?'*) printf 'HTTP 502\n' >&2; exit 1 ;; + down:*) printf 'HTTP 502\n' >&2; exit 1 ;; hang:'api repos/o/r/pulls/8') sleep 4 ;; head:'pr view '*) printf '{"headRefOid":"%s","reviewDecision":"APPROVED"}\n' "$(printf 'b%.0s' $(seq 40))"; exit 0 ;; esac @@ -640,8 +644,152 @@ test_shared_url_observed_once() { pass 'a URL owned by two tasks is observed once and every owner receives the result' } +test_terminal_contribution_settles() { + local mode home out later=2026-09-17T08:00:00Z + for mode in merged closed; do + home=$(new_home "terminal-$mode") + forge_home "$home" + wrap_forge "$home" + printf '%s\n' "$mode" > "$home/forge/state" + mutate_record "$home" delivery '.records[0].checked_at="2026-09-15T08:00:00Z"' + out=$(with_home "$home" "$ROOT/bin/fm-contributions.sh" poll) || fail "terminal observation poll failed ($mode)" + [ -z "$out" ] || fail "a $mode observation printed: $out" + jq -e --arg now "$NOW" --arg mode "$mode" '.records[0] | .checked_at == $now and .error == null and .observation.state == $mode' \ + "$home/data/delivery/contributions.json" >/dev/null || fail "a $mode observation was not recorded once without error" + cp "$home/data/delivery/contributions.json" "$home/prior.json" + : > "$home/forge/calls" + printf 'down\n' > "$home/forge/fault" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW="$later" "$ROOT/bin/fm-contributions.sh" poll) \ + || fail "poll after a $mode observation failed" + [ -z "$out" ] || fail "a $mode contribution woke again when a later read would fail: $out" + [ ! -s "$home/forge/calls" ] || fail "a $mode contribution was re-read: $(cat "$home/forge/calls")" + cmp -s "$home/prior.json" "$home/data/delivery/contributions.json" \ + || fail "a $mode contribution record changed after it settled: $(cat "$home/data/delivery/contributions.json")" + [ ! -s "$home/state/.wake-queue" ] || fail "a $mode contribution enqueued a wake" + NOW=$later bearings "$home" | jq -e '.contributions.checked == 1 and .contributions.counts.nobody == 1 + and .contributions.complete == true' >/dev/null \ + || fail "a settled $mode contribution expired into fleet work" + done + home=$(new_home terminal-legacy-error) + forge_home "$home" + wrap_forge "$home" + mutate_record "$home" delivery '.records[0].observation.state="merged" | .records[0].error="forge observation unavailable or changed during read"' + printf 'down\n' > "$home/forge/fault" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW="$later" "$ROOT/bin/fm-contributions.sh" poll) \ + || fail 'poll of an error-stamped merged record failed' + [ -z "$out" ] || fail "an error-stamped merged record woke again: $out" + [ ! -s "$home/forge/calls" ] || fail 'an error-stamped merged record was re-read' + jq -e --arg at "$NOW" '.records[0] | .error == null and .checked_at == $at and .observation.state == "merged"' \ + "$home/data/delivery/contributions.json" >/dev/null || fail 'an error-stamped merged record did not settle' + pass 'a merged or closed contribution settles once, is not re-read, and never wakes again' +} + +test_late_owner_inherits_terminal_observation() { + local home out later=2026-09-17T08:00:00Z + home=$(new_home terminal-late-owner) + forge_home "$home" + wrap_forge "$home" + printf 'merged\n' > "$home/forge/state" + with_home "$home" "$ROOT/bin/fm-contributions.sh" poll >/dev/null || fail 'initial terminal observation poll failed' + cp "$home/data/delivery/contributions.json" "$home/final.json" + printf -- '- [ ] duplicate - Filed https://github.com/o/r/pull/8 (repo: sample) (kind: ship)\n' >> "$home/data/backlog.md" + : > "$home/forge/calls" + printf 'down\n' > "$home/forge/fault" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW="$later" "$ROOT/bin/fm-contributions.sh" poll) \ + || fail 'late-owner terminal poll failed' + [ -z "$out" ] || fail "a late owner reactivated a terminal contribution: $out" + [ ! -s "$home/forge/calls" ] || fail 'a late owner triggered a terminal forge read' + jq -e --slurpfile final "$home/final.json" ' + .records[0] as $late | $final[0].records[0] as $terminal + | $late.error == null and $late.pending == [] and $late.notified == [] + and $late.checked_at == $terminal.checked_at and $late.observation == $terminal.observation' \ + "$home/data/duplicate/contributions.json" >/dev/null \ + || fail 'a late owner did not inherit the settled terminal observation' + [ ! -s "$home/state/.wake-queue" ] || fail 'a late owner terminal record enqueued a wake' + pass 'a late owner inherits a terminal observation without a forge read or wake' +} + +test_done_task_open_pr_still_observed() { + local home later=2026-09-17T08:00:00Z + home=$(new_home done-open) + forge_home "$home" + wrap_forge "$home" + rm "$home/data/delivery/contributions.json" + printf '# Backlog\n\n## Queued\n\n## Done\n- [x] delivery - Shipped https://github.com/o/r/pull/8 (repo: sample) (kind: ship)\n' \ + > "$home/data/backlog.md" + with_home "$home" "$ROOT/bin/fm-contributions.sh" poll >/dev/null || fail 'poll of a done task failed' + printf '%s\n' "$HEAD_B" > "$home/forge/head" + with_home "$home" env FM_CONTRIBUTIONS_NOW="$later" "$ROOT/bin/fm-contributions.sh" poll >/dev/null \ + || fail 'second poll of a done task failed' + [ "$(grep -cFx 'api repos/o/r/pulls/8' "$home/forge/calls")" = 2 ] \ + || fail 'an open PR linked from a done task was not observed on every poll' + jq -e --arg head "$HEAD_B" --arg at "$later" '.records[0] | .checked_at == $at and .error == null + and .observation.state == "open" and .observation.head == $head' \ + "$home/data/delivery/contributions.json" >/dev/null || fail 'an open PR on a done task did not track its current head' + pass 'an open PR linked from a done task keeps being observed' +} + +test_failure_wakes_once_per_episode() { + local home out line='contributions: observation unavailable for https://github.com/o/r/pull/8' + local error='"forge observation unavailable or changed during read"' + home=$(new_home failure-episode) + forge_home "$home" + wrap_forge "$home" + printf 'down\n' > "$home/forge/fault" + poll_at() { with_home "$home" env FM_CONTRIBUTIONS_NOW="$1" "$ROOT/bin/fm-contributions.sh" poll || fail "poll at $1 failed"; } + out=$(poll_at 2026-09-16T09:00:00Z) + [ "$out" = "$line" ] || fail "the first failure of an episode did not wake: $out" + out=$(poll_at 2026-09-16T10:00:00Z) + [ -z "$out" ] || fail "an unchanged read failure woke again on the next cycle: $out" + jq -e --argjson error "$error" '.records[0] | .checked_at == "2026-09-16T10:00:00Z" and .error == $error' \ + "$home/data/delivery/contributions.json" >/dev/null || fail 'a repeated read failure stopped recording its error' + [ "$(grep -cFx 'api repos/o/r/pulls/8' "$home/forge/calls")" = 2 ] || fail 'a failing open PR stopped being observed' + : > "$home/forge/fault" + out=$(poll_at 2026-09-16T11:00:00Z) + [ -z "$out" ] || fail "a successful read printed: $out" + jq -e '.records[0].error == null' "$home/data/delivery/contributions.json" >/dev/null \ + || fail 'a successful read did not end the failure episode' + printf 'down\n' > "$home/forge/fault" + out=$(poll_at 2026-09-16T12:00:00Z) + [ "$out" = "$line" ] || fail "a new failure after a successful read did not wake: $out" + pass 'a repeated read failure on an open PR records its error but wakes once per episode' +} + +test_late_owner_keeps_failure_episode_suppressed() { + local home out line='contributions: observation unavailable for https://github.com/o/r/pull/8' + local error='forge observation unavailable or changed during read' task + home=$(new_home late-owner-failure-episode) + forge_home "$home" + wrap_forge "$home" + printf 'down\n' > "$home/forge/fault" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW=2026-09-16T09:00:00Z "$ROOT/bin/fm-contributions.sh" poll) \ + || fail 'initial failing poll failed' + [ "$out" = "$line" ] || fail "the initial failure did not wake: $out" + printf -- '- [ ] duplicate - Filed https://github.com/o/r/pull/8 (repo: sample) (kind: ship)\n' >> "$home/data/backlog.md" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW=2026-09-16T10:00:00Z "$ROOT/bin/fm-contributions.sh" poll) \ + || fail 'late-owner failing poll failed' + [ -z "$out" ] || fail "a late owner restarted an unchanged failure episode: $out" + for task in delivery duplicate; do + jq -e --arg error "$error" '.records[0].error == $error' "$home/data/$task/contributions.json" >/dev/null \ + || fail "owner $task did not retain the shared failure evidence" + done + : > "$home/forge/fault" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW=2026-09-16T11:00:00Z "$ROOT/bin/fm-contributions.sh" poll) \ + || fail 'successful shared poll failed' + [ -z "$out" ] || fail "a successful shared poll printed: $out" + for task in delivery duplicate; do + jq -e '.records[0].error == null' "$home/data/$task/contributions.json" >/dev/null \ + || fail "owner $task did not end the shared failure episode" + done + printf 'down\n' > "$home/forge/fault" + out=$(with_home "$home" env FM_CONTRIBUTIONS_NOW=2026-09-16T12:00:00Z "$ROOT/bin/fm-contributions.sh" poll) \ + || fail 'new shared failing poll failed' + [ "$out" = "$line" ] || fail "a failure after shared recovery did not wake: $out" + pass 'a late owner does not restart a shared forge failure episode' +} + failures=0 -for test_name in test_actor_coverage test_stale_verdict test_unchecked_is_not_silence test_newest_check_has_no_verdict test_comment_wake test_review_wake test_inline_wake test_ready_issue_wake test_fresh_issue_requires_maintainer test_missing_lane_remains_missing test_partial_freshness_keeps_measured_rows test_malformed_record_cannot_prove_silence test_issue_timeline_and_exact_ack test_verdict_retains_judged_head test_observed_replacement_refreshes_verdict test_unobserved_head_leaves_verdict_unknown test_away_yolo_is_fleet_work test_away_yolo_cross_home_is_fleet_work test_retired_and_unsupported_coverage test_unsupported_forge_is_not_fleet_work test_held_unsupported_forge_is_not_captain_work test_shared_contribution_signal_wakes_once test_watcher_keeps_diagnostics_separate_from_contribution_wakes test_expired_child_unsupported_forge_stays_unmeasured test_watcher_surfaces_new_contribution_once test_home_summary_coverage test_unreadable_pending_is_not_empty test_budget_refusal_between_calls test_budget_bounded_call_timeout test_genuine_failure_near_deadline_is_unavailable test_shared_url_observed_once; do +for test_name in test_actor_coverage test_stale_verdict test_unchecked_is_not_silence test_newest_check_has_no_verdict test_comment_wake test_review_wake test_inline_wake test_ready_issue_wake test_fresh_issue_requires_maintainer test_missing_lane_remains_missing test_partial_freshness_keeps_measured_rows test_malformed_record_cannot_prove_silence test_issue_timeline_and_exact_ack test_verdict_retains_judged_head test_observed_replacement_refreshes_verdict test_unobserved_head_leaves_verdict_unknown test_away_yolo_is_fleet_work test_away_yolo_cross_home_is_fleet_work test_retired_and_unsupported_coverage test_unsupported_forge_is_not_fleet_work test_held_unsupported_forge_is_not_captain_work test_shared_contribution_signal_wakes_once test_watcher_keeps_diagnostics_separate_from_contribution_wakes test_expired_child_unsupported_forge_stays_unmeasured test_watcher_surfaces_new_contribution_once test_home_summary_coverage test_unreadable_pending_is_not_empty test_budget_refusal_between_calls test_budget_bounded_call_timeout test_genuine_failure_near_deadline_is_unavailable test_shared_url_observed_once test_terminal_contribution_settles test_late_owner_inherits_terminal_observation test_done_task_open_pr_still_observed test_failure_wakes_once_per_episode test_late_owner_keeps_failure_episode_suppressed; do ( "$test_name" ) || failures=$((failures + 1)) done [ "$failures" -eq 0 ] || fail "$failures contribution regressions" From f5d7f5f2484564dd855b76e7e40ef8c40dc7ab2b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Micka=C3=ABl=20R=C3=A9mond?= Date: Thu, 17 Sep 2026 20:34:17 +0200 Subject: [PATCH 05/37] fix: select authoritative no-mistakes runs (#4476) * fix(crew-state): select authoritative validation runs by identity Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row. Refs: https://github.com/kunchenguid/firstmate/issues/3215 * fix(review): Resolve same-branch run identities beyond capped history * fix(review): Fix run-selection compatibility, races, and worker-state fallbacks * fix(review): Limit run validation to the requested branch * fix(test): Anchor AXI fixtures and document remaining live evidence gaps * fix(document): Clarify run selection documentation and capture ownership * fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation --- bin/fm-crew-state.sh | 168 +++-- bin/fm-nm-run-lib.sh | 222 ++++-- docs/architecture.md | 5 +- docs/configuration.md | 2 +- docs/documentation-audiences.json | 4 + tests/captures/no-mistakes-v1.70.1/README.md | 51 ++ .../no-mistakes-v1.70.1/completed.toon | 20 + .../captures/no-mistakes-v1.70.1/failed.toon | 19 + .../no-mistakes-v1.70.1/overview.toon | 12 + .../captures/no-mistakes-v1.70.1/parked.toon | 26 + .../no-mistakes-v1.70.1/replacement.toon | 20 + .../same-branch-inventory.json | 74 ++ .../no-mistakes-v1.70.1/superseded.toon | 20 + .../no-mistakes-v1.70.1/uninitialized.toon | 2 + tests/fm-crew-state.test.sh | 712 ++++++++++++++++-- 15 files changed, 1199 insertions(+), 158 deletions(-) create mode 100644 tests/captures/no-mistakes-v1.70.1/README.md create mode 100644 tests/captures/no-mistakes-v1.70.1/completed.toon create mode 100644 tests/captures/no-mistakes-v1.70.1/failed.toon create mode 100644 tests/captures/no-mistakes-v1.70.1/overview.toon create mode 100644 tests/captures/no-mistakes-v1.70.1/parked.toon create mode 100644 tests/captures/no-mistakes-v1.70.1/replacement.toon create mode 100644 tests/captures/no-mistakes-v1.70.1/same-branch-inventory.json create mode 100644 tests/captures/no-mistakes-v1.70.1/superseded.toon create mode 100644 tests/captures/no-mistakes-v1.70.1/uninitialized.toon diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 8ef77cf25dd..160c729ed67 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -51,10 +51,10 @@ # before it having ended at exactly this worktree's head - so an active fix # round never reads as an older failed run (rule owned by # fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh). -# More than one recorded run can bind to this worktree at once, and -# bin/fm-nm-run-lib.sh also owns which of them wins: a LIVE run always -# outranks a terminal one, so a terminal answer here is provisional until -# the ledger has been asked whether a live sibling run exists. +# fm_nm_select_run in bin/fm-nm-run-lib.sh owns complete run selection +# and ambiguity reporting. The selected run's id-addressed status must +# agree on id, branch, and live/terminal class before attribution; +# disagreement reports unknown with available candidate ids. # The run-step is AUTHORITATIVE: running/fixing -> working, ci -> working, # awaiting_approval/fix_review -> parked (with gate findings), terminal # passed/checks-passed -> done, failed/cancelled -> failed. EXCEPT: while @@ -86,8 +86,9 @@ # running/fixing with recent reported activity: a killed or timed-out drive # call is not daemon death, so that claim is answered by steering the crew # to reattach, not by escalating. -# 4. No run for this crew (pre-validation, or kind=scout): fall back to the -# recorded backend's pane busy state, then the resolved status declaration +# 4. No current run for this crew (pre-validation, uninitialized repository, +# proven historical head, or kind=scout): fall back to the recorded +# backend's pane busy state, then the resolved status declaration # when its verb maps to a recognized run-state. Decision-only events such as # `resolved` never become current state or detail. # 5. Missing meta or torn-down worktree: report unknown · none. If no run is @@ -134,9 +135,9 @@ LOG=${FM_CREW_STATE_STATUS_OVERRIDE:-"$STATE/$ID.status"} NM_TIMEOUT=${FM_CREW_STATE_NM_TIMEOUT:-10} case "$NM_TIMEOUT" in ''|*[!0-9]*) NM_TIMEOUT=10 ;; esac # How many of the most recent `no-mistakes runs` rows each ledger read -# (fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh) scans, whether it is -# the cross-branch fallback or the live-sibling probe behind a terminal `axi -# status` answer (docs/configuration.md owns the setting). Generous enough to +# (fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh) scans for the legacy +# fallback or an unfetched-head continuation (docs/configuration.md owns the +# setting). Generous enough to # still find a branch's own run on a busy multi-crew fleet without listing the # entire history every call. FM_CREW_STATE_RUNS_LIMIT=${FM_CREW_STATE_RUNS_LIMIT:-200} @@ -654,13 +655,13 @@ nm_ci_checks_state() { # has no runs-listing subcommand; tests/fm-crew-state.test.sh owns the # 2026-07-02 dead-code incident history this fallback replaced). # fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh is the ONE owner of -# the ledger format, the newest-row-decides rule, its live-over-terminal -# exception, and the anchored pipeline-continuation recognition +# the ledger format, the newest-row-decides rule, and the anchored +# pipeline-continuation recognition # (model-routing-benchmark-hardening: an active fix round whose head object the # task copy never fetched used to be rejected here, letting the older failed row # answer as current), so both attribution routes share one rule. -# The same reader is also consulted when `axi status` DID bind this branch's run -# but that run is terminal, to find a live sibling run for this worktree. +# The same reader checks for conflicting run records when the AXI overview +# cannot identify this branch's run. nm_runs_list() { nm_run runs --limit "$FM_CREW_STATE_RUNS_LIMIT" } @@ -687,50 +688,106 @@ HAVE_RUN=0 # the TOON field parsing entirely for this crew. RUN_SOURCE=full COARSE_STATUS="" +SELECTED_RUN_ID="" # Scouts and secondmates never drive a no-mistakes validation of their own # worktree, so skip the lookup for them and read state from pane/log directly. if [ "$KIND" = ship ] && [ -n "$CREW_BRANCH" ] && command -v no-mistakes >/dev/null 2>&1; then RUN_OUT=$(nm_run axi status) + if [ "$(strip_quotes "$(printf '%s\n' "$RUN_OUT" | sed -n 's/^error: //p')")" = "repo not initialized (run 'no-mistakes init' first)" ]; then + RUN_OUT="" + fi if [ -n "$RUN_OUT" ]; then - run_branch=$(strip_quotes "$(nm_field branch)") - # Head equality, or the pipeline-owned-active exemption: while the - # pipeline owns this branch, the daemon's own branch attribution is - # authoritative and the lane head need not be a git object here - # (fm_nm_run_is_pipeline_owned_active in bin/fm-nm-run-lib.sh). - if [ -n "$run_branch" ] && [ "$run_branch" = "$CREW_BRANCH" ] \ - && { nm_run_head_matches_worktree || fm_nm_run_is_pipeline_owned_active "$RUN_OUT"; }; then - HAVE_RUN=1 - # Live-over-terminal (bin/fm-nm-run-lib.sh). Bare `axi status` answers - # with the most-recently-touched run, which after a pipeline crash is the - # dead run sitting at this worktree's exact commit while the live run - # that replaced it validates a descendant commit on the same branch. Both - # bind, so a terminal answer is provisional until the ledger has been - # asked whether this worktree also has a live run. Only a live word - # displaces it: a terminal run with no live sibling keeps its full - # `axi status` step and gate detail rather than degrading to the ledger. - if ! fm_nm_run_is_active "$RUN_OUT"; then - live_status=$(fm_nm_runs_status_for_worktree "$WT" "$CREW_BRANCH" "$(nm_runs_list)") - if [ "$(fm_nm_run_status_class "$live_status")" = live ]; then - COARSE_STATUS=$live_status - RUN_SOURCE=coarse + # The overview includes run ids and creation order, which the plain runs + # listing omits. Keep the primary empty-call bound above: a nonresponding + # CLI is not retried. Older CLI surfaces without the table retain the + # coarse fallback below, but cannot turn a replacement into a vague live + # verdict when its identity and gate cannot be read. + overview_ok=1 + run_overview=$(fm_nm_run_checked "$WT" "$NM_TIMEOUT" axi) || overview_ok=0 + [ -n "$run_overview" ] || emit unknown run-step "run inventory unavailable; run id: $(strip_quotes "$(nm_field id)")" + run_choice=$(fm_nm_select_run "$CREW_BRANCH" "$run_overview" "$WT") + [ "$overview_ok" = 1 ] || emit unknown run-step "run inventory unreadable; run ids: $(strip_quotes "$(nm_field id)"), ${run_choice##*|}" + case "$run_choice" in + unknown\|*) + known_run_id="" + if [ "$(strip_quotes "$(nm_field branch)")" = "$CREW_BRANCH" ]; then + known_run_id=$(strip_quotes "$(nm_field id)") fi - fi - else - # The active-or-most-recent run is for another branch, or it names this - # branch with a head this copy cannot verify (a pipeline-advanced fix - # round, or a rewritten tip). Deliberately nested inside - # `[ -n "$RUN_OUT" ]`: an empty/timed-out primary call means the CLI - # itself did not respond, so retrying it immediately with a second - # bounded call would just double the wait for no better answer. - COARSE_STATUS=$(fm_nm_runs_status_for_worktree "$WT" "$CREW_BRANCH" "$(nm_runs_list)") - if [ -n "$COARSE_STATUS" ]; then + emit unknown run-step "${run_choice#*|}${known_run_id:+; last reported run id: $known_run_id}" + ;; + selected\|*) + IFS='|' read -r _ selected_id selected_status candidate_ids <<< "$run_choice" + RUN_OUT=$(fm_nm_run_checked "$WT" "$NM_TIMEOUT" axi status --run "$selected_id") \ + || emit unknown run-step "selected run unreadable; run ids: $candidate_ids" + if [ "$(strip_quotes "$(nm_field id)")" != "$selected_id" ] \ + || [ "$(strip_quotes "$(nm_field branch)")" != "$CREW_BRANCH" ]; then + emit unknown run-step "selected run unavailable or mismatched; run ids: $candidate_ids" + fi + case "$(strip_quotes "$(nm_field status)")" in + pending|running|fixing|ci|awaiting_approval|fix_review|completed|failed|cancelled) ;; + *) emit unknown run-step "selected run status unverified; run ids: $candidate_ids" ;; + esac + if fm_nm_run_is_active "$RUN_OUT"; then current_class=live; else current_class=terminal; fi + if [ "$(fm_nm_run_status_class "$selected_status")" != "$current_class" ]; then + emit unknown run-step "selected run status disagrees with inventory; run ids: $candidate_ids" + fi + if nm_run_head_matches_worktree || fm_nm_run_is_pipeline_owned_active "$RUN_OUT"; then + HAVE_RUN=1 + elif [ -z "$(fm_nm_resolve_commit "$WT" "$(strip_quotes "$(nm_field head)")")" ]; then + if fm_nm_run_is_active "$RUN_OUT" \ + && [ "$(fm_nm_runs_status_for_worktree "$WT" "$CREW_BRANCH" "$(nm_runs_list)" "$(strip_quotes "$(nm_field head)")")" = running ]; then + HAVE_RUN=1 + else + emit unknown run-step "selected run code identity unverified; run ids: $candidate_ids" + fi + fi + SELECTED_RUN_ID=$selected_id + ;; + esac + if [ "$HAVE_RUN" = 0 ] && [ -z "$SELECTED_RUN_ID" ]; then + run_branch=$(strip_quotes "$(nm_field branch)") + # Head equality, or the pipeline-owned-active exemption: while the + # pipeline owns this branch, the daemon's own branch attribution is + # authoritative and the lane head need not be a git object here + # (fm_nm_run_is_pipeline_owned_active in bin/fm-nm-run-lib.sh). + if [ -n "$run_branch" ] && [ "$run_branch" = "$CREW_BRANCH" ] \ + && { nm_run_head_matches_worktree || fm_nm_run_is_pipeline_owned_active "$RUN_OUT"; }; then HAVE_RUN=1 - # A branch-matching answer the strict rule rejected is this branch's - # own current run once the ledger proves the pipeline-owned - # continuation, so its axi TOON is the authoritative run detail - # (RUN_SOURCE stays full); only a foreign-branch answer leaves - # coarse status-word detail. - [ "$run_branch" = "$CREW_BRANCH" ] || RUN_SOURCE=coarse + # Without run ids, contradictory liveness cannot prove precedence. + # A live replacement also needs an id-addressed status read: a bare + # "running" row cannot tell working from waiting at a gate. + ledger_status=$(fm_nm_runs_status_for_worktree "$WT" "$CREW_BRANCH" "$(nm_runs_list)") + if fm_nm_run_is_active "$RUN_OUT"; then + if [ "$(fm_nm_run_status_class "$ledger_status")" = terminal ]; then + emit unknown run-step "run records disagree; run ids: $(strip_quotes "$(nm_field id)"), competing identity unavailable" + fi + else + if [ "$(fm_nm_run_status_class "$ledger_status")" = live ]; then + emit unknown run-step "replacement run identity unavailable; run ids: $(strip_quotes "$(nm_field id)"), replacement unavailable" + elif [ -n "$ledger_status" ] \ + && [ "$ledger_status" != "$(strip_quotes "$(nm_field status)")" ] \ + && [ "$ledger_status" != "$(strip_quotes "$(nm_field outcome)")" ]; then + COARSE_STATUS=$ledger_status + RUN_SOURCE=coarse + fi + fi + else + # The active-or-most-recent run is for another branch, or it names this + # branch with a head this copy cannot verify (a pipeline-advanced fix + # round, or a rewritten tip). Deliberately nested inside + # `[ -n "$RUN_OUT" ]`: an empty/timed-out primary call means the CLI + # itself did not respond, so retrying it immediately with a second + # bounded call would just double the wait for no better answer. + COARSE_STATUS=$(fm_nm_runs_status_for_worktree "$WT" "$CREW_BRANCH" "$(nm_runs_list)") + if [ -n "$COARSE_STATUS" ]; then + HAVE_RUN=1 + # A branch-matching answer the strict rule rejected is this branch's + # own current run once the ledger proves the pipeline-owned + # continuation, so its axi TOON is the authoritative run detail + # (RUN_SOURCE stays full); only a foreign-branch answer leaves + # coarse status-word detail. + [ "$run_branch" = "$CREW_BRANCH" ] || RUN_SOURCE=coarse + fi fi fi fi @@ -746,13 +803,9 @@ if [ "$HAVE_RUN" = 1 ]; then RUN_STATUS="" if [ "$RUN_SOURCE" = coarse ]; then # No step/gate detail is available from the plain runs list - only ever - # true/working, done, or failed. A crew genuinely parked at a gate still - # gets full detail once `axi status` reports its own branch again (e.g. - # once its own step is the most-recently-touched one), and its own - # needs-decision/blocked status-log append (a captain-relevant VERB) is - # surfaced by each supervisor's span classification (fm-classify-lib.sh's - # status_span_first_actionable) regardless of this coarse-vs-full - # distinction, so a real gate is never silently missed. + # working, done, failed, or unknown. Gate detail requires the identity-aware + # read above. The status event span remains independently available to the + # supervisor through fm-classify-lib.sh's status_span_first_actionable. case "$COARSE_STATUS" in running) RUN_STATE=working; RUN_DETAIL="validating (background run)" ;; completed) RUN_STATE="done"; RUN_DETAIL="run completed" ;; @@ -893,6 +946,7 @@ if [ "$HAVE_RUN" = 1 ]; then ;; esac + [ -z "$SELECTED_RUN_ID" ] || RUN_DETAIL="$RUN_DETAIL${SEP}run: $SELECTED_RUN_ID" emit "$RUN_STATE" run-step "$RUN_DETAIL" fi diff --git a/bin/fm-nm-run-lib.sh b/bin/fm-nm-run-lib.sh index ed71d315fd8..3377dffa4cf 100644 --- a/bin/fm-nm-run-lib.sh +++ b/bin/fm-nm-run-lib.sh @@ -86,15 +86,9 @@ fm_nm_resolve_commit() { # # the ancestor rule (observed 2026-08: a crashed validation daemon left a failed # run at the worktree's own commit while the live run that replaced it validated # a descendant commit on the same branch). -# When several runs bind, a LIVE run always outranks a terminal one, whichever -# match rule each one used, because a terminal run can be the corpse of a -# crashed attempt while the live one is what is actually validating this code. -# Within one liveness class the selecting caller's existing precedence is -# unchanged - for the runs ledger, fm_nm_runs_status_for_worktree's -# newest-row-decides rule below. -# fm_nm_run_status_class next classifies a recorded status word for that -# comparison, and a word it cannot classify keeps the caller's own precedence -# rather than being held back for a live row to displace. +# Head compatibility alone does not establish precedence between runs. +# fm_nm_select_run below owns identity-aware selection for current-state reads; +# fm_nm_runs_status_for_worktree owns the coarse ledger fallback. fm_nm_head_matches_worktree() { # local wt=$1 run_head=$2 local_full run_full [ -n "$run_head" ] || return 1 @@ -105,19 +99,177 @@ fm_nm_head_matches_worktree() { # git -C "$wt" merge-base --is-ancestor "$local_full" "$run_full" 2>/dev/null } -# Liveness class of a recorded run's status word, echoed as "terminal", "live", -# or "unknown", for the live-over-terminal selection rule above. -# The coarse `no-mistakes runs` ledger emits exactly these four status words; an +# Liveness class of a recorded ledger status word. +# The coarse `no-mistakes runs` ledger emits database status words; an # `axi status` run object reports its terminal result through its own outcome # field as well, which fm_nm_run_is_active below checks directly. fm_nm_run_status_class() { # case "${1:-}" in completed|failed|cancelled) printf 'terminal' ;; - running) printf 'live' ;; + pending|running) printf 'live' ;; *) printf 'unknown' ;; esac } +# Select from a complete `no-mistakes axi` overview with the existing awk +# toolchain. A capped overview requires an optional Python 3 sqlite3 reader +# for a read-only same-branch query of NM_HOME/state.sqlite (default: +# ~/.no-mistakes/state.sqlite; relative NM_HOME resolves from the worktree). +# If that reader or inventory is unavailable, report unknown with available +# candidate ids rather than treating the displayed window as complete. +# Structural completeness applies to the whole table; semantic validation +# applies only to the requested branch, after complete identity lookup when +# capped. Branch names are matched exactly without a character whitelist. +# Its rows are ordered by creation time descending (not last update), then id. +# The newest same-branch row is the candidate regardless of outcome: an older +# live run must not hide a newer failure. If the newest is live and another +# same-branch live run exists, neither has exclusive authority: report all +# candidate ids as unknown. A newer live row can replace cancelled history, +# but the caller must fetch its full status BY ID and prove branch/head or +# active pipeline custody before using its steps. Never reuse another run's +# gate detail. This is a read-only selection, not teardown authorization. +# +# Prints selected|id|status|candidate-ids, unknown|reason, absent (no row +# for this branch), or unavailable (CLI has no overview table). Malformed or +# structurally truncated tables report unknown, retaining every readable +# same-branch candidate id. +fm_nm_select_run() { # + local selection inventory available_ids + selection=$(printf '%s\n' "$2" | awk -v branch="$1" ' + function scalar(s) { + sub(/^[ \t]+/, "", s); sub(/[ \t]+$/, "", s) + if (s ~ /^".*"$/) s = substr(s, 2, length(s)-2) + return s + } + function row_fields(s, f, i, ch, n, quoted, escaped) { + for (i in f) delete f[i] + n = 1; f[n] = "" + for (i = 1; i <= length(s); i++) { + ch = substr(s, i, 1) + if (escaped) { f[n] = f[n] ch; escaped = 0 } + else if (quoted && ch == "\\") escaped = 1 + else if (ch == "\"") quoted = !quoted + else if (!quoted && ch == ",") { n++; f[n] = "" } + else f[n] = f[n] ch + } + if (quoted || escaped) return 0 + for (i = 1; i <= n; i++) { + sub(/^[ \t]+/, "", f[i]); sub(/[ \t]+$/, "", f[i]) + } + return n + } + /^count: / { + if (counts++) bad = 1 + count = scalar(substr($0, 8)) + if (count !~ /^[0-9]+ of [0-9]+ total$/) bad = 1 + split(count, c, " "); shown = c[1]; total = c[3] + } + /^runs\[[0-9]+\]\{id,branch,status,head,pr\}:$/ { + if (found++) bad = 1 + expected = $0; sub(/^runs\[/, "", expected); sub(/\].*$/, "", expected) + inrows = 1; next + } + /^runs\[/ { bad = 1; found = 1 } + inrows && /^[ \t]+/ { + seen++ + n = row_fields($0, f) + if (n != 5) bad = 1 + id = f[1]; br = f[2]; st = f[3]; head = f[4] + if (br != branch) next + if (id ~ /^[A-Za-z0-9_-]+$/) { + if (known[id]++) invalid_run = 1 + else ids = ids (ids == "" ? "" : ", ") id + } + if (n != 5) next + if (id !~ /^[A-Za-z0-9_-]+$/ || + st !~ /^[a-z_-]+$/ || head !~ /^[a-fA-F0-9]+$/ || length(head) < 7 || length(head) > 40) { + invalid_run = 1; next + } + if (first == "") { first = id; first_status = st } + if (st == "running" || st == "pending") live++ + if (st !~ /^(pending|running|completed|failed|cancelled)$/) unknown_status = 1 + next + } + inrows { inrows = 0 } + END { + if (!found) print "unavailable" + else if (bad || counts != 1 || seen != expected || seen != shown || total < shown) + print "unknown|unreadable runs table; run ids: " ids + else if (shown < total) print "incomplete|" ids + else if (invalid_run) print "unknown|unreadable runs table; run ids: " ids + else if (unknown_status) print "unknown|unrecognized run status; run ids: " ids + else if (first == "") print "absent" + else if ((first_status == "running" || first_status == "pending") && live > 1) + print "unknown|competing live runs; run ids: " ids + else print "selected|" first "|" first_status "|" ids + } + ') + case "$selection" in + incomplete\|*) available_ids=${selection#*|} ;; + *) printf '%s\n' "$selection"; return ;; + esac + if ! inventory=$(python3 - "$1" "$2" "$3" "$available_ids" 2>/dev/null <<'PY' +import json +import os +import re +import sqlite3 +import sys +from contextlib import closing +from pathlib import Path + +branch, overview, worktree, available_ids = sys.argv[1:] +ids = available_ids.split(", ") if available_ids else [] +try: + repos = [line[6:].strip() for line in overview.splitlines() if line.startswith("repo: ")] + if len(repos) != 1: + raise ValueError + repo_path = json.loads(repos[0]) if repos[0].startswith('"') else repos[0] + if not isinstance(repo_path, str) or not os.path.isabs(repo_path): + raise ValueError + root = Path(os.environ.get("NM_HOME") or Path.home() / ".no-mistakes") + if not root.is_absolute(): + root = Path(worktree) / root + with closing(sqlite3.connect((root / "state.sqlite").as_uri() + "?mode=ro", uri=True, timeout=1)) as db: + db.execute("BEGIN") + repo = db.execute("SELECT id FROM repos WHERE working_path = ?", (repo_path,)).fetchall() + if len(repo) != 1: + raise ValueError + rows = db.execute( + "SELECT id, branch, status, head_sha FROM runs WHERE repo_id = ? AND branch = ? " + "ORDER BY created_at DESC, id DESC", (repo[0][0], branch) + ).fetchall() + displayed_ids = set(ids) + for row in rows: + if isinstance(row[0], str) and re.fullmatch(r"[A-Za-z0-9_-]+", row[0]) and row[0] not in ids: + ids.append(row[0]) + if not displayed_ids.issubset(row[0] for row in rows): + raise ValueError + for row in rows: + if (not all(isinstance(value, str) for value in row) + or not re.fullmatch(r"[A-Za-z0-9_-]+", row[0]) or row[1] != branch + or not re.fullmatch(r"[a-z_-]+", row[2]) or not re.fullmatch(r"[a-fA-F0-9]{7,40}", row[3])): + raise ValueError + print("count: %d of %d total" % (len(rows), len(rows))) + print("runs[%d]{id,branch,status,head,pr}:" % len(rows)) + for row in rows: + print(" " + ",".join(json.dumps(value, ensure_ascii=False) for value in row) + ',""') +except (ValueError, OSError, sqlite3.Error): + print("unknown|complete same-branch run inventory unreadable; run ids: " + ", ".join(ids)) +PY + ); then + printf 'unknown|complete same-branch run inventory reader unavailable; run ids: %s\n' "$available_ids" + return + fi + case "$inventory" in + unknown\|*) selection=$inventory ;; + *) selection=$(fm_nm_select_run "$1" "$inventory" "$3") ;; + esac + case "$selection" in + selected\|*|unknown\|*|absent) printf '%s\n' "$selection" ;; + *) printf 'unknown|complete same-branch run inventory unreadable; run ids: %s\n' "$available_ids" ;; + esac +} + # branch_sync.state from captured `axi status` TOON $1: the scalar directly # under the top-level `branch_sync:` block. The first `state:` inside the # block is the direct child (the nested local/pipeline/target/remote @@ -179,32 +331,12 @@ fm_nm_run_is_pipeline_owned_active() { # # printed. Anything else (no anchor row, an anchor that is merely an # ancestor, a terminal unresolvable row) prints nothing, so branch-name # coincidence, arbitrary remote state, and other tasks' runs never match. -# The one exception to newest-row-decides is the live-over-terminal rule stated -# with fm_nm_head_matches_worktree above, and it only ever replaces a TERMINAL -# answer with a LIVE one: when the newest row binds but is terminal, the older -# rows are scanned for a live row that ALSO binds to this worktree, and that -# row's status word is printed instead. A live row whose head resolves in this -# copy binds by fm_nm_head_matches_worktree. A live row whose head does NOT -# resolve (the routine shape: the pipeline's fix-round commits live only in the -# gate repo) binds ONLY when the held terminal row sits at EXACTLY the worktree -# HEAD - the same exact-equality anchor the pipeline-continuation rule above -# requires, so branch-name coincidence and other tasks' runs still never -# match. A terminal newest row is the corpse of a crashed attempt whenever a -# live run for the same worktree is still on the ledger, so it is not the -# present. Nothing else widens: a newest row that does not bind still ends the -# scan, a newest row whose class is live or unclassifiable is still answered -# as-is, the anchored pipeline-continuation path is untouched, and with no live -# sibling the newest terminal word is still what is printed. +# An older live row never displaces a newer terminal result. # Read-only: git reads resolve objects in place; custody never changes. fm_nm_runs_status_for_worktree() { # [expected-head] local wt=$1 branch=$2 list=$3 expected_head=${4:-} local local_full row_full row st br sha day clock pr extra year_num month_num day_num max_day pending_st='' - # Set only by the newest binding row when its status classifies terminal, and - # printed when the scan ends without finding a live row for this worktree. It - # is the sole reason the scan continues past the newest row, and every exit - # below leaves the loop rather than returning, so a malformed older row can - # never swallow an answer the newest row had already decided. - local decided='' decided_exact='' + local decided='' local_full=$(git -C "$wt" rev-parse HEAD 2>/dev/null) || return 0 [ -n "$list" ] || return 0 while IFS= read -r row; do @@ -237,20 +369,6 @@ fm_nm_runs_status_for_worktree() { # [ex esac [ "$day_num" -ge 1 ] && [ "$day_num" -le "$max_day" ] || break [ "$br" = "$branch" ] || continue - if [ -n "$decided" ]; then - # Live-over-terminal: the newest row bound to this worktree but is a - # terminal record, so the older rows are searched for a live run that - # binds to the same worktree by the same head rule. Only such a row - # displaces the held terminal word; anything else leaves it standing. - [ "$(fm_nm_run_status_class "$st")" = live ] || continue - if [ -n "$(fm_nm_resolve_commit "$wt" "$sha")" ]; then - fm_nm_head_matches_worktree "$wt" "$sha" || continue - else - [ -n "$decided_exact" ] || continue - fi - decided=$st - break - fi if [ -n "$pending_st" ]; then # This is the row immediately older than the active unresolvable row: # the only admissible anchor, and only exact head equality proves the @@ -272,12 +390,6 @@ fm_nm_runs_status_for_worktree() { # [ex if [ -n "$row_full" ]; then if fm_nm_head_matches_worktree "$wt" "$sha"; then decided=$st - # A live or unclassifiable word is this worktree's current answer and - # ends the scan; only a terminal one keeps looking for a live sibling. - if [ "$(fm_nm_run_status_class "$st")" = terminal ]; then - [ "$row_full" != "$local_full" ] || decided_exact=1 - continue - fi fi break fi diff --git a/docs/architecture.md b/docs/architecture.md index 8026d38e6b9..e5e55c0be82 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -88,9 +88,8 @@ A turn-ended-only queue row omits its historical status annotation when that sta Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line. `bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed, except that a `blocked:` event reporting a refused or missing daemon socket outranks a potentially stale active run record only while that socket-down declaration is itself the log's latest recognized event, since any later event, including another `blocked:` one, means the crew moved on. For other daemon, timeout, or unreachability claims, a running or fixing run with recent pipeline-reported activity supersedes the event and names reattachment as the recovery instead of surfacing a false block. -[`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules. -It also owns which binding run wins when more than one recorded run binds to the same worktree: a live run outranks a terminal one, so a crashed run sitting at the worktree's own commit never reports a healthy task as failed while its live successor is still validating. -A run head the task copy cannot resolve locally is attributed only when the pipeline's own runs ledger proves it is an active continuation of the submitted head, so a pipeline fix round never reads as an older failed run. +[`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh) owns branch, head, and pipeline-custody attribution, plus complete same-branch run selection, optional inventory lookup, and ambiguity reporting. +[`tests/fm-crew-state.test.sh`](../tests/fm-crew-state.test.sh) covers run selection; its [capture provenance and live-evidence limits](../tests/captures/no-mistakes-v1.70.1/README.md) distinguish recorded inputs from composed scenarios. During no-mistakes' `ci` monitor phase, it also reads the ci step log tail because `axi status` reports both "still waiting on checks" and "checks green, waiting on merge" as `ci,running`. The most recent recognized ci log marker wins, so checks-green monitoring reports done while a later re-arm, failed-check, or issue marker returns the crew to working. `bin/fm-crew-state.sh` owns the evidence guard that recognizes ended CI monitors after green checks, including cancelled runs and skipped rebase steps; a passed run alone never proves a forge merge. diff --git a/docs/configuration.md b/docs/configuration.md index 46949796ef2..15f5f5efb9e 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -1076,7 +1076,7 @@ FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail insid FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh FM_TEARDOWN_NM_TIMEOUT=10 # seconds allowed per no-mistakes query or abort inside fm-teardown.sh -FM_CREW_STATE_RUNS_LIMIT=200 # recent no-mistakes run rows scanned when the runs ledger is consulted: axi status cannot be attributed directly, or its answer is terminal and may have a live sibling run +FM_CREW_STATE_RUNS_LIMIT=200 # plain runs-ledger rows scanned for fallback attribution; does not change the CLI's AXI overview window (selection owner: bin/fm-nm-run-lib.sh) FM_TEARDOWN_NM_RUNS_LIMIT=200 # recent no-mistakes run rows scanned to prove an unresolved-head parked run belongs to teardown's task FM_CREW_STATE_BIN=bin/fm-crew-state.sh # test override for the current-state reader used by working/paused watcher triage FM_MAIL_USER= # mail-plane IMAP/SMTP login, from .env or environment (docs/configuration.md "Mail plane") diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index f639220b551..e459e95006a 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -511,6 +511,10 @@ { "path": "skills/stow/SKILL.md", "audience": "public-product" + }, + { + "path": "tests/captures/no-mistakes-v1.70.1/README.md", + "audience": "maintainer-verification" } ] } diff --git a/tests/captures/no-mistakes-v1.70.1/README.md b/tests/captures/no-mistakes-v1.70.1/README.md new file mode 100644 index 00000000000..75cc474d955 --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/README.md @@ -0,0 +1,51 @@ +# AXI run-state input captures + +These files own recorded serialized CLI inputs and a persisted run-inventory projection for the `test_captured_*` cases in `../../fm-crew-state.test.sh`. +They were captured on 2026-09-14 at 22:06 UTC with `no-mistakes version v1.70.1 (9c380d4) 2026-09-07T20:47:58Z`. +They are replay inputs, not evidence that every composed scenario was driven live. + +## Capture provenance + +Each status file is unchanged stdout from `no-mistakes axi status --run ` executed from the test-phase worktree, without entering the recorded run's checkout. +The command exited zero for all five run-specific captures. +`uninitialized.toon` is unchanged stdout from `no-mistakes axi status` in that worktree, which exited one. +Update notices on stderr are not part of the captured stdout contract. + +| File | Recorded run ID | Observed state | +| --- | --- | --- | +| `replacement.toon` | `01M2GAWMSDQK4B5EA9GZW35RXE` | Live CI step after a rerun | +| `superseded.toon` | `01M2FNFPK984YP0EHFTD1XEF8P` | Same-branch predecessor cancelled with `superseded by new push` | +| `parked.toon` | `01M20NDQH0G96AQYH1EHWGKT5F` | Separate branch parked at the test gate with one finding | +| `failed.toon` | `01M289YXN7V0513AKCF53MJ1BC` | Failed push step | +| `completed.toon` | `01M2FG7SEEP1VBZ3B5SX35BJ6Q` | Completed validation | + +`same-branch-inventory.json` preserves all nine rows for the replacement's branch, selected in a read-only transaction from the real `state.sqlite` database. +The projection is `id, repo_id, branch, status, head_sha, created_at`, ordered by `created_at DESC, id DESC`. +The source contained 78 runs for repository `acf4a767348a`; its `runs` schema confirmed the fixture's text identity/status/head fields and integer creation times. +No branch in that repository had two recorded live runs at capture time. + +`overview.toon` is the unchanged `count` and `runs` section emitted by the real `no-mistakes axi` executable against an isolated database copy of those 78 recorded runs. +Only the copy's repository `working_path` was relocated to the permitted worktree; no pipeline was initialized or controlled. +The copy omitted step data and had no daemon, so the unrelated active-run detail from that output is intentionally excluded. +The retained section demonstrates the actual ten-row cap, row order, quoting, and field layout. +Original stdout, source projections, and SHA-256 digests were retained in the test-phase evidence directory under `real-anchors/`. + +## Replay transformations and limits + +`captured_axi_status` substitutes only the run ID, branch, and head fields so the captures bind to disposable Git repositories. +Status, outcome, steps, findings, and gate bytes remain unchanged. +The inventory replay substitutes its disposable repository key and path, preserves every captured same-branch row, and hashes the database before and after the state read to detect writes. +Its ambiguity case explicitly changes one hidden cancelled row to running; this is a counterfactual, not a captured competing-live history. +The original review-gate, rebased-head, unrelated-metadata, and malformed-input assertions remain unchanged. + +| Required shape | Real anchor used | What remains unproven live | +| --- | --- | --- | +| Superseded cancellation yields to a parked replacement | Genuine cancellation/successor history plus separately captured parked-gate output | The captured successor was in CI, and the captured gate was at test on another branch; a same-rerun replacement parked specifically at review on an unfetched rebased head was not captured | +| Competing live identities beyond the cap and beside unrelated metadata | Real capped overview and complete nine-row branch history | No real branch had two live rows; changing a hidden row to running and injecting unrelated unusual metadata are controlled fixtures | +| Newer failure outranks an older live run | Genuine failed status and genuine live status | This relative ordering with both states on one branch was composed, not observed | +| Changing or unverifiable authority | Genuine live and cancelled status formats | The transition between reads, malformed records, wrong identities, and unreadable inventory are injected; no live race or corrupt production inventory was captured | +| Uninitialized repository preserves worker reporting | Actual uninitialized stdout; earlier live lifecycle-event/pane evidence | The portable test replays stdout and uses the existing pane fake | +| Development continues after completed validation | Genuine completed status | Advancing Git and emitting worker events after completion are disposable-repository actions, not an observed recorded worker sequence | +| Optional inventory dependencies are absent | Genuine gate output and capped inventory | Missing Python/SQLite and complete-inventory compositions are simulated; the captured host had both dependencies | + +Passing replay assertions establish behavior for these explicit inputs, not the absent live scenarios in the final column. diff --git a/tests/captures/no-mistakes-v1.70.1/completed.toon b/tests/captures/no-mistakes-v1.70.1/completed.toon new file mode 100644 index 00000000000..b2e0df67240 --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/completed.toon @@ -0,0 +1,20 @@ +run: + id: "01M2FG7SEEP1VBZ3B5SX35BJ6Q" + branch: fm/fm-installed-timeout-watcher-will-not-stop + status: completed + head: 1129818e + head_sha: 1129818ef720b6827ae75eb690957c2c7393d82b + pr: "https://github.com/kunchenguid/firstmate/pull/4073" + findings: 1 awaiting + steps[9]{step,status,findings,duration_ms}: + intent,completed,0,20 + rebase,completed,0,1309 + review,completed,0,159353 + test,completed,0,2655669 + document,completed,0,119685 + lint,completed,0,6647 + push,completed,0,5691 + pr,completed,0,62239 + ci,completed,1,1845554 +outcome: passed-with-override +ci_override_reason: "live checks for https://github.com/kunchenguid/firstmate/pull/4073 not all passed: Lint (fail)" diff --git a/tests/captures/no-mistakes-v1.70.1/failed.toon b/tests/captures/no-mistakes-v1.70.1/failed.toon new file mode 100644 index 00000000000..8ef2a2a2520 --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/failed.toon @@ -0,0 +1,19 @@ +run: + id: "01M289YXN7V0513AKCF53MJ1BC" + branch: fm/fm-bearings-board-loses-owner-state-and-links + status: failed + head: 9b76c588 + head_sha: 9b76c588dadf8d2e39405526fdbc41bbf8245513 + findings: none + steps[9]{step,status,findings,duration_ms}: + intent,completed,0,207 + rebase,skipped,0,0 + review,completed,0,840023 + test,completed,0,3218433 + document,skipped,0,0 + lint,completed,0,8278 + push,failed,0,6980 + pr,pending,0,0 + ci,pending,0,0 +outcome: failed +error: "step push failed: push to fork: git push https://github.com/mremond/firstmate --force-with-lease=refs/heads/fm/fm-bearings-board-loses-owner-state-and-links:723830dc438374c69661f37fdee6d9606ce744cf 9b76c588dadf8d2e39405526fdbc41bbf8245513:refs/heads/fm/fm-bearings-board-loses-owner-state-and-links: exit status 1: To https://github.com/mremond/firstmate\n ! [remote rejected] 9b76c588dadf8d2e39405526fdbc41bbf8245513 -> fm/fm-bearings-board-loses-owner-state-and-links (refusing to allow an OAuth App to create or update workflow `.github/workflows/ci.yml` without `workflow` scope)\nerror: failed to push some refs to 'https://github.com/mremond/firstmate'" diff --git a/tests/captures/no-mistakes-v1.70.1/overview.toon b/tests/captures/no-mistakes-v1.70.1/overview.toon new file mode 100644 index 00000000000..26839460f7e --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/overview.toon @@ -0,0 +1,12 @@ +count: 10 of 78 total +runs[10]{id,branch,status,head,pr}: + "01M2GV0N1CJ7TGNYN3P76KK9TS",fm/fm-superseded-cancelled-run-outranks-live,running,7feb0272,"" + "01M2GV0HN9PMBPA7YNVGGC9FSS",fm/fm-spawn-tests-borrow-the-checkout,running,a0eb3010,"https://github.com/kunchenguid/firstmate/pull/4056" + "01M2GQTYBGKJGSFTQPP3F8NX0G",fm/fm-origin-credentials-printed-in-errors,running,4e5713f5,"https://github.com/kunchenguid/firstmate/pull/4138" + "01M2GB01EQ2MPCJ200TG71Y9SV",fm/fm-ci-refresh-portable-serial-hints,running,e6be9247,"https://github.com/kunchenguid/firstmate/pull/4144" + "01M2GAWMSDQK4B5EA9GZW35RXE",fm/fm-bearings-board-loses-owner-state-and-links,running,146a90ee,"https://github.com/kunchenguid/firstmate/pull/4019" + "01M2FR2HYYDP8NP8HDDX4D1MXS",fm/fm-absorb-fresh-replacement-wait,running,"71042535","https://github.com/kunchenguid/firstmate/pull/3605" + "01M2FNFPK984YP0EHFTD1XEF8P",fm/fm-bearings-board-loses-owner-state-and-links,cancelled,735a1fc5,"https://github.com/kunchenguid/firstmate/pull/4019" + "01M2FG7SEEP1VBZ3B5SX35BJ6Q",fm/fm-installed-timeout-watcher-will-not-stop,completed,1129818e,"https://github.com/kunchenguid/firstmate/pull/4073" + "01M289YXN7V0513AKCF53MJ1BC",fm/fm-bearings-board-loses-owner-state-and-links,failed,9b76c588,"" + "01M289X89E6CCFQ696MFF8G0MR",fm/fm-bearings-board-loses-owner-state-and-links,cancelled,d5825696,"" diff --git a/tests/captures/no-mistakes-v1.70.1/parked.toon b/tests/captures/no-mistakes-v1.70.1/parked.toon new file mode 100644 index 00000000000..4fb3ea76baf --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/parked.toon @@ -0,0 +1,26 @@ +run: + id: "01M20NDQH0G96AQYH1EHWGKT5F" + branch: fm/fm-codex-hook-trust-dialog-undocumented + status: running + awaiting_agent: parked 2d14h + head: fb90db9f + head_sha: fb90db9f09db9b0cd1a83539959367c612659638 + pr: "https://github.com/kunchenguid/firstmate/pull/3998" + findings: 1 awaiting + steps[9]{step,status,findings,duration_ms}: + intent,completed,0,200 + rebase,completed,0,6521 + review,completed,0,180772 + test,awaiting_approval,1,90408 + document,pending,0,0 + lint,pending,0,0 + push,pending,0,0 + pr,pending,0,0 + ci,pending,0,0 +gate: + step: test + status: awaiting_approval + summary: Documentation-only change with no live test surface. Previously declined missing-evidence findings were not repeated. + findings[1]{id,severity,file,action,description}: + test-1,warning,"",ask-user,"this change has no live-validatable surface; proceed without live validation? (0 of 4 scenarios were driven live against the product); An operator reads the notes and encounters silent loss of supervision as the first warning.: Only documentation changed, with no executable surface. Demonstrating operator response would require a separate operator-driven evaluation.; An operator dismisses hook review with Escape without selecting review, declining hooks, or treating dismissal as approval.: No executable dialog handling changed. Live corroboration requires an operator-controlled isolated sessio… (truncated, 1236 chars total)" +help[2]: The explicitly selected gate for run 01M20NDQH0G96AQYH1EHWGKT5F is inspection-only; no run-scoped response command exists,Run `no-mistakes axi logs --run 01M20NDQH0G96AQYH1EHWGKT5F --step test --full` to read the full step log diff --git a/tests/captures/no-mistakes-v1.70.1/replacement.toon b/tests/captures/no-mistakes-v1.70.1/replacement.toon new file mode 100644 index 00000000000..99491b31a69 --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/replacement.toon @@ -0,0 +1,20 @@ +run: + id: "01M2GAWMSDQK4B5EA9GZW35RXE" + branch: fm/fm-bearings-board-loses-owner-state-and-links + status: running + head: 146a90ee + head_sha: 146a90ee48c675ec7908faf141026efef5158c5a + pr: "https://github.com/kunchenguid/firstmate/pull/4019" + findings: 6 info + steps[9]{step,status,findings,duration_ms}: + intent,completed,0,129 + rebase,skipped,6,2090 + review,completed,0,2396206 + test,completed,0,1245156 + document,completed,0,226600 + lint,completed,0,15165 + push,completed,0,8511 + pr,completed,0,79292 + ci,running,0,0 + active_steps[1]{step,status,active_for,round_active_for,last_activity,agent_pid,round}: + ci,running,4h28m,4h28m,"quiet 2h58m ago: log: all CI checks passed - still monitoring until merged or closed","",starting diff --git a/tests/captures/no-mistakes-v1.70.1/same-branch-inventory.json b/tests/captures/no-mistakes-v1.70.1/same-branch-inventory.json new file mode 100644 index 00000000000..35b2f0c0e90 --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/same-branch-inventory.json @@ -0,0 +1,74 @@ +[ + { + "id": "01M2GAWMSDQK4B5EA9GZW35RXE", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "running", + "head_sha": "146a90ee48c675ec7908faf141026efef5158c5a", + "created_at": 1789402174 + }, + { + "id": "01M2FNFPK984YP0EHFTD1XEF8P", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "cancelled", + "head_sha": "735a1fc5f8a02dbf2cab3851ad0d28ec85a177e3", + "created_at": 1789379730 + }, + { + "id": "01M289YXN7V0513AKCF53MJ1BC", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "failed", + "head_sha": "9b76c588dadf8d2e39405526fdbc41bbf8245513", + "created_at": 1789132764 + }, + { + "id": "01M289X89E6CCFQ696MFF8G0MR", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "cancelled", + "head_sha": "d582569662668edf727b74baefe100be32cf9e9b", + "created_at": 1789132710 + }, + { + "id": "01M21B2PY8J90QVVWWNDZ2WVGJ", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "completed", + "head_sha": "723830dc438374c69661f37fdee6d9606ce744cf", + "created_at": 1788899056 + }, + { + "id": "01M21AXB75DVRYHRRWBC29GSN9", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "cancelled", + "head_sha": "4996127f9d095686e4adb0063626753d68d1837f", + "created_at": 1788898880 + }, + { + "id": "01M20RM7ENZX5ZKSTYK5K3EFEG", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "cancelled", + "head_sha": "d69beeb67488f764128a7e0ef04d583d404fedc2", + "created_at": 1788879707 + }, + { + "id": "01M20MQ02N69VJKXW9N8321SQW", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "cancelled", + "head_sha": "1ecdbcb34f8045fe74ac3acc15d37677179ccc37", + "created_at": 1788875604 + }, + { + "id": "01M20HPP5XR7K31VHP0DC6S5QA", + "repo_id": "acf4a767348a", + "branch": "fm/fm-bearings-board-loses-owner-state-and-links", + "status": "failed", + "head_sha": "1ecdbcb34f8045fe74ac3acc15d37677179ccc37", + "created_at": 1788872448 + } +] diff --git a/tests/captures/no-mistakes-v1.70.1/superseded.toon b/tests/captures/no-mistakes-v1.70.1/superseded.toon new file mode 100644 index 00000000000..8291fcf07fd --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/superseded.toon @@ -0,0 +1,20 @@ +run: + id: "01M2FNFPK984YP0EHFTD1XEF8P" + branch: fm/fm-bearings-board-loses-owner-state-and-links + status: cancelled + head: 735a1fc5 + head_sha: 735a1fc5f8a02dbf2cab3851ad0d28ec85a177e3 + pr: "https://github.com/kunchenguid/firstmate/pull/4019" + findings: "1 awaiting, 2 auto-fix" + steps[9]{step,status,findings,duration_ms}: + intent,completed,0,18 + rebase,completed,0,1205 + review,completed,3,321482 + test,completed,0,906557 + document,completed,0,242319 + lint,completed,0,7346 + push,completed,0,8523 + pr,completed,0,195015 + ci,failed,0,20551480 +outcome: cancelled +error: "cancelled: superseded by new push" diff --git a/tests/captures/no-mistakes-v1.70.1/uninitialized.toon b/tests/captures/no-mistakes-v1.70.1/uninitialized.toon new file mode 100644 index 00000000000..c124abfc66d --- /dev/null +++ b/tests/captures/no-mistakes-v1.70.1/uninitialized.toon @@ -0,0 +1,2 @@ +error: repo not initialized (run 'no-mistakes init' first) +help[1]: Run `no-mistakes init` to set up the gate in this repository diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index 2f3faf3ca3f..a51574f0f14 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -21,10 +21,9 @@ # (d2) terminal failed run whose only failure is an orphaned ci monitor # after checks read green -> done # (e) cross-branch attribution: this branch's own run found via list lookup -# (e2) several runs bound to one worktree: the live one outranks the corpse -# (an unclassifiable status word keeps the ledger's newest-first order) -# (e3) the live sibling's head was never fetched into the task copy: it still -# outranks a terminal row sitting at the worktree's exact commit +# (e2) multiple runs: creation order preserves newer failures, replacement +# gates retain their run identity, and competing live runs read unknown +# (e3) an older live sibling with an unfetched head cannot hide a newer failure # (f) no run + semantic busy -> pane # (g) no run + semantic idle falls to the status-log verb -> status-log # (h) dead pane: no run -> unknown/none; with a run -> run-step (not the shell) @@ -68,7 +67,8 @@ make_repo_on_branch() { # # A fakebin with a fake `no-mistakes` (serves the env-driven run output) and a # fake `tmux` (serves a busy or idle pane). The fake no-mistakes mirrors the real -# command surface the helper uses: `axi status`, `axi status --run ` (the +# command surface the helper uses: `axi` (the identity overview), `axi status`, +# and `axi status --run ` (the # `axi` surface - no runs-listing subcommand exists under it, verified against # the real CLI), and the actual top-level run-listing command, `no-mistakes # runs --limit N`, which is plain text - no run id, no quoting - serving @@ -82,11 +82,20 @@ set -u case "${1:-}" in axi) shift + if [ "$#" = 0 ]; then + printf '%s\n' "${FM_FAKE_AXI_HOME:-${FM_FAKE_AXI_STATUS:-}}" + exit "${FM_FAKE_AXI_HOME_ERROR:-0}" + fi case "${1:-}" in status) shift - if [ "${1:-}" = --run ]; then printf '%s\n' "${FM_FAKE_AXI_STATUS_RUN:-}" - else printf '%s\n' "${FM_FAKE_AXI_STATUS:-}"; fi ;; + if [ "${1:-}" = --run ]; then + printf '%s\n' "${FM_FAKE_AXI_STATUS_RUN:-}" + exit "${FM_FAKE_AXI_STATUS_RUN_ERROR:-0}" + else + printf '%s\n' "${FM_FAKE_AXI_STATUS:-}" + exit "${FM_FAKE_AXI_STATUS_ERROR:-0}" + fi ;; logs) printf '%s\n' "${FM_FAKE_CI_LOGS:-}" ;; esac @@ -267,7 +276,13 @@ arm_idle_record() { # # assignments below stay exported into the fakes without an `export VAR=$(...)` # command-substitution assignment (SC2155). reset_fakes() { + NM_HOME="$TMP_ROOT/no-mistakes-unused" + export NM_HOME FM_FAKE_AXI_STATUS="" + FM_FAKE_AXI_STATUS_ERROR=0 + FM_FAKE_AXI_HOME="" + FM_FAKE_AXI_HOME_ERROR=0 + FM_FAKE_AXI_STATUS_RUN_ERROR=0 FM_FAKE_AXI_STATUS_RUN="" FM_FAKE_RUNS_LIST="" FM_FAKE_BUSY=0 @@ -294,7 +309,8 @@ reset_fakes() { unset FM_FAKE_PR_47_STATE FM_FAKE_PR_47_MERGED FM_FAKE_PR_48_STATE FM_FAKE_PR_48_MERGED export FM_FAKE_AXI_STATUS FM_FAKE_AXI_STATUS_RUN FM_FAKE_RUNS_LIST FM_FAKE_BUSY FM_FAKE_BUSY_TEXT FM_FAKE_TMUX_MISSING FM_FAKE_TMUX_UNREADABLE export FM_FAKE_HERDR_BUSY FM_FAKE_HERDR_MISSING FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK FM_FAKE_HERDR_AGENT_STATUS FM_FAKE_HERDR_PROCESS FM_FAKE_HERDR_SHELL_PID FM_FAKE_CI_LOGS - export FM_FAKE_DAEMON_DOWN + export FM_FAKE_DAEMON_DOWN FM_FAKE_AXI_HOME + export FM_FAKE_AXI_HOME_ERROR FM_FAKE_AXI_STATUS_RUN_ERROR FM_FAKE_AXI_STATUS_ERROR export FM_FAKE_PR_STATE FM_FAKE_PR_MERGED FM_FAKE_PR_READ_FAIL FM_FAKE_PR_READ_LOG FM_FAKE_PR_STATE_AXI export FM_FAKE_GLAB_STATE FM_FAKE_GLAB_READ_FAIL FM_FAKE_GLAB_READ_LOG export FM_FAKE_PR_47_STATE FM_FAKE_PR_47_MERGED FM_FAKE_PR_48_STATE FM_FAKE_PR_48_MERGED @@ -1440,13 +1456,10 @@ EOF pass "cross-branch attribution picks the branch's most recent row" } -# Live-over-terminal selection (bin/fm-nm-run-lib.sh). Reproduces the proven -# 2026-08 case: a crashed validation daemon left a FAILED run at the worktree's -# exact commit, while the live run that replaced it validates a descendant -# commit on the same branch. Both bind - the corpse by the equal-commit rule, -# the live run by the ancestor rule - and bare `axi status` answers with the -# corpse, so every recomputation read a healthy task as failed. -test_terminal_corpse_loses_to_live_run_on_same_branch() { +# The plain ledger is ordered by creation time, not the time a status changed. +# A newer failure must not be hidden by an older live run, even when both heads +# bind to the worktree. These legacy CLI cases lack the AXI identity table. +test_terminal_run_keeps_newer_failure_over_live_sibling() { reset_fakes local d base_head live_head short_base short_live out d=$(new_case live-beats-corpse) @@ -1461,28 +1474,25 @@ test_terminal_corpse_loses_to_live_run_on_same_branch() { [ "$short_base" != "$short_live" ] || fail "live run head did not advance past the worktree" make_fakebin "$d" >/dev/null fm_write_meta "$d/state/corpse.meta" "window=fm:fm-corpse" "worktree=$d/wt" "kind=ship" - # The corpse is the most-recently-touched run, so it is what `axi status` - # reports, at this worktree's own commit. + # The newest run failed at this worktree's own commit. FM_FAKE_RUN_HEAD="$base_head" FM_FAKE_AXI_STATUS="$(run_failed fm/feat-corpse)" - # It is also the newest row in the listing (the crash marked it after the - # live run started), so row order alone still selects the corpse. + # The older live run may have advanced its tip, but it did not replace this run. FM_FAKE_RUNS_LIST="$(cat <