From 3e6899de15fc8ef15a0d9a79009597fc78bcce92 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EB=B0=95=ED=8F=89=EC=8B=9D?= Date: Thu, 1 Oct 2026 17:43:10 +0900 Subject: [PATCH 1/3] fix(spawn): verify the OpenCode model a worker will actually run PR #70 validated an explicit `--model` for OpenCode launches, but only in the launcher's own directory and only when a model was named. Two gaps followed from that, both of which let a worker start on a model nobody checked: - An omitted or `--model default` launch validated nothing, and OpenCode then chose its own default from whatever config the worker directory holds. The gate was bypassed rather than failed. - `opencode models` is project-scoped, so a model supplied by a provider the worktree config declares is listed in the worktree and not in the launcher. Validating an explicit model in the launcher's directory refuses a model the worker can actually run. Both paths now resolve and validate in the settled task worktree, inside the destination pane, under the same environment expression the launch itself gets, so the binary, working directory, allowlisted credentials, and OpenCode server route are the ones the worker will have. `opencode debug config` is read as the source inventory it is (v2) or the resolved object it is (v1), ranked by the documented precedence, and the resulting root model is pinned on the launch. An unresolvable source, an equal-precedence disagreement, or no root model at all fails closed rather than deferring to OpenCode's default. Availability and zero cost are checked the same way for both paths, and `opencode-go/space-bunny-free` is not privileged: any model that clears the gate is accepted. Two things the validation could not see before it saw them: - `opencode models` prints only base ids, so it cannot confirm a `#variant`. The only declaring source is provider metadata in the resolved config, and an available base does not carry an arbitrary variant: the exact `base#variant` must be one the config SELECTS. Selectable variants are resolved through the same precedence order as the root model - union within one precedence step, replaced across steps, including by an empty set, which is how a higher-precedence document disabling every variant clears the model. Comparing declared sets alone would admit two same-step documents that both declare `low` and disagree on whether it is enabled. - The raw-launch escape hatch classified on the executable it could name, so `command opencode`, `exec opencode`, a quoted executable, or a `sh -c opencode` all reached the worker with every model guard skipped. Commands whose whole job is running another program are now walked through to the executable they would run, and a command naming OpenCode anywhere outside the resolved executable position fails closed. A launch that merely mentions OpenCode in a prompt is still an unrelated launch, and `cd && ./probe` is still a legitimate raw command. The shared spawn fixture executes the preflight the way a real pane does, so a suite that spawns OpenCode no longer fails on a timeout instead of on the behavior it means to test, and suites share one OpenCode preflight stub rather than each writing its own. Regression coverage spans the omitted and explicit paths, unavailable and nonzero-cost metadata, declared, disabled-by-precedence, and equal-precedence-conflicting variants, v1 and v2 config shapes, custom config-location overrides, the wrapper and quoting forms above, and the bounded preflight actually cutting off a hung catalog rather than only passing the bound as an argument. Verified against a live OpenCode 2.0.18: `debug config` is a source array carrying `info.model` in object form, `OPENCODE_CONFIG_CONTENT` does not register as a config source, and `opencode models` lists base ids with no `#` forms. --- bin/fm-opencode-probe.sh | 135 ++++ bin/fm-spawn.sh | 664 +++++++++++++++++-- tests/fixtures.sh | 59 ++ tests/fm-busy-adapter-wiring.test.sh | 4 + tests/fm-spawn-dispatch-profile.test.sh | 833 +++++++++++++++++++++++- 5 files changed, 1627 insertions(+), 68 deletions(-) create mode 100755 bin/fm-opencode-probe.sh diff --git a/bin/fm-opencode-probe.sh b/bin/fm-opencode-probe.sh new file mode 100755 index 00000000000..70c1effb8ef --- /dev/null +++ b/bin/fm-opencode-probe.sh @@ -0,0 +1,135 @@ +#!/usr/bin/env bash +# Run OpenCode discovery inside the destination pane's launch environment. +# Only model-related config fields leave the debug-config process. +set -u + +if [ "$#" -ne 5 ]; then + echo "usage: fm-opencode-probe.sh " >&2 + exit 2 +fi + +bin=$1 +worktree=$2 +result_dir=$3 +bound=$4 +resolve_config=$5 +case "$bound" in ''|*[!0-9]*|0*) exit 2 ;; esac +case "$resolve_config" in 0|1) ;; *) exit 2 ;; esac +[ -x "$bin" ] && [ -d "$worktree" ] && [ -d "$result_dir" ] || exit 2 + +# Resolve this script's own directory BEFORE changing directory, from +# BASH_SOURCE rather than $0: fm-spawn invokes it with an absolute path, but a +# relative one would be resolved against the worktree it is about to enter. +script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P) || exit 2 +cd "$worktree" || exit 2 +# shellcheck source=bin/fm-timeout-lib.sh +# shellcheck disable=SC1091 +. "$script_dir/fm-timeout-lib.sh" + +if [ "$resolve_config" = 1 ]; then + if [ -n "${OPENCODE_CONFIG:-}" ] || [ -n "${OPENCODE_CONFIG_DIR:-}" ]; then + printf '%s\n' 'config-location-override' > "$result_dir/config.reason" + printf '78\n' > "$result_dir/config.status" + else + if [ -n "${XDG_CONFIG_HOME:-}" ]; then + global_config="$XDG_CONFIG_HOME/opencode" + elif [ -n "${HOME:-}" ]; then + global_config="$HOME/.config/opencode" + else + printf '78\n' > "$result_dir/config.status" + fi + if [ ! -f "$result_dir/config.status" ]; then + # shellcheck disable=SC2016 # single quotes are deliberate: $provider/$model and $global are jq's, not the shell's + normalize_filter=' + def model_value: + if type == "string" then . + elif type == "object" then + (.providerID // .provider) as $provider + | (.model // .modelID) as $model + | if ($provider | type) == "string" and ($model | type) == "string" then + "\($provider)/\($model)" + + (if (.variant | type) == "string" and .variant != "" then "#\(.variant)" else "" end) + else null end + else null end; + # `opencode models` lists only base ids, never a "#variant" form, so the + # only source that can confirm a selected variant is the provider + # metadata in the resolved config. Both published shapes of that field + # are normalized to the same list, because which one appears is a + # property of the OpenCode version rather than a choice by the caller: + # array of objects v2 - each entry carries its own `id` + # object keyed by id v1 - the variant id is the key + # Any id that cannot be read as a string yields no entry at all, which + # leaves the exact-match requirement in the caller unsatisfied and keeps + # the launch closed. + # Both the DECLARED ids and the SELECTABLE subset are kept, per config + # document and keyed by base id, because a document that disables one + # variant still DECLARES it while selecting none of it. A union across + # documents would let a lower-precedence enabled variant survive a + # higher-precedence disable, so the caller ranks these per document + # rather than merging them here. + def variant_entries: + if type == "array" then [.[] | {id: .id, disabled: .disabled}] + elif type == "object" then [to_entries[] + | {id: .key, disabled: (.value | if type == "object" then .disabled else null end)}] + else [] end; + def variant_map: + if type == "object" then + [ (.providers // {} | to_entries[]) as $provider + | ($provider.value.models // {} | to_entries[]) as $model + | ($model.value.variants // {}) as $declared + # `[]` iterates the normalized entries; without it `as $entry` + # would bind the whole array and the field reads below would fail. + | ($declared | variant_entries)[] + | select((.id | type) == "string" and .id != "") + | {key: "\($provider.key)/\($model.key)", id: .id, off: (.disabled == true)} ] + | group_by(.key) + | map({key: .[0].key, + value: {declared: (map(.id) | unique), + selectable: (map(select(.off | not) | .id) | unique)}}) + | from_entries + else {} end; + def clean_info: + {model: (.model | model_value)} + | with_entries(select(.value != null)); + def entries: + if type == "array" then [.[] | {type, path, info:(.info // {})}] + elif type == "object" then [{type:"document", path:($global + "/opencode.json"), info:.}] + else error("unsupported debug config JSON shape") end; + # The docs list keeps only the decoded model of each source and drops + # everything else: the provider metadata the variant list above reads + # stays behind here, because a config source can carry arbitrary + # provider content the ranking step has no use for. Both passes carry + # their own load. + (entries) as $docs + | {shape:(if type == "array" then "source-array" else "resolved-object" end), + global:$global, + docs:[$docs[] | {type, path, info:(.info | clean_info), variants:(.info | variant_map)}]} + ' + # shellcheck disable=SC2016 # single quotes are deliberate: ${...} expands in the probe's bash -c, not here + if fm_run_timed "$bound" /bin/bash -c ' + "$1" debug config 2>/dev/null | jq -c --arg global "$2" "$3" + pipeline=("${PIPESTATUS[@]}") + [ "${pipeline[0]}" -eq 0 ] || exit "${pipeline[0]}" + exit "${pipeline[1]}" + ' _ "$bin" "$global_config" "$normalize_filter" > "$result_dir/config.json"; then + printf '0\n' > "$result_dir/config.status" + else + status=$? + [ "$status" -ne 0 ] || status=65 + printf '%s\n' "$status" > "$result_dir/config.status" + fi + fi + fi +else + printf '{"shape":"skip"}\n' > "$result_dir/config.json" + printf '0\n' > "$result_dir/config.status" +fi + +if fm_run_timed "$bound" "$bin" models > "$result_dir/models.txt" 2>/dev/null < /dev/null; then + printf '0\n' > "$result_dir/models.status" +else + status=$? + printf '%s\n' "$status" > "$result_dir/models.status" +fi + +: > "$result_dir/done" diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 06b54d5f2e4..5f6cc52c9e4 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -1083,6 +1083,7 @@ SPAWN_TASK_SET_LOCK_HELD=0 SPAWN_TREEHOUSE_PROJECT_LOCK= SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 SPAWN_SLOT_CLAIMED=0 +OPENCODE_PROBE_DIR= RELAUNCH_REPLACEMENT_PENDING=0 RELAUNCH_REPLACEMENT_BUSY_GEN= RELAUNCH_REPLACEMENT_HARNESS= @@ -1120,6 +1121,9 @@ parse_orca_worktree_result() { spawn_abort_cleanup() { local status=$? + if [ -n "${OPENCODE_PROBE_DIR:-}" ] && [ -f "$OPENCODE_PROBE_DIR/done" ]; then + opencode_probe_cleanup || true + fi if [ "$RELAUNCH_REPLACEMENT_PENDING" = 1 ] && [ "$SPAWN_META_PUBLISH_STARTED" = 1 ] && [ -n "$SPAWN_META_TMP" ] && @@ -1530,6 +1534,8 @@ PROJ= ARG3= FIRSTMATE_HOME= RAW_LAUNCH=0 +RAW_OPENCODE=0 +OPENCODE_BIN= # --relaunch adoption: every identity axis comes from the task's own validated # durable record, never from the command line, so a relaunch can only ever @@ -1787,10 +1793,352 @@ agy_model_validate() { # return 1 } +# Build the same destination-pane environment expression used by the launch. +spawn_build_launch_env_prefix() { # + local include_trace=${1:-0} prefix env_name env_arg + [ "$LAUNCH_ENV_ENABLED" = 1 ] || return 0 + prefix='/usr/bin/env -i' + for env_name in HOME PATH USER LOGNAME SHELL TERM COLORTERM LANG LC_ALL LC_CTYPE \ + TMPDIR TMP TEMP GOTMPDIR TMUX TMUX_PANE HERDR_ENV HERDR_SESSION HERDR_SOCKET_PATH \ + HERDR_PANE_ID CMUX_WORKSPACE_ID CMUX_SURFACE_ID CMUX_TAB_ID CMUX_PANEL_ID \ + CMUX_SOCKET_PATH ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION \ + FM_TASK_ID COMPACT_ADVISER_DISABLE \ + $LAUNCH_ENV_NAMES; do + # Only validated names enter shell syntax. Values expand once, quoted, in + # the pane shell and never become source text or spawn-process snapshots. + # shellcheck disable=SC2016 # single quotes are deliberate: ${...} expands in the pane shell, not here + printf -v env_arg '${%s+"%s=$%s"}' "$env_name" "$env_name" "$env_name" + prefix="$prefix $env_arg" + done + prefix="$prefix COMPACT_ADVISER_DISABLE=1" + if [ "$include_trace" = 1 ] && [ -n "${SPAWN_TRACEPARENT:-}" ]; then + # shellcheck disable=SC2016 # single quotes are deliberate: ${...} expands in the pane shell, not here + prefix="$prefix "'${TRACEPARENT+"TRACEPARENT=$TRACEPARENT"}' + fi + printf '%s' "$prefix" +} + +# Probe in the destination pane so the binary, worktree, allowlisted credentials, +# config variables, and OpenCode server route are the same ones the worker gets. +opencode_worker_probe() { # + local bin=$1 dir=$2 resolve_config=$3 bound=${FM_OPENCODE_MODELS_TIMEOUT:-30} + local q_dir q_output q_helper q_bin q_config q_id env_prefix script command max_checks i + case "$bound" in ''|*[!0-9]*|0*) bound=30 ;; esac + case "$resolve_config" in 0|1) ;; *) return 1 ;; esac + OPENCODE_PROBE_DIR="$dir/.fm-opencode-probe-$ID-${BASHPID:-$$}-$RANDOM" + if ! (umask 077 && mkdir "$OPENCODE_PROBE_DIR"); then + echo "error: could not create a private OpenCode probe directory in $dir" >&2 +OPENCODE_PROBE_DIR= + return 1 + fi + q_dir=$(shell_quote "$dir") + q_output=$(shell_quote "$OPENCODE_PROBE_DIR") + q_helper=$(shell_quote "$SCRIPT_DIR/fm-opencode-probe.sh") + q_bin=$(shell_quote "$bin") + q_config=$(shell_quote '{"permission":{"*":"allow"}}') + q_id=$(shell_quote "$ID") + env_prefix=$(spawn_build_launch_env_prefix 0) + script="$OPENCODE_PROBE_DIR/run.sh" + { + printf '#!/bin/sh\ncd -- %s || exit 1\n' "$q_dir" + [ -z "$env_prefix" ] || printf '%s ' "$env_prefix" + printf 'COMPACT_ADVISER_DISABLE=1 ' + [ "$KIND" = secondmate ] || printf 'FM_TASK_ID=%s ' "$q_id" + printf 'OPENCODE_CONFIG_CONTENT=%s /bin/bash %s %s %s %s %s %s\n' \ + "$q_config" "$q_helper" "$q_bin" "$q_dir" "$q_output" "$bound" "$resolve_config" + } > "$script" || return 1 + chmod 700 "$script" || return 1 + command="/bin/sh $(shell_quote "$script")" + spawn_send_text_line "$WT_TARGET" "$command" || { + echo "error: could not run OpenCode model discovery in the destination pane for $dir" >&2 + return 1 + } + max_checks=$((bound * 20 + 100)) + i=0 + while [ "$i" -lt "$max_checks" ] && [ ! -f "$OPENCODE_PROBE_DIR/done" ]; do + sleep 0.1 + i=$((i + 1)) + done + if [ ! -f "$OPENCODE_PROBE_DIR/done" ]; then + echo "error: OpenCode model discovery did not return from the destination pane within $((bound * 2 + 10))s; refusing to launch" >&2 + return 1 + fi + return 0 +} + +opencode_probe_cleanup() { + local dir=${OPENCODE_PROBE_DIR:-} + [ -n "$dir" ] && [ -f "$dir/done" ] || return 1 + rm -f "$dir/run.sh" "$dir/config.json" "$dir/config.status" \ + "$dir/config.reason" "$dir/models.txt" "$dir/models.status" \ + "$dir/effective-variants.json" "$dir/done" || return 1 + rmdir "$dir" || return 1 + OPENCODE_PROBE_DIR= +} + +opencode_model_validate() { # + local result_dir=$1 model=$2 base variant provider id metadata models_status + case "$model" in + *'#'*) + base=${model%%#*} + variant=${model#*#} + case "$variant" in + '' | *'#'*) + echo "error: OpenCode model '$model' has an unsupported variant form; use at most one nonempty #variant" >&2 + return 1 + ;; + esac + ;; + *) + base=$model + variant= + ;; + esac + models_status=$(cat "$result_dir/models.status" 2>/dev/null || true) + if [ "$models_status" != 0 ] || [ ! -f "$result_dir/models.txt" ]; then + echo "error: could not verify OpenCode model '$model' because 'opencode models' failed in the destination pane (exit ${models_status:-unknown}); choose a model listed there" >&2 + return 1 + fi + if ! grep -F -x -- "$base" "$result_dir/models.txt" >/dev/null; then + echo "error: OpenCode model '$model' is not available from 'opencode models' in the destination pane; choose an id listed by that command" >&2 + return 1 + fi + # `opencode models` prints only base ids, so it cannot confirm a variant. The + # declaring source is the provider metadata in the resolved config, so an + # available base does NOT carry an arbitrary variant: require the exact + # "base#variant" the destination config SELECTS. + if [ -n "$variant" ]; then + if [ ! -s "$result_dir/config.json" ]; then + echo "error: could not verify OpenCode model '$model' because the destination pane returned no OpenCode config to declare its variants; pass the base model without #variant" >&2 + return 1 + fi + # The effective variant list is precedence-resolved across every config + # source, so this checks what the worker can actually select rather than + # what any single document mentions. A variant only a lower-precedence + # source declares, or one a higher-precedence source disables, is absent + # here and refused. + if ! jq -e --arg base "$base" --arg variant "$variant" \ + '.effective_variants[$base] | index($variant) != null' \ + "$result_dir/effective-variants.json" >/dev/null 2>&1; then + echo "error: OpenCode model '$model' is not available in the destination pane; no config source there selects the '#$variant' variant of '$base'" >&2 + return 1 + fi + fi + provider=${base%%/*} + id=${base#*/} + if [ "$provider" = "$base" ] || [ -z "$id" ] \ + || ! metadata=$(curl --fail --silent --show-error --location --max-time 10 https://models.dev/api.json); then + echo "error: could not verify OpenCode model '$model' free pricing metadata; refusing dispatch" >&2 + return 1 + fi + if ! printf '%s\n' "$metadata" | jq -e --arg provider "$provider" --arg model "$id" ' + .[$provider].models[$model].cost as $cost + | ($cost | type == "object") + and ($cost.input == 0) + and ($cost.output == 0) + and ([$cost[]] | all(. == 0)) + ' >/dev/null; then + echo "error: OpenCode model '$model' is not classified as free by models.dev; choose a zero-cost model" >&2 + return 1 + fi + return 0 +} + +# OpenCode picks its own model when the launch carries none, so a launch without +# an explicit --model still has an EFFECTIVE model that must clear the same gate. +# `opencode debug config` is a resolved object on v1 and a source inventory on +# v2. The v1 object already reflects the worker's effective config; for v2 this +# reads each listed document and applies the documented precedence itself +# (https://opencode.ai/v2/docs/config): +# the global config is lowest, direct opencode.json(c) files merge from the +# farthest ancestor to the worker directory, and every discovered .opencode +# config overrides every direct config. The root `model` is the default for new +# work; primary-agent selection does not change a session's model, so this +# resolver pins the root model into the canonical launch's --model flag +# (https://opencode.ai/v2/docs/models and https://opencode.ai/v2/docs/agents). +# +# Everything else fails closed rather than guessing: a source this resolver +# cannot rank, a second source disagreeing inside one precedence step, or no +# root model at all. The resolution runs in the worker's own directory with the +# credential and configuration environment the launch itself gets, so a project +# or global config the worker would read is a source this sees too. +# +# The ranked result is the single owner of BOTH answers this gate needs, so they +# can never be computed from different views of the same config: the effective +# root model (`.root`) and the effective selectable variants +# (`effective-variants.json`, which the variant gate reads). Ranking runs even +# when the model was explicit, because an explicit model can still carry a +# "#variant" that has to clear the same precedence. +# +# : prints the ranked object on stdout. +opencode_resolve_effective_model() { # + local probe_dir=$1 dir=$2 need_root=${3:-1} inventory global_config ancestors merged unranked resolved config_status config_reason + command -v jq >/dev/null 2>&1 || { + echo "error: jq is required to resolve the effective OpenCode model; install jq or pass an explicit --model" >&2 + return 1 + } + config_status=$(cat "$probe_dir/config.status" 2>/dev/null || true) + config_reason=$(cat "$probe_dir/config.reason" 2>/dev/null || true) + if [ "$config_reason" = config-location-override ]; then + echo "error: OPENCODE_CONFIG or OPENCODE_CONFIG_DIR is set in the worker environment, but dispatch cannot resolve that custom config location safely; unset the override or pass an explicit --model" >&2 + return 1 + fi + if [ "$config_status" != 0 ] || [ ! -s "$probe_dir/config.json" ]; then + echo "error: could not verify the effective OpenCode model because 'opencode debug config' failed or returned an unsupported shape in the destination pane (exit ${config_status:-unknown}); supported forms are a v1 resolved object and a v2 source array" >&2 + return 1 + fi + global_config=$(jq -r '.global // empty' "$probe_dir/config.json" 2>/dev/null) || global_config= + inventory=$(jq -c '.docs // empty' "$probe_dir/config.json" 2>/dev/null) || inventory= + if [ -z "$global_config" ] || [ -z "$inventory" ] || ! printf '%s' "$inventory" | jq -e 'type == "array"' >/dev/null 2>&1; then + echo "error: 'opencode debug config' did not provide a supported v1 resolved object or v2 config-source list in $dir" >&2 + return 1 + fi + # Farthest ancestor first, so the worker's own directory is the last direct + # config and therefore the highest-precedence one of that step. The upward walk + # is emitted nearest-first and reversed by the jq below, which both orders the + # list and serializes it for --argjson. tac would do the reversal in one step + # but is not a stock macOS command, and jq is already a hard requirement here. + ancestors=$( + walk=$dir + while :; do + printf '%s\n' "$walk" + [ "$walk" = / ] && break + walk=${walk%/*} + [ -n "$walk" ] || walk=/ + done + ) + if ! ancestors_json=$(printf '%s\n' "$ancestors" | jq -Rsc 'split("\n") | map(select(length > 0)) | reverse'); then + echo "error: could not rank OpenCode's config sources in $dir; refusing to launch an OpenCode worker whose effective model cannot be established" >&2 + return 1 + fi + if ! merged=$(printf '%s' "$inventory" | jq -c --argjson ancestors "$ancestors_json" --arg global "$global_config" ' + # The probe already normalized every model into one canonical + # "provider/model#variant" string, so this ranks and compares those strings + # directly rather than decoding the shape a second time. + def rootid: select(type == "string" and length > 0); + def place($path): + if ($path == ($global + "/opencode.json") or $path == ($global + "/opencode.jsonc")) then {tier: 0, step: 0} + else ([$ancestors | to_entries[] + | select($path == (.value + "/opencode.json") or $path == (.value + "/opencode.jsonc")) | .key] | first) as $direct + | if $direct != null then {tier: 1, step: $direct} + else ([$ancestors | to_entries[] + | select($path == (.value + "/.opencode/opencode.json") or $path == (.value + "/.opencode/opencode.jsonc")) | .key] | first) as $dot + | if $dot != null then {tier: 2, step: $dot} else null end + end + end; + [.[] | select(.type == "document") + | {path: .path, info: (.info // {}), variants: (.variants // {}), rank: place(.path)}] as $docs + | ($docs | map(select(.rank != null)) | sort_by([.rank.tier, .rank.step])) as $ranked + # Sources sharing one tier and step are one precedence step, so different + # root models there are unresolvable. Agent model preferences do not select + # the primary session model; the canonical launch pins the root default. + | ($ranked | group_by([.rank.tier, .rank.step])) as $steps + | {unranked: [$docs[] | select(.rank == null) | .path], + root_conflict: (([$steps[] + | [.[] | .info.model | rootid] | unique | length] | max) > 1), + root: ([$ranked[].info.model | rootid] | last // null), + # The ranked documents themselves, grouped by precedence step and already + # in ascending precedence order, so the caller resolves selectable + # variants through the same order without re-deriving the ranking. + steps: [$steps[]]} + ' 2>/dev/null) || ! printf '%s' "$merged" | jq -e 'type == "object"' >/dev/null 2>&1; then + echo "error: could not rank OpenCode's config sources in $dir; refusing to launch an OpenCode worker whose effective model cannot be established" >&2 + return 1 + fi + unranked=$(printf '%s' "$merged" | jq -r '.unranked | join(", ")') + if [ -n "$unranked" ]; then + echo "error: OpenCode reports config source(s) '$unranked' that fall outside the documented precedence this resolver applies, so the effective model in $dir cannot be established; pass an explicit --model" >&2 + return 1 + fi + # Variants resolve through the SAME ranked order as the root model, so a + # higher-precedence source that DISABLES a variant is what decides it: the + # last document that mentions a base id wins for that id, and a document whose + # selectable set is empty therefore clears that id entirely. Two documents in + # ONE precedence step that declare different variant sets for the same base are + # unresolvable, exactly like a root-model disagreement in one step. + if ! printf '%s' "$merged" | jq -S ' + # `.steps` holds the ranked documents grouped by precedence step, in + # ASCENDING precedence order (group_by keeps equal-rank documents adjacent + # and in order). Two rules, because the two situations mean different things: + # + # WITHIN one step the documents are simultaneous alternatives, so their + # selectable sets UNION: two sources of equal precedence are both read, so + # a variant either one selects is selectable. + # ACROSS steps a later step REPLACES the earlier one, INCLUDING with an + # empty selectable set. That empty set is how a higher-precedence document + # that disables every variant of a model clears it, so unioning across + # steps would resurrect exactly what the disable removed. + def union_variants($group): + reduce $group[] as $doc ({}; + reduce (($doc.variants // {}) | to_entries[]) as $kv + (.; .[$kv.key] = (((.[$kv.key] // []) + ($kv.value.selectable // [])) | unique))); + def effective($steps): + reduce $steps[] as $group ({}; + ($group | union_variants($group)) as $step + | reduce ($step | to_entries[]) as $kv (.; .[$kv.key] = $kv.value)); + # Two documents in ONE precedence step must AGREE about a base model, and + # "agree" has to be checked on the SELECTABLE set, not only on the declared + # one. Comparing `declared` alone misses the case that matters: both + # documents declare `low`, one enables it and the other disables it. Their + # declared sets are identical, so that comparison passes, while + # union_variants then hands back the enabled variant as selectable even + # though the effective state is genuinely ambiguous. Requiring identical + # selectable sets refuses that instead. + def step_conflicts($steps): + [ $steps[] as $group + | ( [ $group[] | (.variants // {}) | keys[] ] | unique ) as $keys + | $keys[] as $k + | select( ([ $group[] + | (.variants // {})[$k].selectable // [] ] | unique | length) > 1 ) + | $k ]; + (.steps) as $steps + | {effective_variants: effective($steps), + variant_conflict: ((step_conflicts($steps) | length) > 0), + variant_conflict_base: (step_conflicts($steps) | first)} + ' > "$probe_dir/effective-variants.json" 2>/dev/null; then + echo "error: could not resolve OpenCode's effective model variants in $dir; refusing to launch an OpenCode worker whose selected variant cannot be established" >&2 + return 1 + fi + if jq -e '.variant_conflict' "$probe_dir/effective-variants.json" >/dev/null 2>&1; then + echo "error: two OpenCode config sources of equal precedence declare different variants of '$(jq -r '.variant_conflict_base // "a model"' "$probe_dir/effective-variants.json")' in $dir, so the selectable variants are ambiguous; remove one or pass an explicit --model" >&2 + return 1 + fi + # The root-model answer only matters when dispatch must pick the model itself. + # An explicit model that carries a variant still ranks above for the variant + # gate, but a disagreement between two equal-precedence sources about which + # model to USE is irrelevant to a model the captain already chose. + if [ "$need_root" = 1 ]; then + if printf '%s' "$merged" | jq -e '.root_conflict' >/dev/null; then + echo "error: two OpenCode config sources of equal precedence disagree in $dir, so the effective model is ambiguous; remove one or pass an explicit --model" >&2 + return 1 + fi + resolved=$(printf '%s' "$merged" | jq -r '.root // empty') + if [ -z "$resolved" ]; then + echo "error: no OpenCode root 'model' could be resolved in $dir, so OpenCode would choose an unchecked default; set a root model in its config or pass an explicit --model" >&2 + return 1 + fi + fi + printf '%s\n' "$merged" +} + +# Resolve OpenCode once so the probe and eventual launch use the same executable. +resolve_opencode_binary() { + local candidate dir + candidate=$(command -v opencode 2>/dev/null) || return 1 + [ -x "$candidate" ] || return 1 + case "$candidate" in + /*) printf '%s\n' "$candidate" ;; + *) + dir=$(cd "$(dirname "$candidate")" 2>/dev/null && pwd -P) || return 1 + printf '%s/%s\n' "$dir" "$(basename "$candidate")" + ;; + esac +} + # The verified launch command per adapter. The knowledge half of each adapter # (busy-state source, exit command, dialogs, quirks) lives in the harness-adapters skill. launch_template() { - local harness=$1 kind=${2:-ship} + local harness=$1 kind=${2:-ship} opencode_bin # shellcheck disable=SC2016 # single quotes are deliberate: $(cat ...) expands in the crewmate pane, not here case "$harness" in # CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false disables claude's interactive @@ -1866,11 +2214,13 @@ launch_template() { fi ;; opencode) - mini_help=$(opencode mini --help 2>&1 || :) + opencode_bin=$(resolve_opencode_binary) || return 1 + OPENCODE_BIN=$opencode_bin + mini_help=$("$opencode_bin" mini --help 2>&1 || :) if printf '%s\n' "$mini_help" | grep -Eiq '^[[:space:]]*usage:[[:space:]]*opencode mini([[:space:]]|$)'; then - printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' opencode mini __MODELFLAG__--prompt "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' __OPENCODEBIN__ mini __MODELFLAG__--prompt "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' else - printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' opencode __MODELFLAG__--prompt "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' __OPENCODEBIN__ __MODELFLAG__--prompt "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' fi ;; pi | pi-signed) @@ -2048,14 +2398,194 @@ case "$ARG3" in *' '*) # raw launch command (unverified-adapter escape hatch) RAW_LAUNCH=1 LAUNCH=$ARG3 + # Resolve the EXECUTABLE a raw command actually runs, never arbitrary + # argument text: the structure is leading NAME=value assignments, then + # optionally an `env` wrapper with its own options, then the executable. + # Matching the whole command string instead classified + # `claude --prompt 'review opencode'` as OpenCode and refused an unrelated + # command over prompt text, while stopping at `env` would let + # `env -u FOO opencode ...` skip model validation entirely. HARNESS="" - for word in $LAUNCH; do - case "$word" in [A-Za-z_]*=*) continue ;; *) - HARNESS=$(basename "$word") + raw_ambiguous=0 + raw_quoted=0 + raw_expanded=0 + raw_shell_script=0 + raw_expect_value=0 + raw_in_env=0 + raw_wrapper= + # A shell interpreter is unresolvable by construction: `sh -c opencode ...` + # names a real executable and then runs a different command from its + # arguments. The one legitimate form is running a SCRIPT PATH, where the + # interpreter's argument is a file to read rather than a command string, so + # the script's own contents decide what runs. Detect that here, on the raw + # text, because by the time the walk reaches an executable word the + # interpreter's arguments have not been examined. Note this matches the + # ABSOLUTE interpreter paths the repo's own fixtures use; a bare `sh -c ...` + # carries no absolute path, so it stays refused. + case "$LAUNCH" in + */bin/sh\ * | */bin/bash\ * | */bin/zsh\ * | */bin/dash\ *) + # A shell whose first argument is an absolute script path. + raw_shell_script=1 + ;; + esac + for raw_word in $LAUNCH; do + if [ "$raw_expect_value" = 1 ]; then + raw_expect_value=0 + continue + fi + if [ "$raw_in_env" = 1 ]; then + case "$raw_word" in + # Only value-taking options consume the following word. The no-argument + # flags below consume nothing, so `env -i opencode ...` still reaches the + # executable instead of swallowing it. + -u | --unset | -C | --chdir | -S | --split-string | -n | -p | -g | \ + --nadjust | --adjustment | --priority | -t | --signal | -k | --kill-after) + raw_expect_value=1 + continue + ;; + -i | -0 | --ignore-environment | --null | -- | -v | --verbose | -f | \ + --foreground | -a | --all | --preserve-status | --foreground-only | \ + --kill-when-unfinished | --kill-when-initiated) + continue ;; + # `timeout` is the one wrapper on the list whose DURATION is a bare + # positional rather than an option value, so consume it here. Consuming it + # generically would eat the executable of every other wrapper. + [0-9]*) + if [ "${raw_wrapper:-}" = timeout ]; then + continue + fi + ;; + [A-Za-z_]*=*) continue ;; + -*) + # An option this resolver does not know how to consume. Whether the next + # word is a flag VALUE or the executable decides the whole answer, so it + # refuses rather than guessing. + raw_ambiguous=1 + break + ;; + esac + raw_in_env=0 + fi + case "$raw_word" in + [A-Za-z_]*=*) continue ;; + env | /usr/bin/env | /bin/env | command | exec | nohup | nice | ionice | \ + stdbuf | timeout | setsid | time | sudo | su | xargs | watch | open | caffeinate) + # A wrapper that runs ANOTHER command is transparent here: walk through it + # to the executable it would run. Without this, `command opencode ...` and + # `exec opencode ...` classified as the WRAPPER, walked past every + # `if [ "$HARNESS" = opencode ]` guard, and launched OpenCode unchecked. + # This is a named list on purpose: it is the set of commands whose whole + # job is to run some other command. Each takes its options, and only + # `timeout` also takes a BARE positional (its duration), which is why that + # one is consumed below rather than by a general rule: `timeout 30 + # opencode` is a wrapper plus a value plus the executable, while `nice + # opencode` is a wrapper plus the executable, so a general "wrapper takes + # one bare positional" rule would swallow the executable of every other + # wrapper on the list. + raw_wrapper=$(basename "$raw_word") + raw_in_env=1 + continue + ;; + \'* | \"*) + # This walk splits on whitespace, so a QUOTED executable arrives here with + # its quotes still attached: basename would return `'opencode'`, which + # matches no harness, and the command would run OpenCode anyway behind an + # unrelated-harness classification with every model guard skipped. The real + # spelling is a shell expansion this walk cannot see, so it is + # unclassifiable rather than merely unquoted. + raw_quoted=1 break ;; esac + HARNESS=$(basename "$raw_word") + break + done + # The walk resolves ONE executable, but a compound command can run others: + # `echo x | opencode ...` and `sh -c opencode` both name an executable this + # walk resolves happily (echo, sh) while running OpenCode anyway. Refusing + # every compound command would break the legitimate forms this escape hatch + # exists for (`cd && ./probe`), so the check is narrower and more + # precise: if OpenCode appears as an executable ANYWHERE in the command, the + # command is refused, because then the resolved harness cannot be the whole + # story and the model guard below would apply to the wrong command. A command + # that never mentions OpenCode keeps whatever structure the captain asked for. + # A word only ever RUNS as an executable; inside quotes it is inert text. So the + # question is whether the name `opencode` appears as a BARE word anywhere in + # the command, which means removing every quoted span first and then looking at + # what remains: + # + # claude --prompt 'review opencode' -> quotes removed, no bare name, accepted + # echo x | opencode --prompt -> bare name after a pipe, refused + # $(which opencode) --prompt -> bare name inside a substitution, refused + # cd /tmp && ./probe -> no name at all, accepted + # + # This is what keeps an unrelated launch that merely MENTIONS OpenCode working + # while refusing every form that could actually execute it. + # + # shellcheck disable=SC1003 # the quoted-span delimiters are matched as + # literal characters here; this strips text, it does not emit quoting. + raw_scan_bare=$(printf '%s\n' "$LAUNCH" | sed -e "s/'[^']*'//g" -e 's/"[^"]*"//g') + for raw_scan_word in $raw_scan_bare; do + # Reduce to a command NAME: drop any directory prefix, then strip the + # characters a name cannot contain that word splitting leaves attached, so + # `opencode)` inside `$(which opencode)` still compares equal. Stated as a + # positive test for what may remain rather than a list of punctuation to + # strip, which also avoids a single-quote case pattern (a ShellCheck SC1003 + # hazard). + raw_scan_base=${raw_scan_word##*/} + while [ -n "$raw_scan_base" ]; do + case "${raw_scan_base: -1}" in + [A-Za-z0-9._+-]) break ;; + *) raw_scan_base=${raw_scan_base%?} ;; + esac + done + if [ "$raw_scan_base" = opencode ]; then + # The resolved executable ITSELF is the ordinary case, handled by + # RAW_OPENCODE below with a more specific refusal. This scan is about + # OpenCode appearing somewhere the walk did not resolve as the executable, + # which is what makes the resolved harness an incomplete description of + # what runs. + [ "$HARNESS" = opencode ] || raw_expanded=1 + break + fi done + # A shell INTERPRETER is the same problem with no special character to detect: + # `sh -c opencode --prompt hi` carries a plain executable word, so the walk + # resolves `sh` and stops, while the shell it names runs OpenCode from the very + # next argument. There is no way to read what a shell will do with its + # arguments, so a shell in executable position is unresolvable by construction + # rather than by accident. + raw_shell=0 + case "$HARNESS" in + sh | bash | zsh | dash | ksh | mksh | pdksh | ash | fish | csh | tcsh | busybox) + [ "$raw_shell_script" = 1 ] || raw_shell=1 + ;; + esac + if [ "$raw_quoted" = 1 ]; then + echo "error: raw launch '$LAUNCH' quotes or expands its executable, so dispatch cannot tell which harness and model it runs; use the canonical --harness launch" >&2 + exit 1 + fi + if [ "$raw_ambiguous" = 1 ] || { [ -z "$HARNESS" ] && [ "$raw_in_env" = 1 ]; }; then + # Checked BEFORE the composed-executable refusal so the diagnostic names the + # more specific cause: an unconsumable wrapper option means the executable is + # unknown, which is a narrower and more actionable statement than "this + # command composes something". + # + # Refuse here rather than downstream: every OpenCode model guard lives under + # `if [ "$HARNESS" = opencode ]`, so leaving an unresolved wrapped command + # on the unrelated-harness path would walk straight past all of it. An + # executable this walk cannot name could be OpenCode, so it fails closed + # instead of being waved through as unrelated. + echo "error: raw launch '$LAUNCH' wraps its executable in a form dispatch cannot resolve, so its harness and model cannot be verified; use the canonical --harness launch" >&2 + exit 1 + fi + if [ "$raw_expanded" = 1 ] || [ "$raw_shell" = 1 ]; then + echo "error: raw launch '$LAUNCH' builds or composes its executable, so dispatch cannot tell which harness and model it runs; use the canonical --harness launch" >&2 + exit 1 + fi + if [ "$HARNESS" = opencode ]; then + RAW_OPENCODE=1 + fi ;; '') # No explicit harness: resolve from config. A secondmate AGENT launches on the @@ -2187,37 +2717,13 @@ if [ "$KIND" = secondmate ] && [ -z "$ARG3" ]; then fi fi if [ "$HARNESS" = opencode ]; then - OPENCODE_BIN=$(command -v opencode) || { + OPENCODE_BIN=${OPENCODE_BIN:-$(resolve_opencode_binary)} || { echo "error: opencode executable not found on PATH; install it or select a different verified harness" >&2 exit 1 } - if [ -n "$MODEL" ] && [ "$MODEL" != default ]; then - if ! OPENCODE_MODELS=$("$OPENCODE_BIN" models); then - echo "error: could not verify OpenCode model '$MODEL' because '$OPENCODE_BIN models' failed; rerun 'opencode models' and choose a listed id" >&2 - exit 1 - fi - if ! printf '%s\n' "$OPENCODE_MODELS" | grep -F -x -- "$MODEL" >/dev/null; then - echo "error: OpenCode model '$MODEL' is not available from 'opencode models'; choose an id listed by that command or omit --model" >&2 - exit 1 - fi - OPENCODE_MODEL_PROVIDER=${MODEL%%/*} - OPENCODE_MODEL_ID=${MODEL#*/} - if [ "$OPENCODE_MODEL_PROVIDER" = "$MODEL" ] || [ -z "$OPENCODE_MODEL_ID" ] \ - || ! OPENCODE_MODEL_METADATA=$(curl --fail --silent --show-error --location --max-time 10 https://models.dev/api.json); then - echo "error: could not verify OpenCode model '$MODEL' free pricing metadata; refusing dispatch" >&2 - exit 1 - fi - if ! printf '%s\n' "$OPENCODE_MODEL_METADATA" | jq -e --arg provider "$OPENCODE_MODEL_PROVIDER" --arg model "$OPENCODE_MODEL_ID" ' - .[$provider].models[$model].cost as $cost - | ($cost | type == "object") - and ($cost.input == 0) - and ($cost.output == 0) - and ([$cost[]] | all(. == 0)) - ' >/dev/null; then - echo "error: OpenCode model '$MODEL' is not classified as free by models.dev; choose a zero-cost model" >&2 - exit 1 - fi - fi + # Model resolution and validation both run below, once the worker directory + # has settled: that directory, not this one, is the config and provider scope + # the launched worker actually gets. fi # Ultra is an explicit native capability, never a Pi thinking-level alias. # Validate the fully resolved profile before worktree or endpoint provisioning. @@ -3944,6 +4450,59 @@ if [ "$RELAUNCH" -eq 0 ] && [ "$KIND" != secondmate ]; then freshen_spawn_worktree_base "$WT" || exit 1 fi +# Every OpenCode launch runs on one effective model, and OpenCode decides it +# from the configuration and providers it reads in the directory the worker +# starts in. That directory is known only now, so this is the earliest point +# where the model can be resolved in the same scope the worker will use, and the +# last chance to refuse before any per-task state is created. BOTH paths run +# here, because `opencode models` is itself project-scoped: a model available +# only through a provider the worktree config declares is listed in this +# directory and not in the launcher's, so validating an explicit --model +# anywhere else would refuse a model the worker can actually run. +if [ "$HARNESS" = opencode ]; then + # The config inventory is needed to resolve an omitted/default model AND to + # resolve the SELECTABLE variant set a selected "#variant" is checked against. + # A plain explicit base model needs neither, so it skips both. + # A raw OpenCode command is refused before any probe: fm-spawn cannot inject + # the model it validated into an arbitrary raw command string, so the worker + # would run whatever model that command names. Refusing here rather than after + # resolution keeps the cause the actual one, instead of reporting an unrelated + # missing project model for a launch that is refused either way. + if [ "$RAW_LAUNCH" = 1 ] || [ "$RAW_OPENCODE" = 1 ]; then + echo "error: a raw OpenCode launch cannot be pinned to the model validated for this task; use the canonical --harness opencode launch so the checked model is passed to the worker" >&2 + exit 1 + fi + # The effective-variant gate needs the ranked config pass for an EXPLICIT + # "#variant" too, not only for an omitted model: the variant has to clear the + # same precedence resolution. A plain explicit base model needs neither the + # ranked pass nor the extra read, so it skips them. + OPENCODE_RESOLVE_CONFIG=0 + case "$MODEL" in + '' | default | *'#'*) OPENCODE_RESOLVE_CONFIG=1 ;; + esac + opencode_worker_probe "$OPENCODE_BIN" "$WT" "$OPENCODE_RESOLVE_CONFIG" || exit 1 + if [ "$OPENCODE_RESOLVE_CONFIG" = 1 ]; then + # Ranked once: this writes the effective variant list that + # opencode_model_validate reads for a selected variant, and publishes the + # effective root model the omitted path then pins. An explicit model that + # carries a variant still needs this, so it passes need-root=0 and is not + # refused for an absent root model it does not depend on. + OPENCODE_NEED_ROOT=0 + if [ -z "$MODEL" ] || [ "$MODEL" = default ]; then + OPENCODE_NEED_ROOT=1 + fi + OPENCODE_RANKED=$(opencode_resolve_effective_model "$OPENCODE_PROBE_DIR" "$WT" "$OPENCODE_NEED_ROOT") || exit 1 + if [ "$OPENCODE_NEED_ROOT" = 1 ]; then + MODEL=$(printf '%s' "$OPENCODE_RANKED" | jq -r '.root') + fi + fi + opencode_model_validate "$OPENCODE_PROBE_DIR" "$MODEL" || exit 1 + opencode_probe_cleanup || { + echo "error: could not retire the completed private OpenCode preflight in $WT" >&2 + exit 1 + } +fi + # Pre-register Claude's workspace trust for the directory this launch starts in, # at the first point that directory is known and before any per-task state is # created below. The dialog gates the pane before the brief is ever read, and it @@ -4633,6 +5192,10 @@ sq_ompext=$(shell_quote "$STATE/$ID.omp-ext.ts") sq_ompcfg=$(shell_quote "${OMP_WORKER_CFG:-$FM_ROOT/.omp/fm-worker-overlay.yml}") sq_opinput=$(shell_quote "$FM_ROOT/bin/fm-operational-input.sh") sq_worktree=$(shell_quote "$WT") +if [ "$HARNESS" = opencode ] && [ "$RAW_LAUNCH" = 0 ]; then + sq_opencode_bin=$(shell_quote "$OPENCODE_BIN") + LAUNCH=${LAUNCH//__OPENCODEBIN__/$sq_opencode_bin} +fi MODELFLAG=$(model_flag_for_harness "$HARNESS" "$MODEL") EFFORTFLAG=$(effort_flag_for_harness "$HARNESS" "$EFFORT" "$MODEL") || exit 1 LAUNCH=${LAUNCH//__MODELFLAG__/$MODELFLAG} @@ -4776,36 +5339,7 @@ if [ -n "$SPAWN_TRACEPARENT" ]; then fi fi if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then - LAUNCH_ENV_PREFIX='/usr/bin/env -i' - # COMPACT_ADVISER_DISABLE is the intentional declarative floor-membership - # entry; the explicit COMPACT_ADVISER_DISABLE=1 assignment below is the - # authoritative setter. - for env_name in HOME PATH USER LOGNAME SHELL TERM COLORTERM LANG LC_ALL LC_CTYPE \ - TMPDIR TMP TEMP GOTMPDIR TMUX TMUX_PANE HERDR_ENV HERDR_SESSION HERDR_SOCKET_PATH \ - HERDR_PANE_ID CMUX_WORKSPACE_ID CMUX_SURFACE_ID CMUX_TAB_ID CMUX_PANEL_ID \ - CMUX_SOCKET_PATH ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION \ - FM_TASK_ID COMPACT_ADVISER_DISABLE \ - $LAUNCH_ENV_NAMES; do - # Only validated names enter shell syntax. Values expand once, quoted, in - # the pane shell and never become source text or spawn-process snapshots. - # shellcheck disable=SC2016 - printf -v env_arg '${%s+"%s=$%s"}' "$env_name" "$env_name" "$env_name" - LAUNCH_ENV_PREFIX="$LAUNCH_ENV_PREFIX $env_arg" - done - # COMPACT_ADVISER_DISABLE is retained by the floor loop above, which forwards - # whatever the pane export set, and then pinned here to the one value Firstmate - # launches on. The literal assignment comes last deliberately: `env` applies - # assignments left to right, so this one wins over a forwarded pane value, and - # it still delivers the switch on a pane whose export never landed. Unlike the - # trace carrier below it carries no gate, so it is appended unconditionally. - # Setting it here rather than relying on the assignment already carried by - # $LAUNCH is what gives the wrapping `/bin/sh` itself the switch, not only the - # agent command it runs. - LAUNCH_ENV_PREFIX="$LAUNCH_ENV_PREFIX COMPACT_ADVISER_DISABLE=1" - if [ -n "$SPAWN_TRACEPARENT" ]; then - # shellcheck disable=SC2016 - LAUNCH_ENV_PREFIX="$LAUNCH_ENV_PREFIX "'${TRACEPARENT+"TRACEPARENT=$TRACEPARENT"}' - fi + LAUNCH_ENV_PREFIX=$(spawn_build_launch_env_prefix 1) LAUNCH="$LAUNCH_ENV_PREFIX /bin/sh -c $(shell_quote "$LAUNCH")" fi # Implement the launch-delivery contract in this script's header. The full diff --git a/tests/fixtures.sh b/tests/fixtures.sh index 559cd661eef..50d7cb434ec 100755 --- a/tests/fixtures.sh +++ b/tests/fixtures.sh @@ -120,6 +120,29 @@ case "${1:-}" in ;; has-session|new-session|new-window|kill-window|set-window-option) exit 0 ;; send-keys) + # A real destination pane EXECUTES the OpenCode preflight command fm-spawn + # types into it and waits on, so a fake pane has to execute it too. This is + # the shared fake tmux every spawn suite uses, which is why it runs whenever + # the payload is a probe script: gating it behind an opt-in that only one + # suite sets left every other OpenCode spawn suite hanging on a probe that + # never wrote its `done` marker, and then failing on a timeout instead of on + # the behavior it meant to test. A suite opts OUT with + # FM_FAKE_EXEC_OPENCODE_PROBES=0 when it asserts the refusing path instead. + if [ "${FM_FAKE_EXEC_OPENCODE_PROBES:-1}" = 1 ]; then + for a in "$@"; do + case "$a" in + "/bin/sh '"*) + probe_script=${a#/bin/sh \'} + probe_script=${probe_script%\'} + case "$probe_script" in + */.fm-opencode-probe-*/run.sh) + (cd "${FM_FAKE_PANE_PATH:?}" && /bin/sh "$probe_script") || true + ;; + esac + ;; + esac + done + fi if [ -n "${FM_FAKE_LAUNCH_LOG:-}" ]; then prev= for a in "$@"; do @@ -291,6 +314,42 @@ Exercise the spawn behavior under test. EOF } +# fm_test_make_opencode_preflight_stub +# Replaces a bare exit-0 `opencode` with one that answers the two commands +# fm-spawn's OpenCode model preflight reads before it launches a worker: +# `opencode models` (the catalog) and `opencode debug config` (the config-source +# inventory). Every OpenCode launch verifies its model first, so a suite that +# spawns OpenCode through the fake pane needs both answers or its spawn is +# refused on the preflight rather than on the behavior it meant to test. +# FM_FAKE_OPENCODE_PREFLIGHT_MODEL (default opencode-go/space-bunny-free) names +# the model the inventory declares and the catalog lists; FM_TEST_OPENCODE_PREFLIGHT_MODELS +# overrides the catalog when a case needs a different set. +fm_test_make_opencode_preflight_stub() { + local fakebin=$1 + cat > "$fakebin/opencode" <<'SH' +#!/usr/bin/env bash +model=${FM_FAKE_OPENCODE_PREFLIGHT_MODEL:-opencode-go/space-bunny-free} +case "${1:-}" in +models) + [ -n "${FM_FAKE_OPENCODE_PREFLIGHT_MODELS:-}" ] && printf '%b\n' "$FM_FAKE_OPENCODE_PREFLIGHT_MODELS" + printf '%s\n' "$model" + exit 0 + ;; +debug) + # A v2 source inventory: one document at the directory OpenCode reports it + # read, declaring the model as the object form the v2 debug output publishes. + if [ "${2:-}" = config ]; then + printf '[{"type":"document","path":"%s/opencode.json","info":{"model":{"providerID":"%s","model":"%s"}}}]\n' \ + "$PWD" "${model%%/*}" "${model#*/}" + exit 0 + fi + ;; +esac +exit 0 +SH + chmod +x "$fakebin/opencode" +} + # fm_test_make_spawn_fakebin [extra-exit0-tool...] # Creates /fakebin with the spawn tmux stub, a no-op treehouse, and any # extra exit-0 tools. Echoes the fakebin path. diff --git a/tests/fm-busy-adapter-wiring.test.sh b/tests/fm-busy-adapter-wiring.test.sh index 7a93af888fa..addda6448ba 100755 --- a/tests/fm-busy-adapter-wiring.test.sh +++ b/tests/fm-busy-adapter-wiring.test.sh @@ -24,6 +24,10 @@ make_spawn_case() { # proj="$case_dir/project" wt="$case_dir/wt" fakebin=$(make_spawn_fakebin "$case_dir/fake" pi opencode claude codex gemini) + # The OpenCode cases here spawn without a model, so fm-spawn resolves and + # validates the effective model before launch. A bare exit-0 opencode cannot + # answer that preflight, so give it a catalog and a config inventory. + fm_test_make_opencode_preflight_stub "$fakebin" fm_test_spawn_home "$home" "$harness" fm_git_worktree "$proj" "$wt" "wt-$name" fm_test_spawn_brief "$home" "$id" diff --git a/tests/fm-spawn-dispatch-profile.test.sh b/tests/fm-spawn-dispatch-profile.test.sh index 6aefac037c7..0889c71daac 100755 --- a/tests/fm-spawn-dispatch-profile.test.sh +++ b/tests/fm-spawn-dispatch-profile.test.sh @@ -34,8 +34,29 @@ SH make_spawn_fakebin() { local dir=$1 fakebin fakebin=$(fm_test_make_spawn_fakebin "$dir") - cat > "$fakebin/timeout" <<'SH' +cat > "$fakebin/timeout" <<'SH' #!/usr/bin/env bash +# bin/fm-timeout-lib.sh invokes a real timeout as `-k ` +# (fm_run_external_timeout), so the stub drops the kill-after flag and its value +# plus the duration. A stub that only shifted once left every bounded probe +# failing with a 127 that no assertion could explain. +# +# This stub only RECORDS the arguments it was given and then execs: it cannot +# enforce a deadline, so an assertion about it proves the bound was passed, not +# that a hung command is killed. test_opencode_probe_enforces_the_catalog_deadline +# covers the enforcement itself against the real runner. +[ -n "${FM_FAKE_TIMEOUT_ARGS:-}" ] && printf '%s\n' "$*" >> "$FM_FAKE_TIMEOUT_ARGS" +while [ $# -gt 0 ]; do + case "$1" in + -k | --kill-after) + [ "$#" -ge 3 ] || exit 125 + shift 2 + ;; + -k* | --kill-after=*) shift ;; + *) break ;; + esac +done +case "${1:-}" in ''|*[!0-9]*) exit 125 ;; esac shift exec "$@" SH @@ -49,7 +70,20 @@ exit 0 SH cat > "$fakebin/opencode" <<'SH' #!/usr/bin/env bash +record_probe_context() { + [ -n "${FM_FAKE_OPENCODE_CONTEXT_LOG:-}" ] || return 0 + printf '%s\t%s\t%s\t%s\n' "$*" "$PWD" "${OPENCODE_SERVER-unset}" \ + "${FM_TEST_SHOULD_NOT_LEAK-unset}" >> "$FM_FAKE_OPENCODE_CONTEXT_LOG" +} if [ "${1:-}" = models ]; then + record_probe_context "$@" + # A catalog that never answers, for proving the preflight's deadline is + # enforced rather than merely passed as an argument. The bound kills this + # child, so it only has to outlast the configured timeout. + if [ "${FM_FAKE_OPENCODE_MODELS_HANG:-0}" = 1 ]; then + sleep 600 + exit 0 + fi # OpenCode v2 refuses a provider positional argument ("Unexpected positional # argument") and prints usage text instead of a catalog, so this stub does # the same. A regression to the legacy `opencode models ` form @@ -59,12 +93,55 @@ if [ "${1:-}" = models ]; then exit 1 fi [ -n "${FM_FAKE_OPENCODE_MODELS_ARGS:-}" ] && printf '%s\n' "$@" > "$FM_FAKE_OPENCODE_MODELS_ARGS" + # `opencode models` is project-scoped: a provider the worktree config declares + # is listed only when the probe runs in that worktree. FM_FAKE_OPENCODE_MODELS_IN_DIR + # supplies that per-directory catalog, keyed by the directory the probe ran in, + # so a probe from the wrong directory is distinguishable from a right one. + [ -n "${FM_FAKE_OPENCODE_MODELS_CWD:-}" ] && pwd > "$FM_FAKE_OPENCODE_MODELS_CWD" + if [ -n "${FM_FAKE_OPENCODE_MODELS_IN_DIR:-}" ]; then + scoped=$(printf '%s' "$FM_FAKE_OPENCODE_MODELS_IN_DIR" | jq -r --arg d "$PWD" '.[$d] // empty') + if [ -n "$scoped" ]; then + printf '%b\n' "$scoped" + exit 0 + fi + fi [ "${FM_FAKE_OPENCODE_MODELS_STATUS:-0}" -eq 0 ] || exit "${FM_FAKE_OPENCODE_MODELS_STATUS}" printf '%b\n' "${FM_FAKE_OPENCODE_MODELS:-anthropic/claude-sonnet-4-5\\nopencode-go/space-bunny-free}" fi +if [ "${1:-}" = debug ] && [ "${2:-}" = config ]; then + record_probe_context "$@" + # The source inventory `opencode debug config` reports: one document entry per + # config file, carrying that file's own parsed content. It is an inventory, not + # a merged answer, so a case supplies the sources it wants ranked and fm-spawn + # resolves precedence over them. A case supplies entries already carrying + # absolute paths, because reading real config files would leave the pooled + # worktree dirty and the spawn's own cleanliness guard would refuse first. + [ -n "${FM_FAKE_OPENCODE_DEBUG_ARGS:-}" ] && printf '%s\n' "$@" > "$FM_FAKE_OPENCODE_DEBUG_ARGS" + [ -n "${FM_FAKE_OPENCODE_DEBUG_CWD:-}" ] && pwd > "$FM_FAKE_OPENCODE_DEBUG_CWD" + [ "${FM_FAKE_OPENCODE_DEBUG_STATUS:-0}" -eq 0 ] || exit "${FM_FAKE_OPENCODE_DEBUG_STATUS}" + if [ -n "${FM_FAKE_OPENCODE_DEBUG_OBJECT:-}" ]; then + printf '%s\n' "$FM_FAKE_OPENCODE_DEBUG_OBJECT" + exit 0 + fi + if [ "${FM_FAKE_OPENCODE_DEBUG_NOT_A_LIST:-0}" = 1 ]; then + printf '%s\n' 'not a config source list' + exit 0 + fi + printf '%s\n' "${FM_FAKE_OPENCODE_DEBUG_DOCS:-[]}" +fi exit 0 SH chmod +x "$fakebin/timeout" "$fakebin/cursor-agent" "$fakebin/opencode" + # `tac` is GNU coreutils and is NOT a stock macOS command, so an ancestor + # ordering that shells out to it works on a Homebrew host and breaks on a + # plain one. This stub fails the call, which turns any such dependency into a + # test failure here instead of a launch-time surprise on the captain's machine. + cat > "$fakebin/tac" <<'SH' +#!/usr/bin/env bash +echo 'tac: unavailable, as on stock macOS' >&2 +exit 127 +SH + chmod +x "$fakebin/tac" cat > "$fakebin/curl" <<'SH' #!/usr/bin/env bash [ "${FM_FAKE_MODELS_DEV_STATUS:-0}" -eq 0 ] || exit "${FM_FAKE_MODELS_DEV_STATUS}" @@ -121,11 +198,26 @@ run_spawn() { # A test opts in to the set case via FM_TEST_CLAUDE_CONFIG_DIR. CLAUDE_CONFIG_DIR="${FM_TEST_CLAUDE_CONFIG_DIR:-}" \ FM_FAKE_LAUNCH_LOG="$launchlog" FM_FAKE_PI_VERSION="${FM_TEST_PI_VERSION:-0.84.0}" \ + FM_FAKE_EXEC_OPENCODE_PROBES="${FM_TEST_EXEC_OPENCODE_PROBES:-1}" \ FM_FAKE_CURSOR_MODELS="${FM_TEST_CURSOR_MODELS:-}" \ FM_FAKE_CURSOR_LIST_STATUS="${FM_TEST_CURSOR_LIST_STATUS:-0}" \ FM_FAKE_OPENCODE_MODELS="${FM_TEST_OPENCODE_MODELS:-}" \ FM_FAKE_OPENCODE_MODELS_STATUS="${FM_TEST_OPENCODE_MODELS_STATUS:-0}" \ FM_FAKE_OPENCODE_MODELS_ARGS="${FM_TEST_OPENCODE_MODELS_ARGS:-}" \ + FM_FAKE_OPENCODE_MODELS_HANG="${FM_TEST_OPENCODE_MODELS_HANG:-}" \ + FM_FAKE_OPENCODE_MODELS_CWD="${FM_TEST_OPENCODE_MODELS_CWD:-}" \ + FM_FAKE_OPENCODE_MODELS_IN_DIR="${FM_TEST_OPENCODE_MODELS_IN_DIR:-}" \ + FM_FAKE_OPENCODE_DEBUG_STATUS="${FM_TEST_OPENCODE_DEBUG_STATUS:-0}" \ + FM_FAKE_OPENCODE_DEBUG_NOT_A_LIST="${FM_TEST_OPENCODE_DEBUG_NOT_A_LIST:-0}" \ + FM_FAKE_OPENCODE_DEBUG_DOCS="${FM_TEST_OPENCODE_DEBUG_DOCS:-}" \ + FM_FAKE_OPENCODE_DEBUG_OBJECT="${FM_TEST_OPENCODE_DEBUG_OBJECT:-}" \ + FM_FAKE_OPENCODE_DEBUG_ARGS="${FM_TEST_OPENCODE_DEBUG_ARGS:-}" \ + FM_FAKE_OPENCODE_DEBUG_CWD="${FM_TEST_OPENCODE_DEBUG_CWD:-}" \ + FM_FAKE_OPENCODE_CONTEXT_LOG="${FM_TEST_OPENCODE_CONTEXT_LOG:-}" \ + FM_FAKE_TIMEOUT_ARGS="${FM_TEST_TIMEOUT_ARGS:-}" \ + OPENCODE_CONFIG="${FM_TEST_OPENCODE_CONFIG:-}" \ + OPENCODE_CONFIG_DIR="${FM_TEST_OPENCODE_CONFIG_DIR:-}" \ + OPENCODE_SERVER="${FM_TEST_OPENCODE_SERVER:-}" \ FM_FAKE_MODELS_DEV_JSON="${FM_TEST_MODELS_DEV_JSON:-}" \ FM_FAKE_MODELS_DEV_STATUS="${FM_TEST_MODELS_DEV_STATUS:-0}" \ GROK_HOME="$home/grok-home" \ @@ -412,7 +504,7 @@ test_explicit_opencode_task_model_overrides_configured_fallback() { expect_code 0 "$status" "listed task-specific OpenCode model should override the configured fallback: $out" assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/longcat-2.5-preview-free default launch=$(cat "$LAUNCH_LOG") - assert_contains "$launch" "opencode --model 'opencode-go/longcat-2.5-preview-free' --prompt" \ + assert_contains "$launch" "'$FAKEBIN_DIR/opencode' --model 'opencode-go/longcat-2.5-preview-free' --prompt" \ "task-specific OpenCode model was not passed to the worker" assert_not_contains "$launch" "--model 'opencode-go/space-bunny-free'" \ "configured fallback replaced the explicitly designated task model" @@ -846,6 +938,722 @@ test_opencode_catalog_probe_uses_no_provider_argument() { pass "OpenCode reads the whole model catalog without a provider argument" } +# An OpenCode launch with no --model still has an effective model, decided by the +# config OpenCode reads in the worker's own directory. These cases prove the +# omitted/default path resolves that model, validates it like an explicit one, and +# pins it on the launch, so no unchecked implicit default reaches a worker. +# opencode_config_docs builds the `opencode debug config` inventory +# a case reports: one document entry, at that exact config path, with that file's +# own parsed content. Writing a real file into the pooled worktree is not an +# option because the spawn's own cleanliness guard refuses a dirty worktree first. +opencode_config_docs() { + jq -cn --arg path "$1" --argjson info "$2" \ + '[{"type":"document","path":$path,"info":$info}]' +} + +test_opencode_omitted_model_resolves_validated_project_model() { + local rec id out status launch cwd_file docs + id=profile-opencode-default-model-z7h + rec=$(make_spawn_case profile-opencode-default-model opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/longcat-2.5-preview-free"}') + cwd_file="$CASE_DIR/debug-cwd" + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" FM_TEST_OPENCODE_DEBUG_CWD="$cwd_file" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "an OpenCode launch with no --model should resolve and validate its project model: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/longcat-2.5-preview-free default + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "'$FAKEBIN_DIR/opencode' --model 'opencode-go/longcat-2.5-preview-free' --prompt" \ + "the resolved effective model was not pinned on the launch" + [ "$(cat "$cwd_file" 2>/dev/null)" = "$WT_DIR" ] \ + || fail "effective-model resolution did not read the config in the worker's own directory" + pass "an omitted OpenCode model resolves in the worker directory, is validated, and is pinned on the launch" +} + +test_opencode_default_model_token_resolves_like_an_omitted_one() { + local rec id out status docs + id=profile-opencode-default-token-z7i + rec=$(make_spawn_case profile-opencode-default-token opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/longcat-2.5-preview-free"}') + + # `--model default` is the recorded "no choice made" value, not a model id, so + # it takes the resolution path instead of being validated as a literal id. + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model default) + status=$? + expect_code 0 "$status" "--model default should resolve the effective model, not be validated as an id: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/longcat-2.5-preview-free default + pass "an explicit 'default' model token takes the same resolved-model path as an omitted one" +} + +test_opencode_omitted_model_honors_documented_config_precedence() { + local rec id out status docs + id=profile-opencode-precedence-z7j + rec=$(make_spawn_case profile-opencode-precedence opencode "$id") + read_case_record "$rec" + # Documented precedence: direct configs merge farthest to nearest, and every + # .opencode config outranks every direct config. The nearest direct config here + # is the worker's own, and the .opencode config below outranks it. + docs=$(jq -cn \ + --arg far "$CASE_DIR/opencode.json" --arg near "$WT_DIR/opencode.json" --arg dot "$WT_DIR/.opencode/opencode.json" \ + '[{"type":"document","path":$far,"info":{"model":"opencode-go/space-bunny-free"}}, + {"type":"document","path":$near,"info":{"model":"opencode-go/longcat-2.5-preview-free"}}, + {"type":"document","path":$dot,"info":{"model":"opencode-go/space-bunny-free"}}]') + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "a .opencode config should win over the direct config: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/space-bunny-free default + pass "effective-model resolution applies the documented OpenCode config precedence" +} + +test_opencode_primary_agent_model_does_not_override_the_session_model() { + local rec id out status docs launch + id=profile-opencode-agent-model-z7k + rec=$(make_spawn_case profile-opencode-agent-model opencode "$id") + read_case_record "$rec" + docs=$(jq -cn --arg far "$CASE_DIR/opencode.json" --arg near "$WT_DIR/opencode.json" \ + '[{"type":"document","path":$far,"info":{"model":"opencode-go/longcat-2.5-preview-free","agents":{"review":{"model":{"providerID":"opencode-go","model":"paid-candidate"}}}}}, + {"type":"document","path":$near,"info":{"default_agent":"review","agents":{"review":{"mode":"primary"}}}}]') + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"longcat-2.5-preview-free":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "a custom primary agent should keep the configured session model: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/longcat-2.5-preview-free default + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "--model 'opencode-go/longcat-2.5-preview-free'" \ + "a primary agent's model preference replaced the session model in the launch" + pass "v2 custom primary agent selection leaves the configured session model in force" +} + +test_opencode_omitted_model_refuses_when_no_effective_model_exists() { + local rec id out status docs + id=profile-opencode-no-model-z7l + rec=$(make_spawn_case profile-opencode-no-model opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"shell":"bash"}') + + # No config declares any model, so OpenCode would fall back to its newest + # available supported model. That is exactly the unchecked implicit default this gate + # exists to refuse. + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 1 "$status" "an OpenCode launch with no resolvable effective model must refuse" + assert_contains "$out" "no OpenCode root 'model' could be resolved" \ + "unresolvable-model refusal did not name the resolution failure" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "an unresolvable effective model published metadata" + [ ! -s "$LAUNCH_LOG" ] || fail "an unresolvable effective model launched an agent" + pass "OpenCode dispatch refuses when no effective model can be resolved" +} + +test_opencode_omitted_model_refuses_when_config_sources_are_ambiguous() { + local rec id out status docs + id=profile-opencode-ambiguous-z7m + rec=$(make_spawn_case profile-opencode-ambiguous opencode "$id") + read_case_record "$rec" + # opencode.json and opencode.jsonc in the same directory are the same + # precedence step, so two different root models there are unresolvable. + docs=$(jq -cn --arg json "$WT_DIR/opencode.json" --arg jsonc "$WT_DIR/opencode.jsonc" \ + '[{"type":"document","path":$json,"info":{"model":"opencode-go/space-bunny-free"}}, + {"type":"document","path":$jsonc,"info":{"model":"opencode-go/longcat-2.5-preview-free"}}]') + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 1 "$status" "equal-precedence config sources must refuse rather than pick one" + assert_contains "$out" "of equal precedence disagree" \ + "ambiguity refusal did not explain the disagreement" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "ambiguous config sources published metadata" + [ ! -s "$LAUNCH_LOG" ] || fail "ambiguous config sources launched an agent" + pass "OpenCode dispatch refuses when two equal-precedence config sources disagree" +} + +test_opencode_custom_primary_agent_without_a_model_inherits_the_session_model() { + local rec id out status docs + id=profile-opencode-inherit-agent-model-z7n + rec=$(make_spawn_case profile-opencode-inherit-agent-model opencode "$id") + read_case_record "$rec" + # A custom primary agent without its own model inherits the configured + # session model; the agent ID itself does not change what --prompt receives. + docs=$(opencode_config_docs "$WT_DIR/opencode.json" \ + '{"default_agent":"review","model":"opencode-go/space-bunny-free","agents":{"review":{"mode":"primary"}}}') + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "a custom primary agent should inherit the resolved session model: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/space-bunny-free default + [ -s "$LAUNCH_LOG" ] || fail "a resolvable custom primary agent did not launch" + pass "a custom primary agent without a model inherits the resolved session model" +} + +test_opencode_omitted_model_refuses_paid_and_unavailable_effective_models() { + local rec id out status docs + for case in paid unavailable; do + if [ "$case" = paid ]; then + id=profile-opencode-default-paid-z7o + model=opencode-go/paid-candidate + else + id=profile-opencode-default-unavailable-z7p + model=opencode-go/space-bunny-free + fi + rec=$(make_spawn_case "profile-opencode-default-$case" opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" "{\"model\":\"$model\"}") + + if [ "$case" = paid ]; then + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/paid-candidate' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"paid-candidate":{"cost":{"input":0.15,"output":0.6}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 1 "$status" "a paid effective model must be refused even with no --model" + assert_contains "$out" "is not classified as free by models.dev" \ + "a paid effective model was not refused by pricing" + else + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 1 "$status" "an effective model absent from the live catalog must be refused" + assert_contains "$out" "is not available from 'opencode models'" \ + "an unavailable effective model was not refused by the catalog" + fi + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "a refused effective model ($case) published metadata" + [ ! -s "$LAUNCH_LOG" ] || fail "a refused effective model ($case) launched an agent" + done + pass "the resolved effective model clears the same availability and zero-cost gate as an explicit one" +} + +test_opencode_omitted_model_refuses_when_the_config_inventory_is_unusable() { + local rec id out status docs + for case in failing not-a-list; do + id="profile-opencode-inventory-$case-z7q" + rec=$(make_spawn_case "profile-opencode-inventory-$case" opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/space-bunny-free"}') + + # Each branch captures the spawn's own status before anything else runs, so + # `$?` is the spawn and not the assignment that follows it. + if [ "$case" = failing ]; then + out=$(FM_TEST_OPENCODE_DEBUG_STATUS=1 FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expected="could not verify the effective OpenCode model" + else + out=$(FM_TEST_OPENCODE_DEBUG_NOT_A_LIST=1 FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expected="could not verify the effective OpenCode model" + fi + expect_code 1 "$status" "an unusable config inventory ($case) must refuse the spawn" + assert_contains "$out" "$expected" \ + "an unusable config inventory ($case) did not refuse on its own cause" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "an unusable config inventory ($case) published metadata" + [ ! -s "$LAUNCH_LOG" ] || fail "an unusable config inventory ($case) launched an agent" + done + pass "OpenCode dispatch refuses when debug config fails or has an unsupported shape" +} + +test_opencode_raw_launch_cannot_bypass_a_different_model() { + local rec id out status docs + id=profile-opencode-raw-launch-z7r + rec=$(make_spawn_case profile-opencode-raw-launch opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/space-bunny-free"}') + + # The positional command selects a paid model while the separate --model is + # nonempty, available, and zero cost. The dispatch must refuse the raw launch + # after validating that separate value instead of allowing the mismatch. + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/paid-candidate' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}},"paid-candidate":{"cost":{"input":0.15,"output":0.6}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "opencode --model opencode-go/paid-candidate --prompt hello" \ + --model opencode-go/space-bunny-free --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "a raw OpenCode launch with a different model must refuse" + assert_contains "$out" "raw OpenCode launch cannot be pinned" \ + "raw-launch mismatch refusal did not name the pinning failure" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "a refused raw OpenCode launch published metadata" + [ ! -s "$LAUNCH_LOG" ] || fail "a refused raw OpenCode launch launched an agent" + pass "a raw OpenCode launch cannot bypass validation with a different explicit model" +} + +test_opencode_raw_launch_classification_reads_the_executable_not_its_arguments() { + local rec id out status launch raw wrap_index + # A raw command is classified from the executable it actually runs, never + # from arbitrary argument text: `--prompt 'review opencode'` mentions OpenCode + # inside an unrelated Claude launch and must stay accepted, while an OpenCode + # command hidden behind the supported `env` wrapper prefix must still be + # refused. Matching the whole command string got both of these backwards. + id=profile-opencode-raw-exec-z7q + rec=$(make_spawn_case profile-opencode-raw-exec claude "$id") + read_case_record "$rec" + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "claude --prompt 'review opencode'" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "a Claude launch whose prompt merely mentions opencode must not be classified as OpenCode: $out" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "claude --prompt 'review opencode'" \ + "the unrelated launch did not deliver its own command" + + # Each form gets its own case directory and task id: the fixture builds a + # worktree and a remote clone, so reusing one name across iterations fails on + # the second pass for reasons unrelated to the behavior under test. + wrap_index=0 + for raw in "env -u FM_PANE opencode --prompt hi" "/usr/bin/env -i opencode --prompt hi"; do + wrap_index=$((wrap_index + 1)) + id="profile-opencode-raw-wrapped-z7r$wrap_index" + rec=$(make_spawn_case "profile-opencode-raw-wrapped-z7r$wrap_index" opencode "$id") + read_case_record "$rec" + # An explicit, available, zero-cost model so the dispatch reaches the + # raw-launch pinning refusal: reaching THAT refusal proves the wrapped + # command was classified as OpenCode and validated, not skipped. + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}}}}}' \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "$raw" --model opencode-go/space-bunny-free --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "an OpenCode command behind an env wrapper must refuse, not skip validation: $out ($raw)" + assert_contains "$out" "raw OpenCode launch cannot be pinned" \ + "wrapped raw OpenCode refusal did not name the pinning failure: $raw" + [ ! -s "$LAUNCH_LOG" ] || fail "a refused wrapped raw OpenCode launch launched an agent: $raw" + done + + # A QUOTED executable is the same bypass by another route: this walk splits on + # whitespace, so `'opencode'` keeps its quotes, basename returns `'opencode'`, + # and the command would match no harness and run OpenCode anyway with every + # model guard skipped. Unclassifiable, so it fails closed. Reached with an + # explicit, available, zero-cost model: were the quoted form still classified + # as OpenCode, this launch would instead reach the pinning refusal. + quoted_index=0 + for raw in "'opencode' --prompt hi" "\"\$HOME/bin/opencode\" --prompt hi"; do + quoted_index=$((quoted_index + 1)) + id="profile-opencode-raw-quoted-z7t$quoted_index" + rec=$(make_spawn_case "profile-opencode-raw-quoted-z7t$quoted_index" opencode "$id") + read_case_record "$rec" + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}}}}}' \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "$raw" --model opencode-go/space-bunny-free --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "a quoted OpenCode executable must fail closed, not skip validation: $out ($raw)" + assert_contains "$out" "quotes or expands its executable" \ + "the quoted-executable refusal did not name the unresolvable executable: $raw" + [ ! -s "$LAUNCH_LOG" ] || fail "a refused quoted raw OpenCode launch launched an agent: $raw" + done + + # A wrapper whose whole job is to run ANOTHER command is the same bypass: with + # `command`/`exec`/`nohup` unrecognized, HARNESS became the wrapper, so every + # `if [ "$HARNESS" = opencode ]` guard was skipped while OpenCode still ran. + # Each is walked through to the executable it would run, so reaching the + # PINNING refusal (not a generic refusal) proves the wrapper resolved and the + # model was validated on the way. + wrap2_index=0 + for raw in "command opencode --prompt hi" "exec opencode --prompt hi" \ + "nohup opencode --prompt hi" "nice opencode --prompt hi" "timeout 30 opencode --prompt hi"; do + wrap2_index=$((wrap2_index + 1)) + id="profile-opencode-raw-cmdwrap-z7u$wrap2_index" + rec=$(make_spawn_case "profile-opencode-raw-cmdwrap-z7u$wrap2_index" opencode "$id") + read_case_record "$rec" + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}}}}}' \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "$raw" --model opencode-go/space-bunny-free --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "an OpenCode command behind a command wrapper must refuse, not skip validation: $out ($raw)" + assert_contains "$out" "raw OpenCode launch cannot be pinned" \ + "command-wrapped OpenCode refusal did not name the pinning failure: $raw" + [ ! -s "$LAUNCH_LOG" ] || fail "a refused command-wrapped raw OpenCode launch launched an agent: $raw" + done + + # A command that runs OpenCode somewhere other than the one executable this walk + # resolves cannot have its model validated by that resolution, so it fails + # closed. The pipe case is the important one: its executable word (`echo`) is + # perfectly resolvable and the OpenCode launch is on the far side of it. A + # shell interpreter is the same bypass with no special character at all, since + # `sh -c opencode ...` names a real executable and then runs a different + # command from its arguments. + # + # This must NOT refuse compound commands in general: `cd && ./probe` is a + # legitimate raw launch that mentions no OpenCode, and the companion suite + # tests/fm-spawn-compact-adviser-disable.test.sh covers exactly that form. + compose_index=0 + # shellcheck disable=SC2016 # single quotes are deliberate: these are the literal raw-command text under test, not expansions of this shell + for raw in '$(which opencode) --prompt hi' '`which opencode` --prompt hi' \ + 'echo x | opencode --prompt hi' 'sh -c opencode' 'bash -c "opencode --prompt hi"' \ + 'zsh -c opencode' 'env sh -c opencode' 'cd /tmp && ./opencode' \ + 'true; opencode --prompt hi'; do + compose_index=$((compose_index + 1)) + id="profile-opencode-raw-compose-z7v$compose_index" + rec=$(make_spawn_case "profile-opencode-raw-compose-z7v$compose_index" opencode "$id") + read_case_record "$rec" + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}}}}}' \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "$raw" --model opencode-go/space-bunny-free --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "a composed or expanded raw executable must fail closed: $out ($raw)" + assert_contains "$out" "builds or composes its executable" \ + "the composed-executable refusal did not name the unresolvable executable: $raw" + [ ! -s "$LAUNCH_LOG" ] || fail "a refused composed raw launch launched an agent: $raw" + done + + id=profile-opencode-raw-ambiguous-z7s + rec=$(make_spawn_case profile-opencode-raw-ambiguous opencode "$id") + read_case_record "$rec" + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + "env --not-a-real-env-flag opencode --prompt hi" --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "a wrapped raw command whose executable cannot be resolved must fail closed: $out" + assert_contains "$out" "wraps its executable in a form dispatch cannot resolve" \ + "the ambiguous-wrapper refusal did not name the unresolved executable" + [ ! -s "$LAUNCH_LOG" ] || fail "an ambiguous wrapped raw command launched an agent" + pass "raw OpenCode classification follows the executable through the env wrapper, never prompt text" +} + +test_opencode_validates_the_selected_variant_not_just_the_base_model() { + local rec id out status docs base='opencodex/anthropic/claude-fable-5' + # `opencode models` lists only base ids, so it cannot confirm a variant. An + # available base must NOT carry an arbitrary variant: the exact + # "base#variant" has to be one the destination config actually declares. + id=profile-opencode-variant-bad-z7t + rec=$(make_spawn_case profile-opencode-variant opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" \ + '{"model":"opencodex/anthropic/claude-fable-5", + "providers":{"opencodex":{"models":{"anthropic/claude-fable-5": + {"variants":[{"id":"low"},{"id":"high"}]}}}}}') + + out=$(FM_TEST_OPENCODE_MODELS='opencodex/anthropic/claude-fable-5' \ + FM_TEST_MODELS_DEV_JSON='{"opencodex":{"models":{"anthropic/claude-fable-5":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model 'opencodex/anthropic/claude-fable-5#nonexistent' --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "an available base with a nonexistent variant must refuse: $out" + assert_contains "$out" "no config source there selects the '#nonexistent' variant" \ + "the nonexistent-variant refusal did not name the undeclared variant" + [ ! -s "$LAUNCH_LOG" ] || fail "a launch with an undeclared variant started an agent" + + id=profile-opencode-variant-ok-z7u + rec=$(make_spawn_case profile-opencode-variant-ok opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" \ + '{"model":"opencodex/anthropic/claude-fable-5", + "providers":{"opencodex":{"models":{"anthropic/claude-fable-5": + {"variants":[{"id":"low"},{"id":"high"}]}}}}}') + out=$(FM_TEST_OPENCODE_MODELS='opencodex/anthropic/claude-fable-5' \ + FM_TEST_MODELS_DEV_JSON='{"opencodex":{"models":{"anthropic/claude-fable-5":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model 'opencodex/anthropic/claude-fable-5#low') + status=$? + expect_code 0 "$status" "a variant the destination config declares must be accepted: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode 'opencodex/anthropic/claude-fable-5#low' default + assert_contains "$(cat "$LAUNCH_LOG")" "--model 'opencodex/anthropic/claude-fable-5#low'" \ + "the declared variant was not pinned on the launch" + + # Precedence, not union: a LOWER-precedence config declares `low` and enables + # it, while the worker's OWN higher-precedence config declares the same model + # and disables `low`. Only the effective (highest-precedence) declaration + # decides, so `#low` must be refused even though some config declares it. + id=profile-opencode-variant-precedence-z7v + rec=$(make_spawn_case profile-opencode-variant-precedence opencode "$id") + read_case_record "$rec" + docs=$(jq -cn --arg far "$CASE_DIR/opencode.json" --arg near "$WT_DIR/opencode.json" \ + --arg base 'opencodex/anthropic/claude-fable-5' \ + '[{"type":"document","path":$far,"info":{"model":$base, + "providers":{"opencodex":{"models":{"anthropic/claude-fable-5": + {"variants":[{"id":"low"},{"id":"high"}]}}}}}}, + {"type":"document","path":$near,"info":{"model":$base, + "providers":{"opencodex":{"models":{"anthropic/claude-fable-5": + {"variants":[{"id":"low","disabled":true},{"id":"high"}]}}}}}}]') + out=$(FM_TEST_OPENCODE_MODELS='opencodex/anthropic/claude-fable-5' \ + FM_TEST_MODELS_DEV_JSON='{"opencodex":{"models":{"anthropic/claude-fable-5":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model "$base#low" --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "a variant disabled by the higher-precedence config must refuse: $out" + assert_contains "$out" "no config source there selects the '#low' variant" \ + "the disabled-by-precedence refusal did not name the variant" + [ ! -s "$LAUNCH_LOG" ] || fail "a variant disabled by the higher-precedence config launched an agent" + + # Two EQUAL-precedence sources that DISAGREE about one model are as + # unresolvable as two disagreeing root models. The disagreeing state that is + # easiest to miss is the one where both sources DECLARE the same variant id and + # differ only in whether it is enabled: comparing declared sets alone passes, + # while the effective state is genuinely ambiguous, so that case must refuse + # too rather than select the enabled one. + conflict_index=0 + for conflict in differing-same-declared differing-sets; do + conflict_index=$((conflict_index + 1)) + id="profile-opencode-variant-conflict-z7w$conflict_index" + rec=$(make_spawn_case "profile-opencode-variant-conflict$conflict_index" opencode "$id") + read_case_record "$rec" + if [ "$conflict" = differing-same-declared ]; then + # Both declare `low`; the worker's own config disables it. + near_variants='{"variants":[{"id":"low","disabled":true}]}' + jsonc_variants='{"variants":[{"id":"low"}]}' + requested='low' + else + near_variants='{"variants":[{"id":"low"}]}' + jsonc_variants='{"variants":[{"id":"high"}]}' + requested='low' + fi + docs=$(jq -cn --arg near "$WT_DIR/opencode.json" --arg jsonc "$WT_DIR/opencode.jsonc" \ + --arg base "$base" --argjson nv "$near_variants" --argjson jv "$jsonc_variants" \ + '[{"type":"document","path":$near,"info":{"model":$base, + "providers":{"opencodex":{"models":{"anthropic/claude-fable-5":$nv}}}}}, + {"type":"document","path":$jsonc,"info":{"model":$base, + "providers":{"opencodex":{"models":{"anthropic/claude-fable-5":$jv}}}}}]') + out=$(FM_TEST_OPENCODE_MODELS='opencodex/anthropic/claude-fable-5' \ + FM_TEST_MODELS_DEV_JSON='{"opencodex":{"models":{"anthropic/claude-fable-5":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model "$base#$requested" --mode no-mistakes --yolo off) + status=$? + expect_code 1 "$status" "equal-precedence variant disagreement must refuse rather than pick one ($conflict)" + assert_contains "$out" "declare different variants" \ + "the ambiguous-variant refusal did not explain the disagreement ($conflict)" + [ ! -s "$LAUNCH_LOG" ] || fail "an ambiguous variant selection launched an agent ($conflict)" + done + pass "OpenCode validates the exact selected variant against the effective, precedence-resolved variants" +} + +test_opencode_v1_resolved_debug_config_object_is_supported() { + local rec id out status launch resolved + id=profile-opencode-v1-debug-object-z7u + rec=$(make_spawn_case profile-opencode-v1-debug-object opencode "$id") + read_case_record "$rec" + resolved='{"model":{"providerID":"opencode-go","model":"longcat-2.5-preview-free"},"default_agent":"review","agents":{"review":{"mode":"primary","model":{"providerID":"opencode-go","model":"paid-candidate"}}}}' + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/longcat-2.5-preview-free\nopencode-go/paid-candidate' \ + FM_TEST_OPENCODE_DEBUG_OBJECT="$resolved" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "v1 resolved-object debug output should resolve and validate its root model: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/longcat-2.5-preview-free default + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "--model 'opencode-go/longcat-2.5-preview-free'" \ + "v1 debug output did not pin its root model on the worker launch" + pass "v1 resolved-object debug config output is supported and pins its root model" +} + +test_opencode_v1_map_form_variants_are_read_and_disabled_ones_excluded() { + local rec id out status resolved + # `variants` is published in two shapes: v2 uses an array of objects carrying + # `id`, v1 uses an object keyed by the variant id. Which one appears is a + # property of the OpenCode version, so a launch must not depend on which one + # this machine happens to run. A variant explicitly disabled is not + # selectable, so it must never be the evidence that admits a "#variant" form. + id=profile-opencode-v1-map-variants-z7v + rec=$(make_spawn_case profile-opencode-v1-map-variants opencode "$id") + read_case_record "$rec" + resolved='{"model":"opencode-go/space-bunny-free","providers":{"opencode-go":{"models":{"space-bunny-free":{"variants":{"low":{"options":{}},"disabled-one":{"disabled":true},"high":null}}}}}}' + + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_OBJECT="$resolved" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model 'opencode-go/space-bunny-free#low') + status=$? + expect_code 0 "$status" "a variant declared in the v1 map form must be accepted: $out" + assert_contains "$(cat "$LAUNCH_LOG")" "--model 'opencode-go/space-bunny-free#low'" \ + "the accepted v1 map-form variant was not passed to the worker" + + # A separate case, because the launch log appends and the first half of this + # test deliberately launched a worker. + id=profile-opencode-v1-map-variant-disabled-z7w + rec=$(make_spawn_case profile-opencode-v1-map-variant-disabled opencode "$id") + read_case_record "$rec" + + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_MODELS_DEV_JSON='{"opencode-go":{"models":{"space-bunny-free":{"cost":{"input":0,"output":0}}}}}' \ + FM_TEST_OPENCODE_DEBUG_OBJECT="$resolved" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model 'opencode-go/space-bunny-free#disabled-one') + status=$? + expect_code 1 "$status" "a variant the config disables must refuse even though its base model is valid" + assert_contains "$out" "is not available in the destination pane" \ + "a disabled variant refusal did not name the missing variant" + [ ! -s "$LAUNCH_LOG" ] || fail "a disabled variant launched OpenCode" + pass "v1 map-form variants are read, and an explicitly disabled variant is not accepted as evidence" +} + +test_opencode_custom_config_environment_overrides_fail_closed() { + local rec id out status name + for name in OPENCODE_CONFIG OPENCODE_CONFIG_DIR; do + id="profile-opencode-config-env-${name}" + rec=$(make_spawn_case "profile-opencode-config-env-${name}" opencode "$id") + read_case_record "$rec" + printf '%s\n' OPENCODE_CONFIG OPENCODE_CONFIG_DIR > "$HOME_DIR/config/launch-env-allowlist" + if [ "$name" = OPENCODE_CONFIG ]; then + out=$(FM_TEST_OPENCODE_CONFIG="$CASE_DIR/custom-opencode.json" \ + FM_TEST_OPENCODE_DEBUG_DOCS="$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/space-bunny-free"}')" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + else + out=$(FM_TEST_OPENCODE_CONFIG_DIR="$CASE_DIR/custom-config-dir" \ + FM_TEST_OPENCODE_DEBUG_DOCS="$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/space-bunny-free"}')" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + fi + status=$? + expect_code 1 "$status" "$name must fail closed when effective config sources cannot be ranked" + assert_contains "$out" "cannot resolve that custom config location safely" \ + "$name refusal did not name the unsupported config override" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "$name override published task metadata" + [ ! -s "$LAUNCH_LOG" ] || fail "$name override launched OpenCode" + if find "$WT_DIR" -maxdepth 1 -name ".fm-opencode-probe-$id-*" -print -quit | grep -q .; then + fail "$name refusal left a completed OpenCode probe directory behind" + fi + done + pass "implicit OpenCode model resolution fails closed for OPENCODE_CONFIG and OPENCODE_CONFIG_DIR overrides" +} + +test_opencode_probe_uses_launch_environment_and_bounded_debug_calls() { + local rec id out status docs context timeout_args launch + id=profile-opencode-pane-scope-z7v + rec=$(make_spawn_case profile-opencode-pane-scope opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/space-bunny-free"}') + context="$CASE_DIR/opencode-context.tsv" + timeout_args="$CASE_DIR/timeout-args.log" + cat > "$HOME_DIR/config/launch-env-allowlist" <<'EOF' +FM_FAKE_OPENCODE_MODELS +FM_FAKE_OPENCODE_DEBUG_DOCS +FM_FAKE_OPENCODE_CONTEXT_LOG +FM_FAKE_TIMEOUT_ARGS +OPENCODE_SERVER +EOF + + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + FM_TEST_OPENCODE_CONTEXT_LOG="$context" \ + FM_TEST_TIMEOUT_ARGS="$timeout_args" \ + FM_TEST_OPENCODE_SERVER='https://worker-context.invalid' \ + FM_TEST_SHOULD_NOT_LEAK='must-not-reach-worker' \ + FM_OPENCODE_MODELS_TIMEOUT=3 \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "OpenCode probes should run in the launch allowlist context: $out" + [ "$(wc -l < "$context" | tr -d ' ')" = 2 ] \ + || fail "expected one context record for debug config and one for models" + while IFS="$(printf '\t')" read -r command cwd server leak; do + [ "$cwd" = "$WT_DIR" ] || fail "probe command $command ran in $cwd instead of $WT_DIR" + [ "$server" = 'https://worker-context.invalid' ] || fail "probe command $command lost the worker server route" + [ "$leak" = unset ] || fail "an unallowlisted supervisor variable reached probe command $command" + done < "$context" + [ "$(grep -c '^-k 1 3 bash -c ' "$timeout_args")" -ge 1 ] \ + || fail "bounded debug config did not consume timeout's -k 1 and 3-second arguments" + [ "$(grep -c '^-k 1 3 ' "$timeout_args")" -ge 2 ] \ + || fail "debug config and model catalog were not both bounded" + + launch=$(cat "$LAUNCH_LOG") + # shellcheck disable=SC2016 # single quotes are deliberate: this asserts the literal launch text carries the pane-side expansion, not this shell's + assert_contains "$launch" 'OPENCODE_SERVER=$OPENCODE_SERVER' \ + "the launch did not carry the same allowlisted server context as its probe" + pass "OpenCode debug and catalog probes use the worker cwd, server route, allowlist, and finite timeout" +} + +test_opencode_probe_enforces_the_catalog_deadline() { + local rec id out status docs started elapsed + id=profile-opencode-probe-timeout-z7x + rec=$(make_spawn_case profile-opencode-probe-timeout opencode "$id") + read_case_record "$rec" + docs=$(opencode_config_docs "$WT_DIR/opencode.json" '{"model":"opencode-go/space-bunny-free"}') + # The fake `timeout` in this suite only records the arguments it is given and + # then execs, so it cannot enforce a deadline and an assertion against it would + # prove nothing about enforcement. This case opts out of that stub and uses the + # real bounded runner instead (FM_TIMEOUT_MECHANISM_OVERRIDE=bash selects the + # dependency-free fallback in bin/fm-timeout-lib.sh), against a catalog that + # never answers. The deadline is then real: the spawn must be refused as a + # timeout quickly instead of waiting on the hung command. + cat > "$HOME_DIR/config/launch-env-allowlist" <<'EOF' +FM_TIMEOUT_MECHANISM_OVERRIDE +FM_FAKE_OPENCODE_MODELS +FM_FAKE_OPENCODE_MODELS_HANG +FM_FAKE_OPENCODE_DEBUG_DOCS +EOF + started=$(date +%s) + out=$(FM_TEST_OPENCODE_MODELS_HANG=1 FM_TIMEOUT_MECHANISM_OVERRIDE=bash FM_OPENCODE_MODELS_TIMEOUT=2 \ + FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + elapsed=$(($(date +%s) - started)) + expect_code 1 "$status" "a model catalog that never answers must be refused, not waited on: $out" + assert_contains "$out" "'opencode models' failed in the destination pane (exit 124)" \ + "the bounded catalog refusal did not report the timeout exit status" + [ "$elapsed" -lt 60 ] \ + || fail "the bounded catalog probe waited ${elapsed}s, so its deadline was not enforced" + [ ! -s "$LAUNCH_LOG" ] || fail "a hung catalog probe still launched an agent" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "a hung catalog probe published task metadata" + pass "a hung OpenCode catalog probe is cut off at its bound and refused as a timeout" +} + +test_opencode_explicit_model_is_validated_in_the_worker_worktree() { + local rec id out status cwd_file scoped + id=profile-opencode-explicit-in-wt-z7s + rec=$(make_spawn_case profile-opencode-explicit-in-wt opencode "$id") + read_case_record "$rec" + cwd_file="$CASE_DIR/models-cwd" + # A model a worktree-scoped provider makes available. It is absent from the + # catalog the launcher would see, so validating anywhere but the settled + # worktree refuses a model the worker can actually run. + scoped=$(jq -cn --arg wt "$WT_DIR" '{($wt): "projectscope/local-free"}') + + out=$(FM_TEST_OPENCODE_MODELS='opencode-go/space-bunny-free' \ + FM_TEST_OPENCODE_MODELS_IN_DIR="$scoped" FM_TEST_OPENCODE_MODELS_CWD="$cwd_file" \ + FM_TEST_MODELS_DEV_JSON='{"projectscope":{"models":{"local-free":{"cost":{"input":0,"output":0}}}}}' \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" \ + --harness opencode --model projectscope/local-free) + status=$? + expect_code 0 "$status" "an explicit model a project-scoped provider supplies must validate: $out" + [ "$(cat "$cwd_file" 2>/dev/null)" = "$WT_DIR" ] \ + || fail "an explicit OpenCode model was validated outside the worker worktree: $(cat "$cwd_file" 2>/dev/null)" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode projectscope/local-free default + pass "an explicit OpenCode model is validated against the catalog read in the worker worktree" +} + +test_opencode_effective_model_resolves_without_nonstock_tools() { + local rec id out status docs + id=profile-opencode-no-tac-z7t + rec=$(make_spawn_case profile-opencode-no-tac opencode "$id") + read_case_record "$rec" + # Two direct configs make ancestor ordering load-bearing for the answer. The + # fakebin's `tac` stub + # exits 127, so any ordering that shells out to tac (GNU coreutils, not a + # stock macOS command) fails here instead of on the captain's machine. + docs=$(jq -cn --arg far "$CASE_DIR/opencode.json" --arg near "$WT_DIR/opencode.json" \ + '[{"type":"document","path":$far,"info":{"model":"opencode-go/space-bunny-free"}}, + {"type":"document","path":$near,"info":{"model":"opencode-go/longcat-2.5-preview-free"}}]') + + out=$(FM_TEST_OPENCODE_MODELS=$'opencode-go/space-bunny-free\nopencode-go/longcat-2.5-preview-free' \ + FM_TEST_OPENCODE_DEBUG_DOCS="$docs" \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness opencode) + status=$? + expect_code 0 "$status" "ancestor ordering must not depend on a non-stock command: $out" + assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/longcat-2.5-preview-free default + pass "effective-model ancestor ordering uses no command stock macOS lacks" +} + test_opencode_secondmate_config_model_uses_live_catalog() { local rec id sm out status launch id=profile-opencode-secondmate-config-z7c @@ -860,7 +1668,7 @@ test_opencode_secondmate_config_model_uses_live_catalog() { expect_code 0 "$status" "configured OpenCode secondmate model should pass when listed" assert_meta_profile "$HOME_DIR/state/$id.meta" opencode opencode-go/space-bunny-free default launch=$(cat "$LAUNCH_LOG") - assert_contains "$launch" "opencode --model 'opencode-go/space-bunny-free' --prompt" \ + assert_contains "$launch" "'$FAKEBIN_DIR/opencode' --model 'opencode-go/space-bunny-free' --prompt" \ "configured OpenCode secondmate model was not launched" pass "configured OpenCode secondmate models are checked and launched" } @@ -1721,6 +2529,25 @@ test_opencode_refuses_unreadable_live_catalog test_opencode_refuses_paid_catalog_model test_opencode_refuses_when_free_pricing_metadata_is_unavailable test_opencode_catalog_probe_uses_no_provider_argument +test_opencode_omitted_model_resolves_validated_project_model +test_opencode_default_model_token_resolves_like_an_omitted_one +test_opencode_omitted_model_honors_documented_config_precedence +test_opencode_primary_agent_model_does_not_override_the_session_model +test_opencode_omitted_model_refuses_when_no_effective_model_exists +test_opencode_omitted_model_refuses_when_config_sources_are_ambiguous +test_opencode_custom_primary_agent_without_a_model_inherits_the_session_model +test_opencode_omitted_model_refuses_paid_and_unavailable_effective_models +test_opencode_omitted_model_refuses_when_the_config_inventory_is_unusable +test_opencode_raw_launch_cannot_bypass_a_different_model +test_opencode_raw_launch_classification_reads_the_executable_not_its_arguments +test_opencode_validates_the_selected_variant_not_just_the_base_model +test_opencode_v1_resolved_debug_config_object_is_supported +test_opencode_v1_map_form_variants_are_read_and_disabled_ones_excluded +test_opencode_custom_config_environment_overrides_fail_closed +test_opencode_probe_uses_launch_environment_and_bounded_debug_calls +test_opencode_probe_enforces_the_catalog_deadline +test_opencode_explicit_model_is_validated_in_the_worker_worktree +test_opencode_effective_model_resolves_without_nonstock_tools test_opencode_secondmate_config_model_uses_live_catalog test_opencode_secondmate_config_refuses_model_absent_from_live_catalog test_native_effort_validator_keeps_axes_separate From 87dd889beaeffffc832d05381189012cc8cda640 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EB=B0=95=ED=8F=89=EC=8B=9D?= Date: Thu, 1 Oct 2026 20:28:23 +0900 Subject: [PATCH 2/3] no-mistakes(test): Waiting on background test run before finalizing fix summary --- bin/fm-opencode-probe.sh | 16 +++++++++++----- 1 file changed, 11 insertions(+), 5 deletions(-) diff --git a/bin/fm-opencode-probe.sh b/bin/fm-opencode-probe.sh index 70c1effb8ef..55d7f75baa1 100755 --- a/bin/fm-opencode-probe.sh +++ b/bin/fm-opencode-probe.sh @@ -105,19 +105,25 @@ if [ "$resolve_config" = 1 ]; then global:$global, docs:[$docs[] | {type, path, info:(.info | clean_info), variants:(.info | variant_map)}]} ' + # `debug config` output is redirected to a real file rather than piped + # straight into jq: against the installed OpenCode CLI, piping its stdout + # silently truncates at 64KiB, which breaks jq on any config past that + # size. A regular file does not have that limit. + raw_config="$result_dir/config.raw.json" # shellcheck disable=SC2016 # single quotes are deliberate: ${...} expands in the probe's bash -c, not here if fm_run_timed "$bound" /bin/bash -c ' - "$1" debug config 2>/dev/null | jq -c --arg global "$2" "$3" - pipeline=("${PIPESTATUS[@]}") - [ "${pipeline[0]}" -eq 0 ] || exit "${pipeline[0]}" - exit "${pipeline[1]}" - ' _ "$bin" "$global_config" "$normalize_filter" > "$result_dir/config.json"; then + "$1" debug config > "$2" 2>/dev/null + status=$? + [ "$status" -eq 0 ] || exit "$status" + jq -c --arg global "$3" "$4" < "$2" + ' _ "$bin" "$raw_config" "$global_config" "$normalize_filter" > "$result_dir/config.json"; then printf '0\n' > "$result_dir/config.status" else status=$? [ "$status" -ne 0 ] || status=65 printf '%s\n' "$status" > "$result_dir/config.status" fi + rm -f "$raw_config" 2>/dev/null || true fi fi else From 094a0088dcf255fde840f786e827f42d2c490233 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EB=B0=95=ED=8F=89=EC=8B=9D?= Date: Thu, 1 Oct 2026 20:40:56 +0900 Subject: [PATCH 3/3] no-mistakes(document): docs: update OpenCode model validation to destination-pane config resolution --- .../skills/harness-adapters/references/harness/opencode.md | 2 +- docs/configuration.md | 6 +++--- 2 files changed, 4 insertions(+), 4 deletions(-) diff --git a/.agents/skills/harness-adapters/references/harness/opencode.md b/.agents/skills/harness-adapters/references/harness/opencode.md index 4373fefc561..51358a35ebc 100644 --- a/.agents/skills/harness-adapters/references/harness/opencode.md +++ b/.agents/skills/harness-adapters/references/harness/opencode.md @@ -14,7 +14,7 @@ The worker busy-state lifecycle below describes the generated OpenCode 2 plugin. | Resume | Relaunch with `--continue` to resume the most recent session for the current directory, then send the next instruction after the TUI is ready because `--prompt` does not auto-submit alongside `--continue`. | | Interactive task launch | Firstmate probes `opencode mini --help` for a `Usage: opencode mini` line; OpenCode v2 uses `opencode mini`, while legacy releases use the top-level command. Both receive the requested `--model ` and worker brief through `--prompt`. | | Effort flag | None for the interactive task launch; `opencode run` has `--variant`, but that is a different, non-interactive path. | -| Model discovery | Run `opencode models` with no argument to list every available `provider/model` identifier. The legacy `opencode models ` form is rejected by OpenCode v2 ("Unexpected positional argument") and prints usage text rather than a catalog, so `../../../bin/fm-spawn.sh` reads the whole catalog and matches the exact id. | +| Model discovery | Run `opencode models` with no argument to list every available `provider/model` identifier. The legacy `opencode models ` form is rejected by OpenCode v2 ("Unexpected positional argument") and prints usage text rather than a catalog, so `../../../bin/fm-spawn.sh` reads the whole catalog and matches the exact base id. An omitted or `default` model, and any `#variant` suffix, instead resolve through `opencode debug config` read in the destination pane's settled task worktree, precedence-ranked the same way OpenCode itself ranks config sources; see `opencode_resolve_effective_model` and `opencode_model_validate` in `../../../bin/fm-spawn.sh`. | | Trust dialog | None. | | Marker | None; OpenCode publishes no identity marker, so `../../../bin/fm-harness.sh` identifies it from process ancestry. | diff --git a/docs/configuration.md b/docs/configuration.md index f75dd051f6b..078eb7fa63a 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -499,9 +499,9 @@ Every profile array is an implicit quota-aware choice resolved through `quota-ar If no dispatch rule fits, firstmate resolves `default` through the same object-or-array path before falling back to `config/crew-harness`. Except for `ultra`, which refuses unsupported profiles under the native-effort contract above, an effort value the chosen harness does not accept is recorded as `effort=` in task meta for traceability but omitted from the launch flags. Bootstrap reports unsupported harness/model/effort combinations as a `CREW_DISPATCH` diagnostic when they are visible in the file. -See [`docs/examples/crew-dispatch.json`](examples/crew-dispatch.json) for a starting point to copy into local `config/crew-dispatch.json`; its default pins `opencode-go/space-bunny-free`, declares the `opencode-go` provider, and is rechecked against `opencode models` at spawn time. -An OpenCode spawn with a resolved model fails before launch if `opencode models` cannot be read or does not list that exact ID; refresh the catalog with `opencode models` and choose a listed ID. -An explicitly designated per-task model overrides the configured profile model and is checked against that live catalog before launch. +See [`docs/examples/crew-dispatch.json`](examples/crew-dispatch.json) for a starting point to copy into local `config/crew-dispatch.json`; its default pins `opencode-go/space-bunny-free`, declares the `opencode-go` provider, and is rechecked against `opencode models` and `opencode debug config` at spawn time, inside the destination pane's settled task worktree rather than the launcher's own directory. +An OpenCode spawn fails before launch if that pane's `opencode models` cannot be read or does not list the resolved model's base ID, if a `#variant` suffix is not selected by that pane's resolved config, or if an omitted or `default` model resolves to no root model at all; refresh the catalog with `opencode models` and choose a listed ID, or set a root model in the destination config. +An explicitly designated per-task model overrides the configured profile model and is checked the same way before launch. OpenCode dispatch also requires every pricing field in the model metadata to be zero; the [OpenCode harness reference](../.agents/skills/harness-adapters/references/harness/opencode.md#free-model-selection) documents data-retention limits. When the file exists, bootstrap validates it with `jq`. Valid files stay silent by default; with `FM_BOOTSTRAP_VERBOSE_FACTS=1`, bootstrap emits `BOOTSTRAP_INFO: crew dispatch active config/crew-dispatch.json`, one `BOOTSTRAP_INFO:` fact per rule, and one fact for the optional default profile set.