Skip to content

feat(models): refresh Command Code catalog to 1.79.2 with claude-haiku-5-5 - #145

Open
sandexzx wants to merge 4 commits into
patlux:mainfrom
sandexzx:feat/refresh-command-code-1-79-2
Open

sandexzx wants to merge 4 commits into
patlux:mainfrom
sandexzx:feat/refresh-command-code-1-79-2

Conversation

@sandexzx

@sandexzx sandexzx commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Adds claude-haiku-5-5 from the latest Command Code catalog and clears the drift that had accumulated since the pinned command-code@1.72.4 snapshot.

What changed:

  • Synced the static capability snapshot to command-code@1.79.2: new claude-haiku-5-5, mistral/mistral-large-4 (image input, 262K output limit), and free stealth/glyph-cluster:free; retired stealth/pixel-canary and stealth/space-bunny-alpha.
  • Recognized the off reasoning effort now published for the DeepSeek V4 and V4.1 models. The daily catalog check was failing on it (Unexpected reasoning efforts for deepseek/deepseek-v4-pro: off, high, max) because VALID_EFFORTS and the generated type union did not know off. The generate transport still never forwards off as reasoning_effort; the Provider API and Oh My Pi adapters translate the level on their own wire.
  • Reviewed display pricing against the official page (verified 2026-10-09): claude-haiku-5-5 base plus its 100K long-context tier, mistral/mistral-large-4, and free stealth/glyph-cluster:free; dropped the retired stealth/space-bunny-alpha; corrected claude-sonnet-5-5 cache-read to $0.10/M. The snapshot now covers all 87 advertised models.
  • Refreshed both fixtures, updated the affected test expectations, and replaced a stale README note that claimed MODEL_EFFORT_OVERRIDES still added manual levels (the map is empty).

Validation: npm test (exit 0), npm run check:commandcode-catalog (no drift), npm run check:commandcode-pricing (PASS), npm run typecheck, npm run format:check, git diff --check.

@pierreraby

Copy link
Copy Markdown
Collaborator

Thanks for catching this and refreshing the catalog. I reviewed exact head a2a19ef in an isolated checkout; a separate read-only review also found no introduced defect.

Local validation passed:

  • Full npm test, including real pi 1.1.0 and OMP 18.4.2 against local mock APIs.
  • npm run format:check and git diff --check.
  • Read-only live catalog check: no drift against command-code@1.79.2.
  • Read-only live pricing check: PASS for all 87 advertised API models.

I also independently checked the changed prices against the official pricing page, including Haiku 5.5's >100K tier, Mistral Large 4, Glyph's free preview, and Sonnet 5.5's $0.10/M cache-read rate. A local offline probe accepts off, high, max, still rejects an unknown effort, and confirms the separate pi/OMP thinking metadata. The generate off filtering already existed before this PR, so this is principally a catalog/parser repair rather than a new generate-wire behavior.

One related CI hardening item would help prevent a repeat: the catalog steps use ... | tee ... with the default shell. In the October 8 scheduled run, the parser printed Unexpected reasoning efforts for deepseek/deepseek-v4-pro: off, high, max, but the synchronization step was marked successful, and no sync PR was created. The run failed later on the independent pricing check.

The log shows /usr/bin/bash -e {0}, without pipefail. I reproduced the exit-code behavior locally: false | tee /dev/null exits 0 under that shell, but exits 1 with -e -o pipefail.

Could you add explicit shell: bash to the catalog comparison and synchronization steps in .github/workflows/model-metadata.yml (or defaults.run.shell: bash for the workflow)? GitHub's explicit Bash shell enables pipefail. This is a pre-existing workflow issue, not a regression introduced by your diff; a small follow-up is also reasonable if you prefer to keep it separate.

Two optional coverage improvements, not merge blockers:

  • Build the off-omission stream test with thinkingMetadataForModel("deepseek/deepseek-v4-flash"), so it pins the canonical thinking.effortMap path as well as the older thinkingLevelMap path.
  • Add a small offline pricing-page fixture for Haiku 5.5's labeled ≤100K / >100K bands; the existing HTML fixture predates the refreshed catalog, although the live checker passes today.

The Compare and Synchronize steps piped '... | tee' without pipefail, so
tee's exit 0 masked a failing catalog check and the job stayed green. Run
them with shell: bash so GitHub invokes bash with -o pipefail and a failed
check fails the job. Also trigger the workflow when the new pricing fixture
changes.
Build the off-omission stream test from thinkingMetadataForModel so it pins
the canonical thinking.effortMap path instead of only the legacy level map,
and add a minimal Haiku 5.5 pricing fixture with its <=100K / >100K bands.
The new pricing test compares the fixture against the real MODEL_COSTS, so a
drift between the page bands and the runtime policy fails.
Set contextWindow to 1000000 and individual-goat to true so the snapshot row
mirrors the live props.rows entry verbatim. Neither field feeds the parser,
but a fixture should not carry known-wrong values.
@pierreraby

Copy link
Copy Markdown
Collaborator

Thanks for addressing the three follow-ups on new head 8645883. I re-verified in an isolated checkout.

The shell: bash fix on both catalog steps resolves the masked | tee exit I reported: explicit Bash enables pipefail, so a parser failure can no longer pass silently and block the auto-sync PR. The Haiku fixture mirrors the live page row and its test compares against the real runtime MODEL_COSTS, so drift would actually fail. The stream test now pins the canonical thinking.effortMap path for off.

Local validation on the new head passed:

  • Full npm test, including real pi 1.1.0 and OMP 18.4.2 against local mock APIs.
  • npm run format:check and git diff --check.
  • Read-only live catalog check: no drift against command-code@1.79.2.
  • Read-only live pricing check: PASS for all 87 advertised API models.

Approving — CI is green on this head and the scope matches the reviewed catalog refresh.

@pierreraby pierreraby left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved on exact head 8645883: catalog refresh to command-code@1.79.2, off recognition, and reviewed pricing verified locally (full npm test incl. pi 1.1.0 + OMP 18.4.2 against mocks, live catalog/pricing checks PASS, format/diff-check clean). Follow-up commits address the reported pipefail masking and add the effortMap + Haiku-band coverage.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants