Skip to content

Run-card oracle for the Mate engine, and the card defects it found - #163

Merged
fxck merged 21 commits into
mainfrom
engine/card-oracle
Oct 10, 2026
Merged

fxck merged 21 commits into
mainfrom
engine/card-oracle

Conversation

@fxck

@fxck fxck commented Oct 9, 2026

Copy link
Copy Markdown
Member

A run-card oracle for the Mate engine, and the defects it found between the engine's record and the run card.

The oracle (apps/web/test/engine-oracle)

Seeded trees play on the running engine (live layer, scripted Claude, SQLite, test clock): the Mate's turn starts helpers and background commands, helpers run calls and start their own commands and helpers, each piece of work ends its own way (finished, failed with an exit code, stopped by the Mate or the person, lost to a restart or a crash), an end after the turn wakes the Mate, whose wake starts more; the person sends, steers, stops and reloads in between. The record reaches the card through the client's own adapter and store over the engine's wire (frames and pages through the contract schemas), into the card's logic as ChatView feeds it. After every op the oracle holds the card against the tree (whose each line is and on which card, each piece's true end, the waiting line, counts, the helpers' surface, the dock); at a reload and at the end, a fresh client against the live one. Failures minimise to their smallest tree.

runCard.oracle.test.ts: 12 minimised defect trees + 24 fixed seeds, in chat-gate stage E; ENGINE_PROOF_RANDOM=n / ENGINE_PROOF_SEED=s for deep runs. About 2,100 random seeds were run during the lane.

Fixed (each with its test)

  • engine: background work records whose it is (a helper's own jobs and helpers), the call that started it and its report line; a helper's own call stays on the run that started its helper; work a Stop closed with Claude's session is stopped; a closed card's outcome read carries the calls that started its jobs.
  • client: the card reads owner, call and report from the record (the words heuristic is gone); the waiting line never counts helpers' commands as the Mate's.
  • web: a nested helper is never among the Mate's helpers; the helpers bubble marks a failure at once and says how its helpers ended ("2 stopped"); every call counts in the effort, Zerops calls too; a steered message is drawn once, above its card; a wake's answer appears under the reply the person was reading; "ran in the background"; "the home page" for "/"; a run that only wrote notes "thought" after a reload too; the person's own send/steer glides instead of a 104 px cut.
  • claude adapter: a turn Claude opens on its own starts at its stream's start, so its first command is its own (the "No result" of run 3).

Not covered

Two empty picture boxes in the final reply (run 3, +3:27.5), and the live +130 px send jump (the fake glides already).

Gates

Targeted suites, package typechecks, ci-local green. chat-gate E, C-engine, types green; C fails only composer-overlap (plain, review), which fails on origin/main too.

🤖 Generated with Claude Code

fxck added 21 commits October 9, 2026 22:07
…s Stop closes the CLI

ClaudeAdapter's stopSessionInternal reports every live task stopped (with its
linkage) before the session exits; the scripted provider dropped them, so a
person's Stop met work with no report of its own.
… it and how it ended

The run-card oracle (milo stress runs 2 and 3) found the record losing what the
card needs to tell a helper's work from the Mate's:
- a task a helper's tool started (Claude's owning agentId, Codex's
  parentAgentId) is recorded as that helper's and goes on under the run its
  helper served; a helper's own call goes there too, never onto the next
  message's run, which it made a card of its own after a reload;
- the work names the call that started it (its tool use), and a closed card's
  outcome read carries that call, so its line of jobs reads the same after a
  reload;
- its end keeps its report's first line ("failed with exit code 3");
- work a Stop closed with Claude's session is stopped, not lost: Claude
  reports its tasks only after the turn's end, from a session already closed.
… end from the record

Background work's activities carry the helper the record names (agentId), the
call that started it (toolUseId, so a job is its command's and a helper's
launch is its helper's row) and its report (detail: "Exit code 3"). The
guess by a helper call's words is gone: it held only while the helper's calls
were held, so the dock and the card drew a helper's job as the Mate's after a
reload, and a helper's launch read as a finished one-second step while it
worked (Milo's third stress run). The pipeline test arranges its record as
the engine now writes it.
…mands as the Mate's

Milo's third stress run waited on "its helpers and 2 background commands":
the commands were the nested helper's. A helper's own work counts as its
helpers' (the oracle's row: a card waiting on its helpers).
…tarted

Anything a helper started (agentId) is the helper's, its own helpers too: the
helpers' surface keeps it under its helper (spawnedBy). Milo's second stress
run merged a nested helper started 40 s later into "Started 4 helpers".
…elpers ended

The bubble's words and state are a logic of their own (helpersBubbleOf), as the
card and the oracle read them:
- a helper that failed marks the bubble while the others still work (Milo's
  stress runs: one failed in a batch of three and the bubble said "Working");
- once none works, it says how those that did not finish ended: "2 stopped"
  after a person's Stop (the card read "Milo worked 11s · 2 helpers" and
  nothing more), "1 didn't report back" for one its session took;
- a helper Claude reports stopped reads Stopped (cancelled), no longer Cut
  off, which is the engine's lost.
…erops calls too

Milo's third stress run read "1 tool used" for six Zerops calls: a call whose
result is a row of its card (a deploy, a check, a log read) was left out of the
worked line. It counts as a tool used, held whole or paged alike; a tool the
timeline never draws (the question, a tool search) still does not.

The paging test's row "a deploy left to its result" now reads "a deploy among
the tools used": the owner's run-3 report (every call the Mate made counts).
A message sent into a running run stood above its card and again as a mark in
its chat; a closed card after a reload held no lines, so it drew one copy
(Milo's third stress run, a steer twice live). It is drawn once, on its own row
above the card (the owner's rule), and so is one steered into a turn Claude
opened on its own (engine #158). The timeline tests that read the mark as a
line of the card's chat now read the card without it; the first's sentence
"marked in its chat" becomes "drawn once" by the owner's rule relayed with
run 3.
When a job's end woke the Mate, its run went on in the folded card and the
reply under that card ("…hasn't printed yet") was pulled into the card: the
reply vanished, the content dropped and the wake's answer landed below the
view (Milo's third stress run, 798 px for 10 s). An engine card's earlier
runs' answers stay where they stood under it; the latest follows them. The
engine card journey watches the reply through the wake.
"… finished · in the background" read as if it still ran (Milo's stress runs
2 and 3): an ended job's line says where it ran, "ran in the background"; a
running one still says "running in the background".
"Checking / in the browser" said the path as words (Milo's stress runs 2 and
3). The now line, the check's row and the strip's heading say "the home page";
the caption keeps "/" for what counts pages and names pictures.
…ms before its first message

A turn Claude Code opens to hand the model a background result started at its
first assistant message, but its stream comes first: a command it streamed
before that message had no turn, fell to the turn before, and closed there
unreturned (Milo's third stress run: the wake's first command read "No
result" on the previous run's card). The turn now opens at its stream's
start; its first assistant message still names where it began for the resume
cursor.
… its Mate did

Seeded trees play on the running engine (the live layer over a scripted Claude,
SQLite, the test clock): the Mate's turn starts helpers and background
commands, helpers run calls and start their own commands and helpers, each
piece of work ends its own way (finished, failed with an exit code, stopped by
the Mate or by the person, lost to a restart or a crash), an end after the turn
wakes the Mate, whose wake starts more; the person sends, steers, stops and
reloads between it all. Claude's events are its adapter's shapes (agentId,
parentToolUseId, task linkage).

The record reaches the card through the client's own adapter and store over
the engine's wire, frames and pages through the contract schemas, into the
card's logic as ChatView feeds it. After every op the oracle holds the card
against the tree: whose each line is and on which card, each piece's true end
(and never back), the waiting line, counts, the helpers' surface and the dock;
at a reload and at the end, a fresh client's faces and opened cards against
the live ones.

The table keeps each defect it found as its smallest tree; the gate's 24 fixed
seeds run in the chat gate's engine proof (ENGINE_PROOF_RANDOM=n adds fresh
seeds, ENGINE_PROOF_SEED=s replays one, minimised).
…st's own jump bound

Measured: the reply stays on screen through the wake and the view glides to
the answer in about 250 ms; the list moves it 62 px for one frame as the
answer's row lands (a LegendList insertion frame), under the 120 px the
journey holds the list's own scroll to.
…tion by a glide

The person's send pinned the end with a list jump: a steer queued under the
running card moved the conversation 104 px in one frame (Milo's third stress
run: 105 px), and a send 130 px whenever its row landed before the jump's
frame. Already at the end, the view now follows what the send brings by the
list's own glide; from elsewhere it still goes to the end at once. The engine
card journeys measure both.
# Conflicts:
#	apps/server/src/engine/domain/decide.ts
…ence goes by the owner's rule

The engine card journey's wake keeps its title, its new assertions (the reply
the person read stays through the wake) under it. A message sent into a run is
drawn once, above its card (the owner's rule relayed with Milo's third stress
run), so its mark in the card's chat is gone.

Drops-test: keeps every message the person sent above the whole card, marked in its chat
… a reload

The oracle's deep seeds (11 of 800): a card not held whole "worked" because
its summary counted its notes as work, where live it read "thought". A paged
card worked when its summary counts a call. The timeline test of a card held
unread arranges its summary with the call it stands for.
…age, drawn once above the card

The steered message's picture stood in the card's chat as well as on its
message above the card; drawn once now, the journey holds it there, intact
while the turn continues.
@fxck
fxck merged commit 570a442 into main Oct 10, 2026
20 checks passed
fxck added a commit that referenced this pull request Oct 10, 2026
- #165 test(web): witness the composer's bottom space and its clear top edge
- #163 Run-card oracle for the Mate engine, and the card defects it found
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant