Repository navigation
Run-card oracle for the Mate engine, and the card defects it found - #163
Merged
Merged
Conversation
…s Stop closes the CLI ClaudeAdapter's stopSessionInternal reports every live task stopped (with its linkage) before the session exits; the scripted provider dropped them, so a person's Stop met work with no report of its own.
… it and how it ended
The run-card oracle (milo stress runs 2 and 3) found the record losing what the
card needs to tell a helper's work from the Mate's:
- a task a helper's tool started (Claude's owning agentId, Codex's
parentAgentId) is recorded as that helper's and goes on under the run its
helper served; a helper's own call goes there too, never onto the next
message's run, which it made a card of its own after a reload;
- the work names the call that started it (its tool use), and a closed card's
outcome read carries that call, so its line of jobs reads the same after a
reload;
- its end keeps its report's first line ("failed with exit code 3");
- work a Stop closed with Claude's session is stopped, not lost: Claude
reports its tasks only after the turn's end, from a session already closed.
… end from the record Background work's activities carry the helper the record names (agentId), the call that started it (toolUseId, so a job is its command's and a helper's launch is its helper's row) and its report (detail: "Exit code 3"). The guess by a helper call's words is gone: it held only while the helper's calls were held, so the dock and the card drew a helper's job as the Mate's after a reload, and a helper's launch read as a finished one-second step while it worked (Milo's third stress run). The pipeline test arranges its record as the engine now writes it.
…mands as the Mate's Milo's third stress run waited on "its helpers and 2 background commands": the commands were the nested helper's. A helper's own work counts as its helpers' (the oracle's row: a card waiting on its helpers).
…tarted Anything a helper started (agentId) is the helper's, its own helpers too: the helpers' surface keeps it under its helper (spawnedBy). Milo's second stress run merged a nested helper started 40 s later into "Started 4 helpers".
…elpers ended The bubble's words and state are a logic of their own (helpersBubbleOf), as the card and the oracle read them: - a helper that failed marks the bubble while the others still work (Milo's stress runs: one failed in a batch of three and the bubble said "Working"); - once none works, it says how those that did not finish ended: "2 stopped" after a person's Stop (the card read "Milo worked 11s · 2 helpers" and nothing more), "1 didn't report back" for one its session took; - a helper Claude reports stopped reads Stopped (cancelled), no longer Cut off, which is the engine's lost.
…erops calls too Milo's third stress run read "1 tool used" for six Zerops calls: a call whose result is a row of its card (a deploy, a check, a log read) was left out of the worked line. It counts as a tool used, held whole or paged alike; a tool the timeline never draws (the question, a tool search) still does not. The paging test's row "a deploy left to its result" now reads "a deploy among the tools used": the owner's run-3 report (every call the Mate made counts).
A message sent into a running run stood above its card and again as a mark in its chat; a closed card after a reload held no lines, so it drew one copy (Milo's third stress run, a steer twice live). It is drawn once, on its own row above the card (the owner's rule), and so is one steered into a turn Claude opened on its own (engine #158). The timeline tests that read the mark as a line of the card's chat now read the card without it; the first's sentence "marked in its chat" becomes "drawn once" by the owner's rule relayed with run 3.
When a job's end woke the Mate, its run went on in the folded card and the
reply under that card ("…hasn't printed yet") was pulled into the card: the
reply vanished, the content dropped and the wake's answer landed below the
view (Milo's third stress run, 798 px for 10 s). An engine card's earlier
runs' answers stay where they stood under it; the latest follows them. The
engine card journey watches the reply through the wake.
"… finished · in the background" read as if it still ran (Milo's stress runs 2 and 3): an ended job's line says where it ran, "ran in the background"; a running one still says "running in the background".
"Checking / in the browser" said the path as words (Milo's stress runs 2 and 3). The now line, the check's row and the strip's heading say "the home page"; the caption keeps "/" for what counts pages and names pictures.
…ms before its first message A turn Claude Code opens to hand the model a background result started at its first assistant message, but its stream comes first: a command it streamed before that message had no turn, fell to the turn before, and closed there unreturned (Milo's third stress run: the wake's first command read "No result" on the previous run's card). The turn now opens at its stream's start; its first assistant message still names where it began for the resume cursor.
… its Mate did Seeded trees play on the running engine (the live layer over a scripted Claude, SQLite, the test clock): the Mate's turn starts helpers and background commands, helpers run calls and start their own commands and helpers, each piece of work ends its own way (finished, failed with an exit code, stopped by the Mate or by the person, lost to a restart or a crash), an end after the turn wakes the Mate, whose wake starts more; the person sends, steers, stops and reloads between it all. Claude's events are its adapter's shapes (agentId, parentToolUseId, task linkage). The record reaches the card through the client's own adapter and store over the engine's wire, frames and pages through the contract schemas, into the card's logic as ChatView feeds it. After every op the oracle holds the card against the tree: whose each line is and on which card, each piece's true end (and never back), the waiting line, counts, the helpers' surface and the dock; at a reload and at the end, a fresh client's faces and opened cards against the live ones. The table keeps each defect it found as its smallest tree; the gate's 24 fixed seeds run in the chat gate's engine proof (ENGINE_PROOF_RANDOM=n adds fresh seeds, ENGINE_PROOF_SEED=s replays one, minimised).
…it, so no wake follows for it
…st's own jump bound Measured: the reply stays on screen through the wake and the view glides to the answer in about 250 ms; the list moves it 62 px for one frame as the answer's row lands (a LegendList insertion frame), under the 120 px the journey holds the list's own scroll to.
…tion by a glide The person's send pinned the end with a list jump: a steer queued under the running card moved the conversation 104 px in one frame (Milo's third stress run: 105 px), and a send 130 px whenever its row landed before the jump's frame. Already at the end, the view now follows what the send brings by the list's own glide; from elsewhere it still goes to the end at once. The engine card journeys measure both.
# Conflicts: # apps/server/src/engine/domain/decide.ts
…ence goes by the owner's rule The engine card journey's wake keeps its title, its new assertions (the reply the person read stays through the wake) under it. A message sent into a run is drawn once, above its card (the owner's rule relayed with Milo's third stress run), so its mark in the card's chat is gone. Drops-test: keeps every message the person sent above the whole card, marked in its chat
… a reload The oracle's deep seeds (11 of 800): a card not held whole "worked" because its summary counted its notes as work, where live it read "thought". A paged card worked when its summary counts a call. The timeline test of a card held unread arranges its summary with the call it stands for.
…age, drawn once above the card The steered message's picture stood in the card's chat as well as on its message above the card; drawn once now, the journey holds it there, intact while the turn continues.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A run-card oracle for the Mate engine, and the defects it found between the engine's record and the run card.
The oracle (
apps/web/test/engine-oracle)Seeded trees play on the running engine (live layer, scripted Claude, SQLite, test clock): the Mate's turn starts helpers and background commands, helpers run calls and start their own commands and helpers, each piece of work ends its own way (finished, failed with an exit code, stopped by the Mate or the person, lost to a restart or a crash), an end after the turn wakes the Mate, whose wake starts more; the person sends, steers, stops and reloads in between. The record reaches the card through the client's own adapter and store over the engine's wire (frames and pages through the contract schemas), into the card's logic as ChatView feeds it. After every op the oracle holds the card against the tree (whose each line is and on which card, each piece's true end, the waiting line, counts, the helpers' surface, the dock); at a reload and at the end, a fresh client against the live one. Failures minimise to their smallest tree.
runCard.oracle.test.ts: 12 minimised defect trees + 24 fixed seeds, in chat-gate stage E;ENGINE_PROOF_RANDOM=n/ENGINE_PROOF_SEED=sfor deep runs. About 2,100 random seeds were run during the lane.Fixed (each with its test)
Not covered
Two empty picture boxes in the final reply (run 3, +3:27.5), and the live +130 px send jump (the fake glides already).
Gates
Targeted suites, package typechecks,
ci-localgreen. chat-gate E, C-engine, types green; C fails onlycomposer-overlap(plain, review), which fails on origin/main too.🤖 Generated with Claude Code