Skip to content

record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): the captured-arm verdict is token-clean at three sizes (#2566) - #2567

Open
lu-zero wants to merge 1 commit into
mudler:mainfrom
lu-zero:record/tt-4b-mistral-abc
Open

record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): the captured-arm verdict is token-clean at three sizes (#2566)#2567
lu-zero wants to merge 1 commit into
mudler:mainfrom
lu-zero:record/tt-4b-mistral-abc

Conversation

@lu-zero

@lu-zero lu-zero commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator

record(BACKEND-TENSTORRENT-HOST-FREE-FORWARD): the captured-arm verdict is token-clean at three sizes (#2566)

Closes #2566.

Record-only change: it carries the two measurement paragraphs the repair
PR #2470 could not take without restarting its CI, and one benchmark-record
entry, and nothing else.

The row spec's ## Now gains the Qwen3-4B token-clean A/B/C
re-measurement (hf-abc4b.sh, 2026-09-01, 2 order-alternated triples at
the repaired head 081efabc7: default 9.25 / opt-out 8.89 / captured
13.84 tok/s, 1.50x/1.56x, replays 474 and 0 fatals per leg, capture legs
coherent and deterministic), which supersedes the correctness caveat on
the earlier 13.90 machinery-only figure exactly as the 0.6B re-measurement
superseded the 27.57 one. It also gains the Mistral-7B captured arm
(hf-mist-abc.sh, 2026-09-01, 2 triples, --max-tokens 64 --repeat 5,
card reset first: default 11.51 / opt-out 5.93 / captured 14.23 tok/s,
1.24x/2.40x, 314 replays, 0 fatals, rc 0 both triples), whose capture
legs' output is TOKEN-IDENTICAL to the default arm's in both triples —
the strongest coherence evidence the row has, and on the model with the
sharpest default-vs-opt-out inversion.

.agents/benchmark-record.md gains one entry covering the three-size
campaign with the per-leg recipe, the arm definitions, and the raw-log
naming, so the R5-era 27.1 and the machinery-only 27.57/13.90 figures read
as superseded rather than current.

No code, no checker, no public document; the row itself landed as
4a5935b with both repairs (#2461, #2469) this unit measures behind it.
The captured-default flip stays gated on #1625 plus the recorded
## Owed residuals — this change decides nothing about it.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]

…ct is token-clean at three sizes (mudler#2566)

The row spec's ## Now gains the two measurement paragraphs the repair PR
could not carry without restarting its CI: the Qwen3-4B token-clean A/B/C
re-measurement (default 9.25 / opt-out 8.89 / captured 13.84 tok/s, which
supersedes the 13.90 machinery-only figure's correctness caveat) and the
Mistral-7B captured arm (11.51 / 5.93 / 14.23 tok/s, capture legs
token-identical to eager). The benchmark record gains one entry covering
the three-size campaign, with the per-leg recipe and the raw-log naming,
so the R5-era machinery-only figures read as superseded rather than
current.

Record-only change: no code, no checker, no public document; the spec and
the benchmark record are the two surfaces this unit owes. The row itself
landed as 4a5935b; both repairs it measured (mudler#2461, mudler#2469) are in that
merge, and the captured-default flip stays gated on mudler#1625 plus the
recorded ## Owed residuals.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai-glm-5.3-flash [maki]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Record the token-clean 4B and Mistral-7B A/B/C capture measurements in the row spec and benchmark record

1 participant