Skip to content

Commit 2a91dc8

Browse files
committed
Merge remote-tracking branch 'origin/main' into cl-6906-compaction-preserves-the-loop-and-drops-the-substance
# Conflicts: # CHANGELOG.md
2 parents f75d48a + 6866aef commit 2a91dc8

17 files changed

Lines changed: 321 additions & 87 deletions

‎CHANGELOG.md‎

Lines changed: 87 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,22 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
2121
closures count against `maxAnchorTurns`. The LLM summary is workflow-aware
2222
and skips degenerate assistant text.
2323

24+
### Plugins
25+
26+
- **`run_shell` no longer defaults to a 15s timeout.** Omitted timeout arms no
27+
timer (match Pi). Pass a per-call `timeout`, or set `shell.timeoutMs` in
28+
settings, to bound a command. `shell.maxTimeoutMs` still clamps a resolved
29+
timeout and does not invent one on its own. Abort and the output-byte cap are
30+
unchanged.
31+
32+
### Sub-agents
33+
34+
- **Sub-agent `maxTurns` no longer hard-caps at 100.** Default remains 30 when
35+
unset; values must still be integers ≥1. `task(maxTurns)`, profile
36+
`maxTurns`, and `settings.subagentMaxTurns` may exceed 100 for long jobs.
37+
38+
## [0.2.104] - 2026-08-23
39+
2440
### TUI
2541

2642
- **Taller live chain-of-thought preview.** Parent reasoning still paints
@@ -31,12 +47,81 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
3147
unchanged. Assistant mid-turn text continues to grow the open streaming
3248
assistant row from `inference.text.delta`.
3349

34-
### Fixed
50+
- **Live agents sit in a chrome strip above the prompt.** Running and
51+
finished workers no longer compete with the transcript for vertical
52+
space; the strip stays parked over the input, finished rows linger
53+
briefly, then it clears when idle. Transcript task-row rewrites pause
54+
while the strip owns live status.
55+
56+
### Tools
57+
58+
- **`edit_file` filler args no longer count as a second mode.** Models pad
59+
the unused mode with `start_line: 0` / `end_line: 0` / `old_string: ""`.
60+
Those now count as absent, so substring vs line-range is chosen from the
61+
real fields. Mixed-mode calls still reject, and the error names exactly
62+
which fields to drop so a retry can differ.
63+
64+
- **`task` rejections name only the missing field.** A typed brief that
65+
omitted `prompt` used to be told both `description` and `prompt` were
66+
required, so the model retried the identical call. The error now names
67+
the actual gap and echoes the valid field back.
68+
69+
- **Truncation no longer promises a retrievable remainder.** Tool results
70+
cut at 80,000 chars now say the discarded tail is gone and re-running
71+
yields the same cut, instead of pointing at a `tool-output:///` blob that
72+
only held the truncated text.
73+
74+
### Sub-agents
75+
76+
- **Parents and the TUI see why a child stopped.** Forced stops (repetition,
77+
stall, deadline, turn-budget, no-progress, operator cancel, thrash) carry
78+
a machine-readable `Stopped:` line on the report and a reason on the
79+
child's session. Fleet rows announce `<lane> stopped — <reason>` instead
80+
of a silent done/cancelled.
81+
82+
- **Repetition detection covers short-phrase, counter, emoji, and
83+
zero-width floods.** The periodic window floor drops to 8 chars (with a
84+
higher repeat bar so healthy lists stay quiet). A digit-folded pass
85+
catches incrementing counters and fence/emoji floods; a contentless-growth
86+
check flags streams of invisibles that used to normalize to healthy text.
87+
88+
### Providers
89+
90+
- **Cross-provider replay no longer 400s the rest of the session.**
91+
Switching model/provider mid-session used to replay foreign thinking
92+
signatures and output-only blocks the new adapter cannot encode. Every
93+
adapter now sanitizes persisted history before `buildRequest`: drop
94+
unmappable blocks, strip foreign signatures, and synthesize dangling
95+
`tool_result`s.
96+
97+
- **Grok and OpenAI Responses set `prompt_cache_key` per session.** Codex
98+
already did; xAI and Go Responses did not, so Grok threads cached at
99+
~66–72% versus Codex's 92%+. Parent and each sub-agent thread get a
100+
stable, distinct key.
101+
102+
- **TUI first inference waits for Codex instructions refresh.** The TUI
103+
used to fire the refresh un-awaited, so turn 2's request prefix could
104+
change under a live cache key and force a full miss. Delivery now waits
105+
for that promise (non-Codex profiles skip it; a failed refresh still
106+
falls back to cached/bundled copy).
35107

36108
- **Codex Responses no longer sends `reasoning.summary: "auto"`.** ChatGPT
37109
Codex rejects that value for gpt-5.6-terra / gpt-5.3-codex family models
38110
(HTTP 400 at turn 0). The adapter now sends `{ effort }` only, matching
39-
Codex CLI catalog `default_reasoning_summary=none` (CL-6893).
111+
Codex CLI catalog `default_reasoning_summary=none`.
112+
113+
- **Codex native tools proxy onto Corbits tools.** `apply_patch`,
114+
`exec_command`, and `update_plan` from Codex-family models land on the
115+
real file, shell, and `manage_tasks` handlers instead of being rejected
116+
as unknown names.
117+
118+
### CI
119+
120+
- **Required checks are `prettier`, `eslint`, `typecheck`, and
121+
`build-and-test`.** The old combined `lint` job (cached, continue-on-error)
122+
never reported the status contexts the main ruleset required, so every PR
123+
sat blocked. Lint result caches are gone in CI; local `bun run lint` still
124+
uses `--cache`.
40125

41126
## [0.2.103] - 2026-08-23
42127

‎docs/ARCHITECTURE.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -106,7 +106,7 @@ In TUI chat mode there is no completion gate — the session stays open across t
106106
Two directors, selected by role:
107107

108108
- **ChatDirector** (interactive, `src/agent/director.ts`) — Extends `DefaultDirector` with task list tracking, workflow nudges, LSP auto-activation, and multi-turn chat semantics. It never terminates the session: operator declines are surfaced as replies and the reactor stays alive for the next message. Auto mode is toggled by CLI flags (`--auto` / `--no-auto`); there is currently no in-session key to toggle it (default on; constrained envelope — workspace writes and unconstrained shell auto-allow; installs, recursive rm, force/uncontained worktree changes, sensitive-path and opaque-wrapper shell still ask; contained non-force `git worktree add`/`remove`/`prune` and `list` auto-allow; shell file-mutation denied). It is not a separate edit/plan mode.
109-
- **SubAgentDirector** (delegated work, `src/subagent/index.ts`) — Drives a dispatched worker until a turn arrives with no tool calls, then replies with the final assistant text and ends the run. A tool-less turn **after tools** completes only with the four-heading envelope (Summary, Findings, Blockers, Paths); a missing envelope nudges once then salvages as **incomplete-report**. A tool-less completion with **zero tool calls in the entire run** is returned as a **never-acted** salvage report (not a successful implement). When `task(intent="implement")` is set, a tool-using run that never wrote/edited/deleted a file is returned as **never-edited** instead of complete — so a pure-explore "plan" cannot look shipped to the parent (tracked via `thrashState.editedPaths` from `edit_file` / `write_file` / `delete_file`). Explore/read-only workers that used tools then replied with findings remain normal completes. Hard stops also fire after 5 consecutive identical tool-call fingerprints (**no-progress**, mirroring the director-level `IDENTICAL_REPEAT_MIN` threshold), on progressive re-read pressure (**thrash** — the same path re-read past a limit amid enough tool volume, tracked by `src/subagent/thrash.ts`), or after the leaf turn budget (**turn-budget**, default 30, overridable via `task(maxTurns)`, agent profile `maxTurns`, or `settings.subagentMaxTurns`, capped at 100), each returning a structured salvage report (reason, partial findings, blockers) so a thrashing child cannot burn tokens indefinitely. Before hard thrash, a one-shot **re-read-nudge** fires when re-read pressure crosses a soft threshold (default 3 same-path reads with enough tool volume, still below the hard re-read limit of 4): the director injects an ephemeral redirect — implement leaves are asked to edit or wrap up; explore leaves are asked to expand findings / change approach / report, never forced into edit — then keeps running so hard thrash remains reachable if the leaf ignores it. A fourth hard stop, **repetition**, is detected outside the director entirely:
109+
- **SubAgentDirector** (delegated work, `src/subagent/index.ts`) — Drives a dispatched worker until a turn arrives with no tool calls, then replies with the final assistant text and ends the run. A tool-less turn **after tools** completes only with the four-heading envelope (Summary, Findings, Blockers, Paths); a missing envelope nudges once then salvages as **incomplete-report**. A tool-less completion with **zero tool calls in the entire run** is returned as a **never-acted** salvage report (not a successful implement). When `task(intent="implement")` is set, a tool-using run that never wrote/edited/deleted a file is returned as **never-edited** instead of complete — so a pure-explore "plan" cannot look shipped to the parent (tracked via `thrashState.editedPaths` from `edit_file` / `write_file` / `delete_file`). Explore/read-only workers that used tools then replied with findings remain normal completes. Hard stops also fire after 5 consecutive identical tool-call fingerprints (**no-progress**, mirroring the director-level `IDENTICAL_REPEAT_MIN` threshold), on progressive re-read pressure (**thrash** — the same path re-read past a limit amid enough tool volume, tracked by `src/subagent/thrash.ts`), or after the leaf turn budget (**turn-budget**, default 30, overridable via `task(maxTurns)`, agent profile `maxTurns`, or `settings.subagentMaxTurns`; floor ≥1, no hard upper cap), each returning a structured salvage report (reason, partial findings, blockers) so a thrashing child cannot burn tokens indefinitely. Before hard thrash, a one-shot **re-read-nudge** fires when re-read pressure crosses a soft threshold (default 3 same-path reads with enough tool volume, still below the hard re-read limit of 4): the director injects an ephemeral redirect — implement leaves are asked to edit or wrap up; explore leaves are asked to expand findings / change approach / report, never forced into edit — then keeps running so hard thrash remains reachable if the leaf ignores it. A fourth hard stop, **repetition**, is detected outside the director entirely:
110110
`runSubAgent`'s stream sink watches the streamed text of the in-flight cycle for degenerate token loops (`src/subagent/repetition.ts`) — format chars (ZWSP, BOM, bidi marks, soft hyphen, …) stripped then whitespace-collapsed raw text, a smallest-period KMP check over the probe tail, default window >= 16 chars repeated >= 8 times, evaluated every 256 streamed chars — and on a hit aborts the run controller mid-cycle, returning a `repetition` salvage report that leads with the looped window and warns the parent against re-dispatching the identical brief. `inference.thinking.delta` is sampled the same way on its own buffer, but with digit runs folded to one placeholder and a shorter window (>= 4 chars repeated >= 32 times), gated to periods <= 16 chars once folded: thinking is never rendered to the user, so a monotonic counter (e.g. `0/1 1/2 2/3 …`, which stays non-periodic and escapes the raw-text check) can be caught, but folding still erases real information — a healthy templated enumeration line becomes byte-identical to its neighbors once digits are erased, so the period-length cap only lets counter-shaped folded periods (a handful of chars) through and refuses the much longer periods a folded prose line produces. Because directors only see completed turns, this is the only stop that can catch a loop inside a single turn that never finishes. A one-shot **report-forced** signal fires a few turns before the cap while the leaf is still tooling — it is not a stop: the director injects a wrap-up nudge and lets the leaf finish on its own, so turn-budget stays reachable for a leaf still making progress. When both report-forced and re-read-nudge apply, report-forced wins (near-budget wrap-up is more urgent than a mid-run redirect). Operator/parent cancel after any progress likewise returns a **cancelled** salvage report (partial findings + tool activity) instead of a bare cancel string; cancel before progress still surfaces as cancelled-by-operator.
111111
Optional `task(tier=)` (`fast` | `standard` | `clever`) overrides profile inference, profile tier, and the parent provider for that spawn only, and fails closed when the tier is unconfigured. The parent `task` tool keeps a session-scoped brief-dispatch ledger (`src/subagent/brief-dispatch.ts`): fingerprints cover prompt + agent + intent + success_criteria + do_not (not maxTurns/description/tier). After thrash / no-progress / repetition / never-acted / never-edited salvage, an identical re-dispatch is hard-blocked for the rest of the parent chat; change at least one fingerprint field to force a re-run. Turn-budget salvage still invites a higher maxTurns for a few same-brief retries without a successful complete, then flips the parent hint to stop and change approach (soft — further identical dispatches are still admitted). A successful complete resets the same-brief retry budget.
112112

@@ -354,7 +354,7 @@ tool call
354354
- **Secret Guard** (`secret-guard-plugin.ts`) — Hard-denies path-keyed tool calls (`read_file`, `write_file`, …) that would put a sensitive file into (or write it from) the model context. Runs before the permission plugin, so the path-arg deny holds even under `--dangerously-skip-permissions`. Shell commands that _reference_ a sensitive path (tokenized so `cat .env`, `bun --env-file=.env run …`, and quote/env-assignment forms are detected) are not hard-denied here: they require operator approval via the permission gate, and auto mode forces an ask through the auto-shell policy (`sensitive-path` rule). Once the operator approves, the command runs. Shell detection is best-effort: token matching defeats quoting and env-assignment/redirection forms but not dynamic path construction (variable indirection, `printf` assembly). Tool-result secret scrub still redacts credential-shaped output.
355355
- **Authorization** (`run-shell-authz.ts`, wired by `authz-plugin.ts`) — Denies catastrophic shell command patterns by regex, and hard-blocks shell `find`, head-position `rg`, and recursive `grep -r` (they can walk huge trees and OOM the host). Bounded `grep`/`search_files` tools remain practical alternatives (timeout + output caps); the patterns match those three command shapes only — an `ls -R`, `fd`, or scripted `os.walk` is just as unbounded and is not caught, so the block message tells the model not to substitute one. The permission gate’s shell auto-allow path consults the same policy so it never pre-approves a command authz would reject.
356356
- **Permission** (`permission-plugin.ts`) — Delegates consequential calls to the permission gate.
357-
- **Shell Guard** (`shell-guard-plugin.ts`) — Corbits Code-only replacement for stock `run_shell` (interchange stays unpatched): 15s default timeout, 512KB display cap with head+tail retention (the process keeps running when the cap is hit), process-group kill on timeout/abort only. Also applies a 10s wall-clock budget to `grep`/`search_files`.
357+
- **Shell Guard** (`shell-guard-plugin.ts`) — Corbits Code-only replacement for stock `run_shell` (interchange stays unpatched): no built-in default timeout (optional per-call or `settings.shell.timeoutMs`; `maxTimeoutMs` clamps only a resolved timeout), 512KB display cap with head+tail retention (the process keeps running when the cap is hit), process-group kill on timeout/abort only. Also applies a 10s wall-clock budget to `grep`/`search_files`.
358358
- **Read File Guard** (`read-file-guard-plugin.ts`) — Corbits Code-only short-circuit for `read_file` on real filesystem paths and configured `tool-output://` URIs (interchange stays unpatched): streaming reads that never decode the whole file in one pass, caps model-facing output at 50KB, defaults to 2000 lines, truncates long lines with recovery hints, samples the first chunk to reject binary, and stops at an 8MB scan ceiling. Emits `offset` continuation notices so the model can page without losing file or spill content on disk.
359359
- **Verify** (`verify-plugin.ts`) — Re-reads after `write_file` / `edit_file` and errors on mismatch. Per-path serialization (`file-mutation-lock.ts`) prevents parallel edits on one file from tripping verification.
360360
- **Edit file line range** (`edit-file-line-range-plugin.ts`) — Corbits Code-only short-circuit for `edit_file` mode B (`start_line`/`end_line`/`new_string`), same pattern as shell-guard; schema advertised via `advertiseEditFileLineRange`. Modes are mutually exclusive: a call supplying both `old_string` and `start_line`/`end_line` is rejected with a recoverable error naming which fields to omit (no file-content disambiguation).

‎docs/IMPLEMENTATION.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -246,7 +246,7 @@ Provider and model configuration lives in JSON settings files. The global file h
246246

247247
On expiry the call returns a normal tool-error result ("MCP tool `<name>` timed out after `<n>`s — the server may be wedged; retry or continue without it") that the model can react to; the turn itself is never aborted. `tools.maxTimeoutMs`, if set, still caps `mcp.timeoutMs`.
248248

249-
Optional `subagentMaxTurns` (integer **1–100**, default **30**) sets the default inference-turn budget for dispatched workers (not the parent chat session limit). Per-dispatch `task(maxTurns)` and agent profile `maxTurns` override this default; values above **100** are rejected on `task` and clamped for profiles. Always applies — the primary session is always orchestrator-capable (CL-5814).
249+
Optional `subagentMaxTurns` (integer **≥1**, default **30**) sets the default inference-turn budget for dispatched workers (not the parent chat session limit). Per-dispatch `task(maxTurns)` and agent profile `maxTurns` override this default; there is no hard upper cap (values are floor-sanitized to ≥1). Always applies — the primary session is always orchestrator-capable (CL-5814).
250250

251251
Optional `sessionMode` is **deprecated**. Legacy values (`single` | `orchestrator`) may still appear on disk and load without error; resolve always returns **orchestrator**. There is no first-run mode picker and no Settings row. Both the interactive TUI (`runTUI`) and the non-TUI product path (`runExec` / `corbits exec`) are orchestrator-only. Exec bootstrap is otherwise a forked copy of the TUI path (shared stack, intentional deltas documented under Architecture → Exec Runner).
252252

0 commit comments

Comments
 (0)