Skip to content

Commit 1a0010a

Browse files
committed
Merge remote-tracking branch 'origin/main' into cl-6963-four-tier-eval-suite
2 parents 207fac2 + a3e7b48 commit 1a0010a

82 files changed

Lines changed: 2781 additions & 201 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎CHANGELOG.md‎

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -13,8 +13,17 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
1313

1414
## [Unreleased]
1515

16+
## [0.2.106] - 2026-08-23
17+
1618
### Agent
1719

20+
- **Resuming a session no longer shows a blank error when the saved history
21+
has one corrupted line.** A malformed or schema-invalid line anywhere in the
22+
saved transcript used to abort the entire resume load. The TUI's resume view
23+
now skips just the bad line (logging it) and still shows the rest of the
24+
history; a corrupt file still surfaces as an error during live conversation
25+
loading, where correctness matters more than availability.
26+
1827
- **Every stop and nudge is now logged, and so is what each dispatch produced.**
1928
`interventions.jsonl` in the worker's trace dir records each intervention with
2029
its measured value beside the threshold it crossed, the model family it fired
@@ -47,6 +56,17 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
4756

4857
## [0.2.105] - 2026-08-23
4958

59+
### Permissions
60+
61+
- **Every approval ask and how it settles is now logged.** `approvals.jsonl`
62+
in the session dir records each consequential decision — auto-mode
63+
allow/deny, or an operator prompt's allow-once / allow-with-scope / deny /
64+
timeout / abort — with the classifier rule that triggered it, queued /
65+
displayed / settled timestamps, and shell chain segment count. No command
66+
text, path, or credential is ever recorded; writes are fire-and-forget and
67+
never fail a run. `scripts/approval-forensics.ts` aggregates across local
68+
sessions.
69+
5070
### Agent
5171

5272
- **Context estimate syncs incrementally on append.** `syncFromTurns` keys
@@ -61,6 +81,13 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
6181
closures count against `maxAnchorTurns`. The LLM summary is workflow-aware
6282
and skips degenerate assistant text.
6383

84+
- **Prefix-stable summaries and growth hysteresis.** Existing compacted user
85+
turns stay byte-identical across later passes; new folds become later summary
86+
turns with an assistant spacer so the prompt prefix can stay in the KV cache.
87+
After a compact that remains over the high watermark, the governor waits for
88+
usage to grow by 10% of the window before re-arming. Overflow recovery still
89+
compacts immediately.
90+
6491
### Plugins
6592

6693
- **`run_shell` no longer defaults to a 15s timeout.** Omitted timeout arms no

‎docs/ARCHITECTURE.md‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -175,11 +175,11 @@ The agent maintains an optional **`manage_tasks`** list (create/update via the h
175175

176176
#### Context compaction (the compaction governor)
177177

178-
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The governor covers three cases:
178+
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with an assistant spacer between them so the prompt head can remain in the provider KV cache. The governor covers three cases:
179179

180-
- **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message.
180+
- **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required.
181181
- **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
182-
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever.
182+
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
183183

184184
The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor, after emitting `compact`, self-delivers a content-less inbound message (a host-supplied `requestContinuation` callback). That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history.
185185

@@ -376,6 +376,8 @@ tool call
376376
- **queue** — Headless settle registry (`src/permission/queue.ts`). Surfaces enqueue outstanding requests; `wirePermissionGrantReconciliation` listens for `permission.grant` and drains every queued request the new grant covers, without a second prompt. Teardown calls `drain()` so no awaited resolve is left hanging.
377377
- **types** — `Approval`, `ApprovalScope`, `PermissionRequest`, `ApprovalOutcome`.
378378

379+
**Approval log** (`src/permission/approval-log.ts`, CL-5666): every consequential decision the gate makes — auto-mode allow/deny or an interactive prompt's allow-once/allow-with-scope/deny/timeout/abort — is appended as one JSONL record to `approvals.jsonl` in the session dir, carrying the classifier/auto-shell rule name that fired (the existing `auto-shell-policy.ts`/`classify.ts` rule names, plus a small closed set of additional fixed literals the log itself defines for decisions those modules don't otherwise name — `auto-allowed-tool`, `non-interactive`, `mega-chain` — never model- or user-authored text), whether the decision was `auto` or `interactive`, a shell chain's segment count, and queued/displayed/settled timestamps. `displayedAt` is set by `PermissionRequest.markDisplayed`, called from `gate-wire.ts`'s `open()` the moment a request actually reaches the overlay host — distinct from when it was raised, so the gap it exposes is the CL-5664 signal (a queued gate arming its timeout before the operator could see it). No command text, file content, path, credential, or other free text is ever recorded — only tool name, rule, mode, segment count, and timing; a sub-agent's free-text dispatch label is deliberately left out, even though it would enable a per-agent breakdown, because nothing constrains what a model puts in it. A hard size cap on the serialized line is defense in depth against a future field reintroducing free text. Writes are fire-and-forget and swallow their own errors; the log defaults to a no-op so nothing depends on it being wired. `scripts/approval-forensics.ts` aggregates across local sessions the same way `intervention-forensics.ts` does for stop/nudge events: per-tool counts by outcome and mode, duration/display-delay percentiles, mega-chain counts, and a duplicate-rate proxy (sessions that hit the same rule more than once).
380+
379381
**Tool wall-clock budget vs. permission prompts.** Each tool `run()` is wrapped by an outer execution watchdog (`src/tui/tool-execution-watchdog.ts`). The watchdog arms only when Settings set `tools.timeoutMs` / `tools.maxTimeoutMs`, or when `run_shell` passes a positive timeout (requested plus slack, so this layer cannot beat shell-guard). The `task` tool is always exempt, regardless of Settings — a sub-agent run is bounded by its own limits (maxTurns, no-progress, thrash, opt-in deadlineMs), so the generic per-tool budget never aborts a healthy long-running worker; parent cancel, maxTurns, and eval `--agent-timeout-ms` still bound the run. By default (`tools.waitForApproval`, Settings → Tools, **On**), an armed budget freezes while the operator is deciding on a permission prompt, so a late approve still runs the tool and the agent waits for the decision instead of timing out under the modal. When **Off**, the budget keeps ticking during the prompt; if it expires first the tool is skipped and the permission modal is dismissed via the budget AbortSignal (auto-deny with a timeout message). The TUI permission queue (`src/tui/gate-wire.ts`, backed by `src/permission/queue.ts`) attaches that signal so ghost prompts cannot outlive an already-aborted tool.
380382

381383
`mcp__*` tool calls are the exception to "arms only when Settings set it": they arm unconditionally with a 5-minute default (`DEFAULT_MCP_TOOL_TIMEOUT_MS`), overridable via `mcp.timeoutMs` and still capped by `tools.maxTimeoutMs` (CL-6895). Nothing else bounds an MCP call — the stall watchdog treats an in-flight tool as activity by design, so a wedged MCP server previously hung a tool call, and the turn, forever. On expiry the call returns a normal tool-error result ("MCP tool `<name>` timed out after `<n>`s — the server may be wedged; retry or continue without it"); the turn is never aborted. The MCP client itself (`src/mcp/client.ts`, wrapping `@modelcontextprotocol/sdk`) multiplexes concurrent requests over one connection by JSON-RPC message id with no serial queue or mutex in our code or in the vendored SDK's `Protocol.request()` — so concurrent calls to the same server are not expected to deadlock each other. Live forensics for CL-6895 showed multi-minute MCP calls that eventually completed successfully, consistent with a slow server response rather than a client-side deadlock.

‎docs/PRODUCT.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -92,7 +92,7 @@ is the direct, explicit resume path.
9292
- Paths outside the workspace and writes under the session state root still ask; mutating MCP and unknown tools still prompt.
9393

9494
- **Path sandboxing** — Tool path arguments are resolved against the working directory; paths that escape it are blocked unless `--dangerously-skip-permissions` / `/yolo` is on (secret-guard and authz hard denies still apply).
95-
- **Write verification** — After every write/edit the file is re-read and compared to confirm the change actually landed.
95+
- **Write verification** — After every write/edit the file is re-read and compared to confirm the change actually landed; the result returned to the model (and shown to the operator) includes a bounded diff of the changed region — `write_file`, `edit_file`, `delete_file`, and each op inside `apply_patch` — so a follow-up `read_file` is never needed just to confirm an edit landed. A whole-file rewrite's diff is truncated (and says so) rather than blowing the result size cap.
9696

9797
## Slash Commands (TUI)
9898

‎evals/capability/lib.test.ts‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -667,8 +667,8 @@ describe("resolveRequestedProviderModel", () => {
667667
expect(requested).toEqual({ provider: raw.provider, model: raw.model });
668668

669669
const fallback = detectProviderFallback({
670-
requestedProvider: requested.provider,
671-
requestedModel: requested.model,
670+
...(requested.provider !== undefined ? { requestedProvider: requested.provider } : {}),
671+
...(requested.model !== undefined ? { requestedModel: requested.model } : {}),
672672
resolvedProvider: cell!.provider,
673673
resolvedModel: cell!.model,
674674
});
@@ -695,7 +695,7 @@ describe("resolveRequestedProviderModel", () => {
695695
{},
696696
{ provider: "(default)", model: "(default)" },
697697
);
698-
expect(requested).toEqual({ provider: undefined, model: undefined });
698+
expect(requested).toEqual({});
699699
});
700700
});
701701

‎evals/capability/lib.ts‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -486,9 +486,11 @@ export function resolveRequestedProviderModel(
486486
): { provider?: string; model?: string } {
487487
const requested = (v?: string): string | undefined =>
488488
v === undefined || v === "(default)" ? undefined : v;
489+
const provider = variant.provider ?? requested(labels.provider);
490+
const model = variant.model ?? requested(labels.model);
489491
return {
490-
provider: variant.provider ?? requested(labels.provider),
491-
model: variant.model ?? requested(labels.model),
492+
...(provider !== undefined ? { provider } : {}),
493+
...(model !== undefined ? { model } : {}),
492494
};
493495
}
494496

‎scripts/approval-forensics.ts‎

Lines changed: 170 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,170 @@
1+
// Aggregate scan over the approval logs written by src/permission/approval-log.ts
2+
// (~/.corbits/projects/**/approvals.jsonl) — the data CL-5666 needed to exist
3+
// before approval volume could be measured at all.
4+
//
5+
// Reports: total asks, split by mode (auto vs interactive) and outcome, a
6+
// per-rule breakdown, settle-duration and display-delay percentiles (the
7+
// display delay is the CL-5664 signal — a queued gate arming its timeout
8+
// before the operator could see it), and a mega-chain count (segments >=
9+
// MEGA_CHAIN_SEGMENT_THRESHOLD).
10+
//
11+
// Prints only aggregate counts and timings, never a tool subject or command
12+
// text — the log itself never records either, so there is nothing to leak
13+
// here even by accident.
14+
//
15+
// Run: bun run scripts/approval-forensics.ts
16+
17+
import { readdirSync, lstatSync, readFileSync } from "node:fs";
18+
import { join } from "node:path";
19+
import { homedir } from "node:os";
20+
21+
import { APPROVAL_LOG_FILE, type ApprovalRecord } from "../src/permission/approval-log.js";
22+
import { MEGA_CHAIN_SEGMENT_THRESHOLD } from "../src/permission/classify.js";
23+
24+
// lstat, and skip symlinks: session dirs carry a `latest` symlink to a real
25+
// session, and following it double-counts every record in that session.
26+
function findAll(dir: string, name: string, out: string[]): void {
27+
let entries: string[];
28+
try {
29+
entries = readdirSync(dir);
30+
} catch {
31+
return;
32+
}
33+
for (const entry of entries) {
34+
const path = join(dir, entry);
35+
let info: ReturnType<typeof lstatSync>;
36+
try {
37+
info = lstatSync(path);
38+
} catch {
39+
continue;
40+
}
41+
if (info.isSymbolicLink()) continue;
42+
if (info.isDirectory()) findAll(path, name, out);
43+
else if (entry === name) out.push(path);
44+
}
45+
}
46+
47+
function percentile(sorted: readonly number[], p: number): number {
48+
if (sorted.length === 0) return 0;
49+
const index = Math.min(sorted.length - 1, Math.floor((p / 100) * sorted.length));
50+
return sorted[index]!;
51+
}
52+
53+
interface Bucket {
54+
count: number;
55+
byOutcome: Map<string, number>;
56+
byMode: Map<string, number>;
57+
durations: number[];
58+
displayDelays: number[];
59+
megaChains: number;
60+
}
61+
62+
function emptyBucket(): Bucket {
63+
return {
64+
count: 0,
65+
byOutcome: new Map(),
66+
byMode: new Map(),
67+
durations: [],
68+
displayDelays: [],
69+
megaChains: 0,
70+
};
71+
}
72+
73+
const root = join(homedir(), ".corbits", "projects");
74+
const files: string[] = [];
75+
findAll(root, APPROVAL_LOG_FILE, files);
76+
77+
const buckets = new Map<string, Bucket>();
78+
const sessionsByRule = new Map<string, Set<string>>();
79+
let records = 0;
80+
let malformed = 0;
81+
82+
for (const file of files) {
83+
let lines: string[];
84+
try {
85+
lines = readFileSync(file, "utf8").split("\n");
86+
} catch {
87+
continue;
88+
}
89+
for (const line of lines) {
90+
if (line.trim().length === 0) continue;
91+
let record: ApprovalRecord;
92+
try {
93+
record = JSON.parse(line) as ApprovalRecord;
94+
} catch {
95+
malformed++;
96+
continue;
97+
}
98+
if (typeof record.tool !== "string" || typeof record.outcome !== "string") {
99+
malformed++;
100+
continue;
101+
}
102+
records++;
103+
const key = record.tool;
104+
let bucket = buckets.get(key);
105+
if (bucket === undefined) {
106+
bucket = emptyBucket();
107+
buckets.set(key, bucket);
108+
}
109+
bucket.count++;
110+
bucket.byOutcome.set(record.outcome, (bucket.byOutcome.get(record.outcome) ?? 0) + 1);
111+
bucket.byMode.set(record.mode, (bucket.byMode.get(record.mode) ?? 0) + 1);
112+
if (typeof record.durationMs === "number") bucket.durations.push(record.durationMs);
113+
if (typeof record.displayDelayMs === "number") bucket.displayDelays.push(record.displayDelayMs);
114+
if ((record.segments ?? 0) >= MEGA_CHAIN_SEGMENT_THRESHOLD) bucket.megaChains++;
115+
116+
// Duplicate-rate proxy: how often the same rule fires more than once per
117+
// session file (a session repeatedly asking for something it was already
118+
// told no/yes to under a different subject).
119+
if (record.rule !== undefined) {
120+
const sessions = sessionsByRule.get(record.rule) ?? new Set<string>();
121+
sessions.add(file);
122+
sessionsByRule.set(record.rule, sessions);
123+
}
124+
}
125+
}
126+
127+
console.log(`approval logs: ${files.length}`);
128+
console.log(`records: ${records}${malformed > 0 ? ` (${malformed} malformed, skipped)` : ""}`);
129+
if (records === 0) {
130+
console.log("\nNo approvals logged yet. Run some sessions first.");
131+
process.exit(0);
132+
}
133+
134+
const rows = [...buckets.entries()].sort((a, b) => b[1].count - a[1].count);
135+
console.log(
136+
"\ntool n auto/interactive duration p50/p90/max displayDelay p50/p90/max megaChains",
137+
);
138+
for (const [key, bucket] of rows) {
139+
const durations = [...bucket.durations].sort((a, b) => a - b);
140+
const delays = [...bucket.displayDelays].sort((a, b) => a - b);
141+
const durDist =
142+
durations.length === 0
143+
? "-"
144+
: `${percentile(durations, 50)}/${percentile(durations, 90)}/${durations[durations.length - 1]!}`;
145+
const delayDist =
146+
delays.length === 0
147+
? "-"
148+
: `${percentile(delays, 50)}/${percentile(delays, 90)}/${delays[delays.length - 1]!}`;
149+
const autoCount = bucket.byMode.get("auto") ?? 0;
150+
const interactiveCount = bucket.byMode.get("interactive") ?? 0;
151+
console.log(
152+
`${key.padEnd(26)} ${String(bucket.count).padStart(3)} ${String(autoCount).padStart(4)}/${String(interactiveCount).padEnd(11)} ${durDist.padEnd(24)} ${delayDist.padEnd(24)} ${bucket.megaChains}`,
153+
);
154+
}
155+
156+
console.log("\nby outcome");
157+
for (const [key, bucket] of rows) {
158+
const outcomes = [...bucket.byOutcome.entries()]
159+
.sort((a, b) => b[1] - a[1])
160+
.map(([outcome, count]) => `${outcome}=${count}`)
161+
.join(" ");
162+
console.log(`${key.padEnd(26)} ${outcomes}`);
163+
}
164+
165+
console.log("\nrule -> sessions that hit it at least once (duplicate-rate proxy)");
166+
for (const [rule, sessions] of [...sessionsByRule.entries()].sort(
167+
(a, b) => b[1].size - a[1].size,
168+
)) {
169+
console.log(`${rule.padEnd(26)} ${sessions.size}`);
170+
}

‎scripts/eval-capability.test.ts‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -5,8 +5,8 @@ import { join } from "node:path";
55
import { execFile } from "node:child_process";
66
import { promisify } from "node:util";
77

8-
import { initEvalGitRepo, mapPool, parseArgs, buildEvalDiagnostics } from "./eval-capability.ts";
9-
import type { Config } from "../src/config/index.ts";
8+
import { initEvalGitRepo, mapPool, parseArgs, buildEvalDiagnostics } from "./eval-capability.js";
9+
import type { Config } from "../src/config/index.js";
1010

1111
const execFileAsync = promisify(execFile);
1212

0 commit comments

Comments
 (0)