Skip to content

Commit b14a01c

Browse files
committed
Spill oversized tool results without a name allowlist
Fleet verbs are AgentTools and never entered the posix truncation plugin, so an allowlist of posix names could not cover wait_agents or search_agents. Leisure now applies to every non-error result over the gate, and the same helper wraps AgentTools at mount.
1 parent 07416ba commit b14a01c

5 files changed

Lines changed: 278 additions & 54 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -363,7 +363,7 @@ tool call
363363

364364
**Rejection behavior:** Any plugin can short-circuit by returning a `ToolResult` with `isError: true`; the error propagates to the agent and downstream plugins/execution are skipped.
365365

366-
- **Result truncation / leisure materialization** (`result-truncation-plugin.ts`, `tool-result-materialize.ts`) — Caps model-facing tool results at 10,000 chars (aligned with the reactor size-cap). Over the gate, content is leisure-materialized first (minified JSON → pretty `application/json`; NDJSON preserved; else `text/plain`), then the formatted bytes are spilled to the session blob store under `{callId}:full` and truncated inline with a `tool-output:///` URI plus absolute `contextDir/tool-output/…` path when plumbed. Under-gate results are unchanged (no pretty, no spill). MCP tools apply the same scrub-then-truncate path via `mcpClientToAgentTools` since they skip the posix middleware chain.
366+
- **Result truncation / leisure materialization** (`result-truncation-plugin.ts`, `tool-result-materialize.ts`) — Caps model-facing tool results at 10,000 chars (aligned with the reactor size-cap). Over the gate, any non-error result is leisure-materialized first (minified JSON → pretty `application/json`; NDJSON preserved; else `text/plain`), then the formatted bytes are spilled to the session blob store under `{callId}:full` and truncated inline with a `tool-output:///` URI plus absolute `contextDir/tool-output/…` path when plumbed. Under-gate results are unchanged (no pretty, no spill). Posix tools go through the middleware in `buildCorePosixToolPlugins` (Codex `posixTools.run` included). Fleet AgentTools (`wait_agents`, `search_agents`, …) skip that posix chain, so the same helper wraps them at mount in `createAgentToolset` and nested `runSubAgent`. MCP tools apply the same scrub-then-truncate path via `mcpClientToAgentTools` since they skip both.
367367
- **Path Escape** (`path-escape-plugin.ts`) — Canonicalizes path-like arguments against `cwd` and blocks `..` escapes, except into a root the permission layer's worktree-roots provider allowlists (e.g. a sibling git worktree of the same repo). Runs first so later plugins see resolved paths.
368368
- **Tool-output URI** (`tool-output-uri-plugin.ts`) — Normalizes mistaken `read_file` blob URIs to `tool-output:///id` (corbits-only; interchange stays unpatched).
369369
- **Secret Guard** (`secret-guard-plugin.ts`) — Hard-denies path-keyed tool calls (`read_file`, `write_file`, …) that would put a sensitive file into (or write it from) the model context. Runs before the permission plugin, so the path-arg deny holds even under `--dangerously-skip-permissions`. Shell commands that _reference_ a sensitive path (tokenized so `cat .env`, `bun --env-file=.env run …`, and quote/env-assignment forms are detected) are not hard-denied here: they require operator approval via the permission gate, and auto mode forces an ask through the auto-shell policy (`sensitive-path` rule). Once the operator approves, the command runs. Shell detection is best-effort: token matching defeats quoting and env-assignment/redirection forms but not dynamic path construction (variable indirection, `printf` assembly). Tool-result secret scrub still redacts credential-shaped output.

‎src/agent/tools.ts‎

Lines changed: 16 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,11 @@ import type { PermissionGate } from "../permission/gate.js";
2121
import { buildCorePosixToolPlugins } from "./posix-tool-plugins.js";
2222
import { createLazyBlobReader } from "./lazy-blob-reader.js";
2323
import type { BlobReader } from "@intx/types/runtime";
24-
import type { SpillBlobWriter } from "../plugins/result-truncation-plugin.js";
24+
import {
25+
wrapAgentToolResultTruncation,
26+
wrapAgentToolsWithResultTruncation,
27+
type SpillBlobWriter,
28+
} from "../plugins/result-truncation-plugin.js";
2529
import {
2630
connectMCPServer as connectMCPClient,
2731
type MCPClient,
@@ -413,6 +417,10 @@ export async function createAgentToolset(
413417
);
414418
}
415419

420+
const truncationOptions = {
421+
...(getBlobWriter !== undefined ? { getBlobWriter } : {}),
422+
...(getContextDir !== undefined ? { getContextDir } : {}),
423+
};
416424
const posixTools = createPosixTools({
417425
cwd,
418426
...(sessionBlobReader !== undefined
@@ -426,8 +434,7 @@ export async function createAgentToolset(
426434
...(sessionBlobReader !== undefined
427435
? { readFileGuard: { blobReader: sessionBlobReader } }
428436
: {}),
429-
...(getBlobWriter !== undefined ? { getBlobWriter } : {}),
430-
...(getContextDir !== undefined ? { getContextDir } : {}),
437+
...truncationOptions,
431438
...(shellEnv !== undefined ? { shellEnv } : {}),
432439
getBackgroundShellRegistry: () => backgroundShells,
433440
}),
@@ -698,8 +705,9 @@ export async function createAgentToolset(
698705
}),
699706
);
700707

701-
const primaryTools = baseTools.filter(
702-
(tool) => tool.definition.name !== "apply_patch",
708+
const primaryTools = wrapAgentToolsWithResultTruncation(
709+
baseTools.filter((tool) => tool.definition.name !== "apply_patch"),
710+
truncationOptions,
703711
);
704712

705713
const dynamicRunner = createDynamicToolRunner(primaryTools, toolWatchdog);
@@ -796,9 +804,10 @@ export async function createAgentToolset(
796804
};
797805

798806
const mountWebFetch = (tool: AgentTool): void => {
807+
const wrapped = wrapAgentToolResultTruncation(tool, truncationOptions);
799808
dynamicRunner.removeTools(["web_fetch"]);
800-
dynamicRunner.addTools([tool]);
801-
replaceInheritedTool("web_fetch", tool);
809+
dynamicRunner.addTools([wrapped]);
810+
replaceInheritedTool("web_fetch", wrapped);
802811
};
803812

804813
const swapBuiltinExaToNative = (): void => {

‎src/plugins/result-truncation-plugin.test.ts‎

Lines changed: 159 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -11,6 +11,7 @@ import {
1111
resultTruncationPlugin,
1212
spillBlobKey,
1313
truncateToolResultContent,
14+
wrapAgentToolResultTruncation,
1415
} from "./result-truncation-plugin.js";
1516
import { toolOutputAbsolutePath } from "./tool-result-materialize.js";
1617
import { CREDENTIAL_REDACTION } from "./tool-result-secret-scrub.js";
@@ -373,6 +374,164 @@ describe("resultTruncationPlugin", () => {
373374
expect(result.content).toEqual(record);
374375
expect(store.blobs.size).toBe(0);
375376
});
377+
378+
test("spills oversized minified fleet JSON for wait_agents and list_agents", async () => {
379+
const store = fakeBlobStore();
380+
const obj = {
381+
results: Array.from({ length: 80 }, (_, i) => ({
382+
agent_id: `agent-${i}`,
383+
report: "x".repeat(200),
384+
})),
385+
timed_out: false,
386+
};
387+
const minified = JSON.stringify(obj);
388+
expect(minified.length).toBeGreaterThan(MAX_RESULT_CHARS);
389+
const pretty = JSON.stringify(obj, null, 2);
390+
const plugin = resultTruncationPlugin({ getBlobWriter: () => store.writeBlob });
391+
if (plugin.middleware === undefined) throw new Error("expected middleware");
392+
const middleware = plugin.middleware(async (call) => ({
393+
callId: call.id,
394+
content: minified,
395+
}));
396+
397+
for (const name of ["wait_agents", "list_agents"] as const) {
398+
const callId = `call-${name}`;
399+
const result = await middleware(
400+
{ id: callId, name, arguments: {} },
401+
new AbortController().signal,
402+
);
403+
expect(typeof result.content).toBe("string");
404+
expect(String(result.content).length).toBeLessThanOrEqual(MAX_RESULT_CHARS);
405+
const uri = `tool-output:///${spillBlobKey(callId)}`;
406+
expect(String(result.content)).toContain(uri);
407+
const recovered = new TextDecoder().decode(await createBlobReader(store).read(uri));
408+
expect(recovered).toBe(pretty);
409+
}
410+
});
411+
412+
test("spills an oversized search_agents string payload", async () => {
413+
const store = fakeBlobStore();
414+
const original = `Matching agent profiles:\n\n${"body ".repeat(MAX_RESULT_CHARS)}`;
415+
expect(original.length).toBeGreaterThan(MAX_RESULT_CHARS);
416+
const plugin = resultTruncationPlugin({ getBlobWriter: () => store.writeBlob });
417+
if (plugin.middleware === undefined) throw new Error("expected middleware");
418+
const middleware = plugin.middleware(async (call) => ({
419+
callId: call.id,
420+
content: original,
421+
}));
422+
const result = await middleware(
423+
{ id: "call-search", name: "search_agents", arguments: {} },
424+
new AbortController().signal,
425+
);
426+
expect(String(result.content).length).toBeLessThanOrEqual(MAX_RESULT_CHARS);
427+
const uri = `tool-output:///${spillBlobKey("call-search")}`;
428+
expect(String(result.content)).toContain(uri);
429+
const recovered = new TextDecoder().decode(await createBlobReader(store).read(uri));
430+
expect(recovered).toBe(original);
431+
});
432+
433+
test("does not truncate isError results even when over the gate", async () => {
434+
const store = fakeBlobStore();
435+
const original = `Error: ${"x".repeat(MAX_RESULT_CHARS + 500)}`;
436+
const plugin = resultTruncationPlugin({ getBlobWriter: () => store.writeBlob });
437+
if (plugin.middleware === undefined) throw new Error("expected middleware");
438+
const middleware = plugin.middleware(async (call) => ({
439+
callId: call.id,
440+
content: original,
441+
isError: true,
442+
}));
443+
const result = await middleware(
444+
{ id: "call-err", name: "wait_agents", arguments: {} },
445+
new AbortController().signal,
446+
);
447+
expect(result.content).toBe(original);
448+
expect(result.isError).toBe(true);
449+
expect(store.blobs.size).toBe(0);
450+
});
451+
});
452+
453+
describe("wrapAgentToolResultTruncation", () => {
454+
test("spills oversized wait_agents JSON from a kind:full handler", async () => {
455+
const store = fakeBlobStore();
456+
const payload = { results: [{ report: "x".repeat(MAX_RESULT_CHARS + 500) }] };
457+
const minified = JSON.stringify(payload);
458+
expect(minified.length).toBeGreaterThan(MAX_RESULT_CHARS);
459+
const pretty = JSON.stringify(payload, null, 2);
460+
const wrapped = wrapAgentToolResultTruncation(
461+
{
462+
kind: "full",
463+
definition: {
464+
name: "wait_agents",
465+
description: "wait",
466+
inputSchema: { type: "object" },
467+
},
468+
handler: async (call) => ({ callId: call.id, content: minified }),
469+
},
470+
{ getBlobWriter: () => store.writeBlob },
471+
);
472+
if (wrapped.kind !== "full") throw new Error("expected full tool");
473+
const result = await wrapped.handler(
474+
{ id: "call-wrap-wait", name: "wait_agents", arguments: {} },
475+
new AbortController().signal,
476+
);
477+
expect(String(result.content).length).toBeLessThanOrEqual(MAX_RESULT_CHARS);
478+
const uri = `tool-output:///${spillBlobKey("call-wrap-wait")}`;
479+
expect(String(result.content)).toContain(uri);
480+
const recovered = new TextDecoder().decode(await createBlobReader(store).read(uri));
481+
expect(recovered).toBe(pretty);
482+
});
483+
484+
test("spills oversized search_agents string from a kind:string handler", async () => {
485+
const store = fakeBlobStore();
486+
const original = `Matching agent profiles:\n\n${"z".repeat(MAX_RESULT_CHARS + 500)}`;
487+
const wrapped = wrapAgentToolResultTruncation(
488+
{
489+
kind: "string",
490+
definition: {
491+
name: "search_agents",
492+
description: "search",
493+
inputSchema: { type: "object" },
494+
},
495+
handler: async () => original,
496+
},
497+
{ getBlobWriter: () => store.writeBlob },
498+
);
499+
if (wrapped.kind !== "full") throw new Error("expected full wrapper so spill can use callId");
500+
const result = await wrapped.handler(
501+
{ id: "call-wrap-search", name: "search_agents", arguments: {} },
502+
new AbortController().signal,
503+
);
504+
expect(String(result.content).length).toBeLessThanOrEqual(MAX_RESULT_CHARS);
505+
const uri = `tool-output:///${spillBlobKey("call-wrap-search")}`;
506+
expect(String(result.content)).toContain(uri);
507+
const recovered = new TextDecoder().decode(await createBlobReader(store).read(uri));
508+
expect(recovered).toBe(original);
509+
});
510+
511+
test("does not truncate isError results from a kind:full handler", async () => {
512+
const store = fakeBlobStore();
513+
const original = `Error: ${"e".repeat(MAX_RESULT_CHARS + 500)}`;
514+
const wrapped = wrapAgentToolResultTruncation(
515+
{
516+
kind: "full",
517+
definition: {
518+
name: "wait_agents",
519+
description: "wait",
520+
inputSchema: { type: "object" },
521+
},
522+
handler: async (call) => ({ callId: call.id, content: original, isError: true }),
523+
},
524+
{ getBlobWriter: () => store.writeBlob },
525+
);
526+
if (wrapped.kind !== "full") throw new Error("expected full tool");
527+
const result = await wrapped.handler(
528+
{ id: "call-wrap-err", name: "wait_agents", arguments: {} },
529+
new AbortController().signal,
530+
);
531+
expect(result.content).toBe(original);
532+
expect(result.isError).toBe(true);
533+
expect(store.blobs.size).toBe(0);
534+
});
376535
});
377536

378537
describe("scrub-before-spill", () => {

0 commit comments

Comments
 (0)