Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 32 additions & 1 deletion docs/decisions/0016-speaking-to-agents.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# 0016 — Speaking to agents: what the brain says when an agent connects, measured against the servers that do it best

**Status:** accepted (2026-10-07), with the amendments from the build at the end
**Status:** accepted (2026-10-07), with the amendments from the build at the end; amended (2026-10-09): the tools advertise no output schema

## Context

Expand Down Expand Up @@ -169,3 +169,34 @@ Built on 2026-10-07. Interaction functions (0010) reached main the same day, so
- **Verification.** The replay test calls the tools itself, with what each episode's answer needs, and checks that the served texts hold that answer; it does not run an agent. Episode 4 saves the recall function the `remember` recipe describes over the runs of a function that answers what it posted, runs that function with a scripted model reply and finds the recall function answers what was posted. Episode 6 pins the words of a refused key. Episode 7 is the evening's failure: a request is open through a chat channel, the person says "approve", the expected act is `answer_interaction` with `{"decision": "approve"}` on that request, which ends the workflow run and leaves no request open, and a new run only when the person asks for the next month.
- **Not built.** The recorded run of the asks against `pnpm dev` with a real model, for `docs/engineering/reference/mcp.md`, needs a person with a model provider's key.
- **Tools on a connection, 25 (2026-10-08).** [Decision 0018](0018-testing-a-tool-call.md#7-over-http-and-over-mcp) adds `test_tool_call`, a 25th tool for a key that may call them all, so the bound of §8 is 25; when it is reached, the tools group by endpoint or into a catalogue pair as Sentry's, as before, since a bound that moves by one for each tool is no bound. The instructions gain the clause "; test_tool_call shows what a tool answers" where the tool is listed, and take 1,990 characters on `/mcp`, 1,473 on `/orgs/{org}/mcp` and 1,950 on `/orgs/{org}/brains/{brain}/mcp`.

## Amendment (2026-10-09): the tools advertise no output schema

Decided against `origin/main` at 9203df9d.

### Context

Every operation tool the brain serves over MCP, 24 of the 25 on `/mcp`, carries an `outputSchema`, the operation's output schema made self-contained with its `$defs` and the 2020-12 dialect (`packages/api/src/tools/tool-definition.ts:55`; `tools/tool-schema.ts:69-72`), beside the `inputSchema`, the two text blocks and the `structuredContent` of its result (`tools/tool-result.ts:51-61`). Section 1 of this record counted them as part of the surface: "described schemas, `structuredContent` beside two text blocks".

On 2026-10-08 every call of `list_tool_servers` and of the run tools failed in a desktop client with "The connector returned an error or an invalid response", on a server whose log showed no error and whose answers were valid. The cause, read in the client's own code and its logs: an MCP client compiles a validator from each tool's `outputSchema` when it lists the tools, and checks every result's `structuredContent` against it, refusing the result when it does not match; the official client does this (`@modelcontextprotocol/sdk`, `Client.listTools` caching the output validators and `Client.callTool` throwing "Structured content does not match the tool's output schema"), and so does the desktop client, with a validator of its own. The client had listed the tools once, on 2026-10-07 at 17:49, and never again in twenty-one hours; the server was replaced by a newer main in between, through a bridge process that outlived the restart, so the client kept checking new results against old shapes. Two tools whose result shape had changed failed on every call until the client was restarted.

The protocol gives a server one way to say that its tools changed, `tools.listChanged` with `notifications/tools/list_changed`, which needs a session and a stream to the client. The brain serves MCP without sessions (`packages/api/src/mcp/mcp-routes.ts:80`, `legacy: 'stateless'`) and says `listChanged: false` (`mcp/mcp-connection.ts:43`), so a restart is invisible to a client behind a bridge, and no notification can reach it. The newer protocol revision lets a listing declare its lifetime, which the server package already fills as `ttlMs: 0`, but the clients in use connect through the older one. Our output schemas are closed, 150 structs with `additionalProperties: false` across the 24 tools, so a field added to any result breaks every connected client, not only a rename. The schemas also weigh 113,841 bytes of the `/mcp` listing's 152,633, 13,748 of the org endpoint's 21,905 and 102,882 of the brain endpoint's 133,578. No client of the brain consumes them: agents read the text blocks, the management plane calls the HTTP API, and the MCP reference describes the structured content, not a schema.

### Decision

- The served tools advertise **no `outputSchema`**. Each result keeps its first text block, the plain-language sentence, its second text block, the output as JSON text, and its `structuredContent`, which the protocol allows without a schema. `tool-definition.ts` builds no output schema, and `ToolDefinition` loses the field.
- The input schemas stay, since a client must know what a tool takes; a client holding an old input schema after an upgrade sends old arguments and is told by the brain's refusal what is wrong, a failure that is visible and ends with a restart, where a stale output schema failed opaquely.
- The brain keeps `listChanged: false` and its stateless serving; a schema that clients freeze per connection is not advertised until the brain can tell them when it changes: when it serves a listing with a declared lifetime that the clients in use honour, and the result shapes have stopped moving. Reintroducing the schema is then a decision of its own.
- The MCP reference says that a result carries the summary, the JSON text and the structured content, and no schema; this record's section 1 is read as amended.
- A client already connected before this change still holds the old listing with its schemas, and checks against them until it is restarted once more; the change ends the class of failure for every upgrade after it.

### Verification the build must include

- Every tool served on `/mcp`, the org endpoint and the brain endpoint carries an `inputSchema` and no `outputSchema`, over the real server.
- A tool's result still carries the two text blocks and the `structuredContent`, and an official client calling `list_tool_servers`, `list_brains` and `execute_spec` after listing the tools compiles no output validator and accepts each result.
- The listing's size: the three listings measured, each smaller than today by the schemas' bytes, pinned as an upper bound.
- The MCP reference and the api README say so; the internal-terms check.

### Self-check

Read on 2026-10-09 against `origin/main` at 9203df9d: `tool-definition.ts:55` builds the output schema, `tool-schema.ts:69-72` makes it self-contained, `tool-result.ts:51-61` builds the result, `mcp-routes.ts:80` serves stateless and `mcp-connection.ts:43` announces no changes. The client behaviour was read in the official SDK's client and in the desktop client's bundle, and the single listing in its logs. The counts, 150 closed structs and the schemas' bytes of each listing, were measured over the listings served at 9203df9d. The amendment names no client vendor in its decision, no platform and no customer.
2 changes: 1 addition & 1 deletion docs/decisions/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ Use a numbered Markdown filename for each decision and link to it from the affec
| [11. Waiting calls: a workflow waits for a run that finishes later, and a run can be cancelled](0011-waiting-calls.md) | accepted 2026-10-07 |
| [14. A warm worker pool: a run costs its own work, in a worker kept between jobs and let go of after any bad one](0014-warm-worker-pool.md) | proposed 2026-10-06, built 2026-10-07 |
| [15. Several triggers: a workflow starts on an event and on its schedules, each trigger kept and matched on its own](0015-several-triggers.md) | accepted 2026-10-08, amended 2026-10-08 |
| [16. Speaking to agents: what the brain says when an agent connects](0016-speaking-to-agents.md) | accepted 2026-10-07 |
| [16. Speaking to agents: what the brain says when an agent connects](0016-speaking-to-agents.md) | accepted 2026-10-07, amended 2026-10-09 |
| [18. Testing a tool call: an agent tries one tool of a tool server through the brain, as a run would call it](0018-testing-a-tool-call.md) | accepted 2026-10-08 |
| [19. The answer shape of an open request: the listing shows the schema each request recorded](0019-the-answer-shape-of-an-open-request.md) | accepted 2026-10-08, built 2026-10-08 |
| [20. Replies to what the brain sent: a function that delivers through a tool reads the replies to its message, and the brain takes its party's reply](0020-replies-through-a-channel.md) | proposed 2026-10-08, built 2026-10-08 |
4 changes: 2 additions & 2 deletions docs/engineering/reference/mcp.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,11 +26,11 @@ Most MCP clients take an entry of this shape. In Claude Code, `claude mcp add --
| `POST /orgs/{org}/mcp` | `create_brain`, `list_brains`, `get_brain`, `update_brain`, `retire_brain`, `list_models` and `list_tool_servers` | to manage one org’s brains and discover models and tool servers |
| `POST /orgs/{org}/brains/{brain}/mcp` | the tools inside a brain, acting in that brain, without a `brain` argument | to lock a connection to one brain, such as for one agent |

Every endpoint speaks streamable HTTP without sessions. It serves the current stateless revision (`2026-07-28`) and the earlier ones the SDK supports (`2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`), so agents built on older SDKs connect too. Each tool carries the operation's description and its input and output JSON Schemas, and is marked read-only when it only reads. A tool that cannot do what was asked returns `isError` with the same problem document HTTP would answer with, as text, so the agent can read the `reason` and the `detail`, and correct its arguments when the `reason` is `invalid_input`. The key's permissions and brains hold as they do over HTTP: a read-only key is offered `list_brains` and not `create_brain`, which it would be refused, a key limited to some brains gets `forbidden` for any other, and a brain the org does not have is `not_found`. [`packages/api`](https://github.com/BeOnAuto/auto-brain/blob/main/packages/api/README.md) describes the mappings in full.
Every endpoint speaks streamable HTTP without sessions. It serves the current stateless revision (`2026-07-28`) and the earlier ones the SDK supports (`2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`), so agents built on older SDKs connect too. Each tool carries the operation's description and the JSON Schema of its input, and no output schema, and is marked read-only when it only reads. A tool that cannot do what was asked returns `isError` with the same problem document HTTP would answer with, as text, so the agent can read the `reason` and the `detail`, and correct its arguments when the `reason` is `invalid_input`. The key's permissions and brains hold as they do over HTTP: a read-only key is offered `list_brains` and not `create_brain`, which it would be refused, a key limited to some brains gets `forbidden` for any other, and a brain the org does not have is `not_found`. [`packages/api`](https://github.com/BeOnAuto/auto-brain/blob/main/packages/api/README.md) describes the mappings in full.

## Reading tool results

On success, use `structuredContent` for the operation's output. The first text block is a human-readable summary; the second text block contains the same output as JSON for clients without structured-output support. Do not parse `content[0].text` as JSON.
On success, use `structuredContent` for the operation's output. The first text block is a plain-language summary; the second text block contains the same output as JSON text for clients that do not read structured content. No tool advertises an output schema, so a client takes `structuredContent` as it comes instead of checking it against one. Do not parse `content[0].text` as JSON.

On failure, the result has `isError: true` and no `structuredContent`. The first text block explains what could not be done, and the second contains the JSON problem document. Check its `reason` and `detail` before deciding whether to correct input or retry.

Expand Down
14 changes: 7 additions & 7 deletions docs/reference/mcp.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ Functions and workflows share the definition tools: those tools accept `inferenc
| Learn what a tool answers | `test_tool_call`, with the `server`, the `tool` and its `arguments` |
| Read a guide | `get_guide`, with the name of the guide |

Every tool supplies its description and input and output JSON Schemas. A description says what the tool does and when to use it; the rules of each argument are in its schema, and the format of a definition is in its guide. Each tool's annotations say whether it only reads, whether what it does cannot be undone, as retiring, cancelling and answering a request cannot, whether calling it again with the same input changes nothing, and whether it reaches outside the runtime. Brain-management and model-discovery tools are available at `/mcp` and the org endpoint; function, workflow, brain event and tool test tools are available at `/mcp` and the brain endpoint; `list_tool_servers` and `get_guide` are available on every endpoint.
Every tool supplies its description and the JSON Schema of its input, and no output schema; [Successful results](#successful-results) says what a result carries. A description says what the tool does and when to use it; the rules of each argument are in its schema, and the format of a definition is in its guide. Each tool's annotations say whether it only reads, whether what it does cannot be undone, as retiring, cancelling and answering a request cannot, whether calling it again with the same input changes nothing, and whether it reaches outside the runtime. Brain-management and model-discovery tools are available at `/mcp` and the org endpoint; function, workflow, brain event and tool test tools are available at `/mcp` and the brain endpoint; `list_tool_servers` and `get_guide` are available on every endpoint.

Definition operations identify a function or workflow by `primitive` and `name`. Creating or updating a definition takes its document as `source`. Running it accepts `input` and an optional UUID `execution_id`; inspecting a run requires `execution_id`. Cancelling a run requires its `execution_id` and takes an optional `reason`, 1 to 1,024 characters, which the run keeps as the detail of its ending; [Cancelling a run](http.md#cancelling-a-run) says which runs can be cancelled. Sending an event requires the run's `execution_id` and an `event` with a `type`; [HTTP workflows](http.md#workflows) lists its other fields and limits. Publishing an event to a brain requires an `event` with a `source` and a `type`; [Publishing events](http.md#publishing-events) lists its attributes and limits and what publishing it again returns. See [HTTP operations](http.md) for field limits and retry behavior.

Expand Down Expand Up @@ -92,13 +92,13 @@ This model catalog lists language models, not MCP tools; `list_tool_servers` lis

## Successful results

A successful tool result contains the operation output in `structuredContent`:
A successful tool result carries a plain-language summary, the operation's output as JSON text and the same output as structured content. No tool advertises an output schema, so a client takes `structuredContent` as it comes instead of checking it against one.

| Field | Contents |
| ------------------- | ---------------------------------------------------------------------- |
| `structuredContent` | The operation's structured output |
| `content[0].text` | A human-readable summary |
| `content[1].text` | The same output as JSON, for clients without structured-output support |
| Field | Contents |
| ------------------- | ------------------------------------------------------------------------------ |
| `content[0].text` | A plain-language summary of what was done |
| `content[1].text` | The operation's output as JSON text, for clients that do not read the next one |
| `structuredContent` | The same output as structured content |

Consumers should read `structuredContent`. The first text block is not JSON.

Expand Down
Loading
Loading