diff --git a/docs/decisions/0016-speaking-to-agents.md b/docs/decisions/0016-speaking-to-agents.md index a2a231eb3..e38c6e14d 100644 --- a/docs/decisions/0016-speaking-to-agents.md +++ b/docs/decisions/0016-speaking-to-agents.md @@ -1,6 +1,6 @@ # 0016 — Speaking to agents: what the brain says when an agent connects, measured against the servers that do it best -**Status:** accepted (2026-10-07), with the amendments from the build at the end +**Status:** accepted (2026-10-07), with the amendments from the build at the end; amended (2026-10-09): the tools advertise no output schema ## Context @@ -169,3 +169,34 @@ Built on 2026-10-07. Interaction functions (0010) reached main the same day, so - **Verification.** The replay test calls the tools itself, with what each episode's answer needs, and checks that the served texts hold that answer; it does not run an agent. Episode 4 saves the recall function the `remember` recipe describes over the runs of a function that answers what it posted, runs that function with a scripted model reply and finds the recall function answers what was posted. Episode 6 pins the words of a refused key. Episode 7 is the evening's failure: a request is open through a chat channel, the person says "approve", the expected act is `answer_interaction` with `{"decision": "approve"}` on that request, which ends the workflow run and leaves no request open, and a new run only when the person asks for the next month. - **Not built.** The recorded run of the asks against `pnpm dev` with a real model, for `docs/engineering/reference/mcp.md`, needs a person with a model provider's key. - **Tools on a connection, 25 (2026-10-08).** [Decision 0018](0018-testing-a-tool-call.md#7-over-http-and-over-mcp) adds `test_tool_call`, a 25th tool for a key that may call them all, so the bound of §8 is 25; when it is reached, the tools group by endpoint or into a catalogue pair as Sentry's, as before, since a bound that moves by one for each tool is no bound. The instructions gain the clause "; test_tool_call shows what a tool answers" where the tool is listed, and take 1,990 characters on `/mcp`, 1,473 on `/orgs/{org}/mcp` and 1,950 on `/orgs/{org}/brains/{brain}/mcp`. + +## Amendment (2026-10-09): the tools advertise no output schema + +Decided against `origin/main` at 9203df9d. + +### Context + +Every operation tool the brain serves over MCP, 24 of the 25 on `/mcp`, carries an `outputSchema`, the operation's output schema made self-contained with its `$defs` and the 2020-12 dialect (`packages/api/src/tools/tool-definition.ts:55`; `tools/tool-schema.ts:69-72`), beside the `inputSchema`, the two text blocks and the `structuredContent` of its result (`tools/tool-result.ts:51-61`). Section 1 of this record counted them as part of the surface: "described schemas, `structuredContent` beside two text blocks". + +On 2026-10-08 every call of `list_tool_servers` and of the run tools failed in a desktop client with "The connector returned an error or an invalid response", on a server whose log showed no error and whose answers were valid. The cause, read in the client's own code and its logs: an MCP client compiles a validator from each tool's `outputSchema` when it lists the tools, and checks every result's `structuredContent` against it, refusing the result when it does not match; the official client does this (`@modelcontextprotocol/sdk`, `Client.listTools` caching the output validators and `Client.callTool` throwing "Structured content does not match the tool's output schema"), and so does the desktop client, with a validator of its own. The client had listed the tools once, on 2026-10-07 at 17:49, and never again in twenty-one hours; the server was replaced by a newer main in between, through a bridge process that outlived the restart, so the client kept checking new results against old shapes. Two tools whose result shape had changed failed on every call until the client was restarted. + +The protocol gives a server one way to say that its tools changed, `tools.listChanged` with `notifications/tools/list_changed`, which needs a session and a stream to the client. The brain serves MCP without sessions (`packages/api/src/mcp/mcp-routes.ts:80`, `legacy: 'stateless'`) and says `listChanged: false` (`mcp/mcp-connection.ts:43`), so a restart is invisible to a client behind a bridge, and no notification can reach it. The newer protocol revision lets a listing declare its lifetime, which the server package already fills as `ttlMs: 0`, but the clients in use connect through the older one. Our output schemas are closed, 150 structs with `additionalProperties: false` across the 24 tools, so a field added to any result breaks every connected client, not only a rename. The schemas also weigh 113,841 bytes of the `/mcp` listing's 152,633, 13,748 of the org endpoint's 21,905 and 102,882 of the brain endpoint's 133,578. No client of the brain consumes them: agents read the text blocks, the management plane calls the HTTP API, and the MCP reference describes the structured content, not a schema. + +### Decision + +- The served tools advertise **no `outputSchema`**. Each result keeps its first text block, the plain-language sentence, its second text block, the output as JSON text, and its `structuredContent`, which the protocol allows without a schema. `tool-definition.ts` builds no output schema, and `ToolDefinition` loses the field. +- The input schemas stay, since a client must know what a tool takes; a client holding an old input schema after an upgrade sends old arguments and is told by the brain's refusal what is wrong, a failure that is visible and ends with a restart, where a stale output schema failed opaquely. +- The brain keeps `listChanged: false` and its stateless serving; a schema that clients freeze per connection is not advertised until the brain can tell them when it changes: when it serves a listing with a declared lifetime that the clients in use honour, and the result shapes have stopped moving. Reintroducing the schema is then a decision of its own. +- The MCP reference says that a result carries the summary, the JSON text and the structured content, and no schema; this record's section 1 is read as amended. +- A client already connected before this change still holds the old listing with its schemas, and checks against them until it is restarted once more; the change ends the class of failure for every upgrade after it. + +### Verification the build must include + +- Every tool served on `/mcp`, the org endpoint and the brain endpoint carries an `inputSchema` and no `outputSchema`, over the real server. +- A tool's result still carries the two text blocks and the `structuredContent`, and an official client calling `list_tool_servers`, `list_brains` and `execute_spec` after listing the tools compiles no output validator and accepts each result. +- The listing's size: the three listings measured, each smaller than today by the schemas' bytes, pinned as an upper bound. +- The MCP reference and the api README say so; the internal-terms check. + +### Self-check + +Read on 2026-10-09 against `origin/main` at 9203df9d: `tool-definition.ts:55` builds the output schema, `tool-schema.ts:69-72` makes it self-contained, `tool-result.ts:51-61` builds the result, `mcp-routes.ts:80` serves stateless and `mcp-connection.ts:43` announces no changes. The client behaviour was read in the official SDK's client and in the desktop client's bundle, and the single listing in its logs. The counts, 150 closed structs and the schemas' bytes of each listing, were measured over the listings served at 9203df9d. The amendment names no client vendor in its decision, no platform and no customer. diff --git a/docs/decisions/README.md b/docs/decisions/README.md index e70e1a9d2..8e931a0a9 100644 --- a/docs/decisions/README.md +++ b/docs/decisions/README.md @@ -18,7 +18,7 @@ Use a numbered Markdown filename for each decision and link to it from the affec | [11. Waiting calls: a workflow waits for a run that finishes later, and a run can be cancelled](0011-waiting-calls.md) | accepted 2026-10-07 | | [14. A warm worker pool: a run costs its own work, in a worker kept between jobs and let go of after any bad one](0014-warm-worker-pool.md) | proposed 2026-10-06, built 2026-10-07 | | [15. Several triggers: a workflow starts on an event and on its schedules, each trigger kept and matched on its own](0015-several-triggers.md) | accepted 2026-10-08, amended 2026-10-08 | -| [16. Speaking to agents: what the brain says when an agent connects](0016-speaking-to-agents.md) | accepted 2026-10-07 | +| [16. Speaking to agents: what the brain says when an agent connects](0016-speaking-to-agents.md) | accepted 2026-10-07, amended 2026-10-09 | | [18. Testing a tool call: an agent tries one tool of a tool server through the brain, as a run would call it](0018-testing-a-tool-call.md) | accepted 2026-10-08 | | [19. The answer shape of an open request: the listing shows the schema each request recorded](0019-the-answer-shape-of-an-open-request.md) | accepted 2026-10-08, built 2026-10-08 | | [20. Replies to what the brain sent: a function that delivers through a tool reads the replies to its message, and the brain takes its party's reply](0020-replies-through-a-channel.md) | proposed 2026-10-08, built 2026-10-08 | diff --git a/docs/engineering/reference/mcp.md b/docs/engineering/reference/mcp.md index 8188c1794..640ea5840 100644 --- a/docs/engineering/reference/mcp.md +++ b/docs/engineering/reference/mcp.md @@ -26,11 +26,11 @@ Most MCP clients take an entry of this shape. In Claude Code, `claude mcp add -- | `POST /orgs/{org}/mcp` | `create_brain`, `list_brains`, `get_brain`, `update_brain`, `retire_brain`, `list_models` and `list_tool_servers` | to manage one org’s brains and discover models and tool servers | | `POST /orgs/{org}/brains/{brain}/mcp` | the tools inside a brain, acting in that brain, without a `brain` argument | to lock a connection to one brain, such as for one agent | -Every endpoint speaks streamable HTTP without sessions. It serves the current stateless revision (`2026-07-28`) and the earlier ones the SDK supports (`2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`), so agents built on older SDKs connect too. Each tool carries the operation's description and its input and output JSON Schemas, and is marked read-only when it only reads. A tool that cannot do what was asked returns `isError` with the same problem document HTTP would answer with, as text, so the agent can read the `reason` and the `detail`, and correct its arguments when the `reason` is `invalid_input`. The key's permissions and brains hold as they do over HTTP: a read-only key is offered `list_brains` and not `create_brain`, which it would be refused, a key limited to some brains gets `forbidden` for any other, and a brain the org does not have is `not_found`. [`packages/api`](https://github.com/BeOnAuto/auto-brain/blob/main/packages/api/README.md) describes the mappings in full. +Every endpoint speaks streamable HTTP without sessions. It serves the current stateless revision (`2026-07-28`) and the earlier ones the SDK supports (`2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`), so agents built on older SDKs connect too. Each tool carries the operation's description and the JSON Schema of its input, and no output schema, and is marked read-only when it only reads. A tool that cannot do what was asked returns `isError` with the same problem document HTTP would answer with, as text, so the agent can read the `reason` and the `detail`, and correct its arguments when the `reason` is `invalid_input`. The key's permissions and brains hold as they do over HTTP: a read-only key is offered `list_brains` and not `create_brain`, which it would be refused, a key limited to some brains gets `forbidden` for any other, and a brain the org does not have is `not_found`. [`packages/api`](https://github.com/BeOnAuto/auto-brain/blob/main/packages/api/README.md) describes the mappings in full. ## Reading tool results -On success, use `structuredContent` for the operation's output. The first text block is a human-readable summary; the second text block contains the same output as JSON for clients without structured-output support. Do not parse `content[0].text` as JSON. +On success, use `structuredContent` for the operation's output. The first text block is a plain-language summary; the second text block contains the same output as JSON text for clients that do not read structured content. No tool advertises an output schema, so a client takes `structuredContent` as it comes instead of checking it against one. Do not parse `content[0].text` as JSON. On failure, the result has `isError: true` and no `structuredContent`. The first text block explains what could not be done, and the second contains the JSON problem document. Check its `reason` and `detail` before deciding whether to correct input or retry. diff --git a/docs/reference/mcp.md b/docs/reference/mcp.md index 1d2084eb4..ad0008ac6 100644 --- a/docs/reference/mcp.md +++ b/docs/reference/mcp.md @@ -52,7 +52,7 @@ Functions and workflows share the definition tools: those tools accept `inferenc | Learn what a tool answers | `test_tool_call`, with the `server`, the `tool` and its `arguments` | | Read a guide | `get_guide`, with the name of the guide | -Every tool supplies its description and input and output JSON Schemas. A description says what the tool does and when to use it; the rules of each argument are in its schema, and the format of a definition is in its guide. Each tool's annotations say whether it only reads, whether what it does cannot be undone, as retiring, cancelling and answering a request cannot, whether calling it again with the same input changes nothing, and whether it reaches outside the runtime. Brain-management and model-discovery tools are available at `/mcp` and the org endpoint; function, workflow, brain event and tool test tools are available at `/mcp` and the brain endpoint; `list_tool_servers` and `get_guide` are available on every endpoint. +Every tool supplies its description and the JSON Schema of its input, and no output schema; [Successful results](#successful-results) says what a result carries. A description says what the tool does and when to use it; the rules of each argument are in its schema, and the format of a definition is in its guide. Each tool's annotations say whether it only reads, whether what it does cannot be undone, as retiring, cancelling and answering a request cannot, whether calling it again with the same input changes nothing, and whether it reaches outside the runtime. Brain-management and model-discovery tools are available at `/mcp` and the org endpoint; function, workflow, brain event and tool test tools are available at `/mcp` and the brain endpoint; `list_tool_servers` and `get_guide` are available on every endpoint. Definition operations identify a function or workflow by `primitive` and `name`. Creating or updating a definition takes its document as `source`. Running it accepts `input` and an optional UUID `execution_id`; inspecting a run requires `execution_id`. Cancelling a run requires its `execution_id` and takes an optional `reason`, 1 to 1,024 characters, which the run keeps as the detail of its ending; [Cancelling a run](http.md#cancelling-a-run) says which runs can be cancelled. Sending an event requires the run's `execution_id` and an `event` with a `type`; [HTTP workflows](http.md#workflows) lists its other fields and limits. Publishing an event to a brain requires an `event` with a `source` and a `type`; [Publishing events](http.md#publishing-events) lists its attributes and limits and what publishing it again returns. See [HTTP operations](http.md) for field limits and retry behavior. @@ -92,13 +92,13 @@ This model catalog lists language models, not MCP tools; `list_tool_servers` lis ## Successful results -A successful tool result contains the operation output in `structuredContent`: +A successful tool result carries a plain-language summary, the operation's output as JSON text and the same output as structured content. No tool advertises an output schema, so a client takes `structuredContent` as it comes instead of checking it against one. -| Field | Contents | -| ------------------- | ---------------------------------------------------------------------- | -| `structuredContent` | The operation's structured output | -| `content[0].text` | A human-readable summary | -| `content[1].text` | The same output as JSON, for clients without structured-output support | +| Field | Contents | +| ------------------- | ------------------------------------------------------------------------------ | +| `content[0].text` | A plain-language summary of what was done | +| `content[1].text` | The operation's output as JSON text, for clients that do not read the next one | +| `structuredContent` | The same output as structured content | Consumers should read `structuredContent`. The first text block is not JSON. diff --git a/packages/api/README.md b/packages/api/README.md index 5eca3b0bd..2626c9eaa 100644 --- a/packages/api/README.md +++ b/packages/api/README.md @@ -35,7 +35,7 @@ A `403` `forbidden` also carries `WWW-Authenticate: Bearer error="insufficient_s ## Over MCP -`mcpRoutes` serves the same catalog to agents over the [Model Context Protocol](https://modelcontextprotocol.io), with the official SDK, `@modelcontextprotocol/server`, pinned exactly. It is stateless: each request builds a fresh MCP server for its caller and keeps no session. It serves the current stateless revision, `2026-07-28`, and the earlier ones the SDK supports: `2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`. An `initialize` in any of them is answered in that revision, with the same tools, results and errors in each, and a client that asks for a revision the SDK does not know is offered `2025-11-25`. So agents built on older SDKs, including the 1.x line of the official one, connect out of the box. Revisions before `2025-06-18` predate `structuredContent` and output schemas; the SDK sends them regardless, and a client of such a revision reads the same JSON from the text content. MCP sits outside `/v1` because the protocol versions itself. +`mcpRoutes` serves the same catalog to agents over the [Model Context Protocol](https://modelcontextprotocol.io), with the official SDK, `@modelcontextprotocol/server`, pinned exactly. It is stateless: each request builds a fresh MCP server for its caller and keeps no session. It serves the current stateless revision, `2026-07-28`, and the earlier ones the SDK supports: `2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`. An `initialize` in any of them is answered in that revision, with the same tools, results and errors in each, and a client that asks for a revision the SDK does not know is offered `2025-11-25`. So agents built on older SDKs, including the 1.x line of the official one, connect out of the box. Revisions before `2025-06-18` predate `structuredContent`; the SDK sends it regardless, and a client of such a revision reads the same JSON from the text content. MCP sits outside `/v1` because the protocol versions itself. | Endpoint | Tools | Org and brain | | ------------------------------------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------- | @@ -53,7 +53,7 @@ All three sit behind the same chain as every other path. Before the SDK runs, a **The brain argument.** On `/mcp`, each brain operation's tool is derived from the catalog by its scope, so a new brain operation appears there without new code. Its input schema is the operation's with a required `brain` first, the brain id's pattern and the description "The id of the brain to act in" (added to every member when the input is a union, with any definition no longer referred to removed), and the tool's description is the operation's, since the argument's own description says what it is. A call first decodes `brain` with the brain id schema: a call without it, or with one that is not a well-formed id, gets `invalid_input` pointing at `/brain` before anything else, a format check that reveals nothing. It then takes `brain` out of the arguments and dispatches the rest to that brain, so the dispatcher's authorization holds per call: a key limited to other brains gets `forbidden`, and a brain the org does not have `not_found`, as `isError` tool results. No brain operation can have its own `brain` field: `@beonauto/operations` refuses to define one, as it reserves `org` and `brain` for the scope. A brain operation whose name an org operation also has is not offered on `/mcp`: the org operation, which the catalog accepts under that name only when it takes `brain` itself, answers there in its place, with its own `brain` argument and the description it gives it, as `list_tool_servers` answers for the whole org without a brain and for one brain with it. The brain endpoint offers the brain operation and the org endpoint the org one, under the same name. -**Tools.** Each operation is one tool: its `name`, `title` and `description` are the operation's, and its `inputSchema` and `outputSchema` are the operation's JSON Schemas (draft 2020-12), each self-contained with an object at the root and its definitions under `$defs`. A description says what the tool does, when to use it and when not, the caveats to know before calling and the alternative tool by name, in three to eight sentences and at most 800 characters; the rules of each argument are in its schema's description, at most 300 characters at any depth of the input, and keywords. The annotations follow what the operation declares: +**Tools.** Each operation is one tool: its `name`, `title` and `description` are the operation's, and its `inputSchema` is the operation's input JSON Schema (draft 2020-12), self-contained with an object at the root and its definitions under `$defs`. A tool has no `outputSchema`: a client compiles a validator from each output schema it lists and refuses a result that does not match it, and a server that keeps no session and announces no change to its tools cannot tell a client that listed them before an upgrade that a result's shape changed, so the result alone says what it holds ([decision 0016](../../docs/decisions/0016-speaking-to-agents.md#amendment-2026-10-09-the-tools-advertise-no-output-schema)). A description says what the tool does, when to use it and when not, the caveats to know before calling and the alternative tool by name, in three to eight sentences and at most 800 characters; the rules of each argument are in its schema's description, at most 300 characters at any depth of the input, and keywords. The annotations follow what the operation declares: | Annotation | Value | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | @@ -96,10 +96,10 @@ The HTTP method decides nothing. The server serves these, which `packages/server **Calls.** On a scoped endpoint the org, and the brain of a brain endpoint, come from the URL, and the arguments are the whole input; on `/mcp` the org is the caller's own and the brain an argument. The input is decoded with the `json` encoding. A call goes through the same dispatcher and the same `settle` as an HTTP request, with the caller in the URL's org. -| Outcome | Tool result | -| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| succeeded | `structuredContent` is the output; the first text content says in plain words what happened, and the second is the same output as JSON | -| rejected, failed, cancelled | `isError: true`; the first text content says in plain words what could not be done, and the second is the problem document HTTP would answer with, as JSON; no `structuredContent` | +| Outcome | Tool result | +| --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| succeeded | the first text content says in plain words what happened, the second is the output as JSON text, and `structuredContent` is the same output, with no output schema to check it against | +| rejected, failed, cancelled | `isError: true`; the first text content says in plain words what could not be done, and the second is the problem document HTTP would answer with, as JSON; no `structuredContent` | Human-readable results name reasoning functions, workflows and runs. They omit ids, versions, formats and status codes, which remain in the structured result. The tools take `primitive: inference` for reasoning functions and `primitive: orchestration` for workflows. `get_guide` serves the format of each definition type; no tool description carries one. The words of an outcome take at most 400 characters and those of a refusal at most 600: `withinCharacters` of `@beonauto/operations` keeps the whole sentences that fit and says the rest is in the details below, the JSON beside them. @@ -145,7 +145,7 @@ The SDK reports the errors it answers with JSON-RPC to `reportError`, which the ## The list of models -The server serves `list_models`, an org query of [`@beonauto/inference`](../../primitives/inference/README.md#listing-the-models), like any other operation: `GET /v1/orgs/{org}/models`, with an optional `provider` in the query string, and the read-only tool `list_models` on `/mcp` and `/orgs/{org}/mcp`, whose output schema is self-contained. It answers in the shape of the OpenAI API's list of models, which gateways such as LiteLLM, Portkey and Vercel's serve too: +The server serves `list_models`, an org query of [`@beonauto/inference`](../../primitives/inference/README.md#listing-the-models), like any other operation: `GET /v1/orgs/{org}/models`, with an optional `provider` in the query string, and the read-only tool `list_models` on `/mcp` and `/orgs/{org}/mcp`. It answers in the shape of the OpenAI API's list of models, which gateways such as LiteLLM, Portkey and Vercel's serve too: ```json { @@ -191,7 +191,7 @@ A `RegisterRoutes` function receives `routes.add`, to add a route, and `routes.o `makeAppRuntime(layer)` builds the runtime every call runs in. Its `run(effect, signal?)` answers the effect's value, or `cancelled` when the effect was interrupted, by the runtime's disposal or by the abort of the signal given, which interrupts it so its finalizers run; the server gives the signal of the call a workflow performs, so a call the workflow cancels stops the execution it started. -`src/index.ts` is the entry point: `createApiHandler`, `makeAppRuntime`, `operationRoutes`, `mcpRoutes`, the instructions, the types of a guide and a recipe, the guide's address and media type, the bounds, and their types. `@beonauto/api/testing` exports what the server's tests share: real MCP clients of the current SDK, on either revision, and of the SDK's 1.x line, helpers that read a tool listing, among them `takingBrain`, which gives a brain endpoint's tool the `brain` argument it has on `/mcp`, `operationToolsIn`, the listed tools but `get_guide`, and `schemasOf`, and the `wait_forever` operation, which never finishes on its own. +`src/index.ts` is the entry point: `createApiHandler`, `makeAppRuntime`, `operationRoutes`, `mcpRoutes`, the instructions, the types of a guide and a recipe, the guide's address and media type, the bounds, and their types. `@beonauto/api/testing` exports what the server's tests share: real MCP clients of the current SDK, on either revision, and of the SDK's 1.x line, each of which records in `compiledOutputSchemas` the output schemas it compiles a validator from and refuses every result it checks against one; helpers that read a tool listing, among them `takingBrain`, which gives a brain endpoint's tool the `brain` argument it has on `/mcp`, and `operationToolsIn`, which answers the listed tools but `get_guide`; and the `wait_forever` operation, which never finishes on its own. ## Source diff --git a/packages/api/src/guides/guide-tool.ts b/packages/api/src/guides/guide-tool.ts index c2d84d2bc..7c69c9570 100644 --- a/packages/api/src/guides/guide-tool.ts +++ b/packages/api/src/guides/guide-tool.ts @@ -1,20 +1,14 @@ import { rejected, type Issue } from '@beonauto/operations'; -import type { CallToolResult, McpServer, StandardSchemaWithJSON, ToolAnnotations } from '@modelcontextprotocol/server'; +import type { CallToolResult, McpServer } from '@modelcontextprotocol/server'; import { Result, Schema, SchemaIssue, type SchemaAST } from 'effect'; +import type { ToolDefinition } from '../tools/tool-definition.ts'; import { unsuccessfulResultOf } from '../tools/tool-result.ts'; import { advertisedSchema } from '../tools/tool-schema.ts'; import type { GuideShelf } from './guide-shelf.ts'; export const guideToolName = 'get_guide'; -interface GuideToolDefinition { - readonly title: string; - readonly description: string; - readonly inputSchema: StandardSchemaWithJSON; - readonly annotations: ToolAnnotations; -} - const description = [ "Reads one of the brain's guides: what its words mean, how each kind of definition is written, with its format, examples and bounds, and how the common tasks are done.", 'A format guide is read before a definition of that type is written, and a recipe before the task it names.', @@ -28,7 +22,7 @@ const strictly: SchemaAST.ParseOptions = { onExcessProperty: 'error', errors: 'a const failureOf = SchemaIssue.makeFormatterStandardSchemaV1(); -function guideToolDefinitionOf({ everyGuide }: GuideShelf): GuideToolDefinition { +function guideToolDefinitionOf({ everyGuide }: GuideShelf): ToolDefinition { return { title: 'Get guide', description, diff --git a/packages/api/src/guides/guides.test.ts b/packages/api/src/guides/guides.test.ts index ce9d5a19e..484edf36f 100644 --- a/packages/api/src/guides/guides.test.ts +++ b/packages/api/src/guides/guides.test.ts @@ -72,7 +72,6 @@ describe('the guide tool', () => { additionalProperties: false, }, }); - expect(guideTool?.outputSchema).toBeUndefined(); }, ); }); diff --git a/packages/api/src/mcp/mcp-clients.test.ts b/packages/api/src/mcp/mcp-clients.test.ts index d8fa94739..633c141a0 100644 --- a/packages/api/src/mcp/mcp-clients.test.ts +++ b/packages/api/src/mcp/mcp-clients.test.ts @@ -1,6 +1,8 @@ +import { McpServer, createMcpHandler } from '@modelcontextprotocol/server'; import { afterAll, beforeAll, describe, expect, it } from 'vitest'; -import { instructionsFor } from '../index.ts'; +import { instructionsFor, type RegisterRoutes } from '../index.ts'; +import { createTestHandler } from '../testing/api-calls.ts'; import { testDefinitionTypes, testRecipes } from '../testing/guides.ts'; import { listenOnLoopback, type Listening } from '../testing/listening.ts'; import { @@ -12,18 +14,43 @@ import { type McpConnection, } from '../testing/mcp-clients.ts'; import { acmeAdmin, operationServer, testServerInfo, type OperationServer } from '../testing/operation-server.ts'; -import { outputConformsTo, toolNamesIn } from '../testing/tool-listing.ts'; +import { toolNamesIn } from '../testing/tool-listing.ts'; +import { advertisedSchema } from '../tools/tool-schema.ts'; + +const countSchema = { type: 'object', properties: { count: { type: 'integer' } }, required: ['count'] }; + +function countingServer(): McpServer { + const counting = new McpServer({ name: 'counting', version: '1.0.0' }); + counting.registerTool( + 'count', + { + description: 'Counts to one.', + inputSchema: advertisedSchema({ type: 'object' }), + outputSchema: advertisedSchema(countSchema), + }, + () => ({ content: [{ type: 'text', text: '{"count":1}' }], structuredContent: { count: 1 } }), + ); + return counting; +} + +const advertisingAnOutputSchema: RegisterRoutes = (routes) => { + const counting = createMcpHandler(countingServer, { legacy: 'stateless' }); + routes.add('POST', '/counting/mcp', (c) => counting.fetch(c.req.raw)); + routes.onClose(() => counting.close()); +}; let server: OperationServer; let listening: Listening; +let advertising: Listening; beforeAll(async () => { server = await operationServer(); listening = await listenOnLoopback(server.handler); + advertising = await listenOnLoopback(createTestHandler({ routes: [advertisingAnOutputSchema] }).handler); }); afterAll(async () => { - await listening.close(); + await Promise.all([listening.close(), advertising.close()]); await server.runtime.dispose(); }); @@ -79,14 +106,17 @@ describe.each(mcpClientKinds)('the %s client connecting to a brain endpoint', (k }); describe.each(mcpClientKinds)('the %s client calling tools that succeed', (kind) => { - it('calls a command and a query, whose structured content conforms to the output schema', async () => { + it('lists the tools, then calls a command and a query and accepts each result without compiling an output validator', async () => { const name = `note-${kind.replaceAll(' ', '-')}`; - const outcome = await withMcpSession(kind, alphaEndpoint(), async (session) => ({ - listing: await session.listTools(), - added: await session.callTool('add_note', { name, text: 'hello' }), - read: await session.callTool('get_note', { name }), - })); + const outcome = await withMcpSession(kind, alphaEndpoint(), async (session) => { + await session.listTools(); + return { + added: await session.callTool('add_note', { name, text: 'hello' }), + read: await session.callTool('get_note', { name }), + compiled: session.compiledOutputSchemas(), + }; + }); expect(outcome.added).toEqual({ content: [ @@ -96,7 +126,25 @@ describe.each(mcpClientKinds)('the %s client calling tools that succeed', (kind) structuredContent: { name, text: 'hello' }, }); expect(outcome.read.structuredContent).toEqual({ name, text: 'hello' }); - expect(outputConformsTo(outcome.listing, 'get_note', outcome.read.structuredContent)).toBe(true); + expect(outcome.compiled).toEqual([]); + }); +}); + +describe.each(mcpClientKinds)('the %s client listing a tool with an output schema', (kind) => { + it('records the output schema it compiles a validator from, and refuses the result checked against it', async () => { + const compiled = await withMcpSession( + kind, + { url: `${advertising.origin}/counting/mcp`, headers: {} }, + async (session) => { + await session.listTools(); + await expect(session.callTool('count', {})).rejects.toThrow( + "Structured content does not match the tool's output schema: A result was checked against an advertised output schema", + ); + return session.compiledOutputSchemas(); + }, + ); + + expect(compiled).toMatchObject([countSchema]); }); }); diff --git a/packages/api/src/mcp/mcp-own-org.test.ts b/packages/api/src/mcp/mcp-own-org.test.ts index 3c9ee7264..832b17cf2 100644 --- a/packages/api/src/mcp/mcp-own-org.test.ts +++ b/packages/api/src/mcp/mcp-own-org.test.ts @@ -14,14 +14,7 @@ import { type OperationServer, } from '../testing/operation-server.ts'; import { danglingReferencesIn } from '../testing/self-contained.ts'; -import { - guideToolName, - listedTools, - schemasOf, - takingBrain, - toolNamesIn, - type ListedTool, -} from '../testing/tool-listing.ts'; +import { guideToolName, listedTools, takingBrain, toolNamesIn, type ListedTool } from '../testing/tool-listing.ts'; import { instructionsFor } from './instructions.ts'; const orgTools = ['label_brain', 'list_labels']; @@ -85,11 +78,13 @@ describe('the tools of /mcp', () => { expect([org.at(-1), brain.at(-1)]).toEqual([guide, guide]); }); - it('have self-contained schemas with an object root', async () => { - const schemas = schemasOf(listedTools(await asKey(acmeAdmin.key, (session) => session.listTools()))); + it('have self-contained input schemas with an object root, and no output schema', async () => { + const tools = listedTools(await asKey(acmeAdmin.key, (session) => session.listTools())); + const schemas = tools.map(({ inputSchema }) => inputSchema); expect(schemas.map((schema) => schema['type'])).toEqual(schemas.map(() => 'object')); expect(schemas.flatMap((schema) => danglingReferencesIn(schema))).toEqual([]); + expect(tools.filter(({ outputSchema }) => outputSchema !== undefined)).toEqual([]); }); it("carry the instructions of the caller's own org", async () => { diff --git a/packages/api/src/testing/index.ts b/packages/api/src/testing/index.ts index 19587cab2..314eb2adb 100644 --- a/packages/api/src/testing/index.ts +++ b/packages/api/src/testing/index.ts @@ -18,8 +18,6 @@ export { guideToolName, listedTools, operationToolsIn, - outputConformsTo, - schemasOf, takingBrain, toolNamesIn, type ListedTool, diff --git a/packages/api/src/testing/mcp-clients.ts b/packages/api/src/testing/mcp-clients.ts index dc03173f0..8aa5fec2b 100644 --- a/packages/api/src/testing/mcp-clients.ts +++ b/packages/api/src/testing/mcp-clients.ts @@ -1,4 +1,4 @@ -import { Client, StreamableHTTPClientTransport } from '@modelcontextprotocol/client'; +import { Client, StreamableHTTPClientTransport, type jsonSchemaValidator } from '@modelcontextprotocol/client'; import { Client as PreviousMajorClient } from '@modelcontextprotocol/sdk/client/index.js'; import { StreamableHTTPClientTransport as PreviousMajorTransport } from '@modelcontextprotocol/sdk/client/streamableHttp.js'; import type { Transport } from '@modelcontextprotocol/sdk/shared/transport.js'; @@ -29,6 +29,7 @@ export interface McpSession { readonly readResource: (uri: string) => Promise; readonly listPrompts: () => Promise; readonly getPrompt: (name: string, words?: Readonly>) => Promise; + readonly compiledOutputSchemas: () => readonly unknown[]; readonly close: () => Promise; } @@ -73,6 +74,28 @@ function wordsOf(words: Readonly> | undefined): Record readonly unknown[]; +} + +function recordingValidator(): RecordingValidator { + const compiled: unknown[] = []; + return { + validator: { + getValidator: (schema: unknown) => { + compiled.push(schema); + return () => ({ + valid: false, + data: undefined, + errorMessage: 'A result was checked against an advertised output schema', + }); + }, + }, + compiled: () => [...compiled], + }; +} + class PreviousMajorHttpTransport implements Transport { onclose?: NonNullable; onerror?: NonNullable; @@ -102,7 +125,11 @@ class PreviousMajorHttpTransport implements Transport { } async function connectPreviousMajor(connection: McpConnection): Promise { - const client = new PreviousMajorClient({ name: 'auto-brain-tests', version: '1.0.0' }); + const { validator, compiled } = recordingValidator(); + const client = new PreviousMajorClient( + { name: 'auto-brain-tests', version: '1.0.0' }, + { jsonSchemaValidator: validator }, + ); const transport = new PreviousMajorHttpTransport(connection); await client.connect(transport); return { @@ -115,13 +142,18 @@ async function connectPreviousMajor(connection: McpConnection): Promise client.readResource({ uri }), listPrompts: () => client.listPrompts(), getPrompt: (name, words) => client.getPrompt({ name, arguments: wordsOf(words) }), + compiledOutputSchemas: compiled, close: () => client.close(), }; } async function connectCurrentMajor(kind: McpClientKind, { url, headers }: McpConnection): Promise { const versionNegotiation = kind === 'current revision' ? { mode: { pin: '2026-07-28' } } : {}; - const client = new Client({ name: 'auto-brain-tests', version: '2.0.0' }, { versionNegotiation }); + const { validator, compiled } = recordingValidator(); + const client = new Client( + { name: 'auto-brain-tests', version: '2.0.0' }, + { versionNegotiation, jsonSchemaValidator: validator }, + ); await client.connect(new StreamableHTTPClientTransport(new URL(url), { requestInit: { headers: { ...headers } } })); return { protocolVersion: client.getNegotiatedProtocolVersion(), @@ -133,6 +165,7 @@ async function connectCurrentMajor(kind: McpClientKind, { url, headers }: McpCon readResource: (uri) => client.readResource({ uri }), listPrompts: () => client.listPrompts(), getPrompt: (name, words) => client.getPrompt({ name, arguments: wordsOf(words) }), + compiledOutputSchemas: compiled, close: () => client.close(), }; } diff --git a/packages/api/src/testing/notebook.ts b/packages/api/src/testing/notebook.ts index 24bbd3e82..812c187ab 100644 --- a/packages/api/src/testing/notebook.ts +++ b/packages/api/src/testing/notebook.ts @@ -196,13 +196,15 @@ const listLabels = defineQuery('org', { plainLanguage: plainly('list the labels'), }); +const Line = Schema.String.annotate({ identifier: 'Line' }); + const checkLines = defineCommand('brain', { name: 'check_lines', title: 'Check lines', description: 'Accepts lines that all start with a capital letter, and rejects every other line. Use it to check lines. `lines` are the lines.', route: { method: 'POST', path: '/lines' }, - inputSchema: Schema.Struct({ lines: Schema.Array(Schema.String) }), + inputSchema: Schema.Struct({ lines: Schema.Array(Line) }), outputSchema: Schema.Struct({ accepted: Schema.Int }), reasons: ['invalid_input'], handle: ({ lines }) => { diff --git a/packages/api/src/testing/tool-listing.ts b/packages/api/src/testing/tool-listing.ts index 082d69271..c4aba21e9 100644 --- a/packages/api/src/testing/tool-listing.ts +++ b/packages/api/src/testing/tool-listing.ts @@ -1,4 +1,3 @@ -import { AjvJsonSchemaValidator } from '@modelcontextprotocol/server/validators/ajv'; import { Schema } from 'effect'; import { withBrainArgument } from '../tools/brain-argument.ts'; @@ -15,8 +14,6 @@ export type ListedTool = typeof ListedToolSchema.Type; const toolsOf = Schema.decodeUnknownSync(Schema.Struct({ tools: Schema.Array(ListedToolSchema) })); -const validator = new AjvJsonSchemaValidator(); - export function listedTools(listing: unknown): readonly ListedTool[] { return toolsOf(listing).tools; } @@ -34,15 +31,3 @@ export function operationToolsIn(listing: unknown): readonly ListedTool[] { export function takingBrain({ inputSchema, ...tool }: ListedTool): ListedTool { return { ...tool, inputSchema: withBrainArgument(inputSchema) }; } - -export function schemasOf(tools: readonly ListedTool[]): readonly Readonly>[] { - return tools.flatMap(({ inputSchema, outputSchema }) => - outputSchema === undefined ? [inputSchema] : [inputSchema, outputSchema], - ); -} - -export function outputConformsTo(listing: unknown, name: string, value: unknown): boolean { - return listedTools(listing).some( - (tool) => tool.name === name && validator.getValidator({ ...tool.outputSchema })(value).valid, - ); -} diff --git a/packages/api/src/tools/mcp-tools.test.ts b/packages/api/src/tools/mcp-tools.test.ts index 1c4f73eb7..1d7ad86b3 100644 --- a/packages/api/src/tools/mcp-tools.test.ts +++ b/packages/api/src/tools/mcp-tools.test.ts @@ -7,7 +7,7 @@ import { withMcpSession, type McpConnection, type McpSession } from '../testing/ import { notebookOperations } from '../testing/notebook.ts'; import { acmeAdmin, operationServer, type OperationServer } from '../testing/operation-server.ts'; import { danglingReferencesIn } from '../testing/self-contained.ts'; -import { guideToolName, listedTools, operationToolsIn, schemasOf } from '../testing/tool-listing.ts'; +import { guideToolName, listedTools, operationToolsIn } from '../testing/tool-listing.ts'; const orgOperations = notebookOperations.filter(({ registration }) => registration.scope === 'org'); @@ -62,12 +62,13 @@ describe('the tools of each endpoint', () => { ]); }); - it('publishes self-contained schemas with an object root, and the primitive as a plain enum', async () => { + it('publishes self-contained input schemas with an object root, the primitive as a plain enum, and no output schema', async () => { const tools = listedTools(await onAlpha((session) => session.listTools())); - const schemas = schemasOf(tools); + const schemas = tools.map(({ inputSchema }) => inputSchema); expect(schemas.map((schema) => schema['type'])).toEqual(schemas.map(() => 'object')); expect(schemas.flatMap((schema) => danglingReferencesIn(schema))).toEqual([]); + expect(tools.filter(({ outputSchema }) => outputSchema !== undefined)).toEqual([]); expect(tools.find(({ name }) => name === 'create_spec')?.inputSchema).toMatchObject({ properties: { primitive: { type: 'string', enum: ['echo'] } }, }); diff --git a/packages/api/src/tools/tool-definition.ts b/packages/api/src/tools/tool-definition.ts index c20018067..e44d8f27a 100644 --- a/packages/api/src/tools/tool-definition.ts +++ b/packages/api/src/tools/tool-definition.ts @@ -15,7 +15,6 @@ export interface ToolDefinition { readonly title: string; readonly description: string; readonly inputSchema: StandardSchemaWithJSON; - readonly outputSchema: StandardSchemaWithJSON; readonly annotations: ToolAnnotations; } @@ -52,7 +51,6 @@ function definitionWith(registration: Registration, inputSchema: Readonly { }); }); -describe('advertisedSchemaOf', () => { - it('advertises the self-contained schema for input and output, and accepts any value so the operation validates', () => { - const { '~standard': standard } = advertisedSchemaOf({ schema: { type: 'object' }, definitions: {} }); +describe('advertisedSchema', () => { + it('advertises the schema it is given as its JSON Schema, and accepts any value so the operation validates', () => { + const { '~standard': standard } = advertisedSchema({ $schema: dialect, type: 'object' }); expect({ vendor: standard.vendor, input: standard.jsonSchema.input({ target: 'draft-2020-12' }), diff --git a/packages/api/src/tools/tool-schema.ts b/packages/api/src/tools/tool-schema.ts index 85b215e21..2e5d5189e 100644 --- a/packages/api/src/tools/tool-schema.ts +++ b/packages/api/src/tools/tool-schema.ts @@ -81,7 +81,3 @@ export function advertisedSchema(jsonSchema: Readonly): StandardSche }, }; } - -export function advertisedSchemaOf(document: JsonSchemaDocument): StandardSchemaWithJSON { - return advertisedSchema(selfContainedSchemaOf(document)); -} diff --git a/packages/brains/src/operations/brain-operations.test.ts b/packages/brains/src/operations/brain-operations.test.ts index 8b98c33a6..49665fc97 100644 --- a/packages/brains/src/operations/brain-operations.test.ts +++ b/packages/brains/src/operations/brain-operations.test.ts @@ -29,14 +29,4 @@ describe('the brain operations', () => { ]); expect(new Set(routes).size).toBe(routes.length); }); - - it('describe a brain once, as a shared definition', () => { - expect(catalog.operations.map(({ output }) => Object.keys(output.definitions))).toEqual([ - ['Brain'], - ['Brain'], - ['Brain'], - ['Brain'], - ['Brain'], - ]); - }); }); diff --git a/packages/operations/README.md b/packages/operations/README.md index 3043faabb..f082653ac 100644 --- a/packages/operations/README.md +++ b/packages/operations/README.md @@ -90,7 +90,7 @@ Plain words have bounds: an outcome takes at most `mostOutcomeCharacters`, 400, `InvalidInput`, `Unavailable` and `Conflict` may also carry a `record`, a JSON object of what was done before the rejection, such as the tokens a model call spent. No outcome or problem document shows it; `@beonauto/specs` keeps it on a run its runtime adapter rejected. -`getLabel.registration` is what a catalog stores: the route, the kind, the success status, the reasons, the permissions that let a caller call it, whether the operation targets a brain, whether it reaches outside the server, and JSON Schema for the input and output with their definitions kept apart. Its `run` decodes an input, runs the handler and encodes the output; only the dispatcher calls it, because it checks nothing about the caller. +`getLabel.registration` is what a catalog stores: the route, the kind, the success status, the reasons, the permissions that let a caller call it, whether the operation targets a brain, whether it reaches outside the server, and the JSON Schema of the input with its definitions kept apart. Its `run` decodes an input, runs the handler and encodes the output; only the dispatcher calls it, because it checks nothing about the caller. `getLabel.call(input)` runs the handler in process with typed input and output, checking both against their schemas. This is how one operation calls another. The call runs with the authority of the calling operation: it does not check the permission of the operation it calls, nor run the pipeline steps. diff --git a/packages/operations/src/definition/operation.test.ts b/packages/operations/src/definition/operation.test.ts index bc6646a01..c1ae21f81 100644 --- a/packages/operations/src/definition/operation.test.ts +++ b/packages/operations/src/definition/operation.test.ts @@ -68,14 +68,6 @@ describe('a registration', () => { expect(addNote.registration.input).toEqual({ schema: noteJsonSchema, definitions: {} }); }); - it('keeps named definitions apart so a transport can hoist them', () => { - expect(getNote.registration.output).toEqual({ - schema: { type: 'object', $ref: '#/$defs/Note' }, - definitions: { Note: noteJsonSchema }, - }); - expect(addNote.registration.output.definitions).toEqual({ Note: noteJsonSchema }); - }); - it('answers with 200 unless the definition says otherwise, and names its path parameters', () => { expect(getNote.registration).toMatchObject({ successStatus: 200, pathParameters: ['name'] }); expect(labelBrain.registration).toMatchObject({ scope: 'org', successStatus: 200, pathParameters: ['brain'] }); diff --git a/packages/operations/src/definition/operation.ts b/packages/operations/src/definition/operation.ts index fe561d384..ce464b56e 100644 --- a/packages/operations/src/definition/operation.ts +++ b/packages/operations/src/definition/operation.ts @@ -110,7 +110,6 @@ function defineOperation< permissions, reasons, input: jsonSchemaDocumentOf(inputSchema), - output: jsonSchemaDocumentOf(outputSchema), ...plainLanguageFields(definition.plainLanguage, inputSchema, outputSchema), run: runnerOf(decodeInput, handle, encodeOutput, reasons), }, diff --git a/packages/operations/src/definition/registration.ts b/packages/operations/src/definition/registration.ts index 37c979faf..3c947dd47 100644 --- a/packages/operations/src/definition/registration.ts +++ b/packages/operations/src/definition/registration.ts @@ -29,7 +29,6 @@ export interface RegistrationOf Effect.Effect>; } diff --git a/packages/server/src/inference/inference-over-mcp.test.ts b/packages/server/src/inference/inference-over-mcp.test.ts index 464a60e05..d0f25d21a 100644 --- a/packages/server/src/inference/inference-over-mcp.test.ts +++ b/packages/server/src/inference/inference-over-mcp.test.ts @@ -3,7 +3,6 @@ import { guideToolName, listedTools, problemIn, - schemasOf, textOf, withMcpSession, type McpSession, @@ -130,11 +129,11 @@ describe('the spec tools an agent sees on the endpoint of a brain', () => { ); }); - it('have self-contained input and output schemas with an object root', async () => { + it('have self-contained input schemas with an object root', async () => { const tools = listedTools(await onAlpha([], (session) => session.listTools())); - const schemas = schemasOf(tools); + const schemas = tools.map(({ inputSchema }) => inputSchema); - expect(schemas).toHaveLength(37); + expect(schemas).toHaveLength(19); expect(schemas.map((schema) => schema['type'])).toEqual(schemas.map(() => 'object')); expect(schemas.flatMap((schema) => danglingReferencesIn(schema))).toEqual([]); }); diff --git a/packages/server/src/inference/listed-models.test.ts b/packages/server/src/inference/listed-models.test.ts index 1703b213b..b05aa202a 100644 --- a/packages/server/src/inference/listed-models.test.ts +++ b/packages/server/src/inference/listed-models.test.ts @@ -1,7 +1,6 @@ import { internalTermsIn, listedTools, - outputConformsTo, plainTextIn, technicalTextIn, withMcpSession, @@ -158,7 +157,6 @@ describe('list_models over MCP', () => { expect(own.structuredContent).toMatchObject(listedModels); expect(org.structuredContent).toEqual(own.structuredContent); - expect(outputConformsTo(tools, 'list_models', own.structuredContent)).toBe(true); expect(listedTools(tools).find(({ name }) => name === 'list_models')?.annotations).toMatchObject({ readOnlyHint: true, openWorldHint: true, diff --git a/packages/server/src/inference/tool-servers.test.ts b/packages/server/src/inference/tool-servers.test.ts index 0b70b5e2c..69a147867 100644 --- a/packages/server/src/inference/tool-servers.test.ts +++ b/packages/server/src/inference/tool-servers.test.ts @@ -1,7 +1,6 @@ import { internalTermsIn, listedTools, - outputConformsTo, plainTextIn, problemIn, withMcpSession, @@ -186,7 +185,6 @@ describe('list_tool_servers over MCP', () => { openWorldHint: true, }); expect(listed.structuredContent).toEqual(listedForAlpha); - expect(outputConformsTo(tools, 'list_tool_servers', listed.structuredContent)).toBe(true); expect(plainTextIn(listed)).toBe(listedInWords); expect(internalTermsIn(plainTextIn(listed))).toEqual([]); }); @@ -234,7 +232,6 @@ function listedOnTheOrgEndpoint(server: ReasoningServer) { 'current revision', { url: `${server.origin}/orgs/acme/mcp`, headers: {} }, async (session) => ({ - tools: await session.listTools(), inOrg: await session.callTool('list_tool_servers', {}), inAlpha: await session.callTool('list_tool_servers', { brain: 'alpha' }), }), @@ -242,16 +239,13 @@ function listedOnTheOrgEndpoint(server: ReasoningServer) { } describe('list_tool_servers on the org endpoint, without a brain and with one', () => { - it('answers for the org and for the brain, each answer in the shape the tool lists', async () => { + it('answers for the org and for the brain', async () => { const { server } = await serving(); - const { tools, inOrg, inAlpha } = await listedOnTheOrgEndpoint(server); + const { inOrg, inAlpha } = await listedOnTheOrgEndpoint(server); expect(inOrg.structuredContent).toMatchObject({ tool_servers: [graphInTheOrg, salesInTheOrg, wikiInTheOrg] }); expect(inAlpha.structuredContent).toEqual({ tool_servers: [graphInTheOrg, wikiInTheOrg] }); - expect( - [inOrg, inAlpha].map(({ structuredContent }) => outputConformsTo(tools, 'list_tool_servers', structuredContent)), - ).toEqual([true, true]); expect(plainTextIn(inAlpha)).toBe(listedInWords); expect(plainTextIn(inOrg)).toMatch( /^The brains of this org may use 3 tool servers\. “graph”, for every brain, offers 2 tools: search and echo;/u, diff --git a/packages/server/src/inference/tool-tests-over-mcp.test.ts b/packages/server/src/inference/tool-tests-over-mcp.test.ts index 156638dc0..515e79dac 100644 --- a/packages/server/src/inference/tool-tests-over-mcp.test.ts +++ b/packages/server/src/inference/tool-tests-over-mcp.test.ts @@ -1,13 +1,6 @@ import { setTimeout } from 'node:timers/promises'; -import { - listedTools, - outputConformsTo, - plainTextIn, - problemIn, - withMcpSession, - type McpSession, -} from '@beonauto/api/testing'; +import { listedTools, plainTextIn, problemIn, withMcpSession, type McpSession } from '@beonauto/api/testing'; import { createApiKey } from '@beonauto/identity'; import { serveFakeMcp, type FakeMcpServer } from '@beonauto/mcp/testing'; import { allPermissions } from '@beonauto/operations'; @@ -69,17 +62,15 @@ describe('test_tool_call over MCP', () => { it('is served on /mcp with the brain as an argument and on the brain endpoint, never on the org endpoint', async () => { const { server } = await serving(); - const { all, tested } = await on(server, '/mcp', builder.key, async (session) => ({ - all: await session.listTools(), - tested: await session.callTool('test_tool_call', { brain: 'alpha', ...searched }), - })); + const tested = await on(server, '/mcp', builder.key, (session) => + session.callTool('test_tool_call', { brain: 'alpha', ...searched }), + ); const onTheBrain = await on(server, '/orgs/acme/brains/alpha/mcp', builder.key, (session) => session.callTool('test_tool_call', searched), ); const onTheOrg = await on(server, '/orgs/acme/mcp', builder.key, (session) => session.listTools()); expect([tested.structuredContent, onTheBrain.structuredContent]).toMatchObject([answered, answered]); - expect(outputConformsTo(all, 'test_tool_call', tested.structuredContent)).toBe(true); expect(plainTextIn(tested)).toMatch( /^The tool “search” of “graph” answered in [\d,]+ ms with \d+ bytes; what a reasoning function's model would see is in the details\.$/u, ); diff --git a/packages/server/src/mcp/answering-what-runs-wait-on.test.ts b/packages/server/src/mcp/answering-what-runs-wait-on.test.ts index 9c38919c6..99022d763 100644 --- a/packages/server/src/mcp/answering-what-runs-wait-on.test.ts +++ b/packages/server/src/mcp/answering-what-runs-wait-on.test.ts @@ -97,24 +97,6 @@ const draftAnswerSchema = { properties: { decision: { type: 'string', enum: ['approve', 'revise', 'skip'] }, note: { type: 'string' } }, }; -const Described = Schema.Struct({ description: Schema.String }); - -const decodeListedFields = Schema.decodeUnknownSync( - Schema.Struct({ - $defs: Schema.Struct({ - Interaction: Schema.Struct({ - properties: Schema.Struct({ - answer_schema: Described, - delivery: Described, - conversation: Described, - answerer: Described, - reply_refusals: Described, - }), - }), - }), - }), -); - type OpenRequest = (typeof OpenRequestsSchema.Type)['interactions'][number]; async function openRequestsIn(session: McpSession): Promise { @@ -218,10 +200,8 @@ describe( ); describe('the shape of the answer an open request takes, as the agent reads it', () => { - it('is on each listed request, described within the bounds the server holds texts to', () => { + it('is on each listed request, as the description of list_interactions says within the bounds the server holds texts to', () => { const description = descriptionIn(meetings.surfaces, 'list_interactions'); - const listing = meetings.surfaces.tools.find(({ name }) => name === 'list_interactions'); - const fields = decodeListedFields(listing?.outputSchema).$defs.Interaction.properties; expect(sentenceNaming(description, 'answer_schema')).toBe( 'Each carries its `answer_schema`, the shape answer_interaction checks an answer against, as recorded when it was asked, which get_spec may no longer show; null for a notification.', @@ -229,17 +209,6 @@ describe('the shape of the answer an open request takes, as the agent reads it', expect(descriptionIn(meetings.surfaces, 'answer_interaction')).toContain( "`answer` takes the shape of the request's answer_schema, which list_interactions shows,", ); - expect([ - description.length, - sentencesOf(description).length, - fields.answer_schema.description.length < 300, - ]).toEqual([724, 5, true]); - expect(fields.delivery.description).toBe( - 'The tool the function delivers the request through, as its deliver names it, or null for a request waiting in the inbox', - ); - expect(fields.delivery.description).toHaveLength(119); - expect( - [fields.conversation, fields.answerer, fields.reply_refusals].map(({ description: words }) => words.length < 300), - ).toEqual([true, true, true]); + expect([description.length, sentencesOf(description).length]).toEqual([724, 5]); }); }); diff --git a/packages/server/src/mcp/mcp-endpoint.test.ts b/packages/server/src/mcp/mcp-endpoint.test.ts index 8b1b19c42..7ecb5018d 100644 --- a/packages/server/src/mcp/mcp-endpoint.test.ts +++ b/packages/server/src/mcp/mcp-endpoint.test.ts @@ -6,7 +6,6 @@ import { guideToolName, listedTools, problemIn, - schemasOf, takingBrain, withMcpSession, type ListedTool, @@ -97,7 +96,7 @@ describe('the brain argument of the spec tools on /mcp', () => { }); describe('the tools of /mcp', () => { - it('are the brain tools, list_models, list_tool_servers and the spec tools, the spec tools taking a brain, with self-contained schemas', async () => { + it('are the brain tools, list_models, list_tool_servers and the spec tools, the spec tools taking a brain, with self-contained input schemas', async () => { server = await servingReasoning([]); await server.call('POST', '/v1/orgs/local/brains', { body: { brain: 'alpha', name: 'Alpha' } }); @@ -106,7 +105,7 @@ describe('the tools of /mcp', () => { await listingOn('/orgs/local/mcp'), await listingOn('/orgs/local/brains/alpha/mcp'), ]; - const schemas = schemasOf(own); + const schemas = own.map(({ inputSchema }) => inputSchema); const onlyInBrain = operations(brain).filter(({ name }) => insideABrainAlone.includes(name)); expect(own.map(({ name }) => name)).toEqual([...orgTools, ...insideABrainAlone, guideToolName]); diff --git a/packages/server/src/mcp/mcp-with-keys.test.ts b/packages/server/src/mcp/mcp-with-keys.test.ts index d4458f5f7..fc8046be1 100644 --- a/packages/server/src/mcp/mcp-with-keys.test.ts +++ b/packages/server/src/mcp/mcp-with-keys.test.ts @@ -1,4 +1,4 @@ -import { outputConformsTo, problemIn, toolNamesIn, withMcpSession, type McpConnection } from '@beonauto/api/testing'; +import { problemIn, toolNamesIn, withMcpSession, type McpConnection } from '@beonauto/api/testing'; import { createApiKey } from '@beonauto/identity'; import { allPermissions } from '@beonauto/operations'; import { afterEach, beforeEach, describe, expect, it } from 'vitest'; @@ -102,12 +102,11 @@ describe('the org endpoint of the server', () => { }); describe('the brain tools of the server', () => { - it('create, list, read, update and retire a brain, each output conforming to its schema', async () => { + it('create, list, read, update and retire a brain', async () => { const outcome = await withMcpSession( 'previous major', endpoint('/orgs/acme/mcp', acmeAdmin.key), async (session) => ({ - listing: await session.listTools(), created: await session.callTool('create_brain', { brain: 'alpha', name: 'Alpha' }), listed: await session.callTool('list_brains', {}), read: await session.callTool('get_brain', { brain: 'alpha' }), @@ -115,20 +114,11 @@ describe('the brain tools of the server', () => { retired: await session.callTool('retire_brain', { brain: 'alpha' }), }), ); - const { listing } = outcome; - expect(outcome.created.structuredContent).toMatchObject({ id: 'alpha', name: 'Alpha', status: 'active' }); expect(outcome.listed.structuredContent).toMatchObject({ brains: [{ id: 'alpha' }] }); expect(outcome.read.structuredContent).toEqual(outcome.created.structuredContent); expect(outcome.updated.structuredContent).toMatchObject({ name: 'Alpha prime', description: 'Notes' }); expect(outcome.retired.structuredContent).toMatchObject({ status: 'retired' }); - expect([ - outputConformsTo(listing, 'create_brain', outcome.created.structuredContent), - outputConformsTo(listing, 'list_brains', outcome.listed.structuredContent), - outputConformsTo(listing, 'get_brain', outcome.read.structuredContent), - outputConformsTo(listing, 'update_brain', outcome.updated.structuredContent), - outputConformsTo(listing, 'retire_brain', outcome.retired.structuredContent), - ]).toEqual([true, true, true, true, true]); }); }); diff --git a/packages/server/src/mcp/served-schemas.test.ts b/packages/server/src/mcp/served-schemas.test.ts new file mode 100644 index 000000000..0d9892165 --- /dev/null +++ b/packages/server/src/mcp/served-schemas.test.ts @@ -0,0 +1,152 @@ +import { + internalTermsIn, + listedTools, + mcpClientKinds, + plainTextIn, + technicalTextIn, + withMcpSession, + type ListedTool, + type ToolResult, +} from '@beonauto/api/testing'; +import { answers, textResult } from '@beonauto/inference/testing'; +import { Option, Schema } from 'effect'; +import { afterAll, beforeAll, describe, expect, it } from 'vitest'; + +import { servingReasoning, type ReasoningServer } from '../testing/servers/reasoning-server.ts'; + +const summary = [ + '---', + 'model: anthropic/claude-sonnet-4-5', + 'input:', + ' schema: {type: object, properties: {text: {type: string}}, required: [text]}', + '---', + 'Summarize: {{ input.text }}', +].join('\n'); + +let server: ReasoningServer; + +beforeAll(async () => { + server = await servingReasoning(mcpClientKinds.map(() => answers(textResult('Profits rose.')))); + await withMcpSession('current revision', { url: `${server.origin}/mcp`, headers: {} }, async (session) => { + await session.callTool('create_brain', { brain: 'alpha', name: 'Alpha' }); + await session.callTool('create_spec', { brain: 'alpha', primitive: 'inference', name: 'summary', source: summary }); + }); +}); + +afterAll(async () => { + await server.stop(); +}); + +function listingOn(path: string): Promise { + return withMcpSession('current revision', { url: `${server.origin}${path}`, headers: {} }, async (session) => + listedTools(await session.listTools()), + ); +} + +interface ListingOnTheWire { + readonly status: number; + readonly bytes: number; + readonly toolCount: number | undefined; +} + +const listingMessageOf = Schema.decodeUnknownOption( + Schema.Struct({ result: Schema.Struct({ tools: Schema.Array(Schema.Unknown) }) }), +); + +function toolCountIn(body: string): number | undefined { + const data = body.split('\n').find((line) => line.startsWith('data: ')); + const message: unknown = data === undefined ? undefined : JSON.parse(data.slice('data: '.length)); + return Option.getOrUndefined(Option.map(listingMessageOf(message), ({ result }) => result.tools.length)); +} + +async function listingOnTheWire(path: string): Promise { + const answer = await fetch(`${server.origin}${path}`, { + method: 'POST', + headers: { + 'content-type': 'application/json', + accept: 'application/json, text/event-stream', + 'mcp-protocol-version': '2025-11-25', + }, + body: JSON.stringify({ jsonrpc: '2.0', id: 1, method: 'tools/list', params: {} }), + }); + const body = await answer.text(); + return { status: answer.status, bytes: Buffer.byteLength(body), toolCount: toolCountIn(body) }; +} + +interface Endpoint { + readonly path: string; + readonly tools: number; + readonly mostBytes: number; +} + +const endpoints: readonly Endpoint[] = [ + { path: '/mcp', tools: 25, mostBytes: 40_000 }, + { path: '/orgs/local/mcp', tools: 8, mostBytes: 9000 }, + { path: '/orgs/local/brains/alpha/mcp', tools: 19, mostBytes: 32_000 }, +]; + +describe('the tools of each endpoint', () => { + it.each(endpoints)( + 'on $path each carry an input schema with an object root and no output schema', + async ({ path, tools }) => { + const listed = await listingOn(path); + + expect(listed.map(({ inputSchema }) => inputSchema['type'])).toEqual( + Array.from({ length: tools }, () => 'object'), + ); + expect(listed.filter(({ outputSchema }) => outputSchema !== undefined).map(({ name }) => name)).toEqual([]); + }, + ); +}); + +describe('the listing of each endpoint', () => { + it.each(endpoints)( + 'on $path answers 200 with its $tools tools in at most $mostBytes bytes on the wire', + async ({ path, tools, mostBytes }) => { + const { status, bytes, toolCount } = await listingOnTheWire(path); + + expect({ status, toolCount }).toEqual({ status: 200, toolCount: tools }); + expect(bytes).toBeLessThanOrEqual(mostBytes); + }, + ); +}); + +function outputIn(result: ToolResult): unknown { + const output: unknown = JSON.parse(technicalTextIn(result)); + return output; +} + +describe.each(mcpClientKinds)('the %s client on /mcp', (kind) => { + it('lists the tools, then accepts the results of list_tool_servers, list_brains and execute_spec without compiling an output validator', async () => { + const outcome = await withMcpSession(kind, { url: `${server.origin}/mcp`, headers: {} }, async (session) => { + await session.listTools(); + const results = [ + await session.callTool('list_tool_servers', {}), + await session.callTool('list_brains', {}), + await session.callTool('execute_spec', { + brain: 'alpha', + primitive: 'inference', + name: 'summary', + input: { text: 'the quarter' }, + }), + ]; + return { results, compiled: session.compiledOutputSchemas() }; + }); + + expect(outcome.compiled).toEqual([]); + expect(outcome.results.map((result) => plainTextIn(result))).toEqual([ + 'Whoever runs this server has set up no tool server for this org, so the functions of its brains can call no tools until they set one up; the give-tools guide says what they need.', + 'There is 1 brain: “Alpha”.', + 'Ran the reasoning function “summary”. Its answer: “Profits rose.”', + ]); + expect(outcome.results.flatMap((result) => internalTermsIn(plainTextIn(result)))).toEqual([]); + expect(outcome.results.map(({ structuredContent }) => structuredContent)).toMatchObject([ + { tool_servers: [] }, + { brains: [{ id: 'alpha' }] }, + { status: 'succeeded', output: 'Profits rose.' }, + ]); + expect(outcome.results.map((result) => [result.content.length, outputIn(result)])).toEqual( + outcome.results.map(({ structuredContent }) => [2, structuredContent]), + ); + }); +}); diff --git a/packages/specs/src/operations/spec-operations.test.ts b/packages/specs/src/operations/spec-operations.test.ts index 5ae54e495..4c4cd289c 100644 --- a/packages/specs/src/operations/spec-operations.test.ts +++ b/packages/specs/src/operations/spec-operations.test.ts @@ -66,22 +66,6 @@ describe('a catalog of the spec operations', () => { 'get_brain_analytics', ]); }); - - it('describes a definition and a run once each, as shared definitions', () => { - expect(catalog.operations.map(({ output }) => Object.keys(output.definitions))).toEqual([ - ['Definition'], - ['ListedDefinition'], - ['Definition'], - ['Definition'], - ['Definition'], - ['Run'], - ['RunDetail'], - ['Run'], - ['ListedRun'], - ['PublicEvent'], - ['BrainAnalytics'], - ]); - }); }); describe('the spec operations for a list of primitives', () => {