Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@
- `primitives/*`: one package per brain primitive (interaction, orchestration, inference, prediction, computation, recollection, dream)
- `TODO.md`: setup work that is still outstanding

In text a user reads, a primitive is a brain function and an execution a run: a brain can reason (inference, whose specs are reason functions, each configured by a prompt), interact (interaction), compute (computation), recall (recollection) and predict (prediction), and it coordinates them through workflows (orchestration); code, the API and the architecture keep the primitive names.

## Commands

```bash
Expand Down
36 changes: 20 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,17 +14,21 @@ auto-brain is the server a brain runs on. Auto can host it for you, or you can r

## How a brain works

A brain is made of **primitives** that share one **ledger**.

| Primitive | Today | What you make with it | What it does |
| ----------------------------------------- | ------- | --------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| [Interaction](primitives/interaction) | Planned | | Input and output between the brain and people or machines, in both directions |
| [Orchestration](primitives/orchestration) | Built | A workflow | Runs a workflow of deterministic steps that execute other specs, branch, loop, retry, wait and listen for events, durably, on Temporal. A workflow starts when its spec is executed; starting one from an event or on a schedule is planned |
| [Inference](primitives/inference) | Built | A prompt | Calls a language model with a prompt written in Markdown: front matter sets the model, its settings and the JSON Schemas of the input and output, and a Liquid template renders the prompt from the input. Skills, tools and context from the rest of the brain are planned |
| [Prediction](primitives/prediction) | Planned | | A machine-learning model that makes a prediction, for when an LLM isn't the right tool |
| [Computation](primitives/computation) | Planned | | A deterministic function that workflows and agents can call |
| [Recollection](primitives/recollection) | Planned | | A materialized view of the brain's history, built from the ledger |
| [Dream](primitives/dream) | Planned | | Explores the ledger around a subject to suggest new scenarios and better ways of working, and can iterate towards a goal |
A brain can reason, interact, compute, recall and predict. It coordinates those functions through workflows.

| Brain function | What you define | How it runs | Built or planned |
| -------------- | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------- |
| Reason | A reason function, configured by a prompt | [Inference](primitives/inference) calls a language model with a prompt written in Markdown: front matter sets the model, its settings and the JSON Schemas of the input and output, and a Liquid template renders the prompt from the input. Skills, tools and context from the rest of the brain are planned | Built |
| Interact | | [Interaction](primitives/interaction) carries input and output between the brain and people or machines, in both directions | Planned |
| Compute | | [Computation](primitives/computation) runs a deterministic function that workflows and agents can call | Planned |
| Recall | | [Recollection](primitives/recollection) keeps a materialized view of the brain's history, built from the ledger | Planned |
| Predict | | [Prediction](primitives/prediction) makes a prediction with a machine-learning model, for when an LLM isn't the right tool | Planned |
| Coordinate | A workflow | [Orchestration](primitives/orchestration) runs a workflow of deterministic steps that execute other specs, branch, loop, retry, wait and listen for events, durably, on Temporal. A workflow starts when its spec is executed; starting one from an event or on a schedule is planned | Built |
| Not named yet | | [Dream](primitives/dream) explores the ledger around a subject to suggest new scenarios and better ways of working, and can iterate towards a goal | Planned |

A budget-review workflow could recall previous decisions, compute the remaining budget, predict outcomes, reason about alternatives, and interact with a person for approval. Coordinating those functions is what the workflow does.

Each function runs on a **primitive**, the name the code and the API use for it, and all of them share one **ledger**.

The [ledger](packages/ledger) records every input and output of every primitive. That record lets a brain recall what happened and explain its decisions. It also lets you evaluate and improve the method over time.

Expand Down Expand Up @@ -99,18 +103,18 @@ Next, connect your AI assistant to `http://localhost:8080/mcp`. Local mode needs
- **VS Code**, in `.vscode/mcp.json`: `{"servers": {"auto-brain": {"type": "http", "url": "http://localhost:8080/mcp"}}}`
- **Other assistants** take the entry under [Connecting an agent over MCP](#connecting-an-agent-over-mcp), without its header.

Then ask it, in order. When you ask for a prompt, name a model your provider serves, such as `anthropic/claude-sonnet-4-5` or `gateway/<a model id your gateway serves>`: the assistant learns which providers the server has, but not which models your account offers.
Then ask it, in order. When you ask for a reason function, name a model your provider serves, such as `anthropic/claude-sonnet-4-5` or `gateway/<a model id your gateway serves>`: the assistant learns which providers the server has, but not which models your account offers.

1. "Create a brain called support for our customer support team."
2. "In support, write a prompt that classifies a support ticket by category (billing, bug, account or other) and urgency (low, normal or high), answering in JSON, and run it on: I was charged twice for March and nobody has answered for three days."
2. "In support, create a reason function that classifies a support ticket by category (billing, bug, account or other) and urgency (low, normal or high), answering in JSON, and run it on: I was charged twice for March and nobody has answered for three days."
3. "How many tokens did that run use, and what exactly was sent to the model?"
4. "Change the prompt so that anything about money is billing and at least normal urgency, then run it on the same ticket again."
4. "Change the reason function's prompt so that anything about money is billing and at least normal urgency, then run it on the same ticket again."
5. "Build a workflow that classifies a ticket and, only when it is urgent, drafts a two-sentence note for the on-call lead. Run it on that ticket and on: How do I export my invoices as CSV?"
6. "Start a workflow that waits for a manager to approve a refund, then send it the approval."

You don't need to teach the assistant anything first: each tool's description says how the documents of its primitive are written.

To try it without an assistant, `scripts/try-inference.sh http://localhost:8080 <provider/model>` and `scripts/try-workflows.sh http://localhost:8080 <provider/model>` run a prompt and a workflow over HTTP and print what happened. [How it works](#how-it-works) walks through the same steps.
To try it without an assistant, `scripts/try-inference.sh http://localhost:8080 <provider/model>` and `scripts/try-workflows.sh http://localhost:8080 <provider/model>` run a reason function and a workflow over HTTP and print what happened. [How it works](#how-it-works) walks through the same steps.

### Configuring a model

Expand Down Expand Up @@ -373,7 +377,7 @@ Most MCP clients take an entry of this shape. In Claude Code, `claude mcp add --
| `POST /orgs/{org}/mcp` | `create_brain`, `list_brains`, `get_brain`, `update_brain` and `retire_brain` | to manage the brains of one org and nothing else |
| `POST /orgs/{org}/brains/{brain}/mcp` | the tools inside a brain, acting in that brain, without a `brain` argument | to lock a connection to one brain, such as for one agent |

Every endpoint speaks streamable HTTP without sessions. It serves the current stateless revision (`2026-07-28`) and the earlier ones the SDK supports (`2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`), so agents built on older SDKs connect too. Each tool carries the operation's description and its input and output JSON Schemas, and is marked read-only when it only reads. Each result leads with a sentence or two in plain words for the person the agent works for, then gives the details for follow-up calls. Those words call a spec of the inference primitive a prompt, one of the orchestration primitive a workflow, and an execution a run, while the tools take the primitive's name, `inference` or `orchestration`; each primitive's description, which the spec tools carry, says so. A tool that cannot do what was asked returns `isError`, its plain words saying what could not be done and who can fix it, and then the same problem document HTTP would answer with, as text, so the agent can read the `reason` and the `detail`, and correct its arguments when the `reason` is `invalid_input`. The key's permissions and brains hold as they do over HTTP: a read-only key can call `list_brains` but gets `forbidden` from `create_brain`, a key limited to some brains gets `forbidden` for any other, and a brain the org does not have is `not_found`. [`packages/api`](packages/api) describes the mappings in full.
Every endpoint speaks streamable HTTP without sessions. It serves the current stateless revision (`2026-07-28`) and the earlier ones the SDK supports (`2025-11-25`, `2025-06-18`, `2025-03-26`, `2024-11-05` and `2024-10-07`), so agents built on older SDKs connect too. Each tool carries the operation's description and its input and output JSON Schemas, and is marked read-only when it only reads. Each result leads with a sentence or two in plain words for the person the agent works for, then gives the details for follow-up calls. Those words call a spec of the inference primitive a reason function, one of the orchestration primitive a workflow, and an execution a run, while the tools take the primitive's name, `inference` or `orchestration`; each primitive's description, which the spec tools carry, says so. A tool that cannot do what was asked returns `isError`, its plain words saying what could not be done and who can fix it, and then the same problem document HTTP would answer with, as text, so the agent can read the `reason` and the `detail`, and correct its arguments when the `reason` is `invalid_input`. The key's permissions and brains hold as they do over HTTP: a read-only key can call `list_brains` but gets `forbidden` from `create_brain`, a key limited to some brains gets `forbidden` for any other, and a brain the org does not have is `not_found`. [`packages/api`](packages/api) describes the mappings in full.

### Errors

Expand Down
2 changes: 1 addition & 1 deletion packages/api/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ All three sit behind the same chain as every other path. Before the SDK runs, a
| succeeded | `structuredContent` is the output; the first text content says in plain words what happened, and the second is the same output as JSON |
| rejected, failed, cancelled | `isError: true`; the first text content says in plain words what could not be done, and the second is the problem document HTTP would answer with, as JSON; no `structuredContent` |

The plain words are for the person an agent works for, so they name things as that person knows them, and carry no ids, versions, formats or status codes: a spec of the inference primitive is a prompt, one of the orchestration primitive a workflow, and an execution a run. The tools still take the primitive's name, `inference` or `orchestration`, and each primitive's description opens by saying which word goes with it, so an agent knows that a person's prompt is a spec of `inference`. Each operation says them itself, through the `plainLanguage` of its definition: a `task` and an `attempt` that name what it does, and an `outcome` from its output; each primitive gives the noun for its specs and a sentence for what one of its runs gave back. An endpoint refuses, when it is mounted, to serve an operation without them. The words for what could not be done come from one place, `unsuccessfulWords` in `@beonauto/operations`, by the reason of the rejection and its `kind`, when it has one (`taken`, `retired`, `concurrent_change` or `unworkable` for a conflict, `model_not_offered` for `unavailable`): a rejection the agent can correct says so; one only whoever runs the server can resolve says that, and that nothing on the person's side needs to change; a prompt that names a model of a provider the server is not set up for, while it can use others, is `model_not_offered`, and its words say that the prompt can be switched to one of those, which the details list; one about the request names what is missing, not allowed, taken or retired; and an unexpected failure gives the reference to quote. HTTP answers are unchanged.
The plain words are for the person an agent works for, so they name things as that person knows them, and carry no ids, versions, formats or status codes: a spec of the inference primitive is a reason function, one of the orchestration primitive a workflow, and an execution a run. The tools still take the primitive's name, `inference` or `orchestration`, and each primitive's description opens by saying which word goes with it, so an agent knows that a person's reason function is a spec of `inference`, and that its prompt is the spec's document. Each operation says them itself, through the `plainLanguage` of its definition: a `task` and an `attempt` that name what it does, and an `outcome` from its output; each primitive gives the noun for its specs and a sentence for what one of its runs gave back. An endpoint refuses, when it is mounted, to serve an operation without them. The words for what could not be done come from one place, `unsuccessfulWords` in `@beonauto/operations`, by the reason of the rejection and its `kind`, when it has one (`taken`, `retired`, `concurrent_change` or `unworkable` for a conflict, `model_not_offered` for `unavailable`): a rejection the agent can correct says so; one only whoever runs the server can resolve says that, and that nothing on the person's side needs to change; a reason function whose prompt names a model of a provider the server is not set up for, while it can use others, is `model_not_offered`, and its words say that the prompt can name a model from one of those instead, which the details list; one about the request names what is missing, not allowed, taken or retired; and an unexpected failure gives the reference to quote. HTTP answers are unchanged.

So invalid arguments are a tool result with `isError` and an `invalid_input` problem pointing at each field, as the protocol asks, and an agent can correct them. The SDK's parsing of the arguments drops one named `__proto__`; each endpoint reads the request body before the SDK, within the same 1 MiB, and puts such an argument back before the call is dispatched, so it is rejected as an excess field, `invalid_input` at `/__proto__`, as over HTTP. A tool the endpoint does not list is a JSON-RPC error, `-32602`, from the SDK. A call still running when the server stops gets a `503` `unavailable` problem. A tool that throws instead of settling is reported as an incident, as an HTTP request would be, and answered with the `500` `internal` problem, so the SDK never puts an error message of its own in the result.

Expand Down
2 changes: 1 addition & 1 deletion packages/api/src/testing/internal-terms.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ describe('internalTermsIn', () => {
it('finds none in words for a person', () => {
expect(
internalTermsIn(
'Created the prompt “summary”. What it does: Summarizes a text. It has been saved but has not been run yet.',
'Created the reason function “summary”. What it does: Summarizes a text. It has been saved but has not been run yet.',
),
).toEqual([]);
});
Expand Down
Loading
Loading