Skip to content

feat(interaction): an interaction function calls a tool and answers with its result - #139

Merged
rami-hatoum merged 13 commits into
mainfrom
feat/asking-a-system
Oct 10, 2026
Merged

rami-hatoum merged 13 commits into
mainfrom
feat/asking-a-system

Conversation

@rami-hatoum

@rami-hatoum rami-hatoum commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

What an agent writes

An interaction function can now ask a system and answer at once. Its document names, under call, one tool of a tool server the brain may use, the arguments in with, and in read a JSON Pointer to the part of the tool's answer the function answers with. output.schema gives the shape of that part:

---
description: Read the replies of a thread in the team chat, oldest first
call:
  server: chat
  tool: thread_replies
  with:
    channel: '{{ input.channel }}'
    ts: '{{ input.thread }}'
  read: /messages
input:
  schema: { type: object, required: [channel, thread], properties: { channel: { type: string }, thread: { type: string } } }
output:
  schema:
    type: array
    items: { type: object, required: [user, ts], properties: { user: { type: string }, text: { type: string }, ts: { type: string } } }
---

A run makes one call and records it in its history before the call is sent and again when it is answered. It reads the part read points at, checks it against output.schema and ends with it as its output. It makes no request, sends no message and asks no model. Arguments keep their types: a number, a boolean, null, a list or an object is sent as written. A text is a Liquid template over input, today and now, and a text that is one {{ }} alone sends the value it reads as it is. call is refused beside to, expires, deliver, replies, from, reply or a body, and needs output.schema. A function that asks a person still needs to and expires, and works as before.

The instructions an agent reads on connecting now say "An interaction function asks a system and answers at once, or asks a person and answers started." The interaction-function guide gains a section, "Asking a system", that walks through it in order: list the tool servers, test the tool, read its answer, write call, read and output.schema, save. The give-tools recipe sends an agent to that guide when the brain only needs what a tool answers.

What a workflow gets

A step that calls such a function with run_definition gets what the tool answered at read as its output, within the call. A run of one lasts 70 seconds at most: 10 to open a connection, 10 to list the server's tools, 10 to restart a process that exited, 30 to call and 10 to reopen a session the server forgot. cancel_run and answer_interaction refuse its run with conflict, and list_interactions never lists it. A tool that pages answers one page per run, so the workflow loops with the cursor. A recall function folds the run's ending, and an event trigger reacts to it, as they do for any run.

The brain checks what a call answers itself, against read and output.schema. A call is sent as a plain tools/call request, so the output schema a server declares for its tool is never checked by the client: a tool whose structured content breaks it answers as it answered, and the brain's own check says what does not match.

The endings:

Ending When
succeeded The tool answered, and the part at read matched output.schema
unavailable, tool_not_offered or mcp_server_failed, 503 The server does not serve the brain, the tool is not allowed or not listed, or the server cannot be used; nothing was sent
unavailable, tools_unfinished, 503 The call was sent and the server failed, or the tool answered an error, and the server marks the tool read-only
conflict, effect_unknown, 409 The same for any other tool
conflict, unworkable, 409 The tool refused the arguments, or its answer has nothing to read at read or does not match output.schema. After a tool not marked read-only, the detail ends "The tool ran and may have changed something; a run again calls it again"

The one retry rule

Any run that called tools and could not finish, whether a reasoning function or a call, now ends one of two ways:

  • tools_unfinished, unavailable at 503, when every tool it called is one its server marks readOnlyHint. Its words say that a new run is safe.
  • effect_unknown, a conflict at 409 under its own problem type, https://on.auto/problems/effect_unknown, when any tool it called is not marked read-only. Whether that tool changed something is not known, so a person, or a workflow rule that names the kind, decides.

Both carry a because: server_failed, the new tool_error for a tool that answered an error, or one of a reasoning run's. Neither sends Retry-After, and the same run id answers tools_called after either. A run that ended without success after calling only tools that read answers tools_called with the because only_read, and its words say "Nothing was changed."; a run that called any other tool, or one still started, answers it as before. So a workflow's catch on status 503 retries what is safe to try again and never repeats a tool that may already have posted, paid or written. A catch that names the effect_unknown type, or its kind in a when, retries it deliberately.

On the wire

  • A run that called tools and could not finish is tools_unfinished, unavailable at 503, only when every tool it called is marked read-only by its server; otherwise it is effect_unknown, a conflict at 409 under https://on.auto/problems/effect_unknown. This holds for a reasoning run as for a call, and most tools of today's servers carry no read-only mark.
  • A conflict may carry a because, as an unavailable does: tool_error, server_failed and a reasoning run's becauses with effect_unknown, and only_read with tools_called.
  • tool_call_started carries read_only: true when the server marked the tool read-only, and the history shows it. A test's tool_test_started carries it too.

test_tool_call answers answer

Beside text, which is what a reasoning function's model would see, a test now answers answer: the document a call's read points into. That is the tool's structured content, else its first text block parsed as JSON, else that text, scrubbed of the server's secrets as text is. It is left out when the tool answered neither structured content nor text, and when it takes more than 64 KiB as JSON. A call reads exactly what a test shows.

The run shape of run_definition

run_definition now answers the same run as get_run, with its record. A call's record is { server, tool, read }. The change only adds a field: a client that reads run_definition's answer now gets record too. The words of a run say what a call answered, such as "It asked the thread replies tool of chat; its answer: (user: “member-17”, …)", and the records of a list read in parentheses in every capability's words.

Verification

  • pnpm check passes without PostgreSQL, 5,351 tests with the 194 that need PostgreSQL skipped, and against PostgreSQL 18, 5,545 tests.
  • The image built from this branch passes the eleven smoke steps of ci.yml.
  • The new tests run through the real server, over HTTP and MCP, against the fake MCP server. A table pins status, type, kind, because, detail and Retry-After for every way a call can end, among them a key refused at the opening, a listing that never answers, a call that takes longer than a call may, a start the ledger refuses, and each answer the brain cannot use. The tests give the server short call bounds and hold requests at the fake server, so no test waits on a fixed timer. They also cover a call's history and the content recorded only where the server records content; a caller that goes away mid-call; workflows that retry on 503, that name effect_unknown, that page with a cursor, and that react to a call's ending; and the reasoning runs that end tools_unfinished or effect_unknown by the tools they called.
  • The served texts are pinned. The instructions take 1,514, 1,842 and 1,882 characters on the org, brain and /mcp endpoints, give-tools 2,049 and 2,598 bytes, run_definition's description 765 characters and cancel_run's 671. The interaction guide's size is not pinned, only its bound: it serves 36,618 bytes of its 64 KiB.

rami-hatoum and others added 13 commits October 9, 2026 16:21
… effect_unknown

A run that called tools and could not finish now meets one of two
endings, read over every tool it called. When its server marks each of
them readOnlyHint, the run ends unavailable as tools_unfinished, status
503, and its words say that running it again is safe. When any is not
so marked, it ends conflict as the new kind effect_unknown, status 409,
under a problem type of its own, so no retry on 503 repeats a tool that
may already have posted, paid or written.

A conflict carries a because as an unavailable ending does, the becauses
are one schema, RejectionBecauseSchema, and tool_error joins them; a
workflow settles effect_unknown as a conflict with its because.
run_definition answers the run with its record, as get_run does, and
the five capabilities say their output through one outputInWords, with
the records of a list in parentheses. The reasoning adapter's tool
access is required, so a server without tool servers refuses a tool in
the words of the one naming check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s start where asked

The single-call path that deliveries, reads and tellings share is now
named for what it is, the one-call folder with OneCall and oneCall. It
takes a connection within the open bound, lists the server's tools as a
run's opening lists them and finds the tool there, so a tool the server
does not list is refused before anything is sent, in the words of a
run's opening; a delivery records that as a failed attempt, tool not
offered. A caller that passes its run's journal has the call's start
recorded once the tool is found and before it is sent, and its answer
after, through the same journalledCall a reasoning run's calls use; a
start the run cannot record ends the call path with a defect. The
answer carries the hints the server gave the tool, and the caller gives
the longest wait for a 429. The delivery's own copy of the bounds goes.

What a tool answered is read two ways in one module: resultText, what a
model sees, and answerDocument, the structured content, else the first
text block parsed as JSON, else that text. test_tool_call answers that
document as answer beside text, a reply's reading bounds the document
rather than the whole answer, and a tool is put in words one way,
toolInWords, in the history, a delivery, a reading and a test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ith its result

An interaction function may now name, in call, one tool of a tool server
the brain may use, its arguments in with and, in read, a JSON Pointer to
the part of the tool's answer it answers with. Its run makes one call,
recorded in its history before it is sent and when it is answered, reads
that part, checks it against output.schema and ends with it as its
output within the call: no request, no message, no model. call is
refused beside to, expires, deliver, replies, from, reply or a body, and
needs output.schema; to and expires are required only without it.

A call ends as a run's tool use ends: a tool not offered or a server
that cannot be used is unavailable before anything is sent; arguments
the tool refused, or an answer with nothing to read, are unworkable; a
tool error or a failed server after sending is tools_unfinished, 503,
when the server marks the tool read-only, and effect_unknown, 409,
otherwise. The tool blocks of a delivery and a call share one schema and
one compiler in tool-blocks, and the moment variables are built once.

The instructions say that an interaction function asks a system and
answers at once, or asks a person and answers started, and the run
sentence names the workflow alone. give-tools sends an agent whose brain
only needs what a tool answers to an interaction function that calls it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e one rule for unfinished tools

The interaction reference, served as the interaction-function guide,
names the three shapes of a function and gains "Asking a system": list
the tool servers, test the tool and read its answer, write call, read and
output.schema, save. Its field table, its endings and its bounds take
the call's rows, and a workflow's step takes a call's output within the
call. The HTTP, MCP, workflow and reasoning references, the concepts and
terminology pages and the configuration guide say what a call answers
and the one rule for a run that called tools and could not finish:
tools_unfinished, 503, when every tool it called only reads, and
effect_unknown, 409 under a type of its own, when any may write. The
computation reference reads its figures through an interaction function
that calls the tool.

Decision 0021 records the design in the words of decision 0007, with the
counts measured on the built code; 0018 and 0003 take its amendments,
0020 its edits, and the READMEs of interaction, mcp, definitions,
operations, api, reasoning and the workflow engine say what changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s nothing was changed

tools_called may carry the because only_read, for a run that ended
without success after calling only tools whose servers mark them
read-only. Its words are its own: an attempt under the same id did not
succeed and every tool it called only reads, so the refusal of a command
adds "Nothing was changed." and sends the caller to a new run, where the
words of tools_called otherwise say its tools may have changed something.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rt says the tool only reads

The client's callTool checked a tool's structured content against the
output schema its server declares and threw an invalid-params error when
it did not match, which the brain read as arguments the tool refused,
though the tool had run. A call is now sent as a tools/call request
through the client's request, which takes the call-result schema from
the method and checks nothing of a server's output schema; the brain
checks what it reads itself. The fake server's promised tool declares an
output schema and answers against it.

tool_call_started carries read_only: true when the server marked the
tool readOnlyHint when it was listed, for a run's calls, one call and a
test alike. A reply leaves the answer unread until test_tool_call reads
it, so callReplyOf is as it was. One call finds its tool with unlistedOn
and offeredOn, as a run's opening does, the recording is built by one
recordingOf, and a tool is put in words by one function.

The one-call test helpers are named for one call, the fake server can
hold the next request of one method unanswered, and a listing test that
programmed the next answer of any request now programs the listing's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er its id as only_read

The run's state keeps whether every tool_call_started it recorded says
read_only, across a start recorded again. A run that ended without
success after only such calls answers the same id with tools_called,
because only_read, whose words say nothing was changed; a run still
started, or one that recorded any other call, is refused as before. The
history shows read_only on a call's start. The refusals of a run that
called tools have a file of their own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… reasonOfKind

The settlement of an error whose kind has a problem type of its own tells
a conflict through reasonOfKind of @beonauto/operations, in a type guard
that narrows the kind to those a conflict carries, in place of a second
membership check over the conflict kinds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ModelTools carries calledAny from the run's tools, so a scripted model
can wait until a call it started is recorded instead of for a fixed
time; the scripted tools count their calls the same way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rite says so

After a call of a tool its server does not mark read-only, an answer
with nothing to read, nothing at read, too large, too deep or refused by
output.schema still ends unworkable, and its detail now ends "The tool
ran and may have changed something; a run again calls it again". A call
of a tool whose structured content breaks the output schema its server
declares ends with the brain's own "does not match the output schema".

The run says a refused argument through argumentsFailureWords, and the
save and the run share one remedy clause for a structure among text. A
notification is read from the run's record, a request that takes no
answer or a delivery's delivered_at, not from an empty output. Whether a
document calls a tool is read from its decoded front matter where it
could be read, and issuesOf is written once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aller that leaves over MCP

A table pins status, type, kind, because, detail and Retry-After for
every row of a call's endings through the real server and the fake one:
before anything is sent, after the call was sent, and for each answer
the run cannot use. A test gives the server the timing of its tool
access, so a listing held at the fake server and a call that takes too
long reach their bounds at 300 ms, and a trigger on the ledger's file
refuses the start of a call, which fails the run with an incident.

A caller that leaves a call in flight is tested over HTTP and over MCP,
leaving once the fake server has the call, with the run ended, its four
facts, one call and its session let go. The reasoning test that waited
300 ms for a call it started waits until the call is recorded. The
tests of asking a system have a folder of their own and share their
helpers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decision 0021 records what the review changed: a call is sent as a
request, so a tool's own output schema never refuses its arguments; the
detail of an answer a run cannot use after a tool that may write; the
read_only of a call's start and the only_read refusal under a run's id;
the test's answer built by the test; the record's own functions; one
sentence for a structure among text; the notification read from its
record; and the server's tests of every ending. Its appendix gives the
64 characters of tool_error's words, its recorded facts name read_only,
and the guide's measured size is said to be bounded, not pinned.

The interaction reference says an answer a run cannot use after a tool
that may write still ends unworkable and why its detail says a run
again calls the tool again. The HTTP and engineering references, and
the READMEs of interaction, mcp, definitions and operations, name
read_only, only_read and the request a call is sent as. The index gives
0020 its amendment of 2026-10-09.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nly on the first

The HTTP reference says a tool-using run records tool_call_started and
tool_call_answered, the first with read_only where the server marks the
tool read-only, which keeps the phrase the documentation test reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@rami-hatoum
rami-hatoum merged commit d6998f1 into main Oct 10, 2026
18 checks passed
@rami-hatoum
rami-hatoum deleted the feat/asking-a-system branch October 10, 2026 07:23
rami-hatoum added a commit that referenced this pull request Oct 10, 2026
…l, into the TypeScript branch

Brings in d6998f1 (#139). The conflicts, each resolved so both sides
hold:

- computation-function.ts: main's outputInWords('result') in place of
  the local describeResult, with the branch's check, stripped forms and
  invalidDefinition; InvalidInput is a type only.
- served-functions.ts: main's toolTiming handed to toolUsersServedBy,
  and the branch's awaited computationServedBy, which warms the checks
  before the server listens.
- concepts/functions.md: the computation function is the branch's
  TypeScript function, and reads its figures through main's interaction
  function that calls a tool.
- decisions/README.md: main's table, with records 3, 18 and 20 amended
  and 21 added, and the branch's record 22 after it.
- computation-format.md: main's example reads the rows with an
  interaction function whose output is the rows themselves, so the
  workflow keeps them as `$data` in TypeScript.

Main's new code written in jq for the old language now runs in
TypeScript: the paging test's workflow and recall function in
asking-a-system/pages.test.ts. The computation example's test saves its
first function as the interaction function the page now shows. Both
sides added tests to raised-error.test.ts past its 300 lines, so the
asynchronous catch of a filter's raise moves into filter-verdicts.ts,
where only it is used, tested through filterVerdictsOf with a session
whose test breaks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant