Chat: legible tool activity in the conversation - #190
Merged
Merged
Conversation
Covers the presentation of a turn's tool calls end to end: the sentence a call renders as (no tool identifier, no argument JSON), the plain-text detail it opens onto, a failure that says why, a call still running in the present tense, and a run of consecutive calls folded into one round. The existing live-strip suites move to the same idiom: a tool call now asserts on the phrase a reader sees rather than on the raw tool name, and tool.done carries whether the call actually worked.
A turn's tool calls used to reach the reader as raw material. In the
transcript each call rendered as its own block titled with the tool's
identifier run through a humaniser ("Mcp read", "Post message (Slack)"),
and opening one showed its arguments and its result as JSON.stringify
output. Five consecutive calls stacked as five blocks between the question
and the answer. In the live strip a call showed the bare tool name, and a
call that failed settled into the same quiet marker as one that succeeded
— a failure was invisible.
Now one vocabulary serves both. tool-activity.ts turns a call into a
sentence in the tense its status calls for ("Searching the web for
'pricing'" while it runs, "Searched the web for 'pricing'" once it
settles), resolving the generic MCP dispatch tools to the tool they
actually invoked, and turns a result into plain text — prose when the
tool returned any, a count when it returned a list, and never JSON. A
failure always says something, "No reason given." when the tool offered
nothing.
Consecutive calls fold into one round that collapses to a single line: the
step currently working while the turn is open, then a count once settled,
saying plainly if any of them didn't work. A round with a failure opens
itself; everything else stays closed, so the reply stays above its own
machinery.
The live strip and the transcript now render through the same rows, so a
call does not restyle itself the moment the turn ends. react-ui's
ToolBlock is no longer used here: it renders one call at a time and
stringifies its arguments and output.
Adds the canon the chat surface now implements — sentences over identifiers, plain text over JSON, tense that follows state, consecutive calls folded into a collapsed round, and a failure that says why.
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What a turn with tool calls looked like
Before. In the transcript, each
tool-tracepart rendered as its ownreact-ui
ToolBlock, titled with the tool's identifier run through ahumaniser —
mcp_readbecame "Mcp read",slack__post_messagebecame"Post message (Slack)", so every MCP call in a conversation read alike.
Opening one showed
JSON.stringify(input, null, 2)and the resultstringified the same way. A failed call opened itself by default, so a
raw JSON error blob sat expanded inline in the conversation. Five
consecutive calls stacked as five blocks between the question and the
answer.
In the live strip, a call showed the bare tool name plus elapsed seconds,
and
tool.donemarked it "done" regardless ofisError— a tool thatfailed settled into the same quiet marker as one that succeeded. A failure
was invisible until the turn ended.
After. One vocabulary (
tool-activity.ts) serves both surfaces:Searched the web for "pricing",Wrote a file — report.md,Posted a message in Slack #general,Saved an issue in Linear. The generic MCP dispatch toolsresolve to the tool they actually invoked. No identifier, no namespace,
no underscore reaches the DOM.
dumped. A result becomes plain text: the prose the tool returned, or a
count (
2 results.,Nothing found.) when it returned a list. Anopaque object yields no detail at all rather than a stringified blob.
Searching the web for "x"whileit runs,
Searched the web for "x"once it settles. A running rowcarries elapsed seconds; a settled one drops the timer.
failed, is the only state here that earnscolour, and opens onto the reason in words. A tool that failed without
saying why still says
No reason given.the step currently working while the turn is open, then
3 stepsoncesettled, or
3 steps, 2 didn't workwhen some failed. Collapsed bydefault so the reply stays above its own machinery; a round containing a
failure opens itself.
The live strip and the persisted transcript now render through the same
rows, so a call does not restyle itself the moment its turn ends.
Why
ToolBlockis no longer used herereact-ui's
ToolBlockrenders one call at a time and stringifies both itsarguments and its output — the two things this change exists to stop. The
row idiom it is replaced with is workbench composition (grouping into
rounds, the consumer-language vocabulary) and matches the chat surface's
own radius-0 / Phosphor / quiet-chrome canon.
ToolBlockstays untouchedupstream;
toReactUiToolTraceis removed here rather than left beside thenew path.
toReactUiReasoningand react-ui's reasoning disclosure areunchanged.
Test plan
bun run typecheckinpackages/chat-ui— cleanbun testinpackages/chat-ui— 648 pass, 0 faileslint/prettierclean on touched filessrc/tool-activity.test.ts(28 cases): identity resolution, phrasingin both tenses, output-to-plain-text extraction, failure summaries, round
summaries, part grouping.
test/tool-activity-view.test.tsx(8 cases, mounted for real):success reads as a sentence with no identifier and no
{in the DOM;detail stays closed until clicked; a failure marks itself and opens onto
its reason; a failure with nothing to say still says so; a running call
is present-tense and offers no disclosure; three consecutive calls
collapse to
3 stepsand open onto all three; a still-working roundshows the working step; a round with a failure names it.
Not verified
testing live on :3000; this lane did not touch that stack or its DB).
Everything above is component/state level.
packages/folded-runswas audited (inferenceDoneBlocks,toolDoneResult) and found correct for this work; it is unchanged.streaming-reply.tsis untouched.