Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -154,6 +154,16 @@ one path, deltas to pixels:
synthetic in-progress message alongside the persisted timeline, until
the real message lands and replaces it.

**Provider tool-call quirks.** Some OpenAI-compatible local models emit
a declared tool call as a JSON object in assistant text instead of a
native tool-call field. Workbench's Ollama adapter wraps the platform
OpenAI adapter (composition, not a fork) and reclassifies that exact
object into structured tool-call events when the tool name was
declared for the request. Unknown names, extra keys, arrays, mixed
prose, and fenced JSON stay text so a legitimate JSON or code answer
is not stolen. The reactor and timeline still only see the usual text
versus tool-call events.

## Workbench Definition and hostless onboarding

A picker "template" is a shipped **Workbench Definition**, not a second
Expand Down
19 changes: 19 additions & 0 deletions IMPLEMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,6 +259,25 @@ non-empty owned list restricts candidates to those names; an empty owned
list keeps inherited discovery. Other providers still pin
`CATALOG_SEEDS[provider].models[0]`.

Ollama's OpenAI-compatible endpoint can also emit a declared tool call
as a JSON object in `content` instead of native `tool_calls`.
`@intx/inference`'s OpenAI parser only reads `delta.tool_calls`, so
that JSON would otherwise land on the timeline as assistant text and
the tool would never run. `@corbits/ollama-adapter`
(`reclassifyInlineToolJsonEvents` in
`packages/ollama-adapter/src/inline-tool-json.ts`) rewrites a content
stream that is exactly that object into `inference.tool_call.*`
events, gated on the tools declared for the request. Salvage requires
a plain object whose keys are only `name` plus exactly one of
`parameters` or `arguments` (optional `id`), and a `name` in the
declared set. Unknown names, no-tools requests, extra keys, arrays,
mixed prose, fenced JSON, and incomplete objects stay text.

The assistant definition (`workflows/assistant`) also tells Myra to
invoke tools only through tool calls — never by writing a JSON object
with a tool name into the reply — and not to `memory_search` a bare
greeting.

## Related docs

- [README.md](README.md) — quickstart, local setup, repo layout, e2e detail
Expand Down
5 changes: 4 additions & 1 deletion PRODUCT.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,10 @@ actually pulled, so Myra's first DM message is a real agent turn — even
on a machine that only has models like llama3.2 or qwen3. An inherited
catalog seed the instance never pulled is not the default when the
tenant already owns a pulled completion model. Embedding models never
become that default.
become that default. A local model that writes a tool it meant to
invoke as JSON in the reply still becomes a human turn: the person
sees Myra's words (and ordinary tool activity), not that JSON as the
message.

Create stays on `/new` (`apps/web/src/pages/new-workbench-picker.tsx`):
a prompt box is the primary act: typing a goal and submitting mints an
Expand Down
98 changes: 97 additions & 1 deletion packages/ollama-adapter/src/adapter.test.ts
Original file line number Diff line number Diff line change
@@ -1,7 +1,12 @@
import { describe, expect, test } from "bun:test";
import { createOpenAIAdapter } from "@intx/inference/providers";
import type { LastCycleSource } from "@intx/types/runtime";
import type { ConversationTurn, InferenceOptions } from "@intx/types/runtime";
import type {
ConversationTurn,
InferenceEvent,
InferenceOptions,
ToolDefinition,
} from "@intx/types/runtime";

import { createOllamaAdapter } from "./adapter";

Expand All @@ -21,10 +26,48 @@ const messages: ConversationTurn[] = [

const options: InferenceOptions = {};

const memorySearchTool: ToolDefinition = {
name: "memory_search",
description: "Search firm memory",
inputSchema: { type: "object" },
};

const CL_7186_PAYLOAD =
'{"name":"memory_search","parameters":{"query":"this person"}}';

function bodyOf(built: { body: string }): Record<string, unknown> {
return JSON.parse(built.body) as Record<string, unknown>;
}

function toolStarts(events: readonly InferenceEvent[]): InferenceEvent[] {
return events.filter((event) => event.type === "inference.tool_call.start");
}

function textDeltas(events: readonly InferenceEvent[]): InferenceEvent[] {
return events.filter((event) => event.type === "inference.text.delta");
}

function argumentFragments(events: readonly InferenceEvent[]): string {
return events
.filter((event) => event.type === "inference.tool_call.delta")
.map(
(event) => (event.data as { argumentFragment: string }).argumentFragment,
)
.join("");
}

function expectSharedToolCallId(events: readonly InferenceEvent[]): void {
const ids = events
.filter(
(event) =>
event.type === "inference.tool_call.start" ||
event.type === "inference.tool_call.delta",
)
.map((event) => (event.data as { callId: string }).callId);
expect(ids.length).toBeGreaterThan(0);
expect(new Set(ids)).toEqual(new Set(["ollama-inline-0"]));
}

describe("createOllamaAdapter", () => {
test("no override configured leaves the body equivalent to the built-in adapter's", () => {
const wrapped = createOllamaAdapter(source, undefined);
Expand Down Expand Up @@ -84,4 +127,57 @@ describe("createOllamaAdapter", () => {
expect(typeof wrapped.extractRetryAfterMs).toBe("function");
expect(typeof wrapped.extractPacingDelayMs).toBe("function");
});

test("parseResponse salvages declared inline memory_search JSON into a tool call", () => {
const wrapped = createOllamaAdapter(source, undefined);
wrapped.buildRequest(messages, "gpt-oss:20b", {
tools: [memorySearchTool],
});
const events = [
...wrapped.parseResponse(
JSON.stringify({
choices: [{ delta: { content: CL_7186_PAYLOAD } }],
}),
),
...wrapped.parseResponse(
JSON.stringify({
choices: [{ delta: {}, finish_reason: "stop" }],
}),
),
];
expect(textDeltas(events)).toEqual([]);
expect((toolStarts(events)[0]?.data as { name: string }).name).toBe(
"memory_search",
);
expectSharedToolCallId(events);
expect(JSON.parse(argumentFragments(events))).toEqual({
query: "this person",
});
});

test("parseJSONResponse salvages declared inline memory_search JSON into a tool call", () => {
const wrapped = createOllamaAdapter(source, undefined);
wrapped.buildRequest(messages, "gpt-oss:20b", {
tools: [memorySearchTool],
});
const events = wrapped.parseJSONResponse(
JSON.stringify({
object: "chat.completion",
choices: [
{
message: { role: "assistant", content: CL_7186_PAYLOAD },
},
],
usage: { prompt_tokens: 1, completion_tokens: 1 },
}),
);
expect(textDeltas(events)).toEqual([]);
expect((toolStarts(events)[0]?.data as { name: string }).name).toBe(
"memory_search",
);
expectSharedToolCallId(events);
expect(JSON.parse(argumentFragments(events))).toEqual({
query: "this person",
});
});
});
32 changes: 27 additions & 5 deletions packages/ollama-adapter/src/adapter.ts
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,12 @@ import {
type OllamaAdapterOverride,
} from "./overrides";
import { createThinkSplitState, reclassifyThinkingEvents } from "./think-tags";
import {
createInlineToolJsonState,
reclassifyInlineToolJsonEvents,
responseChunkIsTerminal,
setDeclaredToolNames,
} from "./inline-tool-json";

type OllamaChatBody = {
options?: Record<string, unknown>;
Expand Down Expand Up @@ -85,16 +91,32 @@ export const createOllamaAdapter: AdapterFactory = (
// chunk of one response without leaking state between requests.
const streamThinkState = createThinkSplitState();
const jsonThinkState = createThinkSplitState();
const streamInlineState = createInlineToolJsonState();
const jsonInlineState = createInlineToolJsonState();
return {
...inner,
buildRequest: (messages, model, options) =>
applyOverride(
buildRequest: (messages, model, options) => {
setDeclaredToolNames(streamInlineState, options.tools);
setDeclaredToolNames(jsonInlineState, options.tools);
return applyOverride(
inner.buildRequest(messages, model, options),
resolveOverride(config, model),
),
);
},
parseResponse: (sseData) =>
reclassifyThinkingEvents(inner.parseResponse(sseData), streamThinkState),
reclassifyInlineToolJsonEvents(
reclassifyThinkingEvents(
inner.parseResponse(sseData),
streamThinkState,
),
streamInlineState,
{ flush: responseChunkIsTerminal(sseData) },
),
parseJSONResponse: (body) =>
reclassifyThinkingEvents(inner.parseJSONResponse(body), jsonThinkState),
reclassifyInlineToolJsonEvents(
reclassifyThinkingEvents(inner.parseJSONResponse(body), jsonThinkState),
jsonInlineState,
{ flush: true },
),
};
};
7 changes: 7 additions & 0 deletions packages/ollama-adapter/src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,3 +11,10 @@ export {
reclassifyThinkingEvents,
type ThinkSplitState,
} from "./think-tags";
export {
createInlineToolJsonState,
reclassifyInlineToolJsonEvents,
responseChunkIsTerminal,
setDeclaredToolNames,
type InlineToolJsonState,
} from "./inline-tool-json";
Loading
Loading