Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 26 additions & 25 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,32 +13,31 @@ Callers import from `@corbits/embedding`. The barrel is the public surface:
| `embedTexts` | Batch texts, post sequentially, return `number[][]` in input order |
| `probeEmbedDims` | Embed one probe string; report the vector length that came back |
| `EmbedConfig` / `EmbedConfigSchema` | Endpoint, model, and optional knobs |
| `EmbedOptions` | Harness deps plus optional retry, `Retry-After` override, abort |
| `runJSONRequest` | Classified, retried JSON POST (sibling clients reuse this) |
| `extractRetryAfterMs` | Default `Retry-After` reader (seconds or HTTP-date) |
| `ModelRequestError` | Transport, HTTP, or non-JSON 200 — `reason` is the classified `InferenceError` |
| `EmbedOptions` | Optional harness deps, retry, `Retry-After` override, abort |
| `RequestDependencies` / `RetryAfterExtractor` | Types of the `EmbedOptions` fields |
| `EmbeddingRequestError` | Every request failure — `reason` is the classified `InferenceError` |

The JSON transport (`runJSONRequest`) is private.

Quickstart shape (matches the README):

```ts
import { createDefaultScheduler } from "@intx/inference";
import { embedTexts, probeEmbedDims } from "@corbits/embedding";

const deps = { fetch, scheduler: createDefaultScheduler() };
const [vector] = await embedTexts(
["hello world"],
{ baseURL: "http://localhost:11434/v1", model: "nomic-embed-text" },
{ deps },
);
import { embedTexts } from "@corbits/embedding";

const [vector] = await embedTexts(["hello world"], {
baseURL: "http://localhost:11434/v1",
model: "nomic-embed-text",
});
```

`probeEmbedDims(config, { deps })` is the same config and deps; it is not a
`probeEmbedDims(config)` is the same config and options; it is not a
separate protocol.

## Request dependencies

The transport reads only `fetch` and `scheduler` from the harness
(`RequestDependencies`). That is deliberately narrower than Interchange
(`RequestDependencies`), defaulting to global `fetch` and
`createDefaultScheduler()`. That is deliberately narrower than Interchange
`Dependencies`, whose `adapters` registry a caller should not have to
assemble in order to embed a string. A real `Dependencies` still satisfies
the pick structurally.
Expand Down Expand Up @@ -88,20 +87,22 @@ and `createDefaultRetryPolicy` (back off retryables, abort the rest). A 429
is therefore classified and backed off as a chat 429 would be; a
`credential_failure` aborts immediately.

`runJSONRequest` is duplicated, modulo comments, in `@corbits/reranking`. It
is not shared infrastructure yet — the intended home is `@intx/inference` as
a non-streaming sibling of `runInference`. Until then, `ModelRequestError` is
a distinct class in each package: `instanceof` does not hold across the two.
Catch on `error.name === "ModelRequestError"`.
Following Interchange, the `InferenceError` is carried as data on a thrown
`Error` subclass (`EmbeddingRequestError`), never thrown itself.

`runJSONRequest` is duplicated, modulo comments, in `@corbits/reranking`,
which still throws its own `ModelRequestError`; code using both packages
checks two error classes. The intended home for the transport is
`@intx/inference`, as a non-streaming sibling of `runInference`.

## Failure modes

- Transport, HTTP status, or a 200 whose body is not JSON →
`ModelRequestError` with classified `reason` and the URL.
- Malformed embeddings JSON, wrong count, bad index → thrown `Error` (not
`ModelRequestError`); the protocol succeeded, the payload did not.
- `batchSize` not an integer `>= 1` → `RangeError` before any request
(`i += size` of 0 would never advance).
`EmbeddingRequestError` with classified `reason` and the URL.
- Malformed embeddings JSON, wrong count, bad index →
`EmbeddingRequestError` with a `classifyProtocolMismatch` reason.
- Config failing `EmbedConfigSchema` (e.g. `batchSize` not an integer
`>= 1`) → arktype `TraversalError` before any request.
- Empty `texts` → `[]`, no request.
- Caller `signal` or per-attempt timeout (default 30s) abort the attempt;
aborting mid-delay wakes immediately and the next attempt fails its
Expand Down
9 changes: 3 additions & 6 deletions IMPLEMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,8 @@ Reply: `{ data: { index: number, embedding: number[] | string }[] }`.
Base64 strings decode with `atob` → `Uint8Array` → little-endian
`Float32Array` → `number[]`.

`embedTexts` also range-checks `batchSize` at runtime so a value that
bypassed the schema cannot stall `batches` (`i += size`).
`embedTexts` asserts `config` against `EmbedConfigSchema` first, so a
`batchSize` below 1 cannot stall `batches` (`i += size`).

## Errors and retry

Expand All @@ -92,16 +92,13 @@ bypassed the schema cannot stall `batches` (`i += size`).
3. network throw → `classifyNetworkError` or `classifyAbortError`
4. JSON parse fail → `classifyProtocolMismatch`
5. `retryPolicy` (default `createDefaultRetryPolicy`) → `abort` throws
`ModelRequestError`, else sleep `delayMs` on `deps.scheduler`
`EmbeddingRequestError`, else sleep `delayMs` on `deps.scheduler`

`extractRetryAfterMs` reads `Retry-After` as seconds or HTTP-date, clamped
at zero. Callers override via `EmbedOptions.extractRetryAfterMs` (same
signature as Interchange's unexported extractor; re-declared so this
package depends only on the published surface).

`ModelRequestError.name` is `"ModelRequestError"`. Discriminate on the
name when catching both this package and `@corbits/reranking`.

## Tests

`bun run test` is `bun test ./src`. Coverage is the client contract: empty
Expand Down
11 changes: 1 addition & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,16 +73,7 @@ scheduler), with optional `retryPolicy`, `extractRetryAfterMs`, and `signal`.

## Errors

Every failure mode — transport, HTTP status, or a 200 with an unexpected body — raises `ModelRequestError`, carrying the classified `InferenceError` as `reason` plus the request URL. The embedding and reranking packages each carry their own copy of this class while the shared transport is upstreamed, so code catching both discriminates on `error.name === "ModelRequestError"`.

## Lower-level: transport

The barrel also re-exports the one-shot JSON transport `embedTexts` is built on, for sibling clients that want the same classified, retried request path:

- `runJSONRequest` — one JSON POST with error classification and retry.
- `extractRetryAfterMs` — default `Retry-After` reader, overridable per call.
- `ModelRequestError` — error for every failure mode, with `reason` and URL.
- Types `RunRequestOptions` and `RetryAfterExtractor`.
Every request failure — transport, HTTP status, or a 200 with an unexpected body — throws `EmbeddingRequestError` (`extends Error`), carrying the classified `InferenceError` as `reason` and the request `url`. A config that fails `EmbedConfigSchema` throws arktype's `TraversalError` before any request.

## Interchange

Expand Down
33 changes: 9 additions & 24 deletions src/embed.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ import {
EmbedConfigSchema,
type EmbedConfig,
} from "./embed";
import { extractRetryAfterMs, ModelRequestError } from "./request";
import { EmbeddingRequestError } from "./request";

const CONFIG: EmbedConfig = {
baseURL: "https://embed.example/v1",
Expand Down Expand Up @@ -134,12 +134,12 @@ describe("embedTexts", () => {
it("rejects a batchSize that is not an integer >= 1 instead of looping forever", async () => {
// `batches` advances by `i += size`, so a size of 0 never advances and
// grows the batch list until OOM. Fractional sizes would also stall or
// slice wrongly. Schema and runtime both require integer >= 1.
// slice wrongly. The schema requires integer >= 1.
for (const batchSize of [0, -1, 0.5, Number.NaN, 1.5]) {
const { deps: d, calls } = deps([embedding(3)]);
await expect(
embedTexts(["a"], { ...CONFIG, batchSize }, { deps: d }),
).rejects.toBeInstanceOf(RangeError);
).rejects.toThrow(/batchSize/);
expect(calls).toHaveLength(0);
}
});
Expand All @@ -148,9 +148,9 @@ describe("embedTexts", () => {
const { deps: d } = deps([
Response.json({ data: [{ index: 0, embedding: true }] }),
]);
await expect(embedTexts(["a"], CONFIG, { deps: d })).rejects.toThrow(
/malformed embeddings response/,
);
const rejection = expect(embedTexts(["a"], CONFIG, { deps: d })).rejects;
await rejection.toBeInstanceOf(EmbeddingRequestError);
await rejection.toThrow(/malformed embeddings response/);
});
});

Expand Down Expand Up @@ -195,7 +195,7 @@ describe("retry behaviour inherited from the inference policy", () => {
new Response("unauthorized", { status: 401 }),
]);
await expect(embedTexts(["a"], CONFIG, { deps: d })).rejects.toBeInstanceOf(
ModelRequestError,
EmbeddingRequestError,
);
expect(calls).toHaveLength(1);
});
Expand Down Expand Up @@ -232,21 +232,6 @@ describe("retry-after handling", () => {
});
expect(seen).toEqual(["7"]);
});

it("reads the seconds form and tolerates an absent header", () => {
expect(extractRetryAfterMs(new Headers({ "retry-after": "2" }))).toBe(
2_000,
);
expect(extractRetryAfterMs(new Headers())).toBeUndefined();
});

it("never returns a negative delay for a date already past", () => {
expect(
extractRetryAfterMs(
new Headers({ "retry-after": "Wed, 21 Oct 2015 07:28:00 GMT" }),
),
).toBeGreaterThanOrEqual(0);
});
});

describe("reply integrity", () => {
Expand Down Expand Up @@ -319,7 +304,7 @@ describe("transport failure modes", () => {
}),
]);
await expect(embedTexts(["a"], CONFIG, { deps: d })).rejects.toBeInstanceOf(
ModelRequestError,
EmbeddingRequestError,
);
});

Expand Down Expand Up @@ -347,7 +332,7 @@ describe("transport failure modes", () => {
{ ...CONFIG, timeoutMs: 5 },
{ deps: d, signal: controller.signal },
),
).rejects.toBeInstanceOf(ModelRequestError);
).rejects.toBeInstanceOf(EmbeddingRequestError);
expect(calls).toHaveLength(1);
});
});
Loading
Loading