From 44a8c2ed5c774b4da1c793fa8c52d90c9e2dfda0 Mon Sep 17 00:00:00 2001 From: Sawyer Cutler Date: Thu, 24 Sep 2026 23:03:10 -0700 Subject: [PATCH 1/3] docs: rewrite README quickstart and move dev docs to CONTRIBUTING The quickstart embeds two strings with a literal config and prints the count and width; no env loader, spread, or manual deps. Errors document EmbeddingRequestError, the internal transport section is gone, the @intx/inference paragraph appears once under Using with Interchange, and Versioning/Development move to CONTRIBUTING.md. Also closes CL-9093, CL-9090, CL-9098. --- CONTRIBUTING.md | 18 ++++++++++++ README.md | 78 +++++++++++-------------------------------------- 2 files changed, 35 insertions(+), 61 deletions(-) create mode 100644 CONTRIBUTING.md diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..9379343 --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,18 @@ +# Contributing + +## Development + +```bash +bun install +bun run typecheck +bun run lint +bun run test +``` + +`bun run test` includes a live `/v1/embeddings` round trip that runs when +`EMBEDDING_E2E_BASE_URL` (default `http://localhost:11434/v1`) is reachable +and skips otherwise. `EMBEDDING_E2E_MODEL` defaults to `nomic-embed-text`. + +## Versioning + +Semver. Releases run `bun run build && npm publish` with green CI. diff --git a/README.md b/README.md index 03e37b9..a3c3dcf 100644 --- a/README.md +++ b/README.md @@ -8,8 +8,6 @@ OpenAI, Ollama, TEI, vLLM, and Jina all serve the same wire shape. This package - `encodingFormat` — `float` (default) or compact `base64`. - `batchSize` — inputs per request (default 32). Tune it to provider caps and per-request token ceilings. -Transport, error classification, and retry build on `@intx/inference`: requests go through the shared `deps.fetch` path, failures classify into `InferenceError` just like chat calls, and `createDefaultRetryPolicy` guides backoff. A 429 from an embeddings endpoint behaves like one from a chat endpoint, and credential failures short-circuit. - ## Install ```bash @@ -20,76 +18,34 @@ Runs on Bun >= 1.2 or Node >= 24. The built `dist/` entry is the default import. ## Quickstart -This package never touches credential storage or config sources — it takes an -already-built `EmbedConfig`. Where that config comes from is the host's -choice; `@corbits/memory`, for example, reads it straight from -`EMBED_BASE_URL`/`EMBED_MODEL`/`EMBED_API_KEY` env vars, all-or-nothing on the -first two. A host wraps `embedTexts` in a function of its own that builds the -config once (at startup, or per call) and probes dimensionality up front so a -later model swap is caught as a migration, not a silent width mismatch. - ```ts -import { createDefaultScheduler } from "@intx/inference"; -import { - embedTexts, - probeEmbedDims, - type EmbedConfig, -} from "@corbits/embedding"; - -export function loadEmbedConfig(): EmbedConfig | undefined { - const baseURL = process.env["EMBED_BASE_URL"]; - const model = process.env["EMBED_MODEL"]; - if (baseURL === undefined || model === undefined) return undefined; - - const apiKey = process.env["EMBED_API_KEY"]; - return { baseURL, model, ...(apiKey !== undefined ? { apiKey } : {}) }; -} - -export async function embed( - texts: readonly string[], - config: EmbedConfig, -): Promise { - const deps = { fetch, scheduler: createDefaultScheduler() }; - return embedTexts(texts, config, { deps }); -} - -// At startup: fail fast if the configured model's width doesn't match -// what's already stored. -const config = loadEmbedConfig(); -if (config !== undefined) { - const deps = { fetch, scheduler: createDefaultScheduler() }; - const dims = await probeEmbedDims(config, { deps }); -} +import { embedTexts } from "@corbits/embedding"; + +const config = { + baseURL: "http://localhost:11434/v1", + model: "nomic-embed-text", +}; + +const vectors = await embedTexts(["a", "b"], config); +console.log(vectors.length, vectors[0]?.length); // 2 768 ``` -`embedTexts(texts, config, options)` batches sequentially, posts each batch to -`{baseURL}/embeddings`, and returns vectors in input order. `options` carries -`deps` (a host's real `fetch` plus `createDefaultScheduler()`, the production -scheduler), with optional `retryPolicy`, `extractRetryAfterMs`, and `signal`. +`embedTexts(texts, config, options?)` batches sequentially, posts each batch to +`{baseURL}/embeddings`, and returns vectors in input order. `options` is +optional: `deps` (defaults to global `fetch` and `createDefaultScheduler()`), +`retryPolicy`, `extractRetryAfterMs`, and `signal`. ## Dimensionality -`probeEmbedDims(config, { deps })` embeds one probe string and reports the length received. Dimensionality follows the provider and the `dimensions` setting, so callers that persist vectors discover it at startup and treat a model swap as a migration. Representative widths: 1536 for OpenAI `text-embedding-3-small`, 768 for `nomic-embed-text`, 384 for `bge-small-en-v1.5`. +`probeEmbedDims(config)` embeds one probe string and reports the length received. Dimensionality follows the provider and the `dimensions` setting, so callers that persist vectors discover it at startup and treat a model swap as a migration. Representative widths: 1536 for OpenAI `text-embedding-3-small`, 768 for `nomic-embed-text`, 384 for `bge-small-en-v1.5`. ## Errors -Every request failure — transport, HTTP status, or a 200 with an unexpected body — throws `EmbeddingRequestError` (`extends Error`), carrying the classified `InferenceError` as `reason` and the request `url`. A config that fails `EmbedConfigSchema` throws arktype's `TraversalError` before any request. - -## Interchange +Every failure — transport, HTTP status, or a 200 with an unexpected body — throws `EmbeddingRequestError` (`extends Error`), carrying the classified `InferenceError` as `reason` and the request `url`. A config that fails `EmbedConfigSchema` throws before any request. -Built on the shared `@intx/inference` transport, so an embeddings call fails and retries the same way a chat call does: same `deps.fetch` path, same `InferenceError` classification, same `createDefaultRetryPolicy` backoff. A retrieval pipeline typically pairs `probeEmbedDims` at startup with `embedTexts` batches at ingest and query time, as described above. +## Using with Interchange -## Versioning - -Semver. Releases run `bun run build && npm publish` with green CI. - -## Development - -```bash -bun install -bun test ./src -bunx tsc --noEmit -``` +Transport, error classification, and retry build on `@intx/inference`: requests go through `deps.fetch`, failures classify into `InferenceError` just like chat calls, and `createDefaultRetryPolicy` guides backoff. A 429 from an embeddings endpoint behaves like one from a chat endpoint, and credential failures short-circuit. `@intx/inference` and `@intx/types` are peer dependencies, so the host's Interchange version is the one used. ## License From a74f9818eb6c0ce29f7905ffdb9fecce5a789d72 Mon Sep 17 00:00:00 2001 From: Sawyer Cutler Date: Fri, 25 Sep 2026 06:54:32 -0700 Subject: [PATCH 2/3] docs: install the @intx peers alongside in the README --- CONTRIBUTING.md | 12 +++++++++--- README.md | 11 ++++++----- 2 files changed, 15 insertions(+), 8 deletions(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 9379343..eeb13b5 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -9,9 +9,15 @@ bun run lint bun run test ``` -`bun run test` includes a live `/v1/embeddings` round trip that runs when -`EMBEDDING_E2E_BASE_URL` (default `http://localhost:11434/v1`) is reachable -and skips otherwise. `EMBEDDING_E2E_MODEL` defaults to `nomic-embed-text`. +`bun run test` is hermetic. `bun run test:live` runs a `/v1/embeddings` round +trip and skips unless `EMBEDDING_E2E_BASE_URL` is set; point it at a remote +server rather than loading models locally: + +```bash +EMBEDDING_E2E_BASE_URL=http://100.113.184.123:11434/v1 bun run test:live +``` + +`EMBEDDING_E2E_MODEL` defaults to `nomic-embed-text`. ## Versioning diff --git a/README.md b/README.md index a3c3dcf..eedc611 100644 --- a/README.md +++ b/README.md @@ -11,10 +11,10 @@ OpenAI, Ollama, TEI, vLLM, and Jina all serve the same wire shape. This package ## Install ```bash -bun add @corbits/embedding +bun add @corbits/embedding @intx/inference@^0.4.0 @intx/types@^0.4.0 ``` -Runs on Bun >= 1.2 or Node >= 24. The built `dist/` entry is the default import. +Runs on Bun >= 1.2 or Node >= 24. ## Quickstart @@ -32,7 +32,8 @@ console.log(vectors.length, vectors[0]?.length); // 2 768 `embedTexts(texts, config, options?)` batches sequentially, posts each batch to `{baseURL}/embeddings`, and returns vectors in input order. `options` is -optional: `deps` (defaults to global `fetch` and `createDefaultScheduler()`), +optional: `deps` (defaults to global `fetch` and `@intx/inference`'s +`createDefaultScheduler()`), `retryPolicy`, `extractRetryAfterMs`, and `signal`. ## Dimensionality @@ -41,11 +42,11 @@ optional: `deps` (defaults to global `fetch` and `createDefaultScheduler()`), ## Errors -Every failure — transport, HTTP status, or a 200 with an unexpected body — throws `EmbeddingRequestError` (`extends Error`), carrying the classified `InferenceError` as `reason` and the request `url`. A config that fails `EmbedConfigSchema` throws before any request. +Every request failure — transport, HTTP status, or a 200 with an unexpected body — throws `EmbeddingRequestError` (`extends Error`), carrying the classified `InferenceError` as `reason` and the request `url`. A config that fails `EmbedConfigSchema` throws arktype's `TraversalError` before any request. ## Using with Interchange -Transport, error classification, and retry build on `@intx/inference`: requests go through `deps.fetch`, failures classify into `InferenceError` just like chat calls, and `createDefaultRetryPolicy` guides backoff. A 429 from an embeddings endpoint behaves like one from a chat endpoint, and credential failures short-circuit. `@intx/inference` and `@intx/types` are peer dependencies, so the host's Interchange version is the one used. +Requests go through `deps.fetch`, failures are classified into `@intx/inference`'s `InferenceError`, and `createDefaultRetryPolicy` decides retries: a 429 backs off and retries, a credential failure does not. `@intx/inference` and `@intx/types` are peer dependencies, so the host's Interchange version is the one used. ## License From 605e64e3decfaa3dfea508e5022ac29da976909d Mon Sep 17 00:00:00 2001 From: Sawyer Cutler Date: Fri, 25 Sep 2026 10:09:54 -0700 Subject: [PATCH 3/3] docs: rewrite the README for first-time readers --- README.md | 85 ++++++++++++++++++++++++++++++++++++++++++------------- 1 file changed, 65 insertions(+), 20 deletions(-) diff --git a/README.md b/README.md index eedc611..5258dac 100644 --- a/README.md +++ b/README.md @@ -1,12 +1,14 @@ # @corbits/embedding -One API for every embedding provider. Point at any OpenAI-compatible `/v1/embeddings` endpoint, swap models with a config change, and get ordered vectors back. +Batched text embeddings over any OpenAI-compatible `/v1/embeddings` endpoint, built on `@intx/inference` retries and error classification. A retrieval building block for Corbits and Interchange agents that also works standalone. -OpenAI, Ollama, TEI, vLLM, and Jina all serve the same wire shape. This package is one code path over that shape, with parameters for what varies between providers: +## Why @corbits/embedding? -- `dimensions` — Matryoshka truncation for models that support it, including OpenAI `text-embedding-3` (down to 256), Jina v3 (32), and Jina v4 (128). Sent only when set, so models without MRL stay happy. -- `encodingFormat` — `float` (default) or compact `base64`. -- `batchSize` — inputs per request (default 32). Tune it to provider caps and per-request token ceilings. +1. **One client for every provider.** OpenAI, Ollama, TEI, vLLM and Jina all speak the same wire format. Switching models is a config change. +2. **Order-safe batching.** Inputs are split into batches, and vectors come back in input order, placed by the index each response echoes. Short or duplicate replies throw instead of misaligning results. +3. **Interchange retry semantics.** Requests run through `@intx/inference`, so a 429 backs off and a 401 fails the same way it does for inference calls. Failures throw a typed `EmbeddingRequestError`. + +It does not store or search vectors. For that, pair it with [`@corbits/memory`](https://github.com/corbitsdev/corbits-memory). ## Install @@ -18,36 +20,79 @@ Runs on Bun >= 1.2 or Node >= 24. ## Quickstart +Needs a local Ollama with `nomic-embed-text` pulled. + ```ts import { embedTexts } from "@corbits/embedding"; -const config = { +const vectors = await embedTexts(["a", "b"], { baseURL: "http://localhost:11434/v1", model: "nomic-embed-text", -}; - -const vectors = await embedTexts(["a", "b"], config); +}); console.log(vectors.length, vectors[0]?.length); // 2 768 ``` -`embedTexts(texts, config, options?)` batches sequentially, posts each batch to -`{baseURL}/embeddings`, and returns vectors in input order. `options` is -optional: `deps` (defaults to global `fetch` and `@intx/inference`'s -`createDefaultScheduler()`), -`retryPolicy`, `extractRetryAfterMs`, and `signal`. +## Where it fits + +[Interchange](https://github.com/faremeter/interchange) runs AI agents as principals (accounts that hold their own identity, permissions and credentials). Corbits packages add what an agent product needs around it. + +- **Runs in:** any process: the Interchange hub (the multi-tenant control plane), an agent sidecar (the runtime next to each agent), or a plain script. No hub is required. +- **Plugs into:** [`@intx/inference`](https://github.com/faremeter/interchange) for transport, retries and error classification, and [`@intx/types`](https://github.com/faremeter/interchange). +- **Pairs with:** [`@corbits/reranking`](https://github.com/corbitsdev/corbits-reranking) to reorder results and [`@corbits/memory`](https://github.com/corbitsdev/corbits-memory) to store and search vectors. + +## Reference + +| Export | Description | +| -------------------------------------------- | ----------------------------------------------------------- | +| `embedTexts(texts, config, options?)` | Embeds the texts and returns one vector per text, in order. | +| `probeEmbedDims(config, options?)` | Embeds one probe string and returns the vector dimension. | +| `EmbedConfigSchema`, `EmbedConfig` | Schema and type for `config`. | +| `EmbedOptions` | Type for `options`. | +| `EmbeddingRequestError` | Thrown when a request fails. | +| `RequestDependencies`, `RetryAfterExtractor` | Types for `options.deps` and `options.extractRetryAfterMs`. | -## Dimensionality +### Config -`probeEmbedDims(config)` embeds one probe string and reports the length received. Dimensionality follows the provider and the `dimensions` setting, so callers that persist vectors discover it at startup and treat a model swap as a migration. Representative widths: 1536 for OpenAI `text-embedding-3-small`, 768 for `nomic-embed-text`, 384 for `bge-small-en-v1.5`. +| Field | Type | Description | +| ---------------- | ---------------------- | ------------------------------------------------------------------------ | +| `baseURL` | `string` | Server root with its version path, such as `http://host:11434/v1`. | +| `model` | `string` | Model name. | +| `apiKey` | `string?` | Sent as a bearer token. | +| `dimensions` | `number?` | Requested output dimension. Sent only when set; not all models honor it. | +| `encodingFormat` | `"float" \| "base64"?` | Wire format. Defaults to `float`. Results are always `number[]`. | +| `batchSize` | `number?` | Texts per request. Defaults to 32. | +| `timeoutMs` | `number?` | Per-request timeout. Defaults to 30000. | -## Errors +### Options -Every request failure — transport, HTTP status, or a 200 with an unexpected body — throws `EmbeddingRequestError` (`extends Error`), carrying the classified `InferenceError` as `reason` and the request `url`. A config that fails `EmbedConfigSchema` throws arktype's `TraversalError` before any request. +| Field | Description | +| --------------------- | ------------------------------------------------------------------------------- | +| `deps` | `{ fetch, scheduler }`. Defaults to global `fetch` and Interchange's scheduler. | +| `retryPolicy` | Retry policy. Defaults to Interchange's policy. | +| `extractRetryAfterMs` | Reads the server's retry delay. Defaults to parsing `Retry-After`. | +| `signal` | Aborts all pending requests. | + +### Errors + +A failed request throws `EmbeddingRequestError` with a classified `reason` and the request `url`. An invalid config throws before any request is sent. + +### Dimensions + +Each model has a fixed output dimension: 768 for `nomic-embed-text`, 1536 for OpenAI `text-embedding-3-small`. If you store vectors, call `probeEmbedDims` at startup. Changing models changes the dimension, so treat it as a schema migration. ## Using with Interchange -Requests go through `deps.fetch`, failures are classified into `@intx/inference`'s `InferenceError`, and `createDefaultRetryPolicy` decides retries: a 429 backs off and retries, a credential failure does not. `@intx/inference` and `@intx/types` are peer dependencies, so the host's Interchange version is the one used. +`@intx/inference` and `@intx/types` are peer dependencies, so your host's Interchange version supplies them. + +To share your host's retry scheduler, pass it as `options.deps.scheduler` along with `fetch`. + +## Upgrading from 0.1 + +- Install `@intx/inference` and `@intx/types` (^0.4.0) yourself. They are now peer dependencies. +- `ModelRequestError` is renamed to `EmbeddingRequestError`. +- `runJSONRequest`, `extractRetryAfterMs` and `RunRequestOptions` are no longer exported. To change how retry waits are read, pass `options.extractRetryAfterMs`. +- Config and returned vectors are unchanged. ## License -LGPL-2.1-only — see [`LICENSE`](LICENSE). +[LGPL-2.1-only](https://github.com/corbitsdev/corbits-embedding/blob/main/LICENSE)