A context service for agents. It ingests documents and Agent Skill bundles, digests them, and republishes them as a single addressable virtual filesystem that an agent browses lazily — reading only what it needs, when it needs it.
Run it over stdio or authenticated Streamable HTTP. The HTTP service accepts file uploads; the Docker image includes local Tesseract OCR, and Ollama can provide local embeddings. Full-text search works without an embedding model or API key.
The design goal is progressive disclosure: a cheap always-available catalog at level 0, targeted reads at level 1, deep bundle files at level 2. Nothing enters an agent's context unless the agent asked for it.
cx_ls("/") → every skill and collection, ≤2k tokens
cx_search([...]) → ranked snippets with citable ctx:// links
cx_read(uri) → one file, or one chunk plus its neighbours
MCP Resources are application-driven, not model-driven. The host application
decides how to incorporate resource context, and in Claude Code resources are
pulled in by a human typing @server:resource. An agent cannot autonomously
list or read Resources. A Resources-only VFS would be inert for the primary
use case.
So contextual ships two surfaces over one URI namespace:
| Surface | Controlled by | Purpose |
|---|---|---|
Tools (cx_ls, cx_glob, cx_grep, cx_read, cx_search, cx_skill) |
the model | Agent autonomy — this is how the VFS actually gets used |
Resources (resources/list, /read, /templates/list, subscriptions) |
the host app / human | Stable addressability, @-mentions, citability, change notifications |
They are not two systems. Tool results return MCP resource_link content blocks
pointing back into the Resources namespace, so a search hit is a citation the
agent or the human can follow. Per spec, resource links returned by tools need
not appear in resources/list — which is what lets us cite chunk-level URIs
without listing millions of them.
See USAGE.md for the complete setup guide, MCP tool examples, configuration, and troubleshooting.
An optional .env.example shows settings for local embeddings
and OCR. Copy it to .env after installing Ollama, Tesseract, and Poppler;
the setup steps are in USAGE.md.
Running from source requires Bun 1.4.2 or later (pinned in .bun-version)
and Postgres with pgvector and pg_trgm. The commands below use Docker Compose
to provide Postgres 17. To run the entire service in Docker, use the
container setup; no host Bun installation is needed.
bun install
docker compose up -d # postgres + pgvector on :55432
bun run migrate
bun run src/cli/contextual.ts add ./path/to/skill-bundle
bun run src/cli/contextual.ts add ./report.pdf --collection handbook
bun run src/cli/contextual.ts listRegister it with Claude Code by copying .mcp.json into your project and
replacing /absolute/path/to/contextual with this checkout's real path. MCP
hosts spawn args as literal argv, so shell syntax such as ${VAR:-default} is
not expanded and would be passed through verbatim. Then ask a question the corpus
answers. The agent should reach it on its own via cx_ls → cx_search →
cx_read, with no @-mention.
Semantic search can use a local model through Ollama:
ollama pull qwen3-embedding:0.6b
export CONTEXTUAL_EMBED_PROVIDER=ollama
export CONTEXTUAL_EMBED_MODEL=qwen3-embedding:0.6b
bun run src/cli/contextual.ts reindexKeep Ollama running and set the same environment on the MCP server. The service requests and validates the 1024-dimensional vectors used by the database. No API key is needed. See local setup for startup and configuration.
Alternatively, use Voyage:
export VOYAGE_API_KEY=... # voyage-4, 1024 dims
export CONTEXTUAL_EMBED_PROVIDER=voyage
export CONTEXTUAL_EMBED_MODEL=voyage-4
bun run src/cli/contextual.ts reindexWithout an embedding provider, everything still works — retrieval is full-text only, and skills need no embeddings at all.
The corpus pins the first embedding model that touches it (settings.embed_model).
Changing the embedding provider or CONTEXTUAL_EMBED_MODEL afterwards is refused until reindex --all
re-embeds everything, and a query embedded by a different model is skipped rather
than ranked against vectors from another space. Dimensions are fixed at 1024 by
the schema, so a model with different dimensions needs a migration, not an env var.
Full-text search uses one Postgres text-search configuration, chosen for a
new database with CONTEXTUAL_FTS_LANGUAGE (default english). It is baked
into the generated chunk tsvector column, so changing it later means a fresh database.
To serve clients through a URL:
bun run src/cli/contextual.ts serve --transport http --port 3000Connect to http://127.0.0.1:3000/mcp. Stdio remains the default. HTTP supports
current MCP requests and SSE subscriptions, plus stateless 2025-era clients.
Legacy HTTP clients have no session-based resource subscriptions; GET/DELETE
session operations return 405. All six tools and resource reads are available.
Set CONTEXTUAL_HTTP_TOKEN to require bearer authentication. Non-loopback binds
require a token; wildcard binds such as --host 0.0.0.0 also require
CONTEXTUAL_HTTP_ALLOWED_HOSTS (comma-separated hostnames/IPs without ports).
Host and Origin validation guard the endpoint. For remote use, terminate HTTPS
at a reverse proxy and preserve SSE streaming. See USAGE.md for
host configuration and a Deep Agents HTTP example.
The same HTTP service accepts multipart uploads at POST /api/ingest:
curl --fail-with-body http://127.0.0.1:3000/api/ingest \
-F 'files=@./report.pdf' \
-F 'collection=handbook'Repeat files for a batch; documents and .zip/.skill bundles use the existing
ingest pipeline. Optional fields are collection, force, and lenient.
The synchronous JSON response contains per-source results and citable URIs:
200 for success, 207 for mixed results, and 422 when all sources failed or need
OCR. The HTTP bearer token also authorizes uploads.
Originals are removed after processing; normalized content and assets persist.
Uploads have a default total limit of 32 MiB and 20 files, controlled by
CONTEXTUAL_UPLOAD_MAX_BYTES and CONTEXTUAL_UPLOAD_MAX_FILES. One upload runs
at a time per listener (other uploads receive 429); MCP reads remain available.
See USAGE.md for response fields,
retry behavior, and curl/Python examples.
The service image includes Tesseract 5, English language data, and Poppler.
It runs as UID/GID 10001, enables local OCR, and exposes HTTP on port 3000.
The optional Compose app profile starts it alongside Postgres:
export CONTEXTUAL_HTTP_TOKEN='replace-with-a-long-random-token'
docker compose --profile app build app
docker compose --profile app run --rm app migrate
docker compose --profile app up -d appUpload to http://127.0.0.1:3000/api/ingest with the bearer token:
curl --fail-with-body http://127.0.0.1:3000/api/ingest \
-H "Authorization: Bearer $CONTEXTUAL_HTTP_TOKEN" \
-F 'files=@./scanned.pdf' -F 'collection=handbook'Compose uses full-text search by default unless CONTEXTUAL_EMBED_PROVIDER
is set in your environment or .env. For Ollama, set
CONTEXTUAL_EMBED_PROVIDER=ollama and an address reachable
from the container in CONTEXTUAL_OLLAMA_URL. The Compose default uses
host.docker.internal for Docker Desktop; Linux and Kubernetes deployments
should supply their Ollama service address.
For a pod, provide CONTEXTUAL_DATABASE_URL, CONTEXTUAL_HTTP_TOKEN, and
CONTEXTUAL_HTTP_ALLOWED_HOSTS, run migrations before serving a new database,
and mount writable storage at /app/blobs. A read-only root filesystem also
needs writable /tmp for OCR and native parser files. See
Docker and Kubernetes setup for details.
Install Tesseract and Poppler, then enable local OCR for the CLI or server:
# Debian / Ubuntu
sudo apt-get update
sudo apt-get install -y --no-install-recommends tesseract-ocr tesseract-ocr-eng poppler-utils
# macOS: brew install tesseract poppler
export CONTEXTUAL_OCR=local
bun run src/cli/contextual.ts add ./scanned.pdf --collection handbook --forceEnglish is the default. Additional languages require installed Tesseract data
and CONTEXTUAL_OCR_LANGUAGE, for example eng+deu. OCR language selection is
independent of the database's full-text search language.
ctx://index # level-0 catalog
ctx://skills/{skill}/SKILL.md
ctx://skills/{skill}/references/{path}
ctx://skills/{skill}/scripts/{path}
ctx://docs/{collection}/{path} # normalized markdown
ctx://docs/{collection}/{path}#chunk={id} # link-only, returned by search
RAG corpora and skills share one tree because a skill bundle is a directory and
a corpus is a directory. One cx_read works on both.
This is the product, not an optimization, so it is enforced in code:
- L0 —
ctx://index/cx_ls("/")— every skill's name and description, every collection's summary and doc count. Onlyreadysources count as documents; a PDF still requiring OCR is reported separately as not searchable, so the catalog never promises text it does not have. Hard budget: ≤2k tokens. When the corpus outgrows it, descriptions are clipped before any entry is dropped, and anything held back is stated rather than silently omitted. - L1 —
cx_skill(name)orcx_search(...)— SKILL.md plus a file manifest, or ranked snippets withresource_links. No full documents. - L2 —
cx_read(uri)— one reference file, or one chunk plus neighbours. Binary assets (logos, screenshots, extracted images) come back inline as an image or embedded-resource block up to 3 MiB, because an agent cannot fetch Resources on its own.resources/readhas its own, higher ceiling (CONTEXTUAL_MAX_RESOURCE_BLOB_BYTES), so one@-mention of a large asset cannot exhaust the server.
Every tool has a hard output cap (BUDGET in src/mcp/format.ts). Over-cap
results truncate and return a pointer to read the rest — never a silent dump.
detect → normalize → chunk → write → embed, with a content hash on sources
making re-ingest idempotent. Source writes complete before embedding; database
triggers notify connected clients when changes commit.
- Skills are validated against the real frontmatter spec:
name≤64 chars, lowercase/digits/hyphens, no XML tags, no reserved words;descriptionnon-empty and ≤1024 chars, no XML tags;compatibility≤500 chars, stored and shown bycx_skill. The bundle tree is stored verbatim andSKILL.mdis never chunked — it is authored to be read whole. Reference files are chunked and embedded for search. A.zipis inspected before extraction: entries with.., absolute paths or symlinks, more than 2000 entries, or over 256 MiB uncompressed are rejected. A directory bundle carries the same caps and skips dotfiles at every level, socontextual add .on a repository that happens to have aSKILL.mdcannot publish its.env. Symlinks are never followed. A directory of sources is walked at most 8 levels deep and 1000 entries wide. - Archives with
.zipor.skillextensions ingest every bundle they contain, in deterministic order. Unknown frontmatter keys fail validation by default;--lenientwarns and omits those fields from indexed metadata while preserving the authored SKILL.md. Tool permissions accept whitespace, commas, or YAML lists, including entries such asBash(git add *). - Renames are renames. Re-ingesting a bundle whose
SKILL.mdname changed replaces the old skill; a document with identical bytes whose old path no longer exists is reclaimed under the new name.contextual removeand every re-ingest also delete the binary assets they orphan, so the blob directory is cleaned after successful replacement. New blobs use unique staging directories; a failed transaction removes its staged files and preserves the previous source. - Uploads retain logical identities. HTTP originals are temporary, so stored
origins use
upload://docs/{collection}/{filename}orupload://skills/{name}. Uploading identical bytes under different document names keeps both sources; reuploading the same name updates that source. Filesystem rename detection does not reclaim uploaded sources. - Documents are normalized with
@firecrawl/anydoc, which covers Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF through one document model. Format is detected from the bytes, not the extension. Unknown binaries and invalid text encodings fail explicitly; UTF-8 and BOM-marked UTF-16 are decoded without replacement characters. Empty documents are reported as failed. HTML conversion preserves GFM tables, strikethrough, and task lists. - Chunking walks the document model rather than re-parsing serialized
markdown, so
heading_pathfalls out of the tree. ~800-token windows with ~100 tokens of overlap. Oversize paragraphs and lists split at sentence/item boundaries, with a hard bound for a single very long item. Large tables split into row batches with repeated headers; fenced code remains atomic. - Embedding supports Voyage or a local Ollama endpoint. Ollama defaults to
qwen3-embedding:0.6b, uses small sequential batches, adds the Qwen query instruction, and validates 1024-dimensional outputs. Ollama's stored model identity includes its provider to prevent incompatible vectors being mixed. Voyage sendstruncation: trueand batches by count and estimated tokens. A batch the API rejects is retried one chunk at a time, so one oversize code block leaves only itself without a vector instead of its whole source. The HNSW index is partial (WHERE embedding IS NOT NULL) and is dropped and rebuilt around any load pastCONTEXTUAL_BULK_INDEX_THRESHOLD(default 2000 vectors), which coversreindex --alland all sources in oneaddinvocation. Source writes finish before one aggregate embedding pass. A smalleraddpays the incremental cost rather than dropping an index other queries are using. The rebuild is in afinally, so a failed load cannot leave the corpus unindexed. Both providers have request timeouts. Voyage has bounded retries with jitter, and supports both forms ofRetry-After. Authentication and input failures are classified separately. A failed embedding run leaves full-text retrieval available for a laterreindex.
toDocument() does not support PDF. anydoc converts PDFs with pdf-inspector,
which emits Markdown directly and has no document-model form. document.ts
therefore has two paths, and blocksFromMarkdown reconstructs the same block
stream for PDFs so the chunker never has to care which it is looking at.
HTML is not an anydoc format. It keeps a @mozilla/readability + turndown
path, routed from detect.ts.
Scanned PDFs support local OCR. CONTEXTUAL_OCR=local uses Poppler to render
scanned pages and Tesseract to recognize text on the service's machine. Mixed
PDFs retain their digital text and native Markdown headings in page order.
The Docker image enables this mode. Outside Docker, the default is reject:
scans land as status='needs_ocr' and remain visible in the catalog. Missing
OCR dependencies, timeouts, and unreadable scanned pages fail explicitly without
ingesting partial content. CONTEXTUAL_OCR=hosted is a separate opt-in that
sends the file to Firecrawl Parse; the file leaves your machine. Local mode
never falls back to hosted OCR.
Local OCR processes one document at a time per service process, with one Tesseract thread and one page raster at a time. Pages are grayscale with their longest side scaled to 3500 pixels. Defaults are 60 seconds per subprocess, 5 minutes per document after queue admission, and 200 total pages in a PDF requiring OCR. Native parser calls are checked against the document deadline between steps; they cannot be interrupted mid-call. Temporary files are removed after success or failure. Scanned pages yield text, without reconstructing table cells or heading levels. See OCR configuration for limits and troubleshooting.
One SQL statement per query combines ts_rank_cd over chunks.ts and cosine
<=> over chunks.embedding, fused with reciprocal rank fusion. The shared
scope CTE is inlined, and vector candidates are selected before window ranking.
Transaction-local HNSW settings enable iterative scoped scans. An EXPLAIN
regression checks that both GIN and HNSW remain usable on a representative corpus.
cx_search takes an array of queries and merges the results, so an agent
spends one round trip on a multi-part question. Its scope is strict:
skills, docs, skills/{name} or docs/{collection} — anything else is an
error rather than a silent search of everything. Round-robin reservations keep
independent queries represented; near-duplicate passages are removed and results
are capped at three per document. This can return fewer than the requested limit.
Document/collection names and full heading paths prefix the text indexed for
FTS and embeddings without appearing in cx_read content.
Optional reranking uses rerank-2.5-lite over the fused candidate set. Set
CONTEXTUAL_RERANK=true with a Voyage key to enable it; this sends candidate
passages to Voyage. On API failure the ordinary fused ranking is retained.
cx_grep runs entirely in Postgres: the file-level match is served by a trigram
index on nodes.content, and lines are extracted with regexp_split_to_table,
so no whole document is ever loaded into the server to be scanned. Patterns are
Postgres regular expressions, which do not support named groups.
Postgres resists the classic catastrophic-backtracking patterns, but that is not
a bound: a pattern with no indexable trigram degrades to a sequential scan whose
cost grows with the corpus. So grep runs under a statement timeout
(CONTEXTUAL_GREP_TIMEOUT_MS, default 5s) applied with SET LOCAL inside a
transaction, which cannot leak onto a pooled connection. A cancelled query comes
back as an over-expensive pattern with advice, not as an internal error.
path_glob is matched against the VFS path an agent sees (/skills/pdf/**),
not the source-relative path stored in the database. cx_glob and cx_grep
share one translation, so a pattern copied from one works in the other. Globs
support *, **, ?, braces, and character classes; malformed or nested brace
patterns fail explicitly. Exact matching happens before SQL limits. Directory
listings aggregate immediate children in SQL and expose an offset continuation.
Skills supply instructions to the agent; retrieved documents are treated as untrusted data. The service enforces these boundaries:
- Uploaded scripts are never executed server-side. They are read-only resources; execution stays with the client, in its own sandbox.
- Path traversal is rejected, not normalized (
src/core/vfs/uri.ts):..,~, absolute paths, backslashes, Windows drive letters, null bytes and percent-encoded variants all fail. Rejecting rather than resolving keeps two URIs from addressing one node. - Retrieved documents, search hits and grep hits are framed as data, not
instructions, in an
<contextual-content untrusted>envelope. A body that contains the envelope's closing tag is cut there and the cut is announced. Metadata is escaped too, including headings and filenames. - Skills are instructions, and are returned bare. The split is by content
kind, not by which tool was called:
SKILL.mdis instructions whether it arrives fromcx_skillor fromcx_readof its URI, because an agent that follows aresource_linkmust not be told to disregard bytes it was just told to follow. A skill's reference, script and asset files are ordinary content and stay enveloped. The trust boundary for skills includes the person or client authorized to ingest them through the CLI or upload API; frontmatter validation does not establish that a skill's instructions are trustworthy. allowed-toolsis advisory. It is stored and shown bycx_skillwith a note that enforcement is the client's job; this server cannot stop a skill from asking for Bash.- The corpus is single-tenant. Stdio assumes a trusted local host. HTTP
defaults to loopback, validates Host/Origin, and supports a shared bearer
token (required outside loopback). Remote deployments need HTTPS through a
reverse proxy; OAuth and per-user isolation are not implemented.
docker-compose.ymlpublishes Postgres on127.0.0.1only; its default credentials are for local development.
db/migrations/ 001–005: schema, retrieval context, URIs, notifications
fixtures/ # sample documents and retrieval evaluation data
scripts/ # evaluation, binary smoke checks, fixture generation
Dockerfile # compiled service + Tesseract + Poppler
docker-compose.yml # Postgres; optional app profile for the HTTP service
src/
core/ # transport-agnostic — survives the move to hosted
db.ts # Bun.sql + pgvector + settings + migrations
tokens.ts # BUDGET + estimateTokens (re-exported by mcp/format)
ingest/ detect skill document local-ocr chunk embed pipeline
vfs/ uri resolve list glob grep
search/ hybrid rrf
catalog/ index watch
mcp/ server http http-body uploads resources tools format
cli/ contextual.ts commands.ts
test/ # bun test
Stdio and HTTP share the transport-independent core/;
test/layering.test.ts fails if anything under
src/core/ imports from src/mcp/.
Bun is the runtime and package manager. Development runs directly from source
with bun run serve; bun run build produces the standalone executable used
by the Docker image. tsconfig.json provides editor types and supports
bun run typecheck.
stdout is the JSON-RPC channel. A single stray console.log corrupts the
stream and hangs the client, so all logging goes to stderr via log() in
src/mcp/format.ts, and a test fails the build if console.log appears anywhere
under src/mcp/. The server also attaches an error listener to stdout and
exits cleanly on EPIPE, which the SDK otherwise raises as an unhandled
exception when a client disconnects abruptly.
Migrations are embedded as text imports, so checkout paths with spaces and the
standalone executable use the same SQL. Each migration and its tracking record
commit in one transaction. Fresh installs create pgvector and pg_trgm in the
public schema so separate corpus schemas can resolve their types and operators.
Include public in the database connection's search_path when using a custom
corpus schema. The migration role must be able to install these extensions, or
an administrator can install them beforehand.
The explicit pgArray() literal remains for stable
binding behavior: the Bun 1.4.2 probe still showed double-quoted values from
sql.array(), while the literal helper passes the round-trip regression.
Database triggers update a durable corpus version and issue NOTIFY on commit.
Bun's dedicated sql.listen() connection checks that version on reconnect,
covering writes missed while disconnected. Idle servers do not poll the corpus.
Change reads are paginated; transient failures retry with bounded backoff.
A polling fallback remains for older, unsupported Bun runtimes.
The MCP SDK v2 stdio entry handles both 2025 initialization and actual
2026-07-28 per-request envelopes, including server/discover and resultType.
Legacy clients use resources/subscribe; modern clients opt in through
subscriptions/listen. The SDK filters notifications by the requested URI/type.
resources/list remains paginated. Missing resources use the SDK's typed
ResourceNotFoundError (-32602 with data.uri). All six tools declare read-only
and idempotent annotations. URI segments are percent-encoded canonically;
spaces, Unicode, literal percent signs, #, and ? are addressable on both surfaces.
contextual add <path...> Ingest a skill bundle, a document, or a directory
contextual list Show what is ingested, with the index token cost
contextual reindex [--all] Embed chunks that have no vector yet
contextual remove <kind> <name> Remove a source
contextual migrate Apply database migrations
contextual serve Run MCP over stdio (default) or Streamable HTTP
| Flag | Meaning |
|---|---|
--collection <name> |
Collection for ingested documents |
--lenient |
Warn and ignore unknown skill metadata fields |
--force |
Re-ingest even when the content hash is unchanged |
--all |
reindex: re-embed everything and re-pin the model |
--allow-reserved |
Permit claude/anthropic in a skill name |
--json |
Machine-readable output |
--transport <stdio|http> |
serve: select transport (default stdio) |
--host <address> |
HTTP bind address (default 127.0.0.1) |
--port <number> |
HTTP port (default 3000) |
The spec reserves anthropic and claude in skill names so third-party skills
cannot impersonate first-party ones. Anthropic's own published bundles are exempt
— anthropics/skills ships claude-api — so enforcing the rule blindly makes
real, publisher-authored bundles un-ingestable. The default is spec-conformant;
the flag exists for exactly that case.
bun test # unit + real Postgres + both MCP protocol eras
bun run typecheck
bun run lint
bun run eval # isolated, offline retrieval smoke evaluation
bun run test:binary # compile and test migrations, DOCX/OCR ingest, and stdio/HTTPThe integration and server suites create a throwaway Postgres schema, ingest real files, and drive the server over real stdio pipes and HTTP sockets. They check SQL bindings, stdout purity, both protocol eras, HTTP authentication, and SSE notifications across concurrent requests.
Database suites skip cleanly with a start-Postgres message when unavailable;
CI sets CONTEXTUAL_REQUIRE_DB_TESTS=1 to make unavailable Postgres a failure.
Each suite restores its database environment and preserves existing URL options.
Local OCR tests exercise real scanned and mixed PDFs when Tesseract and Poppler
are installed; CI installs them and sets CONTEXTUAL_REQUIRE_OCR_TESTS=1.
bun run scripts/make-ocr-fixture.ts regenerates the image-only text fixture.
Keep fixtures/ in the repository: parsing, OCR, upload, binary smoke, and
retrieval evaluation checks depend on it. It is excluded from the Docker build
context and npm package, so it is not required in a running service.
bun install
bun run migrate
# If the embedding provider/model changed, configure Voyage or Ollama first:
bun run src/cli/contextual.ts reindex --allMigration 003 removes file-level FTS, adds indexed chunk context, and clears
old embeddings so they can be regenerated consistently. Full-text search remains
available throughout. Migration 004 rewrites existing URIs canonically;
migration 005 installs change notifications. Run reindex to fill missing
vectors with the existing provider/model, or reindex --all to switch providers
or models. Full-text-only installations do not need an embedding reindex.
To apply parser or OCR changes to existing documents, re-add their originals
with --force or reupload them with force=true; reindexing does not reparse
documents. Restart the service after updating its code or image.
bun run build creates dist/contextual for the current platform with its Bun
runtime, native parser, and migrations embedded. bun run test:binary verifies
it outside the checkout, including a directory with spaces. The npm package
ships the Bun launcher and source; bun pm pack previews that distribution.
GitHub Actions runs lint, types, Postgres/MCP tests, retrieval evaluation, and
the compiled-binary smoke test. No package is published automatically.
The default evaluation uses a small synthetic corpus and reports recall@5 and
MRR. bun run eval --voyage additionally compares plain/prefixed voyage-4,
voyage-context-4, and float/halfvec/int8/binary representations on those fixtures.
It requires a key and makes billed API requests. This is an exploratory quality
comparison, not an HNSW index-size or production-corpus benchmark; retain the
float schema until representative measurements justify changing it.
Numeric settings are validated by src/core/config.ts; the CLI help uses the
same defaults. Invalid values fail with the setting name instead of becoming NaN.
| Variable | Default / purpose |
|---|---|
CONTEXTUAL_DATABASE_URL |
Local Docker Postgres on port 55432 |
CONTEXTUAL_HTTP_TOKEN |
Unset; bearer token required when HTTP binds outside loopback |
CONTEXTUAL_HTTP_ALLOWED_HOSTS |
Localhost addresses for loopback binds, otherwise the bind hostname; explicit hostnames/IPs without ports required for wildcard binds |
VOYAGE_API_KEY |
Optional Voyage authentication |
CONTEXTUAL_EMBED_PROVIDER |
voyage (default), ollama, or none; Voyage without a key uses full-text only |
CONTEXTUAL_OLLAMA_URL |
http://127.0.0.1:11434; Ollama base URL |
CONTEXTUAL_VOYAGE_API_KEY |
Alias for VOYAGE_API_KEY |
CONTEXTUAL_EMBED_MODEL |
voyage-4 for Voyage; qwen3-embedding:0.6b for Ollama; 1024 dimensions |
CONTEXTUAL_FTS_LANGUAGE |
english; pinned on a new database |
CONTEXTUAL_OCR |
reject outside Docker; image defaults to local (Tesseract + Poppler); hosted opts into Firecrawl |
CONTEXTUAL_OCR_LANGUAGE |
eng; installed Tesseract languages, e.g. eng+deu |
FIRECRAWL_API_KEY |
Hosted OCR authentication |
CONTEXTUAL_BLOB_DIR |
./blobs; /app/blobs in the Docker image |
CONTEXTUAL_RERANK |
false; enables external passage reranking |
CONTEXTUAL_RERANK_MODEL |
rerank-2.5-lite |
CONTEXTUAL_SEARCH_TIMEOUT_MS |
10000; allowed 1–3600000 |
CONTEXTUAL_GREP_TIMEOUT_MS |
5000; allowed 1–3600000 |
CONTEXTUAL_MAX_DISTANCE |
1; allowed 0–2 |
CONTEXTUAL_BULK_INDEX_THRESHOLD |
2000; allowed 0–1000000000 |
CONTEXTUAL_MAX_RESOURCE_BLOB_BYTES |
12582912; allowed 1–1073741824 |
CONTEXTUAL_RESOURCE_PAGE |
100; allowed 1–1000 |
CONTEXTUAL_WATCH_MS |
3000; allowed 50–3600000 |
CONTEXTUAL_VOYAGE_TIMEOUT_MS |
30000; allowed 1–3600000 |
CONTEXTUAL_VOYAGE_MAX_ATTEMPTS |
6; allowed 1–10 |
CONTEXTUAL_VOYAGE_RETRY_MAX_MS |
120000; allowed 1–3600000 |
CONTEXTUAL_UPLOAD_MAX_BYTES |
33554432; allowed 1024–268435456 |
CONTEXTUAL_UPLOAD_MAX_FILES |
20; allowed 1–200 |
CONTEXTUAL_OLLAMA_TIMEOUT_MS |
30000; allowed 1–3600000 |
CONTEXTUAL_OCR_TIMEOUT_MS |
60000; allowed 1–3600000 |
CONTEXTUAL_OCR_DOCUMENT_TIMEOUT_MS |
300000; allowed 1–3600000 |
CONTEXTUAL_OCR_MAX_PAGES |
200; allowed 1–10000 |