diff --git a/README.md b/README.md index 32084cc..152d817 100644 --- a/README.md +++ b/README.md @@ -1,27 +1,36 @@ # 🪓 whittle -**Carve your agent's context down to what matters, and route every request to the cheapest model that can handle it.** +**Carve your agent's context down to what matters. Lossless or clearly marked, code never touched.** -*Local. Fail-open. Every claim in this README regenerates from a clone.* +*Optionally, route each request to the cheapest model that can handle it. Local, fail-open, and every claim in this README regenerates from a clone.* [![CI](https://github.com/firstops-dev/whittle/actions/workflows/ci.yml/badge.svg)](https://github.com/firstops-dev/whittle/actions/workflows/ci.yml) [![Release](https://img.shields.io/github/v/release/firstops-dev/whittle)](https://github.com/firstops-dev/whittle/releases) [![Go Reference](https://pkg.go.dev/badge/github.com/firstops-dev/whittle.svg)](https://pkg.go.dev/github.com/firstops-dev/whittle) [![License](https://img.shields.io/badge/license-Apache--2.0-blue)](LICENSE) -![whittle: carve your agent's context down to what matters, and route every request to the right model](assets/banner.svg) +![banner: the whittle wordmark carving into green shavings, dense context flowing in on the left, three model tiers flowing out on the right](assets/banner.svg) + +

+ Install · + How it works · + Compression · + Model routing · + Benchmarks · + Why whittle? +

Long agent sessions drown in tokens, and most compressors buy their ratio by silently destroying what agents need: array rows vanish, file reads get gutted, identifiers come back mangled. Whittle holds one line: **lossless or clearly marked, code never touched, every anomaly fails open to the original bytes.** +Built for engineers running long Claude Code sessions who refuse to trade fidelity for token savings. + ## Highlights - 🪓 **Write-time compression**: a Claude Code PostToolUse hook whittles each tool output *before* it enters context, so the savings repeat on every later turn. History is never mutated, so prompt caches stay intact. - 🔒 **Lossless or marked**: JSON reshapes byte-exact (rows are never dropped), logs keep every error plus an honest `[N lines omitted]`, source code passes through untouched. - 🎯 **Token-honest numbers**: savings measured against calibrated tokenizer counts, not byte counts that overstate by up to 4×. - 🧭 **Opt-in model router**: hard reasoning stays on your strongest model, trivia drops to the cheapest, per one auditable policy file driven by trained-classifier signals ([how it works](docs/ROUTER.md)). -- 🛟 **Fail-open everywhere**: whittle down means your agent runs on originals; a rejected rewrite means your original request, retried. Never blocked, never corrupted. -- 🏠 **Runs entirely on your machine**: zero credentials leave your box; the Go binary has zero external dependencies. -- 🧾 **Executable claims**: `make test` and `go run ./bench` regenerate every number in this README from a clone. +- 🛟 **Fail-open, local, yours**: if whittle is down your agent runs on originals, a rejected rewrite retries your original request, zero credentials leave your box, and the Go binary has zero external dependencies. ## Install @@ -60,7 +69,7 @@ Both errors, their context, and the summary survive; 138 lines of noise collapse ## Compression: what happens to each content type -JSON is reshaped **losslessly** (byte-exact reconstruction). Logs and terminal streams are cut lossily but **marked and exactly accounted**. Markdown file reads keep code fences, tables, and headings byte-exact while prose is compressed by an optional local model with fidelity guards. Source code is **never touched**. Every path is wrapped in fail-open guardrails: the worst case is always *not compressed*, never *corrupted*. +JSON is reshaped **losslessly** (byte-exact reconstruction). Logs and terminal streams are cut lossily but **marked and exactly accounted**. Markdown file reads keep code fences, tables, and headings byte-exact while prose is compressed by an optional local model with fidelity guards. Source code is **never touched**. Every path is guarded: the worst case is always *not compressed*, never *corrupted*. The full per-type contract, ML prose path, architecture, and performance tables: [docs/compression.md](docs/compression.md) · what each guarantee is pinned by: [GUARANTEES.md](GUARANTEES.md). @@ -71,14 +80,24 @@ Whittle's second surface: a local proxy on `ANTHROPIC_BASE_URL` that sends each - **Calibrated out of the box**: `whittle policy init` writes a conservative default (hard reasoning → strongest tier, confident chit-chat → cheapest, *everything else untouched*) with your account's real model ids auto-detected. [What it does & how to customize](router/policies/default.md). - **Multi-signal, not keyword-matching**: a trained 14-subject classifier (probability-mass thresholded, so an *uncertain* classification never escalates), a contrastive difficulty score, and your own keywords. Every log line shows each signal's value against its gate. - **Rewrites the model, never your history**: prompt-cache prefixes survive; capabilities the cheaper model rejects are stripped automatically; credentials pass through untouched. -- **Fail-open by construction**: bad policy, dead classifier, or a rejected rewrite all fall back to your original request. Unset the env var and you're direct again. +- **Never blocks you**: bad policy, dead classifier, or a rejected rewrite all fall back to your original request. Unset the env var and you're direct again. - **Savings you can measure**: every request logs requested model, served model, and real token usage. The router itself (engine, policy design, signal composition) is whittle's own; two pretrained models power its ML signals (the 14-subject `domain` classifier and the text embedder, both from [vLLM Semantic Router](https://github.com/vllm-project/semantic-router)). Architecture, signal math, and precise credits: [docs/ROUTER.md](docs/ROUTER.md). ## Why write-time? -Most context compressors are read-time proxies: they rewrite your conversation history on every LLM call, which breaks prompt caches, terminates your API traffic, and makes lossy compression the default. Whittle compresses each output **once, at the moment it's born**, before it enters history: savings compound across every later turn, nothing sits in your request path, and failure costs nothing. The full argument: [docs/why-write-time.md](docs/why-write-time.md). +Most context compressors are read-time proxies: they rewrite your conversation history on every LLM call, which breaks prompt caches, terminates your API traffic, and makes lossy compression the default. Whittle compresses each output **once, at the moment it's born**, before it enters history: savings compound across every later turn, and nothing sits in your request path. + +| | whittle | read-time compressors | +|---|---|---| +| compresses | once, at write time | re-runs on every LLM call | +| prompt cache | intact, history never rewritten | invalidated by history rewrites | +| your request path | nothing resident in it | a proxy that must stay up | +| loss model | lossless or marked | lossy by default, recover on demand | +| cost of failure | original bytes | a broken or blocked call | + +The full argument: [docs/why-write-time.md](docs/why-write-time.md). ## FAQ @@ -106,12 +125,16 @@ python bench/calibrate_tokens.py # reproduces the token-estimator MAE ## Contributing -The bar: guarantees are executable, see [CONTRIBUTING.md](CONTRIBUTING.md). Fidelity bugs (whittle changing an output's meaning) are treated as urgent. Good first issues: agent adapters, Linux packaging, detection corpus cases. +The bar: guarantees are executable, see [CONTRIBUTING.md](CONTRIBUTING.md). Fidelity bugs (whittle changing an output's meaning) are treated as urgent. + +Near-term roadmap: a tagged release that ships the router, agent adapters (Cursor, Codex, OpenCode), Linux packaging. Good first issues: adapters, packaging, detection corpus cases. ## Acknowledgments Whittle's log-selection strategy, several detection heuristics, and the tabular parser were adapted from [Headroom](https://github.com/headroomlabs-ai/headroom) (Apache-2.0); their compaction work is excellent; we wanted the write-time position and the stricter fidelity contract it demands. The router's two pretrained models (the `domain` classifier and the text embedder behind the similarity signals) come from [vLLM Semantic Router](https://github.com/vllm-project/semantic-router) ([whitepaper](https://vllm-semantic-router.com/white-paper)); the routing engine and policy design are whittle's own. See [NOTICE](NOTICE). +If whittle's fidelity contract is the trade you want, a ⭐ helps other agent users find it. + ## License Apache-2.0. diff --git a/cmd/whittle/main.go b/cmd/whittle/main.go index 2e8caea..2698356 100644 --- a/cmd/whittle/main.go +++ b/cmd/whittle/main.go @@ -14,6 +14,8 @@ import ( "io" "net/http" "os" + "runtime/debug" + "strings" "time" "github.com/firstops-dev/whittle" @@ -21,9 +23,22 @@ import ( "github.com/firstops-dev/whittle/server" ) -// version is injected by goreleaser (-X main.version=...); dev builds show the -// last released baseline. -var version = "0.2.1" +// version is injected by goreleaser (-X main.version=...). For plain +// `go install module@vX.Y.Z` builds (no ldflags), the module version from build +// info is used instead — so @latest users see the real tag, not the baseline. +var version = "0.3.0" + +func resolvedVersion() string { + if version != baselineVersion { + return version // goreleaser-injected + } + if bi, ok := debug.ReadBuildInfo(); ok && bi.Main.Version != "" && bi.Main.Version != "(devel)" { + return strings.TrimPrefix(bi.Main.Version, "v") + } + return version +} + +const baselineVersion = "0.3.0" func main() { if len(os.Args) < 2 { @@ -56,7 +71,7 @@ func main() { case "hook": cmdHook(os.Args[2:]) case "version": - fmt.Println("whittle", version) + fmt.Println("whittle", resolvedVersion()) default: usage() os.Exit(2)