Skip to content

Latest commit

 

History

History
152 lines (122 loc) · 19.4 KB

File metadata and controls

152 lines (122 loc) · 19.4 KB

Reason codes

A living reference of every reason_code string the cache plane emits on the wire. The tech spec (§4.2) and the gRPC contract design doc (../design/grpc-contract.md) cover the why; this file is the what is exposed today.

Update this file in the same change that introduces or retires a code. New codes don't break old clients (see "Forward compatibility" below), but an undocumented code in production is a real friction point for the gateway team and for ops dashboards.

Forward compatibility — the load-bearing rule

  • reason_code is a string, not a proto enum. New codes are an additive server-side change with no client recompile.
  • Clients MUST treat any unrecognized code as the no-hint default for that RPC (NO_HINT for LookupRoute / LookupPDRoute, RENDER_ERROR for RenderTemplate). This is what lets the server roll out new codes without coordinating with every gateway.
  • The server emits codes only from the Status column "shipped" below. "Spec'd" codes are reserved in the proto comments but not yet returned; promoting one to "shipped" is the trigger to update this doc.

LookupRoute / LookupPDRoute

LookupPDRoute remains a fail-open stub today (Phase 2 — prefill/decode routing) and unconditionally emits NO_HINT. The codes in the table below that are marked LookupRoute only are emitted by LookupRoute's handler and will surface from LookupPDRoute when its handler is wired in a later module — until then, treat NO_HINT as the only LookupPDRoute answer.

Code Status RPC When the server emits it Response shape What the gateway does
PREFIX_MATCH shipped LookupRoute only The index has at least one replica holding the request's (tenant, model, hash_scheme, adapter) prefix — either the exact prefix_hash (legacy single-blob path) OR the leading run of block_hashes[0..k] (chain longest-prefix path; see ../design/lookuproute-ranking.md §2.5) — the ranker returned a non-empty set, AND at least one replica's realized matched_tokens cleared the per-namespace minimumMatchedTokens floor (default 64) AND the top surviving replica's score cleared the per-namespace routingFloorScore floor (default 0.1). The score includes the distinguishing-power factor 1 − num_matching_replicas / total_replicas so an overlap held by every replica (chat-template framing, RAG corpus headers, custom system prompts) collapses to score 0 and is filtered by the score floor. Both floors can downgrade independently. replica_scores non-empty, ranked best-first by matched_tokens × freshness × pressure_factor × slo_bias × distinguishing_power. The pressure / SLO / distinguishing-power factors collapse appropriately when their inputs are absent (no stats → pressure_factor = 1; no SLO hint → slo_bias = 1; single-replica deployment → distinguishing_power = 1). Every qualifying replica is returned today (no top-K limit); the gateway typically uses the top entry. Route to the top-ranked replica → prefix-cache hit; lower TTFT.
NO_HINT shipped both The fail-open default. LookupPDRoute: every call (the handler is a stub). LookupRoute: the prefix is novel under matching contract keys AND no usable affinity fallback fired (affinityRouting: Disabled on the per-namespace policy, OR no replica known to serve the (tenant, model, hash_scheme) engine domain, OR no usable seed — empty block_hashes and empty prefix_hash, OR structurally malformed input — empty hash_scheme or chain arrays of mismatched length); the ranker found nothing AND the same affinity-disabled / no-replica / no-seed / malformed clause holds; any of tenant_id, model_id, or hash_scheme was unspecified (a contract violation — set-but-wrong values surface as the matching UNKNOWN_* code instead); the index is globally empty (cold-start carve-out); the request was policy-gated below minimumPrefixTokens AND affinityRouting: Disabled (with affinityRouting: Enabled — the default — the same gate surfaces as AFFINITY_HINT); every replica that held the prefix matched fewer tokens than minimumMatchedTokens (result-side per-replica floor, default 64) AND affinityRouting: Disabled (with affinity Enabled the downgrade surfaces as AFFINITY_HINT); the top per-replica score from the distinguishing-power-aware ranker fell below routingFloorScore (result-side score floor, default 0.1) AND affinityRouting: Disabled (with affinity Enabled the downgrade surfaces as AFFINITY_HINT); or an index-disabled state. replica_scores empty. Not an error. Route per the gateway's default policy (round-robin, least-loaded, …). The cache plane is invisible to this request.
POLICY_REQUIRES_CHAIN shipped LookupRoute only The tenant's CachePolicy.spec.strategy.requireChain is true, but the request did not carry a valid block_hashes + block_token_counts chain. The server returns before touching the index. This is a policy-gated empty result, separated from NO_HINT so operators can see legacy/exact callers hitting a chain-only namespace in inferencecache_lookup_route_calls_total{reason_code="POLICY_REQUIRES_CHAIN"}. Empty replica_scores. Treat as NO_HINT; update the caller to send the chain if it should benefit from the cache plane in this namespace.
TENANT_HOT shipped LookupRoute only No exact prefix match for (tenant, model, hash_scheme, adapter, prefix_hash), CachePolicy.spec.strategy.enableTenantHot is not false, and the tenant has at least one replica that (a) has reported stats recently (within ~5 minutes by default), (b) has a hit_rate above a small floor (default 0.1), AND (c) currently has at least one prefix entry in the requested (tenant, model, hash_scheme) in the index — proving the replica serves the requested engine domain. The "in the index" check is sweep-driven (an entry past TTL stays counted until the next sweep removes it), so for at most one sweep interval a recently-stale entry can briefly still satisfy the check; per soft-state semantics that yields at worst a soft hint that turns into a cache miss, never a wrong answer. A coarser locality signal than PREFIX_MATCH — useful when the prefix is novel but the tenant already has servers warm in the cache rotation. When the same fallback is found but the policy disables tenant-hot, the handler downgrades it to NO_HINT. replica_scores non-empty (tenant-hot ranked); matched_tokens is 0 because there is no prefix overlap (the gateway must rely on reason_code, not matched_tokens, to recognize this branch). Shape otherwise unchanged. Treat as a softer hint than PREFIX_MATCH; gateway free to use or ignore.
TIMEOUT shipped LookupRoute only The lookup deadline expired before the index could rank — either the caller's context was already past its deadline on arrival or the per-tenant CachePolicy.spec.lookupTimeoutMs budget elapsed during the lookup. Gateway clients also synthesize this locally when they cancel a slow LookupRoute RPC. Server: empty replica_scores. Client-side synth: same. Treat as NO_HINT.
UNKNOWN_TENANT shipped LookupRoute only After a prefix miss (and the TENANT_HOT fallback when it applies — non-chain requests; chain requests skip TENANT_HOT by design and classify directly), AND the index is not globally empty: the request supplied a non-empty tenant_id and the index has zero prefix entries for that tenant across every model and hash scheme. Cold-start carve-out: a globally empty index (server just started, no ReportCacheState yet) stays on NO_HINT so a fresh deployment does not flood gateways with UNKNOWN_TENANT; the diagnostic resumes the moment any replica has reported state. Canonical asymmetric shape: a gateway-SDK querying with tenant_id="default" while the producer (kvevent-subscriber sidecar) is publishing under tenant_id=$(POD_NAMESPACE). Empty replica_scores. Treat as NO_HINT for routing (still fail-open — the cache plane is hint-only); surface as a configuration error (log line / metric / SDK warning). Do not retry under a different key — the cache plane will not change between calls.
UNKNOWN_MODEL shipped LookupRoute only Same precondition as UNKNOWN_TENANT above. The tenant is known but the (tenant_id, model_id) pair has zero entries. The model has never served traffic in this tenant, or the model identifier disagrees between producer and consumer. Empty replica_scores. Same as UNKNOWN_TENANT: fail-open, surface as configuration error.
UNKNOWN_HASH_SCHEME shipped LookupRoute only Same precondition as UNKNOWN_TENANT above. (tenant_id, model_id) has entries, but none under the request's hash_scheme. Canonical shape: ingest under "vllm", lookup under "vllm-v1". An empty hash_scheme is a contract violation (not a mismatch) and stays on NO_HINT. Empty replica_scores. Same as UNKNOWN_TENANT: fail-open, surface as configuration error.
AFFINITY_HINT shipped LookupRoute only The prefix-match path would otherwise return NO_HINT AND the per-namespace CachePolicy.spec.affinityRouting is Enabled (the kubebuilder default) AND the request has a usable fingerprint AND the index knows at least one replica SERVING the request's (tenant, model, hash_scheme) engine domain. The server returns a single stable replica picked by SHA-256 over the request's length-prefixed block_hashes (fall-back to prefix_hash for legacy callers) modulo the sorted SCHEME-AWARE replica set for (tenant, model, hash_scheme) (read from servingByScope, the same accelerator TENANT_HOT uses, so a vLLM request can never pin to an SGLang replica) — so repeat prompts pin to the same replica and warm T1 on diffuse single-turn workloads. The diagnostic codes (UNKNOWN_TENANT / UNKNOWN_MODEL / UNKNOWN_HASH_SCHEME) and TIMEOUT keep precedence over AFFINITY_HINT; affinity never preempts a real PREFIX_MATCH or TENANT_HOT that cleared the request-side gates (one exception: a tiny request below the per-namespace minimumPrefixTokens gate has its positive-hint result — including TENANT_HOT — downgraded to StrategyNone, so the affinity fallback can still fire on it; the operator intent "tiny prompts don’t surface a positive hint" outranks the TENANT_HOT-vs-affinity precedence); requests with empty hash_scheme or chain arrays of mismatched length stay on NO_HINT (structurally malformed input shouldn't be papered over). replica_scores carries exactly one entry whose score, matched_tokens, and estimated_cache_hit_prob are all 0 — there is no cache-evidence claim, only a stable assignment. Shape otherwise identical to PREFIX_MATCH. Route to the single returned replica. Treat as PREFIX_MATCH for routing purposes; the distinct reason code is the operator-facing signal for measuring the affinity-vs-real-match share via inferencecache_lookup_route_calls_total{reason_code="AFFINITY_HINT"}.

Constants in code: reasonPrefixMatch, reasonTenantHot, reasonNoHint, reasonTimeout, reasonPolicyRequiresChain, reasonAffinityHint, reasonUnknownTenant, reasonUnknownModel, reasonUnknownHashScheme in internal/server/inferencecache_service.go. See also ../design/lookuproute-diagnostics.md for the design rule and gateway-SDK guidance.

Ranking inputs beyond matched_tokens × freshness

The server-side ranker (internal/index) is configurable via RankerConfig (in- binary knobs) and CachePolicy.spec (per-namespace knobs). Each RankerConfig knob defaults to a value that reduces the pressure / SLO layers to the baseline when its supporting signal is absent — so a deployment without replica stats or SLO hints sees pressure_factor = 1 and slo_bias = 1. The cardinality-aware distinguishing-power factor is always on for multi-replica deployments (no per-knob disable; degrades to 1.0 only for single-replica deployments), and the two result-side floors (minimumMatchedTokens, routingFloorScore) still apply on top — see the CachePolicy.spec rows below for their opt-outs. So a deployment without replica stats or SLO hints behaves like the original B6 baseline on the pressure/SLO factors only; the cardinality factor and both floors still run.

Knob What it does Default Off switch
PressureWeight Penalty applied to a replica's score from ReplicaStats.pressure: pressure_factor = max(0, 1 - PressureWeight × pressure). Avoids blindly preferring a saturated cache holder over a fresher, lower-pressure peer. 1.0 0 → no penalty
SLOTightTTFTMs TTFT budget (ms) below which the request is "tight" and the SLO bias kicks in. Uses LookupRouteRequest.slo.ttft_ms. 200 0 → bias never fires
SLOTightBias Coefficient in the freshness boost: slo_bias = 1 + freshness × SLOTightBias when the request is tight. Higher → fresher candidates are favored more aggressively. 1.0 0 → no boost
TenantHotMinHitRate Minimum hit_rate for a replica to count as "warm" for the TENANT_HOT fallback. 0.1 n/a (use TenantHotMaxAge = 0 to disable the fallback)
TenantHotMaxAge Maximum stats age for a replica to count as "warm". 5m 0 → fallback disabled (a prefix miss whose contract keys all populate the index lands at NO_HINT; mismatched-key misses still diagnose as UNKNOWN_* via the miss-classifier)
distinguishing_power factor Cardinality-aware multiplier: 1 − num_matching_replicas / total_replicas, per-replica depth-aware for chain matches. Discounts overlaps every replica holds (chat-template framing, RAG corpus headers, custom system prompts). Always on for multi-replica deployments; degrades to 1.0 for single-replica deployments. See ../design/lookuproute-ranking.md §2.7. always on for multi-replica; 1.0 for single-replica none (operators disable the floor it feeds via CachePolicy.spec.routingFloorScore: "0", not the factor itself)
CachePolicy.spec.minimumMatchedTokens Per-replica matched-tokens floor: filters replicas whose realized matched_tokens falls below the threshold. If no replica survives, the response downgrades to StrategyNone, which surfaces as AFFINITY_HINT under affinityRouting: Enabled (the default) with a usable seed + serving replica or as NO_HINT under affinityRouting: Disabled. 64 (4 KV blocks) 0 on the CR → opt-out for that namespace
CachePolicy.spec.routingFloorScore Per-response score floor on the top surviving replica's score (after the distinguishing-power factor multiplies in). Below the floor → response downgrades to StrategyNone, with the same AFFINITY_HINT vs NO_HINT split as the matched-tokens row above. "0.1" "0" on the CR → opt-out for that namespace
CachePolicy.spec.strategy.enableChainMatching / requireChain / enableTenantHot Per-namespace strategy gates. enableChainMatching=false disables block-hash chain matching and uses exact prefix_hash; requireChain=true rejects non-chain requests with POLICY_REQUIRES_CHAIN; enableTenantHot=false downgrades tenant-hot fallbacks to NO_HINT. true / false / true Set the relevant boolean on the CR; requireChain=true is rejected unless chain matching is enabled
CachePolicy.spec.affinityRouting Per-namespace consistent-hash fallback toggle. When Enabled (the default), StrategyNone responses with a usable seed + serving replica surface as AFFINITY_HINT (single stable replica from servingByScope[(tenant, model, hash_scheme)]). When Disabled, those same responses surface as NO_HINT. Diagnostic codes and TIMEOUT keep precedence over AFFINITY_HINT. Enabled Disabled on the CR → opt-out for that namespace

RenderTemplate

Code Status When the server emits it Notes
OK shipped (stub) The handler currently returns OK unconditionally — the rendering pipeline (Wedge D, D1–D5) isn't wired yet. Becomes "real" when D2 (render pipeline) lands.
TEMPLATE_NOT_FOUND spec'd, not emitted The referenced template_ref doesn't exist. Promoted when D5 (RenderTemplate handler) lands.
RENDER_ERROR spec'd, not emitted Template was found but rendering failed (missing/typed-wrong variables, runtime DSL error). Promoted with D5.

Constants in code: reasonOK in internal/server/inferencecache_service.go.


Ack (ReportCacheState, PublishEvent)

The Ack proto carries an optional reason_code for future use (e.g. partial acceptance, throttling signals). Today the server returns Ack{Accepted: true} with reason_code unset on every code path — there are no Ack reason codes in use.

When the first Ack code ships (e.g. THROTTLED, SCHEMA_DROPPED), add it to the table below and update this paragraph.

Code Status When the server emits it Notes
(none) — — Field reserved for future structured acknowledgments.

How reason codes show up in metrics

reason_code is a label on inferencecache_lookup_route_calls_total, alongside model and hint_used. The cardinality of values is therefore bounded by the table above — adding a new code adds one new label value (cheap), but it does show up as a new time series, so prefer reusing existing codes if the semantic fits.

inferencecache_lookup_route_calls_total{model="...", reason_code="PREFIX_MATCH", hint_used="true"}  42
inferencecache_lookup_route_calls_total{model="...", reason_code="NO_HINT",      hint_used="false"} 318

hint_used="true" ⇔ replica_scores was non-empty in the response. It correlates with PREFIX_MATCH, TENANT_HOT, or AFFINITY_HINT — all three shipped codes return non-empty scores. The diagnostic UNKNOWN_* codes, NO_HINT, and TIMEOUT correlate with hint_used="false".

See metrics.md for the full metric surface.


How to add a new reason code

  1. Reserve the string in the proto comment on the response message in proto/inferencecache/v1alpha1/inferencecache.proto (already done for TENANT_HOT, TIMEOUT, TEMPLATE_NOT_FOUND, RENDER_ERROR). Run make proto-gen if the comment touched the schema.
  2. Add a constant in internal/server/inferencecache_service.go next to reasonPrefixMatch / reasonNoHint / reasonOK. Keep the constant name reason<CamelCase>.
  3. Emit it from the handler at the relevant decision point. Keep handlers side-effect-free apart from metrics; reason_code is the only way the server communicates "what kind of answer is this."
  4. Update the table above — move the row from "spec'd" → "shipped", describe the trigger condition, the response shape, and the gateway action.
  5. Document the metric expectation. If the new code is for LookupRoute, a new reason_code label value will appear automatically in inferencecache_lookup_route_calls_total. Mention this in the PR description so dashboards can be updated.
  6. Backward-compat check. Confirm that an existing client not updated for the new code degrades to its no-hint default (NO_HINT / RENDER_ERROR). This is the contract; verify by reading the client adapter, not by guessing.