Skip to content

fix(master): compact and version-prune the store registries alongside _stats - #274

Merged
beinan merged 1 commit into
lance-format:mainfrom
beinan:fix/registry-maintenance
Sep 29, 2026
Merged

beinan merged 1 commit into
lance-format:mainfrom
beinan:fix/registry-maintenance

Conversation

@beinan

@beinan beinan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Problem

_registry.rollout.lance (and _registry.generic.lance) is the shared directory of stores every worker consults. Workers write it on every store create/delete as a delete plus an append (upsert is delete-then-append so a retried create is idempotent). Nothing ever compacts or version-prunes it.

Lance keeps every manifest until cleaned, and each manifest lists every fragment, so the table grows quadratically. Production today: version 34,504, ~840 KB per manifest, ~29 GB of manifests for a 10k-row three-column table, with one fragment per row.

Every POST /rollouts does contains() (a checkout_latest that reads that head manifest) then upsert() (two commits, each writing a new 840 KB manifest); every GET /rollouts/{name} does another checkout_latest. Measured on prod: creates average 47 s (201), 83 s when the client has already timed out and retried (409), and fail with 500 when ADLS throttles mid-way. GET /rollouts/{name} averages 0.7 s for a row lookup.

Fix

RolloutRegistry gets the same compact() / detached RegistryCleaner::cleanup() / reload() split as StatsStore (#272), plus maintain() for callers without a shared lock.

The master runs it for both registries in the scanner's existing maintenance round (every STATS_MAINTENANCE_EVERY_N_SCANS scans, under the stats-writer lock so only one master rewrites). Compaction is bounded by MAINTENANCE_TIMEOUT; the cleanup runs with the lock released. Workers keep appending throughout: Lance's commit retry carries their appends past the rewrite, and the grace window (STATS_HISTORY_TTL_SECS) keeps any version a worker may still hold open.

The first pass on prod will delete ~34k manifests; with #272 that no longer blocks anything.

Verification

  • maintain_bounds_versions_and_preserves_rows: 20 upserts (≥40 versions) → compaction folds fragments, cleanup removes versions, all 20 rows and further writes work.
  • cleanup_runs_detached_from_the_registry: writes land while the cleaner is outstanding and survive.
  • maintenance_does_not_lose_concurrent_worker_upserts: a second handle on the same URI (the worker) creates stores before, between and after the master's compact/cleanup; both handles see all 26 rows.
  • etcd-backed master suite --include-ignored: 72 passed. Clippy -D warnings, fmt clean.

🤖 Generated with Claude Code

… _stats

Workers write _registry (and _registry.generic) on every store create or
delete as a delete plus an append: two versions and one fragment per
store, never compacted, never pruned. Lance keeps every manifest and
each manifest lists every fragment, so the table grows quadratically.
Production: 34,504 versions of ~840 KB each, ~29 GB of manifests, for
a 10k-row three-column table. Every create re-reads the head manifest
in contains(), then commits two more; every GET /rollouts/{name}
re-reads it too. Creates took 47 s on average and failed outright when
ADLS throttled mid-way.

Give RolloutRegistry the same compact()/RegistryCleaner::cleanup()
split StatsStore has, and run it for both registries in the scanner's
maintenance round under the stats-writer lock, bounded on the
compaction and detached for the cleanup so workers are never blocked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@beinan
beinan merged commit 2759af5 into lance-format:main Sep 29, 2026
10 checks passed
beinan added a commit to beinan/lance-context that referenced this pull request Sep 29, 2026
… a mirrored migration path

The store registries (rollout, generic, datagen) are a name -> uri
directory every worker consults on a cache miss and writes on every
create/delete. As a Lance table each read is a manifest fetch from
object storage and each write is two contended commits; even after lance-format#274
keeps it compact, a create costs seconds and throttling turns it into
failures.

Introduce a StoreRegistry trait (&self, so the RwLock every request
went through goes away) with three impls: LanceRegistry (today's table),
EtcdRegistry (one key per store under <ETCD_PREFIX>/registry/<kind>/,
put-if-absent backfill), and MirroredRegistry (read one, write both,
mirror best-effort). REGISTRY_BACKEND / REGISTRY_MIRROR select them;
defaults keep Lance with no mirror, so this change is behaviour-neutral
until the flags are set.

The master seeds the mirror from the primary at startup and heals it
every maintenance round, and exposes GET /registry/diff and
POST /registry/backfill for verifying a migration step. ETCD_* flags
move to core (EtcdConfig) so the server can share them; the master's
etcd prefix is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
beinan added a commit to beinan/lance-context that referenced this pull request Sep 29, 2026
… a mirrored migration path

The store registries (rollout, generic, datagen) are a name -> uri
directory every worker consults on a cache miss and writes on every
create/delete. As a Lance table each read is a manifest fetch from
object storage and each write is two contended commits; even after lance-format#274
keeps it compact, a create costs seconds and throttling turns it into
failures.

Introduce a StoreRegistry trait (&self, so the RwLock every request
went through goes away) with three impls: LanceRegistry (today's table),
EtcdRegistry (one key per store under <ETCD_PREFIX>/registry/<kind>/,
put-if-absent backfill), and MirroredRegistry (read one, write both,
mirror best-effort). REGISTRY_BACKEND / REGISTRY_MIRROR select them;
defaults keep Lance with no mirror, so this change is behaviour-neutral
until the flags are set.

The master seeds the mirror from the primary at startup and heals it
every maintenance round, and exposes GET /registry/diff and
POST /registry/backfill for verifying a migration step. ETCD_* flags
move to core (EtcdConfig) so the server can share them; the master's
etcd prefix is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
beinan added a commit to beinan/lance-context that referenced this pull request Oct 2, 2026
… a mirrored migration path

The store registries (rollout, generic, datagen) are a name -> uri
directory every worker consults on a cache miss and writes on every
create/delete. As a Lance table each read is a manifest fetch from
object storage and each write is two contended commits; even after lance-format#274
keeps it compact, a create costs seconds and throttling turns it into
failures.

Introduce a StoreRegistry trait (&self, so the RwLock every request
went through goes away) with three impls: LanceRegistry (today's table),
EtcdRegistry (one key per store under <ETCD_PREFIX>/registry/<kind>/,
put-if-absent backfill), and MirroredRegistry (read one, write both,
mirror best-effort). REGISTRY_BACKEND / REGISTRY_MIRROR select them;
defaults keep Lance with no mirror, so this change is behaviour-neutral
until the flags are set.

The master seeds the mirror from the primary at startup and heals it
every maintenance round, and exposes GET /registry/diff and
POST /registry/backfill for verifying a migration step. ETCD_* flags
move to core (EtcdConfig) so the server can share them; the master's
etcd prefix is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant