Skip to content

feat(master): repair base tables whose manifest names missing files, automatically - #276

Merged
beinan merged 1 commit into
lance-format:mainfrom
beinan:feat/repair-missing-fragments
Sep 29, 2026
Merged

beinan merged 1 commit into
lance-format:mainfrom
beinan:feat/repair-missing-fragments

Conversation

@beinan

@beinan beinan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Problem

Six production stores have manifests that reference data files storage no longer has (an early cleanup against a stale listing, or soft-delete expiry; the cause is historical). From then on every scan, merge and compaction of the base table fails with

Not found: rocketkeep/<store>.rollout.lance/data/<file>.lance

No retry changes that, and the cooldown (#267) only spaces the failures out. The rows in those fragments are already gone. Meanwhile the stores' WAL generations accumulate because nothing can merge them (one store is at 7,985 pending), which is a read-amplification and eventually an OOM problem for whoever reads them.

Fix

core: StorageBase::repair_missing_fragments() HEADs every data file and deletion file the current manifest names, and commits an Operation::Delete { deleted_fragment_ids } for the fragments with missing files. That is the same transaction a row delete that empties a fragment commits: it goes through the normal commit path with conflict detection, and touches nothing else. A table with nothing missing commits nothing. Returns a RepairReport (versions, dropped fragment ids, their physical_rows, the missing paths).

master: new TaskKind::Repair (shares the base-table target lock and the general pool with Compact/IndexId). When any task fails and the error matches the missing-fragment signature (is_missing_fragment_error: Not found + a /data/ or /_deletions/ path), the scheduler immediately enqueues a Repair for the target and re-enqueues the failed task depending on it. No failure counting: the condition is deterministic, so waiting is pure waste.

Every repair is persisted in etcd (<prefix>/repairs/<target>/<ts>, kept for TASK_HISTORY_TTL_SECS) as a RepairRecord with what was dropped and which task triggered it, exposed at GET /api/v1/scheduler/repairs and logged at WARN. The decision "these rows are gone" was made by whatever lost the files; the repair records it rather than asking anyone.

Verification

  • core repair_drops_fragments_whose_files_are_missing: 3 one-row fragments; repair on a healthy table commits nothing; delete one data file behind Lance's back; compaction now fails with Not found; repair drops exactly that fragment (id, physical_rows=1, path), commits a new version; 2 rows remain and compaction succeeds again.
  • master missing_fragment_failure_triggers_repair_and_rerun (etcd): compaction of a broken store fails → a Repair and a dependent Compact appear → repair Done with dropped 1 fragments (1 rows) → rerun Done → list_repairs has one record with the right target, trigger and path → 3 rows readable.
  • missing_fragment_error_is_recognised: positive and negative cases (a missing manifest or WAL generation is not repaired).
  • Master suite --include-ignored 74/74, core compaction tests, clippy -D warnings, fmt clean.

🤖 Generated with Claude Code

…automatically

Six production stores have manifests that reference data files storage
no longer has. Every scan, merge and compaction of them fails with
"Not found: .../data/<file>.lance"; no retry changes that, and the
cooldown only spaced the failures out. The rows in those fragments are
already gone. Meanwhile the stores' WAL generations pile up (one is at
7,985) because nothing can merge into the base table.

StorageBase::repair_missing_fragments HEADs every data and deletion
file the current manifest names and commits an Operation::Delete of the
fragments with missing files, the same transaction a row delete that
empties a fragment commits. Everything else, including every WAL
generation, is untouched.

The master enqueues a Repair task the first time any task fails with
that error signature (no counting: the condition is deterministic) and
re-enqueues the failed task behind it. Each repair is recorded in etcd
with the dropped fragment ids, their row counts and the missing paths,
exposed at GET /api/v1/scheduler/repairs and logged at WARN, so "why
does this store have fewer rows" is answerable later.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@beinan
beinan merged commit 5a92cdc into lance-format:main Sep 29, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant