feat(master): repair base tables whose manifest names missing files, automatically - #276
Merged
Merged
Conversation
…automatically Six production stores have manifests that reference data files storage no longer has. Every scan, merge and compaction of them fails with "Not found: .../data/<file>.lance"; no retry changes that, and the cooldown only spaced the failures out. The rows in those fragments are already gone. Meanwhile the stores' WAL generations pile up (one is at 7,985) because nothing can merge into the base table. StorageBase::repair_missing_fragments HEADs every data and deletion file the current manifest names and commits an Operation::Delete of the fragments with missing files, the same transaction a row delete that empties a fragment commits. Everything else, including every WAL generation, is untouched. The master enqueues a Repair task the first time any task fails with that error signature (no counting: the condition is deterministic) and re-enqueues the failed task behind it. Each repair is recorded in etcd with the dropped fragment ids, their row counts and the missing paths, exposed at GET /api/v1/scheduler/repairs and logged at WARN, so "why does this store have fewer rows" is answerable later. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Six production stores have manifests that reference data files storage no longer has (an early cleanup against a stale listing, or soft-delete expiry; the cause is historical). From then on every scan, merge and compaction of the base table fails with
No retry changes that, and the cooldown (#267) only spaces the failures out. The rows in those fragments are already gone. Meanwhile the stores' WAL generations accumulate because nothing can merge them (one store is at 7,985 pending), which is a read-amplification and eventually an OOM problem for whoever reads them.
Fix
core:
StorageBase::repair_missing_fragments()HEADs every data file and deletion file the current manifest names, and commits anOperation::Delete { deleted_fragment_ids }for the fragments with missing files. That is the same transaction a row delete that empties a fragment commits: it goes through the normal commit path with conflict detection, and touches nothing else. A table with nothing missing commits nothing. Returns aRepairReport(versions, dropped fragment ids, theirphysical_rows, the missing paths).master: new
TaskKind::Repair(shares the base-table target lock and the general pool withCompact/IndexId). When any task fails and the error matches the missing-fragment signature (is_missing_fragment_error:Not found+ a/data/or/_deletions/path), the scheduler immediately enqueues aRepairfor the target and re-enqueues the failed task depending on it. No failure counting: the condition is deterministic, so waiting is pure waste.Every repair is persisted in etcd (
<prefix>/repairs/<target>/<ts>, kept forTASK_HISTORY_TTL_SECS) as aRepairRecordwith what was dropped and which task triggered it, exposed atGET /api/v1/scheduler/repairsand logged at WARN. The decision "these rows are gone" was made by whatever lost the files; the repair records it rather than asking anyone.Verification
repair_drops_fragments_whose_files_are_missing: 3 one-row fragments; repair on a healthy table commits nothing; delete one data file behind Lance's back; compaction now fails withNot found; repair drops exactly that fragment (id,physical_rows=1, path), commits a new version; 2 rows remain and compaction succeeds again.missing_fragment_failure_triggers_repair_and_rerun(etcd): compaction of a broken store fails → aRepairand a dependentCompactappear → repairDonewithdropped 1 fragments (1 rows)→ rerunDone→list_repairshas one record with the right target, trigger and path → 3 rows readable.missing_fragment_error_is_recognised: positive and negative cases (a missing manifest or WAL generation is not repaired).--include-ignored74/74, core compaction tests, clippy-D warnings, fmt clean.🤖 Generated with Claude Code