fix: opt-in recovery for stalled WAL merges with progress deadlines - #292
Merged
Merged
Conversation
beinan
marked this pull request as ready for review
October 1, 2026 21:21
beinan
marked this pull request as draft
October 2, 2026 00:54
beinan
marked this pull request as ready for review
October 2, 2026 03:36
beinan
added a commit
that referenced
this pull request
Oct 2, 2026
#293) Operators need to inspect persistent WAL merge failures without probing workers or opening WAL payloads. Add `GET /api/v1/scheduler/merge-failures?after=<cursor>`, returning at most 256 records plus a continuation cursor. Invalid cursors return HTTP 400. Reads do not reset failure counts, change backoff, or trigger retries. This is the observation-only follow-up split from #292. Runtime ownership, fencing, timeouts, worker triggers, and retry budgets remain in #292. **Stacked review:** the base branch mirrors #292 at `c0c0b5a` so this diff contains only the API, its metadata-pagination test, and documentation. Do not merge into the mirror branch. After #292 merges, retarget this PR to `main` and rerun checks. Validation: Clippy with warnings denied passed. The isolated-etcd regression passed: 257-record pagination, invalid cursors, and unchanged failure counts/retry deadlines. Parent runtime validation is documented in #292; CI is rerunning after synchronizing the stack. No production deployment is included. Co-authored-by: Beinan Wang <>
beinan
added a commit
that referenced
this pull request
Oct 2, 2026
Store-directory lookups currently fetch Lance manifests on worker cache misses, and registry writes contend on shared Lance datasets. This PR adds an opt-in etcd registry for rollout, generic and datagen stores. Payload storage and the default Lance backend are unchanged. ## Implementation - Introduce `StoreRegistry`, Lance and etcd implementations, and shared etcd connection configuration. Preserve existing worker caches, Lance registry maintenance, and #292's merge recovery and serial fan-out. - Support Lance authority with an etcd migration mirror. Master startup, maintenance, backfill and diff cover all three registry kinds, including datagen. - Reconcile complete Lance snapshots: repair missed deletions and changed URIs, preserve original timestamps, and compare full entries rather than names alone. Per-entry source versions, retained tombstones and a CAS-protected snapshot floor reject delayed older writes; interrupted passes can resume. - Validate all entries before first etcd activation and seal subsequent mirror writes. Refuse direct authoritative writes to an unsealed migration mirror. Reverse mirroring is rejected at startup. ## Migration constraints Deploying this PR does not select etcd. The supported sequence is Lance → Lance with etcd mirror → authoritative etcd without a mirror. Before switching authority, quiesce and drain **all registry writers**, including discovery and retirement in auxiliary masters, reconcile all three kinds and verify empty diffs. The application does not implement a fleet-wide quiescence barrier; an empty live diff does not authorize a rolling backend flip. After activation, Lance is a historical snapshot, not a live rollback copy. Returning to Lance requires a separate verified export/import while writers are quiescent. Retain migration state/tombstones and do not recreate the source dataset at the same URI. See `docs/src/design/registry-etcd.md` for the procedure and partial-activation recovery. ## Validation Regression coverage includes missed mirror deletions, changed URIs and timestamps, stale snapshot/point-write ordering, interrupted reconciliation, cutover refusal/sealing, tombstone replacement by authoritative discovery, reverse-mirror rejection and datagen startup/migration routes. Validated in the combined tree with #294, #295 and #296 using isolated local etcd: - 20 core registry, 97 master, 92 worker, 15 coordinator, 2 commit-guard and 4 write-scope tests passed, including etcd-backed cases. - All-target Clippy, formatting and diff checks passed. - CI runs the core etcd registry regressions and includes them in coverage. - Adapted the new no-op compaction regressions to the registry/config APIs after rebasing on `211ad24`. No registry backend has been enabled or deployed by this PR update. CI is running on the current PR head. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Beinan Wang <>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A stalled WAL merge can hold a master task indefinitely. Retrying after only an HTTP timeout can overlap old and new base-table writes. This change owns each merge independently of its HTTP connection, closes commit admission on recovery, and fences every admitted manifest version before handing the table to another executor. Worker fan-out remains serial.
MERGE_OWNED_TARGETSandMERGE_DRAIN_TARGETS, empty by default. Deploy capable binaries first, drain only the selected table, then enable it. An owned request never falls back to legacy HTTP. Drain reconciles existing owned work while preventing new maintenance mutations on that target; other tables retain their job pools.Validation for the final runtime changes: master 88, worker 92, and shared protocol 13 tests passed, including isolated-etcd cases for slot waiting, no-progress cancellation, HTTP disconnects, orphan fencing, drain recovery probes, and retained retry budgets. Clippy with warnings denied and formatting passed on the final source. Actual delayed-commit/new-WAL/blob recovery tests passed. The compaction/merge concurrency test now permits only Lance's typed
RetryableCommitConflictfrom index preparation, retries once after compaction completes, and verifies the physical row count plus exact IDs; the corrected test passed once plus 10/10 repeated runs. The broad core run completed: 258 passed / 3 existing benchmarks ignored, with only the concurrency assumption described above failing; that corrected test was then validated in the 11 runs above. CI is running onc0c0b5a.Draft: staging soak is still required. Legacy untracked writes cannot be automatically fenced, and changing flags is not proof that they drained. All replicas/helpers must follow the per-table migration and rollback sequence in
docs/merge-recovery.md. Progress inside opaque Lance operations is coarse: measure realistic large-blob phase latency, queue time, memory/OOM, etcd overhead and p95/p99 before enabling production targets. No production deployment or pod/configuration changes are part of this PR.