Skip to content

fix: opt-in recovery for stalled WAL merges with progress deadlines - #292

Merged
beinan merged 9 commits into
lance-format:mainfrom
beinan:codex/bounded-merge-recovery
Oct 2, 2026
Merged

beinan merged 9 commits into
lance-format:mainfrom
beinan:codex/bounded-merge-recovery

Conversation

@beinan

@beinan beinan commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

A stalled WAL merge can hold a master task indefinitely. Retrying after only an HTTP timeout can overlap old and new base-table writes. This change owns each merge independently of its HTTP connection, closes commit admission on recovery, and fences every admitted manifest version before handing the table to another executor. Worker fan-out remains serial.

  • Default-off, per-table rollout: explicit MERGE_OWNED_TARGETS and MERGE_DRAIN_TARGETS, empty by default. Deploy capable binaries first, drain only the selected table, then enable it. An owned request never falls back to legacy HTTP. Drain reconciles existing owned work while preventing new maintenance mutations on that target; other tables retain their job pools.
  • Keep pending/timer triggers: fix(server): the count-triggered self-merge takes the worker's merge slot #289 configuration remains valid. For enabled rollout/generic tables the worker checks its own manifest and posts coalesced demand; the master picks it up every 15 seconds through normal dedupe, serial ownership, and worker slot/byte limits. Unselected tables retain direct self-merge. Scheduling latency for enabled tables needs measurement.
  • Separate time budgets: slot queue defaults to 600 seconds; execution starts after slot acquisition and has a configurable 3600-second ceiling. A separate 600-second no-progress limit observes completed batch/phase/storage work. Repeated heartbeats do not extend it. No hard-coded 600-second execution cap remains. Timeouts cancel executions, not pods, and handoff still requires completed writes or proven storage fences.
  • Worker etcd connection is lazy: outage does not fail worker startup or disable ingestion/flush. Owned admission and further commit authorization fail closed; already-authorized writes may finish.
  • Bound repeated failures: durable per-target/endpoint budgets survive task recreation and worker demand; three initial transient attempts, then backoff, hourly probes and attention after 15 failures. Healthy shards remain eligible; partial failures stay visible. Missing data never triggers destructive merge repair. The read-only failure inspection API is split into feat(master): inspect merge failures through a read-only paginated API #293 (118 lines across the API, its pagination regression, and documentation).

Validation for the final runtime changes: master 88, worker 92, and shared protocol 13 tests passed, including isolated-etcd cases for slot waiting, no-progress cancellation, HTTP disconnects, orphan fencing, drain recovery probes, and retained retry budgets. Clippy with warnings denied and formatting passed on the final source. Actual delayed-commit/new-WAL/blob recovery tests passed. The compaction/merge concurrency test now permits only Lance's typed RetryableCommitConflict from index preparation, retries once after compaction completes, and verifies the physical row count plus exact IDs; the corrected test passed once plus 10/10 repeated runs. The broad core run completed: 258 passed / 3 existing benchmarks ignored, with only the concurrency assumption described above failing; that corrected test was then validated in the 11 runs above. CI is running on c0c0b5a.

Draft: staging soak is still required. Legacy untracked writes cannot be automatically fenced, and changing flags is not proof that they drained. All replicas/helpers must follow the per-table migration and rollback sequence in docs/merge-recovery.md. Progress inside opaque Lance operations is coarse: measure realistic large-blob phase latency, queue time, memory/OOM, etcd overhead and p95/p99 before enabling production targets. No production deployment or pod/configuration changes are part of this PR.

@beinan beinan changed the title fix: bound owned WAL merges and persist shard recovery budgets fix: fence stalled WAL merges and persist bounded shard retries Oct 1, 2026
@beinan
beinan marked this pull request as ready for review October 1, 2026 21:21
@beinan
beinan marked this pull request as draft October 2, 2026 00:54
@beinan beinan changed the title fix: fence stalled WAL merges and persist bounded shard retries fix: opt-in recovery for stalled WAL merges with progress deadlines Oct 2, 2026
@beinan
beinan marked this pull request as ready for review October 2, 2026 03:36
@beinan
beinan merged commit 772c73b into lance-format:main Oct 2, 2026
15 checks passed
beinan added a commit that referenced this pull request Oct 2, 2026
#293)

Operators need to inspect persistent WAL merge failures without probing
workers or opening WAL payloads. Add `GET
/api/v1/scheduler/merge-failures?after=<cursor>`, returning at most 256
records plus a continuation cursor. Invalid cursors return HTTP 400.
Reads do not reset failure counts, change backoff, or trigger retries.

This is the observation-only follow-up split from #292. Runtime
ownership, fencing, timeouts, worker triggers, and retry budgets remain
in #292.

**Stacked review:** the base branch mirrors #292 at `c0c0b5a` so this
diff contains only the API, its metadata-pagination test, and
documentation. Do not merge into the mirror branch. After #292 merges,
retarget this PR to `main` and rerun checks.

Validation: Clippy with warnings denied passed. The isolated-etcd
regression passed: 257-record pagination, invalid cursors, and unchanged
failure counts/retry deadlines. Parent runtime validation is documented
in #292; CI is rerunning after synchronizing the stack. No production
deployment is included.

Co-authored-by: Beinan Wang <>
beinan added a commit that referenced this pull request Oct 2, 2026
Store-directory lookups currently fetch Lance manifests on worker cache
misses, and registry writes contend on shared Lance datasets. This PR
adds an opt-in etcd registry for rollout, generic and datagen stores.
Payload storage and the default Lance backend are unchanged.

## Implementation

- Introduce `StoreRegistry`, Lance and etcd implementations, and shared
etcd connection configuration. Preserve existing worker caches, Lance
registry maintenance, and #292's merge recovery and serial fan-out.
- Support Lance authority with an etcd migration mirror. Master startup,
maintenance, backfill and diff cover all three registry kinds, including
datagen.
- Reconcile complete Lance snapshots: repair missed deletions and
changed URIs, preserve original timestamps, and compare full entries
rather than names alone. Per-entry source versions, retained tombstones
and a CAS-protected snapshot floor reject delayed older writes;
interrupted passes can resume.
- Validate all entries before first etcd activation and seal subsequent
mirror writes. Refuse direct authoritative writes to an unsealed
migration mirror. Reverse mirroring is rejected at startup.

## Migration constraints

Deploying this PR does not select etcd. The supported sequence is Lance
→ Lance with etcd mirror → authoritative etcd without a mirror. Before
switching authority, quiesce and drain **all registry writers**,
including discovery and retirement in auxiliary masters, reconcile all
three kinds and verify empty diffs. The application does not implement a
fleet-wide quiescence barrier; an empty live diff does not authorize a
rolling backend flip.

After activation, Lance is a historical snapshot, not a live rollback
copy. Returning to Lance requires a separate verified export/import
while writers are quiescent. Retain migration state/tombstones and do
not recreate the source dataset at the same URI. See
`docs/src/design/registry-etcd.md` for the procedure and
partial-activation recovery.

## Validation

Regression coverage includes missed mirror deletions, changed URIs and
timestamps, stale snapshot/point-write ordering, interrupted
reconciliation, cutover refusal/sealing, tombstone replacement by
authoritative discovery, reverse-mirror rejection and datagen
startup/migration routes. Validated in the combined tree with #294, #295
and #296 using isolated local etcd:

- 20 core registry, 97 master, 92 worker, 15 coordinator, 2 commit-guard
and 4 write-scope tests passed, including etcd-backed cases.
- All-target Clippy, formatting and diff checks passed.
- CI runs the core etcd registry regressions and includes them in
coverage.
- Adapted the new no-op compaction regressions to the registry/config
APIs after rebasing on `211ad24`.

No registry backend has been enabled or deployed by this PR update. CI
is running on the current PR head.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Beinan Wang <>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant