Skip to content

feat(master): inspect merge failures through a read-only paginated API - #293

Merged
beinan merged 2 commits into
lance-format:codex/bounded-merge-recoveryfrom
beinan:codex/merge-failure-inspection
Oct 2, 2026
Merged

beinan merged 2 commits into
lance-format:codex/bounded-merge-recoveryfrom
beinan:codex/merge-failure-inspection

Conversation

@beinan

@beinan beinan commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Operators need to inspect persistent WAL merge failures without probing workers or opening WAL payloads. Add GET /api/v1/scheduler/merge-failures?after=<cursor>, returning at most 256 records plus a continuation cursor. Invalid cursors return HTTP 400. Reads do not reset failure counts, change backoff, or trigger retries.

This is the observation-only follow-up split from #292. Runtime ownership, fencing, timeouts, worker triggers, and retry budgets remain in #292.

Stacked review: the base branch mirrors #292 at c0c0b5a so this diff contains only the API, its metadata-pagination test, and documentation. Do not merge into the mirror branch. After #292 merges, retarget this PR to main and rerun checks.

Validation: Clippy with warnings denied passed. The isolated-etcd regression passed: 257-record pagination, invalid cursors, and unchanged failure counts/retry deadlines. Parent runtime validation is documented in #292; CI is rerunning after synchronizing the stack. No production deployment is included.

beinan added a commit that referenced this pull request Oct 2, 2026
…292)

A stalled WAL merge can hold a master task indefinitely. Retrying after
only an HTTP timeout can overlap old and new base-table writes. This
change owns each merge independently of its HTTP connection, closes
commit admission on recovery, and fences every admitted manifest version
before handing the table to another executor. Worker fan-out remains
serial.

- **Default-off, per-table rollout:** explicit `MERGE_OWNED_TARGETS` and
`MERGE_DRAIN_TARGETS`, empty by default. Deploy capable binaries first,
drain only the selected table, then enable it. An owned request never
falls back to legacy HTTP. Drain reconciles existing owned work while
preventing new maintenance mutations on that target; other tables retain
their job pools.
- **Keep pending/timer triggers:** #289 configuration remains valid. For
enabled rollout/generic tables the worker checks its own manifest and
posts coalesced demand; the master picks it up every 15 seconds through
normal dedupe, serial ownership, and worker slot/byte limits. Unselected
tables retain direct self-merge. Scheduling latency for enabled tables
needs measurement.
- **Separate time budgets:** slot queue defaults to 600 seconds;
execution starts after slot acquisition and has a configurable
3600-second ceiling. A separate 600-second no-progress limit observes
completed batch/phase/storage work. Repeated heartbeats do not extend
it. No hard-coded 600-second execution cap remains. Timeouts cancel
executions, not pods, and handoff still requires completed writes or
proven storage fences.
- **Worker etcd connection is lazy:** outage does not fail worker
startup or disable ingestion/flush. Owned admission and further commit
authorization fail closed; already-authorized writes may finish.
- **Bound repeated failures:** durable per-target/endpoint budgets
survive task recreation and worker demand; three initial transient
attempts, then backoff, hourly probes and attention after 15 failures.
Healthy shards remain eligible; partial failures stay visible. Missing
data never triggers destructive merge repair. The read-only failure
inspection API is split into #293 (118 lines across the API, its
pagination regression, and documentation).

Validation for the final runtime changes: master **88**, worker **92**,
and shared protocol **13** tests passed, including isolated-etcd cases
for slot waiting, no-progress cancellation, HTTP disconnects, orphan
fencing, drain recovery probes, and retained retry budgets. Clippy with
warnings denied and formatting passed on the final source. Actual
delayed-commit/new-WAL/blob recovery tests passed. The compaction/merge
concurrency test now permits only Lance's typed
`RetryableCommitConflict` from index preparation, retries once after
compaction completes, and verifies the physical row count plus exact
IDs; the corrected test passed once plus **10/10** repeated runs. The
broad core run completed: **258 passed / 3 existing benchmarks
ignored**, with only the concurrency assumption described above failing;
that corrected test was then validated in the 11 runs above. CI is
running on `c0c0b5a`.

**Draft: staging soak is still required.** Legacy untracked writes
cannot be automatically fenced, and changing flags is not proof that
they drained. All replicas/helpers must follow the per-table migration
and rollback sequence in `docs/merge-recovery.md`. Progress inside
opaque Lance operations is coarse: measure realistic large-blob phase
latency, queue time, memory/OOM, etcd overhead and p95/p99 before
enabling production targets. No production deployment or
pod/configuration changes are part of this PR.

---------

Co-authored-by: Beinan Wang <>
@beinan
beinan marked this pull request as ready for review October 2, 2026 03:36
@beinan
beinan merged commit 421b69e into lance-format:codex/bounded-merge-recovery Oct 2, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant