Skip to content

Sync master-aligned delayed delivery & cursor persistence work into branch-as-4.0 (CI verification #2) - #99

Closed
nodece wants to merge 5 commits into
branch-as-4.0from
branch-as-4.0-08282121
Closed

nodece wants to merge 5 commits into
branch-as-4.0from
branch-as-4.0-08282121

Conversation

@nodece

@nodece nodece commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator

Summary

Test plan

  • CI green (purpose of this PR)
  • Local: 49/49 managed-ledger (checkpoint suite + PositionRangeSet), prior CI run on the same tree: all suites green except the known flaky-runner false positives

Dirty bits are stored at ledgerId + 1 (markDirty adds [L1+1, L2+2) and
snapshotAndClearDirtyLedgers shifts back by one), but isDirtyLedgers
probed dirtyLedgers.contains(ledgerId), under-reporting by one ledger:
a ledger marked dirty returned false. Align the probe with the +1
storage offset and fix the test that was written against the buggy
output (a lower-bound ledger may also be extended, so markDirty covers
[L1, L2] and isDirtyLedgers must agree with snapshotAndClearDirtyLedgers).
…t fails

A cursor close has no next flush to retry with; swallowing the final
checkpoint failure left the last ack state neither in BK nor ZK while
close still reported success. When the per-msgLedger checkpoint fails
while closing, persist the legacy metadata-store snapshot (ack ranges
truncated to the configured limit) instead — same durability as the
legacy close path. An md-only tombstone is deliberately not used here:
with cursorsLedgerId=-1 recovery would skip the BK checkpoints and drop
every persisted hole, replaying more than this fallback. Also document
the known non-atomic flush window in CursorCheckpointPersistence.
A reset's mark-delete persist snapshotted the pre-reset individual acks
(memory is only cleared after the persist succeeds), so a crash right
after a successful reset recovered the pre-reset holes: messages the
reset rewound to were filtered as already-deleted and never redelivered.

Persist reset entries as hole-free state on both paths: the per-msgLedger
path writes a checkpoint with mark-delete and properties only (no ack
states, no refs) and clears lastCheckpointPos so every earlier checkpoint
becomes unrecoverable; the legacy PositionInfo path skips the ack ranges
and batch-deletion indexes for reset entries. The metadata-store fallback
was already md-only, so all three forms now agree.

Verified with testCheckpointResetCursorCrashRecoveryNoHoles: ack 4..7,
reset to 2, recover with a fresh factory (no close) — holes 4..7 must be
gone. Legacy behavior also benefits since the fix is path-independent.
…and document the downgrade constraint

PersistentTopic.getInternalStats rebuilt per-cursor CursorStats without
copying individualDeletedMessagesCount / firstIndividualDeletedMessage,
so topic-level internal-stats always reported 0/null while the
managed-ledger-level stats had the real values.

Also document on persistentUnackedRangesWithPerLedgerEntryEnabled that
the feature changes the on-disk cursor-ledger format: brokers without
the feature cannot parse those ledgers, so clusters that may need to
roll back must not enable it.
When a legacy-path reset's cursor-ledger append fails, the metadata-store
fallback persisted the ack ranges read from memory — but the reset's
align (which clears the in-memory holes) only runs after the persist
succeeds, so the fallback wrote the pre-reset holes into the ZK
tombstone and a crash right after the reset recovered them.

Pass persistIndividualDeletedMessageRanges=!propagatePersistFailure so a
reset entry falls back md-only, matching the per-msgLedger path's
fallback semantics. Covered by testLegacyResetBkFailureZkFallbackIsMdOnly.
@nodece nodece closed this Sep 8, 2026
@nodece
nodece deleted the branch-as-4.0-08282121 branch September 8, 2026 07:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant