Skip to content

After a promotion, pg_control_checkpoint() reports the old timeline for minutes, so #1738 doesn't separate the two primaries yet #1739

Description

@roshangara

PgDog version

main at 784f0b2 (#1738).

Description

Follow-up to #1731 (issue 1) and #1738. The LSN query reads the timeline from (pg_control_checkpoint()).timeline_id. This is the timeline of the last completed checkpoint. A promotion starts the new timeline at once, but its checkpoint is requested without CHECKPOINT_IMMEDIATE. So it is spread over checkpoint_completion_target × checkpoint_timeout, which is 0.9 × 5 min with the defaults. Until that checkpoint completes, the promoted server reports the old timeline. An old primary that is still running reports the same timeline, so the election falls back to the freshest stats and flips between the two, as before #1738.

Measured on PostgreSQL 18.3 with default checkpoint settings. A primary and a streaming standby, 200k rows updated on the primary just before pg_promote():

t=0s    promoted: pg_control_checkpoint tl=1, pg_walfile_name tl=2 | old primary: tl=1
t=131s  promoted: pg_control_checkpoint tl=1, pg_walfile_name tl=2 | old primary: tl=1
t=262s  promoted: pg_control_checkpoint tl=1, pg_walfile_name tl=2 | old primary: tl=1
t=272s  promoted: pg_control_checkpoint tl=2, pg_walfile_name tl=2 | old primary: tl=1

LOG:  checkpoint starting: force
LOG:  checkpoint complete: wrote 11770 buffers (71.8%); ... write=269.336 s, sync=0.006 s, total=269.353 s

The LSN query from main and from the branch below, right after pg_promote() (columns: replica, lsn, offset_bytes, timestamp, timeline):

standby, main:        t 0/B95CFE8 194367464 ... 1
standby, branch:      t 0/B95CFE8 194367464 ... 1
promoted, main:       f 0/B95D020 194367520 ... 1
promoted, branch:     f 0/B95D020 194367520 ... 2
old primary, main:    f 0/B95CFE8 194367464 ... 1
old primary, branch:  f 0/B95CFE8 194367464 ... 1

A manager may run its own CHECKPOINT after promoting, which shortens the window. A plain pg_promote() or pg_ctl promote doesn't.

Expected

The promoted server reports the new timeline as soon as it accepts writes.

Fix branch

https://github.com/roshangara/pgdog/tree/fix/lb-timeline-from-wal-file (one commit on 784f0b2, +10/−1 in lsn_monitor.rs)

  • For a primary, the timeline is read from the name of its current WAL file, as Patroni does: ('x' || substr(pg_walfile_name(lsn), 1, 8))::bit(32)::int. It changes at the promotion.
  • Replicas still report pg_control_checkpoint(), because pg_walfile_name() fails during recovery. The CASE keeps it from being called there, and the query runs without error on the standby above.
  • test_run_check_detects_non_aurora now also checks that a live primary reports a timeline of at least 1. A unit test can't promote a server, so the measurement above is the reproduction.
  • The lb, lsn_monitor and shard::monitor tests pass. cargo fmt --check and clippy -D warnings are clean.

We couldn't check whether Aurora allows pg_control_checkpoint(), which #1738 added to AURORA_LSN_QUERY. This branch doesn't change that query.

Configuration

Any role = "auto" setup with default checkpoint_timeout and checkpoint_completion_target.

Note

Written with Claude.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions