PgDog version
main at 784f0b2 (#1738).
Description
Follow-up to #1731 (issue 1) and #1738. The LSN query reads the timeline from (pg_control_checkpoint()).timeline_id. This is the timeline of the last completed checkpoint. A promotion starts the new timeline at once, but its checkpoint is requested without CHECKPOINT_IMMEDIATE. So it is spread over checkpoint_completion_target × checkpoint_timeout, which is 0.9 × 5 min with the defaults. Until that checkpoint completes, the promoted server reports the old timeline. An old primary that is still running reports the same timeline, so the election falls back to the freshest stats and flips between the two, as before #1738.
Measured on PostgreSQL 18.3 with default checkpoint settings. A primary and a streaming standby, 200k rows updated on the primary just before pg_promote():
t=0s promoted: pg_control_checkpoint tl=1, pg_walfile_name tl=2 | old primary: tl=1
t=131s promoted: pg_control_checkpoint tl=1, pg_walfile_name tl=2 | old primary: tl=1
t=262s promoted: pg_control_checkpoint tl=1, pg_walfile_name tl=2 | old primary: tl=1
t=272s promoted: pg_control_checkpoint tl=2, pg_walfile_name tl=2 | old primary: tl=1
LOG: checkpoint starting: force
LOG: checkpoint complete: wrote 11770 buffers (71.8%); ... write=269.336 s, sync=0.006 s, total=269.353 s
The LSN query from main and from the branch below, right after pg_promote() (columns: replica, lsn, offset_bytes, timestamp, timeline):
standby, main: t 0/B95CFE8 194367464 ... 1
standby, branch: t 0/B95CFE8 194367464 ... 1
promoted, main: f 0/B95D020 194367520 ... 1
promoted, branch: f 0/B95D020 194367520 ... 2
old primary, main: f 0/B95CFE8 194367464 ... 1
old primary, branch: f 0/B95CFE8 194367464 ... 1
A manager may run its own CHECKPOINT after promoting, which shortens the window. A plain pg_promote() or pg_ctl promote doesn't.
Expected
The promoted server reports the new timeline as soon as it accepts writes.
Fix branch
https://github.com/roshangara/pgdog/tree/fix/lb-timeline-from-wal-file (one commit on 784f0b2, +10/−1 in lsn_monitor.rs)
- For a primary, the timeline is read from the name of its current WAL file, as Patroni does:
('x' || substr(pg_walfile_name(lsn), 1, 8))::bit(32)::int. It changes at the promotion.
- Replicas still report
pg_control_checkpoint(), because pg_walfile_name() fails during recovery. The CASE keeps it from being called there, and the query runs without error on the standby above.
test_run_check_detects_non_aurora now also checks that a live primary reports a timeline of at least 1. A unit test can't promote a server, so the measurement above is the reproduction.
- The
lb, lsn_monitor and shard::monitor tests pass. cargo fmt --check and clippy -D warnings are clean.
We couldn't check whether Aurora allows pg_control_checkpoint(), which #1738 added to AURORA_LSN_QUERY. This branch doesn't change that query.
Configuration
Any role = "auto" setup with default checkpoint_timeout and checkpoint_completion_target.
PgDog version
mainat 784f0b2 (#1738).Description
Follow-up to #1731 (issue 1) and #1738. The LSN query reads the timeline from
(pg_control_checkpoint()).timeline_id. This is the timeline of the last completed checkpoint. A promotion starts the new timeline at once, but its checkpoint is requested withoutCHECKPOINT_IMMEDIATE. So it is spread overcheckpoint_completion_target×checkpoint_timeout, which is 0.9 × 5 min with the defaults. Until that checkpoint completes, the promoted server reports the old timeline. An old primary that is still running reports the same timeline, so the election falls back to the freshest stats and flips between the two, as before #1738.Measured on PostgreSQL 18.3 with default checkpoint settings. A primary and a streaming standby, 200k rows updated on the primary just before
pg_promote():The LSN query from
mainand from the branch below, right afterpg_promote()(columns: replica, lsn, offset_bytes, timestamp, timeline):A manager may run its own
CHECKPOINTafter promoting, which shortens the window. A plainpg_promote()orpg_ctl promotedoesn't.Expected
The promoted server reports the new timeline as soon as it accepts writes.
Fix branch
https://github.com/roshangara/pgdog/tree/fix/lb-timeline-from-wal-file (one commit on 784f0b2, +10/−1 in
lsn_monitor.rs)('x' || substr(pg_walfile_name(lsn), 1, 8))::bit(32)::int. It changes at the promotion.pg_control_checkpoint(), becausepg_walfile_name()fails during recovery. TheCASEkeeps it from being called there, and the query runs without error on the standby above.test_run_check_detects_non_auroranow also checks that a live primary reports a timeline of at least 1. A unit test can't promote a server, so the measurement above is the reproduction.lb,lsn_monitorandshard::monitortests pass.cargo fmt --checkandclippy -D warningsare clean.We couldn't check whether Aurora allows
pg_control_checkpoint(), which #1738 added toAURORA_LSN_QUERY. This branch doesn't change that query.Configuration
Any
role = "auto"setup with defaultcheckpoint_timeoutandcheckpoint_completion_target.Note
Written with Claude.