You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
With role = "auto", failover goes wrong in three ways: two primaries, a demoted primary that keeps the role, and writes stuck on a dead primary #1731
main at 298565e (2026-10-06). Reproduced first on 5d516b5.
Description
We run PgDog in front of a Patroni cluster (PostgreSQL 18.6) with every server on role = "auto". Failover tests showed three bugs in LoadBalancer. All the code below is unchanged on current main.
Two primaries after a failover.redetect_roles() elects the non-recovery server whose LSN stats were fetched most recently (targets.sort_by_cached_key(|target| target.0.lsn_age(now))). The old primary can come back before it is turned into a replica: a crash and restart, a network partition, a promotion while the old primary still runs. Then both servers report pg_is_in_recovery() = false. Whichever answered last wins, the election flips between them, and writes go to both. The old primary's writes are lost when it is rewound.
Reproduction: a primary and a streaming standby, lsn_check_interval = 100, one client inserting every 20 ms. Promote the standby while the old primary keeps running. In three runs, the old primary acknowledged 62, 115 and 117 of about 250 writes after the promotion, and the elected primary changed 7 to 14 times.
Reproduction in a 3-node Patroni lab: stop the primary, promote a replica, restart the old primary on its old timeline for 15 s. Writes split 151/143, and pg_rewind later discarded the 151.
A primary back in recovery keeps the role. When no server reports being a primary, roles change only if every server has valid stats (} else if targets.iter().all(|target| target.0.valid()) {). If another server hasn't answered an LSN check yet (e.g. it is down), a primary that comes back as a replica stays Primary. Writes then fail with cannot execute ... in a read-only transaction. This is one way to get Pgdog sometimes do not switch primary #1255.
A write waits on the dead primary for the whole checkout_timeout.get_primary() checks out from the pool that is primary when the write arrives and stays on it. If that server goes down, the write waits in its pool's queue even after another server has been elected. It fails with checkout timeout. If the old server comes back as a replica first, the write fails with cannot execute ... in a read-only transaction instead, which is the sequence logged in fix(pool): fence stale automatic-primary checkouts #1494.
Reproduction: checkout_timeout = 10000, one client inserting every 20 ms. Stop the primary with pg_ctl stop -m immediate and promote the standby 1.1 s later. The next write waited the full 10 s and failed with checkout timeout. Writes resumed 8.8 s after the promotion. Two runs gave the same result.
Expected
A promoted server always wins the election over the old primary. A promotion starts a new timeline, so the timeline orders them whatever their LSNs and whenever they were last checked.
A server that reports being in recovery stops being the primary, even while another server has no stats.
A write waiting for the primary moves to the newly elected one within checkout_timeout.
Fix branches
Both branches are on main5d516b5 and merge cleanly with 298565e. Each change comes with tests that fail on main.
The LSN query also returns the primary's timeline: ('x' || substr(pg_walfile_name(pg_current_wal_lsn()), 1, 8))::bit(32)::int. It is 0 on replicas and on Aurora.
The election orders servers by timeline, then LSN, then stats age.
A primary that reports being in recovery is demoted even while another server has no stats.
With this branch, no write after the promotion reached the old primary (0 of about 254, three runs), and the elected primary changed once.
PgDog version
mainat 298565e (2026-10-06). Reproduced first on 5d516b5.Description
We run PgDog in front of a Patroni cluster (PostgreSQL 18.6) with every server on
role = "auto". Failover tests showed three bugs inLoadBalancer. All the code below is unchanged on currentmain.Two primaries after a failover.
redetect_roles()elects the non-recovery server whose LSN stats were fetched most recently (targets.sort_by_cached_key(|target| target.0.lsn_age(now))). The old primary can come back before it is turned into a replica: a crash and restart, a network partition, a promotion while the old primary still runs. Then both servers reportpg_is_in_recovery() = false. Whichever answered last wins, the election flips between them, and writes go to both. The old primary's writes are lost when it is rewound.lsn_check_interval = 100, one client inserting every 20 ms. Promote the standby while the old primary keeps running. In three runs, the old primary acknowledged 62, 115 and 117 of about 250 writes after the promotion, and the elected primary changed 7 to 14 times.pg_rewindlater discarded the 151.A primary back in recovery keeps the role. When no server reports being a primary, roles change only if every server has valid stats (
} else if targets.iter().all(|target| target.0.valid()) {). If another server hasn't answered an LSN check yet (e.g. it is down), a primary that comes back as a replica staysPrimary. Writes then fail withcannot execute ... in a read-only transaction. This is one way to get Pgdog sometimes do not switch primary #1255.A write waits on the dead primary for the whole
checkout_timeout.get_primary()checks out from the pool that is primary when the write arrives and stays on it. If that server goes down, the write waits in its pool's queue even after another server has been elected. It fails withcheckout timeout. If the old server comes back as a replica first, the write fails withcannot execute ... in a read-only transactioninstead, which is the sequence logged in fix(pool): fence stale automatic-primary checkouts #1494.checkout_timeout = 10000, one client inserting every 20 ms. Stop the primary withpg_ctl stop -m immediateand promote the standby 1.1 s later. The next write waited the full 10 s and failed withcheckout timeout. Writes resumed 8.8 s after the promotion. Two runs gave the same result.Expected
checkout_timeout.Fix branches
Both branches are on
main5d516b5 and merge cleanly with 298565e. Each change comes with tests that fail onmain.('x' || substr(pg_walfile_name(pg_current_wal_lsn()), 1, 8))::bit(32)::int. It is 0 on replicas and on Aurora.get_primary()follows the election for up tocheckout_timeout.send_if_modified).checkout timeout.The two branches edit neighbouring lines of
redetect_roles(), so the second one needs a trivial rebase.Configuration
Note
Written with Claude.