Skip to content

With role = "auto", failover goes wrong in three ways: two primaries, a demoted primary that keeps the role, and writes stuck on a dead primary #1731

Description

@roshangara

PgDog version

main at 298565e (2026-10-06). Reproduced first on 5d516b5.

Description

We run PgDog in front of a Patroni cluster (PostgreSQL 18.6) with every server on role = "auto". Failover tests showed three bugs in LoadBalancer. All the code below is unchanged on current main.

  1. Two primaries after a failover. redetect_roles() elects the non-recovery server whose LSN stats were fetched most recently (targets.sort_by_cached_key(|target| target.0.lsn_age(now))). The old primary can come back before it is turned into a replica: a crash and restart, a network partition, a promotion while the old primary still runs. Then both servers report pg_is_in_recovery() = false. Whichever answered last wins, the election flips between them, and writes go to both. The old primary's writes are lost when it is rewound.

    • Reproduction: a primary and a streaming standby, lsn_check_interval = 100, one client inserting every 20 ms. Promote the standby while the old primary keeps running. In three runs, the old primary acknowledged 62, 115 and 117 of about 250 writes after the promotion, and the elected primary changed 7 to 14 times.
    • Reproduction in a 3-node Patroni lab: stop the primary, promote a replica, restart the old primary on its old timeline for 15 s. Writes split 151/143, and pg_rewind later discarded the 151.
  2. A primary back in recovery keeps the role. When no server reports being a primary, roles change only if every server has valid stats (} else if targets.iter().all(|target| target.0.valid()) {). If another server hasn't answered an LSN check yet (e.g. it is down), a primary that comes back as a replica stays Primary. Writes then fail with cannot execute ... in a read-only transaction. This is one way to get Pgdog sometimes do not switch primary #1255.

  3. A write waits on the dead primary for the whole checkout_timeout. get_primary() checks out from the pool that is primary when the write arrives and stays on it. If that server goes down, the write waits in its pool's queue even after another server has been elected. It fails with checkout timeout. If the old server comes back as a replica first, the write fails with cannot execute ... in a read-only transaction instead, which is the sequence logged in fix(pool): fence stale automatic-primary checkouts #1494.

    • Reproduction: checkout_timeout = 10000, one client inserting every 20 ms. Stop the primary with pg_ctl stop -m immediate and promote the standby 1.1 s later. The next write waited the full 10 s and failed with checkout timeout. Writes resumed 8.8 s after the promotion. Two runs gave the same result.

Expected

  • A promoted server always wins the election over the old primary. A promotion starts a new timeline, so the timeline orders them whatever their LSNs and whenever they were last checked.
  • A server that reports being in recovery stops being the primary, even while another server has no stats.
  • A write waiting for the primary moves to the newly elected one within checkout_timeout.

Fix branches

Both branches are on main 5d516b5 and merge cleanly with 298565e. Each change comes with tests that fail on main.

The two branches edit neighbouring lines of redetect_roles(), so the second one needs a trivial rebase.

Configuration

[general]
lsn_check_interval = 100
checkout_timeout = 10000

[[databases]]
name = "app"
host = "10.0.0.1"
role = "auto"

[[databases]]
name = "app"
host = "10.0.0.2"
role = "auto"

Note

Written with Claude.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions