Skip to content

fix(lb): a write waiting for the automatic primary follows the election - #1679

Open
roshangara wants to merge 1 commit into
pgdogdev:mainfrom
roshangara:fix/lb-primary-follows-election
Open

roshangara wants to merge 1 commit into
pgdogdev:mainfrom
roshangara:fix/lb-primary-follows-election

Conversation

@roshangara

Copy link
Copy Markdown

Problem

With role = "auto", LoadBalancer::get_primary() checks out from the pool that is primary when the write arrives, and stays on it:

if let Some(pool) = self.primary() {
    return pool.get(request).await;
}

If that primary goes down, the write waits in its pool's queue for the whole checkout_timeout, even after the shard monitor has elected another server a few hundred milliseconds later. The write then fails with checkout timeout. If the old server comes back as a replica first, the write gets a new connection to it and fails with cannot execute ... in a read-only transaction. That is the sequence logged in #1494: the checkout starts, the new primary is chosen 0.5 s later, and the checkout completes 4.4 s after that on the old server, now in recovery.

Reproduced on main with PostgreSQL 18.6: a primary and a streaming standby, both role = "auto", lsn_check_interval = 100, checkout_timeout = 10000, and one client inserting every 20 ms through PgDog. The primary is stopped (pg_ctl stop -m immediate) and the standby promoted 1.1 s later. The next write waited the full 10 s and failed with checkout timeout. Writes resumed 8.8 s after the promotion. Two runs gave the same result.

Fix

  • With automatic roles, get_primary() follows the election for up to checkout_timeout. It waits while no server is primary. When the elected pool changes, it drops its pending checkout and checks out from the new primary. Static roles are unchanged.
  • The election channel is now written only when the elected pool changes (send_if_modified). Before, every monitor run sent None and then the primary, and with this change each such send would restart the waiting checkouts.
  • move_conns_to() publishes the primary carried over to the new load balancer, so a write after a configuration reload doesn't wait for the next election.
  • wait_primary() is folded into get_primary().

Tests

  • lb::test::test_waiting_write_moves_to_new_primary: a write waits on an elected primary that refuses connections. After a replica is promoted, the write gets a connection to it within 1 s, well inside checkout_timeout. Fails on main, where the write stays on the dead pool.
  • lb::test::test_unchanged_election_does_not_wake_waiting_writes: an election that confirms the same primary doesn't notify waiting writes, and losing the primary does. Fails on main.
  • lb::test::test_write_after_reload_goes_to_elected_primary: after move_conns_to(), a write goes to the carried-over primary without waiting for an election. It passes on main and fails if the publish_primary() call in move_conns_to() is removed.
  • test_election_channel_no_race_condition now goes through get_primary() instead of the removed wait_primary().
  • The end-to-end run above, with this branch: writes stopped for 1.22 s, 1.1 s of which came before the promotion. They resumed as soon as the standby was promoted, with no checkout timeout, in two runs. The only error, in one run, was the write in flight when the primary was killed.
  • cargo nextest run --profile dev for pgdog, pgdog-config, pgdog-stats, pgdog-vector and pgdog-postgres-types passes. Three backend::replication tests shell out to pg_dump. On this machine they need PostgreSQL 18's pg_dump first in PATH (the system one is 16), and with it they pass on this branch and on main.
  • cargo fmt --all -- --check and cargo clippy --all-targets -- -D warnings are clean.

This change and "fix(lb): elect the automatic primary by timeline, demote a primary in recovery" edit neighbouring lines of redetect_roles(), so whichever merges second needs a trivial rebase. Otherwise they are independent.

Refs #1494, #1255

Note

The patch has been written by Claude.

🤖 Generated with Claude Code

@CLAassistant

CLAassistant commented Sep 29, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codecov

codecov Bot commented Sep 29, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

With `role = "auto"`, get_primary() checked out from the pool elected
when the write arrived and waited there for up to checkout_timeout. If
that primary went down, the write stayed in its queue after the shard
monitor had elected another server, and failed with a checkout timeout,
or got a new connection to the old primary if it came back as a replica
("cannot execute ... in a read-only transaction").

A write now follows the election for up to checkout_timeout: while no
server is the primary it waits, and when the election changes it drops
its pending checkout and checks out from the new primary.

The election channel is written only when the elected pool changes, so
a monitor tick that confirms the same primary doesn't wake waiting
writes, and a configuration reload publishes the primary carried over
to the new pools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants