Repository navigation
fix(lb): a write waiting for the automatic primary follows the election - #1679
Open
roshangara wants to merge 1 commit into
Open
roshangara wants to merge 1 commit into
roshangara wants to merge 1 commit into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
With `role = "auto"`, get_primary() checked out from the pool elected
when the write arrived and waited there for up to checkout_timeout. If
that primary went down, the write stayed in its queue after the shard
monitor had elected another server, and failed with a checkout timeout,
or got a new connection to the old primary if it came back as a replica
("cannot execute ... in a read-only transaction").
A write now follows the election for up to checkout_timeout: while no
server is the primary it waits, and when the election changes it drops
its pending checkout and checks out from the new primary.
The election channel is written only when the elected pool changes, so
a monitor tick that confirms the same primary doesn't wake waiting
writes, and a configuration reload publishes the primary carried over
to the new pools.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
roshangara
force-pushed
the
fix/lb-primary-follows-election
branch
from
October 6, 2026 19:41
913e778 to
5f9e451
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
With
role = "auto",LoadBalancer::get_primary()checks out from the pool that is primary when the write arrives, and stays on it:If that primary goes down, the write waits in its pool's queue for the whole
checkout_timeout, even after the shard monitor has elected another server a few hundred milliseconds later. The write then fails withcheckout timeout. If the old server comes back as a replica first, the write gets a new connection to it and fails withcannot execute ... in a read-only transaction. That is the sequence logged in #1494: the checkout starts, the new primary is chosen 0.5 s later, and the checkout completes 4.4 s after that on the old server, now in recovery.Reproduced on
mainwith PostgreSQL 18.6: a primary and a streaming standby, bothrole = "auto",lsn_check_interval = 100,checkout_timeout = 10000, and one client inserting every 20 ms through PgDog. The primary is stopped (pg_ctl stop -m immediate) and the standby promoted 1.1 s later. The next write waited the full 10 s and failed withcheckout timeout. Writes resumed 8.8 s after the promotion. Two runs gave the same result.Fix
get_primary()follows the election for up tocheckout_timeout. It waits while no server is primary. When the elected pool changes, it drops its pending checkout and checks out from the new primary. Static roles are unchanged.send_if_modified). Before, every monitor run sentNoneand then the primary, and with this change each such send would restart the waiting checkouts.move_conns_to()publishes the primary carried over to the new load balancer, so a write after a configuration reload doesn't wait for the next election.wait_primary()is folded intoget_primary().Tests
lb::test::test_waiting_write_moves_to_new_primary: a write waits on an elected primary that refuses connections. After a replica is promoted, the write gets a connection to it within 1 s, well insidecheckout_timeout. Fails onmain, where the write stays on the dead pool.lb::test::test_unchanged_election_does_not_wake_waiting_writes: an election that confirms the same primary doesn't notify waiting writes, and losing the primary does. Fails onmain.lb::test::test_write_after_reload_goes_to_elected_primary: aftermove_conns_to(), a write goes to the carried-over primary without waiting for an election. It passes onmainand fails if thepublish_primary()call inmove_conns_to()is removed.test_election_channel_no_race_conditionnow goes throughget_primary()instead of the removedwait_primary().checkout timeout, in two runs. The only error, in one run, was the write in flight when the primary was killed.cargo nextest run --profile devforpgdog,pgdog-config,pgdog-stats,pgdog-vectorandpgdog-postgres-typespasses. Threebackend::replicationtests shell out topg_dump. On this machine they need PostgreSQL 18'spg_dumpfirst inPATH(the system one is 16), and with it they pass on this branch and onmain.cargo fmt --all -- --checkandcargo clippy --all-targets -- -D warningsare clean.This change and "fix(lb): elect the automatic primary by timeline, demote a primary in recovery" edit neighbouring lines of
redetect_roles(), so whichever merges second needs a trivial rebase. Otherwise they are independent.Refs #1494, #1255
Note
The patch has been written by Claude.
🤖 Generated with Claude Code