Skip to content

Limit lock acquisition attempt time - #13

Open
slawwan wants to merge 4 commits into
Pliner:mainfrom
slawwan:lock-acquire-timeout
Open

slawwan wants to merge 4 commits into
Pliner:mainfrom
slawwan:lock-acquire-timeout

Conversation

@slawwan

@slawwan slawwan commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

pg_try_advisory_lock is called without a timeout. If the connection dies silently while the guard waits for the lock (e.g. a standby replica or a new pod during a rolling update), the guard hangs forever and never acquires the lock.

The fix adds acquire_timeout (5 s by default) for each lock acquisition attempt. On timeout:

  • fetchval raises TimeoutError, asyncpg sends a CancelRequest for the query over a separate connection;
  • the attempt is logged as a warning with the TimeoutError traceback;
  • the guard closes the connection and reconnects after reconnect_delay, the same way as on any other connection failure;
  • nothing propagates from run(), func is never started without the lock.

If the server granted the lock but the response was lost, the lock stays with the old session until the session ends. Closing such a connection reliably depends on MagicStack/asyncpg#1361.

Tests: test_acquire_lock_after_silent_disruption_while_waiting fails on main, passes with the fix; test_log_failed_lock_acquisition_attempt covers the warning.

@slawwan
slawwan marked this pull request as ready for review October 6, 2026 14:38
@slawwan
slawwan requested a review from Pliner October 6, 2026 14:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant