Skip to content

web: fix 90s shutdown hang on every deploy - #423

Merged
espadonne merged 3 commits into
trunkfrom
fix/web-shutdown-hang
Sep 2, 2026
Merged

espadonne merged 3 commits into
trunkfrom
fix/web-shutdown-hang

Conversation

@espadonne

Copy link
Copy Markdown
Contributor

Every deploy restart produced ~90 s of 502s: Run() logged "shutdown signal received" and then hung until systemd's default TimeoutStopSec SIGKILLed it. The page-cache LISTEN goroutine holds a pool connection and waits on the server's root context, which the web command never cancels, so the deferred pool.Close() blocked forever.

  • f4ac673 web: cancel background context on shutdown so pool close cannot hang (also srv.Close() after the drain window so long polls/SSE release)
  • a898a2b systemd: bound web stop timeout at 30s
  • pagecache: pin that Listen releases its connection on cancel

Test targets:

  • CI Integration step: internal/cache/pagecache (new TestListen_ReleasesConnOnCancel)
  • After deploy: journalctl -u shithubd-web shows Stopping→Stopped in seconds, not 90 s; 502 count per deploy near zero

@espadonne
espadonne merged commit 58f17e0 into trunk Sep 2, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants