Skip to content

Answer 503 while the database is down, also behind HTTP Basic (B02) - #17

Closed
AsyncAssassin wants to merge 3 commits into
mainfrom
fix/db-outage-503
Closed

AsyncAssassin wants to merge 3 commits into
mainfrom
fix/db-outage-503

Conversation

@AsyncAssassin

@AsyncAssassin AsyncAssassin commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

B02 from the merged review of v0.4.0, with D05 and D10 from the same review; 0.4.1 is released once all fixes are merged. docs/api.md promises 503 database-unavailable while PostgreSQL is unavailable, but the service answered otherwise:

  • local and test: 7 of the 9 endpoints answered 500 internal-error. A transaction that cannot get a connection fails with CannotCreateTransactionException, and a rollback on a connection the outage broke fails with TransactionSystemException; neither is a DataAccessException. An SQL error Spring cannot classify, such as a write to a read-only database, stayed a jOOQ exception and got 500 as well.
  • demo, prod, and every other protected profile: a request with valid credentials got 401 with a Basic challenge, because HTTP Basic could not read its user store (InternalAuthenticationServiceException).
  • A server that stops answering (paused container, network partition): a query in flight waited forever, because pgjdbc has no socket timeout by default and statement_timeout needs a live server.

What changes:

  • One rule, isDatabaseFailure(), decides what a database failure is for the API, the credential check, and the sync worker: a Spring DataAccessException, a transaction that could not begin or roll back, or an exception whose direct cause is an SQLException. Other transaction exceptions, such as UnexpectedRollbackException, stay 500. A sync run that meets such a failure records Database error (<class>). instead of Unexpected error (<class>)..
  • The entry point answers a user store failure caused by the database with the same 503, and any other with the generic 500, both without a challenge. A request without credentials keeps its 401, and so does a username the store cannot hold, such as one with a NUL character. Both paths write the same database_operation_failed line.
  • The pool sets socketTimeout=40, above the 30-second statement_timeout, and connectTimeout=5. The 30-second Hikari connection-timeout stays: a shorter one would give false 503s under load.
  • Docs (D05, D10):
    • docs/failure-modes.md section 6 is the one full description of the outage. It covers the timeouts, and a COMMIT that the socket timeout cuts off but that may still commit, with what that means for a retry. It also notes that every request with credentials waits for the pool during the outage. The other files link to it.
    • The immutable-conflict examples show the fields the service writes, in its order.
    • The error list names the 404 for an unsupported asset and the 409 for a duplicate account.
    • docs/failure-modes.md names the meters and log lines that exist instead of promising them "when metrics are implemented".

CHANGELOG.md lists the changes under [Unreleased] → Fixed.

Verification

  • ./gradlew clean check: 280 tests, 0 failures, 2 skipped (the env-gated Alchemy live smoke).

  • DatabaseOutageIntegrationTests starts its own PostgreSQL and boots a test and an e2e context against it:

    • with the database up, the operator's credentials work, and a username with a NUL character gets 401 with the challenge;
    • the pool carries the connect timeout and a socket timeout above the statement_timeout its connections run with;
    • after the container stops, all 9 endpoints answer 503 database-unavailable without a challenge in both contexts, the anonymous request keeps its 401, and health turns 503.

    On the first version's handlers the same test reports 7 × 500 in test and 9 × 401 in e2e. It now runs in 9 s instead of 29 s.

  • Unit tests use the real exception classes:

    • CannotCreateTransactionException, TransactionSystemException, and a jOOQ exception around an SQLException give 503; UnexpectedRollbackException and jOOQ's TooManyRowsException give 500;
    • the entry point gives 503 for a lookup that failed on CannotGetJdbcConnectionException, 401 with the challenge for one that failed on a DataIntegrityViolationException, and 500 otherwise.
  • SyncApiIntegrationTests: a TransactionSystemException during a run leaves it QUEUED with Database error (TransactionSystemException)..

  • Mutation check: dropping the NUL guard, the SQLException rule, or the worker rule, a socket timeout below statement_timeout, or a misspelt connectTimeout each turns a test red.

  • Live demo against its own PostgreSQL container, the main jar against this branch:

    Scenario main This branch
    database up, username with a NUL character 401 with challenge 401 with challenge
    docker stop, demo-reader GET 401 with challenge 503 database-unavailable, no challenge
    docker stop, demo-operator POST 401 with challenge 503 database-unavailable, no challenge
    docker stop, no credentials 401 with challenge 401 with challenge
    docker start again 404 (recovered) 404 (recovered)
    docker pause with a query in flight no answer after 100 s 503 after 40 s

    The docker pause row bypasses Hikari's connection check (-Dcom.zaxxer.hikari.aliveBypassWindowMs=3600000) to reproduce a connection taken right before the pause. A request that waits for a new connection gets its answer after the 30-second Hikari connection-timeout in both versions.

A regression review of the first version (/code-review) found that a NUL username on a healthy database got the outage answer. It also found untranslated jOOQ errors, the worker's run error, the COMMIT caveat, and weaker tests and docs; all are fixed in 0665722. Left as a documented limit: every request with credentials reads the user store, so during an outage even authenticated metrics requests wait for the pool. A credential cache or a separate management port would change that, and it is a decision of its own.

This PR, #15, #16, #18, and #19 each add a section under [Unreleased] in CHANGELOG.md, so each later merge needs a one-file rebase. The other files merge cleanly with all of them.

The API promised 503 database-unavailable while PostgreSQL is down, but
only a DataAccessException got it. A transaction that cannot get a
connection fails with CannotCreateTransactionException, and a rollback
on a connection the outage broke fails with TransactionSystemException.
Both are TransactionExceptions, so seven of the nine endpoints answered
500 in local. In demo and prod, HTTP Basic could not read its user
store, and the entry point answered 401 with a challenge, blaming valid
credentials for the outage.

Both transaction failures now get the same 503; other transaction
exceptions, such as an unexpected rollback, stay 500. The entry point
answers a user store failure caused by the database with that 503 and
any other with the generic 500, both without a challenge, and logs
database_operation_failed like the API. pgjdbc gets a 40-second socket
timeout, above the 30-second statement_timeout, and a 5-second connect
timeout, so a query to a server that stopped answering ends instead of
holding the request thread and its connection forever.

A test stops a dedicated PostgreSQL under running test and e2e contexts
and checks every endpoint, the anonymous 401, health, and the socket
timeout of pooled connections.
The immutable-conflict ProblemDetail example in docs/api.md and
docs/architecture.md still said "amount or direction did not match" and
lacked the address, asset, and conflictingFields the service writes.
The error list in docs/architecture.md left out the 404 for an
unsupported asset and the 409 for a duplicate account.

docs/failure-modes.md promised metrics "when metrics are implemented"
three times, although the ingest and transition counters exist, and a
stale-confirmation log line with stored and incoming counts that the
service never writes. It now names the meters and log lines as they
are.
Review of the outage fix found that any database error behind the user
store lookup became the outage answer: a Basic username with a NUL
character, which PostgreSQL refuses as a data error, got 503 and an
ERROR database_operation_failed line on a healthy database, where it
used to get 401. Such a username is a bad credential again.

One rule now decides what a database failure is, for the API, the
credential check, and the sync worker: a Spring data access exception,
a transaction that could not begin or roll back, or an SQL error Spring
could not translate, such as a write to a read-only database, which
answered 500. A sync run that meets one records "Database error" instead
of "Unexpected error". Both entry paths write the same log line.

The outage test checks that the credentials worked before the stop, a
NUL username on the healthy database, the connect timeout, and that the
socket timeout stays above statement_timeout; it sends the requests in
parallel with the shortest pool wait, from 29 to 9 seconds. The docs
describe the outage once, in docs/failure-modes.md, with the COMMIT
that the socket timeout can cut off and its retry, and the examples show
the fields in the order the service writes them.
@AsyncAssassin

Copy link
Copy Markdown
Owner Author

Combined into #23 together with the other batch-1 fixes, so they merge without a rebase per PR. This description keeps the detailed evidence for its fix.

@AsyncAssassin
AsyncAssassin deleted the fix/db-outage-503 branch September 24, 2026 05:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant