Repository navigation
Conversation
The server waits for open requests before it stops the hosted services. An open dashboard tab held its live stream until the host shutdown timeout ended. The worker groups then had no time to give their leases back on a clean stop, and the leases expired later. The stream now also ends on ApplicationStopping.
An abandoned handler can unwind on a thread-pool thread and log while the test reads the records. The read then failed with "Collection was modified". The scope stack was also shared across threads. The capture now locks its records and returns a copy. Scopes use LoggerExternalScopeProvider, which keeps one scope stack per async flow. LeaseReclaimLogTests failed 3 times in 25 runs before and 0 times in 40 runs after.
A Scheduled job did not tell an operator if an attempt went wrong. A job with Attempt > 0 is not a good signal, because a clean-stop relinquish also gives back a claimed job. A deploy then looks like a wave of problems. Each job now records a retry cause: - HandlerFailed: the handler failed and the retry policy rescheduled the job (single and batched outcome reports). - LeaseExpired: the lease expired and the sweep rescheduled the job. A requeue clears the cause. A claim, a relinquish, a terminal outcome, a cancel and a parent latch keep it. A new job has no cause. Retrying means Scheduled with a retry cause. It is available as: - JobQuery.Retrying and JobSnapshot.RetryCause on the Monitor API. - A Retrying tab on the Failures page, and a Retrying option in the Jobs state filter. Both show the next attempt and the cause. - A Retry cause row on the job detail page. - A "retrying" filter on the MCP search_jobs tool. Each SQL adapter gets one additive migration: a nullable retry_cause column and an index for the Retrying query. Postgres and Oracle go to version 2, SQL Server and SQLite to version 3. There is no backfill. Jobs that are Scheduled before the upgrade have no cause, so they do not show as Retrying. The Oracle migrator now sets ddl_lock_timeout on its own unpooled session. Without it, the new ALTER TABLE failed with ORA-00054 when other nodes ran the first script at the same time on a cold boot.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A Scheduled job now records why it went back to Scheduled. Retrying means Scheduled with a retry cause. An operator sees the jobs that have a problem, and a deploy does not show as a wave of problems.
Why not
Attempt > 0: a clean-stop relinquish gives back claimed jobs withAttempt > 0. Every deploy would look like a problem.Surfaces:
Schema: one additive migration for each SQL adapter. There is no backfill. A job that is Scheduled before the upgrade has no cause, so it does not show as Retrying.
0002_retry_cause.sqlix_backwave_jobs_retrying0003_retry_cause.sqlix_backwave_jobs_retrying0003_retry_cause.sqlix_backwave_jobs_retrying0002_retry_cause.sqlix_bw_jobs_retrying (retry_cause)Two fixes that the work found:
LeaseExpired. That is a false alarm on deploy. The stream now also ends onApplicationStopping.ALTER TABLEfailed with ORA-00054 while another node ran the v1 script. The migrator now setsddl_lock_timeout = 30on its own unpooled session.There is also one flake fix:
CapturingLoggerwas not safe when two threads wrote to it, soLeaseReclaimLogTestsfailed 3 times in 25 runs.Evidence
Tests (local, private compose ports):
BackWave.Tests)LiveView_EndsTheSseStream_WhenTheApplicationStartsToStopreads the stream until the 15 s timeout.After: the stream ends right after
StopApplication().LeaseReclaimLogTestsfailed 3 of 25 runs ("Collection was modified").After: 0 of 40 runs.
E2E (Sample.Api on SQLite, real workers, dashboard clicks):
bd6f407bflakye1a720c6processfacdfd89greet40bc05c0greetee6c80a9flaky/backwave/executingopen, the app stopped in under 1 s.app.log: "Worker group 'strict-priority' relinquished 2 lease(s) on shutdown."search_jobs {retrying: true}returned onlybd6f407bande1a720c6.Screenshots (kept local, not uploaded):
failures-retrying: Retrying tab, count 2, rows for the 2 problem jobs with Next Attempt and Retry Cause.failures-retrying-light: the same page in the light theme.jobs-scheduled: the Scheduled filter shows all 4 Scheduled jobs.jobs-retrying: the Retrying filter shows only the 2 problem jobs.detail-e1a720c6/detail-bd6f407b: "Retry cause" row reads "Lease expired" / "Handler failed".detail-facdfd89: the relinquished job has no Retry cause row.Merge Danger
Door: one-way
Each SQL adapter gets a schema version bump. The column is additive and nullable, and the N-1 fleet and in-place upgrade tests pass. A rollback of the code still leaves the column and the new version stamp. An older build then refuses the newer schema version.
Blast Radius: schema
All four SQL adapters migrate on the next start. Writes to
retry_causeare on the outcome, expiry and requeue paths. The Oracle round-trip budget is unchanged except for the job page window bytes. The dashboard Failures page gets a third tab. The Jobs table gets two different columns only when the Retrying filter is set.