Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -226,6 +226,14 @@ _Avoid_: Release (that word belongs to the Concurrency-Limit slot), return the L
One execution try of a job, numbered and visible to the handler. A lease expiry counts as an attempt, the same as a thrown exception.
_Avoid_: Retry (retry is attempts after the first; counting "retries" invites off-by-one ambiguity)

**Retry Cause**:
Why a job last went back to Scheduled because an Attempt went wrong: the handler failed and the retry policy scheduled another Attempt (**handler failed**), or the Lease lapsed before the worker reported an outcome (**lease expired**). Recorded on the job row by the store, atomically with the reschedule. Sticky: a claim, a Relinquish, a Cancel and a terminal outcome leave it as it is (a terminal job keeps it as a record of its last retry); only an operator requeue clears it, because the requeued job starts over. A new job has none. It names the kind of problem only; the message is the Failure Detail on the Transition Log.
_Avoid_: Retry reason, last error (it is a kind, not a message), Attempt > 0 (a relinquished or requeued job has Attempts but no problem)

**Retrying**:
A Scheduled job that carries a Retry Cause: still live, waiting for another Attempt because the last one went wrong. Not a state of its own but a view of Scheduled, so an operator sees trouble before it ends as Dead-Lettered. A new job, a requeued job and a job a stopping worker handed back are Scheduled but not Retrying, so a deploy raises no false alarm. Leaves the view when the job is claimed again.
_Avoid_: Failed (the job is not terminal), Retried (it has not run again yet), Retry state (there is no such state)

**At-Least-Once Execution**:
BackWave's delivery contract: a job's handler body may run more than once, and idempotency is the handler author's responsibility. The field standard — Hangfire, Sidekiq, River, Celery (acks-late), RabbitMQ-with-acks, and Temporal *activities* all land here. Exactly-once *body* execution is not offered, because the only two roads to it are both rejected: at-most-once (accept job loss on crash) or durable execution. What BackWave still guarantees exactly once is the Effect-Once property.
_Avoid_: Exactly-once execution, at-most-once, deliver-once
Expand Down
190 changes: 190 additions & 0 deletions src/BackWave.Conformance/ConformanceSuite.cs
Original file line number Diff line number Diff line change
Expand Up @@ -1878,6 +1878,196 @@ public async Task Clause_5_9_CountMatchingJobs_EmptyStore_IsZero()
new JobQuery { State = JobState.Scheduled, TagPredicates = [JobTagPredicate.HasLabel("urgent")] }));
}

// ── §5.9 Retrying: the jobs an attempt went wrong for ───────────────────────

private static async Task<IReadOnlyList<JobRecord>> ListRetryingAsync(IJobStore store, string? queue = null)
=> await store.ListJobsAsync(new JobQuery { Retrying = true, Queue = queue, SortDirection = JobSortDirection.OldestFirst });

/// <summary>
/// Certifies that a handler failure the retry policy reschedules records the HandlerFailed cause, so
/// the Scheduled job lists as Retrying, and that the Retrying filter ANDs with the other filters.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_AFailureRetry_IsRetrying_WithTheHandlerFailedCause()
{
var store = await CreateStoreAsync();
await store.EnqueueAsync(Job(), now: T0);
var claimed = Assert.Single(await ClaimAsync(store, T0));

var retryAt = T0.AddMinutes(5);
await store.ReportOutcomeAsync(claimed.JobId, "w1", claimed.Attempt, new JobOutcome.Failure(retryAt, "transient"), T0);

var job = await store.GetJobAsync(claimed.JobId);
Assert.Equal(JobState.Scheduled, job!.State);
Assert.Equal(RetryCause.HandlerFailed, job.RetryCause);
var listed = Assert.Single(await ListRetryingAsync(store));
Assert.Equal(claimed.JobId, listed.JobId);
Assert.Equal(RetryCause.HandlerFailed, listed.RetryCause);
Assert.Equal(retryAt, listed.DueTime);
Assert.Equal(1, listed.Attempt);
Assert.Empty(await ListRetryingAsync(store, queue: "other"));
}

/// <summary>
/// Certifies that the batched outcome path records the same HandlerFailed cause as a single report
/// for a failure it reschedules, and records none for a success or a dead-letter in the same batch.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_ABatchedFailureRetry_IsRetrying_LikeASingleReport()
{
var store = await CreateStoreAsync();
var retried = Job();
var succeeded = Job();
var dead = Job();
await store.EnqueueAsync(retried, now: T0);
await store.EnqueueAsync(succeeded, now: T0);
await store.EnqueueAsync(dead, now: T0);
var claimed = await ClaimAsync(store, T0);
OutcomeReport Report(NewJob job, JobOutcome outcome)
=> new(job.JobId, "w1", claimed.Single(j => j.JobId == job.JobId).Attempt, outcome);

await store.ReportOutcomesAsync(
[
Report(retried, new JobOutcome.Failure(T0.AddMinutes(5), "transient")),
Report(succeeded, new JobOutcome.Success()),
Report(dead, new JobOutcome.Failure(null, "fatal")),
], T0);

Assert.Equal(RetryCause.HandlerFailed, (await store.GetJobAsync(retried.JobId))!.RetryCause);
Assert.Null((await store.GetJobAsync(succeeded.JobId))!.RetryCause);
Assert.Null((await store.GetJobAsync(dead.JobId))!.RetryCause);
Assert.Equal(retried.JobId, Assert.Single(await ListRetryingAsync(store)).JobId);
}

/// <summary>
/// Certifies that a lapsed lease the expiry sweep reschedules records the LeaseExpired cause, so the
/// Scheduled job lists as Retrying.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_ALeaseExpiry_IsRetrying_WithTheLeaseExpiredCause()
{
var store = await CreateStoreAsync();
await store.EnqueueAsync(Job(), now: T0);
var claimed = Assert.Single(await ClaimAsync(store, T0));

var afterExpiry = T0 + Lease + TimeSpan.FromSeconds(1);
Assert.Equal(1, await store.ExpireLeasesAsync(afterExpiry, maxJobs: 32, DefaultQueues, TwoAttempts));

var job = await store.GetJobAsync(claimed.JobId);
Assert.Equal(JobState.Scheduled, job!.State);
Assert.Equal(RetryCause.LeaseExpired, job.RetryCause);
var listed = Assert.Single(await ListRetryingAsync(store));
Assert.Equal(claimed.JobId, listed.JobId);
Assert.Equal(RetryCause.LeaseExpired, listed.RetryCause);
}

/// <summary>
/// Certifies that a clean-stop hand-back is not a problem: a job that never went wrong comes back
/// Scheduled without a cause and does not list as Retrying, while a job that was already Retrying
/// keeps its cause through the hand-back, so a deploy neither raises nor hides an alarm.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_ARelinquish_IsNotRetrying_AndKeepsAnEarlierCause()
{
var store = await CreateStoreAsync();
var healthy = Job();
var failing = Job();
await store.EnqueueAsync(healthy, now: T0);
await store.EnqueueAsync(failing, now: T0);
var first = await ClaimAsync(store, T0);
var failed = first.Single(j => j.JobId == failing.JobId);
await store.ReportOutcomeAsync(failed.JobId, "w1", failed.Attempt, new JobOutcome.Failure(T0, "transient"), T0);
Assert.Single(await ClaimAsync(store, T0)); // the failing job's second attempt; both are now leased

// Three attempts, so the failing job's second attempt is below the ceiling and is handed back
// rather than dead-lettered.
var threeAttempts = new RetryPolicy { MaxAttempts = 3, Backoff = _ => TimeSpan.FromMinutes(1) }.ToDisposition();
var handBack = T0.AddSeconds(5);
if (!Declares(ConformanceCapabilities.LeaseRelinquish))
{
await AssertRelinquishIsANoOpAsync(store, handBack, threeAttempts, healthy.JobId, failing.JobId);
return;
}
Assert.Equal(2, await store.RelinquishLeasesAsync("w1", handBack, threeAttempts));

var handedBack = await store.GetJobAsync(healthy.JobId);
Assert.Equal(JobState.Scheduled, handedBack!.State);
Assert.Null(handedBack.RetryCause);
Assert.Equal(RetryCause.HandlerFailed, (await store.GetJobAsync(failing.JobId))!.RetryCause);
Assert.Equal(failing.JobId, Assert.Single(await ListRetryingAsync(store)).JobId);
}

/// <summary>
/// Certifies that a terminal outcome keeps the last retry cause as a record, and that an operator
/// requeue clears it: the requeued job starts over and does not list as Retrying.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_ARequeue_IsNotRetrying_AndClearsTheCause()
{
var store = await CreateStoreAsync();
await store.EnqueueAsync(Job(), now: T0);
var first = Assert.Single(await ClaimAsync(store, T0));
await store.ReportOutcomeAsync(first.JobId, "w1", first.Attempt, new JobOutcome.Failure(T0, "transient"), T0);
var second = Assert.Single(await ClaimAsync(store, T0));
await store.ReportOutcomeAsync(second.JobId, "w1", second.Attempt, new JobOutcome.Failure(null, "fatal"), T0);

var dead = await store.GetJobAsync(first.JobId);
Assert.Equal(JobState.DeadLettered, dead!.State);
Assert.Equal(RetryCause.HandlerFailed, dead.RetryCause); // kept as the record of the last retry
Assert.Empty(await ListRetryingAsync(store)); // terminal, so not Retrying

var requeueTime = T0.AddMinutes(1);
Assert.Equal(RequeueResult.Requeued, await store.RequeueAsync(first.JobId, "alice", requeueTime));

var requeued = await store.GetJobAsync(first.JobId);
Assert.Equal(JobState.Scheduled, requeued!.State);
Assert.Null(requeued.RetryCause);
Assert.Empty(await ListRetryingAsync(store));
}

/// <summary>
/// Certifies that new work is never Retrying: a job enqueued due now, one enqueued for later, and one
/// awaiting a parent all carry no cause.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_AFreshEnqueue_IsNotRetrying()
{
var store = await CreateStoreAsync();
var dueNow = Job();
var later = Job(dueTime: T0.AddHours(1));
await store.EnqueueAsync(dueNow, now: T0);
await store.EnqueueAsync(later, now: T0);
var child = Job() with { Parents = [dueNow.JobId] };
await store.EnqueueAsync(child, now: T0);

foreach (var id in (Guid[])[dueNow.JobId, later.JobId, child.JobId])
{
Assert.Null((await store.GetJobAsync(id))!.RetryCause);
}
Assert.Empty(await ListRetryingAsync(store));
}

/// <summary>
/// Certifies that a Retrying job leaves the Retrying list once it is claimed again, and stays gone
/// when that attempt succeeds.
/// </summary>
[Fact]
public async Task Clause_5_9_Retrying_ClaimedAgainAndSucceeded_LeavesTheRetryingList()
{
var store = await CreateStoreAsync();
await store.EnqueueAsync(Job(), now: T0);
var first = Assert.Single(await ClaimAsync(store, T0));
await store.ReportOutcomeAsync(first.JobId, "w1", first.Attempt, new JobOutcome.Failure(T0, "transient"), T0);
Assert.Single(await ListRetryingAsync(store));

var second = Assert.Single(await ClaimAsync(store, T0));
Assert.Empty(await ListRetryingAsync(store)); // running, not waiting

await store.ReportOutcomeAsync(second.JobId, "w1", second.Attempt, new JobOutcome.Success(), T0);
Assert.Equal(JobState.Succeeded, (await store.GetJobAsync(first.JobId))!.State);
Assert.Empty(await ListRetryingAsync(store));
}

// ── Mutation teeth: boundary/misc contract facts (issue 0235) ────────────────

/// <summary>
Expand Down
4 changes: 4 additions & 0 deletions src/BackWave.Dashboard/Components/JobDetailPanel.razor
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,10 @@
{
<tr><th>Terminal cause</th><td>@cause</td></tr>
}
@if (Job.RetryCause is { } retryCause)
{
<tr><th>Retry cause</th><td>@DashboardGlossary.RetryCauseName(retryCause)</td></tr>
}
@if (Job.ScheduleId is { } scheduleId)
{
<tr>
Expand Down
30 changes: 26 additions & 4 deletions src/BackWave.Dashboard/Components/JobTable.razor
Original file line number Diff line number Diff line change
Expand Up @@ -15,8 +15,16 @@ else
<th>Queue</th>
<th>State</th>
<th class="num">Attempt</th>
<th>Due Time</th>
<th>Terminal At</th>
@if (ShowRetryCause)
{
<th>Next Attempt</th>
<th>Retry Cause</th>
}
else
{
<th>Due Time</th>
<th>Terminal At</th>
}
@if (ShowTags)
{
<th>Tags</th>
Expand All @@ -42,8 +50,16 @@ else
</span>
</td>
<td class="num" data-label="Attempt">@job.Attempt</td>
<td class="mono nowrap" data-label="Due Time">@DashboardGlossary.Instant(job.DueTime)</td>
<td class="mono nowrap" data-label="Terminal At">@(job.TerminalAt is { } terminalAt ? DashboardGlossary.Instant(terminalAt) : "—")</td>
@if (ShowRetryCause)
{
<td class="mono nowrap" data-label="Next Attempt">@DashboardGlossary.Instant(job.DueTime)</td>
<td class="nowrap" data-label="Retry Cause">@(job.RetryCause is { } retryCause ? DashboardGlossary.RetryCauseName(retryCause) : "-")</td>
}
else
{
<td class="mono nowrap" data-label="Due Time">@DashboardGlossary.Instant(job.DueTime)</td>
<td class="mono nowrap" data-label="Terminal At">@(job.TerminalAt is { } terminalAt ? DashboardGlossary.Instant(terminalAt) : "—")</td>
}
@if (TagHref is { } tagHref)
{
@* Tag pills (ADR 0022, issue 0113): an empty tag set renders NOTHING — no empty
Expand Down Expand Up @@ -95,5 +111,11 @@ else
/// </summary>
[Parameter] public Func<JobTag, string>? TagHref { get; set; }

/// <summary>
/// Show the Retrying columns: the due time reads as the next attempt, and the retry cause (handler
/// failed or lease expired) takes the place of the terminal instant, which a live job never has.
/// </summary>
[Parameter] public bool ShowRetryCause { get; set; }

private bool ShowTags => TagHref is not null;
}
Loading
Loading