Skip to content

Add OpenTelemetry metrics for retry and archive operations - #5836

Merged
johnsimons merged 14 commits into
masterfrom
john/open_telemetry_error_management
Sep 1, 2026
Merged

johnsimons merged 14 commits into
masterfrom
john/open_telemetry_error_management

Conversation

@johnsimons

Copy link
Copy Markdown
Member

Adds telemetry for error management, the follow-up to the ingestion (#5816) and retention (#5831) metrics. Retry and archive were the last long-running operations in the error instance with no signal at all: a retry of a large group can run for hours across four stages and two polling loops, and until now the only way to see it was the log.

What is instrumented

sc.retry.* covers the whole retry pipeline: an end-to-end operation duration from request to completion, a duration histogram per stage (prepare, stage, forward), a message counter by outcome, an in-progress gauge, and the depth of the bulk request queue. The abandoned message outcome is the one worth alerting on, since it counts messages that hit the staging retry limit and were dropped from their batch, which previously existed only as a log line. The in-progress gauge is what surfaces a stuck operation, because an operation that never completes never records a duration.

sc.archive.* covers group archive and unarchive: whole-operation duration, per-batch duration, a message counter, and an in-progress gauge. The instrumentation lives once in the shared InMemoryArchive/InMemoryUnarchive state machines, but only the EF persistence registers ArchiveMetrics, so on RavenDB the family does not exist rather than being half-populated. No RavenDB file is touched.

The host also gains the standard ASP.NET Core, HTTP client and runtime instrumentation packages, behind the existing OTEL_EXPORTER_OTLP_ENDPOINT gate. That covers the read APIs per route, the scatter-gather calls to remote instances, and GC/thread pool context. AddIngestionMetrics is renamed AddTelemetry now that it wires more than ingestion.

All tag values come from enums, so series counts are bounded; group ids, endpoint names and queue addresses stay out and are recovered from the audit log when a specific operation needs investigating. The full instrument contract is documented in docs/telemetry.md.

@johnsimons johnsimons self-assigned this Sep 1, 2026
Comment on lines +44 to +51
static string ResultTag(ScopeOutcome outcome) => outcome switch
{
ScopeOutcome.Success => "success",
ScopeOutcome.Empty => "empty",
ScopeOutcome.Cancelled => "cancelled",
ScopeOutcome.Failed => "failed",
_ => "failed"
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessarily better but you could avoid all the string constants using enum.ToString()

Suggested change
static string ResultTag(ScopeOutcome outcome) => outcome switch
{
ScopeOutcome.Success => "success",
ScopeOutcome.Empty => "empty",
ScopeOutcome.Cancelled => "cancelled",
ScopeOutcome.Failed => "failed",
_ => "failed"
};
static string ResultTag(ScopeOutcome outcome) => (Enum.IsDefined(outcome) ? outcome : ScopeOutcome.Failed).ToString().ToLower();

RetryingManager now depends on RetryMetrics, so tests that register it must also provide a test double to satisfy the dependency.
@johnsimons
johnsimons enabled auto-merge September 1, 2026 03:50
@johnsimons
johnsimons merged commit 20fc6be into master Sep 1, 2026
36 checks passed
@johnsimons
johnsimons deleted the john/open_telemetry_error_management branch September 1, 2026 04:03
@rbev rbev added this to the 6.20.0 milestone Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants