Skip to content

[redis_enterprise_prometheus] Fix dashboard queries and collect two missing metrics - #3193

Open
slorello89 wants to merge 2 commits into
DataDog:masterfrom
redis-field-engineering:redis-enterprise-prometheus-dashboard-metric-fixes
Open

slorello89 wants to merge 2 commits into
DataDog:masterfrom
redis-field-engineering:redis-enterprise-prometheus-dashboard-metric-fixes

Conversation

@slorello89

Copy link
Copy Markdown
Contributor

What does this PR do?

Fixes default-dashboard widgets in redis_enterprise_prometheus that return no data, and collects two metrics that were documented in metadata.csv but missing from the check's metric map.

Check

  • Collect rdse2.node_uname_info and rdse2.node_config. Both were already in metadata.csv but absent from metrics.py, so they never emitted. node_uname_info carries the per-node nodename label — the internal hostname the Node dashboard needs — and node_config carries rs_version.

Dashboards

  • Database List: database_syncer_config → db_config. The former is a configuration-label placeholder that only exists on Active-Active clusters, and it lives in the opt-in REDIS2.REPLICATION group, so the widget was empty on a default install.
  • Database Input/Output: add the missing .count suffix to endpoint_ingress / endpoint_egress. These are Prometheus counters, which OpenMetrics v2 emits as .count.
  • Connections: compute current connections from the .total gauges and subtract endpoint_proxy_disconnections.total, rather than differencing two monotonic .count deltas (which yields net change over the interval, not a current count).
  • Double rate division: Database Input/Output, Shard Process CPU, and Proxy Threads CLI Session each wrapped an already-rated .as_rate() query in per_second() / derivative(), dividing by time twice.
  • Ingress/egress labels were swapped in Database Input/Output — alias egress pointed at the ingress query. The accompanying note text had the definitions backwards too.
  • Node Latency: rebuilt as a calculated field over the endpoint_*_requests_latency_histogram sum/count pairs. It previously queried rdse.node_avg_latency — a V1 metric, under the rdse. namespace rather than rdse2.. Histogram buckets run 0.25–16908, so the unit is microseconds and the formula divides by 1000 for milliseconds.
  • Cluster Nodes: group by nodename from node_uname_info instead of the internal-hostname tag, which is not present on any /v2 metric.
  • Node Network Traffic: removed. Both rdse.node_egress_bytes_median and rdse.node_ingress_bytes_median are V1-only with no V2 equivalent.
  • Note widgets: dropped title, folding it into the content as a markdown heading. The Dashboards API rejects title on note widgets, which made every one of these dashboard assets fail to import.

Motivation

Datadog raised a list of redis_enterprise_prometheus dashboard widgets showing no data. Each item was verified against a live Redis Enterprise 8.0.10-76 /v2 endpoint rather than against the repo fixture, which turned out to be stale in places — for example process_resident_memory_bytes and process_virtual_memory_bytes exist in tests/data/metrics.txt but are gone from RS 8.0.10 entirely.

Two findings worth calling out:

  • The internal-hostname tag genuinely does not exist on any /v2 metric, but node_uname_info already exposes the same information as nodename, so no upstream Redis change is needed.
  • ddev validate dashboards passes both before and after the note-widget fix, so it does not currently catch assets that the Dashboards API will reject on import.

Review checklist

  • PR has a meaningful title or PR has the no-changelog label attached
  • Feature or bugfix has tests
  • Git history is clean
  • If PR impacts documentation, docs team has been notified or an issue has been opened on the documentation repo
  • If this PR includes a log pipeline, please add a description describing the remappers and processors.

Additional Notes

Dashboard asset changes are not unit-testable, so no tests are added; each query was instead validated against a live cluster. Two caveats on that validation:

  • The cluster used had a single database, so the Node Latency formula could not be checked against a real latency value — the microsecond assumption comes from the histogram bucket boundaries. Worth a second look from someone with a busier cluster.
  • Three widgets on the Shard dashboard (Resident Memory, Virtual Memory, Process CPU) are still broken and are not addressed here. Resident/Virtual Memory query rdse2.redis_process_*, which has never existed under any name; the viable source is namedprocess_namegroup_memory_bytes{memtype="resident"|"virtual"}. Process CPU has data but groups by {db,redis} on namedprocess_* series that carry neither tag (shards appear as groupname="redis-<n>"). Fixing those changes the widgets' semantics from shard-scoped to process-group-scoped, which felt like a separate discussion.

Replacement for #3136, which was automatically closed as stale.

slorello89 and others added 2 commits August 26, 2026 10:24
…sing metrics

Datadog reported several default-dashboard widgets returning no data. Verified
each against a live Redis Enterprise 8.0.10 /v2 endpoint.

Check:
- Collect node_uname_info and node_config. Both were already documented in
  metadata.csv but absent from the metric map, so they never emitted.
  node_uname_info carries the per-node `nodename` label, which is the internal
  hostname the Node dashboard needs; node_config carries `rs_version`.

Dashboards:
- Database List: database_syncer_config -> db_config. The former is a config
  placeholder that only exists on Active-Active clusters and sits in the opt-in
  REDIS2.REPLICATION group, so the widget was empty by default.
- endpoint_ingress / endpoint_egress: add the missing `.count` suffix. These are
  Prometheus counters and OpenMetrics v2 emits them as `.count`.
- Connections: compute from the `.total` gauges and subtract
  endpoint_proxy_disconnections.total, instead of differencing monotonic deltas.
- Remove double rate division in Database Input/Output, Shard Process CPU and
  Proxy Threads CLI Session, which wrapped `.as_rate()` queries in
  per_second()/derivative().
- Un-swap the ingress/egress series labels and fix the note text describing them.
- Node Latency: rebuild as a calculated field over the
  endpoint_*_requests_latency_histogram sum/count pairs (microseconds -> ms).
  It previously queried rdse.node_avg_latency, a V1 metric under the wrong
  namespace.
- Cluster Nodes: group by `nodename` rather than the non-existent
  `internal-hostname` tag.
- Remove Node Network Traffic; both rdse.node_{in,e}gress_bytes_median are
  V1-only with no V2 equivalent.
- Drop `title` from note widgets, folding it into the content as a markdown
  heading. The Dashboards API rejects `title` on notes, which made every one of
  these dashboard assets fail to import.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@slorello89
slorello89 requested a review from a team as a code owner September 30, 2026 13:50
@slorello89
slorello89 requested review from JoshPatel13 and removed request for a team September 30, 2026 13:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant