[redis_enterprise_prometheus] Fix dashboard queries and collect two missing metrics - #3193
Open
slorello89 wants to merge 2 commits into
Conversation
…sing metrics
Datadog reported several default-dashboard widgets returning no data. Verified
each against a live Redis Enterprise 8.0.10 /v2 endpoint.
Check:
- Collect node_uname_info and node_config. Both were already documented in
metadata.csv but absent from the metric map, so they never emitted.
node_uname_info carries the per-node `nodename` label, which is the internal
hostname the Node dashboard needs; node_config carries `rs_version`.
Dashboards:
- Database List: database_syncer_config -> db_config. The former is a config
placeholder that only exists on Active-Active clusters and sits in the opt-in
REDIS2.REPLICATION group, so the widget was empty by default.
- endpoint_ingress / endpoint_egress: add the missing `.count` suffix. These are
Prometheus counters and OpenMetrics v2 emits them as `.count`.
- Connections: compute from the `.total` gauges and subtract
endpoint_proxy_disconnections.total, instead of differencing monotonic deltas.
- Remove double rate division in Database Input/Output, Shard Process CPU and
Proxy Threads CLI Session, which wrapped `.as_rate()` queries in
per_second()/derivative().
- Un-swap the ingress/egress series labels and fix the note text describing them.
- Node Latency: rebuild as a calculated field over the
endpoint_*_requests_latency_histogram sum/count pairs (microseconds -> ms).
It previously queried rdse.node_avg_latency, a V1 metric under the wrong
namespace.
- Cluster Nodes: group by `nodename` rather than the non-existent
`internal-hostname` tag.
- Remove Node Network Traffic; both rdse.node_{in,e}gress_bytes_median are
V1-only with no V2 equivalent.
- Drop `title` from note widgets, folding it into the content as a markdown
heading. The Dashboards API rejects `title` on notes, which made every one of
these dashboard assets fail to import.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
slorello89
requested review from
JoshPatel13
and removed request for
a team
September 30, 2026 13:50
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes default-dashboard widgets in
redis_enterprise_prometheusthat return no data, and collects two metrics that were documented inmetadata.csvbut missing from the check's metric map.Check
rdse2.node_uname_infoandrdse2.node_config. Both were already inmetadata.csvbut absent frommetrics.py, so they never emitted.node_uname_infocarries the per-nodenodenamelabel — the internal hostname the Node dashboard needs — andnode_configcarriesrs_version.Dashboards
database_syncer_config→db_config. The former is a configuration-label placeholder that only exists on Active-Active clusters, and it lives in the opt-inREDIS2.REPLICATIONgroup, so the widget was empty on a default install..countsuffix toendpoint_ingress/endpoint_egress. These are Prometheus counters, which OpenMetrics v2 emits as.count..totalgauges and subtractendpoint_proxy_disconnections.total, rather than differencing two monotonic.countdeltas (which yields net change over the interval, not a current count)..as_rate()query inper_second()/derivative(), dividing by time twice.egresspointed at the ingress query. The accompanying note text had the definitions backwards too.endpoint_*_requests_latency_histogramsum/count pairs. It previously queriedrdse.node_avg_latency— a V1 metric, under therdse.namespace rather thanrdse2.. Histogram buckets run 0.25–16908, so the unit is microseconds and the formula divides by 1000 for milliseconds.nodenamefromnode_uname_infoinstead of theinternal-hostnametag, which is not present on any/v2metric.rdse.node_egress_bytes_medianandrdse.node_ingress_bytes_medianare V1-only with no V2 equivalent.title, folding it into the content as a markdown heading. The Dashboards API rejectstitleonnotewidgets, which made every one of these dashboard assets fail to import.Motivation
Datadog raised a list of
redis_enterprise_prometheusdashboard widgets showing no data. Each item was verified against a live Redis Enterprise8.0.10-76/v2endpoint rather than against the repo fixture, which turned out to be stale in places — for exampleprocess_resident_memory_bytesandprocess_virtual_memory_bytesexist intests/data/metrics.txtbut are gone from RS 8.0.10 entirely.Two findings worth calling out:
internal-hostnametag genuinely does not exist on any/v2metric, butnode_uname_infoalready exposes the same information asnodename, so no upstream Redis change is needed.ddev validate dashboardspasses both before and after the note-widget fix, so it does not currently catch assets that the Dashboards API will reject on import.Review checklist
no-changeloglabel attachedAdditional Notes
Dashboard asset changes are not unit-testable, so no tests are added; each query was instead validated against a live cluster. Two caveats on that validation:
rdse2.redis_process_*, which has never existed under any name; the viable source isnamedprocess_namegroup_memory_bytes{memtype="resident"|"virtual"}. Process CPU has data but groupsby {db,redis}onnamedprocess_*series that carry neither tag (shards appear asgroupname="redis-<n>"). Fixing those changes the widgets' semantics from shard-scoped to process-group-scoped, which felt like a separate discussion.Replacement for #3136, which was automatically closed as stale.