fix(dashboard): read GPU memory from per-card series, not per-container - #615
Merged
Conversation
The cluster GPU Memory card and the node table's GPU columns were built on HAMi's per-container series (hami_container_device_memory_bytes, hami_vgpu_memory_used_bytes, hami_vgpu_memory_limit_bytes). HAMi-core only emits those from inside an actively running vGPU container, so they vanish on an idle cluster — even while the scheduler holds an allocation for a pod. In PromQL `sum(<empty>) / sum(x)` is an empty vector, so both panels rendered "no data" while Prometheus reported success. Point every GPU memory reading at per-card series instead: hami_gpu_memory_limit_bytes for capacity, hami_gpu_memory_allocated_bytes for request, and the device plugin's hami_host_gpu_memory_used_bytes for usage. Pod attribution in the node table moves to hami_vgpu_memory_allocated_bytes, which is scheduler-side and still carries exported_namespace/exported_pod, so the per-card consumer breakdown survives. Also harden the card: gate "no data" on capacity rather than usage, and wrap each numerator in `or vector(0)` so an absent series reads 0% instead of collapsing the panel.
ZhangEnYao
approved these changes
Aug 6, 2026
KUASWoodyLIN
deleted the
614-cluster-gpu-memory-card-and-node-table-gpu-columns-show-no-data-while-gpus-are-allocated
branch
August 6, 2026 13:55
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The cluster GPU Memory card and the node table's GPU columns were built on HAMi's per-container series (hami_container_device_memory_bytes, hami_vgpu_memory_used_bytes, hami_vgpu_memory_limit_bytes). HAMi-core only emits those from inside an actively running vGPU container, so they vanish on an idle cluster — even while the scheduler holds an allocation for a pod. In PromQL
sum(<empty>) / sum(x)is an empty vector, so both panels rendered "no data" while Prometheus reported success.Point every GPU memory reading at per-card series instead: hami_gpu_memory_limit_bytes for capacity, hami_gpu_memory_allocated_bytes for request, and the device plugin's hami_host_gpu_memory_used_bytes for usage. Pod attribution in the node table moves to hami_vgpu_memory_allocated_bytes, which is scheduler-side and still carries exported_namespace/exported_pod, so the per-card consumer breakdown survives.
Also harden the card: gate "no data" on capacity rather than usage, and wrap each numerator in
or vector(0)so an absent series reads 0% instead of collapsing the panel.