Skip to content

Complete the metric surface described in the telemetry design #476

Description

@dennis-upbound

design/telemetry.md describes a surface of 50 metrics. #470 composes the collector and the renames that produce 23 of them: the engine, GPU and gateway-latency core, everything reachable by renaming one metric to one name and converting its unit.

The other 27 each need something that doesn't exist yet. Grouped by what actually blocks them, because the groups are independent and not equally sized.

Per-object state, 11 metrics

Needs resource-state-metrics on the control plane reading Modelplane's own objects, plus kube-state-metrics for container restarts and gang status. None of these come from scraping a workload; they're the state of a ModelReplica, a ModelDeployment or an InferenceCluster.

  • modelplane_replica_gpus, modelplane_replica_gpu{gpu_uuid}, modelplane_replica_allocated_time_seconds
  • modelplane_replicas_desired, modelplane_replicas_ready, modelplane_replica_ready_duration_seconds
  • modelplane_replica_cache_hit, modelplane_cluster_gpus_allocatable, modelplane_cluster_connected
  • modelplane_gang_incomplete (LWS/Grove status), modelplane_engine_restarts_total (kube-state-metrics)

This is the largest group, and modelplane_gpu_seconds_total — the metric a cost question actually needs — is derived from modelplane_replica_allocated_time_seconds, so it's here too.

Folding several series under one label, 8 metrics

Needs MetricMapping to grow the two fields left out of #470: a labels form that folds two metrics into one distinguished by a fixed value, or renames a label and remaps its value vocabulary; and a part form that takes a histogram's count as a counter.

  • modelplane_requests_total{status}, modelplane_tokens_total{direction} — fleet gateway
  • modelplane_responses_total{reason}, modelplane_tool_call_parses_total{outcome} — engine
  • modelplane_route_requests_total{decision}, modelplane_route_pd_pairings_total{status} — picker
  • modelplane_gpu_ecc_errors_total{type}, modelplane_gpu_interconnect_errors_total{link} — DCGM, and also gated on fields that are off by default

The API shape is sketched in #470 (comment). It was left out of #470 because six of these eight depend on source metric names at the gateway, the picker and DCGM that nobody has read off a running deployment yet, and the seventh and eighth need DCGM reconfigured. modelplane_responses_total{reason} is the one implementable today, from vllm:request_success_total{finished_reason}.

First step: capture the gateway's and the picker's actual metric names from a live cluster. The API follows from those, not the other way round.

Merging replicas twice, 2 metrics

A saturation gauge needs a mean and a maximum, because capacity planning reads the mean and an alert fires on the hot replica. The collector merges replicas once today.

  • modelplane_kv_cache_utilization_ratio_max, modelplane_requests_waiting_max

A pipeline change in compose-serving-stack, not an API change.

Instrumentation that doesn't emit yet, 6 metrics

  • modelplane_replica_staging_seconds, modelplane_replica_warmup_seconds — ModelExpress
  • modelplane_requests_throttled_total — Envoy rate limits
  • modelplane_dra_allocation_errors_total — DRA driver
  • modelplane_gpu_thermal_throttle_seconds_total — GPU exporter
  • modelplane_request_kv_transfer_seconds — engine, disaggregated only

Order

Per-object state first: it's eleven metrics, it's the only group blocking cost questions, and it needs no live measurement to design. The label folding wants a cluster to read names off first. The double merge is small and can go any time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions