design/telemetry.md describes a surface of 50 metrics. #470 composes the collector and the renames that produce 23 of them: the engine, GPU and gateway-latency core, everything reachable by renaming one metric to one name and converting its unit.
The other 27 each need something that doesn't exist yet. Grouped by what actually blocks them, because the groups are independent and not equally sized.
Per-object state, 11 metrics
Needs resource-state-metrics on the control plane reading Modelplane's own objects, plus kube-state-metrics for container restarts and gang status. None of these come from scraping a workload; they're the state of a ModelReplica, a ModelDeployment or an InferenceCluster.
modelplane_replica_gpus, modelplane_replica_gpu{gpu_uuid}, modelplane_replica_allocated_time_seconds
modelplane_replicas_desired, modelplane_replicas_ready, modelplane_replica_ready_duration_seconds
modelplane_replica_cache_hit, modelplane_cluster_gpus_allocatable, modelplane_cluster_connected
modelplane_gang_incomplete (LWS/Grove status), modelplane_engine_restarts_total (kube-state-metrics)
This is the largest group, and modelplane_gpu_seconds_total — the metric a cost question actually needs — is derived from modelplane_replica_allocated_time_seconds, so it's here too.
Folding several series under one label, 8 metrics
Needs MetricMapping to grow the two fields left out of #470: a labels form that folds two metrics into one distinguished by a fixed value, or renames a label and remaps its value vocabulary; and a part form that takes a histogram's count as a counter.
modelplane_requests_total{status}, modelplane_tokens_total{direction} — fleet gateway
modelplane_responses_total{reason}, modelplane_tool_call_parses_total{outcome} — engine
modelplane_route_requests_total{decision}, modelplane_route_pd_pairings_total{status} — picker
modelplane_gpu_ecc_errors_total{type}, modelplane_gpu_interconnect_errors_total{link} — DCGM, and also gated on fields that are off by default
The API shape is sketched in #470 (comment). It was left out of #470 because six of these eight depend on source metric names at the gateway, the picker and DCGM that nobody has read off a running deployment yet, and the seventh and eighth need DCGM reconfigured. modelplane_responses_total{reason} is the one implementable today, from vllm:request_success_total{finished_reason}.
First step: capture the gateway's and the picker's actual metric names from a live cluster. The API follows from those, not the other way round.
Merging replicas twice, 2 metrics
A saturation gauge needs a mean and a maximum, because capacity planning reads the mean and an alert fires on the hot replica. The collector merges replicas once today.
modelplane_kv_cache_utilization_ratio_max, modelplane_requests_waiting_max
A pipeline change in compose-serving-stack, not an API change.
Instrumentation that doesn't emit yet, 6 metrics
modelplane_replica_staging_seconds, modelplane_replica_warmup_seconds — ModelExpress
modelplane_requests_throttled_total — Envoy rate limits
modelplane_dra_allocation_errors_total — DRA driver
modelplane_gpu_thermal_throttle_seconds_total — GPU exporter
modelplane_request_kv_transfer_seconds — engine, disaggregated only
Order
Per-object state first: it's eleven metrics, it's the only group blocking cost questions, and it needs no live measurement to design. The label folding wants a cluster to read names off first. The double merge is small and can go any time.
design/telemetry.mddescribes a surface of 50 metrics. #470 composes the collector and the renames that produce 23 of them: the engine, GPU and gateway-latency core, everything reachable by renaming one metric to one name and converting its unit.The other 27 each need something that doesn't exist yet. Grouped by what actually blocks them, because the groups are independent and not equally sized.
Per-object state, 11 metrics
Needs
resource-state-metricson the control plane reading Modelplane's own objects, plus kube-state-metrics for container restarts and gang status. None of these come from scraping a workload; they're the state of aModelReplica, aModelDeploymentor anInferenceCluster.modelplane_replica_gpus,modelplane_replica_gpu{gpu_uuid},modelplane_replica_allocated_time_secondsmodelplane_replicas_desired,modelplane_replicas_ready,modelplane_replica_ready_duration_secondsmodelplane_replica_cache_hit,modelplane_cluster_gpus_allocatable,modelplane_cluster_connectedmodelplane_gang_incomplete(LWS/Grove status),modelplane_engine_restarts_total(kube-state-metrics)This is the largest group, and
modelplane_gpu_seconds_total— the metric a cost question actually needs — is derived frommodelplane_replica_allocated_time_seconds, so it's here too.Folding several series under one label, 8 metrics
Needs
MetricMappingto grow the two fields left out of #470: alabelsform that folds two metrics into one distinguished by a fixed value, or renames a label and remaps its value vocabulary; and apartform that takes a histogram's count as a counter.modelplane_requests_total{status},modelplane_tokens_total{direction}— fleet gatewaymodelplane_responses_total{reason},modelplane_tool_call_parses_total{outcome}— enginemodelplane_route_requests_total{decision},modelplane_route_pd_pairings_total{status}— pickermodelplane_gpu_ecc_errors_total{type},modelplane_gpu_interconnect_errors_total{link}— DCGM, and also gated on fields that are off by defaultThe API shape is sketched in #470 (comment). It was left out of #470 because six of these eight depend on source metric names at the gateway, the picker and DCGM that nobody has read off a running deployment yet, and the seventh and eighth need DCGM reconfigured.
modelplane_responses_total{reason}is the one implementable today, fromvllm:request_success_total{finished_reason}.First step: capture the gateway's and the picker's actual metric names from a live cluster. The API follows from those, not the other way round.
Merging replicas twice, 2 metrics
A saturation gauge needs a mean and a maximum, because capacity planning reads the mean and an alert fires on the hot replica. The collector merges replicas once today.
modelplane_kv_cache_utilization_ratio_max,modelplane_requests_waiting_maxA pipeline change in
compose-serving-stack, not an API change.Instrumentation that doesn't emit yet, 6 metrics
modelplane_replica_staging_seconds,modelplane_replica_warmup_seconds— ModelExpressmodelplane_requests_throttled_total— Envoy rate limitsmodelplane_dra_allocation_errors_total— DRA drivermodelplane_gpu_thermal_throttle_seconds_total— GPU exportermodelplane_request_kv_transfer_seconds— engine, disaggregated onlyOrder
Per-object state first: it's eleven metrics, it's the only group blocking cost questions, and it needs no live measurement to design. The label folding wants a cluster to read names off first. The double merge is small and can go any time.