#470 composes a collector on every inference cluster but deliberately keeps it out of the ServingStack's readiness: nothing serving depends on telemetry, and gating on it meant one unreachable destination could stop the scheduler placing replicas fleet-wide.
That leaves the other half unbuilt. Nothing reports whether the fleet is actually exporting, so a collector that cannot start is invisible unless you go looking at the workload cluster.
The TelemetryDestination is where that belongs — it is the object the operator wrote, and it is cluster-scoped, so it can speak for the whole fleet. Something like:
status:
conditions:
- type: Exporting
status: "False"
reason: CollectorsNotReady
message: "Collector not ready on 2 of 5 clusters: gke-prod (0/1 replicas ready), nebius-1 (0/1 replicas ready)"
Reading the composed collector Deployments' readiness is enough for a first cut. Whether telemetry is actually flowing is a further question — the healthcheckv2 extension covers it but is still marked development, so it can wait.
Raised by @lsviben on #470 (comment).
#470 composes a collector on every inference cluster but deliberately keeps it out of the
ServingStack's readiness: nothing serving depends on telemetry, and gating on it meant one unreachable destination could stop the scheduler placing replicas fleet-wide.That leaves the other half unbuilt. Nothing reports whether the fleet is actually exporting, so a collector that cannot start is invisible unless you go looking at the workload cluster.
The
TelemetryDestinationis where that belongs — it is the object the operator wrote, and it is cluster-scoped, so it can speak for the whole fleet. Something like:Reading the composed collector Deployments' readiness is enough for a first cut. Whether telemetry is actually flowing is a further question — the healthcheckv2 extension covers it but is still marked development, so it can wait.
Raised by @lsviben on #470 (comment).