Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 65 additions & 0 deletions .github/workflows/e2e.yml
Original file line number Diff line number Diff line change
Expand Up @@ -97,3 +97,68 @@ jobs:
- name: Teardown
if: always()
run: nix run .#e2e -- --clean

# The Provided-mode variant: the substrate is pre-installed on the
# workload cluster and Modelplane only checks it (see e2e/run.sh
# --provided). Its own label, because it roughly doubles the e2e
# minutes and the RequirementsMet flow only needs exercising on
# changes near the serving stack.
e2e-provided:
if: >-
github.event_name == 'workflow_dispatch'
|| (github.event.action == 'labeled' && github.event.label.name == 'test-e2e-provided')
|| (github.event.action != 'labeled' && contains(github.event.pull_request.labels.*.name, 'test-e2e-provided'))
runs-on: ubuntu-24.04
timeout-minutes: 60

steps:
- name: Cleanup Disk
uses: jlumbroso/free-disk-space@54081f138730dfa15788a46383842cd2f914a1be # v1.3.1
with:
android: true
dotnet: true
haskell: true
tool-cache: true
swap-storage: false
large-packages: false
docker-images: false

- name: Checkout
uses: actions/checkout@v4

- name: Install Nix
uses: cachix/install-nix-action@v31
with:
extra_nix_config: log-lines = 500

- name: Cache Nix store
uses: DeterminateSystems/magic-nix-cache-action@908b263ff629f4cc17666315b7fd3ec127c6244d # v14
with:
use-flakehub: false

- name: Run e2e (provided)
run: nix run .#e2e --print-build-logs -- --provided --verify

- name: Collect diagnostics
if: failure()
run: |
mkdir -p e2e-logs
nix develop --command bash -c '
kind export logs e2e-logs/control-plane --name modelplane-e2e-local || true
kind export logs e2e-logs/workload --name modelplane-e2e-workload || true
kubectl --context kind-modelplane-e2e-local get composite -o wide > e2e-logs/composites.txt 2>&1 || true
kubectl --context kind-modelplane-e2e-local get inferencecluster local -o yaml > e2e-logs/inferencecluster.yaml 2>&1 || true
kubectl --context kind-modelplane-e2e-local -n ml-team get modeldeployment,modelreplica,modelendpoint,modelservice -o yaml > e2e-logs/model.yaml 2>&1 || true
'

- name: Upload diagnostics
if: failure()
uses: actions/upload-artifact@v4
with:
name: e2e-provided-diagnostics
path: e2e-logs
if-no-files-found: ignore

- name: Teardown
if: always()
run: nix run .#e2e -- --clean
26 changes: 26 additions & 0 deletions apis/inferenceclusters/definition.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,32 @@ spec:
ReadWriteMany dynamic provisioning).
minLength: 1
maxLength: 253
components:
type: string
default: Managed
x-kubernetes-validations:
- rule: self == oldSelf
message: >-
spec.cluster.existing.components is immutable;
recreate the cluster to change who supplies the
serving substrate.
description: >-
Who supplies the serving substrate on this cluster.
Managed (the default) has Modelplane install every
serving stack component. Provided installs no
substrate: the cluster already runs cert-manager,
the gateway stack, Prometheus, the GPU DRA driver
and the stack's workload controller, and Modelplane
verifies they are present (the RequirementsMet
condition reports what's missing) and composes only
its own configuration on top. All or nothing; the
cluster provides the whole substrate or none of it.
Immutable because flipping it would uninstall a
live cluster's substrate or adopt one Modelplane
doesn't own.
enum:
- Managed
- Provided
gke:
type: object
description: >-
Expand Down
18 changes: 18 additions & 0 deletions apis/servingstacks/definition.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,25 @@ spec:
x-kubernetes-validations:
- rule: "self.secrets.exists(s, s.type == 'Kubeconfig')"
message: spec.secrets must include a Kubeconfig entry.
- rule: "!has(self.components) || self.components != 'Provided' || self.cloud == 'Existing'"
message: spec.components Provided is only supported when cloud is Existing.
properties:
components:
type: string
default: Managed
description: >-
Who supplies the serving substrate. Managed (the default)
has Modelplane install every component. Provided, valid
only on an Existing cluster, installs no substrate charts
or vendored CRDs: Modelplane verifies the cluster supplies
them - reported through the RequirementsMet condition - and
composes only its own configuration on top. All or
nothing. Mirrors
InferenceCluster.spec.cluster.existing.components; the
cluster composition sets it.
enum:
- Managed
- Provided
cloud:
type: string
description: >-
Expand Down
12 changes: 12 additions & 0 deletions docs/content/platform/inference-cluster.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,11 @@ The `cluster.source` discriminator picks one of two models:
provision infrastructure, and each pool's `InferenceClass` provides hardware
capabilities for scheduling only. You're responsible for the cluster meeting
[Modelplane's requirements](#requirements-for-an-existing-cluster).
If the cluster already runs the serving substrate - cert-manager, a gateway
stack, Prometheus - set `existing.components: Provided` and Modelplane
installs none of it, verifying the cluster meets the
[serving stack requirements]({{< ref "/platform/serving-stack-requirements.md" >}})
instead and reporting what's missing through the `RequirementsMet` condition.

## Requirements for an existing cluster

Expand Down Expand Up @@ -77,12 +82,19 @@ An existing cluster must meet what Modelplane would otherwise set up for you:
address.
- **No conflicting Gateway controller.** Modelplane installs Envoy Gateway and
owns its `GatewayClass`. Don't run another controller claiming the same class.
With `components: Provided` this inverts: the cluster runs its own Envoy
Gateway, which serves the `GatewayClass` Modelplane composes.
- **A `ReadWriteMany` StorageClass**, if you use a `ModelCache`. See
[Cache storage](#cache-storage).
- **Any multi-node fabric you need.** For multi-node serving you provide and
configure the RDMA or InfiniBand fabric and its drivers. Modelplane installs
those only on the clouds it provisions.

With `components: Provided` the cluster provides the serving stack itself on
top of all this; the
[serving stack requirements]({{< ref "/platform/serving-stack-requirements.md" >}})
page lists what that adds, per component.

## Serving stack

`spec.stack` selects the serving layer the cluster runs: `Standard` (the default)
Expand Down
238 changes: 238 additions & 0 deletions docs/content/platform/serving-stack-requirements.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,238 @@
---
title: Serving Stack Requirements
weight: 35
description: What an existing cluster must provide to run the serving stack itself.
---
<!-- Generated by functions/compose-serving-stack/requirements_doc.py
from the serving stack component lists. Do not edit; run
`nix run .#requirements-doc` after changing the stack data. -->
An `InferenceCluster` with `spec.cluster.existing.components: Provided`
installs no serving stack components. The cluster provides the whole
substrate itself, and Modelplane only verifies it and composes its own
configuration on top. This page lists what that cluster must provide,
per component Modelplane would otherwise install. It applies on top of
the general [requirements for an existing
cluster]({{< ref "/platform/inference-cluster.md" >}}).

Modelplane checks the **checked** entries continuously and reports what
is missing through the `RequirementsMet` condition on the
`InferenceCluster`. A check verifies presence and served API versions,
not the installed release: the **not checked** entries, including
component versions, are yours to meet. The version listed per component
is the one Modelplane installs in Managed mode and tests against; stay
close to it.

## On every stack

### `cert-manager`

Managed mode installs chart `cert-manager` `v1.20.2`.

Checked:

- CRD `certificates.cert-manager.io` serving `v1`

Not checked:

- The cert-manager controller and webhook are running and issue Certificates.

### `kube-prometheus-stack`

Managed mode installs chart `kube-prometheus-stack` `84.4.0`.

Checked:

- CRD `podmonitors.monitoring.coreos.com` serving `v1`
- CRD `servicemonitors.monitoring.coreos.com` serving `v1`

Not checked:

- Prometheus discovers `PodMonitor` objects in every namespace: with the chart, set `podMonitorSelectorNilUsesHelmValues` to false and `podMonitorNamespaceSelector` to empty. Modelplane's scrape targets are `PodMonitor` objects in workload namespaces, and the chart's default release label selector never matches them.
- A scrape job for the Envoy Gateway proxy pods' stats endpoint, if you want request metrics at the proxy level.

Not needed:

- Modelplane disables Grafana and Alertmanager.

### `node-feature-discovery`

Managed mode installs chart `node-feature-discovery` `0.19.0`.

Checked:

- CRD `nodefeatures.nfd.k8s-sigs.io` serving `v1alpha1`

Not checked:

- The NFD worker runs on the GPU nodes and labels them with `feature.node.kubernetes.io/pci-10de` and friends. If GPU nodes are tainted, the worker must tolerate the taint, or the DRA driver's `kubelet` plugin never schedules there and every GPU ResourceClaim stays pending with all components looking healthy.

### `nvidia-dra-driver-gpu`

Managed mode installs chart `dra-driver-nvidia-gpu` `0.4.1`.

Checked:

- `DeviceClass` `gpu.nvidia.com` (`resource.k8s.io/v1`)

Not checked:

- The NVIDIA kernel driver and Container Toolkit on every GPU node, from the node image or the GPU Operator. The requirements for an existing cluster state the versions.
- The driver's `kubelet` plugin publishes each GPU node's devices as ResourceSlices.
- No device plugin advertising `nvidia.com/gpu`: a second allocator would hand out the same GPUs behind DRA's back.

Not needed:

- Modelplane disables `ComputeDomains` (multi-node NVLink) and their prerequisites.

### `envoy-gateway`

Managed mode installs chart `gateway-helm` `v1.8.4`.

Checked:

- CRD `gatewayclasses.gateway.networking.k8s.io` serving `v1`
- CRD `gateways.gateway.networking.k8s.io` serving `v1`
- CRD `httproutes.gateway.networking.k8s.io` serving `v1`
- CRD `envoyproxies.gateway.envoyproxy.io` serving `v1alpha1`
- CRD `backends.gateway.envoyproxy.io` serving `v1alpha1`
- CRD `clienttrafficpolicies.gateway.envoyproxy.io` serving `v1alpha1`

Not checked:

- The Envoy Gateway controller runs with the `extensionManager` wired exactly as the values above: external processing delegated to the AI Gateway controller's Service, with the Backend API enabled and InferencePool declared a backend resource. Without it, HTTPRoute to InferencePool `backendRefs` never route, with every component looking healthy.

The exact values Modelplane installs the chart with. The `extensionManager` wiring is the one coupling no check can verify. Without it, routes to an `InferencePool` never route while every component looks healthy:

```yaml
config:
envoyGateway:
extensionApis:
enableBackend: true
extensionManager:
hooks:
xdsTranslator:
translation:
listener:
includeAll: true
route:
includeAll: true
cluster:
includeAll: true
secret:
includeAll: true
post:
- Translation
- Cluster
- Route
service:
fqdn:
hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local
port: 1063
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
```

### `ai-gateway-crds`

Managed mode installs chart `ai-gateway-crds-helm` `v1.1.0`.

Checked:

- CRD `aigatewayroutes.aigateway.envoyproxy.io` serving `v1alpha1`

Not checked:

- The AI Gateway APIs are v1alpha1 and move with the controller. A provided install tracks the pinned v1.1.0 release.

### `ai-gateway`

Managed mode installs chart `ai-gateway-helm` `v1.1.0`.

Not checked:

- The AI Gateway controller is reachable at `ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local:1063`, the address Modelplane's Envoy Gateway `extensionManager` values point at. A controller installed elsewhere never receives the external processing traffic.

### `gaie-crds`

Managed mode applies manifests Modelplane vendors from the upstream release.

Checked:

- CRD `inferencepools.inference.networking.k8s.io` serving `v1`

### `trust-manager`

Managed mode installs chart `trust-manager` `v0.25.0`.

Checked:

- CRD `bundles.trust.cert-manager.io` serving `v1alpha1`

Not checked:

- trust-manager watches modelplane-system as its trust namespace, where the gateway PKI composes its Bundle. An install watching another namespace never syncs the Bundle, and the control plane can't read the cluster CA.

### `dra-driver-critical-pods-quota`

Managed mode applies manifests Modelplane vendors from the upstream release.

Not checked:

- On clusters that restrict `system-node-critical` pods by namespace quota (GKE does), the DRA driver's namespace needs a `ResourceQuota` admitting them, or the `kubelet` plugin `DaemonSet` never starts.

## Standard stack

### `leader-worker-set`

Managed mode installs chart `lws` `v0.8.0`.

Checked:

- CRD `leaderworkersets.leaderworkerset.x-k8s.io` serving `v1`

Not checked:

- The LeaderWorkerSet controller is running. Modelplane composes against the v0.8 line.

## Dynamo stack

### `grove`

Managed mode installs chart `grove-charts` `v0.1.0-alpha.12-rc2`.

Checked:

- CRD `podcliquesets.grove.io` serving `v1alpha1`

Not checked:

- Grove's API is v1alpha1 with no compatibility promise between versions, so a provided install must run the exact pinned version. A CRD check can't tell alpha revisions apart. Prefer Managed for the Dynamo stack until Grove stabilizes.

### `kai-scheduler`

Managed mode installs chart `kai-scheduler` `v0.16.8`.

Checked:

- CRD `queues.scheduling.run.ai` serving `v2`

Not checked:

- The scheduler answers to the `schedulerName` value `kai-scheduler`, the name Modelplane's engine pods request. Its Queue webhook must be serving. Modelplane composes against the v0.16 line.

## What Modelplane still installs

Provided mode only skips the substrate. Modelplane still composes its
own configuration and workloads: the `modelplane-system` namespace, the
`EnvoyProxy`, `GatewayClass` and `Gateway` for the inference gateway,
and on the Dynamo stack its KAI `Queue` hierarchy and the ModelExpress
server with its CRDs. Their substrate dependencies gate on the checks
above, so none of them is applied before the cluster serves the APIs
they need.

During deletion, Modelplane can't order its configuration ahead of a
substrate it doesn't own. Keep your controllers (the gateway
controller, KAI) running while an `InferenceCluster` deletes, so they
can process finalizers on Modelplane's configuration.
Loading
Loading