From e848dcb15a28f0da63a53709465762d9ff3e5a8e Mon Sep 17 00:00:00 2001 From: Nic Cope Date: Tue, 6 Oct 2026 15:21:39 -0700 Subject: [PATCH 1/2] Bump vale-ai-tells to v1.37.0 v1.19.0 has 60 rules and v1.37.0 has 137. The new rules raise 172 errors in docs that pass the docs-vale check today, and this commit leaves them unfixed. Unlike v1.19.0, v1.37.0 ships a config/ directory. Copied out of the store read-only, it blocked the lint step from merging the repo's vocabularies into config/. Signed-off-by: Nic Cope --- docs/utils/vale/.vale.ini | 2 +- nix/docs.nix | 6 ++++-- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/docs/utils/vale/.vale.ini b/docs/utils/vale/.vale.ini index d10d4ab01..23f79887b 100644 --- a/docs/utils/vale/.vale.ini +++ b/docs/utils/vale/.vale.ini @@ -1,6 +1,6 @@ StylesPath = styles MinAlertLevel = warning -Packages = Google, Microsoft, write-good, alex, proselint, https://github.com/tbhb/vale-ai-tells/releases/download/v1.19.0/ai-tells.zip +Packages = Google, Microsoft, write-good, alex, proselint, https://github.com/tbhb/vale-ai-tells/releases/download/v1.37.0/ai-tells.zip Vocab = Modelplane diff --git a/nix/docs.nix b/nix/docs.nix index daf2412e9..3b3f0bf55 100644 --- a/nix/docs.nix +++ b/nix/docs.nix @@ -23,7 +23,7 @@ let ]; outputHashMode = "recursive"; outputHashAlgo = "sha256"; - outputHash = "sha256-WbPAE0+fnV7gqU7P4c9qKWETRHTR7uP5OpsHCj+l9r4="; + outputHash = "sha256-2eRYliSjeVajgP5xCpXjyc62DgZ9PBLfXqBCdOdL/SQ="; } '' export HOME=$TMPDIR @@ -63,7 +63,9 @@ in # picks up both the synced packages and the repo's local styles. cp ${self}/docs/utils/vale/.vale.ini .vale.ini mkdir styles - cp -r ${valeStyles}/* styles/ + # The synced packages come out of the store read-only, and ai-tells + # ships a config/ directory that the repo's vocabularies merge into. + cp -r --no-preserve=mode ${valeStyles}/* styles/ cp -r ${self}/docs/utils/vale/styles/* styles/ find ${self}/docs/content -name '*.md' -print0 | \ xargs -0 --no-run-if-empty \ From 565b345696b9814e72660fe86f579e83acb3288a Mon Sep 17 00:00:00 2001 From: Nic Cope Date: Tue, 6 Oct 2026 16:38:16 -0700 Subject: [PATCH 2/2] Rewrite docs prose flagged by vale-ai-tells v1.37.0 The v1.37.0 rules flag 172 sentences across the docs. This rewrites each one rather than suppressing the rule, so docs-vale passes without new inline exceptions. A few flagged sentences were also inaccurate, such as a CEL selector described as a single line and a scheduler rule limited to healthy replicas, and the rewrites correct them. Three headings change, and so do their anchors: "How a service reaches its gateways", "How a deployment is composed", and "What the control plane reconciles". Nothing in this repo links to them. Signed-off-by: Nic Cope --- docs/content/architecture/_index.md | 28 +++---- docs/content/architecture/scheduling.md | 51 +++++++------ docs/content/getting-started/_index.md | 6 +- .../getting-started/build-the-platform.md | 10 +-- docs/content/getting-started/clean-up.md | 16 ++-- .../getting-started/deploying-a-model.md | 3 +- .../getting-started/scale-the-model.md | 10 +-- .../getting-started/scale-the-platform.md | 2 +- docs/content/guides/anthropic-messages-api.md | 2 +- .../guides/serving-multi-node-on-dynamo.md | 8 +- docs/content/install/_index.md | 4 +- docs/content/models/model-cache.md | 30 ++++---- docs/content/models/model-deployment.md | 73 ++++++++++--------- docs/content/models/model-endpoint.md | 2 +- docs/content/models/model-service.md | 57 +++++++-------- docs/content/overview/_index.md | 2 +- docs/content/overview/ai-tools.md | 13 ++-- docs/content/overview/faq.md | 10 +-- docs/content/overview/glossary.md | 8 +- docs/content/overview/how-it-works.md | 26 +++---- docs/content/overview/why.md | 8 +- docs/content/platform/drain-cluster.md | 16 ++-- docs/content/platform/inference-class.md | 10 +-- docs/content/platform/inference-cluster.md | 8 +- docs/content/platform/inference-gateway.md | 19 +++-- docs/content/platform/providers.md | 4 +- docs/content/platform/telemetry.md | 51 +++++++------ docs/content/recipes/glm-4.5-air.md | 14 ++-- docs/content/recipes/kimi-k2.md | 4 +- docs/content/recipes/laguna.md | 15 ++-- docs/content/recipes/llama-3.1-8b.md | 14 ++-- .../content/recipes/nemotron-3.5-lightning.md | 10 +-- docs/content/recipes/qwen2.5-72b.md | 26 +++---- docs/content/recipes/qwen2.5-7b.md | 15 ++-- docs/content/recipes/qwen3-8b.md | 16 ++-- docs/content/recipes/qwen3-coder.md | 14 ++-- 36 files changed, 302 insertions(+), 303 deletions(-) diff --git a/docs/content/architecture/_index.md b/docs/content/architecture/_index.md index 41deff015..870c1041b 100644 --- a/docs/content/architecture/_index.md +++ b/docs/content/architecture/_index.md @@ -26,8 +26,8 @@ matter here: - **Composition functions** are that controller logic. A function is a small gRPC service handed the observed XR and the resources it depends on, which returns the desired child resources. An XR runs a pipeline of one or more functions - every reconcile; in Modelplane each is typically a single function, so the rest - of this section says "the function" for short. + every reconcile; in Modelplane each pipeline typically has one function, so + the rest of this section says "the function" for short. - **Providers** are controllers that manage external systems through their own managed resources: `provider-gcp` and `provider-aws` for cloud APIs, `provider-helm` for Helm releases, `provider-kubernetes` for arbitrary objects @@ -43,7 +43,7 @@ The resource model mirrors Kubernetes core, one scope up: rather than within one. A `ModelDeployment` composes a `ModelReplica` per replica, a `ModelReplica` composes the serving workload on its target cluster, and a `ModelService` routes across the `ModelEndpoint`s. If you know how those core -objects relate, you already know the shape of Modelplane's. +objects relate, you already know how Modelplane's fit together. ## Why Crossplane? @@ -54,8 +54,8 @@ ways: providers and functions. **Providers** give us reach. Modelplane has to provision Kubernetes clusters and all the infrastructure they need across different clouds, then install software -onto them. That's an enormous surface, and providers cover it without us rolling -our own controllers for each cloud API and Helm release. +onto them. That spans many cloud APIs and Helm releases, and providers cover +them without us rolling our own controller for each. **Functions** are where Modelplane's own logic lives, and writing it as composition functions buys several things: @@ -74,7 +74,7 @@ composition functions buys several things: for contributors. The performance-sensitive distributed-systems core stays in Go, where Crossplane and its providers already are. -The bet underneath both is that inference infrastructure is the same shape of +The bet underneath both is that inference infrastructure is the same kind of problem as cloud infrastructure, which Crossplane manages well. Building on it lets Modelplane spend its effort on the part that's actually inference-specific. @@ -82,18 +82,18 @@ lets Modelplane spend its effort on the part that's actually inference-specific. Modelplane runs on a **control cluster** and manages a fleet of **workload clusters**, the `InferenceCluster`s. The split is deliberate: the control plane -holds no GPUs and serves no tokens. It schedules and composes, and the +has no GPUs and doesn't serve tokens. It schedules and composes, and the workload clusters do the serving. The control cluster runs Crossplane, the Modelplane composition functions (one per resource, each a pod Crossplane calls per reconcile), and the providers. It -also holds every Modelplane resource and the `ProviderConfig`s that let the -providers reach each workload cluster, built from that cluster's kubeconfig. +also holds every Modelplane resource and the `ProviderConfig`s the providers use +to connect to each workload cluster, built from that cluster's kubeconfig. -Crossplane core drives everything. Each reconcile it asks a function what a -resource should compose and gets back the desired resources. Core then reconciles -them, applying the provider resources that the providers act on. A function only -computes desired state. It never reaches a provider or a cluster itself. +Crossplane core drives everything. Each reconcile it calls a resource's function +and gets back the desired resources. Core then reconciles them, applying the +provider resources that the providers act on. A function only computes desired +state. It never reaches a provider or a cluster itself. ```mermaid flowchart TB @@ -119,7 +119,7 @@ The exact components evolve, but Modelplane composes and owns all of them. For provisioned clusters the providers also create the cluster and its node pools first. -## How a deployment is composed +## Composing a deployment A resource composes others, which compose others, until the tree bottoms out in provider resources and plain Kubernetes objects. A `ModelDeployment` is the diff --git a/docs/content/architecture/scheduling.md b/docs/content/architecture/scheduling.md index 5ea50c07a..55ce26cbd 100644 --- a/docs/content/architecture/scheduling.md +++ b/docs/content/architecture/scheduling.md @@ -6,8 +6,8 @@ description: How Modelplane places a deployment's replicas across the fleet, and **API:** [`modelplane.ai/v1alpha1` · ModelDeployment]({{< ref "/reference/modeldeployments" >}}) When an ML team creates a [ModelDeployment]({{< ref "/models/model-deployment.md" >}}), -the fleet scheduler decides which cluster each replica runs on and which node -pool each engine uses. Platform teams don't drive it directly, but what they +the fleet scheduler picks the cluster each replica runs on and the node pool +each engine uses. Platform teams don't drive it directly, but what they publish, the clusters, their labels, and each pool's [InferenceClass]({{< ref "/platform/inference-class.md" >}}), is exactly what the scheduler matches against. This page explains how it places work and where it @@ -22,9 +22,9 @@ every existing `ModelReplica`, and returns a placement. Given the same inputs it returns the same placement, so it's safe to run continuously. The key consequence is stability. Existing replicas are *inputs*, not decisions. -A healthy replica is never moved to improve the global picture, even if a better -cluster appears later. This keeps placement from churning underneath a running -deployment. +The scheduler never moves a replica to improve the global picture, even if a +better cluster appears later. This keeps placement from churning underneath a +running deployment. ## Two-level matching @@ -39,8 +39,8 @@ against what the platform team published. `nodeSelector.devices` against the devices a pool's `InferenceClass` publishes. A request is a real DRA request: a `count` and CEL selectors over a device's attributes and capacity, such as "a GPU with at least 141Gi of memory." A pool - fits a member when it has devices satisfying every request, with `count` to - cover them. + fits a member when it satisfies each request, with enough matching devices to + cover its `count`. The CEL is the same expression an ML engineer would write in a DRA `ResourceClaim`, evaluated against the devices the `InferenceClass` declares. The @@ -50,9 +50,9 @@ pool only if the class publishes the attributes and capacity it asks for. ## Co-scheduling and pools A replica is a set of engines placed together on one cluster. Within a replica, -every member of a single engine is placed on **one** pool: each member carries -its own `nodeSelector`, but the scheduler requires a single pool that satisfies -them all. +every member of an engine is placed on **one** pool: each member has its own +`nodeSelector`, but the scheduler only places the engine on a pool that +satisfies them all. It works this way because a gang's members coordinate over their pool's interconnect fabric, and the scheduler can't reason about fabric. Pool identity @@ -77,13 +77,13 @@ graph TD R --> D ``` -A member with no `nodeSelector` claims no devices. It matches the engine's pool -at no node cost and rides along on the gang's nodes, packed there by the -cluster's own scheduler. +A member with no `nodeSelector` doesn't claim any devices. It matches the +engine's pool at no node cost, and the cluster's own scheduler packs its pods +onto the gang's nodes. ## Counting capacity in nodes -Capacity is gated on **nodes**, not on individual GPUs. The only number the +Capacity is counted in **nodes**, not individual GPUs. The only number the scheduler reads from a member is its node cost: ```text @@ -91,11 +91,10 @@ nodes = pods × copies pods = 1 for a Standalone or Leader, or worker.nodes for a Worker ``` -A member that resolves no `claim: DRA` device, because it carried no -`nodeSelector` or matched only synthetic devices, costs zero nodes. The scheduler -sums the cost of a replica's members and places the replica only where every -engine's pool has enough free nodes, tracking a running ledger so it never -overcommits a cluster. +A member that resolves no `claim: DRA` device, because it has no `nodeSelector` +or matches only synthetic devices, costs zero nodes. The scheduler sums the cost +of a replica's members and places the replica only where every engine's pool has +enough free nodes, tracking a running ledger so it never overcommits a cluster. This accounting is deliberately coarse. The control-plane scheduler answers "could this cluster plausibly host this replica," not "exactly which GPU does @@ -106,8 +105,8 @@ state. ## Pinning placement to a pool -The scheduler's pool choice is enforced, not advisory. Each scheduled pod carries -a Kubernetes `nodeSelector` on the `modelplane.ai/pool` node label, so it can only +The scheduler's pool choice is enforced, not advisory. Each scheduled pod has a +Kubernetes `nodeSelector` on the `modelplane.ai/pool` node label, so it can only land on the pool the scheduler chose. Without it, the cluster's scheduler could place a pod on any pool whose devices match its DRA claim, and the fleet's per-pool accounting would drift from where pods actually run. @@ -127,21 +126,21 @@ Scheduling runs in two phases each reconcile: `nodeSelector`. A degraded cluster, one that's not Ready or has no gateway address, is still retained; transient outages surface through the deployment's conditions, not re-placement. -- **Fill.** If the deployment wants more replicas than were retained, the +- **Fill.** If the deployment specifies more replicas than were retained, the shortfall is placed one at a time, each onto the eligible cluster hosting the - fewest of this deployment's replicas, spreading before packing. If it wants - fewer, the highest-index replicas are dropped first. + fewest of this deployment's replicas, spreading before packing. If it + specifies fewer, the highest-index replicas are dropped first. A replica never changes cluster. If its cluster is deleted, the replica stops -being emitted, Crossplane garbage-collects it, and the fill phase mints a fresh +being emitted, Crossplane garbage-collects it, and the fill phase creates a new replica elsewhere. Moving is always delete-plus-create, mirroring how Kubernetes treats a pod whose node is gone. ## Known limitations The scheduler is built to be conservative and predictable rather than optimal. -Two limits follow from that, both tracked for future work: +The limits below follow from that, and each is tracked for future work: - **A whole node is charged per pod** ([#172](https://github.com/modelplaneai/modelplane/issues/172)). A pod that diff --git a/docs/content/getting-started/_index.md b/docs/content/getting-started/_index.md index 7cf53f2a9..538061e32 100644 --- a/docs/content/getting-started/_index.md +++ b/docs/content/getting-started/_index.md @@ -11,10 +11,10 @@ against it. Without it, every change on one side creates work for the other. When the platform team updates infrastructure, ML teams have to react. When model requirements change, the platform team gets a request. -With Modelplane, the platform team publishes hardware without knowing what +With Modelplane, the platform team publishes its hardware without knowing what models will run on it. The ML team declares what a model needs without knowing -what clusters exist. The control plane resolves it and keeps it current as -both sides change. +what clusters exist. The control plane places the model on matching hardware +and moves it only when a change on either side breaks the match. In this tour, you'll switch between provisioning infrastructure and declaring a model to see how they interact. By the end you'll have a GPU fleet across three regions and one OpenAI-compatible endpoint routing to a model served across two of them. diff --git a/docs/content/getting-started/build-the-platform.md b/docs/content/getting-started/build-the-platform.md index 86aae966a..885244cd8 100644 --- a/docs/content/getting-started/build-the-platform.md +++ b/docs/content/getting-started/build-the-platform.md @@ -4,9 +4,9 @@ weight: 20 description: Set up the gateway, give the control plane cloud credentials, and provision your first GPU cluster. --- This is the platform team's side of Modelplane. You set up the gateway that -fronts your models, give the control plane cloud credentials, and register your -first GPU cluster: a hardware profile published as an `InferenceClass` and an -`InferenceCluster` that offers it. +fronts your models, give the control plane credentials for your cloud account, +and register your first GPU cluster: a hardware profile published as an +`InferenceClass` and an `InferenceCluster` that offers it. In the next step, the ML team will create a model deployment that schedules against this capacity without knowing which cluster it runs on. @@ -31,8 +31,8 @@ against this capacity without knowing which cluster it runs on. | `roles/iam.serviceAccountUser` | attaching that account to the nodes | | `roles/resourcemanager.projectIamAdmin` | granting the node account `container.admin` | - The last one is worth a look before you hand the key over. Modelplane grants - the node service account `roles/container.admin`, so the credential doing the + Review the last one before you hand the key over. Modelplane grants the node + service account `roles/container.admin`, so the credential doing the provisioning has to be able to set project IAM policy. {{< /tab >}} {{< tab "AKS" >}} diff --git a/docs/content/getting-started/clean-up.md b/docs/content/getting-started/clean-up.md index 095f6bf85..4bb03ac0c 100644 --- a/docs/content/getting-started/clean-up.md +++ b/docs/content/getting-started/clean-up.md @@ -9,9 +9,9 @@ plane. ## Delete model resources Delete model resources before clusters. A cluster refuses deletion while -anything still runs on it. Foreground cascading deletion holds each resource -until what it composed on the clusters is gone, so a cluster isn't released -while that's still being removed: +anything still runs on it. Foreground cascading deletion removes a resource +only after what it composed on the clusters is gone. That stops a cluster being +released while its workloads are still being torn down: ```bash kubectl delete md --all -n ml-team --cascade=foreground @@ -21,9 +21,9 @@ kubectl delete ms --all -n ml-team --cascade=foreground ## Delete the gateway Delete the gateway before its cluster. The `InferenceGateway` runs a load balancer -on the cluster it names; deleting it while that cluster is still up lets the load -balancer be removed, rather than leaking it when the cluster goes. Foreground -deletion holds the gateway until its objects on the cluster are deleted: +on the cluster it names. If that cluster is deleted first, the load balancer is +orphaned. Foreground deletion removes the gateway only after its objects on the +cluster are deleted: ```bash kubectl delete ig --all --cascade=foreground @@ -33,8 +33,8 @@ kubectl delete ig --all --cascade=foreground Delete all clusters with foreground cascading deletion. The serving stack on each workload cluster must uninstall while that cluster's API server is still -reachable. Foreground deletion holds each cluster object until its stack -finishes. Background deletion can orphan cloud resources. +reachable. Foreground deletion removes each cluster object only after its stack +finishes uninstalling. Background deletion can orphan cloud resources. ```bash kubectl delete ic --all --cascade=foreground diff --git a/docs/content/getting-started/deploying-a-model.md b/docs/content/getting-started/deploying-a-model.md index 729f52a21..1639bc4bb 100644 --- a/docs/content/getting-started/deploying-a-model.md +++ b/docs/content/getting-started/deploying-a-model.md @@ -120,7 +120,8 @@ You should get a response in a few seconds: ## Next step The platform team declared capacity and in this guide the ML team deployed a -model behind a stable endpoint. Neither team needed to know what the other was doing. Modelplane matched them. +model behind a stable endpoint. Each team worked without needing to know what +the other was doing. Modelplane matched them. In the next step, the platform team grows the fleet. [Scale the platform]({{< ref "getting-started/scale-the-platform.md" >}}) to add more clusters across regions. diff --git a/docs/content/getting-started/scale-the-model.md b/docs/content/getting-started/scale-the-model.md index c4ec69891..afd1cb340 100644 --- a/docs/content/getting-started/scale-the-model.md +++ b/docs/content/getting-started/scale-the-model.md @@ -1,10 +1,10 @@ --- title: Scale the model weight: 50 -description: Serve the model from two regions behind a single endpoint. +description: Serve the model from two regions behind one endpoint. --- A `ModelService` can front more than one `ModelDeployment`. Here you add a second -deployment, pinned to a different region, and point the same service at both. The +deployment in a different region and point the same service at both. The endpoint you already curled stays the same. Behind it, traffic now load-balances across two regions. @@ -72,7 +72,7 @@ Update the `ModelService` to select both deployments. Each entry in The model name doesn't change. Callers that had it before still have it; they don't know the fleet changed. The gateway load-balances across both regions, and if one region's replicas fail it sends every request to the other. The gateway -itself runs in one region, so surviving the loss of that region takes a second +itself runs in one region. Tolerating the loss of that region requires a second gateway on a cluster in the other. Send the same request as before: ```bash @@ -91,10 +91,10 @@ kubectl run -i --rm curl-test \ ## That's the tour -You stood up a control plane, built a multi-region GPU fleet, deployed a model +You installed a control plane, built a multi-region GPU fleet, deployed a model across it, and ended with one stable endpoint serving requests. The platform team published hardware. The ML team described what the model needs. Modelplane -placed them and served behind a single endpoint. +placed the model and served it behind that endpoint. [Clean up]({{< ref "getting-started/clean-up.md" >}}) tears everything down when you're done. diff --git a/docs/content/getting-started/scale-the-platform.md b/docs/content/getting-started/scale-the-platform.md index f8909981c..9178f3654 100644 --- a/docs/content/getting-started/scale-the-platform.md +++ b/docs/content/getting-started/scale-the-platform.md @@ -80,4 +80,4 @@ its deployment changes in a way that no longer fits where it runs. ## Next step -The fleet has grown with larger-GPU capacity. The ML team is next. [Scale the model]({{< ref "getting-started/scale-the-model.md" >}}) to serve it across the fleet behind a single endpoint. +The fleet has grown with larger-GPU capacity. The ML team is next. [Scale the model]({{< ref "getting-started/scale-the-model.md" >}}) to serve it across the fleet behind one endpoint. diff --git a/docs/content/guides/anthropic-messages-api.md b/docs/content/guides/anthropic-messages-api.md index 2892adc31..7e6588c19 100644 --- a/docs/content/guides/anthropic-messages-api.md +++ b/docs/content/guides/anthropic-messages-api.md @@ -11,7 +11,7 @@ client that speaks the Messages API, including Claude Code via `ANTHROPIC_BASE_URL`, names the service as the model in its request. See [Alternate APIs]({{< ref "/models/model-service.md" >}}) for the detail. -This recipe serves Qwen3-8B on a single NVIDIA H100 on Nebius, with tool calling +This recipe serves Qwen3-8B on one NVIDIA H100 on Nebius, with tool calling on: `--enable-auto-tool-choice` and `--tool-call-parser=hermes` are what let Claude Code's tool use work. An 8B model needs a fraction of an H100, so the GPU has ample headroom. Apply the platform side first, then the ML side. diff --git a/docs/content/guides/serving-multi-node-on-dynamo.md b/docs/content/guides/serving-multi-node-on-dynamo.md index 7abc0b5b7..a308f144f 100644 --- a/docs/content/guides/serving-multi-node-on-dynamo.md +++ b/docs/content/guides/serving-multi-node-on-dynamo.md @@ -48,8 +48,8 @@ kubectl create namespace ml-team The `Leader` and `Worker` run the same `vllm serve`, differing only in node rank. `$(MODELPLANE_LEADER_ADDRESS)` resolves to the leader on either stack, but -`$(MODELPLANE_RANK)` isn't injected on Dynamo yet, so the worker derives its rank -from Grove's `GROVE_PCLQ_POD_INDEX`. +`$(MODELPLANE_RANK)` isn't injected on Dynamo yet. The worker derives its rank +from Grove's `GROVE_PCLQ_POD_INDEX` instead. [Multi-node deployments]({{< ref "/models/model-deployment.md#multi-node" >}}) covers this. Both load with `--load-format modelexpress`, so they read the cached weights through the Dynamo stack's ModelExpress server. @@ -98,8 +98,8 @@ seeds its weights from the cache and publishes itself as a source; the second loads them straight from the first, peer-to-peer, rather than reading the cache again. -Each replica is a gang of two nodes, so a second replica needs two more nodes. -Grow the pool to four, then scale the deployment: +Each replica is a gang of two nodes. A second replica needs two more, so grow +the pool to four, then scale the deployment: ```bash kubectl patch ic/eks-us-east-dynamo --type=json -p '[ diff --git a/docs/content/install/_index.md b/docs/content/install/_index.md index 93be07b5c..4428ddbc1 100644 --- a/docs/content/install/_index.md +++ b/docs/content/install/_index.md @@ -2,7 +2,7 @@ title: Install weight: 8 navLanding: "Install the control plane" -description: Stand up the Modelplane control plane on a Kubernetes cluster you run. +description: Install the Modelplane control plane on a Kubernetes cluster you run. --- Modelplane's control plane is where everything runs: the Crossplane runtime, the providers it provisions infrastructure through, and the composition functions @@ -74,7 +74,7 @@ functions that reconcile them: {{< manifests "install/configuration.yaml" >}} -Wait until the configuration is healthy: +Wait for the configuration to become `Healthy`: ```bash kubectl wait configuration/modelplane --for=condition=Healthy --timeout=5m diff --git a/docs/content/models/model-cache.md b/docs/content/models/model-cache.md index 2f642d384..ce788c5f7 100644 --- a/docs/content/models/model-cache.md +++ b/docs/content/models/model-cache.md @@ -22,8 +22,8 @@ The required `source` enum names the kind, with the matching source object set alongside it. Setting `source: HuggingFace` selects `spec.huggingFace`, which carries the `repo` to fetch, an optional `revision` (branch, tag, or commit), and `sizeGiB`, how much storage the weights get on each cluster. Size it to the -model, since a value below the model's size leaves no room to stage the weights. -`HuggingFace` is the only source today. +model, since a value below the model's size doesn't leave room to stage the +weights. `HuggingFace` is the only source today. The engine's args name the model the same way with or without a cache. A `HuggingFace` source stages into HuggingFace's own cache layout on the mount, and @@ -35,13 +35,13 @@ itself: naming the model belongs to the engine command, like every other flag. Name the same `revision` the cache staged. A bare repository ID resolves at the default branch, which finds a cache staged without a `revision` or with `revision: main`. A cache pinned to a commit or tag needs the engine to pass that -revision too (`--revision` for vLLM). An engine that asks for the default branch -finds nothing staged under it, and downloads the model a second time. +revision too (`--revision` for vLLM). If the engine uses the default branch +instead, it finds nothing staged there and downloads the model a second time. ## Authenticating A gated or private model needs a credential to fetch. When a cache stages the -weights, the credential lives on the cache: set `authSecret` to name a Secret in +weights, the credential goes on the cache: set `authSecret` to name a Secret in the cache's namespace, and Modelplane propagates it to every cluster the cache stages to, for the hydration to read. @@ -72,7 +72,7 @@ container's `env`. An optional `clusterSelector` scopes where the cache is staged. Omitting it stages the cache on every cluster in the fleet; setting `matchLabels` restricts -it to clusters carrying those labels. Either way, a cluster with no +it to clusters with those labels. Either way, a cluster with no [cache storage](#storage-prerequisites) is skipped. A `ModelDeployment` that references the cache places replicas only on clusters the cache stages to, and moves a replica off a cluster the cache stops staging to, since the cache's @@ -80,11 +80,11 @@ volume is removed from it. ## Loading from cache -A cache only pays off if the engine reads from it quickly. With its default -loader an engine can read a large model from shared storage slowly enough that -the cache makes cold starts *worse* than fetching the model directly, since you -pay to hydrate the cache and then wait on a slow read. Choose a fast loader with -your engine flags. +A cache only shortens cold starts if the engine reads from it quickly. With its +default loader an engine can read a large model from shared storage slowly +enough that the cache makes cold starts *worse* than fetching the model +directly, since you pay to hydrate the cache and then wait on a slow read. +Choose a fast loader with your engine flags. For vLLM on EKS, `--load-format=runai_streamer` reads from the EFS-backed cache dramatically faster than the default loader (minutes rather than tens of @@ -112,10 +112,10 @@ Modelplane injects ModelExpress env into every engine pod that references a cache. An engine opts in with `--load-format modelexpress`: the first replica loads from its PVC seed and publishes itself as a source, and later replicas pull from a peer over RDMA rather than reading storage again. A replica that finds no -compatible peer, or no fabric to reach one over, falls back to the PVC, so the -cache still has to be sized and kept for every replica. The env is inert unless -the engine opts in, so a cache still works unchanged on a Standard cluster and a -deployment is portable between the two. +compatible peer, or no fabric connecting it to one, falls back to the PVC, so +the cache still has to be sized and kept for every replica. Because the env is +inert unless the engine opts in, a cache still works unchanged on a Standard +cluster and a deployment is portable between the two. Modelplane injects no `--load-format` flag: the ML team's engine command decides whether to use ModelExpress's loader, the same as it decides diff --git a/docs/content/models/model-deployment.md b/docs/content/models/model-deployment.md index ba8b8c8e5..c5ab2a8e4 100644 --- a/docs/content/models/model-deployment.md +++ b/docs/content/models/model-deployment.md @@ -1,13 +1,13 @@ --- title: Deploy a Model weight: 10 -description: Deploy a model to the fleet, from a single pod to disaggregated prefill and decode. +description: Deploy a model to the fleet, from one pod to disaggregated prefill and decode. --- **API:** [`modelplane.ai/v1alpha1` · ModelDeployment]({{< ref "/reference/modeldeployments" >}}) A `ModelDeployment` is the ML team's primary interface. You describe the model you want served, the hardware it needs, and how many copies to run; Modelplane -schedules it onto matching clusters and keeps it running. You never name a +schedules it onto matching clusters and keeps it running. You never pick a cluster. Modelplane is unopinionated about the engine itself. You bring the container and @@ -15,13 +15,13 @@ its flags, and Modelplane shapes a serving topology around it. The engine flags you write carry parallelism, quantization, and KV transfer, never injected by Modelplane. -A deployment's replica shape lives under `spec.template` (mirroring a -Kubernetes Deployment): the replica count is `spec.replicas`, labels for the -composed ModelReplicas and ModelEndpoints go on `spec.template.metadata.labels`, -and everything else lives in `spec.template.spec`. Its -`spec.template.spec.engines` array describes the topology through two choices: +As with a Kubernetes Deployment, `spec.replicas` sets the replica count and +`spec.template` describes each replica. Labels for the composed ModelReplicas +and ModelEndpoints go on `spec.template.metadata.labels`, and everything else +goes in `spec.template.spec`. The `spec.template.spec.engines` array describes +the topology through two choices: -- **One pod or a gang**: whether an engine is a single `Standalone` pod or a +- **One pod or a gang**: whether an engine is one `Standalone` pod or a `Leader` with one or more `Worker` pods coordinating across nodes. - **Unified or disaggregated**: whether `spec.template.spec.serving.mode` keeps prefill and decode together (`Unified`, the default) or splits them across two @@ -35,7 +35,7 @@ How many of each to run is a separate question, covered in The default, and what the [getting started tour]({{< ref "/getting-started" >}}) deploys. One `Standalone` member is one pod on one node, claiming that node's GPUs through its `nodeSelector`. It's usually the right choice when a model fits -on a single node. Within a node, tensor parallelism is an engine flag +on one node. Within a node, tensor parallelism is an engine flag (`--tensor-parallel-size`), not a Modelplane concept. ```yaml {nocopy=true} @@ -53,8 +53,8 @@ node. The pods serve the model together; how the model splits across them (tensor, pipeline, data, or expert parallelism) is up to your engine flags. A gang should use a [`ModelCache`]({{< ref "model-cache.md" >}}) via -`spec.template.spec.modelCacheRef`, so every pod mounts the same weights instead -of each pulling its own. +`spec.template.spec.modelCacheRef`. Every pod in the gang then mounts the same +weights instead of each pulling its own. ```yaml {nocopy=true} modelCacheRef: @@ -114,13 +114,13 @@ so this is a prerequisite Modelplane does not bundle for you. ## Requesting GPUs -You don't name a cluster or a GPU model. Instead each member's `nodeSelector` +You don't specify a cluster or a GPU model. Instead each member's `nodeSelector` lists the hardware its pods need, and Modelplane finds a node pool that has it. The platform team publishes node pools as `InferenceClass` resources, each -describing the devices its nodes carry. Your request is matched against them. +describing the devices on its nodes. Your request is matched against them. -A request names a device (`gpu`), how many of it each pod needs (`count`), and -one or more `selectors` the device must match: +A request has a name (`gpu`), the number of devices each pod needs (`count`), +and one or more `selectors` those devices must match: ```yaml {nocopy=true} nodeSelector: @@ -132,19 +132,19 @@ nodeSelector: device.capacity["gpu.nvidia.com"].memory.compareTo(quantity("40Gi")) >= 0 ``` -Each selector is a single line of [CEL](https://cel.dev/), a small expression -language, that returns true or false for one device. The part in brackets, `"gpu.nvidia.com"`, is the -GPU vendor's driver. The fields after it, like `memory` or `architecture`, are -what the platform team published for that device. This one says "match a GPU -whose memory is at least 40Gi." A device has to match every selector in the -request. Give two selectors to mean "Hopper, with at least 80Gi." +Each selector is a [CEL](https://cel.dev/) expression that returns true or +false for one device. The part in brackets, `"gpu.nvidia.com"`, is the GPU +vendor's driver. The fields after it, like `memory` or `architecture`, are what +the platform team published for that device. This one says "match a GPU whose +memory is at least 40Gi." A device has to match every selector in the request. +Give two selectors to mean "Hopper, with at least 80Gi." ### Requesting more than one device -`devices` is a list, so a member can ask for distinct kinds of hardware at once, -each its own entry with its own `count` and `selectors`. A node pool matches the -member only when it satisfies every entry. This is how you ask for both a GPU and -a fast NIC on the same node: +`devices` is a list, so a member can request distinct kinds of hardware at +once, each as its own entry with its own `count` and `selectors`. A node pool +matches the member only when it satisfies each entry. This is how you ask for +both a GPU and a fast NIC on the same node: ```yaml {nocopy=true} nodeSelector: @@ -169,10 +169,10 @@ device exposes three things: or version), such as `architecture` or `cudaComputeCapability`. - `device.capacity[""].`: a capacity quantity, such as `memory`. -Two helpers build comparable values: `quantity()` parses Kubernetes quantities -like `"40Gi"`, and `semver()` parses versions like `"9.0.0"`. Both support -`compareTo` (which orders two values), `isGreaterThan`, and `isLessThan`. Combine -selectors with the usual CEL operators (`==`, `!=`, `>=`, `&&`, `||`). +Use `quantity()` to parse Kubernetes quantities like `"40Gi"`, and `semver()` to +parse versions like `"9.0.0"`. Both return values that support `compareTo` +(which orders two values), `isGreaterThan`, and `isLessThan`. Combine selectors +with the usual CEL operators (`==`, `!=`, `>=`, `&&`, `||`). ```yaml {nocopy=true} selectors: @@ -192,10 +192,10 @@ selectors: device.capacity["gpu.nvidia.com"].memory.compareTo(quantity("80Gi")) >= 0 ``` -This is the Kubernetes DRA device selector expression surface. The -Kubernetes-specific CEL extension libraries (such as regular expressions and IP -address helpers) aren't available. Selectors in practice are attribute and -capacity comparisons like those above. +Selectors use the CEL that Kubernetes DRA accepts for device selectors, +including `quantity()` and `semver()` but not Kubernetes' other extension +libraries, such as regular expressions and IP address helpers. Selectors in +practice are attribute and capacity comparisons like those above. ### Seeing what's available @@ -209,12 +209,13 @@ kubectl describe inferenceclass gke-l4-1x-g2 The `describe` output shows each device's driver, attributes (like `architecture`), and capacity (like `memory`), which are exactly the keys your -selectors read. If a selector asks for something no published class offers, the +selectors read. If no published class has a device that matches a selector, the deployment won't schedule. ## Sizing a deployment -Three independent numbers control how many pods a deployment runs: +A deployment's pod count comes from separate settings that you size +independently: - **`spec.replicas`** stamps out whole copies of the entire topology. Each replica is a complete serving instance, and replicas usually land on different @@ -225,7 +226,7 @@ Three independent numbers control how many pods a deployment runs: failure drops one copy instead of taking the whole replica out of service. In disaggregated serving they also set the prefill-to-decode ratio. - **`worker.nodes`** sets how many nodes one gang spans: a `Leader` plus that - many `Worker` pods. It's how big a single multi-node engine is. + many `Worker` pods. It's how big one multi-node engine is. ## Scaling diff --git a/docs/content/models/model-endpoint.md b/docs/content/models/model-endpoint.md index fb726cd25..268d3ca92 100644 --- a/docs/content/models/model-endpoint.md +++ b/docs/content/models/model-endpoint.md @@ -5,7 +5,7 @@ description: A reachable inference endpoint, composed per replica or created man --- **API:** [`modelplane.ai/v1alpha1` · ModelEndpoint]({{< ref "/reference/modelendpoints" >}}) -A `ModelEndpoint` is a single reachable inference endpoint that a +A `ModelEndpoint` is a reachable inference endpoint that a [`ModelService`]({{< ref "model-service.md" >}}) can route to. Modelplane creates one for each of your replicas automatically, but you can also create one by hand to point at an inference endpoint Modelplane doesn't run, most often a SaaS diff --git a/docs/content/models/model-service.md b/docs/content/models/model-service.md index a366694c8..7002040fb 100644 --- a/docs/content/models/model-service.md +++ b/docs/content/models/model-service.md @@ -6,14 +6,14 @@ description: Expose a deployment's replicas as one model a caller can name. **API:** [`modelplane.ai/v1alpha1` · ModelService]({{< ref "/reference/modelservices" >}}) A [`ModelDeployment`]({{< ref "model-deployment.md" >}}) serves a model, but its -replicas are scattered across the fleet with no single name. A `ModelService` +replicas are scattered across the fleet with no shared name. A `ModelService` gives them one: a stable name that load-balances across every replica, wherever it runs. A caller names it as the model in an ordinary OpenAI or Anthropic request to a gateway that serves it. A service selects what to route to by label. Behind the scenes, Modelplane -creates one `ModelEndpoint`, a single reachable backend, for each replica of a -deployment and labels it. Two of those labels carry routing intent: +creates one `ModelEndpoint`, a reachable backend, for each replica of a +deployment and sets these routing labels on it: - `modelplane.ai/deployment`: the deployment the replica belongs to. - `modelplane.ai/cluster`: the cluster the replica runs on. @@ -33,7 +33,7 @@ endpoint that any entry matches. The patterns below build on that. ## Route to a whole deployment -The common case: one selector matching a deployment's name reaches every replica, +In the common case, one selector on a deployment's name matches every replica, wherever in the fleet they run. ```yaml {nocopy=true} @@ -67,7 +67,7 @@ spec: ## Route across several deployments Give more than one entry to front several deployments under the same model name. Each -entry contributes its matched endpoints. By default every entry carries equal +entry contributes its matched endpoints. By default every entry has equal weight, so traffic splits evenly between entries and then spreads as evenly as possible across the endpoints each one matches. @@ -92,8 +92,8 @@ takes 80% of requests. The weight applies to the entry as a whole and spreads as evenly as possible across the endpoints it matches, so scaling a deployment up or down doesn't change its share. An entry without a `weight` defaults to 1. -This is the shape of a canary rollout: send most traffic to the stable deployment -and a sliver to the new one, then shift the ratio as confidence grows. +Use this for a canary rollout: send most traffic to the stable deployment and a +sliver to the new one, then shift the ratio as confidence grows. ```yaml {nocopy=true} spec: @@ -112,7 +112,7 @@ spec: The entries don't have to be deployments. One can select a manually created [ModelEndpoint]({{< ref "model-endpoint.md" >}}) that points at an external -provider, so one model name covers both your own replicas and a SaaS endpoint. +provider, so one model name covers your own replicas and a SaaS endpoint. At equal priority the two share traffic by weight. To send the provider only the traffic your replicas can't serve, see [failover tiers](#failover-tiers) below. @@ -156,20 +156,19 @@ spec: ## Timeouts `timeouts` sets how long a gateway waits on the service's endpoints. `request` -bounds a whole request, retries included. `idle` is how long an endpoint may -send nothing. Before the first byte, the gateway gives up on the endpoint, which -counts against its health, and retries the request, on another endpoint if -there is one. Each retry starts the response again, so an `idle` shorter than a -response that isn't streamed makes the backend generate it up to four times -before the caller gets a 504. After the first byte, the stream is cut short. -They default to `300s` and `60s`. +bounds a whole request, retries included. `idle` is how long an endpoint may go +without sending anything. Before the first byte, the gateway gives up on the +endpoint, which counts against its health, and retries the request, on another +endpoint if there is one. Each retry starts the response again, so an `idle` +shorter than a response that isn't streamed makes the backend generate it up to +four times before the caller gets a 504. After the first byte, the stream is cut +short. They default to `300s` and `60s`. Whether a response sends anything early depends on whether the caller streams. A streamed response starts after prefill, so `idle` bounds time to first token and -every gap between chunks after it. A response that isn't streamed sends nothing -until it's complete. If any of a -service's callers don't stream, set `idle` at least as long as `request`, or to -`0s` to disable it. +every gap between chunks after it. A response that isn't streamed doesn't send +anything until it's complete. If any of a service's callers don't stream, set +`idle` at least as long as `request`, or to `0s` to disable it. ```yaml {nocopy=true} spec: @@ -182,20 +181,20 @@ Tune both from what the gateway measures. AI Gateway's `gen_ai.server.time_to_first_token` and `gen_ai.server.request.duration` metrics give each model's latencies, and Envoy's `envoy_cluster_upstream_rq_per_try_idle_timeout` counts idle timeouts. If that -count rises while the endpoints are healthy, `idle` is too short. A restarting -gateway lets requests already in flight run for five minutes, so a restart can -cut off a response allowed longer than that. +count rises while the endpoints' other error counts don't, `idle` is too short. +A restarting gateway lets requests already in flight run for five minutes, so a +restart can cut off a response allowed longer than that. -## How a service reaches its gateways +## Gateways and routes An `InferenceGateway` names the services it serves, through a `serviceSelector` that matches a service's labels. A gateway with no selector serves every service. Label a service for a region and give that region's gateways a matching selector, and only they serve it. -For every gateway that serves it, Modelplane composes a `ModelRoute` that renders -the routing onto that gateway's cluster. You don't write `ModelRoute`s. -`status.routes` counts them, and `kubectl get modelroutes -l +For each gateway that serves the service, Modelplane composes a `ModelRoute` +that renders the routing onto that gateway's cluster. You don't write +`ModelRoute`s. `status.routes` counts them, and `kubectl get modelroutes -l modelplane.ai/service=` shows each one, its gateway, and whether the route is ready there. Look there when a service is Ready but a gateway isn't serving it. @@ -209,9 +208,9 @@ publishes a base URL per API it speaks: ADDRESS=$(kubectl get ig public -o jsonpath='{.status.endpoints.openAI}') ``` -Send a request naming the service. The gateway rewrites the name to whatever -each endpoint's engine or provider expects, so one name reaches replicas and -third-party providers alike: +Send a request with the service as its model. The gateway rewrites the name to +whatever each endpoint's engine or provider expects, so one name covers replicas +and third-party providers alike: ```bash curl "$ADDRESS/chat/completions" \ diff --git a/docs/content/overview/_index.md b/docs/content/overview/_index.md index f1e3ed89b..c512df6d6 100644 --- a/docs/content/overview/_index.md +++ b/docs/content/overview/_index.md @@ -11,7 +11,7 @@ install and run in your own environment, and it orchestrates the models, serving stack, and infrastructure across cloud, neocloud, and on-premise. Modelplane supports running any model and any engine on any infrastructure, with the frontier-level serving topologies and performance the largest models demand, -from a single GPU to disaggregated, multi-node deployments. +from one GPU to disaggregated, multi-node deployments. Modelplane operates across the whole fleet: provisioning inference clusters, scheduling model deployments on compatible clusters, autoscaling model replicas diff --git a/docs/content/overview/ai-tools.md b/docs/content/overview/ai-tools.md index ff139153e..f4346458c 100644 --- a/docs/content/overview/ai-tools.md +++ b/docs/content/overview/ai-tools.md @@ -6,18 +6,17 @@ description: Connect AI assistants and coding agents to the Modelplane docs thro The Modelplane docs are built to be read by AI assistants as well as people. You can connect a coding agent directly to this site, pull any page as Markdown, or -point a model at a single index file that lists the whole documentation set. -Every page also carries a **Copy page** menu next to its title with the same -shortcuts. +point a model at an index file that lists the whole documentation set. A +**Copy page** menu next to each page's title offers the same shortcuts. ## Connect to the MCP server The documentation MCP server lets an assistant search these docs and read any -page in real time, so its answers track the current content instead of its -training data. It exposes two tools: +page in real time, rather than answer from its training data. It exposes two +tools: - `search_modelplane_docs`: search the docs and get back the most relevant sections with their titles, URLs, and snippets. -- `get_modelplane_doc`: fetch the full Markdown of a single page. +- `get_modelplane_doc`: fetch the full Markdown of one page. The server URL is: @@ -64,7 +63,7 @@ Create `.vscode/mcp.json` in your workspace: ``` {{< /tab >}} {{< tab "Other" >}} -Any MCP client that speaks the streamable HTTP transport can connect to the server URL directly. No authentication is required. +Any MCP client that speaks the streamable HTTP transport can connect to the server URL directly, without authenticating. {{< /tab >}} {{< /tabs >}} diff --git a/docs/content/overview/faq.md b/docs/content/overview/faq.md index e52b2b56d..a9b42478f 100644 --- a/docs/content/overview/faq.md +++ b/docs/content/overview/faq.md @@ -1,7 +1,7 @@ --- title: FAQ weight: 35 -description: Short answers to the questions practitioners ask about Modelplane first. +description: Short answers to the questions people ask first about Modelplane. --- @@ -19,7 +19,7 @@ it, routes to it, scales it, and caches its weights across your inference fleet. {{< /qa >}} {{< qa "Does Modelplane replace vLLM or SGLang?" >}} -No, they run the model; Modelplane runs the fleet. A `ModelDeployment` carries +No, they run the model; Modelplane runs the fleet. A `ModelDeployment` specifies your engine container and its flags, and Modelplane composes it onto the right cluster. Switching or upgrading engines is a change to your deployment, not to Modelplane. @@ -27,7 +27,7 @@ Modelplane. {{< qa "How is Modelplane different from KServe or NVIDIA Dynamo?" >}} Scope. KServe and Dynamo are cluster orchestrators: they schedule, scale, route, -and cache within a single Kubernetes cluster. Modelplane runs those operations +and cache within one Kubernetes cluster. Modelplane runs those operations across a fleet of clusters, clouds, and regions. It uses llm-d for inference-aware routing, and installs a per-cluster [serving stack]({{< ref "/platform/inference-cluster.md#serving-stack" >}}) that's @@ -116,8 +116,8 @@ replica on a cluster and pool that fits and has free capacity. {{< /qa >}} {{< qa "Can I serve across regions and clusters behind one endpoint?" >}} -Yes, that's the point. A `ModelService` gives callers one model name and -load-balances across every replica of a deployment, wherever they run. +Yes, a `ModelService` gives callers one model name and load-balances across +every replica of a deployment, wherever they run. {{< /qa >}} {{< qa "Can I route to a managed provider?" >}} diff --git a/docs/content/overview/glossary.md b/docs/content/overview/glossary.md index 8a99b7601..224220f5e 100644 --- a/docs/content/overview/glossary.md +++ b/docs/content/overview/glossary.md @@ -6,9 +6,9 @@ description: Terms used throughout the Modelplane docs and what they mean. ## Modelplane -The open source control plane software. You install Modelplane on a Kubernetes -cluster (the **control cluster**). Modelplane never serves tokens itself; it -orchestrates the clusters and engines that do. +The open source control plane for AI inference. You install Modelplane on a +Kubernetes cluster (the **control cluster**). Modelplane never serves tokens +itself. It orchestrates the clusters and engines that do. ## Control cluster @@ -24,7 +24,7 @@ you can bring your own through an `InferenceCluster` with `source: Existing`. ## Fleet -All inference clusters managed by a single Modelplane control cluster. +All inference clusters managed by one Modelplane control cluster. ## Serving stack diff --git a/docs/content/overview/how-it-works.md b/docs/content/overview/how-it-works.md index 028f613f6..ef59546cf 100644 --- a/docs/content/overview/how-it-works.md +++ b/docs/content/overview/how-it-works.md @@ -66,9 +66,9 @@ model, and Modelplane composes the rest. The hierarchy mirrors Kubernetes core one scope up: `ModelDeployment` → `ModelReplica` → `ModelService` → `ModelEndpoint` parallels `Deployment` → `Pod` → `Service` → -`Endpoint`, across a fleet instead of within a single cluster. +`Endpoint`, across a fleet instead of within one cluster. -## What the control plane reconciles +## Reconciliation Once the resources exist, Modelplane keeps the fleet matching them. Five concerns run continuously: @@ -127,14 +127,13 @@ covers the placement rules and their limits in full. ## Deploying a model -Creating a `ModelDeployment` kicks off the loop end to end. The scheduler -discovers the ready clusters (filtered by your label selector if you set one), -matches each engine's device requests against their pools, and pins each replica -to a cluster that fits. Modelplane composes a `ModelReplica` on each chosen -cluster, turns it into the right serving workload there, creates a `ModelEndpoint` -per replica, and your `ModelService` routes traffic across them under one stable -model name on the gateway. Scale the deployment up or down and the same loop -re-converges. +Creating a `ModelDeployment` starts the whole loop. The scheduler discovers the +ready clusters (filtered by your label selector if you set one) and pins each +replica to a cluster whose pools fit its engines' device requests. Modelplane +composes a `ModelReplica` on each chosen cluster, turns it into the right +serving workload there, creates a `ModelEndpoint` per replica, and your +`ModelService` routes traffic across them under one stable model name on the +gateway. Scale the deployment up or down and the same loop re-converges. ## Serving topologies @@ -144,13 +143,14 @@ service. When a model is too large for one node, an engine becomes a gang: a across nodes. How Modelplane composes and schedules the gang depends on the cluster's [serving stack]({{< ref "/platform/inference-cluster.md#serving-stack" >}}), Standard or Dynamo. Gang deployments should stage their weights through a -`ModelCache`, so the pods share one copy instead of each pulling the same model. +`ModelCache` so that the pods share one copy instead of each pulling the same +model. Disaggregated serving splits prefill and decode into separate engines (`serving.mode: PrefillDecode`) that run on the same cluster and hand off the KV cache between them. Modelplane wires up the cluster-edge routing that pairs each -request's prefill and decode; the engines carry the KV-transfer flags. Both are -described in full in the [model deployment docs]({{< ref "/models/model-deployment" >}}). +request's prefill and decode; you set the KV-transfer flags on the engines. Both +are described in full in the [model deployment docs]({{< ref "/models/model-deployment" >}}). ## Next steps diff --git a/docs/content/overview/why.md b/docs/content/overview/why.md index d14c956a9..30a5a473f 100644 --- a/docs/content/overview/why.md +++ b/docs/content/overview/why.md @@ -20,7 +20,7 @@ AI workloads, adding device-aware scheduling, multi-node inference, distributed serving, and accelerator management. The major open source inference projects are converging on it; among them are vLLM, SGLang, NVIDIA Dynamo, llm-d, Ray, Slurm, KubeAI, and Kueue. Neoclouds like Baseten and CoreWeave have standardized on -Kubernetes for their own operations. Inside a single cluster, the open source +Kubernetes for their own operations. Inside one cluster, the open source stack is now strong. ## Inference is a fleet problem @@ -32,11 +32,11 @@ across multiple clouds and on-premise environments. Large clusters concentrate failure and risk, so fleets of smaller clusters are often preferable, and inference workloads don't bin-pack the way other workloads do. -Inference grows into a fleet, and a new set of problems appears above -any single cluster: +Inference grows into a fleet, and a new set of problems appears above the +cluster level: - Deciding where each model runs across available capacity. -- Optimizing placement across heterogeneous accelerators. +- Making the best use of heterogeneous accelerators. - Failing over across clouds and regions. - Routing by cost, latency, and sovereignty requirements. - Provisioning new capacity as demand grows. diff --git a/docs/content/platform/drain-cluster.md b/docs/content/platform/drain-cluster.md index 769c3522e..1fd641c6c 100644 --- a/docs/content/platform/drain-cluster.md +++ b/docs/content/platform/drain-cluster.md @@ -37,17 +37,17 @@ spec: # cluster source and node pools unchanged ``` -Removing the taint lets the cluster take work again. Nothing reschedules back on -its own: a taint only governs where new replicas can land, so replicas that -moved away stay where they went. +Removing the taint lets the cluster take work again, but Modelplane doesn't move +replicas back. A taint only governs where new replicas can land, so replicas +that moved away stay where they went. ## What happens to running replicas Under `NoExecute`, Modelplane reschedules each replica on the cluster the way it schedules a new one, onto another cluster whose hardware satisfies the deployment's device selectors and that isn't repelling the replica. The move -deletes the replica here and recreates it there, so the model reloads on the new -cluster and any requests still in flight to the old replica are dropped. The +deletes the replica here and recreates it there. The model reloads on the new +cluster, and any requests still in flight to the old replica are dropped. The deployment's other replicas keep serving while one moves. When no other cluster can take a replica, because every candidate is full or @@ -55,14 +55,14 @@ tainted, the deployment runs below its `spec.replicas` until capacity frees up. Its `ReplicasScheduled` condition reports the shortfall, so a drain that can't finish is visible rather than silent. -Under `NoSchedule`, running replicas stay put and only new placement is blocked. +Under `NoSchedule`, only new placement is blocked. Running replicas don't move. ## Keep a deployment through a drain An ML team pins a critical deployment to a cluster through a drain by giving it a matching toleration under `spec.template.spec.tolerations`. A replica that tolerates a cluster's `NoSchedule` taint can still be placed there; one that -tolerates a `NoExecute` taint stays put when that taint is applied. +tolerates a `NoExecute` taint isn't moved when that taint is applied. ```yaml {nocopy=true} apiVersion: modelplane.ai/v1alpha1 @@ -81,7 +81,7 @@ A toleration matches a taint by `key` and `effect`. `operator: Exists` matches any value for the key, while the default `Equal` matches key and value together; an empty `key` with `Exists` tolerates every taint on the cluster. An empty `effect` matches both effects. A replica is placed on, or left on, a tainted -cluster only when it tolerates every taint the cluster carries. +cluster only when it tolerates every taint on the cluster. ## Confirm the drain diff --git a/docs/content/platform/inference-class.md b/docs/content/platform/inference-class.md index 9c3a5323a..7ebc63308 100644 --- a/docs/content/platform/inference-class.md +++ b/docs/content/platform/inference-class.md @@ -27,13 +27,13 @@ A class's `devices` follow Kubernetes (DRA), the mechanism modern Kubernetes uses to match GPUs to pods. Each device has a `driver` (the vendor that owns it, such as `gpu.nvidia.com`), a `count` (how many a node has), typed `attributes` (such as `architecture`), and -`capacity` (quantities, such as `memory`). This mirrors the shape the GPU's DRA -driver publishes on a real node, so what you declare here is what an ML team's -`nodeSelector` matches against and what DRA binds at runtime. +`capacity` (quantities, such as `memory`). This mirrors what the GPU's DRA +driver publishes on a real node, so what you declare in a class is what an ML +team's `nodeSelector` matches against and what DRA binds at runtime. You author the attribute and capacity keys, and there's no fixed list. Pick the -ones an ML team would reasonably select on, the GPU memory, the architecture, the -compute capability, using the same names the driver reports. +ones an ML team would reasonably select on, such as the GPU memory, +architecture, and compute capability, and use the same keys the driver reports. ## DRA and synthetic devices diff --git a/docs/content/platform/inference-cluster.md b/docs/content/platform/inference-cluster.md index 652ade6e9..a99b05138 100644 --- a/docs/content/platform/inference-cluster.md +++ b/docs/content/platform/inference-cluster.md @@ -73,8 +73,8 @@ An existing cluster must meet what Modelplane would otherwise set up for you: - **The `nvidia.com/gpu` taint key, if you taint GPU nodes.** Modelplane's GPU workloads tolerate that key. A different taint keeps them off the nodes. - **A load balancer.** Modelplane exposes the cluster's serving gateway through a - `LoadBalancer` Service, so the cluster needs one that assigns it an external - address. + `LoadBalancer` Service. The cluster needs a load balancer that assigns the + Service an external address. - **No conflicting Gateway controller.** Modelplane installs Envoy Gateway and owns its `GatewayClass`. Don't run another controller claiming the same class. - **A `ReadWriteMany` StorageClass**, if you use a `ModelCache`. See @@ -99,8 +99,8 @@ run a multi-node gang and distribute its weights: LeaderWorkerSet controller, and composes a gang as a Grove `PodCliqueSet` that they gang-schedule all-or-nothing and topology-aware. It also runs a [ModelExpress](https://github.com/ai-dynamo/modelexpress) server that moves - weights between replicas over the fabric, so a later replica pulls a model from - a peer's GPU rather than reading storage again. + weights between replicas over the fabric. A later replica pulls a model from a + peer's GPU rather than reading storage again. A `ModelDeployment` looks the same on either stack. On `Dynamo` an engine can opt into peer-to-peer weight loading with `--load-format modelexpress`. An diff --git a/docs/content/platform/inference-gateway.md b/docs/content/platform/inference-gateway.md index 7df40bcbd..c594f691f 100644 --- a/docs/content/platform/inference-gateway.md +++ b/docs/content/platform/inference-gateway.md @@ -12,14 +12,13 @@ to a cluster serving the model it asked for. It runs on an `InferenceCluster`, named by `spec.clusterName`. The cluster it names can serve models too, or run the gateway alone. -Create as many as you need, one per cluster. A second gateway naming a cluster -that already has one reports `ClusterAlreadyHasGateway` and doesn't become -ready. A gateway is where a request enters your fleet, so run one per place -requests should enter from. `spec.serviceSelector` decides -which `ModelService`s each one serves. Left unset, a gateway serves every -service. Scoping a gateway to a region is how you express residency: label a -service for the EU and it reaches only EU gateways, and from there only the -endpoints it selects. +Create as many as you need, one per cluster. A second gateway on a cluster that +already has one reports `ClusterAlreadyHasGateway` and doesn't become ready. A +gateway is where a request enters your fleet, so run one per place requests +should enter from. `spec.serviceSelector` decides which `ModelService`s each one +serves. Left unset, a gateway serves every service. Scoping a gateway to a +region is how you express residency: label a service for the EU and it reaches +only EU gateways, and from there only the endpoints it selects. Set `spec.tls.certificateRefs` to serve over TLS, with certificates for the names callers will use. The names are yours: point your DNS at the address the gateway @@ -65,8 +64,8 @@ stringData: ## Run behind another gateway -Without `spec.auth` the gateway authenticates nobody. That's the shape for -running behind a gateway that already does: the upstream sets the +Without `spec.auth` the gateway authenticates nobody. That's the configuration +for running behind a gateway that already does: the upstream sets the `x-modelplane-caller` header to name the caller it authenticated, and the gateway trusts it. You have to ensure traffic reaches this gateway only through that front, so nothing else can set the header. diff --git a/docs/content/platform/providers.md b/docs/content/platform/providers.md index 681b29953..af4f38a62 100644 --- a/docs/content/platform/providers.md +++ b/docs/content/platform/providers.md @@ -19,8 +19,8 @@ A provider can show up here in three ways: version), so you can run on the providers below now, ahead of native provisioning. - **Crossplane provider exists.** A Crossplane provider is published for the - cloud. That provider is the path by which native provisioning lands, so it - marks where Modelplane can grow next. + cloud. Native provisioning would build on that provider, so it marks where + Modelplane can grow next. {{< /hint >}} ## Clouds and neoclouds diff --git a/docs/content/platform/telemetry.md b/docs/content/platform/telemetry.md index 02527c69b..227e4dbc2 100644 --- a/docs/content/platform/telemetry.md +++ b/docs/content/platform/telemetry.md @@ -11,27 +11,27 @@ description: Collect normalized metrics across the fleet and send them to any ba Modelplane runs an OpenTelemetry collector on every inference cluster. It collects from every component Modelplane installs. This includes the inference server engine, inference gateway and Envoy proxy, router, and the GPU exporter -your cloud provides. It renames each component's series to a single +your cloud provides. It renames each component's series to a common `modelplane_*` vocabulary and exports them to wherever you say - any backend the collector has an exporter for, not only OTLP. Modelplane allows you to write one destination for your metrics. You don't need to manage per-deployment configurations or update your configuration when a -deployment changes. The collector finds pods itself, so a leader/worker split or a -prefill/decode pair is collected the same as a single pod. +deployment changes. The collector finds pods itself, so a leader/worker split or +a prefill/decode pair is collected the same way as one pod. ## Telemetry workflow -Every series carries `cluster`, `job`, and `instance` labels of the target -resource. A series about a deployment also carries `deployment`, `replica`, -`namespace`, `engine`, and `role` labels. +Every series has `cluster`, `job`, and `instance` labels of the target resource. +A series about a deployment also has `deployment`, `replica`, `namespace`, +`engine`, and `role` labels. Each replica publishes its own series, so combine them in your query. Which combination is right follows from what the metric measures: - `sum by (deployment)`, for anything counted, such as requests, tokens, or queue depth. - `avg by (deployment)`, for a ratio. - - `max by (deployment)`, for a saturation figure an alert fires on. + - `max by (deployment)`, for a saturation figure you alert on. To combine: @@ -40,7 +40,8 @@ sum by (deployment) (rate(modelplane_frontend_request_duration_seconds_count[5m] ``` The replica is an index rather than a pod, so it's bounded by the replica count -and survives a restart and a rolling update. Group by `replica`, not by `instance`: +and doesn't change across a restart or a rolling update. Group by `replica`, not +by `instance`: ```promql # One line per replica, stable across rolling updates @@ -109,7 +110,7 @@ kubectl create secret generic telemetry-credentials \ ``` Reference the Secret from the sink with `secretRef`, and set `auth.bearerTokenKey` to -the key that holds the token: +the token's key in the Secret: ```yaml spec: @@ -124,7 +125,7 @@ spec: ``` Modelplane configures the collector to send the token with every export. Rotating the -token needs no restart. +token doesn't require a restart. If you run Prometheus, export to your Prometheus endpoint instead and query the fleet there: @@ -136,8 +137,8 @@ spec: endpoint: https://prom.example.internal/api/v1/write ``` -If you create more than one sink, all get the entire stream. Each sink carries its -own credential, so a vendor and your own Prometheus don't have to share a Secret. +If you create more than one sink, all get the entire stream. Credentials are per +sink, so a vendor and your own Prometheus don't have to share a Secret. ```yaml spec: @@ -211,14 +212,15 @@ Creating a destination turns collection on everywhere at once, and there's no pe opt-out. Each cluster's collector exports to your backend itself. A cluster needs a route to that -backend to report. Where a sink names a `secretRef`, you create that Secret once on the +backend to report. Where a sink sets a `secretRef`, you create that Secret once on the control plane and Modelplane copies it to every cluster running a collector, so the credential is held on each of them. ## Computing rates, quantiles, and ratios -A collector transforms each measurement as it passes it on. It holds no history, so it -produces no rates and no quantiles. Your backend does that. A fleet-wide p99: +A collector transforms each measurement as it passes it on. It doesn't store past +measurements, so it can't compute rates or quantiles. Your backend does that. A +fleet-wide p99: ```promql histogram_quantile(0.99, sum by (le) ( @@ -235,7 +237,7 @@ Modelplane already knows vLLM's and SGLang's metric names and renames them for y neither needs anything from you here. SGLang needs two flags: `--enable-metrics` to publish `/metrics` at all, and `--collect-tokens-histogram` for the prompt and generation histograms behind `modelplane_request_input_tokens` and `modelplane_request_output_tokens`. Without the -second it publishes those as plain counters and both series stay empty. vLLM needs +second it publishes those as plain counters and both series are empty. vLLM needs nothing. ```yaml @@ -268,13 +270,11 @@ spec: --collect-tokens-histogram ``` -SGLang publishes no queue time per request and no preemption counters, so -`modelplane_request_queue_seconds` and `modelplane_requests_preempted_total` carry vLLM -only. +SGLang doesn't publish request queue time or preemption counters, so only vLLM +populates `modelplane_request_queue_seconds` and `modelplane_requests_preempted_total`. -Any other OpenAI-compatible engine reports its frontend numbers with no configuration. -The gateway measures those, not the engine, so `modelplane_frontend_*` works for an -engine Modelplane has never seen. +Because the gateway, not the engine, measures the `modelplane_frontend_*` series, they +work without configuration for an engine Modelplane has never seen. To normalize that engine's own metrics as well, create a `MetricMapping`: @@ -298,16 +298,15 @@ Modelplane renders every mapping into every cluster's collector, so you write on Modelplane leaves the combining to your backend. A scrape of one replica is one batch, so a collector that added them up would be summing readings taken at different moments, and two readings of one cumulative counter come to twice the traffic that happened. Your -backend holds every replica's series and combines them at query time. +backend stores every replica's series and combines them at query time. Say `fromUnit` whenever the engine measures in something other than the unit the name claims, and Modelplane converts to the base one. Skipping this is the expensive mistake here: a series named `_seconds` that holds milliseconds reads a thousand times fast, and nothing downstream can tell. -Rename only where the measurements agree. Two engines' histograms under one name are worth -less than nothing if their buckets disagree, because a quantile over them is wrong rather -than approximate. +Rename only where the measurements agree. If two engines' histograms share a name and +their buckets disagree, a quantile over them is wrong rather than approximate. ### Examples diff --git a/docs/content/recipes/glm-4.5-air.md b/docs/content/recipes/glm-4.5-air.md index 5987f2a05..3b48c0a9b 100644 --- a/docs/content/recipes/glm-4.5-air.md +++ b/docs/content/recipes/glm-4.5-air.md @@ -1,7 +1,7 @@ --- title: GLM-4.5-Air weight: 25 -description: A 106B MoE served from a GGUF checkpoint via llama.cpp on a single A100. +description: A 106B MoE served from a GGUF checkpoint via llama.cpp on one A100. model: unsloth/GLM-4.5-Air-GGUF:IQ4_XS vendors: [Z.ai] clouds: [GKE] @@ -17,12 +17,12 @@ gpuNote: 1× per node --- A 106B MoE served from an Unsloth GGUF checkpoint via llama.cpp instead of -vLLM, on a single A100 40 GB. Modelplane treats the engine as any -OpenAI-compatible container, so the only changes from a vLLM deployment are -the image and args: the container is still named `engine` and listens on -`:8000`. vLLM can't load this Unsloth quantization format. llama.cpp can, and -`-hf` pulls the checkpoint straight from Hugging Face at startup, so a -one-time deployment needs no `ModelCache`. +vLLM, on one A100 40 GB. Because Modelplane treats llama.cpp like any other +OpenAI-compatible container, the only changes from a vLLM deployment are the +image and args: the container is still named `engine` and listens on `:8000`. +vLLM can't load this Unsloth quantization format. llama.cpp can, and `-hf` +pulls the checkpoint straight from Hugging Face at startup, so a one-time +deployment needs no `ModelCache`. The model is bigger than one A100's VRAM, so `--n-cpu-moe` offloads the MoE expert tensors to host RAM and the GPU runs the active path and KV cache. diff --git a/docs/content/recipes/kimi-k2.md b/docs/content/recipes/kimi-k2.md index 71079c9f5..2c6e4b624 100644 --- a/docs/content/recipes/kimi-k2.md +++ b/docs/content/recipes/kimi-k2.md @@ -23,8 +23,8 @@ model; the native FP8 weights need four such nodes. This recipe was run end to end; the `InferenceClass` and `ModelDeployment` are the exact manifests from that run. Apply the platform side first, then the ML -side. The `InferenceCluster` carries an EC2 capacity reservation placeholder to -edit before applying. +side. Edit the EC2 capacity reservation placeholder in the `InferenceCluster` +before applying it. ## Validated deployments diff --git a/docs/content/recipes/laguna.md b/docs/content/recipes/laguna.md index 85518f6a8..4eed01b33 100644 --- a/docs/content/recipes/laguna.md +++ b/docs/content/recipes/laguna.md @@ -1,7 +1,7 @@ --- title: Laguna-S-2.1 weight: 35 -description: A 118B code MoE served FP8 on a single 8x H100 node on Nebius. +description: A 118B code MoE served FP8 on one 8x H100 node on Nebius. model: poolside/Laguna-S-2.1-FP8 vendors: [Poolside] clouds: [Nebius] @@ -16,14 +16,15 @@ engineImages: [vllm/vllm-openai:v0.25.1, lmsysorg/sglang:v0.5.12.post1-cu129] gpuNote: 8× per node --- -Poolside's Laguna-S-2.1 (118B total, 8B active MoE) served FP8 as a single -`Standalone` vLLM engine on one 8x H100 node on Nebius. The FP8 weights (~121 GiB) -fit one node with headroom for KV cache, so the engine is tensor-parallel across -the 8 GPUs over NVLink, with no gang and no prefill/decode disaggregation. Weights -stage once to a `ModelCache` on a Nebius shared filesystem and mount at `/mnt/models`. +Poolside's Laguna-S-2.1 (118B total, 8B active MoE) served FP8 as a +`Standalone` vLLM engine on one 8x H100 node on Nebius. The FP8 weights +(~121 GiB) fit one node with headroom for KV cache, so the engine is +tensor-parallel across the 8 GPUs over NVLink and doesn't need a gang or +prefill/decode disaggregation. Weights stage once to a `ModelCache` on a Nebius +shared filesystem and mount at `/mnt/models`. This recipe was run end to end on Nebius (`eu-north`): serving and tool calling -validated on a single 8x H100 node. `poolside/Laguna-S-2.1-FP8` is a public +validated on one 8x H100 node. `poolside/Laguna-S-2.1-FP8` is a public repository, so no Hugging Face token or Secret is needed. Apply the platform side first, then the ML side. diff --git a/docs/content/recipes/llama-3.1-8b.md b/docs/content/recipes/llama-3.1-8b.md index ffae4f068..67eed5a7e 100644 --- a/docs/content/recipes/llama-3.1-8b.md +++ b/docs/content/recipes/llama-3.1-8b.md @@ -1,7 +1,7 @@ --- title: Llama-3.1-8B weight: 40 -description: An 8B dense chat model on a single NVIDIA L4. +description: An 8B dense chat model on one NVIDIA L4. model: NousResearch/Meta-Llama-3.1-8B-Instruct vendors: [Meta] clouds: [EKS, GKE] @@ -16,16 +16,16 @@ engineImages: [vllm/vllm-openai:v0.7.3] gpuNote: 1× per node --- -An 8B dense chat model on a single NVIDIA L4. The entry recipe: one `Standalone` -engine, no cache, public weights from a Hugging Face mirror. It carries no -`clusterSelector`, so device capacity alone matches it to any compatible L4 in -the fleet. +An 8B dense chat model on one NVIDIA L4. It's the entry recipe, with one +`Standalone` engine, no cache, and public weights from a Hugging Face mirror. +The deployment has no `clusterSelector`, so device capacity alone matches it to +any compatible L4 in the fleet. This recipe was run end to end on GKE; the `InferenceClass`, `InferenceCluster`, and `ModelDeployment` are the exact manifests from that run. The EKS platform shape is the standard single-L4 recipe. It passes server validation but was not -served in this run. Apply the platform side first, then the ML side. The GKE -`InferenceCluster` carries a GCP project placeholder to edit before applying. +served in this run. Apply the platform side first, then the ML side. Edit the +GCP project placeholder in the GKE `InferenceCluster` before applying it. ## Validated deployments diff --git a/docs/content/recipes/nemotron-3.5-lightning.md b/docs/content/recipes/nemotron-3.5-lightning.md index d0cd8d62d..e777d28e5 100644 --- a/docs/content/recipes/nemotron-3.5-lightning.md +++ b/docs/content/recipes/nemotron-3.5-lightning.md @@ -1,7 +1,7 @@ --- title: Nemotron-3.5-Lightning weight: 15 -description: An open 30B MoE with 3B active parameters served NVFP4 on a single H100 on Nebius. +description: An open 30B MoE with 3B active parameters served NVFP4 on one H100 on Nebius. model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 vendors: [NVIDIA] clouds: [Nebius] @@ -18,14 +18,14 @@ gpuNote: 1× per node NVIDIA's Nemotron-3.5-Lightning, an open 30B mixture-of-experts model with 3B active parameters built for the execution layer of long-running agents, served -NVFP4 as a single `Standalone` vLLM engine on one H100 node on Nebius. -The NVFP4 checkpoint (~20 GiB) fits a single GPU with headroom for the KV and -Mamba caches, so the engine needs no tensor parallelism, no gang, and no +NVFP4 as a `Standalone` vLLM engine on one H100 node on Nebius. +Because the NVFP4 checkpoint (~20 GiB) fits one GPU with headroom for the KV +and Mamba caches, the engine doesn't need tensor parallelism, a gang, or prefill/decode disaggregation. Weights stage once to a `ModelCache` on a Nebius shared filesystem and mount at `/mnt/models`. This recipe was run end to end on Nebius (`eu-north`): serving and tool -calling validated on a single H100 node. +calling validated on one H100 node. `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4` is a public repository (OpenMDW-1.1), so no Hugging Face token or Secret is needed. Apply the platform side first, then the ML side. diff --git a/docs/content/recipes/qwen2.5-72b.md b/docs/content/recipes/qwen2.5-72b.md index e9bee702f..fec544dda 100644 --- a/docs/content/recipes/qwen2.5-72b.md +++ b/docs/content/recipes/qwen2.5-72b.md @@ -1,7 +1,7 @@ --- title: Qwen2.5-72B weight: 37 -description: A 72B dense chat model (AWQ INT4) on a single 80 GB GPU, on AKS and Nebius. +description: A 72B dense chat model (AWQ INT4) on one 80 GB GPU, on AKS and Nebius. model: Qwen/Qwen2.5-72B-Instruct-AWQ vendors: [Qwen] clouds: [AKS, Nebius] @@ -16,14 +16,14 @@ engineImages: [vllm/vllm-openai:v0.23.0] gpuNote: 1× per node --- -A 72B dense chat model served from an AWQ INT4 quantization on a single 80 GB -GPU per replica: one `Standalone` engine fed by a `ModelCache`. The platform -side comes in two shapes - an A100 on AKS and an H100 on Nebius - and the ML -side is the same manifest for both. The deployment carries no -`clusterSelector` and two replicas, and each pool holds exactly one GPU, so -with both platforms applied one replica runs on each. The service then splits -traffic between the two GPUs by weight, which the last section uses to compare -them. To serve on just one platform, apply one tab and drop `replicas` to 1. +A 72B dense chat model served from an AWQ INT4 quantization on one 80 GB GPU +per replica: one `Standalone` engine fed by a `ModelCache`. The platform side +covers an A100 on AKS and an H100 on Nebius, and the ML side is the same +manifest for both. The deployment has two replicas and no +`clusterSelector`. Each pool has exactly one GPU, so with both platforms +applied one replica runs on each. The service then splits traffic between the +two GPUs by weight, which the last section uses to compare them. To serve on +just one platform, apply one tab and drop `replicas` to 1. These manifests mirror the repository's AKS and Nebius demos. Apply the platform side first, then the ML side. @@ -58,7 +58,7 @@ platform side first, then the ML side. ## Compare the A100 and the H100 Replicas are fleet-wide, not per-cluster: the deployment's `replicas: 2` means -two complete serving instances, and because each pool holds a single 80 GB GPU +two complete serving instances, and because each pool has exactly one 80 GB GPU they land one on the A100 and one on the H100. Modelplane labels each replica's endpoint with the cluster it runs on, so the service can split traffic between the platforms by weight. This service pairs the deployment @@ -70,7 +70,7 @@ under the same model name: Both GPUs now serve the same workload, so their engine metrics give a direct performance comparison: scrape each replica's latency and throughput as in [Monitor the Fleet]({{< ref "/platform/telemetry.md" >}}) and -read the two side by side. Weights are relative, so once one platform wins, -shift the 50/50 toward it - 80/20, and as far as 100/0 - without touching the -deployment. +read the two side by side. Weights are relative, so you can shift traffic +toward whichever platform performs better (80/20, or as far as 100/0) without +touching the deployment. diff --git a/docs/content/recipes/qwen2.5-7b.md b/docs/content/recipes/qwen2.5-7b.md index e44ba551c..6d750c99c 100644 --- a/docs/content/recipes/qwen2.5-7b.md +++ b/docs/content/recipes/qwen2.5-7b.md @@ -1,7 +1,7 @@ --- title: Qwen2.5-7B weight: 12 -description: A 7B dense chat model (AWQ INT4) on a single NVIDIA A16 on Vultr. +description: A 7B dense chat model (AWQ INT4) on one NVIDIA A16 on Vultr. model: Qwen/Qwen2.5-7B-Instruct-AWQ vendors: [Qwen] clouds: [Vultr] @@ -16,17 +16,18 @@ engineImages: [vllm/vllm-openai:v0.9.2] gpuNote: 1× per node --- -A 7B dense chat model served from an AWQ INT4 quantization on a single NVIDIA -A16 on Vultr: one `Standalone` engine, no cache, weights pulled straight from -Hugging Face. The A16 slice on the `vcg-a16-6c-64g-16vram` plan carries 16 GiB -of VRAM, so the INT4 weights (~5 GiB) fit with headroom for KV cache; +A 7B dense chat model served from an AWQ INT4 quantization on one NVIDIA A16 +on Vultr: one `Standalone` engine, no cache, weights pulled straight from +Hugging Face. The A16 slice on the `vcg-a16-6c-64g-16vram` plan has 16 GiB of +VRAM, so the INT4 weights (~5 GiB) fit with headroom for KV cache; `--gpu-memory-utilization=0.85` and `--enforce-eager` keep the engine inside the small card. This recipe was run end to end on Vultr (`ewr`); the `InferenceClass`, `InferenceCluster`, and `ModelDeployment` are the exact manifests from that -run. GPU plans are region-gated on Vultr, so check the plan is offered in your -region before applying. Apply the platform side first, then the ML side. +run. GPU plan availability varies by Vultr region, so check the plan is offered +in your region before applying. Apply the platform side first, then the ML +side. ## Validated deployments diff --git a/docs/content/recipes/qwen3-8b.md b/docs/content/recipes/qwen3-8b.md index 645b0e67f..9dd76d057 100644 --- a/docs/content/recipes/qwen3-8b.md +++ b/docs/content/recipes/qwen3-8b.md @@ -1,7 +1,7 @@ --- title: Qwen3-8B weight: 10 -description: An 8.2B dense chat model on a single NVIDIA L4. +description: An 8.2B dense chat model on one NVIDIA L4. model: Qwen/Qwen3-8B vendors: [Qwen] clouds: [EKS] @@ -20,8 +20,8 @@ features: note: for low latency and small batch sizes --- -An 8.2B dense chat model on a single NVIDIA L4. The smallest recipe: one -`Standalone` engine, no cache, weights pulled straight from Hugging Face. +An 8.2B dense chat model on one NVIDIA L4. It's the smallest recipe, with one +`Standalone` engine, no cache, and weights pulled straight from Hugging Face. This recipe was run end to end; the `InferenceClass` and `ModelDeployment` are the exact manifests from that run. Apply the platform side first, then the ML @@ -46,17 +46,17 @@ side. ## Speculative decoding The same model and platform also serve with n-gram (prompt-lookup) speculative -decoding, which proposes tokens by matching the prompt and so needs no draft -model or second set of weights. On copy-heavy output, editing a pasted code -block where most output tokens are copied from the prompt, it roughly doubles -decode throughput and halves the time per output token: +decoding, which proposes tokens by matching the prompt and so doesn't need a +draft model or a second set of weights. On copy-heavy output, editing a pasted +code block where most output tokens are copied from the prompt, it roughly +doubles decode throughput and halves the time per output token: | Metric | Without speculation | With n-gram speculation | |---|---|---| | Output token throughput (tok/s) | 16.10 | 39.01 | | Mean TPOT (ms/token) | 60.20 | 24.21 | -Measured on a single L4 (`vllm/vllm-openai:v0.23.0`, Qwen3-8B, 30 copy-heavy +Measured on one L4 (`vllm/vllm-openai:v0.23.0`, Qwen3-8B, 30 copy-heavy prompts at concurrency 1) against the same model without `--speculative-config`; the speculative run accepted 65% of drafted tokens, a mean acceptance length of 4.27 of 5. Speculation proposes several tokens per decode step and verifies them in diff --git a/docs/content/recipes/qwen3-coder.md b/docs/content/recipes/qwen3-coder.md index 64fd86c66..60c94cbae 100644 --- a/docs/content/recipes/qwen3-coder.md +++ b/docs/content/recipes/qwen3-coder.md @@ -16,15 +16,15 @@ engineImages: [vllm/vllm-openai:v0.23.0, lmsysorg/sglang:v0.5.10.post1-runtime] gpuNote: 8× per node --- -A 480B code MoE (35B active). Two validated shapes: the BF16 weights span two -H200 nodes as a gang over EFA, served from a `ModelCache`; the FP8 checkpoint -fits one node, so it runs as a single `Standalone` engine on SGLang with no +A 480B code MoE (35B active), validated in two deployments. The BF16 weights +span two H200 nodes as a gang over EFA, served from a `ModelCache`. The FP8 +checkpoint fits one node, so it runs as a `Standalone` engine on SGLang with no cache. -Both shapes were run end to end; the `InferenceClass` and `ModelDeployment` are -the exact manifests from those runs. Apply the platform side first, then the ML -side. The `InferenceCluster` carries an EC2 capacity reservation placeholder to -edit before applying. +Both deployments were run end to end; the `InferenceClass` and +`ModelDeployment` are the exact manifests from those runs. Apply the platform +side first, then the ML side. Edit the EC2 capacity reservation placeholder in +the `InferenceCluster` before applying it. ## Validated deployments