Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 14 additions & 14 deletions docs/content/architecture/_index.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,8 @@ matter here:
- **Composition functions** are that controller logic. A function is a small gRPC
service handed the observed XR and the resources it depends on, which returns
the desired child resources. An XR runs a pipeline of one or more functions
every reconcile; in Modelplane each is typically a single function, so the rest
of this section says "the function" for short.
every reconcile; in Modelplane each pipeline typically has one function, so
the rest of this section says "the function" for short.
- **Providers** are controllers that manage external systems through their own
managed resources: `provider-gcp` and `provider-aws` for cloud APIs,
`provider-helm` for Helm releases, `provider-kubernetes` for arbitrary objects
Expand All @@ -43,7 +43,7 @@ The resource model mirrors Kubernetes core, one scope up:
rather than within one. A `ModelDeployment` composes a `ModelReplica` per replica,
a `ModelReplica` composes the serving workload on its target cluster, and a
`ModelService` routes across the `ModelEndpoint`s. If you know how those core
objects relate, you already know the shape of Modelplane's.
objects relate, you already know how Modelplane's fit together.

## Why Crossplane?

Expand All @@ -54,8 +54,8 @@ ways: providers and functions.

**Providers** give us reach. Modelplane has to provision Kubernetes clusters and
all the infrastructure they need across different clouds, then install software
onto them. That's an enormous surface, and providers cover it without us rolling
our own controllers for each cloud API and Helm release.
onto them. That spans many cloud APIs and Helm releases, and providers cover
them without us rolling our own controller for each.

**Functions** are where Modelplane's own logic lives, and writing it as
composition functions buys several things:
Expand All @@ -74,26 +74,26 @@ composition functions buys several things:
for contributors. The performance-sensitive distributed-systems core stays in
Go, where Crossplane and its providers already are.

The bet underneath both is that inference infrastructure is the same shape of
The bet underneath both is that inference infrastructure is the same kind of
problem as cloud infrastructure, which Crossplane manages well. Building on it
lets Modelplane spend its effort on the part that's actually inference-specific.

## The control cluster and the fleet

Modelplane runs on a **control cluster** and manages a fleet of **workload
clusters**, the `InferenceCluster`s. The split is deliberate: the control plane
holds no GPUs and serves no tokens. It schedules and composes, and the
has no GPUs and doesn't serve tokens. It schedules and composes, and the
workload clusters do the serving.

The control cluster runs Crossplane, the Modelplane composition functions (one
per resource, each a pod Crossplane calls per reconcile), and the providers. It
also holds every Modelplane resource and the `ProviderConfig`s that let the
providers reach each workload cluster, built from that cluster's kubeconfig.
also holds every Modelplane resource and the `ProviderConfig`s the providers use
to connect to each workload cluster, built from that cluster's kubeconfig.

Crossplane core drives everything. Each reconcile it asks a function what a
resource should compose and gets back the desired resources. Core then reconciles
them, applying the provider resources that the providers act on. A function only
computes desired state. It never reaches a provider or a cluster itself.
Crossplane core drives everything. Each reconcile it calls a resource's function
and gets back the desired resources. Core then reconciles them, applying the
provider resources that the providers act on. A function only computes desired
state. It never reaches a provider or a cluster itself.

```mermaid
flowchart TB
Expand All @@ -119,7 +119,7 @@ The exact components evolve, but Modelplane composes and owns all of them. For
provisioned clusters the providers also create the cluster and its node pools
first.

## How a deployment is composed
## Composing a deployment

A resource composes others, which compose others, until the tree bottoms out in
provider resources and plain Kubernetes objects. A `ModelDeployment` is the
Expand Down
51 changes: 25 additions & 26 deletions docs/content/architecture/scheduling.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@ description: How Modelplane places a deployment's replicas across the fleet, and
**API:** [`modelplane.ai/v1alpha1` · ModelDeployment]({{< ref "/reference/modeldeployments" >}})
<!-- vale write-good.Passive = NO -->
When an ML team creates a [ModelDeployment]({{< ref "/models/model-deployment.md" >}}),
the fleet scheduler decides which cluster each replica runs on and which node
pool each engine uses. Platform teams don't drive it directly, but what they
the fleet scheduler picks the cluster each replica runs on and the node pool
each engine uses. Platform teams don't drive it directly, but what they
publish, the clusters, their labels, and each pool's
[InferenceClass]({{< ref "/platform/inference-class.md" >}}), is exactly what the
scheduler matches against. This page explains how it places work and where it
Expand All @@ -22,9 +22,9 @@ every existing `ModelReplica`, and returns a placement. Given the same inputs it
returns the same placement, so it's safe to run continuously.

The key consequence is stability. Existing replicas are *inputs*, not decisions.
A healthy replica is never moved to improve the global picture, even if a better
cluster appears later. This keeps placement from churning underneath a running
deployment.
The scheduler never moves a replica to improve the global picture, even if a
better cluster appears later. This keeps placement from churning underneath a
running deployment.

## Two-level matching

Expand All @@ -39,8 +39,8 @@ against what the platform team published.
`nodeSelector.devices` against the devices a pool's `InferenceClass` publishes.
A request is a real DRA request: a `count` and CEL selectors over a device's
attributes and capacity, such as "a GPU with at least 141Gi of memory." A pool
fits a member when it has devices satisfying every request, with `count` to
cover them.
fits a member when it satisfies each request, with enough matching devices to
cover its `count`.

The CEL is the same expression an ML engineer would write in a DRA
`ResourceClaim`, evaluated against the devices the `InferenceClass` declares. The
Expand All @@ -50,9 +50,9 @@ pool only if the class publishes the attributes and capacity it asks for.
## Co-scheduling and pools

A replica is a set of engines placed together on one cluster. Within a replica,
every member of a single engine is placed on **one** pool: each member carries
its own `nodeSelector`, but the scheduler requires a single pool that satisfies
them all.
every member of an engine is placed on **one** pool: each member has its own
`nodeSelector`, but the scheduler only places the engine on a pool that
satisfies them all.

It works this way because a gang's members coordinate over their pool's
interconnect fabric, and the scheduler can't reason about fabric. Pool identity
Expand All @@ -77,25 +77,24 @@ graph TD
R --> D
```

A member with no `nodeSelector` claims no devices. It matches the engine's pool
at no node cost and rides along on the gang's nodes, packed there by the
cluster's own scheduler.
A member with no `nodeSelector` doesn't claim any devices. It matches the
engine's pool at no node cost, and the cluster's own scheduler packs its pods
onto the gang's nodes.

## Counting capacity in nodes

Capacity is gated on **nodes**, not on individual GPUs. The only number the
Capacity is counted in **nodes**, not individual GPUs. The only number the
scheduler reads from a member is its node cost:

```text
nodes = pods × copies
pods = 1 for a Standalone or Leader, or worker.nodes for a Worker
```

A member that resolves no `claim: DRA` device, because it carried no
`nodeSelector` or matched only synthetic devices, costs zero nodes. The scheduler
sums the cost of a replica's members and places the replica only where every
engine's pool has enough free nodes, tracking a running ledger so it never
overcommits a cluster.
A member that resolves no `claim: DRA` device, because it has no `nodeSelector`
or matches only synthetic devices, costs zero nodes. The scheduler sums the cost
of a replica's members and places the replica only where every engine's pool has
enough free nodes, tracking a running ledger so it never overcommits a cluster.

This accounting is deliberately coarse. The control-plane scheduler answers
"could this cluster plausibly host this replica," not "exactly which GPU does
Expand All @@ -106,8 +105,8 @@ state.

## Pinning placement to a pool

The scheduler's pool choice is enforced, not advisory. Each scheduled pod carries
a Kubernetes `nodeSelector` on the `modelplane.ai/pool` node label, so it can only
The scheduler's pool choice is enforced, not advisory. Each scheduled pod has a
Kubernetes `nodeSelector` on the `modelplane.ai/pool` node label, so it can only
land on the pool the scheduler chose. Without it, the cluster's scheduler could
place a pod on any pool whose devices match its DRA claim, and the fleet's
per-pool accounting would drift from where pods actually run.
Expand All @@ -127,21 +126,21 @@ Scheduling runs in two phases each reconcile:
`nodeSelector`. A degraded cluster, one that's not Ready or has no gateway
address, is still retained; transient outages surface through the deployment's
conditions, not re-placement.
- **Fill.** If the deployment wants more replicas than were retained, the
- **Fill.** If the deployment specifies more replicas than were retained, the
shortfall is placed one at a time, each onto the eligible cluster hosting the
fewest of this deployment's replicas, spreading before packing. If it wants
fewer, the highest-index replicas are dropped first.
fewest of this deployment's replicas, spreading before packing. If it
specifies fewer, the highest-index replicas are dropped first.
<!-- vale write-good.TooWordy = YES -->

A replica never changes cluster. If its cluster is deleted, the replica stops
being emitted, Crossplane garbage-collects it, and the fill phase mints a fresh
being emitted, Crossplane garbage-collects it, and the fill phase creates a new
replica elsewhere. Moving is always delete-plus-create, mirroring how Kubernetes
treats a pod whose node is gone.

## Known limitations

The scheduler is built to be conservative and predictable rather than optimal.
Two limits follow from that, both tracked for future work:
The limits below follow from that, and each is tracked for future work:

- **A whole node is charged per pod**
([#172](https://github.com/modelplaneai/modelplane/issues/172)). A pod that
Expand Down
6 changes: 3 additions & 3 deletions docs/content/getting-started/_index.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,10 +11,10 @@ against it. Without it, every change on one side creates work for the other.
When the platform team updates infrastructure, ML teams have to react. When
model requirements change, the platform team gets a request.

With Modelplane, the platform team publishes hardware without knowing what
With Modelplane, the platform team publishes its hardware without knowing what
models will run on it. The ML team declares what a model needs without knowing
what clusters exist. The control plane resolves it and keeps it current as
both sides change.
what clusters exist. The control plane places the model on matching hardware
and moves it only when a change on either side breaks the match.

In this tour, you'll switch between provisioning infrastructure and declaring a
model to see how they interact. By the end you'll have a GPU fleet across three regions and one OpenAI-compatible endpoint routing to a model served across two of them.
Expand Down
10 changes: 5 additions & 5 deletions docs/content/getting-started/build-the-platform.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@ weight: 20
description: Set up the gateway, give the control plane cloud credentials, and provision your first GPU cluster.
---
This is the platform team's side of Modelplane. You set up the gateway that
fronts your models, give the control plane cloud credentials, and register your
first GPU cluster: a hardware profile published as an `InferenceClass` and an
`InferenceCluster` that offers it.
fronts your models, give the control plane credentials for your cloud account,
and register your first GPU cluster: a hardware profile published as an
`InferenceClass` and an `InferenceCluster` that offers it.

In the next step, the ML team will create a model deployment that schedules
against this capacity without knowing which cluster it runs on.
Expand All @@ -31,8 +31,8 @@ against this capacity without knowing which cluster it runs on.
| `roles/iam.serviceAccountUser` | attaching that account to the nodes |
| `roles/resourcemanager.projectIamAdmin` | granting the node account `container.admin` |

The last one is worth a look before you hand the key over. Modelplane grants
the node service account `roles/container.admin`, so the credential doing the
Review the last one before you hand the key over. Modelplane grants the node
service account `roles/container.admin`, so the credential doing the
provisioning has to be able to set project IAM policy.
{{< /tab >}}
{{< tab "AKS" >}}
Expand Down
16 changes: 8 additions & 8 deletions docs/content/getting-started/clean-up.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,9 @@ plane.
## Delete model resources

Delete model resources before clusters. A cluster refuses deletion while
anything still runs on it. Foreground cascading deletion holds each resource
until what it composed on the clusters is gone, so a cluster isn't released
while that's still being removed:
anything still runs on it. Foreground cascading deletion removes a resource
only after what it composed on the clusters is gone. That stops a cluster being
released while its workloads are still being torn down:

```bash
kubectl delete md --all -n ml-team --cascade=foreground
Expand All @@ -21,9 +21,9 @@ kubectl delete ms --all -n ml-team --cascade=foreground
## Delete the gateway

Delete the gateway before its cluster. The `InferenceGateway` runs a load balancer
on the cluster it names; deleting it while that cluster is still up lets the load
balancer be removed, rather than leaking it when the cluster goes. Foreground
deletion holds the gateway until its objects on the cluster are deleted:
on the cluster it names. If that cluster is deleted first, the load balancer is
orphaned. Foreground deletion removes the gateway only after its objects on the
cluster are deleted:

```bash
kubectl delete ig --all --cascade=foreground
Expand All @@ -33,8 +33,8 @@ kubectl delete ig --all --cascade=foreground

Delete all clusters with foreground cascading deletion. The serving stack on each
workload cluster must uninstall while that cluster's API server is still
reachable. Foreground deletion holds each cluster object until its stack
finishes. Background deletion can orphan cloud resources.
reachable. Foreground deletion removes each cluster object only after its stack
finishes uninstalling. Background deletion can orphan cloud resources.

```bash
kubectl delete ic --all --cascade=foreground
Expand Down
3 changes: 2 additions & 1 deletion docs/content/getting-started/deploying-a-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,7 +120,8 @@ You should get a response in a few seconds:
## Next step

The platform team declared capacity and in this guide the ML team deployed a
model behind a stable endpoint. Neither team needed to know what the other was doing. Modelplane matched them.
model behind a stable endpoint. Each team worked without needing to know what
the other was doing. Modelplane matched them.

In the next step, the platform team grows the fleet. [Scale the platform]({{< ref "getting-started/scale-the-platform.md" >}}) to add more clusters across regions.

10 changes: 5 additions & 5 deletions docs/content/getting-started/scale-the-model.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
---
title: Scale the model
weight: 50
description: Serve the model from two regions behind a single endpoint.
description: Serve the model from two regions behind one endpoint.
---
A `ModelService` can front more than one `ModelDeployment`. Here you add a second
deployment, pinned to a different region, and point the same service at both. The
deployment in a different region and point the same service at both. The
endpoint you already curled stays the same. Behind it, traffic now load-balances
across two regions.

Expand Down Expand Up @@ -72,7 +72,7 @@ Update the `ModelService` to select both deployments. Each entry in
The model name doesn't change. Callers that had it before still have it; they
don't know the fleet changed. The gateway load-balances across both regions, and
if one region's replicas fail it sends every request to the other. The gateway
itself runs in one region, so surviving the loss of that region takes a second
itself runs in one region. Tolerating the loss of that region requires a second
gateway on a cluster in the other. Send the same request as before:

```bash
Expand All @@ -91,10 +91,10 @@ kubectl run -i --rm curl-test \

## That's the tour

You stood up a control plane, built a multi-region GPU fleet, deployed a model
You installed a control plane, built a multi-region GPU fleet, deployed a model
across it, and ended with one stable endpoint serving requests. The platform
team published hardware. The ML team described what the model needs. Modelplane
placed them and served behind a single endpoint.
placed the model and served it behind that endpoint.

[Clean up]({{< ref "getting-started/clean-up.md" >}}) tears everything down
when you're done.
Expand Down
2 changes: 1 addition & 1 deletion docs/content/getting-started/scale-the-platform.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,4 +80,4 @@ its deployment changes in a way that no longer fits where it runs.

## Next step

The fleet has grown with larger-GPU capacity. The ML team is next. [Scale the model]({{< ref "getting-started/scale-the-model.md" >}}) to serve it across the fleet behind a single endpoint.
The fleet has grown with larger-GPU capacity. The ML team is next. [Scale the model]({{< ref "getting-started/scale-the-model.md" >}}) to serve it across the fleet behind one endpoint.
2 changes: 1 addition & 1 deletion docs/content/guides/anthropic-messages-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ client that speaks the Messages API, including Claude Code via
`ANTHROPIC_BASE_URL`, names the service as the model in its request. See
[Alternate APIs]({{< ref "/models/model-service.md" >}}) for the detail.

This recipe serves Qwen3-8B on a single NVIDIA H100 on Nebius, with tool calling
This recipe serves Qwen3-8B on one NVIDIA H100 on Nebius, with tool calling
on: `--enable-auto-tool-choice` and `--tool-call-parser=hermes` are what let
Claude Code's tool use work. An 8B model needs a fraction of an H100, so the GPU
has ample headroom. Apply the platform side first, then the ML side.
Expand Down
8 changes: 4 additions & 4 deletions docs/content/guides/serving-multi-node-on-dynamo.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,8 +48,8 @@ kubectl create namespace ml-team

The `Leader` and `Worker` run the same `vllm serve`, differing only in node rank.
`$(MODELPLANE_LEADER_ADDRESS)` resolves to the leader on either stack, but
`$(MODELPLANE_RANK)` isn't injected on Dynamo yet, so the worker derives its rank
from Grove's `GROVE_PCLQ_POD_INDEX`.
`$(MODELPLANE_RANK)` isn't injected on Dynamo yet. The worker derives its rank
from Grove's `GROVE_PCLQ_POD_INDEX` instead.
[Multi-node deployments]({{< ref "/models/model-deployment.md#multi-node" >}})
covers this. Both load with `--load-format modelexpress`, so they read the cached
weights through the Dynamo stack's ModelExpress server.
Expand Down Expand Up @@ -98,8 +98,8 @@ seeds its weights from the cache and publishes itself as a source; the second
loads them straight from the first, peer-to-peer, rather than reading the cache
again.

Each replica is a gang of two nodes, so a second replica needs two more nodes.
Grow the pool to four, then scale the deployment:
Each replica is a gang of two nodes. A second replica needs two more, so grow
the pool to four, then scale the deployment:

```bash
kubectl patch ic/eks-us-east-dynamo --type=json -p '[
Expand Down
4 changes: 2 additions & 2 deletions docs/content/install/_index.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
title: Install
weight: 8
navLanding: "Install the control plane"
description: Stand up the Modelplane control plane on a Kubernetes cluster you run.
description: Install the Modelplane control plane on a Kubernetes cluster you run.
---
Modelplane's control plane is where everything runs: the Crossplane runtime, the
providers it provisions infrastructure through, and the composition functions
Expand Down Expand Up @@ -74,7 +74,7 @@ functions that reconcile them:

{{< manifests "install/configuration.yaml" >}}

Wait until the configuration is healthy:
Wait for the configuration to become `Healthy`:

```bash
kubectl wait configuration/modelplane --for=condition=Healthy --timeout=5m
Expand Down
Loading
Loading