GPUStack Operator makes GPUs, NPUs and other accelerators available to Kubernetes workloads. It discovers the hardware on each node, groups similar capacity into pools and uses Kueue to admit workloads when a pool has room.
Read the architecture overview to see how the pieces fit together, or browse the documentation.
- Heterogeneous Devices — discover accelerators across manufacturers and allocate whole devices, shared devices, logical slices or hardware partitions.
- RDMA Networking — give workloads RDMA network interfaces alongside their accelerators.
- Topology Management — place related workloads within the requested network or location domain.
- KV Cache — share inference cache across workloads.
- Model Delivery — fetch, verify and cache model weights on nodes before workloads need them.
- Model Deployment — run model serving replicas with routing and optional prefill/decode roles.
- Accelerated Instances — enter accelerator-backed container workspaces over SSH.
Use Kubernetes 1.29 or newer for scheduling and node model delivery, Helm 3.8 or newer, and cluster-admin access. The chart accepts Kubernetes 1.23 and newer, but simple allocation on 1.23–1.28 is untested; see version requirements.
Install the manufacturer's driver on each accelerator node. Some devices also need a container toolkit; see Vendor Prerequisites.
helm repo add gpustack https://docs.gpustack.ai/gpustack-operator/charts
helm install gpustack-operator gpustack/gpustack-operator \
--namespace gpustack-system --create-namespaceThe chart includes Node Feature Discovery, Kueue and the storage drivers. See Installation Modes if your cluster already runs any of them. For upgrades from v0.7.x or earlier, follow the subchart migration guide.
helm uninstall gpustack-operator --namespace gpustack-systemRemoving the bundled Kueue chart also removes its CRDs and the queues and workloads stored in
them. To use a separately managed Kueue installation, install GPUStack with
--set kueue.enabled=false. To remove resources created by the worker during uninstall, set
cleanupOnUninstall=true when installing the chart.
The examples below assume a namespace named team-a, accelerator nodes with their drivers
installed, and an InstanceType with available capacity. Replace the queue and InstanceType
names with values from kubectl get instancetypes before applying a workload.
kubectl get devices shows the discovered hardware. The ENTRANCE column of
kubectl get instancetypes gives the queue name for a Pod. This one asks for a whole NVIDIA GPU:
apiVersion: v1
kind: Pod
metadata:
name: gpu-workload
namespace: team-a
labels:
kueue.x-k8s.io/queue-name: gpustack-fnv64-c4680bb149644f1c
spec:
restartPolicy: Never
containers:
- name: worker
image: ubuntu:24.04
command:
- "sleep"
- "3600"
resources:
limits:
nvidia.com/gpu: "1"The queue name is an example; use the ENTRANCE of your chosen pool. See
Accelerator Requests for shared, sliced and partitioned Pods.
To give that Pod a network endpoint, put both requests in the same container:
resources:
limits:
nvidia.com/gpu: "1"
device.gpustack.ai/rdma.shared: "1"This replaces the resources block above. A shared endpoint suits a container that needs one
interface; use an exclusive request when it must be the adapter's only tenant. The node needs a
usable endpoint, and NUMA alignment depends on its kubelet policy. See
RDMA Operations for the other resource keys and checks.
When nodes already carry region and zone labels, an administrator can publish their hierarchy:
apiVersion: worker.gpustack.ai/v1
kind: TopologySource
metadata:
name: cloud-zones
spec:
nodeSelector: {}
levels:
- topology.kubernetes.io/region
- topology.kubernetes.io/zone
nodeLabels: {}For a role whose replica spans multiple Pods, request one zone for each replica group:
roles:
- name: server
replicas: 2
size: 4
topology:
requiredLevel: topology.kubernetes.io/zoneThis is a role fragment; the engine still needs its own distributed execution settings. Use Topology Operations to verify the source and the Kueue topology before depending on that placement.
An administrator creates the store, its capacity pool and a grant for team-a, in that order:
apiVersion: worker.gpustack.ai/v1
kind: KVCacheBackend
metadata:
name: mooncake-dram
spec:
type: Mooncake
connection:
managed:
leader: {}
members:
- nodeSelector:
kubernetes.io/os: linux
medium: DRAM
capacityPerMember: 8Gi
---
apiVersion: worker.gpustack.ai/v1
kind: KVCachePool
metadata:
name: shared-dram
spec:
backends:
- mooncake-dram
quota:
total: 8Gi
---
apiVersion: worker.gpustack.ai/v1
kind: KVCachePoolBinding
metadata:
name: qwen-cache
namespace: team-a
spec:
poolRef:
name: shared-dram
quota:
ceiling: 4Gi
domain:
blockSize: 64
dtype: bfloat16The ModelDeployment below names the binding, which grants this namespace access to the cache.
Choose member nodes, quotas, block size and dtype for your engine before applying these resources.
The KV Cache Walkthrough explains what to check at each step.
A ModelArtifact names the weights that the serving workload will use:
apiVersion: worker.gpustack.ai/v1
kind: ModelArtifact
metadata:
name: qwen-7b
namespace: team-a
spec:
source:
huggingFace:
repository: Qwen/Qwen2.5-7B-Instruct
revision: mainThe deployment below refers to this artifact. The operator resolves its revision before starting the serving Pods. Model Artifact covers PVC and image sources; Model Store Operations covers node delivery.
This deployment uses the artifact and cache binding above. Replace instanceType with one from your
cluster:
apiVersion: worker.gpustack.ai/v1
kind: ModelDeployment
metadata:
name: qwen-chat
namespace: team-a
spec:
model:
name: Qwen/Qwen2.5-7B-Instruct
artifactRef:
name: qwen-7b
engine:
name: vLLM
version: "0.29.0"
kvCache:
poolRef:
name: qwen-cache
roles:
- name: server
replicas: 1
instanceType: gpustack--nvidia-a10g-linux-amd64
resources:
accelerator: 1For a standalone serving example, read Model Deployment. Prefill and Decode shows how to split the serving roles.
Choose an InstanceType from the pool list whose SL capacity has room for a slice. Replace
spec.type with that name and spec.data with your SSH public key. This workspace gives you SSH
access to 20% of one accelerator's memory:
apiVersion: worker.gpustack.ai/v1
kind: InstanceSSHPublicKey
metadata:
name: workspace-key
namespace: team-a
spec:
data: ssh-ed25519 <paste-your-public-key-here>
---
apiVersion: worker.gpustack.ai/v1
kind: Instance
metadata:
name: sliced-demo
namespace: team-a
spec:
type: gpustack--nvidia-a10g-linux-amd64
image: ubuntu:24.04
command:
- "tail"
- "-f"
- "/dev/null"
sshPublicKey:
name: workspace-key
resources:
accelerator: "1"
acceleratorSlicedMemoryPercentage: 20
acceleratorSlicedCoresPercentage: 100
volume:
ephemeral:
capacity: 1GiSave it as instance.yaml, then run kubectl apply -f instance.yaml. Once the Pod is ready,
forward its SSH port in one terminal and connect with the matching private key in another:
kubectl -n team-a wait --for=condition=Ready pod/sliced-demo --timeout=5m
kubectl -n team-a port-forward pod/sliced-demo 2222:22ssh -i ~/.ssh/id_ed25519 -p 2222 root@127.0.0.1The SSH session enters the workload container, where the accelerator and workspace are available. The walkthrough shows the resulting resources. Pods can request accelerators directly; see Accelerator Requests for complete manifests.
All manufacturers support whole-device and shared requests. Logical slicing shares a device through the manufacturer's software facilities. Physical partitioning uses hardware partitions enabled by an administrator; its availability depends on the device and its current mode.
| Manufacturer | Class | Kubernetes resource | Logical slicing | Physical partitioning |
|---|---|---|---|---|
| AMD | GPU | amd.com/gpu |
Yes | — |
| Cambricon | MLU | cambricon.com/mlu |
Yes | — |
| Huawei Ascend | NPU | huawei.com/npu |
Yes | — |
| Hygon | DCU | hygon.com/dcu |
Yes | Yes (MIG) |
| Iluvatar | GPU | iluvatar.com/gpu |
Yes | — |
| MetaX | GPU | metax-tech.com/gpu |
Yes | — |
| Moore Threads | GPU | mthreads.com/gpu |
Yes | — |
| NVIDIA | GPU | nvidia.com/gpu |
Yes | Yes (MIG) |
| T-Head | PPU | alibabacloud.com/ppu |
Yes | Yes (MIG) |
A Pod requests one resource family in one container group. For NVIDIA, the request forms are:
| Request | resources.limits |
|---|---|
| Whole device | nvidia.com/gpu: "1" |
| Shared device | nvidia.com/gpu.shared: "1" |
| Logical slice | nvidia.com/gpu.sliced: "1"nvidia.com/gpu.sliced.memory-percentage: "20"nvidia.com/gpu.sliced.cores-percentage: "40" |
| MIG partition | nvidia.com/gpu.partitioned: "1"nvidia.com/gpu.partitioned.mig-3g.40gb: "1" |
See Accelerator Requests for the exact keys and rules, and Device Discovery for each manufacturer's slicing behavior.
Copyright (c) 2026 The GPUStack Authors. Licensed under the Apache License 2.0.
