What problem are you facing?
The fleet scheduler gates capacity on nodes, not on individual GPUs. Every pod that claims a device charges a whole node, regardless of how many of the node's devices it actually requests. A pod that claims 1 GPU of a 4-GPU node consumes the whole node in the scheduler's accounting.
The accounting lives in _member_cost (scheduling.py:348-369): a member that resolves any claim: DRA device costs pods × copies nodes, where each pod is one node. The module docstring spells out the model (scheduling.py:50-60): capacity is "could this cluster plausibly host this replica," deliberately coarse, leaving device-level contention to DRA admission on the workload cluster.
This is safe — charging a whole node per pod is strictly conservative, so the scheduler can only ever under-count a pool's capacity, never overcommit it. But it's wasteful. A deployment of small single-GPU engines on 8-GPU nodes pins one node per engine and strands seven GPUs per node. GPUs are the most expensive thing in the fleet, and paying for nodes you've only partly used is the most expensive way to be conservative.
The waste is concentrated exactly where it hurts: sub-node claims (a 1-GPU engine, a throughput deployment of many single-GPU copies, a disaggregated prefill phase of single-GPU pods). A deployment whose pods each claim a whole node loses nothing.
How could Modelplane help solve your problem?
The scheduler should be able to place more than one pod on a node when the node has the devices for it, without giving up the two properties that make it tractable today: it stays a pure function of observed state, and it never overcommits a cluster.
The constraint is that the fleet scheduler is a predictor, not an authority. DRA admission on the workload cluster decides what actually runs. So the scheduler doesn't need to bin-pack perfectly — it needs to predict well enough that wrong guesses are rare and self-healing. That's a much smaller problem than reimplementing the Kubernetes scheduler from a fleet-wide vantage point, which we explicitly don't want to do.
The challenges
Three things make device-granular accounting harder than the current node-granular model:
-
Fragmentation is invisible from the control plane. "This pool has 6 free GPUs" doesn't tell you whether a pod claiming 2 GPUs fits: the 6 could be one free GPU on each of 6 nodes. The scheduler can see a pool's total devices (nodes × count) and subtract what observed replicas claim, but it can't see how free devices are distributed across nodes — and that distribution is exactly what decides whether a multi-GPU pod fits.
-
Tracking per-device usage is untenable at fleet scale . Carrying it means an InferenceCluster's status enumerating thousands of nodes' free-device counts: a status object that's huge and churns on every pod admission, and a scheduler that processes thousands of node entries per pool per reconcile across the whole fleet.
-
Self-heal is at odds with the scheduler being a pure function of its inputs. If we place optimistically and a pod can't admit, the obvious recovery (re-place the replica next reconcile) loops: a pure function of the same inputs makes the same placement again, back onto the cluster that just rejected it.
A relatively simple direction: optimistic placement with fast self-heal
Schedule a little more optimistically using only what the control plane already observes, and lean on DRA plus a feedback loop to catch the mistakes:
-
Count devices, not nodes, against a pool's total. A pool has nodes × count devices; subtract what observed replicas claim. Place a replica when the pool appears to have enough free devices. This needs no new status API and no in-cluster component — it's an accounting change to the existing pure function, derived from the same observed state the scheduler already reads. It will sometimes mispredict (fragmentation), which is what the next two points handle.
-
Detect an unschedulable replica quickly. When a placed replica's pods can't admit (DRA rejects the claim, the pod sits Pending/Unschedulable), surface that back into the scheduler's observed state.
-
Record the failure as a scheduler input so we don't retry the same cluster forever. Persist the failed (replica, cluster) as observed state (e.g. a status field on the replica), and have the scheduler exclude it on the next fill, re-placing the replica elsewhere. This keeps the scheduler a pure function.
This trades GPU waste for some replica churn and scheduling latency on a mis-prediction.
More involved: publish a per-node free-device histogram
Instead of scheduling on dumb pool-total device counts and recovering from mispredictions, the control plane could carry just enough of the per-node distribution to predict fit directly: a free-devices-per-node histogram per pool (how many nodes have N free devices), bounded by devices-per-node rather than node count, so it stays compact regardless of pool size. The scheduler could then tell whether a pod claiming k devices fits without guessing, eliminating most mispredictions. (Not all though: it'd probably need to reschedule, just rarely.)
The catch is that the histogram needs something in-cluster to compute and publish it: the reporting component the optimistic approach avoids.
What problem are you facing?
The fleet scheduler gates capacity on nodes, not on individual GPUs. Every pod that claims a device charges a whole node, regardless of how many of the node's devices it actually requests. A pod that claims 1 GPU of a 4-GPU node consumes the whole node in the scheduler's accounting.
The accounting lives in
_member_cost(scheduling.py:348-369): a member that resolves anyclaim: DRAdevice costspods × copiesnodes, where each pod is one node. The module docstring spells out the model (scheduling.py:50-60): capacity is "could this cluster plausibly host this replica," deliberately coarse, leaving device-level contention to DRA admission on the workload cluster.This is safe — charging a whole node per pod is strictly conservative, so the scheduler can only ever under-count a pool's capacity, never overcommit it. But it's wasteful. A deployment of small single-GPU engines on 8-GPU nodes pins one node per engine and strands seven GPUs per node. GPUs are the most expensive thing in the fleet, and paying for nodes you've only partly used is the most expensive way to be conservative.
The waste is concentrated exactly where it hurts: sub-node claims (a 1-GPU engine, a throughput deployment of many single-GPU copies, a disaggregated prefill phase of single-GPU pods). A deployment whose pods each claim a whole node loses nothing.
How could Modelplane help solve your problem?
The scheduler should be able to place more than one pod on a node when the node has the devices for it, without giving up the two properties that make it tractable today: it stays a pure function of observed state, and it never overcommits a cluster.
The constraint is that the fleet scheduler is a predictor, not an authority. DRA admission on the workload cluster decides what actually runs. So the scheduler doesn't need to bin-pack perfectly — it needs to predict well enough that wrong guesses are rare and self-healing. That's a much smaller problem than reimplementing the Kubernetes scheduler from a fleet-wide vantage point, which we explicitly don't want to do.
The challenges
Three things make device-granular accounting harder than the current node-granular model:
Fragmentation is invisible from the control plane. "This pool has 6 free GPUs" doesn't tell you whether a pod claiming 2 GPUs fits: the 6 could be one free GPU on each of 6 nodes. The scheduler can see a pool's total devices (
nodes × count) and subtract what observed replicas claim, but it can't see how free devices are distributed across nodes — and that distribution is exactly what decides whether a multi-GPU pod fits.Tracking per-device usage is untenable at fleet scale . Carrying it means an InferenceCluster's status enumerating thousands of nodes' free-device counts: a status object that's huge and churns on every pod admission, and a scheduler that processes thousands of node entries per pool per reconcile across the whole fleet.
Self-heal is at odds with the scheduler being a pure function of its inputs. If we place optimistically and a pod can't admit, the obvious recovery (re-place the replica next reconcile) loops: a pure function of the same inputs makes the same placement again, back onto the cluster that just rejected it.
A relatively simple direction: optimistic placement with fast self-heal
Schedule a little more optimistically using only what the control plane already observes, and lean on DRA plus a feedback loop to catch the mistakes:
Count devices, not nodes, against a pool's total. A pool has
nodes × countdevices; subtract what observed replicas claim. Place a replica when the pool appears to have enough free devices. This needs no new status API and no in-cluster component — it's an accounting change to the existing pure function, derived from the same observed state the scheduler already reads. It will sometimes mispredict (fragmentation), which is what the next two points handle.Detect an unschedulable replica quickly. When a placed replica's pods can't admit (DRA rejects the claim, the pod sits
Pending/Unschedulable), surface that back into the scheduler's observed state.Record the failure as a scheduler input so we don't retry the same cluster forever. Persist the failed
(replica, cluster)as observed state (e.g. a status field on the replica), and have the scheduler exclude it on the next fill, re-placing the replica elsewhere. This keeps the scheduler a pure function.This trades GPU waste for some replica churn and scheduling latency on a mis-prediction.
More involved: publish a per-node free-device histogram
Instead of scheduling on dumb pool-total device counts and recovering from mispredictions, the control plane could carry just enough of the per-node distribution to predict fit directly: a free-devices-per-node histogram per pool (how many nodes have N free devices), bounded by devices-per-node rather than node count, so it stays compact regardless of pool size. The scheduler could then tell whether a pod claiming
kdevices fits without guessing, eliminating most mispredictions. (Not all though: it'd probably need to reschedule, just rarely.)The catch is that the histogram needs something in-cluster to compute and publish it: the reporting component the optimistic approach avoids.