Skip to content

Design: ModelCache sources and the mount contract - #505

Closed
dennis-upbound wants to merge 1 commit into
modelplaneai:mainfrom
dennis-upbound:dennis/design-modelcache-sources
Closed

dennis-upbound wants to merge 1 commit into
modelplaneai:mainfrom
dennis-upbound:dennis/design-modelcache-sources

Conversation

@dennis-upbound

Copy link
Copy Markdown
Collaborator

Description of your changes

A ModelCache reads from HuggingFace and nowhere else, stages what it reads onto a ReadWriteMany volume on every matched cluster, and publishes a phase. A consumer works out how to read what was staged by deriving the volume's name — a derivation compose-model-replica and compose-model-cache each hold a copy of, with a comment telling both to change together, and which is true for one source.

This proposes two more sources and one status field.

source: OCI reads an artifact where it already is. The kubelet pulls the reference at pod start, mounts it read-only, and serves the next pod on that node from what the first pulled. Modelplane composes nothing per cluster for it. source: Existing mounts a claim the caller populated, which the same machinery makes nearly free.

status.clusters[].mount says how to read the artifact on that cluster — volumes, mounts and environment, in corev1's own types. A consumer copies the three lists without reading them. What does the mounting depends on what the artifact is: Kubernetes mounts a container image as an image volume, and a model artifact, which containerd will not mount, is read by a CSI driver. Modelplane resolves the reference once to tell them apart, so a reference it cannot serve becomes a Failed cluster rather than an empty directory at pod start.

What changed since the last look

The design now carries the evidence from Pablo's benchmark work, which arrived after it was drafted and sharpened the argument from consistency to correctness:

§6 is new, and is the part I'd most like a view on. This design argues OCI delivery wins on warm pulls — 11.7s against 40.7 minutes — and Modelplane publishes no series that measures either number. §6.1 proposes the two that make the claim falsifiable, with a cached label, because warm and cold differ by two orders of magnitude and one histogram holding both describes neither. §6.2 proposes four for cache state and marks them blocked by #476, where I've posted the research on how that group should be built.

On scope

This is the design alone. The implementation on dennis/oci-source-poc predates it and stages through modctl onto the cache volume — the approach this design replaces — so it would only confuse a review of the proposal. It needs rewriting against whatever lands here.

The doc supersedes #362 and #210, and revises design/modelcache.md, which will contradict it until that page is updated.

I have:

  • Read and followed Modelplane's contribution process.
  • Run nix flake check (or ./nix.sh flake check) and made sure it passes.
  • Added or updated tests covering any composition function changes. Design only; no function changes.
  • Signed off every commit with git commit -s.

🤖 Generated with Claude Code

A ModelCache reads from HuggingFace only, stages to a volume on each matched
cluster, and publishes a phase. A consumer calculates the name of that volume,
and two functions hold a copy of a calculation that one source makes true.

This design adds an OCI source and an Existing source, and a mount contract on
each cluster entry. The contract holds the volumes, the volume mounts and the
environment that a consumer adds to a pod, in the Kubernetes types of those
names. The consumer copies it and never learns the source.

Section 4.4 gives the faults the current design permits, from the benchmark
work: modelplaneai#495, where an engine starts while the cache stages and reads a 464 GB
model a second time; modelplaneai#494, where a replaced node strands the staging volume;
and modelplaneai#498, which asks for the cause on the control plane. Section 6 says what
model delivery should measure, because the performance claim this design rests
on has no series behind it today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Dennis Ramdass <dennis@upbound.io>
@dennis-upbound

Copy link
Copy Markdown
Collaborator Author

Closing until the design is reviewed in upbound/inference#7, which holds this document and the Upbound Inference one together. The split between them is the thing worth a reviewer's judgement, and it is only visible with both side by side.

This reopens, or comes back as a fresh PR, once that review settles. The branch stays at dennis/design-modelcache-sources.

Nothing here is blocked: #504 (the telemetry guide correction) and the research on #476 are independent.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant