Design: ModelCache sources and the mount contract - #505
Closed
dennis-upbound wants to merge 1 commit into
Closed
dennis-upbound wants to merge 1 commit into
dennis-upbound wants to merge 1 commit into
Conversation
A ModelCache reads from HuggingFace only, stages to a volume on each matched cluster, and publishes a phase. A consumer calculates the name of that volume, and two functions hold a copy of a calculation that one source makes true. This design adds an OCI source and an Existing source, and a mount contract on each cluster entry. The contract holds the volumes, the volume mounts and the environment that a consumer adds to a pod, in the Kubernetes types of those names. The consumer copies it and never learns the source. Section 4.4 gives the faults the current design permits, from the benchmark work: modelplaneai#495, where an engine starts while the cache stages and reads a 464 GB model a second time; modelplaneai#494, where a replaced node strands the staging volume; and modelplaneai#498, which asks for the cause on the control plane. Section 6 says what model delivery should measure, because the performance claim this design rests on has no series behind it today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Dennis Ramdass <dennis@upbound.io>
Collaborator
Author
|
Closing until the design is reviewed in upbound/inference#7, which holds this document and the Upbound Inference one together. The split between them is the thing worth a reviewer's judgement, and it is only visible with both side by side. This reopens, or comes back as a fresh PR, once that review settles. The branch stays at Nothing here is blocked: #504 (the telemetry guide correction) and the research on #476 are independent. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description of your changes
A
ModelCachereads from HuggingFace and nowhere else, stages what it reads onto a ReadWriteMany volume on every matched cluster, and publishes a phase. A consumer works out how to read what was staged by deriving the volume's name — a derivationcompose-model-replicaandcompose-model-cacheeach hold a copy of, with a comment telling both to change together, and which is true for one source.This proposes two more sources and one status field.
source: OCIreads an artifact where it already is. The kubelet pulls the reference at pod start, mounts it read-only, and serves the next pod on that node from what the first pulled. Modelplane composes nothing per cluster for it.source: Existingmounts a claim the caller populated, which the same machinery makes nearly free.status.clusters[].mountsays how to read the artifact on that cluster — volumes, mounts and environment, in corev1's own types. A consumer copies the three lists without reading them. What does the mounting depends on what the artifact is: Kubernetes mounts a container image as an image volume, and a model artifact, which containerd will not mount, is read by a CSI driver. Modelplane resolves the reference once to tell them apart, so a reference it cannot serve becomes aFailedcluster rather than an empty directory at pod start.What changed since the last look
The design now carries the evidence from Pablo's benchmark work, which arrived after it was drafted and sharpened the argument from consistency to correctness:
compose-model-deploymentreadsstatus.cache.storageClassName, which says a cluster has storage, never that the artifact is on it. §5.6 is what stops that.Hydrating. GKE ModelCache PVC never binds: Filestore CSI driver addon not enabled on the cluster #221 is the same fault on GKE. §4.4 treats it as a class: each source that stages to a volume has it; a source that mounts from a registry does not.§6 is new, and is the part I'd most like a view on. This design argues OCI delivery wins on warm pulls — 11.7s against 40.7 minutes — and Modelplane publishes no series that measures either number. §6.1 proposes the two that make the claim falsifiable, with a
cachedlabel, because warm and cold differ by two orders of magnitude and one histogram holding both describes neither. §6.2 proposes four for cache state and marks them blocked by #476, where I've posted the research on how that group should be built.On scope
This is the design alone. The implementation on
dennis/oci-source-pocpredates it and stages throughmodctlonto the cache volume — the approach this design replaces — so it would only confuse a review of the proposal. It needs rewriting against whatever lands here.The doc supersedes #362 and #210, and revises
design/modelcache.md, which will contradict it until that page is updated.I have:
nix flake check(or./nix.sh flake check) and made sure it passes.Added or updated tests covering any composition function changes.Design only; no function changes.git commit -s.🤖 Generated with Claude Code