Repository navigation
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
This comment was marked as resolved.
This comment was marked as resolved.
|
Another thing to keep in mind is that we can actually make good use of git, we just don't currently. Repacking the tarball cache yields very good results in terms of delta compressing multiple similar blobs between nixpkgs versions or such. |
|
Also the issue with probing multiple index files is solved with MDIX (and newer git has incremental mdix too https://github.blog/open-source/git/highlights-from-git-2-55/). |
|
note: nix-tarmac seems to have had some success with LMDB |
…helpers The helpers for parsing and dumping Git blobs and trees are not user-facing. The git-hashing feature is still checked where the Git content-address method is selected. Assisted-by: Claude Fable 5.1 <noreply@anthropic.com>
Tarballs are now unpacked into ~/.cache/nix/tarball-cache-v3.sqlite instead of a bare Git repository written via libgit2. Files and directories are still identified by their Git blob and tree hashes, so tree hashes and accessor fingerprints are unchanged. Blobs are compressed individually using zstd. This avoids the accumulation of packfiles (one per import), which made object lookups slow, and libgit2's concurrency issues. There is no garbage collection yet, and every file is stored as a single blob. Assisted-by: Claude Fable 5.1 <noreply@anthropic.com>
edolstra
force-pushed
the
tarball-cache
branch
from
October 6, 2026 09:13
cc1cfbb to
f436270
Compare
Blobs are now stored as a sequence of chunks of at most 1 MiB, each compressed independently. Files larger than a chunk are spilled to a temporary file during import. So memory use no longer scales with file size, and blobs larger than SQLite's blob size limit (1 GB) can be stored. Importing a tarball with a 300 MB file now takes 184 MB of RSS, compared to 1.2 GB with the Git-based cache. The cost for Nixpkgs is about 5% on cold imports and 4 MB of database size, since every blob now has two rows. Assisted-by: Claude Fable 5.1 <noreply@anthropic.com>
edolstra
force-pushed
the
tarball-cache
branch
from
October 6, 2026 10:30
7ebf90e to
81c78da
Compare
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The tarball cache (
~/.cache/nix/tarball-cache-v2) is a bare Git repository written through libgit2. This has two problems:objects/pack/pack-*files, and object lookups have to probe all of their indices.This PR replaces it with a SQLite database (
~/.cache/nix/tarball-cache-v3.sqlite).Context
Files and directories are still identified by their Git blob and tree hashes (SHA-1), so tree hashes, accessor fingerprints (
git:<tree>) and NAR hashes are unchanged. Importing a Nixpkgs tarball with the old and new implementation yields the same tree hash and NAR hash.Schema:
Implementation notes:
insert or ignoretransactions of about 16 MiB. Blobs that are already in the cache are skipped, so importing a new revision of a repository only writes the files that changed.Treesrow is only committed after all of its children, so the existence of a tree implies that everything reachable from it is present. Interrupted imports only leave unreferenced blobs behind.nix/util/compression.hhmade cold imports about 60% slower, since it sets up a new context for every blob.tarball-cache-v2directory is left untouched. Existing fetcher cache entries refer to trees that are not in the new cache, so those inputs are simply refetched.GitRepo::getFileSystemObjectSink()) is kept.Limitations of this first version:
The commits are best reviewed separately:
git: Don't require the git-hashing experimental feature in low-level helpers: the blob/tree dumping helpers in libutil are reused to compute Git hashes. The experimental feature is still checked where the Git content-address method is selected.Replace the Git-based tarball cache by a SQLite database: the newTarballCache(src/libfetchers/tarball-cache.{hh,cc}) and the switch of the tarball and GitHub fetchers to it.SQLite: Only apply the ZFS -shm workaround on the first open of a database: fixes a SIGBUS found while benchmarking. The workaround opened and closeddb.sqlite-shmon everySQLiteconstruction; closing a file descriptor drops all POSIX locks the process holds on that file, including SQLite's lock on behalf of other connections to the same database. Another process could then truncate the file. The tarball cache is the first user that opens several connections to one database per process.Benchmarks with a Nixpkgs tarball (93k entries, 53 MB compressed, 191 MB of file contents) on a 24-core machine with ZFS, fresh cache per cold run:
nix flake metadata)nix search … fizzbuzz --no-eval-cache--refresh)Note that the Git numbers are for a cache containing a single packfile, so they don't show the slowdown from accumulated packfiles.
🤖 Generated with Claude Code