Skip to content

Replace the Git-based tarball cache by a SQLite database - #646

Draft
edolstra wants to merge 3 commits into
mainfrom
tarball-cache
Draft

edolstra wants to merge 3 commits into
mainfrom
tarball-cache

Conversation

@edolstra

Copy link
Copy Markdown
Collaborator

Motivation

The tarball cache (~/.cache/nix/tarball-cache-v2) is a bare Git repository written through libgit2. This has two problems:

  • Every import creates a new packfile, so the cache accumulates thousands of objects/pack/pack-* files, and object lookups have to probe all of their indices.
  • libgit2 has concurrency issues.

This PR replaces it with a SQLite database (~/.cache/nix/tarball-cache-v3.sqlite).

Context

Files and directories are still identified by their Git blob and tree hashes (SHA-1), so tree hashes, accessor fingerprints (git:<tree>) and NAR hashes are unchanged. Importing a Nixpkgs tarball with the old and new implementation yields the same tree hash and NAR hash.

Schema:

create table Blobs (
    oid         blob primary key not null,  -- Git blob hash
    size        integer not null,           -- uncompressed size
    compression integer not null,           -- 0 = none, 1 = zstd
    data        blob not null
);
create table Trees (
    oid blob primary key not null
) without rowid;
create table TreeEntries (
    tree  blob not null,
    name  text not null,
    mode  integer not null,
    child blob not null,
    primary key (tree, name)
) without rowid;

Implementation notes:

  • The database uses WAL mode. Hashing and compression happen in worker threads outside of any transaction; blobs are written in batched insert or ignore transactions of about 16 MiB. Blobs that are already in the cache are skipped, so importing a new revision of a repository only writes the files that changed.
  • A Trees row is only committed after all of its children, so the existence of a tree implies that everything reachable from it is present. Interrupted imports only leave unreferenced blobs behind.
  • Blobs are compressed individually with zstd (level 3), calling libzstd directly. Going through nix/util/compression.hh made cold imports about 60% slower, since it sets up a new context for every blob.
  • The old tarball-cache-v2 directory is left untouched. Existing fetcher cache entries refer to trees that are not in the new cache, so those inputs are simply refetched.
  • The libgit2-based sink (GitRepo::getFileSystemObjectSink()) is kept.

Limitations of this first version:

  • There is no garbage collection (as before).
  • Every file is stored as a single blob, so a file is held in memory while it is imported or read, and cannot exceed SQLite's blob size limit (1 GB by default).
  • There are no unit tests for the new cache yet.

The commits are best reviewed separately:

  1. git: Don't require the git-hashing experimental feature in low-level helpers: the blob/tree dumping helpers in libutil are reused to compute Git hashes. The experimental feature is still checked where the Git content-address method is selected.
  2. Replace the Git-based tarball cache by a SQLite database: the new TarballCache (src/libfetchers/tarball-cache.{hh,cc}) and the switch of the tarball and GitHub fetchers to it.
  3. SQLite: Only apply the ZFS -shm workaround on the first open of a database: fixes a SIGBUS found while benchmarking. The workaround opened and closed db.sqlite-shm on every SQLite construction; closing a file descriptor drops all POSIX locks the process holds on that file, including SQLite's lock on behalf of other connections to the same database. Another process could then truncate the file. The tarball cache is the first user that opens several connections to one database per process.

Benchmarks with a Nixpkgs tarball (93k entries, 53 MB compressed, 191 MB of file contents) on a 24-core machine with ZFS, fresh cache per cold run:

Test Git SQLite
Cold import (nix flake metadata) 2.81–3.06 s 2.10–2.21 s
nix search … fizzbuzz --no-eval-cache 4.31–4.59 s 4.01–4.30 s
Re-import (--refresh) 0.81–0.83 s 0.69–0.70 s
4 concurrent imports into one cache 3.26–3.64 s 2.32–2.44 s
Cache size 68 MB 82 MB (plus WAL)

Note that the Git numbers are for a cache containing a single packfile, so they don't show the slowdown from accumulated packfiles.

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

@github-actions
github-actions Bot temporarily deployed to pull request September 30, 2026 14:56 Inactive
@cole-h

This comment was marked as resolved.

@xokdvium

xokdvium commented Oct 3, 2026

Copy link
Copy Markdown

Another thing to keep in mind is that we can actually make good use of git, we just don't currently. Repacking the tarball cache yields very good results in terms of delta compressing multiple similar blobs between nixpkgs versions or such.

@xokdvium

xokdvium commented Oct 3, 2026

Copy link
Copy Markdown

Also the issue with probing multiple index files is solved with MDIX (and newer git has incremental mdix too https://github.blog/open-source/git/highlights-from-git-2-55/).

@tomberek

tomberek commented Oct 5, 2026

Copy link
Copy Markdown

note: nix-tarmac seems to have had some success with LMDB

…helpers

The helpers for parsing and dumping Git blobs and trees are not
user-facing. The git-hashing feature is still checked where the Git
content-address method is selected.

Assisted-by: Claude Fable 5.1 <noreply@anthropic.com>
Tarballs are now unpacked into ~/.cache/nix/tarball-cache-v3.sqlite
instead of a bare Git repository written via libgit2. Files and
directories are still identified by their Git blob and tree hashes, so
tree hashes and accessor fingerprints are unchanged. Blobs are
compressed individually using zstd.

This avoids the accumulation of packfiles (one per import), which made
object lookups slow, and libgit2's concurrency issues.

There is no garbage collection yet, and every file is stored as a
single blob.

Assisted-by: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions
github-actions Bot temporarily deployed to pull request October 6, 2026 09:22 Inactive
Blobs are now stored as a sequence of chunks of at most 1 MiB, each
compressed independently. Files larger than a chunk are spilled to a
temporary file during import. So memory use no longer scales with
file size, and blobs larger than SQLite's blob size limit (1 GB) can
be stored.

Importing a tarball with a 300 MB file now takes 184 MB of RSS,
compared to 1.2 GB with the Git-based cache. The cost for Nixpkgs is
about 5% on cold imports and 4 MB of database size, since every blob
now has two rows.

Assisted-by: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions
github-actions Bot temporarily deployed to pull request October 6, 2026 10:06 Inactive
@github-actions
github-actions Bot temporarily deployed to pull request October 6, 2026 10:32 Inactive

This branch was previously deployed

1 inactive deployment
pull request — 81c78da4 Deployed Oct 6, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants