Skip to content

nix nario: Support binary diffs and compression - #674

Draft
edolstra wants to merge 12 commits into
mainfrom
eelcodolstra/nix-492
Draft

edolstra wants to merge 12 commits into
mainfrom
eelcodolstra/nix-492

Conversation

@edolstra

@edolstra edolstra commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

This makes it cheap to deploy closure updates using narios. nix nario export --base <installable> assumes the destination already has the closure of <installable>. Paths in that closure are left out, and other paths are exported as binary diffs against a path in the base closure where possible. This is similar to the bsdiff patches Nix had in the manifest era. In addition, --compression <method> (with --compression-level) compresses the full NARs in a nario. Diffs and compressed NARs are computed in parallel.

For example, the update from one NixOS system generation to the next is a ~2.6 MB nario, compared with a 19.7 GB full export (6.56 GB with --compression zstd).

Diff algorithms

--diff-algorithm selects how diffs are computed:

  • zstd (default): zstd level 19, using the base NAR as a prefix dictionary.
  • hdiffpatch (experimental): HDiffPatch, linked as a library, with zstd-compressed patches.

Comparison: exporting 35 generations of a NixOS system profile (a full export of the first generation, plus a diff of each later generation against its predecessor, all with --compression zstd) on a 12-core/24-thread machine. bsdiff was tested by calling the external bsdiff program, and isn't part of this PR.

Algorithm Total size Diffs only Elapsed CPU Peak RSS
zstd 8.06 GB 1.50 GB 1957 s 17562 s 42.6 GB
bsdiff 7.92 GB 1.36 GB (−9%) 5859 s 37758 s 14.5 GB
hdiffpatch 7.82 GB 1.26 GB (−16%) 1164 s 4112 s 16.3 GB

"Diffs only" excludes the 6.56 GB full export, which is the same in all three runs. HDiffPatch gives the smallest diffs, is the fastest, and uses far less memory than zstd. zstd's peak memory comes from many large level-19 diffs running at once during mass rebuilds.

Context

New NARIO2 entry types, so there's no format bump. Older Nix rejects narios that use them.

  • 2: binary diff against a base path, with the diff algorithm and the base's NAR hash.
  • 3: a path that must already be present, with its path info so the NAR hash can be checked.
  • 4: compressed NAR.

Base paths are chosen by matching store path names without the version. Reading narios now goes through parseNario() with a visitor, which is used by both importPaths() and nix nario list. See the commit messages for details.

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

`nix nario export --format 2 --base <installable>...` assumes that the
destination already has the closure of the given installables (the
"base closure"). Paths in the base closure are not exported, and other
paths are exported as binary diffs against a path in the base closure
where possible. This makes it cheap to deploy closure updates, similar
to the bsdiff-based patches that Nix supported in the manifest era
(removed in 8679672).

Diffs are computed using zstd's "patch-from" mode (i.e. the base NAR
is used as a prefix dictionary, with long-distance matching and a
window covering both NARs), using the libzstd we already link
against. On import, the patched NAR is streamed into the store, so
only the base NAR and the patch are held in memory.

The base for a path is selected by matching the store path name
without the version (and with the output name kept separate, so
`openssl-*-dev` only matches `-dev`), skipping bases that differ in
NAR size by more than a factor of 3. Among candidates, the one with
the longest common version prefix wins, then the one closest in
size. A diff is only used if it's smaller than 0.9 times the
standalone zstd compression of the target NAR. NARs larger than 1
GiB are not diffed. Diffs are computed in parallel.

Format: no version bump. NARIO2 entries are tagged; besides the
existing tag 1 (full NAR), there are now:

* Tag 2: diff entry: the target's ValidPathInfo (so narHash/narSize
  and signatures describe the reconstructed NAR), the base path, the
  base's NAR hash and the zstd patch.

* Tag 3: the ValidPathInfo of a path that is expected to already be
  present at the destination. Import fails if it's missing or has a
  different NAR hash.

Import also fails with a clear error if a diff base is missing or
doesn't have the expected NAR hash. Older versions of Nix will reject
narios containing these entries.

Reading narios is now done by `parseNario()` with a `NarioVisitor`,
which is used by both `importPaths()` and `nix nario list`. The
latter shows diff and present entries (also in `--json`).

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
At high compression levels, the single-threaded zstd compressor makes
poor use of the prefix (base). For example, for two byte-identical 65
MB NARs (noto-fonts-cjk-sans), it produced a 13 MB patch at level 19,
versus 6 KB with the multi-threaded compressor (which `zstd
--patch-from` uses by default, since it defaults to 1 worker). Patch
sizes for other NARs also improved somewhat (e.g. nodejs-slim
24.15.0 -> 24.16.0: 5.6 MB -> 4.8 MB), though a few got slightly
bigger.

So set ZSTD_c_nbWorkers to 1. More workers don't improve patch size,
and we already compute diffs in parallel.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
Binary diffs in narios are already zstd-compressed, while full NARs
are not, so compressing a nario as a whole (as users would previously
do) is wasteful. Instead, `nix nario export --compression <method>`
now compresses the NARs of full entries.

These are stored in a new NARIO2 entry type (tag 4): like tag 1, it
has the ValidPathInfo (describing the uncompressed NAR), followed by
the compression method (e.g. `zstd`) and the compressed NAR as a
length-prefixed string. Since the length must be written first, the
exporter buffers each compressed NAR in memory. The importer streams
the compressed data through a decompressor, so it uses constant
memory. Any compression method supported by Nix can be used; zstd
uses the multi-threaded compressor.

The default is still no compression, so that the output of `nix
nario export` without `--base` remains readable by older versions of
Nix.

`NarioVisitor::fullPath()` now gets the compression method and size
of compressed NARs, which `nix nario list` shows (and includes in its
JSON output as `compression`).

For example, for a NixOS system closure, `--compression zstd` reduced
the nario from 19.7 GB to 7.15 GB (in 112 s, with a peak RSS of 1.6
GB).

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
@edolstra
edolstra force-pushed the eelcodolstra/nix-492 branch from 955c078 to f8f470b Compare October 7, 2026 20:50
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

@github-actions
github-actions Bot temporarily deployed to pull request October 7, 2026 21:00 Inactive
Diff entries (tag 2) now contain the name of the diff algorithm,
directly after the ValidPathInfo (like the compression method in tag 4
entries). This allows other diff algorithms to be added in the future.
The current algorithm (zstd with the base NAR as a prefix dictionary)
is called `zstd`. Unknown algorithms are rejected when parsing.

`nix nario list` shows the algorithm, and includes it in the JSON
output as `diff.algorithm`.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
This prepares for adding other methods for selecting the base path
against which to diff a path (e.g. based on content similarity). The
only method for now is `by-name` (the default), which is the existing
heuristic of matching store path names without the version.

selectDiffBases() now applies the constraints that are independent of
the selection method (the NAR size limit, and base + target fitting in
the zstd window) and dispatches to selectDiffBasesByName() for the
actual matching.

Also, exportPaths() now takes a NarioExportOptions struct rather than
an ever-growing list of arguments.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
This sets the compression level for the method given by
`--compression`. The default for zstd is now 9 rather than zstd's own
default of 3, since it's still fast but compresses significantly
better. Other methods keep their own defaults.

The baseline against which binary diffs are compared (a diff is only
used if it's sufficiently smaller than the compressed NAR) now uses
the same zstd level that full NARs are compressed with, or 9 if they
aren't compressed with zstd. Previously it was always level 3.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
Since the size of a compressed NAR must be written before its
contents, we buffered the entire compressed NAR in memory, which uses
unbounded memory for large NARs.

This adds `SpillingStringSink`, a sink that accumulates data in
memory up to a limit, and spills to an anonymous temporary file
beyond that. `getSource()` returns a source that reads the data back
from either. `nix nario export` now uses it with a 32 MiB limit. The
nario format is unchanged.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions
github-actions Bot temporarily deployed to pull request October 8, 2026 11:13 Inactive
Previously, binary diffs were all computed up front in a thread pool
before anything was written, and full NARs were compressed one at a
time on the writing thread (relying only on zstd's internal
multithreading, which doesn't help much for small NARs).

This adds `processOrdered()`, which runs a "produce" function for a
range of items in a thread pool and calls a "consume" function on the
results strictly in order. At most `window` items (by default twice
the number of threads) are in progress or waiting to be consumed, so
memory and temporary disk usage stay bounded.

`exportPaths()` now uses this to compute diffs and compressed NARs in
parallel, falling back to compression if a diff isn't worth it, and
writes entries as soon as they're ready. Only paths that need
expensive work go through the thread pool; cheap entries (paths in
the base closure and uncompressed full NARs) are written in between.
Otherwise, a diff export where most paths are in the base closure
would get little parallelism from the window.

The output is byte-identical to before. Exporting a 6.6 GB NixOS
system closure with `--compression zstd` went from 240s to 101s on a
24-core machine. Diff exports take about the same time as before,
since they already computed diffs in parallel.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions
github-actions Bot temporarily deployed to pull request October 8, 2026 18:16 Inactive
processOrdered() previously allowed at most `window` (by default
twice the number of threads) items to be in progress or waiting to
be consumed. This caused head-of-line blocking: when a slow item
(e.g. a binary diff of a large NAR) is the next one to be consumed,
the other threads quickly finish the remaining items in the window
and then sit idle. In an export of a mass rebuild of a NixOS system
(2314 diffs, of which ~15 are NARs over 250 MB), the large diffs
were effectively serialized, so the export took 1181s instead of
437s before the parallelism changes.

Instead, processOrdered() now stops starting new items only while
the total size of the buffered results (as returned by a `getSize`
callback) exceeds `maxBuffered`. Since diffs are small, this allows
the thread pool to work far ahead of a slow item, while still
bounding the amount of compressed NAR data that piles up.
`nix nario export` uses a limit of 1 GiB.

Up to twice as many items as threads are active at any time, since
ThreadPool only starts a new thread if there are more pending items
than threads.

Timings on a 12-core/24-thread machine (output is identical):

  mass rebuild diff:  417s before parallelism, 1181s with window, 424s now
  full system export: 226s before parallelism,  105s with window,  67s now

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
This adds `--diff-algorithm {zstd,hdiffpatch}`. With `hdiffpatch`,
binary diffs are computed using the external `hdiffz` program from
HDiffPatch (https://github.com/sisong/HDiffPatch) and applied using
`hpatchz`. These must be in `PATH` on export and import. The nario
format is unchanged, since diff entries already record the diff
algorithm.

hdiffz is run with the options used by HDiffPatch's own zstd
benchmark (`-m-6 -SD`), but with zstd level 19 to match our zstd
diffs, single-threaded since diffs are already computed in parallel,
and without its patch self-check since the NAR hash is verified on
import.

HDiffPatch isn't in Nixpkgs, so this adds a package for it
(packaging/hdiffpatch.nix), which is used by the functional tests and
exposed by the nix-make flake for experimentation.

This is an experiment to compare against zstd diffs. (A similar
experiment with bsdiff produced slightly smaller diffs than zstd but
used much more memory.)

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
Instead of running the external `hdiffz` and `hpatchz` programs
(which required them to be in `PATH` and went through temporary
files), link HDiffPatch into libnixutil and call it directly via the
new `makeHdiffPatch()` and `applyHdiffPatch()`, analogous to
`makeZstdPatch()` and `applyZstdPatch()`.

Upstream doesn't install a library, so the HDiffPatch package now
installs the static library built by its Makefile (single-threaded
and with -fPIC), its headers, and a pkg-config file that includes the
configuration macros that affect the headers.

The settings are the same as before (match score 6, fast block
matching, zstd level 19 with a 2^24 dictionary, single-threaded).
`applyHdiffPatch()` streams the output to a sink and validates the
patch header (base size, patch length and memory requirements)
before applying it. Since HDiffPatch overwrites its inputs during
matching, `makeDiff()` now moves the NARs into `makeHdiffPatch()`
rather than copying them.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>
HDiffPatch's diff code and its zstd compression plugin print progress
messages (e.g. "create suffix string ...") to stdout. Since `nix nario
export` writes the nario to stdout, these messages ended up in the
nario. Because stdout is buffered, they were appended after the end
marker when the process exited, so the narios could still be read,
but they were larger than necessary (and not a multiple of 8 bytes).

Build HDiffPatch with `_IS_OUT_DIFF_INFO=0` (also exported via the
pkg-config file since it affects the headers), and disable the zstd
plugin's notices with `IS_NOTICE_compress_canceled=0`.

Assisted-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
pull request — a88bd558 Deployed Oct 11, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant