Skip to content

create --sparse: fix stale help text, sparse input works with all chunkers - #5

Closed
ThomasWaldmann wants to merge 210 commits into
masterfrom
claude/priceless-mclean-b34733
Closed

ThomasWaldmann wants to merge 210 commits into
masterfrom
claude/priceless-mclean-b34733

Conversation

@ThomasWaldmann

Copy link
Copy Markdown
Owner

What changed

  • borg create --sparse help text: "detect sparse holes in input (supported only by fixed chunker)" → "detect sparse holes in input and seek over them instead of reading them" (src/borg/archiver/create_cmd.py).
  • Regenerated the affected generated docs: docs/usage/create.rst.inc and docs/man/borg-create.1 (their diffs are exactly the sparse line, plus the man page generation date).
  • docs/internals/data-structures.rst: moved the sparse-processing note from the "fixed" chunker section into the chunker overview and stated it generically.

Why

The "supported only by fixed chunker" claim is stale. Sparse input processing is generic now: --sparse flows through FilesystemObjectProcessors into get_chunker(..., sparse=...) for every chunker, and ChunkerBase.chunkify() hands it to FileReader/FileFMAPReader, which builds the hole map via SEEK_HOLE/SEEK_DATA and seeks over hole ranges instead of reading them, for all chunkers alike.

Verification

Ran all 7 chunkers (buzhash, buzhash64, fastcdc, rabin-aes, goldilocks-aes, toeplitz-aes, fixed) over a 58 MiB data/hole/data layout in two modes:

  • poisoned holes: the file contains random data in ranges an explicitly passed fmap declares as holes — every chunker reproduced the expected content (zeros in the hole ranges), proving the reader seeks over holes rather than reading them;
  • real sparse file with sparse=True: sparsemap() reported the exact holes and every chunker reproduced the file content exactly, with 40+ MiB of the holes returned as zero-chunks (CH_ALLOC).

Notes for the reviewer

  • Regenerating build_usage/build_man also surfaced unrelated drift in other commands' generated docs (analyze, check, benchmark_cpu, repo-create, environment, help) — left untouched here, to be picked up at the next release-time regen.
  • Existing sparse-file chunking tests only cover the fixed chunker; a parametrized all-chunker sparse test (using an explicit fmap so it runs on filesystems without SEEK_HOLE support) could be a follow-up.

🤖 Generated with Claude Code

MannXo and others added 30 commits August 1, 2026 22:56
Every other status-type command (info, repo-info, repo-list, list, diff,
prune) can emit JSON, so tooling does not have to scrape their text. borg
analyze was the odd one out, although its numbers are exactly what
monitoring wants.

--json emits the numbers the text report is rendered from, as raw byte
values, for the default mode (dedup_size, hotspots) as well as for
--by-name (by_name). The compression factor is left out: it is
stored_size / source_size, and "n/a" is not a useful JSON value.

To keep one source of truth, the analysis methods now return their
numbers and the printing moved into report_*() methods that format them.
The text output is unchanged, byte for byte.

hotspots is null rather than empty when fewer than two archives matched:
the hot spots were then not computed at all, which is different from
having computed them and found nothing.
The test rebuilt the expected hot-spot path from the input directory,
stripping a leading slash. Archived paths are normalized, and on Windows
that also drops the drive colon (C:\Users -> C/Users), so the expectation
read D:/a/... where borg had stored D/a/....

Assert the size of the input directory's hot spot by path suffix, like
the text-report test above already does, and check the paths against
what the text report prints instead of rebuilding them.
Only the intact-pack records are cycle progress; the corrupt ones are
kept for repair until the pack verifies intact or is gone from packs/,
and their ids are now reported in the check summary. Refs borgbackup#9696.
…ls their reuse

Records are pruned only for packs no longer listed in packs/, so --max-duration now requires --max-age to make progress.
Parse --max-age with the calendar-aware relative time marker (m, y measured
against now), tolerate MAX_CLOCK_SKEW at both ends of the reuse window, and
correct the stale --repository-only comment.
…orgbackup#9218

Cap distinct chunks and file refs kept for the end-of-run report, combine the two collection dicts into one, and add tests for the grouping and truncation.
…borgbackup#9989

fish completions are now generated by shtab, like the bash, zsh and tcsh
ones, extended with a fish preamble that provides the same dynamic
completions the bash/zsh scripts have (archive names, aid: archive IDs,
tags, sort keys, files cache modes, compression specs, chunker params,
relative time markers, timestamps, file sizes, help topics).

This needs shtab >= 1.9.3, which brings the fixed fish backend (see
tqdm/shtab#227, borgbackup#228, borgbackup#229, borgbackup#230 and borgbackup#231), so the pin was bumped.

Arguments taking a filesystem path are now typed, so that all shells can
complete them: such a path was either untyped or just a str, so the
shells relied on their default completion (bash/zsh/tcsh) or offered
nothing at all (fish). FilesystemPathSpec is used for arguments taking a
file (--exclude-from, --patterns-from, TARFILE, the borg debug file
arguments, borg key export/import PATH), the new FilesystemDirSpec for
arguments taking a directory (MOUNTPOINT, borg benchmark crud PATH).
Besides the completions, they now also reject empty paths and slashify
them on Windows, like the other filesystem path arguments do.
Note that borg key export/import PATH completed directories before
(it takes a file) and borg benchmark crud PATH completed files (it takes
a directory) - both are fixed by this.

shtab's fish FILE completion needs a little help to also complete
absolute paths, see tqdm/shtab#245 (fix: tqdm/shtab#246).

The completion script is now always generated for the "borg" command,
no matter how borg was invoked while generating it - completions
generated by "python -m borg completion ..." were useless before, as
the shells would complete a "__main__.py" command.

The hand-written, unmaintained fish completions in
scripts/shell_completions/ are removed, as are their tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ns-9989

completion: generate fish completions, remove hand-written ones
`borg completion tcsh` now generates a usable completion script:

- add a tcsh preamble with the dynamic completion helpers (as aliases, since tcsh has no
  functions) and wire the tcsh patterns up for sort keys, files-cache mode, compression specs,
  chunker params, relative times, timestamps, file sizes and help topics.
- complete archive names, archive IDs (when the token starts with "aid:") and tags. tcsh can
  neither define functions nor use backquotes there (the completion rule calling the helper is
  backquoted already), so both helpers run one POSIX sh script that parses $COMMAND_LINE for
  --repo/-r and queries `borg repo-list`. tcsh has no completion descriptions, so unlike zsh
  and fish these are plain candidate lists.

The generator fixes this needed are all upstream in shtab now (tqdm/shtab#213, released in
1.9.3, which we already require): positional completion under subcommands at any depth, no
out-of-range `$cmd` indexing, custom `.complete` patterns in multi-requirement rules, `--opt=`
completion, and rule deduplication.

One upstream fix is merged but not yet released (tqdm/shtab#241): completion patterns (`f`,
`d`, ...) for a positional of a subcommand end up inside a `p@N@` rule, where tcsh runs the
clauses as commands and only uses their output, so they do nothing - e.g. `borg umount <TAB>`
would not complete a mountpoint. `_tcsh_anchor_positional_patterns` rewrites those into `n/`
rules keyed off the preceding (sub)command word, producing exactly what a shtab with borgbackup#241
generates. It is a no-op with such a shtab and can then be removed.

Note that tcsh matches completions for positional arguments by word position, so options
before an archive name shift it out of place and it is not completed - the `borg completion`
epilog points at BORG_REPO for this.
LRUCache refused that with an assertion, so callers had to write

    if key in cache:
        cache.replace(key, value)
    else:
        cache[key] = value

and a caller that just assigned crashed - which is a nasty failure mode for a
cache: it is a correctness-neutral operation that suddenly is not. It bites
especially when several threads use one cache (see the FUSE data cache), where
two threads can miss on the same key and both insert the value they fetched.

Assignment now replaces the value, disposing the old one (it left the cache,
just like a deleted one would) unless it is the identical object, and counts as
a use, so the entry moves to the most-recently-used end.

replace() stays for the one case that really needs it: updating an entry while
keeping its old value alive - the legacy repository's fd cache re-timestamps its
entries that way and must not have the open file closed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Several threads share one LRUCache in borg: mfusepy serves the FUSE requests of
one mount from a thread pool, and "borg webdav" from one thread per connection.
Until now every caller had to serialize the access itself - a rule that is easy
to miss, and missing it does not fail loudly at the cache, it fails somewhere
far away (mfusepy turns any exception from a FUSE operation into EINVAL, so a
mount just reports "Invalid argument" for a healthy file).

Getting, setting, deleting, popping, replacing and clearing now each hold a lock,
so they are atomic. A *sequence* of them still is not - two threads can miss on
the same key and both store a value, which for a cache is wasteful, not wrong.

pop() is implemented here now instead of inheriting it: MutableMapping builds it
from __getitem__ + __delitem__, which is not atomic (two threads popping the same
key raise KeyError, see the new test), and it would dispose the value it hands to
the caller - for e.g. a cache of open files it returned a closed one.

Cost, measured with timeit: about 75 ns per operation (get 70 -> 148 ns, set 148
-> 222 ns). That is nothing next to what these caches guard - a chunk decrypt is
in the milliseconds, an item unpack in the microseconds - and a borg mount read
benchmark (1 GiB sequential, and stat+read of 200 small files, mfusepy) shows no
difference at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ng in docs

The sort provides partial-check progress on its own, so --max-duration no
longer requires --max-age; bare --max-duration works as documented again.
Credit the least-recently-checked ordering (not --max-age) for resumption in
the epilog and data-structures.rst, and rework the partial-ordering test to
use intact packs so it exercises the skip path together with the sort.
…ad-safe

lrucache: allow replacing an entry, make it thread-safe
…0020

borg materialized an archive as a browsable tree in three independent places:
fuse.py (llfuse/pyfuse3, low-level FUSE), hlfuse.py (mfusepy, high-level FUSE)
and webdav.py (ArchiveVFS + WebDAV/HTTP server). All three re-implemented tree
building, hardlink handling, the versions view, uid/gid/mode/time mapping and
reading file content from chunk lists - so every behaviour fix had to be applied
N times (e.g. the ACL/xattr exposure fix borgbackup#9954 touched both FUSE variants).

New module vfs.py has that logic exactly once:

- ArchiveVFS: archive selection and name deduplication, lazily built per-archive
  trees, the versions view, hardlinks (nodes sharing one inode), item storage
  (msgpacked, path-less, as hlfuse did it), attribute mapping, xattrs/ACLs.
- DataReader: reads byte ranges out of chunk lists, with the decrypted-chunk
  cache (BORG_MOUNT_DATA_CACHE_ENTRIES) and the sequential-read position hint.
- parse_mount_options(): the "borg mount -o ..." parsing both mounts duplicated.

fuse.py, hlfuse.py and webdav.py are now thin protocol adapters over it (2604 ->
2009 lines in total). Behaviour changes that fell out of the unification:

- webdav reads now go through DownloadPipeline.fetch_many(), so the all-zero
  chunk shortcut and the parsed-chunk cache (borgbackup#1678) apply to mounts as well.
- the mounts get webdav's Unicode NFC lookup fallback (macOS decomposes names).
- directories report st_nlink >= 2 (hlfuse behaviour) in both mounts.
- a chunk that is read to its end is no longer put into the data cache, so a
  full download does not evict the chunks partial (range) reads need - this was
  the FUSE behaviour, now webdav shares it.

Threading: mfusepy serves the FUSE requests of one mount from a thread pool and
webdav one request per thread, so the VFS is used from several threads. It
serializes the repository access (borgstore connections are not thread-safe) and
leaves the caches to LRUCache, which is thread-safe itself. Two threads can miss
on the same inode and both unpack the item, or store a read position over each
other - they get equal items, and the position is only a hint.

The ACL emulation and the NFC lookup are now tested against the core
(testsuite/vfs_test.py, no FUSE dependency); fuse_test.py keeps testing what is
left in the adapters: the errno mapping.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
unify the three archive-as-filesystem implementations
LZ4 used one module-level Buffer for both directions. Two threads
(de)compressing lz4 chunks concurrently got the same bytearray and wrote
into it at the same time, silently corrupting each other's output - a
blocker for parallel decompression, and lz4 is not an exotic setting.

Thread-locals are the right granularity, not per compressor instance:
compressor instances are shared between threads (LZ4_COMPRESSOR, used by
Auto), a thread's buffer is by definition only used by that thread. The
buffer reuse that avoids a fresh allocation per chunk is kept.
fastcdc, buzhash64: AVX-512 variants of the 8-lane candidate test
(one 512-bit vector, vptestnmq fusing the AND and the ==0 test into
a mask register), runtime-detected on x86-64 above the AVX2 kernels.
BORG_FASTCDC_NO_AVX512 / BORG_BUZHASH64_NO_AVX512 cap dispatch at
AVX2 for benchmarking the kernels against each other.

rabin-aes/goldilocks-aes/toeplitz-aes: VAES/AVX-512 variant of the
x86-64 hardware path, kind "vaes": groups of 32 positions encrypted
as 8 zmm vectors of 4 AES blocks each - 4x fewer AES instructions,
register-resident round keys (no per-round reloads), 8 independent
chains to hide the vaesenc latency, vpexpandq digest placement and
a masked vptestnmq cut test without extracts. BORG_PHTE_NO_VAES
caps the AES chunkers at the 128-bit AES-NI path. The VAES path
needs GCC >= 11 / clang >= 14 for __builtin_cpu_supports("vaes");
older compilers keep the AES-NI path.

All kernels return bit-identical cut points; the existing
kernel-identity tests cover the new paths where the CPU has them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ThomasWaldmann and others added 27 commits August 15, 2026 00:30
…rs-default

compression: cap the default zstd MT workers at 4 for chunks
…selftest

tests: keep pytest out of the chunkers testsuite package __init__
All our CI machines are little-endian, so the big-endian code paths in the
native code (e.g. the __builtin_bswap64 calls in the chunker kernels) are
never executed and the architecture independence of the repository format
is not tested at all.

Build and test borg on s390x in an emulated container, running the tests of
the endianness sensitive parts (chunkers, crypto, compression, hashindex,
item, repository format).

Additionally run a cross-architecture interoperability test: a repository
written on the big-endian side is checked, extracted and compared on the
little-endian side and the other way round.  Both sides also have to produce
identical chunk ids for the same data - if they did not, the archives would
still be correct, but deduplication between machines of different endianness
would silently not work.

Emulating everything is slow (~35min), so this does not run for every pull
request, but only for pull requests touching native or format relevant code,
plus weekly and on demand.
…s390x

CI: add big-endian (s390x) test under qemu emulation
It was the only workflow without a permissions declaration, so it got
whatever the repository default is. It only checks out and runs black.
…issions

CI: give the black workflow read-only token permissions
Pushing a release tag so far only built the standalone binaries and uploaded
them as workflow artifacts - everything else was manual: build and sign the
sdist, upload it to PyPI, fetch the binary artifacts and create the GitHub
release with them.

Add a release job (.github/workflows/release.yml), called by ci.yml for tags
after the tests and the binary builds succeeded - it runs inside the same
workflow run, because that is where the binary artifacts are. It builds the
sdist, verifies that the sdist is complete by installing it into a clean venv
and running borg, attests its provenance and drafts the GitHub release with
the sdist and all standalone binaries attached.

A pypi job then uploads the sdist to PyPI via trusted publishing, so no API
token has to be stored anywhere. It lives in ci.yml because PyPI trusted
publishing does not work from a reusable workflow. It uses the "pypi"
environment: configuring required reviewers for it makes the irreversible
upload wait for an approval.

The Windows binaries now use the same naming and layout as the binaries of
the other platforms (borg-windows-x86_64-gh.exe and .tgz, with the
single-directory variant and a provenance attestation), so they become
release assets, too.

The GitHub release is created as a draft: the release notes want a human, the
detached GPG signature of the sdist can only be made locally, and the binaries
should be tried out before the release becomes visible.
CI: release automation for PyPI and GitHub releases
…1648

create: support --stats with --dry-run
os.utime is a bad fit for restoring timestamps on Windows:

- it can not set the birthtime (creation time), so the birthtime was
  archived (Python >= 3.12 has st_birthtime_ns there), but never restored.
- it does not support follow_symlinks=False there, so the timestamps of an
  extracted symlink were set on the symlink's target.
- it does not support file descriptors there, so we had to use the path
  even though we had the file open already.

Add platform.set_times() which does all that: the POSIX implementation is
the os.utime code moved from Archive.restore_attrs (incl. the utimes trick
to set the birthtime on the BSDs and Darwin), the win32 implementation uses
CreateFileW / SetFileTime.

Also, failing to set the timestamps is not silently ignored on win32 anymore,
but warns and sets the warning exit code.
…-7269

windows: restore timestamps via SetFileTime, fixes borgbackup#7269
If a recursion root (a path given on the command line or via a patterns
file) is a symlink, borg now follows it and archives what it points to,
using the path as given.

Symlinks encountered while recursing are not affected, they are archived
as symlinks as before. Symlinks given via --paths-from-* are never
followed either.

Archiving the symlink target under the symlink's path also keeps the files
cache working if the symlink target changes its name (e.g. a 'current'
symlink pointing to the latest of a series of directories).

A recursion root that is a symlink with a non-existing target is skipped
with a warning (new BackupBrokenSymlinkError, rc 112).

As a followed symlink must be opened without O_NOFOLLOW, add flags_dir_follow
and flags_normal_follow and let the callers choose the flags.
The BackupErrors raised while backing up a fs object do not reach the
top level, they are wrapped into a BackupWarning and only reported as a
warning. Thus, the expected exit code must come from the BackupWarning
(giving EXIT_WARNING for BORG_EXIT_CODES=classic and the error's specific
warning rc for modern), not from the wrapped error (which would give
EXIT_ERROR for classic).

Removes the "workaround, TODO: fix it" that patched up EXIT_ERROR after
the fact.
…ymlinks-4737

create: follow symlinks given as recursion roots, fixes borgbackup#4737
…rebuild-10026

check --repair: rebuild a corrupt repository index from the packs, borgbackup#10026
shtab 1.11.0 (released 2026-08-15) labels a zsh option argument with the
argument's metavar instead of its dest, so the generated spec for
"borg list --sort-by" (metavar KEYS, dest sort_by) changed from

  "--sort-by[...]:sort_by:_borg_complete_item_sortby"

to

  "--sort-by[...]:KEYS:_borg_complete_item_sortby"

test_zsh_sortby_wiring matched the ":sort_by:" form literally and started
failing for list and repo-list (borg diff --sort-by has no metavar, so
shtab still falls back to the dest there and that entry kept passing).

The completions are fine: the label is only what zsh displays while
completing, and --sort-by is still wired to our helper function. Match the
helper and leave the label open, so either shtab behaviour passes.

Note that shtab is not pinned in requirements.d/, it comes in via the
shtab>=1.9.3 dependency in pyproject.toml, so CI picks up new shtab
releases as they appear.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…y-test-shtab111

tests: do not pin the zsh --sort-by label, shtab 1.11.0 changed it
…d process

The foreground process installs SIGTERM/SIGHUP/SIGINT handlers (via
archiver.run) before it reaches the point where it actually waits for the
background (grandchild) process to notify it (via os.kill). If the
background process started up and signalled fast enough, the signal was
delivered to the foreground while it was still between os.fork() and the
waiting code, so the globally installed handler raised at an unexpected,
uncaught place. The signal then escaped daemonizing(), bubbled up through
repository teardown (NotLocked) and made "borg mount" exit with rc 74.

This was observed flaky in CI with coverage's sys.monitoring backend on
Python 3.14 (its first-branch lazy source parse widens the window) and the
pyfuse3 backend (faster grandchild startup).

Fix the race by blocking the notify signals before the fork in
_daemonize() and waiting for them atomically in the foreground. An early
signal then stays pending and is reliably picked up by the wait. The
background process restores the original signal mask so it keeps normal
signal handling.

Use signal.sigwait() plus a SIGALRM timer (signal.setitimer) for the
wait, rather than signal.sigtimedwait(): the latter does not exist on
macOS, where it would raise AttributeError in the foreground and let it
die before the background migrated the lock (breaking test_migrate_lock_alive).

signal.SIGALRM in turn does not exist on Windows, where referencing it at
module import time would raise AttributeError and break the import chain
(and the Windows PyInstaller build), so guard it with hasattr, matching the
defensive getattr pattern already used for the notify signals. Daemonizing
is not supported on Windows anyway (no os.fork), so the empty SIGALRM list
has no functional effect there.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The archiver fixture's rmtree onerror handler called os.lchflags(path, 0)
when has_lchflags was True. But has_lchflags is also True on Linux (flags
are cleared via ioctl there), where os.lchflags does not exist, raising an
uncaught AttributeError and turning teardown into an ERROR. Use borg's
cross-platform platform.set_flags instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
borg2 does not have the asymmetric hard link model any more (all members of a
hard link group have their own chunks list, they are identified by hlid), so
the remaining master/slave wording only described borg1 archives.

Rename the borg1 compat predicates by what they actually test:

- borg1_hardlink_master -> borg1_hardlink_with_content
- borg1_hardlink_slave -> borg1_hardlink_reference

The distinction is not a hierarchy, but whether an item carries the chunks
list or only a .source back reference to the item that does. Since borgbackup#7175,
borg1 set the hardlink_master key for all hardlinkable items, so that key
does not tell them apart - only the absence of .source does. Added a test
pinning that.

The 'hardlink_master' item key itself is borg1 on-disk data and has to keep
its name, it is now commented as such.

Also fix two places that were not only using the old wording, but were also
stale: borg2 hard links do have their own chunks list and thus their real
size, they are not size 0.

While at it, fix the Item.get_size stub, it still had the long gone
hardlink_masters and compressed parameters.
…rmtree-linux

conftest: fix rmtree cleanup crash on Linux (os.lchflags does not exist)
…inology-5248

get rid of master/slave terminology for hard links, borgbackup#5248
…ng-early-signal

fix daemonizing: don't lose an early notify signal from the background process
…nkers

The help text still claimed sparse hole detection is "supported only by
fixed chunker", which is stale: --sparse is passed to every chunker via
get_chunker(), and ChunkerBase/FileFMAPReader implement sparse processing
generically (build the hole map via SEEK_HOLE/SEEK_DATA, seek over holes
instead of reading them, process their content as all-zero).

Verified for all chunkers (buzhash, buzhash64, fastcdc, rabin-aes,
goldilocks-aes, toeplitz-aes, fixed) by chunking a real sparse file with
sparse=True and a hole-poisoned file with an explicit fmap: content is
reproduced exactly and hole ranges are seeked over, not read.

Also move the sparse processing note in docs/internals/data-structures.rst
from the "fixed" chunker section to the chunker overview, and regenerate
the affected generated docs (docs/usage/create.rst.inc, borg-create.1).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ThomasWaldmann

Copy link
Copy Markdown
Owner Author

Opened against the wrong repo - this should target borgbackup/borg. Superseded by a PR there.

@ThomasWaldmann
ThomasWaldmann deleted the claude/priceless-mclean-b34733 branch August 17, 2026 02:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants