gather: read multiple byte ranges with one call, fixes #211 - #231
Merged
Merged
Conversation
ThomasWaldmann
force-pushed
the
gather-211
branch
4 times, most recently
from
September 25, 2026 22:46
b129214 to
8c24d4d
Compare
Store.gather(sources, namespace=, deleted=) / BackendBase.gather(sources) take a list of (name, offset, size) tuples (like defrag) and return the contents of all ranges concatenated, in the order given. A short read raises ReadRangeError, the caller splits the result knowing the sizes it requested. - BackendBase.gather: one partial load per range, works for all backends. defrag is now implemented on top of it (gather, then store the new item). - posixfs: checks the read permission of all sources before reading anything. defrag only checks the target permission itself, the sources are checked by gather. - rest: POST /?cmd=gather with the sources as JSON body, the server gathers with one roundtrip and answers with the concatenated bytes (application/octet-stream). A server that does not know the command gives a clear error. - Store.gather: ranges of items in a cached namespace are served like load does it, all other ranges go to the backend with one gather call. Stats: gather_calls/time/volume, backend_gather_calls/volume. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ckend In a cached namespace, a partial load that has to load the whole item from the primary backend (mirror mode, or a cache miss in writethrough mode) sliced the range out of the item as value[offset:offset + size]. With a negative offset, that end index is wrong, e.g. offset=-3, size=3 gave value[-3:0] == b"" instead of the last 3 bytes. A writethrough cache hit was correct, so the result depended on the cache state. gather has the same issue for ranges of cached items. Make a negative offset absolute (relative to the item size) before slicing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ThomasWaldmann
force-pushed
the
gather-211
branch
from
September 26, 2026 09:01
8c24d4d to
6c56bef
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #211.
Store.gather(sources, namespace=, deleted=)/BackendBase.gather(sources)take(name, offset, size)tuples (same format asdefrag) and return the contents of all ranges concatenated, in the order given. A short read raisesReadRangeError; the caller knows the sizes it requested, so it can split the result (e.g. into memoryview slices).defragis now implemented on top of it (gather, then store the new item) - a behavior-neutral refactor, the existing defrag tests pass unchanged.validate_sourceschecks each tuple and returns a new list, which the callers use. So a generator works, too: validating consumes it, iterating over the original sources again would silently yield nothing.gather, to check the read permission of all sources before reading anything.defragonly checks the target permission itself now, the sources are checked by gather. This also fixes which names are checked:defragcheckednamespace + name, but the names it gets already include the namespace, so it checked a nonexistent path (only the parent-directory fallback made that pass).POST /?cmd=gatherwith the sources as JSON body; the server gathers with one roundtrip and answers with the concatenated bytes (application/octet-stream, 416 on a short read). The client keeps its own server-sidedefrag(building it on the client's gather would download and re-upload the data).loaddoes it, all other ranges go to the backend with one call. New stats:gather_calls/time/volume,backend_gather_calls/volume.Separate commit, a fix for an existing bug that gather also hits:
value[offset:offset + size]. With a negative offset, that end index is wrong, e.g. offset=-3, size=3 gaveb""instead of the last 3 bytes. A writethrough cache hit was correct, so the result depended on the cache state. Now a negative offset is made absolute before slicing.Tests: store (nesting, soft-deleted, short read, invalid sources, generator sources, stats keys), all backends, cache interaction, negative offsets in both cache modes, posixfs permissions, REST server (raw HTTP) and REST client.
🤖 Generated with Claude Code