Conversation
|
Thanks for opening a pull request! This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format. If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project. Then could you also rename the pull request title in the following format? or After updating the title, you can mark the pull request as ready for review. See also: |
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 2, 2026 15:18
e737aaf to
da4ffbd
Compare
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
4 times, most recently
from
October 3, 2026 01:25
633ba32 to
199407c
Compare
Add an optional bias to the scalar and SIMD unpack paths so frame-of-reference decoders can apply it before each store. Preserve the full-width memcpy path and cover biased decoding across widths, offsets, epilogues, and dispatched instruction sets.
Implement frame selection, bit packing, exception handling, endian-independent wire serialization, and validated vector and page decoding for int32 and int64. Build the codec with CMake and Meson and include focused unit and benchmark coverage.
Register PFOR for INT32 and INT64, connect its encoder and decoder to the Parquet factories, and support dense, nullable, and vector-at-a-time page reads. Add the codec sources and integration targets to both build systems.
Add end-to-end file round trips, null handling, batched reads, malformed-input cases, and encoding-factory coverage. Compare PFOR with existing integer encodings over modeled analytical column distributions.
PFOR codec definitions are built into libarrow but used by Parquet and its tests. Mark the instantiated template classes for export so shared builds do not depend on ELF default visibility.
List the supported integer types and Preview status, and make explicit that writers use PFOR only when a column selects it.
Choose between raw values and adjacent differences per vector, and search for a frame that permits exceptions on either side of the packed window. Build the search histograms in one pass, retain the minimum-frame result as the fallback, and move the selected frame to the smallest covered value.
Add timestamp, sawtooth, bounded-rate, monotonic, and clustered distributions at both integer widths. These inputs distinguish value packing from delta packing and exercise frames that need not equal the minimum.
Estimate the cost of adjacent differences from a strided sample before running the full delta frame search. Reject delta mode when the estimate cannot beat the incumbent raw plan, avoiding unnecessary work on uncorrelated vectors.
Require writers to opt in before emitting the Preview PFOR encoding. Add a separate per-column setting for delta planning and pass the resulting options through the encoder factory. Readers continue to accept either representation.
Reject inconsistent delta flags, truncated start values, and vector bounds that extend beyond the page. Compute the unpacker's readable range from the complete delta header and clarify the related invariants in tests and comments.
Add an off-by-default force_delta option that bypasses planner rejection while emitting an ordinary delta vector. Use it to compare raw and delta payloads for the same input, and cover the forced path with round-trip tests.
Implement a portable 32-lane bit-packing kernel and record the selected layout in the PFOR page header. Use interleaving for complete 1024-value int32 vectors and retain sequential packing for unsupported types and tails.
Expose an experimental writer option and verify every bit width, exception patching, sequential tails, encoded-size equivalence, and round trips through a written Parquet file.
Register paired decode variants, add a page-sized destination, and document the scope required for a valid comparison. Hold value generation and encoded size constant so the benchmark isolates layout cost.
Implement the paper's contiguous 32-value lane assignment, configurable base coding, and a fused prefix-sum transpose back to file order. Add benchmarks that separate base coding, dependency layout, and permutation cost.
Add an int32 page decoder, validate the layout marker and stream length, and stop reads at the declared value count. Compare the complete decoder path with Arrow's existing delta decoder.
Store a little-endian value count and validate every encoded section before decoding. Support both lane-delta and transposed-delta payloads, handle alignment safely, and gate their tests behind the PFOR Preview option.
Compare sequential and lane-interleaved PFOR using the same encoded column shapes while holding delta selection fixed.
Exclude the generated AVX-512 unpacker from runtime dispatch because it builds vectors from scalar loads and uses out-of-line calls. Retain AVX2 as the highest dispatched target until the AVX-512 kernels use vector loads directly.
Build transpose tiles directly from unpacked registers to avoid a 4 KiB scratch grid. Apply the same structure to lane-parallel deltas, benchmark it across working-set sizes, and document valid comparisons between output orders.
Compile baseline and vectorized kernel tables in separate translation units and select between them with Arrow's runtime dispatch. Parameterize the kernels by architecture so separately compiled instances remain distinct at link time.
Replace boxed headings and informal benchmark terminology, clarify the validity criteria, and centralize the shared benchmark corpus. No codec behavior changes.
Add a four-lane NEON transpose that writes file order directly instead of using the portable scratch-grid fallback. Verify every packed width with and without a frame bias.
Remove convenience constructors that prevented designated initialization and caused force_delta to be dropped at layout call sites. Pass complete option objects through the wrapper and benchmarks.
Report unavailable ZSTD or LZ4 codecs through SkipWithError instead of aborting the benchmark process. Remove an unregistered duplicate DBP benchmark body.
Store the page's packing mode in VectorReader and pass it to each vector decode. Add a regression test covering multiple interleaved vectors and a sequential tail through the vector-at-a-time path.
Replace the informal term "ladder" with "shuffle sequence" in comments. No code changes.
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 3, 2026 22:09
199407c to
59a378b
Compare
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 4, 2026 19:35
59a378b to
3dd10b0
Compare
The interleaved block kernels were written only for uint32_t output: the grid was fixed at 32 rows of 32 lanes and the shift arithmetic carried the element width implicitly. A bit width of 3 that decodes into a byte therefore had no in-tree kernel at all, even though the container has nothing 32-bit about it. BlockGeometry<T> derives the grid from the element instead, so a block of 1024 values becomes 8 rows of 128 lanes for uint8_t and 16 rows of 64 lanes for uint16_t, and the packed payload stays 128 * w bytes at every element width, the same size the sequential layout needs for the same values. The unqualified kLanes and kRowsPerBlock now name that same 32-bit grid through BlockGeometry<uint32_t>, so the delta and PFOR wire formats, whose payload arithmetic is written on it, keep their meaning and their existing uses unchanged. The round-trip tests cover every bit width for all three element types, with and without a folded bias, and pack into a buffer that already holds another block, since a page encoder walks one output buffer block by block and relies on a kernel fully defining its own payload. They also pin the layout contract the format depends on: within a lane, successive rows occupy successive bit positions LSB-first, which is what lets a lane be read as an ordinary sequentially packed stream. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measure the interleaved block decoder against the three sequential bit-unpack paths Arrow already has: the unpack_width template driving Kernel<T, w, xsimd::avx2>, the entry point libarrow ships, and the scalar reference. Two sweeps share the harness. One holds the decoded output at 16 KiB and moves the bit width from 1 to 32, decoding into the next whole byte, so a row moves only with the width and the element it lands in; both buffers count against L1, so the pin keeps source plus destination inside this core's 48 KiB at every width. The other fixes three widths and walks the footprint from 16 KiB to 4 MiB, which already exceeds a Parquet page. The interleaved decoder is measured in both of the output orders a reader can ask for. UnpackBlock writes the values where the container holds them, which leaves a column still owing the 32x32 permutation, so a second row pays it: through UnpackBlockFlToFileOrder where that kernel exists, otherwise through a scratch grid and Transpose32x32, the same two paths InterleavedPforDecode chooses between. Both are 32-bit only, so the file-order row exists at u32 and not at the narrower elements. The sequential decoders are driven 16384 values at a time rather than in one call over the whole footprint, because batch_size is an int that the unpacker multiplies by the bit width. That page size was measured rather than assumed, so the sequential side carries no handicap from call granularity. Every buffer is 64-byte aligned, the way Arrow aligns a column's buffers; on an unaligned destination the interleaved decoder reads half speed. The footprint sizes are not labelled by cache level. The packed source adds w bits per value on top of the decoded destination, so the level a row lands in depends on the width as well as the size: 48 KiB of bytes at three bits totals 66 KiB and is already past this core's L1. The benchmark needs the xsimd headers to name the kernel it drives, so the target links ARROW_XSIMD. The adjacent test entry's continuation line is realigned to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
UnpackBlockFlToFileOrder claims to equal an UnpackBlock into a scratch grid followed by Transpose32x32, with the permutation carried out in registers instead of through memory. Nothing checked that claim, and it is the one a caller depends on when it picks the fused kernel over the grid. Compare the two at every width. They are separate implementations, so this is not a kernel agreeing with itself. The test is 32-bit only, like both kernels it compares, which is why it sits outside the typed tests above it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The footprint ladder carried 3, 11 and 21, one width per element size, so a row at u8 could not be told apart from a row at a width where Arrow's sequential kernel collapses. Add 4, 7, 12, 14 and 18 so one-byte and two-byte output each hold both a healthy width and a collapsing one, and record against each the column shape a Parquet writer produces there. Four-byte output collapses at no width the ladder covers, so both of its widths are healthy. No behaviour change outside this benchmark: the ladder registers more rows at the same footprints and the width sweep is untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The comment above DecodePageFileOrder read "a reader that wants a column in file order still owes the 32x32 permutation". That is true only of bytes already packed with the FastLanes lane assignment. The order is settled when the grid is filled: kFileOrder fills it in input order and the same UnpackBlock then writes file order owing nothing, which interleaved_pfor.h:396 states directly. Production's kForBitPackInterleaved carries no lane-assignment mode, so a Parquet reader is in that case and owes no permutation either. Say so, and note that the u8 and u16 gap here limits this column rather than what a reader can consume. Comment only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The large kernel aligns the high part of each value with a variable left shift. AVX2 has one for 32 and 64 bit lanes and none for byte or short lanes, so for those two output sizes xsimd widened and narrowed around its own shift, paying a lane extract and a pack on either side of it. Arrow had already met this in the other direction and solved it; right_shift_by_excess shifts narrow lanes as the next integer size up, masking the even and odd lanes apart. The left shift beside it did not. left_shift_by_excess mirrors it, and large unpacking to a byte now routes through the medium kernel on shorts on AVX2, as it already did on SSE4.2, because a byte needs two rounds of that widening on each shift where a short needs one. Only the large kernel needs a variable left shift, and it takes the packed widths whose remainder modulo eight is 3, 5, 6 or 7. In L1 with 256-bit registers those four widths now read 4.23x to 4.28x of the same code before the change at short output and 2.02x to 2.13x at byte output. Every other width moves within 0.94x to 1.06x, and four-byte output is untouched because a native instruction already covered it. AVX-512 keeps the old routing: the same measurement reads 0.16x to 0.23x there. Checked against a scalar reference over every kernel instantiation x86 can build, four output sizes and packed widths 1 to 63 on SSE4.2, AVX2 and AVX-512, in both the plain and the biased form. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Parquet uses a continuous bit-packed stream, while FastLanes uses a
lane-interleaved grid. This PR measures whether the grid's cheaper unpacking
can offset the cost of returning values in Parquet file order.
The answer depends on packed width, output element width, decode call shape,
memory footprint, instruction set, and the permutation back to file order.
This branch keeps those variables separate instead of reducing the comparison
to one throughput number.
What changes are included in this PR?
This benchmarking-only proof of concept adds:
Run
./layoutsinfl5_corpuswithout arguments to see the measurement index.Each study names the variable it changes and the assumptions required to
interpret its output.
Are these changes tested?
Yes. The codecs and experimental layouts are checked for bit-exact round trips
before they are timed. Unit tests cover packing widths, output order, delta
paths, tails, exceptions, and boundary values.
The benchmark harness includes timing controls and rejects results that do not
meet its validity checks. Results are recorded for multiple compilers,
instruction sets, and working-set sizes.
Are there any user-facing changes?
No production encoding or default writer behavior is proposed by this PR. It
is a benchmarking and research branch used to evaluate layout choices.
The experimental APIs and on-disk layouts in this branch should not be treated
as stable or interoperable formats.