Skip to content

[WIP][POC][BenchmarkingOnly] PFOR interleaved and fastlanes numbers - #51296

Draft
prtkgaur wants to merge 34 commits into
apache:mainfrom
prtkgaur:pgaur_interleavedPlusFastLanesDelta
Draft

prtkgaur wants to merge 34 commits into
apache:mainfrom
prtkgaur:pgaur_interleavedPlusFastLanesDelta

Conversation

@prtkgaur

@prtkgaur prtkgaur commented Sep 10, 2026 •

Copy link
Copy Markdown

Rationale for this change

Parquet uses a continuous bit-packed stream, while FastLanes uses a
lane-interleaved grid. This PR measures whether the grid's cheaper unpacking
can offset the cost of returning values in Parquet file order.

The answer depends on packed width, output element width, decode call shape,
memory footprint, instruction set, and the permutation back to file order.
This branch keeps those variables separate instead of reducing the comparison
to one throughput number.

What changes are included in this PR?

This benchmarking-only proof of concept adds:

  • Portable lane-interleaved and FastLanes-style bit-packing kernels.
  • AVX2 and NEON paths that perform the file-order permutation in registers.
  • Runtime dispatch for the interleaved kernels.
  • PFOR experiments using interleaved, lane-delta, and transposed-delta layouts.
  • Benchmarks for packed width, output width, corpus shape, and memory footprint.
  • Controls for call shape, output address, bytes moved, lane order, and stores.
  • Reproducible build and analysis scripts for x86-64 and AArch64.
  • Checked-in benchmark results and documentation of their limitations.

Run ./layouts in fl5_corpus without arguments to see the measurement index.
Each study names the variable it changes and the assumptions required to
interpret its output.

Are these changes tested?

Yes. The codecs and experimental layouts are checked for bit-exact round trips
before they are timed. Unit tests cover packing widths, output order, delta
paths, tails, exceptions, and boundary values.

The benchmark harness includes timing controls and rejects results that do not
meet its validity checks. Results are recorded for multiple compilers,
instruction sets, and working-set sizes.

Are there any user-facing changes?

No production encoding or default writer behavior is proposed by this PR. It
is a benchmarking and research branch used to evaluate layout choices.

The experimental APIs and on-disk layouts in this branch should not be treated
as stable or interoperable formats.

@github-actions

Copy link
Copy Markdown

Thanks for opening a pull request!

This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format.

If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose

Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project.

Then could you also rename the pull request title in the following format?

GH-${GITHUB_ISSUE_ID}: [${COMPONENT}] ${SUMMARY}

or

MINOR: [${COMPONENT}] ${SUMMARY}

After updating the title, you can mark the pull request as ready for review.

See also:

@prtkgaur prtkgaur changed the title [WIP][POC][BenchmarkingOnly] PFOR interleaved and flanes numbers [WIP][POC][BenchmarkingOnly] PFOR interleaved and fastlanes numbers Sep 11, 2026
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from e737aaf to da4ffbd Compare October 2, 2026 15:18
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch 4 times, most recently from 633ba32 to 199407c Compare October 3, 2026 01:25
Add an optional bias to the scalar and SIMD unpack paths so
frame-of-reference decoders can apply it before each store. Preserve the
full-width memcpy path and cover biased decoding across widths, offsets,
epilogues, and dispatched instruction sets.
Implement frame selection, bit packing, exception handling,
endian-independent wire serialization, and validated vector and page
decoding for int32 and int64. Build the codec with CMake and Meson and
include focused unit and benchmark coverage.
Register PFOR for INT32 and INT64, connect its encoder and decoder to
the Parquet factories, and support dense, nullable, and vector-at-a-time
page reads. Add the codec sources and integration targets to both build
systems.
Add end-to-end file round trips, null handling, batched reads,
malformed-input cases, and encoding-factory coverage. Compare PFOR with
existing integer encodings over modeled analytical column distributions.
PFOR codec definitions are built into libarrow but used by Parquet and
its tests. Mark the instantiated template classes for export so shared
builds do not depend on ELF default visibility.
List the supported integer types and Preview status, and make explicit
that writers use PFOR only when a column selects it.
Choose between raw values and adjacent differences per vector, and
search for a frame that permits exceptions on either side of the packed
window. Build the search histograms in one pass, retain the
minimum-frame result as the fallback, and move the selected frame to the
smallest covered value.
Add timestamp, sawtooth, bounded-rate, monotonic, and clustered
distributions at both integer widths. These inputs distinguish value
packing from delta packing and exercise frames that need not equal the
minimum.
Estimate the cost of adjacent differences from a strided sample before
running the full delta frame search. Reject delta mode when the estimate
cannot beat the incumbent raw plan, avoiding unnecessary work on
uncorrelated vectors.
Require writers to opt in before emitting the Preview PFOR encoding. Add
a separate per-column setting for delta planning and pass the resulting
options through the encoder factory. Readers continue to accept either
representation.
Reject inconsistent delta flags, truncated start values, and vector
bounds that extend beyond the page. Compute the unpacker's readable
range from the complete delta header and clarify the related invariants
in tests and comments.
Add an off-by-default force_delta option that bypasses planner rejection
while emitting an ordinary delta vector. Use it to compare raw and delta
payloads for the same input, and cover the forced path with round-trip
tests.
Implement a portable 32-lane bit-packing kernel and record the selected
layout in the PFOR page header. Use interleaving for complete 1024-value
int32 vectors and retain sequential packing for unsupported types and
tails.
Expose an experimental writer option and verify every bit width,
exception patching, sequential tails, encoded-size equivalence, and
round trips through a written Parquet file.
Register paired decode variants, add a page-sized destination, and
document the scope required for a valid comparison. Hold value
generation and encoded size constant so the benchmark isolates layout
cost.
Implement the paper's contiguous 32-value lane assignment, configurable
base coding, and a fused prefix-sum transpose back to file order. Add
benchmarks that separate base coding, dependency layout, and permutation
cost.
Add an int32 page decoder, validate the layout marker and stream length,
and stop reads at the declared value count. Compare the complete decoder
path with Arrow's existing delta decoder.
Store a little-endian value count and validate every encoded section
before decoding. Support both lane-delta and transposed-delta payloads,
handle alignment safely, and gate their tests behind the PFOR Preview
option.
Compare sequential and lane-interleaved PFOR using the same encoded
column shapes while holding delta selection fixed.
Exclude the generated AVX-512 unpacker from runtime dispatch because it
builds vectors from scalar loads and uses out-of-line calls. Retain AVX2
as the highest dispatched target until the AVX-512 kernels use vector
loads directly.
Build transpose tiles directly from unpacked registers to avoid a 4 KiB
scratch grid. Apply the same structure to lane-parallel deltas,
benchmark it across working-set sizes, and document valid comparisons
between output orders.
Compile baseline and vectorized kernel tables in separate translation
units and select between them with Arrow's runtime dispatch.
Parameterize the kernels by architecture so separately compiled
instances remain distinct at link time.
Replace boxed headings and informal benchmark terminology, clarify the
validity criteria, and centralize the shared benchmark corpus. No codec
behavior changes.
Add a four-lane NEON transpose that writes file order directly instead
of using the portable scratch-grid fallback. Verify every packed width
with and without a frame bias.
Remove convenience constructors that prevented designated initialization
and caused force_delta to be dropped at layout call sites. Pass complete
option objects through the wrapper and benchmarks.
Report unavailable ZSTD or LZ4 codecs through SkipWithError instead of
aborting the benchmark process. Remove an unregistered duplicate DBP
benchmark body.
Store the page's packing mode in VectorReader and pass it to each vector
decode. Add a regression test covering multiple interleaved vectors and
a sequential tail through the vector-at-a-time path.
Replace the informal term "ladder" with "shuffle sequence" in comments.
No code changes.
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from 199407c to 59a378b Compare October 3, 2026 22:09
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from 59a378b to 3dd10b0 Compare October 4, 2026 19:35
sfc-gh-pgaur and others added 6 commits October 4, 2026 20:56
The interleaved block kernels were written only for uint32_t output:
the grid was fixed at 32 rows of 32 lanes and the shift arithmetic
carried the element width implicitly. A bit width of 3 that decodes
into a byte therefore had no in-tree kernel at all, even though the
container has nothing 32-bit about it. BlockGeometry<T> derives the
grid from the element instead, so a block of 1024 values becomes 8
rows of 128 lanes for uint8_t and 16 rows of 64 lanes for uint16_t,
and the packed payload stays 128 * w bytes at every element width,
the same size the sequential layout needs for the same values. The
unqualified kLanes and kRowsPerBlock now name that same 32-bit grid
through BlockGeometry<uint32_t>, so the delta and PFOR wire formats,
whose payload arithmetic is written on it, keep their meaning and
their existing uses unchanged.

The round-trip tests cover every bit width for all three element
types, with and without a folded bias, and pack into a buffer that
already holds another block, since a page encoder walks one output
buffer block by block and relies on a kernel fully defining its own
payload. They also pin the layout contract the format depends on:
within a lane, successive rows occupy successive bit positions
LSB-first, which is what lets a lane be read as an ordinary
sequentially packed stream.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measure the interleaved block decoder against the three sequential
bit-unpack paths Arrow already has: the unpack_width template driving
Kernel<T, w, xsimd::avx2>, the entry point libarrow ships, and the
scalar reference. Two sweeps share the harness. One holds the decoded
output at 16 KiB and moves the bit width from 1 to 32, decoding into
the next whole byte, so a row moves only with the width and the
element it lands in; both buffers count against L1, so the pin keeps
source plus destination inside this core's 48 KiB at every width. The
other fixes three widths and walks the footprint from 16 KiB to 4 MiB,
which already exceeds a Parquet page.

The interleaved decoder is measured in both of the output orders a
reader can ask for. UnpackBlock writes the values where the container
holds them, which leaves a column still owing the 32x32 permutation,
so a second row pays it: through UnpackBlockFlToFileOrder where that
kernel exists, otherwise through a scratch grid and Transpose32x32,
the same two paths InterleavedPforDecode chooses between. Both are
32-bit only, so the file-order row exists at u32 and not at the
narrower elements.

The sequential decoders are driven 16384 values at a time rather than
in one call over the whole footprint, because batch_size is an int
that the unpacker multiplies by the bit width. That page size was
measured rather than assumed, so the sequential side carries no
handicap from call granularity. Every buffer is 64-byte aligned, the
way Arrow aligns a column's buffers; on an unaligned destination the
interleaved decoder reads half speed.

The footprint sizes are not labelled by cache level. The packed source
adds w bits per value on top of the decoded destination, so the level
a row lands in depends on the width as well as the size: 48 KiB of
bytes at three bits totals 66 KiB and is already past this core's L1.

The benchmark needs the xsimd headers to name the kernel it drives, so
the target links ARROW_XSIMD. The adjacent test entry's continuation
line is realigned to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
UnpackBlockFlToFileOrder claims to equal an UnpackBlock into a scratch
grid followed by Transpose32x32, with the permutation carried out in
registers instead of through memory. Nothing checked that claim, and
it is the one a caller depends on when it picks the fused kernel over
the grid. Compare the two at every width. They are separate
implementations, so this is not a kernel agreeing with itself.

The test is 32-bit only, like both kernels it compares, which is why
it sits outside the typed tests above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The footprint ladder carried 3, 11 and 21, one width per element size,
so a row at u8 could not be told apart from a row at a width where
Arrow's sequential kernel collapses. Add 4, 7, 12, 14 and 18 so one-byte
and two-byte output each hold both a healthy width and a collapsing one,
and record against each the column shape a Parquet writer produces
there. Four-byte output collapses at no width the ladder covers, so both
of its widths are healthy.

No behaviour change outside this benchmark: the ladder registers more
rows at the same footprints and the width sweep is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The comment above DecodePageFileOrder read "a reader that wants a column
in file order still owes the 32x32 permutation". That is true only of
bytes already packed with the FastLanes lane assignment. The order is
settled when the grid is filled: kFileOrder fills it in input order and
the same UnpackBlock then writes file order owing nothing, which
interleaved_pfor.h:396 states directly. Production's
kForBitPackInterleaved carries no lane-assignment mode, so a Parquet
reader is in that case and owes no permutation either.

Say so, and note that the u8 and u16 gap here limits this column rather
than what a reader can consume. Comment only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The large kernel aligns the high part of each value with a variable
left shift. AVX2 has one for 32 and 64 bit lanes and none for byte or
short lanes, so for those two output sizes xsimd widened and narrowed
around its own shift, paying a lane extract and a pack on either side
of it. Arrow had already met this in the other direction and solved
it; right_shift_by_excess shifts narrow lanes as the next integer size
up, masking the even and odd lanes apart. The left shift beside it did
not. left_shift_by_excess mirrors it, and large unpacking to a byte
now routes through the medium kernel on shorts on AVX2, as it already
did on SSE4.2, because a byte needs two rounds of that widening on
each shift where a short needs one.

Only the large kernel needs a variable left shift, and it takes the
packed widths whose remainder modulo eight is 3, 5, 6 or 7. In L1 with
256-bit registers those four widths now read 4.23x to 4.28x of the
same code before the change at short output and 2.02x to 2.13x at byte
output. Every other width moves within 0.94x to 1.06x, and four-byte
output is untouched because a native instruction already covered it.
AVX-512 keeps the old routing: the same measurement reads 0.16x to
0.23x there.

Checked against a scalar reference over every kernel instantiation x86
can build, four output sizes and packed widths 1 to 63 on SSE4.2, AVX2
and AVX-512, in both the plain and the biased form.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants