[WIP][POC] GH-51268: [C++][Parquet] Unpack equal-width DELTA_BINARY_PACKED miniblocks in one call - #51250
Draft
prtkgaur wants to merge 2 commits into
Draft
[WIP][POC] GH-51268: [C++][Parquet] Unpack equal-width DELTA_BINARY_PACKED miniblocks in one call#51250prtkgaur wants to merge 2 commits into
prtkgaur wants to merge 2 commits into
Conversation
|
Thanks for opening a pull request! This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format. If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project. Then could you also rename the pull request title in the following format? or After updating the title, you can mark the pull request as ready for review. See also: |
prtkgaur
force-pushed
the
delta-binary-packed-coalesce-miniblocks
branch
from
September 9, 2026 03:07
c16db5a to
fa243dd
Compare
|
|
prtkgaur
force-pushed
the
delta-binary-packed-coalesce-miniblocks
branch
from
September 10, 2026 19:32
fa243dd to
9174897
Compare
prtkgaur
force-pushed
the
delta-binary-packed-coalesce-miniblocks
branch
5 times, most recently
from
September 18, 2026 21:09
0cd12d7 to
e686ed5
Compare
The output buffer may alias decoder members of the same type, forcing repeated reloads after stores. Keep the frame and running value in locals, write back once, and preserve unsigned wrapping during reconstruction.
prtkgaur
force-pushed
the
delta-binary-packed-coalesce-miniblocks
branch
from
October 4, 2026 00:13
443a6d4 to
4e3c46f
Compare
Adjacent equal-width miniblocks form one packed stream. Decode each run with one unpack call, bounded by the current block and output request; retain the zero-width path and cover width changes, partial reads, block boundaries, and both integer widths.
prtkgaur
force-pushed
the
delta-binary-packed-coalesce-miniblocks
branch
from
October 4, 2026 02:36
4e3c46f to
b710c74
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Two thins that impact DELTA_BINARY_PACKED decoder per value:
min_delta_andlast_value_have the sametype as the output buffer, so the compiler cannot prove the store does not alias them
and reloads both for every value.
default geometry means a call per 32 values for 32-bit and per 64 for 64-bit. At narrow
widths, where unpacking a value is cheap, that setup is a large share of the work.
What changes are included in this PR?
The prefix-sum loop now keeps the running value and the frame of reference in locals and
writes the running value back once, after the loop.
The decoder also unpacks consecutive miniblocks of equal bit width in one call.
Are these changes tested?
A new typed test covers the width patterns that decide where a run starts and stops, and
a read batch size that stops partway through a coalesced run.
Benchmark
Graviton4, GCC 11.5,
Release, one core, 9 repetitions, medians, 65,536 values. Bothsides were built twice, with the builds interleaved; the two builds agree within 0.6%.
Decode_Int32_NarrowDecode_Int32_WideDecode_Int64_NarrowDecode_Int64_WideDecode_Int32_FixedDecode_Int64_FixedAre there any user-facing changes?
No. No API change, no format change, and decoded values are identical.