Skip to content

GH-51268: [C++][Parquet] Scan DELTA_BINARY_PACKED deltas a vector at a time - #51295

Draft
prtkgaur wants to merge 3 commits into
apache:mainfrom
prtkgaur:gh51268-dbp-vector-prefix-sum
Draft

prtkgaur wants to merge 3 commits into
apache:mainfrom
prtkgaur:gh51268-dbp-vector-prefix-sum

Conversation

@prtkgaur

@prtkgaur prtkgaur commented Sep 10, 2026 •

Copy link
Copy Markdown

Stacked on #51250.

Rationale for this change

Prefix sum: a chain of dependent additions. A log-step inclusive scan computes it a register at a
time, in log2(lanes) shifted additions. Adding the frame of reference before the scan turns
its running multiple into a term the scan produces, rather than a multiply per lane.

What changes are included in this PR?

I replaced the sequential addition loop with a vectorized prefix sum (scan) to speed up decoding.

Are these changes tested?

Yes. I added a new typed test that explicitly forces the decoder through the vector loop, the scalar remainder loop, and the hand-off between them to ensure totals and wrapping behave correctly.

Benchmark

Same setup as #51250. NarrowSorted is added by this PR and holds non-decreasing values,
the shape DELTA_BINARY_PACKED is usually chosen for.

benchmark #51250 this PR vs. main
Decode_Int32_NarrowSorted 73.7 us 45.7 us 1.61x 2.18x
Decode_Int32_Narrow 74.8 us 46.4 us 1.61x 2.15x
Decode_Int32_Wide 83.7 us 53.6 us 1.56x 1.90x
Decode_Int64_NarrowSorted 71.4 us 55.0 us 1.30x 1.47x
Decode_Int64_Narrow 70.9 us 55.2 us 1.28x 1.49x
Decode_Int64_Wide 318.6 us 301.3 us 1.06x 1.07x
Decode_Int32_Fixed 20.4 us 20.5 us 1.00x 0.97x
Decode_Int64_Fixed 31.3 us 31.3 us 1.00x 0.98x

Are there any user-facing changes?

No. No API change, no format change, and decoded values are identical.

@prtkgaur
prtkgaur force-pushed the gh51268-dbp-vector-prefix-sum branch 6 times, most recently from 83a5289 to 5f91b9b Compare September 18, 2026 21:09
The output buffer may alias decoder members of the same type, forcing repeated
reloads after stores. Keep the frame and running value in locals, write back
once, and preserve unsigned wrapping during reconstruction.
@prtkgaur
prtkgaur force-pushed the gh51268-dbp-vector-prefix-sum branch from 8692d8c to 4e3aaeb Compare October 4, 2026 00:13
Adjacent equal-width miniblocks form one packed stream. Decode each run with one
unpack call, bounded by the current block and output request; retain the
zero-width path and cover width changes, partial reads, block boundaries, and
both integer widths.
Replace the dependent scalar additions with an inclusive SIMD scan for batches
of at least four lanes. Retain scalar handling for narrower batches and tails,
preserve unsigned wrapping, and cover every residual width, both block sizes,
and wrapped totals.
@prtkgaur
prtkgaur force-pushed the gh51268-dbp-vector-prefix-sum branch from 4e3aaeb to 3e46ef0 Compare October 4, 2026 02:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants