Avoid zeroing full NumPy input blocks in miniexpr gather - #728
Merged
FrancescAlted merged 1 commit intoSep 29, 2026
Merged
Conversation
Member
|
Nice patch! Thanks @Johnny-Kao ! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
memsetredundant.valid_nitems == blocknitems. Partial blocks keep the existing zero-fill behavior so padding remains unchanged.What was happening
For each raw NumPy operand, the current path effectively does:
For a full block, the gather writes the entire destination buffer, so the preceding zero-fill performs memory writes that are immediately overwritten.
Change
The zero-fill is now conditional:
Full blocks skip the redundant initialization. Partial blocks still take the existing zero-filled path because unwritten padding must remain zero.
No expression evaluation, arithmetic, dtype handling, input validation, or non-NumPy input path is changed.
Performance validation
I benchmarked the change using repeated:
builds rather than relying on a single before/after run.
The main measurements used an Apple M5 MacBook Air and 4 threads.
These numbers apply specifically to the raw NumPy miniexpr input-gather path; they are not intended as an overall miniexpr performance claim.
Partial-heavy regression check
I also tested a deliberately partial-heavy 2-D layout three times using paired baseline/patched runs:
This workload is effectively performance-neutral, which is consistent with the optimization: partial blocks still require zero-initialization, so most of the work removed for full blocks remains necessary.
For the tested
65×65chunk /64×64block layout, a chunk produces approximately:Only the
64×64block can skip initialization. The other three blocks still contain padding and therefore retain the zero-fill.I also considered zeroing only the exact padding bytes in partial blocks. For this geometry, however, that would remove very little additional memory traffic while requiring more complicated multidimensional and row-level handling, so I did not include it.
Why this is safe
valid_nitemsis already computed for the current block before input processing.When
valid_nitems == blocknitems, the raw NumPy gather covers the complete block buffer before it is passed to the evaluator. The observable buffer contents are therefore the same:For partial blocks, the existing zero-fill remains unchanged.
The change introduces no shared mutable state or reusable scratch buffer, and it does not alter arithmetic, dtype conversion, input acceptance, or expression semantics.
Alternatives considered
Zero only the padding of partial blocks
This could theoretically reduce some additional writes, but partial padding is often fragmented across rows in multidimensional blocks. That would require more branches or multiple small
memsetoperations for very little expected benefit in the tested partial-heavy geometry.Not included because the added complexity is disproportionate to the likely gain.
Reuse input scratch buffers
This could also reduce allocation overhead, but it introduces buffer ownership, resizing, lifetime, and concurrency considerations. That is a substantially larger change and is intentionally outside the scope of this patch.
Avoid staging and use direct input pointers
Potentially a larger optimization, but it would change more assumptions around layout, lifetime, and evaluator inputs. Also outside the scope of this patch.
The current change is intentionally limited to removing work that is provably redundant.
Validation
Relevant ndarray/miniexpr tests:
Additional stress testing covered signed and unsigned integer dtypes, float32/float64, complex64/complex128, NaN/infinities/subnormal-adjacent values/negative zero, 1-D/2-D/3-D shapes, full and padded blocks, multiple thread counts, and concurrent evaluations.
Full local test suite:
The two float32
sinh/coshfailures were reproduced unchanged on the unmodified baseline.This patch introduced 0 new test failures.