Skip to content

Avoid zeroing full NumPy input blocks in miniexpr gather - #728

Merged
FrancescAlted merged 1 commit into
Blosc:mainfrom
Johnny-Kao:perf/skip-full-numpy-block-zeroing
Sep 29, 2026
Merged

FrancescAlted merged 1 commit into
Blosc:mainfrom
Johnny-Kao:perf/skip-full-numpy-block-zeroing

Conversation

@Johnny-Kao

Copy link
Copy Markdown
Contributor

TL;DR

  • Why: raw NumPy inputs in the miniexpr gather path zero-initialize every newly allocated block buffer before copying data into it. For full blocks, the subsequent gather overwrites every byte, making that memset redundant.
  • Change: skip zero-initialization when valid_nitems == blocknitems. Partial blocks keep the existing zero-fill behavior so padding remains unchanged.
  • Performance: repeated A/B/A/B measurements on an Apple M5 showed about 2–15% lower execution time in the tested multi-input full-block workloads, with the largest gain for small blocks. A separate partial-heavy workload showed no measurable regression.

What was happening

For each raw NumPy operand, the current path effectively does:

malloc(block)
memset(block, 0)
gather NumPy data into block
evaluate expression

For a full block, the gather writes the entire destination buffer, so the preceding zero-fill performs memory writes that are immediately overwritten.

Change

The zero-fill is now conditional:

if valid_nitems < blocknitems:
    memset(input_buffers[i], 0, block_nbytes)

Full blocks skip the redundant initialization. Partial blocks still take the existing zero-filled path because unwritten padding must remain zero.

No expression evaluation, arithmetic, dtype handling, input validation, or non-NumPy input path is changed.

Performance validation

I benchmarked the change using repeated:

baseline → patched → baseline → patched

builds rather than relying on a single before/after run.

The main measurements used an Apple M5 MacBook Air and 4 threads.

Workload A1 → B1 A2 → B2
2 NumPy operands, full blocks 5.1% faster 8.9% faster
4 operands, small blocks 14.4% faster 15.0% faster
8 operands, micro blocks 5.3% faster 2.2% faster

These numbers apply specifically to the raw NumPy miniexpr input-gather path; they are not intended as an overall miniexpr performance claim.

Partial-heavy regression check

I also tested a deliberately partial-heavy 2-D layout three times using paired baseline/patched runs:

Run Baseline Patched
1 143.851 ms 143.308 ms
2 141.980 ms 141.916 ms
3 142.522 ms 142.385 ms

This workload is effectively performance-neutral, which is consistent with the optimization: partial blocks still require zero-initialization, so most of the work removed for full blocks remains necessary.

For the tested 65×65 chunk / 64×64 block layout, a chunk produces approximately:

64×64   64×1
 1×64    1×1

Only the 64×64 block can skip initialization. The other three blocks still contain padding and therefore retain the zero-fill.

I also considered zeroing only the exact padding bytes in partial blocks. For this geometry, however, that would remove very little additional memory traffic while requiring more complicated multidimensional and row-level handling, so I did not include it.

Why this is safe

valid_nitems is already computed for the current block before input processing.

When valid_nitems == blocknitems, the raw NumPy gather covers the complete block buffer before it is passed to the evaluator. The observable buffer contents are therefore the same:

before: zero entire buffer → overwrite every byte
after:  overwrite every byte

For partial blocks, the existing zero-fill remains unchanged.

The change introduces no shared mutable state or reusable scratch buffer, and it does not alter arithmetic, dtype conversion, input acceptance, or expression semantics.

Alternatives considered

Zero only the padding of partial blocks

This could theoretically reduce some additional writes, but partial padding is often fragmented across rows in multidimensional blocks. That would require more branches or multiple small memset operations for very little expected benefit in the tested partial-heavy geometry.

Not included because the added complexity is disproportionate to the likely gain.

Reuse input scratch buffers

This could also reduce allocation overhead, but it introduces buffer ownership, resizing, lifetime, and concurrency considerations. That is a substantially larger change and is intentionally outside the scope of this patch.

Avoid staging and use direct input pointers

Potentially a larger optimization, but it would change more assumptions around layout, lifetime, and evaluator inputs. Also outside the scope of this patch.

The current change is intentionally limited to removing work that is provably redundant.

Validation

Relevant ndarray/miniexpr tests:

1596 passed

Additional stress testing covered signed and unsigned integer dtypes, float32/float64, complex64/complex128, NaN/infinities/subnormal-adjacent values/negative zero, 1-D/2-D/3-D shapes, full and padded blocks, multiple thread counts, and concurrent evaluations.

Full local test suite:

10714 passed, 39 skipped, 2 failed

The two float32 sinh / cosh failures were reproduced unchanged on the unmodified baseline.

This patch introduced 0 new test failures.

@FrancescAlted
FrancescAlted merged commit 3ce2a39 into Blosc:main Sep 29, 2026
36 checks passed
@FrancescAlted

Copy link
Copy Markdown
Member

Nice patch! Thanks @Johnny-Kao !

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants