Skip to content

Make strided AxisArray data contiguous before serializing - #269

Draft
cboulay wants to merge 1 commit into
cboulay/json-schema-settings-stackfrom
cboulay/contiguous-before-serialize
Draft

cboulay wants to merge 1 commit into
cboulay/json-schema-settings-stackfrom
cboulay/contiguous-before-serialize

Conversation

@cboulay

@cboulay cboulay commented Sep 4, 2026 •

Copy link
Copy Markdown
Member

Restacked: now based on #271 (v3.10.0b4) rather than dev, as the first PR after the b4 release in the stack for the next beta. The only conflict was in tests/messages/test_axisarray.py (both sides appended tests).

Draft, same as #268 — parking it for review rather than merge. Independent of #268; they touch different files and can land in either order.

The problem

numpy hands a C- or F-contiguous array to a protocol-5 pickler out-of-band, and ezmsg's marshal writes those buffers straight into shared memory without copying. An array that is neither falls back to an in-band copy through the pickle stream.

That matters because it is the one size-dependent cost in serialization. Everything else is a flat ~7 us regardless of payload. Any transformer emitting a decimated or channel-sliced view (data[:, ::2], data[::2, :]) hits this silently, and it scales with the array.

The fix

ArrayWithNamedDims.__getstate__ makes data contiguous when it is neither C- nor F-contiguous, which puts the payload back on the out-of-band path.

layout before after out-of-band buffers
C-contiguous (4 MB) 6.5 us 7.1 us 1 → 1
F-order transposed 6.6 us 6.8 us 1 → 1
strided [:, ::2] (2 MB) 311.9 us 141.5 us 0 → 1
strided [::2, :] (2 MB) 253.5 us 41.1 us 0 → 1

In-band bytes for the strided cases drop from 2,097,501 to 307.

Why this is transparent, not a behaviour change

Unpickling a strided array already yields a C-contiguous one — numpy's reduce does tobytes() on the way out. So the receiver sees exactly the same array either way; it just gets there for a lot less. Verified explicitly in the tests.

Why F-order is excluded

This is the part worth a second pair of eyes. F-contiguous arrays are already on the out-of-band path, and forcing C order on them measured ~80x worse (7 us → 582 us on a 4 MB transposed array). A blanket np.ascontiguousarray would have been a serious pessimization on a common case — a transposed AxisArray is not unusual. The guard is not (c_contiguous or f_contiguous), costing two flag reads on the common path.

Scope

  • Only numpy arrays are touched. torch/cupy have their own serialization and no flags attribute to consult.
  • Lives on ArrayWithNamedDims, the shared base, so CoordinateAxis gets it too — a ch axis with a strided data hit the same cliff.
  • The caller's array is never mutated; the copy goes into the returned state dict only.

Testing

7 new tests: all four layouts round-trip out-of-band with correct values, F-order is confirmed not reordered, the source message is confirmed unmutated, and CoordinateAxis is covered.

Full suite green: 431 passed, 1 skipped. The 2 remaining ruff errors in these files (E731 at axisarray.py, E402 in the test) are pre-existing on dev — verified against the pristine versions.

Measurements are single-machine (Darwin, arm64, Python 3.13).

🤖 Generated with Claude Code

numpy hands a C- or F-contiguous array to a protocol-5 pickler out-of-band,
and ezmsg's marshal writes those buffers straight into shared memory without
copying them. An array that is neither falls back to an in-band copy through
the pickle stream. Unlike the rest of serialization -- which is a flat ~7us
regardless of payload -- that copy scales with size, and any transformer
emitting a decimated or channel-sliced view hits it silently.

ArrayWithNamedDims.__getstate__ now materializes a contiguous copy in that
case, which puts the payload back on the out-of-band path:

    layout                    before     after   out-of-band buffers
    C-contiguous (4 MB)       6.5us      7.1us         1 -> 1
    F-order transposed        6.6us      6.8us         1 -> 1
    strided [:, ::2] (2 MB) 311.9us    141.5us         0 -> 1
    strided [::2, :] (2 MB) 253.5us     41.1us         0 -> 1

In-band bytes for the strided cases drop from 2,097,501 to 307.

This is transparent rather than a behaviour change: unpickling a strided
array already yields a C-contiguous one, so the receiver sees exactly the
same array either way.

F-contiguous arrays are deliberately excluded. They are already out-of-band,
and forcing C order on them measured ~80x worse -- a blanket
ascontiguousarray would have been a serious pessimization. The guard costs
two flag reads on the common path.

Only numpy arrays are touched; torch/cupy arrays have their own
serialization and no `flags` to consult. The fix lives on the shared base
class, so CoordinateAxis gets it too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@cboulay
cboulay force-pushed the cboulay/contiguous-before-serialize branch from db328dd to 50d477f Compare October 1, 2026 16:22
cboulay added a commit that referenced this pull request Oct 1, 2026
Stacked on #269, ArrayWithNamedDims now defines __getstate__ (making strided
data contiguous before pickling), and returning self.__dict__ here bypassed
it for coordinate axes. super().__getstate__() reaches it through the MRO,
so it is also safe on Python 3.10, where object has no __getstate__.
@cboulay
cboulay changed the base branch from dev to cboulay/json-schema-settings-stack October 1, 2026 16:22

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant