Skip to content

Declare chunk_dim and prime the channel fingerprint - #5

Merged
cboulay merged 6 commits into
devfrom
cboulay/chunk-dim-and-fingerprint
Sep 4, 2026
Merged

cboulay merged 6 commits into
devfrom
cboulay/chunk-dim-and-fingerprint

Conversation

@cboulay

@cboulay cboulay commented Sep 3, 2026 •

Copy link
Copy Markdown
Member

Two things every consumer of a replayed file needs, and only the reader can supply. Both are set once per stream; neither costs anything per message.

Two commits. The first is ruff-format only: iter.py predates the repo's current ruff config and had never been run through it, so the pre-commit hook reflows it wholesale on first touch. Separated so the second commit reads as the 18 lines it actually is.

chunk_dim

A stateful consumer caches state — filter coefficients, per-channel history, resolved channel indices — against the stream's configuration: channel count, labels, sample rate. It has to exclude the one dimension whose length is just however much of the file this chunk covered.

It cannot reliably infer which that is. It's time here, but win downstream of a windowing stage, so a consumer that assumes time either thrashes on chunk-size jitter or, worse, stops noticing a real change:

post-Window (win, time, ch)   0=first  1,2=win-count jitter  3=relabel  4=same  5=window-len change
  consumer assumes "time":  resets at [0, 1, 2, 3]
  message declares "win":   resets at [0, 3, 5]      correct

Only the producer, which named the dims, knows. Both iterators declare it, and it holds whether the stream is regular or carries per-sample timestamps. Both emit paths already go through replace(template, ..., axes={**template.axes, "time": ...}), so fast_replace carries the field for free.

CoordinateAxis.fingerprint

A content digest, computed on first access and cached on the instance — it's what lets a consumer notice that channels were relabelled at a fixed channel count, which is otherwise silent and numerically destructive:

first 4 samples of the new channel 'armB-1':
  filter state carried from armA: [-5.919  -4.772  -0.5114  1.064]
  with a correct reset:           [ 0.     -0.0013 -0.0023 -0.0026]
  max |difference| = 11.12   vs new-data amplitude 0.024

Each template builds its ch axis once and every message reuses that object, so priming costs one checksum per stream. Left cold it's computed by the first stateful consumer in this process — and, because unpickling builds a new axis object per message, by the first consumer in every other process, on every message.

Testing

20 passed (was 1 placeholder). A third commit adds the XDF fixture this repo never had, because tests/test_iter.py was still test_dummy with its TODO and nothing here was covered.

tests/create_test_xdf.py writes the XDF chunked format directly, with each field pinned against what pyxdf's reader actually consumes. Generated rather than checked in as a binary, so a test needing a different stream layout changes an argument instead of asking someone for a new recording. Three choices are deliberate:

  • every sample carries an explicit timestamp rather than relying on delta decompression, which encodes "same as last plus 1/srate" and would make the fixture silently agree with any reader that got the nominal rate wrong;
  • streams start at t=10 s, not 0, so a reader honouring rezero is distinguishable from one ignoring it;
  • ClockOffset chunks are written despite there being no skew to model, because without them pyxdf logs Segments and clock-segments differ on every load.

Beyond the two fields above, the tests cover sample values and ordering, chunk_dur, rezero, the nominal rate reaching the axis gain, per-sample timestamps on the irregular stream, and force_single_sample. A final class checks the fixture itself, since a wrong fixture would make every other assertion agree with the wrong thing.

Verified by mutation rather than by passing: dropping chunk_dim fails 3 tests, dropping the fingerprint priming fails 3, and rebuilding the channel axis per message instead of reusing it fails 2.

Notes

  • src/ezmsg/xdf/source.py still fails ruff on trailing whitespace. Pre-existing and unrelated, so left alone.
  • Dependency: ezmsg>=3.10.0b2 for both fields. This is a pre-release pin until 3.10.0 final ships.

🤖 Generated with Claude Code


Update: the tests now actually run in CI

Adding tests exposed that this repo never ran any. Three commits follow from that, all pre-existing defects that were invisible while the workflow was dormant:

  1. The triggers were commented out. push and pull_request were both disabled, leaving workflow_dispatch as the only route — reasonable when the suite was one test_dummy, not now. Every sibling package triggers on push to main and on pull requests; blackrock, neo and nwb also include dev, which is where these PRs are based, so that's matched here.

  2. setup-uv was keyed on a lock file that doesn't exist. cache-dependency-glob: "uv.lock", but no uv.lock is committed here — nor in any sibling package. The glob matched nothing and the action failed the job before installing a single dependency:

    ##[error]No file in /home/runner/work/ezmsg-xdf/ezmsg-xdf matched to [uv.lock]
    

    ezmsg-lsl already keys on pyproject.toml; this matches it.

  3. source.py:65 had trailing whitespace that fails the lint step. I'd noted this as pre-existing and left it alone; with CI on it blocks every matrix entry, so it's fixed. Touching the file meant the pre-commit hook reflowed it, same as iter.py — both are in that one commit since the fix can't be staged without the reformat.

13 checks passing across Python 3.10–3.13 on Linux, macOS and Windows.

The file predates the repo's current ruff config and had never been run
through it, so the pre-commit hook reflows it wholesale on first touch.
Separated here so the change that follows is readable.
Two things every consumer of a replayed file needs, and only the reader can
supply.

`chunk_dim` names the dimension messages accumulate along. A consumer caches
state against the stream's configuration -- channel count, labels, sample rate
-- and must exclude the one dimension whose length is just however much of the
file this chunk covered. It cannot reliably infer which that is: it is `time`
here but `win` downstream of a windowing stage, so a guess either thrashes on
chunk-size jitter or stops noticing real changes. Both iterators declare it,
and it holds whether the stream is regular or carries per-sample timestamps.

`CoordinateAxis.fingerprint` is a content digest, computed on first access and
cached on the instance. Each template builds its `ch` axis once and every
message reuses that object, so priming costs one checksum per stream. Left
cold it is computed by the first stateful consumer in this process -- and,
because unpickling builds a new axis object per message, by the first consumer
in every other process, on every message.

Unverified by tests: this repo has no XDF fixture, only the placeholder in
tests/test_iter.py. The change mirrors ezmsg-neo and ezmsg-nwb, where it is
covered.

Requires ezmsg 3.10.0b2 for both fields.
`tests/test_iter.py` was a `test_dummy` placeholder with a TODO asking for a
small XDF source, so nothing in this package was covered -- including the
chunk_dim and fingerprint change in the previous commit.

The fixture is generated rather than checked in as a binary, so it stays
readable and adjustable: a test needing an irregular stream, a string stream or
a different channel layout changes an argument instead of asking someone to
produce a new recording. `create_test_xdf.py` writes the XDF chunked format
directly, with each field pinned against what pyxdf's reader actually consumes
-- `_read_varlen_int`, `_read_chunk3` and the tag dispatch in `load_xdf` --
since that is the only reader these files ever meet.

Three choices in the generator are deliberate and would otherwise look
arbitrary:

* Every sample carries an explicit timestamp rather than relying on delta
  decompression. The compressed form encodes "same as last plus 1/srate",
  which would make the fixture silently agree with any reader that got the
  nominal rate wrong.
* Streams start at t=10 s, not 0, so a reader honouring `rezero` is
  distinguishable from one ignoring it.
* ClockOffset chunks are written even though there is no clock skew to model.
  Without them pyxdf logs "Segments and clock-segments differ" on every load,
  and a fixture that warns every time trains readers to ignore warnings.

The default file has two streams -- a 100 Hz 4-channel float32 ramp and an
irregular string marker stream -- split across several sample chunks so the
reader's chunk stitching is exercised rather than arriving as one block.

20 tests cover sample values and ordering, chunk_dur, rezero, the nominal rate
reaching the axis gain, per-sample timestamps on the irregular stream,
force_single_sample, and both fields the previous commit added. A class at the
end checks the fixture itself, since a wrong fixture would make every other
assertion agree with the wrong thing.

Verified by mutation rather than by the tests merely passing: dropping
chunk_dim fails 3, dropping the fingerprint priming fails 3, and rebuilding
the channel axis per message instead of reusing it fails 2.
The push and pull_request triggers were commented out, leaving
workflow_dispatch as the only way to run the suite. That was reasonable when
the suite was a single `test_dummy` placeholder, and is not now that it covers
the iterators, the producers and both units.

Every other ezmsg source package triggers on push to main and on pull requests;
blackrock, neo and nwb also include dev, which is where these PRs are based, so
that is matched here. Without it a PR to dev reports only the publish
workflow's build job, which does not run a single test.
setup-uv is configured with `cache-dependency-glob: "uv.lock"`, but no
uv.lock is committed here -- nor in any sibling ezmsg package. The glob matches
nothing and the action fails the job before a single dependency is installed:

    ##[error]No file in /home/runner/work/ezmsg-xdf/ezmsg-xdf matched to
    [uv.lock], make sure you have checked out the target repository

Latent since the workflow was written, and invisible until the previous commit
turned the triggers on. ezmsg-lsl already keys on pyproject.toml; this matches
it.
ruff has flagged W291 on line 65 since before any of this work. It went
unnoticed because the test workflow, which runs the lint step, was never
triggered; with the triggers on it fails every matrix entry.

Fixing it means touching the file, and this one predates the repo's current
ruff config just as iter.py did, so the pre-commit hook reflows it wholesale on
first touch. Both are in one commit here because the whitespace fix cannot be
staged without the reformat.
@cboulay
cboulay merged commit 2cdcd2d into dev Sep 4, 2026
14 checks passed
@cboulay
cboulay deleted the cboulay/chunk-dim-and-fingerprint branch September 4, 2026 06:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant