Conversation
cboulay
added a commit
that referenced
this pull request
Oct 1, 2026
…272) What remains of the #272 work after the uncontroversial parts moved to #275, kept here for review of the policy it implies. - Channel: a crashed publisher connection records its exception and notifies its clients; Subscriber.recv_zero_copy raises ChannelFailed (chained from the cause) instead of waiting forever. handle_subscriber logs it at ERROR and keeps serving the stream's other publishers. - GraphServer: a segment whose last lease has ended is marked unlinked and forgotten, so a late SHM_ATTACH is refused (ValueError, which channels already treat as a stale generation) instead of handing back a name the client cannot open. Unlinked entries are also pruned on SHM_CREATE, and a segment is never unlinked twice. - Channel: FileNotFoundError on attach is treated as stale too, for older GraphServers.
cboulay
added a commit
that referenced
this pull request
Oct 1, 2026
…272) What remains of the #272 work after the uncontroversial parts moved to #275, kept here for review of the policy it implies. - Channel: a crashed publisher connection records its exception and notifies its clients; Subscriber.recv_zero_copy raises ChannelFailed (chained from the cause) instead of waiting forever. handle_subscriber logs it at ERROR and keeps serving the stream's other publishers. - GraphServer: a segment whose last lease has ended is marked unlinked and forgotten, so a late SHM_ATTACH is refused (ValueError, which channels already treat as a stale generation) instead of handing back a name the client cannot open. Unlinked entries are also pruned on SHM_CREATE, and a segment is never unlinked twice. - Channel: FileNotFoundError on attach is treated as stale too, for older GraphServers.
cboulay
marked this pull request as draft
October 1, 2026 15:58
Split from #273: the parts that stand on their own, without the ChannelFailed policy that is still under discussion. - SHMContext: if closing the mapping raises BufferError because views into it are still alive (a unit kept a zero-copy array), drop the handle and let the mmap be unmapped when the last view is collected, instead of killing the channel when the publisher grows its segment. - SHMContext.wait_closed() also closes the segment: a monitor task cancelled before it first ran never executed its finally, leaking the mapping (the "Exception ignored in SharedMemory.__del__" noise). - Channel: SHM-delivered messages are read-only, matching the TCP path, so a subscriber cannot write into memory shared with the publisher and every other subscriber. - Tests: closing with a live view, and a cross-process grow with a receiver that retains every message.
On a grow the publisher closed the old segment at once. A message already sent names that segment, and a channel that has not processed it yet still has to attach it; if the publisher's lease was the last one, the GraphServer unlinked it first and the channel's attach failed with FileNotFoundError, killing the channel. Two quick grows (num_buffers > 1, no wait for acks between them) lost that race whenever timing shifted. The publisher now retires the old segment with the last msg_id sent under its name, and closes it once the backpressure wait for msg_id L + num_buffers is over: by then every channel has released message L, and a channel attaches the segment a message names before it can release it.
cboulay
force-pushed
the
cboulay/fingerprint-on-pickle
branch
from
October 1, 2026 16:22
014f440 to
8f43997
Compare
cboulay
force-pushed
the
cboulay/shm-hardening
branch
from
October 1, 2026 16:22
bb120ed to
d71704f
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #274. Split from #273: the parts of the #272 fix that stand on their own. The
ChannelFailedpolicy and the dead-segment race stay in draft #273.Changes
shm.py): ifSharedMemory.close()raisesBufferErrorbecause views into the segment are still alive (a unit kept a zero-copy array),SHMContextdrops its handle to the mmap and closes the fd. The mapping is unmapped when the last view is collected, and the GraphServer still unlinks the segment when its leases end. Before this, the channel died for good when the publisher grew its SHM. Relies onSharedMemory's private_mmap/_fd;close()is unchanged from Python 3.8 through 3.14.SHMContext.wait_closed()also closes the segment. A monitor task cancelled before it first ran never executed itsfinally(the "Exception ignored inSharedMemory.__del__" noise at shutdown).FileNotFoundError. Two quick grows lost this race whenever timing shifted (it broke Surface a dead channel to subscribers; refuse attaching unlinked SHM (#272) #273's race test, and later the existing grow test under Elide coordinate axes a receiving channel already holds #276). The publisher now retires the old segment with the lastmsg_idsent under its name and closes it once the backpressure wait formsg_id L + num_buffersis over. By then every channel has released message L, and a channel attaches the segment a message names before it can release it.shm[buf_idx].toreadonly(), matching the TCP path. Subscribers could previously write into memory shared with the publisher and every other subscriber.Behaviour change
Editing a received array in place (
msg.data *= 2) now fails with "assignment destination is read-only" on SHM deliveries, as it already did over TCP. Worth a release note for the beta.Not fixed here
A message kept past its callback is still overwritten when the publisher reuses its slot. This PR only stops the crash; units that keep received data must copy it. The baseproc hash witness, which did this implicitly, is fixed in ezmsg-org/ezmsg-baseproc#18.
Tests
test_shm.py::test_close_with_live_viewtest_shm_grow.py::test_cross_process_grow_with_retained_views: a receiver keeps every message across two grows. Without the fix it hangs.Full suite against a private GraphServer: only the 10
test_settings_json_schema/test_inspect_commandfailures that also fail on the base branch.test_shm_resize_race_repro_completespasses, and the grow tests passed 8/8 under the timing that used to lose the race.Cost: the read-only view is ~20 ns/msg; the tolerant close runs only when a segment is closed.