Repository navigation
Conversation
…ting it zmwatch's startup grace counted from zmwatch's own start, so a zmc started later (a monitor added or saved through the web UI or API, a zmdc restart) had none. When a zmwatch pass landed in the seconds between that zmc being forked and it mapping its shared memory, zmwatch logged "shared data not valid" and asked zmdc to restart it, replacing a zmc that was about to come up. Count ZM_WATCH_MAX_DELAY per monitor from the pass that first found its shared data missing or invalid, and forget it once the data is valid again or after the restart. The zmwatch-start grace and the heartbeat checks are unchanged, so a zmc that hangs after mapping its memory is restarted as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zmdc forks each daemon and, in the child, connects to the database, closes every descriptor up to _SC_OPEN_MAX and only then resets TERM, INT and ABRT to their defaults before exec. Until then the child runs zmdc's own shutdown handler, which only sets $zm_terminate in the child, so a TERM sent in that window is lost. The daemon then starts as if nothing happened, zmdc waits KILL_DELAY (30 s) and KILLs it, and an event that zmc opened in the meantime is left with no EndDateTime. zmwatch's restart of a zmc that hasn't mapped its shared memory yet lands in this window. Closing descriptors one by one takes about 1.7 s with a 1048576 open-file limit (Docker's default), which makes the window wide. Block TERM, INT and ABRT across the fork along with SIGCHLD, and in the child reset them to their defaults and unblock them first thing, so a stop sent before exec ends the child and the reaper restarts or forgets it as asked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The child closed every descriptor number from 3 up to _SC_OPEN_MAX one by one. With a 1048576 open-file limit (Docker's default) that is about 1.7 s of close() calls before each daemon starts, and it is the window in which zmwatch can find a just-started zmc without shared memory. Close the descriptors /proc/self/fd lists instead, and keep the old loop where /proc isn't available. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
ZM_WATCH_MAX_DELAYfrom when it first finds its shared data missing. zmdc's child resets the stop signals before anything else, so a TERM before exec ends it. The child closes only open descriptors. 2 files, +44/−12.start(). zmwatch now restarts a zmc that has no shared memory after 45 s instead of at the next pass. Checked with the script tests, smoke and the event checks.Details: reproduction, change, checks
Before
Upstream master
6c0f976c2, clean build:Change
scripts/zmwatch.pl.in: a per-monitor%no_shared_data_since. A monitor without valid shared data is restarted onlyZM_WATCH_MAX_DELAYafter the first pass that saw it; the timer clears when the data is valid or after the restart. The startup grace and heartbeat checks are unchanged. The Info line addsfor <N>s.scripts/zmdc.pl.instart(): TERM, INT and ABRT are blocked across the fork with SIGCHLD. The child sets them to DEFAULT and unblocks them before connecting to the database, so a stop sent before exec ends it and the reaper restarts or forgets it as asked.scripts/zmdc.pl.instart(): close the fds listed in/proc/self/fdinstead of every number up to_SC_OPEN_MAX; the old loop stays where /proc is missing.tests/perl, because it needs a running zmdc; covered by the lab probe below.Behaviour changes
ZM_WATCH_MAX_DELAYbefore restarting a monitor with no shared datashared data not valid for 50sRestarting capture daemon ..., shared data not validline gainsfor <N>sexited, signal 14) and the restart follows at onceSide fixes in this PR: the descriptor loop (the
perf:commit).After
Branch
fix/b1b2atcfdbc4312(base6c0f976c2); build.json headcfdbc43125, dirty False.With only the signal commit (fd loop still slow), the TERM lands before exec and the child goes at once:
Not run: clients (no API or shared-memory change). Backport to release-1.38: the two zmdc commits apply cleanly; zmwatch needs an adapted commit (1.38 has a fixed START_DELAY sleep instead of
$starting_up), prepared but not run.refs #5198
🤖 Generated with Claude Code