Skip to content

fix: keep zmwatch from restarting a starting zmc and stop zmdc losing the stop - #5199

Open
nabbi wants to merge 3 commits into
ZoneMinder:masterfrom
nabbi:fix/b1b2-zmwatch-zmdc-stop
Open

nabbi wants to merge 3 commits into
ZoneMinder:masterfrom
nabbi:fix/b1b2-zmwatch-zmdc-stop

Conversation

@nabbi

@nabbi nabbi commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Automated PR, tested in a zm-holodeck sandbox. An AI agent (Claude Opus 5.5) reproduced the bug on a disposable ZoneMinder built from upstream master 6c0f976c2, wrote this fix and ran the checks below on the same lab. Minimal human review before opening: please review it as you would any contributor's PR.

TL;DR

  • Fixes: zmwatch restarts a zmc that is still starting, and zmdc loses the stop and KILLs it 30 s later, leaving its event open #5198: zmwatch restarts a zmc that is still starting; zmdc loses that stop, KILLs the zmc 30 s later and its event stays open.
  • Change: zmwatch gives each monitor ZM_WATCH_MAX_DELAY from when it first finds its shared data missing. zmdc's child resets the stop signals before anything else, so a TERM before exec ends it. The child closes only open descriptors. 2 files, +44/−12.
  • Risk: every daemon zmdc starts goes through the changed start(). zmwatch now restarts a zmc that has no shared memory after 45 s instead of at the next pass. Checked with the script tests, smoke and the event checks.
  • Tested: reproduced before, passes after; the checks covering the diff are green apart from one known XFAIL.
Details: reproduction, change, checks

Before

Upstream master 6c0f976c2, clean build:

zmdc[220]: INF ['zmc -m 9' sending stop to pid 3995 at 26/10/10 18:15:27]
zmc_m9[3995]: INF [zmc_m9] [Starting Capture version 1.39.40]
zmdc[220]: WAR ['zmc -m 9' has not stopped at 26/10/10 18:15:58 after 30 seconds. Sending KILL to pid 3995]

Change

  • scripts/zmwatch.pl.in: a per-monitor %no_shared_data_since. A monitor without valid shared data is restarted only ZM_WATCH_MAX_DELAY after the first pass that saw it; the timer clears when the data is valid or after the restart. The startup grace and heartbeat checks are unchanged. The Info line adds for <N>s.
  • scripts/zmdc.pl.in start(): TERM, INT and ABRT are blocked across the fork with SIGCHLD. The child sets them to DEFAULT and unblocks them before connecting to the database, so a stop sent before exec ends it and the reaper restarts or forgets it as asked.
  • scripts/zmdc.pl.in start(): close the fds listed in /proc/self/fd instead of every number up to _SC_OPEN_MAX; the old loop stays where /proc is missing.
  • Test: none in tests/perl, because it needs a running zmdc; covered by the lab probe below.

Behaviour changes

Change Who sees it Intended / collateral Alternatives looked into Chosen, why Tested by
zmwatch waits ZM_WATCH_MAX_DELAY before restarting a monitor with no shared data admins: a zmc stopped by hand, or stuck before mapping memory, comes back after about 45–55 s instead of about 10 s intended ask zmdc for the process start time (new IPC); Monitor_Status (written before the shared memory exists) per-monitor timer, no new IPC stop a zmc through zmdc: shared data not valid for 50s
the Restarting capture daemon ..., shared data not valid line gains for <N>s log readers intended — prefix kept events check
a TERM before exec ends zmdc's child (exited, signal 14) and the restart follows at once zmdc log intended re-send TERM after exec block + reset, no timing guess start + restart 0.2 s later
daemons start about 1.7 s sooner with a high open-file limit none side fix — — all daemons start in smoke and events

Side fixes in this PR: the descriptor loop (the perf: commit).

After

Branch fix/b1b2 at cfdbc4312 (base 6c0f976c2); build.json head cfdbc43125, dirty False.

zmdc[220]: INF ['zmc -m 10' sending stop to pid 929 at 26/10/10 18:27:56]
zmdc[220]: INF ['zmc -m 10' exited normally]

With only the signal commit (fd loop still slow), the TERM lands before exec and the child goes at once:

zmdc[220]: INF ['zmc -m 26' sending stop to pid 435 at 26/10/10 19:01:51]
zmdc[220]: INF ['zmc -m 26' exited, signal 14]
Check Result
holodeck checks for the diff (script-tests, smoke, events) script-tests PASS=28; smoke PASS=11 XFAIL=1 (stock zmonvif-probe.pl picks the interface after connecting, known); events PASS=10
zmwatch restarts / zmdc KILLs during that run 0 / 0

Not run: clients (no API or shared-memory change). Backport to release-1.38: the two zmdc commits apply cleanly; zmwatch needs an adapted commit (1.38 has a fixed START_DELAY sleep instead of $starting_up), prepared but not run.

refs #5198

🤖 Generated with Claude Code

nabbi and others added 3 commits October 10, 2026 12:26
…ting it

zmwatch's startup grace counted from zmwatch's own start, so a zmc started
later (a monitor added or saved through the web UI or API, a zmdc restart)
had none. When a zmwatch pass landed in the seconds between that zmc being
forked and it mapping its shared memory, zmwatch logged "shared data not
valid" and asked zmdc to restart it, replacing a zmc that was about to come
up.

Count ZM_WATCH_MAX_DELAY per monitor from the pass that first found its
shared data missing or invalid, and forget it once the data is valid again
or after the restart. The zmwatch-start grace and the heartbeat checks are
unchanged, so a zmc that hangs after mapping its memory is restarted as
before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zmdc forks each daemon and, in the child, connects to the database, closes
every descriptor up to _SC_OPEN_MAX and only then resets TERM, INT and ABRT
to their defaults before exec. Until then the child runs zmdc's own
shutdown handler, which only sets $zm_terminate in the child, so a TERM sent
in that window is lost. The daemon then starts as if nothing happened, zmdc
waits KILL_DELAY (30 s) and KILLs it, and an event that zmc opened in the
meantime is left with no EndDateTime.

zmwatch's restart of a zmc that hasn't mapped its shared memory yet lands
in this window. Closing descriptors one by one takes about 1.7 s with a
1048576 open-file limit (Docker's default), which makes the window wide.

Block TERM, INT and ABRT across the fork along with SIGCHLD, and in the
child reset them to their defaults and unblock them first thing, so a stop
sent before exec ends the child and the reaper restarts or forgets it as
asked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The child closed every descriptor number from 3 up to _SC_OPEN_MAX one by
one. With a 1048576 open-file limit (Docker's default) that is about 1.7 s
of close() calls before each daemon starts, and it is the window in which
zmwatch can find a just-started zmc without shared memory.

Close the descriptors /proc/self/fd lists instead, and keep the old loop
where /proc isn't available.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant