fix(status): runtime status agrees with the heartbeat; current-status republishes on wake - #726
Conversation
… republishes on wake /api/v3/plugins/installed said the display's plugins were "live" while /api/v3/health said display_loop: stalled. The runtime snapshot is written from its own thread, which keeps ticking while the render loop is hung. The reader now also checks the render loop's heartbeat: a fresh snapshot whose process's heartbeat is HEARTBEAT_STALE_SECONDS old is "stalled", with no per-plugin facts. A running snapshot from a process that no longer exists (after a watchdog kill systemd removes the heartbeat directory) is "stale" at once. Reading side only: no new files or writes. display_current_state was republished only on a mode change or every CURRENT_STATE_REFRESH_SECONDS, so waking from scheduled-off (or on-demand starting or ending) with the same mode left is_display_active / on_demand_active up to 30 s out of date. A change to either flag now republishes at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 35 minutes. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (7)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Up to standards ✅🟢 Issues
|
| Metric | Results |
|---|---|
| Complexity | 21 |
NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.
Fixes two places where the status the web interface reports disagrees with what the display is doing.
1. Plugin runtime said "live" while the render loop was stalled
Root cause. The plugin runtime snapshot from #690 is written by the
plugin-runtime-publisherdaemon thread. That thread keeps ticking when the render thread hangs inside a plugin. So/api/v3/plugins/installedreporteddata.runtime.status: "live"while/api/v3/healthreportedchecks.display_loop: "stalled"from the #687 heartbeat. After a watchdog kill, systemd removes/run/ledmatrix(RuntimeDirectory=), so no heartbeat is left behind. The dead process's last snapshot then stayed "live" until its 180 sstale_afterran out.Fix: on the reading side (
src/plugin_system/plugin_runtime.py). There are no new files and no extra writes. The heartbeat is already written to tmpfs.read_plugin_runtime()also reads the render-loop heartbeat (display_watchdog.read_heartbeat). The status becomesstalledwhen the snapshot is fresh, the heartbeat came from the same pid, and the heartbeat is at leastHEARTBEAT_STALE_SECONDS(60 s) old. That is the same threshold/healthuses. A stalled view reports no per-plugin facts, just likestale.staleat once. The check is POSIXkill(pid, 0), which sends no signal. It returns EPERM when the web user is not root and the display process is alive, and that counts as alive. On Windows the check is skipped, becauseos.killterminates the process there. This closes the window after a watchdog kill.data.runtimegainsheartbeat_age_seconds.docs/REST_API_REFERENCE.mdand the route docstring are updated.I chose the reader option over having the publisher include render-loop liveness. The publisher option would mean another write, or the publisher thread tracking the render thread. The reader already has both signals, and doing the check there keeps the
/healthand runtime verdicts on one threshold. Reconciliation seesstalledas not-live and skips observed checks, which is the same handling asstale.2. current-status kept is_display_active stale for up to 30 s after a wake
Root cause.
_publish_current_mode_state_if_changed()republisheddisplay_current_stateonly when the mode name changed orCURRENT_STATE_REFRESH_SECONDS(30 s) had passed. Waking from scheduled-off, blanking, or an on-demand session starting or ending usually leaves the mode name unchanged. The blank path publishes once and then sleeps through_service_pending_changes(). On wake, the next publish point saw the same mode and a recent publish, so it skipped. This is what ledpi showed.Fix (
src/display_controller.py, the publish helpers only, away from the plugin-reload code that #723 touches): record(is_display_active, on_demand_active)at each publish, and republish as soon as either one changes. When nothing changes there are no extra writes.Tests
test/test_plugin_runtime_snapshot.py:stalledand reports no plugin facts;live;stale;stale;process_exists;stalled, a fresh heartbeat giveslive.test/test_display_controller_shutdown_and_state.py:_service_pending_changes();src/reverted toorigin/main, 15 of the new tests fail (the route-levelstalledtest among them).test/test_run_loop_golden.py): 16 passed, byte-identical.origin/mainbaseline worktree (Windows, Pillow 12.3.0 before and after): the FAILED/ERROR test IDs are identical (61 failed, 6 errors, all pre-existing). The only diff is three captured log lines whose line number moved from 272 to 273 after a one-line docstring edit. There are 17 more passes.scripts/check_types.pywith mypy 1.20.2 (pure-Python build): clean, 93 modules.🤖 Generated with Claude Code