Skip to content

fix: prevent away-mode escalation injection wedges - #49

Merged
knowttl merged 7 commits into
mainfrom
fm/fm-afk-inject-wedge-claude
Sep 23, 2026
Merged

knowttl merged 7 commits into
mainfrom
fm/fm-afk-inject-wedge-claude

Conversation

@knowttl

@knowttl knowttl commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Intent

just scout on the root cause.

Context the ask refers to: overnight 2026-09-22 to 2026-09-23 the primary firstmate ran as Claude Code (claude-opus-5-5) inside Herdr (backend herdr, pane w3:p7R, workspace w3), and entered away mode with /afk, which on Claude launches the away supervise daemon through the native background bash (bin/fm-afk-launch.sh start-native then FM_AFK_STATE_PREPARED=1 bin/fm-afk-start.sh).
The daemon started correctly (state/.supervise-daemon.log: daemon starting ... target=default:w3:p7R; target_source=HERDR_ENV(HERDR_PANE_ID); backend=herdr), and the target was the right pane.
For about 10 hours not one escalation reached the primary: the log holds 2016 lines inject failed: submit unconfirmed after 3 retries (verdict=send-failed, text may be in composer), 95 wedge alarm: no OS-level alert channel on Linux, and ERROR: away-mode escalation undelivered ...; the return brief reported fm away-mode inject WEDGED: 37178s undelivered.
Captain decisions, finished lanes waiting to start validation, and review findings all sat unhandled until the captain returned.
Normal attended supervision on the same pane (Claude Stop-hook-owned rewake, no text injection) worked all day.
A similar overnight wedge is on file from 2026-08-30 (backlog item fm-afk-inject-wedge-herdr).
The idle Claude Code composer in this pane renders as an unbordered ❯ prompt between two horizontal rules, above a status line, which may matter to the composer classifier.
The captain chose only this scout for now (not a phone-alert setup); the goal, as with today's other firstmate scouts, is a verified root cause that can become a fix pull request on the fork and then upstream.

yes also help me build the outage fix

The scout's finding, in substance: the away daemon never reached the pane. escalate_flush in bin/fm-supervise-daemon.sh joins every buffered escalation into one unbounded digest (about 448 KB overnight, inflated by the catch-all scan replaying a day of already-handled status lines), and inject_msg passes it to the backend send as a single argument; Linux refuses any single exec argument of 128 KiB or more, so herdr pane send-text never ran (tmux fails at about 16 KB, upstream issue kunchenguid#4382), the buffer was never cleared, and every retry sent the same ever-growing digest for ten hours while the log misleadingly blamed a busy pane and text left in the composer.
The fix: cap each buffered item with an explicit truncation marker pointing at the task's status log, build each digest from the oldest lines that fit one byte budget below every transport ceiling (about 1,000 bytes, which also clears the Claude-on-Herdr head-truncation limit in upstream issue kunchenguid#3473), remove only the delivered lines on a confirmed submit so the rest go in later batches, and log a backend send refusal honestly with the digest size.

What Changed

  • Cap each buffered escalation item with a truncation marker that points to its source status log.
  • Send the oldest queued events in digests of at most 1,000 bytes, removing only events whose submission is confirmed. Log backend send failures with the digest size.
  • Keep source-log metadata out of return briefs and add daemon, tmux, and Herdr regression coverage for oversized backlogs.

Risk Assessment

✅ Low: The bounded digest, per-status source records, confirmed-delivery removal, and return catch-up path satisfy the stated fix without a substantiated remaining failure in intended use.

Testing

The daemon tests and return tests passed. Private tmux delivered oversized events in bounded digests; real Herdr delivered a 150 KB backlog in two digests of 972 and 878 bytes, preserved the correct status pointer, deferred on pending input, and raised the persistent wedge alarm. The throwaway Herdr lab was torn down. A genuine backend refusal was not induced live.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
An away captain receives a backlog containing a 150 KB event as bounded Herdr digests, with the oldest and later events delivered ✅ pass live tests/fm-afk-inject-herdr-e2e.test.sh Scenario E; real Herdr submission transcript
A status excerpt that names an existing b.status stays one a.status event with an a.status recovery pointer ✅ pass live tests/fm-afk-inject-herdr-e2e.test.sh Scenario E; second submitted digest points to a.status
A pending Herdr composer defers escalation; persistent pending input raises a wedge alarm and preserves the buffer ✅ pass live tests/fm-afk-inject-herdr-e2e.test.sh Scenarios A and D
A genuine Herdr send refusal logs the digest size and retains undelivered events ⏸️ untested no The lab had no reproducible genuine Herdr send refusal; provide a Herdr transport failpoint or reproducible refusal to drive this live.
Evidence: Real Herdr submission transcript
Real Herdr submitted two digests: 972 bytes for the 150 KB backlog and 878 bytes for a.status. The submitted messages contained “full text in .../big.status” and “full text in .../a.status”; the latter kept “b.status:” inside the a.status excerpt.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 4 issues found → auto-fixed (4) ✅
  • 🚨 bin/fm-supervise-daemon.sh:763 - A signal can combine several task statuses into one buffered item. If that item exceeds the budget, truncation points only to the first status log, then a confirmed submit removes the entire item. Later tasks' decisions can disappear from the digest with no pointer to recover them. Preserve a recoverable reference for every task before removing the item. The same invariant must hold at bin/fm-supervise-daemon.sh:802 (truncation) and bin/fm-supervise-daemon.sh:811 (removal).
  • 🚨 bin/fm-supervise-daemon.sh:802 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” The changed code still appends the full item at bin/fm-supervise-daemon.sh:706 and truncates only during a flush at line 802. Thus the buffer itself remains unbounded, contrary to that criterion. Apply the cap when buffering each item.
  • ⚠️ bin/fm-supervise-daemon.sh:719 - The new FM_INJECT_MAX_BYTES option is not required by the stated fixed budget of about 1,000 bytes. It also accepts values above transport ceilings, so setting it to 200000 recreates the oversized-send wedge. Remove the option and use a fixed safe budget; the option is also exposed at docs/configuration.md:1245 and .agents/skills/afk/SKILL.md:182.
  • ⚠️ bin/fm-supervise-daemon.sh:1404 - Herdr also returns send-failed when sending the text succeeded but every Enter attempt failed (bin/backends/herdr.sh:3268,3306). This new log then incorrectly says the backend refused the send, although the digest may remain in the composer. Distinguish text-send refusal from submit-key failure, or use wording accurate for both outcomes.

🔧 Fix applied.
7 issues (4 errors, 3 warnings) still open:

  • 🚨 bin/fm-supervise-daemon.sh:763 - A signal can combine several task statuses into one buffered item. If that item exceeds the budget, truncation points only to the first status log, then a confirmed submit removes the entire item. Later tasks' decisions can disappear from the digest with no pointer to recover them. Preserve a recoverable reference for every task before removing the item. The same invariant must hold at bin/fm-supervise-daemon.sh:802 (truncation) and bin/fm-supervise-daemon.sh:811 (removal).
  • 🚨 bin/fm-supervise-daemon.sh:802 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” The changed code still appends the full item at bin/fm-supervise-daemon.sh:706 and truncates only during a flush at line 802. Thus the buffer itself remains unbounded, contrary to that criterion. Apply the cap when buffering each item.
  • ⚠️ bin/fm-supervise-daemon.sh:719 - The new FM_INJECT_MAX_BYTES option is not required by the stated fixed budget of about 1,000 bytes. It also accepts values above transport ceilings, so setting it to 200000 recreates the oversized-send wedge. Remove the option and use a fixed safe budget; the option is also exposed at docs/configuration.md:1245 and .agents/skills/afk/SKILL.md:182.
  • ⚠️ bin/fm-supervise-daemon.sh:1404 - Herdr also returns send-failed when sending the text succeeded but every Enter attempt failed (bin/backends/herdr.sh:3268,3306). This new log then incorrectly says the backend refused the send, although the digest may remain in the composer. Distinguish text-send refusal from submit-key failure, or use wording accurate for both outcomes.
  • 🚨 bin/fm-supervise-daemon.sh:703 - Fix round 1 introduced a splitter that stops at the first | unless the following text begins with a status filename. For a real signal such as a.status: needs-decision: A | B | b.status: done: ..., line 703 leaves both statuses in one item; the cap at line 711 then truncates it with only a.status as a recovery pointer. The flush-time truncation at line 803 has the same single-pointer assumption. Split at each actual status boundary so every task remains recoverable.
  • 🚨 bin/fm-supervise-daemon.sh:711 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” A stale wake with an actionable status produces stale + actionable status: <event> at line 426, so the new cap at line 711 truncates a long event without a log pointer: line 765 recognizes only items beginning with *.status: . The same omission applies if line 803 truncates that item again. Include the task's status log in this status-producing path.
  • ⚠️ bin/fm-supervise-daemon.sh:778 - The new more queued count is an additional digest output component at lines 776-779. The stated requirement is to deliver the remaining lines in later batches; it does not require announcing their count. Remove this suffix to keep the change within the requested behavior.

🔧 Fix applied.
9 issues (5 errors, 4 warnings) still open:

  • 🚨 bin/fm-supervise-daemon.sh:763 - A signal can combine several task statuses into one buffered item. If that item exceeds the budget, truncation points only to the first status log, then a confirmed submit removes the entire item. Later tasks' decisions can disappear from the digest with no pointer to recover them. Preserve a recoverable reference for every task before removing the item. The same invariant must hold at bin/fm-supervise-daemon.sh:802 (truncation) and bin/fm-supervise-daemon.sh:811 (removal).
  • 🚨 bin/fm-supervise-daemon.sh:802 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” The changed code still appends the full item at bin/fm-supervise-daemon.sh:706 and truncates only during a flush at line 802. Thus the buffer itself remains unbounded, contrary to that criterion. Apply the cap when buffering each item.
  • ⚠️ bin/fm-supervise-daemon.sh:719 - The new FM_INJECT_MAX_BYTES option is not required by the stated fixed budget of about 1,000 bytes. It also accepts values above transport ceilings, so setting it to 200000 recreates the oversized-send wedge. Remove the option and use a fixed safe budget; the option is also exposed at docs/configuration.md:1245 and .agents/skills/afk/SKILL.md:182.
  • ⚠️ bin/fm-supervise-daemon.sh:1404 - Herdr also returns send-failed when sending the text succeeded but every Enter attempt failed (bin/backends/herdr.sh:3268,3306). This new log then incorrectly says the backend refused the send, although the digest may remain in the composer. Distinguish text-send refusal from submit-key failure, or use wording accurate for both outcomes.
  • 🚨 bin/fm-supervise-daemon.sh:703 - Fix round 1 introduced a splitter that stops at the first | unless the following text begins with a status filename. For a real signal such as a.status: needs-decision: A | B | b.status: done: ..., line 703 leaves both statuses in one item; the cap at line 711 then truncates it with only a.status as a recovery pointer. The flush-time truncation at line 803 has the same single-pointer assumption. Split at each actual status boundary so every task remains recoverable.
  • 🚨 bin/fm-supervise-daemon.sh:711 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” A stale wake with an actionable status produces stale + actionable status: <event> at line 426, so the new cap at line 711 truncates a long event without a log pointer: line 765 recognizes only items beginning with *.status: . The same omission applies if line 803 truncates that item again. Include the task's status log in this status-producing path.
  • ⚠️ bin/fm-supervise-daemon.sh:778 - The new more queued count is an additional digest output component at lines 776-779. The stated requirement is to deliver the remaining lines in later batches; it does not require announcing their count. Remove this suffix to keep the change within the requested behavior.
  • 🚨 bin/fm-supervise-daemon.sh:707 - Round 2's splitter treats status text as a task boundary. A real status such as real.status: needs-decision: choose | missing.status: <long explanation> is split at line 707; the long second part is truncated with a pointer to nonexistent missing.status, then removed after delivery. Only split boundaries established by the signal's actual status files. This also affects items produced by classify_stale (line 426) and the catch-all scan (line 1306).
  • ⚠️ bin/fm-supervise-daemon.sh:815 - After a confirmed submit, a failed tail write or mv leaves delivered lines in the buffer, but the function still returns success. The next flush can submit those events again. Check both the copy at line 815 and replacement at line 816, and preserve an explicit failure result if either fails.

🔧 Fix applied.
10 issues (6 errors, 4 warnings) still open:

  • 🚨 bin/fm-supervise-daemon.sh:763 - A signal can combine several task statuses into one buffered item. If that item exceeds the budget, truncation points only to the first status log, then a confirmed submit removes the entire item. Later tasks' decisions can disappear from the digest with no pointer to recover them. Preserve a recoverable reference for every task before removing the item. The same invariant must hold at bin/fm-supervise-daemon.sh:802 (truncation) and bin/fm-supervise-daemon.sh:811 (removal).
  • 🚨 bin/fm-supervise-daemon.sh:802 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” The changed code still appends the full item at bin/fm-supervise-daemon.sh:706 and truncates only during a flush at line 802. Thus the buffer itself remains unbounded, contrary to that criterion. Apply the cap when buffering each item.
  • ⚠️ bin/fm-supervise-daemon.sh:719 - The new FM_INJECT_MAX_BYTES option is not required by the stated fixed budget of about 1,000 bytes. It also accepts values above transport ceilings, so setting it to 200000 recreates the oversized-send wedge. Remove the option and use a fixed safe budget; the option is also exposed at docs/configuration.md:1245 and .agents/skills/afk/SKILL.md:182.
  • ⚠️ bin/fm-supervise-daemon.sh:1404 - Herdr also returns send-failed when sending the text succeeded but every Enter attempt failed (bin/backends/herdr.sh:3268,3306). This new log then incorrectly says the backend refused the send, although the digest may remain in the composer. Distinguish text-send refusal from submit-key failure, or use wording accurate for both outcomes.
  • 🚨 bin/fm-supervise-daemon.sh:703 - Fix round 1 introduced a splitter that stops at the first | unless the following text begins with a status filename. For a real signal such as a.status: needs-decision: A | B | b.status: done: ..., line 703 leaves both statuses in one item; the cap at line 711 then truncates it with only a.status as a recovery pointer. The flush-time truncation at line 803 has the same single-pointer assumption. Split at each actual status boundary so every task remains recoverable.
  • 🚨 bin/fm-supervise-daemon.sh:711 - The required fix says to “cap each buffered item with an explicit truncation marker pointing at the task's status log.” A stale wake with an actionable status produces stale + actionable status: <event> at line 426, so the new cap at line 711 truncates a long event without a log pointer: line 765 recognizes only items beginning with *.status: . The same omission applies if line 803 truncates that item again. Include the task's status log in this status-producing path.
  • ⚠️ bin/fm-supervise-daemon.sh:778 - The new more queued count is an additional digest output component at lines 776-779. The stated requirement is to deliver the remaining lines in later batches; it does not require announcing their count. Remove this suffix to keep the change within the requested behavior.
  • 🚨 bin/fm-supervise-daemon.sh:707 - Round 2's splitter treats status text as a task boundary. A real status such as real.status: needs-decision: choose | missing.status: <long explanation> is split at line 707; the long second part is truncated with a pointer to nonexistent missing.status, then removed after delivery. Only split boundaries established by the signal's actual status files. This also affects items produced by classify_stale (line 426) and the catch-all scan (line 1306).
  • ⚠️ bin/fm-supervise-daemon.sh:815 - After a confirmed submit, a failed tail write or mv leaves delivered lines in the buffer, but the function still returns success. The next flush can submit those events again. Check both the copy at line 815 and replacement at line 816, and preserve an explicit failure result if either fails.
  • 🚨 bin/fm-supervise-daemon.sh:707 - Round 3 left a false boundary reachable when the named status file exists. A signal for only a.status can contain needs-decision: inspect excerpt | b.status: <long text> while b.status also exists. The splitter treats the excerpt as a second event; the cap then truncates it with a pointer to b.status, although the full text is in a.status. The affected sites are bin/fm-supervise-daemon.sh:707 (boundary), :718 (cap), :772 (pointer), and :812 (delivered item removal). The new test covers only a missing filename. Correcting this requires boundary information from classify_signal rather than the existing-file heuristic chosen in round 3, so the remedy needs authorization.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
An away captain receives a backlog containing a 150 KB event as bounded Herdr digests, with the oldest and later events delivered ✅ pass live tests/fm-afk-inject-herdr-e2e.test.sh Scenario E; real Herdr submission transcript
A status excerpt that names an existing b.status stays one a.status event with an a.status recovery pointer ✅ pass live tests/fm-afk-inject-herdr-e2e.test.sh Scenario E; second submitted digest points to a.status
A pending Herdr composer defers escalation; persistent pending input raises a wedge alarm and preserves the buffer ✅ pass live tests/fm-afk-inject-herdr-e2e.test.sh Scenarios A and D
A genuine Herdr send refusal logs the digest size and retains undelivered events ⏸️ untested no The lab had no reproducible genuine Herdr send refusal; provide a Herdr transport failpoint or reproducible refusal to drive this live.
  • tests/fm-afk-inject-herdr-e2e.test.sh
  • tests/fm-afk-inject-e2e.test.sh
  • tests/fm-daemon.test.sh
  • tests/fm-afk-return.test.sh
  • git status --short and named Herdr session cleanup check
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…deliver

escalate_flush joined the whole escalation buffer into one digest and
typed it as a single backend argument. A catch-all replay of a long
status span made that digest hundreds of kilobytes, which exceeds the
kernel's 128 KiB single-argument limit (herdr pane send-text never
execs) and tmux's ~16 KB command limit. The send failed before reaching
the pane, the buffer was kept, and every retry resent the same growing
digest until the captain returned.

Each flush now sends only the oldest items that fit FM_INJECT_MAX_BYTES
(default 1000, below tmux, the kernel, and the Claude-on-Herdr
head-truncation window), truncates an item too long to fit alone with a
marker naming its status log, and removes only the delivered lines so
the rest go in later batches. A backend send refusal is now logged as
such with the digest size instead of blaming the composer.
…n.sh by using the metadata check already used by the daemon. The return test suite, repository lint, and diff check pass locally. Stock macOS Bash 3.2 is unavailable here, so CI must confirm that parser check
@knowttl
knowttl merged commit fad8567 into main Sep 23, 2026
19 checks passed
knowttl added a commit that referenced this pull request Sep 26, 2026
This reverts commit fad8567.

Superseded by upstream 683b3eb (bound the away digest and log why a
delivery failed) and 1d3ac67 (refuse a Herdr Claude submit that would
send only a message tail), which fix the same oversized-digest wedge.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant