Skip to content

chat: close wakeFoldedRun concurrent-wake race via lifecycle.ensureAwake - #492

Merged
TheGreatAxios merged 2 commits into
mainfrom
cl-7214-serialize-wakefoldedrun
Aug 30, 2026
Merged

TheGreatAxios merged 2 commits into
mainfrom
cl-7214-serialize-wakefoldedrun

Conversation

@TheGreatAxios

Copy link
Copy Markdown
Contributor

Summary

wakeFoldedRun (packages/folded-runs) clears a run's session_asset manifest rows and then redeploys, with no lock spanning the two steps — two concurrent calls for the same instance race on the same primary key and the same git ref.

@corbits/agent-lifecycle's ensureAwake already coalesces concurrent wakes onto one in-flight promise per address (and, as of the stacked CL-7217 PR below this one, keeps that dedup alive until the real wake settles, not just until a caller's timeout fires). But packages/chat/src/platform-adapter.ts's wakeByAddressBounded bypassed that coalescing for one caller: sendFoldedMailWithReclaimRetry's reclaim-retry loop called wakeByAddress directly, with its own ad hoc timeout, never touching lifecycle's pendingWakes map — even though lifecycle is always configured in production (apps/hub/src/index.ts).

Fix

wakeByAddressBounded now routes through lifecycle.ensureAwake(address) whenever a lifecycle is configured, converging the reclaim retry onto the exact same coalescing map sendMail's lifecycle branch and the exported ensureAwake hook already use — no new lock or map added anywhere in packages/chat. Only the no-lifecycle fallback still calls wakeByAddress + withTimeout directly, unchanged.

Confirmed before implementing that wakeFoldedRun has exactly one real caller in the repo (wakeByAddress, in this file), and that every wake entry point here resolves the same live address before touching wake logic — so address-keyed coalescing in agent-lifecycle closes the same race a per-instanceId lock in folded-runs would, without a second primitive.

Known, accepted side effect: the reclaim retry now calls reconcileDriftedRun after every wake attempt (up to ~5 retries) — an extra DB round-trip per retry that didn't happen before, mirroring the CL-6588 pattern already used at the other two call sites (lifecycle.ensureAwake no-ops when already routable, so a staleness check needs to run unconditionally alongside it). Harmless, but flagging explicitly so it doesn't read as a regression in a wake-storm trace.

Test

New test in packages/chat/test/platform-adapter.test.ts ("wakeByAddressBounded reclaim-retry coalescing"): drives a sendMail call whose first delivery attempt fails "agent is unreachable" (forcing the reclaim retry), gates the resulting redeploy open, and fires a concurrent ensureAwake call for the same address while it's in flight. Asserts only one session_asset delete+redeploy cycle runs, not two racing ones. Manually confirmed this test fails (a third, racing deploy) against the unfixed adapter.

Stack

Stacked on cl-7217-agent-lifecycle-wake-timeout (#491) — this PR's fix depends on that one's coalescing primitive being correct. Merge #491 first.

Review

  • Approach reviewed with Greybeard before implementation.
  • Diff reviewed with Critique after implementation: no blocking findings.

Closes CL-7214.

Base automatically changed from cl-7217-agent-lifecycle-wake-timeout to main August 30, 2026 21:29
sendFoldedMailWithReclaimRetry's reclaim path calls wakeByAddressBounded
directly, bypassing lifecycle.ensureAwake's per-address dedup even when
a lifecycle is configured. Drives a sendMail call whose first delivery
attempt fails "agent is unreachable" (forcing the reclaim retry), gates
its redeploy open, and fires a concurrent ensureAwake call for the same
address -- asserting only one wakeFoldedRun delete+redeploy cycle runs,
not two racing ones. Confirmed the test fails (deployCallCount reaches
3) against the unfixed platform-adapter.ts.
wakeFoldedRun clears a run's session_asset rows and redeploys with no
lock spanning the two steps, so two concurrent calls for the same
instance race on the same primary key and git ref.
sendFoldedMailWithReclaimRetry's reclaim path called wakeByAddress
directly, bypassing lifecycle.ensureAwake's per-address pendingWakes
coalescing even when a lifecycle was configured -- the one caller that
could still race a lifecycle-driven wake for the same address.

wakeByAddressBounded now routes through lifecycle.ensureAwake whenever
a lifecycle is configured, converging every wake path in this adapter
(sendMail, the exported ensureAwake hook, and the reclaim retry) onto
the one coalescing map @corbits/agent-lifecycle already owns, rather
than adding a second one here. reconcileDriftedRun runs unconditionally
afterward, matching the CL-6588 pattern already used at the other two
call sites, since lifecycle.ensureAwake no-ops on an address that is
already routable. Only the no-lifecycle fallback still calls
wakeByAddress directly, unchanged.
@TheGreatAxios
TheGreatAxios force-pushed the cl-7214-serialize-wakefoldedrun branch from 9e613c8 to 840950d Compare August 30, 2026 21:29
@TheGreatAxios
TheGreatAxios merged commit ea06ea3 into main Aug 30, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant