Skip to content

Turn claims: token-scoped release, bound the CL-6670 wait - #450

Merged
TheGreatAxios merged 4 commits into
mainfrom
cl-7129-turn-claim-ttl-backstop-lets-a-second-dispatch-loop-run
Aug 29, 2026
Merged

TheGreatAxios merged 4 commits into
mainfrom
cl-7129-turn-claim-ttl-backstop-lets-a-second-dispatch-loop-run

Conversation

@TheGreatAxios

@TheGreatAxios TheGreatAxios commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Fixes CL-7129 — https://linear.app/abklabs/issue/CL-7129

Problem

createInMemoryTurnClaimStore.tryClaim (packages/chat/src/turn-claims.ts, pre-fix) let tryClaim return true again for a workbench whose claim was still held, once now() - claimedAt >= ttlMs. createWorkbenchTurnQueue.run (packages/chat/src/turn-queue.ts, pre-fix) holds a claim across its whole batch-drain loop, so a second run() call winning a fresh claim while the first was still draining started a second, concurrent drain loop over the same pendingByWorkbench queue — both loops able to pop and dispatch.

Separately, dispatchTurnBatch's agentTurns.waitUntilFree call (packages/chat/src/workbench-service.ts, pre-fix, CL-6670) ran with no bound of its own while holding this same claim. A slow prior turn for the same agent could itself hold the claim past its TTL and open exactly this window.

An earlier version of this PR bounded that wait with DEFAULT_WAIT_UNTIL_FREE_TIMEOUT_MS = CHAT_TURN_TIMEOUT_MS - DEFAULT_TURN_DISPATCH_TIMEOUT_MS - 30_000 (150s), below CHAT_TURN_TIMEOUT_MS (300s) — the documented max length a legitimate turn is allowed to run. A message queued behind a prior turn taking, say, 4 minutes would then time out waiting at 2.5 minutes and drop as undelivered: the exact bug CL-6670 fixed, reopened for any prior turn longer than ~2.5 minutes.

Change

  • TurnClaimStore.tryClaim now returns an opaque token (or false) instead of a bare boolean. release and a new holds query both require the token, so a stale holder's release (after the TTL reassigns the claim) is a no-op and can never evict a live claim.
  • createWorkbenchTurnQueue.run's drain loop calls holds() after every dispatch() and stops the moment it no longer holds the claim, leaving whatever queued behind it for the new holder to drain — never two loops draining the same workbench at once.
  • dispatchTurnBatch wraps agentTurns.waitUntilFree in its own waitUntilFreeTimeoutMs (new, injectable). Its default, DEFAULT_WAIT_UNTIL_FREE_TIMEOUT_MS, is now CHAT_TURN_TIMEOUT_MS + 30_000 — it adds grace on top of the full legitimate turn length instead of subtracting from it, so a prior turn running the whole CHAT_TURN_TIMEOUT_MS no longer times out the wait.
  • The claim TTL can no longer just be CHAT_TURN_TIMEOUT_MS: it has to clear this larger wait bound plus the dispatch deadline (DEFAULT_TURN_DISPATCH_TIMEOUT_MS) to stay an unreachable backstop. It's now its own constant, DEFAULT_TURN_CLAIM_TTL_MS = DEFAULT_WAIT_UNTIL_FREE_TIMEOUT_MS + DEFAULT_TURN_DISPATCH_TIMEOUT_MS + 30_000 (480s), derived from both rather than a second independent literal. apps/hub/src/index.ts's claim store and createChatRoutes's turnTimeoutMs both use it in place of CHAT_TURN_TIMEOUT_MS.
  • AGENT_SECTION_MODE.turnTimeoutMs (standalone-launch.ts) is unaffected — it stays CHAT_TURN_TIMEOUT_MS, the platform's own per-occurrence turn timeout, unrelated to the claim TTL.
  • TurnClaimStore has one implementer (createInMemoryTurnClaimStore); no other implementer needed updating.

Tests

  • packages/chat/src/turn-claims.test.ts: tryClaim returns a token; release/holds are token-scoped; a release with a stale token (one the TTL already reassigned) is a no-op and never evicts the live holder.
  • packages/chat/src/turn-queue.test.ts: a new test reproduces the reported shape — a message queues behind an in-flight turn, the claim TTL elapses mid-dispatch, a second run() wins a fresh claim and dispatches its own turn, and the first loop's own dispatch settling afterward no longer pops and re-dispatches what queued behind it; the second loop picks it up exactly once.
  • packages/chat/test/agent-turn-dispatch.test.ts: a new test proves a prior turn that closes just under the injected wait budget still lets the queued send through, never posting an undelivered notice.
  • packages/chat/src/workbench-service.timeouts.test.ts: new constants tests — DEFAULT_WAIT_UNTIL_FREE_TIMEOUT_MS clears CHAT_TURN_TIMEOUT_MS, and DEFAULT_TURN_CLAIM_TTL_MS clears DEFAULT_WAIT_UNTIL_FREE_TIMEOUT_MS + DEFAULT_TURN_DISPATCH_TIMEOUT_MS.
  • bun test packages/chat/src/turn-queue.test.ts packages/chat/src/turn-claims.test.ts packages/chat/test/agent-turn-dispatch.test.ts packages/chat/test/turn-dispatch-deadline.test.ts packages/chat/src/workbench-service.timeouts.test.ts — all pass.
  • bunx tsc --noEmit clean for packages/chat and apps/hub; eslint clean on changed files.

Covers the CL-7129 fix ahead of the implementation: a token issued by
tryClaim must be the only thing that can release or confirm a claim,
so a stale holder's release (after the TTL reassigns the claim to a
second winner) is a no-op, and a drain loop can tell it lost its claim
mid-dispatch.
The TTL backstop in createInMemoryTurnClaimStore let tryClaim return
true again for a workbench whose claim was still held by an in-flight
dispatch — createWorkbenchTurnQueue's run() holds that claim across
its whole batch-drain loop, so a second run() winning a fresh claim
while the first was still draining started a second, concurrent drain
loop over the same queue.

tryClaim now returns an opaque token instead of a bare boolean;
release and the new holds query require it, so a stale holder's
release can never evict a second winner's live claim. The drain loop
checks holds() after every dispatch and stops the moment it no longer
holds the claim, leaving whatever queued behind it for the new holder
to drain — never two loops draining the same workbench at once.

Separately, dispatchTurnBatch's wait on agentTurns.waitUntilFree
(CL-6670) ran with no bound of its own while holding this same claim,
so a slow prior turn for an agent could itself hold the claim past its
TTL and open exactly this window. It now runs under its own
waitUntilFreeTimeoutMs, defaulting to a budget that leaves the
existing turnDispatchTimeoutMs (CL-6644) room to still fit inside the
claim TTL — CL-6670's serialization is unchanged, only the promise it
awaits is now bounded.

Fixes CL-7129.
A prior turn is allowed to run the full CHAT_TURN_TIMEOUT_MS; the wait
dispatchTurnBatch puts on agentTurns.waitUntilFree has to tolerate that,
or a message queued behind a merely-slow (not hung) prior turn times
out and drops as undelivered instead of waiting it out. Covers both the
behavior (a prior turn closing just under the injected wait budget
still lets the queued send through) and the constants' own arithmetic
(the wait bound clears CHAT_TURN_TIMEOUT_MS; the claim TTL clears the
wait bound plus the dispatch deadline).
DEFAULT_WAIT_UNTIL_FREE_TIMEOUT_MS was derived by subtracting from
CHAT_TURN_TIMEOUT_MS, landing below it (150s vs. the 300s a turn is
legitimately allowed to run) — a message queued behind a prior turn
that takes, say, 4 minutes now timed out waiting and dropped as
undelivered, reopening the exact bug CL-6670 fixed. The wait bound now
adds to CHAT_TURN_TIMEOUT_MS instead of subtracting from it.

That in turn means the claim TTL can no longer just be
CHAT_TURN_TIMEOUT_MS — it has to clear the new, larger wait bound plus
the dispatch deadline to stay an unreachable backstop. It's now its own
constant, DEFAULT_TURN_CLAIM_TTL_MS, derived from both.

AGENT_SECTION_MODE.turnTimeoutMs (standalone-launch.ts) is unaffected:
it stays CHAT_TURN_TIMEOUT_MS, the platform's own per-occurrence turn
timeout, unrelated to the claim TTL.

Fixes CL-7129.
@TheGreatAxios
TheGreatAxios force-pushed the cl-7129-turn-claim-ttl-backstop-lets-a-second-dispatch-loop-run branch from faf25d2 to 9affe03 Compare August 29, 2026 04:46
@TheGreatAxios
TheGreatAxios merged commit 8299cf9 into main Aug 29, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant