From 693d3647348a625a8d381fe4c844b668709cf414 Mon Sep 17 00:00:00 2001 From: Phil Merrell Date: Sun, 27 Sep 2026 11:42:25 -0600 Subject: [PATCH] docs(specs): background handoffs with in-place review and notifications Revise the agent handoffs spec after a UX pass: - Rename "task chips" to handoffs. Shared Projects already uses "task" for a conversation, so HANDOFF# rows, /handoffs routes and HANDOFFS_ENABLED replace the TASK# names. - A mid-thread mention opens a review sheet in place, with a full-prompt preview. On Send the target conversation runs in the background as a second stream; the user stays in the original thread. - A handoff row tracks running / needs your input / done / stopped and offers Bring back result. - Notifications: the Projects in-app inbox gains handoff kinds, plus local browser notifications while the page is hidden. Web Push is a follow-up shared with scheduled runs. - The Phase 1 gate now asks whether users bring results back without opening the target, since background handoffs cover most of a consult. Co-Authored-By: Claude Opus 5.5 --- docs/specs/agent-handoffs-and-consults.md | 465 ++++++++++++++++------ 1 file changed, 349 insertions(+), 116 deletions(-) diff --git a/docs/specs/agent-handoffs-and-consults.md b/docs/specs/agent-handoffs-and-consults.md index 7f3481ac5..ae00b0a24 100644 --- a/docs/specs/agent-handoffs-and-consults.md +++ b/docs/specs/agent-handoffs-and-consults.md @@ -1,9 +1,14 @@ # Agent handoffs and consults — carrying an `@`-mention's work across conversations -**Status:** Draft proposal. Not scheduled. **Phase 0 (task chips) ships first; its numbers decide whether Phase 1 +**Status:** Draft proposal. Not scheduled. **Phase 0 (handoffs) ships first; its numbers decide whether Phase 1 (consults) is built at all.** **Author:** Phil Merrell (drafted with Claude) -**Date:** 2026-09-24 +**Date:** 2026-09-24, revised 2026-09-27 +**Revision (2026-09-27):** A handoff now runs **in the background** while the user stays in the original thread, +after a **review step in place** (§4.2). It reports back through a handoff row with *running / needs your input / +done* states, the in-app inbox and browser notifications (§4.3). The feature is renamed from "task chips" to +**handoffs** (§4.0), because Shared Projects already uses "task" for a conversation. §4.6 and the §7 gate are +updated to match. **Targets branch:** `develop` **Related:** - `docs/specs/agent-marketplace.md` §D11 — the 2026-09-14 revision that made a mention *bind* (#1115) @@ -39,7 +44,7 @@ Two things are missing. They are separate, and this spec delivers them in order: We shipped a mid-thread mention once as a per-turn borrow (Phase 7, PR #738) and removed it (#1115). It failed in three ways, and both phases here are built to avoid all three: -| Failure | Mechanism | Phase 0 (chips) | Phase 1 (consult) | +| Failure | Mechanism | Phase 0 (handoffs) | Phase 1 (consult) | |---|---|---|---| | **History fork** | The mention changed the agent cache key, so a second `Agent` served the same session. Each instance held its own stale message list (#741, #751) | The work happens in an **ordinary new conversation** | The Agent runs in **its own hidden session** | | **Prefix re-write** | The Agent's bindings changed `toolConfig`, the first cachePoint, so system prompt and history were re-written behind it (~$0.12 per mention, measured) | The parent is never touched. A send-back is an append | The parent's `toolConfig`, system prompt and model never change. The report is an append | @@ -58,20 +63,25 @@ send-back is a row we can count. **§7 turns those counts into the go/no-go for | Phase | What the user gets | New plumbing | Status | |---|---|---|---| -| **0a — Handoff with context** | A mid-thread mention opens the Agent's conversation with a **context card** from the old thread, reviewed before anything is sent | `TASK#` rows, excerpt renderer, `handoff_context` on the request | Proposed | -| **0b — Send back** | "Send to *original conversation*" on a reply in a spawned conversation; lands as a context card on the parent's composer | Reuses 0a in the other direction | Proposed | -| **0c — Suggested tasks** | The model can offer a chip ("Build this rubric with Rubric Builder →") without stopping its turn | `suggest_task` tool, `task_suggested` SSE event | Proposed, opt-in | +| **0a — Handoff with context** | A mid-thread mention opens a **review sheet** in place. On Send, the Agent's conversation starts **in the background** with a **context card** from this thread; the user stays put | `HANDOFF#` rows, excerpt renderer, `handoff_context` on the request, a no-navigate send | Proposed | +| **0b — Handoff row + send back** | A row in the original thread tracks the handoff (*running / needs your input / done*) and offers **Bring back result** when it finishes; the new conversation offers "Send to *original conversation*" too | Status on `HANDOFF#`, reuses 0a in the other direction | Proposed | +| **0c — Suggested handoffs** | The model can offer a chip ("Hand off to Rubric Builder →") without stopping its turn; Start opens the same review sheet | `suggest_handoff` tool, `handoff_suggested` SSE event | Proposed, opt-in | | **Gate** | — | — | §7 | | **1 — Consult** | A mid-thread mention runs the Agent **in place** and its report comes back automatically | Hidden consult sessions, nested streaming, agent-resolution refactor | Only if the gate passes | -Out of both phases: **running a spawned task in the background** while the user stays in the original thread -(§4.6). +Background runs are **client-driven** in Phase 0: they live as long as the tab that started them (§4.6). +Handoffs that outlive the tab need the headless run entrypoint and are a follow-up. ## 3. What already exists - **A way to hand text to the composer.** `ComposerDraftService` puts text into the composer for a given session for the user to edit and send. Nothing is sent. Its first user is feedback retry-with-correction. Phase 0 is its second. +- **Concurrent streams, and a stream that keeps going when the user looks away.** The SPA's streaming layer is + keyed per session (`ChatStateService`, `StreamParserService`), and navigating away from a streaming conversation + does not abort it. When a stream finishes in a conversation the user isn't viewing, `setChatLoading` marks it + unread and the session list shows the blue dot. A background handoff reuses all of this. What's missing is a + send that doesn't navigate: `submitChatRequest` always calls `setViewedSession` and `navigateToSession`. - **Display text vs. model text.** `original_message` → `displayText` already separates what the user typed from the augmented prompt the model sees. That is how RAG chunks stay out of the rendered bubble. Context cards use the same split. @@ -101,41 +111,93 @@ Out of both phases: **running a spawned task in the background** while the user --- -## 4. Phase 0 — Task chips +## 4. Phase 0 — Handoffs + +### 4.0 Naming + +An earlier draft called Phase 0 "task chips", after Claude Code's `spawn_task` chip. That name is dropped from +copy, code and row keys, for two reasons: + +- It wasn't intuitive. In 0a and 0b the user never sees a chip, and in 0c the chip is only the entry point. +- **Shared Projects already uses "task" for a conversation in a project.** Users see a **Tasks** tab, "New task", + "Your tasks", "Shared with the project" and "Continue in my own task". The backend has `SHARED_TASK#` rows and + `/projects/{id}/tasks` routes. A handoff started from inside a project task would put two different "tasks" on + one screen. + +| Term | Means | In code | +|---|---|---| +| **Handoff** | Sending a request, plus reviewed context, to an Agent in a new conversation | `HANDOFF#`, `/sessions/{id}/handoffs`, `HANDOFFS_ENABLED`, `FEATURES.handoffs` | +| **Handoff row** | The line in the original thread that tracks a handoff (§4.3) | Replayed as `handoffs` on `GET /messages` | +| **Bring back result** / **Send to *{title}*** | Returning a handoff's result to the original thread, from either end | `kind: "send_back"` | +| **Suggested handoff** | A chip the model offers (§4.4) | `suggest_handoff`, `handoff_suggested` | +| **Consult** | Phase 1: the Agent runs in place | `CONSULT#` | +| **Task** | Reserved for Shared Projects. Not used by this feature | — | + +User-facing copy: "Hand off to *Rubric Builder*", "Handed off to *Rubric Builder*", "Bring back result", +"Send to *{title}*". Avoid "delegate", which already means delegated admin scopes in this codebase. ### 4.1 Decision summary | Question | Decision | |---|---| -| Unit of work | A **task**: a `TASK#` row on the *source* session. `kind` is `mention_handoff`, `suggested` or `send_back` | -| How context moves | A **context card**: a bounded, deterministic **excerpt** attached to the first message of the target conversation. **Visible and editable before sending** | +| Unit of work | A **handoff**: a `HANDOFF#` row on the *source* session. `kind` is `mention`, `suggested` or `send_back` | +| How context moves | A **context card**: a bounded, deterministic **excerpt** attached to the first message of the target conversation | +| Review | **Required, in place.** A review sheet opens in the source thread before anything is sent. The full prompt the Agent will receive is one click away | +| Where the user ends up | **Where they were.** The target conversation starts in the background and appears in the conversation list. The user opens it when they choose | +| How progress is shown | A **handoff row** in the source thread (*running*, *needs your input*, *done*, *stopped*), the existing unread dot, a toast, an entry in the **in-app inbox** (the bell), and a **browser notification** while the page is hidden. Never a model message | | Where the excerpt is rendered | **Server-side, once**, by a pure renderer in `apis.shared.handoff`. The SPA never builds it, because Phase 1 needs the same renderer in Python | -| Is anything sent automatically? | **No.** A handoff navigates and waits for the user's Enter. That costs one keystroke more than today, which is what makes the context reviewable | | How the model sees a card | Wrapped block **in the user message**; `displayText` = what the user typed; the card renders as a collapsible panel on the bubble | -| Model-suggested tasks | `suggest_task` tool, **non-blocking** (returns immediately; the turn continues). Constant spec. No Agent list in the schema | -| Linking | The target session records `spawnedFrom: {sessionId, taskId}`. The source's `TASK#` row records `targetSessionId` | -| Background execution | **Not in Phase 0** | -| Flags | Backend `TASK_CHIPS_ENABLED`, SPA `FEATURES.taskChips` (both default **off** in development); `suggest_task` separately gated by `TASK_SUGGESTIONS_ENABLED` | +| Background execution | **Client-driven.** The SPA runs the target's turn as a second stream. It lives as long as the tab (§4.6) | +| Model-suggested handoffs | `suggest_handoff` tool, **non-blocking** (returns immediately; the turn continues). Constant spec. No Agent list in the schema | +| Linking | The target session records `spawnedFrom: {sessionId, handoffId}`. The source's `HANDOFF#` row records `targetSessionId` and `status` | +| Flags | Backend `HANDOFFS_ENABLED`, SPA `FEATURES.handoffs` (both default **off** in development); `suggest_handoff` separately gated by `HANDOFF_SUGGESTIONS_ENABLED` | ### 4.2 Phase 0a — Handoff with context -**Flow.** The user types `@Rubric Builder draft a rubric for this` in a thread with history. +**Flow.** The user types `@Rubric Builder draft a rubric for this` in a thread with history and presses Enter. -1. The SPA calls app-api `POST /sessions/{id}/tasks` with `{kind: "mention_handoff", agentId, request}`, using - cookie auth (`get_current_user_from_session`). The server checks that the user owns the session, then: +1. **Nothing is sent yet.** The SPA calls app-api `POST /sessions/{id}/handoffs` with + `{kind: "mention", agentId, request}`, using cookie auth (`get_current_user_from_session`). The server checks + that the user owns the session, then: - renders the excerpt from persisted messages (§4.5) - - writes `TASK#` with `status: "open"` - - returns `{taskId, excerpt, excerptHash, request}` + - writes `HANDOFF#` with `status: "open"` + - returns `{handoffId, excerpt, excerptHash, excerptTokens, turnsCovered, contextBlock, request}`, where + `contextBlock` is the excerpt already wrapped as the Agent will receive it It does **not** create the target session. Nothing hides empty sessions from the list, so a row created here - would leave a blank "New Conversation" behind every abandoned handoff. -2. The SPA opens a new conversation with `?assistantId={agentId}`, using the staged-session path it already uses - for new chats, and hands `request` to that composer via `ComposerDraftService`. The **context card** sits above - the composer, like an attachment chip. It shows the source conversation's title, can be expanded to read in - full, can be edited, and can be removed. -3. On Enter, the invocation carries `handoff_context: {taskId, text}`, where `text` is the card as sent (possibly - edited), or it is absent if the card was removed. The server: - - checks that the task belongs to the user and is still `open` + would leave a blank "New Conversation" behind every cancelled handoff. +2. **A review sheet opens above the composer, in the source thread.** It is a labelled region, not a modal: the + thread stays readable above it, and Esc cancels. + + ``` + ┌ Hand off to [icon] Rubric Builder ──────────────────────────────────┐ + │ draft a rubric for this (edit) │ + │ ▸ Context from this conversation · turns 7–12 · ~3k tokens │ + │ Preview full prompt · Edit · Remove │ + │ Attached: syllabus-notes.docx │ + │ Cancel [ Send to Rubric Builder ] │ + └──────────────────────────────────────────────────────────────────────┘ + ``` + + - **Request**: what the user typed, minus the mention, editable. + - **Context card**: collapsed by default, showing the source range and size. **Edit** turns it into a text + area; **Remove** drops it. + - **Preview full prompt**: a read-only view of exactly what the Agent will receive: `contextBlock` (or the + edited card, wrapped the same way) followed by the request. The server performs the same join in PR-0.2; a + shared fixture pins the two so the preview can't drift from what is sent. + - **Attachments and invoked skills** from the mention message are listed; they follow the request, as they do + today. +3. **On Send**, the SPA generates the target session id and starts its turn **without navigating** (§3's no-navigate + send). The invocation carries `handoff_context: {handoffId, text}`, where `text` is the card as sent, or is + absent if the card was removed. In the source thread: + - the review sheet closes and the composer clears + - a **handoff row** appears (§4.3) + - a quiet toast confirms it: "Handed off to Rubric Builder · Open". It auto-dismisses and is announced politely + (`aria-live="polite"`), not assertively + + The new conversation appears at the top of the conversation list with the existing streaming dot. +4. **The server**, on the target's first turn: + - checks that the handoff belongs to the user and is still `open` - applies the excerpt cap to `text` again (the SPA can't raise it) - prepends the wrapped block to the model message: @@ -149,55 +211,142 @@ Out of both phases: **running a spawned task in the background** while the user ``` `original_message` = the user's message, so the bubble shows only what they typed. - - writes `spawnedFrom: {sessionId, taskId}` onto the new session's metadata as it is created by this first + - writes `spawnedFrom: {sessionId, handoffId}` onto the new session's metadata as it is created by this first turn - - marks the task `accepted` with `targetSessionId`, and records `contextEdited` (the sent text's hash ≠ - `excerptHash`) and `contextRemoved`. -4. The target is an ordinary bound launch from this point on. Binding, tools, interrupts, MCP Apps, artifacts, + - marks the handoff `running` with `targetSessionId`, and records `contextEdited` (the sent text's hash ≠ + `excerptHash`) and `contextRemoved` +5. The target is an ordinary bound launch from this point on. Binding, tools, interrupts, MCP Apps, artifacts, compaction and cost rows all behave as they do for any Agent conversation. +**The user's typed message never enters the source thread's history.** The source's model is not called, and its +prefix is untouched. The handoff row is the record of what happened (§4.3). + +**Why review in place.** The source thread is where the user can judge what should leave it: the conversation the +excerpt came from is on screen. Reviewing it after navigating away would mean judging it out of context, and it +would take the user out of the thread they meant to stay in. + **Why a card and not composer text.** An excerpt can be 4k tokens. Put in the textarea, it would bury the user's own sentence and could be sent half-edited by accident. A card keeps the request editable as a sentence and the context editable as a document, and lets the SPA show "Context from *Course planning*" on the sent bubble without parsing the prompt. -**Attachments and invoked skills** from the mention message follow the request to the target, as they do today. -Attachments in the *source* thread are named in the excerpt, not carried over (§4.5). +**Attachments** in the *source* thread are named in the excerpt, not carried over (§4.5). -**Cancelled handoffs.** If the user navigates away without sending, no session exists yet. The task stays `open`, -and one older than 24h reads as `abandoned` (derived when read, so no sweep job is needed). For a `mention_handoff`, -a removed context card still accepts the task, because the user did send; `contextRemoved` records the choice. +**Cancelled handoffs.** Cancel calls `POST /sessions/{id}/handoffs/{handoffId}/dismiss`, and no session is created. +If the user navigates away with the sheet open, the handoff stays `open`, and one older than 24h reads as +`abandoned` (derived when read, so no sweep job is needed). A removed context card still sends the handoff; +`contextRemoved` records the choice. -### 4.3 Phase 0b — Send back +**Inside a Shared Project.** A mention in a project task hands off to an ordinary conversation with the Agent. It +does not inherit `projectId` and doesn't appear in the project's Tasks list; the handoff row in the project task +links to it. -A conversation with `spawnedFrom` shows **"Send to *{source title}*"** in the action row of each assistant message. +### 4.3 Phase 0b — The handoff row and send back -1. `POST /sessions/{targetId}/tasks` with `{kind: "send_back", messageIndex}`. The server renders a **return - excerpt**: that assistant message's text, bounded at `SEND_BACK_MAX_TOKENS = 2000`, cut at a paragraph boundary - with a `[truncated]` marker. It writes a `TASK#` row on the **target**, and returns `{taskId, sourceSessionId, - excerpt, excerptHash}`. -2. The SPA navigates to the source conversation with the card on its composer (`from_conversation` = the spawned - conversation's title) and an empty request, and focuses the composer. -3. On send, the same `handoff_context` path appends the block to the source's next user message. +**The handoff row** sits in the source thread after the message that was last when the handoff was sent +(`sourceMessageIndex`). It is **UI only**: it comes from the `HANDOFF#` row, is replayed on +`GET /sessions/{id}/messages` as `handoffs` (like `toolSummaries`), and never enters the model's history. + +| State | Row reads | Set by | +|---|---|---| +| `running` | "*Rubric Builder* is working on *draft a rubric for this* · Open" | Acceptance on the first turn | +| `needs_input` | "*Rubric Builder* needs your input · Open" (amber) | The turn ends in an interrupt: OAuth consent, a tool approval, `ask_user_question`, a browser sign-in | +| `done` | "*Rubric Builder* finished · Open · **Bring back result**" | The turn completes | +| `stopped` | "*Rubric Builder* stopped · Open" | A `stream_error`, a Stop, or a dropped stream (§4.6) | + +While the target is streaming in this tab, the row follows the SPA's live per-session state. The persisted status +is written by inference-api **after** the turn's `done` block, on the same side of the stream as the `TSUM#` write, +so none of it is on the TTFT path. It follows the target's turns until the first `done` (a `needs_input` handoff +goes back to `running` when the user answers), then stays `done`: from there the conversation is the user's own. + +**Notifications.** The user hears about a handoff's state in three layers. Each one covers a case the one before +it can't. + +| Layer | Reaches the user when | What it shows | +|---|---|---| +| **The row, the unread dot and a toast** | They're in the app, on any page | Toast: "*Rubric Builder* finished · Open", or amber "*Rubric Builder* needs your input · Open", which stays until dismissed | +| **The in-app inbox (the bell)** | They missed the toast, reloaded, or come back later | An inbox entry per state change that needs them | +| **A browser notification** | The tab is open but hidden: another tab, another app | The same one line, from the operating system | + +A handoff that is waiting on the user and says nothing is the worst outcome of running in the background, which is +why *needs your input* uses all three. + +**The in-app inbox.** Shared Projects built a general-purpose inbox (`apis/shared/notifications/`, +`PK=INBOX#{email}`, `SK=NOTIF#{id}` on the projects table, 90-day TTL, `GET /notifications`, the bell in the +sidenav). It isn't tied to projects; handoffs become its second producer. + +- **New kinds:** `handoff_needs_input`, `handoff_done` and `handoff_stopped`. The payload is `{handoffId, + sourceSessionId, targetSessionId, agentId, agentName, sourceTitle}`. The request text is left out. +- **Written by inference-api**, beside the status write after the `done` block (§4.3's table), so no notification + is on the TTFT path. A tab that closes mid-turn still gets its `handoff_stopped`, because `_persist_interruption` + runs server-side. That entry is what the user sees on their next visit, since no toast was ever shown. +- **Read state follows the work, not the bell.** Opening the target conversation marks its `needs_input` and + `stopped` entries read. Opening it or using Bring back result marks `done` read. If the state change lands while + the user is viewing the source or the target, the SPA marks the entry read at once, so the bell counts only what + the user missed. +- ⚠️ Two changes this needs: + - **An IAM grant.** inference-api can `UpdateItem` on the projects table but not `PutItem`, and its policy comment + says it never creates rows there. Add a separate `PutItem` statement limited to inbox rows with a + `dynamodb:LeadingKeys` condition (`INBOX#*`), rather than widening the existing one. + - **The bell's visibility.** The sidenav shows the bell only when `features.projects` is on, because projects were + its only producer. Show it when `features.projects || features.handoffs`. + +**Browser notifications** use the **Notification API** (local): no service worker and no push server. + +- **Only when the page is hidden** (`document.visibilityState === "hidden"`) and a handoff turns *needs your input*, + *done* or *stopped* in this tab. When the page is visible, the toast is enough. +- **Permission is asked at a moment that makes sense.** The first time a handoff needs input, its row and toast offer + "Get a browser notification next time". There's also a setting. Never on page load: browsers downrank sites that + ask before the user has a reason. +- **Generic text:** "*Rubric Builder* needs your input". No request or conversation text, because these can show on + a lock screen. +- **One per handoff:** tagged with the `handoffId`, so a later state replaces the earlier one instead of stacking. + Clicking focuses the tab and opens the target. +- Without permission, or on a browser that doesn't support it (iOS Safari outside an installed web app), nothing + is lost: the inbox has the same entry. + +**Web Push** (notifications with the tab closed) is not in Phase 0. A Phase 0 handoff stops when its tab closes, so +Web Push would only ever announce a stopped run. It becomes worth building with durable handoffs (§4.6), and +scheduled runs have the same "your result is ready and you're not here" problem, so it should be built once for +both. + +**Interrupt prompts stay with the target.** The notification tells the user where to go; the prompt itself is +answered there. A consent, approval, question or sign-in prompt from the target's stream is held for that session +and shown when the user opens it. It must never render over the source thread. ⚠️ Verify first: the +concurrent-streaming refactor gated artifacts, compaction and MCP App state to the viewed session, but left the +OAuth, tool-approval, quota and error services session-tagged or global. Confirm each one scopes to its own +session before 0a ships. + +**Bring back result.** On a `done` row, **Bring back result** calls `POST /sessions/{targetId}/handoffs` with +`{kind: "send_back", messageIndex}` for the target's last assistant message. The server renders a **return +excerpt**: that message's text, bounded at `SEND_BACK_MAX_TOKENS = 2000`, cut at a paragraph boundary with a +`[truncated]` marker, and returns `{handoffId, sourceSessionId, excerpt, excerptHash}`. The card lands on the +source composer (`from_conversation` = the target's title) with an empty request. The user is already there, so +nothing navigates. + +**Send to *{title}*** does the same from the other end: a conversation with `spawnedFrom` shows it in the action row +of each assistant message, and it navigates to the source with the card on its composer. + +On send, the same `handoff_context` path appends the block to the source's next user message. **Cache.** The block goes into a *new* user message on the source. That is an append past the history cachePoint: the source's `toolConfigHash` and `systemPromptHash` don't move. Nothing in Phase 0 writes into existing history, the system prompt or the tool list. -**Trust.** A send-back is model output that may have been shaped by web pages or tool results in the spawned +**Trust.** A returned result is model output that may have been shaped by web pages or tool results in the target conversation. The framing sentence in the wrapper labels it. The user reads it and decides to send it, which is the strongest protection available. It is still not a guarantee, and the spec does not claim it is. -### 4.4 Phase 0c — Suggested tasks (`suggest_task`) +### 4.4 Phase 0c — Suggested handoffs (`suggest_handoff`) -The model can offer a task without being asked: "This rubric deserves its own conversation with Rubric Builder." -It works like Claude Code's `spawn_task` chip. +The model can offer a handoff without being asked: "This rubric deserves its own conversation with Rubric +Builder." It works like Claude Code's `spawn_task` chip. **Tool spec (constant):** ```json { - "name": "suggest_task", + "name": "suggest_handoff", "description": "Offer the user a separate conversation for a piece of work that would sidetrack this one …", "inputSchema": { "type": "object", @@ -212,41 +361,40 @@ It works like Claude Code's `spawn_task` chip. } ``` -- **Non-blocking.** The handler writes `TASK#` (`kind: "suggested"`), queues a `task_suggested` event, and +- **Non-blocking.** The handler writes `HANDOFF#` (`kind: "suggested"`), queues a `handoff_suggested` event, and returns `"Suggestion shown to the user."` The turn continues. It never pauses through `ToolContext.interrupt`; this is not `ask_user_question`. - **No Agent list in the schema.** An `enum` of the user's Agents would change whenever pins change, and each change would re-write `toolConfig`, the first cachePoint. `agent` is free text, matched on the server against the - user's own and pinned Agents (exact name, then case-insensitive). If nothing matches, the chip shows an Agent - picker. The model can see Agent names only if the user mentions them. That is fine: a suggestion without an Agent - is still a useful chip. -- **The prompt is the brief.** Unlike 0a, the model writes the brief, which is the best brief this spec can get. - The chip's **Start** button fetches the task (`GET /sessions/{id}/tasks/{taskId}`) and opens the new - conversation exactly as 0a does, with `prompt` as the card text and an empty request. The user still sees and - can edit it before sending, and the first send accepts the task. -- **Rate limits.** At most one `suggest_task` per turn and three per session. Further calls return `"Suggestion - limit reached; mention it in your reply instead."` rather than failing. + user's own and pinned Agents (exact name, then case-insensitive). If nothing matches, the review sheet shows an + Agent picker. The model can see Agent names only if the user mentions them. That is fine: a suggestion without an + Agent is still a useful chip. +- **The prompt is the request.** The chip's **Start** button fetches the handoff + (`GET /sessions/{id}/handoffs/{handoffId}`) and opens the **same review sheet** as 0a, with the model's `prompt` + as the editable request and the excerpt as the context card. Send runs it in the background exactly as 0a does. +- **Rate limits.** At most one `suggest_handoff` per turn and three per session. Further calls return + `"Suggestion limit reached; mention it in your reply instead."` rather than failing. - **Cost, per the token cost tenet.** The spec is in `toolConfig` for every session that has it. It is read from cache on every call and written once for everyone when it ships. Size it by **marginal CountTokens** (tool list with and without it), not by eyeballing the JSON. Expected: 200–400 tokens. - **Steering.** `ask_user_question` needed a **system-prompt clause** before the model used it well. Expect the same here, and budget for it: a base-prompt change is a one-time cache re-write for every user. Ship the tool without the clause first, measure how often it's used, and add the clause only if it's under-used. -- **Gating.** `TASK_SUGGESTIONS_ENABLED` defaults to off. The catalog entry is `enabledByDefault: false` until the - feedback cohort has lived with it, because a chip-happy model is worse than no chips. +- **Gating.** `HANDOFF_SUGGESTIONS_ENABLED` defaults to off. The catalog entry is `enabledByDefault: false` until + the feedback cohort has lived with it, because a chip-happy model is worse than no chips. **SSE event** (added to the `CLAUDE.md` table in the same PR): | Event | Payload | When | |---|---|---| -| `task_suggested` | `{type, sessionId, taskId, title, tldr, agentId?, agentName?}` | Drained after the tool's `tool_result`, like `steering_applied` | +| `handoff_suggested` | `{type, sessionId, handoffId, title, tldr, agentId?, agentName?}` | Drained after the tool's `tool_result`, like `steering_applied` | -The prompt is **not** on the event. It is fetched when the user presses Start (from the `TASK#` row), so a long +The prompt is **not** on the event. It is fetched when the user presses Start (from the `HANDOFF#` row), so a long brief isn't streamed for chips the user dismisses. **Chip UI.** The chip renders inline after the assistant message that produced it, with **Start** and -**Dismiss**. Chips replay on reload from `GET /sessions/{id}/messages`, which returns them as `tasks` (like -`toolSummaries`). Dismiss is `POST /sessions/{id}/tasks/{taskId}/dismiss`. +**Dismiss**. Once started, it becomes that handoff's row (§4.3). Chips replay on reload with the rest of +`handoffs`. Dismiss is `POST /sessions/{id}/handoffs/{handoffId}/dismiss`. ### 4.5 The excerpt renderer (shared with Phase 1) @@ -276,36 +424,65 @@ A **model-written** brief for 0a (a Nova Micro side-channel like `session_title` uncached 50k-token thread) is a deliberate follow-up, not v1. Build the deterministic version first. Replace it only if `contextEdited` shows users rewriting it. -### 4.6 What Phase 0 leaves out - -- **Background execution.** "Start this and let me keep working here" is appealing, and most of the parts exist: - the headless run entrypoint (scheduled runs F1), the unattended-run unread dot, and pause prompts persisted as - breadcrumbs. But it spends quota with nobody watching, and it inherits F1's per-owner token mint. A follow-up, - after Phase 0's numbers. +### 4.6 Background runs, and what Phase 0 leaves out + +**A background handoff is a second stream in the same tab.** That makes it cheap: no headless run entrypoint, no +per-owner token mint, no server-side SSE reader. The user's session cookie and the ordinary chat path do the work. +It also sets the limits: + +- **It lives as long as the tab.** Closing the tab, or losing the network mid-turn, drops the SSE stream. + Starlette cancels the response and the turn is persisted as **interrupted** (`connection_lost`); it does **not** + finish on its own. The row reads *stopped*, and the target shows the existing interrupted-turn marker with + Continue. While any handoff is running, the SPA's existing "is anything streaming" check (`ChatStateService`) + also guards a page unload, so the browser asks before the tab closes. +- **The quota is the user's, and they started it.** It is one reviewed turn, not an unattended loop, so the + original objection to background work (quota spent with nobody watching) doesn't apply. `quota_warning` and + `quota_exceeded` behave as they do for any turn. +- **No cap on concurrent handoffs in v1.** Each target has its own session and lease, so two handoffs can't + collide. Measure how many run at once before adding a cap. +- **TTFT is unchanged.** The source's next turn doesn't wait for anything, and the target's first-token path adds + only a string prepend. + +**Left out of Phase 0:** +- **Handoffs that outlive the tab.** The durable version runs the target through the headless run entrypoint + (scheduled runs F1), with the server-side `unread` flag and persisted pause breadcrumbs. It inherits F1's + per-owner token mint. Build it only if §4.7's `HandoffStopped` shows dropped streams are common. +- **Web Push.** Notifications with the tab closed need a service worker, VAPID keys, stored subscriptions and a + backend sender. They pay off only with durable handoffs, and should be built once for handoffs and scheduled runs + (§4.3). - **Automatic results.** That is Phase 1. -- **Chips that cross users.** A task is always started by the user who owns the source conversation. +- **Handoffs that cross users.** A handoff is always started by the user who owns the source conversation. +- **Telling the source's model.** The source's model doesn't know a handoff happened until a result is brought + back (§12 Q2). ### 4.7 Phase 0 measurement EMF, under the existing PromptCache namespace so the cost dashboard can show it: -- `TaskCreated` by `kind` -- `TaskAccepted` / `TaskDismissed` / `TaskAbandoned` by `kind` +- `HandoffCreated` by `kind` +- `HandoffSent` / `HandoffCancelled` / `HandoffAbandoned` by `kind` +- `HandoffPromptPreviewed`: the review sheet's full-prompt preview was opened - `HandoffContextEdited` / `HandoffContextRemoved` -- `SendBack` +- `HandoffCompleted` / `HandoffNeedsInput` / `HandoffStopped` +- `HandoffOpenedWhileRunning`: the user opened the target before it finished +- `HandoffNotified` by channel (`inbox`, `browser`), and `HandoffNotificationOpened` by channel +- `HandoffBrowserPermission` by outcome (`granted`, `denied`, `dismissed`) +- `SendBack` by origin (`row` for Bring back result, `target` for Send to *{title}*), with `openedTargetFirst` -The same facts are on the `TASK#` rows, so the per-user view can be queried without the metrics. +The same facts are on the `HANDOFF#` rows, so the per-user view can be queried without the metrics. --- ## 5. Goals and non-goals, by phase **Phase 0 goals** -- A mid-thread mention starts the Agent's conversation **with context the user has reviewed**. -- A spawned conversation's results can go back to the original thread in **one action**. -- The model can offer a task without derailing its turn. +- A mid-thread mention starts the Agent's conversation **with context the user has reviewed**, without taking the + user out of the thread they were in. +- The user can see whether a handoff is running, needs them, or is done, from the original thread. +- A handoff's result can go back to the original thread in **one action**. +- The model can suggest a handoff without derailing its turn. - No prefix is ever re-written; no session is ever served by two configurations. -- Every handoff, edit and send-back is counted. +- Every handoff, review action, completion and send-back is counted. **Phase 1 goals** (only if the gate passes) - A mid-thread mention can run the Agent **in place**, with a **bounded report** that the conversation's own agent @@ -313,7 +490,7 @@ The same facts are on the `TASK#` rows, so the per-user view can be queried with - Repeated consults of the same Agent in the same thread continue that Agent's own session. **Non-goals (both phases)** -- Nesting: a consulted Agent or spawned task cannot itself consult. Depth is 1. +- Nesting: a consulted Agent or a handoff's target cannot itself consult. Depth is 1. - Changing launch semantics: a mention on an empty thread still launches and binds. - An A2A server. If this ever runs across a Runtime boundary, the `streaming=True` rule in `CLAUDE.md` applies. @@ -321,8 +498,8 @@ The same facts are on the `TASK#` rows, so the per-user view can be queried with | Flag | Where | Default in development | Kill switch once shipped | |---|---|---|---| -| `TASK_CHIPS_ENABLED` / `FEATURES.taskChips` | backend / SPA | off | yes; off = today's immediate handoff | -| `TASK_SUGGESTIONS_ENABLED` | backend | off | yes; off = `suggest_task` never registered | +| `HANDOFFS_ENABLED` / `FEATURES.handoffs` | backend / SPA | off | yes; off = today's immediate handoff | +| `HANDOFF_SUGGESTIONS_ENABLED` | backend | off | yes; off = `suggest_handoff` never registered | | `AGENT_CONSULT_ENABLED` / `FEATURES.agentConsult` | backend / SPA | off | yes; off = Phase 0 behaviour | With a backend flag off, a request field it would have consumed (`handoff_context`, `agent_consult`) is @@ -330,13 +507,20 @@ With a backend flag off, a request field it would have consumed (`handoff_contex ## 7. The gate: does Phase 1 get built? -After **six weeks** of Phase 0 in prod at default-on, read the `TASK#` rows and EMF counts: +A background handoff with **Bring back result** already delivers most of what a consult does. Phase 1 adds only +the last step: the result comes back and is answered from **without** the user pressing anything. So the gate asks +whether users already do that step by hand, every time, without looking. + +After **six weeks** of Phase 0 in prod at default-on, read the `HANDOFF#` rows and EMF counts: | Signal | Reading | Implication | |---|---|---| -| `mention_handoff` volume | Small share of Agent conversations | **Stop.** Mid-thread use of Agents is rare; chips are enough | -| `SendBack` / accepted handoffs | **High** (proposed bar: ≥ 25%) | People want the result *in the original thread*. **Build Phase 1** | -| `SendBack` / accepted handoffs | Low | People are happy to keep working in the new conversation. **Stop** | +| `mention` volume | Small share of Agent conversations | **Stop.** Mid-thread use of Agents is rare; handoffs are enough | +| `SendBack` / completed handoffs | **High** (proposed bar: ≥ 25%), mostly **without** `openedTargetFirst` | Users bring the result back unread, so the manual step is overhead. **Build Phase 1** | +| `SendBack` / completed handoffs | High, mostly **after** opening the target | Users work in the target before bringing anything back. Background handoffs fit; a consult would cut that work short. **Stop** | +| `SendBack` / completed handoffs | Low | People are happy to keep working in the new conversation. **Stop** | +| `HandoffOpenedWhileRunning` | High | Users don't want to wait in the background. Consider opening the target on Send as a per-user preference, not Phase 1 | +| `HandoffNeedsInput` / `HandoffStopped` | High | Background runs stall on the user or on the tab. Fix that (interrupt UX, or the durable F1 path in §4.6) before anything else | | `HandoffContextEdited` | High | Improve the brief (model-written, §4.5) **before** Phase 1, which uses the same renderer | | `HandoffContextRemoved` | High | The excerpt is unwanted or distrusted. Investigate before sending excerpts *automatically* in Phase 1 | @@ -347,7 +531,7 @@ already reached. ## 8. Phase 1 — Consults -Everything below is built only if §7 says so. It reuses Phase 0's excerpt renderer, `TASK#`-style linking and +Everything below is built only if §7 says so. It reuses Phase 0's excerpt renderer, `HANDOFF#`-style linking and wrapper conventions. ### 8.1 Decision summary @@ -371,7 +555,7 @@ wrapper conventions. With both features on, a mid-thread mention offers two actions in the `@` menu's confirm step: - **Ask here**: a consult (this phase). -- **Open with context**: a Phase 0a handoff. +- **Hand off**: a Phase 0a handoff, reviewed and run in the background. The gate's data picks the default. `routeMention` gains the consult outcome; the empty-thread and same-Agent branches stay as they are: @@ -381,7 +565,7 @@ if (threadHasMessages) { if (choice === 'consult' && FEATURES.agentConsult) { return { sessionId, assistantId: boundAssistantId, handedOff: false, consultAgentId: mentionedAgentId }; } - return { sessionId: null, assistantId: mentionedAgentId, handedOff: true }; // Phase 0a path + return { sessionId: null, assistantId: mentionedAgentId, handedOff: true }; // Phase 0a: review sheet, background send } ``` @@ -558,15 +742,33 @@ v2 would add a `consult` `PendingInterrupt` kind that points at the consult sess **Phase 0** - **Excerpt renderer.** Deterministic output for the same input, whole-turn truncation, compaction-summary substitution, attachments named not inlined, and a `since_index` that only moves forward. -- **Task lifecycle.** Owner scoping on every `/tasks` route (cookie auth; another user's task id → 404), - `open → accepted | dismissed | abandoned`, and no double-accept. +- **Handoff lifecycle.** Owner scoping on every `/handoffs` route (cookie auth; another user's handoff id → 404), + `open → running → needs_input ⇄ running → done | stopped`, `open → dismissed | abandoned`, and no double-send. - **`handoff_context`.** The server applies the cap again; `displayText` is the user's text only; `contextEdited` / `contextRemoved` are recorded correctly; a flag-off server ignores the field. +- **Prompt preview parity.** A shared fixture: the SPA's preview of `contextBlock` + request equals the model message + the server builds, byte for byte. +- **Status writes.** Written after the `done` block, never before the first token; `needs_input` for each interrupt + kind; `stopped` on `connection_lost`; frozen after the first `done`. +- **Inbox writes.** One entry per state change that needs the user, written after `done`; `handoff_stopped` on + `connection_lost`; no request text in the payload. The IAM statement allows `PutItem` only on `INBOX#` keys + (CDK assertion test). - **Send-back prefix stability.** The source's `toolConfigHash` and `systemPromptHash` are byte-identical before and after a send-back turn. -- **`suggest_task`.** Non-blocking return; per-turn and per-session limits; unmatched `agent` still produces a +- **`suggest_handoff`.** Non-blocking return; per-turn and per-session limits; unmatched `agent` still produces a chip; the spec is byte-identical across users with different pins. -- **SPA.** Extend the mirrored `routeMention` tests; the card renders, edits and removes; replay from `tasks`. +- **SPA.** + - Extend the mirrored `routeMention` tests: a mid-thread mention opens the review sheet and sends nothing. + - The no-navigate send leaves `viewedSessionId` and the route on the source. + - The sheet's card renders, previews, edits and removes; Cancel dismisses. + - The handoff row renders each state live and on replay from `handoffs`. + - A target's interrupt prompt never renders over the source thread. + - The bell shows with `features.handoffs` alone, and renders the three handoff kinds. + - Entries are marked read by opening the target or bringing a result back, and at once when the change lands in + view. + - A browser notification fires only while the page is hidden, never before permission, and a later state + replaces the earlier one. + - Axe clean on the sheet and the row, in both themes. **Phase 1** - A consult turn leaves the parent's `toolConfig` and system prompt byte-identical (fingerprint hashes). @@ -575,29 +777,44 @@ v2 would add a `consult` `PendingInterrupt` kind that points at the consult sess that still runs. The next consult supersedes the pause (drive the real `_InterruptState`). - The excerpt watermark moves forward only. -**Both:** import boundaries. `apis.shared.handoff` holds the renderer and the `TASK#` store; `app_api` and +**Both:** import boundaries. `apis.shared.handoff` holds the renderer and the `HANDOFF#` store; `app_api` and `inference_api` import only from there. ## 10. Delivery plan **Phase 0** -- **PR-0.1 — Excerpt renderer + `TASK#` store** in `apis.shared.handoff`, and the app-api `/tasks` routes - (create, get, dismiss). Acceptance happens on the first send (PR-0.2), not as a route. Flags off. -- **PR-0.2 — `handoff_context` on the invocation path.** Wrapper, `displayText`, task acceptance, EMF. -- **PR-0.3 — SPA handoff.** `routeMention`'s handoff calls `/tasks`, navigates, and shows the context card; the - card on sent bubbles; the toast copy is updated. -- **PR-0.4 — Send back.** Action on assistant messages in spawned conversations, and a card on the source composer. -- **PR-0.5 — `suggest_task`.** Tool, `task_suggested` event, chip UI, replay, `CLAUDE.md` SSE table. Marginal - CountTokens measured and recorded in the PR. -- **Dev validation.** A mid-thread mention with an edited card; a send-back; a suggested chip that is started and - one that is dismissed. The source's `C#` rows show no prefix change on the send-back turn. Then flip defaults on - and start the §7 clock. +- **PR-0.1 — Excerpt renderer + `HANDOFF#` store** in `apis.shared.handoff`, and the app-api `/handoffs` routes + (create, get, dismiss). The create response includes `contextBlock` for the preview. Acceptance happens on the + first send (PR-0.2), not as a route. Flags off. +- **PR-0.2 — `handoff_context` on the invocation path.** Wrapper, `displayText`, acceptance, status writes after + `done`, EMF. +- **PR-0.3 — SPA handoff.** A mid-thread mention opens the review sheet (request, context card, full-prompt + preview); a no-navigate send; the handoff row with live and replayed states; completion and needs-input toasts; + the card on sent bubbles; the unload guard. Verify interrupt-service scoping (§4.3) here. +- **PR-0.3b — Notifications.** The three `handoff_*` inbox kinds, written by inference-api after `done`; the scoped + `PutItem` grant; the bell shown with `features.handoffs` and rendering the new kinds; read-state rules; browser + notifications through the Notification API, with the in-context permission ask. +- **PR-0.4 — Send back.** Bring back result on a `done` row, Send to *{title}* on assistant messages in spawned + conversations, and the card on the source composer. +- **PR-0.5 — `suggest_handoff`.** Tool, `handoff_suggested` event, chip UI that opens the review sheet, replay, + `CLAUDE.md` SSE table. Marginal CountTokens measured and recorded in the PR. +- **Dev validation.** + - A mid-thread mention with an edited card, previewed, sent, and finished while the user stays in the source. + - A handoff that needs OAuth consent: *needs your input* on the row and in the bell, a browser notification with + the tab hidden, and the prompt only in the target. + - A tab closed mid-run: *stopped*, a `handoff_stopped` entry in the bell on the next visit, and Continue works + in the target. + - Bring back result, and a send-back from the target. + - A suggested chip that is started, and one that is dismissed. + - The source's `C#` rows show no prefix change on the send-back turn. + + Then flip defaults on and start the §7 clock. **Phase 1** (after the gate) - **PR-1.1 — Extract `resolve_turn_agent`.** Pure refactor. - **PR-1.2 — Backend consult** (flag off): session prefix + list filter, orchestration, report, `CONSULT#` rows, SSE events, metering attributes, interrupt handling, `CLAUDE.md` SSE table. -- **PR-1.3 — SPA**: Ask here / Open with context choice, consult card, disclosure, replay. +- **PR-1.3 — SPA**: Ask here / Hand off choice, consult card, disclosure, replay. - **PR-1.4 — Transcript view + delete/export integration.** - **Dev validation:** a consult, a follow-up consult continuing the consult session, an OAuth-blocked consult, a Stop mid-consult; parent rows show `hit` with unchanged hashes; measured costs replace §8.10's estimates. @@ -606,7 +823,7 @@ v2 would add a `consult` `PendingInterrupt` kind that points at the consult sess - **A. Parent-initiated `consult_agent` tool** (automatic, not a chip). Needs a constant `toolConfig` spec with no Agent list; Strands re-runs an interrupted tool from the top; it costs an extra parent call. Phase 0c's - `suggest_task` gets most of its value without the nested execution, and a later `consult_agent` could reuse the + `suggest_handoff` gets most of its value without the nested execution, and a later `consult_agent` could reuse the Phase 1 consult function as its handler. - **B. Hand the other Agent the full parent history.** Best answer quality, but the most expensive option: about $0.21 per handoff or consult on a 50k-token thread on Sonnet 4.6. Rejected; the excerpt cap can be tuned upward. @@ -616,19 +833,35 @@ v2 would add a `consult` `PendingInterrupt` kind that points at the consult sess Phase 0 answers the question and reviews the context, for a fraction of the work. - **E. Run consults through the headless run entrypoint (F1)** as a separate Runtime invocation. Real isolation, but it brings in the per-owner token mint, a second container and server-side SSE reading. Revisit if consults, or - background tasks (§4.6), ever need to outlive the request. + background handoffs (§4.6), ever need to outlive the request or the tab. +- **F. Navigate to the target on Send** (this spec's first draft of 0a). Simple, but it takes the user out of the + thread they meant to stay in, and reviewing the context after leaving it means judging it out of context. + Replaced by the in-place review and background run. §7's `HandoffOpenedWhileRunning` would show if users want it + back as a preference. +- **G. Have the source's model announce the handoff.** It costs a model call on the source, adds latency before + anything shows, and has the model describe an action it didn't take. The handoff row says the same thing, + accurately, for nothing. ## 12. Open questions 1. **Excerpt size.** Is 4k tokens enough in practice? Measure `HandoffContextEdited` before tuning. -2. **Auto-send vs. review.** Phase 0a adds one Enter compared with today's immediate handoff. If - `HandoffContextEdited` and `HandoffContextRemoved` are both near zero, should the handoff send immediately with - the card attached? +2. **Should the source's model learn about a handoff?** Today it doesn't until a result is brought back, so "what + did Rubric Builder say?" gets a blank. One option: on the source's next user turn, append a one-line note + (`[Handed off "draft a rubric for this" to Rubric Builder]`). It is append-only, so the cached prefix is safe, but + it is a small per-handoff token cost and a note the user didn't write. 3. **Direct mode (Phase 1).** For Agents whose output *is* the answer, should a consult end the turn with the report as the assistant message and skip the parent call? 4. **Cohort.** An RBAC capability for a feedback cohort first, as scheduled runs did, or flags alone? -5. **Suggestion steering.** Does `suggest_task` need a system-prompt clause, and is its one-time fleet-wide cache +5. **Suggestion steering.** Does `suggest_handoff` need a system-prompt clause, and is its one-time fleet-wide cache write worth paying? 6. **Gate thresholds.** Agree §7's bars before Phase 0 ships. 7. **Model pinning (Phase 1).** An Agent that pins its own model runs the consult on it. Should the card say so, like the composer's "set by this agent" label? +8. **Review sheet default.** Review is required (decided 2026-09-27). Open is only whether the context card + should start expanded for a user's first few handoffs. `HandoffPromptPreviewed` will say whether people look. +9. **Should `handoff_done` go to the inbox at all?** It is the noisiest kind. If `HandoffNotificationOpened` shows + done entries are rarely opened from the bell, keep `done` to the toast and browser notification, and reserve the + inbox for states that need the user. + +**Decided (2026-09-27):** review before sending is required, and the handoff never sends immediately. This replaces +the earlier open question about auto-send.