From d22f8533f61e0bcdf6511591462bd16d460f43ae Mon Sep 17 00:00:00 2001 From: Phil Merrell Date: Sat, 26 Sep 2026 08:54:42 -0600 Subject: [PATCH] Release/1.25.1 (#1363) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(agents): degrade-with-notice binding resolution for project harnesses shared-projects §9.6. A project can have 200 members; blocking the whole turn because one member lacks one bound tool makes the project unusable to them. resolve_agent_invocation(..., degrade=True) drops a missing capability instead of raising, and records it in plan.unavailable: - tools: a denied SERVER drops every ref of it (the gate is per base id, so skipping only the checked ref would let a later scoped ref through); an all-dropped toolset is still an empty toolset, never a fall-through to the request's enabled_tools - model: no override, so the route's chain picks the member's default - skills: filtered; none left means no skill binding - memory space: skipped AgentNoticeEvent (SSE agent_notice) serializes the drops. Ordinary shared agents keep block-with-message (D5) unchanged — degrade defaults off. Co-Authored-By: Claude Opus 5.5 * feat(projects): run a project's harness in chat, with per-project cost (PR-1.4a) shared-projects PR-1.4a. Membership already gates the harness (1.2's access delegation), so the route adds what a project turn needs on top: - Archived projects refuse new turns with a conversational error naming the project. - The resolver runs with degrade=True for a harness; anything dropped is streamed as agent_notice before message_start (SSE only, never the prompt, not persisted). Documented in CLAUDE.md's event table. - "## Project Instructions" replaces "## Assistant-Specific Instructions" for a harness, via compose_agent_system_prompt, which takes no user argument, so every member of a project renders a byte-identical prefix. Every other agent's text is unchanged byte for byte. - preferences.projectId is written at session binding. - projectId rides every C# cost row next to turnAgentId (route → ChatAgent → StreamCoordinator), and the metadata writer adds the call to PROJECT#{id}/COST#{YYYY-MM} plus COST#{YYYY-MM}#USER#{userId}: one atomic ADD each, both UpdateItem (the runtime has no PutItem). The spec's byUser map became per-member rows. Best-effort like every aggregate. - kb-sync already copies apis/shared/projects; the scheduled-runs image now does too (plus dynamo_errors.py) because sessions/metadata.py reaches the repository lazily. Import closure only. Personal instructions (1.4b) are split out: they change every user's system prompt, not only project turns. Co-Authored-By: Claude Opus 5.5 * docs(specs): link the Projects UI mockup from shared-projects §6 Adds the clickable prototype link (claude.ai/artifact/2syz1g6zY2ZsCAkgsWVp7F) with its two deliberate departures from §6, the open questions it raises, and the convention that a PR changing a behavior the mockup shows notes it in its as-built entry for the 1.8 re-sync. Backfills those notes for the two merged PRs that already made such decisions (1.2 delete/transfer/ invite shape, 1.4a agent_notice). Requested from the "Projects UI mockups" session on Phil's behalf. Co-Authored-By: Claude Opus 5.5 * feat(infra): ProjectSessionIndex on sessions-metadata (PR-1.6-infra) shared-projects §3.4. A member's own tasks in one Shared Project, newest first: GSI5_PK = PROJECT#{projectId}#USER#{userId}, GSI5_SK = {lastMessageAt}#{sessionId}, projection ALL. The recency key of SessionRecencyIndex (GSI4), sparse the same way, so the 1.6 backend writes and removes it wherever GSI4 is. Lands alone and ahead of any writer: it is the one GSI this plan adds to sessions-metadata (one GSI per UpdateTable), and it is inert until rows carry GSI5 keys. gsi-inventory.json regenerated; tables-detailed pins the key shape. Co-Authored-By: Claude Opus 5.5 * feat(infra): enable CloudFront access logs on the SPA distribution A user-reported 404 on prod (2026-09-24) could not be traced: app-api and the ALB had no matching request, CloudFront's 4xx metric showed errors in the window, and logging was disabled on the distribution, so nothing could name the URI or client. Requests the edge answers itself (a lazy chunk a deploy removed, an S3 403) never reach the ALB. - Standard logging (legacy) to a dedicated S3 bucket under `spa/`. v2 is configured via CloudWatch vended-log deliveries that must be created in us-east-1, which the single app-region stack cannot do. - Bucket: ACLs enabled (BUCKET_OWNER_PREFERRED; legacy delivery writes through the ACL and fails against BucketOwnerEnforced), SSE-S3, block public access, SSL-only, 90-day expiry. - Cookies are never logged (the SPA session is an httpOnly cookie). - CDK_FRONTEND_ACCESS_LOGS_ENABLED / frontend.accessLogsEnabled: default ON with a kill switch. The bucket is provisioned unconditionally so a disable/re-enable never collides with a RETAINed bucket name. Co-Authored-By: Claude Opus 5.5 * feat(projects): project tasks, project shares and forks (PR-1.6 backend) - Write ProjectSessionIndex (GSI5) keys beside GSI4 for sessions with preferences.projectId: on store, on activity, and removed on soft-delete. Reads now strip all recency keys, so a read-modify-write no longer replays stale GSI4_* extras (which could SET and REMOVE the same attribute). - GET /projects/{id}/tasks: the caller's own sessions in the project, newest first, value-cursor paginated; a missing index degrades to empty. - accessLevel "project" on shares: requires the task's projectId, the caller's membership and an active project; the read check is membership. - SHARED_TASK#{sessionId} pointer on the projects table, rebuilt from the share rows on create/revoke/update/session delete (newest project share wins); GET /projects/{id}/shared-tasks lists them without user ids. - A fork keeps its project (projectId + the project's current harness) when the requester can work in it; otherwise it stays a plain session. Co-Authored-By: Claude Opus 5.5 * fix(spa): recover from stale lazy chunks after a frontend deploy deploy.sh syncs with --delete, so a tab opened before a deploy asks for lazy chunks that no longer exist the first time it navigates to a route it has not loaded (reported on prod 2026-09-24: Settings -> API Keys showed an error; no request ever reached app-api). - isChunkLoadError matches the Chromium, Safari, Firefox and webpack wording (plus HTML-served-as-module), following cause/rejection/error. - A failed navigation becomes a full load of its destination, guarded to one auto-reload per tab per 5 minutes (sessionStorage, fails closed). It is skipped, in favour of a "new version available - Refresh" toast, when a chat stream or upload is in flight, since in-app navigation keeps those alive and a page load would not. - Chunk failures outside navigation (@defer, lazy libraries) reach a custom ErrorHandler and only prompt. - Proactive check (environment.versionCheckEnabled): on tab focus, at most every 10 minutes, compare the running main-.js with the one the no-cache index.html names. Inert on the dev server. - ToastService gains an optional action button. Composer drafts need no change: the composer mirrors to localStorage on every edit, so there is nothing to flush before the reload. Co-Authored-By: Claude Opus 5.5 * fix(users): resolve an email to its live profile, not an arbitrary duplicate Some emails own more than one PROFILE row: the pre-Cognito login keyed users by a numeric employee ID, the current one by the Cognito sub, and nothing retired the old rows. get_user_by_email queried EmailIndex (no sort key) with Limit=1, so which row it returned was arbitrary. - UserRepository.get_users_by_email pages through every match and ranks them: most recent lastLoginAt (compared as instants), then non-numeric ids over legacy numeric ones, then id for determinism. - get_user_by_email returns the first and logs a warning naming the ignored duplicates. - Admin email search returns every profile, live first, so support sees the ambiguity instead of landing on the stale row. - The sharing user search dedupes by email, not id; previously a stale exact match plus the live row from the name scan listed one person twice. - upsert_user logs when a newly created profile's email already belongs to another id (creation only; returning users pay nothing). Co-Authored-By: Claude Opus 5.5 * feat(admin): flag emails with more than one profile in user lookup When an email search matches several profiles, show how many, badge the one that signs in, and print each row's id, so a quota override or tier is not assigned to a legacy duplicate nobody can log in to. Co-Authored-By: Claude Opus 5.5 * chore(scripts): audit legacy duplicate user profiles, retire them in two phases backend/scripts/audit_user_duplicates.py groups users-table PROFILE rows by email, ranks each group with the same live_profile_rank the API uses, and checks every stale id against each place a user id is stored: one Query per indexed key path across 20 tables, an S3 prefix probe per user-keyed bucket, AgentCore Memory sessions for the actor, and, with --deep, a single Scan per table with un-indexed references. It writes a JSON and a Markdown report. Read-only by default. --apply mark sets an unreferenced legacy row to status=inactive with mergedInto/mergedAt (reversible; status=merged would not parse). --apply delete removes only rows marked at least --min-soak-days ago for today's live id. Only numeric-id + uuid pairs whose every check came back zero are eligible; a failed check makes a row incomplete, not unreferenced. Every write is conditional on the lastLoginAt the audit saw, and --apply needs --confirm-prefix. A coverage test fails when a CDK table is neither checked nor listed as holding no user ids. live_profile_rank and item_to_profile become public module functions in the users repository so the script cannot drift from the API's choice of live row. Co-Authored-By: Claude Opus 5.5 * fix(spa): reset the refresh prompt's destination after a successful navigation Found validating #1262 in dev: a navigation that fails on a stale chunk while a stream is in flight shows the Refresh prompt aimed at the failed destination. If the user then opened another view (e.g. went back to the conversation) and clicked Refresh, they were sent to the view that failed minutes earlier instead of reloading where they were. AppUpdateService now clears the remembered destination on NavigationEnd. A failed or cancelled navigation doesn't end in NavigationEnd, so the failure that set it, or a guard refusing the next one, leaves it in place. Co-Authored-By: Claude Opus 5.5 * feat(projects): people directory for inviting members (PR-1.3) - apis/shared/directory: DirectoryAdapter port, DIRECTORY_PROVIDER switch (default users_table; unknown values fall back), and UsersTableDirectory, which pages the whole active partition of StatusLoginIndex instead of the 100 most recent sign-ins, matching email prefix and name in one pass, ranked by match quality then recency. The active list is held 60s per process so a typeahead filters in memory per keystroke. - GET /projects/{id}/directory?q=&limit= (viewer): email, name, hasSignedIn and the person's current memberRole, no user ids; a well-formed unknown email is appended so it can always be invited. Co-Authored-By: Claude Opus 5.5 * chore(scripts): flag legacy user ids that are still in use lastLoginAt cannot show that an old id is idle: an API key minted under it still authenticates as that id, the api-converse path never updates the profile's lastLoginAt, and it 401s a key whose profile row is gone. So deleting such a row would break a live integration outright. The audit now reads each stale id's API keys and sessions-metadata partition and marks the row in_use when it owns an unexpired key or an active scheduled prompt, or when any model call, message or key use is newer than the day its live twin was created (that person's cutover). in_use rows are never retired and lead the Markdown report; every stale row also reports its newest activity so "history only" is visible. Co-Authored-By: Claude Opus 5.5 * feat(spa): shrink the agent-state orb by 20% The 14px core with a 26px halo outweighed the status line beside it. Scale it to an 11px core with a 21px halo (inset -5px) and a 10px glow, keeping the same motion, proportions and timing. The finished-turn recap dot follows so the live-to-settled transition still reads as one element. Co-Authored-By: Claude Opus 5.5 * fix(scripts): count deep-scan references per row, not per attribute The first read-only dev run reported "deep:app-roles 2" for a single TOOL_PREFERENCES row that names the id in both PK and userId. Count one hit per item and list the matching attributes in the sample instead. Co-Authored-By: Claude Opus 5.5 * feat(projects): project settings with version history (PR-1.5a) - GET/PUT /projects/{id}/instructions, /model, /tools, /skills: viewer reads, editor writes on an active project. Tools/skills replace only their own kind. Only what a save adds is validated against the saver (validate_agent_write); a no-op save cuts nothing. - Every change cuts an AgentVersion (createdBy + createdByEmail); the first save also records the starting state. GET .../instructions/versions and .../versions/{n} return history and a diff against the previous version. - The project is the harness's only write path: PUT /assistants and PUT /agents return 409 on a harness. - An archived project's harness resolves every member as viewer, so no agent-level write route (documents, web sources, sync) can change it. - Move diff wire helpers into version_diff; label-able instructions_diff; public ProjectService.authorize. Co-Authored-By: Claude Opus 5.5 * chore(deps): bump dependencies to clear 21 Dependabot alerts Resolves all open Dependabot alerts (2 critical, 5 high, 14 medium) across the four dependency manifests. Every resolved version meets or exceeds its advisory fix version. backend (pyproject.toml + uv.lock) - anyio 4.12.1 -> 4.14.2 (#409 critical GHSA-82r6-8w77-94w6, #408). Transitive via starlette/httpx/fastapi, so added as an explicit pin in the existing "Security: pin transitive deps" block. - soupsieve 2.8.4 -> 2.9.0 (#403, #404, #405, #406) tests/load (pyproject.toml + uv.lock) - pytest 8.4.2 -> 9.0.3 (#395) frontend/ai.client (package.json + package-lock.json) - sharp 0.33.0 -> 0.35.4 (#385, #394 high) - @angular/{common,compiler,core,forms,platform-browser,router, compiler-cli} 21.2.19 -> 21.2.20 (#398, #400, #402). The whole family moves together because Angular packages peer-depend on exact sibling versions. - vitest + @vitest/coverage-v8 4.1.5 -> 4.1.11, pulling @vitest/mocker to 4.1.11 (#392, #393) docs-site (package.json + package-lock.json) - astro 7.1.3 -> 7.2.8 (#390 critical, #389) - sharp 0.35.3 -> 0.35.4 (#388 high) - svgo 4.0.2 -> 4.1.0 (#386 high, #387) - smol-toml 1.7.0 -> 1.9.0 (#391 high) - devalue 5.8.2 -> 5.9.4 (#407) These three are transitive but already sat under permissive caret ranges admitting the patched versions, so no overrides were needed. Verified: backend pytest 9902 passed / 3 skipped; frontend npm ci + build + full suite (301 files, 3823 tests) passed; docs-site npm ci + build (52 pages) clean. npm audit reports 0 vulnerabilities in both JS projects. * feat(projects): project files over the harness's documents (PR-1.5b) - /projects/{id}/knowledge: members list, read and download the project's files (the agent document routes are editor-only); editors on an active project upload, import, crawl and delete through the agent's own route handlers, so provisioning, the byte cap and cleanup are not duplicated. - addedByUserId on every document create path: create_document defaults it to the provenance importer (import, crawl, sync) and device upload passes the uploader. Project responses show addedByEmail from the member list. - Upload, import and crawl responses carry the "everyone in the project can open this file" notice. Co-Authored-By: Claude Opus 5.5 * feat(projects): audit trail, notification inbox, admin.projects (PR-1.7) - project.* audit actions recorded from the project service (lifecycle and membership), settings saves, the knowledge routes and project shares, on target `project`; GET /projects/{id}/audit for editors (no user ids). - Per-user inbox on the projects table (INBOX#{email}, 90-day ttl), keyed by email so an invitation reaches someone who has never signed in: project_invited / role_changed / removed / ownership_transferred. GET /notifications, POST /notifications/{id}/read, POST /notifications/read-all. - admin.projects (delegable): /admin/projects list, detail, force-archive or restore, and the full trail. Scope registry, coverage test and the SPA's scope id list updated. Co-Authored-By: Claude Opus 5.5 * docs(models): add managed-model retirement runbook Sibling of mcp-server-retirement.md. Models differ from tools in three ways that reshape the stages: the runtime model check reads role grants only (not the catalog row or `enabled`), AWS owns the EOL clock, and a model has a successor that can be invoked in its place — so the cutover is a redirect, and a model row is tombstoned rather than deleted (a deleted row plus a surviving `*` grant runs unmetered and uncached). Co-Authored-By: Claude Opus 5.5 * fix(agent): stop truncating agent instructions at ~1,400 characters The 8 KiB cap in SystemPromptBuilder.from_user_prompt was sized for an agent's instructions alone, but it is applied to the whole composed block: the ~6.8K-character default prompt and date, then the instructions. Every agent's instructions were silently cut after ~1,400 characters (133 of 232 prod agents with instructions, 2026-09-24; the longest is 32,887). - MAX_AGENT_INSTRUCTIONS_CHARS (100,000) is the author's cap, enforced at save (Create/UpdateAssistantRequest, project settings) and on a request's system_prompt (the Agent Designer preview, which 422'd past 8 KiB). - The runtime bound becomes that cap plus PLATFORM_PROMPT_HEADROOM (64 KiB) for the platform text composed around it, so a saved agent never truncates. - Tests: a 33K-character agent keeps its last rule, the default prompt plus a default memory block must fit the headroom, and save/preview share a cap. Co-Authored-By: Claude Opus 5.5 * feat(settings): personal instructions in every conversation (PR-1.4b) - UserSettings.personalInstructions (max 4,000 chars) via PUT /users/me/settings; trimmed, blank clears. - Appended last in the instructions block as "## Personal Instructions", below an agent's or project's instructions, with a precedence sentence only when there is something to defer to. A user without them keeps a byte-identical prompt (and cached prefix); project members still share the project part of the prompt. - Applied on plain, agent, project and @-mention turns (not previews); the MCP App dispatch paths add the same text so they reuse the turn's cached agent. One settings read per turn, shared with the default-model lookup. Co-Authored-By: Claude Opus 5.5 * feat(sidenav): scroll the nav entries with the sessions under a pinned New Session Only the session list scrolled; New Session, Agents, Artifacts and Customize sat frozen above it, eating the column on short windows. Now only New Session is pinned and everything below it scrolls as one region. A fade under New Session appears once the region has scrolled, so rows slide out beneath the button instead of being sliced by a hard edge. It is drawn from the sidebar's own background token, so it reads the same in light and dark mode. Co-Authored-By: Claude Opus 5.5 * feat(sidenav): compact the user bar to the nav rows' rhythm 73px -> 57px. The trigger now matches a nav row (40px, rounded-md, the same hover tint), with a 28px avatar that lines up with the nav icons above it. The unread-announcements dot moves with the smaller avatar's corner. Co-Authored-By: Claude Opus 5.5 * feat(sidenav): page the session list as the user scrolls The sidebar fetched a user's entire history on every load (`GET /sessions` with no limit). It now asks for 30 and appends the next page as the end of the list nears view, using the cursor the endpoint already returned. Reloads re-fetch the whole loaded window in one request (`limit` = rows held) rather than page one alone. The list reloads constantly — new session, delete, rename, tab focus, mark unread — and a first-page refetch would drop rows: a new session pushes page one's last row past the boundary, where the pages appended below it do not have it either. A page that lands after a reload started is dropped rather than written, because `resource.set()` aborts the in-flight load. Each sighting of the end-of-list sentinel pays for one page and is re-measured after the page lays out, so a scroll to the bottom never fetches a page too many and a tall window keeps filling. A failed page shows a Retry and stops the loop. Removes the unused `updateSessionsParams` / `resetSessionsParams`. Co-Authored-By: Claude Opus 5.5 * feat(models): lifecycle status with redirect-to-successor on retirement Adds status (active/deprecated/retired), replacedBy, retiresOn and retirementNote to managed models, and resolves a retired model before the grant-only runtime access check at every entry point — /invocations (model and provider swapped at the top of the handler), the saved-default fallback, Agent modelConfig (live record and published snapshot) and /chat/api-converse. A retired model with a successor runs as it, access-checked and priced on the successor; with none it is refused for everyone, wildcard holders included, as a conversational message (410 on the API-key route). Deprecation stays a picker concern: nothing at runtime reads it, and the Designer write check does not filter on it. Admin writes validate the successor (exists, active, enabled, not itself), refuse a non-active default, and DELETE refuses a model another row names as its replacement. See docs/specs/model-retirement.md §7. Co-Authored-By: Claude Opus 5.5 * feat(spa): model retirement guards in the pickers and admin form A deprecated model can be kept but not newly chosen: the chat picker hides it unless selected (then badges it and names the successor), the Designer hides it unless it is the agent's model (then shows a notice), and Settings disables it as a default option. A selection on a retired model follows replacedBy to the successor with a once-per-tab toast instead of a silent fallback. The admin form gains Status / Replaced by / Retires on / Note, and the list badges non-active rows. The retirement helpers move to shared/utils/retirement.ts, re-exported from tool.service.ts so no tool surface changes. Co-Authored-By: Claude Opus 5.5 * docs(models): record the built §7 in the model retirement runbook Co-Authored-By: Claude Opus 5.5 * feat(projects): Projects SPA shell, Overview, Members and Settings (PR-1.8a) - /projects: your own and shared projects, filtered All / Mine / Shared with me / Archived, with each card naming your role; "not available" when the kill switch is off. - /projects/:id/:tab (the bare URL redirects to overview): header with role, archived notice, pill tabs. - Overview: a composer that starts a task bound to the project's agent, and a rail with the instructions and the people. - Members: people picker over /directory (existing members marked, pasted lists staged at once, unknown emails allowed), role changes, remove, leave, and "Make owner" for editors who have opened the project. - Settings: details, instructions, model, tools and skills (each save is a version), history with per-version diffs, and owner controls (editors manage members, archive/restore, delete). - Sidenav entry. Specs for the API client, facade, picker and both pages; axe clean in light and dark. Co-Authored-By: Claude Opus 5.5 * feat(spa): chime when dictation starts and stops listening Two short synthesized tones bracket a dictation: rising (E5 -> B5) once the mic is live, falling (B5 -> E5) when it goes off. They track the mic, not the buttons: start plays on the transition into `listening` (the first moment speech is captured), stop on any way out of it — Done, Cancel, the time limit or a dropped socket. An attempt that never reached `listening` plays neither. Web Audio, no asset files. The AudioContext is primed synchronously inside `start()` — before the first await — because browsers only allow audio after a user gesture and the start chime lands well after the click. Every failure path (no Web Audio, a refused context) is a silent no-op. Co-Authored-By: Claude Opus 5.5 * feat(spa): compact send-less composer, meta line below, agent breadcrumb Drop the send button (Enter sends) and give its other jobs their own homes: Stop takes voice mode's slot while a response streams (Escape also stops, once no menu or dictation claims the key), touch devices get a Send button only once there is a draft (return adds a line there), and an upload that blocks Enter says so on the status line. In a conversation the composer is one row — attach, text, speech controls — with a line beneath it: cost and context on the left (or a short status), the model picker on the right. A draft that wraps unfolds to text-on-top with the controls in a bar beneath, and folds back only once empty so the layout cannot flip at the wrap point. The empty state keeps the tall composer. The first send swaps one for the other across the `/` -> `/s/:id` rebuild, so ComposerHandoffService carries the tall composer's height and the compact one animates down from it. The `@`-mention and `/skill` chips go: the tokens are tinted in place by a mirror layer behind the textarea, and deleting an `@Name` now un-mentions (the binding used to outlive its text). The bound agent leaves the composer for the top nav, as a breadcrumb before the conversation title (`Agent / Title`) projected through a new `[topnavCrumb]` slot, with its menu opening downward. The model dropdown gains a compact size for the meta line, and the cost/context popovers anchor left now that the badge sits at the composer's left edge. Co-Authored-By: Claude Opus 5.5 * feat(projects): Tasks, Files, share-to-project, sidebar grouping, agent_notice (PR-1.8b) - Tasks tab: the caller's tasks (paged, open on the project's agent) and tasks shared with the project (open, continue as a fork, revoke). - Share modal: "Project members" access level for a task in a project, with the API's 403/409 detail shown inline. - Files tab over /projects/{id}/knowledge: list with "Added by", upload with the server's sharing notice, polling, download, delete; viewers read and download only. - Sidebar: project tasks grouped under a project heading inside each time bucket; optimistic rows adopt the API's preferences. - agent_notice SSE event: parser type/validator/callback, per-session notice service, dismissible banner above the composer. Co-Authored-By: Claude Opus 5.5 * fix(inference): fall back to the catalog's default model, not an unpriced id A turn that names no model (scheduled runs, Agents without a modelConfig, model_id: null) fell through to the hard-coded Defaults.MODEL_ID, a us.* Haiku id with no catalog row where prod registers global.* ids. It priced to None, so those turns were unmetered and free against quota. The fallback is now: the user's saved default, then the catalog's enabled isDefault row (the model the SPA already pre-selects), then Defaults.MODEL_ID only when the catalog has no default. A saved default with no catalog row is treated as unset. The MCP App dispatch paths share the resolver so they keep hitting the turn's cached agent. Co-Authored-By: Claude Opus 5.5 * feat(observability): alarm on model calls that price to nothing Emit an UnmeteredModelCall EMF metric (chat and voice) whenever a call spends tokens but gets no cost, because it then reaches neither the cost rollups nor quota. Adds a threshold-0 alarm and dashboard widgets that name the model id missing a catalog row. Co-Authored-By: Claude Opus 5.5 * Platform self-service: system tool tier + account tools (flag-gated) (#1279) * feat(tools): system/hidden tool tier (platform self-service Phase 1) Adds a platform-shipped 'system' tool tier on top of the existing always_on primitive, plus a 'hidden' display flag, per .kiro/specs/platform-self-service. - ToolDefinition gains system + hidden (default false; backward-compatible dynamo round-trip). system implies always_on via the normalizing validator. - freshness: new system-tool-ids snapshot filled by the same single catalog pass; cleared on invalidate. - always_on.py: resolve_system_tool_ids — like resolve_always_on_tool_ids but NOT gated on ADMIN_ALWAYS_ON_TOOLS_ENABLED (system is part of the app). - inference routes + voice: system tools unioned into every turn, including agent-bound turns (author scopes user-facing tools, not platform plumbing). - app_api tools service: surface hidden on UserToolAccess so the SPA excludes it from the toggle list but keeps it for tool-use event labeling. - Tests: model round-trip/validator, freshness snapshot, and the invocation seam (system applies on bound turns; system+always_on combine; dedup). 63 targeted tests green; changed regions ruff-clean. * feat(tools): account self-service tools, read-only pilot (platform self-service Phase 2-3) Phase 2 (identity binding) + Phase 3 (Account & Usage read tools): whoami, get_my_quota, get_my_settings. Identity is captured by closure (make_*_tool factories injected as extra_tools), matching the six existing per-request tool families -- the runtime does not populate Strands ToolContext, so this is the proven pattern and satisfies Req 3 more strongly (tools take no user arg). - get_my_quota is read-only: computes via resolver + cost aggregator, NOT QuotaChecker.check_quota (which records warning/block events). - Gated behind PLATFORM_SELF_SERVICE_ENABLED (default off) per the rollout plan, so agent-cache eligibility and the cacheable prefix are unchanged when off; when on the tools are key-described (close over user_id) and cacheable. - Catalog metadata (system/hidden) added for transcript labeling; whoami hidden, quota/settings visible-but-locked. DynamoDB row seeding and SPA panel-hiding deferred (not needed for the pilot). - 18 new tests; targeted suites green. * feat(tools): admin-governed runtime off-switch for system account tools Reverses the earlier 'system tools are not an admin knob' call after local testing showed code-only delivery needs a redeploy to disable a tool. - get_system_tool_ids now filters on status: a system row set to disabled/ deprecated drops out of the injected set within the freshness TTL -- an admin off-switch with no redeploy (via the existing admin Tools panel). - _build_account_tools injects a tool only when its id is already in the turn's effective set (row present, system, active, RBAC-granted), so a disabled row or ungranted role removes it live. - Seed whoami/get_my_quota/get_my_settings rows in seed_bootstrap_data.py (system + hidden per rule + isPublic). Fixes test_seed_matches_tool_catalog parity (Phase 3 catalogued them without seed rows) and makes them appear in the admin panel. Run the seeder once per env; PLATFORM_SELF_SERVICE_ENABLED remains the per-env master gate. - Tests: freshness excludes non-active system rows; builder injects only ids in the effective set. 46 targeted + 183 broader green. * fix(tools): add 'account' to ToolCategory enum Seeded account tools (whoami, get_my_quota, get_my_settings) use category='account', but the DynamoDB-backed ToolDefinition.ToolCategory enum lacked that member, so list_tools() 500'd when reading the rows back and GET /tools/ failed on load. Adds ACCOUNT to the enum (the static tool_catalog enum already had it) plus a from_dynamo_item round-trip regression test. * feat(self-service): add set_default_model confirmed-write tool Phase 6 of platform-self-service: the first WRITE in the Account & Usage tier, completing the about-the-user capability. - set_default_model is a system/always-on injected tool (like the reads), admin-governable via its catalog row, gated by PLATFORM_SELF_SERVICE_ENABLED. - Two-step/confirmation-gated: previews the change first, applies only on confirm=true after the user agrees. - Only ever changes the caller's own defaultModelId (identity closure-bound, never an argument) and only to a model they may use — accessibility mirrors ModelAccessService.filter_accessible_models via shared RBAC (AppRoleService) + list_all_managed_models, staying inside the agents->shared import boundary. - Persists the record UUID (what defaultModelId is keyed on) via the user-settings repository's idempotent update_settings. - Catalog + seed rows added (visible-but-locked); 12 new tests (handshake, accessibility reject, ambiguity, write, never-raise, identity). * docs(spec): record phase re-order + Phase 6 done in platform-self-service tasks * feat(projects): notification bell, Activity tab, personal instructions (PR-1.8c) - Notification bell beside the user menu: unread badge from GET /notifications, panel of the newest 20, open marks read and goes to the project, mark all as read; refreshes on open and on tab focus. - Activity tab (editors and owner): the project's audit trail as sentences, paged, with links to the settings version a save cut. - Settings > Chat: personal instructions (4,000 characters), saved and cleared inline. Co-Authored-By: Claude Opus 5.5 * docs(projects): user guide, admin page and env vars for Shared Projects (PR-1.9) - docs-site features/projects.md: roles, tabs, tasks and sharing, files, settings history, notifications, personal instructions, lifecycle. - docs-site admin/projects.md: admin.projects routes, the kill switch and what it doesn't stop, configuration, where the data lives. - .env.example and the docs-site env-var page: Shared Projects variables. Co-Authored-By: Claude Opus 5.5 * fix(projects): make PROJECTS_ENABLED a full stop; drop the Projects nav item - A project harness now resolves no role for anyone while PROJECTS_ENABLED=false, so the kill switch stops project chat turns on inference-api and the agent document/sync routes on app-api, not only the /projects surface. - A turn in a project task while Projects are off gets a conversational message saying so, instead of a bare 403. - The sidenav no longer has a Projects entry: a menu item must not depend on a request to know whether the feature is on. /projects stays reachable by URL, from sidebar project headings and from notifications. Co-Authored-By: Claude Opus 5.5 * docs(projects): kill switch is a full stop; no Projects nav item Describe the behavior after fix/projects-kill-switch: a project's assistant refuses everyone while PROJECTS_ENABLED=false, a project task turn gets a conversational message, and /projects is reached by URL, sidebar project headings or notifications rather than a menu item. Co-Authored-By: Claude Opus 5.5 * fix(spa): tighten the gap beneath the compact composer's meta line The cost/model line's 32px row already leaves slack beneath its text, so the footer's full 1rem bottom padding left ~26px below the line against ~14px above it. Compact footers (conversation + embedded preview) now pad 0.5rem; the empty-state composer, which has no meta line, keeps 1rem. Co-Authored-By: Claude Opus 5.5 * fix(spa): drop the top nav's below-lg separator It sat between the conversation title and an empty slot, so it divided nothing. Co-Authored-By: Claude Opus 5.5 * test(projects): assert the agent routes' permission check also refuses while Projects are off Co-Authored-By: Claude Opus 5.5 * fix(agents): show a retiring model on the agent detail page and in runnability The detail page's Details panel named an agent's pinned model even after it was retired, and both "will it run?" checks (per viewer, and per role for pins) answered ready for an agent on a retired model with no successor — which the runtime now refuses for everyone. GET /agents/{id} now carries modelRetirement (status, successor display name, date, note) and the page renders a line under the Model row. resolve_runnability and the role-pin diff resolve retirement first, like agent_binding_resolver: a redirect is checked against the successor, and a retired model with no successor is missing regardless of grants. Co-Authored-By: Claude Opus 5.5 * feat(admin): simplify Manage Models rows; move delete to the edit page Each row is now drag handle, icon, name/id, price, edit, toggle. - Drop the provider chip: the model icon already names the provider. - Drop the detail accordion. Pricing, the one detail admins compare across rows, moves inline as an input/output per-1M column; roles and modalities stay on the edit page. - Move the enable switch to the far right so state reads down one column. Drop the Enabled/Disabled label; disabled rows dim instead. - Move delete off the row to a Danger zone on the edit page. Deleting a model's row leaves its usage unpriced and breaks retirement redirects, so it should not be one click away on every row. The section points at Status: Retired as the safer alternative. Co-Authored-By: Claude Opus 5.5 * fix(agents): draw the detail-page hero as our chat composer The marketplace detail hero showed the @-mention prompt in a generic search pill — rounded-full, blue bold mention, black send-arrow button — none of which our composer has. Redraw it as the compact composer: rounded-2xl shell with the same outline and shadow, leading attach, trailing dictate and voice glyphs, the composer's @mention tint, and no send button (Enter sends on desktop). Still display-only. Co-Authored-By: Claude Opus 5.5 * style(tooltip): tighten tooltip sizing Tooltips read as oversized next to the icon buttons they label. Drop to 12px medium text with px-2/py-1 padding, a smaller arrow and shadow, and halve the overlay offset so the tooltip sits closer to its trigger. Co-Authored-By: Claude Opus 5.5 * feat(flags): in-development features default off; SPA feature switches in Angular environments - Policy (CLAUDE.MD "Feature Flags"): new and in-development features are opt-in per deployment; finished features move to default-on deliberately. - Shared Projects is the first: PROJECTS_ENABLED / CDK_PROJECTS_ENABLED now enable only on "true" (unset = off). - SPA: `features` in src/environments/environment*.ts (typed by feature-flags.ts, read via the FEATURES token). New environment.development.ts + `dev-deploy` build configuration; build.sh builds SPA_BUILD_CONFIGURATION, which frontend-deploy.yml and nightly-deploy-pipeline.yml set to match the environment they deploy to. - Projects nav item and notification bell return behind features.projects; /projects routes canMatch on it; project grouping in the sidebar and the "Project members" share option follow it too. Dev: on. Prod and local: off. - Breadcrumbs for agents: root CLAUDE.MD, frontend .claude/CLAUDE.md, feature_flags.py, feature-flags.ts and each environment file. Co-Authored-By: Claude Opus 5.5 * docs(projects): Projects are opt-in while in development; SPA switch in environments Follow feature/feature-flags-default-off: PROJECTS_ENABLED / CDK_PROJECTS_ENABLED enable only on "true", and the SPA's features.projects (src/environments) decides whether the UI offers Projects, including the nav item. Co-Authored-By: Claude Opus 5.5 * docs(specs): agent handoffs and consults for mid-thread @-mentions Phase 0 (task chips): hand off to a new conversation with a reviewable context card, send results back, and let the model suggest tasks. Phase 1 (consults), where an @-mentioned Agent runs in place and reports back, is built only if Phase 0's measured send-back rate passes a gate. Co-Authored-By: Claude Opus 5.5 * feat(memory-audit): Phase 0.1 AgentCore Memory baseline audit + decision record scripts/memory-audit/audit.py: read-only inventory (strategies and namespace templates vs the backend's, actor-id classes incl. the session-id fallback, record census by strategy/namespace shape, extraction jobs, runtime log signals) and a dev-only probe that replays retrieval with the runtime's top-k and relevance cut against a synthetic actor, then cleans up. raw/ output is personal data and stays out of the repo; summary.json is aggregates only. docs/specs/memory-baseline-decision.md: dev evidence. Memory is written, extracted and correctly namespaced, but the 0.7 relevance cut discards every realistic hit (questions score 0.57-0.67), so no dev turn received context in 7 days and the two-chat test failed. Recommends C (hybrid), conditional on re-testing after 0.2 lowers the cut. tests/supply_chain/test_memory_audit.py pins the audit's copies of the backend namespace templates and retrieval parameters. Co-Authored-By: Claude Opus 5.5 * fix(memory): Phase 0.2 AgentCore Memory fixes from the baseline audit - Relevance cut 0.7 -> 0.5 (constants default). In dev, questions about a stored fact scored 0.57-0.67 and unrelated records <= 0.40, so 0.7 dropped every realistic hit and no turn ever received memory context. - Session delete also purges that session's SUMMARIZATION records (exact per-session namespace, batches of 100), even after its events expired. Semantic facts/preferences carry no source session and are left alone. - Share-fork writes events with extractionMode="SKIP", so another user's messages stay in the fork's history but never feed the forker's long-term extraction. - Runtime role: both memory statements share RUNTIME_MEMORY_ACTIONS (the scoped one lacked GetMemory under a wrong comment). - Stale "write-only" lines corrected in two specs and the app_context_dispatch docstring (which cited a deleted analysis doc). - CDK test pins strategy names, no custom namespaces, 90-day expiry and identical Runtime memory action sets. Co-Authored-By: Claude Opus 5.5 * fix(projects): purge tears down the managed KB, rename reaches the harness, bell passes axe Three bugs found validating Shared Projects Phase 1 on dev (2026-09-25). - Deleting a project, or an ordinary agent, left its KB# record and the born-managed Bedrock knowledge base ACTIVE and billing. Both delete paths now queue the record in a new `teardown` migration state (generation bump fences any in-flight worker); the kb-migration worker, which holds the delete grant app-api lacks, runs the tombstoned saga under its lease, deletes the data source and knowledge base, then removes the record. Failures re-queue instead of going to `failed`. `assert_deletable` now refuses a harness before DELETE /assistants soft-deletes its files. - PATCH /projects/{id} renamed META only; the harness (and the chat breadcrumb) kept the creation-time name. update_project now carries name and description onto the harness without cutting a version, and project history reports only settings fields. - The notification panel's loading/error/empty states were a role="menu" with no menu items (critical axe aria-required-children). They are now a disabled menuitem; the header

is a div. - Activity tab names models, tools and skills from the bindable palettes. Co-Authored-By: Claude Opus 5.5 * fix(projects): show a project task's crumb as the project, not an agent A project task binds the project's hidden harness Agent, so the chat's breadcrumb rendered it as an ordinary Agent — offering Edit agent and Share settings, both of which the harness refuses (409 on edit; share mutations return not-found). - GET /assistants/{id} now carries `kind` and `projectId` (omitted for every other agent). - The agent indicator takes a `projectId`: project mark, "Project" in place of the owner line (the harness keeps its creator after a transfer), New task + Open project, no Edit / Share, and "fixed by this project". - The crumb reads the project's name from the project list when loaded, since renaming a project does not rename its harness. Co-Authored-By: Claude Opus 5.5 * docs(memory): record the passing Phase 0.2 re-test; decision C (hybrid) On dev, after 0.2 deployed (relevance_score=0.5 in the agent-build log), chat B recalled a fact stated in chat A and the runtime logged "Retrieved 1 customer context items" for that turn. Deleting chat A purged its summary record. The decision record moves from recommendation to decided: C (hybrid). Co-Authored-By: Claude Opus 5.5 * feat(chat): context meter with itemized attribution; fold cost into its panel Composer meta row: the running cost figure is gone. A context ring now sits to the right of the model picker (percentage appears inline only at >=70%). Hovering, tapping or pressing it opens a panel with: - what filled the window: System instructions, Skills, Memory, Tools, Messages, Free space, as a stacked bar plus rows with sub-items; - this conversation's cost and the user's quota (replaces the cost tooltip). Replaces SessionCostBadgeComponent. Backend: the measured system/tools totals are itemized by character share (context_itemization.py), at zero extra CountTokens calls: - system -> System instructions (platform / agent / project / personal / mode sections) + Skills (per-skill rows) + Memory (bound Memory Space); - tools -> children by origin (built-in, each Gateway target, each external MCP server by catalog display name, skill tools, memory tools). Each row set sums to the measured total, so prefixTokens and compaction are unchanged; the long-TTL static-prefix size now sums every non-messages partition. No added pre-model-call latency: the hook still stores only the three measured totals; itemizing happens on read (final metadata SSE, persisted row), after the model has answered, memoized per agent. The itemized breakdown is persisted on the turn's last message row (contextBreakdown) and re-seeds the meter on reload. It stays out of the admin CALL_ROW_PROJECTION: labels can name skills/servers and `label` is a content-bearing path segment. Co-Authored-By: Claude Opus 5.5 * feat(agent-cache): cache agents bound to a Memory Space (Shared Projects 2.1) Memory-Space tools close over the resolved binding (space id, name, access) and read the space live on every call, so the binding becomes an agent cache key element instead of a cache veto: - _create_cache_key gains memory_binding (digest; "" without a binding), placed ahead of the document/assistant/skills elements so their positions do not move. - injected_tools_are_key_described drops the has_memory_binding veto. - get_agent takes memory_binding, stamps it on the construction snapshot; PausedTurnSnapshot.memory_binding persists it and resume replays it, so a paused memory-bound agent is found again. - Resume passes cache_write=False. Resume builds no injected tools, so a resume *miss* used to write a tool-less agent into the slot the next plain turn hits, silently dropping artifact/document/spreadsheet (and now memory) tools. Pre-existing for every injected family. AGENT_CACHE_INJECTED_TOOLS_ENABLED=false still restores the blanket bypass. Co-Authored-By: Claude Opus 5.5 * docs: make time to first token a hard budget in CLAUDE.MD Nothing may add latency between a request arriving and the first streamed token (request handling, agent build, MCP pre-flight, BeforeInvocation / BeforeModelCall hooks). Hooks capture raw facts; derived work runs after the model answers, memoized. Anything unavoidable must be measured at realistic scale, stated in the PR, and explicitly accepted by the developer. Prompted by this PR: the first cut of context itemization ran inside BeforeModelCallEvent at ~1.4ms per model call before it moved to on-read. Co-Authored-By: Claude Opus 5.5 * fix(chat): render the context meter from the start so the model picker never shifts The meter only mounted once a turn had reported cost or a context window, so it popped in beside the model picker after the first response and pushed the picker sideways. It now always renders under the compact composer: an empty ring until the first response is measured (arc hidden while empty, so the round cap doesn't paint a stray dot), and a panel that says the window is "Not measured yet" above the conversation's cost and quota. Co-Authored-By: Claude Opus 5.5 * feat(chat): context meter signals urgency by colour only Drop the percentage that appeared beside the ring at >=70% full. It widened the trigger and nudged the model picker left; the ring's colour (success -> info -> warning -> danger) now carries urgency on its own, and the exact figure stays in the panel and the trigger's aria-label. Co-Authored-By: Claude Opus 5.5 * feat(memory): memory block behind its own prompt-cache point (Shared Projects 2.2) The bound Memory-Space block used to be appended to system_prompt, so it sat inside and inside the single cached system block: every index edit rewrote the whole static prefix, and member-written text was wrapped as "instructions". - Routes hydrate the block into `memory_context`, passed through get_agent -> BaseAgent -> AgentFactory, which sends [static, cachePoint, memory, cachePoint]: the fourth and last point (tools, system, memory, message). Turns without memory send today's bytes. - Rendered as ; the name is attribute-escaped and a closing tag in member text is defused. - Budget is token-based (MEMORY_INJECTION_MAX_TOKENS, default 6,000 = the old 24 KB; MEMORY_INJECTION_MAX_BYTES still honored). - Cache key folds memory_context into the prompt hash; PausedTurnSnapshot persists it and resume replays it. Voice keeps memory in its prompt. - contextBreakdown's memory marker moves to the new tag. Memory no longer shares the user-instructions truncation headroom. Baseline measured on dev before the change; the after-run follows deploy. Co-Authored-By: Claude Opus 5.5 * chore(kaizen): weekly research scan 2026-09-25 Generated by the kaizen-research skill. Top 5 ideas appended to docs/kaizen/review-queue.md for the kaizen-review-prep run later this morning. Co-Authored-By: Claude Opus 5.5 (1M context) * fix(agents): DELETE /agents runs the full cleanup; shares go with the agent #1293 made agent deletion queue the managed-KB teardown and clean up documents, but only on DELETE /assistants/{id}. The SPA's Agents page calls DELETE /agents/{id}, which deleted only the record, leaving DOC# rows (and their S3 objects and vectors), sync policies, and the KB# record with its Bedrock knowledge base behind. Both routes now call one sequence, agent_designer/services/agent_deletion: refuse or 404 before touching anything, queue the KB teardown, soft-delete every page of documents (was first 1,000 only), delete sync policies, delete the record, clean up documents in the background. _delete_assistant_cloud also removes SHARE# rows. Co-Authored-By: Claude Opus 5.5 * chore(kaizen): weekly review prep 2026-09-25 Generated by kaizen-review-prep. Ranked agenda for the 10-15 min decision pass; queue updated with prior-review outcomes. Co-Authored-By: Claude Opus 5.5 (1M context) * chore(agents): script to clear rows deleted agents left behind; fix vector probe batch backend/scripts/cleanup_orphaned_agent_rows.py finds AST# partitions with no METADATA row (older than --min-age-hours) and, with --apply and a matching --confirm-prefix, clears them through the app's own paths: KB# queued for teardown, DOC# through cleanup_document_resources, leftover S3 objects, sync policies, versions, reports, share and crawl rows. Report-only by default. Run on dev 2026-09-25: 10 orphaned partitions cleared. delete_vectors_for_document probed 500 keys per GetVectors call; the API allows 100, so every probe failed and each delete listed the whole index. Co-Authored-By: Claude Opus 5.5 * docs(projects): record the 2.2 prompt-cache after-measurement (A wins) Same agent, index and procedure as the baseline: the turn after a memory edit now reads 12,552 and writes 1,775 (was 10,477 / 3,824). The static system prompt is read instead of rewritten; the A-vs-B spike closes on A. Co-Authored-By: Claude Opus 5.5 * chore(tooling): align npm version pins and harden lockfile-sync check Three different npm versions were in play, which is why bumping the frontend dependencies in #1268 required working around an npm crash: dev container 11.2.0 (.devcontainer/Dockerfile ARG + packageManager) GitHub Actions 10.9.x (bundled with Node 22; nothing pinned it) authored #1268 12.1.0 (npx npm@latest, as a workaround) npm 11.2.0 cannot resolve this repo's frontend dependency graph: it crashes with "Cannot read properties of null (reading 'edgesOut')" in arborist's #loadPeerSet on node_modules/vitest, from a vitest <-> @vitest/coverage-v8 peerOptional cycle. That is a resolution-time bug, so `npm ci` (which replays the lockfile rather than resolving) was never affected and no build, test, or deploy path was broken — but authoring a lockfile was impossible without reaching for a newer npm out-of-band. Changes: - Raise the pin to 12.1.0 in .devcontainer/Dockerfile and in frontend/ai.client/package.json "packageManager". 12.1.0 is chosen over the more conservative 11.20.0 because the frontend lockfile now on develop was authored by 12.1.0, so aligning on it makes lockfile regeneration a no-op instead of re-running the resolution that was crashing. Engines are satisfied: npm 12.1.0 requires Node ^22.22.2 || ^24.15.0 || >=26.0.0; the dev container pins 22.22.3 and CI's node-version '22' currently resolves to 22.23.3. - Pin npm in version-check.yml, the only CI job that resolves rather than replays a lockfile. It regenerates the frontend and infrastructure lockfiles and asserts they are byte-identical to what is committed, so authoring and validating with different npm majors can fail the gate on a correct lockfile. The version is read from the "packageManager" field rather than hardcoded, keeping one source of truth. - Stop discarding npm's stderr in that same step and check its exit status. Previously a resolution crash left the lockfile untouched, so the following `git diff --exit-code` came back clean and the check passed vacuously — the failure mode it exists to catch was the one it could not report. Not verified (Docker unavailable in this environment): whether infrastructure/package-lock.json is byte-stable under npm 12.1.0, and whether npm 12 changes `npm ci` behavior for the three npm projects. See the PR description for the follow-up this needs before the next release PR to main. * fix(compaction): never offload skill instructions from the skills tool The Strands AgentSkills activation tool (`skills`) returns a skill's SKILL.md body: instructions the model must follow verbatim. A prod readout on 2026-09-25 (ToolResultOffloaded, toolName=skills) showed one 4,647-token activation swapped for a preview plus a retrieval handle, which quietly degrades skill-following. Add `skills` to OFFLOAD_EXEMPT_TOOLS beside document_read. Its size is bounded by the skill author, not an external payload. read_skill_file stays offloadable: reference files run up to 1 MiB of arbitrary text, which is what preview plus retrieval is for. Memory-space tools return user/agent-written data, not instructions, and also stay offloadable. Per-turn tool payload only; the cacheable prefix is unaffected. Co-Authored-By: Claude Opus 5.5 * fix(agents): stop Strands printing streamed model text to stdout AgentFactory built Agent without a callback_handler, so Strands 1.55 installed PrintingCallbackHandler, which print()s every streamed text delta to stdout with end="". The runtime ships stdout to CloudWatch, so response text reached the AgentCore Runtime log group and, unterminated, was glued onto the front of the next EMF line. Pass callback_handler=None. Nothing consumes callback events: the stream processor reads agent.stream_async() directly. The test builds a real Agent and streams a turn, so it fails if the SDK default comes back. Co-Authored-By: Claude Opus 5.5 * fix(observability): start every EMF record on a fresh line stdout is shared with anything in the process that writes to it. An unterminated write turns into a prefix of the next EMF JSON line, and CloudWatch silently declines to extract it as a metric. Prefix each record with a newline so a stray write can't corrupt a metric. Co-Authored-By: Claude Opus 5.5 * feat(compaction): offline quality-veto harness (slice 1) Adds backend/scripts/compaction_quality_harness.py and the compaction_quality package: a seeded corpus of long editing sessions with planted, exactly-scorable facts (constraint, decision, reference, superseded), replayed through the production TurnBasedSessionManager under each arm (full, model_relative, legacy, floor_50) and pace (restore, cold, warm). Only I/O is stubbed; the cut, the deferred apply, the restore slice, anchor truncation and bound_summary are the real code. Steps: cut (free, writes probe-time histories plus a plant-availability table), records and ask (Bedrock spend, estimate-first, --yes to run, a cachePoint after the shared history), score (exact match, paired per family against the full-history control, exact McNemar, n on every row). The clock stub stamps real time on save and inserts the pace gap only between turns; backdating every save hid the restore-path missed free apply, which a tripwire test now pins until it is fixed. Co-Authored-By: Claude Opus 5.5 * docs(compaction): record the harness, the control, and the prod readout Spec §5 gains a status note above the waiver: the harness exists, the control is full history rather than the kill switch, and the first full run waits for the missed-free-apply fix. The scoping doc marks slice 1 built, answers its open questions, and records the aggregate prod readout. Co-Authored-By: Claude Opus 5.5 * fix(compaction): read the head-of-turn cache gap from before any save Every _save_compaction_state stamps updated_at. On a restore, _apply_compaction advanced the truncation anchor (a save) before the coordinator called apply_pending_compaction, which then read a ~0s gap and left a parked cut waiting on a cache that was already cold; it landed only at the paid hard-ceiling apply, after the session had also paid the anchor re-write. apply_document_offload had the same defect after any apply that itself saved. The restore slice now captures the previous turn's stamp before the anchor advance (_restore_turn_stamp); apply_pending_compaction consumes it every turn into _turn_start_stamp, and _document_offload_reason reads that first. Both stamps are cleared at turn end so an old stamp can never make a warm turn look cold. Co-Authored-By: Claude Opus 5.5 * test(compaction): flip the harness tripwire now the restore path applies The quality harness's restore pace pinned the missed free apply; with the turn-start gap fix it applies the parked cut on cache_expired. Spec §5 and the scoping doc record the fix. Co-Authored-By: Claude Opus 5.5 * fix: teardown polls long enough for one run; name the Agents page icon actions - kb_migration teardown passes its own 780 s poll window to the delete saga. Managed KBs on dev took ~10 minutes to delete, so at the shared 480 s every teardown needed two worker runs (~30 min). The shared default stays 480 s for the reconciler, which deletes several per run. - Agents page Chat/Edit/Share/Delete icon buttons (grid and list) get aria-labels naming the agent; the tooltip only set aria-describedby. - Kaizen queue: the production orphaned-agent cleanup chore, blocked on #1293 and #1301 reaching main. Co-Authored-By: Claude Opus 5.5 * feat(memory): canonical memory-file format and save validation (Shared Projects 2.3a) Pure library for the canonical memory-file format (spec §4.2) and the manifest-aware save checks (§4.3 steps 1-4), plus a bounded CountTokens counter in apis.shared for save-time token accounting. No callers yet; 2.3b wires these into the save path. - format.py: system-rendered frontmatter (YAML subset, no new dependency), items with ULID anchors, links, aliases, slug rules, reserved MEMORY.md. - validation.py: locked-field echo rule, name/alias collisions, anchor known/unique/minted, link resolution incl. archived targets; only new dead links fail. - tokens.py: one-attempt, 2 s CountTokens with a labelled chars/4 fallback; MEMORY_TOKEN_COUNT_MODEL_ID="" turns it off (tests do). - Hypothesis properties for round trips, anchor stability and link resolution. Co-Authored-By: Claude Opus 5.5 * feat(memory): save pipeline and FILEVER history (Shared Projects 2.3b) Wires the 2.3a format and validation into every Memory Space write path and adds per-file version history. - MemorySpaceService.save_entry: validate, count tokens once (CountTokens), write the object, conditional manifest swap, then the FILEVER row. The swap is the commit point, so version numbers cannot collide; replaced objects are kept because history references them. - fileFormat freeform|canonical on the space, set at create. Existing spaces stay freeform (stored as written; warnings only). Canonical spaces enforce items, anchors, links, collisions and the 8,000-token hard cap; a concurrent save of the same file is a 409. - Pre-history entries get a baseline version on first replacement. delete_entry purges its history (dedup-aware); delete_space pages its rows and purges every object; consolidate never collects history. - MEMORY.md is reserved in the service on every path. - API: PUT entries returns version/tokens/warnings/anchors; new GET /memory/spaces/{id}/history?slug= and /history/{n}?slug=. - memory_write records reason "save" and no longer sends an empty description (tool spec text unchanged, so toolConfig is unchanged). - app-api task role gains bedrock:CountTokens. Co-Authored-By: Claude Opus 5.5 * fix(observability): keep conversation content out of runtime OTEL logs Under AgentCore, ADOT writes every prompt and reply into the runtime log group's otel-rt-logs stream from two scopes: - strands.telemetry.tracer: Strands puts messages on span events and ADOT's LLOHandler re-emits them as log records. Setting OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_unredacted_attributes= turns on Strands' redaction with nothing exempt. - botocore bedrock-runtime: ADOT defaults OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT to true with setdefault, so the image sets it to false first. Both are image ENV rather than Runtime variables, because the Runtime's 50-variable cap is nearly spent. A Runtime variable still overrides them. The new test drives a Strands turn through the real tracer and LLOHandler, and the botocore event builders, against the Dockerfile's values. Each case has a capture-on control. Co-Authored-By: Claude Opus 5.5 * Wire PLATFORM_SELF_SERVICE_ENABLED flag into infra (dev-only, opt-in) (#1314) Adds the CDK plumbing so the platform self-service flag can be turned on per deployed environment (it was code-only before, readable from os.environ but never set on the runtime). - config.ts: PlatformSelfServiceConfig + opt-in parse (CDK_PLATFORM_SELF_SERVICE_ENABLED, only literal 'true' enables; mirrors projects). - inference-agentcore-construct.ts: sets PLATFORM_SELF_SERVICE_ENABLED on the AgentCore Runtime (read only there). 47->48 of 50 env slots; guard test reports 2 free. - platform.yml: forwards the CDK var from the GitHub environment. - mock-config.ts: default off. Default OFF everywhere; prod stays dark unless its own env var is set. Not wired into app-api (it never reads the flag). Verified: tsc clean, env-var-limit + config tests pass (198). * test(infra): find app-api's CountTokens grant across every IAM resource type Co-Authored-By: Claude Opus 5.5 * feat(compaction): harness arms and knobs for scoring summary compression - raw_summary arm: the production cut with the summary budget lifted, so comparing it with model_relative isolates what bound_summary's Nova Micro compression costs, separately from the cut. - records: --record-words sets the per-record word cap, and --workers runs transcripts in parallel. One record per turn at 600 words reproduces prod's ~20k-token LTM records at cut time, which is what makes the compression run. - ask: --workers parallelizes the questions on one history after its first call has written the cache. Co-Authored-By: Claude Opus 5.5 * fix(observability): widen unhealthy-host alarm to a sustained 20-min window (#1319) The alb-unhealthy-hosts alarm paged on any single App API target failing one health check for 10 min (evaluationPeriods:2). ECS self-heals such a target in ~5 min while the other targets keep serving, so this fired ~daily on benign, zero-impact blips. Require 4 consecutive 5-min periods (20 min) instead. Threshold stays 0 and treatMissingData stays BREACHING, so a stuck target and a total outage both still page. Window-not-threshold per observability.md section 11. * fix(ingestion): skip agent icon uploads instead of ingesting them as documents Agent icons live in the RAG documents bucket at assistants/{id}/icons/{digest}.{ext}. The legacy ingestion parser read that 4-part key as document_id="icons" and left a failed DOC#icons row on the agent. Both ingestion paths now parse strictly and skip (info log, success, no row written) anything that is not assistants/{id}/documents/{docId}/{file}. The managed-KB EventBridge rule is narrowed to assistants/*/documents/*. Adds backend/scripts/cleanup_stray_doc_rows.py (report-only by default) to remove the stray rows already written on live agents. Co-Authored-By: Claude Opus 5.5 * docs(memory): production read-path findings for the Phase 0 baseline The ~1,300 retrieval failure lines in the production runtime log are 745 events (each logged twice) and none is a throttle. 476 turns lost retrieval to dead pooled connections on cached agents after ~6 min idle, masked by a handler bug that discards every namespace's results; 84 more sent a query over the 10,000-character limit. Documents the evidence and a recommended fix (not implemented). Co-Authored-By: Claude Opus 5.5 * fix(kb): stop the reconciler resurrecting torn-down KB records The daily reconciler snapshots every KB_Record and then walks AWS. A teardown that finished in between left the run holding a record whose knowledge base was gone, and mark_vector_state_missing's unconditional UpdateItem recreated it as a ghost KB# item holding only vectorState/updatedAt. That's the orphaned row the teardown exists to remove. refresh_stored_bytes had the same upsert hazard, and both run in report-only mode too. - records.update_if_present: an update_item guarded on attribute_exists(PK) that returns False (logged at info) instead of creating a removed record. - reconciler: both record-side writes go through it and are only counted when they land. Records in migrationState=teardown are neither marked missing nor refreshed, and are listed in ReconcileReport.tearing_down (tearingDown in the serialized report). - Same guard on the other unconditional KB# writes that could outlive a teardown: set_resource_policy_state (both branches; REMOVE-only upserts too), byte_cap commit/release/reserve (reserve tells "record gone" apart from "over the cap" via ReturnValuesOnConditionCheckFailure), and the worker's _set_retain_until and _record_progress. Co-Authored-By: Claude Opus 5.5 * fix(a11y): name icon-only buttons on admin connector pages (#1318) appTooltip only sets aria-describedby, so the icon-only buttons on the connector list and form were announced as just "button" (axe button-name, critical). Add aria-labels following the Agents page pattern: row actions name their connector ("Delete Example Drive"), the secret toggle names its current action, and the callback-URL copy buttons say what they copy. Also names the row Edit link (same defect, axe link-name) and adds type="button" where missing. Specs assert every icon-only control has an accessible name. Co-authored-by: Claude Opus 5.5 * docs(compaction): the quality veto ran — the cut passes, the compression fails Spec §5's waiver is replaced by the first full harness result (Haiku 4.5, 12 transcripts x 9 planted facts, k=3, prod-sized LTM-style records). The model-relative cut with an uncompressed summary matches the full history; the same cut with bound_summary's Nova Micro compression (~15k -> ~730 tokens, reproducing prod's ~20k -> ~760) loses standing instructions (1.00 -> 0.72, p=0.002) and exact identifiers (0.96 -> 0.58, p=0.004). The tuning rule's FLOOR_RATIO lever is scoped to retained-turn loss. The scoping doc's §8 records method, spend (~$16) and the candidate levers. Co-Authored-By: Claude Opus 5.5 * docs(shared-projects): record 2.3 dev validation and carry 2.4 follow-ups Adds a "Dev-validated 2026-09-25" entry to the 2.3b as-built notes with the observed numbers (counted tokens, versions, rejections, purge), and a "Carried from 2.3" list under 2.4: canonical project spaces, the anchor-stripping decision for injected memory, the counter cold start, and the deferred items. Co-Authored-By: Claude Opus 5.5 * Fix: grant runtime write access to user-settings for set_default_model (#1325) The platform self-service set_default_model tool calls UserSettingsRepository.update_settings (PutItem + conditional UpdateItem), but the inference/AgentCore runtime IAM role granted only GetItem on the user-settings table. The write AccessDenied'd and the tool swallowed it into a generic error. - inference-api-iam-roles.ts: broaden the user-settings grant to GetItem + PutItem + UpdateItem (renamed sid to UserSettingsTableReadWriteAccess), scoped to the one table ARN, no DeleteItem, no wildcard. - security-policy.test.ts: update the regression guard to assert read+write (no delete, no wildcard). - account_tools.py: set_default_model now surfaces the real error type (e.g. AccessDenied) instead of a generic 'try again', so a hidden infra gap is diagnosable from the tool result. * feat(compaction): screen candidate summarizers in the quality harness Arm gains an optional summarizer with bound_summary's signature, patched into turn_based_session_manager for that arm only and restored after, so a candidate runs behind the production cut. summarizers.py adds two prototypes: - compress_only: production's prompt, temperature only. Production's compress_with_model also sends topP, which Claude 4.5+ rejects. - extract_then_compress: a verbatim pinned block of instructions, decisions, identifiers and changed values, then the narrative compressed into the rest of the budget. New arms: nova2lite_compress (production code, other model id), haiku_compress, extract_nova_micro, extract_nova2lite and extract_haiku. Co-Authored-By: Claude Opus 5.5 * docs(compaction): screening results — Nova 2 Lite, and extract-then-compress Scoping doc §9: three repetitions per candidate on the free availability check. Production's Nova Micro compression keeps 76.5% of planted facts (identifiers 58%). Nova 2 Lite on the unchanged production path keeps 98.8%, and extract-then-compress on Nova 2 Lite keeps 100% in every run, at ~1.9k summary tokens and about $0.02 per cut. Also records that compress_with_model cannot run a Claude summarizer (temperature plus topP is rejected, and it silently falls back to truncation). Co-Authored-By: Claude Opus 5.5 * fix(inference-api): keep MCP tool arguments and results out of trace spans ADOT 0.19 loads its own MCP instrumentor (entry point aws_mcp), which sets gen_ai.tool.call.arguments and gen_ai.tool.call.result on every tools/call span with no capture gate. Those spans go to aws/spans, so neither switch from #1317 reaches them. Set AWS_AGENTIC_INSTRUMENTATION=disabled as image ENV beside #1317's two switches, so the distro skips ADOT's native agentic instrumentors at startup. Only mcp applies to this app. Strands' execute_tool spans (already redacted) and its own _meta trace-context injection are unaffected. Extend test_otel_content_capture.py to load aws_mcp through AwsOpenTelemetryDistro.load_instrumentor with the Dockerfile's value and drive a real in-memory MCP tools/call: no sentinel and no MCP-scope span. The control keeps #1317's switches on and drops only the new one, and the sentinels appear on the MCP spans' arguments and result. Co-Authored-By: Claude Opus 5.5 * feat(compaction): default the summary model to Nova 2 Lite Nova Micro compressed LTM summary records to ~760 tokens and dropped about a quarter of standing instructions and 42% of exact identifiers on the quality harness. Nova 2 Lite runs on the unchanged bound_summary (Nova accepts topP) and keeps ~99% of planted facts for ~$0.01 per cut. Co-Authored-By: Claude Opus 5.5 * chore(harness): add a pinned Nova Micro compression arm model_relative now follows the new Nova 2 Lite default, so the old baseline needs its own arm to stay comparable. Co-Authored-By: Claude Opus 5.5 * fix(kb): stop late writes recreating removed KBTOMB# and DOC# rows UpdateItem is an upsert, so a write aimed at a removed row recreates it as a ghost holding only the attributes it set. #1322 guarded KB# records; this extends the guard to the two row types it left out. - records.update_if_exists(key, ...): key-taking form of update_if_present, which now delegates to it (signature unchanged). - tombstones.record_tombstone_error: guarded. A saga failing after a concurrent saga cleared the tombstone no longer leaves a ghost with no intent/AWS id and no TTL in iter_tombstones. _write_tombstone stays an intentional upsert. - byte_cap.settle_once: ANDs RECORD_EXISTS in and returns False for a removed DOC# row (ALL_OLD distinguishes it in the log). Previously the claim succeeded on a missing row, recreated it, and let a late ingestion commit bytes the document delete had already released. - ingestion_consumer.set_document_terminal: guarded, returns bool. A removed row is a logged skip: no retry, no raise, no DLQ. - provisioner._ingest_waiting skips ingesting a row deleted mid-provisioning; document_reconciler._plan does not count a skipped write as performed. Co-Authored-By: Claude Opus 5.5 * fix: send temperature without topP on side-channel converse calls Claude 4.5+ rejects `temperature` and `topP` in the same inferenceConfig with a ValidationException. compress_with_model swallows every exception and returns None, so a Claude AGENTCORE_MEMORY_COMPACTION_SUMMARY_MODEL_ID silently degraded every compaction summary to newest-first truncation of the raw records (measured in the compaction quality harness). Drop topP from all four side channels that hardcoded the pair: the compaction summary, tool-batch summaries, document abstracts and session titles. The latter three run on Nova today, but each has the same fail-quiet shape, and the document digest's model id is env-configurable. Tests stub Bedrock with Claude's rejection rule so the old config fails them and the fixed one passes. Co-Authored-By: Claude Opus 5.5 * feat(memory): project-scoped memory spaces (Shared Projects 2.4a) A project now owns a canonical shared Memory Space, and each member can have a personal_in_project space. Both take every role from the project: resolve_permission reads project membership for project scopes, no MEMBER# rows are written, and nobody resolves owner on a project space. - META gains scope/projectId/userId; project scopes stay out of OwnerIndex, so they never show in personal space lists or the binding picker. - create_project creates the shared space (rolled back with the harness); GET /projects/{id}/memory backfills older projects through a conditional, version-bumping pointer write. - POST /projects/{id}/memory/mine creates the caller's own space on first use (PERSONAL_SPACE# pointer via an UpdateItem upsert). - Rename and purge follow the project; direct share/leave/delete are refused; agent bindings to a project space are refused at design time and dropped by the runtime resolver. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): queue the prod check for the reconciler teardown guard Co-Authored-By: Claude Opus 5.5 * fix(agents): delete an agent's icon objects when the agent is deleted Agent delete (DELETE /agents/{id}, DELETE /assistants/{id}) and the project-harness purge never deleted assistants/{id}/icons/*, so every deleted agent with an icon left its objects in S3 with no row left to find them by. - IconStore.delete_prefix + AgentIconStore.delete_all + delete_agent_icons (best effort, never raises), called after the record delete on both paths, and on a retried harness purge. - cleanup_orphaned_agent_rows.py --s3-prefixes: finds assistants/{id}/ prefixes whose agent has no row at all; report-only unless --apply --confirm-prefix, keeps any key a row still names, re-checks the partition before deleting. Co-Authored-By: Claude Opus 5.5 * fix(infra): sweep retention onto every AgentCore Runtime log group generation The RuntimeLogRetention custom resource sets retention on the live runtime's -DEFAULT group once per deploy. A replaced Runtime gets a new id and a new log group, and the old group drops out of CDK's view with whatever retention it had. Groups orphaned before the custom resource existed never expire, and they hold conversation text. In an account whose landing zone stamps retention on every CreateLogGroup event, the deploy-time value is also overwritten minutes later, even on the live group. Add a daily Lambda that lowers every group under /aws/bedrock-agentcore/runtimes/- to observability.logRetentionDays when its retention is unset or longer. Runtime names cannot contain '-', so the prefix matches every generation of this deployment's Runtime and no other deployment's. PutRetentionPolicy is IAM-scoped to that prefix; the function cannot delete groups or read events. On by default with a kill switch, CDK_OBSERVABILITY_RUNTIME_LOG_RETENTION_SWEEP_ENABLED=false. Co-Authored-By: Claude Opus 5.5 * docs(memory): production record census for the Phase 0 baseline Adds the operator-run audit.py inventory aggregates to the decision record: prod writes and extracts like dev at ~35x the scale (2,770 actors, 85,382 records, 0 session-id fallback actors). Notes 112 unreachable records under 13 legacy-format actor ids and 7 extraction jobs failed with LTM_RATE_EXCEEDED; both are operator decisions. Co-Authored-By: Claude Opus 5.5 * docs(compaction): record the Nova 2 Lite summary-model confirmation run Co-Authored-By: Claude Opus 5.5 * chore(infra): expire noncurrent versions in the RAG documents bucket The documents bucket is versioned with no lifecycle rules, so every delete the app makes (document cleanup, icon replace/remove, agent-delete icon cleanup, the orphaned-row cleanup script) only adds a delete marker and the bytes stay forever as noncurrent versions. Add one rule: noncurrent versions expire after 35 days (matching the assistants table's PITR window, so a PITR-restorable DOC# row still has its source bytes), orphaned delete markers are removed, and incomplete multipart uploads abort after 7 days. Nothing reads by VersionId, and both ingestion paths address current objects by key. Co-Authored-By: Claude Opus 5.5 * fix(compaction): attach post-turn cut events to the call that triggered them A cut is decided by update_after_turn after the turn's last model call and waited in the session manager's in-memory queue for the next call on the same instance. Prod runtime logs show 32 of 81 cuts never got one: a new microVM or agent-cache miss (fresh, empty queue) or an abandoned session. The coordinator now hands them to ContextLedgerHook before the turn's C# rows are written. The fleet readout's distance-from-cut now keys on the 'applied' event, the first call sent on the sliced history. Co-Authored-By: Claude Opus 5.5 * fix(costs): build the prefix split from native counts only Claude Sonnet 5 has no CountTokens, so every count is Strands' heuristic (JSON at chars/2) and the tools residual overshot the whole prompt. The attribution hook now requires token_count_is_authoritative (which reads the SDK skip list) and compares heuristic_count_fallbacks around every count it relies on, deferring a tainted split to the next turn. Co-Authored-By: Claude Opus 5.5 * feat(observability): carry sessionId on compaction EMF records Every AgentCoreStack/Compaction record the session manager emits (cut, apply, truncation anchor, document offload) now has the conversation id as a log property, the way the prompt-cache records do. It is not a dimension, so no per-session metric streams. The 2026-09-25 readout had to join cuts to sessions by matching input tokens within 30 minutes. Co-Authored-By: Claude Opus 5.5 * docs(compaction): record root causes of the 2026-09-25 telemetry gaps Co-Authored-By: Claude Opus 5.5 * fix(memory): reconnect long-term memory retrieval on a dead pooled connection Production lost long-term memory retrieval on ~10% of turns (Phase 0 audit, docs/specs/memory-baseline-decision.md "Production (read-only)"): - A cached agent keeps its retrieval client between turns. After ~6 min idle its pooled connection is dropped on the network path, and with total_max_attempts=1 (#1157) the next request failed within ms. Retry once, immediately, on a fresh client for ConnectionClosedError and SSLError. Throttles are still never retried. - ConnectionClosedError/ReadTimeoutError carry response=None; the per-namespace handler raised on it and the outer handler discarded every namespace's results. Guard it and log the exception class. - Cut searchQuery to the API's 10,000-character limit (counted in UTF-16 code units) instead of failing with ValidationException. audit.py counts the new "memory retrieval reconnected" log line. Co-Authored-By: Claude Opus 5.5 * docs(compaction): fix stale topP note on the summary model default #1330 dropped topP from compress_with_model, so a Claude override no longer fails; the comment merged in #1334 predates that. Co-Authored-By: Claude Opus 5.5 * feat(compaction): pin verbatim facts ahead of the compressed summary bound_summary(extract_enabled=True) makes one extraction call (standing instructions, decisions, identifiers, changed values at their latest value) into a pinned block capped at half the budget, then compresses the narrative into the rest with temperature only. Extraction failure falls back to plain compression, narrative failure to truncation; it never raises. The result is persisted verbatim, so restore bytes stay stable. Gated by COMPACTION_SUMMARY_EXTRACT_ENABLED, in development and default off. Co-Authored-By: Claude Opus 5.5 * feat(infra): wire CDK_COMPACTION_SUMMARY_EXTRACT_ENABLED to the runtime Opt-in while in development: only "true" enables. Runtime only, since compaction runs in the AgentCore Runtime. Takes the runtime to 49 of its 50 environment variables in the worst case the ceiling test builds. Co-Authored-By: Claude Opus 5.5 * test(compaction-harness): screen extract-then-compress from production code The extract_* arms set summary_extract_enabled on the config instead of swapping in the summarizers.py prototype, which is deleted. compress_only now calls production's compress_with_model with top_p=None. Co-Authored-By: Claude Opus 5.5 * docs(compaction): extract-then-compress clears the quality veto Paid harness run: extract + Nova 2 Lite matches the full history in every family (0 losses), while today's Nova Micro compression loses constraints, decisions and identifiers. Records the cut-turn latency (11-23 s, after the final metadata event, so no TTFT change). Co-Authored-By: Claude Opus 5.5 * fix(scripts): forward the runtime log retention sweep kill switch as CDK context #1332 added observability.runtimeLogRetentionSweepEnabled but skipped build_cdk_context_params(), the fourth of the five places a new observability tunable must touch. The env var still took effect in CI because config.ts reads it first; a local deploy through load-env.sh did not pass it on. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): queue the runtime log retention follow-ups Records the dev validation of #1332 and queues the four follow-ups: the production one-off, verifying the sweep in production after release, the false 'PutRetentionPolicy creates the group' claim, and the development account's 10-year retention default. Co-Authored-By: Claude Opus 5.5 * fix(kb): don't ingest or revive a document deleted before its event lands A late or redelivered ingestion event can outlive the document it is for. #1327 stopped it recreating a removed DOC# row; two gaps remained: - handle_object still submitted a deleted document to Bedrock, putting its content back into the knowledge base after cleanup removed it. It now returns early when the row is gone or `deleting`. Every producer writes the row before the S3 object, and the read is now strongly consistent. - set_document_terminal could flip a soft-deleted (`deleting`) row back to complete/failed/uploading. The retrieval filter joins on that row, so a deleted document whose cleanup failed became retrievable again for its whole TTL. The write is now also conditioned on status <> deleting. The document reconciler's NOT_FOUND re-ingest re-reads the row before submitting, closing the same window between its scan and the ingest. Co-Authored-By: Claude Opus 5.5 * feat(memory): scope-labelled project blocks, anchor stripping and a lean project index read render_memory_block varies its intro by scope and may omit the space name; render_project_memory builds the project (2,000-token) and mine (1,000-token) blocks with anchors stripped, and renders nothing for an empty or starter index. MemorySpaceService.read_project_space_index reads a project space's MEMORY.md for a caller whose project role is already settled. MemoryEntryNotFoundError tells a missing entry from an unreachable space. An Agent's block bytes are unchanged (hash-pinned). Co-Authored-By: Claude Opus 5.5 * feat(agents): scope-addressed memory tools for the project harness memory_list/read/query/save with scope in {project, mine}. A save to mine creates the member's personal space on first use; a viewer saving to project is pointed at mine; an archived project says it is archived. Ordinary Agents' memory_list/read/write specs stay byte-identical (SHA-256 pinned). memory_query and memory_save count as Memory tools in the context breakdown. Co-Authored-By: Claude Opus 5.5 * feat(inference-api): wire project memory into harness turns The harness turn reads its project's shared and personal MEMORY.md concurrently, started as soon as the project is known and awaited at prompt assembly, and sends them as memory_context behind the fourth cache point. The multi-scope memory_binding digest keys harness agents; the 2.1 binding digest is unchanged, and PausedTurnSnapshot needs no new field. Tool construction is memoized per member and spaces. Co-Authored-By: Claude Opus 5.5 * docs(projects): record 2.4a dev validation and 2.4b as built Co-Authored-By: Claude Opus 5.5 * perf(compaction): run extraction and narrative compression concurrently Split the summary budget up front (pinned block at most half, narrative the rest) so the two calls run side by side. A cut now costs the slower call rather than the sum: on prod-sized records the median fell from 11.6 s to 8.4 s, against 7.6 s for the single-call default. An extraction failure keeps the narrative as a plain compression instead of making a third call, so a cut never makes more than two. Co-Authored-By: Claude Opus 5.5 * fix(compaction): keep each identifier's label in the extraction The harness caught the extractor copying a value verbatim under a new name ("Budget reference" for the budget workbook version tag), which the answering model could not tie back to the question. Ask for each identifier with the conversation's own name for it. Co-Authored-By: Claude Opus 5.5 * fix(infra): tolerate a missing runtime log group in RuntimeLogRetention PutRetentionPolicy does not create the log group: on a missing group it returns ResourceNotFoundException and creates nothing. Deploys succeed only because the Runtime's execution role creates the group first; if that ever loses the race, the custom resource fails and the stack update rolls back. Both calls now set ignoreErrorCodesMatching ResourceNotFoundException; the daily retention sweep sets retention on a group missed this way. Drop the unused logs:CreateLogGroup grant, correct the comment and observability.md section 9, and add jest assertions for both. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): record #1345 against the runtime log retention race Co-Authored-By: Claude Opus 5.5 * docs(compaction): record the concurrent extract-then-compress rescore Concurrent calls bring a cut to the default's latency (median 8.4 s vs 7.6 s, 11.6 s sequential) and keep the no-loss result against the full history. Records the identifier-labelling defect the paid run caught, that the prompt fix is unproven, and the narrative's output-cap discard. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): queue the RAG documents bucket lifecycle follow-up Dev-validated #1336 (rule live, matches the template). Queue the checks that can't happen on deploy day: confirm the dev drain after S3's asynchronous lifecycle pass, and confirm plus inventory production after the release. Co-Authored-By: Claude Opus 5.5 * fix(kb): refund managed-KB bytes on delete and settle migrated corpora Deleting a document that reached `complete` never returned its bytes: the delete path only released reservations under settle_once, and a completed document was already settled. The ingestion consumer now stamps `committedBytes` on the DOC# row before the KB# commit (refused on a gone or `deleting` row, in which case it releases instead), and the delete claims `byteCapRefunded` once and refunds exactly that amount from storedBytes and totalBytes. Marker writes are guarded on the row existing, so nothing recreates a ghost DOC# row. The reconciler's refresh re-anchored storedBytes but never totalBytes, the accumulator the cap guard reads. It now sets totalBytes = stored + reservedBytes in the same write, repairs a drifted total, and skips records the migration worker is part-way through. Migration reserved the whole corpus in shadow and nothing ever committed it, so every migrated knowledge base carried its corpus in reservedBytes with storedBytes=0. reserve_snapshot now records snapshotReservedBytes and run_promote adopts each complete document (settled + committedBytes, counted into storedBytes) and returns the snapshot reservation once. Also: deleting an unsettled `failed` row no longer releases bytes it never reserved, and release/refund refuse to drive a counter below zero (KbByteCapAccountingSkew). scripts/repair_managed_kb_byte_counters.py repairs existing records (report-only by default). Co-Authored-By: Claude Opus 5.5 * docs: note compaction extract-then-compress is on in development Co-Authored-By: Claude Opus 5.5 * fix(kb): let the reconciler read the documents bucket; anchor storedBytes to the ledger The managed-KB reconciler's storedBytes refresh has never run: its Lambda had no read grant on the documents bucket, so every ListObjectsV2 was AccessDenied. Grant it read-only access under `assistants/*`. Before turning it on, change what the anchor counts. It used to total every object under `assistants/{id}/documents/`, which would charge owners for failed rows, stuck legacy rows and orphaned uploads, and count in-flight uploads twice (also in reservedBytes). It now sums S3 sizes only for DOC# rows the ledger counts: `committedBytes` stamped and `byteCapRefunded` not claimed. That follows the commit-before-`complete` and refund-after-`deleting` ordering, which a status filter would not. Co-Authored-By: Claude Opus 5.5 * fix(compaction): salvage a summary generation that hits its token ceiling compress_with_model discarded any generation that stopped on max_tokens, so the cut fell back to newest-first truncation of the raw records, which drops the oldest standing instructions first. Nova 2 Lite varies its narrative several times over on the same records, so the ceiling is hit in normal use. Keep the complete lines instead (the last, possibly partial, line is dropped) and hold them to the budget from the end, since the prompt puts standing instructions first. The plain path labels such a cut model_salvaged so it can be counted apart from a clean compression. maxTokens now follows the model's ceiling: Nova 2 Lite's card gives 64K output, so the budget binds and matches the prompt's word limit. Unlisted models keep 4k, under Nova Micro's 5K, so an override is never rejected. Co-Authored-By: Claude Opus 5.5 * docs: note compaction extract-then-compress is on in production Co-Authored-By: Claude Opus 5.5 * feat(memory): calibrate the relevance cut on a labelled eval set; default 0.4 audit.py calibrate: one synthetic memory-audit-probe-* actor states 16 synthetic facts and preferences (calibration_set.json) across 5 conversations, waits for extraction to settle, asks each fact three ways (direct, indirect, wrapped in filler) plus 20 unrelated questions, raw and with filler stripped, labels every returned record from its text, and writes score distributions and per-policy recall/precision to summary.json. Record text stays in raw/. Cleans up in finally, with late sweeps. --reanalyze recomputes from raw/ without AWS calls. Two dev runs: realistic questions score 0.34-0.56 and noise sits at 0.34-0.40, so the 0.5 cut kept the right record for 7% of questions. 0.4 keeps it for 59% at 94% precision and injects on 5% of unrelated turns (~21 tokens/turn on average). The default moves to 0.4; findings are in docs/specs/memory-baseline-decision.md. Co-Authored-By: Claude Opus 5.5 * feat(memory): log per-namespace retrieval scores, no ids or text retrieve_customer_context now logs one line per namespace per turn: top score, records returned, records kept, and the cut in force. The namespace is the template ({actorId} unresolved) and no record text is logged. Computed from the response already in hand: no extra calls, and the added cost before the first token is one log line. audit.py inventory picks the lines up (LOG_PATTERNS retrieval_scores) and histograms the top score per strategy type, counting the plain runtime stream only (the OTEL copy duplicates every line). Co-Authored-By: Claude Opus 5.5 * docs(memory): link the score-logging PR Co-Authored-By: Claude Opus 5.5 * docs(kaizen): track harness-sdk#4618 as a Strands pin-bump hazard Upstream #4618 proposes normalizing every provider's Usage to the semconv subset convention, folding Bedrock cache tokens into inputTokens. Our cost, quota and context math assume the disjoint Converse shape, so a bump that ships it would double-price every cached Bedrock token silently. - review-queue: new watch entry with the bump gate; amend the 2026-09-05 caching watch (#4193 merged 2026-09-25, #3546 closing via #4617) - kaizen-research SKILL §2a: #4618 warning + per-run check Co-Authored-By: Claude Opus 5.5 * docs(projects): record 2.4b dev validation Project memory blocks, scope tools, the personal space created on first save, the viewer refusal and the archived gate all behaved as built on dev. The load's waitedMs was 0.0-1.1 ms, so it stays off the first-token path. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): record the posted #4618 comment in the watch entry Co-Authored-By: Claude Opus 5.5 * fix(count-tokens): strip global., au. and jp. inference-profile prefixes base_foundation_model_id only knew us./eu./apac./us-gov., so a global.* profile id reached CountTokens unchanged. Bedrock rejected it as unsupported, the model landed on Strands' skip list, and every count for every prod global.* model fell back to the chars/4 heuristic, which also means no prefixTokens / contextBreakdown after #1337. Adds global plus the geo prefixes the AWS model cards document today (au., jp.). Undocumented codes stay untouched. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): queue the production managed-KB byte-counter repair for after release Co-Authored-By: Claude Opus 5.5 * perf(context-attribution): take native token counts off the model-call path Stripping global. from the CountTokens id makes prod global.* models count natively, and a native count awaited before a model call adds a CountTokens round trip to time to first token: measured +519 ms on a freshly built agent's first call and +74 ms on each later call (+56 ms per tool round, for a count that always failed). Nothing needs those counts before the call: - CountTokensBedrockModel(native_projection=False), as the factory now builds every Converse model, keeps Strands' pre-call count_tokens on the local heuristic. Strands' proactive compression is off in our config and our compaction does not read the projection. - native_count_tokens() is a blocking counter that answers natively or None, never a heuristic, with the same skip-list bookkeeping. - ContextAttributionHook's BeforeModelCall callback is now synchronous: it records the projection and, while the agent has no split, starts one background task that makes four native counts of a snapshot concurrently with the model call. AfterModelCall places the messages partition against the call's billed prompt (input + cache read + cache write). - Interrupted turns read get_projected_input_tokens(): the native count of the exact request when the split measured one, else a usage-anchored projection, else nothing. Same dev harness (40 tools, 32k-token prefix, n=8 per arm): the pre-stream window matches today's heuristic path, and the split was ready before call 1 finished in 8/8 sessions. Co-Authored-By: Claude Opus 5.5 * feat(compaction): turn extract-then-compress on by default COMPACTION_SUMMARY_EXTRACT_ENABLED now defaults on with =false as the kill switch. The quality harness showed no loss against the full history and a forced cut on dev pinned and answered every planted fact. Harness model_relative follows the production default; the plain-compression arms pin extraction off as baselines. Co-Authored-By: Claude Opus 5.5 * chore(infra): drop the CDK wiring for the extract-then-compress flag A default-on flag needs no AgentCore Runtime env var, so the config, construct entry, workflow variable and their tests go. The Runtime is back to 48 of its 50 variables. Co-Authored-By: Claude Opus 5.5 * docs(compaction): record the dev validation and the default-on flip Co-Authored-By: Claude Opus 5.5 * test(compaction): pin the salvage test to the plain compression path It asserts the single-call path's outcome and bytes, which the default-on extract flag would otherwise route through extraction. Co-Authored-By: Claude Opus 5.5 * docs(kaizen): queue the production check for #1343's native token counts Dev can't exercise a global.* id (the SCP denies it), so the last open question on #1343 is production: prefixTokens and contextBreakdown on global.* Haiku 4.5 sessions, and no CountTokens span in front of a model call. Read-only steps, blocked on the release. Co-Authored-By: Claude Opus 5.5 * fix(agents): run every agent-stream step in one context `_merge_agent_status` drives the agent stream by scheduling each `__anext__` as its own task, and a task runs in a copy of the context it was created from. Strands holds its spans attached across yields (the invocation, each cycle, each model stream), so a span attached in one step was detached in a later step's copy. OpenTelemetry logged that as `Failed to detach context` — 3 per tool-less turn, 5 per one-tool turn, ~2.3k in dev over a week, starting with the first deploy of the merge. The same split mis-parented spans: anything opened from the ambient context after a yield — the AgentCore Memory CreateEvent calls the session manager makes, DynamoDB reads from hooks — parented to the request span instead of the cycle that caused it. Strands' own spans were unaffected (explicit parent, explicit end), so nothing was lost. Every step now runs in one Context copied once per turn, which is what a plain `async for` gave the generator. No added latency: one copy per turn instead of one per step. Co-Authored-By: Claude Opus 5.5 * fix(spa): forget the new-conversation draft on the first send The first send from the empty state swaps its composer for the compact one, which destroys it before the draft effect runs again. The new-conversation draft kept the message just sent, so every later New Session opened with that message already in the composer. submitChatRequest now writes the cleared draft at once instead of relying on the effect. Co-Authored-By: Claude Opus 5.5 * Release/1.25.1 Patch release carrying one SPA fix: New Session no longer reopens with the previous conversation's first message (#1362). - Notes lead with a pointer to v1.25.0 so the feature release and its deployment steps stay visible to anyone upgrading from 1.24.x - SPA only: no CDK change, no scripts Co-Authored-By: Claude Opus 5.5 --------- Co-authored-by: Claude Opus 5.5 Co-authored-by: Colin Co-authored-by: Colin Smith <7762103+colinmxs@users.noreply.github.com> Co-authored-by: Derrick Fink --- CHANGELOG.md | 8 ++++++ README.md | 4 +-- RELEASE_NOTES.md | 28 +++++++++++++++++++ VERSION | 2 +- backend/pyproject.toml | 2 +- backend/uv.lock | 2 +- frontend/ai.client/package-lock.json | 4 +-- frontend/ai.client/package.json | 2 +- .../chat-input/chat-input.component.spec.ts | 14 ++++++++++ .../chat-input/chat-input.component.ts | 21 ++++++++++++-- infrastructure/package-lock.json | 4 +-- infrastructure/package.json | 2 +- tui/pyproject.toml | 2 +- tui/src/agentcore_tui/__init__.py | 2 +- tui/uv.lock | 2 +- 15 files changed, 83 insertions(+), 16 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c895bed3a..2412bf560 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,14 @@ All notable changes to this project are documented in this file. Format follows For narrative release notes written for operators and product owners, see [RELEASE_NOTES.md](RELEASE_NOTES.md). +## [1.25.1] - 2026-09-26 + +A single SPA fix on top of 1.25.0, which carries this week's features. **Upgrading from 1.24.x? Follow the 1.25.0 deployment notes** (CDK deploy plus post-deploy scripts). 1.25.1 adds no steps of its own. + +### 🐛 Fixed + +- New Session reopened with the previous conversation's first message, because the empty-state composer was destroyed before its draft effect could clear the saved `composer-draft:new` (#1362) + ## [1.25.0] - 2026-09-25 The assistant remembers more, and it is cheaper to see why. **Long-term memory reaches the model for the first time**: the relevance cut drops from 0.7 to 0.4, retrieval survives dead pooled connections, and it logs its scores. **Compaction keeps the facts that matter**: extract-then-compress is on by default, the summary model moves to Nova 2 Lite, and a truncated summary is salvaged rather than discarded. Users get **personal instructions**, **dictation**, a **redesigned compact composer**, a **context meter** that itemizes the window, and a paged sidebar. Admins get **model retirement with redirect to a successor** and drag-to-order for the model picker. **Shared Projects** lands as an in-development preview that is **off by default** (`CDK_PROJECTS_ENABLED=true` to opt in). The release also closes three channels that exported conversation text to logs and traces, stops agent deletes leaking documents and knowledge bases, and repairs managed-KB byte accounting. **A CDK deploy is required, and three post-deploy scripts should be run** (see the release notes). diff --git a/README.md b/README.md index cd29a6c3d..cc80e1ca2 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ **An open-source, production-ready Generative AI platform for institutions** *Built by Boise State University, designed for everyone.* -[![Release](https://img.shields.io/badge/Release-v1.25.0-6366f1?style=flat&logo=github&logoColor=white)](RELEASE_NOTES.md) +[![Release](https://img.shields.io/badge/Release-v1.25.1-6366f1?style=flat&logo=github&logoColor=white)](RELEASE_NOTES.md) [![Nightly](https://github.com/Boise-State-Development/agentcore-public-stack/actions/workflows/nightly.yml/badge.svg)](https://github.com/Boise-State-Development/agentcore-public-stack/actions/workflows/nightly.yml) ![Python](https://img.shields.io/badge/Python-3.13+-3776AB?style=flat&logo=python&logoColor=white) @@ -296,7 +296,7 @@ agentcore-public-stack/ See [RELEASE_NOTES.md](RELEASE_NOTES.md) for the full changelog, including new features, bug fixes, platform upgrades, and deployment notes for each release. -**Current release:** v1.25.0 +**Current release:** v1.25.1 --- diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index f2294d8b7..e37d65d47 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -1,3 +1,31 @@ +# Release Notes — v1.25.1 + +**Release Date:** September 26, 2026 +**Previous Release:** v1.25.0 (September 25, 2026) + +--- + +> 📣 **This is a one-fix patch on top of [v1.25.0](https://github.com/Boise-State-Development/agentcore-public-stack/releases/tag/v1.25.0), which is where this week's features are.** 1.25.0 made long-term memory reach the model and made compaction keep the facts that matter. It also added personal instructions, dictation, the compact composer, the context meter and model retirement. +> +> **Upgrading from 1.24.x? Follow the v1.25.0 Deployment notes.** 1.25.1 does not change them. That upgrade still needs the CDK deploy, the required managed-KB byte-counter repair and the recommended cleanups listed there. + +--- + +## Highlights + +A new conversation no longer opens with a message that was already sent. The first message sent from the empty-state composer stayed in the saved new-conversation draft, so every later **New Session** opened with it already typed in. This is an SPA-only fix, with no backend or infrastructure changes. + +## 🐛 Bug fixes + +- **New Session reopened with the previous first message.** Composer drafts are saved to `localStorage` by an effect in `ChatInputComponent`, which removes the draft once a send empties the input. On the first send, the view swaps to the compact composer and destroys the empty-state composer before that effect runs again. The effect therefore never removed the draft, and `composer-draft:new` kept the sent text. `submitChatRequest` now writes the cleared draft immediately (`persistDraftNow()`), and an empty draft is removed. A new spec covers the destroy-without-change-detection path that the existing spec missed (#1362). + +## 🚀 Deployment notes + +- **From 1.25.0:** no special steps. Only the SPA changes; there is no CDK change and no script to run. +- **From 1.24.x or earlier:** follow the v1.25.0 Deployment notes first. + +--- + # Release Notes — v1.25.0 **Release Date:** September 25, 2026 diff --git a/VERSION b/VERSION index ad2191947..d905a6d1d 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -1.25.0 +1.25.1 diff --git a/backend/pyproject.toml b/backend/pyproject.toml index f82465b69..7956a0b39 100644 --- a/backend/pyproject.toml +++ b/backend/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "agentcore-stack" -version = "1.25.0" +version = "1.25.1" requires-python = ">=3.10" description = "Multi-agent conversational AI system with AWS Bedrock AgentCore" readme = "README.md" diff --git a/backend/uv.lock b/backend/uv.lock index 6696a9d66..967c3d633 100644 --- a/backend/uv.lock +++ b/backend/uv.lock @@ -15,7 +15,7 @@ constraints = [{ name = "mcp", specifier = "<2" }] [[package]] name = "agentcore-stack" -version = "1.25.0" +version = "1.25.1" source = { editable = "." } dependencies = [ { name = "aiofiles" }, diff --git a/frontend/ai.client/package-lock.json b/frontend/ai.client/package-lock.json index 7f34c2fbe..0d2c4e0f5 100644 --- a/frontend/ai.client/package-lock.json +++ b/frontend/ai.client/package-lock.json @@ -1,12 +1,12 @@ { "name": "ai.client", - "version": "1.25.0", + "version": "1.25.1", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "ai.client", - "version": "1.25.0", + "version": "1.25.1", "dependencies": { "@angular/cdk": "21.2.14", "@angular/common": "21.2.20", diff --git a/frontend/ai.client/package.json b/frontend/ai.client/package.json index af0b8357a..52f2ec08f 100644 --- a/frontend/ai.client/package.json +++ b/frontend/ai.client/package.json @@ -1,6 +1,6 @@ { "name": "ai.client", - "version": "1.25.0", + "version": "1.25.1", "scripts": { "ng": "ng", "prestart": "tsx scripts/branding/generate-brand-theme.ts && tsx scripts/branding/generate-surface-theme.ts && tsx scripts/branding/generate-surface-colors.ts && tsx scripts/branding/generate-favicons.ts", diff --git a/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.spec.ts b/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.spec.ts index b40840fcd..d013b830b 100644 --- a/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.spec.ts +++ b/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.spec.ts @@ -1549,6 +1549,20 @@ describe('ChatInputComponent — unsent text survives leaving the conversation', expect(component.userInput()).toBe('question for s1'); }); + it('forgets the new-conversation draft when the first send tears the composer down', async () => { + await mount(NEW_CONVERSATION_DRAFT_KEY); + type('first message'); + + // The empty state's composer is replaced by the compact one on the first + // send, so it is destroyed before another change detection pass can run + // its draft effect. + component.submitChatRequest(); + fixture.destroy(); + + await mount(NEW_CONVERSATION_DRAFT_KEY); + expect(component.userInput()).toBe(''); + }); + it('forgets the draft once the message is sent', async () => { await mount('s1'); type('about to send'); diff --git a/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.ts b/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.ts index 3effe78d3..26b03f620 100644 --- a/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.ts +++ b/frontend/ai.client/src/app/session/components/chat-input/chat-input.component.ts @@ -1036,8 +1036,9 @@ export class ChatInputComponent { // // Every path that empties the composer — send, queue-as-follow-up, the // user deleting it — flows through `userInput` and so forgets the draft - // without naming it, which is why none of those call sites mention - // persistence at all. + // without naming it. The one exception is send: it also writes at once + // (`persistDraftNow`), because the first send can destroy this instance + // before the effect runs again. effect(() => { const key = this.draftKey(); const draft = this.composerDraftSnapshot(); @@ -1273,6 +1274,21 @@ export class ChatInputComponent { if (key !== null) this.draftStorage.write(key, this.composerDraftSnapshot()); } + /** + * Write the composer's state under its conversation now, instead of on the + * draft effect's next run. + * + * The first send from the empty state swaps this composer for the compact + * one, and the swap destroys this instance before its effect runs again. The + * mirrored draft would then still hold the message just sent, and every later + * New Session would open with it in the composer. + */ + private persistDraftNow(): void { + if (this.mirroredDraftKey) { + this.draftStorage.write(this.mirroredDraftKey, this.composerDraftSnapshot()); + } + } + /** * Replace the composer's contents with a conversation's remembered draft * (or empty it, when that conversation has none). @@ -1519,6 +1535,7 @@ export class ChatInputComponent { this.closeSkillMenu(); this.resetTextareaHeight(); this.clearAttachments(); + this.persistDraftNow(); } cancelChatRequest() { diff --git a/infrastructure/package-lock.json b/infrastructure/package-lock.json index b39e0ad64..038de6bea 100644 --- a/infrastructure/package-lock.json +++ b/infrastructure/package-lock.json @@ -1,12 +1,12 @@ { "name": "infrastructure", - "version": "1.25.0", + "version": "1.25.1", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "infrastructure", - "version": "1.25.0", + "version": "1.25.1", "dependencies": { "aws-cdk-lib": "2.265.0", "constructs": "10.6.0" diff --git a/infrastructure/package.json b/infrastructure/package.json index e2e3910fc..19310c420 100644 --- a/infrastructure/package.json +++ b/infrastructure/package.json @@ -1,6 +1,6 @@ { "name": "infrastructure", - "version": "1.25.0", + "version": "1.25.1", "bin": { "infrastructure": "bin/infrastructure.js" }, diff --git a/tui/pyproject.toml b/tui/pyproject.toml index bdfa99e71..4b679bf73 100644 --- a/tui/pyproject.toml +++ b/tui/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "agentcore-tui" -version = "1.25.0" +version = "1.25.1" requires-python = ">=3.11" description = "Terminal client for the AgentCore Public Stack — streaming AI chat in your terminal" readme = "README.md" diff --git a/tui/src/agentcore_tui/__init__.py b/tui/src/agentcore_tui/__init__.py index da5bdd5d6..367e3d40e 100644 --- a/tui/src/agentcore_tui/__init__.py +++ b/tui/src/agentcore_tui/__init__.py @@ -7,6 +7,6 @@ from __future__ import annotations -__version__ = "1.25.0" +__version__ = "1.25.1" __all__ = ["__version__"] diff --git a/tui/uv.lock b/tui/uv.lock index d752b9bf8..1adcfde87 100644 --- a/tui/uv.lock +++ b/tui/uv.lock @@ -8,7 +8,7 @@ resolution-markers = [ [[package]] name = "agentcore-tui" -version = "1.25.0" +version = "1.25.1" source = { editable = "." } dependencies = [ { name = "httpx" },