chore: backmerge main into develop after 1.25.1 - #1364
Merged
Merged
Conversation
* feat(agents): degrade-with-notice binding resolution for project harnesses
shared-projects §9.6. A project can have 200 members; blocking the whole
turn because one member lacks one bound tool makes the project unusable to
them. resolve_agent_invocation(..., degrade=True) drops a missing
capability instead of raising, and records it in plan.unavailable:
- tools: a denied SERVER drops every ref of it (the gate is per base id,
so skipping only the checked ref would let a later scoped ref through);
an all-dropped toolset is still an empty toolset, never a fall-through
to the request's enabled_tools
- model: no override, so the route's chain picks the member's default
- skills: filtered; none left means no skill binding
- memory space: skipped
AgentNoticeEvent (SSE agent_notice) serializes the drops. Ordinary shared
agents keep block-with-message (D5) unchanged — degrade defaults off.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): run a project's harness in chat, with per-project cost (PR-1.4a)
shared-projects PR-1.4a. Membership already gates the harness (1.2's access
delegation), so the route adds what a project turn needs on top:
- Archived projects refuse new turns with a conversational error naming
the project.
- The resolver runs with degrade=True for a harness; anything dropped is
streamed as agent_notice before message_start (SSE only, never the
prompt, not persisted). Documented in CLAUDE.md's event table.
- "## Project Instructions" replaces "## Assistant-Specific Instructions"
for a harness, via compose_agent_system_prompt, which takes no user
argument, so every member of a project renders a byte-identical prefix.
Every other agent's text is unchanged byte for byte.
- preferences.projectId is written at session binding.
- projectId rides every C# cost row next to turnAgentId (route →
ChatAgent → StreamCoordinator), and the metadata writer adds the call to
PROJECT#{id}/COST#{YYYY-MM} plus COST#{YYYY-MM}#USER#{userId}: one
atomic ADD each, both UpdateItem (the runtime has no PutItem). The spec's
byUser map became per-member rows. Best-effort like every aggregate.
- kb-sync already copies apis/shared/projects; the scheduled-runs image now
does too (plus dynamo_errors.py) because sessions/metadata.py reaches the
repository lazily. Import closure only.
Personal instructions (1.4b) are split out: they change every user's
system prompt, not only project turns.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(specs): link the Projects UI mockup from shared-projects §6
Adds the clickable prototype link (claude.ai/artifact/2syz1g6zY2ZsCAkgsWVp7F)
with its two deliberate departures from §6, the open questions it raises,
and the convention that a PR changing a behavior the mockup shows notes it
in its as-built entry for the 1.8 re-sync. Backfills those notes for the
two merged PRs that already made such decisions (1.2 delete/transfer/
invite shape, 1.4a agent_notice).
Requested from the "Projects UI mockups" session on Phil's behalf.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(infra): ProjectSessionIndex on sessions-metadata (PR-1.6-infra)
shared-projects §3.4. A member's own tasks in one Shared Project, newest
first: GSI5_PK = PROJECT#{projectId}#USER#{userId}, GSI5_SK =
{lastMessageAt}#{sessionId}, projection ALL. The recency key of
SessionRecencyIndex (GSI4), sparse the same way, so the 1.6 backend
writes and removes it wherever GSI4 is.
Lands alone and ahead of any writer: it is the one GSI this plan adds to
sessions-metadata (one GSI per UpdateTable), and it is inert until rows
carry GSI5 keys. gsi-inventory.json regenerated; tables-detailed pins the
key shape.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(infra): enable CloudFront access logs on the SPA distribution
A user-reported 404 on prod (2026-09-24) could not be traced: app-api and
the ALB had no matching request, CloudFront's 4xx metric showed errors in
the window, and logging was disabled on the distribution, so nothing could
name the URI or client. Requests the edge answers itself (a lazy chunk a
deploy removed, an S3 403) never reach the ALB.
- Standard logging (legacy) to a dedicated S3 bucket under `spa/`. v2 is
configured via CloudWatch vended-log deliveries that must be created in
us-east-1, which the single app-region stack cannot do.
- Bucket: ACLs enabled (BUCKET_OWNER_PREFERRED; legacy delivery writes
through the ACL and fails against BucketOwnerEnforced), SSE-S3, block
public access, SSL-only, 90-day expiry.
- Cookies are never logged (the SPA session is an httpOnly cookie).
- CDK_FRONTEND_ACCESS_LOGS_ENABLED / frontend.accessLogsEnabled: default
ON with a kill switch. The bucket is provisioned unconditionally so a
disable/re-enable never collides with a RETAINed bucket name.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): project tasks, project shares and forks (PR-1.6 backend)
- Write ProjectSessionIndex (GSI5) keys beside GSI4 for sessions with
preferences.projectId: on store, on activity, and removed on soft-delete.
Reads now strip all recency keys, so a read-modify-write no longer replays
stale GSI4_* extras (which could SET and REMOVE the same attribute).
- GET /projects/{id}/tasks: the caller's own sessions in the project, newest
first, value-cursor paginated; a missing index degrades to empty.
- accessLevel "project" on shares: requires the task's projectId, the caller's
membership and an active project; the read check is membership.
- SHARED_TASK#{sessionId} pointer on the projects table, rebuilt from the
share rows on create/revoke/update/session delete (newest project share
wins); GET /projects/{id}/shared-tasks lists them without user ids.
- A fork keeps its project (projectId + the project's current harness) when
the requester can work in it; otherwise it stays a plain session.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(spa): recover from stale lazy chunks after a frontend deploy
deploy.sh syncs with --delete, so a tab opened before a deploy asks for
lazy chunks that no longer exist the first time it navigates to a route it
has not loaded (reported on prod 2026-09-24: Settings -> API Keys showed an
error; no request ever reached app-api).
- isChunkLoadError matches the Chromium, Safari, Firefox and webpack
wording (plus HTML-served-as-module), following cause/rejection/error.
- A failed navigation becomes a full load of its destination, guarded to
one auto-reload per tab per 5 minutes (sessionStorage, fails closed).
It is skipped, in favour of a "new version available - Refresh" toast,
when a chat stream or upload is in flight, since in-app navigation keeps
those alive and a page load would not.
- Chunk failures outside navigation (@defer, lazy libraries) reach a
custom ErrorHandler and only prompt.
- Proactive check (environment.versionCheckEnabled): on tab focus, at most
every 10 minutes, compare the running main-<hash>.js with the one the
no-cache index.html names. Inert on the dev server.
- ToastService gains an optional action button.
Composer drafts need no change: the composer mirrors to localStorage on
every edit, so there is nothing to flush before the reload.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(users): resolve an email to its live profile, not an arbitrary duplicate
Some emails own more than one PROFILE row: the pre-Cognito login keyed
users by a numeric employee ID, the current one by the Cognito sub, and
nothing retired the old rows. get_user_by_email queried EmailIndex (no
sort key) with Limit=1, so which row it returned was arbitrary.
- UserRepository.get_users_by_email pages through every match and ranks
them: most recent lastLoginAt (compared as instants), then non-numeric
ids over legacy numeric ones, then id for determinism.
- get_user_by_email returns the first and logs a warning naming the
ignored duplicates.
- Admin email search returns every profile, live first, so support sees
the ambiguity instead of landing on the stale row.
- The sharing user search dedupes by email, not id; previously a stale
exact match plus the live row from the name scan listed one person twice.
- upsert_user logs when a newly created profile's email already belongs
to another id (creation only; returning users pay nothing).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(admin): flag emails with more than one profile in user lookup
When an email search matches several profiles, show how many, badge the
one that signs in, and print each row's id, so a quota override or tier
is not assigned to a legacy duplicate nobody can log in to.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(scripts): audit legacy duplicate user profiles, retire them in two phases
backend/scripts/audit_user_duplicates.py groups users-table PROFILE rows
by email, ranks each group with the same live_profile_rank the API uses,
and checks every stale id against each place a user id is stored: one
Query per indexed key path across 20 tables, an S3 prefix probe per
user-keyed bucket, AgentCore Memory sessions for the actor, and, with
--deep, a single Scan per table with un-indexed references. It writes a
JSON and a Markdown report.
Read-only by default. --apply mark sets an unreferenced legacy row to
status=inactive with mergedInto/mergedAt (reversible; status=merged
would not parse). --apply delete removes only rows marked at least
--min-soak-days ago for today's live id. Only numeric-id + uuid pairs
whose every check came back zero are eligible; a failed check makes a
row incomplete, not unreferenced. Every write is conditional on the
lastLoginAt the audit saw, and --apply needs --confirm-prefix.
A coverage test fails when a CDK table is neither checked nor listed as
holding no user ids. live_profile_rank and item_to_profile become public
module functions in the users repository so the script cannot drift
from the API's choice of live row.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(spa): reset the refresh prompt's destination after a successful navigation
Found validating #1262 in dev: a navigation that fails on a stale chunk
while a stream is in flight shows the Refresh prompt aimed at the failed
destination. If the user then opened another view (e.g. went back to
the conversation) and clicked Refresh, they were sent to the view that
failed minutes earlier instead of reloading where they were.
AppUpdateService now clears the remembered destination on NavigationEnd.
A failed or cancelled navigation doesn't end in NavigationEnd, so the
failure that set it, or a guard refusing the next one, leaves it in place.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): people directory for inviting members (PR-1.3)
- apis/shared/directory: DirectoryAdapter port, DIRECTORY_PROVIDER switch
(default users_table; unknown values fall back), and UsersTableDirectory,
which pages the whole active partition of StatusLoginIndex instead of the
100 most recent sign-ins, matching email prefix and name in one pass,
ranked by match quality then recency. The active list is held 60s per
process so a typeahead filters in memory per keystroke.
- GET /projects/{id}/directory?q=&limit= (viewer): email, name,
hasSignedIn and the person's current memberRole, no user ids; a
well-formed unknown email is appended so it can always be invited.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(scripts): flag legacy user ids that are still in use
lastLoginAt cannot show that an old id is idle: an API key minted under
it still authenticates as that id, the api-converse path never updates
the profile's lastLoginAt, and it 401s a key whose profile row is gone.
So deleting such a row would break a live integration outright.
The audit now reads each stale id's API keys and sessions-metadata
partition and marks the row in_use when it owns an unexpired key or an
active scheduled prompt, or when any model call, message or key use is
newer than the day its live twin was created (that person's cutover).
in_use rows are never retired and lead the Markdown report; every stale
row also reports its newest activity so "history only" is visible.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(spa): shrink the agent-state orb by 20%
The 14px core with a 26px halo outweighed the status line beside it.
Scale it to an 11px core with a 21px halo (inset -5px) and a 10px glow,
keeping the same motion, proportions and timing. The finished-turn recap
dot follows so the live-to-settled transition still reads as one element.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(scripts): count deep-scan references per row, not per attribute
The first read-only dev run reported "deep:app-roles 2" for a single
TOOL_PREFERENCES row that names the id in both PK and userId. Count one
hit per item and list the matching attributes in the sample instead.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): project settings with version history (PR-1.5a)
- GET/PUT /projects/{id}/instructions, /model, /tools, /skills: viewer
reads, editor writes on an active project. Tools/skills replace only
their own kind. Only what a save adds is validated against the saver
(validate_agent_write); a no-op save cuts nothing.
- Every change cuts an AgentVersion (createdBy + createdByEmail); the first
save also records the starting state. GET .../instructions/versions and
.../versions/{n} return history and a diff against the previous version.
- The project is the harness's only write path: PUT /assistants and
PUT /agents return 409 on a harness.
- An archived project's harness resolves every member as viewer, so no
agent-level write route (documents, web sources, sync) can change it.
- Move diff wire helpers into version_diff; label-able instructions_diff;
public ProjectService.authorize.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(deps): bump dependencies to clear 21 Dependabot alerts
Resolves all open Dependabot alerts (2 critical, 5 high, 14 medium)
across the four dependency manifests. Every resolved version meets or
exceeds its advisory fix version.
backend (pyproject.toml + uv.lock)
- anyio 4.12.1 -> 4.14.2 (#409 critical GHSA-82r6-8w77-94w6, #408).
Transitive via starlette/httpx/fastapi, so added as an explicit pin
in the existing "Security: pin transitive deps" block.
- soupsieve 2.8.4 -> 2.9.0 (#403, #404, #405, #406)
tests/load (pyproject.toml + uv.lock)
- pytest 8.4.2 -> 9.0.3 (#395)
frontend/ai.client (package.json + package-lock.json)
- sharp 0.33.0 -> 0.35.4 (#385, #394 high)
- @angular/{common,compiler,core,forms,platform-browser,router,
compiler-cli} 21.2.19 -> 21.2.20 (#398, #400, #402). The whole
family moves together because Angular packages peer-depend on exact
sibling versions.
- vitest + @vitest/coverage-v8 4.1.5 -> 4.1.11, pulling @vitest/mocker
to 4.1.11 (#392, #393)
docs-site (package.json + package-lock.json)
- astro 7.1.3 -> 7.2.8 (#390 critical, #389)
- sharp 0.35.3 -> 0.35.4 (#388 high)
- svgo 4.0.2 -> 4.1.0 (#386 high, #387)
- smol-toml 1.7.0 -> 1.9.0 (#391 high)
- devalue 5.8.2 -> 5.9.4 (#407)
These three are transitive but already sat under permissive caret
ranges admitting the patched versions, so no overrides were needed.
Verified: backend pytest 9902 passed / 3 skipped; frontend npm ci +
build + full suite (301 files, 3823 tests) passed; docs-site npm ci +
build (52 pages) clean. npm audit reports 0 vulnerabilities in both
JS projects.
* feat(projects): project files over the harness's documents (PR-1.5b)
- /projects/{id}/knowledge: members list, read and download the project's
files (the agent document routes are editor-only); editors on an active
project upload, import, crawl and delete through the agent's own route
handlers, so provisioning, the byte cap and cleanup are not duplicated.
- addedByUserId on every document create path: create_document defaults it
to the provenance importer (import, crawl, sync) and device upload passes
the uploader. Project responses show addedByEmail from the member list.
- Upload, import and crawl responses carry the "everyone in the project can
open this file" notice.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): audit trail, notification inbox, admin.projects (PR-1.7)
- project.* audit actions recorded from the project service (lifecycle and
membership), settings saves, the knowledge routes and project shares, on
target `project`; GET /projects/{id}/audit for editors (no user ids).
- Per-user inbox on the projects table (INBOX#{email}, 90-day ttl), keyed by
email so an invitation reaches someone who has never signed in:
project_invited / role_changed / removed / ownership_transferred.
GET /notifications, POST /notifications/{id}/read, POST /notifications/read-all.
- admin.projects (delegable): /admin/projects list, detail, force-archive or
restore, and the full trail. Scope registry, coverage test and the SPA's
scope id list updated.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(models): add managed-model retirement runbook
Sibling of mcp-server-retirement.md. Models differ from tools in three ways
that reshape the stages: the runtime model check reads role grants only (not
the catalog row or `enabled`), AWS owns the EOL clock, and a model has a
successor that can be invoked in its place — so the cutover is a redirect,
and a model row is tombstoned rather than deleted (a deleted row plus a
surviving `*` grant runs unmetered and uncached).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(agent): stop truncating agent instructions at ~1,400 characters
The 8 KiB cap in SystemPromptBuilder.from_user_prompt was sized for an
agent's instructions alone, but it is applied to the whole composed block:
the ~6.8K-character default prompt and date, then the instructions. Every
agent's instructions were silently cut after ~1,400 characters (133 of 232
prod agents with instructions, 2026-09-24; the longest is 32,887).
- MAX_AGENT_INSTRUCTIONS_CHARS (100,000) is the author's cap, enforced at
save (Create/UpdateAssistantRequest, project settings) and on a request's
system_prompt (the Agent Designer preview, which 422'd past 8 KiB).
- The runtime bound becomes that cap plus PLATFORM_PROMPT_HEADROOM (64 KiB)
for the platform text composed around it, so a saved agent never truncates.
- Tests: a 33K-character agent keeps its last rule, the default prompt plus
a default memory block must fit the headroom, and save/preview share a cap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(settings): personal instructions in every conversation (PR-1.4b)
- UserSettings.personalInstructions (max 4,000 chars) via PUT
/users/me/settings; trimmed, blank clears.
- Appended last in the instructions block as "## Personal Instructions",
below an agent's or project's instructions, with a precedence sentence
only when there is something to defer to. A user without them keeps a
byte-identical prompt (and cached prefix); project members still share
the project part of the prompt.
- Applied on plain, agent, project and @-mention turns (not previews); the
MCP App dispatch paths add the same text so they reuse the turn's cached
agent. One settings read per turn, shared with the default-model lookup.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(sidenav): scroll the nav entries with the sessions under a pinned New Session
Only the session list scrolled; New Session, Agents, Artifacts and Customize
sat frozen above it, eating the column on short windows. Now only New Session
is pinned and everything below it scrolls as one region.
A fade under New Session appears once the region has scrolled, so rows slide
out beneath the button instead of being sliced by a hard edge. It is drawn
from the sidebar's own background token, so it reads the same in light and
dark mode.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(sidenav): compact the user bar to the nav rows' rhythm
73px -> 57px. The trigger now matches a nav row (40px, rounded-md, the same
hover tint), with a 28px avatar that lines up with the nav icons above it.
The unread-announcements dot moves with the smaller avatar's corner.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(sidenav): page the session list as the user scrolls
The sidebar fetched a user's entire history on every load (`GET /sessions`
with no limit). It now asks for 30 and appends the next page as the end of
the list nears view, using the cursor the endpoint already returned.
Reloads re-fetch the whole loaded window in one request (`limit` = rows
held) rather than page one alone. The list reloads constantly — new session,
delete, rename, tab focus, mark unread — and a first-page refetch would drop
rows: a new session pushes page one's last row past the boundary, where the
pages appended below it do not have it either.
A page that lands after a reload started is dropped rather than written,
because `resource.set()` aborts the in-flight load. Each sighting of the
end-of-list sentinel pays for one page and is re-measured after the page
lays out, so a scroll to the bottom never fetches a page too many and a
tall window keeps filling. A failed page shows a Retry and stops the loop.
Removes the unused `updateSessionsParams` / `resetSessionsParams`.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(models): lifecycle status with redirect-to-successor on retirement
Adds status (active/deprecated/retired), replacedBy, retiresOn and
retirementNote to managed models, and resolves a retired model before the
grant-only runtime access check at every entry point — /invocations (model and
provider swapped at the top of the handler), the saved-default fallback, Agent
modelConfig (live record and published snapshot) and /chat/api-converse. A
retired model with a successor runs as it, access-checked and priced on the
successor; with none it is refused for everyone, wildcard holders included, as
a conversational message (410 on the API-key route). Deprecation stays a
picker concern: nothing at runtime reads it, and the Designer write check does
not filter on it.
Admin writes validate the successor (exists, active, enabled, not itself),
refuse a non-active default, and DELETE refuses a model another row names as
its replacement. See docs/specs/model-retirement.md §7.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(spa): model retirement guards in the pickers and admin form
A deprecated model can be kept but not newly chosen: the chat picker hides it
unless selected (then badges it and names the successor), the Designer hides it
unless it is the agent's model (then shows a notice), and Settings disables it
as a default option. A selection on a retired model follows replacedBy to the
successor with a once-per-tab toast instead of a silent fallback. The admin form
gains Status / Replaced by / Retires on / Note, and the list badges non-active
rows. The retirement helpers move to shared/utils/retirement.ts, re-exported
from tool.service.ts so no tool surface changes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(models): record the built §7 in the model retirement runbook
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): Projects SPA shell, Overview, Members and Settings (PR-1.8a)
- /projects: your own and shared projects, filtered All / Mine / Shared with
me / Archived, with each card naming your role; "not available" when the
kill switch is off.
- /projects/:id/:tab (the bare URL redirects to overview): header with role,
archived notice, pill tabs.
- Overview: a composer that starts a task bound to the project's agent, and
a rail with the instructions and the people.
- Members: people picker over /directory (existing members marked, pasted
lists staged at once, unknown emails allowed), role changes, remove, leave,
and "Make owner" for editors who have opened the project.
- Settings: details, instructions, model, tools and skills (each save is a
version), history with per-version diffs, and owner controls (editors
manage members, archive/restore, delete).
- Sidenav entry. Specs for the API client, facade, picker and both pages;
axe clean in light and dark.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(spa): chime when dictation starts and stops listening
Two short synthesized tones bracket a dictation: rising (E5 -> B5) once the
mic is live, falling (B5 -> E5) when it goes off. They track the mic, not
the buttons: start plays on the transition into `listening` (the first
moment speech is captured), stop on any way out of it — Done, Cancel, the
time limit or a dropped socket. An attempt that never reached `listening`
plays neither.
Web Audio, no asset files. The AudioContext is primed synchronously inside
`start()` — before the first await — because browsers only allow audio
after a user gesture and the start chime lands well after the click. Every
failure path (no Web Audio, a refused context) is a silent no-op.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(spa): compact send-less composer, meta line below, agent breadcrumb
Drop the send button (Enter sends) and give its other jobs their own
homes: Stop takes voice mode's slot while a response streams (Escape also
stops, once no menu or dictation claims the key), touch devices get a Send
button only once there is a draft (return adds a line there), and an
upload that blocks Enter says so on the status line.
In a conversation the composer is one row — attach, text, speech
controls — with a line beneath it: cost and context on the left (or a
short status), the model picker on the right. A draft that wraps unfolds
to text-on-top with the controls in a bar beneath, and folds back only
once empty so the layout cannot flip at the wrap point. The empty state
keeps the tall composer. The first send swaps one for the other across
the `/` -> `/s/:id` rebuild, so ComposerHandoffService carries the tall
composer's height and the compact one animates down from it.
The `@`-mention and `/skill` chips go: the tokens are tinted in place by
a mirror layer behind the textarea, and deleting an `@Name` now
un-mentions (the binding used to outlive its text).
The bound agent leaves the composer for the top nav, as a breadcrumb
before the conversation title (`Agent / Title`) projected through a new
`[topnavCrumb]` slot, with its menu opening downward. The model dropdown
gains a compact size for the meta line, and the cost/context popovers
anchor left now that the badge sits at the composer's left edge.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(projects): Tasks, Files, share-to-project, sidebar grouping, agent_notice (PR-1.8b)
- Tasks tab: the caller's tasks (paged, open on the project's agent) and
tasks shared with the project (open, continue as a fork, revoke).
- Share modal: "Project members" access level for a task in a project,
with the API's 403/409 detail shown inline.
- Files tab over /projects/{id}/knowledge: list with "Added by", upload
with the server's sharing notice, polling, download, delete; viewers
read and download only.
- Sidebar: project tasks grouped under a project heading inside each
time bucket; optimistic rows adopt the API's preferences.
- agent_notice SSE event: parser type/validator/callback, per-session
notice service, dismissible banner above the composer.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(inference): fall back to the catalog's default model, not an unpriced id
A turn that names no model (scheduled runs, Agents without a modelConfig,
model_id: null) fell through to the hard-coded Defaults.MODEL_ID, a us.*
Haiku id with no catalog row where prod registers global.* ids. It priced
to None, so those turns were unmetered and free against quota.
The fallback is now: the user's saved default, then the catalog's enabled
isDefault row (the model the SPA already pre-selects), then Defaults.MODEL_ID
only when the catalog has no default. A saved default with no catalog row is
treated as unset. The MCP App dispatch paths share the resolver so they keep
hitting the turn's cached agent.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(observability): alarm on model calls that price to nothing
Emit an UnmeteredModelCall EMF metric (chat and voice) whenever a call
spends tokens but gets no cost, because it then reaches neither the cost
rollups nor quota. Adds a threshold-0 alarm and dashboard widgets that name
the model id missing a catalog row.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Platform self-service: system tool tier + account tools (flag-gated) (#1279)
* feat(tools): system/hidden tool tier (platform self-service Phase 1)
Adds a platform-shipped 'system' tool tier on top of the existing always_on primitive, plus a 'hidden' display flag, per .kiro/specs/platform-self-service.
- ToolDefinition gains system + hidden (default false; backward-compatible dynamo round-trip). system implies always_on via the normalizing validator.
- freshness: new system-tool-ids snapshot filled by the same single catalog pass; cleared on invalidate.
- always_on.py: resolve_system_tool_ids — like resolve_always_on_tool_ids but NOT gated on ADMIN_ALWAYS_ON_TOOLS_ENABLED (system is part of the app).
- inference routes + voice: system tools unioned into every turn, including agent-bound turns (author scopes user-facing tools, not platform plumbing).
- app_api tools service: surface hidden on UserToolAccess so the SPA excludes it from the toggle list but keeps it for tool-use event labeling.
- Tests: model round-trip/validator, freshness snapshot, and the invocation seam (system applies on bound turns; system+always_on combine; dedup). 63 targeted tests green; changed regions ruff-clean.
* feat(tools): account self-service tools, read-only pilot (platform self-service Phase 2-3)
Phase 2 (identity binding) + Phase 3 (Account & Usage read tools): whoami,
get_my_quota, get_my_settings. Identity is captured by closure (make_*_tool
factories injected as extra_tools), matching the six existing per-request tool
families -- the runtime does not populate Strands ToolContext, so this is the
proven pattern and satisfies Req 3 more strongly (tools take no user arg).
- get_my_quota is read-only: computes via resolver + cost aggregator, NOT
QuotaChecker.check_quota (which records warning/block events).
- Gated behind PLATFORM_SELF_SERVICE_ENABLED (default off) per the rollout
plan, so agent-cache eligibility and the cacheable prefix are unchanged when
off; when on the tools are key-described (close over user_id) and cacheable.
- Catalog metadata (system/hidden) added for transcript labeling; whoami hidden,
quota/settings visible-but-locked. DynamoDB row seeding and SPA panel-hiding
deferred (not needed for the pilot).
- 18 new tests; targeted suites green.
* feat(tools): admin-governed runtime off-switch for system account tools
Reverses the earlier 'system tools are not an admin knob' call after local
testing showed code-only delivery needs a redeploy to disable a tool.
- get_system_tool_ids now filters on status: a system row set to disabled/
deprecated drops out of the injected set within the freshness TTL -- an
admin off-switch with no redeploy (via the existing admin Tools panel).
- _build_account_tools injects a tool only when its id is already in the
turn's effective set (row present, system, active, RBAC-granted), so a
disabled row or ungranted role removes it live.
- Seed whoami/get_my_quota/get_my_settings rows in seed_bootstrap_data.py
(system + hidden per rule + isPublic). Fixes test_seed_matches_tool_catalog
parity (Phase 3 catalogued them without seed rows) and makes them appear in
the admin panel. Run the seeder once per env; PLATFORM_SELF_SERVICE_ENABLED
remains the per-env master gate.
- Tests: freshness excludes non-active system rows; builder injects only ids
in the effective set. 46 targeted + 183 broader green.
* fix(tools): add 'account' to ToolCategory enum
Seeded account tools (whoami, get_my_quota, get_my_settings) use category='account', but the DynamoDB-backed ToolDefinition.ToolCategory enum lacked that member, so list_tools() 500'd when reading the rows back and GET /tools/ failed on load. Adds ACCOUNT to the enum (the static tool_catalog enum already had it) plus a from_dynamo_item round-trip regression test.
* feat(self-service): add set_default_model confirmed-write tool
Phase 6 of platform-self-service: the first WRITE in the Account & Usage tier, completing the about-the-user capability.
- set_default_model is a system/always-on injected tool (like the reads), admin-governable via its catalog row, gated by PLATFORM_SELF_SERVICE_ENABLED.
- Two-step/confirmation-gated: previews the change first, applies only on confirm=true after the user agrees.
- Only ever changes the caller's own defaultModelId (identity closure-bound, never an argument) and only to a model they may use — accessibility mirrors ModelAccessService.filter_accessible_models via shared RBAC (AppRoleService) + list_all_managed_models, staying inside the agents->shared import boundary.
- Persists the record UUID (what defaultModelId is keyed on) via the user-settings repository's idempotent update_settings.
- Catalog + seed rows added (visible-but-locked); 12 new tests (handshake, accessibility reject, ambiguity, write, never-raise, identity).
* docs(spec): record phase re-order + Phase 6 done in platform-self-service tasks
* feat(projects): notification bell, Activity tab, personal instructions (PR-1.8c)
- Notification bell beside the user menu: unread badge from
GET /notifications, panel of the newest 20, open marks read and goes to
the project, mark all as read; refreshes on open and on tab focus.
- Activity tab (editors and owner): the project's audit trail as
sentences, paged, with links to the settings version a save cut.
- Settings > Chat: personal instructions (4,000 characters), saved and
cleared inline.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(projects): user guide, admin page and env vars for Shared Projects (PR-1.9)
- docs-site features/projects.md: roles, tabs, tasks and sharing, files,
settings history, notifications, personal instructions, lifecycle.
- docs-site admin/projects.md: admin.projects routes, the kill switch and
what it doesn't stop, configuration, where the data lives.
- .env.example and the docs-site env-var page: Shared Projects variables.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(projects): make PROJECTS_ENABLED a full stop; drop the Projects nav item
- A project harness now resolves no role for anyone while
PROJECTS_ENABLED=false, so the kill switch stops project chat turns on
inference-api and the agent document/sync routes on app-api, not only the
/projects surface.
- A turn in a project task while Projects are off gets a conversational
message saying so, instead of a bare 403.
- The sidenav no longer has a Projects entry: a menu item must not depend on
a request to know whether the feature is on. /projects stays reachable by
URL, from sidebar project headings and from notifications.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(projects): kill switch is a full stop; no Projects nav item
Describe the behavior after fix/projects-kill-switch: a project's assistant
refuses everyone while PROJECTS_ENABLED=false, a project task turn gets a
conversational message, and /projects is reached by URL, sidebar project
headings or notifications rather than a menu item.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(spa): tighten the gap beneath the compact composer's meta line
The cost/model line's 32px row already leaves slack beneath its text, so the
footer's full 1rem bottom padding left ~26px below the line against ~14px
above it. Compact footers (conversation + embedded preview) now pad 0.5rem;
the empty-state composer, which has no meta line, keeps 1rem.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(spa): drop the top nav's below-lg separator
It sat between the conversation title and an empty slot, so it divided
nothing.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(projects): assert the agent routes' permission check also refuses while Projects are off
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(agents): show a retiring model on the agent detail page and in runnability
The detail page's Details panel named an agent's pinned model even after it was
retired, and both "will it run?" checks (per viewer, and per role for pins)
answered ready for an agent on a retired model with no successor — which the
runtime now refuses for everyone.
GET /agents/{id} now carries modelRetirement (status, successor display name,
date, note) and the page renders a line under the Model row. resolve_runnability
and the role-pin diff resolve retirement first, like agent_binding_resolver: a
redirect is checked against the successor, and a retired model with no
successor is missing regardless of grants.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(admin): simplify Manage Models rows; move delete to the edit page
Each row is now drag handle, icon, name/id, price, edit, toggle.
- Drop the provider chip: the model icon already names the provider.
- Drop the detail accordion. Pricing, the one detail admins compare
across rows, moves inline as an input/output per-1M column; roles and
modalities stay on the edit page.
- Move the enable switch to the far right so state reads down one
column. Drop the Enabled/Disabled label; disabled rows dim instead.
- Move delete off the row to a Danger zone on the edit page. Deleting a
model's row leaves its usage unpriced and breaks retirement redirects,
so it should not be one click away on every row. The section points
at Status: Retired as the safer alternative.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(agents): draw the detail-page hero as our chat composer
The marketplace detail hero showed the @-mention prompt in a generic
search pill — rounded-full, blue bold mention, black send-arrow button —
none of which our composer has. Redraw it as the compact composer:
rounded-2xl shell with the same outline and shadow, leading attach,
trailing dictate and voice glyphs, the composer's @mention tint, and no
send button (Enter sends on desktop). Still display-only.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* style(tooltip): tighten tooltip sizing
Tooltips read as oversized next to the icon buttons they label. Drop to
12px medium text with px-2/py-1 padding, a smaller arrow and shadow, and
halve the overlay offset so the tooltip sits closer to its trigger.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(flags): in-development features default off; SPA feature switches in Angular environments
- Policy (CLAUDE.MD "Feature Flags"): new and in-development features are
opt-in per deployment; finished features move to default-on deliberately.
- Shared Projects is the first: PROJECTS_ENABLED / CDK_PROJECTS_ENABLED now
enable only on "true" (unset = off).
- SPA: `features` in src/environments/environment*.ts (typed by
feature-flags.ts, read via the FEATURES token). New environment.development.ts
+ `dev-deploy` build configuration; build.sh builds SPA_BUILD_CONFIGURATION,
which frontend-deploy.yml and nightly-deploy-pipeline.yml set to match the
environment they deploy to.
- Projects nav item and notification bell return behind features.projects;
/projects routes canMatch on it; project grouping in the sidebar and the
"Project members" share option follow it too. Dev: on. Prod and local: off.
- Breadcrumbs for agents: root CLAUDE.MD, frontend .claude/CLAUDE.md,
feature_flags.py, feature-flags.ts and each environment file.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(projects): Projects are opt-in while in development; SPA switch in environments
Follow feature/feature-flags-default-off: PROJECTS_ENABLED / CDK_PROJECTS_ENABLED
enable only on "true", and the SPA's features.projects (src/environments)
decides whether the UI offers Projects, including the nav item.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(specs): agent handoffs and consults for mid-thread @-mentions
Phase 0 (task chips): hand off to a new conversation with a reviewable
context card, send results back, and let the model suggest tasks.
Phase 1 (consults), where an @-mentioned Agent runs in place and reports
back, is built only if Phase 0's measured send-back rate passes a gate.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(memory-audit): Phase 0.1 AgentCore Memory baseline audit + decision record
scripts/memory-audit/audit.py: read-only inventory (strategies and namespace
templates vs the backend's, actor-id classes incl. the session-id fallback,
record census by strategy/namespace shape, extraction jobs, runtime log
signals) and a dev-only probe that replays retrieval with the runtime's
top-k and relevance cut against a synthetic actor, then cleans up. raw/
output is personal data and stays out of the repo; summary.json is
aggregates only.
docs/specs/memory-baseline-decision.md: dev evidence. Memory is written,
extracted and correctly namespaced, but the 0.7 relevance cut discards every
realistic hit (questions score 0.57-0.67), so no dev turn received context
in 7 days and the two-chat test failed. Recommends C (hybrid), conditional
on re-testing after 0.2 lowers the cut.
tests/supply_chain/test_memory_audit.py pins the audit's copies of the
backend namespace templates and retrieval parameters.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(memory): Phase 0.2 AgentCore Memory fixes from the baseline audit
- Relevance cut 0.7 -> 0.5 (constants default). In dev, questions about a
stored fact scored 0.57-0.67 and unrelated records <= 0.40, so 0.7 dropped
every realistic hit and no turn ever received memory context.
- Session delete also purges that session's SUMMARIZATION records (exact
per-session namespace, batches of 100), even after its events expired.
Semantic facts/preferences carry no source session and are left alone.
- Share-fork writes events with extractionMode="SKIP", so another user's
messages stay in the fork's history but never feed the forker's
long-term extraction.
- Runtime role: both memory statements share RUNTIME_MEMORY_ACTIONS (the
scoped one lacked GetMemory under a wrong comment).
- Stale "write-only" lines corrected in two specs and the
app_context_dispatch docstring (which cited a deleted analysis doc).
- CDK test pins strategy names, no custom namespaces, 90-day expiry and
identical Runtime memory action sets.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(projects): purge tears down the managed KB, rename reaches the harness, bell passes axe
Three bugs found validating Shared Projects Phase 1 on dev (2026-09-25).
- Deleting a project, or an ordinary agent, left its KB# record and the
born-managed Bedrock knowledge base ACTIVE and billing. Both delete paths
now queue the record in a new `teardown` migration state (generation bump
fences any in-flight worker); the kb-migration worker, which holds the
delete grant app-api lacks, runs the tombstoned saga under its lease,
deletes the data source and knowledge base, then removes the record.
Failures re-queue instead of going to `failed`. `assert_deletable` now
refuses a harness before DELETE /assistants soft-deletes its files.
- PATCH /projects/{id} renamed META only; the harness (and the chat
breadcrumb) kept the creation-time name. update_project now carries name
and description onto the harness without cutting a version, and project
history reports only settings fields.
- The notification panel's loading/error/empty states were a role="menu"
with no menu items (critical axe aria-required-children). They are now a
disabled menuitem; the header <p> is a div.
- Activity tab names models, tools and skills from the bindable palettes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(projects): show a project task's crumb as the project, not an agent
A project task binds the project's hidden harness Agent, so the chat's
breadcrumb rendered it as an ordinary Agent — offering Edit agent and Share
settings, both of which the harness refuses (409 on edit; share mutations
return not-found).
- GET /assistants/{id} now carries `kind` and `projectId` (omitted for
every other agent).
- The agent indicator takes a `projectId`: project mark, "Project" in place
of the owner line (the harness keeps its creator after a transfer),
New task + Open project, no Edit / Share, and "fixed by this project".
- The crumb reads the project's name from the project list when loaded,
since renaming a project does not rename its harness.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(memory): record the passing Phase 0.2 re-test; decision C (hybrid)
On dev, after 0.2 deployed (relevance_score=0.5 in the agent-build log),
chat B recalled a fact stated in chat A and the runtime logged
"Retrieved 1 customer context items" for that turn. Deleting chat A purged
its summary record. The decision record moves from recommendation to
decided: C (hybrid).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(chat): context meter with itemized attribution; fold cost into its panel
Composer meta row: the running cost figure is gone. A context ring now sits
to the right of the model picker (percentage appears inline only at >=70%).
Hovering, tapping or pressing it opens a panel with:
- what filled the window: System instructions, Skills, Memory, Tools,
Messages, Free space, as a stacked bar plus rows with sub-items;
- this conversation's cost and the user's quota (replaces the cost tooltip).
Replaces SessionCostBadgeComponent.
Backend: the measured system/tools totals are itemized by character share
(context_itemization.py), at zero extra CountTokens calls:
- system -> System instructions (platform / agent / project / personal /
mode sections) + Skills (per-skill rows) + Memory (bound Memory Space);
- tools -> children by origin (built-in, each Gateway target, each external
MCP server by catalog display name, skill tools, memory tools).
Each row set sums to the measured total, so prefixTokens and compaction are
unchanged; the long-TTL static-prefix size now sums every non-messages
partition.
No added pre-model-call latency: the hook still stores only the three
measured totals; itemizing happens on read (final metadata SSE, persisted
row), after the model has answered, memoized per agent.
The itemized breakdown is persisted on the turn's last message row
(contextBreakdown) and re-seeds the meter on reload. It stays out of the
admin CALL_ROW_PROJECTION: labels can name skills/servers and `label` is a
content-bearing path segment.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(agent-cache): cache agents bound to a Memory Space (Shared Projects 2.1)
Memory-Space tools close over the resolved binding (space id, name, access)
and read the space live on every call, so the binding becomes an agent cache
key element instead of a cache veto:
- _create_cache_key gains memory_binding (digest; "" without a binding),
placed ahead of the document/assistant/skills elements so their positions
do not move.
- injected_tools_are_key_described drops the has_memory_binding veto.
- get_agent takes memory_binding, stamps it on the construction snapshot;
PausedTurnSnapshot.memory_binding persists it and resume replays it, so a
paused memory-bound agent is found again.
- Resume passes cache_write=False. Resume builds no injected tools, so a
resume *miss* used to write a tool-less agent into the slot the next plain
turn hits, silently dropping artifact/document/spreadsheet (and now
memory) tools. Pre-existing for every injected family.
AGENT_CACHE_INJECTED_TOOLS_ENABLED=false still restores the blanket bypass.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: make time to first token a hard budget in CLAUDE.MD
Nothing may add latency between a request arriving and the first streamed
token (request handling, agent build, MCP pre-flight, BeforeInvocation /
BeforeModelCall hooks). Hooks capture raw facts; derived work runs after the
model answers, memoized. Anything unavoidable must be measured at realistic
scale, stated in the PR, and explicitly accepted by the developer.
Prompted by this PR: the first cut of context itemization ran inside
BeforeModelCallEvent at ~1.4ms per model call before it moved to on-read.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(chat): render the context meter from the start so the model picker never shifts
The meter only mounted once a turn had reported cost or a context window, so
it popped in beside the model picker after the first response and pushed the
picker sideways. It now always renders under the compact composer: an empty
ring until the first response is measured (arc hidden while empty, so the
round cap doesn't paint a stray dot), and a panel that says the window is
"Not measured yet" above the conversation's cost and quota.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(chat): context meter signals urgency by colour only
Drop the percentage that appeared beside the ring at >=70% full. It widened
the trigger and nudged the model picker left; the ring's colour
(success -> info -> warning -> danger) now carries urgency on its own, and
the exact figure stays in the panel and the trigger's aria-label.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(memory): memory block behind its own prompt-cache point (Shared Projects 2.2)
The bound Memory-Space block used to be appended to system_prompt, so it sat
inside <user_instructions> and inside the single cached system block: every
index edit rewrote the whole static prefix, and member-written text was
wrapped as "instructions".
- Routes hydrate the block into `memory_context`, passed through get_agent ->
BaseAgent -> AgentFactory, which sends [static, cachePoint, memory,
cachePoint]: the fourth and last point (tools, system, memory, message).
Turns without memory send today's bytes.
- Rendered as <memory_space scope="agent" name=... note="...data...">; the
name is attribute-escaped and a closing tag in member text is defused.
- Budget is token-based (MEMORY_INJECTION_MAX_TOKENS, default 6,000 = the old
24 KB; MEMORY_INJECTION_MAX_BYTES still honored).
- Cache key folds memory_context into the prompt hash; PausedTurnSnapshot
persists it and resume replays it. Voice keeps memory in its prompt.
- contextBreakdown's memory marker moves to the new tag. Memory no longer
shares the user-instructions truncation headroom.
Baseline measured on dev before the change; the after-run follows deploy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(kaizen): weekly research scan 2026-09-25
Generated by the kaizen-research skill. Top 5 ideas appended to
docs/kaizen/review-queue.md for the kaizen-review-prep run later this morning.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(agents): DELETE /agents runs the full cleanup; shares go with the agent
#1293 made agent deletion queue the managed-KB teardown and clean up
documents, but only on DELETE /assistants/{id}. The SPA's Agents page
calls DELETE /agents/{id}, which deleted only the record, leaving DOC#
rows (and their S3 objects and vectors), sync policies, and the KB# record
with its Bedrock knowledge base behind.
Both routes now call one sequence, agent_designer/services/agent_deletion:
refuse or 404 before touching anything, queue the KB teardown, soft-delete
every page of documents (was first 1,000 only), delete sync policies,
delete the record, clean up documents in the background.
_delete_assistant_cloud also removes SHARE# rows.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(kaizen): weekly review prep 2026-09-25
Generated by kaizen-review-prep. Ranked agenda for the 10-15 min decision pass;
queue updated with prior-review outcomes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* chore(agents): script to clear rows deleted agents left behind; fix vector probe batch
backend/scripts/cleanup_orphaned_agent_rows.py finds AST# partitions with no
METADATA row (older than --min-age-hours) and, with --apply and a matching
--confirm-prefix, clears them through the app's own paths: KB# queued for
teardown, DOC# through cleanup_document_resources, leftover S3 objects,
sync policies, versions, reports, share and crawl rows. Report-only by
default. Run on dev 2026-09-25: 10 orphaned partitions cleared.
delete_vectors_for_document probed 500 keys per GetVectors call; the API
allows 100, so every probe failed and each delete listed the whole index.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(projects): record the 2.2 prompt-cache after-measurement (A wins)
Same agent, index and procedure as the baseline: the turn after a memory
edit now reads 12,552 and writes 1,775 (was 10,477 / 3,824). The static
system prompt is read instead of rewritten; the A-vs-B spike closes on A.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(tooling): align npm version pins and harden lockfile-sync check
Three different npm versions were in play, which is why bumping the
frontend dependencies in #1268 required working around an npm crash:
dev container 11.2.0 (.devcontainer/Dockerfile ARG + packageManager)
GitHub Actions 10.9.x (bundled with Node 22; nothing pinned it)
authored #1268 12.1.0 (npx npm@latest, as a workaround)
npm 11.2.0 cannot resolve this repo's frontend dependency graph: it
crashes with "Cannot read properties of null (reading 'edgesOut')" in
arborist's #loadPeerSet on node_modules/vitest, from a vitest <->
@vitest/coverage-v8 peerOptional cycle. That is a resolution-time bug, so
`npm ci` (which replays the lockfile rather than resolving) was never
affected and no build, test, or deploy path was broken — but authoring a
lockfile was impossible without reaching for a newer npm out-of-band.
Changes:
- Raise the pin to 12.1.0 in .devcontainer/Dockerfile and in
frontend/ai.client/package.json "packageManager". 12.1.0 is chosen
over the more conservative 11.20.0 because the frontend lockfile now
on develop was authored by 12.1.0, so aligning on it makes lockfile
regeneration a no-op instead of re-running the resolution that was
crashing. Engines are satisfied: npm 12.1.0 requires Node
^22.22.2 || ^24.15.0 || >=26.0.0; the dev container pins 22.22.3 and
CI's node-version '22' currently resolves to 22.23.3.
- Pin npm in version-check.yml, the only CI job that resolves rather
than replays a lockfile. It regenerates the frontend and
infrastructure lockfiles and asserts they are byte-identical to what
is committed, so authoring and validating with different npm majors
can fail the gate on a correct lockfile. The version is read from the
"packageManager" field rather than hardcoded, keeping one source of
truth.
- Stop discarding npm's stderr in that same step and check its exit
status. Previously a resolution crash left the lockfile untouched, so
the following `git diff --exit-code` came back clean and the check
passed vacuously — the failure mode it exists to catch was the one it
could not report.
Not verified (Docker unavailable in this environment): whether
infrastructure/package-lock.json is byte-stable under npm 12.1.0, and
whether npm 12 changes `npm ci` behavior for the three npm projects.
See the PR description for the follow-up this needs before the next
release PR to main.
* fix(compaction): never offload skill instructions from the skills tool
The Strands AgentSkills activation tool (`skills`) returns a skill's
SKILL.md body: instructions the model must follow verbatim. A prod
readout on 2026-09-25 (ToolResultOffloaded, toolName=skills) showed one
4,647-token activation swapped for a preview plus a retrieval handle,
which quietly degrades skill-following.
Add `skills` to OFFLOAD_EXEMPT_TOOLS beside document_read. Its size is
bounded by the skill author, not an external payload. read_skill_file
stays offloadable: reference files run up to 1 MiB of arbitrary text,
which is what preview plus retrieval is for. Memory-space tools return
user/agent-written data, not instructions, and also stay offloadable.
Per-turn tool payload only; the cacheable prefix is unaffected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(agents): stop Strands printing streamed model text to stdout
AgentFactory built Agent without a callback_handler, so Strands 1.55
installed PrintingCallbackHandler, which print()s every streamed text
delta to stdout with end="". The runtime ships stdout to CloudWatch, so
response text reached the AgentCore Runtime log group and, unterminated,
was glued onto the front of the next EMF line.
Pass callback_handler=None. Nothing consumes callback events: the stream
processor reads agent.stream_async() directly. The test builds a real
Agent and streams a turn, so it fails if the SDK default comes back.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(observability): start every EMF record on a fresh line
stdout is shared with anything in the process that writes to it. An
unterminated write turns into a prefix of the next EMF JSON line, and
CloudWatch silently declines to extract it as a metric. Prefix each
record with a newline so a stray write can't corrupt a metric.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(compaction): offline quality-veto harness (slice 1)
Adds backend/scripts/compaction_quality_harness.py and the
compaction_quality package: a seeded corpus of long editing sessions with
planted, exactly-scorable facts (constraint, decision, reference,
superseded), replayed through the production TurnBasedSessionManager under
each arm (full, model_relative, legacy, floor_50) and pace (restore, cold,
warm). Only I/O is stubbed; the cut, the deferred apply, the restore slice,
anchor truncation and bound_summary are the real code.
Steps: cut (free, writes probe-time histories plus a plant-availability
table), records and ask (Bedrock spend, estimate-first, --yes to run, a
cachePoint after the shared history), score (exact match, paired per family
against the full-history control, exact McNemar, n on every row).
The clock stub stamps real time on save and inserts the pace gap only
between turns; backdating every save hid the restore-path missed free
apply, which a tripwire test now pins until it is fixed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(compaction): record the harness, the control, and the prod readout
Spec §5 gains a status note above the waiver: the harness exists, the
control is full history rather than the kill switch, and the first full run
waits for the missed-free-apply fix. The scoping doc marks slice 1 built,
answers its open questions, and records the aggregate prod readout.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(compaction): read the head-of-turn cache gap from before any save
Every _save_compaction_state stamps updated_at. On a restore,
_apply_compaction advanced the truncation anchor (a save) before the
coordinator called apply_pending_compaction, which then read a ~0s gap and
left a parked cut waiting on a cache that was already cold; it landed only
at the paid hard-ceiling apply, after the session had also paid the anchor
re-write. apply_document_offload had the same defect after any apply that
itself saved.
The restore slice now captures the previous turn's stamp before the anchor
advance (_restore_turn_stamp); apply_pending_compaction consumes it every
turn into _turn_start_stamp, and _document_offload_reason reads that first.
Both stamps are cleared at turn end so an old stamp can never make a warm
turn look cold.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(compaction): flip the harness tripwire now the restore path applies
The quality harness's restore pace pinned the missed free apply; with the
turn-start gap fix it applies the parked cut on cache_expired. Spec §5 and
the scoping doc record the fix.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: teardown polls long enough for one run; name the Agents page icon actions
- kb_migration teardown passes its own 780 s poll window to the delete
saga. Managed KBs on dev took ~10 minutes to delete, so at the shared
480 s every teardown needed two worker runs (~30 min). The shared
default stays 480 s for the reconciler, which deletes several per run.
- Agents page Chat/Edit/Share/Delete icon buttons (grid and list) get
aria-labels naming the agent; the tooltip only set aria-describedby.
- Kaizen queue: the production orphaned-agent cleanup chore, blocked on
#1293 and #1301 reaching main.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(memory): canonical memory-file format and save validation (Shared Projects 2.3a)
Pure library for the canonical memory-file format (spec §4.2) and the
manifest-aware save checks (§4.3 steps 1-4), plus a bounded CountTokens
counter in apis.shared for save-time token accounting. No callers yet;
2.3b wires these into the save path.
- format.py: system-rendered frontmatter (YAML subset, no new dependency),
items with ULID anchors, links, aliases, slug rules, reserved MEMORY.md.
- validation.py: locked-field echo rule, name/alias collisions, anchor
known/unique/minted, link resolution incl. archived targets; only new
dead links fail.
- tokens.py: one-attempt, 2 s CountTokens with a labelled chars/4
fallback; MEMORY_TOKEN_COUNT_MODEL_ID="" turns it off (tests do).
- Hypothesis properties for round trips, anchor stability and link
resolution.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(memory): save pipeline and FILEVER history (Shared Projects 2.3b)
Wires the 2.3a format and validation into every Memory Space write path
and adds per-file version history.
- MemorySpaceService.save_entry: validate, count tokens once
(CountTokens), write the object, conditional manifest swap, then the
FILEVER row. The swap is the commit point, so version numbers cannot
collide; replaced objects are kept because history references them.
- fileFormat freeform|canonical on the space, set at create. Existing
spaces stay freeform (stored as written; warnings only). Canonical
spaces enforce items, anchors, links, collisions and the 8,000-token
hard cap; a concurrent save of the same file is a 409.
- Pre-history entries get a baseline version on first replacement.
delete_entry purges its history (dedup-aware); delete_space pages its
rows and purges every object; consolidate never collects history.
- MEMORY.md is reserved in the service on every path.
- API: PUT entries returns version/tokens/warnings/anchors; new
GET /memory/spaces/{id}/history?slug= and /history/{n}?slug=.
- memory_write records reason "save" and no longer sends an empty
description (tool spec text unchanged, so toolConfig is unchanged).
- app-api task role gains bedrock:CountTokens.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(observability): keep conversation content out of runtime OTEL logs
Under AgentCore, ADOT writes every prompt and reply into the runtime log
group's otel-rt-logs stream from two scopes:
- strands.telemetry.tracer: Strands puts messages on span events and
ADOT's LLOHandler re-emits them as log records. Setting
OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_unredacted_attributes= turns on
Strands' redaction with nothing exempt.
- botocore bedrock-runtime: ADOT defaults
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT to true with
setdefault, so the image sets it to false first.
Both are image ENV rather than Runtime variables, because the Runtime's
50-variable cap is nearly spent. A Runtime variable still overrides them.
The new test drives a Strands turn through the real tracer and
LLOHandler, and the botocore event builders, against the Dockerfile's
values. Each case has a capture-on control.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Wire PLATFORM_SELF_SERVICE_ENABLED flag into infra (dev-only, opt-in) (#1314)
Adds the CDK plumbing so the platform self-service flag can be turned on per deployed environment (it was code-only before, readable from os.environ but never set on the runtime).
- config.ts: PlatformSelfServiceConfig + opt-in parse (CDK_PLATFORM_SELF_SERVICE_ENABLED, only literal 'true' enables; mirrors projects).
- inference-agentcore-construct.ts: sets PLATFORM_SELF_SERVICE_ENABLED on the AgentCore Runtime (read only there). 47->48 o…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reconciles
main→developafter the 1.25.1 release squash-merge (#1363).Brings
VERSION, the manifests and lockfiles at 1.25.1, and the 1.25.1CHANGELOG.md/RELEASE_NOTES.mdentries into develop. The merge was clean: develop had nothing after #1362, andmainhas been an ancestor since the 1.25.0 backmerge (#1360). The resulting tree is identical toorigin/main,git merge-base --is-ancestor origin/main HEADpasses, andsync-version.sh --checkpasses.🤖 Generated with Claude Code