Skip to content

Release/1.25.1 - #1363

Merged
philmerrell merged 3133 commits into
mainfrom
release/1.25.1
Sep 26, 2026
Merged

philmerrell merged 3133 commits into
mainfrom
release/1.25.1

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Release 1.25.1

Patch release with one SPA fix. New Session no longer reopens with the previous conversation's first message (#1362). The empty-state composer was destroyed before its draft effect could clear the saved composer-draft:new entry.

Keeping 1.25.0 visible

The Release workflow publishes only this version's section as the GitHub release body, and v1.25.1 becomes "Latest". So the 1.25.1 notes open with a callout that links v1.25.0 and names what it shipped. The callout also tells anyone upgrading from 1.24.x to follow the 1.25.0 Deployment notes (CDK deploy, KB byte-counter repair, cleanups), because 1.25.1 adds no steps of its own. The 1.25.0 entry is untouched. I ran the workflow's extraction awk locally, and the callout is the first thing on the release page.

Gates

  • GSI update limit: PASS (no infrastructure change)
  • Pending backfills: PASS (none)
  • Version sync: PASS (1.25.1)

Deploy

  • From 1.25.0: no special steps; only the SPA changes.
  • After this squash-merges, backmerge main → develop with a merge commit.

🤖 Generated with Claude Code

philmerrell and others added 30 commits September 23, 2026 15:04
…s-table-infra

feat(infra): provision the Shared Projects table and flag (PR-1.1)
…ests

ci: run restore-data script tests on every PR
…dead-oauth-env-app-api

chore(infra): retire dead OAuth env vars from app-api
…ackend-core

# Conflicts:
#	tests/supply_chain/test_env_var_contract.py
…s-backend-core

feat(projects): backend core — projects, members, and a project-owned harness agent (PR-1.2)
…esses

shared-projects §9.6. A project can have 200 members; blocking the whole
turn because one member lacks one bound tool makes the project unusable to
them. resolve_agent_invocation(..., degrade=True) drops a missing
capability instead of raising, and records it in plan.unavailable:

- tools: a denied SERVER drops every ref of it (the gate is per base id,
  so skipping only the checked ref would let a later scoped ref through);
  an all-dropped toolset is still an empty toolset, never a fall-through
  to the request's enabled_tools
- model: no override, so the route's chain picks the member's default
- skills: filtered; none left means no skill binding
- memory space: skipped

AgentNoticeEvent (SSE agent_notice) serializes the drops. Ordinary shared
agents keep block-with-message (D5) unchanged — degrade defaults off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t (PR-1.4a)

shared-projects PR-1.4a. Membership already gates the harness (1.2's access
delegation), so the route adds what a project turn needs on top:

- Archived projects refuse new turns with a conversational error naming
  the project.
- The resolver runs with degrade=True for a harness; anything dropped is
  streamed as agent_notice before message_start (SSE only, never the
  prompt, not persisted). Documented in CLAUDE.md's event table.
- "## Project Instructions" replaces "## Assistant-Specific Instructions"
  for a harness, via compose_agent_system_prompt, which takes no user
  argument, so every member of a project renders a byte-identical prefix.
  Every other agent's text is unchanged byte for byte.
- preferences.projectId is written at session binding.
- projectId rides every C# cost row next to turnAgentId (route →
  ChatAgent → StreamCoordinator), and the metadata writer adds the call to
  PROJECT#{id}/COST#{YYYY-MM} plus COST#{YYYY-MM}#USER#{userId}: one
  atomic ADD each, both UpdateItem (the runtime has no PutItem). The spec's
  byUser map became per-member rows. Best-effort like every aggregate.
- kb-sync already copies apis/shared/projects; the scheduled-runs image now
  does too (plus dynamo_errors.py) because sessions/metadata.py reaches the
  repository lazily. Import closure only.

Personal instructions (1.4b) are split out: they change every user's
system prompt, not only project turns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s-invocation

feat(projects): run a project's harness in chat, with degrade-with-notice and per-project cost (PR-1.4a)
Adds the clickable prototype link (claude.ai/artifact/2syz1g6zY2ZsCAkgsWVp7F)
with its two deliberate departures from §6, the open questions it raises,
and the convention that a PR changing a behavior the mockup shows notes it
in its as-built entry for the 1.8 re-sync. Backfills those notes for the
two merged PRs that already made such decisions (1.2 delete/transfer/
invite shape, 1.4a agent_notice).

Requested from the "Projects UI mockups" session on Phil's behalf.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
shared-projects §3.4. A member's own tasks in one Shared Project, newest
first: GSI5_PK = PROJECT#{projectId}#USER#{userId}, GSI5_SK =
{lastMessageAt}#{sessionId}, projection ALL. The recency key of
SessionRecencyIndex (GSI4), sparse the same way, so the 1.6 backend
writes and removes it wherever GSI4 is.

Lands alone and ahead of any writer: it is the one GSI this plan adds to
sessions-metadata (one GSI per UpdateTable), and it is inert until rows
carry GSI5 keys. gsi-inventory.json regenerated; tables-detailed pins the
key shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s-session-index

feat(infra): ProjectSessionIndex on sessions-metadata (PR-1.6-infra)
A user-reported 404 on prod (2026-09-24) could not be traced: app-api and
the ALB had no matching request, CloudFront's 4xx metric showed errors in
the window, and logging was disabled on the distribution, so nothing could
name the URI or client. Requests the edge answers itself (a lazy chunk a
deploy removed, an S3 403) never reach the ALB.

- Standard logging (legacy) to a dedicated S3 bucket under `spa/`. v2 is
  configured via CloudWatch vended-log deliveries that must be created in
  us-east-1, which the single app-region stack cannot do.
- Bucket: ACLs enabled (BUCKET_OWNER_PREFERRED; legacy delivery writes
  through the ACL and fails against BucketOwnerEnforced), SSE-S3, block
  public access, SSL-only, 90-day expiry.
- Cookies are never logged (the SPA session is an httpOnly cookie).
- CDK_FRONTEND_ACCESS_LOGS_ENABLED / frontend.accessLogsEnabled: default
  ON with a kill switch. The bucket is provisioned unconditionally so a
  disable/re-enable never collides with a RETAINed bucket name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Write ProjectSessionIndex (GSI5) keys beside GSI4 for sessions with
  preferences.projectId: on store, on activity, and removed on soft-delete.
  Reads now strip all recency keys, so a read-modify-write no longer replays
  stale GSI4_* extras (which could SET and REMOVE the same attribute).
- GET /projects/{id}/tasks: the caller's own sessions in the project, newest
  first, value-cursor paginated; a missing index degrades to empty.
- accessLevel "project" on shares: requires the task's projectId, the caller's
  membership and an active project; the read check is membership.
- SHARED_TASK#{sessionId} pointer on the projects table, rebuilt from the
  share rows on create/revoke/update/session delete (newest project share
  wins); GET /projects/{id}/shared-tasks lists them without user ids.
- A fork keeps its project (projectId + the project's current harness) when
  the requester can work in it; otherwise it stays a plain session.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
deploy.sh syncs with --delete, so a tab opened before a deploy asks for
lazy chunks that no longer exist the first time it navigates to a route it
has not loaded (reported on prod 2026-09-24: Settings -> API Keys showed an
error; no request ever reached app-api).

- isChunkLoadError matches the Chromium, Safari, Firefox and webpack
  wording (plus HTML-served-as-module), following cause/rejection/error.
- A failed navigation becomes a full load of its destination, guarded to
  one auto-reload per tab per 5 minutes (sessionStorage, fails closed).
  It is skipped, in favour of a "new version available - Refresh" toast,
  when a chat stream or upload is in flight, since in-app navigation keeps
  those alive and a page load would not.
- Chunk failures outside navigation (@defer, lazy libraries) reach a
  custom ErrorHandler and only prompt.
- Proactive check (environment.versionCheckEnabled): on tab focus, at most
  every 10 minutes, compare the running main-<hash>.js with the one the
  no-cache index.html names. Inert on the dev server.
- ToastService gains an optional action button.

Composer drafts need no change: the composer mirrors to localStorage on
every edit, so there is nothing to flush before the reload.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s-tasks

feat(projects): project tasks, project shares and forks (PR-1.6 backend)
…plicate

Some emails own more than one PROFILE row: the pre-Cognito login keyed
users by a numeric employee ID, the current one by the Cognito sub, and
nothing retired the old rows. get_user_by_email queried EmailIndex (no
sort key) with Limit=1, so which row it returned was arbitrary.

- UserRepository.get_users_by_email pages through every match and ranks
  them: most recent lastLoginAt (compared as instants), then non-numeric
  ids over legacy numeric ones, then id for determinism.
- get_user_by_email returns the first and logs a warning naming the
  ignored duplicates.
- Admin email search returns every profile, live first, so support sees
  the ambiguity instead of landing on the stale row.
- The sharing user search dedupes by email, not id; previously a stale
  exact match plus the live row from the name scan listed one person twice.
- upsert_user logs when a newly created profile's email already belongs
  to another id (creation only; returning users pay nothing).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When an email search matches several profiles, show how many, badge the
one that signs in, and print each row's id, so a quota override or tier
is not assigned to a legacy duplicate nobody can log in to.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…le-chunk-recovery

fix(spa): recover from stale lazy chunks after a frontend deploy
…two phases

backend/scripts/audit_user_duplicates.py groups users-table PROFILE rows
by email, ranks each group with the same live_profile_rank the API uses,
and checks every stale id against each place a user id is stored: one
Query per indexed key path across 20 tables, an S3 prefix probe per
user-keyed bucket, AgentCore Memory sessions for the actor, and, with
--deep, a single Scan per table with un-indexed references. It writes a
JSON and a Markdown report.

Read-only by default. --apply mark sets an unreferenced legacy row to
status=inactive with mergedInto/mergedAt (reversible; status=merged
would not parse). --apply delete removes only rows marked at least
--min-soak-days ago for today's live id. Only numeric-id + uuid pairs
whose every check came back zero are eligible; a failed check makes a
row incomplete, not unreferenced. Every write is conditional on the
lastLoginAt the audit saw, and --apply needs --confirm-prefix.

A coverage test fails when a CDK table is neither checked nor listed as
holding no user ids. live_profile_rank and item_to_profile become public
module functions in the users repository so the script cannot drift
from the API's choice of live row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…avigation

Found validating #1262 in dev: a navigation that fails on a stale chunk
while a stream is in flight shows the Refresh prompt aimed at the failed
destination. If the user then opened another view (e.g. went back to
the conversation) and clicked Refresh, they were sent to the view that
failed minutes earlier instead of reloading where they were.

AppUpdateService now clears the remembered destination on NavigationEnd.
A failed or cancelled navigation doesn't end in NavigationEnd, so the
failure that set it, or a guard refusing the next one, leaves it in place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…udfront-access-logs

feat(infra): enable CloudFront access logs on the SPA distribution
- apis/shared/directory: DirectoryAdapter port, DIRECTORY_PROVIDER switch
  (default users_table; unknown values fall back), and UsersTableDirectory,
  which pages the whole active partition of StatusLoginIndex instead of the
  100 most recent sign-ins, matching email prefix and name in one pass,
  ranked by match quality then recency. The active list is held 60s per
  process so a typeahead filters in memory per keystroke.
- GET /projects/{id}/directory?q=&limit= (viewer): email, name,
  hasSignedIn and the person's current memberRole, no user ids; a
  well-formed unknown email is appended so it can always be invited.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…resh-target-reset

fix(spa): reset the refresh prompt's destination after a successful navigation
lastLoginAt cannot show that an old id is idle: an API key minted under
it still authenticates as that id, the api-converse path never updates
the profile's lastLoginAt, and it 401s a key whose profile row is gone.
So deleting such a row would break a live integration outright.

The audit now reads each stale id's API keys and sessions-metadata
partition and marks the row in_use when it owns an unexpired key or an
active scheduled prompt, or when any model call, message or key use is
newer than the day its live twin was created (that person's cutover).
in_use rows are never retired and lead the Markdown report; every stale
row also reports its newest activity so "history only" is visible.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 14px core with a 26px halo outweighed the status line beside it.
Scale it to an 11px core with a 21px halo (inset -5px) and a 10px glow,
keeping the same motion, proportions and timing. The finished-turn recap
dot follows so the live-to-settled transition still reads as one element.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s-directory

feat(projects): people directory for inviting members (PR-1.3)
The first read-only dev run reported "deep:app-roles 2" for a single
TOOL_PREFERENCES row that names the id in both PK and userId. Count one
hit per item and list the matching attributes in the sample instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- GET/PUT /projects/{id}/instructions, /model, /tools, /skills: viewer
  reads, editor writes on an active project. Tools/skills replace only
  their own kind. Only what a save adds is validated against the saver
  (validate_agent_write); a no-op save cuts nothing.
- Every change cuts an AgentVersion (createdBy + createdByEmail); the first
  save also records the starting state. GET .../instructions/versions and
  .../versions/{n} return history and a diff against the previous version.
- The project is the harness's only write path: PUT /assistants and
  PUT /agents return 409 on a harness.
- An archived project's harness resolves every member as viewer, so no
  agent-level write route (documents, web sources, sync) can change it.
- Move diff wire helpers into version_diff; label-able instructions_diff;
  public ProjectService.authorize.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tate-orb-sizing

feat(spa): shrink the agent-state orb by 20%
philmerrell and others added 26 commits September 25, 2026 18:27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
base_foundation_model_id only knew us./eu./apac./us-gov., so a global.*
profile id reached CountTokens unchanged. Bedrock rejected it as
unsupported, the model landed on Strands' skip list, and every count for
every prod global.* model fell back to the chars/4 heuristic, which also
means no prefixTokens / contextBreakdown after #1337.

Adds global plus the geo prefixes the AWS model cards document today
(au., jp.). Undocumented codes stay untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rocess-tracking-da4ea8

docs(kaizen): track harness-sdk#4618 as a Strands pin-bump hazard
…-4b-dev-validation

docs(projects): record 2.4b dev validation
… after release

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l path

Stripping global. from the CountTokens id makes prod global.* models count
natively, and a native count awaited before a model call adds a CountTokens
round trip to time to first token: measured +519 ms on a freshly built
agent's first call and +74 ms on each later call (+56 ms per tool round, for
a count that always failed).

Nothing needs those counts before the call:
- CountTokensBedrockModel(native_projection=False), as the factory now
  builds every Converse model, keeps Strands' pre-call count_tokens on the
  local heuristic. Strands' proactive compression is off in our config and
  our compaction does not read the projection.
- native_count_tokens() is a blocking counter that answers natively or
  None, never a heuristic, with the same skip-list bookkeeping.
- ContextAttributionHook's BeforeModelCall callback is now synchronous: it
  records the projection and, while the agent has no split, starts one
  background task that makes four native counts of a snapshot concurrently
  with the model call. AfterModelCall places the messages partition
  against the call's billed prompt (input + cache read + cache write).
- Interrupted turns read get_projected_input_tokens(): the native count of
  the exact request when the split measured one, else a usage-anchored
  projection, else nothing.

Same dev harness (40 tools, 32k-token prefix, n=8 per arm): the pre-stream
window matches today's heuristic path, and the split was ready before call
1 finished in 8/8 sessions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…retrieval-score-logging

feat(memory): log per-namespace retrieval scores (no ids, no text)
…relevance-calibration

feat(memory): calibrate the relevance cut on a labelled eval set; default 0.5 → 0.4
Resolve against extract-then-compress: keep both docstring sections, both
constants and both outcome lists, and keep one _keep_head_lines (the two
copies had identical bodies). Add a test that a cut-off narrative in the
extract path is salvaged instead of falling back to extract_then_truncate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…summary-max-tokens-salvage

fix(compaction): salvage a summary generation that hits its token ceiling
…byte-cap-prod-repair

docs(kaizen): queue the production managed-KB byte-counter repair
…okens-global-cris-prefix

perf(count-tokens): native counts for global.* models, off the model-call path
COMPACTION_SUMMARY_EXTRACT_ENABLED now defaults on with =false as the kill
switch. The quality harness showed no loss against the full history and a
forced cut on dev pinned and answered every planted fact. Harness
model_relative follows the production default; the plain-compression arms
pin extraction off as baselines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A default-on flag needs no AgentCore Runtime env var, so the config,
construct entry, workflow variable and their tests go. The Runtime is back
to 48 of its 50 variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It asserts the single-call path's outcome and bytes, which the default-on
extract flag would otherwise route through extraction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion-extract-default-on

feat(compaction): turn extract-then-compress on by default
Dev can't exercise a global.* id (the SCP denies it), so the last open
question on #1343 is production: prefixTokens and contextBreakdown on
global.* Haiku 4.5 sessions, and no CountTokens span in front of a
model call. Read-only steps, blocked on the release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`_merge_agent_status` drives the agent stream by scheduling each
`__anext__` as its own task, and a task runs in a copy of the context it
was created from. Strands holds its spans attached across yields (the
invocation, each cycle, each model stream), so a span attached in one
step was detached in a later step's copy. OpenTelemetry logged that as
`Failed to detach context` — 3 per tool-less turn, 5 per one-tool turn,
~2.3k in dev over a week, starting with the first deploy of the merge.

The same split mis-parented spans: anything opened from the ambient
context after a yield — the AgentCore Memory CreateEvent calls the
session manager makes, DynamoDB reads from hooks — parented to the
request span instead of the cycle that caused it. Strands' own spans
were unaffected (explicit parent, explicit end), so nothing was lost.

Every step now runs in one Context copied once per turn, which is what a
plain `async for` gave the generator. No added latency: one copy per turn
instead of one per step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nt-tokens-prod-check

docs(kaizen): queue the production check for #1343's native token counts
…ntext-status-merge

fix(agents): run every agent-stream step in one context
…evelop-1.25.0

# Conflicts:
#	backend/scripts/compaction_quality/ask.py
#	backend/src/apis/app_api/admin/projects/routes.py
#	backend/src/apis/app_api/projects/harness_settings.py
#	backend/src/apis/app_api/projects/routes.py
…into-develop-1.25.0

chore: backmerge main into develop after 1.25.0
The first send from the empty state swaps its composer for the compact
one, which destroys it before the draft effect runs again. The
new-conversation draft kept the message just sent, so every later New
Session opened with that message already in the composer.

submitChatRequest now writes the cleared draft at once instead of
relying on the effect.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aft-sent-from-empty-state

fix(spa): forget the new-conversation draft on the first send
Patch release carrying one SPA fix: New Session no longer reopens with
the previous conversation's first message (#1362).

- Notes lead with a pointer to v1.25.0 so the feature release and its
  deployment steps stay visible to anyone upgrading from 1.24.x
- SPA only: no CDK change, no scripts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@philmerrell
philmerrell requested a review from a team September 26, 2026 14:15
@philmerrell
philmerrell merged commit d22f853 into main Sep 26, 2026
16 checks passed
@philmerrell
philmerrell deleted the release/1.25.1 branch September 26, 2026 14:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants