Skip to content

release: Memmy v1.1.6 - #441

Merged
syzsunshine219 merged 241 commits into
mainfrom
release/v1.1.6
Sep 20, 2026
Merged

syzsunshine219 merged 241 commits into
mainfrom
release/v1.1.6

Conversation

@syzsunshine219

Copy link
Copy Markdown
Collaborator

Release candidate

Promote the main-repository release/v1.1.6 branch into main, following the established release PR flow used by PR #382.

This release branch contains:

Validation

  • Post-merge PR integration/release v1.1.6 cu final #439 Full regression passed on the merged release candidate.
  • npm run version:sync -- --check passed for 1.1.6.
  • Targeted backend project-version test passed.
  • Packaging/install testing remains a release-owner step after this PR and its checks pass.

Please run the main-repository Harness/CI and release preflight before merging.

Wang-Daoji and others added 30 commits September 1, 2026 17:55
…gent-wy

# Conflicts:
#	App/shell/desktop/src/main/windows-launch-at-login.ts
#	scripts/internal/linux/build-cli-archive.sh
fix(frontend): refine sidebar alignment and modal interaction
Synthetic checkpoint for safe three-way migration onto upstream computer_use; source worktree remains untouched.
Include recording, summarization, workflow extraction, replay scripts, schema, and demo fixtures.
…dinates

The recorder resolved every click by hit-testing the cursor position and then
sleeping 300ms hoping the application had rebuilt its accessibility tree. That
races the renderer, and its own comment admitted Chrome often still returned
stale geometry on the retry.

Keep the focused element continuously up to date from AXObserver notifications
so a click can be attributed immediately; the positional hit test is now only a
fallback for controls that never take focus. Coordinates are no longer emitted.

Event kinds now match the shape Codex/Skysight produces, which the diff engine
and the layered summaries will build on:

  selection.changed  new, and the highest-volume semantic signal
  keyboard.submit    new, a cheap and reliable task-boundary marker
  mouse.drag         new, carrying origin and destination elements
  mouse.context_menu new
  scroll             removed; the AX tree diff carries that state instead

Also add app.secureInput from IsSecureEventInputEnabled() so keystroke text is
suppressed while a password field owns focus, rather than relying only on the
redaction regexes as a second line of defence.

The history JSONL output shape is unchanged so summarize-history and its
fixtures keep working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Codex's equivalent tool has no replay: it answers questions about what the user
did, and reproducing a behavior is Computer Use's job. Splitting the two the
same way stops a retrieval result from turning into desktop control on its own,
and matches the conclusion that Computer History is searchable evidence rather
than a library of replayable templates.

Drop the `action` enum and the replay branch, add an explicit result limit, and
label every returned field as untrusted observed evidence — the event stream
records whatever appeared on screen, including text written by third parties.

The service keeps prepareReplayUserRequest for the desktop UI; only the
agent-facing surface loses it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ools

Always-on capture needs a policy before it needs a daemon. This adds the model
Codex uses: two orthogonal axes rather than one list, so "record everything
except my bank" and "record nothing except my IDE" are both expressible.

  defaultApplicationBehavior  applications matching no app rule, by bundle id
  defaultURLBehavior          websites matching no URL rule, by bare domain

Allow and block rules coexist; a block rule always wins inside its own axis, so
an allow entry can never re-enable something the user excluded. A browser
record with a usable URL must pass both axes, a record without one is judged by
its application alone, and private browsing is excluded unconditionally.

The default is do-not-observe: a fresh install records only what the user has
explicitly allowed.

Adds computer_history_status, _get_settings and _update_settings. Updates
replace the whole document rather than merging, so the update tool says so and
tells the agent to read first — that is what stops it from silently dropping
rules the user never mentioned.

The recorder still takes its allowlist from CLI flags; wiring it to this policy
lands with the resident daemon.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every event now carries the state of the focused window, which is only
affordable if consecutive events do not repeat the whole tree. Keep the
previous snapshot per window and emit just the added and removed nodes,
falling back to the full tree when there is nothing to diff against or when
the change is large enough that a diff would not be smaller.

Measured on a live session: the first snapshot is 6775 characters, the next
is a 44-character diff.

Snapshots are throttled to at most one every 400ms and bounded at 400 nodes,
so a busy window cannot turn the recorder into a tree-walking loop. The state
rides through to the history JSONL for the layered summaries to consume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Observation is now continuous rather than a set of named recordings the user
starts by hand. A recorder runs in the background and its output is sliced into
ten-minute segments, each a directory holding events.jsonl and metadata.json,
with segment ids aligned to the ten-minute grid so they sort and group cleanly.

The state machine is three-valued, not two:

  running   capturing according to the observation settings
  paused    current segment kept, nothing new written to it
  stopped   recording nothing; completed segments stay searchable

Pause is a first-class state rather than a weaker stop, because "stop watching
while I do something private" must not cost the user the arc they were in the
middle of. The desktop toggle gains a matching third state.

Recording is scoped to the app: the server's stop path finalizes the open
segment, so closing the app ends observation instead of leaving a recorder
running behind the user's back.

Segments carry no title or starting URL — those described a single recording,
and a continuous stream has neither.

Also fixes two defects in the preceding commits: appendEvent destructured its
argument and silently dropped the ax field, and flushKeys still read the
pre-AXObserver application shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Looking back over a working day should not cost a re-read of the whole event
stream. A six-hour summary is now built from the ten-minute summaries covering
its window, so the cost of the wider view scales with the number of summaries
rather than the size of the stream.

That is also why the six-hour file cites the ten-minute files it reused instead
of the segments beneath them: a reader follows one level down, not all the way
down. Applications and context lines are merged without repetition, and a
six-hour file is never folded into another six-hour file.

Segment summaries are now named `<id>-10min-summary.md` so the two layers are
distinguishable on disk and the rollup can find its inputs. The rollup runs
whenever a segment is finalized, which is cheap precisely because it reuses
what is already summarized.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ion policy

Making observation resident dropped the hardcoded six-app allowlist that the
old named-recording path passed, and an empty allowlist means "no filter" in
the consumer. The recorder would have observed every application, with no
policy in front of it — the exact combination the settings model exists to
prevent.

The recorder now takes --observation-settings and evaluates the policy per
event, which it has to do itself because the website axis depends on the URL
each event carries. Starting observation is refused outright when the policy
would record nothing, so a fresh install says what is missing instead of
looking like it is recording while producing empty segments.

The rule table is asserted on both sides, in observation-settings.test.ts and
in record-human-history.test.mjs, so the capture path and the agent tools
cannot drift apart on what a policy means.

Also lets the service take an observation settings path, so tests stop writing
the user's real policy file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The timeline was showing another product's records as if they were Memmy's own,
from two directions: Codex's live Skysight directory was read on every snapshot,
and 155 copies of its summaries had been dropped into Memmy's own history
directory during earlier experiments, where nothing distinguished them from a
real capture.

Codex writes `<utc>-<4 random chars>-10min-memory-summary.md`; Memmy names its
segments `<segment id>-10min-summary.md`, with no random component and no
"memory-", so the copies can be recognized by name and left out. They stay on
disk — this hides them, it does not delete anything.

Reading Codex's live directory is now opt-in rather than the default, through a
constructor flag or MEMMY_COMPUTER_HISTORY_CODEX_SYNC=1. The two tests covering
that path opt in explicitly instead of relying on the old default.

On the current machine this takes the timeline from 163 entries to the 8 Memmy
actually captured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reading Codex's Skysight directory made another product's records show up as
Memmy's own. In a real deployment the timeline should come from Memmy's own
capture, so the sync is gone rather than merely defaulted off: the
codex_synced source type, the directory reader, the opt-in flag, and the
sync-state file that existed only to hide synced entries from the timeline.

Trimming codex_synced also simplifies the paths that branched on it — replay
plans, workflow creation and deletion now have one fewer source to reason
about, and deleteHistory no longer has a branch that hides instead of deletes.

isCodexSkysightCopy stays: copies of Codex summaries are still sitting in the
history directory from earlier experiments and must not be shown as captures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The timeline showed a title and a timestamp, and the pane beside it dumped the
raw markdown including its YAML frontmatter. Worse, since observation became
resident there is no user-supplied title, so every entry read "Computer History
<segment id>" above the same templated sentence.

Each finished summary is now narrated by the model: a specific title and two or
three sentences addressed to the user, written from the segment's own evidence.
The call is fire-and-forget after the mechanical summary is already on disk, so
an unreachable model costs the better wording and never the recording, and the
prompt states that the recorded screen content is evidence rather than
instructions.

The timeline groups by day and each entry carries its own account. A day that
has a six-hour rollup is headed by it as a collapsible overview rather than
letting it compete with the ten-minute entries. The detail pane renders time,
title, description and applications as a header and drops the frontmatter from
the body.

Entries now carry description, applications and summaryWindow so the timeline
does not have to parse markdown to render itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding description, applications and summaryWindow to the shared markdown
reader put them on workflows too, because the workflow list was built by
spreading the directory entry. The desktop client parses the snapshot with a
strict schema, so every workflow failed validation and the whole page rendered
as a wall of unrecognized_keys errors instead of the timeline.

Build workflows by naming their fields. A workflow is not a summary, and
spreading meant the next field added to the reader would leak the same way.

Nothing caught this: the backend never asserted the snapshot's shape, and the
desktop test mocks the client, so it never runs the schema. Both gaps are now
covered — the workflow test fails if the spread comes back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
feat: integrate Computer History and Computer Use
…owlist

The default was do-not-observe, which meant a fresh install recorded nothing
until the user named applications one by one. That fails silently in the worst
way available: the UI reads "recording" while nothing is written, and the gap
surfaces days later when the history is asked for and turns out to be empty.

Codex, measured on this machine, recorded 27 distinct applications including
Dock, Finder and System Settings — it observes by default and uses the
blocklist for exceptions. Match that.

What protects the user here was never the direction of this default: it is the
blocklist, pause, the unconditional private-browsing exclusion, secure-input
suppression, no screenshots, local-only storage and the retention window. All
of those hold either way.

The service now writes the policy out on first start, so the recorder parses
one explicit document instead of inferring a policy from a missing file, and
the user has something to edit. A file that exists but does not parse stops the
start rather than falling back to a policy they did not choose.

No default blocklist ships: which applications would go on it is a judgement
better left to the user.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Moving recordings under `segments/` left the cleanup scanning one level too
high, so it judged the container by its own mtime. That broke retention in both
directions: while recording continued the container stayed fresh and nothing
inside was ever removed, and once recording stopped for long enough the whole
container expired at once, taking every segment with it. In practice it meant
the 48-hour window silently never applied and raw event streams accumulated
without limit.

Judge each segment directory on its own age, never the container, and leave the
open segment alone because it is still being written to. Recordings captured
before segments existed still sit at the top level and are still expired there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
memory-lee and others added 22 commits September 17, 2026 10:47
….1.6

fix(release): integrate Computer History and Computer Use updates for v1.1.6
Co-authored-by: Cursor <cursoragent@cursor.com>
…ent-hash

fix(memory): dedupe agent skills by content hash
…ase-v1.1.6

Backport Computer Use permission and Windows fixes to v1.1.6
…-release-v1.1.6

Fix Computer History action tooltips for v1.1.6
…-pr439-isolated

fix: stabilize v1.1.6 release regression checks
…-release-v1.1.6

Backport native npm Computer Use and macOS permission onboarding to v1.1.6
…-pr439-isolated

fix: cover PR494 computer use contracts and tests

This branch was successfully deployed

1 active deployment
release — fb2caae5 Deployed Sep 21, 2026 by syzsunshine219 via release #40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.