release: Memmy v1.1.6 - #441
Merged
Merged
Conversation
…gent-wy # Conflicts: # App/shell/desktop/src/main/windows-launch-at-login.ts # scripts/internal/linux/build-cli-archive.sh
fix(frontend): refine sidebar alignment and modal interaction
Synthetic checkpoint for safe three-way migration onto upstream computer_use; source worktree remains untouched.
Include recording, summarization, workflow extraction, replay scripts, schema, and demo fixtures.
Feature/agent wy
…dinates The recorder resolved every click by hit-testing the cursor position and then sleeping 300ms hoping the application had rebuilt its accessibility tree. That races the renderer, and its own comment admitted Chrome often still returned stale geometry on the retry. Keep the focused element continuously up to date from AXObserver notifications so a click can be attributed immediately; the positional hit test is now only a fallback for controls that never take focus. Coordinates are no longer emitted. Event kinds now match the shape Codex/Skysight produces, which the diff engine and the layered summaries will build on: selection.changed new, and the highest-volume semantic signal keyboard.submit new, a cheap and reliable task-boundary marker mouse.drag new, carrying origin and destination elements mouse.context_menu new scroll removed; the AX tree diff carries that state instead Also add app.secureInput from IsSecureEventInputEnabled() so keystroke text is suppressed while a password field owns focus, rather than relying only on the redaction regexes as a second line of defence. The history JSONL output shape is unchanged so summarize-history and its fixtures keep working. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Codex's equivalent tool has no replay: it answers questions about what the user did, and reproducing a behavior is Computer Use's job. Splitting the two the same way stops a retrieval result from turning into desktop control on its own, and matches the conclusion that Computer History is searchable evidence rather than a library of replayable templates. Drop the `action` enum and the replay branch, add an explicit result limit, and label every returned field as untrusted observed evidence — the event stream records whatever appeared on screen, including text written by third parties. The service keeps prepareReplayUserRequest for the desktop UI; only the agent-facing surface loses it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ools Always-on capture needs a policy before it needs a daemon. This adds the model Codex uses: two orthogonal axes rather than one list, so "record everything except my bank" and "record nothing except my IDE" are both expressible. defaultApplicationBehavior applications matching no app rule, by bundle id defaultURLBehavior websites matching no URL rule, by bare domain Allow and block rules coexist; a block rule always wins inside its own axis, so an allow entry can never re-enable something the user excluded. A browser record with a usable URL must pass both axes, a record without one is judged by its application alone, and private browsing is excluded unconditionally. The default is do-not-observe: a fresh install records only what the user has explicitly allowed. Adds computer_history_status, _get_settings and _update_settings. Updates replace the whole document rather than merging, so the update tool says so and tells the agent to read first — that is what stops it from silently dropping rules the user never mentioned. The recorder still takes its allowlist from CLI flags; wiring it to this policy lands with the resident daemon. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every event now carries the state of the focused window, which is only affordable if consecutive events do not repeat the whole tree. Keep the previous snapshot per window and emit just the added and removed nodes, falling back to the full tree when there is nothing to diff against or when the change is large enough that a diff would not be smaller. Measured on a live session: the first snapshot is 6775 characters, the next is a 44-character diff. Snapshots are throttled to at most one every 400ms and bounded at 400 nodes, so a busy window cannot turn the recorder into a tree-walking loop. The state rides through to the history JSONL for the layered summaries to consume. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Observation is now continuous rather than a set of named recordings the user starts by hand. A recorder runs in the background and its output is sliced into ten-minute segments, each a directory holding events.jsonl and metadata.json, with segment ids aligned to the ten-minute grid so they sort and group cleanly. The state machine is three-valued, not two: running capturing according to the observation settings paused current segment kept, nothing new written to it stopped recording nothing; completed segments stay searchable Pause is a first-class state rather than a weaker stop, because "stop watching while I do something private" must not cost the user the arc they were in the middle of. The desktop toggle gains a matching third state. Recording is scoped to the app: the server's stop path finalizes the open segment, so closing the app ends observation instead of leaving a recorder running behind the user's back. Segments carry no title or starting URL — those described a single recording, and a continuous stream has neither. Also fixes two defects in the preceding commits: appendEvent destructured its argument and silently dropped the ax field, and flushKeys still read the pre-AXObserver application shape. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Looking back over a working day should not cost a re-read of the whole event stream. A six-hour summary is now built from the ten-minute summaries covering its window, so the cost of the wider view scales with the number of summaries rather than the size of the stream. That is also why the six-hour file cites the ten-minute files it reused instead of the segments beneath them: a reader follows one level down, not all the way down. Applications and context lines are merged without repetition, and a six-hour file is never folded into another six-hour file. Segment summaries are now named `<id>-10min-summary.md` so the two layers are distinguishable on disk and the rollup can find its inputs. The rollup runs whenever a segment is finalized, which is cheap precisely because it reuses what is already summarized. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ion policy Making observation resident dropped the hardcoded six-app allowlist that the old named-recording path passed, and an empty allowlist means "no filter" in the consumer. The recorder would have observed every application, with no policy in front of it — the exact combination the settings model exists to prevent. The recorder now takes --observation-settings and evaluates the policy per event, which it has to do itself because the website axis depends on the URL each event carries. Starting observation is refused outright when the policy would record nothing, so a fresh install says what is missing instead of looking like it is recording while producing empty segments. The rule table is asserted on both sides, in observation-settings.test.ts and in record-human-history.test.mjs, so the capture path and the agent tools cannot drift apart on what a policy means. Also lets the service take an observation settings path, so tests stop writing the user's real policy file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The timeline was showing another product's records as if they were Memmy's own, from two directions: Codex's live Skysight directory was read on every snapshot, and 155 copies of its summaries had been dropped into Memmy's own history directory during earlier experiments, where nothing distinguished them from a real capture. Codex writes `<utc>-<4 random chars>-10min-memory-summary.md`; Memmy names its segments `<segment id>-10min-summary.md`, with no random component and no "memory-", so the copies can be recognized by name and left out. They stay on disk — this hides them, it does not delete anything. Reading Codex's live directory is now opt-in rather than the default, through a constructor flag or MEMMY_COMPUTER_HISTORY_CODEX_SYNC=1. The two tests covering that path opt in explicitly instead of relying on the old default. On the current machine this takes the timeline from 163 entries to the 8 Memmy actually captured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reading Codex's Skysight directory made another product's records show up as Memmy's own. In a real deployment the timeline should come from Memmy's own capture, so the sync is gone rather than merely defaulted off: the codex_synced source type, the directory reader, the opt-in flag, and the sync-state file that existed only to hide synced entries from the timeline. Trimming codex_synced also simplifies the paths that branched on it — replay plans, workflow creation and deletion now have one fewer source to reason about, and deleteHistory no longer has a branch that hides instead of deletes. isCodexSkysightCopy stays: copies of Codex summaries are still sitting in the history directory from earlier experiments and must not be shown as captures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The timeline showed a title and a timestamp, and the pane beside it dumped the raw markdown including its YAML frontmatter. Worse, since observation became resident there is no user-supplied title, so every entry read "Computer History <segment id>" above the same templated sentence. Each finished summary is now narrated by the model: a specific title and two or three sentences addressed to the user, written from the segment's own evidence. The call is fire-and-forget after the mechanical summary is already on disk, so an unreachable model costs the better wording and never the recording, and the prompt states that the recorded screen content is evidence rather than instructions. The timeline groups by day and each entry carries its own account. A day that has a six-hour rollup is headed by it as a collapsible overview rather than letting it compete with the ten-minute entries. The detail pane renders time, title, description and applications as a header and drops the frontmatter from the body. Entries now carry description, applications and summaryWindow so the timeline does not have to parse markdown to render itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding description, applications and summaryWindow to the shared markdown reader put them on workflows too, because the workflow list was built by spreading the directory entry. The desktop client parses the snapshot with a strict schema, so every workflow failed validation and the whole page rendered as a wall of unrecognized_keys errors instead of the timeline. Build workflows by naming their fields. A workflow is not a summary, and spreading meant the next field added to the reader would leak the same way. Nothing caught this: the backend never asserted the snapshot's shape, and the desktop test mocks the client, so it never runs the schema. Both gaps are now covered — the workflow test fails if the spread comes back. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
feat: integrate Computer History and Computer Use
…owlist The default was do-not-observe, which meant a fresh install recorded nothing until the user named applications one by one. That fails silently in the worst way available: the UI reads "recording" while nothing is written, and the gap surfaces days later when the history is asked for and turns out to be empty. Codex, measured on this machine, recorded 27 distinct applications including Dock, Finder and System Settings — it observes by default and uses the blocklist for exceptions. Match that. What protects the user here was never the direction of this default: it is the blocklist, pause, the unconditional private-browsing exclusion, secure-input suppression, no screenshots, local-only storage and the retention window. All of those hold either way. The service now writes the policy out on first start, so the recorder parses one explicit document instead of inferring a policy from a missing file, and the user has something to edit. A file that exists but does not parse stops the start rather than falling back to a policy they did not choose. No default blocklist ships: which applications would go on it is a judgement better left to the user. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Moving recordings under `segments/` left the cleanup scanning one level too high, so it judged the container by its own mtime. That broke retention in both directions: while recording continued the container stayed fresh and nothing inside was ever removed, and once recording stopped for long enough the whole container expired at once, taking every segment with it. In practice it meant the 48-hour window silently never applied and raw event streams accumulated without limit. Judge each segment directory on its own age, never the container, and leave the open segment alone because it is still being written to. Recordings captured before segments existed still sit at the top level and are still expired there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
….1.6 fix(release): integrate Computer History and Computer Use updates for v1.1.6
Co-authored-by: Cursor <cursoragent@cursor.com>
…ent-hash fix(memory): dedupe agent skills by content hash
…ase-v1.1.6 Backport Computer Use permission and Windows fixes to v1.1.6
…-release-v1.1.6 Fix Computer History action tooltips for v1.1.6
…n/release-v1.1.6-pr439-isolated
…-pr439-isolated fix: stabilize v1.1.6 release regression checks
…-release-v1.1.6 Backport native npm Computer Use and macOS permission onboarding to v1.1.6
…-pr439-isolated fix: cover PR494 computer use contracts and tests
syzsunshine219
marked this pull request as ready for review
September 20, 2026 12:29
syzsunshine219
had a problem deploying
to
release
September 20, 2026 12:29 — with
GitHub Actions
Failure
syzsunshine219
had a problem deploying
to
release
September 20, 2026 12:36 — with
GitHub Actions
Failure
syzsunshine219
had a problem deploying
to
release
September 21, 2026 00:59 — with
GitHub Actions
Failure
syzsunshine219
had a problem deploying
to
release
September 21, 2026 02:04 — with
GitHub Actions
Failure
syzsunshine219
had a problem deploying
to
release
September 21, 2026 02:50 — with
GitHub Actions
Failure
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release candidate
Promote the main-repository
release/v1.1.6branch intomain, following the established release PR flow used by PR #382.This release branch contains:
Validation
npm run version:sync -- --checkpassed for 1.1.6.Please run the main-repository Harness/CI and release preflight before merging.