From 6e89e263b444f7daae679d860942ff393fa9fa98 Mon Sep 17 00:00:00 2001 From: Ned Date: Thu, 1 Oct 2026 21:02:32 -0700 Subject: [PATCH 01/10] fix: preserve durable standalone state across transfer and close failures --- docs/specs/auto-update.md | 4 +- docs/specs/auto-update.rationale.md | 4 + docs/specs/security-remote.md | 2 +- docs/specs/standalone.md | 160 ++-- docs/specs/standalone.rationale.md | 20 +- lib/src/host/remote/burrow-state-store.ts | 3 +- scripts/spec-word-budgets.json | 2 +- standalone/scripts/dev-agent-browser.test.mjs | 15 +- standalone/scripts/dev-standalone.test.mjs | 4 +- standalone/src-tauri/src/close_commit.rs | 69 ++ standalone/src-tauri/src/lib.rs | 699 +++++++++++++++--- standalone/src-tauri/src/quit_state.rs | 2 + standalone/src-tauri/src/routing.rs | 70 +- standalone/src/WorkspaceTeardownModal.tsx | 23 +- standalone/src/quit-confirm-store.ts | 24 +- standalone/src/quit.ts | 2 +- standalone/src/tauri-adapter.ts | 7 +- standalone/src/tauri-session-store.test.ts | 53 +- standalone/src/tauri-session-store.ts | 13 + standalone/src/updater.test.ts | 47 ++ standalone/src/updater.ts | 3 + standalone/src/window-close.test.ts | 197 ++++- standalone/src/window-close.ts | 65 +- standalone/src/workspace-move.test.ts | 37 +- standalone/src/workspace-move.ts | 8 +- 25 files changed, 1294 insertions(+), 239 deletions(-) create mode 100644 standalone/src-tauri/src/close_commit.rs diff --git a/docs/specs/auto-update.md b/docs/specs/auto-update.md index 3da2af03e..562d5278a 100644 --- a/docs/specs/auto-update.md +++ b/docs/specs/auto-update.md @@ -8,7 +8,9 @@ The standalone app checks for updates on launch, where the network policy allows **Must read and clear the post-install marker on launch** (§localStorage) and show its banner; a reported failure suppresses this launch's check. Otherwise wait 5 seconds, then read the network policy with `networkPolicy` over the Burrow link and, where it allows (`docs/specs/remote-network.md` → "Updates"), `check()` — no update is silent, an update raises the approval prompt; then the reminder, if due: `check-due`, recording `remindedAt`. **The reminder is re-evaluated hourly while the app runs**, reading no policy and never checking; **never over an undismissed notice, nor while the clock reads before 2026-09**, not yet set. Version-lookup and check failures are logged. **Only approval starts the background `download()`**; a failed one is logged and the prompt returns. -**Check now** — the `check-due` and `check-failed` links, and the `updates` port — shows `checking`, then `available`, `up-to-date`, or `check-failed`. **A second ask joins the check in flight. An update already approved is shown again, `downloading` or `downloaded`, instead of checked for**, which would offer it for approval twice. **Every successful check, automatic or asked for, records `checkedAt`** (§localStorage). +**Check now** — the `check-due` and `check-failed` links, and the `updates` port — shows `checking`, then `available`, `up-to-date`, or `check-failed`. **Must join a check already in flight and preserve an approved update as `downloading` or `downloaded` instead of checking again.** This applies to both manual and delayed automatic checks, including approval while network-policy lookup is pending. **Every successful check, automatic or asked for, records `checkedAt`** (§localStorage). `standalone/src/updater.test.ts` pins the approval races. + +Source of truth: `runUpdateCheck` and `approveUpdate` in `standalone/src/updater.ts`. **A self-host build never checks** (`docs/specs/relay.md` → "Relay origin"): `startUpdateCheck()` returns at once and `checkNow()` does nothing unless the webview's own baked mode, `bakedRelayMode()`, is `hosted`. diff --git a/docs/specs/auto-update.rationale.md b/docs/specs/auto-update.rationale.md index 5ae132d41..8b0a4b6a4 100644 --- a/docs/specs/auto-update.rationale.md +++ b/docs/specs/auto-update.rationale.md @@ -2,6 +2,10 @@ > Informative companion to [auto-update.md](auto-update.md): evidence and design history keyed by that spec's headings. Nothing here is normative. +## How it works + +On Windows, 2026-10-01, a delayed launch check ran after a manual check had already approved and downloaded the update; it called `check()` again and replaced the downloaded notice with another approval prompt. The same race occurred while network-policy lookup was pending. Checking approved-update ownership after that await preserves the approval throughout both automatic and manual paths. + ## Quit-time install **Why install runs last.** On Windows `install()` starts NSIS and then calls `std::process::exit` itself (`tauri-plugin-updater-2.11.0/src/updater.rs`, checked 2026-09), so starting it early interrupts teardown. This ordering originally protected persisted scrollback; what it protects now is the window's structure, which standalone does persist. The retained save/drain hooks and their completion semantics are explained in `docs/specs/standalone.rationale.md` → Quit flow. diff --git a/docs/specs/security-remote.md b/docs/specs/security-remote.md index abb14b771..852fa6ee1 100644 --- a/docs/specs/security-remote.md +++ b/docs/specs/security-remote.md @@ -111,7 +111,7 @@ per-Burrow browser storage follows `docs/specs/remote-security-model.md` -> - **FAIL IF** `relay/src/state.ts` stops creating `$DORMOUSE_STATE_DIR` mode `0o700`, or stops writing every file through `writeAtomic` at mode `0o600`. The "every file" clause is a negative search over `relay/src/`: no `writeFile`, `appendFile`, or `createWriteStream` may target the state directory outside `writeAtomic`. A cheap default, not a cross-platform guarantee; the installer's directory permissions below protect the installed Relay's state (rationale). - **FAIL IF** `FileBurrowStateStore` (`lib/src/host/remote/burrow-state-store.ts`) stops creating its directory `0o700` and writing `0o600` on non-Windows platforms, or if `VsCodeBurrowStateStore` stops keeping the **enrollment** in `SecretStorage`. The ACL's home in `globalState` is deliberate and is not a finding; the enrollment's is what carries `burrowToken`. - **FAIL IF** a credential the Host→Burrow rename retired stops being deleted unread at boot: `state/hosts.json` on the Relay (`forgetRetiredState` in `relay/src/state.ts`, called from `relay/src/start.ts`), `remote-host.json` on a Node-resident Burrow (`forgetRetiredState` in `lib/src/host/remote/burrow-state-store.ts`, called from `sidecar-entry.ts`), and `dormouse.remote-host.enrollment`, `dormouse.remote-host.acl.*`, `remote-host.peer-token` in VS Code (`vscode-ext/src/retired-state.ts`, called from `activate()`). Each held a live `burrowToken` or peer secret, and `SecretStorage` cannot be enumerated — a key nothing removes *by name* outlives every build that knew it. Pinned by `relay/test/state-records.test.mjs`, `lib/src/host/remote/burrow-state-store.test.ts` and `vscode-ext/test/retired-state.test.ts`. -- **FAIL IF** `burrow_state_dir` in `standalone/src-tauri/src/lib.rs` stops calling `restrict_to_owner` on the state directory **before** spawning the sidecar — on Windows those Node modes are no-ops and Node cannot set an ACL, so the guarantee is held one layer down. That call carries both legs: a newly written enrollment file *inherits* the owner-only entry, and one a prior version already left under the `%LOCALAPPDATA%` ACL — with a live `burrowToken` in it — has that entry *propagated* onto it, the half `restrict_to_owner_leaves_one_owner_only_ace` covers with its pre-existing `before.json`. +- **FAIL IF** `burrow_state_dir` in `standalone/src-tauri/src/lib.rs` publishes durable storage before `restrict_to_owner` succeeds. Windows ignores Node modes: newly written enrollment files must inherit the owner-only DACL, and existing files holding `burrowToken` must receive it by propagation. `restrict_to_owner_leaves_one_owner_only_ace` pins the existing-file case; `burrow_directory_permission_failure_disables_durable_state` pins refusal without changing existing enrollment bytes. A refused restriction must select the in-memory fallback. - **FAIL IF** `relay/src/start.ts` stops obtaining the setup password from `SetupPasswordStore.loadOrCreate(generateSetupPassword)`, `generateSetupPassword` stops using `crypto.randomBytes(32)`, `readConfig` reads `DORMOUSE_SETUP_PASSWORD` or any other setup-password input, or `SetupPasswordStore` stops refusing a persisted or generated value outside 64 lowercase hexadecimal characters. Pinned by `relay/test/config.test.mjs` and `relay/test/setup-password-store.test.mjs`. - **FAIL IF** `createApp` accepts anything but 64 lowercase hexadecimal characters as the setup password injected by the entrypoint; pinned by `relay/test/app.test.mjs`. - **FAIL IF** any installer stops making `config/`, `state/`, and `config/relay.env` reachable only by the installing user — the effective property `manage verify` tests: no principal other than that user may appear in the effective permissions. macOS and Linux achieve it with `0700`/`0600` under `umask 077`; Windows with a single owner-only ACE, whether the path carries it directly or inherits it from an already-locked parent. The Windows and Linux installers create `relay.env` and lock it before writing its contents (rationale). diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index c54888564..141d6c08d 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -101,7 +101,7 @@ constants in `lib/src/lib/platform/types.ts` (and `standalone/sidecar/pty-core.j `#[tauri::command]` over an `async fn`, which the guard below accepts equally. Tauri runs a *sync* command on the main thread, where the `recv_timeout` inside `request_from_sidecar` / `request_from_sidecar_timeout` stops the webview painting -for the whole round trip, up to `AGENT_BROWSER_TIMEOUT` (30s) (rationale). **The +for the whole round trip, up to `BROWSER_REQUEST_TIMEOUT` (40s) (rationale). **The three clipboard readers included**: their non-Windows branches round-trip through the sidecar, and the declaration is per command, not per branch. A unit test in `lib.rs` scans the source and fails on any command that reaches the blocking @@ -138,7 +138,7 @@ side"): the same as `DORMOUSE_STATE_DIR` (§Persistence, "Rust file store"); `FileBurrowStateStore` keeps enrollment and ACL there as **one** `burrow.json`, so a write is one atomic rename (rationale), and the network policy beside it (`docs/specs/remote-network.md` -→ "Policy"), both 0600 in a 0700 directory via temp-then-rename. `burrowToken` is +→ "Policy"), both committed by owner-only temp-then-rename (POSIX modes; inherited Windows DACL). `burrowToken` is a bearer credential and **never enters a webview realm**. Against the shared store contract (`docs/specs/relay.md` → "Burrow side"): - **Reads fail closed.** Only `ENOENT` and a read-but-unparseable file answer @@ -150,7 +150,7 @@ a bearer credential and **never enters a webview realm**. Against the shared sto directory Rust already created is best-effort; failing the save over it would lose the Burrow instead. - **`persistent` is declared, never inferred.** With no state directory — Rust - passes an empty value when it cannot create one — the fallback store still + passes an empty value when it cannot create or restrict one — the fallback store still *holds* both values in memory, warns once, and reports `persistent: false`. The browser dev harness is *not* this case: its per-run temp directory makes a dev enrollment live and die with the run. @@ -571,11 +571,13 @@ checks in the debounce flush. main thread may be driving while it waits inside `window_at_cursor` for that same lock. The flush reads both first and hands the scale to `refresh_rect`, whose signature takes no window at all. +- **Must serialize geometry writes with snapshot removal and recheck the closing fence after platform reads.** A debounce captured before close must not recreate the removed geometry; `a_geometry_flush_captured_before_close_cannot_recreate_removed_geometry` pins the fence. - **The flush slot is released in the same step as the drain.** A `Moved` landing between the two was marked dirty with no thread left to write it — and that move is exactly a window's final position. Source of truth: `CachedRect` / `GeometryState` / `note_geometry` / +`write_open_window_geometry` / `restore_windows` in `standalone/src-tauri/src/lib.rs`; the sequencing is pinned by `the_geometry_flush_slot_is_released_with_the_drain`. @@ -601,55 +603,41 @@ Source of truth: `CleanupGate` and `WindowEvent::Destroyed` in ### Per-window close -**Closing a window with siblings alive ends that window alone**; only the last -window's close is the quit. Rust prevents the close and emits -`dormouse://window-close-requested`; the webview acks (a ~2 s watchdog closes it -anyway if that listener is dead), asks about *its own* running work, removes its snapshot, kills the PTYs it owns, and calls back -`close_window`. - -- **A close is deliberate, so it removes the blob** — geometry - and temp sibling included — and the next launch does not reopen the window - (`docs/specs/transport.md` → "The governing rule"). -- **It runs no agent-recovery capture**: nothing is coming back. -- **A cancelled close retires its watchdog's token and never reuses it**: the - next close on that window is a fresh seq, so a watchdog still sleeping on the - cancelled one cannot destroy the window under the second dialog - (`a_cleared_close_never_hands_its_seq_to_the_next_request`). -- **It confirms on a pending download as well as on running work.** An approved, - downloaded update lives in this webview's memory, so closing the window throws - it away and nothing else can install it (`docs/specs/auto-update.md`). -- **The snapshot is removed before the kill**, and Rust refuses every later save - for that label, so a PTY exit's save cannot write it back. **Both close paths - set that refusal** — the webview's own `remove_window_session`, and - `finish_window_close` for the ack-timeout path, where the webview never ran at - all. It is dropped when the webview is destroyed and can no longer save. -- **`close_window` is the one Rust half both endings share** — a deliberate close - and a window whose last Workspace moved away (§Transfer) — because what - separates them is entirely what the webview did before calling it. -- **macOS keeps its rule**: closing the last window quits. - -**The ack and confirm gates are one shared flow with the quit** -(`createTeardownFlow` in `standalone/src/teardown-flow.ts`); what differs is only -the step past them — a quit votes and waits its turn in the walk, a close tears -down at once. - -**Arbitration.** They are two machines over one window, one dialog and one -human, and **a second flow is never refused in silence**: an unsettled context -parks its own flow and leaves its host waiting out a decision that cannot come. +**Must close only the invoking Window when siblings remain; the last Window's close is a quit.** + +1. Rust prevents native close and emits `dormouse://window-close-requested` to that Window. +2. The webview acknowledges, obtains consent, removes its snapshot, tears down its PTYs, then invokes `close_window`. +3. Without acknowledgement, Rust attempts close after `CLOSE_ACK_TIMEOUT_MS` (rationale). + +- **Must remove the deliberately closed Window's snapshot, geometry and temp sibling** (`docs/specs/transport.md` -> "The governing rule"). +- **Never capture agent recovery for a deliberate close.** +- **Must retire cancelled watchdog tokens without reuse** (`a_cleared_close_never_hands_its_seq_to_the_next_request`; rationale). +- **Must confirm on pending downloads and running work** (`docs/specs/auto-update.md`). +- **Must commit snapshot removal before killing PTYs and refuse saves for successfully closed labels for the process lifetime**, including saves dispatched before `Destroyed`. +- **Must bound preparation and journal-lock waits to eight seconds after consent, cancelling only before the first mutation.** Cancelled generations cannot remove retained or re-saved Workspaces. **Never time out consent or abandon entered kernel IO**; await its result. +- **Must retain the Window and live PTYs on removal failure, restoring changed journal/geometry before releasing save refusal.** Incomplete rollback retains the committed flow and refusal with a visible error; uncertain commits never resave (rationale). +- **Must confirm native cancellation before resetting the flow**, then drain refused saves, republish the current Window aggregate and drain again with bounded waits before offering Retry / Keep window open. **Must remember one latest-value retry when a refused write outlives those waits.** Failed cancellation keeps the flow guarded (rationale). +- **Must reject skipped saves instead of acknowledging persistence**, permitting unchanged-cache retries (`skipped_close_save_is_rejected_and_retained_window_can_retry`; `standalone/src/tauri-session-store.test.ts`). +- **Must bound PTY teardown after successful removal to eight seconds.** `standalone/src/window-close.test.ts`, `close_commit` tests and `failed_close_rolls_back_journal_and_geometry_before_retry` pin these boundaries. +- **Must finish both deliberate close and last-Workspace departure through `close_window`** (see Transfer). + +**Must share acknowledgement and consent gates with quit** through `createTeardownFlow`: quit votes and waits its walk turn; close tears down immediately. + +**Never leave a competing teardown flow unanswered** (rationale): | Arriving | Holder | Outcome | |---|---|---| -| quit | a close still on its dialog | the close is cancelled (`window_close_cancel`); the quit takes over | -| quit | any committed flow | the quit acks and **votes** — this window is ending anyway, and a window that never votes holds the machine in `Voting` with no dialog to answer | -| close | a quit, in any state | refused at once with `window_close_cancel`; the window stays | +| quit | close awaiting consent | cancel close with `window_close_cancel`; quit takes over | +| quit | committed flow | acknowledge and vote | +| close | quit in any state | immediately `window_close_cancel`; retain Window | -**A quit cancelled elsewhere drops only a quit's dialog**, never this window's -own close question. The confirm store cancels any context it cannot open, as the -backstop. Both orderings are pinned by -`standalone/src/teardown-arbiter.test.ts`. +**Must drop only quit's dialog when another Window cancels quit, and cancel any context the confirm store cannot open.** `standalone/src/teardown-arbiter.test.ts` pins both orderings. -Source of truth: `standalone/src/window-close.ts`; `request_window_close` / -`finish_window_close` in `standalone/src-tauri/src/lib.rs`. +Source of truth: `runCloseTeardown` / `retainWindowAfterFailure` in `standalone/src/window-close.ts`; +`retryLatest` in `standalone/src/tauri-session-store.ts`; +`createTeardownFlow` in `standalone/src/teardown-flow.ts`; +`CloseCommit` in `standalone/src-tauri/src/close_commit.rs`; +`remove_window_session_bounded` / `close_window_snapshot_locked` / `request_window_close` / `finish_window_close` in `standalone/src-tauri/src/lib.rs`. ### Transfer @@ -700,7 +688,7 @@ below reads that record rather than inferring itself from the suppression map. `transfer_workspace` / `open_workspace_window`. **Must return preparation refusals as `{ moved: false, reason }` without changing ownership.** On `Ok` it marks the Workspace **transferring**: the Wall stays mounted, nothing is released, and `getWindowSnapshot` omits it. -2. **Rust** reassigns `terminalIds` to the target, keeps routing their output to +2. **Rust** reserves both endpoints, writes the initial arrival journal, then reassigns `terminalIds` to the target, keeps routing their output to the source, and asks the sidecar to stamp a `pty:marked` line per id; at that line the id's suppression begins, until its replay has been emitted to the target. The source serializes each buffer at its mark and invokes @@ -709,17 +697,16 @@ below reads that record rather than inferring itself from the suppression map. for a tear-out, builds the new window, whose boot pulls a payload that is complete (`docs/specs/transport.md` → "Transferring a Workspace"; rationale). **An arrival without content is not drainable** - (`an_arrival_is_drainable_only_once_its_content_landed`). + (`an_arrival_is_drainable_only_once_its_content_landed`). **Must keep pending journal admission source-owned and save-visible, exclude it from target drains, and count it as an in-flight close/quit blocker.** A failed initial write leaves ownership unchanged; a cancelled generation cannot promote late ownership (rationale). 3. **Target** drains with `take_arrivals` and, per arrival, arms its collector - *before* calling `adopt_ready(workspaceId)` — the hop that removes the whole - "arrived before armed" class of bug (rationale). Rust answers + *before* calling `adopt_ready(workspaceId)` (rationale). Rust answers `pty:requestInit` with **that arrival's ids and no others**; `pty:list` and each `pty:replay` echo the collector's token. The target resumes over them, mounts the Workspace at the payload’s index, else the drop index. **Never spawns or kills**: the Sessions' alert state never left the sidecar, whose answer to that `pty:requestInit` re-sends it (§Alerts; `docs/specs/alert.md` → Live Workspace transfer). -4. **Target adopted** invokes `adopt_done(workspaceId)`. Rust retires the record, +4. **Target adopted** invokes `adopt_done(workspaceId)`. Rust durably marks the journal settled, then retires the in-memory record, clears pending marks and suppression so live output routes to the target, and emits `workspace-departed` for **that Workspace alone** to its own source. @@ -732,28 +719,26 @@ below reads that record rather than inferring itself from the suppression map. had no Sessions and no window that owned them. - **A transferring Workspace is in no snapshot its source writes, and neither end's teardown kills or interrupts its shells.** They belong to the target by - ownership from the invoke, and the target's `pty_graceful_kill` and + ownership after journal admission, and the target's `pty_graceful_kill` and `capture_agent_recovery` exclude every id an arrival claims (`boot_list_ids`), - since the source is still showing them. A quit or a close in the gap would - otherwise persist the same Workspace in two windows, or kill it under the - source. + since the source is still showing them (rationale). - **Must await `adopt_done` before installing a torn-out Window; refusal releases its resumed Sessions and boots fresh.** -- **A refused `adopt_done` unwinds the mount.** The `ARRIVAL_MAX` watchdog has - already handed the shells back and the source kept the Workspace, so the +- **Must unwind a refused `adopt_done` mount and request handback.** Refusal can be a failed settled-journal write or an arrival already returned by its watchdog; the target releases its Sessions (never kills them) and closes the Workspace rather than leaving it live and persisted in two windows. **Must unwind from the received payload without preparing another move.** - **Must remove this window's unmounted semantic and Activity state when arrival collection times out**, and kill nothing: the source goes on showing those Sessions. Source of truth: `discardArrival` in `standalone/src/workspace-move.ts`. -- **A refused arrival hands the shells back.** The target's `adopt_failed` +- **Must durably reverse a refused arrival before returning ownership, retiring its reservation or notifying the source.** The target's `adopt_failed` (a `planArrival` timeout, a missing list, a mount error) and a target `Destroyed` with the arrival still queued both return `terminalIds` to the - source unsuppressed (`a_target_closing_mid_arrival_hands_its_shells_back`), - drop the record, and emit `workspace-arrival-failed`; the + source (`a_target_closing_mid_arrival_hands_its_shells_back`), + drop the in-memory record, and emit `workspace-arrival-failed`; the source clears **transferring** and the Workspace is simply still there. With - both ends gone the shells are reaped rather than left owned by a dead label. + both ends gone the shells are reaped rather than left owned by a dead label, + and the durable recovery record remains. **Must retain a failed return as an in-flight blocker, show its retry status and retry its journal write without adopting or reaping the claimed PTYs.** A late worker acts only on its own phase and generation; `a_failed_reverse_journal_retains_the_return_and_retries_without_losing_recovery` pins this boundary (rationale). - **Must change transfer ownership and source routing under one routing lock**, so output before the mark always reaches the source (`transfer_ownership_and_source_routing_change_together`). @@ -785,20 +770,14 @@ below reads that record rather than inferring itself from the suppression map. rationale). - **`planArrival` never throws into `bootstrap()`.** A refused sole arrival on the boot path renders a fresh one-pane Workspace, never a blank window. -- **`take_arrivals` does not consume.** The record settles at `adopt_done`, so a - webview that drains at boot and again when its listener is installed cannot - lose a Workspace to a drain that happened too early; the webview dedupes by id. +- **Must keep `take_arrivals` non-consuming and deduplicate target adoption by id.** Settlement owns retirement (rationale). - **A window with a snapshot boots as itself**, mounting whatever was dropped on it mid-boot over the restore rather than instead of it. - **`AWAITING_REPLAY_MAX` fails open only for suppressions no arrival claims.** A cold boot slower than it would otherwise have a real arrival's shells unsilenced into a window that has not resumed them yet. -- **A boot's `pty_request_init` excludes every id an arrival claims.** Ownership - moves at the invoke, so those shells would otherwise be listed as top-level - panes beside the Workspace about to mount them. -- **`begin_arrival` records the arrival in `sessions/arrivals.json`** — a JSON - array of `{ workspaceId, from, to, workspace, settled }`, never an entry in - either window's snapshot (rationale); the tombstone rules below read `settled`. **Must retain an adopted record until target +- **Must exclude transferred arrival ids from boot's `pty_request_init`, while retaining pending admission's source-owned ids.** +- **Must journal an arrival in `sessions/arrivals.json` before ownership moves, never stage it in either window's snapshot** (rationale). **Must retain an adopted record until target and source snapshots both reflect the move**, marking it settled at `adopt_done` and checking after each `save_session` or source-window close (`adoption_keeps_the_journal_until_both_snapshots_are_durable`). **Must reverse @@ -813,24 +792,17 @@ below reads that record rather than inferring itself from the suppression map. and successful records are deleted; **must retain failed records for retry and roll back the target if trimming the source fails**. **Must preserve a settled arrival’s newer target record during boot recovery** - (`an_arrival_record_round_trips_until_it_is_forgotten`, - `a_leftover_arrival_boots_into_an_existing_target_snapshot`, - `a_leftover_arrival_boots_into_a_tear_out_targets_new_snapshot`, - `a_leftover_arrival_leaves_a_source_snapshot_that_still_names_it`, - `the_arrivals_file_is_gone_after_the_boot_merge`). + (`standalone/src-tauri/src/lib.rs` recovery tests). - **Must run journal I/O and its lock waits off the main thread**, including transfer/settlement/close commands and destroyed-window cleanup (`blocking_commands_run_off_the_main_thread`). -- **An arrival unadopted after `ARRIVAL_MAX` is handed back** by a watchdog armed - at `begin_arrival`, retiring only the record it was armed for (`queued_at`): - a target alive but wedged never reaches `adopt_failed` or `Destroyed`, and the - source would otherwise stay transferring with its shells silent for good +- **Must request return after `ARRIVAL_MAX` only for the watchdog's own generation (`queued_at`)**, including admission still waiting for disk. Journal failure delays settlement under the retained blocker (`an_expiry_retires_only_the_record_it_was_armed_for`). Source of truth: `Arrival` / `sweep_awaiting` / `expire_arrival` / `boot_list_ids` in `standalone/src-tauri/src/routing.rs`; `begin_arrival` / `adopt_ready` / -`adopt_done` / `adopt_failed` / `hand_back_arrival` / `record_arrival_on_disk` / -`forget_arrival_on_disk` / `restore_arrivals` in +`adopt_done`, `adopt_failed`, `hand_back_arrival`, `commit_initial_arrival_with`, +`commit_arrival_adoption_with`, `commit_arrival_return_with` and `restore_arrivals` in `standalone/src-tauri/src/lib.rs`; `standalone/src/workspace-move.ts`; `markWorkspaceTransferring` in `lib/src/lib/window-session-aggregator.ts`. Pinned by `standalone/src/workspace-move.test.ts`, the disk tests in @@ -976,22 +948,11 @@ record cannot reach the installed app). The browser-dev harness sets both to its own per-run temp directory. Source of truth: `recovery_state_dir` in `standalone/src-tauri/src/lib.rs`. -**Must restrict the session store to the owner before any bytes are written** -(`docs/specs/security-local.md` -> "Persisted state"; rationale). -`restrict_to_owner` sets `0700` on the directory and `0600` on the temp file -*first*, since the rename preserves its mode; on Windows, where a unix mode is a -silent no-op, it applies a protected single-entry DACL instead (mechanism in its -doc comment). `burrow_state_dir` locks the sidecar's state directory with the -same call and relies on it reaching a file that already *existed*, which -`restrict_to_owner_leaves_one_owner_only_ace` pins (rationale). **Must abort a snapshot save if either permission change fails**, preserving the previous snapshot. The state-directory call remains nonfatal and logs a `WARNING` naming the path. Pinned by `session_permission_failures_preserve_previous_snapshot_without_writing_bytes` and `session_write_tightens_directory_and_existing_temp_file`. - -**Boot + the synchronous-read constraint.** `getState()` is synchronous — -cold-start restore reads it before React mounts — but a Tauri `invoke` is async, so -`TauriSessionStore` keeps an in-memory write-through cache: `TauriAdapter.init()` -`hydrate`s it from `load_session` (§Boot sequence), `getItem` reads it -synchronously, `setItem` updates it and forwards to `save_session` asynchronously, -coalescing bursts to at most one in-flight write (latest value wins). Mirrors the -VS Code adapter's host-injected seed (`docs/specs/vscode.md`). +**Must restrict the session directory and temp file to the owner before writing bytes; permission failure must preserve the previous snapshot** (`docs/specs/security-local.md` → Persisted state; rationale). Windows uses a protected owner-only DACL, including existing files; Unix uses `0700` / `0600`. `session_permission_failures_preserve_previous_snapshot_without_writing_bytes` and `restrict_to_owner_leaves_one_owner_only_ace` pin this. **Must disable recovery storage when its directory cannot be hardened**, logging the path; Burrow state-directory failure warns and selects the in-memory fallback. + +**Must hydrate synchronous `getState()` before cold restore and coalesce asynchronous writes with latest-value-wins ordering.** + +Source of truth: `restrict_to_owner`, `recovery_state_dir` and `write_file_with_permissions` in `standalone/src-tauri/src/lib.rs`; `TauriSessionStore` in `standalone/src/tauri-session-store.ts`. Dirty tracking is shared frontend behavior (`docs/specs/layout.md` → Session persistence). **Must skip an unchanged store write only when that value is queued or saved.** @@ -1007,6 +968,9 @@ write; failed writes are logged. With Session persistence disabled, the pipeline is already idle. Drain is a completion barrier, not a guarantee of successful disk persistence. +**Must distinguish process-crash recovery from hardware power-loss durability.** +Directory fsync is best-effort after writes; unlink helpers do not fsync their directory (rationale). + ### Agent recovery **Must run shared capture and record ownership in the sidecar**, under `docs/compatible-agents.md`. Rust bridges `capture_agent_recovery` and `take_recovery_commands` asynchronously. **Must answer both through `respondAsync`**, returning `{ error }` on throws rather than stranding the invoke. diff --git a/docs/specs/standalone.rationale.md b/docs/specs/standalone.rationale.md index a22e726dd..2b4ceb3f0 100644 --- a/docs/specs/standalone.rationale.md +++ b/docs/specs/standalone.rationale.md @@ -28,6 +28,8 @@ ## Burrow service +A refused state-directory DACL leaves the Node sidecar with no privacy control on Windows: POSIX modes do nothing there. Falling back to in-memory enrollment and ACL keeps the Burrow usable without writing new credentials into an unrestricted directory. Existing durable state stays untouched until a later launch can establish the boundary. + **Why one file rather than one per value.** Per-value files leave a window between two writes in which the enrollment can end up describing a different Burrow than the ACL records approved under it. **Why the correlation field cannot be `requestId`.** A `burrow:*` payload reusing that field has its results consumed by the invoke table, vanishing at random. @@ -108,6 +110,18 @@ enumeration — which is already Rust's job, beside the snapshots it reads. The cap is a ceiling on one launch, not a limit on how many windows may exist: the excess snapshots stay on disk untouched, so raising the cap restores them. +## Per-window close + +The October 2026 Windows review found that a skipped close-time save acknowledged as successful advanced the webview's saved-value cache. Wall dirty tracking had already completed at aggregate publication, so an unchanged retained Window could stay unsaved indefinitely. Explicit refusal plus one remembered latest-value retry covers writes outliving the bounded drains; the real store and aggregate regression holds the refused write past both timeouts. The retry is consumed once and never loops on disk failure. A failed native cancellation keeps the frontend flow guarded rather than letting its next close race the old handshake. + +A cancelled watchdog needs a fresh token on retry so the earlier sleeping watchdog cannot destroy the Window under another consent dialog. Pending approved downloads live only in their webview: closing loses the installer held there. Competing teardown flows share one dialog; leaving a request unanswered parks its native machine with no decision to receive. + +An audit on Windows, 2026-10-01, held `remove_window_session` pending for nine seconds: the frontend's eight-second ceiling never started, since it covered only the later PTY kill. Native completion then took the same journal lock a second time. A disk error was logged while the code destroyed the window anyway, leaving its old snapshot eligible for relaunch. + +Cancellation before the mutation boundary can retire a worker waiting for the journal lock without touching disk. Cancellation after that boundary cannot safely release save refusal: a late unlink could delete a newer snapshot from the retained window. The token therefore gives that entered commit exclusive ownership until completion. Normal deletion errors restore the earlier journal/geometry before exposing retry; rollback failure remains guarded. An arbitrary synchronous filesystem syscall has no safe forced-cancellation mechanism, so the preparation deadline is distinct from entered kernel IO and the later PTY deadline. + +Queued native commands also survive the window's `Destroyed` event. Keeping refusal for successfully closed labels, which are not reused in a process, prevents such a save from recreating its snapshot after native teardown. + ## Transfer The `adopt_ready` hop exists because the alternative is a race with no safe @@ -133,7 +147,9 @@ ends at `adopt_failed` or at the target's `Destroyed`, not at a timer. ## Arrival queue -A target whose `adopt_done` is refused already has the arrival payload needed to release its Sessions. Preparing a new transfer first re-entered Tool startup checks and could throw while ownership was already back at the source; unwinding directly also avoids sending `adopt_failed` for that retired arrival. +A target whose `adopt_done` is refused already has the arrival payload needed to release its Sessions. Preparing a new transfer first re-entered Tool startup checks and could throw while ownership was already back at the source. Direct unwind also handles a failed settled-journal write; an explicit `adopt_failed` then requests the durable return instead of waiting for the watchdog. A return already completed makes that request stale and harmless. + +The Windows audit on 2026-10-01 found that a failed initial journal write was logged before ownership moved anyway, while the source's next snapshot omitted the Workspace. A failed return write likewise removed the close/quit blocker before its durable destination changed. Admission and return reservations now hold the blockers across disk IO and release source state only after the matching generation's journal commit. Native failure tests preserve earlier journal bytes and replay the return into exactly one source snapshot. **Why the mark is stamped in the stream rather than asked for.** A mark fetched by request answers at some instant the sidecar chose, while the source's xterm @@ -224,6 +240,8 @@ stays the webview's throughout and no polling loop is needed. **Why the sessions directory is fsynced after the rename.** Fsyncing only the temp file leaves the new name recoverable-but-absent after a power loss; the directory-entry fsync is what makes the rename itself durable. Windows has no equivalent concept, hence unix-only. +Checked on 2026-10-01: `write_file_with_permissions` attempts that Unix directory fsync best-effort, and session/journal unlink helpers do not fsync the directory. Their completed operations and journal ordering support process-crash recovery; they do not establish a guarantee for abrupt hardware power loss. + **Why the mode is set before the bytes.** Under the bare umask the transcript-bearing blob lands `0644` in a `0755` directory any other local account can read, and tightening after the write would leave a window in which it was readable. Continuing after a permission failure would contradict the owner-only guarantee; aborting before writing preserves the previous snapshot and leaves at most an empty temp file. **Why the ACE test asserts an already-existing file.** On an upgrade the Burrow enrollment file is already there, so what tightens it is propagation onto an existing entry rather than create-time inheritance. `FileBurrowStateStore`'s own `0700`/`0600` cannot help on Windows: Node has no ACL API. diff --git a/lib/src/host/remote/burrow-state-store.ts b/lib/src/host/remote/burrow-state-store.ts index 7b4438a43..41d5137f4 100644 --- a/lib/src/host/remote/burrow-state-store.ts +++ b/lib/src/host/remote/burrow-state-store.ts @@ -8,7 +8,8 @@ * The interface is async because the hosts that implement it are: files the * sidecar owns here, `VsCodeBurrowStateStore` there (enrollment in * `SecretStorage`, ACL in `globalState` — `docs/specs/vscode.md`). {@link FileBurrowStateStore} - * is the sidecar's: two files, 0600, under a directory the app passes in. + * is the sidecar's: private JSON state under a directory the app passes in + * only after establishing owner-only access (POSIX modes or a Windows DACL). */ import { readFile, rm } from 'node:fs/promises'; diff --git a/scripts/spec-word-budgets.json b/scripts/spec-word-budgets.json index b80e5b64c..698474a5c 100644 --- a/scripts/spec-word-budgets.json +++ b/scripts/spec-word-budgets.json @@ -30,7 +30,7 @@ "docs/specs/security-supply-chain.md": 1250, "docs/specs/security.md": 2150, "docs/specs/shortcuts.md": 1100, - "docs/specs/standalone.md": 11950, + "docs/specs/standalone.md": 11850, "docs/specs/terminal-context.md": 1100, "docs/specs/terminal-escapes.md": 4050, "docs/specs/terminal-state.md": 2400, diff --git a/standalone/scripts/dev-agent-browser.test.mjs b/standalone/scripts/dev-agent-browser.test.mjs index bc7916c17..ddac0e687 100644 --- a/standalone/scripts/dev-agent-browser.test.mjs +++ b/standalone/scripts/dev-agent-browser.test.mjs @@ -2,7 +2,7 @@ import test from 'node:test'; import assert from 'node:assert/strict'; import { access, copyFile, mkdir, readFile, rm, writeFile } from 'node:fs/promises'; import path from 'node:path'; -import { fileURLToPath } from 'node:url'; +import { fileURLToPath, pathToFileURL } from 'node:url'; import { spawn } from 'node:child_process'; import { get } from 'node:http'; import { setTimeout as delay } from 'node:timers/promises'; @@ -30,6 +30,10 @@ async function fixture(t) { }); `); const cli = path.join(bin, 'cli.cjs'); + // Windows kill('SIGTERM') bypasses JS handlers. Exercise the same shutdown + // handler over IPC there; POSIX continues exercising the actual signal. + const signals = path.join(bin, 'signals.mjs'); + await writeFile(signals, "process.on('message', signal => process.emit(signal));"); await writeFile(cli, ` if (process.argv[2] === 'list' && process.env.TEST_DOR_LIST) { console.log(process.env.TEST_DOR_LIST); @@ -54,8 +58,8 @@ async function fixture(t) { return { root, start(overrides = {}) { - const child = spawn(process.execPath, [path.join(standalone, 'scripts/dev-agent-browser.mjs')], { - cwd: root, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe'], + const child = spawn(process.execPath, ['--import', pathToFileURL(signals).href, path.join(standalone, 'scripts/dev-agent-browser.mjs')], { + cwd: root, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe', 'ipc'], }); // Object.assign, not a spread: `runner`'s `output`/`closed` are getters // over live state, and spreading would snapshot them once. @@ -73,7 +77,10 @@ async function fixture(t) { return this; }, async stop() { - if (child.exitCode === null && child.signalCode === null) child.kill('SIGTERM'); + if (child.exitCode === null && child.signalCode === null) { + if (process.platform === 'win32' && child.connected) child.send('SIGTERM', () => {}); + else child.kill('SIGTERM'); + } const timer = setTimeout(() => child.kill('SIGKILL'), 5000); try { return await this.exited; } finally { clearTimeout(timer); } }, diff --git a/standalone/scripts/dev-standalone.test.mjs b/standalone/scripts/dev-standalone.test.mjs index 4894deb2e..039a7c07a 100644 --- a/standalone/scripts/dev-standalone.test.mjs +++ b/standalone/scripts/dev-standalone.test.mjs @@ -2,7 +2,7 @@ import test from 'node:test'; import assert from 'node:assert/strict'; import { copyFile, mkdir, rm, writeFile } from 'node:fs/promises'; import path from 'node:path'; -import { fileURLToPath } from 'node:url'; +import { fileURLToPath, pathToFileURL } from 'node:url'; import { spawn } from 'node:child_process'; import { setTimeout as delay } from 'node:timers/promises'; import { cleanEnv, devWorkspace, runner, writeShims } from './dev-fixture.mjs'; @@ -54,7 +54,7 @@ async function fixture(t) { return { root, start(args = ['dev'], overrides = {}) { - const child = spawn(process.execPath, ['--import', signals, path.join(standalone, 'scripts/tauri.mjs'), ...args], { + const child = spawn(process.execPath, ['--import', pathToFileURL(signals).href, path.join(standalone, 'scripts/tauri.mjs'), ...args], { cwd: standalone, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe', 'ipc'], }); // Object.assign, not a spread: `runner`'s `output`/`closed` are getters diff --git a/standalone/src-tauri/src/close_commit.rs b/standalone/src-tauri/src/close_commit.rs new file mode 100644 index 000000000..257a0b15f --- /dev/null +++ b/standalone/src-tauri/src/close_commit.rs @@ -0,0 +1,69 @@ +//! Cancellation owns only preparation. Once a worker crosses the mutation +//! boundary, its caller must await the result rather than release save refusal +//! while a late unlink still owns the window's files. +use std::sync::atomic::{AtomicU8, Ordering}; + +const PREPARING: u8 = 0; +const COMMITTING: u8 = 1; +const DONE: u8 = 2; +const CANCELLED: u8 = 3; + +#[derive(Debug, Default)] +pub struct CloseCommit(AtomicU8); + +impl CloseCommit { + /// Called under the journal lock, immediately before the first mutation. + pub fn begin(&self) -> Result<(), String> { + self.0.compare_exchange(PREPARING, COMMITTING, Ordering::AcqRel, Ordering::Acquire) + .map(|_| ()).map_err(|_| "close preparation was cancelled".to_string()) + } + + pub fn cancel(&self) -> bool { + self.0.compare_exchange(PREPARING, CANCELLED, Ordering::AcqRel, Ordering::Acquire).is_ok() + } + + pub fn finish(&self) { self.0.store(DONE, Ordering::Release); } + pub fn done(&self) -> bool { self.0.load(Ordering::Acquire) == DONE } +} + +#[cfg(test)] +mod tests { + use super::*; + use std::sync::{Arc, Mutex}; + + #[test] + fn cancellation_while_waiting_for_journal_fences_the_late_worker() { + let disk = Arc::new(Mutex::new(())); + let blocked = disk.lock().unwrap(); + let token = Arc::new(CloseCommit::default()); + let worker_token = token.clone(); + let worker_disk = disk.clone(); + let worker = std::thread::spawn(move || { + let _disk = worker_disk.lock().unwrap(); + worker_token.begin() + }); + assert!(token.cancel()); + drop(blocked); + assert!(worker.join().unwrap().is_err()); + assert!(!token.done()); + // A retry owns a distinct generation; cancelling the old token cannot + // cancel or authorize the new worker. + let retry = CloseCommit::default(); + retry.begin().unwrap(); + assert!(!token.cancel()); + retry.finish(); + assert!(retry.done()); + } + + #[test] + fn entered_commit_cannot_be_cancelled_or_reentered() { + let token = CloseCommit::default(); + token.begin().unwrap(); + assert!(!token.cancel()); + assert!(!token.done()); + assert!(token.begin().is_err()); + token.finish(); + assert!(token.done()); + assert!(!token.cancel()); + } +} diff --git a/standalone/src-tauri/src/lib.rs b/standalone/src-tauri/src/lib.rs index 3b799c9dd..1d3431d49 100644 --- a/standalone/src-tauri/src/lib.rs +++ b/standalone/src-tauri/src/lib.rs @@ -5,6 +5,7 @@ mod panic_policy; mod quit_state; mod routing; mod workspaces; +mod close_commit; // The Dock's Quit, an `osascript` quit and a logout reach AppKit without ever // raising `RunEvent::ExitRequested` (docs/specs/standalone.md §Trigger // interception). @@ -140,8 +141,9 @@ struct WindowState { hover_target: Mutex>, /// Labels whose snapshot has been deliberately removed. A save arriving /// from a webview that is going away must not put the file back; the entry - /// is dropped once that webview is destroyed and can no longer save. + /// survives destruction: a previously dispatched save can still arrive. closing: Mutex>, + close_commits: Mutex>>, /// The next `ws-`, seeded above every live and saved label at setup. next_ws: AtomicU64, /// Every window's Workspaces under their stable refs (§Workspace registry). @@ -182,7 +184,8 @@ impl WindowState { } /// Refuse every later `save_session` for `label` (a deliberate close removed - /// its snapshot). Cleared by `Destroyed`, after which no save can arrive. + /// its snapshot). Kept for the process lifetime: an already dispatched + /// save may acquire the journal lock after `Destroyed`. fn begin_closing(&self, label: &str) { guard(&self.closing).insert(label.to_string()); } @@ -274,14 +277,20 @@ impl WindowState { // reads it: an arriving shell belongs to its source again, which // `hand_back` returns it to, and reaping it here would kill a terminal // the source is still showing. - let lost = { + let (lost, claimed) = { let mut arrivals = guard(&self.arrivals); arrivals.forget_deferred_close(label); - routing::take_arrivals_to(&mut arrivals, label) + let lost = routing::take_arrivals_to(&mut arrivals, label); + // A return whose journal is still being retried already has a + // worker, but its source's shells must not become orphan kills. + let claimed: Vec = arrivals.iter() + .filter(|arrival| arrival.to == label && arrival.transferred) + .flat_map(|arrival| arrival.terminal_ids.clone()).collect(); + (lost, claimed) }; let owned = { let mut routing = guard(&self.routing); - for id in lost.iter().flat_map(|arrival| &arrival.terminal_ids) { + for id in &claimed { routing.owners.remove(id); routing.awaiting_replay.remove(id); } @@ -798,20 +807,19 @@ fn request_window_close(app: &AppHandle, label: &str) { /// Tauri has actually taken the label out of `webview_windows()`. fn finish_window_close(app: &AppHandle, label: &str) { append_log(format!("[window] closing {label} and removing its snapshot")); - if let Some(state) = app.try_state::() { - // Before the removal, not after: a save already in flight from this - // webview would otherwise put the snapshot back. Reached from the - // watchdog too, where the webview never called `remove_window_session`. - state.begin_closing(label); + // The webview already committed removal on the normal path; the cached + // completion avoids another journal wait after its bounded PTY teardown. + if let Err(err) = remove_window_session_bounded(app, label) { + append_log(format!("[window] retaining {label}: {err}")); + if !err.starts_with("close-commit-uncertain:") { + if let Some(state) = app.try_state::() { guard(&state.close).clear(label); } + } + let _ = app.emit_to(label, "dormouse://window-close-failed", err); + return; } if let Some(state) = app.try_state::() { guard(&state.close).clear(label); } - if let Ok(dir) = sessions_dir(app) { - if let Err(err) = close_window_snapshot(&dir, label) { - append_log(format!("[session] {err}")); - } - } if let Some(window) = app.get_webview_window(label) { let _ = window.destroy(); } @@ -1894,14 +1902,19 @@ async fn save_session(window: tauri::Window, state: String) -> Result<(), String // A deliberate close removes the snapshot; a save still in flight from the // webview that is going away must not put it back // (docs/specs/standalone.md §Per-window close). - if let Some(windows) = window.app_handle().try_state::() { - if windows.refuses_save(window.label()) { - return Ok(()); - } - } + let windows = window.app_handle().try_state::(); let dir = sessions_dir(window.app_handle())?; - write_session_to(&dir, window.label(), &state)?; - retire_saved_arrivals(&dir) + save_open_window_session(windows.as_deref(), &dir, window.label(), &state) +} + +/// Caller owns ARRIVAL_DISK_LOCK. Refusal is an error, never a persistence +/// acknowledgment: a retained window must retry its unchanged cached value. +fn save_open_window_session(windows: Option<&WindowState>, dir: &Path, label: &str, state: &str) -> Result<(), String> { + if windows.is_some_and(|windows| windows.refuses_save(label)) { + return Err("window close is holding its snapshot; no session was saved".to_string()); + } + write_session_to(dir, label, state)?; + retire_saved_arrivals(dir) } /// The suffix `write_file_atomically` leaves on a session snapshot's temp @@ -2165,6 +2178,16 @@ fn geometry_path(dir: &Path, label: &str) -> PathBuf { dir.join(format!("{stem}.geometry.json")) } +/// Platform reads and rect-cache access precede this disk lock. Recheck the +/// closing fence here: a live webview/rect captured before close can outlive +/// snapshot removal and must not recreate geometry or share its temp writer. +fn write_open_window_geometry(windows: Option<&WindowState>, dir: &Path, label: &str, json: &str) -> Result { + let _disk = guard(&ARRIVAL_DISK_LOCK); + if windows.is_some_and(|windows| windows.refuses_save(label)) { return Ok(false); } + write_file_atomically(&geometry_path(dir, label), json)?; + Ok(true) +} + fn read_geometry(dir: &Path, label: &str) -> Option { let raw = std::fs::read_to_string(geometry_path(dir, label)).ok()?; serde_json::from_str(&raw).ok() @@ -2245,7 +2268,8 @@ fn note_geometry(app: &AppHandle, label: &str, origin: Option<(i32, i32)>, size: let Ok(json) = serde_json::to_string(&rect.to_logical()) else { continue; }; - if let Err(err) = write_file_atomically(&geometry_path(&dir, &label), &json) { + let windows = app.try_state::(); + if let Err(err) = write_open_window_geometry(windows.as_deref(), &dir, &label, &json) { append_log(format!("[window] geometry write for {label}: {err}")); } } @@ -2414,6 +2438,10 @@ fn retire_saved_arrivals(dir: &Path) -> Result<(), String> { fn mark_arrival_adopted_on_disk(dir: &Path, workspace_id: &str) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); + mark_arrival_adopted_in_locked_journal(dir, workspace_id) +} + +fn mark_arrival_adopted_in_locked_journal(dir: &Path, workspace_id: &str) -> Result<(), String> { let mut records = read_arrivals_from(dir)?; for record in &mut records { if record_workspace_id(record) == Some(workspace_id) { record["settled"] = JsonValue::Bool(true); } @@ -2425,6 +2453,10 @@ fn mark_arrival_adopted_on_disk(dir: &Path, workspace_id: &str) -> Result<(), St /// Refusal reverses the durable destination before the source is told to save. fn return_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); + return_arrival_in_locked_journal(dir, arrival) +} + +fn return_arrival_in_locked_journal(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let mut records = read_arrivals_from(dir)?; records.retain(|r| record_workspace_id(r) != Some(&arrival.workspace_id)); records.push(serde_json::json!({ @@ -2440,7 +2472,18 @@ fn return_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result<(), /// removes that stale copy instead of resurrecting it in either window. fn close_window_snapshot(dir: &Path, label: &str) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); + close_window_snapshot_locked(dir, label, None) +} + +fn close_window_snapshot_locked(dir: &Path, label: &str, token: Option<&close_commit::CloseCommit>) -> Result<(), String> { let mut records = read_arrivals_from(dir)?; + let previous_records = records.clone(); + let geometry = geometry_path(dir, label); + let previous_geometry = match std::fs::read_to_string(&geometry) { + Ok(value) => Some(value), + Err(err) if err.kind() == std::io::ErrorKind::NotFound => None, + Err(err) => return Err(format!("read {}: {err}", geometry.display())), + }; let mut changed = false; for record in &mut records { if record.get("to").and_then(JsonValue::as_str) == Some(label) @@ -2450,9 +2493,38 @@ fn close_window_snapshot(dir: &Path, label: &str) -> Result<(), String> { changed = true; } } + // All reads and lock waits precede the cancellation fence. Neither a late + // worker nor a second request may mutate a retained window after it loses. + if let Some(token) = token { token.begin()?; } + let session = dir.join(session_file_name(label)); + let remove = |path: &Path| match std::fs::remove_file(path) { + Ok(()) => Ok(()), + Err(err) if err.kind() == std::io::ErrorKind::NotFound => Ok(()), + Err(err) => Err(format!("remove {}: {err}", path.display())), + }; + // Fail before touching recoverable state when an orphan temp is blocked. + remove(&temp_write_path(&session))?; if changed { write_arrivals_to(dir, &records)?; } - remove_session_from(dir, label)?; - retire_saved_arrivals(dir) + let result = remove(&geometry).and_then(|_| remove(&session)); + if let Err(err) = result { + // The live snapshot was the final mutation. A failed removal leaves + // it present; restore the preceding journal/geometry before permitting + // the retained window to save or retry. An incomplete rollback keeps + // save refusal and the committed flow rather than claiming safety. + let rollback = (|| { + if changed { write_arrivals_to(dir, &previous_records)?; } + if let Some(value) = previous_geometry { write_file_atomically(&geometry, &value)?; } + Ok::<(), String>(()) + })(); + return match rollback { + Ok(()) => Err(err), + Err(rollback) => Err(format!("close-commit-uncertain: {err}; rollback failed: {rollback}")), + }; + } + // Retirement is cleanup: retaining a settled tombstone is recoverable and + // must not turn a successful deletion into a retry of a retained window. + if let Err(err) = retire_saved_arrivals(dir) { append_log(format!("[session] close journal cleanup: {err}")); } + Ok(()) } fn remove_workspace_from_disk(dir: &Path, label: &str, id: &str) -> Result<(), String> { @@ -2495,6 +2567,13 @@ fn record_workspace_id(record: &JsonValue) -> Option<&str> { /// Append one arrival's record, replacing any earlier record of the same id. fn record_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); + record_arrival_in_locked_journal(dir, arrival) +} + +/// Caller owns ARRIVAL_DISK_LOCK. Initial admission checks its reservation +/// under this same lock, so a cancelled initial write cannot overwrite a +/// return that already committed its reverse destination. +fn record_arrival_in_locked_journal(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let Some(workspace) = arrival.payload.get("workspace") else { return Err("arrival payload carries no workspace".to_string()); }; @@ -2648,6 +2727,8 @@ fn arrival_from(from: &str, to: &str, payload: JsonValue) -> Result Result<(), String>, +) -> Result<(), String> { + let _disk = guard(&ARRIVAL_DISK_LOCK); + let matches = |arrival: &routing::Arrival| { + arrival.workspace_id == expected.workspace_id && arrival.to == expected.to + && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Preparing + }; + if !guard(&windows.arrivals).iter().any(matches) { + return Err("the transfer reservation was retired before its journal write".to_string()); + } + write()?; + let mut arrivals = guard(&windows.arrivals); + let arrival = arrivals.iter_mut().find(|arrival| matches(arrival)) + .ok_or_else(|| "the transfer reservation was retired during its journal write".to_string())?; + windows.begin_transfer(&arrival.terminal_ids, &arrival.from, &arrival.to); + arrival.transferred = true; + arrival.phase = routing::ArrivalPhase::Active; + Ok(()) +} + +/// The queue keeps Returning arrivals as close/quit blockers until their +/// reverse journal write succeeds. Failed writes change neither ownership nor +/// the old durable destination, and retries cannot settle a later generation. +fn commit_arrival_return_with( + windows: &WindowState, + expected: &routing::Arrival, + write: impl FnOnce() -> Result<(), String>, +) -> Result { + let _disk = guard(&ARRIVAL_DISK_LOCK); + let matches = |arrival: &routing::Arrival| { + arrival.workspace_id == expected.workspace_id && arrival.to == expected.to + && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning + }; + if !guard(&windows.arrivals).iter().any(matches) { + return Err("the arrival return was retired before its journal write".to_string()); + } + write()?; + let mut arrivals = guard(&windows.arrivals); + if !arrivals.iter().any(matches) { + return Err("the arrival return was retired during its journal write".to_string()); + } + let marks = if expected.transferred { windows.hand_back(&expected.terminal_ids, &expected.from) } + else { serde_json::json!({}) }; + routing::retire_arrival(&mut arrivals, expected); + Ok(marks) +} + +/// Adoption releases the source only after its settled journal marker is +/// durable. A watchdog or dead target may request a return during that write; +/// the final phase/generation check fences the stale adoption. +fn commit_arrival_adoption_with( + windows: &WindowState, + expected: &routing::Arrival, + write: impl FnOnce() -> Result<(), String>, +) -> Result<(), String> { + let _disk = guard(&ARRIVAL_DISK_LOCK); + let matches = |arrival: &routing::Arrival| { + arrival.workspace_id == expected.workspace_id && arrival.to == expected.to + && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Active + }; + if !guard(&windows.arrivals).iter().any(matches) { + return Err("the arrival was returned before its adoption write".to_string()); + } + write()?; + let mut arrivals = guard(&windows.arrivals); + if !arrivals.iter().any(matches) { + return Err("the arrival was returned during its adoption write".to_string()); + } + windows.clear_suppression(&expected.terminal_ids); + routing::retire_arrival(&mut arrivals, expected); + Ok(()) +} + +/// Reserve both endpoints, durably journal the move, then reassign/suppress +/// the shells before either window is told. A failed initial write leaves the +/// source owner and save-visible; an expired reservation cannot move ownership. fn begin_arrival( app: &AppHandle, windows: &WindowState, @@ -2709,14 +2868,29 @@ fn begin_arrival( arrival.workspace_id )); } - // Ownership moves now; suppression waits for each id's `marked` line. - windows.begin_transfer(&arrival.terminal_ids, &arrival.from, &arrival.to); + // Reserve admission while the worker writes the journal. Both endpoint + // teardowns wait, but the source still owns and saves every Session. routing::queue_arrival(&mut arrivals, arrival.clone()); } + spawn_arrival_watchdog(app.clone(), &arrival); + let written = sessions_dir(app).and_then(|dir| { + commit_initial_arrival_with(windows, &arrival, || record_arrival_in_locked_journal(&dir, &arrival)) + }); + if let Err(error) = written { + let retired = routing::retire_arrival(&mut guard(&windows.arrivals), &arrival); + if retired { redrive_deferred_teardown(app); } + return Err(format!("cannot commit the Workspace transfer: {error}")); + } // The split point, stamped in the stream by the sidecar and routed to the // source, which serializes what it holds when it sees it (§Transfer). A // Workspace of browser panes alone has no ids to mark; its source sends // content at once. + let arrivals = guard(&windows.arrivals); + if !arrivals.iter().any(|current| current.workspace_id == arrival.workspace_id + && current.queued_at == arrival.queued_at && current.phase == routing::ArrivalPhase::Active) + { + return Err("the transfer was returned before its stream marks".to_string()); + } if let Some(sidecar) = app.try_state::() { let msg = serde_json::json!({ "event": "pty:mark", @@ -2724,27 +2898,15 @@ fn begin_arrival( }); send_to_sidecar(&sidecar, msg.to_string()); } - // Recorded on disk here, in neither window's snapshot: the source omits a - // transferring Workspace from its saves and the target writes only after - // adoption, so a crash in the gap would otherwise restore it nowhere - // (§Arrival queue). Never fatal: a failed write is logged and the transfer - // proceeds. - match sessions_dir(app) { - Ok(dir) => { - if let Err(e) = record_arrival_on_disk(&dir, &arrival) { - append_log(format!("[window] could not record {} on disk: {e}", arrival.workspace_id)); - } - } - Err(e) => append_log(format!("[window] {e}")), - } - spawn_arrival_watchdog(app.clone(), &arrival); + drop(arrivals); Ok(()) } /// Bound an arrival: a target that never settles it — alive but wedged, so /// `Destroyed` never hands it back either — would leave the Workspace marked /// transferring in the source and its shells silent for good. Past -/// `ARRIVAL_MAX` the record is retired and handed back like any refusal. +/// `ARRIVAL_MAX` requests a return. The reservation retires only after the +/// reverse journal is durable, like any refusal. fn spawn_arrival_watchdog(app: AppHandle, arrival: &routing::Arrival) { let workspace_id = arrival.workspace_id.clone(); let to = arrival.to.clone(); @@ -2779,7 +2941,8 @@ fn spawn_arrival_watchdog(app: AppHandle, arrival: &routing::Arrival) { /// routing state at `pty:marked`, never from the later serialized content. Only /// an id the sidecar never stamped goes straight back. /// -/// The record must already be out of the queue; the caller took it. +/// The caller has marked this generation Returning in the queue. It remains +/// there until its reverse destination is durable; no teardown may overtake it. fn hand_back_arrival( app: &AppHandle, windows: &WindowState, @@ -2791,12 +2954,30 @@ fn hand_back_arrival( arrival.workspace_id, arrival.to, arrival.from )); if app.get_webview_window(&arrival.from).is_some() { - if let Ok(dir) = sessions_dir(app) { - if let Err(e) = return_arrival_on_disk(&dir, arrival) { - append_log(format!("[window] could not record hand-back: {e}")); + let returned = sessions_dir(app).and_then(|dir| { + commit_arrival_return_with(windows, arrival, || return_arrival_in_locked_journal(&dir, arrival)) + }); + let marks = match returned { + Ok(marks) => marks, + Err(error) => { + append_log(format!("[window] hand-back is waiting for its journal: {error}")); + let _ = app.emit_to(arrival.from.as_str(), "dormouse://workspace-arrival-retry", + serde_json::json!({ "workspaceId": arrival.workspace_id, + "reason": format!("The Workspace could not be returned safely yet. Retrying its recovery write: {error}") })); + retry_arrival_return(app.clone(), arrival.clone(), reason.to_string()); + return; } + }; + // The source can disappear while the reverse journal waits for disk. + // If Destroyed ran before hand_back, its owner sweep could not see the + // newly returned ids; settle those orphans here instead of assigning + // invisible shells to a label that no longer exists. + if app.get_webview_window(&arrival.from).is_none() { + for id in &arrival.terminal_ids { windows.forget_pty(id); } + reap_orphaned_ptys(app, &arrival.from, arrival.terminal_ids.clone()); + redrive_deferred_teardown(app); + return; } - let marks = windows.hand_back(&arrival.terminal_ids, &arrival.from); let (marked, _) = routing::hand_back_ids(arrival, &marks); // Told before the replay is asked for, so the source is listening for // it (`acceptHandBackReplay` in `standalone/src/workspace-move.ts`). @@ -2820,10 +3001,10 @@ fn hand_back_arrival( } } } else { - if let Ok(dir) = sessions_dir(app) { - if let Err(e) = forget_arrival_on_disk(&dir, &arrival.workspace_id) { - append_log(format!("[window] could not forget orphaned arrival: {e}")); - } + // No source can continue this Workspace. Keep the durable destination + // instead of deleting its only recoverable record. + if !routing::retire_arrival(&mut guard(&windows.arrivals), arrival) { + return; } // Both ends are gone, so these shells belong to no window and nothing // would ever paint them (`routing::owner`). @@ -2835,6 +3016,20 @@ fn hand_back_arrival( redrive_deferred_teardown(app); } +/// One retry chain per Returning generation. A later drop of the same id is +/// independent; it is never adopted, retired or written by this worker. +fn retry_arrival_return(app: AppHandle, expected: routing::Arrival, reason: String) { + std::thread::spawn(move || { + std::thread::sleep(Duration::from_secs(1)); + let Some(windows) = app.try_state::() else { return; }; + let live = guard(&windows.arrivals).iter().any(|arrival| { + arrival.workspace_id == expected.workspace_id && arrival.to == expected.to + && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning + }); + if live { hand_back_arrival(&app, &windows, &expected, &reason); } + }); +} + /// Tear a Workspace out into a brand-new window under the cursor. /// /// The payload is *queued*, never emitted: an `emit_to` a window that does not @@ -2899,7 +3094,8 @@ fn transfer_workspace_content( let mut arrivals = guard(&windows.arrivals); let arrival = arrivals .iter_mut() - .find(|arrival| arrival.workspace_id == workspace_id && arrival.from == window.label()) + .find(|arrival| arrival.workspace_id == workspace_id && arrival.from == window.label() + && arrival.phase == routing::ArrivalPhase::Active) .ok_or_else(|| format!("no arrival of '{workspace_id}' from {}", window.label()))?; arrival.content = Some(content); (arrival.to.clone(), arrival.pending_window.take()) @@ -2917,9 +3113,8 @@ fn transfer_workspace_content( // already marked the Workspace transferring — `handOff` set // it when the invoke returned. So the ids go back *and* the // source is told: `arrival-failed` is what clears that mark. - if let Some(arrival) = - routing::take_arrival(&mut guard(&windows.arrivals), &workspace_id, &to) - { + let returned = routing::return_arrival(&mut guard(&windows.arrivals), &workspace_id, &to); + if let Some(arrival) = returned { hand_back_arrival(&app, &windows, &arrival, "the new window could not be built"); } return Err(err); @@ -2996,7 +3191,7 @@ fn adopt_ready( let (ids, marks) = { let arrivals = guard(&windows.arrivals); let arrival = routing::find_arrival(&arrivals, &workspace_id) - .filter(|arrival| arrival.to == label) + .filter(|arrival| arrival.to == label && arrival.phase == routing::ArrivalPhase::Active) .ok_or_else(|| format!("no arrival of '{workspace_id}' into {label}"))?; (arrival.terminal_ids.clone(), routing::arrival_marks(arrival)) }; @@ -3020,19 +3215,15 @@ fn adopt_done( windows: tauri::State<'_, WindowState>, workspace_id: String, ) -> Result<(), String> { - let arrival = routing::take_arrival( - &mut guard(&windows.arrivals), - &workspace_id, - window.label(), - ) - .ok_or_else(|| format!("no arrival of '{workspace_id}' into {}", window.label()))?; - windows.clear_suppression(&arrival.terminal_ids); + let arrival = { + let arrivals = guard(&windows.arrivals); + routing::find_arrival(&arrivals, &workspace_id) + .filter(|arrival| arrival.to == window.label() && arrival.phase == routing::ArrivalPhase::Active) + .cloned().ok_or_else(|| format!("no active arrival of '{workspace_id}' into {}", window.label()))? + }; // Keep the journal until source and target saves both reflect the move. - if let Ok(dir) = sessions_dir(&app) { - if let Err(e) = mark_arrival_adopted_on_disk(&dir, &workspace_id) { - append_log(format!("[window] could not mark {workspace_id} adopted on disk: {e}")); - } - } + let dir = sessions_dir(&app)?; + commit_arrival_adoption_with(&windows, &arrival, || mark_arrival_adopted_in_locked_journal(&dir, &workspace_id))?; append_log(format!( "[window] {workspace_id} adopted by {}; telling {}", arrival.to, arrival.from @@ -3055,7 +3246,7 @@ fn adopt_failed( workspace_id: String, reason: Option, ) -> Result<(), String> { - let arrival = routing::take_arrival( + let arrival = routing::return_arrival( &mut guard(&windows.arrivals), &workspace_id, window.label(), @@ -3163,11 +3354,53 @@ fn take_arrivals(window: tauri::Window, windows: tauri::State<'_, WindowState>) /// Remove this window's persisted snapshot and stop it being written again. #[tauri::command] async fn remove_window_session(window: tauri::Window) -> Result<(), String> { - let app = window.app_handle(); - if let Some(windows) = app.try_state::() { - windows.begin_closing(window.label()); + remove_window_session_bounded(window.app_handle(), window.label()) +} + +/// Bound preparation/lock waits, never an irreversible kernel operation. The +/// timeout can retire only the worker's own token before its first mutation. +fn remove_window_session_bounded(app: &AppHandle, label: &str) -> Result<(), String> { + const PREPARATION_TIMEOUT: Duration = Duration::from_millis(8000); + let dir = sessions_dir(app)?; + let windows = app.state::(); + let token = { + let mut commits = guard(&windows.close_commits); + if let Some(previous) = commits.get(label) { + if previous.done() { return Ok(()); } + return Err("close-commit-uncertain: an earlier close is still finishing".to_string()); + } + let token = Arc::new(close_commit::CloseCommit::default()); + commits.insert(label.to_string(), token.clone()); + windows.begin_closing(label); + token + }; + let (send, receive) = mpsc::channel(); + let worker_token = token.clone(); + let worker_label = label.to_string(); + std::thread::spawn(move || { + let result = { + let _disk = guard(&ARRIVAL_DISK_LOCK); + close_window_snapshot_locked(&dir, &worker_label, Some(&worker_token)) + }; + if result.is_ok() { worker_token.finish(); } + let _ = send.send(result); + }); + let result = match receive.recv_timeout(PREPARATION_TIMEOUT) { + Ok(result) => result, + Err(mpsc::RecvTimeoutError::Timeout) if token.cancel() => + Err("close preparation timed out; workspaces were retained".to_string()), + // Commit owns the files; retain the progress modal and save refusal + // until the result is known. Arbitrary kernel IO cannot be cancelled. + Err(mpsc::RecvTimeoutError::Timeout) => receive.recv() + .map_err(|_| "close-commit-uncertain: close worker stopped".to_string())?, + Err(mpsc::RecvTimeoutError::Disconnected) => + Err("close-commit-uncertain: close worker stopped".to_string()), + }; + if result.as_ref().is_err_and(|err| !err.starts_with("close-commit-uncertain:")) { + guard(&windows.close_commits).remove(label); + guard(&windows.closing).remove(label); } - close_window_snapshot(&sessions_dir(app)?, window.label()) + result } /// Which window is under the cursor, in that window's own logical client space. @@ -3330,6 +3563,13 @@ fn forget_restart(app: &AppHandle) { // ── Per-window close (docs/specs/standalone.md §Per-window close) ───────────── // This window's close orchestrator is alive; stand its ack watchdog down. +#[tauri::command] +fn retry_window_close(app: AppHandle, window: tauri::Window) { + // Retry starts a fresh native handshake too: close admission must block + // transfers while the new human gate is open, before removal begins. + request_window_close(&app, window.label()); +} + #[tauri::command] fn window_close_ack(window: tauri::Window, state: tauri::State<'_, QuitState>) { guard(&state.close).ack(window.label()); @@ -3612,6 +3852,17 @@ fn resolve_dor_cli_paths(sidecar_path: &Path, manifest_dir: &Path) -> DorCliPath // (lib/src/host/remote/burrow-state-store.ts). Created here so a first launch // hands the sidecar a directory that exists; if it can't be made, the sidecar is // told nothing and runs without persistence rather than not at all. +// The sidecar must receive a durable directory only after its privacy boundary +// exists. On Windows the Node store cannot repair a refused DACL installation. +fn prepare_burrow_state_dir( + dir: &Path, + restrict: impl FnOnce(&Path, u32) -> Result<(), String>, +) -> Result { + create_dir_all(dir).map_err(|e| format!("create state dir: {e}"))?; + restrict(dir, 0o700).map_err(|e| format!("restrict state dir {}: {e}", dir.display()))?; + Ok(dir.to_string_lossy().into_owned()) +} + fn burrow_state_dir(app: &AppHandle) -> Option { let dir = match app.path().app_data_dir() { Ok(dir) => dir, @@ -3620,10 +3871,6 @@ fn burrow_state_dir(app: &AppHandle) -> Option { return None; } }; - if let Err(e) = create_dir_all(&dir) { - append_log(format!("[sidecar] create state dir: {e}")); - return None; - } // The Node sidecar writes the Burrow enrollment here, and that record carries // `burrowToken` — a bearer credential for `/ws/burrow`. `FileBurrowStateStore` // asks for `0700`/`0600`, which Windows ignores entirely, so on Windows this @@ -3634,17 +3881,13 @@ fn burrow_state_dir(app: &AppHandle) -> Option { // `restrict_to_owner_leaves_one_owner_only_ace` covers with `before.json`. // On unix the store's own modes already do the job and this is a harmless // re-assert of the same intent. - if let Err(e) = restrict_to_owner(&dir, 0o700) { - // Not fatal — a Burrow that cannot start is worse than one whose state - // directory kept the OS default — but never silent: on Windows this - // call is the only thing restricting `burrowToken`, so its failure is a - // downgrade of the sole control and has to be visible. - append_log(format!( - "[sidecar] WARNING could not restrict state dir {}: {e}", - dir.display() - )); + match prepare_burrow_state_dir(&dir, restrict_to_owner) { + Ok(path) => Some(path), + Err(error) => { + append_log(format!("[sidecar] WARNING {error}; Burrow state is ephemeral")); + None + } } - Some(dir.to_string_lossy().into_owned()) } /// Where the sidecar writes the single-use agent-recovery record. Under the @@ -3669,6 +3912,9 @@ fn recovery_state_dir(app: &AppHandle) -> Option { "[recovery] WARNING could not restrict state dir {}: {e}", dir.display() )); + // Recovery commands must never be persisted to a directory whose + // owner-only boundary could not be established. + return None; } Some(dir.to_string_lossy().into_owned()) } @@ -3996,7 +4242,8 @@ pub fn run() { // Drop label-keyed ownership synchronously; only the // returned arrivals need the blocking journal worker. let (lost, orphaned) = state.drop_window(&label); - guard(&state.closing).remove(&label); + // Keep successful-close save refusal: a queued async + // save can still arrive after this native event. reap_orphaned_ptys(app, &label, orphaned); let changed = workspaces::forget_window(&mut guard(&state.registry), &label); if changed { broadcast_registry(app, &state); } @@ -4161,6 +4408,7 @@ pub fn run() { quit_proceed, quit_restart, window_close_ack, + retry_window_close, window_close_cancel, close_window, open_workspace_window, @@ -4239,6 +4487,9 @@ mod tests { }; use super::routing; use super::guard; + use super::{WindowState, QuitIntent, commit_initial_arrival_with, + commit_arrival_return_with, commit_arrival_adoption_with, + record_arrival_in_locked_journal, return_arrival_in_locked_journal}; use std::collections::HashSet; use std::fs; use std::path::{Path, PathBuf}; @@ -4295,6 +4546,53 @@ mod tests { } } + #[test] + fn burrow_directory_creation_failure_never_attempts_permissions() { + let root = TempDir::new("burrow-state-create-failure"); + let occupied = root.path().join("not-a-directory"); + fs::write(&occupied, b"previous").unwrap(); + let called = std::cell::Cell::new(false); + let result = super::prepare_burrow_state_dir(&occupied, |_, _| { + called.set(true); + Ok(()) + }); + assert!(result.is_err()); + assert!(!called.get()); + assert_eq!(fs::read(occupied).unwrap(), b"previous"); + } + + #[test] + fn burrow_directory_permission_failure_disables_durable_state() { + let root = TempDir::new("burrow-state-restrict-failure"); + let target = root.path().join("state"); + fs::create_dir(&target).unwrap(); + fs::write(target.join("burrow.json"), b"existing enrollment").unwrap(); + let result = super::prepare_burrow_state_dir(&target, |path, mode| { + assert_eq!(path, target); + assert_eq!(mode, 0o700); + Err("DACL refused".into()) + }); + assert!(result.unwrap_err().contains("DACL refused")); + assert_eq!(fs::read(target.join("burrow.json")).unwrap(), b"existing enrollment"); + assert_eq!(fs::read_dir(target).unwrap().count(), 1); + } + + #[test] + fn burrow_directory_is_created_and_restricted_before_publication() { + let root = TempDir::new("burrow-state-order"); + let target = root.path().join("nested").join("state"); + let called = std::cell::Cell::new(false); + let result = super::prepare_burrow_state_dir(&target, |path, mode| { + assert!(path.is_dir()); + assert_eq!(fs::read_dir(path).unwrap().count(), 0); + assert_eq!(mode, 0o700); + called.set(true); + Ok(()) + }); + assert!(called.get()); + assert_eq!(result.unwrap(), target.to_string_lossy()); + } + // --- Pending arrivals on disk (§Arrival queue) --------------------------- fn workspace_json(id: &str, name: &str) -> JsonValue { @@ -4303,6 +4601,8 @@ mod tests { fn arrival_of(id: &str, from: &str, to: &str) -> routing::Arrival { routing::Arrival { + phase: routing::ArrivalPhase::Active, + transferred: true, workspace_id: id.to_string(), from: from.to_string(), to: to.to_string(), @@ -4321,6 +4621,133 @@ mod tests { .to_string() } + fn reserve_test_arrival(windows: &WindowState) -> routing::Arrival { + let mut arrival = arrival_of("workspace-7", "main", "ws-2"); + arrival.phase = routing::ArrivalPhase::Preparing; + arrival.transferred = false; + arrival.terminal_ids = vec!["pane-a".to_string()]; + windows.mint("pane-a", "main"); + guard(&windows.arrivals).push(arrival.clone()); + arrival + } + + #[test] + fn a_pending_journal_reserves_teardown_without_hiding_or_moving_source_ptys() { + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + let mut arrivals = guard(&windows.arrivals); + arrivals[0].content = Some(serde_json::json!({})); + assert!(arrivals.defer_close("main")); + assert!(arrivals.defer_close("ws-2")); + assert_eq!(arrivals.defer_quit(&QuitIntent::default()), Some(false)); + assert!(routing::arrival_payloads(&arrivals, "ws-2").is_empty()); + assert_eq!(routing::boot_list_ids(windows.owned_by("main"), &arrivals), expected.terminal_ids); + assert!(guard(&windows.routing).marking.is_empty()); + } + + #[test] + fn failed_initial_journal_preserves_source_ownership_and_previous_durable_bytes() { + let dir = TempDir::new("arrival-initial-failure"); + let old = arrival_of("workspace-3", "main", "ws-3"); + record_arrival_on_disk(dir.path(), &old).unwrap(); + let previous = fs::read(arrivals_path(dir.path())).unwrap(); + write_session_to(dir.path(), "main", &snapshot_json(&[("workspace-7", "Source")], "workspace-7")).unwrap(); + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + + assert!(commit_initial_arrival_with(&windows, &expected, || Err("disk full".to_string())).is_err()); + assert_eq!(windows.owned_by("main"), expected.terminal_ids); + assert!(guard(&windows.routing).marking.is_empty()); + assert!(guard(&windows.routing).awaiting_replay.is_empty()); + assert!(routing::retire_arrival(&mut guard(&windows.arrivals), &expected)); + assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); + assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "main").unwrap()), vec!["workspace-7"]); + } + + #[test] + fn a_target_destroyed_before_initial_commit_never_drops_source_owned_shells() { + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + let (lost, orphaned) = windows.drop_window("ws-2"); + assert!(orphaned.is_empty()); + assert_eq!(lost.len(), 1); + assert!(!lost[0].transferred); + assert_eq!(windows.owned_by("main"), expected.terminal_ids); + assert!(commit_initial_arrival_with(&windows, &expected, || panic!("cancelled reservation must not write")).is_err()); + } + + #[test] + fn a_return_requested_during_initial_write_fences_late_ownership_and_recovers_source() { + let dir = TempDir::new("arrival-initial-return-race"); + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + write_session_to(dir.path(), "main", &snapshot_json(&[("workspace-7", "Source")], "workspace-7")).unwrap(); + let result = commit_initial_arrival_with(&windows, &expected, || { + record_arrival_in_locked_journal(dir.path(), &expected)?; + routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); + Ok(()) + }); + assert!(result.is_err()); + assert_eq!(windows.owned_by("main"), expected.terminal_ids); + assert!(guard(&windows.routing).marking.is_empty()); + let returning = guard(&windows.arrivals)[0].clone(); + commit_arrival_return_with(&windows, &returning, || return_arrival_in_locked_journal(dir.path(), &returning)).unwrap(); + restore_arrivals(dir.path()).unwrap(); + assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "main").unwrap()), vec!["workspace-7"]); + assert!(read_snapshot(dir.path(), "ws-2").is_none()); + } + + #[test] + fn a_failed_reverse_journal_retains_the_return_and_retries_without_losing_recovery() { + let dir = TempDir::new("arrival-return-failure"); + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + commit_initial_arrival_with(&windows, &expected, || record_arrival_in_locked_journal(dir.path(), &expected)).unwrap(); + let returning = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); + let previous = fs::read(arrivals_path(dir.path())).unwrap(); + assert!(commit_arrival_return_with(&windows, &returning, || Err("permission denied".to_string())).is_err()); + assert_eq!(windows.owned_by("ws-2"), expected.terminal_ids); + assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); + assert_eq!(guard(&windows.arrivals).defer_quit(&QuitIntent::default()), Some(false)); + assert!(routing::arrival_payloads(&guard(&windows.arrivals), "ws-2").is_empty()); + + commit_arrival_return_with(&windows, &returning, || return_arrival_in_locked_journal(dir.path(), &returning)).unwrap(); + assert_eq!(windows.owned_by("main"), expected.terminal_ids); + assert!(guard(&windows.arrivals).is_empty()); + restore_arrivals(dir.path()).unwrap(); + assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "main").unwrap()), vec!["workspace-7"]); + assert!(read_snapshot(dir.path(), "ws-2").is_none()); + } + + #[test] + fn a_return_retry_survives_target_destruction_without_reaping_source_shells() { + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + commit_initial_arrival_with(&windows, &expected, || Ok(())).unwrap(); + let returned = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); + let (lost, orphaned) = windows.drop_window("ws-2"); + assert!(lost.is_empty(), "one return worker is already responsible"); + assert!(orphaned.is_empty(), "a return's PTYs are not sibling-window orphans"); + commit_arrival_return_with(&windows, &returned, || Ok(())).unwrap(); + assert_eq!(windows.owned_by("main"), expected.terminal_ids); + } + + #[test] + fn adoption_write_failure_or_a_concurrent_return_never_releases_the_source() { + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + commit_initial_arrival_with(&windows, &expected, || Ok(())).unwrap(); + let active = guard(&windows.arrivals)[0].clone(); + assert!(commit_arrival_adoption_with(&windows, &active, || Err("journal unavailable".to_string())).is_err()); + assert_eq!(guard(&windows.arrivals)[0].phase, routing::ArrivalPhase::Active); + assert!(commit_arrival_adoption_with(&windows, &active, || { + routing::return_arrival(&mut guard(&windows.arrivals), &active.workspace_id, &active.to).unwrap(); + Ok(()) + }).is_err()); + assert_eq!(guard(&windows.arrivals)[0].phase, routing::ArrivalPhase::Returning); + assert!(commit_arrival_adoption_with(&windows, &active, || panic!("stale adopter must not write")).is_err()); + } + fn read_snapshot(dir: &Path, label: &str) -> Option { read_session_from(dir, label) .unwrap() @@ -5300,6 +5727,27 @@ mod tests { ); } + #[test] + fn a_geometry_flush_captured_before_close_cannot_recreate_removed_geometry() { + let dir = TempDir::new("geometry-close-fence"); + let windows = WindowState::default(); + write_session_to(dir.path(), "ws-2", "snapshot").unwrap(); + assert!(super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", "previous").unwrap()); + let captured_before_close = "new position"; + windows.begin_closing("ws-2"); + super::close_window_snapshot(dir.path(), "ws-2").unwrap(); + assert!(!super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", captured_before_close).unwrap()); + assert!(!super::geometry_path(dir.path(), "ws-2").exists()); + windows.drop_window("ws-2"); + assert!(!super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", captured_before_close).unwrap()); + assert!(!super::geometry_path(dir.path(), "ws-2").exists()); + // A recoverable refusal clears closing while the window remains live. + windows.begin_closing("ws-3"); + guard(&windows.closing).remove("ws-3"); + assert!(super::write_open_window_geometry(Some(&windows), dir.path(), "ws-3", "retained window").unwrap()); + assert_eq!(fs::read_to_string(super::geometry_path(dir.path(), "ws-3")).unwrap(), "retained window"); + } + /// The cached box is fed by the window events alone, and it is what both the /// debounced write and the cross-window drag read (§Boot and geometry). #[test] @@ -5400,16 +5848,69 @@ mod tests { /// paths set it: the webview's own `remove_window_session`, and /// `finish_window_close` for the ack-timeout path where it never ran. #[test] - fn a_closing_window_refuses_every_later_save_until_it_is_destroyed() { + fn a_successfully_closed_window_refuses_even_saves_dispatched_before_destroyed() { let state = super::WindowState::default(); assert!(!state.refuses_save("ws-2")); state.begin_closing("ws-2"); assert!(state.refuses_save("ws-2")); // Never a sibling's. assert!(!state.refuses_save("main")); - // `Destroyed` drops the refusal: no save can arrive under a dead label. - guard(&state.closing).remove("ws-2"); - assert!(!state.refuses_save("ws-2")); + state.drop_window("ws-2"); + assert!(state.refuses_save("ws-2")); + } + + #[test] + fn cancelled_close_never_mutates_retained_or_resaved_window() { + let dir = TempDir::new("close-cancelled"); + write_session_to(dir.path(), "ws-2", "retained").unwrap(); + let token = super::close_commit::CloseCommit::default(); + assert!(token.cancel()); + assert!(super::close_window_snapshot_locked(dir.path(), "ws-2", Some(&token)).is_err()); + write_session_to(dir.path(), "ws-2", "newer save").unwrap(); + assert!(super::close_window_snapshot_locked(dir.path(), "ws-2", Some(&token)).is_err()); + assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("newer save")); + } + + #[test] + fn failed_close_rolls_back_journal_and_geometry_before_retry() { + let dir = TempDir::new("close-rollback"); + let arrival = arrival_of("ws-7", "main", "ws-2"); + record_arrival_on_disk(dir.path(), &arrival).unwrap(); + super::mark_arrival_adopted_on_disk(dir.path(), "ws-7").unwrap(); + let previous = fs::read(arrivals_path(dir.path())).unwrap(); + let geometry = super::geometry_path(dir.path(), "ws-2"); + fs::write(&geometry, "geometry").unwrap(); + // A directory cannot be removed as a session file on any platform. + let session = dir.path().join(session_file_name("ws-2")); + fs::create_dir(&session).unwrap(); + assert!(super::close_window_snapshot(dir.path(), "ws-2").is_err()); + assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); + assert_eq!(fs::read_to_string(&geometry).unwrap(), "geometry"); + fs::remove_dir(&session).unwrap(); + write_session_to(dir.path(), "ws-2", &snapshot_json(&[("ws-7", "Moved")], "ws-7")).unwrap(); + write_session_to(dir.path(), "main", &snapshot_json(&[("ws-7", "Moved")], "ws-7")).unwrap(); + super::close_window_snapshot(dir.path(), "ws-2").unwrap(); + restore_arrivals(dir.path()).unwrap(); + assert!(read_session_from(dir.path(), "ws-2").unwrap().is_none()); + assert!(read_session_from(dir.path(), "main").unwrap().is_none()); + } + + #[test] + fn skipped_close_save_is_rejected_and_retained_window_can_retry() { + let dir = TempDir::new("close-cache-refusal"); + let windows = WindowState::default(); + write_session_to(dir.path(), "ws-2", "previous").unwrap(); + windows.begin_closing("ws-2"); + assert!(super::save_open_window_session(Some(&windows), dir.path(), "ws-2", "newer").is_err()); + assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("previous")); + assert!(read_session_from(dir.path(), "main").unwrap().is_none()); + guard(&windows.closing).remove("ws-2"); // recoverable close preparation failure + super::save_open_window_session(Some(&windows), dir.path(), "ws-2", "newer").unwrap(); + assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("newer")); + windows.begin_closing("ws-2"); + windows.drop_window("ws-2"); + assert!(super::save_open_window_session(Some(&windows), dir.path(), "ws-2", "stale queued save").is_err()); + assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("newer")); } fn queue_test_suppression(state: &super::WindowState, id: &str) { @@ -5674,6 +6175,8 @@ mod tests { "record_arrival_on_disk", "mark_arrival_adopted_on_disk", "return_arrival_on_disk", "forget_arrival_on_disk", "read_arrivals_from", "write_arrivals_to", "restore_arrivals", "close_window_snapshot", "finish_window_close", "begin_arrival", "hand_back_arrival", + "remove_window_session_bounded", "commit_initial_arrival_with", + "commit_arrival_adoption_with", "commit_arrival_return_with", ].iter().any(|helper| body.contains(helper)); if reaches_blocking && !(is_async_attr || is_async_fn) { offenders.push(name.to_string()); diff --git a/standalone/src-tauri/src/quit_state.rs b/standalone/src-tauri/src/quit_state.rs index 62e43c222..2795359b6 100644 --- a/standalone/src-tauri/src/quit_state.rs +++ b/standalone/src-tauri/src/quit_state.rs @@ -462,6 +462,8 @@ mod tests { fn arrival(from: &str, to: &str) -> crate::routing::Arrival { crate::routing::Arrival { + phase: crate::routing::ArrivalPhase::Active, + transferred: true, workspace_id: format!("{from}-to-{to}"), from: from.to_string(), to: to.to_string(), diff --git a/standalone/src-tauri/src/routing.rs b/standalone/src-tauri/src/routing.rs index 3c460670a..4474b2fcb 100644 --- a/standalone/src-tauri/src/routing.rs +++ b/standalone/src-tauri/src/routing.rs @@ -245,6 +245,13 @@ pub fn route<'a>(event: &str, data: &'a JsonValue, view: &RouteView<'a>) -> Rout /// drop target, and an `emit_to` it would simply be lost. #[derive(Debug, Clone, PartialEq)] pub struct Arrival { + /// Preparing reserves both endpoints while disk I/O runs without the + /// arrival lock. Returning retains that reservation until its reverse + /// destination is durable. Neither phase is adoptable. + pub phase: ArrivalPhase, + /// A Preparing arrival that is cancelled has never moved the source's + /// shells; its return must not hide or reap those source-owned PTYs. + pub transferred: bool, pub workspace_id: String, /// The window that still shows the Workspace until the target adopts it. pub from: String, @@ -266,6 +273,9 @@ pub struct Arrival { pub pending_window: Option, } +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum ArrivalPhase { Preparing, Active, Returning } + /// Every arrival in flight, oldest first. A Vec, not a map: there are a handful /// at most, and both the per-window drain and the by-Workspace lookup want the /// order the drops happened in. @@ -299,7 +309,7 @@ pub fn take_arrival(arrivals: &mut Arrivals, workspace_id: &str, to: &str) -> Op Some(arrivals.remove(position)) } -/// Retire an arrival that outlived `ARRIVAL_MAX`, but **only the exact record +/// Request return of an arrival that outlived `ARRIVAL_MAX`, but only the exact record /// the watchdog was armed for**: one adopted and re-dropped since would carry a /// later `queued_at`, and belongs to its own watchdog. pub fn expire_arrival( @@ -308,23 +318,46 @@ pub fn expire_arrival( to: &str, queued_at: Instant, ) -> Option { - let position = arrivals.iter().position(|arrival| { + let arrival = arrivals.iter_mut().find(|arrival| { arrival.workspace_id == workspace_id && arrival.to == to && arrival.queued_at == queued_at + && arrival.phase != ArrivalPhase::Returning })?; - Some(arrivals.remove(position)) + arrival.phase = ArrivalPhase::Returning; + Some(arrival.clone()) +} + +/// Reserve settlement before releasing the lock. A close/quit must continue +/// waiting while the return destination is being written or retried. +pub fn return_arrival(arrivals: &mut Arrivals, workspace_id: &str, to: &str) -> Option { + let arrival = arrivals.iter_mut().find(|arrival| { + arrival.workspace_id == workspace_id && arrival.to == to + && arrival.phase != ArrivalPhase::Returning + })?; + arrival.phase = ArrivalPhase::Returning; + Some(arrival.clone()) +} + +/// Retire only the generation whose durable settlement finished. +pub fn retire_arrival(arrivals: &mut Arrivals, expected: &Arrival) -> bool { + let Some(position) = arrivals.iter().position(|arrival| { + arrival.workspace_id == expected.workspace_id && arrival.to == expected.to + && arrival.queued_at == expected.queued_at && arrival.phase == expected.phase + }) else { return false; }; + arrivals.remove(position); + true } -/// Every arrival `label` will never take, removed: its window is gone. +/// Reserve returns for arrivals whose target is gone, retaining their teardown +/// blockers until the reverse journal commits. Existing return workers own +/// their generation and are not started a second time. pub fn take_arrivals_to(arrivals: &mut Arrivals, label: &str) -> Vec { let mut lost = Vec::new(); - arrivals.retain(|arrival| { - if arrival.to == label { + for arrival in arrivals { + if arrival.to == label && arrival.phase != ArrivalPhase::Returning { + arrival.phase = ArrivalPhase::Returning; lost.push(arrival.clone()); - false - } else { - true } - }); + } lost } @@ -335,7 +368,7 @@ pub fn take_arrivals_to(arrivals: &mut Arrivals, label: &str) -> Vec { pub fn arrival_payloads(arrivals: &Arrivals, label: &str) -> Vec { arrivals .iter() - .filter(|arrival| arrival.to == label) + .filter(|arrival| arrival.to == label && arrival.phase == ArrivalPhase::Active) .filter_map(|arrival| { let content = arrival.content.as_ref()?; let mut payload = arrival.payload.clone(); @@ -386,13 +419,14 @@ pub fn hand_back_ids(arrival: &Arrival, marks: &JsonValue) -> (Vec, Vec< pub fn arrival_ids(arrivals: &Arrivals) -> HashSet { arrivals .iter() + .filter(|arrival| arrival.transferred) .flat_map(|arrival| arrival.terminal_ids.iter().cloned()) .collect() } /// What a window's own `pty:requestInit` may name — and what its teardown may /// kill or interrupt: the ids it owns, **minus every id an arrival claims**. -/// Ownership moves at the source's invoke, so a window with a Workspace queued +/// Ownership moves after journal admission, so a window with a Workspace queued /// for it owns those shells while the source is still showing them; listed at /// boot they would be placed as top-level panes beside the Workspace about to /// mount them, and in a teardown's kill set they would die under the source. @@ -865,6 +899,8 @@ mod tests { fn arrival(workspace_id: &str, from: &str, to: &str, ids: &[&str]) -> Arrival { Arrival { + phase: ArrivalPhase::Active, + transferred: true, workspace_id: workspace_id.to_string(), from: from.to_string(), to: to.to_string(), @@ -933,10 +969,12 @@ mod tests { assert_eq!(expire_arrival(&mut arrivals, "ws-a", "ws-2", armed_for), None); assert_eq!(arrivals, vec![second.clone()]); - assert_eq!( - expire_arrival(&mut arrivals, "ws-a", "ws-2", second.queued_at), - Some(second) - ); + let mut returned = second; + returned.phase = ArrivalPhase::Returning; + assert_eq!(expire_arrival(&mut arrivals, "ws-a", "ws-2", returned.queued_at), Some(returned.clone())); + assert_eq!(arrivals, vec![returned.clone()], "return keeps blocking teardown until its journal commits"); + assert_eq!(expire_arrival(&mut arrivals, "ws-a", "ws-2", returned.queued_at), None); + assert!(retire_arrival(&mut arrivals, &returned)); assert!(arrivals.is_empty()); } diff --git a/standalone/src/WorkspaceTeardownModal.tsx b/standalone/src/WorkspaceTeardownModal.tsx index 40795f7df..5eacab1da 100644 --- a/standalone/src/WorkspaceTeardownModal.tsx +++ b/standalone/src/WorkspaceTeardownModal.tsx @@ -2,7 +2,7 @@ import { useCallback, useRef, useSyncExternalStore } from 'react'; // Standalone reaches into the lib source directly (same relative form as the // sibling UpdateDebugModal.tsx). The terminal registry comes in via the // `dormouse-lib` alias, matching quit.ts. -import { ModalFrame } from '../../lib/src/components/design'; +import { ModalFrame, modalActionButton } from '../../lib/src/components/design'; import { WorkspaceKillConfirm } from '../../lib/src/components/WorkspaceKillConfirm'; import { subscribeToTerminalPaneState } from 'dormouse-lib/lib/terminal-registry'; import { @@ -12,6 +12,8 @@ import { getQuitConfirmChar, getQuitConfirmWorkspaceNames, getQuitConfirmPhase, + getCloseFailure, + getQuitProgressDetail, quitRunningWork, subscribeQuitConfirm, type QuitConfirmIntent, @@ -28,6 +30,8 @@ import { export function WorkspaceTeardownModalHost() { const phase = useSyncExternalStore(subscribeQuitConfirm, getQuitConfirmPhase); const intent = useSyncExternalStore(subscribeQuitConfirm, getQuitConfirmIntent); + const failure = useSyncExternalStore(subscribeQuitConfirm, getCloseFailure); + const progressDetail = useSyncExternalStore(subscribeQuitConfirm, getQuitProgressDetail); if (!phase) return null; return ( @@ -36,6 +40,8 @@ export function WorkspaceTeardownModalHost() { workspaceNames={getQuitConfirmWorkspaceNames()} confirming={phase === 'quitting'} intent={intent} + failure={phase === 'close-failed' ? failure : null} + progressDetail={progressDetail} /> ); } @@ -47,6 +53,8 @@ export function WorkspaceTeardownModal({ char = 'q', workspaceNames = [], intent = { kind: 'quit' }, + failure = null, + progressDetail = null, }: { confirming: boolean; char?: string; @@ -54,6 +62,8 @@ export function WorkspaceTeardownModal({ /** Whether this tears down the whole app or one window, and whether that * discards a downloaded update. A quit and a restart read the same. */ intent?: QuitConfirmIntent; + failure?: ReturnType; + progressDetail?: string | null; }) { const progressRef = useRef(null); // Live count — the dialog stays open even if it drops to 0 (see spec). @@ -62,12 +72,21 @@ export function WorkspaceTeardownModal({ const runningCount = useSyncExternalStore(subscribeToTerminalPaneState, getRunningCount); const hasRunning = runningCount > 0; + if (failure) { + return +

Could not close window

+

{failure.reason}

+ + +
; + } + if (confirming) { return (

Confirm kill workspace

- {intent.kind === 'quit' ? 'Waiting for all windows, then closing…' : 'Closing workspaces…'} + {progressDetail ?? (intent.kind === 'quit' ? 'Waiting for all windows, then closing…' : 'Closing workspaces…')}

); diff --git a/standalone/src/quit-confirm-store.ts b/standalone/src/quit-confirm-store.ts index 846fce41d..0f21c420c 100644 --- a/standalone/src/quit-confirm-store.ts +++ b/standalone/src/quit-confirm-store.ts @@ -12,7 +12,23 @@ import type { TeardownConfirmContext } from "./teardown-flow"; * docs/specs/standalone.md §Quit flow, "Confirmation dialog". */ -export type QuitConfirmPhase = "open" | "quitting"; +export type QuitConfirmPhase = "open" | "quitting" | "close-failed"; +export interface CloseFailure { reason: string; retry: () => void; stay: () => void; } +let closeFailure: CloseFailure | null = null; +let progressDetail: string | null = null; +export function getCloseFailure(): CloseFailure | null { return closeFailure; } +export function getQuitProgressDetail(): string | null { return progressDetail; } +export function showCloseFailure(failure: CloseFailure): void { + closeFailure = failure; + intent = { kind: 'close-window' }; + phase = 'close-failed'; + ownDialog(); + emit(); +} +export function showCloseCommitUncertain(reason: string): void { + progressDetail = reason; + emit(); +} /** * What the dialog is asking about. A quit tears every window down; a @@ -127,6 +143,8 @@ export function openQuitConfirm(ctx: TeardownConfirmContext, next: QuitConfirmIn /** Own the window before voting, including an all-idle request. */ export function beginQuitProgress(next: QuitConfirmIntent): void { + closeFailure = null; + progressDetail = null; stopWatchingWorkspaces(); activeCtx = null; intent = next; @@ -171,6 +189,8 @@ export function dismissQuitConfirm(kind?: QuitConfirmIntent["kind"]): void { activeCtx = null; phase = null; intent = QUIT_INTENT; + closeFailure = null; + progressDetail = null; emit(); } @@ -183,4 +203,6 @@ export function _resetQuitConfirmForTesting(): void { intent = QUIT_INTENT; activeCtx = null; listeners.clear(); + closeFailure = null; + progressDetail = null; } diff --git a/standalone/src/quit.ts b/standalone/src/quit.ts index 70d15a0b6..96770e804 100644 --- a/standalone/src/quit.ts +++ b/standalone/src/quit.ts @@ -136,7 +136,7 @@ async function runQuitTeardown(last: boolean): Promise { `[quit] teardown exceeded ${QUIT_TEARDOWN_CEILING_MS}ms; proceeding to exit`, ); } - // Install strictly after the completed final save, in the window the walk + // Install after bounded flush/drain attempts, in the window the walk // tears down last. Only `main` ever holds a pending download and the // updater capability (docs/specs/auto-update.md), so this is `main` or a // no-op. A fresh `quit_progress` gives install its own watchdog budget diff --git a/standalone/src/tauri-adapter.ts b/standalone/src/tauri-adapter.ts index e7e97a1fd..7ea0b6d81 100644 --- a/standalone/src/tauri-adapter.ts +++ b/standalone/src/tauri-adapter.ts @@ -594,12 +594,17 @@ export class TauriAdapter implements PlatformAdapter { this.pendingFlushRequests.set(requestId, resolve); // Timeout is a synthetic completion; a stale timer after a real completion // hits notify's map-miss guard. Fan out after registering so a synchronous - // completion still finds the entry (first notify wins — one Wall ships). + // completion still finds the entry. WorkspaceWindow's single listener + // waits for every Wall before reporting completion for this window. setTimeout(() => this.notifySessionFlushComplete(requestId), timeoutMs); for (const handler of this.flushHandlers) handler({ requestId, ...options }); }); } + /** Keep one latest-value retry after close rollback, including a refused + * write that outlives the bounded drain. Never retries a successful save. */ + retrySessionSave(): void { this.sessionStore.retryLatest(); } + // Await the session store's in-flight/pending save_session pipeline (the Rust // temp+fsync+rename that actually reaches disk). Bounded: on timeout resolve // anyway rather than wedge quit. diff --git a/standalone/src/tauri-session-store.test.ts b/standalone/src/tauri-session-store.test.ts index 0f1d4479d..88dec4f97 100644 --- a/standalone/src/tauri-session-store.test.ts +++ b/standalone/src/tauri-session-store.test.ts @@ -1,4 +1,4 @@ -import { describe, it, expect } from "vitest"; +import { describe, it, expect, vi } from "vitest"; import { TauriSessionStore } from "./tauri-session-store"; const tick = () => new Promise((r) => setTimeout(r, 0)); @@ -128,6 +128,26 @@ describe("TauriSessionStore", () => { expect(saved).toEqual(["a", "a"]); }); + it("persists an unchanged cache after close preparation refused its write and the window stayed open", async () => { + let closing = true; + let persisted = "previous"; + const store = new TauriSessionStore(async (value) => { + // Rust save_open_window_session refuses rather than acknowledging a + // skipped write while close preparation owns the snapshot. + if (closing) throw new Error("window close is holding its snapshot; no session was saved"); + persisted = value; + }); + store.hydrate(persisted); + store.setItem("k", "newer"); + await store.drain(); + expect(persisted).toBe("previous"); + expect(store.getItem("k")).toBe("newer"); + closing = false; + store.setItem("k", "newer"); + await store.drain(); + expect(persisted).toBe("newer"); + }); + it("drain resolves only after an in-flight save settles", async () => { let release!: () => void; const store = new TauriSessionStore( @@ -193,4 +213,35 @@ describe("TauriSessionStore", () => { await tick(); expect(drained).toBe(true); }); + it("remembers only one idle retry and uses the latest queued value", async () => { + const log = vi.spyOn(console, 'error').mockImplementation(() => {}); + try { + let rejectFirst!: (error: Error) => void; + const values: string[] = []; + const store = new TauriSessionStore(async (value) => { + values.push(value); + if (values.length === 1) await new Promise((_, reject) => { rejectFirst = reject; }); + else throw new Error('still unavailable'); + }); + store.hydrate('previous'); + store.setItem('', 'first'); + store.retryLatest(); + store.setItem('', 'newest'); + rejectFirst(new Error('refused')); + await store.drain(); + expect(values).toEqual(['first', 'newest']); + } finally { log.mockRestore(); } + }); + + it("does not duplicate an in-flight successful save when retry is remembered", async () => { + let complete!: () => void; + const save = vi.fn(() => new Promise((resolve) => { complete = resolve; })); + const store = new TauriSessionStore(save); + store.setItem('', 'latest'); + store.retryLatest(); + complete(); + await store.drain(); + expect(save).toHaveBeenCalledOnce(); + }); + }); diff --git a/standalone/src/tauri-session-store.ts b/standalone/src/tauri-session-store.ts index 89b8d9d8c..134936d32 100644 --- a/standalone/src/tauri-session-store.ts +++ b/standalone/src/tauri-session-store.ts @@ -32,6 +32,7 @@ export class TauriSessionStore implements SessionKeyValueStore { // queued. A queued `value` is always a JSON string (never JS null), so null is // a safe "nothing pending" sentinel — even an empty-string blob is distinct. private pending: string | null = null; + private retryLatestWhenIdle = false; // Resolvers for pending drain() calls, fired when the pipeline next goes idle. private drainWaiters: Array<() => void> = []; @@ -49,6 +50,14 @@ export class TauriSessionStore implements SessionKeyValueStore { }); } + /** Remember one retry of the latest unsaved value, even if a bounded drain + * returned before the current save settled. Newer setItem values replace it; + * a successful current write needs no duplicate, and failure never loops. */ + retryLatest(): void { + if (this.saveInFlight) this.retryLatestWhenIdle = true; + else if (this.cache !== null && this.cache !== this.savedValue) this.setItem('', this.cache); + } + /** Seed the cache from the host's persisted blob (or null) at boot. */ hydrate(seed: string | null): void { this.cache = seed; @@ -81,10 +90,14 @@ export class TauriSessionStore implements SessionKeyValueStore { .then(() => { this.savedValue = value; }) .catch((err) => console.error("[tauri-session-store] save_session failed:", err)) .finally(() => { + const retryLatest = this.retryLatestWhenIdle; + this.retryLatestWhenIdle = false; if (this.pending !== null) { const next = this.pending; this.pending = null; this.flush(next); + } else if (retryLatest && this.cache !== null && this.cache !== this.savedValue) { + this.flush(this.cache); } else { this.saveInFlight = false; // Pipeline idle: release drain waiters. diff --git a/standalone/src/updater.test.ts b/standalone/src/updater.test.ts index 05571c864..1ee5a6e24 100644 --- a/standalone/src/updater.test.ts +++ b/standalone/src/updater.test.ts @@ -121,6 +121,53 @@ describe('updater', () => { mocks.platform = { requestAppRestart: mocks.requestAppRestart, burrow: { command: mocks.burrowCommand } }; }); + it('does not reoffer an approved download when the delayed launch check begins', async () => { + mocks.check.mockResolvedValue(makeUpdate('0.5.0')); + startUpdateCheck(); + await vi.advanceTimersByTimeAsync(0); + checkNow(); + await vi.advanceTimersByTimeAsync(0); + approveUpdate(); + await vi.advanceTimersByTimeAsync(0); + expect(hasPendingUpdate()).toBe(true); + + await vi.advanceTimersByTimeAsync(5_000); + expect(mocks.check).toHaveBeenCalledOnce(); + expect(readBannerState()).toEqual({ status: 'downloaded', version: '0.5.0' }); + }); + + it('does not reoffer an approval while its download is pending at the launch check', async () => { + const update = makeUpdate('0.5.0'); + update.download.mockImplementation(() => new Promise(() => {})); + mocks.check.mockResolvedValue(update); + startUpdateCheck(); + await vi.advanceTimersByTimeAsync(0); + checkNow(); + await vi.advanceTimersByTimeAsync(0); + approveUpdate(); + await vi.advanceTimersByTimeAsync(5_000); + + expect(mocks.check).toHaveBeenCalledOnce(); + expect(readBannerState()).toEqual({ status: 'downloading', version: '0.5.0' }); + }); + + it('keeps an approval made while the delayed launch policy read is pending', async () => { + let answerPolicy!: (value: typeof CHECKS_ON) => void; + mocks.burrowCommand.mockImplementation(() => new Promise(resolve => { answerPolicy = resolve; })); + mocks.check.mockResolvedValue(makeUpdate('0.5.0')); + startUpdateCheck(); + await vi.advanceTimersByTimeAsync(5_000); + checkNow(); + await vi.advanceTimersByTimeAsync(0); + approveUpdate(); + await vi.advanceTimersByTimeAsync(0); + answerPolicy(CHECKS_ON); + await vi.advanceTimersByTimeAsync(0); + + expect(mocks.check).toHaveBeenCalledOnce(); + expect(readBannerState()).toEqual({ status: 'downloaded', version: '0.5.0' }); + }); + // Drive check → approve → download so an approved, downloaded update is pending. async function reachDownloadedUpdate(update: ReturnType) { mocks.check.mockResolvedValue(update); diff --git a/standalone/src/updater.ts b/standalone/src/updater.ts index 4c6c91715..61a2ef758 100644 --- a/standalone/src/updater.ts +++ b/standalone/src/updater.ts @@ -414,6 +414,9 @@ async function runUpdateCheck(): Promise { // Read at the check, so a change made meanwhile counts // (`docs/specs/remote-network.md` → "Updates"). const policy = await readNetworkPolicy(); + // A manual approval during the launch delay or policy lookup owns this + // session's update. Rechecking would offer it for approval a second time. + if (pendingUpdate || downloadPromise) return; if (policy && checksForUpdates(policy)) { // An update found is offered by `performCheck`. await performCheck().catch((e) => console.error('[updater] Check failed:', e)); diff --git a/standalone/src/window-close.test.ts b/standalone/src/window-close.test.ts index 40d8eadd4..07f3d3f74 100644 --- a/standalone/src/window-close.test.ts +++ b/standalone/src/window-close.test.ts @@ -1,4 +1,8 @@ import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { flushWindowSession, installWindowSessionWriter, resetWindowSessionAggregator, seedWindowSession } from "dormouse-lib/lib/window-session-aggregator"; +import { withTimeout } from "./with-timeout"; +import { TauriSessionStore } from "./tauri-session-store"; +import type { PersistedWindow } from "dormouse-lib/lib/session-types"; import type { TauriAdapter } from "./tauri-adapter"; /** @@ -12,7 +16,7 @@ const mocks = vi.hoisted(() => ({ invoke: vi.fn(async (_cmd: string) => undefined as unknown), listen: vi.fn(), countRunningSessions: vi.fn(() => 0), - getWorkspacesSnapshot: vi.fn(() => ({ workspaces: [{ id: "w1", name: "Deploys" }], activeId: "w1" })), + getWorkspacesSnapshot: vi.fn(() => ({ workspaces: [{ id: "w1", name: "Deploys", nameIsAuto: false }], activeId: "w1" })), hasPendingUpdate: vi.fn(() => false), })); @@ -37,6 +41,9 @@ import { cancelQuit as dismissDialog, getQuitConfirmIntent, getQuitConfirmPhase, + getCloseFailure, + getQuitProgressDetail, + confirmQuit, _resetQuitConfirmForTesting, } from "./quit-confirm-store"; @@ -48,6 +55,8 @@ const commands = () => mocks.invoke.mock.calls.map((call) => call[0]); function fakeAdapter(order: string[] = []): TauriAdapter { return { gracefulKillPtys: vi.fn(async () => void order.push("gracefulKill")), + retrySessionSave: vi.fn(), + drainSessionSaves: vi.fn(async () => void order.push('drain')), captureAgentRecovery: vi.fn(async () => void order.push("captureRecovery")), } as unknown as TauriAdapter; } @@ -55,6 +64,7 @@ function fakeAdapter(order: string[] = []): TauriAdapter { describe("per-window close", () => { beforeEach(() => { vi.clearAllMocks(); + resetWindowSessionAggregator(); _resetWindowCloseForTesting(); _resetQuitConfirmForTesting(); listeners.clear(); @@ -67,7 +77,7 @@ describe("per-window close", () => { mocks.hasPendingUpdate.mockReturnValue(false); }); - afterEach(() => _resetWindowCloseForTesting()); + afterEach(() => { _resetWindowCloseForTesting(); resetWindowSessionAggregator(); }); it("acks, removes the snapshot, kills, and proceeds — with no recovery capture", async () => { const order: string[] = []; @@ -166,4 +176,187 @@ describe("per-window close", () => { expect(commands()).toContain("close_window"); }); + + it("retains live PTYs after failed snapshot removal and permits a fresh retry", async () => { + let refusals = 1; + mocks.invoke.mockImplementation(async (cmd) => { + if (cmd === 'remove_window_session' && refusals-- > 0) throw new Error('disk full'); + if (cmd === 'retry_window_close') closeRequested(); + }); + const adapter = fakeAdapter(); + initWindowClose(adapter); + closeRequested(); + await settle(); + expect(getQuitConfirmPhase()).toBe('close-failed'); + expect(getCloseFailure()?.reason).toContain('disk full'); + expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); + expect(commands()).not.toContain('close_window'); + expect(commands()).toContain('window_close_cancel'); + getCloseFailure()?.retry(); + await settle(); + expect(commands().filter((cmd) => cmd === 'remove_window_session')).toHaveLength(2); + expect(commands()).toContain('retry_window_close'); + expect(adapter.gracefulKillPtys).toHaveBeenCalledOnce(); + expect(commands()).toContain('close_window'); + }); + + it("drains a refused save and republishes unchanged aggregate state after close rollback", async () => { + const errorLog = vi.spyOn(console, 'error').mockImplementation(() => {}); + try { + let rejectRefused!: (error: Error) => void; + let disk = 'previous'; + let saveCalls = 0; + const store = new TauriSessionStore(async (value) => { + if (++saveCalls === 1) await new Promise((_, reject) => { rejectRefused = reject; }); + disk = value; + }); + store.hydrate(disk); + const latest: PersistedWindow = { version: 1, activeWorkspaceId: 'w1', workspaces: [ + { id: 'w1', name: 'Deploys', nameIsAuto: false, session: { version: 3, panes: [] } }, + ] }; + seedWindowSession(latest); + installWindowSessionWriter((snapshot) => store.setItem('state', JSON.stringify(snapshot))); + await flushWindowSession(); // the native close fence has refused this pending save + const order: string[] = []; + mocks.invoke.mockImplementation(async (cmd) => { + order.push(cmd); + if (cmd === 'remove_window_session') throw new Error('disk full'); + }); + const adapter = fakeAdapter(order); + adapter.drainSessionSaves = vi.fn(async () => { order.push('drain'); await store.drain(); }); + adapter.retrySessionSave = () => store.retryLatest(); + initWindowClose(adapter); + closeRequested(); + await settle(); + expect(order).toEqual(['window_close_ack', 'remove_window_session', 'window_close_cancel', 'drain']); + expect(saveCalls).toBe(1); + expect(getQuitConfirmPhase()).toBe('quitting'); + rejectRefused(new Error('window close is holding its snapshot; no session was saved')); + await settle(); + expect(saveCalls).toBe(2); + expect(JSON.parse(disk)).toMatchObject(latest); + expect(adapter.drainSessionSaves).toHaveBeenCalledTimes(2); + expect(getQuitConfirmPhase()).toBe('close-failed'); + expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); + } finally { errorLog.mockRestore(); } + }); + + it("retries after a refused save outlives both bounded recovery drains", async () => { + vi.useFakeTimers(); + const errorLog = vi.spyOn(console, 'error').mockImplementation(() => {}); + const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); + try { + let rejectFirst!: (error: Error) => void; + let disk = 'previous'; + let saves = 0; + const store = new TauriSessionStore(async (value) => { + if (++saves === 1) await new Promise((_, reject) => { rejectFirst = reject; }); + disk = value; + }); + store.hydrate(disk); + seedWindowSession({ version: 1, activeWorkspaceId: 'w1', workspaces: [ + { id: 'w1', name: 'Deploys', nameIsAuto: false, session: { version: 3, panes: [] } }, + ] }); + installWindowSessionWriter((snapshot) => store.setItem('', JSON.stringify(snapshot))); + await flushWindowSession(); + mocks.invoke.mockImplementation(async (cmd) => { + if (cmd === 'remove_window_session') throw new Error('disk full'); + }); + const adapter = fakeAdapter(); + adapter.drainSessionSaves = (ms) => withTimeout(store.drain(), ms, 'test drain timeout'); + adapter.retrySessionSave = () => store.retryLatest(); + initWindowClose(adapter); + closeRequested(); + await vi.advanceTimersByTimeAsync(4001); + expect(getQuitConfirmPhase()).toBe('close-failed'); + expect(saves).toBe(1); + expect(disk).toBe('previous'); + rejectFirst(new Error('close refused the earlier write')); + await vi.advanceTimersByTimeAsync(0); + expect(saves).toBe(2); + expect(JSON.parse(disk).workspaces[0].name).toBe('Deploys'); + await vi.advanceTimersByTimeAsync(60_000); + expect(saves).toBe(2); // one remembered retry, no loop + } finally { errorLog.mockRestore(); warn.mockRestore(); vi.useRealTimers(); } + }); + + it("keeps the flow guarded if native cancellation does not confirm release", async () => { + mocks.invoke.mockImplementation(async (cmd) => { + if (cmd === 'remove_window_session') throw new Error('disk full'); + if (cmd === 'window_close_cancel') throw new Error('native bridge unavailable'); + }); + const adapter = fakeAdapter(); + initWindowClose(adapter); + closeRequested(); + await settle(); + expect(getQuitConfirmPhase()).toBe('quitting'); + expect(getQuitProgressDetail()).toContain('save refusal could not be released'); + expect(getCloseFailure()).toBeNull(); + expect(adapter.drainSessionSaves).not.toHaveBeenCalled(); + closeRequested(); + await settle(); + expect(commands().filter((cmd) => cmd === 'remove_window_session')).toHaveLength(1); + }); + + it("bounds the aggregate resave without discarding retained live PTYs", async () => { + vi.useFakeTimers(); + const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); + try { + installWindowSessionWriter(() => new Promise(() => {})); + mocks.invoke.mockImplementation(async (cmd) => { + if (cmd === 'remove_window_session') throw new Error('disk full'); + }); + const adapter = fakeAdapter(); + initWindowClose(adapter); + closeRequested(); + await vi.advanceTimersByTimeAsync(1001); + expect(getQuitConfirmPhase()).toBe('close-failed'); + expect(adapter.drainSessionSaves).toHaveBeenCalledTimes(2); + expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); + expect(commands()).not.toContain('close_window'); + } finally { warn.mockRestore(); vi.useRealTimers(); } + }); + + it("leaves an entered uncertain commit guarded without offering a duplicate close", async () => { + mocks.invoke.mockImplementation(async (cmd) => { + if (cmd === 'remove_window_session') throw new Error('close-commit-uncertain: rollback failed'); + }); + const adapter = fakeAdapter(); + initWindowClose(adapter); + closeRequested(); + await settle(); + expect(getQuitConfirmPhase()).toBe('quitting'); + expect(getQuitProgressDetail()).toContain('rollback failed'); + expect(getCloseFailure()).toBeNull(); + closeRequested(); + await settle(); + expect(commands().filter((cmd) => cmd === 'remove_window_session')).toHaveLength(1); + expect(commands()).not.toContain('window_close_cancel'); + expect(adapter.drainSessionSaves).not.toHaveBeenCalled(); + expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); + }); + + it("does not time out a human confirmation or an entered native commit", async () => { + vi.useFakeTimers(); + try { + mocks.countRunningSessions.mockReturnValue(1); + let resolveRemoval!: () => void; + mocks.invoke.mockImplementation((cmd) => cmd === 'remove_window_session' + ? new Promise((resolve) => { resolveRemoval = resolve; }) : Promise.resolve()); + const adapter = fakeAdapter(); + initWindowClose(adapter); + closeRequested(); + await vi.advanceTimersByTimeAsync(60_000); + expect(getQuitConfirmPhase()).toBe('open'); + expect(commands()).not.toContain('remove_window_session'); + confirmQuit(); + await vi.advanceTimersByTimeAsync(60_000); + expect(getQuitConfirmPhase()).toBe('quitting'); + expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); + expect(commands()).not.toContain('close_window'); + resolveRemoval(); + await vi.advanceTimersByTimeAsync(0); + expect(adapter.gracefulKillPtys).toHaveBeenCalledOnce(); + } finally { vi.useRealTimers(); } + }); }); diff --git a/standalone/src/window-close.ts b/standalone/src/window-close.ts index 119290b00..fd205eb06 100644 --- a/standalone/src/window-close.ts +++ b/standalone/src/window-close.ts @@ -1,6 +1,7 @@ import { invoke } from "@tauri-apps/api/core"; +import { flushWindowSession } from "dormouse-lib/lib/window-session-aggregator"; import { countRunningSessions } from "dormouse-lib/lib/terminal-registry"; -import { openQuitConfirm } from "./quit-confirm-store"; +import { dismissQuitConfirm, openQuitConfirm, showCloseFailure, showCloseCommitUncertain } from "./quit-confirm-store"; import { createTeardownFlow } from "./teardown-flow"; import type { TauriAdapter } from "./tauri-adapter"; import { hasPendingUpdate } from "./updater"; @@ -23,8 +24,10 @@ import { listenToWindow } from "./window-label"; */ const GRACEFUL_KILL_MS = 2000; -/** The whole teardown, past the human decision. Well under Rust's own budget. */ +/** PTY teardown after durable removal; native preparation has its own bound. */ const CLOSE_TEARDOWN_CEILING_MS = 8000; +const RETAINED_WINDOW_DRAIN_MS = 2000; +const RETAINED_WINDOW_WRITE_MS = 1000; let closeAdapter: TauriAdapter | null = null; @@ -49,15 +52,65 @@ export function initWindowClose(adapter: TauriAdapter): void { ...(hasPendingUpdate() ? { discardsUpdate: true } : {}), }); }); + void listenToWindow("dormouse://window-close-failed", (event) => { + void retainWindowAfterFailure(event.payload); + }); +} + +async function retainWindowAfterFailure(error: unknown): Promise { + const reason = String(error); + if (reason.includes('close-commit-uncertain:')) { + // The native commit still owns its files: never release the arbiter or + // enable a resave/second close while a late unlink may finish. + showCloseCommitUncertain(`Closing could not finish safely: ${reason}`); + return; + } + // Retire the previous native handshake before offering a retry. A late + // cancel must never clear the retry's fresh token while its dialog is open. + try { + await invoke('window_close_cancel'); + } catch (cancelError) { + showCloseCommitUncertain('The Window is retained, but its save refusal could not be released: ' + String(cancelError)); + return; + } + // A refused in-flight write must settle before publishing the same value: + // the synchronous cache coalesces identical values while its save is pending. + // Wall dirty tracking already ended when it published into the aggregate; + // a heartbeat may never republish this retained Window without this retry. + try { + if (closeAdapter) await closeAdapter.drainSessionSaves(RETAINED_WINDOW_DRAIN_MS); + await withTimeout( + flushWindowSession(), + RETAINED_WINDOW_WRITE_MS, + '[window-close] retained Window write timed out; Window stays open', + ); + closeAdapter?.retrySessionSave(); + if (closeAdapter) await closeAdapter.drainSessionSaves(RETAINED_WINDOW_DRAIN_MS); + } catch (saveError) { + console.warn('[window-close] retained Window resave failed:', saveError); + } + flow.reset(); + showCloseFailure({ + reason: `Workspaces were retained. ${reason}`, + retry: () => { + dismissQuitConfirm('close-window'); + void invoke('retry_window_close').catch(retainWindowAfterFailure); + }, + stay: () => dismissQuitConfirm('close-window'), + }); } async function runCloseTeardown(): Promise { const adapter = closeAdapter; try { - // Remove the snapshot BEFORE the kill, so an exit-triggered save cannot - // write it back: Rust refuses every later save for this label. - await invoke("remove_window_session").catch((err) => - console.warn("[window-close] remove_window_session failed; proceeding", err)); + // Cancellation can win only before native disk mutation. A failure retains + // this window and its live PTYs; the old best-effort path lost recovery. + await invoke("remove_window_session"); + } catch (error) { + await retainWindowAfterFailure(error); + return; + } + try { // No `ids`: Rust scopes the kill to this window's own PTYs, and a sibling's // terminals must never be reachable from here. if (adapter) { diff --git a/standalone/src/workspace-move.test.ts b/standalone/src/workspace-move.test.ts index e679fd352..0361de2a2 100644 --- a/standalone/src/workspace-move.test.ts +++ b/standalone/src/workspace-move.test.ts @@ -89,6 +89,7 @@ import { disposeAllSessions, getOrCreateTerminal } from "dormouse-lib/lib/termin import { FakePtyAdapter } from "dormouse-lib/lib/platform/fake-adapter"; import { createAlertEpisode } from "dormouse-lib/lib/alert-episode"; import { clearTerminalActivity, getActivitySnapshot, setTerminalActivity } from "dormouse-lib/lib/session-activity-store"; +import { getWorkspaceUiSnapshot, resetWorkspaceUi } from 'dormouse-lib/lib/workspace-ui-store'; const WORKSPACE_ID = "ws-moving"; @@ -234,6 +235,19 @@ const emit = async (event: string, data: unknown) => { }; describe("the source half", () => { + it('keeps a failed durable handback pending while making its retry visible', async () => { + resetWorkspaceUi(); + registerWallHandle(stubWallHandle(WORKSPACE_ID, { prepareWorkspaceTransfer: async () => prepared() })); + initWorkspaceMoves(fakePlatform()); + const moved = transferWorkspaceTo(WORKSPACE_ID, 'ws-2'); + await contentSent(); + await emit('dormouse://workspace-arrival-retry', { workspaceId: WORKSPACE_ID, reason: 'Waiting for recovery write' }); + expect(getWorkspaceUiSnapshot().moveError).toEqual({ id: WORKSPACE_ID, reason: 'Waiting for recovery write' }); + expect(isWorkspaceTransferPending(WORKSPACE_ID)).toBe(true); + await emit('dormouse://workspace-arrival-failed', { workspaceId: WORKSPACE_ID, reason: 'returned', replayIds: [] }); + await expect(moved).resolves.toEqual({ moved: false, reason: 'returned' }); + expect(isWorkspaceTransferPending(WORKSPACE_ID)).toBe(false); + }); it.each([true, false])("refuses dirty editors without explicit discard (existing window: %s)", async (existing) => { const prepare = vi.fn(async () => prepared()); registerWallHandle(stubWallHandle(WORKSPACE_ID, { @@ -632,6 +646,23 @@ describe("the source half", () => { }); describe("the target half", () => { + it('requests handback promptly when durable adoption fails', async () => { + const platform = fakePlatform(); + const kill = vi.spyOn(platform, 'killPty'); + const host = mocks.invoke.getMockImplementation()!; + mocks.invoke.mockImplementation(async (command, args) => { + if (command === 'adopt_done') throw new Error('journal unavailable'); + return host(command, args); + }); + arrivals = [payload()]; + initWorkspaceMoves(platform); + await settle(); + await settle(); + expect(mocks.invoke).toHaveBeenCalledWith('adopt_failed', { workspaceId: WORKSPACE_ID, reason: 'journal unavailable' }); + expect(arrivals).toHaveLength(0); + expect(getWorkspacesSnapshot().workspaces.map(workspace => workspace.id)).not.toContain(WORKSPACE_ID); + expect(kill).not.toHaveBeenCalled(); + }); // The Sessions' alert state never left the sidecar's one manager, which // re-sends it to whichever window collects them. A spawn starts it over and a @@ -747,7 +778,11 @@ describe("the target half", () => { expect(prepare).not.toHaveBeenCalled(); expect(getTerminalInstance("pane-a")).toBeNull(); expect(killPty).not.toHaveBeenCalled(); - expect(mocks.invoke).not.toHaveBeenCalledWith("adopt_failed", expect.anything()); + // A refused settled write explicitly requests durable handback. The + // watchdog may already have returned it; that stale request is harmless. + expect(mocks.invoke).toHaveBeenCalledWith("adopt_failed", { + workspaceId: WORKSPACE_ID, reason: `no arrival of '${WORKSPACE_ID}'`, + }); }); it("discards this window's copy of the alert state when adopt_done is refused before the Wall mounts", async () => { diff --git a/standalone/src/workspace-move.ts b/standalone/src/workspace-move.ts index 710b7b1b0..07d8791b9 100644 --- a/standalone/src/workspace-move.ts +++ b/standalone/src/workspace-move.ts @@ -1,6 +1,6 @@ import { clearToolDirty, recordToolDirty } from 'dormouse-lib/lib/tool-dirty-store'; import { UNSAVED_TOOL_MOVE_REFUSAL } from 'dormouse-lib/lib/tool-editor'; -import { dismissWorkspaceUi } from 'dormouse-lib/lib/workspace-ui-store'; +import { dismissWorkspaceUi, setWorkspaceMoveError } from 'dormouse-lib/lib/workspace-ui-store'; import { restoreToolParams } from 'dormouse-lib/components/wall/tool-transfer'; import { recordToolAnnounce } from 'dormouse-lib/lib/tool-announce-store'; import { invoke } from "@tauri-apps/api/core"; @@ -511,6 +511,7 @@ async function adoptWorkspace(platform: PlatformAdapter, payload: MovePayload): // new move here can fail on a Tool that is still starting. console.error("[workspace-move] adopt_done refused; unwinding the mount", err); discardArrival(payload); + settle("adopt_failed", id, reasonOf(err)); } } catch (err) { console.error("[workspace-move] adoption failed; handing the Workspace back", err); @@ -556,6 +557,10 @@ export function initWorkspaceMoves(platform: PlatformAdapter): void { "dormouse://workspace-arrival-failed", (event) => handleArrivalFailed(event.payload.workspaceId, event.payload.reason ?? "no reason given", event.payload.replayIds ?? []), ); + void listenToWindow<{ workspaceId: WorkspaceId; reason: string }>( + "dormouse://workspace-arrival-retry", + (event) => setWorkspaceMoveError({ id: event.payload.workspaceId, reason: event.payload.reason }), + ); // Immediately, and not only on the nudge: a Workspace dropped on this window // while it was still booting is already in the queue, and its `emit_to` // reached no listener. @@ -596,6 +601,7 @@ export async function bootFromTearOut(platform: PlatformAdapter): Promise Date: Thu, 1 Oct 2026 21:34:32 -0700 Subject: [PATCH 02/10] Bound failed Workspace returns without trapping global quit --- docs/specs/standalone.md | 12 +-- docs/specs/standalone.rationale.md | 2 + standalone/src-tauri/src/lib.rs | 139 +++++++++++++++++++++++-- standalone/src-tauri/src/quit_state.rs | 50 +++++++-- standalone/src-tauri/src/routing.rs | 17 +-- 5 files changed, 189 insertions(+), 31 deletions(-) diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index 141d6c08d..a33d5c425 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -738,7 +738,9 @@ below reads that record rather than inferring itself from the suppression map. drop the in-memory record, and emit `workspace-arrival-failed`; the source clears **transferring** and the Workspace is simply still there. With both ends gone the shells are reaped rather than left owned by a dead label, - and the durable recovery record remains. **Must retain a failed return as an in-flight blocker, show its retry status and retry its journal write without adopting or reaping the claimed PTYs.** A late worker acts only on its own phase and generation; `a_failed_reverse_journal_retains_the_return_and_retries_without_losing_recovery` pins this boundary (rationale). + and the durable recovery record remains. **Must bound failed-return retries and then reserve the arrival for cold recovery without returning ownership, adopting it or reaping its claimed PTYs.** Show recovery status; permit global quit/restart through normal confirmation and watchdogs while individual closes and transfers involving either endpoint remain blocked, including after a quit cancellation. Preserve the journal and snapshots on that quit. Workers act only on their phase and generation; `permanent_return_failure_bounds_quit_and_cold_restores_exactly_once` pins this boundary (rationale). + + Source of truth: `failed_arrival_return` in `standalone/src-tauri/src/lib.rs`, `ArrivalQueue` in `standalone/src-tauri/src/quit_state.rs`. - **Must change transfer ownership and source routing under one routing lock**, so output before the mark always reaches the source (`transfer_ownership_and_source_routing_change_together`). @@ -1010,13 +1012,7 @@ teardown at a time. Source of truth: `QuitMachine` in exits** rather than leaving a process with none, and so does a trigger that finds none: parked in `Voting` it would have no window to vote and refuse every later exit. -- **A quit keeps every window's snapshot on disk** — that is what a relaunch - restores from, and the whole difference from a per-window close. **A Workspace - in transfer is in no snapshot until its target publishes it, and neither end's - teardown kills its shells** (§Arrival queue). A quit mid-transfer restores it - at most once: from the target once it has published it, or from a source - handed it back because the target was destroyed first; a source torn down - before its target adopts leaves it in no snapshot. +- **Must keep every Window snapshot on disk during quit and leave in-flight arrivals' claimed shells out of each endpoint's teardown.** Cold recovery follows §Arrival queue. ### Trigger interception diff --git a/docs/specs/standalone.rationale.md b/docs/specs/standalone.rationale.md index 2b4ceb3f0..17fe908e7 100644 --- a/docs/specs/standalone.rationale.md +++ b/docs/specs/standalone.rationale.md @@ -151,6 +151,8 @@ A target whose `adopt_done` is refused already has the arrival payload needed to The Windows audit on 2026-10-01 found that a failed initial journal write was logged before ownership moved anyway, while the source's next snapshot omitted the Workspace. A failed return write likewise removed the close/quit blocker before its durable destination changed. Admission and return reservations now hold the blockers across disk IO and release source state only after the matching generation's journal commit. Native failure tests preserve earlier journal bytes and replay the return into exactly one source snapshot. +Review on 2026-10-02 found that unlimited reverse-write retries also swallowed every quit before its watchdog started. A parked return retains live ownership and the durable journal; global quit preserves snapshots, so it can use normal confirmation and bounded teardown and let cold restore settle that record at its durable destination. Individual close deletes snapshots and remains blocked, as do new transfers at those endpoints. Canceling quit leaves the reservation in place. The initial attempt and five one-second retries allow transient failures to recover without requiring a restart; exhaustion performs no further write or false hand-back. + **Why the mark is stamped in the stream rather than asked for.** A mark fetched by request answers at some instant the sidecar chose, while the source's xterm stands at whatever `pty:data` had reached it — two clocks nothing aligns, so a diff --git a/standalone/src-tauri/src/lib.rs b/standalone/src-tauri/src/lib.rs index 1d3431d49..1b2fc4a4a 100644 --- a/standalone/src-tauri/src/lib.rs +++ b/standalone/src-tauri/src/lib.rs @@ -682,6 +682,7 @@ fn request_quit(app: &AppHandle, intent: QuitIntent) -> bool { append_log("[quit] transfer in progress; quit queued until settlement"); return relaunches; } + let intent = arrivals.take_quit_intent(intent); let labels = window_labels(app); let mut machine = guard(&state.machine); let (seq, actions) = machine.request(&labels, intent); @@ -2785,7 +2786,7 @@ fn commit_initial_arrival_with( Ok(()) } -/// The queue keeps Returning arrivals as close/quit blockers until their +/// The queue keeps Returning arrivals as close/quit blockers during bounded retries until their /// reverse journal write succeeds. Failed writes change neither ownership nor /// the old durable destination, and retries cannot settle a later generation. fn commit_arrival_return_with( @@ -2928,6 +2929,29 @@ fn spawn_arrival_watchdog(app: AppHandle, arrival: &routing::Arrival) { }); } +// Initial attempt plus five one-second retries. Exhaustion retains the live +// ownership fence and durable record for cold recovery, while admitting the +// non-destructive global quit flow. Individual close still deletes snapshots +// and must not pass it. A parked generation never resumes an old writer. +const ARRIVAL_RETURN_RETRIES: u8 = 5; + +#[derive(Debug, PartialEq, Eq)] +enum ArrivalReturnFailure { Retry(u8), RecoveryPending, Stale } + +fn failed_arrival_return( + windows: &WindowState, expected: &routing::Arrival, retries: u8, +) -> ArrivalReturnFailure { + let _disk = guard(&ARRIVAL_DISK_LOCK); + let mut arrivals = guard(&windows.arrivals); + let Some(arrival) = arrivals.iter_mut().find(|arrival| { + arrival.workspace_id == expected.workspace_id && arrival.to == expected.to + && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning + }) else { return ArrivalReturnFailure::Stale; }; + if retries > 0 { return ArrivalReturnFailure::Retry(retries - 1); } + arrival.phase = routing::ArrivalPhase::RecoveryPending; + ArrivalReturnFailure::RecoveryPending +} + /// One arrival will never be adopted: give its shells back to the source and /// tell the source so it clears the Workspace's transferring mark. **The /// Workspace simply stays where it is** — nothing was released, so there is @@ -2942,12 +2966,20 @@ fn spawn_arrival_watchdog(app: AppHandle, arrival: &routing::Arrival) { /// an id the sidecar never stamped goes straight back. /// /// The caller has marked this generation Returning in the queue. It remains -/// there until its reverse destination is durable; no teardown may overtake it. +/// there until its reverse destination is durable or bounded retries park it +/// for cold recovery; only global quit may pass the parked reservation. fn hand_back_arrival( app: &AppHandle, windows: &WindowState, arrival: &routing::Arrival, reason: &str, +) { + hand_back_arrival_with_retries(app, windows, arrival, reason, ARRIVAL_RETURN_RETRIES); +} + +fn hand_back_arrival_with_retries( + app: &AppHandle, windows: &WindowState, arrival: &routing::Arrival, + reason: &str, retries: u8, ) { append_log(format!( "[window] {} never arrived in {} ({reason}); handing it back to {}", @@ -2960,11 +2992,20 @@ fn hand_back_arrival( let marks = match returned { Ok(marks) => marks, Err(error) => { - append_log(format!("[window] hand-back is waiting for its journal: {error}")); + let status = match failed_arrival_return(windows, arrival, retries) { + ArrivalReturnFailure::Retry(remaining) => { + retry_arrival_return(app.clone(), arrival.clone(), reason.to_string(), remaining); + format!("The Workspace could not be returned safely yet. Retrying its recovery write: {error}") + } + ArrivalReturnFailure::RecoveryPending => { + redrive_deferred_teardown(app); + format!("The Workspace recovery write failed. Quit or restart to recover it from its saved transfer record: {error}") + } + ArrivalReturnFailure::Stale => return, + }; + append_log(format!("[window] {status}")); let _ = app.emit_to(arrival.from.as_str(), "dormouse://workspace-arrival-retry", - serde_json::json!({ "workspaceId": arrival.workspace_id, - "reason": format!("The Workspace could not be returned safely yet. Retrying its recovery write: {error}") })); - retry_arrival_return(app.clone(), arrival.clone(), reason.to_string()); + serde_json::json!({ "workspaceId": arrival.workspace_id, "reason": status })); return; } }; @@ -3018,7 +3059,7 @@ fn hand_back_arrival( /// One retry chain per Returning generation. A later drop of the same id is /// independent; it is never adopted, retired or written by this worker. -fn retry_arrival_return(app: AppHandle, expected: routing::Arrival, reason: String) { +fn retry_arrival_return(app: AppHandle, expected: routing::Arrival, reason: String, remaining: u8) { std::thread::spawn(move || { std::thread::sleep(Duration::from_secs(1)); let Some(windows) = app.try_state::() else { return; }; @@ -3026,7 +3067,7 @@ fn retry_arrival_return(app: AppHandle, expected: routing::Arrival, reason: Stri arrival.workspace_id == expected.workspace_id && arrival.to == expected.to && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning }); - if live { hand_back_arrival(&app, &windows, &expected, &reason); } + if live { hand_back_arrival_with_retries(&app, &windows, &expected, &reason, remaining); } }); } @@ -3562,7 +3603,6 @@ fn forget_restart(app: &AppHandle) { // ── Per-window close (docs/specs/standalone.md §Per-window close) ───────────── -// This window's close orchestrator is alive; stand its ack watchdog down. #[tauri::command] fn retry_window_close(app: AppHandle, window: tauri::Window) { // Retry starts a fresh native handshake too: close admission must block @@ -3570,6 +3610,7 @@ fn retry_window_close(app: AppHandle, window: tauri::Window) { request_window_close(&app, window.label()); } +// This window's close orchestrator is alive; stand its ack watchdog down. #[tauri::command] fn window_close_ack(window: tauri::Window, state: tauri::State<'_, QuitState>) { guard(&state.close).ack(window.label()); @@ -4487,7 +4528,7 @@ mod tests { }; use super::routing; use super::guard; - use super::{WindowState, QuitIntent, commit_initial_arrival_with, + use super::{WindowState, QuitIntent, failed_arrival_return, ArrivalReturnFailure, ARRIVAL_RETURN_RETRIES, commit_initial_arrival_with, commit_arrival_return_with, commit_arrival_adoption_with, record_arrival_in_locked_journal, return_arrival_in_locked_journal}; use std::collections::HashSet; @@ -4711,7 +4752,9 @@ mod tests { assert_eq!(guard(&windows.arrivals).defer_quit(&QuitIntent::default()), Some(false)); assert!(routing::arrival_payloads(&guard(&windows.arrivals), "ws-2").is_empty()); + assert_eq!(failed_arrival_return(&windows, &returning, ARRIVAL_RETURN_RETRIES), ArrivalReturnFailure::Retry(ARRIVAL_RETURN_RETRIES - 1)); commit_arrival_return_with(&windows, &returning, || return_arrival_in_locked_journal(dir.path(), &returning)).unwrap(); + assert_eq!(failed_arrival_return(&windows, &returning, 0), ArrivalReturnFailure::Stale); assert_eq!(windows.owned_by("main"), expected.terminal_ids); assert!(guard(&windows.arrivals).is_empty()); restore_arrivals(dir.path()).unwrap(); @@ -4719,6 +4762,82 @@ mod tests { assert!(read_snapshot(dir.path(), "ws-2").is_none()); } + #[test] + fn permanent_return_failure_bounds_quit_and_cold_restores_exactly_once() { + use super::quit_state::{QuitMachine, QuitAction}; + let dir = TempDir::new("arrival-return-exhausted"); + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + write_session_to(dir.path(), "main", &snapshot_json(&[("workspace-7", "Source")], "workspace-7")).unwrap(); + write_session_to(dir.path(), "ws-2", &snapshot_json(&[], "")).unwrap(); + commit_initial_arrival_with(&windows, &expected, || record_arrival_in_locked_journal(dir.path(), &expected)).unwrap(); + let returning = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); + let previous = fs::read(arrivals_path(dir.path())).unwrap(); + let intent = QuitIntent::restart(None); + assert_eq!(guard(&windows.arrivals).defer_quit(&intent), Some(true)); + for remaining in (0..=ARRIVAL_RETURN_RETRIES).rev() { + assert!(commit_arrival_return_with(&windows, &returning, || Err("read-only volume".into())).is_err()); + let result = failed_arrival_return(&windows, &returning, remaining); + assert_eq!(result, if remaining > 0 { ArrivalReturnFailure::Retry(remaining - 1) } + else { ArrivalReturnFailure::RecoveryPending }); + } + assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); + assert_eq!(windows.owned_by("ws-2"), expected.terminal_ids); + assert!(routing::arrival_payloads(&guard(&windows.arrivals), "ws-2").is_empty()); + assert!(commit_arrival_return_with(&windows, &returning, || panic!("parked writer must not touch disk")).is_err()); + assert!(commit_arrival_adoption_with(&windows, &expected, || panic!("late adoption must not write")).is_err()); + let live = || HashSet::from(["main".to_string(), "ws-2".to_string()]); + assert_eq!(guard(&windows.arrivals).take_ready(live), (Some(intent.clone()), vec![])); + assert_eq!(guard(&windows.arrivals).defer_quit(&intent), None); + // Actual global-quit machine takes its normal confirmation, then emits + // destroys without the per-window snapshot commit/removal path. + let mut machine = QuitMachine::default(); + let labels = vec!["main".to_string(), "ws-2".to_string()]; + assert_eq!(machine.request(&labels, intent.clone()).1, vec![QuitAction::RequestAll { requester: None }]); + assert!(!machine.all_acked(), "the normal no-ack watchdog is now reachable"); + assert_eq!(machine.cancel(), vec![QuitAction::CancelAll]); + guard(&windows.arrivals).cancel_deferred(); + assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); + assert_eq!(windows.owned_by("ws-2"), expected.terminal_ids); + assert!(guard(&windows.arrivals).defer_close("main")); + assert!(guard(&windows.arrivals).defer_close("ws-2")); + assert_eq!(guard(&windows.arrivals).take_ready(live), (None, vec![])); + assert_eq!(guard(&windows.arrivals).defer_quit(&intent), None); + assert_eq!(machine.request(&labels, intent.clone()).1, vec![QuitAction::RequestAll { requester: None }]); + machine.ack("main"); machine.ack("ws-2"); + assert!(machine.vote("main").is_empty()); + assert_eq!(machine.vote("ws-2"), vec![QuitAction::Teardown { label: "ws-2".into(), last: false }]); + assert_eq!(machine.window_done("ws-2"), vec![QuitAction::Destroy { label: "ws-2".into() }, + QuitAction::Teardown { label: "main".into(), last: true }]); + let (lost, orphaned) = windows.drop_window("ws-2"); + assert!(lost.is_empty(), "Destroyed must not restart a parked reverse write"); + assert!(orphaned.is_empty(), "claimed PTYs must not be sibling-close kills"); + assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); + assert!(read_snapshot(dir.path(), "main").is_some()); + assert!(read_snapshot(dir.path(), "ws-2").is_some()); + assert!(machine.proceed().contains(&QuitAction::Exit)); + restore_arrivals(dir.path()).unwrap(); + assert!(read_snapshot(dir.path(), "main").is_none(), "empty source is removed by cold restore"); + assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "ws-2").unwrap()), vec!["workspace-7"]); + restore_arrivals(dir.path()).unwrap(); + assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "ws-2").unwrap()), vec!["workspace-7"]); + } + + #[test] + fn exhausted_return_cannot_park_or_write_a_newer_generation() { + let windows = WindowState::default(); + let expected = reserve_test_arrival(&windows); + commit_initial_arrival_with(&windows, &expected, || Ok(())).unwrap(); + let old = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); + assert!(routing::retire_arrival(&mut guard(&windows.arrivals), &old)); + let mut newer = old.clone(); + newer.queued_at += std::time::Duration::from_millis(1); + guard(&windows.arrivals).push(newer.clone()); + assert_eq!(failed_arrival_return(&windows, &old, 0), ArrivalReturnFailure::Stale); + assert_eq!(guard(&windows.arrivals)[0], newer); + assert!(commit_arrival_return_with(&windows, &old, || panic!("stale generation must not write")).is_err()); + } + #[test] fn a_return_retry_survives_target_destruction_without_reaping_source_shells() { let windows = WindowState::default(); diff --git a/standalone/src-tauri/src/quit_state.rs b/standalone/src-tauri/src/quit_state.rs index 2795359b6..6b8d9ef0a 100644 --- a/standalone/src-tauri/src/quit_state.rs +++ b/standalone/src-tauri/src/quit_state.rs @@ -6,7 +6,7 @@ //! transitions live here, free of Tauri, and hand the caller a list of actions //! to perform; `lib.rs` owns the emitting, destroying and exiting. -use crate::routing::{quit_order, Arrivals}; +use crate::routing::{quit_order, ArrivalPhase, Arrivals}; use std::collections::{HashMap, HashSet}; use std::ops::{Deref, DerefMut}; @@ -35,14 +35,21 @@ impl DerefMut for ArrivalQueue { } impl ArrivalQueue { - /// Queue a quit while anything is in flight. `None` means nothing is in - /// flight and the quit runs now; otherwise whether the queued quit relaunches. + /// Only unsettled live work delays global quit. Exhausted returns retain + /// their recovery/individual-close fences, but global quit preserves the + /// journal and snapshots and may proceed through its normal human gate. pub fn defer_quit(&mut self, intent: &QuitIntent) -> Option { - if self.records.is_empty() { return None; } + if !self.blocks_quit() { return None; } let restart = self.quit.get_or_insert_with(|| intent.clone()).restart; self.closes.clear(); Some(restart) } + /// Global quit admission consumes the first queued intent even when a new + /// trigger arrives before the scheduled redrive; individual closes yield. + pub fn take_quit_intent(&mut self, fallback: QuitIntent) -> QuitIntent { + self.closes.clear(); + self.quit.take().unwrap_or(fallback) + } /// Queue `label`'s close while it is either end of a transfer; whether it /// was queued (a queued quit absorbs it). pub fn defer_close(&mut self, label: &str) -> bool { @@ -50,20 +57,25 @@ impl ArrivalQueue { if self.quit.is_none() { self.closes.insert(label.to_string()); } true } + fn blocks_quit(&self) -> bool { + self.records.iter().any(|arrival| arrival.phase != ArrivalPhase::RecoveryPending) + } pub fn blocks_transfer(&self, from: &str, to: &str) -> bool { self.quit.is_some() || self.closes.contains(from) || self.closes.contains(to) + || self.records.iter().any(|arrival| arrival.phase == ArrivalPhase::RecoveryPending + && [from, to].into_iter().any(|label| arrival.from == label || arrival.to == label)) } pub fn cancel_deferred(&mut self) { self.quit = None; self.closes.clear(); } pub fn forget_deferred_close(&mut self, label: &str) { self.closes.remove(label); } /// Take the requests no transfer holds any more: the quit (with its intent) - /// once nothing is in flight, else every close whose window is no endpoint. + /// once no live settlement is in flight, else every close whose window is no endpoint. /// `live` (the open window labels) is read only when something is queued. pub fn take_ready(&mut self, live: impl FnOnce() -> HashSet) -> (Option, Vec) { if self.quit.is_none() && self.closes.is_empty() { return (None, Vec::new()); } let live = live(); self.closes.retain(|label| live.contains(label)); if self.quit.is_some() { - if self.records.is_empty() { return (self.quit.take(), Vec::new()); } + if !self.blocks_quit() { return (self.quit.take(), Vec::new()); } return (None, Vec::new()); } let endpoints: HashSet<&str> = @@ -517,6 +529,32 @@ mod tests { assert!(!queue.blocks_transfer("main", "ws-2")); } + #[test] + fn parked_recovery_allows_only_global_quit_and_keeps_fences_after_cancel() { + let mut queue = ArrivalQueue::default(); + let live = || HashSet::from(["main".to_string(), "ws-2".to_string(), "ws-3".to_string()]); + let mut parked = arrival("main", "ws-2"); + parked.phase = ArrivalPhase::RecoveryPending; + queue.push(parked); + queue.push(arrival("ws-3", "main")); + let intent = QuitIntent::restart(Some("requester".into())); + assert_eq!(queue.defer_quit(&intent), Some(true), "live settlements still block"); + assert_eq!(queue.take_ready(live), (None, vec![])); + queue.pop(); + assert_eq!(queue.defer_quit(&QuitIntent::default()), None, "repeated quit must reach its watchdog"); + assert_eq!(queue.take_quit_intent(QuitIntent::default()), intent, "first queued restart owns intent"); + assert_eq!(queue.take_ready(live), (None, vec![])); + queue.cancel_deferred(); + assert!(queue.defer_close("main")); + assert!(queue.defer_close("ws-2")); + assert_eq!(queue.take_ready(live), (None, vec![]), "no destructive individual close"); + assert!(queue.blocks_transfer("main", "ws-3")); + assert!(queue.blocks_transfer("ws-3", "ws-2")); + assert!(crate::routing::arrival_payloads(&queue, "ws-2").is_empty()); + assert!(crate::routing::return_arrival(&mut queue, "main-to-ws-2", "ws-2").is_none()); + assert!(crate::routing::take_arrivals_to(&mut queue, "ws-2").is_empty()); + } + #[test] fn stalled_cleanup_cannot_block_an_approved_exit_forever() { let mut gate = CleanupGate::default(); diff --git a/standalone/src-tauri/src/routing.rs b/standalone/src-tauri/src/routing.rs index 4474b2fcb..5335c7207 100644 --- a/standalone/src-tauri/src/routing.rs +++ b/standalone/src-tauri/src/routing.rs @@ -247,7 +247,8 @@ pub fn route<'a>(event: &str, data: &'a JsonValue, view: &RouteView<'a>) -> Rout pub struct Arrival { /// Preparing reserves both endpoints while disk I/O runs without the /// arrival lock. Returning retains that reservation until its reverse - /// destination is durable. Neither phase is adoptable. + /// destination is durable. RecoveryPending parks an exhausted return for + /// cold recovery: only global quit may pass it. None is adoptable. pub phase: ArrivalPhase, /// A Preparing arrival that is cancelled has never moved the source's /// shells; its return must not hide or reap those source-owned PTYs. @@ -274,7 +275,7 @@ pub struct Arrival { } #[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum ArrivalPhase { Preparing, Active, Returning } +pub enum ArrivalPhase { Preparing, Active, Returning, RecoveryPending } /// Every arrival in flight, oldest first. A Vec, not a map: there are a handful /// at most, and both the per-window drain and the by-Workspace lookup want the @@ -320,18 +321,19 @@ pub fn expire_arrival( ) -> Option { let arrival = arrivals.iter_mut().find(|arrival| { arrival.workspace_id == workspace_id && arrival.to == to && arrival.queued_at == queued_at - && arrival.phase != ArrivalPhase::Returning + && matches!(arrival.phase, ArrivalPhase::Preparing | ArrivalPhase::Active) })?; arrival.phase = ArrivalPhase::Returning; Some(arrival.clone()) } /// Reserve settlement before releasing the lock. A close/quit must continue -/// waiting while the return destination is being written or retried. +/// waiting while the return destination is being written or retried. A parked +/// RecoveryPending record cannot start a second return chain. pub fn return_arrival(arrivals: &mut Arrivals, workspace_id: &str, to: &str) -> Option { let arrival = arrivals.iter_mut().find(|arrival| { arrival.workspace_id == workspace_id && arrival.to == to - && arrival.phase != ArrivalPhase::Returning + && matches!(arrival.phase, ArrivalPhase::Preparing | ArrivalPhase::Active) })?; arrival.phase = ArrivalPhase::Returning; Some(arrival.clone()) @@ -348,12 +350,13 @@ pub fn retire_arrival(arrivals: &mut Arrivals, expected: &Arrival) -> bool { } /// Reserve returns for arrivals whose target is gone, retaining their teardown -/// blockers until the reverse journal commits. Existing return workers own +/// blockers until the reverse journal commits. Parked recovery stays reserved, +/// and existing return workers own /// their generation and are not started a second time. pub fn take_arrivals_to(arrivals: &mut Arrivals, label: &str) -> Vec { let mut lost = Vec::new(); for arrival in arrivals { - if arrival.to == label && arrival.phase != ArrivalPhase::Returning { + if arrival.to == label && matches!(arrival.phase, ArrivalPhase::Preparing | ArrivalPhase::Active) { arrival.phase = ArrivalPhase::Returning; lost.push(arrival.clone()); } From ac10f2fe3e3b0157068a76272aa4c35cd45b85b8 Mon Sep 17 00:00:00 2001 From: Ned Twigg Date: Thu, 1 Oct 2026 22:43:32 -0700 Subject: [PATCH 03/10] Back out the transfer and close state machines The original change answered two local disk-write failures with about 1,000 lines of new state: four arrival phases, commit_*_with helpers, 5x1s return retries, RecoveryPending parking, CloseCommit with snapshot rollback, retainWindowAfterFailure/retryLatest, and new teardown-modal states. Both failures need the app-data write itself to fail (disk full, antivirus holding the rename), and the transfer case also needs a crash inside a sub-second window. Meanwhile the machinery added ways to get stuck: an undismissable close-commit-uncertain modal with saves refused, RecoveryPending blocking closes and transfers at both windows until restart, a drag that now failed when the settled-journal write did, and (until review caught it) a quit that could hang forever. Close retention also kept a window open behind a Retry modal after the user asked to close it, where the old outcome was only that window reopening on the next launch with nothing lost. Return to the base behaviour; the following commits keep the small, independent fixes and replace the transfer guard with journal-first. Co-Authored-By: Claude Opus 5.5 --- docs/specs/auto-update.md | 4 +- docs/specs/auto-update.rationale.md | 4 - docs/specs/security-remote.md | 2 +- docs/specs/standalone.md | 170 ++-- docs/specs/standalone.rationale.md | 22 +- lib/src/host/remote/burrow-state-store.ts | 3 +- scripts/spec-word-budgets.json | 2 +- standalone/scripts/dev-agent-browser.test.mjs | 15 +- standalone/scripts/dev-standalone.test.mjs | 4 +- standalone/src-tauri/src/close_commit.rs | 69 -- standalone/src-tauri/src/lib.rs | 818 +++--------------- standalone/src-tauri/src/quit_state.rs | 52 +- standalone/src-tauri/src/routing.rs | 73 +- standalone/src/WorkspaceTeardownModal.tsx | 23 +- standalone/src/quit-confirm-store.ts | 24 +- standalone/src/quit.ts | 2 +- standalone/src/tauri-adapter.ts | 7 +- standalone/src/tauri-session-store.test.ts | 53 +- standalone/src/tauri-session-store.ts | 13 - standalone/src/updater.test.ts | 47 - standalone/src/updater.ts | 3 - standalone/src/window-close.test.ts | 197 +---- standalone/src/window-close.ts | 65 +- standalone/src/workspace-move.test.ts | 37 +- standalone/src/workspace-move.ts | 8 +- 25 files changed, 252 insertions(+), 1465 deletions(-) delete mode 100644 standalone/src-tauri/src/close_commit.rs diff --git a/docs/specs/auto-update.md b/docs/specs/auto-update.md index 562d5278a..3da2af03e 100644 --- a/docs/specs/auto-update.md +++ b/docs/specs/auto-update.md @@ -8,9 +8,7 @@ The standalone app checks for updates on launch, where the network policy allows **Must read and clear the post-install marker on launch** (§localStorage) and show its banner; a reported failure suppresses this launch's check. Otherwise wait 5 seconds, then read the network policy with `networkPolicy` over the Burrow link and, where it allows (`docs/specs/remote-network.md` → "Updates"), `check()` — no update is silent, an update raises the approval prompt; then the reminder, if due: `check-due`, recording `remindedAt`. **The reminder is re-evaluated hourly while the app runs**, reading no policy and never checking; **never over an undismissed notice, nor while the clock reads before 2026-09**, not yet set. Version-lookup and check failures are logged. **Only approval starts the background `download()`**; a failed one is logged and the prompt returns. -**Check now** — the `check-due` and `check-failed` links, and the `updates` port — shows `checking`, then `available`, `up-to-date`, or `check-failed`. **Must join a check already in flight and preserve an approved update as `downloading` or `downloaded` instead of checking again.** This applies to both manual and delayed automatic checks, including approval while network-policy lookup is pending. **Every successful check, automatic or asked for, records `checkedAt`** (§localStorage). `standalone/src/updater.test.ts` pins the approval races. - -Source of truth: `runUpdateCheck` and `approveUpdate` in `standalone/src/updater.ts`. +**Check now** — the `check-due` and `check-failed` links, and the `updates` port — shows `checking`, then `available`, `up-to-date`, or `check-failed`. **A second ask joins the check in flight. An update already approved is shown again, `downloading` or `downloaded`, instead of checked for**, which would offer it for approval twice. **Every successful check, automatic or asked for, records `checkedAt`** (§localStorage). **A self-host build never checks** (`docs/specs/relay.md` → "Relay origin"): `startUpdateCheck()` returns at once and `checkNow()` does nothing unless the webview's own baked mode, `bakedRelayMode()`, is `hosted`. diff --git a/docs/specs/auto-update.rationale.md b/docs/specs/auto-update.rationale.md index 8b0a4b6a4..5ae132d41 100644 --- a/docs/specs/auto-update.rationale.md +++ b/docs/specs/auto-update.rationale.md @@ -2,10 +2,6 @@ > Informative companion to [auto-update.md](auto-update.md): evidence and design history keyed by that spec's headings. Nothing here is normative. -## How it works - -On Windows, 2026-10-01, a delayed launch check ran after a manual check had already approved and downloaded the update; it called `check()` again and replaced the downloaded notice with another approval prompt. The same race occurred while network-policy lookup was pending. Checking approved-update ownership after that await preserves the approval throughout both automatic and manual paths. - ## Quit-time install **Why install runs last.** On Windows `install()` starts NSIS and then calls `std::process::exit` itself (`tauri-plugin-updater-2.11.0/src/updater.rs`, checked 2026-09), so starting it early interrupts teardown. This ordering originally protected persisted scrollback; what it protects now is the window's structure, which standalone does persist. The retained save/drain hooks and their completion semantics are explained in `docs/specs/standalone.rationale.md` → Quit flow. diff --git a/docs/specs/security-remote.md b/docs/specs/security-remote.md index 852fa6ee1..abb14b771 100644 --- a/docs/specs/security-remote.md +++ b/docs/specs/security-remote.md @@ -111,7 +111,7 @@ per-Burrow browser storage follows `docs/specs/remote-security-model.md` -> - **FAIL IF** `relay/src/state.ts` stops creating `$DORMOUSE_STATE_DIR` mode `0o700`, or stops writing every file through `writeAtomic` at mode `0o600`. The "every file" clause is a negative search over `relay/src/`: no `writeFile`, `appendFile`, or `createWriteStream` may target the state directory outside `writeAtomic`. A cheap default, not a cross-platform guarantee; the installer's directory permissions below protect the installed Relay's state (rationale). - **FAIL IF** `FileBurrowStateStore` (`lib/src/host/remote/burrow-state-store.ts`) stops creating its directory `0o700` and writing `0o600` on non-Windows platforms, or if `VsCodeBurrowStateStore` stops keeping the **enrollment** in `SecretStorage`. The ACL's home in `globalState` is deliberate and is not a finding; the enrollment's is what carries `burrowToken`. - **FAIL IF** a credential the Host→Burrow rename retired stops being deleted unread at boot: `state/hosts.json` on the Relay (`forgetRetiredState` in `relay/src/state.ts`, called from `relay/src/start.ts`), `remote-host.json` on a Node-resident Burrow (`forgetRetiredState` in `lib/src/host/remote/burrow-state-store.ts`, called from `sidecar-entry.ts`), and `dormouse.remote-host.enrollment`, `dormouse.remote-host.acl.*`, `remote-host.peer-token` in VS Code (`vscode-ext/src/retired-state.ts`, called from `activate()`). Each held a live `burrowToken` or peer secret, and `SecretStorage` cannot be enumerated — a key nothing removes *by name* outlives every build that knew it. Pinned by `relay/test/state-records.test.mjs`, `lib/src/host/remote/burrow-state-store.test.ts` and `vscode-ext/test/retired-state.test.ts`. -- **FAIL IF** `burrow_state_dir` in `standalone/src-tauri/src/lib.rs` publishes durable storage before `restrict_to_owner` succeeds. Windows ignores Node modes: newly written enrollment files must inherit the owner-only DACL, and existing files holding `burrowToken` must receive it by propagation. `restrict_to_owner_leaves_one_owner_only_ace` pins the existing-file case; `burrow_directory_permission_failure_disables_durable_state` pins refusal without changing existing enrollment bytes. A refused restriction must select the in-memory fallback. +- **FAIL IF** `burrow_state_dir` in `standalone/src-tauri/src/lib.rs` stops calling `restrict_to_owner` on the state directory **before** spawning the sidecar — on Windows those Node modes are no-ops and Node cannot set an ACL, so the guarantee is held one layer down. That call carries both legs: a newly written enrollment file *inherits* the owner-only entry, and one a prior version already left under the `%LOCALAPPDATA%` ACL — with a live `burrowToken` in it — has that entry *propagated* onto it, the half `restrict_to_owner_leaves_one_owner_only_ace` covers with its pre-existing `before.json`. - **FAIL IF** `relay/src/start.ts` stops obtaining the setup password from `SetupPasswordStore.loadOrCreate(generateSetupPassword)`, `generateSetupPassword` stops using `crypto.randomBytes(32)`, `readConfig` reads `DORMOUSE_SETUP_PASSWORD` or any other setup-password input, or `SetupPasswordStore` stops refusing a persisted or generated value outside 64 lowercase hexadecimal characters. Pinned by `relay/test/config.test.mjs` and `relay/test/setup-password-store.test.mjs`. - **FAIL IF** `createApp` accepts anything but 64 lowercase hexadecimal characters as the setup password injected by the entrypoint; pinned by `relay/test/app.test.mjs`. - **FAIL IF** any installer stops making `config/`, `state/`, and `config/relay.env` reachable only by the installing user — the effective property `manage verify` tests: no principal other than that user may appear in the effective permissions. macOS and Linux achieve it with `0700`/`0600` under `umask 077`; Windows with a single owner-only ACE, whether the path carries it directly or inherits it from an already-locked parent. The Windows and Linux installers create `relay.env` and lock it before writing its contents (rationale). diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index c4f415e9f..709cc8056 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -101,7 +101,7 @@ constants in `lib/src/lib/platform/types.ts` (and `standalone/sidecar/pty-core.j `#[tauri::command]` over an `async fn`, which the guard below accepts equally. Tauri runs a *sync* command on the main thread, where the `recv_timeout` inside `request_from_sidecar` / `request_from_sidecar_timeout` stops the webview painting -for the whole round trip, up to `BROWSER_REQUEST_TIMEOUT` (40s) (rationale). **The +for the whole round trip, up to `AGENT_BROWSER_TIMEOUT` (30s) (rationale). **The three clipboard readers included**: their non-Windows branches round-trip through the sidecar, and the declaration is per command, not per branch. A unit test in `lib.rs` scans the source and fails on any command that reaches the blocking @@ -138,7 +138,7 @@ side"): the same as `DORMOUSE_STATE_DIR` (§Persistence, "Rust file store"); `FileBurrowStateStore` keeps enrollment and ACL there as **one** `burrow.json`, so a write is one atomic rename (rationale), and the network policy beside it (`docs/specs/remote-network.md` -→ "Policy"), both committed by owner-only temp-then-rename (POSIX modes; inherited Windows DACL). `burrowToken` is +→ "Policy"), both 0600 in a 0700 directory via temp-then-rename. `burrowToken` is a bearer credential and **never enters a webview realm**. Against the shared store contract (`docs/specs/relay.md` → "Burrow side"): - **Reads fail closed.** Only `ENOENT` and a read-but-unparseable file answer @@ -150,7 +150,7 @@ a bearer credential and **never enters a webview realm**. Against the shared sto directory Rust already created is best-effort; failing the save over it would lose the Burrow instead. - **`persistent` is declared, never inferred.** With no state directory — Rust - passes an empty value when it cannot create or restrict one — the fallback store still + passes an empty value when it cannot create one — the fallback store still *holds* both values in memory, warns once, and reports `persistent: false`. The browser dev harness is *not* this case: its per-run temp directory makes a dev enrollment live and die with the run. @@ -571,13 +571,11 @@ checks in the debounce flush. main thread may be driving while it waits inside `window_at_cursor` for that same lock. The flush reads both first and hands the scale to `refresh_rect`, whose signature takes no window at all. -- **Must serialize geometry writes with snapshot removal and recheck the closing fence after platform reads.** A debounce captured before close must not recreate the removed geometry; `a_geometry_flush_captured_before_close_cannot_recreate_removed_geometry` pins the fence. - **The flush slot is released in the same step as the drain.** A `Moved` landing between the two was marked dirty with no thread left to write it — and that move is exactly a window's final position. Source of truth: `CachedRect` / `GeometryState` / `note_geometry` / -`write_open_window_geometry` / `restore_windows` in `standalone/src-tauri/src/lib.rs`; the sequencing is pinned by `the_geometry_flush_slot_is_released_with_the_drain`. @@ -603,41 +601,55 @@ Source of truth: `CleanupGate` and `WindowEvent::Destroyed` in ### Per-window close -**Must close only the invoking Window when siblings remain; the last Window's close is a quit.** - -1. Rust prevents native close and emits `dormouse://window-close-requested` to that Window. -2. The webview acknowledges, obtains consent, removes its snapshot, tears down its PTYs, then invokes `close_window`. -3. Without acknowledgement, Rust attempts close after `CLOSE_ACK_TIMEOUT_MS` (rationale). - -- **Must remove the deliberately closed Window's snapshot, geometry and temp sibling** (`docs/specs/transport.md` -> "The governing rule"). -- **Never capture agent recovery for a deliberate close.** -- **Must retire cancelled watchdog tokens without reuse** (`a_cleared_close_never_hands_its_seq_to_the_next_request`; rationale). -- **Must confirm on pending downloads and running work** (`docs/specs/auto-update.md`). -- **Must commit snapshot removal before killing PTYs and refuse saves for successfully closed labels for the process lifetime**, including saves dispatched before `Destroyed`. -- **Must bound preparation and journal-lock waits to eight seconds after consent, cancelling only before the first mutation.** Cancelled generations cannot remove retained or re-saved Workspaces. **Never time out consent or abandon entered kernel IO**; await its result. -- **Must retain the Window and live PTYs on removal failure, restoring changed journal/geometry before releasing save refusal.** Incomplete rollback retains the committed flow and refusal with a visible error; uncertain commits never resave (rationale). -- **Must confirm native cancellation before resetting the flow**, then drain refused saves, republish the current Window aggregate and drain again with bounded waits before offering Retry / Keep window open. **Must remember one latest-value retry when a refused write outlives those waits.** Failed cancellation keeps the flow guarded (rationale). -- **Must reject skipped saves instead of acknowledging persistence**, permitting unchanged-cache retries (`skipped_close_save_is_rejected_and_retained_window_can_retry`; `standalone/src/tauri-session-store.test.ts`). -- **Must bound PTY teardown after successful removal to eight seconds.** `standalone/src/window-close.test.ts`, `close_commit` tests and `failed_close_rolls_back_journal_and_geometry_before_retry` pin these boundaries. -- **Must finish both deliberate close and last-Workspace departure through `close_window`** (see Transfer). - -**Must share acknowledgement and consent gates with quit** through `createTeardownFlow`: quit votes and waits its walk turn; close tears down immediately. - -**Never leave a competing teardown flow unanswered** (rationale): +**Closing a window with siblings alive ends that window alone**; only the last +window's close is the quit. Rust prevents the close and emits +`dormouse://window-close-requested`; the webview acks (a ~2 s watchdog closes it +anyway if that listener is dead), asks about *its own* running work, removes its snapshot, kills the PTYs it owns, and calls back +`close_window`. + +- **A close is deliberate, so it removes the blob** — geometry + and temp sibling included — and the next launch does not reopen the window + (`docs/specs/transport.md` → "The governing rule"). +- **It runs no agent-recovery capture**: nothing is coming back. +- **A cancelled close retires its watchdog's token and never reuses it**: the + next close on that window is a fresh seq, so a watchdog still sleeping on the + cancelled one cannot destroy the window under the second dialog + (`a_cleared_close_never_hands_its_seq_to_the_next_request`). +- **It confirms on a pending download as well as on running work.** An approved, + downloaded update lives in this webview's memory, so closing the window throws + it away and nothing else can install it (`docs/specs/auto-update.md`). +- **The snapshot is removed before the kill**, and Rust refuses every later save + for that label, so a PTY exit's save cannot write it back. **Both close paths + set that refusal** — the webview's own `remove_window_session`, and + `finish_window_close` for the ack-timeout path, where the webview never ran at + all. It is dropped when the webview is destroyed and can no longer save. +- **`close_window` is the one Rust half both endings share** — a deliberate close + and a window whose last Workspace moved away (§Transfer) — because what + separates them is entirely what the webview did before calling it. +- **macOS keeps its rule**: closing the last window quits. + +**The ack and confirm gates are one shared flow with the quit** +(`createTeardownFlow` in `standalone/src/teardown-flow.ts`); what differs is only +the step past them — a quit votes and waits its turn in the walk, a close tears +down at once. + +**Arbitration.** They are two machines over one window, one dialog and one +human, and **a second flow is never refused in silence**: an unsettled context +parks its own flow and leaves its host waiting out a decision that cannot come. | Arriving | Holder | Outcome | |---|---|---| -| quit | close awaiting consent | cancel close with `window_close_cancel`; quit takes over | -| quit | committed flow | acknowledge and vote | -| close | quit in any state | immediately `window_close_cancel`; retain Window | +| quit | a close still on its dialog | the close is cancelled (`window_close_cancel`); the quit takes over | +| quit | any committed flow | the quit acks and **votes** — this window is ending anyway, and a window that never votes holds the machine in `Voting` with no dialog to answer | +| close | a quit, in any state | refused at once with `window_close_cancel`; the window stays | -**Must drop only quit's dialog when another Window cancels quit, and cancel any context the confirm store cannot open.** `standalone/src/teardown-arbiter.test.ts` pins both orderings. +**A quit cancelled elsewhere drops only a quit's dialog**, never this window's +own close question. The confirm store cancels any context it cannot open, as the +backstop. Both orderings are pinned by +`standalone/src/teardown-arbiter.test.ts`. -Source of truth: `runCloseTeardown` / `retainWindowAfterFailure` in `standalone/src/window-close.ts`; -`retryLatest` in `standalone/src/tauri-session-store.ts`; -`createTeardownFlow` in `standalone/src/teardown-flow.ts`; -`CloseCommit` in `standalone/src-tauri/src/close_commit.rs`; -`remove_window_session_bounded` / `close_window_snapshot_locked` / `request_window_close` / `finish_window_close` in `standalone/src-tauri/src/lib.rs`. +Source of truth: `standalone/src/window-close.ts`; `request_window_close` / +`finish_window_close` in `standalone/src-tauri/src/lib.rs`. ### Transfer @@ -688,7 +700,7 @@ below reads that record rather than inferring itself from the suppression map. `transfer_workspace` / `open_workspace_window`. **Must return preparation refusals as `{ moved: false, reason }` without changing ownership.** On `Ok` it marks the Workspace **transferring**: the Wall stays mounted, nothing is released, and `getWindowSnapshot` omits it. -2. **Rust** reserves both endpoints, writes the initial arrival journal, then reassigns `terminalIds` to the target, keeps routing their output to +2. **Rust** reassigns `terminalIds` to the target, keeps routing their output to the source, and asks the sidecar to stamp a `pty:marked` line per id; at that line the id's suppression begins, until its replay has been emitted to the target. The source serializes each buffer at its mark and invokes @@ -697,16 +709,17 @@ below reads that record rather than inferring itself from the suppression map. for a tear-out, builds the new window, whose boot pulls a payload that is complete (`docs/specs/transport.md` → "Transferring a Workspace"; rationale). **An arrival without content is not drainable** - (`an_arrival_is_drainable_only_once_its_content_landed`). **Must keep pending journal admission source-owned and save-visible, exclude it from target drains, and count it as an in-flight close/quit blocker.** A failed initial write leaves ownership unchanged; a cancelled generation cannot promote late ownership (rationale). + (`an_arrival_is_drainable_only_once_its_content_landed`). 3. **Target** drains with `take_arrivals` and, per arrival, arms its collector - *before* calling `adopt_ready(workspaceId)` (rationale). Rust answers + *before* calling `adopt_ready(workspaceId)` — the hop that removes the whole + "arrived before armed" class of bug (rationale). Rust answers `pty:requestInit` with **that arrival's ids and no others**; `pty:list` and each `pty:replay` echo the collector's token. The target resumes over them, mounts the Workspace at the payload’s index, else the drop index. **Never spawns or kills**: the Sessions' alert state never left the sidecar, whose answer to that `pty:requestInit` re-sends it (§Alerts; `docs/specs/alert.md` → Live Workspace transfer). -4. **Target adopted** invokes `adopt_done(workspaceId)`. Rust durably marks the journal settled, then retires the in-memory record, +4. **Target adopted** invokes `adopt_done(workspaceId)`. Rust retires the record, clears pending marks and suppression so live output routes to the target, and emits `workspace-departed` for **that Workspace alone** to its own source. @@ -719,28 +732,28 @@ below reads that record rather than inferring itself from the suppression map. had no Sessions and no window that owned them. - **A transferring Workspace is in no snapshot its source writes, and neither end's teardown kills or interrupts its shells.** They belong to the target by - ownership after journal admission, and the target's `pty_graceful_kill` and + ownership from the invoke, and the target's `pty_graceful_kill` and `capture_agent_recovery` exclude every id an arrival claims (`boot_list_ids`), - since the source is still showing them (rationale). + since the source is still showing them. A quit or a close in the gap would + otherwise persist the same Workspace in two windows, or kill it under the + source. - **Must await `adopt_done` before installing a torn-out Window; refusal releases its resumed Sessions and boots fresh.** -- **Must unwind a refused `adopt_done` mount and request handback.** Refusal can be a failed settled-journal write or an arrival already returned by its watchdog; the +- **A refused `adopt_done` unwinds the mount.** The `ARRIVAL_MAX` watchdog has + already handed the shells back and the source kept the Workspace, so the target releases its Sessions (never kills them) and closes the Workspace rather than leaving it live and persisted in two windows. **Must unwind from the received payload without preparing another move.** - **Must remove this window's unmounted semantic and Activity state when arrival collection times out**, and kill nothing: the source goes on showing those Sessions. Source of truth: `discardArrival` in `standalone/src/workspace-move.ts`. -- **Must durably reverse a refused arrival before returning ownership, retiring its reservation or notifying the source.** The target's `adopt_failed` +- **A refused arrival hands the shells back.** The target's `adopt_failed` (a `planArrival` timeout, a missing list, a mount error) and a target `Destroyed` with the arrival still queued both return `terminalIds` to the - source (`a_target_closing_mid_arrival_hands_its_shells_back`), - drop the in-memory record, and emit `workspace-arrival-failed`; the + source unsuppressed (`a_target_closing_mid_arrival_hands_its_shells_back`), + drop the record, and emit `workspace-arrival-failed`; the source clears **transferring** and the Workspace is simply still there. With - both ends gone the shells are reaped rather than left owned by a dead label, - and the durable recovery record remains. **Must bound failed-return retries and then reserve the arrival for cold recovery without returning ownership, adopting it or reaping its claimed PTYs.** Show recovery status; permit global quit/restart through normal confirmation and watchdogs while individual closes and transfers involving either endpoint remain blocked, including after a quit cancellation. Preserve the journal and snapshots on that quit. Workers act only on their phase and generation; `permanent_return_failure_bounds_quit_and_cold_restores_exactly_once` pins this boundary (rationale). - - Source of truth: `failed_arrival_return` in `standalone/src-tauri/src/lib.rs`, `ArrivalQueue` in `standalone/src-tauri/src/quit_state.rs`. + both ends gone the shells are reaped rather than left owned by a dead label. - **Must change transfer ownership and source routing under one routing lock**, so output before the mark always reaches the source (`transfer_ownership_and_source_routing_change_together`). @@ -772,14 +785,20 @@ below reads that record rather than inferring itself from the suppression map. rationale). - **`planArrival` never throws into `bootstrap()`.** A refused sole arrival on the boot path renders a fresh one-pane Workspace, never a blank window. -- **Must keep `take_arrivals` non-consuming and deduplicate target adoption by id.** Settlement owns retirement (rationale). +- **`take_arrivals` does not consume.** The record settles at `adopt_done`, so a + webview that drains at boot and again when its listener is installed cannot + lose a Workspace to a drain that happened too early; the webview dedupes by id. - **A window with a snapshot boots as itself**, mounting whatever was dropped on it mid-boot over the restore rather than instead of it. - **`AWAITING_REPLAY_MAX` fails open only for suppressions no arrival claims.** A cold boot slower than it would otherwise have a real arrival's shells unsilenced into a window that has not resumed them yet. -- **Must exclude transferred arrival ids from boot's `pty_request_init`, while retaining pending admission's source-owned ids.** -- **Must journal an arrival in `sessions/arrivals.json` before ownership moves, never stage it in either window's snapshot** (rationale). **Must retain an adopted record until target +- **A boot's `pty_request_init` excludes every id an arrival claims.** Ownership + moves at the invoke, so those shells would otherwise be listed as top-level + panes beside the Workspace about to mount them. +- **`begin_arrival` records the arrival in `sessions/arrivals.json`** — a JSON + array of `{ workspaceId, from, to, workspace, settled }`, never an entry in + either window's snapshot (rationale); the tombstone rules below read `settled`. **Must retain an adopted record until target and source snapshots both reflect the move**, marking it settled at `adopt_done` and checking after each `save_session` or source-window close (`adoption_keeps_the_journal_until_both_snapshots_are_durable`). **Must reverse @@ -794,17 +813,24 @@ below reads that record rather than inferring itself from the suppression map. and successful records are deleted; **must retain failed records for retry and roll back the target if trimming the source fails**. **Must preserve a settled arrival’s newer target record during boot recovery** - (`standalone/src-tauri/src/lib.rs` recovery tests). + (`an_arrival_record_round_trips_until_it_is_forgotten`, + `a_leftover_arrival_boots_into_an_existing_target_snapshot`, + `a_leftover_arrival_boots_into_a_tear_out_targets_new_snapshot`, + `a_leftover_arrival_leaves_a_source_snapshot_that_still_names_it`, + `the_arrivals_file_is_gone_after_the_boot_merge`). - **Must run journal I/O and its lock waits off the main thread**, including transfer/settlement/close commands and destroyed-window cleanup (`blocking_commands_run_off_the_main_thread`). -- **Must request return after `ARRIVAL_MAX` only for the watchdog's own generation (`queued_at`)**, including admission still waiting for disk. Journal failure delays settlement under the retained blocker +- **An arrival unadopted after `ARRIVAL_MAX` is handed back** by a watchdog armed + at `begin_arrival`, retiring only the record it was armed for (`queued_at`): + a target alive but wedged never reaches `adopt_failed` or `Destroyed`, and the + source would otherwise stay transferring with its shells silent for good (`an_expiry_retires_only_the_record_it_was_armed_for`). Source of truth: `Arrival` / `sweep_awaiting` / `expire_arrival` / `boot_list_ids` in `standalone/src-tauri/src/routing.rs`; `begin_arrival` / `adopt_ready` / -`adopt_done`, `adopt_failed`, `hand_back_arrival`, `commit_initial_arrival_with`, -`commit_arrival_adoption_with`, `commit_arrival_return_with` and `restore_arrivals` in +`adopt_done` / `adopt_failed` / `hand_back_arrival` / `record_arrival_on_disk` / +`forget_arrival_on_disk` / `restore_arrivals` in `standalone/src-tauri/src/lib.rs`; `standalone/src/workspace-move.ts`; `markWorkspaceTransferring` in `lib/src/lib/window-session-aggregator.ts`. Pinned by `standalone/src/workspace-move.test.ts`, the disk tests in @@ -952,11 +978,22 @@ record cannot reach the installed app). The browser-dev harness sets both to its own per-run temp directory. Source of truth: `recovery_state_dir` in `standalone/src-tauri/src/lib.rs`. -**Must restrict the session directory and temp file to the owner before writing bytes; permission failure must preserve the previous snapshot** (`docs/specs/security-local.md` → Persisted state; rationale). Windows uses a protected owner-only DACL, including existing files; Unix uses `0700` / `0600`. `session_permission_failures_preserve_previous_snapshot_without_writing_bytes` and `restrict_to_owner_leaves_one_owner_only_ace` pin this. **Must disable recovery storage when its directory cannot be hardened**, logging the path; Burrow state-directory failure warns and selects the in-memory fallback. - -**Must hydrate synchronous `getState()` before cold restore and coalesce asynchronous writes with latest-value-wins ordering.** - -Source of truth: `restrict_to_owner`, `recovery_state_dir` and `write_file_with_permissions` in `standalone/src-tauri/src/lib.rs`; `TauriSessionStore` in `standalone/src/tauri-session-store.ts`. +**Must restrict the session store to the owner before any bytes are written** +(`docs/specs/security-local.md` -> "Persisted state"; rationale). +`restrict_to_owner` sets `0700` on the directory and `0600` on the temp file +*first*, since the rename preserves its mode; on Windows, where a unix mode is a +silent no-op, it applies a protected single-entry DACL instead (mechanism in its +doc comment). `burrow_state_dir` locks the sidecar's state directory with the +same call and relies on it reaching a file that already *existed*, which +`restrict_to_owner_leaves_one_owner_only_ace` pins (rationale). **Must abort a snapshot save if either permission change fails**, preserving the previous snapshot. The state-directory call remains nonfatal and logs a `WARNING` naming the path. Pinned by `session_permission_failures_preserve_previous_snapshot_without_writing_bytes` and `session_write_tightens_directory_and_existing_temp_file`. + +**Boot + the synchronous-read constraint.** `getState()` is synchronous — +cold-start restore reads it before React mounts — but a Tauri `invoke` is async, so +`TauriSessionStore` keeps an in-memory write-through cache: `TauriAdapter.init()` +`hydrate`s it from `load_session` (§Boot sequence), `getItem` reads it +synchronously, `setItem` updates it and forwards to `save_session` asynchronously, +coalescing bursts to at most one in-flight write (latest value wins). Mirrors the +VS Code adapter's host-injected seed (`docs/specs/vscode.md`). Dirty tracking is shared frontend behavior (`docs/specs/layout.md` → Session persistence). **Must skip an unchanged store write only when that value is queued or saved.** @@ -972,9 +1009,6 @@ write; failed writes are logged. With Session persistence disabled, the pipeline is already idle. Drain is a completion barrier, not a guarantee of successful disk persistence. -**Must distinguish process-crash recovery from hardware power-loss durability.** -Directory fsync is best-effort after writes; unlink helpers do not fsync their directory (rationale). - ### Agent recovery **Must run shared capture and record ownership in the sidecar**, under `docs/compatible-agents.md`. Rust bridges `capture_agent_recovery` and `take_recovery_commands` asynchronously. **Must answer both through `respondAsync`**, returning `{ error }` on throws rather than stranding the invoke. @@ -1014,7 +1048,13 @@ teardown at a time. Source of truth: `QuitMachine` in exits** rather than leaving a process with none, and so does a trigger that finds none: parked in `Voting` it would have no window to vote and refuse every later exit. -- **Must keep every Window snapshot on disk during quit and leave in-flight arrivals' claimed shells out of each endpoint's teardown.** Cold recovery follows §Arrival queue. +- **A quit keeps every window's snapshot on disk** — that is what a relaunch + restores from, and the whole difference from a per-window close. **A Workspace + in transfer is in no snapshot until its target publishes it, and neither end's + teardown kills its shells** (§Arrival queue). A quit mid-transfer restores it + at most once: from the target once it has published it, or from a source + handed it back because the target was destroyed first; a source torn down + before its target adopts leaves it in no snapshot. ### Trigger interception diff --git a/docs/specs/standalone.rationale.md b/docs/specs/standalone.rationale.md index 17fe908e7..a22e726dd 100644 --- a/docs/specs/standalone.rationale.md +++ b/docs/specs/standalone.rationale.md @@ -28,8 +28,6 @@ ## Burrow service -A refused state-directory DACL leaves the Node sidecar with no privacy control on Windows: POSIX modes do nothing there. Falling back to in-memory enrollment and ACL keeps the Burrow usable without writing new credentials into an unrestricted directory. Existing durable state stays untouched until a later launch can establish the boundary. - **Why one file rather than one per value.** Per-value files leave a window between two writes in which the enrollment can end up describing a different Burrow than the ACL records approved under it. **Why the correlation field cannot be `requestId`.** A `burrow:*` payload reusing that field has its results consumed by the invoke table, vanishing at random. @@ -110,18 +108,6 @@ enumeration — which is already Rust's job, beside the snapshots it reads. The cap is a ceiling on one launch, not a limit on how many windows may exist: the excess snapshots stay on disk untouched, so raising the cap restores them. -## Per-window close - -The October 2026 Windows review found that a skipped close-time save acknowledged as successful advanced the webview's saved-value cache. Wall dirty tracking had already completed at aggregate publication, so an unchanged retained Window could stay unsaved indefinitely. Explicit refusal plus one remembered latest-value retry covers writes outliving the bounded drains; the real store and aggregate regression holds the refused write past both timeouts. The retry is consumed once and never loops on disk failure. A failed native cancellation keeps the frontend flow guarded rather than letting its next close race the old handshake. - -A cancelled watchdog needs a fresh token on retry so the earlier sleeping watchdog cannot destroy the Window under another consent dialog. Pending approved downloads live only in their webview: closing loses the installer held there. Competing teardown flows share one dialog; leaving a request unanswered parks its native machine with no decision to receive. - -An audit on Windows, 2026-10-01, held `remove_window_session` pending for nine seconds: the frontend's eight-second ceiling never started, since it covered only the later PTY kill. Native completion then took the same journal lock a second time. A disk error was logged while the code destroyed the window anyway, leaving its old snapshot eligible for relaunch. - -Cancellation before the mutation boundary can retire a worker waiting for the journal lock without touching disk. Cancellation after that boundary cannot safely release save refusal: a late unlink could delete a newer snapshot from the retained window. The token therefore gives that entered commit exclusive ownership until completion. Normal deletion errors restore the earlier journal/geometry before exposing retry; rollback failure remains guarded. An arbitrary synchronous filesystem syscall has no safe forced-cancellation mechanism, so the preparation deadline is distinct from entered kernel IO and the later PTY deadline. - -Queued native commands also survive the window's `Destroyed` event. Keeping refusal for successfully closed labels, which are not reused in a process, prevents such a save from recreating its snapshot after native teardown. - ## Transfer The `adopt_ready` hop exists because the alternative is a race with no safe @@ -147,11 +133,7 @@ ends at `adopt_failed` or at the target's `Destroyed`, not at a timer. ## Arrival queue -A target whose `adopt_done` is refused already has the arrival payload needed to release its Sessions. Preparing a new transfer first re-entered Tool startup checks and could throw while ownership was already back at the source. Direct unwind also handles a failed settled-journal write; an explicit `adopt_failed` then requests the durable return instead of waiting for the watchdog. A return already completed makes that request stale and harmless. - -The Windows audit on 2026-10-01 found that a failed initial journal write was logged before ownership moved anyway, while the source's next snapshot omitted the Workspace. A failed return write likewise removed the close/quit blocker before its durable destination changed. Admission and return reservations now hold the blockers across disk IO and release source state only after the matching generation's journal commit. Native failure tests preserve earlier journal bytes and replay the return into exactly one source snapshot. - -Review on 2026-10-02 found that unlimited reverse-write retries also swallowed every quit before its watchdog started. A parked return retains live ownership and the durable journal; global quit preserves snapshots, so it can use normal confirmation and bounded teardown and let cold restore settle that record at its durable destination. Individual close deletes snapshots and remains blocked, as do new transfers at those endpoints. Canceling quit leaves the reservation in place. The initial attempt and five one-second retries allow transient failures to recover without requiring a restart; exhaustion performs no further write or false hand-back. +A target whose `adopt_done` is refused already has the arrival payload needed to release its Sessions. Preparing a new transfer first re-entered Tool startup checks and could throw while ownership was already back at the source; unwinding directly also avoids sending `adopt_failed` for that retired arrival. **Why the mark is stamped in the stream rather than asked for.** A mark fetched by request answers at some instant the sidecar chose, while the source's xterm @@ -242,8 +224,6 @@ stays the webview's throughout and no polling loop is needed. **Why the sessions directory is fsynced after the rename.** Fsyncing only the temp file leaves the new name recoverable-but-absent after a power loss; the directory-entry fsync is what makes the rename itself durable. Windows has no equivalent concept, hence unix-only. -Checked on 2026-10-01: `write_file_with_permissions` attempts that Unix directory fsync best-effort, and session/journal unlink helpers do not fsync the directory. Their completed operations and journal ordering support process-crash recovery; they do not establish a guarantee for abrupt hardware power loss. - **Why the mode is set before the bytes.** Under the bare umask the transcript-bearing blob lands `0644` in a `0755` directory any other local account can read, and tightening after the write would leave a window in which it was readable. Continuing after a permission failure would contradict the owner-only guarantee; aborting before writing preserves the previous snapshot and leaves at most an empty temp file. **Why the ACE test asserts an already-existing file.** On an upgrade the Burrow enrollment file is already there, so what tightens it is propagation onto an existing entry rather than create-time inheritance. `FileBurrowStateStore`'s own `0700`/`0600` cannot help on Windows: Node has no ACL API. diff --git a/lib/src/host/remote/burrow-state-store.ts b/lib/src/host/remote/burrow-state-store.ts index 41d5137f4..7b4438a43 100644 --- a/lib/src/host/remote/burrow-state-store.ts +++ b/lib/src/host/remote/burrow-state-store.ts @@ -8,8 +8,7 @@ * The interface is async because the hosts that implement it are: files the * sidecar owns here, `VsCodeBurrowStateStore` there (enrollment in * `SecretStorage`, ACL in `globalState` — `docs/specs/vscode.md`). {@link FileBurrowStateStore} - * is the sidecar's: private JSON state under a directory the app passes in - * only after establishing owner-only access (POSIX modes or a Windows DACL). + * is the sidecar's: two files, 0600, under a directory the app passes in. */ import { readFile, rm } from 'node:fs/promises'; diff --git a/scripts/spec-word-budgets.json b/scripts/spec-word-budgets.json index bab45db83..19f48d5fe 100644 --- a/scripts/spec-word-budgets.json +++ b/scripts/spec-word-budgets.json @@ -30,7 +30,7 @@ "docs/specs/security-supply-chain.md": 1250, "docs/specs/security.md": 2150, "docs/specs/shortcuts.md": 1100, - "docs/specs/standalone.md": 11850, + "docs/specs/standalone.md": 11950, "docs/specs/terminal-context.md": 1100, "docs/specs/terminal-escapes.md": 4050, "docs/specs/terminal-state.md": 2400, diff --git a/standalone/scripts/dev-agent-browser.test.mjs b/standalone/scripts/dev-agent-browser.test.mjs index ddac0e687..bc7916c17 100644 --- a/standalone/scripts/dev-agent-browser.test.mjs +++ b/standalone/scripts/dev-agent-browser.test.mjs @@ -2,7 +2,7 @@ import test from 'node:test'; import assert from 'node:assert/strict'; import { access, copyFile, mkdir, readFile, rm, writeFile } from 'node:fs/promises'; import path from 'node:path'; -import { fileURLToPath, pathToFileURL } from 'node:url'; +import { fileURLToPath } from 'node:url'; import { spawn } from 'node:child_process'; import { get } from 'node:http'; import { setTimeout as delay } from 'node:timers/promises'; @@ -30,10 +30,6 @@ async function fixture(t) { }); `); const cli = path.join(bin, 'cli.cjs'); - // Windows kill('SIGTERM') bypasses JS handlers. Exercise the same shutdown - // handler over IPC there; POSIX continues exercising the actual signal. - const signals = path.join(bin, 'signals.mjs'); - await writeFile(signals, "process.on('message', signal => process.emit(signal));"); await writeFile(cli, ` if (process.argv[2] === 'list' && process.env.TEST_DOR_LIST) { console.log(process.env.TEST_DOR_LIST); @@ -58,8 +54,8 @@ async function fixture(t) { return { root, start(overrides = {}) { - const child = spawn(process.execPath, ['--import', pathToFileURL(signals).href, path.join(standalone, 'scripts/dev-agent-browser.mjs')], { - cwd: root, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe', 'ipc'], + const child = spawn(process.execPath, [path.join(standalone, 'scripts/dev-agent-browser.mjs')], { + cwd: root, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe'], }); // Object.assign, not a spread: `runner`'s `output`/`closed` are getters // over live state, and spreading would snapshot them once. @@ -77,10 +73,7 @@ async function fixture(t) { return this; }, async stop() { - if (child.exitCode === null && child.signalCode === null) { - if (process.platform === 'win32' && child.connected) child.send('SIGTERM', () => {}); - else child.kill('SIGTERM'); - } + if (child.exitCode === null && child.signalCode === null) child.kill('SIGTERM'); const timer = setTimeout(() => child.kill('SIGKILL'), 5000); try { return await this.exited; } finally { clearTimeout(timer); } }, diff --git a/standalone/scripts/dev-standalone.test.mjs b/standalone/scripts/dev-standalone.test.mjs index 039a7c07a..4894deb2e 100644 --- a/standalone/scripts/dev-standalone.test.mjs +++ b/standalone/scripts/dev-standalone.test.mjs @@ -2,7 +2,7 @@ import test from 'node:test'; import assert from 'node:assert/strict'; import { copyFile, mkdir, rm, writeFile } from 'node:fs/promises'; import path from 'node:path'; -import { fileURLToPath, pathToFileURL } from 'node:url'; +import { fileURLToPath } from 'node:url'; import { spawn } from 'node:child_process'; import { setTimeout as delay } from 'node:timers/promises'; import { cleanEnv, devWorkspace, runner, writeShims } from './dev-fixture.mjs'; @@ -54,7 +54,7 @@ async function fixture(t) { return { root, start(args = ['dev'], overrides = {}) { - const child = spawn(process.execPath, ['--import', pathToFileURL(signals).href, path.join(standalone, 'scripts/tauri.mjs'), ...args], { + const child = spawn(process.execPath, ['--import', signals, path.join(standalone, 'scripts/tauri.mjs'), ...args], { cwd: standalone, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe', 'ipc'], }); // Object.assign, not a spread: `runner`'s `output`/`closed` are getters diff --git a/standalone/src-tauri/src/close_commit.rs b/standalone/src-tauri/src/close_commit.rs deleted file mode 100644 index 257a0b15f..000000000 --- a/standalone/src-tauri/src/close_commit.rs +++ /dev/null @@ -1,69 +0,0 @@ -//! Cancellation owns only preparation. Once a worker crosses the mutation -//! boundary, its caller must await the result rather than release save refusal -//! while a late unlink still owns the window's files. -use std::sync::atomic::{AtomicU8, Ordering}; - -const PREPARING: u8 = 0; -const COMMITTING: u8 = 1; -const DONE: u8 = 2; -const CANCELLED: u8 = 3; - -#[derive(Debug, Default)] -pub struct CloseCommit(AtomicU8); - -impl CloseCommit { - /// Called under the journal lock, immediately before the first mutation. - pub fn begin(&self) -> Result<(), String> { - self.0.compare_exchange(PREPARING, COMMITTING, Ordering::AcqRel, Ordering::Acquire) - .map(|_| ()).map_err(|_| "close preparation was cancelled".to_string()) - } - - pub fn cancel(&self) -> bool { - self.0.compare_exchange(PREPARING, CANCELLED, Ordering::AcqRel, Ordering::Acquire).is_ok() - } - - pub fn finish(&self) { self.0.store(DONE, Ordering::Release); } - pub fn done(&self) -> bool { self.0.load(Ordering::Acquire) == DONE } -} - -#[cfg(test)] -mod tests { - use super::*; - use std::sync::{Arc, Mutex}; - - #[test] - fn cancellation_while_waiting_for_journal_fences_the_late_worker() { - let disk = Arc::new(Mutex::new(())); - let blocked = disk.lock().unwrap(); - let token = Arc::new(CloseCommit::default()); - let worker_token = token.clone(); - let worker_disk = disk.clone(); - let worker = std::thread::spawn(move || { - let _disk = worker_disk.lock().unwrap(); - worker_token.begin() - }); - assert!(token.cancel()); - drop(blocked); - assert!(worker.join().unwrap().is_err()); - assert!(!token.done()); - // A retry owns a distinct generation; cancelling the old token cannot - // cancel or authorize the new worker. - let retry = CloseCommit::default(); - retry.begin().unwrap(); - assert!(!token.cancel()); - retry.finish(); - assert!(retry.done()); - } - - #[test] - fn entered_commit_cannot_be_cancelled_or_reentered() { - let token = CloseCommit::default(); - token.begin().unwrap(); - assert!(!token.cancel()); - assert!(!token.done()); - assert!(token.begin().is_err()); - token.finish(); - assert!(token.done()); - assert!(!token.cancel()); - } -} diff --git a/standalone/src-tauri/src/lib.rs b/standalone/src-tauri/src/lib.rs index 1f2c410de..8fda5925c 100644 --- a/standalone/src-tauri/src/lib.rs +++ b/standalone/src-tauri/src/lib.rs @@ -5,7 +5,6 @@ mod panic_policy; mod quit_state; mod routing; mod workspaces; -mod close_commit; // The Dock's Quit, an `osascript` quit and a logout reach AppKit without ever // raising `RunEvent::ExitRequested` (docs/specs/standalone.md §Trigger // interception). @@ -141,9 +140,8 @@ struct WindowState { hover_target: Mutex>, /// Labels whose snapshot has been deliberately removed. A save arriving /// from a webview that is going away must not put the file back; the entry - /// survives destruction: a previously dispatched save can still arrive. + /// is dropped once that webview is destroyed and can no longer save. closing: Mutex>, - close_commits: Mutex>>, /// The next `ws-`, seeded above every live and saved label at setup. next_ws: AtomicU64, /// Every window's Workspaces under their stable refs (§Workspace registry). @@ -184,8 +182,7 @@ impl WindowState { } /// Refuse every later `save_session` for `label` (a deliberate close removed - /// its snapshot). Kept for the process lifetime: an already dispatched - /// save may acquire the journal lock after `Destroyed`. + /// its snapshot). Cleared by `Destroyed`, after which no save can arrive. fn begin_closing(&self, label: &str) { guard(&self.closing).insert(label.to_string()); } @@ -277,20 +274,14 @@ impl WindowState { // reads it: an arriving shell belongs to its source again, which // `hand_back` returns it to, and reaping it here would kill a terminal // the source is still showing. - let (lost, claimed) = { + let lost = { let mut arrivals = guard(&self.arrivals); arrivals.forget_deferred_close(label); - let lost = routing::take_arrivals_to(&mut arrivals, label); - // A return whose journal is still being retried already has a - // worker, but its source's shells must not become orphan kills. - let claimed: Vec = arrivals.iter() - .filter(|arrival| arrival.to == label && arrival.transferred) - .flat_map(|arrival| arrival.terminal_ids.clone()).collect(); - (lost, claimed) + routing::take_arrivals_to(&mut arrivals, label) }; let owned = { let mut routing = guard(&self.routing); - for id in &claimed { + for id in lost.iter().flat_map(|arrival| &arrival.terminal_ids) { routing.owners.remove(id); routing.awaiting_replay.remove(id); } @@ -682,7 +673,6 @@ fn request_quit(app: &AppHandle, intent: QuitIntent) -> bool { append_log("[quit] transfer in progress; quit queued until settlement"); return relaunches; } - let intent = arrivals.take_quit_intent(intent); let labels = window_labels(app); let mut machine = guard(&state.machine); let (seq, actions) = machine.request(&labels, intent); @@ -808,19 +798,20 @@ fn request_window_close(app: &AppHandle, label: &str) { /// Tauri has actually taken the label out of `webview_windows()`. fn finish_window_close(app: &AppHandle, label: &str) { append_log(format!("[window] closing {label} and removing its snapshot")); - // The webview already committed removal on the normal path; the cached - // completion avoids another journal wait after its bounded PTY teardown. - if let Err(err) = remove_window_session_bounded(app, label) { - append_log(format!("[window] retaining {label}: {err}")); - if !err.starts_with("close-commit-uncertain:") { - if let Some(state) = app.try_state::() { guard(&state.close).clear(label); } - } - let _ = app.emit_to(label, "dormouse://window-close-failed", err); - return; + if let Some(state) = app.try_state::() { + // Before the removal, not after: a save already in flight from this + // webview would otherwise put the snapshot back. Reached from the + // watchdog too, where the webview never called `remove_window_session`. + state.begin_closing(label); } if let Some(state) = app.try_state::() { guard(&state.close).clear(label); } + if let Ok(dir) = sessions_dir(app) { + if let Err(err) = close_window_snapshot(&dir, label) { + append_log(format!("[session] {err}")); + } + } if let Some(window) = app.get_webview_window(label) { let _ = window.destroy(); } @@ -1903,19 +1894,14 @@ async fn save_session(window: tauri::Window, state: String) -> Result<(), String // A deliberate close removes the snapshot; a save still in flight from the // webview that is going away must not put it back // (docs/specs/standalone.md §Per-window close). - let windows = window.app_handle().try_state::(); - let dir = sessions_dir(window.app_handle())?; - save_open_window_session(windows.as_deref(), &dir, window.label(), &state) -} - -/// Caller owns ARRIVAL_DISK_LOCK. Refusal is an error, never a persistence -/// acknowledgment: a retained window must retry its unchanged cached value. -fn save_open_window_session(windows: Option<&WindowState>, dir: &Path, label: &str, state: &str) -> Result<(), String> { - if windows.is_some_and(|windows| windows.refuses_save(label)) { - return Err("window close is holding its snapshot; no session was saved".to_string()); + if let Some(windows) = window.app_handle().try_state::() { + if windows.refuses_save(window.label()) { + return Ok(()); + } } - write_session_to(dir, label, state)?; - retire_saved_arrivals(dir) + let dir = sessions_dir(window.app_handle())?; + write_session_to(&dir, window.label(), &state)?; + retire_saved_arrivals(&dir) } /// The suffix `write_file_atomically` leaves on a session snapshot's temp @@ -2179,16 +2165,6 @@ fn geometry_path(dir: &Path, label: &str) -> PathBuf { dir.join(format!("{stem}.geometry.json")) } -/// Platform reads and rect-cache access precede this disk lock. Recheck the -/// closing fence here: a live webview/rect captured before close can outlive -/// snapshot removal and must not recreate geometry or share its temp writer. -fn write_open_window_geometry(windows: Option<&WindowState>, dir: &Path, label: &str, json: &str) -> Result { - let _disk = guard(&ARRIVAL_DISK_LOCK); - if windows.is_some_and(|windows| windows.refuses_save(label)) { return Ok(false); } - write_file_atomically(&geometry_path(dir, label), json)?; - Ok(true) -} - fn read_geometry(dir: &Path, label: &str) -> Option { let raw = std::fs::read_to_string(geometry_path(dir, label)).ok()?; serde_json::from_str(&raw).ok() @@ -2269,8 +2245,7 @@ fn note_geometry(app: &AppHandle, label: &str, origin: Option<(i32, i32)>, size: let Ok(json) = serde_json::to_string(&rect.to_logical()) else { continue; }; - let windows = app.try_state::(); - if let Err(err) = write_open_window_geometry(windows.as_deref(), &dir, &label, &json) { + if let Err(err) = write_file_atomically(&geometry_path(&dir, &label), &json) { append_log(format!("[window] geometry write for {label}: {err}")); } } @@ -2439,10 +2414,6 @@ fn retire_saved_arrivals(dir: &Path) -> Result<(), String> { fn mark_arrival_adopted_on_disk(dir: &Path, workspace_id: &str) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); - mark_arrival_adopted_in_locked_journal(dir, workspace_id) -} - -fn mark_arrival_adopted_in_locked_journal(dir: &Path, workspace_id: &str) -> Result<(), String> { let mut records = read_arrivals_from(dir)?; for record in &mut records { if record_workspace_id(record) == Some(workspace_id) { record["settled"] = JsonValue::Bool(true); } @@ -2454,10 +2425,6 @@ fn mark_arrival_adopted_in_locked_journal(dir: &Path, workspace_id: &str) -> Res /// Refusal reverses the durable destination before the source is told to save. fn return_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); - return_arrival_in_locked_journal(dir, arrival) -} - -fn return_arrival_in_locked_journal(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let mut records = read_arrivals_from(dir)?; records.retain(|r| record_workspace_id(r) != Some(&arrival.workspace_id)); records.push(serde_json::json!({ @@ -2473,18 +2440,7 @@ fn return_arrival_in_locked_journal(dir: &Path, arrival: &routing::Arrival) -> R /// removes that stale copy instead of resurrecting it in either window. fn close_window_snapshot(dir: &Path, label: &str) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); - close_window_snapshot_locked(dir, label, None) -} - -fn close_window_snapshot_locked(dir: &Path, label: &str, token: Option<&close_commit::CloseCommit>) -> Result<(), String> { let mut records = read_arrivals_from(dir)?; - let previous_records = records.clone(); - let geometry = geometry_path(dir, label); - let previous_geometry = match std::fs::read_to_string(&geometry) { - Ok(value) => Some(value), - Err(err) if err.kind() == std::io::ErrorKind::NotFound => None, - Err(err) => return Err(format!("read {}: {err}", geometry.display())), - }; let mut changed = false; for record in &mut records { if record.get("to").and_then(JsonValue::as_str) == Some(label) @@ -2494,38 +2450,9 @@ fn close_window_snapshot_locked(dir: &Path, label: &str, token: Option<&close_co changed = true; } } - // All reads and lock waits precede the cancellation fence. Neither a late - // worker nor a second request may mutate a retained window after it loses. - if let Some(token) = token { token.begin()?; } - let session = dir.join(session_file_name(label)); - let remove = |path: &Path| match std::fs::remove_file(path) { - Ok(()) => Ok(()), - Err(err) if err.kind() == std::io::ErrorKind::NotFound => Ok(()), - Err(err) => Err(format!("remove {}: {err}", path.display())), - }; - // Fail before touching recoverable state when an orphan temp is blocked. - remove(&temp_write_path(&session))?; if changed { write_arrivals_to(dir, &records)?; } - let result = remove(&geometry).and_then(|_| remove(&session)); - if let Err(err) = result { - // The live snapshot was the final mutation. A failed removal leaves - // it present; restore the preceding journal/geometry before permitting - // the retained window to save or retry. An incomplete rollback keeps - // save refusal and the committed flow rather than claiming safety. - let rollback = (|| { - if changed { write_arrivals_to(dir, &previous_records)?; } - if let Some(value) = previous_geometry { write_file_atomically(&geometry, &value)?; } - Ok::<(), String>(()) - })(); - return match rollback { - Ok(()) => Err(err), - Err(rollback) => Err(format!("close-commit-uncertain: {err}; rollback failed: {rollback}")), - }; - } - // Retirement is cleanup: retaining a settled tombstone is recoverable and - // must not turn a successful deletion into a retry of a retained window. - if let Err(err) = retire_saved_arrivals(dir) { append_log(format!("[session] close journal cleanup: {err}")); } - Ok(()) + remove_session_from(dir, label)?; + retire_saved_arrivals(dir) } fn remove_workspace_from_disk(dir: &Path, label: &str, id: &str) -> Result<(), String> { @@ -2568,13 +2495,6 @@ fn record_workspace_id(record: &JsonValue) -> Option<&str> { /// Append one arrival's record, replacing any earlier record of the same id. fn record_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let _disk = guard(&ARRIVAL_DISK_LOCK); - record_arrival_in_locked_journal(dir, arrival) -} - -/// Caller owns ARRIVAL_DISK_LOCK. Initial admission checks its reservation -/// under this same lock, so a cancelled initial write cannot overwrite a -/// return that already committed its reverse destination. -fn record_arrival_in_locked_journal(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { let Some(workspace) = arrival.payload.get("workspace") else { return Err("arrival payload carries no workspace".to_string()); }; @@ -2728,8 +2648,6 @@ fn arrival_from(from: &str, to: &str, payload: JsonValue) -> Result Result<(), String>, -) -> Result<(), String> { - let _disk = guard(&ARRIVAL_DISK_LOCK); - let matches = |arrival: &routing::Arrival| { - arrival.workspace_id == expected.workspace_id && arrival.to == expected.to - && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Preparing - }; - if !guard(&windows.arrivals).iter().any(matches) { - return Err("the transfer reservation was retired before its journal write".to_string()); - } - write()?; - let mut arrivals = guard(&windows.arrivals); - let arrival = arrivals.iter_mut().find(|arrival| matches(arrival)) - .ok_or_else(|| "the transfer reservation was retired during its journal write".to_string())?; - windows.begin_transfer(&arrival.terminal_ids, &arrival.from, &arrival.to); - arrival.transferred = true; - arrival.phase = routing::ArrivalPhase::Active; - Ok(()) -} - -/// The queue keeps Returning arrivals as close/quit blockers during bounded retries until their -/// reverse journal write succeeds. Failed writes change neither ownership nor -/// the old durable destination, and retries cannot settle a later generation. -fn commit_arrival_return_with( - windows: &WindowState, - expected: &routing::Arrival, - write: impl FnOnce() -> Result<(), String>, -) -> Result { - let _disk = guard(&ARRIVAL_DISK_LOCK); - let matches = |arrival: &routing::Arrival| { - arrival.workspace_id == expected.workspace_id && arrival.to == expected.to - && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning - }; - if !guard(&windows.arrivals).iter().any(matches) { - return Err("the arrival return was retired before its journal write".to_string()); - } - write()?; - let mut arrivals = guard(&windows.arrivals); - if !arrivals.iter().any(matches) { - return Err("the arrival return was retired during its journal write".to_string()); - } - let marks = if expected.transferred { windows.hand_back(&expected.terminal_ids, &expected.from) } - else { serde_json::json!({}) }; - routing::retire_arrival(&mut arrivals, expected); - Ok(marks) -} - -/// Adoption releases the source only after its settled journal marker is -/// durable. A watchdog or dead target may request a return during that write; -/// the final phase/generation check fences the stale adoption. -fn commit_arrival_adoption_with( - windows: &WindowState, - expected: &routing::Arrival, - write: impl FnOnce() -> Result<(), String>, -) -> Result<(), String> { - let _disk = guard(&ARRIVAL_DISK_LOCK); - let matches = |arrival: &routing::Arrival| { - arrival.workspace_id == expected.workspace_id && arrival.to == expected.to - && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Active - }; - if !guard(&windows.arrivals).iter().any(matches) { - return Err("the arrival was returned before its adoption write".to_string()); - } - write()?; - let mut arrivals = guard(&windows.arrivals); - if !arrivals.iter().any(matches) { - return Err("the arrival was returned during its adoption write".to_string()); - } - windows.clear_suppression(&expected.terminal_ids); - routing::retire_arrival(&mut arrivals, expected); - Ok(()) -} - -/// Reserve both endpoints, durably journal the move, then reassign/suppress -/// the shells before either window is told. A failed initial write leaves the -/// source owner and save-visible; an expired reservation cannot move ownership. +/// Open one arrival: reassign its shells to the target and suppress them, then +/// queue the record. **Ownership moves synchronously here**, before either +/// window is told anything — the single Rust reader thread processes sidecar +/// lines in order, so every byte after this point is either dropped (and present +/// in the replay the target is about to get) or delivered to the target +/// (docs/specs/standalone.md §Transfer). fn begin_arrival( app: &AppHandle, windows: &WindowState, @@ -2869,29 +2709,14 @@ fn begin_arrival( arrival.workspace_id )); } - // Reserve admission while the worker writes the journal. Both endpoint - // teardowns wait, but the source still owns and saves every Session. + // Ownership moves now; suppression waits for each id's `marked` line. + windows.begin_transfer(&arrival.terminal_ids, &arrival.from, &arrival.to); routing::queue_arrival(&mut arrivals, arrival.clone()); } - spawn_arrival_watchdog(app.clone(), &arrival); - let written = sessions_dir(app).and_then(|dir| { - commit_initial_arrival_with(windows, &arrival, || record_arrival_in_locked_journal(&dir, &arrival)) - }); - if let Err(error) = written { - let retired = routing::retire_arrival(&mut guard(&windows.arrivals), &arrival); - if retired { redrive_deferred_teardown(app); } - return Err(format!("cannot commit the Workspace transfer: {error}")); - } // The split point, stamped in the stream by the sidecar and routed to the // source, which serializes what it holds when it sees it (§Transfer). A // Workspace of browser panes alone has no ids to mark; its source sends // content at once. - let arrivals = guard(&windows.arrivals); - if !arrivals.iter().any(|current| current.workspace_id == arrival.workspace_id - && current.queued_at == arrival.queued_at && current.phase == routing::ArrivalPhase::Active) - { - return Err("the transfer was returned before its stream marks".to_string()); - } if let Some(sidecar) = app.try_state::() { let msg = serde_json::json!({ "event": "pty:mark", @@ -2899,15 +2724,27 @@ fn begin_arrival( }); send_to_sidecar(&sidecar, msg.to_string()); } - drop(arrivals); + // Recorded on disk here, in neither window's snapshot: the source omits a + // transferring Workspace from its saves and the target writes only after + // adoption, so a crash in the gap would otherwise restore it nowhere + // (§Arrival queue). Never fatal: a failed write is logged and the transfer + // proceeds. + match sessions_dir(app) { + Ok(dir) => { + if let Err(e) = record_arrival_on_disk(&dir, &arrival) { + append_log(format!("[window] could not record {} on disk: {e}", arrival.workspace_id)); + } + } + Err(e) => append_log(format!("[window] {e}")), + } + spawn_arrival_watchdog(app.clone(), &arrival); Ok(()) } /// Bound an arrival: a target that never settles it — alive but wedged, so /// `Destroyed` never hands it back either — would leave the Workspace marked /// transferring in the source and its shells silent for good. Past -/// `ARRIVAL_MAX` requests a return. The reservation retires only after the -/// reverse journal is durable, like any refusal. +/// `ARRIVAL_MAX` the record is retired and handed back like any refusal. fn spawn_arrival_watchdog(app: AppHandle, arrival: &routing::Arrival) { let workspace_id = arrival.workspace_id.clone(); let to = arrival.to.clone(); @@ -2929,29 +2766,6 @@ fn spawn_arrival_watchdog(app: AppHandle, arrival: &routing::Arrival) { }); } -// Initial attempt plus five one-second retries. Exhaustion retains the live -// ownership fence and durable record for cold recovery, while admitting the -// non-destructive global quit flow. Individual close still deletes snapshots -// and must not pass it. A parked generation never resumes an old writer. -const ARRIVAL_RETURN_RETRIES: u8 = 5; - -#[derive(Debug, PartialEq, Eq)] -enum ArrivalReturnFailure { Retry(u8), RecoveryPending, Stale } - -fn failed_arrival_return( - windows: &WindowState, expected: &routing::Arrival, retries: u8, -) -> ArrivalReturnFailure { - let _disk = guard(&ARRIVAL_DISK_LOCK); - let mut arrivals = guard(&windows.arrivals); - let Some(arrival) = arrivals.iter_mut().find(|arrival| { - arrival.workspace_id == expected.workspace_id && arrival.to == expected.to - && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning - }) else { return ArrivalReturnFailure::Stale; }; - if retries > 0 { return ArrivalReturnFailure::Retry(retries - 1); } - arrival.phase = routing::ArrivalPhase::RecoveryPending; - ArrivalReturnFailure::RecoveryPending -} - /// One arrival will never be adopted: give its shells back to the source and /// tell the source so it clears the Workspace's transferring mark. **The /// Workspace simply stays where it is** — nothing was released, so there is @@ -2965,60 +2779,24 @@ fn failed_arrival_return( /// routing state at `pty:marked`, never from the later serialized content. Only /// an id the sidecar never stamped goes straight back. /// -/// The caller has marked this generation Returning in the queue. It remains -/// there until its reverse destination is durable or bounded retries park it -/// for cold recovery; only global quit may pass the parked reservation. +/// The record must already be out of the queue; the caller took it. fn hand_back_arrival( app: &AppHandle, windows: &WindowState, arrival: &routing::Arrival, reason: &str, -) { - hand_back_arrival_with_retries(app, windows, arrival, reason, ARRIVAL_RETURN_RETRIES); -} - -fn hand_back_arrival_with_retries( - app: &AppHandle, windows: &WindowState, arrival: &routing::Arrival, - reason: &str, retries: u8, ) { append_log(format!( "[window] {} never arrived in {} ({reason}); handing it back to {}", arrival.workspace_id, arrival.to, arrival.from )); if app.get_webview_window(&arrival.from).is_some() { - let returned = sessions_dir(app).and_then(|dir| { - commit_arrival_return_with(windows, arrival, || return_arrival_in_locked_journal(&dir, arrival)) - }); - let marks = match returned { - Ok(marks) => marks, - Err(error) => { - let status = match failed_arrival_return(windows, arrival, retries) { - ArrivalReturnFailure::Retry(remaining) => { - retry_arrival_return(app.clone(), arrival.clone(), reason.to_string(), remaining); - format!("The Workspace could not be returned safely yet. Retrying its recovery write: {error}") - } - ArrivalReturnFailure::RecoveryPending => { - redrive_deferred_teardown(app); - format!("The Workspace recovery write failed. Quit or restart to recover it from its saved transfer record: {error}") - } - ArrivalReturnFailure::Stale => return, - }; - append_log(format!("[window] {status}")); - let _ = app.emit_to(arrival.from.as_str(), "dormouse://workspace-arrival-retry", - serde_json::json!({ "workspaceId": arrival.workspace_id, "reason": status })); - return; + if let Ok(dir) = sessions_dir(app) { + if let Err(e) = return_arrival_on_disk(&dir, arrival) { + append_log(format!("[window] could not record hand-back: {e}")); } - }; - // The source can disappear while the reverse journal waits for disk. - // If Destroyed ran before hand_back, its owner sweep could not see the - // newly returned ids; settle those orphans here instead of assigning - // invisible shells to a label that no longer exists. - if app.get_webview_window(&arrival.from).is_none() { - for id in &arrival.terminal_ids { windows.forget_pty(id); } - reap_orphaned_ptys(app, &arrival.from, arrival.terminal_ids.clone()); - redrive_deferred_teardown(app); - return; } + let marks = windows.hand_back(&arrival.terminal_ids, &arrival.from); let (marked, _) = routing::hand_back_ids(arrival, &marks); // Told before the replay is asked for, so the source is listening for // it (`acceptHandBackReplay` in `standalone/src/workspace-move.ts`). @@ -3042,10 +2820,10 @@ fn hand_back_arrival_with_retries( } } } else { - // No source can continue this Workspace. Keep the durable destination - // instead of deleting its only recoverable record. - if !routing::retire_arrival(&mut guard(&windows.arrivals), arrival) { - return; + if let Ok(dir) = sessions_dir(app) { + if let Err(e) = forget_arrival_on_disk(&dir, &arrival.workspace_id) { + append_log(format!("[window] could not forget orphaned arrival: {e}")); + } } // Both ends are gone, so these shells belong to no window and nothing // would ever paint them (`routing::owner`). @@ -3057,20 +2835,6 @@ fn hand_back_arrival_with_retries( redrive_deferred_teardown(app); } -/// One retry chain per Returning generation. A later drop of the same id is -/// independent; it is never adopted, retired or written by this worker. -fn retry_arrival_return(app: AppHandle, expected: routing::Arrival, reason: String, remaining: u8) { - std::thread::spawn(move || { - std::thread::sleep(Duration::from_secs(1)); - let Some(windows) = app.try_state::() else { return; }; - let live = guard(&windows.arrivals).iter().any(|arrival| { - arrival.workspace_id == expected.workspace_id && arrival.to == expected.to - && arrival.queued_at == expected.queued_at && arrival.phase == routing::ArrivalPhase::Returning - }); - if live { hand_back_arrival_with_retries(&app, &windows, &expected, &reason, remaining); } - }); -} - /// Tear a Workspace out into a brand-new window under the cursor. /// /// The payload is *queued*, never emitted: an `emit_to` a window that does not @@ -3135,8 +2899,7 @@ fn transfer_workspace_content( let mut arrivals = guard(&windows.arrivals); let arrival = arrivals .iter_mut() - .find(|arrival| arrival.workspace_id == workspace_id && arrival.from == window.label() - && arrival.phase == routing::ArrivalPhase::Active) + .find(|arrival| arrival.workspace_id == workspace_id && arrival.from == window.label()) .ok_or_else(|| format!("no arrival of '{workspace_id}' from {}", window.label()))?; arrival.content = Some(content); (arrival.to.clone(), arrival.pending_window.take()) @@ -3154,8 +2917,9 @@ fn transfer_workspace_content( // already marked the Workspace transferring — `handOff` set // it when the invoke returned. So the ids go back *and* the // source is told: `arrival-failed` is what clears that mark. - let returned = routing::return_arrival(&mut guard(&windows.arrivals), &workspace_id, &to); - if let Some(arrival) = returned { + if let Some(arrival) = + routing::take_arrival(&mut guard(&windows.arrivals), &workspace_id, &to) + { hand_back_arrival(&app, &windows, &arrival, "the new window could not be built"); } return Err(err); @@ -3232,7 +2996,7 @@ fn adopt_ready( let (ids, marks) = { let arrivals = guard(&windows.arrivals); let arrival = routing::find_arrival(&arrivals, &workspace_id) - .filter(|arrival| arrival.to == label && arrival.phase == routing::ArrivalPhase::Active) + .filter(|arrival| arrival.to == label) .ok_or_else(|| format!("no arrival of '{workspace_id}' into {label}"))?; (arrival.terminal_ids.clone(), routing::arrival_marks(arrival)) }; @@ -3256,15 +3020,19 @@ fn adopt_done( windows: tauri::State<'_, WindowState>, workspace_id: String, ) -> Result<(), String> { - let arrival = { - let arrivals = guard(&windows.arrivals); - routing::find_arrival(&arrivals, &workspace_id) - .filter(|arrival| arrival.to == window.label() && arrival.phase == routing::ArrivalPhase::Active) - .cloned().ok_or_else(|| format!("no active arrival of '{workspace_id}' into {}", window.label()))? - }; + let arrival = routing::take_arrival( + &mut guard(&windows.arrivals), + &workspace_id, + window.label(), + ) + .ok_or_else(|| format!("no arrival of '{workspace_id}' into {}", window.label()))?; + windows.clear_suppression(&arrival.terminal_ids); // Keep the journal until source and target saves both reflect the move. - let dir = sessions_dir(&app)?; - commit_arrival_adoption_with(&windows, &arrival, || mark_arrival_adopted_in_locked_journal(&dir, &workspace_id))?; + if let Ok(dir) = sessions_dir(&app) { + if let Err(e) = mark_arrival_adopted_on_disk(&dir, &workspace_id) { + append_log(format!("[window] could not mark {workspace_id} adopted on disk: {e}")); + } + } append_log(format!( "[window] {workspace_id} adopted by {}; telling {}", arrival.to, arrival.from @@ -3287,7 +3055,7 @@ fn adopt_failed( workspace_id: String, reason: Option, ) -> Result<(), String> { - let arrival = routing::return_arrival( + let arrival = routing::take_arrival( &mut guard(&windows.arrivals), &workspace_id, window.label(), @@ -3395,53 +3163,11 @@ fn take_arrivals(window: tauri::Window, windows: tauri::State<'_, WindowState>) /// Remove this window's persisted snapshot and stop it being written again. #[tauri::command] async fn remove_window_session(window: tauri::Window) -> Result<(), String> { - remove_window_session_bounded(window.app_handle(), window.label()) -} - -/// Bound preparation/lock waits, never an irreversible kernel operation. The -/// timeout can retire only the worker's own token before its first mutation. -fn remove_window_session_bounded(app: &AppHandle, label: &str) -> Result<(), String> { - const PREPARATION_TIMEOUT: Duration = Duration::from_millis(8000); - let dir = sessions_dir(app)?; - let windows = app.state::(); - let token = { - let mut commits = guard(&windows.close_commits); - if let Some(previous) = commits.get(label) { - if previous.done() { return Ok(()); } - return Err("close-commit-uncertain: an earlier close is still finishing".to_string()); - } - let token = Arc::new(close_commit::CloseCommit::default()); - commits.insert(label.to_string(), token.clone()); - windows.begin_closing(label); - token - }; - let (send, receive) = mpsc::channel(); - let worker_token = token.clone(); - let worker_label = label.to_string(); - std::thread::spawn(move || { - let result = { - let _disk = guard(&ARRIVAL_DISK_LOCK); - close_window_snapshot_locked(&dir, &worker_label, Some(&worker_token)) - }; - if result.is_ok() { worker_token.finish(); } - let _ = send.send(result); - }); - let result = match receive.recv_timeout(PREPARATION_TIMEOUT) { - Ok(result) => result, - Err(mpsc::RecvTimeoutError::Timeout) if token.cancel() => - Err("close preparation timed out; workspaces were retained".to_string()), - // Commit owns the files; retain the progress modal and save refusal - // until the result is known. Arbitrary kernel IO cannot be cancelled. - Err(mpsc::RecvTimeoutError::Timeout) => receive.recv() - .map_err(|_| "close-commit-uncertain: close worker stopped".to_string())?, - Err(mpsc::RecvTimeoutError::Disconnected) => - Err("close-commit-uncertain: close worker stopped".to_string()), - }; - if result.as_ref().is_err_and(|err| !err.starts_with("close-commit-uncertain:")) { - guard(&windows.close_commits).remove(label); - guard(&windows.closing).remove(label); + let app = window.app_handle(); + if let Some(windows) = app.try_state::() { + windows.begin_closing(window.label()); } - result + close_window_snapshot(&sessions_dir(app)?, window.label()) } /// Which window is under the cursor, in that window's own logical client space. @@ -3603,13 +3329,6 @@ fn forget_restart(app: &AppHandle) { // ── Per-window close (docs/specs/standalone.md §Per-window close) ───────────── -#[tauri::command] -fn retry_window_close(app: AppHandle, window: tauri::Window) { - // Retry starts a fresh native handshake too: close admission must block - // transfers while the new human gate is open, before removal begins. - request_window_close(&app, window.label()); -} - // This window's close orchestrator is alive; stand its ack watchdog down. #[tauri::command] fn window_close_ack(window: tauri::Window, state: tauri::State<'_, QuitState>) { @@ -3893,17 +3612,6 @@ fn resolve_dor_cli_paths(sidecar_path: &Path, manifest_dir: &Path) -> DorCliPath // (lib/src/host/remote/burrow-state-store.ts). Created here so a first launch // hands the sidecar a directory that exists; if it can't be made, the sidecar is // told nothing and runs without persistence rather than not at all. -// The sidecar must receive a durable directory only after its privacy boundary -// exists. On Windows the Node store cannot repair a refused DACL installation. -fn prepare_burrow_state_dir( - dir: &Path, - restrict: impl FnOnce(&Path, u32) -> Result<(), String>, -) -> Result { - create_dir_all(dir).map_err(|e| format!("create state dir: {e}"))?; - restrict(dir, 0o700).map_err(|e| format!("restrict state dir {}: {e}", dir.display()))?; - Ok(dir.to_string_lossy().into_owned()) -} - fn burrow_state_dir(app: &AppHandle) -> Option { let dir = match app.path().app_data_dir() { Ok(dir) => dir, @@ -3912,6 +3620,10 @@ fn burrow_state_dir(app: &AppHandle) -> Option { return None; } }; + if let Err(e) = create_dir_all(&dir) { + append_log(format!("[sidecar] create state dir: {e}")); + return None; + } // The Node sidecar writes the Burrow enrollment here, and that record carries // `burrowToken` — a bearer credential for `/ws/burrow`. `FileBurrowStateStore` // asks for `0700`/`0600`, which Windows ignores entirely, so on Windows this @@ -3922,13 +3634,17 @@ fn burrow_state_dir(app: &AppHandle) -> Option { // `restrict_to_owner_leaves_one_owner_only_ace` covers with `before.json`. // On unix the store's own modes already do the job and this is a harmless // re-assert of the same intent. - match prepare_burrow_state_dir(&dir, restrict_to_owner) { - Ok(path) => Some(path), - Err(error) => { - append_log(format!("[sidecar] WARNING {error}; Burrow state is ephemeral")); - None - } + if let Err(e) = restrict_to_owner(&dir, 0o700) { + // Not fatal — a Burrow that cannot start is worse than one whose state + // directory kept the OS default — but never silent: on Windows this + // call is the only thing restricting `burrowToken`, so its failure is a + // downgrade of the sole control and has to be visible. + append_log(format!( + "[sidecar] WARNING could not restrict state dir {}: {e}", + dir.display() + )); } + Some(dir.to_string_lossy().into_owned()) } /// Where the sidecar writes the single-use agent-recovery record. Under the @@ -3953,9 +3669,6 @@ fn recovery_state_dir(app: &AppHandle) -> Option { "[recovery] WARNING could not restrict state dir {}: {e}", dir.display() )); - // Recovery commands must never be persisted to a directory whose - // owner-only boundary could not be established. - return None; } Some(dir.to_string_lossy().into_owned()) } @@ -4283,8 +3996,7 @@ pub fn run() { // Drop label-keyed ownership synchronously; only the // returned arrivals need the blocking journal worker. let (lost, orphaned) = state.drop_window(&label); - // Keep successful-close save refusal: a queued async - // save can still arrive after this native event. + guard(&state.closing).remove(&label); reap_orphaned_ptys(app, &label, orphaned); let changed = workspaces::forget_window(&mut guard(&state.registry), &label); if changed { broadcast_registry(app, &state); } @@ -4449,7 +4161,6 @@ pub fn run() { quit_proceed, quit_restart, window_close_ack, - retry_window_close, window_close_cancel, close_window, open_workspace_window, @@ -4528,9 +4239,6 @@ mod tests { }; use super::routing; use super::guard; - use super::{WindowState, QuitIntent, failed_arrival_return, ArrivalReturnFailure, ARRIVAL_RETURN_RETRIES, commit_initial_arrival_with, - commit_arrival_return_with, commit_arrival_adoption_with, - record_arrival_in_locked_journal, return_arrival_in_locked_journal}; use std::collections::HashSet; use std::fs; use std::path::{Path, PathBuf}; @@ -4587,53 +4295,6 @@ mod tests { } } - #[test] - fn burrow_directory_creation_failure_never_attempts_permissions() { - let root = TempDir::new("burrow-state-create-failure"); - let occupied = root.path().join("not-a-directory"); - fs::write(&occupied, b"previous").unwrap(); - let called = std::cell::Cell::new(false); - let result = super::prepare_burrow_state_dir(&occupied, |_, _| { - called.set(true); - Ok(()) - }); - assert!(result.is_err()); - assert!(!called.get()); - assert_eq!(fs::read(occupied).unwrap(), b"previous"); - } - - #[test] - fn burrow_directory_permission_failure_disables_durable_state() { - let root = TempDir::new("burrow-state-restrict-failure"); - let target = root.path().join("state"); - fs::create_dir(&target).unwrap(); - fs::write(target.join("burrow.json"), b"existing enrollment").unwrap(); - let result = super::prepare_burrow_state_dir(&target, |path, mode| { - assert_eq!(path, target); - assert_eq!(mode, 0o700); - Err("DACL refused".into()) - }); - assert!(result.unwrap_err().contains("DACL refused")); - assert_eq!(fs::read(target.join("burrow.json")).unwrap(), b"existing enrollment"); - assert_eq!(fs::read_dir(target).unwrap().count(), 1); - } - - #[test] - fn burrow_directory_is_created_and_restricted_before_publication() { - let root = TempDir::new("burrow-state-order"); - let target = root.path().join("nested").join("state"); - let called = std::cell::Cell::new(false); - let result = super::prepare_burrow_state_dir(&target, |path, mode| { - assert!(path.is_dir()); - assert_eq!(fs::read_dir(path).unwrap().count(), 0); - assert_eq!(mode, 0o700); - called.set(true); - Ok(()) - }); - assert!(called.get()); - assert_eq!(result.unwrap(), target.to_string_lossy()); - } - // --- Pending arrivals on disk (§Arrival queue) --------------------------- fn workspace_json(id: &str, name: &str) -> JsonValue { @@ -4642,8 +4303,6 @@ mod tests { fn arrival_of(id: &str, from: &str, to: &str) -> routing::Arrival { routing::Arrival { - phase: routing::ArrivalPhase::Active, - transferred: true, workspace_id: id.to_string(), from: from.to_string(), to: to.to_string(), @@ -4662,211 +4321,6 @@ mod tests { .to_string() } - fn reserve_test_arrival(windows: &WindowState) -> routing::Arrival { - let mut arrival = arrival_of("workspace-7", "main", "ws-2"); - arrival.phase = routing::ArrivalPhase::Preparing; - arrival.transferred = false; - arrival.terminal_ids = vec!["pane-a".to_string()]; - windows.mint("pane-a", "main"); - guard(&windows.arrivals).push(arrival.clone()); - arrival - } - - #[test] - fn a_pending_journal_reserves_teardown_without_hiding_or_moving_source_ptys() { - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - let mut arrivals = guard(&windows.arrivals); - arrivals[0].content = Some(serde_json::json!({})); - assert!(arrivals.defer_close("main")); - assert!(arrivals.defer_close("ws-2")); - assert_eq!(arrivals.defer_quit(&QuitIntent::default()), Some(false)); - assert!(routing::arrival_payloads(&arrivals, "ws-2").is_empty()); - assert_eq!(routing::boot_list_ids(windows.owned_by("main"), &arrivals), expected.terminal_ids); - assert!(guard(&windows.routing).marking.is_empty()); - } - - #[test] - fn failed_initial_journal_preserves_source_ownership_and_previous_durable_bytes() { - let dir = TempDir::new("arrival-initial-failure"); - let old = arrival_of("workspace-3", "main", "ws-3"); - record_arrival_on_disk(dir.path(), &old).unwrap(); - let previous = fs::read(arrivals_path(dir.path())).unwrap(); - write_session_to(dir.path(), "main", &snapshot_json(&[("workspace-7", "Source")], "workspace-7")).unwrap(); - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - - assert!(commit_initial_arrival_with(&windows, &expected, || Err("disk full".to_string())).is_err()); - assert_eq!(windows.owned_by("main"), expected.terminal_ids); - assert!(guard(&windows.routing).marking.is_empty()); - assert!(guard(&windows.routing).awaiting_replay.is_empty()); - assert!(routing::retire_arrival(&mut guard(&windows.arrivals), &expected)); - assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); - assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "main").unwrap()), vec!["workspace-7"]); - } - - #[test] - fn a_target_destroyed_before_initial_commit_never_drops_source_owned_shells() { - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - let (lost, orphaned) = windows.drop_window("ws-2"); - assert!(orphaned.is_empty()); - assert_eq!(lost.len(), 1); - assert!(!lost[0].transferred); - assert_eq!(windows.owned_by("main"), expected.terminal_ids); - assert!(commit_initial_arrival_with(&windows, &expected, || panic!("cancelled reservation must not write")).is_err()); - } - - #[test] - fn a_return_requested_during_initial_write_fences_late_ownership_and_recovers_source() { - let dir = TempDir::new("arrival-initial-return-race"); - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - write_session_to(dir.path(), "main", &snapshot_json(&[("workspace-7", "Source")], "workspace-7")).unwrap(); - let result = commit_initial_arrival_with(&windows, &expected, || { - record_arrival_in_locked_journal(dir.path(), &expected)?; - routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); - Ok(()) - }); - assert!(result.is_err()); - assert_eq!(windows.owned_by("main"), expected.terminal_ids); - assert!(guard(&windows.routing).marking.is_empty()); - let returning = guard(&windows.arrivals)[0].clone(); - commit_arrival_return_with(&windows, &returning, || return_arrival_in_locked_journal(dir.path(), &returning)).unwrap(); - restore_arrivals(dir.path()).unwrap(); - assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "main").unwrap()), vec!["workspace-7"]); - assert!(read_snapshot(dir.path(), "ws-2").is_none()); - } - - #[test] - fn a_failed_reverse_journal_retains_the_return_and_retries_without_losing_recovery() { - let dir = TempDir::new("arrival-return-failure"); - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - commit_initial_arrival_with(&windows, &expected, || record_arrival_in_locked_journal(dir.path(), &expected)).unwrap(); - let returning = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); - let previous = fs::read(arrivals_path(dir.path())).unwrap(); - assert!(commit_arrival_return_with(&windows, &returning, || Err("permission denied".to_string())).is_err()); - assert_eq!(windows.owned_by("ws-2"), expected.terminal_ids); - assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); - assert_eq!(guard(&windows.arrivals).defer_quit(&QuitIntent::default()), Some(false)); - assert!(routing::arrival_payloads(&guard(&windows.arrivals), "ws-2").is_empty()); - - assert_eq!(failed_arrival_return(&windows, &returning, ARRIVAL_RETURN_RETRIES), ArrivalReturnFailure::Retry(ARRIVAL_RETURN_RETRIES - 1)); - commit_arrival_return_with(&windows, &returning, || return_arrival_in_locked_journal(dir.path(), &returning)).unwrap(); - assert_eq!(failed_arrival_return(&windows, &returning, 0), ArrivalReturnFailure::Stale); - assert_eq!(windows.owned_by("main"), expected.terminal_ids); - assert!(guard(&windows.arrivals).is_empty()); - restore_arrivals(dir.path()).unwrap(); - assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "main").unwrap()), vec!["workspace-7"]); - assert!(read_snapshot(dir.path(), "ws-2").is_none()); - } - - #[test] - fn permanent_return_failure_bounds_quit_and_cold_restores_exactly_once() { - use super::quit_state::{QuitMachine, QuitAction}; - let dir = TempDir::new("arrival-return-exhausted"); - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - write_session_to(dir.path(), "main", &snapshot_json(&[("workspace-7", "Source")], "workspace-7")).unwrap(); - write_session_to(dir.path(), "ws-2", &snapshot_json(&[], "")).unwrap(); - commit_initial_arrival_with(&windows, &expected, || record_arrival_in_locked_journal(dir.path(), &expected)).unwrap(); - let returning = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); - let previous = fs::read(arrivals_path(dir.path())).unwrap(); - let intent = QuitIntent::restart(None); - assert_eq!(guard(&windows.arrivals).defer_quit(&intent), Some(true)); - for remaining in (0..=ARRIVAL_RETURN_RETRIES).rev() { - assert!(commit_arrival_return_with(&windows, &returning, || Err("read-only volume".into())).is_err()); - let result = failed_arrival_return(&windows, &returning, remaining); - assert_eq!(result, if remaining > 0 { ArrivalReturnFailure::Retry(remaining - 1) } - else { ArrivalReturnFailure::RecoveryPending }); - } - assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); - assert_eq!(windows.owned_by("ws-2"), expected.terminal_ids); - assert!(routing::arrival_payloads(&guard(&windows.arrivals), "ws-2").is_empty()); - assert!(commit_arrival_return_with(&windows, &returning, || panic!("parked writer must not touch disk")).is_err()); - assert!(commit_arrival_adoption_with(&windows, &expected, || panic!("late adoption must not write")).is_err()); - let live = || HashSet::from(["main".to_string(), "ws-2".to_string()]); - assert_eq!(guard(&windows.arrivals).take_ready(live), (Some(intent.clone()), vec![])); - assert_eq!(guard(&windows.arrivals).defer_quit(&intent), None); - // Actual global-quit machine takes its normal confirmation, then emits - // destroys without the per-window snapshot commit/removal path. - let mut machine = QuitMachine::default(); - let labels = vec!["main".to_string(), "ws-2".to_string()]; - assert_eq!(machine.request(&labels, intent.clone()).1, vec![QuitAction::RequestAll { requester: None }]); - assert!(!machine.all_acked(), "the normal no-ack watchdog is now reachable"); - assert_eq!(machine.cancel(), vec![QuitAction::CancelAll]); - guard(&windows.arrivals).cancel_deferred(); - assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); - assert_eq!(windows.owned_by("ws-2"), expected.terminal_ids); - assert!(guard(&windows.arrivals).defer_close("main")); - assert!(guard(&windows.arrivals).defer_close("ws-2")); - assert_eq!(guard(&windows.arrivals).take_ready(live), (None, vec![])); - assert_eq!(guard(&windows.arrivals).defer_quit(&intent), None); - assert_eq!(machine.request(&labels, intent.clone()).1, vec![QuitAction::RequestAll { requester: None }]); - machine.ack("main"); machine.ack("ws-2"); - assert!(machine.vote("main").is_empty()); - assert_eq!(machine.vote("ws-2"), vec![QuitAction::Teardown { label: "ws-2".into(), last: false }]); - assert_eq!(machine.window_done("ws-2"), vec![QuitAction::Destroy { label: "ws-2".into() }, - QuitAction::Teardown { label: "main".into(), last: true }]); - let (lost, orphaned) = windows.drop_window("ws-2"); - assert!(lost.is_empty(), "Destroyed must not restart a parked reverse write"); - assert!(orphaned.is_empty(), "claimed PTYs must not be sibling-close kills"); - assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); - assert!(read_snapshot(dir.path(), "main").is_some()); - assert!(read_snapshot(dir.path(), "ws-2").is_some()); - assert!(machine.proceed().contains(&QuitAction::Exit)); - restore_arrivals(dir.path()).unwrap(); - assert!(read_snapshot(dir.path(), "main").is_none(), "empty source is removed by cold restore"); - assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "ws-2").unwrap()), vec!["workspace-7"]); - restore_arrivals(dir.path()).unwrap(); - assert_eq!(snapshot_ids(&read_snapshot(dir.path(), "ws-2").unwrap()), vec!["workspace-7"]); - } - - #[test] - fn exhausted_return_cannot_park_or_write_a_newer_generation() { - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - commit_initial_arrival_with(&windows, &expected, || Ok(())).unwrap(); - let old = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); - assert!(routing::retire_arrival(&mut guard(&windows.arrivals), &old)); - let mut newer = old.clone(); - newer.queued_at += std::time::Duration::from_millis(1); - guard(&windows.arrivals).push(newer.clone()); - assert_eq!(failed_arrival_return(&windows, &old, 0), ArrivalReturnFailure::Stale); - assert_eq!(guard(&windows.arrivals)[0], newer); - assert!(commit_arrival_return_with(&windows, &old, || panic!("stale generation must not write")).is_err()); - } - - #[test] - fn a_return_retry_survives_target_destruction_without_reaping_source_shells() { - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - commit_initial_arrival_with(&windows, &expected, || Ok(())).unwrap(); - let returned = routing::return_arrival(&mut guard(&windows.arrivals), &expected.workspace_id, &expected.to).unwrap(); - let (lost, orphaned) = windows.drop_window("ws-2"); - assert!(lost.is_empty(), "one return worker is already responsible"); - assert!(orphaned.is_empty(), "a return's PTYs are not sibling-window orphans"); - commit_arrival_return_with(&windows, &returned, || Ok(())).unwrap(); - assert_eq!(windows.owned_by("main"), expected.terminal_ids); - } - - #[test] - fn adoption_write_failure_or_a_concurrent_return_never_releases_the_source() { - let windows = WindowState::default(); - let expected = reserve_test_arrival(&windows); - commit_initial_arrival_with(&windows, &expected, || Ok(())).unwrap(); - let active = guard(&windows.arrivals)[0].clone(); - assert!(commit_arrival_adoption_with(&windows, &active, || Err("journal unavailable".to_string())).is_err()); - assert_eq!(guard(&windows.arrivals)[0].phase, routing::ArrivalPhase::Active); - assert!(commit_arrival_adoption_with(&windows, &active, || { - routing::return_arrival(&mut guard(&windows.arrivals), &active.workspace_id, &active.to).unwrap(); - Ok(()) - }).is_err()); - assert_eq!(guard(&windows.arrivals)[0].phase, routing::ArrivalPhase::Returning); - assert!(commit_arrival_adoption_with(&windows, &active, || panic!("stale adopter must not write")).is_err()); - } - fn read_snapshot(dir: &Path, label: &str) -> Option { read_session_from(dir, label) .unwrap() @@ -5846,27 +5300,6 @@ mod tests { ); } - #[test] - fn a_geometry_flush_captured_before_close_cannot_recreate_removed_geometry() { - let dir = TempDir::new("geometry-close-fence"); - let windows = WindowState::default(); - write_session_to(dir.path(), "ws-2", "snapshot").unwrap(); - assert!(super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", "previous").unwrap()); - let captured_before_close = "new position"; - windows.begin_closing("ws-2"); - super::close_window_snapshot(dir.path(), "ws-2").unwrap(); - assert!(!super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", captured_before_close).unwrap()); - assert!(!super::geometry_path(dir.path(), "ws-2").exists()); - windows.drop_window("ws-2"); - assert!(!super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", captured_before_close).unwrap()); - assert!(!super::geometry_path(dir.path(), "ws-2").exists()); - // A recoverable refusal clears closing while the window remains live. - windows.begin_closing("ws-3"); - guard(&windows.closing).remove("ws-3"); - assert!(super::write_open_window_geometry(Some(&windows), dir.path(), "ws-3", "retained window").unwrap()); - assert_eq!(fs::read_to_string(super::geometry_path(dir.path(), "ws-3")).unwrap(), "retained window"); - } - /// The cached box is fed by the window events alone, and it is what both the /// debounced write and the cross-window drag read (§Boot and geometry). #[test] @@ -5967,69 +5400,16 @@ mod tests { /// paths set it: the webview's own `remove_window_session`, and /// `finish_window_close` for the ack-timeout path where it never ran. #[test] - fn a_successfully_closed_window_refuses_even_saves_dispatched_before_destroyed() { + fn a_closing_window_refuses_every_later_save_until_it_is_destroyed() { let state = super::WindowState::default(); assert!(!state.refuses_save("ws-2")); state.begin_closing("ws-2"); assert!(state.refuses_save("ws-2")); // Never a sibling's. assert!(!state.refuses_save("main")); - state.drop_window("ws-2"); - assert!(state.refuses_save("ws-2")); - } - - #[test] - fn cancelled_close_never_mutates_retained_or_resaved_window() { - let dir = TempDir::new("close-cancelled"); - write_session_to(dir.path(), "ws-2", "retained").unwrap(); - let token = super::close_commit::CloseCommit::default(); - assert!(token.cancel()); - assert!(super::close_window_snapshot_locked(dir.path(), "ws-2", Some(&token)).is_err()); - write_session_to(dir.path(), "ws-2", "newer save").unwrap(); - assert!(super::close_window_snapshot_locked(dir.path(), "ws-2", Some(&token)).is_err()); - assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("newer save")); - } - - #[test] - fn failed_close_rolls_back_journal_and_geometry_before_retry() { - let dir = TempDir::new("close-rollback"); - let arrival = arrival_of("ws-7", "main", "ws-2"); - record_arrival_on_disk(dir.path(), &arrival).unwrap(); - super::mark_arrival_adopted_on_disk(dir.path(), "ws-7").unwrap(); - let previous = fs::read(arrivals_path(dir.path())).unwrap(); - let geometry = super::geometry_path(dir.path(), "ws-2"); - fs::write(&geometry, "geometry").unwrap(); - // A directory cannot be removed as a session file on any platform. - let session = dir.path().join(session_file_name("ws-2")); - fs::create_dir(&session).unwrap(); - assert!(super::close_window_snapshot(dir.path(), "ws-2").is_err()); - assert_eq!(fs::read(arrivals_path(dir.path())).unwrap(), previous); - assert_eq!(fs::read_to_string(&geometry).unwrap(), "geometry"); - fs::remove_dir(&session).unwrap(); - write_session_to(dir.path(), "ws-2", &snapshot_json(&[("ws-7", "Moved")], "ws-7")).unwrap(); - write_session_to(dir.path(), "main", &snapshot_json(&[("ws-7", "Moved")], "ws-7")).unwrap(); - super::close_window_snapshot(dir.path(), "ws-2").unwrap(); - restore_arrivals(dir.path()).unwrap(); - assert!(read_session_from(dir.path(), "ws-2").unwrap().is_none()); - assert!(read_session_from(dir.path(), "main").unwrap().is_none()); - } - - #[test] - fn skipped_close_save_is_rejected_and_retained_window_can_retry() { - let dir = TempDir::new("close-cache-refusal"); - let windows = WindowState::default(); - write_session_to(dir.path(), "ws-2", "previous").unwrap(); - windows.begin_closing("ws-2"); - assert!(super::save_open_window_session(Some(&windows), dir.path(), "ws-2", "newer").is_err()); - assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("previous")); - assert!(read_session_from(dir.path(), "main").unwrap().is_none()); - guard(&windows.closing).remove("ws-2"); // recoverable close preparation failure - super::save_open_window_session(Some(&windows), dir.path(), "ws-2", "newer").unwrap(); - assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("newer")); - windows.begin_closing("ws-2"); - windows.drop_window("ws-2"); - assert!(super::save_open_window_session(Some(&windows), dir.path(), "ws-2", "stale queued save").is_err()); - assert_eq!(read_session_from(dir.path(), "ws-2").unwrap().as_deref(), Some("newer")); + // `Destroyed` drops the refusal: no save can arrive under a dead label. + guard(&state.closing).remove("ws-2"); + assert!(!state.refuses_save("ws-2")); } fn queue_test_suppression(state: &super::WindowState, id: &str) { @@ -6294,8 +5674,6 @@ mod tests { "record_arrival_on_disk", "mark_arrival_adopted_on_disk", "return_arrival_on_disk", "forget_arrival_on_disk", "read_arrivals_from", "write_arrivals_to", "restore_arrivals", "close_window_snapshot", "finish_window_close", "begin_arrival", "hand_back_arrival", - "remove_window_session_bounded", "commit_initial_arrival_with", - "commit_arrival_adoption_with", "commit_arrival_return_with", ].iter().any(|helper| body.contains(helper)); if reaches_blocking && !(is_async_attr || is_async_fn) { offenders.push(name.to_string()); diff --git a/standalone/src-tauri/src/quit_state.rs b/standalone/src-tauri/src/quit_state.rs index 445fa0287..d627d14ae 100644 --- a/standalone/src-tauri/src/quit_state.rs +++ b/standalone/src-tauri/src/quit_state.rs @@ -6,7 +6,7 @@ //! transitions live here, free of Tauri, and hand the caller a list of actions //! to perform; `lib.rs` owns the emitting, destroying and exiting. -use crate::routing::{quit_order, ArrivalPhase, Arrivals}; +use crate::routing::{quit_order, Arrivals}; use std::collections::{HashMap, HashSet}; use std::ops::{Deref, DerefMut}; @@ -35,21 +35,14 @@ impl DerefMut for ArrivalQueue { } impl ArrivalQueue { - /// Only unsettled live work delays global quit. Exhausted returns retain - /// their recovery/individual-close fences, but global quit preserves the - /// journal and snapshots and may proceed through its normal human gate. + /// Queue a quit while anything is in flight. `None` means nothing is in + /// flight and the quit runs now; otherwise whether the queued quit relaunches. pub fn defer_quit(&mut self, intent: &QuitIntent) -> Option { - if !self.blocks_quit() { return None; } + if self.records.is_empty() { return None; } let restart = self.quit.get_or_insert_with(|| intent.clone()).restart; self.closes.clear(); Some(restart) } - /// Global quit admission consumes the first queued intent even when a new - /// trigger arrives before the scheduled redrive; individual closes yield. - pub fn take_quit_intent(&mut self, fallback: QuitIntent) -> QuitIntent { - self.closes.clear(); - self.quit.take().unwrap_or(fallback) - } /// Queue `label`'s close while it is either end of a transfer; whether it /// was queued (a queued quit absorbs it). pub fn defer_close(&mut self, label: &str) -> bool { @@ -57,25 +50,20 @@ impl ArrivalQueue { if self.quit.is_none() { self.closes.insert(label.to_string()); } true } - fn blocks_quit(&self) -> bool { - self.records.iter().any(|arrival| arrival.phase != ArrivalPhase::RecoveryPending) - } pub fn blocks_transfer(&self, from: &str, to: &str) -> bool { self.quit.is_some() || self.closes.contains(from) || self.closes.contains(to) - || self.records.iter().any(|arrival| arrival.phase == ArrivalPhase::RecoveryPending - && [from, to].into_iter().any(|label| arrival.from == label || arrival.to == label)) } pub fn cancel_deferred(&mut self) { self.quit = None; self.closes.clear(); } pub fn forget_deferred_close(&mut self, label: &str) { self.closes.remove(label); } /// Take the requests no transfer holds any more: the quit (with its intent) - /// once no live settlement is in flight, else every close whose window is no endpoint. + /// once nothing is in flight, else every close whose window is no endpoint. /// `live` (the open window labels) is read only when something is queued. pub fn take_ready(&mut self, live: impl FnOnce() -> HashSet) -> (Option, Vec) { if self.quit.is_none() && self.closes.is_empty() { return (None, Vec::new()); } let live = live(); self.closes.retain(|label| live.contains(label)); if self.quit.is_some() { - if !self.blocks_quit() { return (self.quit.take(), Vec::new()); } + if self.records.is_empty() { return (self.quit.take(), Vec::new()); } return (None, Vec::new()); } let endpoints: HashSet<&str> = @@ -475,8 +463,6 @@ mod tests { fn arrival(from: &str, to: &str) -> crate::routing::Arrival { crate::routing::Arrival { - phase: crate::routing::ArrivalPhase::Active, - transferred: true, workspace_id: format!("{from}-to-{to}"), from: from.to_string(), to: to.to_string(), @@ -530,32 +516,6 @@ mod tests { assert!(!queue.blocks_transfer("main", "ws-2")); } - #[test] - fn parked_recovery_allows_only_global_quit_and_keeps_fences_after_cancel() { - let mut queue = ArrivalQueue::default(); - let live = || HashSet::from(["main".to_string(), "ws-2".to_string(), "ws-3".to_string()]); - let mut parked = arrival("main", "ws-2"); - parked.phase = ArrivalPhase::RecoveryPending; - queue.push(parked); - queue.push(arrival("ws-3", "main")); - let intent = QuitIntent::restart(Some("requester".into())); - assert_eq!(queue.defer_quit(&intent), Some(true), "live settlements still block"); - assert_eq!(queue.take_ready(live), (None, vec![])); - queue.pop(); - assert_eq!(queue.defer_quit(&QuitIntent::default()), None, "repeated quit must reach its watchdog"); - assert_eq!(queue.take_quit_intent(QuitIntent::default()), intent, "first queued restart owns intent"); - assert_eq!(queue.take_ready(live), (None, vec![])); - queue.cancel_deferred(); - assert!(queue.defer_close("main")); - assert!(queue.defer_close("ws-2")); - assert_eq!(queue.take_ready(live), (None, vec![]), "no destructive individual close"); - assert!(queue.blocks_transfer("main", "ws-3")); - assert!(queue.blocks_transfer("ws-3", "ws-2")); - assert!(crate::routing::arrival_payloads(&queue, "ws-2").is_empty()); - assert!(crate::routing::return_arrival(&mut queue, "main-to-ws-2", "ws-2").is_none()); - assert!(crate::routing::take_arrivals_to(&mut queue, "ws-2").is_empty()); - } - #[test] fn stalled_cleanup_cannot_block_an_approved_exit_forever() { let mut gate = CleanupGate::default(); diff --git a/standalone/src-tauri/src/routing.rs b/standalone/src-tauri/src/routing.rs index 5335c7207..3c460670a 100644 --- a/standalone/src-tauri/src/routing.rs +++ b/standalone/src-tauri/src/routing.rs @@ -245,14 +245,6 @@ pub fn route<'a>(event: &str, data: &'a JsonValue, view: &RouteView<'a>) -> Rout /// drop target, and an `emit_to` it would simply be lost. #[derive(Debug, Clone, PartialEq)] pub struct Arrival { - /// Preparing reserves both endpoints while disk I/O runs without the - /// arrival lock. Returning retains that reservation until its reverse - /// destination is durable. RecoveryPending parks an exhausted return for - /// cold recovery: only global quit may pass it. None is adoptable. - pub phase: ArrivalPhase, - /// A Preparing arrival that is cancelled has never moved the source's - /// shells; its return must not hide or reap those source-owned PTYs. - pub transferred: bool, pub workspace_id: String, /// The window that still shows the Workspace until the target adopts it. pub from: String, @@ -274,9 +266,6 @@ pub struct Arrival { pub pending_window: Option, } -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum ArrivalPhase { Preparing, Active, Returning, RecoveryPending } - /// Every arrival in flight, oldest first. A Vec, not a map: there are a handful /// at most, and both the per-window drain and the by-Workspace lookup want the /// order the drops happened in. @@ -310,7 +299,7 @@ pub fn take_arrival(arrivals: &mut Arrivals, workspace_id: &str, to: &str) -> Op Some(arrivals.remove(position)) } -/// Request return of an arrival that outlived `ARRIVAL_MAX`, but only the exact record +/// Retire an arrival that outlived `ARRIVAL_MAX`, but **only the exact record /// the watchdog was armed for**: one adopted and re-dropped since would carry a /// later `queued_at`, and belongs to its own watchdog. pub fn expire_arrival( @@ -319,48 +308,23 @@ pub fn expire_arrival( to: &str, queued_at: Instant, ) -> Option { - let arrival = arrivals.iter_mut().find(|arrival| { + let position = arrivals.iter().position(|arrival| { arrival.workspace_id == workspace_id && arrival.to == to && arrival.queued_at == queued_at - && matches!(arrival.phase, ArrivalPhase::Preparing | ArrivalPhase::Active) - })?; - arrival.phase = ArrivalPhase::Returning; - Some(arrival.clone()) -} - -/// Reserve settlement before releasing the lock. A close/quit must continue -/// waiting while the return destination is being written or retried. A parked -/// RecoveryPending record cannot start a second return chain. -pub fn return_arrival(arrivals: &mut Arrivals, workspace_id: &str, to: &str) -> Option { - let arrival = arrivals.iter_mut().find(|arrival| { - arrival.workspace_id == workspace_id && arrival.to == to - && matches!(arrival.phase, ArrivalPhase::Preparing | ArrivalPhase::Active) })?; - arrival.phase = ArrivalPhase::Returning; - Some(arrival.clone()) -} - -/// Retire only the generation whose durable settlement finished. -pub fn retire_arrival(arrivals: &mut Arrivals, expected: &Arrival) -> bool { - let Some(position) = arrivals.iter().position(|arrival| { - arrival.workspace_id == expected.workspace_id && arrival.to == expected.to - && arrival.queued_at == expected.queued_at && arrival.phase == expected.phase - }) else { return false; }; - arrivals.remove(position); - true + Some(arrivals.remove(position)) } -/// Reserve returns for arrivals whose target is gone, retaining their teardown -/// blockers until the reverse journal commits. Parked recovery stays reserved, -/// and existing return workers own -/// their generation and are not started a second time. +/// Every arrival `label` will never take, removed: its window is gone. pub fn take_arrivals_to(arrivals: &mut Arrivals, label: &str) -> Vec { let mut lost = Vec::new(); - for arrival in arrivals { - if arrival.to == label && matches!(arrival.phase, ArrivalPhase::Preparing | ArrivalPhase::Active) { - arrival.phase = ArrivalPhase::Returning; + arrivals.retain(|arrival| { + if arrival.to == label { lost.push(arrival.clone()); + false + } else { + true } - } + }); lost } @@ -371,7 +335,7 @@ pub fn take_arrivals_to(arrivals: &mut Arrivals, label: &str) -> Vec { pub fn arrival_payloads(arrivals: &Arrivals, label: &str) -> Vec { arrivals .iter() - .filter(|arrival| arrival.to == label && arrival.phase == ArrivalPhase::Active) + .filter(|arrival| arrival.to == label) .filter_map(|arrival| { let content = arrival.content.as_ref()?; let mut payload = arrival.payload.clone(); @@ -422,14 +386,13 @@ pub fn hand_back_ids(arrival: &Arrival, marks: &JsonValue) -> (Vec, Vec< pub fn arrival_ids(arrivals: &Arrivals) -> HashSet { arrivals .iter() - .filter(|arrival| arrival.transferred) .flat_map(|arrival| arrival.terminal_ids.iter().cloned()) .collect() } /// What a window's own `pty:requestInit` may name — and what its teardown may /// kill or interrupt: the ids it owns, **minus every id an arrival claims**. -/// Ownership moves after journal admission, so a window with a Workspace queued +/// Ownership moves at the source's invoke, so a window with a Workspace queued /// for it owns those shells while the source is still showing them; listed at /// boot they would be placed as top-level panes beside the Workspace about to /// mount them, and in a teardown's kill set they would die under the source. @@ -902,8 +865,6 @@ mod tests { fn arrival(workspace_id: &str, from: &str, to: &str, ids: &[&str]) -> Arrival { Arrival { - phase: ArrivalPhase::Active, - transferred: true, workspace_id: workspace_id.to_string(), from: from.to_string(), to: to.to_string(), @@ -972,12 +933,10 @@ mod tests { assert_eq!(expire_arrival(&mut arrivals, "ws-a", "ws-2", armed_for), None); assert_eq!(arrivals, vec![second.clone()]); - let mut returned = second; - returned.phase = ArrivalPhase::Returning; - assert_eq!(expire_arrival(&mut arrivals, "ws-a", "ws-2", returned.queued_at), Some(returned.clone())); - assert_eq!(arrivals, vec![returned.clone()], "return keeps blocking teardown until its journal commits"); - assert_eq!(expire_arrival(&mut arrivals, "ws-a", "ws-2", returned.queued_at), None); - assert!(retire_arrival(&mut arrivals, &returned)); + assert_eq!( + expire_arrival(&mut arrivals, "ws-a", "ws-2", second.queued_at), + Some(second) + ); assert!(arrivals.is_empty()); } diff --git a/standalone/src/WorkspaceTeardownModal.tsx b/standalone/src/WorkspaceTeardownModal.tsx index 5eacab1da..40795f7df 100644 --- a/standalone/src/WorkspaceTeardownModal.tsx +++ b/standalone/src/WorkspaceTeardownModal.tsx @@ -2,7 +2,7 @@ import { useCallback, useRef, useSyncExternalStore } from 'react'; // Standalone reaches into the lib source directly (same relative form as the // sibling UpdateDebugModal.tsx). The terminal registry comes in via the // `dormouse-lib` alias, matching quit.ts. -import { ModalFrame, modalActionButton } from '../../lib/src/components/design'; +import { ModalFrame } from '../../lib/src/components/design'; import { WorkspaceKillConfirm } from '../../lib/src/components/WorkspaceKillConfirm'; import { subscribeToTerminalPaneState } from 'dormouse-lib/lib/terminal-registry'; import { @@ -12,8 +12,6 @@ import { getQuitConfirmChar, getQuitConfirmWorkspaceNames, getQuitConfirmPhase, - getCloseFailure, - getQuitProgressDetail, quitRunningWork, subscribeQuitConfirm, type QuitConfirmIntent, @@ -30,8 +28,6 @@ import { export function WorkspaceTeardownModalHost() { const phase = useSyncExternalStore(subscribeQuitConfirm, getQuitConfirmPhase); const intent = useSyncExternalStore(subscribeQuitConfirm, getQuitConfirmIntent); - const failure = useSyncExternalStore(subscribeQuitConfirm, getCloseFailure); - const progressDetail = useSyncExternalStore(subscribeQuitConfirm, getQuitProgressDetail); if (!phase) return null; return ( @@ -40,8 +36,6 @@ export function WorkspaceTeardownModalHost() { workspaceNames={getQuitConfirmWorkspaceNames()} confirming={phase === 'quitting'} intent={intent} - failure={phase === 'close-failed' ? failure : null} - progressDetail={progressDetail} /> ); } @@ -53,8 +47,6 @@ export function WorkspaceTeardownModal({ char = 'q', workspaceNames = [], intent = { kind: 'quit' }, - failure = null, - progressDetail = null, }: { confirming: boolean; char?: string; @@ -62,8 +54,6 @@ export function WorkspaceTeardownModal({ /** Whether this tears down the whole app or one window, and whether that * discards a downloaded update. A quit and a restart read the same. */ intent?: QuitConfirmIntent; - failure?: ReturnType; - progressDetail?: string | null; }) { const progressRef = useRef(null); // Live count — the dialog stays open even if it drops to 0 (see spec). @@ -72,21 +62,12 @@ export function WorkspaceTeardownModal({ const runningCount = useSyncExternalStore(subscribeToTerminalPaneState, getRunningCount); const hasRunning = runningCount > 0; - if (failure) { - return -

Could not close window

-

{failure.reason}

- - -
; - } - if (confirming) { return (

Confirm kill workspace

- {progressDetail ?? (intent.kind === 'quit' ? 'Waiting for all windows, then closing…' : 'Closing workspaces…')} + {intent.kind === 'quit' ? 'Waiting for all windows, then closing…' : 'Closing workspaces…'}

); diff --git a/standalone/src/quit-confirm-store.ts b/standalone/src/quit-confirm-store.ts index 0f21c420c..846fce41d 100644 --- a/standalone/src/quit-confirm-store.ts +++ b/standalone/src/quit-confirm-store.ts @@ -12,23 +12,7 @@ import type { TeardownConfirmContext } from "./teardown-flow"; * docs/specs/standalone.md §Quit flow, "Confirmation dialog". */ -export type QuitConfirmPhase = "open" | "quitting" | "close-failed"; -export interface CloseFailure { reason: string; retry: () => void; stay: () => void; } -let closeFailure: CloseFailure | null = null; -let progressDetail: string | null = null; -export function getCloseFailure(): CloseFailure | null { return closeFailure; } -export function getQuitProgressDetail(): string | null { return progressDetail; } -export function showCloseFailure(failure: CloseFailure): void { - closeFailure = failure; - intent = { kind: 'close-window' }; - phase = 'close-failed'; - ownDialog(); - emit(); -} -export function showCloseCommitUncertain(reason: string): void { - progressDetail = reason; - emit(); -} +export type QuitConfirmPhase = "open" | "quitting"; /** * What the dialog is asking about. A quit tears every window down; a @@ -143,8 +127,6 @@ export function openQuitConfirm(ctx: TeardownConfirmContext, next: QuitConfirmIn /** Own the window before voting, including an all-idle request. */ export function beginQuitProgress(next: QuitConfirmIntent): void { - closeFailure = null; - progressDetail = null; stopWatchingWorkspaces(); activeCtx = null; intent = next; @@ -189,8 +171,6 @@ export function dismissQuitConfirm(kind?: QuitConfirmIntent["kind"]): void { activeCtx = null; phase = null; intent = QUIT_INTENT; - closeFailure = null; - progressDetail = null; emit(); } @@ -203,6 +183,4 @@ export function _resetQuitConfirmForTesting(): void { intent = QUIT_INTENT; activeCtx = null; listeners.clear(); - closeFailure = null; - progressDetail = null; } diff --git a/standalone/src/quit.ts b/standalone/src/quit.ts index 96770e804..70d15a0b6 100644 --- a/standalone/src/quit.ts +++ b/standalone/src/quit.ts @@ -136,7 +136,7 @@ async function runQuitTeardown(last: boolean): Promise { `[quit] teardown exceeded ${QUIT_TEARDOWN_CEILING_MS}ms; proceeding to exit`, ); } - // Install after bounded flush/drain attempts, in the window the walk + // Install strictly after the completed final save, in the window the walk // tears down last. Only `main` ever holds a pending download and the // updater capability (docs/specs/auto-update.md), so this is `main` or a // no-op. A fresh `quit_progress` gives install its own watchdog budget diff --git a/standalone/src/tauri-adapter.ts b/standalone/src/tauri-adapter.ts index 7ea0b6d81..e7e97a1fd 100644 --- a/standalone/src/tauri-adapter.ts +++ b/standalone/src/tauri-adapter.ts @@ -594,17 +594,12 @@ export class TauriAdapter implements PlatformAdapter { this.pendingFlushRequests.set(requestId, resolve); // Timeout is a synthetic completion; a stale timer after a real completion // hits notify's map-miss guard. Fan out after registering so a synchronous - // completion still finds the entry. WorkspaceWindow's single listener - // waits for every Wall before reporting completion for this window. + // completion still finds the entry (first notify wins — one Wall ships). setTimeout(() => this.notifySessionFlushComplete(requestId), timeoutMs); for (const handler of this.flushHandlers) handler({ requestId, ...options }); }); } - /** Keep one latest-value retry after close rollback, including a refused - * write that outlives the bounded drain. Never retries a successful save. */ - retrySessionSave(): void { this.sessionStore.retryLatest(); } - // Await the session store's in-flight/pending save_session pipeline (the Rust // temp+fsync+rename that actually reaches disk). Bounded: on timeout resolve // anyway rather than wedge quit. diff --git a/standalone/src/tauri-session-store.test.ts b/standalone/src/tauri-session-store.test.ts index 88dec4f97..0f1d4479d 100644 --- a/standalone/src/tauri-session-store.test.ts +++ b/standalone/src/tauri-session-store.test.ts @@ -1,4 +1,4 @@ -import { describe, it, expect, vi } from "vitest"; +import { describe, it, expect } from "vitest"; import { TauriSessionStore } from "./tauri-session-store"; const tick = () => new Promise((r) => setTimeout(r, 0)); @@ -128,26 +128,6 @@ describe("TauriSessionStore", () => { expect(saved).toEqual(["a", "a"]); }); - it("persists an unchanged cache after close preparation refused its write and the window stayed open", async () => { - let closing = true; - let persisted = "previous"; - const store = new TauriSessionStore(async (value) => { - // Rust save_open_window_session refuses rather than acknowledging a - // skipped write while close preparation owns the snapshot. - if (closing) throw new Error("window close is holding its snapshot; no session was saved"); - persisted = value; - }); - store.hydrate(persisted); - store.setItem("k", "newer"); - await store.drain(); - expect(persisted).toBe("previous"); - expect(store.getItem("k")).toBe("newer"); - closing = false; - store.setItem("k", "newer"); - await store.drain(); - expect(persisted).toBe("newer"); - }); - it("drain resolves only after an in-flight save settles", async () => { let release!: () => void; const store = new TauriSessionStore( @@ -213,35 +193,4 @@ describe("TauriSessionStore", () => { await tick(); expect(drained).toBe(true); }); - it("remembers only one idle retry and uses the latest queued value", async () => { - const log = vi.spyOn(console, 'error').mockImplementation(() => {}); - try { - let rejectFirst!: (error: Error) => void; - const values: string[] = []; - const store = new TauriSessionStore(async (value) => { - values.push(value); - if (values.length === 1) await new Promise((_, reject) => { rejectFirst = reject; }); - else throw new Error('still unavailable'); - }); - store.hydrate('previous'); - store.setItem('', 'first'); - store.retryLatest(); - store.setItem('', 'newest'); - rejectFirst(new Error('refused')); - await store.drain(); - expect(values).toEqual(['first', 'newest']); - } finally { log.mockRestore(); } - }); - - it("does not duplicate an in-flight successful save when retry is remembered", async () => { - let complete!: () => void; - const save = vi.fn(() => new Promise((resolve) => { complete = resolve; })); - const store = new TauriSessionStore(save); - store.setItem('', 'latest'); - store.retryLatest(); - complete(); - await store.drain(); - expect(save).toHaveBeenCalledOnce(); - }); - }); diff --git a/standalone/src/tauri-session-store.ts b/standalone/src/tauri-session-store.ts index 134936d32..89b8d9d8c 100644 --- a/standalone/src/tauri-session-store.ts +++ b/standalone/src/tauri-session-store.ts @@ -32,7 +32,6 @@ export class TauriSessionStore implements SessionKeyValueStore { // queued. A queued `value` is always a JSON string (never JS null), so null is // a safe "nothing pending" sentinel — even an empty-string blob is distinct. private pending: string | null = null; - private retryLatestWhenIdle = false; // Resolvers for pending drain() calls, fired when the pipeline next goes idle. private drainWaiters: Array<() => void> = []; @@ -50,14 +49,6 @@ export class TauriSessionStore implements SessionKeyValueStore { }); } - /** Remember one retry of the latest unsaved value, even if a bounded drain - * returned before the current save settled. Newer setItem values replace it; - * a successful current write needs no duplicate, and failure never loops. */ - retryLatest(): void { - if (this.saveInFlight) this.retryLatestWhenIdle = true; - else if (this.cache !== null && this.cache !== this.savedValue) this.setItem('', this.cache); - } - /** Seed the cache from the host's persisted blob (or null) at boot. */ hydrate(seed: string | null): void { this.cache = seed; @@ -90,14 +81,10 @@ export class TauriSessionStore implements SessionKeyValueStore { .then(() => { this.savedValue = value; }) .catch((err) => console.error("[tauri-session-store] save_session failed:", err)) .finally(() => { - const retryLatest = this.retryLatestWhenIdle; - this.retryLatestWhenIdle = false; if (this.pending !== null) { const next = this.pending; this.pending = null; this.flush(next); - } else if (retryLatest && this.cache !== null && this.cache !== this.savedValue) { - this.flush(this.cache); } else { this.saveInFlight = false; // Pipeline idle: release drain waiters. diff --git a/standalone/src/updater.test.ts b/standalone/src/updater.test.ts index 1ee5a6e24..05571c864 100644 --- a/standalone/src/updater.test.ts +++ b/standalone/src/updater.test.ts @@ -121,53 +121,6 @@ describe('updater', () => { mocks.platform = { requestAppRestart: mocks.requestAppRestart, burrow: { command: mocks.burrowCommand } }; }); - it('does not reoffer an approved download when the delayed launch check begins', async () => { - mocks.check.mockResolvedValue(makeUpdate('0.5.0')); - startUpdateCheck(); - await vi.advanceTimersByTimeAsync(0); - checkNow(); - await vi.advanceTimersByTimeAsync(0); - approveUpdate(); - await vi.advanceTimersByTimeAsync(0); - expect(hasPendingUpdate()).toBe(true); - - await vi.advanceTimersByTimeAsync(5_000); - expect(mocks.check).toHaveBeenCalledOnce(); - expect(readBannerState()).toEqual({ status: 'downloaded', version: '0.5.0' }); - }); - - it('does not reoffer an approval while its download is pending at the launch check', async () => { - const update = makeUpdate('0.5.0'); - update.download.mockImplementation(() => new Promise(() => {})); - mocks.check.mockResolvedValue(update); - startUpdateCheck(); - await vi.advanceTimersByTimeAsync(0); - checkNow(); - await vi.advanceTimersByTimeAsync(0); - approveUpdate(); - await vi.advanceTimersByTimeAsync(5_000); - - expect(mocks.check).toHaveBeenCalledOnce(); - expect(readBannerState()).toEqual({ status: 'downloading', version: '0.5.0' }); - }); - - it('keeps an approval made while the delayed launch policy read is pending', async () => { - let answerPolicy!: (value: typeof CHECKS_ON) => void; - mocks.burrowCommand.mockImplementation(() => new Promise(resolve => { answerPolicy = resolve; })); - mocks.check.mockResolvedValue(makeUpdate('0.5.0')); - startUpdateCheck(); - await vi.advanceTimersByTimeAsync(5_000); - checkNow(); - await vi.advanceTimersByTimeAsync(0); - approveUpdate(); - await vi.advanceTimersByTimeAsync(0); - answerPolicy(CHECKS_ON); - await vi.advanceTimersByTimeAsync(0); - - expect(mocks.check).toHaveBeenCalledOnce(); - expect(readBannerState()).toEqual({ status: 'downloaded', version: '0.5.0' }); - }); - // Drive check → approve → download so an approved, downloaded update is pending. async function reachDownloadedUpdate(update: ReturnType) { mocks.check.mockResolvedValue(update); diff --git a/standalone/src/updater.ts b/standalone/src/updater.ts index 61a2ef758..4c6c91715 100644 --- a/standalone/src/updater.ts +++ b/standalone/src/updater.ts @@ -414,9 +414,6 @@ async function runUpdateCheck(): Promise { // Read at the check, so a change made meanwhile counts // (`docs/specs/remote-network.md` → "Updates"). const policy = await readNetworkPolicy(); - // A manual approval during the launch delay or policy lookup owns this - // session's update. Rechecking would offer it for approval a second time. - if (pendingUpdate || downloadPromise) return; if (policy && checksForUpdates(policy)) { // An update found is offered by `performCheck`. await performCheck().catch((e) => console.error('[updater] Check failed:', e)); diff --git a/standalone/src/window-close.test.ts b/standalone/src/window-close.test.ts index 07f3d3f74..40d8eadd4 100644 --- a/standalone/src/window-close.test.ts +++ b/standalone/src/window-close.test.ts @@ -1,8 +1,4 @@ import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; -import { flushWindowSession, installWindowSessionWriter, resetWindowSessionAggregator, seedWindowSession } from "dormouse-lib/lib/window-session-aggregator"; -import { withTimeout } from "./with-timeout"; -import { TauriSessionStore } from "./tauri-session-store"; -import type { PersistedWindow } from "dormouse-lib/lib/session-types"; import type { TauriAdapter } from "./tauri-adapter"; /** @@ -16,7 +12,7 @@ const mocks = vi.hoisted(() => ({ invoke: vi.fn(async (_cmd: string) => undefined as unknown), listen: vi.fn(), countRunningSessions: vi.fn(() => 0), - getWorkspacesSnapshot: vi.fn(() => ({ workspaces: [{ id: "w1", name: "Deploys", nameIsAuto: false }], activeId: "w1" })), + getWorkspacesSnapshot: vi.fn(() => ({ workspaces: [{ id: "w1", name: "Deploys" }], activeId: "w1" })), hasPendingUpdate: vi.fn(() => false), })); @@ -41,9 +37,6 @@ import { cancelQuit as dismissDialog, getQuitConfirmIntent, getQuitConfirmPhase, - getCloseFailure, - getQuitProgressDetail, - confirmQuit, _resetQuitConfirmForTesting, } from "./quit-confirm-store"; @@ -55,8 +48,6 @@ const commands = () => mocks.invoke.mock.calls.map((call) => call[0]); function fakeAdapter(order: string[] = []): TauriAdapter { return { gracefulKillPtys: vi.fn(async () => void order.push("gracefulKill")), - retrySessionSave: vi.fn(), - drainSessionSaves: vi.fn(async () => void order.push('drain')), captureAgentRecovery: vi.fn(async () => void order.push("captureRecovery")), } as unknown as TauriAdapter; } @@ -64,7 +55,6 @@ function fakeAdapter(order: string[] = []): TauriAdapter { describe("per-window close", () => { beforeEach(() => { vi.clearAllMocks(); - resetWindowSessionAggregator(); _resetWindowCloseForTesting(); _resetQuitConfirmForTesting(); listeners.clear(); @@ -77,7 +67,7 @@ describe("per-window close", () => { mocks.hasPendingUpdate.mockReturnValue(false); }); - afterEach(() => { _resetWindowCloseForTesting(); resetWindowSessionAggregator(); }); + afterEach(() => _resetWindowCloseForTesting()); it("acks, removes the snapshot, kills, and proceeds — with no recovery capture", async () => { const order: string[] = []; @@ -176,187 +166,4 @@ describe("per-window close", () => { expect(commands()).toContain("close_window"); }); - - it("retains live PTYs after failed snapshot removal and permits a fresh retry", async () => { - let refusals = 1; - mocks.invoke.mockImplementation(async (cmd) => { - if (cmd === 'remove_window_session' && refusals-- > 0) throw new Error('disk full'); - if (cmd === 'retry_window_close') closeRequested(); - }); - const adapter = fakeAdapter(); - initWindowClose(adapter); - closeRequested(); - await settle(); - expect(getQuitConfirmPhase()).toBe('close-failed'); - expect(getCloseFailure()?.reason).toContain('disk full'); - expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); - expect(commands()).not.toContain('close_window'); - expect(commands()).toContain('window_close_cancel'); - getCloseFailure()?.retry(); - await settle(); - expect(commands().filter((cmd) => cmd === 'remove_window_session')).toHaveLength(2); - expect(commands()).toContain('retry_window_close'); - expect(adapter.gracefulKillPtys).toHaveBeenCalledOnce(); - expect(commands()).toContain('close_window'); - }); - - it("drains a refused save and republishes unchanged aggregate state after close rollback", async () => { - const errorLog = vi.spyOn(console, 'error').mockImplementation(() => {}); - try { - let rejectRefused!: (error: Error) => void; - let disk = 'previous'; - let saveCalls = 0; - const store = new TauriSessionStore(async (value) => { - if (++saveCalls === 1) await new Promise((_, reject) => { rejectRefused = reject; }); - disk = value; - }); - store.hydrate(disk); - const latest: PersistedWindow = { version: 1, activeWorkspaceId: 'w1', workspaces: [ - { id: 'w1', name: 'Deploys', nameIsAuto: false, session: { version: 3, panes: [] } }, - ] }; - seedWindowSession(latest); - installWindowSessionWriter((snapshot) => store.setItem('state', JSON.stringify(snapshot))); - await flushWindowSession(); // the native close fence has refused this pending save - const order: string[] = []; - mocks.invoke.mockImplementation(async (cmd) => { - order.push(cmd); - if (cmd === 'remove_window_session') throw new Error('disk full'); - }); - const adapter = fakeAdapter(order); - adapter.drainSessionSaves = vi.fn(async () => { order.push('drain'); await store.drain(); }); - adapter.retrySessionSave = () => store.retryLatest(); - initWindowClose(adapter); - closeRequested(); - await settle(); - expect(order).toEqual(['window_close_ack', 'remove_window_session', 'window_close_cancel', 'drain']); - expect(saveCalls).toBe(1); - expect(getQuitConfirmPhase()).toBe('quitting'); - rejectRefused(new Error('window close is holding its snapshot; no session was saved')); - await settle(); - expect(saveCalls).toBe(2); - expect(JSON.parse(disk)).toMatchObject(latest); - expect(adapter.drainSessionSaves).toHaveBeenCalledTimes(2); - expect(getQuitConfirmPhase()).toBe('close-failed'); - expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); - } finally { errorLog.mockRestore(); } - }); - - it("retries after a refused save outlives both bounded recovery drains", async () => { - vi.useFakeTimers(); - const errorLog = vi.spyOn(console, 'error').mockImplementation(() => {}); - const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); - try { - let rejectFirst!: (error: Error) => void; - let disk = 'previous'; - let saves = 0; - const store = new TauriSessionStore(async (value) => { - if (++saves === 1) await new Promise((_, reject) => { rejectFirst = reject; }); - disk = value; - }); - store.hydrate(disk); - seedWindowSession({ version: 1, activeWorkspaceId: 'w1', workspaces: [ - { id: 'w1', name: 'Deploys', nameIsAuto: false, session: { version: 3, panes: [] } }, - ] }); - installWindowSessionWriter((snapshot) => store.setItem('', JSON.stringify(snapshot))); - await flushWindowSession(); - mocks.invoke.mockImplementation(async (cmd) => { - if (cmd === 'remove_window_session') throw new Error('disk full'); - }); - const adapter = fakeAdapter(); - adapter.drainSessionSaves = (ms) => withTimeout(store.drain(), ms, 'test drain timeout'); - adapter.retrySessionSave = () => store.retryLatest(); - initWindowClose(adapter); - closeRequested(); - await vi.advanceTimersByTimeAsync(4001); - expect(getQuitConfirmPhase()).toBe('close-failed'); - expect(saves).toBe(1); - expect(disk).toBe('previous'); - rejectFirst(new Error('close refused the earlier write')); - await vi.advanceTimersByTimeAsync(0); - expect(saves).toBe(2); - expect(JSON.parse(disk).workspaces[0].name).toBe('Deploys'); - await vi.advanceTimersByTimeAsync(60_000); - expect(saves).toBe(2); // one remembered retry, no loop - } finally { errorLog.mockRestore(); warn.mockRestore(); vi.useRealTimers(); } - }); - - it("keeps the flow guarded if native cancellation does not confirm release", async () => { - mocks.invoke.mockImplementation(async (cmd) => { - if (cmd === 'remove_window_session') throw new Error('disk full'); - if (cmd === 'window_close_cancel') throw new Error('native bridge unavailable'); - }); - const adapter = fakeAdapter(); - initWindowClose(adapter); - closeRequested(); - await settle(); - expect(getQuitConfirmPhase()).toBe('quitting'); - expect(getQuitProgressDetail()).toContain('save refusal could not be released'); - expect(getCloseFailure()).toBeNull(); - expect(adapter.drainSessionSaves).not.toHaveBeenCalled(); - closeRequested(); - await settle(); - expect(commands().filter((cmd) => cmd === 'remove_window_session')).toHaveLength(1); - }); - - it("bounds the aggregate resave without discarding retained live PTYs", async () => { - vi.useFakeTimers(); - const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); - try { - installWindowSessionWriter(() => new Promise(() => {})); - mocks.invoke.mockImplementation(async (cmd) => { - if (cmd === 'remove_window_session') throw new Error('disk full'); - }); - const adapter = fakeAdapter(); - initWindowClose(adapter); - closeRequested(); - await vi.advanceTimersByTimeAsync(1001); - expect(getQuitConfirmPhase()).toBe('close-failed'); - expect(adapter.drainSessionSaves).toHaveBeenCalledTimes(2); - expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); - expect(commands()).not.toContain('close_window'); - } finally { warn.mockRestore(); vi.useRealTimers(); } - }); - - it("leaves an entered uncertain commit guarded without offering a duplicate close", async () => { - mocks.invoke.mockImplementation(async (cmd) => { - if (cmd === 'remove_window_session') throw new Error('close-commit-uncertain: rollback failed'); - }); - const adapter = fakeAdapter(); - initWindowClose(adapter); - closeRequested(); - await settle(); - expect(getQuitConfirmPhase()).toBe('quitting'); - expect(getQuitProgressDetail()).toContain('rollback failed'); - expect(getCloseFailure()).toBeNull(); - closeRequested(); - await settle(); - expect(commands().filter((cmd) => cmd === 'remove_window_session')).toHaveLength(1); - expect(commands()).not.toContain('window_close_cancel'); - expect(adapter.drainSessionSaves).not.toHaveBeenCalled(); - expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); - }); - - it("does not time out a human confirmation or an entered native commit", async () => { - vi.useFakeTimers(); - try { - mocks.countRunningSessions.mockReturnValue(1); - let resolveRemoval!: () => void; - mocks.invoke.mockImplementation((cmd) => cmd === 'remove_window_session' - ? new Promise((resolve) => { resolveRemoval = resolve; }) : Promise.resolve()); - const adapter = fakeAdapter(); - initWindowClose(adapter); - closeRequested(); - await vi.advanceTimersByTimeAsync(60_000); - expect(getQuitConfirmPhase()).toBe('open'); - expect(commands()).not.toContain('remove_window_session'); - confirmQuit(); - await vi.advanceTimersByTimeAsync(60_000); - expect(getQuitConfirmPhase()).toBe('quitting'); - expect(adapter.gracefulKillPtys).not.toHaveBeenCalled(); - expect(commands()).not.toContain('close_window'); - resolveRemoval(); - await vi.advanceTimersByTimeAsync(0); - expect(adapter.gracefulKillPtys).toHaveBeenCalledOnce(); - } finally { vi.useRealTimers(); } - }); }); diff --git a/standalone/src/window-close.ts b/standalone/src/window-close.ts index fd205eb06..119290b00 100644 --- a/standalone/src/window-close.ts +++ b/standalone/src/window-close.ts @@ -1,7 +1,6 @@ import { invoke } from "@tauri-apps/api/core"; -import { flushWindowSession } from "dormouse-lib/lib/window-session-aggregator"; import { countRunningSessions } from "dormouse-lib/lib/terminal-registry"; -import { dismissQuitConfirm, openQuitConfirm, showCloseFailure, showCloseCommitUncertain } from "./quit-confirm-store"; +import { openQuitConfirm } from "./quit-confirm-store"; import { createTeardownFlow } from "./teardown-flow"; import type { TauriAdapter } from "./tauri-adapter"; import { hasPendingUpdate } from "./updater"; @@ -24,10 +23,8 @@ import { listenToWindow } from "./window-label"; */ const GRACEFUL_KILL_MS = 2000; -/** PTY teardown after durable removal; native preparation has its own bound. */ +/** The whole teardown, past the human decision. Well under Rust's own budget. */ const CLOSE_TEARDOWN_CEILING_MS = 8000; -const RETAINED_WINDOW_DRAIN_MS = 2000; -const RETAINED_WINDOW_WRITE_MS = 1000; let closeAdapter: TauriAdapter | null = null; @@ -52,65 +49,15 @@ export function initWindowClose(adapter: TauriAdapter): void { ...(hasPendingUpdate() ? { discardsUpdate: true } : {}), }); }); - void listenToWindow("dormouse://window-close-failed", (event) => { - void retainWindowAfterFailure(event.payload); - }); -} - -async function retainWindowAfterFailure(error: unknown): Promise { - const reason = String(error); - if (reason.includes('close-commit-uncertain:')) { - // The native commit still owns its files: never release the arbiter or - // enable a resave/second close while a late unlink may finish. - showCloseCommitUncertain(`Closing could not finish safely: ${reason}`); - return; - } - // Retire the previous native handshake before offering a retry. A late - // cancel must never clear the retry's fresh token while its dialog is open. - try { - await invoke('window_close_cancel'); - } catch (cancelError) { - showCloseCommitUncertain('The Window is retained, but its save refusal could not be released: ' + String(cancelError)); - return; - } - // A refused in-flight write must settle before publishing the same value: - // the synchronous cache coalesces identical values while its save is pending. - // Wall dirty tracking already ended when it published into the aggregate; - // a heartbeat may never republish this retained Window without this retry. - try { - if (closeAdapter) await closeAdapter.drainSessionSaves(RETAINED_WINDOW_DRAIN_MS); - await withTimeout( - flushWindowSession(), - RETAINED_WINDOW_WRITE_MS, - '[window-close] retained Window write timed out; Window stays open', - ); - closeAdapter?.retrySessionSave(); - if (closeAdapter) await closeAdapter.drainSessionSaves(RETAINED_WINDOW_DRAIN_MS); - } catch (saveError) { - console.warn('[window-close] retained Window resave failed:', saveError); - } - flow.reset(); - showCloseFailure({ - reason: `Workspaces were retained. ${reason}`, - retry: () => { - dismissQuitConfirm('close-window'); - void invoke('retry_window_close').catch(retainWindowAfterFailure); - }, - stay: () => dismissQuitConfirm('close-window'), - }); } async function runCloseTeardown(): Promise { const adapter = closeAdapter; try { - // Cancellation can win only before native disk mutation. A failure retains - // this window and its live PTYs; the old best-effort path lost recovery. - await invoke("remove_window_session"); - } catch (error) { - await retainWindowAfterFailure(error); - return; - } - try { + // Remove the snapshot BEFORE the kill, so an exit-triggered save cannot + // write it back: Rust refuses every later save for this label. + await invoke("remove_window_session").catch((err) => + console.warn("[window-close] remove_window_session failed; proceeding", err)); // No `ids`: Rust scopes the kill to this window's own PTYs, and a sibling's // terminals must never be reachable from here. if (adapter) { diff --git a/standalone/src/workspace-move.test.ts b/standalone/src/workspace-move.test.ts index 0361de2a2..e679fd352 100644 --- a/standalone/src/workspace-move.test.ts +++ b/standalone/src/workspace-move.test.ts @@ -89,7 +89,6 @@ import { disposeAllSessions, getOrCreateTerminal } from "dormouse-lib/lib/termin import { FakePtyAdapter } from "dormouse-lib/lib/platform/fake-adapter"; import { createAlertEpisode } from "dormouse-lib/lib/alert-episode"; import { clearTerminalActivity, getActivitySnapshot, setTerminalActivity } from "dormouse-lib/lib/session-activity-store"; -import { getWorkspaceUiSnapshot, resetWorkspaceUi } from 'dormouse-lib/lib/workspace-ui-store'; const WORKSPACE_ID = "ws-moving"; @@ -235,19 +234,6 @@ const emit = async (event: string, data: unknown) => { }; describe("the source half", () => { - it('keeps a failed durable handback pending while making its retry visible', async () => { - resetWorkspaceUi(); - registerWallHandle(stubWallHandle(WORKSPACE_ID, { prepareWorkspaceTransfer: async () => prepared() })); - initWorkspaceMoves(fakePlatform()); - const moved = transferWorkspaceTo(WORKSPACE_ID, 'ws-2'); - await contentSent(); - await emit('dormouse://workspace-arrival-retry', { workspaceId: WORKSPACE_ID, reason: 'Waiting for recovery write' }); - expect(getWorkspaceUiSnapshot().moveError).toEqual({ id: WORKSPACE_ID, reason: 'Waiting for recovery write' }); - expect(isWorkspaceTransferPending(WORKSPACE_ID)).toBe(true); - await emit('dormouse://workspace-arrival-failed', { workspaceId: WORKSPACE_ID, reason: 'returned', replayIds: [] }); - await expect(moved).resolves.toEqual({ moved: false, reason: 'returned' }); - expect(isWorkspaceTransferPending(WORKSPACE_ID)).toBe(false); - }); it.each([true, false])("refuses dirty editors without explicit discard (existing window: %s)", async (existing) => { const prepare = vi.fn(async () => prepared()); registerWallHandle(stubWallHandle(WORKSPACE_ID, { @@ -646,23 +632,6 @@ describe("the source half", () => { }); describe("the target half", () => { - it('requests handback promptly when durable adoption fails', async () => { - const platform = fakePlatform(); - const kill = vi.spyOn(platform, 'killPty'); - const host = mocks.invoke.getMockImplementation()!; - mocks.invoke.mockImplementation(async (command, args) => { - if (command === 'adopt_done') throw new Error('journal unavailable'); - return host(command, args); - }); - arrivals = [payload()]; - initWorkspaceMoves(platform); - await settle(); - await settle(); - expect(mocks.invoke).toHaveBeenCalledWith('adopt_failed', { workspaceId: WORKSPACE_ID, reason: 'journal unavailable' }); - expect(arrivals).toHaveLength(0); - expect(getWorkspacesSnapshot().workspaces.map(workspace => workspace.id)).not.toContain(WORKSPACE_ID); - expect(kill).not.toHaveBeenCalled(); - }); // The Sessions' alert state never left the sidecar's one manager, which // re-sends it to whichever window collects them. A spawn starts it over and a @@ -778,11 +747,7 @@ describe("the target half", () => { expect(prepare).not.toHaveBeenCalled(); expect(getTerminalInstance("pane-a")).toBeNull(); expect(killPty).not.toHaveBeenCalled(); - // A refused settled write explicitly requests durable handback. The - // watchdog may already have returned it; that stale request is harmless. - expect(mocks.invoke).toHaveBeenCalledWith("adopt_failed", { - workspaceId: WORKSPACE_ID, reason: `no arrival of '${WORKSPACE_ID}'`, - }); + expect(mocks.invoke).not.toHaveBeenCalledWith("adopt_failed", expect.anything()); }); it("discards this window's copy of the alert state when adopt_done is refused before the Wall mounts", async () => { diff --git a/standalone/src/workspace-move.ts b/standalone/src/workspace-move.ts index 07d8791b9..710b7b1b0 100644 --- a/standalone/src/workspace-move.ts +++ b/standalone/src/workspace-move.ts @@ -1,6 +1,6 @@ import { clearToolDirty, recordToolDirty } from 'dormouse-lib/lib/tool-dirty-store'; import { UNSAVED_TOOL_MOVE_REFUSAL } from 'dormouse-lib/lib/tool-editor'; -import { dismissWorkspaceUi, setWorkspaceMoveError } from 'dormouse-lib/lib/workspace-ui-store'; +import { dismissWorkspaceUi } from 'dormouse-lib/lib/workspace-ui-store'; import { restoreToolParams } from 'dormouse-lib/components/wall/tool-transfer'; import { recordToolAnnounce } from 'dormouse-lib/lib/tool-announce-store'; import { invoke } from "@tauri-apps/api/core"; @@ -511,7 +511,6 @@ async function adoptWorkspace(platform: PlatformAdapter, payload: MovePayload): // new move here can fail on a Tool that is still starting. console.error("[workspace-move] adopt_done refused; unwinding the mount", err); discardArrival(payload); - settle("adopt_failed", id, reasonOf(err)); } } catch (err) { console.error("[workspace-move] adoption failed; handing the Workspace back", err); @@ -557,10 +556,6 @@ export function initWorkspaceMoves(platform: PlatformAdapter): void { "dormouse://workspace-arrival-failed", (event) => handleArrivalFailed(event.payload.workspaceId, event.payload.reason ?? "no reason given", event.payload.replayIds ?? []), ); - void listenToWindow<{ workspaceId: WorkspaceId; reason: string }>( - "dormouse://workspace-arrival-retry", - (event) => setWorkspaceMoveError({ id: event.payload.workspaceId, reason: event.payload.reason }), - ); // Immediately, and not only on the nudge: a Workspace dropped on this window // while it was still booting is already in the queue, and its `emit_to` // reached no listener. @@ -601,7 +596,6 @@ export async function bootFromTearOut(platform: PlatformAdapter): Promise Date: Thu, 1 Oct 2026 22:32:29 -0700 Subject: [PATCH 04/10] Keep Burrow and recovery state in memory when their directory cannot be locked On Windows the Node stores' file modes are no-ops, so a refused DACL on the state directory used to leave burrowToken written under the inherited ACL with only a warning. prepare_owner_only_dir now publishes a directory only after restrict_to_owner succeeds; otherwise the sidecar gets no directory and keeps that state in memory. Co-Authored-By: Claude Opus 5.5 --- docs/specs/security-remote.md | 2 +- docs/specs/standalone.md | 4 +- lib/src/host/remote/burrow-state-store.ts | 3 +- standalone/src-tauri/src/lib.rs | 108 ++++++++++++++++------ 4 files changed, 83 insertions(+), 34 deletions(-) diff --git a/docs/specs/security-remote.md b/docs/specs/security-remote.md index abb14b771..e8db112b1 100644 --- a/docs/specs/security-remote.md +++ b/docs/specs/security-remote.md @@ -111,7 +111,7 @@ per-Burrow browser storage follows `docs/specs/remote-security-model.md` -> - **FAIL IF** `relay/src/state.ts` stops creating `$DORMOUSE_STATE_DIR` mode `0o700`, or stops writing every file through `writeAtomic` at mode `0o600`. The "every file" clause is a negative search over `relay/src/`: no `writeFile`, `appendFile`, or `createWriteStream` may target the state directory outside `writeAtomic`. A cheap default, not a cross-platform guarantee; the installer's directory permissions below protect the installed Relay's state (rationale). - **FAIL IF** `FileBurrowStateStore` (`lib/src/host/remote/burrow-state-store.ts`) stops creating its directory `0o700` and writing `0o600` on non-Windows platforms, or if `VsCodeBurrowStateStore` stops keeping the **enrollment** in `SecretStorage`. The ACL's home in `globalState` is deliberate and is not a finding; the enrollment's is what carries `burrowToken`. - **FAIL IF** a credential the Host→Burrow rename retired stops being deleted unread at boot: `state/hosts.json` on the Relay (`forgetRetiredState` in `relay/src/state.ts`, called from `relay/src/start.ts`), `remote-host.json` on a Node-resident Burrow (`forgetRetiredState` in `lib/src/host/remote/burrow-state-store.ts`, called from `sidecar-entry.ts`), and `dormouse.remote-host.enrollment`, `dormouse.remote-host.acl.*`, `remote-host.peer-token` in VS Code (`vscode-ext/src/retired-state.ts`, called from `activate()`). Each held a live `burrowToken` or peer secret, and `SecretStorage` cannot be enumerated — a key nothing removes *by name* outlives every build that knew it. Pinned by `relay/test/state-records.test.mjs`, `lib/src/host/remote/burrow-state-store.test.ts` and `vscode-ext/test/retired-state.test.ts`. -- **FAIL IF** `burrow_state_dir` in `standalone/src-tauri/src/lib.rs` stops calling `restrict_to_owner` on the state directory **before** spawning the sidecar — on Windows those Node modes are no-ops and Node cannot set an ACL, so the guarantee is held one layer down. That call carries both legs: a newly written enrollment file *inherits* the owner-only entry, and one a prior version already left under the `%LOCALAPPDATA%` ACL — with a live `burrowToken` in it — has that entry *propagated* onto it, the half `restrict_to_owner_leaves_one_owner_only_ace` covers with its pre-existing `before.json`. +- **FAIL IF** `burrow_state_dir` in `standalone/src-tauri/src/lib.rs` passes the sidecar a state directory `restrict_to_owner` did not lock — on Windows those Node modes are no-ops and Node cannot set an ACL, so the guarantee is held one layer down; a refusal keeps the Burrow in memory (`burrow_directory_permission_failure_disables_durable_state`). That call carries both legs: a newly written enrollment file *inherits* the owner-only entry, and one a prior version already left under the `%LOCALAPPDATA%` ACL — with a live `burrowToken` in it — has that entry *propagated* onto it, the half `restrict_to_owner_leaves_one_owner_only_ace` covers with its pre-existing `before.json`. - **FAIL IF** `relay/src/start.ts` stops obtaining the setup password from `SetupPasswordStore.loadOrCreate(generateSetupPassword)`, `generateSetupPassword` stops using `crypto.randomBytes(32)`, `readConfig` reads `DORMOUSE_SETUP_PASSWORD` or any other setup-password input, or `SetupPasswordStore` stops refusing a persisted or generated value outside 64 lowercase hexadecimal characters. Pinned by `relay/test/config.test.mjs` and `relay/test/setup-password-store.test.mjs`. - **FAIL IF** `createApp` accepts anything but 64 lowercase hexadecimal characters as the setup password injected by the entrypoint; pinned by `relay/test/app.test.mjs`. - **FAIL IF** any installer stops making `config/`, `state/`, and `config/relay.env` reachable only by the installing user — the effective property `manage verify` tests: no principal other than that user may appear in the effective permissions. macOS and Linux achieve it with `0700`/`0600` under `umask 077`; Windows with a single owner-only ACE, whether the path carries it directly or inherits it from an already-locked parent. The Windows and Linux installers create `relay.env` and lock it before writing its contents (rationale). diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index 709cc8056..a03cd6125 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -975,7 +975,7 @@ owner-only first and each an empty string when it could not be: `DORMOUSE_STATE_DIR` (the Burrow store, `app_data_dir`) and `DORMOUSE_RECOVERY_DIR` (the recovery record, the state root — so a dev run's record cannot reach the installed app). The browser-dev harness sets both to its -own per-run temp directory. Source of truth: `recovery_state_dir` in +own per-run temp directory. Source of truth: `prepare_owner_only_dir` / `recovery_state_dir` in `standalone/src-tauri/src/lib.rs`. **Must restrict the session store to the owner before any bytes are written** @@ -985,7 +985,7 @@ own per-run temp directory. Source of truth: `recovery_state_dir` in silent no-op, it applies a protected single-entry DACL instead (mechanism in its doc comment). `burrow_state_dir` locks the sidecar's state directory with the same call and relies on it reaching a file that already *existed*, which -`restrict_to_owner_leaves_one_owner_only_ace` pins (rationale). **Must abort a snapshot save if either permission change fails**, preserving the previous snapshot. The state-directory call remains nonfatal and logs a `WARNING` naming the path. Pinned by `session_permission_failures_preserve_previous_snapshot_without_writing_bytes` and `session_write_tightens_directory_and_existing_temp_file`. +`restrict_to_owner_leaves_one_owner_only_ace` pins (rationale). **Must abort a snapshot save if either permission change fails**, preserving the previous snapshot. **Must withhold a state directory whose restriction fails**, logging a `WARNING` naming the path; the Burrow store and recovery record then stay in memory (`burrow_directory_permission_failure_disables_durable_state`). Pinned by `session_permission_failures_preserve_previous_snapshot_without_writing_bytes` and `session_write_tightens_directory_and_existing_temp_file`. **Boot + the synchronous-read constraint.** `getState()` is synchronous — cold-start restore reads it before React mounts — but a Tauri `invoke` is async, so diff --git a/lib/src/host/remote/burrow-state-store.ts b/lib/src/host/remote/burrow-state-store.ts index 7b4438a43..41d5137f4 100644 --- a/lib/src/host/remote/burrow-state-store.ts +++ b/lib/src/host/remote/burrow-state-store.ts @@ -8,7 +8,8 @@ * The interface is async because the hosts that implement it are: files the * sidecar owns here, `VsCodeBurrowStateStore` there (enrollment in * `SecretStorage`, ACL in `globalState` — `docs/specs/vscode.md`). {@link FileBurrowStateStore} - * is the sidecar's: two files, 0600, under a directory the app passes in. + * is the sidecar's: private JSON state under a directory the app passes in + * only after establishing owner-only access (POSIX modes or a Windows DACL). */ import { readFile, rm } from 'node:fs/promises'; diff --git a/standalone/src-tauri/src/lib.rs b/standalone/src-tauri/src/lib.rs index 8fda5925c..6ccc9b803 100644 --- a/standalone/src-tauri/src/lib.rs +++ b/standalone/src-tauri/src/lib.rs @@ -3610,8 +3610,8 @@ fn resolve_dor_cli_paths(sidecar_path: &Path, manifest_dir: &Path) -> DorCliPath // Where the sidecar's Burrow persists its enrollment (a bearer credential) // and its ACL, as one 0600 file it writes itself // (lib/src/host/remote/burrow-state-store.ts). Created here so a first launch -// hands the sidecar a directory that exists; if it can't be made, the sidecar is -// told nothing and runs without persistence rather than not at all. +// hands the sidecar a directory that exists; if it can't be made or locked, the +// sidecar is told nothing and keeps that state in memory rather than not at all. fn burrow_state_dir(app: &AppHandle) -> Option { let dir = match app.path().app_data_dir() { Ok(dir) => dir, @@ -3620,10 +3620,6 @@ fn burrow_state_dir(app: &AppHandle) -> Option { return None; } }; - if let Err(e) = create_dir_all(&dir) { - append_log(format!("[sidecar] create state dir: {e}")); - return None; - } // The Node sidecar writes the Burrow enrollment here, and that record carries // `burrowToken` — a bearer credential for `/ws/burrow`. `FileBurrowStateStore` // asks for `0700`/`0600`, which Windows ignores entirely, so on Windows this @@ -3634,24 +3630,33 @@ fn burrow_state_dir(app: &AppHandle) -> Option { // `restrict_to_owner_leaves_one_owner_only_ace` covers with `before.json`. // On unix the store's own modes already do the job and this is a harmless // re-assert of the same intent. - if let Err(e) = restrict_to_owner(&dir, 0o700) { - // Not fatal — a Burrow that cannot start is worse than one whose state - // directory kept the OS default — but never silent: on Windows this - // call is the only thing restricting `burrowToken`, so its failure is a - // downgrade of the sole control and has to be visible. - append_log(format!( - "[sidecar] WARNING could not restrict state dir {}: {e}", - dir.display() - )); + match prepare_owner_only_dir(&dir, restrict_to_owner) { + Ok(path) => Some(path), + Err(e) => { + append_log(format!("[sidecar] WARNING {e}; Burrow state stays in memory")); + None + } } - Some(dir.to_string_lossy().into_owned()) +} + +/// Create `dir` and lock it owner-only, publishing its path only once both +/// succeeded. A refused restriction publishes nothing: on Windows the Node +/// stores' modes are no-ops, so an unlocked directory would hold their +/// credentials and commands under the inherited ACL. +fn prepare_owner_only_dir( + dir: &Path, + restrict: impl FnOnce(&Path, u32) -> Result<(), String>, +) -> Result { + create_dir_all(dir).map_err(|e| format!("create state dir: {e}"))?; + restrict(dir, 0o700).map_err(|e| format!("could not restrict state dir {}: {e}", dir.display()))?; + Ok(dir.to_string_lossy().into_owned()) } /// Where the sidecar writes the single-use agent-recovery record. Under the -/// state root, so a dev run never consumes the installed app's. Created here so -/// a first launch hands the sidecar a directory that exists; owner-only for the -/// same reason the Burrow's is — the record holds command lines the user typed, -/// and a unix mode is a silent no-op on Windows. +/// state root, so a dev run never consumes the installed app's. Owner-only for +/// the same reason the Burrow's is — the record holds command lines the user +/// typed, and a unix mode is a silent no-op on Windows — so a directory that +/// cannot be locked is withheld and the record stays in memory. fn recovery_state_dir(app: &AppHandle) -> Option { let dir = match state_root(app) { Ok(dir) => dir, @@ -3660,17 +3665,13 @@ fn recovery_state_dir(app: &AppHandle) -> Option { return None; } }; - if let Err(e) = create_dir_all(&dir) { - append_log(format!("[recovery] create state dir: {e}")); - return None; - } - if let Err(e) = restrict_to_owner(&dir, 0o700) { - append_log(format!( - "[recovery] WARNING could not restrict state dir {}: {e}", - dir.display() - )); + match prepare_owner_only_dir(&dir, restrict_to_owner) { + Ok(path) => Some(path), + Err(e) => { + append_log(format!("[recovery] WARNING {e}; recovery stays in memory")); + None + } } - Some(dir.to_string_lossy().into_owned()) } fn start_sidecar(app: &AppHandle) -> Result { @@ -4295,6 +4296,53 @@ mod tests { } } + #[test] + fn state_directory_creation_failure_never_attempts_permissions() { + let root = TempDir::new("state-dir-create-failure"); + let occupied = root.path().join("not-a-directory"); + fs::write(&occupied, b"previous").unwrap(); + let called = std::cell::Cell::new(false); + let result = super::prepare_owner_only_dir(&occupied, |_, _| { + called.set(true); + Ok(()) + }); + assert!(result.is_err()); + assert!(!called.get()); + assert_eq!(fs::read(occupied).unwrap(), b"previous"); + } + + #[test] + fn burrow_directory_permission_failure_disables_durable_state() { + let root = TempDir::new("burrow-state-restrict-failure"); + let target = root.path().join("state"); + fs::create_dir(&target).unwrap(); + fs::write(target.join("burrow.json"), b"existing enrollment").unwrap(); + let result = super::prepare_owner_only_dir(&target, |path, mode| { + assert_eq!(path, target); + assert_eq!(mode, 0o700); + Err("DACL refused".into()) + }); + assert!(result.unwrap_err().contains("DACL refused")); + assert_eq!(fs::read(target.join("burrow.json")).unwrap(), b"existing enrollment"); + assert_eq!(fs::read_dir(target).unwrap().count(), 1); + } + + #[test] + fn state_directory_is_created_and_restricted_before_publication() { + let root = TempDir::new("state-dir-order"); + let target = root.path().join("nested").join("state"); + let called = std::cell::Cell::new(false); + let result = super::prepare_owner_only_dir(&target, |path, mode| { + assert!(path.is_dir()); + assert_eq!(fs::read_dir(path).unwrap().count(), 0); + assert_eq!(mode, 0o700); + called.set(true); + Ok(()) + }); + assert!(called.get()); + assert_eq!(result.unwrap(), target.to_string_lossy()); + } + // --- Pending arrivals on disk (§Arrival queue) --------------------------- fn workspace_json(id: &str, name: &str) -> JsonValue { From 6156e87f098634ed3e1dcd02b8e97534644ed573 Mon Sep 17 00:00:00 2001 From: Ned Twigg Date: Thu, 1 Oct 2026 22:33:06 -0700 Subject: [PATCH 05/10] Skip the delayed launch check once an update is approved runUpdateCheck awaited the 5 s delay and the network-policy read, then called check() regardless, so an update approved through Check now in that window was offered for approval a second time. Co-Authored-By: Claude Opus 5.5 --- docs/specs/auto-update.md | 2 +- standalone/src/updater.test.ts | 47 ++++++++++++++++++++++++++++++++++ standalone/src/updater.ts | 5 +++- 3 files changed, 52 insertions(+), 2 deletions(-) diff --git a/docs/specs/auto-update.md b/docs/specs/auto-update.md index 3da2af03e..858c6544f 100644 --- a/docs/specs/auto-update.md +++ b/docs/specs/auto-update.md @@ -6,7 +6,7 @@ The standalone app checks for updates on launch, where the network policy allows ## How it works -**Must read and clear the post-install marker on launch** (§localStorage) and show its banner; a reported failure suppresses this launch's check. Otherwise wait 5 seconds, then read the network policy with `networkPolicy` over the Burrow link and, where it allows (`docs/specs/remote-network.md` → "Updates"), `check()` — no update is silent, an update raises the approval prompt; then the reminder, if due: `check-due`, recording `remindedAt`. **The reminder is re-evaluated hourly while the app runs**, reading no policy and never checking; **never over an undismissed notice, nor while the clock reads before 2026-09**, not yet set. Version-lookup and check failures are logged. **Only approval starts the background `download()`**; a failed one is logged and the prompt returns. +**Must read and clear the post-install marker on launch** (§localStorage) and show its banner; a reported failure suppresses this launch's check. Otherwise wait 5 seconds, then read the network policy with `networkPolicy` over the Burrow link and, where it allows (`docs/specs/remote-network.md` → "Updates"), `check()` — no update is silent, an update raises the approval prompt; then the reminder, if due: `check-due`, recording `remindedAt`. **Must skip that `check()` once an update is approved**, including through Check now during the wait or the policy read. **The reminder is re-evaluated hourly while the app runs**, reading no policy and never checking; **never over an undismissed notice, nor while the clock reads before 2026-09**, not yet set. Version-lookup and check failures are logged. **Only approval starts the background `download()`**; a failed one is logged and the prompt returns. **Check now** — the `check-due` and `check-failed` links, and the `updates` port — shows `checking`, then `available`, `up-to-date`, or `check-failed`. **A second ask joins the check in flight. An update already approved is shown again, `downloading` or `downloaded`, instead of checked for**, which would offer it for approval twice. **Every successful check, automatic or asked for, records `checkedAt`** (§localStorage). diff --git a/standalone/src/updater.test.ts b/standalone/src/updater.test.ts index 05571c864..1ee5a6e24 100644 --- a/standalone/src/updater.test.ts +++ b/standalone/src/updater.test.ts @@ -121,6 +121,53 @@ describe('updater', () => { mocks.platform = { requestAppRestart: mocks.requestAppRestart, burrow: { command: mocks.burrowCommand } }; }); + it('does not reoffer an approved download when the delayed launch check begins', async () => { + mocks.check.mockResolvedValue(makeUpdate('0.5.0')); + startUpdateCheck(); + await vi.advanceTimersByTimeAsync(0); + checkNow(); + await vi.advanceTimersByTimeAsync(0); + approveUpdate(); + await vi.advanceTimersByTimeAsync(0); + expect(hasPendingUpdate()).toBe(true); + + await vi.advanceTimersByTimeAsync(5_000); + expect(mocks.check).toHaveBeenCalledOnce(); + expect(readBannerState()).toEqual({ status: 'downloaded', version: '0.5.0' }); + }); + + it('does not reoffer an approval while its download is pending at the launch check', async () => { + const update = makeUpdate('0.5.0'); + update.download.mockImplementation(() => new Promise(() => {})); + mocks.check.mockResolvedValue(update); + startUpdateCheck(); + await vi.advanceTimersByTimeAsync(0); + checkNow(); + await vi.advanceTimersByTimeAsync(0); + approveUpdate(); + await vi.advanceTimersByTimeAsync(5_000); + + expect(mocks.check).toHaveBeenCalledOnce(); + expect(readBannerState()).toEqual({ status: 'downloading', version: '0.5.0' }); + }); + + it('keeps an approval made while the delayed launch policy read is pending', async () => { + let answerPolicy!: (value: typeof CHECKS_ON) => void; + mocks.burrowCommand.mockImplementation(() => new Promise(resolve => { answerPolicy = resolve; })); + mocks.check.mockResolvedValue(makeUpdate('0.5.0')); + startUpdateCheck(); + await vi.advanceTimersByTimeAsync(5_000); + checkNow(); + await vi.advanceTimersByTimeAsync(0); + approveUpdate(); + await vi.advanceTimersByTimeAsync(0); + answerPolicy(CHECKS_ON); + await vi.advanceTimersByTimeAsync(0); + + expect(mocks.check).toHaveBeenCalledOnce(); + expect(readBannerState()).toEqual({ status: 'downloaded', version: '0.5.0' }); + }); + // Drive check → approve → download so an approved, downloaded update is pending. async function reachDownloadedUpdate(update: ReturnType) { mocks.check.mockResolvedValue(update); diff --git a/standalone/src/updater.ts b/standalone/src/updater.ts index 4c6c91715..25043ab1a 100644 --- a/standalone/src/updater.ts +++ b/standalone/src/updater.ts @@ -414,7 +414,10 @@ async function runUpdateCheck(): Promise { // Read at the check, so a change made meanwhile counts // (`docs/specs/remote-network.md` → "Updates"). const policy = await readNetworkPolicy(); - if (policy && checksForUpdates(policy)) { + // A manual approval during the delay or the policy read owns this session's + // update; checking again would offer it for approval a second time. + const approved = pendingUpdate !== null || downloadPromise !== null; + if (policy && checksForUpdates(policy) && !approved) { // An update found is offered by `performCheck`. await performCheck().catch((e) => console.error('[updater] Check failed:', e)); } From 8bbee41bce88d2f06d71c01c585918cc01d4fc9d Mon Sep 17 00:00:00 2001 From: Ned Twigg Date: Thu, 1 Oct 2026 22:33:26 -0700 Subject: [PATCH 06/10] Fix the dev-script tests on Windows and the spec's browser timeout name Node's --import needs a file URL on Windows, and kill('SIGTERM') there terminates the child without running its JS handler, so the agent-browser harness test delivers the signal over IPC on win32. standalone.md named the retired AGENT_BROWSER_TIMEOUT (30s); the constant is BROWSER_REQUEST_TIMEOUT (40s). Co-Authored-By: Claude Opus 5.5 --- docs/specs/standalone.md | 2 +- standalone/scripts/dev-agent-browser.test.mjs | 15 +++++++++++---- standalone/scripts/dev-standalone.test.mjs | 4 ++-- 3 files changed, 14 insertions(+), 7 deletions(-) diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index a03cd6125..ab196c6ce 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -101,7 +101,7 @@ constants in `lib/src/lib/platform/types.ts` (and `standalone/sidecar/pty-core.j `#[tauri::command]` over an `async fn`, which the guard below accepts equally. Tauri runs a *sync* command on the main thread, where the `recv_timeout` inside `request_from_sidecar` / `request_from_sidecar_timeout` stops the webview painting -for the whole round trip, up to `AGENT_BROWSER_TIMEOUT` (30s) (rationale). **The +for the whole round trip, up to `BROWSER_REQUEST_TIMEOUT` (40s) (rationale). **The three clipboard readers included**: their non-Windows branches round-trip through the sidecar, and the declaration is per command, not per branch. A unit test in `lib.rs` scans the source and fails on any command that reaches the blocking diff --git a/standalone/scripts/dev-agent-browser.test.mjs b/standalone/scripts/dev-agent-browser.test.mjs index bc7916c17..ddac0e687 100644 --- a/standalone/scripts/dev-agent-browser.test.mjs +++ b/standalone/scripts/dev-agent-browser.test.mjs @@ -2,7 +2,7 @@ import test from 'node:test'; import assert from 'node:assert/strict'; import { access, copyFile, mkdir, readFile, rm, writeFile } from 'node:fs/promises'; import path from 'node:path'; -import { fileURLToPath } from 'node:url'; +import { fileURLToPath, pathToFileURL } from 'node:url'; import { spawn } from 'node:child_process'; import { get } from 'node:http'; import { setTimeout as delay } from 'node:timers/promises'; @@ -30,6 +30,10 @@ async function fixture(t) { }); `); const cli = path.join(bin, 'cli.cjs'); + // Windows kill('SIGTERM') bypasses JS handlers. Exercise the same shutdown + // handler over IPC there; POSIX continues exercising the actual signal. + const signals = path.join(bin, 'signals.mjs'); + await writeFile(signals, "process.on('message', signal => process.emit(signal));"); await writeFile(cli, ` if (process.argv[2] === 'list' && process.env.TEST_DOR_LIST) { console.log(process.env.TEST_DOR_LIST); @@ -54,8 +58,8 @@ async function fixture(t) { return { root, start(overrides = {}) { - const child = spawn(process.execPath, [path.join(standalone, 'scripts/dev-agent-browser.mjs')], { - cwd: root, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe'], + const child = spawn(process.execPath, ['--import', pathToFileURL(signals).href, path.join(standalone, 'scripts/dev-agent-browser.mjs')], { + cwd: root, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe', 'ipc'], }); // Object.assign, not a spread: `runner`'s `output`/`closed` are getters // over live state, and spreading would snapshot them once. @@ -73,7 +77,10 @@ async function fixture(t) { return this; }, async stop() { - if (child.exitCode === null && child.signalCode === null) child.kill('SIGTERM'); + if (child.exitCode === null && child.signalCode === null) { + if (process.platform === 'win32' && child.connected) child.send('SIGTERM', () => {}); + else child.kill('SIGTERM'); + } const timer = setTimeout(() => child.kill('SIGKILL'), 5000); try { return await this.exited; } finally { clearTimeout(timer); } }, diff --git a/standalone/scripts/dev-standalone.test.mjs b/standalone/scripts/dev-standalone.test.mjs index 4894deb2e..039a7c07a 100644 --- a/standalone/scripts/dev-standalone.test.mjs +++ b/standalone/scripts/dev-standalone.test.mjs @@ -2,7 +2,7 @@ import test from 'node:test'; import assert from 'node:assert/strict'; import { copyFile, mkdir, rm, writeFile } from 'node:fs/promises'; import path from 'node:path'; -import { fileURLToPath } from 'node:url'; +import { fileURLToPath, pathToFileURL } from 'node:url'; import { spawn } from 'node:child_process'; import { setTimeout as delay } from 'node:timers/promises'; import { cleanEnv, devWorkspace, runner, writeShims } from './dev-fixture.mjs'; @@ -54,7 +54,7 @@ async function fixture(t) { return { root, start(args = ['dev'], overrides = {}) { - const child = spawn(process.execPath, ['--import', signals, path.join(standalone, 'scripts/tauri.mjs'), ...args], { + const child = spawn(process.execPath, ['--import', pathToFileURL(signals).href, path.join(standalone, 'scripts/tauri.mjs'), ...args], { cwd: standalone, env: { ...cleanEnv(bin), ...overrides }, stdio: ['ignore', 'pipe', 'pipe', 'ipc'], }); // Object.assign, not a spread: `runner`'s `output`/`closed` are getters From 996e06593104da4020aceefade91411ff3185888 Mon Sep 17 00:00:00 2001 From: Ned Twigg Date: Thu, 1 Oct 2026 22:34:44 -0700 Subject: [PATCH 07/10] Refuse a closed window's saves and geometry writes for the whole process Destroyed used to drop the save refusal, but a save_session dispatched before Destroyed can take the disk lock after it and recreate the removed snapshot. Labels are never reused within a process, so the refusal now lasts until exit. The debounced geometry flush now rechecks that refusal under the same disk lock close removal holds, so a flush captured before the close cannot write the geometry file back. Co-Authored-By: Claude Opus 5.5 --- docs/specs/standalone.md | 11 +++++-- scripts/spec-word-budgets.json | 2 +- standalone/src-tauri/src/lib.rs | 52 ++++++++++++++++++++++++++------- 3 files changed, 52 insertions(+), 13 deletions(-) diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index ab196c6ce..3021a3c2e 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -574,8 +574,12 @@ checks in the debounce flush. - **The flush slot is released in the same step as the drain.** A `Moved` landing between the two was marked dirty with no thread left to write it — and that move is exactly a window's final position. +- **Must recheck the save refusal under the journal lock before writing + geometry.** The flush reads the window and rect first; a close in between + removed the file (§Per-window close), and the write would put it back + (`a_geometry_flush_captured_before_close_cannot_recreate_removed_geometry`). -Source of truth: `CachedRect` / `GeometryState` / `note_geometry` / +Source of truth: `CachedRect` / `GeometryState` / `note_geometry` / `write_open_window_geometry` / `restore_windows` in `standalone/src-tauri/src/lib.rs`; the sequencing is pinned by `the_geometry_flush_slot_is_released_with_the_drain`. @@ -622,7 +626,10 @@ anyway if that listener is dead), asks about *its own* running work, removes its for that label, so a PTY exit's save cannot write it back. **Both close paths set that refusal** — the webview's own `remove_window_session`, and `finish_window_close` for the ack-timeout path, where the webview never ran at - all. It is dropped when the webview is destroyed and can no longer save. + all. **Must keep that refusal for the process lifetime**, geometry writes + included: a save dispatched before `Destroyed` can reach the disk lock after + it, and no label is reused within a process + (`a_closed_window_refuses_saves_for_the_process_lifetime`). - **`close_window` is the one Rust half both endings share** — a deliberate close and a window whose last Workspace moved away (§Transfer) — because what separates them is entirely what the webview did before calling it. diff --git a/scripts/spec-word-budgets.json b/scripts/spec-word-budgets.json index 19f48d5fe..324a726dc 100644 --- a/scripts/spec-word-budgets.json +++ b/scripts/spec-word-budgets.json @@ -30,7 +30,7 @@ "docs/specs/security-supply-chain.md": 1250, "docs/specs/security.md": 2150, "docs/specs/shortcuts.md": 1100, - "docs/specs/standalone.md": 11950, + "docs/specs/standalone.md": 12000, "docs/specs/terminal-context.md": 1100, "docs/specs/terminal-escapes.md": 4050, "docs/specs/terminal-state.md": 2400, diff --git a/standalone/src-tauri/src/lib.rs b/standalone/src-tauri/src/lib.rs index 6ccc9b803..b8cefba0c 100644 --- a/standalone/src-tauri/src/lib.rs +++ b/standalone/src-tauri/src/lib.rs @@ -139,8 +139,9 @@ struct WindowState { /// one can be told to clear it. hover_target: Mutex>, /// Labels whose snapshot has been deliberately removed. A save arriving - /// from a webview that is going away must not put the file back; the entry - /// is dropped once that webview is destroyed and can no longer save. + /// from a webview that is going away must not put the file back. Kept for + /// the process lifetime: a save dispatched before `Destroyed` can still + /// reach the disk lock after it, and labels are never reused in a process. closing: Mutex>, /// The next `ws-`, seeded above every live and saved label at setup. next_ws: AtomicU64, @@ -181,8 +182,8 @@ impl WindowState { .store(routing.awaiting_replay.len(), Ordering::Relaxed); } - /// Refuse every later `save_session` for `label` (a deliberate close removed - /// its snapshot). Cleared by `Destroyed`, after which no save can arrive. + /// Refuse every later `save_session` and geometry write for `label` (a + /// deliberate close removed its snapshot), for the rest of the process. fn begin_closing(&self, label: &str) { guard(&self.closing).insert(label.to_string()); } @@ -2165,6 +2166,16 @@ fn geometry_path(dir: &Path, label: &str) -> PathBuf { dir.join(format!("{stem}.geometry.json")) } +/// The window and rect were read before this lock, so a close may have removed +/// the geometry since: recheck the save refusal under the lock that removal +/// holds, or a debounce captured before the close writes the file back. +fn write_open_window_geometry(windows: Option<&WindowState>, dir: &Path, label: &str, json: &str) -> Result { + let _disk = guard(&ARRIVAL_DISK_LOCK); + if windows.is_some_and(|windows| windows.refuses_save(label)) { return Ok(false); } + write_file_atomically(&geometry_path(dir, label), json)?; + Ok(true) +} + fn read_geometry(dir: &Path, label: &str) -> Option { let raw = std::fs::read_to_string(geometry_path(dir, label)).ok()?; serde_json::from_str(&raw).ok() @@ -2245,7 +2256,8 @@ fn note_geometry(app: &AppHandle, label: &str, origin: Option<(i32, i32)>, size: let Ok(json) = serde_json::to_string(&rect.to_logical()) else { continue; }; - if let Err(err) = write_file_atomically(&geometry_path(&dir, &label), &json) { + let windows = app.try_state::(); + if let Err(err) = write_open_window_geometry(windows.as_deref(), &dir, &label, &json) { append_log(format!("[window] geometry write for {label}: {err}")); } } @@ -3997,7 +4009,6 @@ pub fn run() { // Drop label-keyed ownership synchronously; only the // returned arrivals need the blocking journal worker. let (lost, orphaned) = state.drop_window(&label); - guard(&state.closing).remove(&label); reap_orphaned_ptys(app, &label, orphaned); let changed = workspaces::forget_window(&mut guard(&state.registry), &label); if changed { broadcast_registry(app, &state); } @@ -5448,16 +5459,37 @@ mod tests { /// paths set it: the webview's own `remove_window_session`, and /// `finish_window_close` for the ack-timeout path where it never ran. #[test] - fn a_closing_window_refuses_every_later_save_until_it_is_destroyed() { + fn a_closed_window_refuses_saves_for_the_process_lifetime() { let state = super::WindowState::default(); assert!(!state.refuses_save("ws-2")); state.begin_closing("ws-2"); assert!(state.refuses_save("ws-2")); // Never a sibling's. assert!(!state.refuses_save("main")); - // `Destroyed` drops the refusal: no save can arrive under a dead label. - guard(&state.closing).remove("ws-2"); - assert!(!state.refuses_save("ws-2")); + // A save dispatched before `Destroyed` can reach the disk lock after + // it, so neither the label sweep nor the arm itself drops the refusal. + state.drop_window("ws-2"); + assert!(state.refuses_save("ws-2")); + let src = include_str!("lib.rs").split("#[cfg(test)]").next().unwrap(); + let destroyed = src.split("WindowEvent::Destroyed => {").nth(1).unwrap() + .split("if let Some(state) = app.try_state::()").next().unwrap(); + assert!(!destroyed.contains(".closing"), "Destroyed must not clear save refusal"); + } + + #[test] + fn a_geometry_flush_captured_before_close_cannot_recreate_removed_geometry() { + let dir = TempDir::new("geometry-close-fence"); + let windows = super::WindowState::default(); + write_session_to(dir.path(), "ws-2", "snapshot").unwrap(); + assert!(super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", "previous").unwrap()); + assert_eq!(fs::read_to_string(super::geometry_path(dir.path(), "ws-2")).unwrap(), "previous"); + windows.begin_closing("ws-2"); + super::close_window_snapshot(dir.path(), "ws-2").unwrap(); + assert!(!super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", "captured before close").unwrap()); + assert!(!super::geometry_path(dir.path(), "ws-2").exists()); + windows.drop_window("ws-2"); + assert!(!super::write_open_window_geometry(Some(&windows), dir.path(), "ws-2", "captured before close").unwrap()); + assert!(!super::geometry_path(dir.path(), "ws-2").exists()); } fn queue_test_suppression(state: &super::WindowState, id: &str) { From def6e872174fde3d9644ce0ef33a40b91930147b Mon Sep 17 00:00:00 2001 From: Ned Twigg Date: Thu, 1 Oct 2026 22:38:32 -0700 Subject: [PATCH 08/10] Journal a Workspace transfer before its ownership moves begin_arrival moved the shells to the target and only then wrote sessions/arrivals.json, logging a failed write and carrying on, so a crash after that failure restored the in-flight Workspace in neither window. journal_and_open_arrival now writes the record first and refuses the move with nothing changed if the write fails. The write runs outside the arrivals lock, whose waits reach the main thread, so admission is checked again under that lock and a refusal there withdraws the record. Co-Authored-By: Claude Opus 5.5 --- docs/specs/standalone.md | 13 ++- scripts/spec-word-budgets.json | 2 +- standalone/src-tauri/src/lib.rs | 167 ++++++++++++++++++++++---------- 3 files changed, 126 insertions(+), 56 deletions(-) diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index 3021a3c2e..40061612a 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -707,7 +707,7 @@ below reads that record rather than inferring itself from the suppression map. `transfer_workspace` / `open_workspace_window`. **Must return preparation refusals as `{ moved: false, reason }` without changing ownership.** On `Ok` it marks the Workspace **transferring**: the Wall stays mounted, nothing is released, and `getWindowSnapshot` omits it. -2. **Rust** reassigns `terminalIds` to the target, keeps routing their output to +2. **Rust** journals the arrival (below), then reassigns `terminalIds` to the target, keeps routing their output to the source, and asks the sidecar to stamp a `pty:marked` line per id; at that line the id's suppression begins, until its replay has been emitted to the target. The source serializes each buffer at its mark and invokes @@ -803,9 +803,12 @@ below reads that record rather than inferring itself from the suppression map. - **A boot's `pty_request_init` excludes every id an arrival claims.** Ownership moves at the invoke, so those shells would otherwise be listed as top-level panes beside the Workspace about to mount them. -- **`begin_arrival` records the arrival in `sessions/arrivals.json`** — a JSON - array of `{ workspaceId, from, to, workspace, settled }`, never an entry in - either window's snapshot (rationale); the tombstone rules below read `settled`. **Must retain an adopted record until target +- **`begin_arrival` records the arrival in `sessions/arrivals.json` before + ownership moves** — a JSON array of `{ workspaceId, from, to, workspace, + settled }`, never an entry in either window's snapshot (rationale). **A failed + write must refuse the move with nothing changed**; the write runs outside + `arrivals`, so admission is rechecked after it and a refusal withdraws the + record (`a_failed_arrival_journal_refuses_the_move_with_nothing_changed`); the tombstone rules below read `settled`. **Must retain an adopted record until target and source snapshots both reflect the move**, marking it settled at `adopt_done` and checking after each `save_session` or source-window close (`adoption_keeps_the_journal_until_both_snapshots_are_durable`). **Must reverse @@ -1147,7 +1150,7 @@ asks before discarding a pending download (§Per-window close). `deferred_quit_and_close_requests_wait_for_membership_then_run_once` in `standalone/src-tauri/src/quit_state.rs`, and `transfers_cannot_change_membership_after_close_or_quit_confirmation_begins` and - `begin_arrival_admits_under_the_arrivals_lock_before_queueing` in + `begin_arrival_journals_then_admits_under_the_arrivals_lock_before_queueing` in `standalone/src-tauri/src/lib.rs`. - **Must collect votes before killing any window's Sessions.** Confirmation consumes its callback once; a noninteractive full-window progress overlay diff --git a/scripts/spec-word-budgets.json b/scripts/spec-word-budgets.json index 324a726dc..d00c4cefa 100644 --- a/scripts/spec-word-budgets.json +++ b/scripts/spec-word-budgets.json @@ -30,7 +30,7 @@ "docs/specs/security-supply-chain.md": 1250, "docs/specs/security.md": 2150, "docs/specs/shortcuts.md": 1100, - "docs/specs/standalone.md": 12000, + "docs/specs/standalone.md": 12050, "docs/specs/terminal-context.md": 1100, "docs/specs/terminal-escapes.md": 4050, "docs/specs/terminal-state.md": 2400, diff --git a/standalone/src-tauri/src/lib.rs b/standalone/src-tauri/src/lib.rs index b8cefba0c..2cdc4d1e2 100644 --- a/standalone/src-tauri/src/lib.rs +++ b/standalone/src-tauri/src/lib.rs @@ -2504,20 +2504,35 @@ fn record_workspace_id(record: &JsonValue) -> Option<&str> { record.get("workspaceId").and_then(JsonValue::as_str) } -/// Append one arrival's record, replacing any earlier record of the same id. -fn record_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result<(), String> { +/// Append one arrival's record, replacing and returning any earlier record of +/// the same id. +fn record_arrival_on_disk(dir: &Path, arrival: &routing::Arrival) -> Result, String> { let _disk = guard(&ARRIVAL_DISK_LOCK); let Some(workspace) = arrival.payload.get("workspace") else { return Err("arrival payload carries no workspace".to_string()); }; let mut records = read_arrivals_from(dir)?; - records.retain(|record| record_workspace_id(record) != Some(&arrival.workspace_id)); + let previous = records.iter().position(|record| record_workspace_id(record) == Some(&arrival.workspace_id)) + .map(|at| records.remove(at)); records.push(serde_json::json!({ "workspaceId": arrival.workspace_id, "from": arrival.from, "to": arrival.to, "workspace": workspace, })); + write_arrivals_to(dir, &records)?; + Ok(previous) +} + +/// Undo `record_arrival_on_disk` for an arrival refused after its write, +/// unless a later drop of the same id has since replaced the record. +fn withdraw_arrival_on_disk(dir: &Path, arrival: &routing::Arrival, previous: Option) -> Result<(), String> { + let _disk = guard(&ARRIVAL_DISK_LOCK); + let mut records = read_arrivals_from(dir)?; + let ours = |r: &JsonValue| record_workspace_id(r) == Some(&arrival.workspace_id) + && r["from"] == arrival.from.as_str() && r["to"] == arrival.to.as_str() && r.get("settled").is_none(); + let Some(at) = records.iter().position(ours) else { return Ok(()); }; + match previous { Some(record) => records[at] = record, None => { records.remove(at); } } write_arrivals_to(dir, &records) } @@ -2688,8 +2703,65 @@ fn transfer_admitted( && [from, to].into_iter().all(|label| !close.active(label) && !closing.contains(label)) } -/// Open one arrival: reassign its shells to the target and suppress them, then -/// queue the record. **Ownership moves synchronously here**, before either +/// Whether `arrival` may begin: no quit or close of either end (see +/// `transfer_admitted`) and its Workspace not already in flight. +fn admit_arrival( + quit: Option<&QuitState>, + windows: &WindowState, + arrivals: &ArrivalQueue, + arrival: &routing::Arrival, +) -> Result<(), String> { + if let Some(state) = quit { + // Lock order: arrivals, quit machine, close machine, closing. + let admitted = transfer_admitted( + arrivals, + &guard(&state.machine), + &guard(&state.close), + &guard(&windows.closing), + &arrival.from, + &arrival.to, + ); + if !admitted { + return Err("cannot transfer a Workspace while its window is closing or Dormouse is quitting".to_string()); + } + } + if routing::has_arrival(arrivals, &arrival.workspace_id) { + return Err(format!("Workspace '{}' is already in flight", arrival.workspace_id)); + } + Ok(()) +} + +/// Journal one arrival, then reassign its shells to the target and queue it. +/// **The journal comes first**: the source omits a transferring Workspace +/// from its saves and the target writes only after adoption, so a crash in +/// the gap is recovered from this record alone (§Arrival queue). A failed +/// write refuses the move with nothing changed. The write runs outside +/// `arrivals`, whose waits reach the main thread, so admission is checked +/// again under it and a refusal then withdraws the record. +fn journal_and_open_arrival( + quit: Option<&QuitState>, + windows: &WindowState, + dir: &Path, + arrival: &routing::Arrival, +) -> Result<(), String> { + admit_arrival(quit, windows, &guard(&windows.arrivals), arrival)?; + let previous = record_arrival_on_disk(dir, arrival) + .map_err(|e| format!("could not record the Workspace transfer: {e}"))?; + let mut arrivals = guard(&windows.arrivals); + if let Err(refused) = admit_arrival(quit, windows, &arrivals, arrival) { + drop(arrivals); + if let Err(e) = withdraw_arrival_on_disk(dir, arrival, previous) { + append_log(format!("[window] could not withdraw {}'s record: {e}", arrival.workspace_id)); + } + return Err(refused); + } + // Ownership moves now; suppression waits for each id's `marked` line. + windows.begin_transfer(&arrival.terminal_ids, &arrival.from, &arrival.to); + routing::queue_arrival(&mut arrivals, arrival.clone()); + Ok(()) +} + +/// Open one arrival. **Ownership moves synchronously here**, before either /// window is told anything — the single Rust reader thread processes sidecar /// lines in order, so every byte after this point is either dropped (and present /// in the replay the target is about to get) or delivered to the target @@ -2699,32 +2771,8 @@ fn begin_arrival( windows: &WindowState, arrival: routing::Arrival, ) -> Result<(), String> { - { - let mut arrivals = guard(&windows.arrivals); - if let Some(state) = app.try_state::() { - // Lock order: arrivals, quit machine, close machine, closing. - let admitted = transfer_admitted( - &arrivals, - &guard(&state.machine), - &guard(&state.close), - &guard(&windows.closing), - &arrival.from, - &arrival.to, - ); - if !admitted { - return Err("cannot transfer a Workspace while its window is closing or Dormouse is quitting".to_string()); - } - } - if routing::has_arrival(&arrivals, &arrival.workspace_id) { - return Err(format!( - "Workspace '{}' is already in flight", - arrival.workspace_id - )); - } - // Ownership moves now; suppression waits for each id's `marked` line. - windows.begin_transfer(&arrival.terminal_ids, &arrival.from, &arrival.to); - routing::queue_arrival(&mut arrivals, arrival.clone()); - } + let quit = app.try_state::(); + journal_and_open_arrival(quit.as_deref(), windows, &sessions_dir(app)?, &arrival)?; // The split point, stamped in the stream by the sidecar and routed to the // source, which serializes what it holds when it sees it (§Transfer). A // Workspace of browser panes alone has no ids to mark; its source sends @@ -2736,19 +2784,6 @@ fn begin_arrival( }); send_to_sidecar(&sidecar, msg.to_string()); } - // Recorded on disk here, in neither window's snapshot: the source omits a - // transferring Workspace from its saves and the target writes only after - // adoption, so a crash in the gap would otherwise restore it nowhere - // (§Arrival queue). Never fatal: a failed write is logged and the transfer - // proceeds. - match sessions_dir(app) { - Ok(dir) => { - if let Err(e) = record_arrival_on_disk(&dir, &arrival) { - append_log(format!("[window] could not record {} on disk: {e}", arrival.workspace_id)); - } - } - Err(e) => append_log(format!("[window] {e}")), - } spawn_arrival_watchdog(app.clone(), &arrival); Ok(()) } @@ -4726,15 +4761,47 @@ mod tests { } #[test] - fn begin_arrival_admits_under_the_arrivals_lock_before_queueing() { + fn begin_arrival_journals_then_admits_under_the_arrivals_lock_before_queueing() { let src = include_str!("lib.rs").split("#[cfg(test)]").next().unwrap(); - let body = src.split("fn begin_arrival(").nth(1).unwrap().split("\n}").next().unwrap(); + let body = src.split("fn journal_and_open_arrival(").nth(1).unwrap().split("\n}").next().unwrap(); + let journal = body.find("record_arrival_on_disk(").unwrap(); let lock = body.find("let mut arrivals = guard(&windows.arrivals)").unwrap(); - let check = body.find("transfer_admitted(").unwrap(); - let refuse = body.find("if !admitted {").unwrap(); + let check = lock + body[lock..].find("admit_arrival(").unwrap(); + let refuse = body.find("return Err(refused)").unwrap(); + let transfer = body.find("windows.begin_transfer(").unwrap(); let queue = body.find("routing::queue_arrival").unwrap(); - assert!(lock < check && check < refuse && refuse < queue); - assert!(body[refuse..queue].contains("return Err(")); + assert!(journal < lock && lock < check && check < refuse && refuse < transfer && transfer < queue); + } + + #[test] + fn a_failed_arrival_journal_refuses_the_move_with_nothing_changed() { + let dir = TempDir::new("arrival-journal-first"); + let windows = super::WindowState::default(); + let mut arrival = arrival_of("workspace-7", "main", "ws-2"); + arrival.terminal_ids = vec!["pane-a".to_string()]; + windows.mint("pane-a", "main"); + // A directory where the journal belongs fails its read on every platform. + fs::create_dir(arrivals_path(dir.path())).unwrap(); + assert!(super::journal_and_open_arrival(None, &windows, dir.path(), &arrival).is_err()); + assert_eq!(windows.owned_by("main"), vec!["pane-a"]); + assert!(guard(&windows.arrivals).is_empty()); + assert!(guard(&windows.routing).marking.is_empty()); + fs::remove_dir(arrivals_path(dir.path())).unwrap(); + super::journal_and_open_arrival(None, &windows, dir.path(), &arrival).unwrap(); + assert_eq!(windows.owned_by("ws-2"), vec!["pane-a"]); + assert_eq!(guard(&windows.arrivals).len(), 1); + // A refusal after the write puts back the record it replaced, and + // never withdraws a newer drop's record. + let mut later = arrival_of("workspace-7", "ws-2", "ws-3"); + let previous = record_arrival_on_disk(dir.path(), &later).unwrap(); + super::withdraw_arrival_on_disk(dir.path(), &arrival, None).unwrap(); + assert_eq!(read_arrivals_from(dir.path()).unwrap()[0]["to"], "ws-3"); + super::withdraw_arrival_on_disk(dir.path(), &later, previous).unwrap(); + assert_eq!(read_arrivals_from(dir.path()).unwrap()[0]["to"], "ws-2"); + later.to = "ws-4".to_string(); + record_arrival_on_disk(dir.path(), &later).unwrap(); + super::withdraw_arrival_on_disk(dir.path(), &later, None).unwrap(); + assert!(read_arrivals_from(dir.path()).unwrap().is_empty()); } #[test] From 24fc850e8210ada1fea555af3686762aebb39c02 Mon Sep 17 00:00:00 2001 From: Ned Date: Thu, 1 Oct 2026 23:37:09 -0700 Subject: [PATCH 09/10] Describe logged native persistence failures without overstating durability --- docs/specs/standalone.md | 18 ++++++++---------- 1 file changed, 8 insertions(+), 10 deletions(-) diff --git a/docs/specs/standalone.md b/docs/specs/standalone.md index 40061612a..73d011e82 100644 --- a/docs/specs/standalone.md +++ b/docs/specs/standalone.md @@ -608,12 +608,12 @@ Source of truth: `CleanupGate` and `WindowEvent::Destroyed` in **Closing a window with siblings alive ends that window alone**; only the last window's close is the quit. Rust prevents the close and emits `dormouse://window-close-requested`; the webview acks (a ~2 s watchdog closes it -anyway if that listener is dead), asks about *its own* running work, removes its snapshot, kills the PTYs it owns, and calls back +anyway if that listener is dead), asks about *its own* running work, attempts snapshot removal, kills its PTYs, and calls back `close_window`. -- **A close is deliberate, so it removes the blob** — geometry - and temp sibling included — and the next launch does not reopen the window - (`docs/specs/transport.md` → "The governing rule"). +- **Must attempt to remove the blob before killing this Window's PTYs** — geometry + and temp sibling included; successful removal prevents reopening + (`docs/specs/transport.md` → "The governing rule"). Removal failures are logged and close proceeds; an old snapshot may reopen. - **It runs no agent-recovery capture**: nothing is coming back. - **A cancelled close retires its watchdog's token and never reuses it**: the next close on that window is a fresh seq, so a watchdog still sleeping on the @@ -622,8 +622,7 @@ anyway if that listener is dead), asks about *its own* running work, removes its - **It confirms on a pending download as well as on running work.** An approved, downloaded update lives in this webview's memory, so closing the window throws it away and nothing else can install it (`docs/specs/auto-update.md`). -- **The snapshot is removed before the kill**, and Rust refuses every later save - for that label, so a PTY exit's save cannot write it back. **Both close paths +- **Must refuse every later save for a closing label**, so a PTY exit cannot recreate its snapshot. **Both close paths set that refusal** — the webview's own `remove_window_session`, and `finish_window_close` for the ack-timeout path, where the webview never ran at all. **Must keep that refusal for the process lifetime**, geometry writes @@ -813,7 +812,7 @@ below reads that record rather than inferring itself from the suppression map. `adopt_done` and checking after each `save_session` or source-window close (`adoption_keeps_the_journal_until_both_snapshots_are_durable`). **Must reverse the durable destination on hand-back and retain the record until both - snapshots reflect the return** (`a_hand_back_is_recovered_in_the_source_before_its_next_flush`). + snapshots reflect the return** (`a_hand_back_is_recovered_in_the_source_before_its_next_flush`). Failed settlement-marker or hand-back writes are logged and do not block live adoption or return. **Must tombstone settled arrivals into a deliberately closed Window until both snapshots omit them**, including during boot recovery (`closing_an_adopted_target_never_resurrects_either_copy`). @@ -953,7 +952,7 @@ written. - **The label is sanitized** so it cannot escape the directory. - **Temp-then-rename**, so a crash cannot truncate the previous snapshot. The temp file is fsynced before the rename and, on unix only, the sessions directory - *after* it (rationale). + *after* it, best-effort (rationale). - **Window identity is implicit**: each command keys by the invoking `tauri::Window`'s `label()`, so the frontend stays window-agnostic and every window (`ws-2`, …) persists to its own file rather than rewriting a sibling's. @@ -961,8 +960,7 @@ written. blob (rationale). - **The writer removes its own temp file on every error path**, so only a crash can leave one behind. -- **A per-window close removes the blob, its temp sibling and its geometry** - (§Per-window close); nothing else deletes a snapshot but the boot merge +- Per-window cleanup follows §Per-window close; only the boot merge otherwise deletes a snapshot (§Arrival queue). - **Must sweep orphan session temp files once at boot** in the active sessions directory and, for debug builds, the legacy `/sessions` directory. From 17718def6d839c2852a391cbea690c43096487d1 Mon Sep 17 00:00:00 2001 From: Ned Date: Thu, 1 Oct 2026 23:39:10 -0700 Subject: [PATCH 10/10] Qualify best-effort directory fsync evidence --- docs/specs/standalone.rationale.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/specs/standalone.rationale.md b/docs/specs/standalone.rationale.md index a22e726dd..3132c8ba0 100644 --- a/docs/specs/standalone.rationale.md +++ b/docs/specs/standalone.rationale.md @@ -222,7 +222,7 @@ stays the webview's throughout and no polling loop is needed. **The WKWebView WAL measurement.** WKWebView stores `localStorage` as SQLite in WAL mode, and WebKit pins that WAL with a long-lived reader that never advances during a running session — so it is never checkpointed, and an external checkpoint is blocked by the same reader. Rewriting the multi-MB scrollback-bearing session blob on every save grew the WAL to ~1 GB within a few hours (recorded 2026-07); a days-long session made it pathological. The Rust file store that replaced it has no WAL and rewrites the same file each time. -**Why the sessions directory is fsynced after the rename.** Fsyncing only the temp file leaves the new name recoverable-but-absent after a power loss; the directory-entry fsync is what makes the rename itself durable. Windows has no equivalent concept, hence unix-only. +**Why the sessions directory is fsynced after the rename.** Fsyncing only the temp file leaves the new name recoverable-but-absent after a power loss; a successful directory-entry fsync makes the rename durable. Its failure is ignored, so this step is best-effort. Windows has no equivalent concept, hence unix-only. **Why the mode is set before the bytes.** Under the bare umask the transcript-bearing blob lands `0644` in a `0755` directory any other local account can read, and tightening after the write would leave a window in which it was readable. Continuing after a permission failure would contradict the owner-only guarantee; aborting before writing preserves the previous snapshot and leaves at most an empty temp file.