Skip to content

Make the board responsive: connection lanes, deduped reads, audio and Desk fixes - #10

Merged
cailmdaley merged 15 commits into
mainfrom
perf/ws-perf
Oct 5, 2026
Merged

cailmdaley merged 15 commits into
mainfrom
perf/ws-perf

Conversation

@cailmdaley

Copy link
Copy Markdown
Owner

Shuttle had become close to unusable on the desktop. Opening a channel took 10–20 s, mp3s took 9–17 s to start, Desk cards jumped under the pointer, and the board fetched the same documents over and over. This PR fixes the causes found in round 3's bug and performance lanes. The reader's layout changes come separately.

Claude Opus 5.5 on behalf of Cail.

The main cause: six connections

On 127.0.0.1:4000 the board speaks HTTP/1.1, so Chrome opens at most six connections to the daemon. Background reads held all six for seconds at a time:

  • the overview resolving fibers;
  • the 8.7 MB wikilink index;
  • slow per-host agents reads;
  • one file-info per document every 15 s;
  • 36 preloading <audio> elements on Music.

The page you'd just opened queued behind them. The phone goes through tailscale serve (HTTP/2), which is why it felt better there.

  • Background daemon reads take two lanes, so an opened page always finds a connection (9f7cb3ba).
  • A channel opens on its documents before owner file times are read (223d4e83).
    • Its body refresh runs in the slow lane (1bedbfc3).
    • Title peeks skip pages the stage is already reading (b709c751).
    • The probe and thumbnails share one 64 KiB peek per document (9a074428).
  • The daemon names large documents by digest, so an unchanged poll is a 304 (e60f79fa). A live document's first read revalidates the browser's copy (18b452b7).
  • Audio:
    • Background audio reads share two slots, and siblings' durations come from one short metadata read each (a8ef1710). Play reaches sound in 0.1–1.5 s, down from 9–17 s in WebKit.
    • Waveforms decode in an OfflineAudioContext at 8 kHz, never a realtime AudioContext. Before, every decode opened a coreaudiod session that leaked when a tab or browser died; Cail's coreaudiod had grown to 13 GB. A test asserts zero realtime contexts across 13 recordings (1c61e51b).
    • Waveforms decode only for the selected page (f1d82fda).

Bugs

  • Clicking a Desk card's centre opened another card. Every poll render smooth-scrolled the last-opened card back into view under the pointer, and on the phone it also paged the Desk back to that column. A redraw now repaints the selection without scrolling (ded8fd35).
  • The Board re-read missing fibers on every 15 s poll, from every tab. That produced thousands of no fiber found daemon calls, which fed the hub's fd exhaustion. A missing fiber now backs off across polls (9970ec38).

Measured

The numbers come from ui/scripts/measure-workspace.mjs (4429c3b2): headless Chrome against the live daemon over HTTP/1.1. Each cell is the median of 3 runs; "phone" is a 390×844 viewport.

before after
open a channel, warm (desktop) 15.2 s 1.0 s
open a channel, warm (phone viewport) 9.4 s 0.15 s
worst page step, cold (desktop) 15.3 s 20 ms
longest a request waited before sending 10.0 s 0.11 s
KB fetched to open, warm 2535 32
Music open, cold 1.3 s 0.24 s

Caveats:

  • The live data was noisy: the machine was under heavy load, and the daemon restarted during the run.
  • Switching to a channel whose fiber lives on a slow remote host (cineca, nibi) still takes up to 25 s. That host's response time, not client queueing, is the limit.

Gates

  • tsc --noEmit passes.
  • vitest: 1556/1556 under America/Los_Angeles and Europe/Paris.
  • e2e: 93/93 on a low-load rerun.
  • mix test: 1680, 0 failures.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3

cailmdaley and others added 15 commits October 5, 2026 20:23
Every mounted listening page (the selected one and both receded
neighbours) gave each sibling in its Compare list its own preloading
<audio>, and decoded its waveform from a whole-file read at the same
time. On Music that is 36 metadata players and three 6 MB downloads at
once. Over HTTP/1.1 (the desktop board at 127.0.0.1:4000) a browser
holds six connections per host, so the selected page's own media
request queued behind them: Safari took 9-17 s from Play to sound on a
local daemon. The phone reaches the board through the tailnet's HTTP/2
front, which multiplexes, so it never queued.

Background listening reads now share two slots, newest request first,
and each sibling's duration comes from one transient metadata read
cached by path and provenance. Play to sound measured 0.08-1.5 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every Desk render ended with the keyboard selection's scroll into view,
and the selection is the card last opened. So each poll that changed
the feed scrolled that card back into view, smoothly. On the phone it
also paged the Desk back to that card's column. A person who had
scrolled or paged on to another card and pressed it during that scroll
opened whatever card had slid under the pointer. A redraw now only
repaints the selection. Moving the selection, and returning from the
reader, still reveal it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
While the Board was the active view, each 15 s poll refreshed the
overview, and every refresh set every metadata retry back to zero. A
receipt whose owner had already answered "not found" was read again on
every poll, by every open tab, forever. Each of those reads made the
daemon ask each host in turn: 1.2 s of daemon work per read. Measured
on the live board before the fix: 7 such fibers, each read 7 times in
95 s by one tab. The daemon's log held 13k "no fiber found" lines for
the day.

A poll now retries only transient failures at once. A confirmed miss
waits 30 s, doubling to 5 min, and a fresh visit to the Board retries
it at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A playwright-core script that drives a running board read-only: Desk card to
first settled page, page steps, channel switches, the Board overview's first
thumbnails, the phone reader, every /api request per phase (count, bytes,
duplicate reads of one document, time queued before send) and long tasks,
cold and warm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
…onnection

Over HTTP/1.1 the browser opens six connections to the daemon. At load the
overview's fiber preloads (four at a time, 2-20 s each when relayed to a
remote host), the 8 MB wikilink index, one agent registry per host and the
body-link probes took all six, and the page someone opened queued 2-20 s
behind them. Preloads and the index now share a slow lane of one connection,
registries and probes a quiet lane of two; over HTTP/2 the lanes widen. A
click on a folio still reads at once, taking over its queued preload, and a
host that refused its agent registry is asked again after a minute instead
of on every render.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
The channel waited for one file-info read per document (four at a time)
before it selected its report, and did again every 15 s. The reads now
follow the first show in the quiet lane, at most one pass per channel, and
redraw only when a time moved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
A report mounted on the stage reads its whole body and names itself, so its
64 KiB peek was a second read of the same bytes racing the first. Peeks also
run in the quiet lane, and a first owner file time for unchanged receipts no
longer counts as a new version to peek again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
Files above one MiB carried a weak size/mtime/inode ETag, which the reader's
live poll will not trust, so a selected 4-21 MB report was downloaded whole
every 4 s, through the relay when its owner was remote. A whole GET of a
document up to 64 MiB (not media) now carries its SHA-256, read once per
settled version: FileDigests remembers it under {path, size, mtime, ctime,
inode} only after the file has been still for two seconds, so a later write's
new ctime can never meet a stale digest. Ranged, HEAD and media reads stay
bounded and never read a file to make one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
The first read of a watched document used no-store, so reopening the board
downloaded every report again. It now uses no-cache: the browser sends its
stored validator and an unchanged report comes back as a 304 from its cache.
Later polls and forced refreshes still go to the owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
The reader re-reads its channel's fiber every 15 s. A body relayed from a
remote host takes up to 20 s, so the refresh held a connection nearly all
the time. It now waits in the slow lane. Opening a channel still reads at
once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
A stalled open or switch dropped its whole run, so the worst stalls vanished
from the medians. It now counts as 30 s, and a crashed browser is relaunched
for the next run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
Every waveform decode constructed a realtime AudioContext, which opens a
session with the system audio server (coreaudiod on macOS) for a file nobody
was playing, and a closed tab, killed browser or abandoned decode never closed
it. Decoding now uses an OfflineAudioContext, which opens no device, at
8 kHz: a 26-minute recording decodes to about a sixth of the samples it did
at 48 kHz, and its 1000 peaks differ from a 44.1 kHz decode by a median of
0.02 (in Chrome, on three real mp3s). Playback stays on the <audio> element.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
A receded neighbour downloaded and decoded its whole recording for a poster
nobody was listening to. A neighbour now draws peaks only when this session
already decoded them, and a page decodes when it is first selected. Both
reads wait in the listening page's shared read slots.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
The title probe, each thumbnail family and a hover preview each read a
document's first 64 KiB on their own, so one report was peeked two or three
times within seconds. documentResources.peekDocument reads it once for
everyone asking at the same time, keeps it 30 s, and queues peeks in the
quiet lane by what they serve (selected, neighbours, titles, thumbnails,
durations). A version that moved past an earlier peek reads fresh. It is
the seed of the document resource cache the next refactor grows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
…eRefresh

The probe asks whether a page is already reading a document; the mocks
lacked that export and the suite reported unhandled errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
@cailmdaley
cailmdaley merged commit ff6afc2 into main Oct 5, 2026
3 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant