Skip to content

Daemon: stop a missed fiber read from wedging the hub at its fd limit - #11

Merged
cailmdaley merged 4 commits into
mainfrom
fix/hub-wedge
Oct 5, 2026
Merged

cailmdaley merged 4 commits into
mainfrom
fix/hub-wedge

Conversation

@cailmdaley

@cailmdaley cailmdaley commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

The hub wedged five times on 5 Oct with its 256-descriptor soft limit full of CLOSED sockets. This PR fixes the cause, which an isolated daemon reproduced with request load alone. No macOS privacy (TCC) prompt was involved: the reproduction ran from a shell against a store outside the protected folders.

What wedged it

Measured on a release of origin/main serving a copy of a 6,671-fiber store on its own port, with ulimit -n 256. A load generator mimicked the board: composite feeds, fiber documents, files, and the six missing UUIDs from the hub's log, with clients abandoning requests mid-flight. Handler stacks were read over the release's remote shell while the daemon was wedged.

  1. A missing fiber cost a full store dump. GET /api/v1/fibers/<id> that felt show cannot resolve falls through to scan_lookup. That ran felt ls -s all -j --body (26 MB), decoded it, and stat'ed report.html for each of the ~6,500 rows that lack a report. Every stat went through the VM's single file_server_2, so concurrent misses queued behind one another. A request abandoned by its client kept running and kept its socket, which shows as CLOSED in lsof. On the hub this path was the leading :emfile site in the 5 Oct log: FiberDocuments.show_store and scan_lookup appear in 505 and 53 traces.
  2. At the limit the store list read as empty and was cached. PathListConfig.registered/1 returns [] for any read failure, :emfile included. FeltStores.cached_expansion/1 cached the empty base. From then on, every request saw a changed base and re-walked the store tree on its own process (a File.ls and File.lstat per directory, all through the same file server). In one capture 110 handlers were walking at once, and a single configured_stores() call took 133 s. The walks outlived the load by minutes, so the daemon stayed at the limit after clients stopped. That is the hub's signature: the process is alive, CPU is low and CLOSED sockets pile up.
  3. The limit's fallout. Under EMFILE the Poller crashed and restarted, re-arming the boot quarantine and logging a false contract skew because shuttle contract could not spawn. In one baseline run the application then exceeded its restart intensity and the VM exited. The beam also gained up to ~100 write descriptors on its own log file across EMFILE episodes, and they were never released.

The change

  • Missed reads stay cheap. The scan lists metadata only (ls without --body) and stats nothing, because both felt ls and shuttle ls carry the native report_path. The listing only picks the fiber; the answer comes from a fresh show of its traversal id, which resolves directly. The match semantics, wire path and error shapes are unchanged. The metadata listing behind GET /api/v1/fibers stops stat'ing too, which matches the invariant report_present?/3 already documents. ls --body omits report_path, so ?body=true still finds each report by stat.
  • An unreadable registry is not an empty one. PathListConfig.read_configured/1 distinguishes an absent file from one that cannot be read. When the read fails, FeltStores keeps serving its last good expansion. configured/1 and registered/1 keep their Go-parity [] contract.
  • Concurrent identical reads share one run. Shuttle.SingleFlight coalesces concurrent GET /api/v1/fibers/:id reads of the same fiber, and concurrent per-store listings for distinct misses. It caches nothing. FiberDocuments.get/2 itself stays uncoalesced, so the Poller and action resolution never receive a lookup that began before their own write.
  • An unreadable project list is not an empty one either. Adding a project reads projects.json strictly and fails the request on a read error, instead of overwriting the list with the one new path.
  • Descriptor headroom. The launchd template raises the soft NumberOfFiles to 8192 and leaves the hard limit at launchd's default, so the daemon's children (a tmux server and its workers) keep their headroom; the systemd unit sets LimitNOFILE=8192 (cherry-picked from 4645df48 on fix/hub-emfile). Installed supervisors need shuttle daemon install to pick this up; the docs say so.

Before and after

Protocol: a 60 s abusive burst (48 clients, half abandoning requests, 40% missing ids), then 120 s of normal board use (6 clients), with descriptor counts every 2 s and a fresh-connection /api/v1/version probe every second. The machine's load average was 90–150 from other work, so read the ratios, not the absolute latencies.

run (fd limit) peak fds :emfile lines normal phase: requests OK afterwards
main (256), run 1 260 yes 0 the application exceeded its restart intensity and the VM exited
main (256), run 2 264 552 24 of 31 in 2 min 54 CLOSED held, 53 handlers still walking the store tree
main (256), run 3 263 78 567 of 582 recovered
this PR (256), two runs 162, 171 0 753 of 774, 1189 of 1209 back to 34–36 fds
main (8192) 327 0 438 of 449 recovered

Whether main latches after reaching the limit depends on whether a store-registry read lands on EMFILE, so outcomes vary run to run. In all three baseline runs it reached the limit; with this PR it stayed well below.

Twenty concurrent requests for distinct missing fibers took a median of 62–95 s on main and 9–13 s with this PR. Twenty concurrent opens of one fiber took 24 s and 5 s.

The 8192 limit alone keeps main from exhausting under this load. The code change removes the latch and the miss path's cost, so a burst that does reach the limit no longer keeps the daemon there.

Not in this PR

From fix/hub-emfile, only the supervisor limits are kept. Its raw FileAccess pool, request-lifetime fuses, relay admission, blocked-prefix memory and version diagnostics answer a blocked open() under TCC. That hazard is real but separate, and was not needed to reproduce this wedge.

Follow-ups this work surfaced:

  • felt show <uid> walks the whole store (about 1.8 s here), so every document open by uid pays a full tree walk in felt.
  • Shuttle.Runner raises on a failed spawn, so EMFILE crashes the Poller. A contract check that fails to spawn reports skew rather than "could not run".
  • The board re-requests missing fibers in a loop; lane bugs3 is fixing that client side.

Gates

  • After review: make mix-test 1687 tests, 0 failures; go test ./internal/shuttlecli passes; plutil -lint on the plist template is OK. The new body-listing and project-list tests each fail against the reverted code.

  • make mix-test: 1682 tests, 0 failures before rebasing. At the head commit, 1685 tests, 1 failure: DispatchIntegrationTest "closed fiber is refused" recorded another test's tmux has-session call. That file passes 3 of 3 runs alone. These suites are flaky under load on origin/main too: four PollerTest runs there each failed one to three tests, and OriginRouterTest's "never logs" assertion failed once on this branch from another test's log line.

  • go test ./...: pass.

  • make test-linux: pass.

  • The new tests fail against the reverted code: the unreadable-registry test and the metadata-only scan test.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3

cailmdaley and others added 2 commits October 5, 2026 22:32
Cherry-picked from 4645df48 on fix/hub-emfile.

Co-Authored-By: GPT-6.1 Sol <noreply@openai.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
…non-latching

A GET for a fiber no store resolves fell through to a scan that listed every
fiber with bodies (`felt ls --body`) and stat'ed a report sibling for each row
through the VM's single file server. Concurrent misses convoyed for minutes,
holding their sockets after clients left, until the process ran out of
descriptors. At EMFILE the store registry read as `[]`; the store expansion
cached that empty base, so every later request re-walked the store tree on its
own process, which kept the daemon at the limit after the load had gone.

- The scan lists metadata only and stats nothing (both `ls` surfaces carry the
  native `report_path`), and takes the match's body from a `show` of its
  traversal id. The list endpoint stops stat'ing too.
- An unreadable registry keeps the last good store expansion.
- `Shuttle.SingleFlight` shares one run among concurrent identical fiber reads
  and per-store listings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
cailmdaley and others added 2 commits October 5, 2026 22:47
… a fresh show

A coalesced read can return a lookup that began before the caller's last
write, so `get/2` itself stays uncoalesced for the Poller and action
resolution; the controller's read path, which a board hammers, coalesces. The
scan's listing only picks the fiber, and the answer comes from a `show` of its
traversal id, so a shared listing never serves stale metadata.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
…he hard fd limit

Review of the wedge fix:
- `felt ls --body` carries no native `report_path`, so a body listing
  stats for it again; the metadata listing still never stats.
- Appending a project reads projects.json strictly: an unreadable file
  fails the request instead of being overwritten with one path.
- The plist raises only the soft open-file limit, so the daemon's
  children keep launchd's default hard limit.
- The controller's coalescing comment says what internal reads share.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UPZ4Bmrtmhgm6yUBjFm7CK
@cailmdaley
cailmdaley merged commit 381d576 into main Oct 5, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant