Daemon: stop a missed fiber read from wedging the hub at its fd limit - #11
Merged
Merged
Conversation
Cherry-picked from 4645df48 on fix/hub-emfile. Co-Authored-By: GPT-6.1 Sol <noreply@openai.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
…non-latching A GET for a fiber no store resolves fell through to a scan that listed every fiber with bodies (`felt ls --body`) and stat'ed a report sibling for each row through the VM's single file server. Concurrent misses convoyed for minutes, holding their sockets after clients left, until the process ran out of descriptors. At EMFILE the store registry read as `[]`; the store expansion cached that empty base, so every later request re-walked the store tree on its own process, which kept the daemon at the limit after the load had gone. - The scan lists metadata only and stats nothing (both `ls` surfaces carry the native `report_path`), and takes the match's body from a `show` of its traversal id. The list endpoint stops stat'ing too. - An unreadable registry keeps the last good store expansion. - `Shuttle.SingleFlight` shares one run among concurrent identical fiber reads and per-store listings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
cailmdaley
force-pushed
the
fix/hub-wedge
branch
from
October 5, 2026 20:35
bf601b6 to
69f5538
Compare
… a fresh show A coalesced read can return a lookup that began before the caller's last write, so `get/2` itself stays uncoalesced for the Poller and action resolution; the controller's read path, which a board hammers, coalesces. The scan's listing only picks the fiber, and the answer comes from a `show` of its traversal id, so a shared listing never serves stale metadata. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3
…he hard fd limit Review of the wedge fix: - `felt ls --body` carries no native `report_path`, so a body listing stats for it again; the metadata listing still never stats. - Appending a project reads projects.json strictly: an unreadable file fails the request instead of being overwritten with one path. - The plist raises only the soft open-file limit, so the daemon's children keep launchd's default hard limit. - The controller's coalescing comment says what internal reads share. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UPZ4Bmrtmhgm6yUBjFm7CK
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The hub wedged five times on 5 Oct with its 256-descriptor soft limit full of
CLOSEDsockets. This PR fixes the cause, which an isolated daemon reproduced with request load alone. No macOS privacy (TCC) prompt was involved: the reproduction ran from a shell against a store outside the protected folders.What wedged it
Measured on a release of
origin/mainserving a copy of a 6,671-fiber store on its own port, withulimit -n 256. A load generator mimicked the board: composite feeds, fiber documents, files, and the six missing UUIDs from the hub's log, with clients abandoning requests mid-flight. Handler stacks were read over the release's remote shell while the daemon was wedged.GET /api/v1/fibers/<id>thatfelt showcannot resolve falls through toscan_lookup. That ranfelt ls -s all -j --body(26 MB), decoded it, and stat'edreport.htmlfor each of the ~6,500 rows that lack a report. Every stat went through the VM's singlefile_server_2, so concurrent misses queued behind one another. A request abandoned by its client kept running and kept its socket, which shows asCLOSEDinlsof. On the hub this path was the leading:emfilesite in the 5 Oct log:FiberDocuments.show_storeandscan_lookupappear in 505 and 53 traces.PathListConfig.registered/1returns[]for any read failure,:emfileincluded.FeltStores.cached_expansion/1cached the empty base. From then on, every request saw a changed base and re-walked the store tree on its own process (aFile.lsandFile.lstatper directory, all through the same file server). In one capture 110 handlers were walking at once, and a singleconfigured_stores()call took 133 s. The walks outlived the load by minutes, so the daemon stayed at the limit after clients stopped. That is the hub's signature: the process is alive, CPU is low andCLOSEDsockets pile up.contract skewbecauseshuttle contractcould not spawn. In one baseline run the application then exceeded its restart intensity and the VM exited. The beam also gained up to ~100 write descriptors on its own log file across EMFILE episodes, and they were never released.The change
lswithout--body) and stats nothing, because bothfelt lsandshuttle lscarry the nativereport_path. The listing only picks the fiber; the answer comes from a freshshowof its traversal id, which resolves directly. The match semantics, wirepathand error shapes are unchanged. The metadata listing behindGET /api/v1/fibersstops stat'ing too, which matches the invariantreport_present?/3already documents.ls --bodyomitsreport_path, so?body=truestill finds each report by stat.PathListConfig.read_configured/1distinguishes an absent file from one that cannot be read. When the read fails,FeltStoreskeeps serving its last good expansion.configured/1andregistered/1keep their Go-parity[]contract.Shuttle.SingleFlightcoalesces concurrentGET /api/v1/fibers/:idreads of the same fiber, and concurrent per-store listings for distinct misses. It caches nothing.FiberDocuments.get/2itself stays uncoalesced, so the Poller and action resolution never receive a lookup that began before their own write.projects.jsonstrictly and fails the request on a read error, instead of overwriting the list with the one new path.NumberOfFilesto 8192 and leaves the hard limit at launchd's default, so the daemon's children (a tmux server and its workers) keep their headroom; the systemd unit setsLimitNOFILE=8192(cherry-picked from 4645df48 onfix/hub-emfile). Installed supervisors needshuttle daemon installto pick this up; the docs say so.Before and after
Protocol: a 60 s abusive burst (48 clients, half abandoning requests, 40% missing ids), then 120 s of normal board use (6 clients), with descriptor counts every 2 s and a fresh-connection
/api/v1/versionprobe every second. The machine's load average was 90–150 from other work, so read the ratios, not the absolute latencies.:emfilelinesmain(256), run 1main(256), run 2CLOSEDheld, 53 handlers still walking the store treemain(256), run 3main(8192)Whether
mainlatches after reaching the limit depends on whether a store-registry read lands on EMFILE, so outcomes vary run to run. In all three baseline runs it reached the limit; with this PR it stayed well below.Twenty concurrent requests for distinct missing fibers took a median of 62–95 s on
mainand 9–13 s with this PR. Twenty concurrent opens of one fiber took 24 s and 5 s.The 8192 limit alone keeps
mainfrom exhausting under this load. The code change removes the latch and the miss path's cost, so a burst that does reach the limit no longer keeps the daemon there.Not in this PR
From
fix/hub-emfile, only the supervisor limits are kept. Its rawFileAccesspool, request-lifetime fuses, relay admission, blocked-prefix memory and version diagnostics answer a blockedopen()under TCC. That hazard is real but separate, and was not needed to reproduce this wedge.Follow-ups this work surfaced:
felt show <uid>walks the whole store (about 1.8 s here), so every document open by uid pays a full tree walk in felt.Shuttle.Runnerraises on a failed spawn, so EMFILE crashes the Poller. A contract check that fails to spawn reports skew rather than "could not run".bugs3is fixing that client side.Gates
After review:
make mix-test1687 tests, 0 failures;go test ./internal/shuttleclipasses;plutil -linton the plist template is OK. The new body-listing and project-list tests each fail against the reverted code.make mix-test: 1682 tests, 0 failures before rebasing. At the head commit, 1685 tests, 1 failure:DispatchIntegrationTest"closed fiber is refused" recorded another test'stmux has-sessioncall. That file passes 3 of 3 runs alone. These suites are flaky under load onorigin/maintoo: fourPollerTestruns there each failed one to three tests, andOriginRouterTest's "never logs" assertion failed once on this branch from another test's log line.go test ./...: pass.make test-linux: pass.The new tests fail against the reverted code: the unreadable-registry test and the metadata-only scan test.
🤖 Generated with Claude Code
https://claude.ai/code/session_01XTQ9CSGaUs1DbFHxME9BB3