Repository navigation
vm: inotify: sync file events in batches - #1649
Open
fredericmorin-flare wants to merge 4 commits into
Open
fredericmorin-flare wants to merge 4 commits into
fredericmorin-flare wants to merge 4 commits into
Conversation
Every file event was synced with two `limactl shell` commands run in sequence (a `stat`, then the sync), about 100ms each. Events were handled at roughly 4 per second while the watcher waited, so changing a few hundred files (switching branches, formatting a project) took close to a minute to reach the containers. Events above 50 unique files per 500ms were dropped. - Collect events for 100ms, deduplicated by path, and sync them with a single command. The command is split when its arguments would exceed 64KiB. - Check that the file exists in the VM within the sync command instead of running a separate `stat`. - Run the sync in the background so events keep being received while it runs. Events received meanwhile are synced right after. - Remove the rate limit, which dropped events: batching bounds the number of commands run in the VM. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Frederic Morin <frederic.morin@flare.io>
On macOS, a chmod run in the VM on a file that was previously opened for writing from the VM emits a write event on the host. Since abiosoft#1643 every sync does both, so syncing a file emits an event for that same file, which is synced again, endlessly. The rate limit only slowed this loop down; batching removes it. Remember the state of the files on the host (inode, size, mode and modification time) when they are synced, and skip events for files whose state has not changed since. These events are either caused by the sync itself, or carry nothing that the VM has not already been notified of. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Frederic Morin <frederic.morin@flare.io>
notify drops events when the channel it sends to is full. The channel had a buffer of one, so most events of a burst of file changes were lost: rewriting 2000 files from the host delivered fewer than half of them to the containers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Frederic Morin <frederic.morin@flare.io>
Syncing the 27k files of a project checkout took 36s after the last change, and some events were lost on the host. - The watcher only forwards the paths of the events. Checking the files on the host (stat, directory, unchanged since the last sync) moves to the sync, once per path and batch. Events are received as fast as notify sends them, and it drops those that are not received in time. - chmod runs once per batch and file mode instead of once per file: starting a process for each file is what took most of the time in the VM. - Up to 4 sync commands run at the same time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Frederic Morin <frederic.morin@flare.io>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Related: #1543, #1643.
Problem
With
--mount-inotify, every host file event is synced by running twolimactl shellcommands in sequence: astat, then the sync. Each command takes about 100ms, and the watcher waits while they run. Events are handled at roughly 4 per second. Changing a few hundred files at once, for example switching branches or formatting a project, takes close to a minute to reach the containers. Watchers such aspants --loopor pantsd keep running on stale files until then. The rate limit also drops events above 50 unique files per 500ms.Making this faster exposes a feedback loop. On macOS, a
chmodrun in the VM on a file that was previously opened for writing from the VM emits a write event on the host. Since #1643, every sync does both (chmod, then: >>). So syncing a file emits an event for that same file, which is synced again, endlessly. The rate limit only slowed this loop down. It shows in the daemon log as the same files being synced every second or so, long after they were last changed.Change
Batching (
sync file events in batches):[ -e "$f" ]), which removes the separatestatcommand.Loop fix (
skip events for files unchanged since their last sync):Event buffer (
buffer the file events channel):Large bursts (
keep up with large bursts of file events):chmodruns once per batch and file mode instead of once per file. Starting a process for each file is what took most of the time in the VM (about 680µs per file, down to 260µs).Results
macOS 26 (Apple M3 Pro), vz + virtiofs, docker runtime. The test rewrites N existing files from the host in one burst, then measures how long until every file's
IN_CLOSE_WRITEis seen in a container watching the directory:The time is measured from the end of the burst on the host. Writing the 27,469 files took 7–9s.
Before the loop fix, one host write to a file made the daemon sync it ~25 times over 3 seconds, and it kept going. With the fix, it syncs once; the echo event is received and skipped.
Known limitation
In bursts of tens of thousands of files, macOS can still drop events before notify sees them. FSEvents then reports
UserDroppedandMustScanSubDirson the root of the watched directory, flags that notify discards unless a watcher subscribes to them. This is how the second 27k run lost 94 files. Fixing it needs changes in how events are received from FSEvents, which I'd like to address in a follow-up.Tests
events_test.go: generated commands (one per file mode, split when too large), batching (deduplication, events received during a sync form the next batch), and the syncer (directories and missing files skipped, unchanged files skipped, changed files synced again).events_linux_test.go(Linux only, runs in CI): runs one batch command over several files, one of them missing, and checks that each existing file getsIN_ATTRIBandIN_CLOSE_WRITE, noIN_MODIFY, and keeps its content and modification time.All tests pass on macOS and on Linux, with
-race, andgolangci-lintreports no issues.LLM usage disclosure
This change was written with the help of an LLM (Claude): the investigation, the fix and the tests. I reviewed every line and ran the tests and measurements above myself.
🤖 Generated with Claude Code