Speed up path matching and add label-only traversal - #4
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reduce path-matching overhead by storing pointers to corpus rules and rejecting invalid prefix boundaries before case folding. On Go 1.27.1, darwin/arm64, the median million-path benchmark fell from 274 ms to 209 ms across three runs. Matching the brief path inventory fell from 228 to 177 ns/path, with zero allocations. Corpus labels and evidence ordering are unchanged.
Add
WalkMatchandClassifier.WalkMatchfor visitors that only need a compactSet, plus CLI-labels-only. Both walk modes share traversal, limits and pruning semantics. For 10,000 files in a deep vendor/source/test layout, label-only traversal used 6.65 MB versus 34.05 MB for owned evidence results. Wide and monorepo disk timings were similar despite lower allocation, so this does not imply a speedup for every filesystem scan.An isolated brief integration called
Matchon accepted paths inside its existing indexing loop, without adding file reads or another traversal. Six alternating pairs of fullEngine.Runmeasurements over 10,000 synthetic files showed median scan overhead of 0.37%, with individual pairs between -0.37% and +1.03%. Import initialization was measured separately at roughly 1 ms. This experiment used minimal aggregation; storing richer per-file results would have additional costs. No brief integration is shipped here.Update API documentation and CLI examples, add concurrent matching benchmarks, and cover label-only traversal through public APIs and the CLI, including vendor-root context, pruning, limits and errors.