perf(import): memoize file-match recall and ranking by identical evidence within one manual import preview - #253
Open
jordanfelle wants to merge 2 commits into
Conversation
The manual import preview matches every file of a multi-file group independently (PerFileMatching), and each file ran a full staged match: an FTS recall query and a rank query. Files of one audiobook normally produce identical recall and rank inputs, and the only cache was by path, so a 655-file preview paid the search 655 times (about 0.9 s per file on a repeat run, with tags already cached). Within one MatchFilesToLibraryAsync call with PerFileMatching, memoize RecallBooks and RankEditions by their exact inputs (author scope, media type, tokens; gated books and field queries). Values are stored serialized and returned as fresh copies so no instance is shared, the memo lives in an AsyncLocal scoped to the call and restored afterwards, and it is skipped when matching trace is enabled. Everything that depends on the individual file (duration checks, path fallback, evidence scoring) still runs per file, so results are unchanged.
jordanfelle
added a commit
to jordanfelle/chaptarr
that referenced
this pull request
Sep 27, 2026
jordanfelle
added a commit
to jordanfelle/chaptarr
that referenced
this pull request
Sep 27, 2026
Author
|
Live measurement with this change deployed: the same 655-file audiobook preview went from 577 s to 512 s (about 11%), so recall and ranking are not the dominant per-file cost. The change is result-preserving and still removes redundant work (one recall and rank per distinct evidence instead of per file), but the remaining ~0.75 s per file is elsewhere in the per-file preview path; profiling that is the next step. |
…e keyed or serialized System.Text.Json refuses NaN and Infinity, so a non-finite score in a recall or rank result made the memo (or the key built from it) throw and fail the whole match, where the unmemoized code worked. Build the key lazily and fall back to computing directly on any key, serialize or deserialize failure.
jordanfelle
added a commit
to jordanfelle/chaptarr
that referenced
this pull request
Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #244.
Problem
The manual import preview sets
PerFileMatching, so every file of a multi-file group is matched independently, and each file runs a full staged match: an FTS recall query (RecallBooks) and a rank query (RankEditions). Files of one audiobook normally produce identical recall/rank inputs, but the only cache is by path, so the same search is repeated for every file. Measured on a live library: a 655-file audiobook preview took ~577 s on a repeat run (~0.9 s per file) even though tags were already cached (see #220 review: tag reads are not the cost).Fix
Within one
MatchFilesToLibraryAsynccall withPerFileMatching,RecallBooksandRankEditionsare memoized by their exact inputs:Design points that keep results unchanged:
AsyncLocalscoped to the call and restored afterwards (no cross-request or static cache), and values are stored serialized and returned as fresh copies, so callers never share or mutate instances.Verification
per_file_memo_should_run_the_staged_search_once_per_distinct_evidence_and_not_change_results: three files with identical evidence run one recall and one rank (three each without the memo), with identical matches.per_file_memo_should_not_share_results_between_files_with_different_evidence: two distinct evidence sets across three files run two recalls, results equal to the un-memoized run. Both fail (2 expected, 3 was) when the memo lookup is disabled.Not yet measured: the live speed-up. The searches are the repeated work by inspection, but how much of the ~0.9 s per file they account for has not been timed on a real folder.
Review update: NaN/Infinity scores made
System.Text.Jsonthrow while building the memo key or value, failing a match that worked unmemoized; the key is now built lazily and any key/serialize/deserialize failure falls back to computing directly (two tests, 3,024 core tests pass). Live result: the 655-file audiobook preview went 577 s -> 512 s (~11%), so recall/ranking are not the dominant per-file cost (see the linked issue for the remaining suspects).