Skip to content

perf(import): memoize file-match recall and ranking by identical evidence within one manual import preview - #253

Open
jordanfelle wants to merge 2 commits into
Chaptarr:developfrom
jordanfelle:perf-manualimport-per-file-match-memo
Open

jordanfelle wants to merge 2 commits into
Chaptarr:developfrom
jordanfelle:perf-manualimport-per-file-match-memo

Conversation

@jordanfelle

@jordanfelle jordanfelle commented Sep 27, 2026 •

Copy link
Copy Markdown

Fixes #244.

Problem

The manual import preview sets PerFileMatching, so every file of a multi-file group is matched independently, and each file runs a full staged match: an FTS recall query (RecallBooks) and a rank query (RankEditions). Files of one audiobook normally produce identical recall/rank inputs, but the only cache is by path, so the same search is repeated for every file. Measured on a live library: a 655-file audiobook preview took ~577 s on a repeat run (~0.9 s per file) even though tags were already cached (see #220 review: tag reads are not the cost).

Fix

Within one MatchFilesToLibraryAsync call with PerFileMatching, RecallBooks and RankEditions are memoized by their exact inputs:

  • recall key: author scope, media type, tokens;
  • rank key: media type, the gated books and the field queries.

Design points that keep results unchanged:

  • Only the search is memoized. Everything that depends on the individual file (duration checks, path fallback, evidence scoring, candidate gating) still runs per file, so files whose evidence differs never share a match; identical inputs give identical search results.
  • The memo lives in an AsyncLocal scoped to the call and restored afterwards (no cross-request or static cache), and values are stored serialized and returned as fresh copies, so callers never share or mutate instances.
  • It is skipped when matching trace is enabled, so trace output is unchanged.

Verification

  • per_file_memo_should_run_the_staged_search_once_per_distinct_evidence_and_not_change_results: three files with identical evidence run one recall and one rank (three each without the memo), with identical matches.
  • per_file_memo_should_not_share_results_between_files_with_different_evidence: two distinct evidence sets across three files run two recalls, results equal to the un-memoized run. Both fail (2 expected, 3 was) when the memo lookup is disabled.
  • 3,022 core tests pass; a dry merge into an integration branch that carries other import PRs is clean.

Not yet measured: the live speed-up. The searches are the repeated work by inspection, but how much of the ~0.9 s per file they account for has not been timed on a real folder.

Review update: NaN/Infinity scores made System.Text.Json throw while building the memo key or value, failing a match that worked unmemoized; the key is now built lazily and any key/serialize/deserialize failure falls back to computing directly (two tests, 3,024 core tests pass). Live result: the 655-file audiobook preview went 577 s -> 512 s (~11%), so recall/ranking are not the dominant per-file cost (see the linked issue for the remaining suspects).

The manual import preview matches every file of a multi-file group independently (PerFileMatching), and each file ran a full staged match: an FTS recall query and a rank query. Files of one audiobook normally produce identical recall and rank inputs, and the only cache was by path, so a 655-file preview paid the search 655 times (about 0.9 s per file on a repeat run, with tags already cached).

Within one MatchFilesToLibraryAsync call with PerFileMatching, memoize RecallBooks and RankEditions by their exact inputs (author scope, media type, tokens; gated books and field queries). Values are stored serialized and returned as fresh copies so no instance is shared, the memo lives in an AsyncLocal scoped to the call and restored afterwards, and it is skipped when matching trace is enabled. Everything that depends on the individual file (duration checks, path fallback, evidence scoring) still runs per file, so results are unchanged.
jordanfelle added a commit to jordanfelle/chaptarr that referenced this pull request Sep 27, 2026
jordanfelle added a commit to jordanfelle/chaptarr that referenced this pull request Sep 27, 2026
@jordanfelle

Copy link
Copy Markdown
Author

Live measurement with this change deployed: the same 655-file audiobook preview went from 577 s to 512 s (about 11%), so recall and ranking are not the dominant per-file cost. The change is result-preserving and still removes redundant work (one recall and rank per distinct evidence instead of per file), but the remaining ~0.75 s per file is elsewhere in the per-file preview path; profiling that is the next step.

…e keyed or serialized

System.Text.Json refuses NaN and Infinity, so a non-finite score in a recall or rank result made the memo (or the key built from it) throw and fail the whole match, where the unmemoized code worked. Build the key lazily and fall back to computing directly on any key, serialize or deserialize failure.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[PERF] Manual import preview matches every file of a multi-file group independently; memoize the match by identical tag evidence

1 participant