Skip to content

perf(import): read tags for a multi-file manual import preview in parallel - #220

Draft
jordanfelle wants to merge 2 commits into
Chaptarr:developfrom
jordanfelle:perf-manualimport-preview
Draft

jordanfelle wants to merge 2 commits into
Chaptarr:developfrom
jordanfelle:perf-manualimport-preview

Conversation

@jordanfelle

@jordanfelle jordanfelle commented Sep 26, 2026 •

Copy link
Copy Markdown

DRAFT.

Problem

The Manual Import preview reads tags and duration (ffprobe or TagLib over network storage) one file at a time, which took about 340 s for a 3.5 GB, roughly 100-file audiobook on a cold cache (a 655-file audiobook took 577 s).

Fix

Warm the existing tag cache with four concurrent reads before the per-file loop; the loop then reads every file from the cache, so results are unchanged. Files that fail tag extraction are read twice.

Verification

  • New SimpleImportDecisionMaker test (fails without the change).
  • Core tests pass (3,021 on top of develop).
  • Speedup not measured: there is no before/after timing on the same folder, so this stays a draft.

Review update:

  • The warm-up only helps a cold folder: tags are also persisted in a database cache keyed by path, modification time and size, so a re-preview already skips tag reads. A 655-file audiobook preview that took ~577 s on a repeat run is therefore dominated by per-file matching (PerFileMatching runs a full staged match with FTS recall for every file; see the linked issue), not by tag reads. This PR stays a draft until a cold folder has been timed before and after.
  • The tag cache is capped at 1,000 entries and clears itself when it overflows, so prefetching a bigger folder wiped the cache mid-prefetch. The prefetch is now capped at 800 files (test: multi_file_preview_should_not_prefetch_more_files_than_the_tag_cache_can_hold, failed with 887 reads against a cap of 800) (3,022 core tests pass).

The preview read tags and duration (ffprobe or TagLib over network storage) one file at a time, which took about 340 s for a 3.5 GB, roughly 100-file audiobook on a cold cache. Warm the existing tag cache with four concurrent reads before the per-file loop; the loop then reads every file from the cache, so results are unchanged.
jordanfelle added a commit to jordanfelle/chaptarr that referenced this pull request Sep 26, 2026
jordanfelle added a commit to jordanfelle/chaptarr that referenced this pull request Sep 26, 2026
jordanfelle added a commit to jordanfelle/chaptarr that referenced this pull request Sep 26, 2026
The tag extraction cache clears itself entirely once it passes 1,000 entries. Warming a folder with more files than that filled the cache, wiped it mid-prefetch, and the per-file loop then read every file again. Prefetch at most 800 files; the rest are read by the loop as before.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant