Skip to content

QVAC-25501 fix: support header-only safetensors fitting - #46

Merged
DmitryMalishev merged 5 commits into
2026-08-11from
codex/QVAC-25501-safetensors-fit
Sep 25, 2026
Merged

DmitryMalishev merged 5 commits into
2026-08-11from
codex/QVAC-25501-safetensors-fit

Conversation

@DmitryMalishev

@DmitryMalishev DmitryMalishev commented Sep 23, 2026 •

Copy link
Copy Markdown

🎯 What problem does this PR solve?

sd_fit_params rejects header-only safetensors because the shared reader requires tensor payloads to exist. This prevents metadata-only fitting of models with safetensors components.

📝 How does it solve it?

  • Scope metadata-only reads to fit initialization and measurement, restoring the previous thread-local mode on every exit.
  • Preserve normal loading checks and metadata consistency validation, including offset overflow and invalid offset types.
  • Reject non-integer/negative dimensions and bound intermediate shape products before tensor creation, including five-dimensional collapse, zero dimensions and F64/I64 conversion.
  • Avoid loading LoRA weights or scale scalars during measurement; adapters loaded after initialization use the same metadata-only scope.
  • Initialize the GGML timer before timing a cold-start fit call on Windows.
  • Document that a successful fit does not validate weight-file completeness.

🧪 How was it tested?

  • Windows Release and Linux Vulkan builds: all 12 CTests passed at 8268f58.
  • Reproduced the review finding against 8b9c313: all three supplied malformed headers and F64/I64 overflow cases were accepted. 8268f58 rejects all five; the real FLUX header produces identical metadata for all 1,438 parsed tensors. New regression cases fail against the old reader and pass after the fix. Coverage also includes zero-dimension intermediate overflow, invalid shape types, valid scalar/zero-sized/five-dimensional shapes, strict loading, and SD_FIT_ERROR cleanup.
  • Regression coverage: complete/header-only/partial payloads, metadata equivalence, nested scopes, thread isolation, invalid headers and offsets, converted dtypes, shards, cold-start fit, and exception cleanup. LoRA alpha/scale tests verify identical graph structure and preserve real scalar reads outside measurement.
  • RTX 5090: the previous engine rejects the FLUX.1 Schnell FP8 header; this revision fits it successfully. Reported module memory matches a full-sized sparse reference with the same header. Normal generation still rejects the header-only file.
  • RTX 5090: header-only runtime LoRA with alpha fits successfully and has the same memory report as an adapter with payload bytes.
  • Windows and Linux addons build successfully through the temporary vcpkg overlay; 121 Bare tests (467 assertions) and 47 Node tests pass on each host.
  • Linux addon C++: 270 passed using the repository's existing CI LeakSanitizer configuration; 3 optional FLUX/Wan model cases lacked assets. RTX 5090 SD2 generation: 25 assertions passed. All 7 ABot integration tests passed (206 assertions), including scene creation and resident/streamed/disk parity.
  • Current revision: engine CI passed all build lanes, including CUDA and ROCm; its test-enabled Linux lane passed all 12 CTests. Addon CI passed all nine prebuild targets, C++ suites and eight desktop integration jobs.
  • The addon's separate Nx workflow still has one Linux integration job queued for a runner. It has not started or reported a failure; the diffusion-specific workflow above already passed the Linux integration suite. The queued job has not been cancelled or bypassed.
  • Current revision: Android, Pixel 9 Pro and iOS, iPhone 17 Pro passed against addon 1266832 / engine 8268f58: 23/23 runner groups on each device, no crashes, 48 non-skipped cases and 31 existing skips per device. Direct header-only fitting is covered by native engine tests. Existing mobile settings skip ABot and several large-model generation cases.
  • git diff --check passed.

The review fix and revalidation are complete; this PR is ready for re-review. Merge order: merge this PR after review, publish the merged engine pin in the native dependency registry, then replace the addon overlay with that registry version and validate the final addon revision.

Comment thread src/model_io/safetensors_io.cpp
@DmitryMalishev
DmitryMalishev merged commit 60bd9de into 2026-08-11 Sep 25, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants