You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add two opt-in storage APIs for loading external initializers into CUDA memory:
Linux NVIDIA GPUDirect Storage:external_data_loader_use_gds=1 dynamically loads cuFile, requests native GDS with compatibility mode disabled, reads aligned blocks into a registered GPU staging buffer, and copies them into CUDA arena allocations.
Windows Microsoft DirectStorage:external_data_loader_use_directstorage=1 uses a D3D12 buffer and fence shared with the same CUDA device (matched by LUID). Build with onnxruntime_USE_CUDA_DIRECTSTORAGE=ON and deploy the matching DirectStorage runtime DLLs. Files are identity-checked against the already validated Windows handle before reads.
Both options default to disabled. Setup/read failures are logged and use the configured host-memory fallback: external_data_loader_reading_threads=0 selects pageable loading; 1..64 selects pinned-buffer loading.
Successful loads log their actual path and byte count, so benchmarks cannot mislabel fallback as direct storage.
The Windows backend is included intentionally to provide the corresponding Microsoft API alongside Linux GDS in the same CUDA external-initializer feature. Microsoft DirectStorage's uncompressed flow can use host/upload staging; using its API is not a claim of zero-host-copy NVMe-to-VRAM DMA. This does not add a DirectML backend or compressed-weight support.
Benchmark
The benchmark compares pageable, pinned, and platform-specific direct storage on the same CUDA GPU using deterministic aligned external weights, fresh processes, independent correctness checks, and actual-path byte accounting. It reports distributions and effective end-to-end initialization throughput, not raw storage bandwidth.
Windows 11, RTX 4060 Laptop GPU (8 GiB), local WD 2 TB SSD, driver 591.55, CUDA 13.0.2, MSVC 2022, DirectStorage 1.2.3. Five measured repetitions after one warmup per path; OS-managed/potentially warm caches:
Path
1 GiB median initialization
4 GiB median initialization
CPU pageable
1.517 s
6.201 s
Pinned buffers
1.105 s
2.292 s
Microsoft DirectStorage
1.853 s
7.870 s
All 30 measured samples passed output verification and accounted for the full weight byte count through the requested path. Pinned buffers are faster on this machine; the current uncompressed DirectStorage implementation does not improve these measurements. Linux native GDS was not benchmarked here.
Validation
Built source ORT Python bindings and the CUDA provider on Windows with DirectStorage enabled, using a reduced MatMul/Gather Release build with contrib support enabled.
Verified CUDA MatMul/Gather dispatch with CPU EP fallback disabled.
Loaded two unaligned 64 MiB + 17 byte tensors through each of pageable, pinned, and DirectStorage, checking every byte against the independent source pattern; logs confirmed all three actual paths.
Ran 16 CPU-only benchmark tests, including mixed UTF-8/UTF-16LE Windows log parsing, path/fallback accounting, fixture generation, and correctness checks.
Added C++ coverage for DirectStorage option validation, aligned/unaligned/repeated/multi-buffer loads, and a backend-direct test that cannot pass through host fallback. The backend-direct test explicitly skips unavailable D3D12/CUDA/DirectStorage environments.
C++ and Python formatting/lint checks passed.
Configuration, deployment, benchmark commands, limitations, and results are documented in docs/Model_Loading_Performance.md.
Disable cuFile compatibility mode so unsupported systems use the configured ONNX Runtime pinned-buffer fallback instead of cuFile's internal POSIX path.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Model: qwen3.5-35b-cuda-bf16/model.onnx with a 69,451,776,000-byte external-data file.
Environment: NVIDIA H200 (physical GPU 2), 96 intra-op threads, spinning disabled, fresh process per sample. Before every sample both model files received POSIX_FADV_DONTNEED. Two cold-cache samples were collected per configuration.
The pinned loader is 7.43x faster than the classic path on this machine (86.5% lower session-creation time).
Native GDS could not be measured: nvidia-fs is not installed for kernel 6.6.141.1-1.azl3 (modinfo nvidia_fs reports module not found, /dev/nvidia-fs is absent). The GDS-requested samples therefore exercised the intended pinned fallback and logged cuFileDriverOpen failed: nvidia-fs driver is not loaded. Their 2.9% mean difference from pinned-only is run-to-run/storage variation plus the one-time failed GDS initialization.
I also disabled cuFile compatibility mode in 606d6a20ce, ensuring the GDS option measures/uses native GDS only; unsupported systems now fall back to the ORT pinned loader rather than silently using cuFile POSIX compatibility mode.
Request the driverless PCI P2PDMA path before opening cuFile while retaining pinned-buffer fallback when the host kernel driver or PCIe topology is unsupported.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Tested the CUDA 12.8+ PCI P2PDMA path on the current H200 host by setting CUFILE_PARAM_USE_PCIP2PDMA=true before opening cuFile. The build succeeded, but native GDS still could not initialize: this host uses the proprietary R580 kernel module (modinfo nvidia reports license: NVIDIA), whereas driverless P2PDMA requires the open NVIDIA kernel module. cuFileDriverOpen therefore returned error 5001 and ORT correctly used the pinned fallback. The full cold Qwen BF16 load completed in 152.391 s.
Commit HEAD keeps the P2PDMA request so compatible hosts can use native GDS without nvidia-fs; hosts like this one continue to fall back explicitly.
Avoid changing the RandomAccessFile vtable, guard optional cuFile headers, make GDS test counters atomic, and clarify diagnostics and documentation.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Probe the required cuFile configuration API before enabling native GDS, so older CUDA toolkits compile and retain the logged host-memory fallback. Guard the core Tensor forward declaration from SHARED_PROVIDER builds, where Tensor is a struct.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Linux TensorRT/full and CUDA-minimal builds use cuFile headers without cuFileSetParameterBool and the required configuration enums. CMake now probes that API before compiling native GDS; older toolkits retain the existing warning and host-memory fallback.
Windows CUDA and the documentation build hit MSVC C4099 because the GDS header forward-declared class Tensor after the shared-provider bridge declared struct Tensor. The forward declaration is now guarded by SHARED_PROVIDER, matching the existing external-data-loader interface.
Documentation validation was failing because its Windows build failed, not because generated operator documentation was stale. Fresh CI is triggered by this push; no old runs were rerun.
Because active_driver is weak, releasing the last shared_ptr expires it before ~CuFileDriver acquires GlobalMutex(). A concurrent Acquire can take the mutex first, see no active owner, and call cuFileDriverOpen while the old driver is still open; that initialization fails and permanently sends the new loader to fallback. Keep ownership and close/reopen serialization in one process-wide lifetime manager (or retain the driver for process lifetime), and cover concurrent final release/acquire with a regression test.
Document pinned or pageable host-buffer fallback
docs/CUDA_GPU_Direct_Storage.md:5
This fallback is not always pinned: when external_data_loader_reading_threads is 0, LoadTensor uses the pageable-buffer path. Describe this as the configured pinned or pageable host-buffer fallback so the introductory behavior matches the configuration section.
The fallback is pageable when external_data_loader_reading_threads is 0, so this public option comment is inaccurate for a valid configuration. Refer to the configured host-memory loader instead.
Encapsulate shared driver ownership in a handle that acquires the lifetime mutex before dropping its strong reference. Driver teardown therefore completes before another loader can configure a replacement. Keep the mutex state alive with each handle, preserve shared reuse and initialization-error cleanup, and cover final close versus reacquire without requiring native GDS.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🟡 Changes recommended
The native cuFile read implementation lacks direct hardware-independent tests, and the fallback overview is inconsistent with pageable fallback behavior.
Get a fresh assessment by requesting another Copilot review.
Clarify fallback host memory can be pinned or pageable
docs/CUDA_GPU_Direct_Storage.md:5
This says fallback is always pinned, but the documented and implemented external_data_loader_reading_threads=0 behavior uses pageable host memory (see line 53). Describe this as the configured pinned/pageable host-memory fallback so the overview does not contradict the configuration section.
Extract the existing native descriptor and chunked-read implementation behind cuFile and device-operation callbacks. Exercise that same implementation with host-only tests for multi-chunk reads, error decoding, copy/synchronization failures, and descriptor/handle cleanup. Wire coverage into regular and plugin CUDA internal tests.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Call CUDA allocation, stream creation, and GdsLoader::Create directly. Remove the injectable native GDS read API and restore the private native implementation. Replace injected loader tests with real-provider integration coverage while retaining shared-driver lifetime safety.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address the documentation note embedded in the Copilot review summary: fallback can use pinned or pageable memory, depending on the configured reading-thread count.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fixed the remaining documentation note embedded under “Previously missed” in #32713 (review) in fdea334. Both the opening overview and the compatibility-mode paragraph now describe the configured host-memory fallback, including pageable memory when reading_threads=0.
The inline review threads are resolved. The earlier Copilot overview still contains its historical “Open (1)” heading; that does not reflect the current thread state. The injection-based testing recommendation was explicitly withdrawn at the author’s request in 3704b43, as documented in the thread, rather than reintroduced into the production API.
Added two tests using the existing public CUDA provider-options API, with no production hooks or replacement parser:
Omitted option defaults to 0; string updates 0 → 1 → 0 update the actual struct field. GetCUDAProviderOptionsAsString output is decoded and fed back into Update, verifying the GDS value and full serialized round trip.
Values 2, -1, invalid text, and an empty value are rejected with an appropriate diagnostic, without changing the previous valid setting.
Update exercises CUDAExecutionProviderInfo::FromProviderOptions through the real provider factory, and serialization exercises ToProviderOptions(OrtCUDAProviderOptionsV2). Rebuilt onnxruntime_shared_lib_test and ran both new CApiTest cases successfully for two repetitions. Targeted formatting also passed.
Copilot resolve the merge conflicts in this pull request, also move the markdown documentation to the file docs/Model_Loading_Performance.md
Resolved in 65049bd. Merged the latest main, resolved the FAQ conflict, moved the GDS guidance into docs/Model_Loading_Performance.md, and removed the standalone GDS document.
Xavier Dupré (xadupre)
changed the title
[DRAFT] Add GPUDirect Storage model loading
[DRAFT] Add CUDA weight loading with GPUDirect Storage and Microsoft DirectStorage
Sep 23, 2026
Zero-byte external initializers satisfy the alignment test, so enabling GDS attempts driver/buffer/file registration even though there is nothing to read. A setup or file-registration failure on such an initializer then sets gds_disabled_ and prevents later non-empty aligned weights from using GDS. Skip the GDS attempt for zero-length tensors, as the DirectStorage branch already does.
Clarify the behavior of CUDAExecutionProviderInfo and its options. Introduce a helper to check cuDNN status in kernels.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Declare BUILD_UNIT_TESTS before the dependent CUDA internal-test option so a fresh configure honors ENABLE_CUDA_EP_INTERNAL_TESTS=ON and writes the DirectStorage SDK manifest required by Windows CI.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Add two opt-in storage APIs for loading external initializers into CUDA memory:
external_data_loader_use_gds=1dynamically loads cuFile, requests native GDS with compatibility mode disabled, reads aligned blocks into a registered GPU staging buffer, and copies them into CUDA arena allocations.external_data_loader_use_directstorage=1uses a D3D12 buffer and fence shared with the same CUDA device (matched by LUID). Build withonnxruntime_USE_CUDA_DIRECTSTORAGE=ONand deploy the matching DirectStorage runtime DLLs. Files are identity-checked against the already validated Windows handle before reads.external_data_loader_reading_threads=0selects pageable loading;1..64selects pinned-buffer loading.The Windows backend is included intentionally to provide the corresponding Microsoft API alongside Linux GDS in the same CUDA external-initializer feature. Microsoft DirectStorage's uncompressed flow can use host/upload staging; using its API is not a claim of zero-host-copy NVMe-to-VRAM DMA. This does not add a DirectML backend or compressed-weight support.
Benchmark
The benchmark compares pageable, pinned, and platform-specific direct storage on the same CUDA GPU using deterministic aligned external weights, fresh processes, independent correctness checks, and actual-path byte accounting. It reports distributions and effective end-to-end initialization throughput, not raw storage bandwidth.
Windows 11, RTX 4060 Laptop GPU (8 GiB), local WD 2 TB SSD, driver 591.55, CUDA 13.0.2, MSVC 2022, DirectStorage 1.2.3. Five measured repetitions after one warmup per path; OS-managed/potentially warm caches:
All 30 measured samples passed output verification and accounted for the full weight byte count through the requested path. Pinned buffers are faster on this machine; the current uncompressed DirectStorage implementation does not improve these measurements. Linux native GDS was not benchmarked here.
Validation
Configuration, deployment, benchmark commands, limitations, and results are documented in
docs/Model_Loading_Performance.md.