Skip to content

[DRAFT] Add CUDA weight loading with GPUDirect Storage and Microsoft DirectStorage - #32713

Open
Xavier Dupré (xadupre) wants to merge 21 commits into
mainfrom
feature/cuda-gds-external-data
Open

Xavier Dupré (xadupre) wants to merge 21 commits into
mainfrom
feature/cuda-gds-external-data

Conversation

@xadupre

@xadupre Xavier Dupré (xadupre) commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Description

Add two opt-in storage APIs for loading external initializers into CUDA memory:

  • Linux NVIDIA GPUDirect Storage: external_data_loader_use_gds=1 dynamically loads cuFile, requests native GDS with compatibility mode disabled, reads aligned blocks into a registered GPU staging buffer, and copies them into CUDA arena allocations.
  • Windows Microsoft DirectStorage: external_data_loader_use_directstorage=1 uses a D3D12 buffer and fence shared with the same CUDA device (matched by LUID). Build with onnxruntime_USE_CUDA_DIRECTSTORAGE=ON and deploy the matching DirectStorage runtime DLLs. Files are identity-checked against the already validated Windows handle before reads.
  • Both options default to disabled. Setup/read failures are logged and use the configured host-memory fallback: external_data_loader_reading_threads=0 selects pageable loading; 1..64 selects pinned-buffer loading.
  • Successful loads log their actual path and byte count, so benchmarks cannot mislabel fallback as direct storage.

The Windows backend is included intentionally to provide the corresponding Microsoft API alongside Linux GDS in the same CUDA external-initializer feature. Microsoft DirectStorage's uncompressed flow can use host/upload staging; using its API is not a claim of zero-host-copy NVMe-to-VRAM DMA. This does not add a DirectML backend or compressed-weight support.

Benchmark

The benchmark compares pageable, pinned, and platform-specific direct storage on the same CUDA GPU using deterministic aligned external weights, fresh processes, independent correctness checks, and actual-path byte accounting. It reports distributions and effective end-to-end initialization throughput, not raw storage bandwidth.

Windows 11, RTX 4060 Laptop GPU (8 GiB), local WD 2 TB SSD, driver 591.55, CUDA 13.0.2, MSVC 2022, DirectStorage 1.2.3. Five measured repetitions after one warmup per path; OS-managed/potentially warm caches:

Path 1 GiB median initialization 4 GiB median initialization
CPU pageable 1.517 s 6.201 s
Pinned buffers 1.105 s 2.292 s
Microsoft DirectStorage 1.853 s 7.870 s

All 30 measured samples passed output verification and accounted for the full weight byte count through the requested path. Pinned buffers are faster on this machine; the current uncompressed DirectStorage implementation does not improve these measurements. Linux native GDS was not benchmarked here.

Validation

  • Built source ORT Python bindings and the CUDA provider on Windows with DirectStorage enabled, using a reduced MatMul/Gather Release build with contrib support enabled.
  • Verified CUDA MatMul/Gather dispatch with CPU EP fallback disabled.
  • Loaded two unaligned 64 MiB + 17 byte tensors through each of pageable, pinned, and DirectStorage, checking every byte against the independent source pattern; logs confirmed all three actual paths.
  • Ran 16 CPU-only benchmark tests, including mixed UTF-8/UTF-16LE Windows log parsing, path/fallback accounting, fixture generation, and correctness checks.
  • Added C++ coverage for DirectStorage option validation, aligned/unaligned/repeated/multi-buffer loads, and a backend-direct test that cannot pass through host fallback. The backend-direct test explicitly skips unavailable D3D12/CUDA/DirectStorage environments.
  • C++ and Python formatting/lint checks passed.

Configuration, deployment, benchmark commands, limitations, and results are documented in docs/Model_Loading_Performance.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings September 21, 2026 12:09

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Process-wide cuFile lifecycle, path reopening, and unaligned direct-I/O handling remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 2 Medium severity

Open (4)
What changed in this PR

Adds opt-in GPUDirect Storage loading for CUDA external initializers with configurable host-buffer fallback.

Changes:

  • Dynamically loads cuFile and stages external data through reusable GPU memory.
  • Adds CUDA provider options, validation, fallback behavior, and tests.
  • Documents configuration and prerequisites.
File Description
onnxruntime/​test/​providers/​cuda/​test_cases/​cuda_external_data_loader_test.cc Tests GDS selection and fallback behavior.
onnxruntime/​core/​session/​provider_bridge_ort.cc Validates the public GDS option.
onnxruntime/​core/​providers/​cuda/​cuda_provider_factory.cc Transfers GDS configuration into the provider.
onnxruntime/​core/​providers/​cuda/​cuda_external_data_loader.h Extends the loader with GDS state.
onnxruntime/​core/​providers/​cuda/​cuda_external_data_loader.cc Selects GDS or host-buffer fallback.
onnxruntime/​core/​providers/​cuda/​cuda_external_data_loader_gds.h Defines the GDS loader interface.
onnxruntime/​core/​providers/​cuda/​cuda_external_data_loader_gds.cc Implements dynamic cuFile loading and reads.
onnxruntime/​core/​providers/​cuda/​cuda_execution_provider.cc Creates GDS-enabled external-data loaders.
onnxruntime/​core/​providers/​cuda/​cuda_execution_provider_info.h Stores and hashes the GDS option.
onnxruntime/​core/​providers/​cuda/​cuda_execution_provider_info.cc Parses and serializes the GDS option.
include/​onnxruntime/​core/​providers/​cuda/​cuda_provider_options.h Exposes the public CUDA option.
docs/​FAQ.md Links to the GDS documentation.
docs/​CUDA_GPU_Direct_Storage.md Documents setup, behavior, and examples.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread onnxruntime/core/providers/cuda/cuda_external_data_loader_gds.cc Outdated
Comment thread onnxruntime/core/providers/cuda/cuda_external_data_loader_gds.cc
Comment thread onnxruntime/core/providers/cuda/cuda_external_data_loader.cc
Comment thread onnxruntime/core/providers/cuda/cuda_external_data_loader_gds.cc Outdated
Disable cuFile compatibility mode so unsupported systems use the configured ONNX Runtime pinned-buffer fallback instead of cuFile's internal POSIX path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@xadupre

Copy link
Copy Markdown
Member Author

CUDA model-loading benchmark

Model: qwen3.5-35b-cuda-bf16/model.onnx with a 69,451,776,000-byte external-data file.

Environment: NVIDIA H200 (physical GPU 2), 96 intra-op threads, spinning disabled, fresh process per sample. Before every sample both model files received POSIX_FADV_DONTNEED. Two cold-cache samples were collected per configuration.

Configuration Run 1 Run 2 Mean
Classic framework path (reading_threads=0, GDS off) 1,106.970 s 1,110.118 s 1,108.544 s
Pinned loader (reading_threads=4, GDS off) 148.712 s 149.554 s 149.133 s
GDS requested, pinned fallback 152.224 s 154.770 s 153.497 s

The pinned loader is 7.43x faster than the classic path on this machine (86.5% lower session-creation time).

Native GDS could not be measured: nvidia-fs is not installed for kernel 6.6.141.1-1.azl3 (modinfo nvidia_fs reports module not found, /dev/nvidia-fs is absent). The GDS-requested samples therefore exercised the intended pinned fallback and logged cuFileDriverOpen failed: nvidia-fs driver is not loaded. Their 2.9% mean difference from pinned-only is run-to-run/storage variation plus the one-time failed GDS initialization.

I also disabled cuFile compatibility mode in 606d6a20ce, ensuring the GDS option measures/uses native GDS only; unsupported systems now fall back to the ORT pinned loader rather than silently using cuFile POSIX compatibility mode.

Request the driverless PCI P2PDMA path before opening cuFile while retaining pinned-buffer fallback when the host kernel driver or PCIe topology is unsupported.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@xadupre

Copy link
Copy Markdown
Member Author

Tested the CUDA 12.8+ PCI P2PDMA path on the current H200 host by setting CUFILE_PARAM_USE_PCIP2PDMA=true before opening cuFile. The build succeeded, but native GDS still could not initialize: this host uses the proprietary R580 kernel module (modinfo nvidia reports license: NVIDIA), whereas driverless P2PDMA requires the open NVIDIA kernel module. cuFileDriverOpen therefore returned error 5001 and ORT correctly used the pinned fallback. The full cold Qwen BF16 load completed in 152.391 s.

Commit HEAD keeps the P2PDMA request so compatible hosts can use native GDS without nvidia-fs; hosts like this one continue to fall back explicitly.

Reuse validated file descriptors, share the process-wide cuFile driver, route unaligned ranges to fallback, and cover partial GDS failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@xadupre
Xavier Dupré (xadupre) requested a balanced review from Copilot September 21, 2026 15:35
@xadupre Xavier Dupré (xadupre) changed the title Add GPUDirect Storage model loading [DRAFT] Add GPUDirect Storage model loading Sep 21, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Comment thread onnxruntime/core/platform/env.h Outdated
Comment thread onnxruntime/core/providers/cuda/cuda_external_data_loader_gds.cc Outdated
Comment thread onnxruntime/test/providers/cuda/test_cases/cuda_external_data_loader_test.cc Outdated
Comment thread docs/CUDA_GPU_Direct_Storage.md Outdated
Comment thread onnxruntime/core/session/provider_bridge_ort.cc Outdated
Avoid changing the RandomAccessFile vtable, guard optional cuFile headers, make GDS test counters atomic, and clarify diagnostics and documentation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Probe the required cuFile configuration API before enabling native GDS, so older CUDA toolkits compile and retain the logged host-memory fallback. Guard the core Tensor forward declaration from SHARED_PROVIDER builds, where Tensor is a struct.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@xadupre

Copy link
Copy Markdown
Member Author

Fixed both CI compilation causes in 857b793:

  • Linux TensorRT/full and CUDA-minimal builds use cuFile headers without cuFileSetParameterBool and the required configuration enums. CMake now probes that API before compiling native GDS; older toolkits retain the existing warning and host-memory fallback.
  • Windows CUDA and the documentation build hit MSVC C4099 because the GDS header forward-declared class Tensor after the shared-provider bridge declared struct Tensor. The forward declaration is now guarded by SHARED_PROVIDER, matching the existing external-data-loader interface.

Documentation validation was failing because its Windows build failed, not because generated operator documentation was stale. Fresh CI is triggered by this push; no old runs were rerun.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The process-wide cuFile driver can still race during concurrent final release and reacquisition.

Review effort: Balanced
Findings: 1 High severity

Open (1)
Resolved since last review (8)
Previously missed (3)

In code that hasn't changed since last review

Medium severity Serialize driver lifetime with close and reopen operations

onnxruntime/​core/​providers/​cuda/​cuda_external_data_loader_gds.cc:81

Because active_driver is weak, releasing the last shared_ptr expires it before ~CuFileDriver acquires GlobalMutex(). A concurrent Acquire can take the mutex first, see no active owner, and call cuFileDriverOpen while the old driver is still open; that initialization fails and permanently sends the new loader to fallback. Keep ownership and close/reopen serialization in one process-wide lifetime manager (or retain the driver for process lifetime), and cover concurrent final release/acquire with a regression test.

Low severity Document pinned or pageable host-buffer fallback

docs/​CUDA_GPU_Direct_Storage.md:5

This fallback is not always pinned: when external_data_loader_reading_threads is 0, LoadTensor uses the pageable-buffer path. Describe this as the configured pinned or pageable host-buffer fallback so the introductory behavior matches the configuration section.

Low severity Correct host-memory loader option documentation

include/​onnxruntime/​core/​providers/​cuda/​cuda_provider_options.h:46

The fallback is pageable when external_data_loader_reading_threads is 0, so this public option comment is inaccurate for a valid configuration. Refer to the configured host-memory loader instead.

Encapsulate shared driver ownership in a handle that acquires the lifetime mutex before dropping its strong reference. Driver teardown therefore completes before another loader can configure a replacement. Keep the mutex state alive with each handle, preserve shared reuse and initialization-error cleanup, and cover final close versus reacquire without requiring native GDS.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The native cuFile read implementation lacks direct hardware-independent tests, and the fallback overview is inconsistent with pageable fallback behavior.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Resolved since last review (1)
Previously missed (1)

In code that hasn't changed since last review

Low severity Clarify fallback host memory can be pinned or pageable

docs/​CUDA_GPU_Direct_Storage.md:5

This says fallback is always pinned, but the documented and implemented external_data_loader_reading_threads=0 behavior uses pageable host memory (see line 53). Describe this as the configured pinned/pageable host-memory fallback so the overview does not contradict the configuration section.

Comment thread onnxruntime/core/providers/cuda/cuda_external_data_loader_gds.cc
Extract the existing native descriptor and chunked-read implementation behind cuFile and device-operation callbacks. Exercise that same implementation with host-only tests for multi-chunk reads, error decoding, copy/synchronization failures, and descriptor/handle cleanup. Wire coverage into regular and plugin CUDA internal tests.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Call CUDA allocation, stream creation, and GdsLoader::Create directly. Remove the injectable native GDS read API and restore the private native implementation. Replace injected loader tests with real-provider integration coverage while retaining shared-driver lifetime safety.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address the documentation note embedded in the Copilot review summary: fallback can use pinned or pageable memory, depending on the configured reading-thread count.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@xadupre

Copy link
Copy Markdown
Member Author

Fixed the remaining documentation note embedded under “Previously missed” in #32713 (review) in fdea334. Both the opening overview and the compatibility-mode paragraph now describe the configured host-memory fallback, including pageable memory when reading_threads=0.

The inline review threads are resolved. The earlier Copilot overview still contains its historical “Open (1)” heading; that does not reflect the current thread state. The injection-based testing recommendation was explicitly withdrawn at the author’s request in 3704b43, as documented in the thread, rather than reintroduced into the production API.

@xadupre

Copy link
Copy Markdown
Member Author

Addressed the provider-options coverage note from #32713 (review) in 0a3a7f9.

Added two tests using the existing public CUDA provider-options API, with no production hooks or replacement parser:

  • Omitted option defaults to 0; string updates 0 → 1 → 0 update the actual struct field. GetCUDAProviderOptionsAsString output is decoded and fed back into Update, verifying the GDS value and full serialized round trip.
  • Values 2, -1, invalid text, and an empty value are rejected with an appropriate diagnostic, without changing the previous valid setting.

Update exercises CUDAExecutionProviderInfo::FromProviderOptions through the real provider factory, and serialization exercises ToProviderOptions(OrtCUDAProviderOptionsV2). Rebuilt onnxruntime_shared_lib_test and ran both new CApiTest cases successfully for two repetitions. Targeted formatting also passed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Extending OrtCUDAProviderOptionsV2 violates the repository’s public C ABI guidance and can cause layout incompatibility.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity

Open (1)

Comment thread include/onnxruntime/core/providers/cuda/cuda_provider_options.h
@xadupre

Copy link
Copy Markdown
Member Author

Copilot resolve the merge conflicts in this pull request, also move the markdown documentation to the file docs/Model_Loading_Performance.md

…rnal-data

# Conflicts:
#	docs/FAQ.md

Co-authored-by: xadupre <22452781+xadupre@users.noreply.github.com>

Copilot AI commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Copilot resolve the merge conflicts in this pull request, also move the markdown documentation to the file docs/Model_Loading_Performance.md

Resolved in 65049bd. Merged the latest main, resolved the FAQ conflict, moved the GDS guidance into docs/Model_Loading_Performance.md, and removed the standalone GDS document.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 8af39836-82f0-478b-b1a0-338c518c1828
@xadupre
Xavier Dupré (xadupre) requested a balanced review from Copilot September 23, 2026 16:08

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The DirectStorage option silently lacks plugin-build support and its native configuration is not exercised by CI.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 2 Medium severity · 1 Low severity

Open (4)

Comment thread cmake/CMakeLists.txt
Comment thread include/onnxruntime/core/providers/cuda/cuda_provider_options.h
@xadupre Xavier Dupré (xadupre) changed the title [DRAFT] Add GPUDirect Storage model loading [DRAFT] Add CUDA weight loading with GPUDirect Storage and Microsoft DirectStorage Sep 23, 2026
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 8af39836-82f0-478b-b1a0-338c518c1828

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Zero-byte initializers can unnecessarily disable GDS for subsequent eligible weights, and one documentation path is incorrect.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
Resolved since last review (4)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Skip GDS registration for zero-length external initializers

onnxruntime/​core/​providers/​cuda/​cuda_external_data_loader.cc:245

Zero-byte external initializers satisfy the alignment test, so enabling GDS attempts driver/buffer/file registration even though there is nothing to read. A setup or file-registration failure on such an initializer then sets gds_disabled_ and prevents later non-empty aligned weights from using GDS. Skip the GDS attempt for zero-length tensors, as the DirectStorage branch already does.

Comment thread docs/CUDA_cuDNN_Optional_Design.md Outdated
Clarify the behavior of CUDAExecutionProviderInfo and its options. Introduce a helper to check cuDNN status in kernels.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Minimal CUDA configurations currently accept the DirectStorage build option while compiling out its backend.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Resolved since last review (1)

Comment on lines +27 to +31
if(onnxruntime_MINIMAL_BUILD OR onnxruntime_CUDA_MINIMAL)
list(REMOVE_ITEM onnxruntime_providers_cuda_cc_srcs
"${ONNXRUNTIME_ROOT}/core/providers/cuda/cuda_external_data_loader_directstorage.cc"
"${ONNXRUNTIME_ROOT}/core/providers/cuda/cuda_external_data_loader_gds.cc"
)
Declare BUILD_UNIT_TESTS before the dependent CUDA internal-test option so a fresh configure honors ENABLE_CUDA_EP_INTERNAL_TESTS=ON and writes the DirectStorage SDK manifest required by Windows CI.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants