Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
123044b
Add GPUDirect Storage model loading
xadupre Sep 21, 2026
606d6a2
Require native GPUDirect Storage support
xadupre Sep 21, 2026
de48496
Try PCI P2PDMA for native GDS
xadupre Sep 21, 2026
72ed26c
Address GDS loader review feedback
xadupre Sep 21, 2026
df468f7
Address GDS follow-up review comments
xadupre Sep 21, 2026
857b793
Fix GDS builds with older cuFile headers and MSVC
xadupre Sep 22, 2026
4c171d6
Serialize final GDS driver release before reacquisition
xadupre Sep 22, 2026
d78f750
Test native GDS reads with injectable I/O callbacks
xadupre Sep 22, 2026
3704b43
Remove test-only CUDA external loader injection APIs
xadupre Sep 22, 2026
fdea334
Clarify configured GDS host-memory fallback
xadupre Sep 22, 2026
9efeabc
Preserve close-on-exec on duplicated GDS descriptors
xadupre Sep 22, 2026
0a3a7f9
Cover GDS provider option parsing and serialization
xadupre Sep 22, 2026
65049bd
Merge remote-tracking branch 'origin/main' into feature/cuda-gds-exte…
Copilot Sep 23, 2026
990945a
Add Windows DirectStorage CUDA loading and three-path benchmarks
xadupre Sep 23, 2026
c40fe68
Address DirectStorage review feedback and enable Windows CI coverage
xadupre Sep 23, 2026
94341de
Update CUDA cuDNN design documentation
xadupre Sep 23, 2026
126728e
Exclude direct storage from minimal builds
xadupre Sep 25, 2026
d5be8fb
Merge remote-tracking branch 'origin/main' into feature/cuda-gds-exte…
xadupre Sep 25, 2026
7116be9
Evaluate CUDA internal-test option after its prerequisites
xadupre Sep 25, 2026
0af4cc4
Merge main into feature/cuda-gds-external-data
xadupre Sep 25, 2026
e2e1197
Merge remote-tracking branch 'origin/main' into feature/cuda-gds-exte…
xadupre Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 24 additions & 3 deletions .github/workflows/windows_cuda.yml
Original file line number Diff line number Diff line change
Expand Up @@ -116,13 +116,34 @@ jobs:
exit $lastExitCode
}
# Execute the build process
python.exe ${{ github.workspace }}\tools\ci_build\build.py --update --build --config RelWithDebInfo --build_dir build --skip_submodule_sync --build_csharp --parallel --nvcc_threads 4 --flash_nvcc_threads 4 --use_binskim_compliant_compile_flags --cmake_generator "Visual Studio 17 2022" --build_shared_lib --build_wheel --build_java --use_cuda --cuda_home="$env:RUNNER_TEMP\v12.8" --enable_cuda_profiling --use_vcpkg --use_vcpkg_ms_internal_asset_cache --enable_transformers_tool_test --cmake_extra_defines onnxruntime_QUICK_BUILD=ON --cmake_extra_defines CMAKE_CUDA_ARCHITECTURES=86 --cmake_extra_defines onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS=ON
python.exe ${{ github.workspace }}\tools\ci_build\build.py --update --build --config RelWithDebInfo --build_dir build --skip_submodule_sync --build_csharp --parallel --nvcc_threads 4 --flash_nvcc_threads 4 --use_binskim_compliant_compile_flags --cmake_generator "Visual Studio 17 2022" --build_shared_lib --build_wheel --build_java --use_cuda --cuda_home="$env:RUNNER_TEMP\v12.8" --enable_cuda_profiling --use_vcpkg --use_vcpkg_ms_internal_asset_cache --enable_transformers_tool_test --cmake_extra_defines onnxruntime_QUICK_BUILD=ON --cmake_extra_defines CMAKE_CUDA_ARCHITECTURES=86 --cmake_extra_defines onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS=ON onnxruntime_USE_CUDA_DIRECTSTORAGE=ON
if ($lastExitCode -ne 0) {
exit $lastExitCode
}

# Clean up the output directory before uploading artifacts
$outputDir = "${{ runner.temp }}\build\RelWithDebInfo"
# DirectStorage resolves dstoragecore.dll relative to the process executable.
$testDir = Join-Path $outputDir "RelWithDebInfo"
$sdkManifest = Join-Path $outputDir "directstorage-source-dir.txt"
if (!(Test-Path -LiteralPath $sdkManifest -PathType Leaf)) {
throw "Missing DirectStorage SDK source manifest: $sdkManifest"
}
if (!(Test-Path -LiteralPath (Join-Path $testDir "onnxruntime_provider_test.exe") -PathType Leaf)) {
throw "Missing CUDA provider test executable in $testDir"
}
$sdkDir = (Get-Content -LiteralPath $sdkManifest -Raw -ErrorAction Stop).Trim()
if ([string]::IsNullOrWhiteSpace($sdkDir)) {
throw "Empty DirectStorage SDK source manifest: $sdkManifest"
}
foreach ($dll in @("dstorage.dll", "dstoragecore.dll")) {
$source = Join-Path $sdkDir "native\bin\x64\$dll"
if (!(Test-Path -LiteralPath $source -PathType Leaf)) {
throw "Missing DirectStorage runtime: $source"
}
Copy-Item -LiteralPath $source -Destination $testDir -ErrorAction Stop
}

# Clean up the output directory before uploading artifacts
Write-Host "Cleaning up files from $outputDir..."

Remove-Item -Path "$outputDir\onnxruntime" -Recurse -Force -ErrorAction SilentlyContinue
Expand Down Expand Up @@ -243,7 +264,7 @@ jobs:
exit $lastExitCode
}

python.exe ${{ github.workspace }}\tools\ci_build\build.py --test --config RelWithDebInfo --build_dir build --skip_submodule_sync --build_csharp --parallel --nvcc_threads 4 --flash_nvcc_threads 4 --use_binskim_compliant_compile_flags --cmake_generator "Visual Studio 17 2022" --build_shared_lib --build_wheel --build_java --use_cuda --cuda_home="$env:RUNNER_TEMP\v12.8" --enable_cuda_profiling --use_vcpkg --use_vcpkg_ms_internal_asset_cache --enable_transformers_tool_test --cmake_extra_defines onnxruntime_QUICK_BUILD=ON --cmake_extra_defines CMAKE_CUDA_ARCHITECTURES=86 --cmake_extra_defines onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS=ON
python.exe ${{ github.workspace }}\tools\ci_build\build.py --test --config RelWithDebInfo --build_dir build --skip_submodule_sync --build_csharp --parallel --nvcc_threads 4 --flash_nvcc_threads 4 --use_binskim_compliant_compile_flags --cmake_generator "Visual Studio 17 2022" --build_shared_lib --build_wheel --build_java --use_cuda --cuda_home="$env:RUNNER_TEMP\v12.8" --enable_cuda_profiling --use_vcpkg --use_vcpkg_ms_internal_asset_cache --enable_transformers_tool_test --cmake_extra_defines onnxruntime_QUICK_BUILD=ON --cmake_extra_defines CMAKE_CUDA_ARCHITECTURES=86 --cmake_extra_defines onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS=ON onnxruntime_USE_CUDA_DIRECTSTORAGE=ON
if ($lastExitCode -ne 0) {
exit $lastExitCode
}
Expand Down
33 changes: 28 additions & 5 deletions cmake/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -70,13 +70,12 @@ option(onnxruntime_ENABLE_PYTHON "Enable python bindings" OFF)
option(onnxruntime_ENABLE_MEMLEAK_CHECKER "Experimental: Enable memory leak checker in Windows debug build" OFF)
option(onnxruntime_ENABLE_CONVSYMKERNELAVX2_SAT_CHECKER "Experimental: Enable ConvSymKernelAvx2 assembly saturation checker in build" OFF)
option(onnxruntime_USE_CUDA "Build with CUDA support" OFF)
# Enable ONNX Runtime CUDA EP's internal unit tests that directly access the EP's internal functions instead of through
# OpKernels. When the option is ON, we will have two copies of GTest library in the same process. It is not a typical
# use. If you hit any problem with that, please do not report it to GTest. Turn OFF the following build option instead.
cmake_dependent_option(onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS "Build with CUDA unit tests" OFF "onnxruntime_USE_CUDA;onnxruntime_BUILD_UNIT_TESTS" OFF)

cmake_dependent_option(onnxruntime_USE_CUDA_DIRECTSTORAGE "Build Microsoft DirectStorage CUDA loading support" OFF "onnxruntime_USE_CUDA;WIN32" OFF)
Comment thread
xadupre marked this conversation as resolved.
cmake_dependent_option(onnxruntime_USE_CUDA_NHWC_OPS "Build CUDA with NHWC op support" ON "onnxruntime_USE_CUDA" OFF)
cmake_dependent_option(onnxruntime_BUILD_CUDA_EP_AS_PLUGIN "Build CUDA EP as a separate plugin shared library instead of the legacy in-tree provider" OFF "onnxruntime_USE_CUDA" OFF)
if(onnxruntime_USE_CUDA_DIRECTSTORAGE AND onnxruntime_BUILD_CUDA_EP_AS_PLUGIN)
message(FATAL_ERROR "onnxruntime_USE_CUDA_DIRECTSTORAGE is not supported with onnxruntime_BUILD_CUDA_EP_AS_PLUGIN.")
endif()
option(onnxruntime_BUILD_CUDA_QUANT_PREPROCESS "Build CUDA weight-packing module onnxruntime_cuda_quant_preprocess.so" OFF)
option(onnxruntime_CUDA_MINIMAL "Build CUDA without any operations apart from memcpy ops. Useful for a very minimal TRT build" OFF)
option(onnxruntime_ENABLE_CUDA_LINE_NUMBER_INFO "When building with CUDA support, generate device code line number information." OFF)
Expand All @@ -97,6 +96,12 @@ option(onnxruntime_USE_ARM_NEON_NCHWC "Build with ARM Neon NCHWc kernels in MLAS
option(onnxruntime_USE_KLEIDIAI "Build with KleidiAI integration in MLAS" OFF)
option(onnxruntime_USE_QMX_KLEIDIAI_COEXIST "Build with QMX and Arm KLEIDIAI libraries" OFF)
option(onnxruntime_BUILD_UNIT_TESTS "Build ONNXRuntime unit tests" ON)
# Declare both prerequisites before evaluating this dependent option on a fresh configure.
# Enable ONNX Runtime CUDA EP's internal unit tests that directly access the EP's internal functions instead of through
# OpKernels. When the option is ON, we will have two copies of GTest library in the same process. It is not a typical
# use. If you hit any problem with that, please do not report it to GTest. Turn OFF the following build option instead.
cmake_dependent_option(onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS "Build with CUDA unit tests" OFF "onnxruntime_USE_CUDA;onnxruntime_BUILD_UNIT_TESTS" OFF)

# Materialize the ONNX node-test corpus from ONNX's Python generators into the build tree at
# configure/build time, instead of depending on the on-disk corpus shipped in the ONNX source
# archive. This detaches ORT from ONNX PR #7959 (which deletes onnx/backend/test/data/node).
Expand Down Expand Up @@ -1545,6 +1550,24 @@ if (onnxruntime_USE_CUDA)
endif()
find_package(CUDAToolkit REQUIRED)

if(CMAKE_SYSTEM_NAME STREQUAL "Linux" AND NOT onnxruntime_MINIMAL_BUILD AND NOT onnxruntime_CUDA_MINIMAL)
include(CheckCXXSourceCompiles)
include(CMakePushCheckState)
cmake_push_check_state(RESET)
set(CMAKE_REQUIRED_INCLUDES ${CUDAToolkit_INCLUDE_DIRS})
check_cxx_source_compiles("
#include <cufile.h>
using SetBoolParameterFn = decltype(&cuFileSetParameterBool);
int main() {
return CUFILE_PARAM_USE_PCIP2PDMA == CUFILE_PARAM_PROPERTIES_ALLOW_COMPAT_MODE;
}" onnxruntime_CUFILE_CONFIG_API_SUPPORTED)
cmake_pop_check_state()
if(onnxruntime_CUFILE_CONFIG_API_SUPPORTED)
set_property(SOURCE "${ONNXRUNTIME_ROOT}/core/providers/cuda/cuda_external_data_loader_gds.cc"
APPEND PROPERTY COMPILE_DEFINITIONS ORT_CUDA_GDS_AVAILABLE)
endif()
endif()

if(MSVC AND CMAKE_CUDA_COMPILER_VERSION VERSION_GREATER_EQUAL 12.9
AND CMAKE_CUDA_COMPILER_VERSION VERSION_LESS 13.0)
foreach(_cuda_include_dir IN LISTS CUDAToolkit_INCLUDE_DIRS)
Expand Down
1 change: 1 addition & 0 deletions cmake/deps.txt
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,7 @@ cutlass;https://github.com/NVIDIA/cutlass/archive/refs/tags/v4.7.0.zip;51d4f1ba4
deep_gemm;https://github.com/deepseek-ai/DeepGEMM/archive/559d79fb6994a58b8a15b4b93bf13ccc16edf247.tar.gz;76a0076386991cac8e5d32c3e7e74d9bb8102115
extensions;https://github.com/microsoft/onnxruntime-extensions/archive/c24b7bab0c12f53da76d0c31b03b9f0f8ec8f3b4.zip;239063aee4946a9af147b473a4c3da78ba7413b4
directx_headers;https://github.com/microsoft/DirectX-Headers/archive/refs/tags/v1.613.1.zip;47653509a3371eabb156360f42faf582f314bf2e
directstorage;https://www.nuget.org/api/v2/package/Microsoft.Direct3D.DirectStorage/1.2.3;be08f099a75c54997a753f444224af96ed0ee3d9
cudnn_frontend;https://github.com/NVIDIA/cudnn-frontend/archive/refs/tags/v1.27.0.zip;1e4c9a464d3437e388ab0163f3be068dba783c08
dawn;https://github.com/google/dawn/archive/refs/tags/v20260916.214021.zip;514e914f23a0787213c89eed879f74b5e275690d
dawn_agility_sdk;https://www.nuget.org/api/v2/package/Microsoft.Direct3D.D3D12/1.721.3-preview;fa5f5fc8d0c8c209cfbb530be634960d20595841
Expand Down
16 changes: 16 additions & 0 deletions cmake/external/directstorage.cmake
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Copyright (c) Microsoft Corporation. All rights reserved.
# Licensed under the MIT License.

include_guard(GLOBAL)
onnxruntime_fetchcontent_declare(
directstorage
URL ${DEP_URL_directstorage}
URL_HASH SHA1=${DEP_SHA1_directstorage}
DOWNLOAD_NAME directstorage.zip
)
onnxruntime_fetchcontent_makeavailable(directstorage)

if(onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS)
# CI deploys the SDK runtime only beside test executables, including when the SDK source is overridden.
file(WRITE "${CMAKE_BINARY_DIR}/directstorage-source-dir.txt" "${directstorage_SOURCE_DIR}\n")
endif()
17 changes: 17 additions & 0 deletions cmake/onnxruntime_providers_cuda.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,12 @@
endif()
# Exclude plugin directory if it was picked up by GLOB_RECURSE
list(FILTER onnxruntime_providers_cuda_cc_srcs EXCLUDE REGEX "core/providers/cuda/plugin/.*")
if(onnxruntime_MINIMAL_BUILD OR onnxruntime_CUDA_MINIMAL)
list(REMOVE_ITEM onnxruntime_providers_cuda_cc_srcs
"${ONNXRUNTIME_ROOT}/core/providers/cuda/cuda_external_data_loader_directstorage.cc"
"${ONNXRUNTIME_ROOT}/core/providers/cuda/cuda_external_data_loader_gds.cc"
)
Comment on lines +27 to +31
endif()

# Remove pch files
list(REMOVE_ITEM onnxruntime_providers_cuda_cc_srcs
Expand Down Expand Up @@ -236,7 +242,18 @@

# config_cuda_provider_shared_module can be used to config onnxruntime_providers_cuda_obj, onnxruntime_providers_cuda & onnxruntime_providers_cuda_ut.
# This function guarantees that all 3 targets have the same configurations.
if(onnxruntime_USE_CUDA_DIRECTSTORAGE AND
NOT onnxruntime_MINIMAL_BUILD AND NOT onnxruntime_CUDA_MINIMAL)
include(external/directstorage.cmake)
endif()

function(config_cuda_provider_shared_module target)
if(onnxruntime_USE_CUDA_DIRECTSTORAGE AND
NOT onnxruntime_MINIMAL_BUILD AND NOT onnxruntime_CUDA_MINIMAL)
target_compile_definitions(${target} PRIVATE ORT_CUDA_DIRECTSTORAGE_AVAILABLE)
target_include_directories(${target} PRIVATE "${directstorage_SOURCE_DIR}/native/include")
target_link_libraries(${target} PRIVATE d3d12 dxgi)
endif()
if (onnxruntime_REDUCED_OPS_BUILD)
add_op_reduction_include_dirs(${target})
endif()
Expand Down
1 change: 1 addition & 0 deletions cmake/onnxruntime_unittests.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -1075,6 +1075,7 @@ if (onnxruntime_ENABLE_CUDA_EP_INTERNAL_TESTS AND onnxruntime_BUILD_CUDA_EP_AS_P
NOT onnxruntime_MINIMAL_BUILD AND NOT onnxruntime_REDUCED_OPS_BUILD)
set(onnxruntime_test_providers_cuda_plugin_internal_test_src
"${TEST_SRC_DIR}/providers/cuda/test_cases/allocator_cuda_test.cc"
"${TEST_SRC_DIR}/providers/cuda/test_cases/cuda_external_data_loader_gds_test.cc"
"${TEST_SRC_DIR}/providers/cuda/test_cases/cuda_utils_test.cc"
"${TEST_SRC_DIR}/providers/cuda/test_cases/group_query_attention_workspace_header_test.cc"
"${TEST_SRC_DIR}/providers/cuda/test_cases/packed_attention_workspace_header_test.cc"
Expand Down
7 changes: 4 additions & 3 deletions docs/CUDA_cuDNN_Optional_Design.md
Original file line number Diff line number Diff line change
Expand Up @@ -292,9 +292,10 @@ Implementation details:
`CUDAExecutionProviderInfo::FromProviderOptions(...)`.
- Emit it from `CUDAExecutionProviderInfo::ToProviderOptions(...)`.
- Include it in `std::hash<CUDAExecutionProviderInfo>` because it changes the EP behavior.
- Do **not** add a field to `OrtCUDAProviderOptionsV2` for Phase 1. That struct is public C
ABI surface; string-key provider options are sufficient and can be set through existing
provider-options APIs.
- Phase 1 keeps this policy in `CUDAExecutionProviderInfo`. `OrtCUDAProviderOptionsV2` is opaque in the
public C API: callers obtain it through `CreateCUDAProviderOptions` and configure it through string keys.
Its definition in `include/onnxruntime/core/providers/cuda/cuda_provider_options.h` is internal and may be extended for new
options. This differs from the publicly defined `OrtCUDAProviderOptions`, whose layout must remain stable.
- Add an EP helper such as `CUDAExecutionProvider::IsCudnnEnabled()` or
`CudaKernel::IsCudnnEnabled()` so kernels can distinguish:
- cuDNN disabled by user (`enable_cudnn=0`), and
Expand Down
3 changes: 2 additions & 1 deletion docs/FAQ.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,8 @@ The default CUDA build supports 3 standard quantization operators: QuantizeLinea
## How can I reduce model loading time for large CPU or CUDA models?

See [Accelerate model loading](Model_Loading_Performance.md) for parallel CPU weight prepacking and CUDA external-data
loading through pinned host buffers, including the session and execution provider options that control them.
loading through GPUDirect Storage, Microsoft DirectStorage, or pinned host buffers, including the session and execution provider options that
control them.

## How do I change the severity level of the default logger to something other than the default (WARNING)?
Setting the severity level to VERBOSE is most useful when debugging errors.
Expand Down
Loading
Loading