Fix TensorRT I/O binding synchronization - #32860
Open
Silu Panda (SiluPanda) wants to merge 1 commit into
Open
Silu Panda (SiluPanda) wants to merge 1 commit into
Silu Panda (SiluPanda) wants to merge 1 commit into
Conversation
Implement the existing Sync contract with a device-wide wait on the provider device and restoration of the caller device. Add CUDA provider regression coverage for external streams, device selection, and errors. NVIDIA compilation and GPU execution are pending; local style and source review only. This change is being published as a draft. Co-authored-by: OpenAI Codex Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com>
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The CUDA/TensorRT code and concurrency tests still require compilation and execution on NVIDIA hardware.
Review effort: Balanced
Findings: None
What changed in this PR
Adds TensorRT EP synchronization for I/O-bound CUDA buffers.
Changes:
- Implements device-wide synchronization with device restoration.
- Adds regression coverage for external streams, multi-GPU restoration, and graph capture errors.
- No actionable defects identified in static review.
| File | Description |
|---|---|
tensorrt_execution_provider.h |
Declares the Sync() override. |
tensorrt_execution_provider.cc |
Implements synchronization and device restoration. |
tensorrt_sync_test.cc |
Adds five provider synchronization tests. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Silu Panda (SiluPanda)
marked this pull request as ready for review
September 29, 2026 05:22
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Add the missing TensorRT execution-provider
Sync()override used by I/O binding synchronization. It selects the provider device, waits for all preceding device work withcudaDeviceSynchronize(), and restores the caller's device. Device restoration is attempted even when synchronization fails; synchronization or restoration failures are propagated as ORT status errors.A compute-stream-only wait would not cover the externally populated input stream from the issue. This does not add synchronization to ordinary
Runor alterOnRunEnd, user-stream ownership, or graph capture/replay behavior. Explicit synchronization during an active capture reports CUDA's error instead of silently succeeding.Add five direct provider-contract regression cases covering external nonblocking producer streams with default/user compute streams, conditional multi-GPU device selection/restoration, and active-capture error propagation. Worker CUDA results are asserted on the joining test thread. The intentional producer gate has timeout/release cleanup; the overall test still depends on a responsive CUDA driver.
Motivation and Context
Related to #32673.
IOBinding::SynchronizeInputs()reaches the base no-opIExecutionProvider::Sync()for TensorRT. The reported asynchronous producer path can therefore expose inputs before device work finishes. The issue discussion identifies the missing override; CUDA EP already performs device-wide synchronization.Fresh direct issue and TensorRT synchronization searches found no matching active implementation. This patch changes an existing internal virtual override, not a public API.
Validation
Draft: NVIDIA compilation and execution are still required. This was prepared on arm64 macOS without the CUDA/TensorRT toolchain or a NVIDIA GPU. Neither the changed TensorRT C++ nor the new tests has been compiled or executed locally. No unrelated CPU build is offered as validation of these files.
Passed local checks:
source .venv/bin/activate lintrunner onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.h \ onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.cc \ onnxruntime/test/providers/tensorrt/tensorrt_sync_test.cc clang-format --dry-run --Werror \ onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.h \ onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.cc \ onnxruntime/test/providers/tensorrt/tensorrt_sync_test.cc git diff --checkIndependent source review checked device restoration, status propagation, callback/resource lifetimes, and CMake test inclusion. This is not a substitute for compilation/runtime testing.
On a CUDA/TensorRT builder, regenerate CMake to discover the new test file, build
onnxruntime_provider_test, and run from its build directory:./onnxruntime_provider_test \ --gtest_filter='*TensorrtExecutionProviderSyncTest.*:TensorrtExecutionProviderTest.SyncReportsActiveCaptureError'Expect five discovered tests; the two multi-GPU cases explicitly skip with fewer than two CUDA devices. A zero-test run is not validation. Existing TensorRT I/O-binding/graph tests and the reporter's numerical reproduction should also pass before this is marked ready.
Prepared with OpenAI Codex assistance. No GPU correctness, full-model parity, or performance result is claimed.