Skip to content

Fix TensorRT I/O binding synchronization - #32860

Open
Silu Panda (SiluPanda) wants to merge 1 commit into
microsoft:mainfrom
SiluPanda:codex/onnx-32673-tensorrt-sync
Open

Silu Panda (SiluPanda) wants to merge 1 commit into
microsoft:mainfrom
SiluPanda:codex/onnx-32673-tensorrt-sync

Conversation

@SiluPanda

Copy link
Copy Markdown
Contributor

Description

Add the missing TensorRT execution-provider Sync() override used by I/O binding synchronization. It selects the provider device, waits for all preceding device work with cudaDeviceSynchronize(), and restores the caller's device. Device restoration is attempted even when synchronization fails; synchronization or restoration failures are propagated as ORT status errors.

A compute-stream-only wait would not cover the externally populated input stream from the issue. This does not add synchronization to ordinary Run or alter OnRunEnd, user-stream ownership, or graph capture/replay behavior. Explicit synchronization during an active capture reports CUDA's error instead of silently succeeding.

Add five direct provider-contract regression cases covering external nonblocking producer streams with default/user compute streams, conditional multi-GPU device selection/restoration, and active-capture error propagation. Worker CUDA results are asserted on the joining test thread. The intentional producer gate has timeout/release cleanup; the overall test still depends on a responsive CUDA driver.

Motivation and Context

Related to #32673. IOBinding::SynchronizeInputs() reaches the base no-op IExecutionProvider::Sync() for TensorRT. The reported asynchronous producer path can therefore expose inputs before device work finishes. The issue discussion identifies the missing override; CUDA EP already performs device-wide synchronization.

Fresh direct issue and TensorRT synchronization searches found no matching active implementation. This patch changes an existing internal virtual override, not a public API.

Validation

Draft: NVIDIA compilation and execution are still required. This was prepared on arm64 macOS without the CUDA/TensorRT toolchain or a NVIDIA GPU. Neither the changed TensorRT C++ nor the new tests has been compiled or executed locally. No unrelated CPU build is offered as validation of these files.

Passed local checks:

source .venv/bin/activate
lintrunner onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.h \
  onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.cc \
  onnxruntime/test/providers/tensorrt/tensorrt_sync_test.cc
clang-format --dry-run --Werror \
  onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.h \
  onnxruntime/core/providers/tensorrt/tensorrt_execution_provider.cc \
  onnxruntime/test/providers/tensorrt/tensorrt_sync_test.cc
git diff --check

Independent source review checked device restoration, status propagation, callback/resource lifetimes, and CMake test inclusion. This is not a substitute for compilation/runtime testing.

On a CUDA/TensorRT builder, regenerate CMake to discover the new test file, build onnxruntime_provider_test, and run from its build directory:

./onnxruntime_provider_test \
  --gtest_filter='*TensorrtExecutionProviderSyncTest.*:TensorrtExecutionProviderTest.SyncReportsActiveCaptureError'

Expect five discovered tests; the two multi-GPU cases explicitly skip with fewer than two CUDA devices. A zero-test run is not validation. Existing TensorRT I/O-binding/graph tests and the reporter's numerical reproduction should also pass before this is marked ready.

Prepared with OpenAI Codex assistance. No GPU correctness, full-model parity, or performance result is claimed.

Implement the existing Sync contract with a device-wide wait on the
provider device and restoration of the caller device. Add CUDA provider
regression coverage for external streams, device selection, and errors.

NVIDIA compilation and GPU execution are pending; local style and
source review only. This change is being published as a draft.

Co-authored-by: OpenAI Codex
Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com>
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The CUDA/TensorRT code and concurrency tests still require compilation and execution on NVIDIA hardware.

Review effort: Balanced
Findings: None

What changed in this PR

Adds TensorRT EP synchronization for I/O-bound CUDA buffers.

Changes:

  • Implements device-wide synchronization with device restoration.
  • Adds regression coverage for external streams, multi-GPU restoration, and graph capture errors.
  • No actionable defects identified in static review.
File Description
tensorrt_execution_provider.h Declares the Sync() override.
tensorrt_execution_provider.cc Implements synchronization and device restoration.
tensorrt_sync_test.cc Adds five provider synchronization tests.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@SiluPanda
Silu Panda (SiluPanda) marked this pull request as ready for review September 29, 2026 05:22

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants