You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Introduce explicit WebGPU MatMul algorithm enums for subgroup matrix, naive, Intel subgroup, packed, and packed Split-K implementations.
Add a MatMul scheduler with this priority: forced test algorithm, zero-K correctness guard, vendor policy, then a private common fallback.
Allow vendor schedulers to choose any algorithm with vendor-specific thresholds and independently tune the selected algorithm's execution configuration.
Represent algorithm and typed configuration together in a MatMulExecutionPlan. Packed tuning includes workgroup size, elements per thread, inner tile size, and Split-K size.
Move the existing Intel Split-K architecture profiles and measured threshold tables under vendor/intel while keeping eligibility evaluation in the generic SplitKConfig.
Route adapter-specific Split-K profiles through a generic factory owned by WebGpuContext, so GEMM and MatMul share the same immutable policy.
Keep forced-algorithm tests on the real vendor tuning path through the test-only ep.webgpuexecutionprovider.forceMatmulAlgorithm session option.
Route MatMul, pointwise Conv, and contrib Attention through MatMulComputeDispatcher, dispatch directly from the selected enum, and validate algorithm prerequisites and packed tuning before execution.
Document the architecture in docs/design/webgpu_matmul_algorithm_scheduler.md.
Motivation and Context
WebGPU MatMul chooses among several implementations using shape, capability, and vendor-specific conditions. Keeping these choices as distributed conditionals made individual implementations difficult to force in tests and made vendor thresholds difficult to evolve.
This refactor separates immutable problem facts, algorithm policy, execution tuning, and dispatch. A vendor can override performance boundaries, supply a Split-K profile, or tune execution parameters without changing MatMulComputeDispatcher, while returning no selection or configuration delegates to the common defaults.
Behavior and Compatibility
The common scheduler retains the existing subgroup-matrix, small-matrix naive, Split-K, and packed rules as its fallback.
Intel preserves its original priority: subgroup matrix, Intel subgroup, then common fallback.
Intel's existing legacy, discrete/Lunar Lake, Xe3, and default Split-K boundaries are unchanged.
Automatic selection always uses the Naive path when K is zero before consulting vendor policy. Forced algorithms still take precedence, and the dispatcher rejects forced algorithms whose hard prerequisites exclude zero-K inputs.
Existing subgroup-matrix vendor tiling remains in its dedicated selector.
The PR description still claims this change adds Vulkan-only Windows Dawn/system-Vulkan build support, but this diff changes no build or dependency files; this line only records a validation configuration, and forcing dawnBackendType=Vulkan requires that backend to already be built. Remove that feature claim from the PR description (or restore the corresponding build implementation) so the stated scope matches the changes.
In ORT_USE_EP_API_ADAPTERS builds, WebGpuExecutionProviderWithOptions returns nullptr before parsing options when the dynamic plugin is uninitialized (test/util/default_providers.cc:379-388). Consequently this EXPECT_THROW observes no exception and fails in configurations where WebGPU is legitimately unavailable. Probe provider availability with a valid configuration and GTEST_SKIP() before asserting that the invalid option throws.
The reason will be displayed to describe this comment to others. Learn more.
Review frame
Problem/feature validity: Validated. Before this PR, ComputeMatMul selected subgroup-matrix, naive, Intel subgroup, Split-K, and packed implementations through distributed conditionals, which prevented a test from selecting one implementation independently and coupled vendor thresholds to dispatch. The diff centralizes that policy while preserving the existing automatic order and keeps execution-specific validation at the dispatch boundary.
Risk/scope: Deep. This changes algorithm selection for MatMul and the pointwise-Conv caller, exposes a test-only provider option, changes subgroup-matrix applicability/dispatch, and threads packed tuning into shader generation and cache keys across Metal, D3D12, and Vulkan.
Direction gate: Pass. As the owner solution, I would keep immutable problem facts, vendor policy, hard prerequisites, typed execution configuration, and final dispatch separate. The PR now follows that boundary: vendor policy can decline to the private common fallback, forced selection still uses real vendor tuning, and unsupported forced paths fail rather than silently changing algorithms.
Confirmed findings
T1: Skip the invalid-option test when the plugin EP is unavailable
This test assumes WebGpuExecutionProviderWithOptions() reaches WebGPU option parsing. In an ORT_USE_EP_API_ADAPTERS Windows build where the dynamic plugin has not been initialized, however, onnxruntime/test/util/default_providers.cc:366-389 returns nullptr before forwarding any configuration to the plugin. EXPECT_THROW therefore observes a normal null return and fails even though WebGPU is legitimately unavailable. This PR introduces that failure in a supported build configuration, and the plugin build check does not execute this provider test.
Please first construct the provider with a valid configuration and GTEST_SKIP() when that returns null, then run the invalid-option assertion. That preserves parser coverage when the plugin is present without turning provider unavailability into a test failure.
Q1: Disable copy and move on the new polymorphic scheduler
MatMulAlgorithmScheduler is a new polymorphic base class but remains implicitly copyable and movable. Repository guidance requires new classes to use ORT_DISALLOW_COPY_ASSIGNMENT_AND_MOVE until those operations are shown to be necessary. Applying the macro to this base class also covers the Intel and test subclasses and prevents accidental value copies or slicing as more vendor schedulers are added.
Clarifications
None.
Test coverage
The device-independent tests cover parser behavior, selection precedence, common thresholds, typed packed configuration, Split-K prerequisites, overflow-safe dispatch arithmetic, and Intel low-M vec4 eligibility. Hardware tests compare deterministic results with a CPU reference and gate Intel subgroup, Split-K, and subgroup-matrix success cases on the selected adapter's actual capabilities. The automatic zero-K execution test also exercises the new no-input-binding shader path.
At the reviewed head, the macOS-arm64 WebGPU Release legs passed and the Linux WebGPU build passed; several unrelated or additional matrix jobs, including macOS WebGPU Debug and Windows plugin/build variants, were still pending. Linux WebGPU remains build-only, so it is not execution evidence. T1 remains uncovered because the plugin build leg does not run this test.
Verdict
The feature is valid and the scheduler/typed-plan direction is appropriate. T1 blocks merge because the new test fails in a supported plugin configuration. There are no remaining clarification requests. Q1 is mechanical cleanup required by the repository's new-class convention and does not change runtime behavior.
The reason will be displayed to describe this comment to others. Learn more.
What about making the scheduler not return a valid plan if configuration is not compatible or prerequisites are not met?
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
MatMulComputeDispatcher, dispatch directly from the selected enum, and validate algorithm prerequisites and packed tuning before execution.Motivation and Context
WebGPU MatMul chooses among several implementations using shape, capability, and vendor-specific conditions. Keeping these choices as distributed conditionals made individual implementations difficult to force in tests and made vendor thresholds difficult to evolve.
This refactor separates immutable problem facts, algorithm policy, execution tuning, and dispatch. A vendor can override performance boundaries, supply a Split-K profile, or tune execution parameters without changing
MatMulComputeDispatcher, while returning no selection or configuration delegates to the common defaults.Behavior and Compatibility
Validation
onnxruntime_provider_testtargets on Windows.