Skip to content

[WebGPU] Refactor MatMul algorithm selection - #32630

Open
Jiajia Qin (qjia7) wants to merge 25 commits into
mainfrom
codex/webgpu-matmul-algorithm-scheduler
Open

Jiajia Qin (qjia7) wants to merge 25 commits into
mainfrom
codex/webgpu-matmul-algorithm-scheduler

Conversation

@qjia7

@qjia7 Jiajia Qin (qjia7) commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Description

  • Introduce explicit WebGPU MatMul algorithm enums for subgroup matrix, naive, Intel subgroup, packed, and packed Split-K implementations.
  • Add a MatMul scheduler with this priority: forced test algorithm, zero-K correctness guard, vendor policy, then a private common fallback.
  • Allow vendor schedulers to choose any algorithm with vendor-specific thresholds and independently tune the selected algorithm's execution configuration.
  • Represent algorithm and typed configuration together in a MatMulExecutionPlan. Packed tuning includes workgroup size, elements per thread, inner tile size, and Split-K size.
  • Move the existing Intel Split-K architecture profiles and measured threshold tables under vendor/intel while keeping eligibility evaluation in the generic SplitKConfig.
  • Route adapter-specific Split-K profiles through a generic factory owned by WebGpuContext, so GEMM and MatMul share the same immutable policy.
  • Keep forced-algorithm tests on the real vendor tuning path through the test-only ep.webgpuexecutionprovider.forceMatmulAlgorithm session option.
  • Route MatMul, pointwise Conv, and contrib Attention through MatMulComputeDispatcher, dispatch directly from the selected enum, and validate algorithm prerequisites and packed tuning before execution.
  • Document the architecture in docs/design/webgpu_matmul_algorithm_scheduler.md.

Motivation and Context

WebGPU MatMul chooses among several implementations using shape, capability, and vendor-specific conditions. Keeping these choices as distributed conditionals made individual implementations difficult to force in tests and made vendor thresholds difficult to evolve.

This refactor separates immutable problem facts, algorithm policy, execution tuning, and dispatch. A vendor can override performance boundaries, supply a Split-K profile, or tune execution parameters without changing MatMulComputeDispatcher, while returning no selection or configuration delegates to the common defaults.

Behavior and Compatibility

  • The common scheduler retains the existing subgroup-matrix, small-matrix naive, Split-K, and packed rules as its fallback.
  • Intel preserves its original priority: subgroup matrix, Intel subgroup, then common fallback.
  • Intel's existing legacy, discrete/Lunar Lake, Xe3, and default Split-K boundaries are unchanged.
  • Automatic selection always uses the Naive path when K is zero before consulting vendor policy. Forced algorithms still take precedence, and the dispatcher rejects forced algorithms whose hard prerequisites exclude zero-K inputs.
  • Existing subgroup-matrix vendor tiling remains in its dedicated selector.
  • NVIDIA Split-K policy from [WebGPU] Enable Split-K on reported-Pascal NVIDIA adapters #32333 is intentionally not included; this PR only provides the extension point it can use.
  • The PR does not change WebGPU build configuration or dependency configuration.

Validation

  • Built the static and plugin-adapter onnxruntime_provider_test targets on Windows.
  • All 37 targeted scheduler, Split-K profile, prerequisite, forced-algorithm, zero-K, and MatMul execution tests passed with the Vulkan WebGPU build.
  • Forced FP16 subgroup-matrix execution passed on Intel Arc 140V with Vulkan cooperative-matrix support.
  • Hardware-backed MatMul tests use the build's default WebGPU backend and inspect the selected adapter before running capability-specific algorithms.
  • Added regression coverage for vendor-first selection, vendor tuning of forced algorithms, independent Split-K thresholds, Intel profile routing and boundaries, packed Z-axis invariants, Split-K tile alignment, overflow-safe dispatch arithmetic, and per-invocation reselection for dynamic shapes.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Intel eligibility misses a shape constraint, and hardware-specific tests do not skip unsupported adapters.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Centralizes WebGPU MatMul algorithm selection and enables explicit path testing.

Changes:

  • Adds enum-based common and Intel-specific scheduling.
  • Adds forced-algorithm configuration and coverage.
  • Supports Vulkan-only Dawn builds on Windows.
File summaries
File Description
onnxruntime/test/providers/webgpu/matmul_large_test.cc Adds forced-path integration tests.
onnxruntime/test/providers/webgpu/matmul_algorithm_scheduler_test.cc Tests parsing, scheduling, and prerequisites.
onnxruntime/core/providers/webgpu/webgpu_provider_options.h Declares the forcing option.
onnxruntime/core/providers/webgpu/webgpu_provider_factory.cc Parses and validates the option.
onnxruntime/core/providers/webgpu/webgpu_execution_provider.h Stores forced selection state.
onnxruntime/core/providers/webgpu/webgpu_execution_provider.cc Initializes forced selection state.
onnxruntime/core/providers/webgpu/vendor/intel/math/matmul.h Exposes Intel capability detection.
onnxruntime/core/providers/webgpu/vendor/intel/math/matmul.cc Separates capability from heuristics.
onnxruntime/core/providers/webgpu/vendor/intel/math/matmul_algorithm_scheduler.h Implements Intel scheduling policy.
onnxruntime/core/providers/webgpu/math/subgroup_matrix_matmul.cc Separates applicability from execution.
onnxruntime/core/providers/webgpu/math/matmul.h Extends implementation and scheduler caches.
onnxruntime/core/providers/webgpu/math/matmul.cc Implements scheduling and direct dispatch.
onnxruntime/core/providers/webgpu/math/matmul_algorithm.h Defines algorithm identities and conversion.
onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h Defines common scheduling and prerequisites.
onnxruntime/core/providers/webgpu/compute_context.h Exposes forced selection to kernels.
docs/superpowers/specs/2026-09-16-webgpu-matmul-algorithm-scheduler-design.md Documents the design.
docs/superpowers/plans/2026-09-16-webgpu-matmul-algorithm-scheduler.md Records the implementation plan.
cmake/external/onnxruntime_external_deps.cmake Configures Vulkan-only Windows Dawn builds.
Review details
  • Files reviewed: 15/18 changed files
  • Comments generated: 1
  • Review effort level: Balanced (auto)

Note

Copilot is running an experiment and ran this review at Balanced.


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread onnxruntime/test/providers/webgpu/matmul_large_test.cc

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Forced Intel dispatch can accept small shapes that violate the subgroup kernel’s vec4-load invariant.

Get a fresh assessment by requesting another Copilot review.

Review details
  • Files reviewed: 15/18 changed files
  • Comments generated: 2
  • Review effort level: Balanced (auto)

Note

Copilot is running an experiment and ran this review at Balanced.

Comment thread onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h Outdated
Comment thread cmake/external/onnxruntime_external_deps.cmake

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Forced Packed zero-K inputs and low-M Intel execution expose unsupported shader paths.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 1 Low severity

Open (3)

Comment thread onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Forced Intel subgroup dispatch mishandles zero-K inputs, and the new zero-K test invokes undefined pointer arithmetic.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity

Open (2)
Resolved since last review (3)
Previously missed (1)

In code that hasn't changed since last review

Low severity Align PR description with implemented Vulkan build changes

docs/​design/​webgpu_matmul_algorithm_scheduler.md:62

The PR description still claims this change adds Vulkan-only Windows Dawn/system-Vulkan build support, but this diff changes no build or dependency files; this line only records a validation configuration, and forcing dawnBackendType=Vulkan requires that backend to already be built. Remove that feature claim from the PR description (or restore the corresponding build implementation) so the stated scope matches the changes.

Comment thread onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h Outdated
Comment thread onnxruntime/test/providers/webgpu/matmul_large_test.cc

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The subgroup-matrix capability probe can run an expected-success test when the runtime implementation remains unavailable.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)
Resolved since last review (2)

Comment thread onnxruntime/test/providers/webgpu/matmul_large_test.cc
Comment thread docs/design/webgpu_matmul_algorithm_scheduler.md Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Zero-contraction MatMul dispatch can bind null buffers, and one new test fails when the plugin EP is unavailable.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity

Open (1)
Resolved since last review (2)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Skip WebGPU invalid-option test when provider is unavailable

onnxruntime/​test/​providers/​webgpu/​matmul_large_test.cc:233

In ORT_USE_EP_API_ADAPTERS builds, WebGpuExecutionProviderWithOptions returns nullptr before parsing options when the dynamic plugin is uninitialized (test/util/default_providers.cc:379-388). Consequently this EXPECT_THROW observes no exception and fails in configurations where WebGPU is legitimately unavailable. Probe provider availability with a valid configuration and GTEST_SKIP() before asserting that the invalid option throws.

Comment thread onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Hardware-specific GPU dispatch and capability handling merit final maintainer review despite comprehensive focused coverage.

Review effort: Balanced
Findings: None

Resolved since last review (1)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The hardware-dependent dispatch refactor spans multiple kernels and requires final maintainer validation across supported GPU backends.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

Comment thread onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h

@qjia7 Jiajia Qin (qjia7) left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review frame

  • Problem/feature validity: Validated. Before this PR, ComputeMatMul selected subgroup-matrix, naive, Intel subgroup, Split-K, and packed implementations through distributed conditionals, which prevented a test from selecting one implementation independently and coupled vendor thresholds to dispatch. The diff centralizes that policy while preserving the existing automatic order and keeps execution-specific validation at the dispatch boundary.
  • Risk/scope: Deep. This changes algorithm selection for MatMul and the pointwise-Conv caller, exposes a test-only provider option, changes subgroup-matrix applicability/dispatch, and threads packed tuning into shader generation and cache keys across Metal, D3D12, and Vulkan.
  • Direction gate: Pass. As the owner solution, I would keep immutable problem facts, vendor policy, hard prerequisites, typed execution configuration, and final dispatch separate. The PR now follows that boundary: vendor policy can decline to the private common fallback, forced selection still uses real vendor tuning, and unsupported forced paths fail rather than silently changing algorithms.

Confirmed findings

T1: Skip the invalid-option test when the plugin EP is unavailable

onnxruntime/test/providers/webgpu/matmul_large_test.cc:233-236

This test assumes WebGpuExecutionProviderWithOptions() reaches WebGPU option parsing. In an ORT_USE_EP_API_ADAPTERS Windows build where the dynamic plugin has not been initialized, however, onnxruntime/test/util/default_providers.cc:366-389 returns nullptr before forwarding any configuration to the plugin. EXPECT_THROW therefore observes a normal null return and fails even though WebGPU is legitimately unavailable. This PR introduces that failure in a supported build configuration, and the plugin build check does not execute this provider test.

Please first construct the provider with a valid configuration and GTEST_SKIP() when that returns null, then run the invalid-option assertion. That preserves parser coverage when the plugin is present without turning provider unavailability into a test failure.

Q1: Disable copy and move on the new polymorphic scheduler

onnxruntime/core/providers/webgpu/math/matmul_algorithm_scheduler.h:150-153

MatMulAlgorithmScheduler is a new polymorphic base class but remains implicitly copyable and movable. Repository guidance requires new classes to use ORT_DISALLOW_COPY_ASSIGNMENT_AND_MOVE until those operations are shown to be necessary. Applying the macro to this base class also covers the Intel and test subclasses and prevents accidental value copies or slicing as more vendor schedulers are added.

Clarifications

None.

Test coverage

The device-independent tests cover parser behavior, selection precedence, common thresholds, typed packed configuration, Split-K prerequisites, overflow-safe dispatch arithmetic, and Intel low-M vec4 eligibility. Hardware tests compare deterministic results with a CPU reference and gate Intel subgroup, Split-K, and subgroup-matrix success cases on the selected adapter's actual capabilities. The automatic zero-K execution test also exercises the new no-input-binding shader path.

At the reviewed head, the macOS-arm64 WebGPU Release legs passed and the Linux WebGPU build passed; several unrelated or additional matrix jobs, including macOS WebGPU Debug and Windows plugin/build variants, were still pending. Linux WebGPU remains build-only, so it is not execution evidence. T1 remains uncovered because the plugin build leg does not run this test.

Verdict

The feature is valid and the scheduler/typed-plan direction is appropriate. T1 blocks merge because the new test fails in a supported plugin configuration. There are no remaining clarification requests. Q1 is mechanical cleanup required by the repository's new-class convention and does not change runtime behavior.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can commit the suggested changes from lintrunner.

Comment thread onnxruntime/core/providers/webgpu/vendor/intel/math/split_k_config.cc Outdated
Comment thread onnxruntime/core/providers/webgpu/vendor/intel/math/split_k_config.cc Outdated
Comment thread onnxruntime/core/providers/webgpu/vendor/intel/math/split_k_config.cc Outdated
@qjia7
Jiajia Qin (qjia7) marked this pull request as ready for review September 28, 2026 06:11
@qjia7

Copy link
Copy Markdown
Contributor Author

Jie Chen (@jchen10) Yang Gu (@gyagp), please take a look, thanks.

enum class MatMulAlgorithm {
SubgroupMatrix,
Naive,
IntelSubgroup,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IntelSubgroup is slightly weird here. Can we use Subgroup instead?

PackedSplitK,
};

inline std::optional<MatMulAlgorithm> ParseMatMulAlgorithm(std::string_view name) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

algorithm, scheduler, and dispatcher headers all use inline. It seems unusual.

return false;
}

class MatMulAlgorithmScheduler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A brief comment may improve readability for classes of new concept, like scheduler, plan, configuration, prerequisite, etc.

Comment on lines +412 to +432
const MatMulExecutionPlan plan = scheduler_->CreateExecutionPlan(
selection_params, context.ForcedMatMulAlgorithm());
const MatMulAlgorithm algorithm = plan.algorithm;
ORT_RETURN_IF_NOT(IsMatMulAlgorithmConfigurationCompatible(plan),
"MatMul algorithm ", MatMulAlgorithmName(algorithm),
" received an incompatible vendor configuration.");
const auto* packed_configuration = std::get_if<MatMulPackedConfiguration>(&plan.configuration);

MatMulAlgorithmPrerequisites prerequisites{};
prerequisites.can_use_subgroup_matrix = can_use_subgroup_matrix;
prerequisites.has_intel_subgroup_capability = has_intel_subgroup_capability;
prerequisites.has_nonzero_k = helper.K() > 0;
prerequisites.split_k_configured =
packed_configuration != nullptr && packed_configuration->split_dim_inner > 1;
prerequisites.deterministic_compute = selection_params.deterministic_compute;
prerequisites.is_vec4 = selection_params.is_vec4;
prerequisites.has_fused_activation = selection_params.has_fused_activation;
prerequisites.split_k_bias_layout_supported = !has_bias || is_channels_last;
ORT_RETURN_IF_NOT(MeetsMatMulAlgorithmPrerequisites(algorithm, prerequisites),
"MatMul algorithm ", MatMulAlgorithmName(algorithm),
" does not support these inputs or this device.");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What about making the scheduler not return a valid plan if configuration is not compatible or prerequisites are not met?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants