Skip to content

[New GPU Codegen Prep] Add NestedGPUDeviceMapLowering - #2534

Open
ThrudPrimrose wants to merge 14 commits into
mainfrom
gpu-prep/lower-nested-gpu-maps
Open

ThrudPrimrose wants to merge 14 commits into
mainfrom
gpu-prep/lower-nested-gpu-maps

Conversation

@ThrudPrimrose

Copy link
Copy Markdown
Collaborator

Lowers a GPU_Device map nested inside another into a single kernel: the outer map absorbs the inner maps' parameters and each inner body becomes a nested SDFG guarded by the condition selecting the iterations it owns. The body is moved with nest_state_subgraph and the scope dissolved through memlet paths rather than rebuilt by hand, the bound check reproduces the range step, and the absorbed range is the bounding box of the inner ranges alone.

Lowers a GPU_Device map nested inside another into a single kernel: the outer
map absorbs the inner maps' parameters and each inner body becomes a nested
SDFG guarded by the condition selecting the iterations it owns.

The body is moved with nest_state_subgraph and the scope dissolved through
memlet paths, rather than rebuilt by hand. The bound check reproduces the
range step, and the absorbed range is the bounding box of the inner ranges
alone instead of being seeded at zero.
@ThrudPrimrose ThrudPrimrose changed the title [WIP] Add NestedGPUDeviceMapLowering [New GPU Codegen Prep] Add NestedGPUDeviceMapLowering Aug 26, 2026
… removed

Moving the body with nest_state_subgraph lost two things the hand-written
nesting did. An empty memlet leaving the map scope is an ordering edge, and
dropping it with the map detached the guarded body from its enclosing kernel,
so the scope traversal could no longer place the nodes below it. The nested
SDFG node was also left without the enclosing scope's symbols, which are
declared on the nested SDFG but have to be bound on the node as well.

A body directly nested in a kernel scope, and the binding of the outer
kernel's parameter, are both now tested.
…orbed params unique

Ported from extended and generalized:
- translate each hoisted inner range through every NestedSDFG symbol_mapping, applied
  simultaneously (a swapping mapping no longer maps a name back to itself);
- refuse a hoisted bound naming anything the host cannot evaluate where the grid is sized
  (a kernel param, or a loop variable defined inside the kernel);
- rename an inner param that repeats a kernel param or a symbol of any SDFG between the
  kernel and the inner map, so the absorbed dimension is declared once and a nested
  binding of the same name does not shadow it;
- type the threaded symbols from the map ranges instead of a fixed default.
…refusal cases, share one kernel lookup

The guarded-body symbol check joins the flattening test it shared a setup with.
@ThrudPrimrose
ThrudPrimrose marked this pull request as ready for review September 28, 2026 16:33
ThrudPrimrose and others added 3 commits October 1, 2026 15:41
…p duplicate helpers, share one traversal

The pass walks every scope and nested SDFG below a kernel with one recursive generator, so a GPU map behind a sequential scope or at any nesting depth is absorbed into the outermost kernel. The tests parametrize the bound, symbol-mapping and step cases and check each guard against the iterations of its own map.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant