Skip to content

Add fabrics FI_HMEM support for CUDA grain payloads - #654

Open
thomastrue wants to merge 5 commits into
dmf-mxl:mainfrom
thomastrue:feature/fabrics-cuda-hmem
Open

thomastrue wants to merge 5 commits into
dmf-mxl:mainfrom
thomastrue:feature/fabrics-cuda-hmem

Conversation

@thomastrue

@thomastrue thomastrue commented Aug 5, 2026 •

Copy link
Copy Markdown

Summary

  • Advertise/require MXL_FABRICS_IFACE_CAP_HMEM for device grain flows on verbs
  • Register split host-header + CUDA-payload regions; use process-local CUDA bounce buffers on egress for IPC imports (not peermem-registerable)
  • Fix RC initiator/target CQ binding for HMEM MSG endpoints
  • Two-domain CUDA path: each side owns local cudaMalloc buffers and fabrics RDMA copies grains (fabrics-cuda-two-domain.sh, kube-fabrics-cuda-two-domain.yaml)

Stacked on #653 — this branch includes the CUDA grain-payload commits. Please merge #653 first; those commits drop out of this PR’s diff after that.

Test plan

  • mxl-fabrics-ofi-tests "[ofi]" (ProviderConfig HMEM cases)
  • Fabrics demo on tmpfs domain: target → mxl-gst-testsrc (CUDA options) → initiator; confirm ~30 grains/s both sides
  • Domain on ext4//tmp is expected to fail MR pin on this host; use /dev/shm
  • SRC_GPU=0 DST_GPU=0 ./examples/scripts/fabrics-cuda-two-domain.sh on a verbs/HMEM NIC

Introduce host/device payload allocators, flow-options payload location, and OpenGrain/GetGrain Ex APIs so discrete flows can back grains with cuda-linear device memory, including gst testsrc/sink paths.

Signed-off-by: Thomas True <ttrue@nvidia.com>
@thomastrue
thomastrue force-pushed the feature/fabrics-cuda-hmem branch from e16833e to 7181c8d Compare August 5, 2026 21:28
@felixpou felixpou linked an issue Aug 6, 2026 that may be closed by this pull request
19 tasks
Readers on the same node can open the writer's cudaMalloc buffers on another device when peer access is available, with same-GPU and two-GPU example skeletons.

Signed-off-by: Thomas True <ttrue@nvidia.com>
@thomastrue
thomastrue force-pushed the feature/fabrics-cuda-hmem branch from 7181c8d to 44bbccb Compare September 27, 2026 14:55
Move the CUDA allocator into libmxl-payload-cuda.so so libmxl can dlopen backends by name without linking cudart, while host and placeholder stay built in.

Signed-off-by: Thomas True <ttrue@nvidia.com>
Advertise and require HMEM-capable verbs domains for device flows, register split host header + CUDA payload regions, and use process-local CUDA bounce buffers on egress when payloads are IPC-imported.

Signed-off-by: Thomas True <ttrue@nvidia.com>
Each side keeps local cudaMalloc buffers and fabrics RDMA copies grains, with a same-machine rehearsal script and a two-pod skeleton.

Signed-off-by: Thomas True <ttrue@nvidia.com>
@thomastrue
thomastrue force-pushed the feature/fabrics-cuda-hmem branch from 44bbccb to a61be1b Compare September 27, 2026 15:44

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Architecture : <device hosted> grain mechanism.

1 participant