Skip to content

cuda: Q4_0_ROCMFP4 + Q4_0_ROCMFP4_FAST support and MMQ tiles on HIP (gfx1151) - #40

Merged
dzannotti merged 6 commits into
halo-box:masterfrom
baraxnaxgaming-commits:feat/hip-fp4-mmq
Sep 30, 2026
Merged

dzannotti merged 6 commits into
halo-box:masterfrom
baraxnaxgaming-commits:feat/hip-fp4-mmq

Conversation

@baraxnaxgaming-commits

@baraxnaxgaming-commits baraxnaxgaming-commits commented Sep 10, 2026 •

Copy link
Copy Markdown

What

Adds ggml-cuda/HIP support for the two ROCmFP4 types this tree ships for CPU+Vulkan:
Q4_0_ROCMFP4, Q4_0_ROCMFP4_FAST. The Q2/Q3/Q6/Q8_0_ROCMFPX reference layouts stay
CPU+Vulkan-only; on HIP they remain not supported and fall back to CPU (no kernels added here).

  1. Dispatch and support - dequant, convert (f32 and FP-to-FP), get_rows, MMVQ template instances, backend glue, so these GGUFs load and run on HIP at all. Also
    restores the *_hip_* helper headers (scale LUT + codebooks) that the format import pruned while
    common.cuh still includes them.
  2. MMQ tile path for Q4_0_ROCMFP4_FAST - mirrors the existing MXFP4 tiles: SRAM_LAYOUT_Q8_1
    loads, q8_0_q8_1 dp4a/WMMA vec_dot, D4 y-layout, UE4M3 scale decoder without the e8m0 x0.5
    factor, kvalues_rocmfp4 codebook (max level 10, not NVFP4's 12). Config rows are added only for
    RDNA3.5 (gfx1151) and Ampere - the two architectures validated below; should_use_mmq gates the
    type to exactly those, so every other architecture keeps today's dequant + BLAS fallback with
    zero behavior change (per review: no unvalidated policy changes on pascal/rdna2/rdna3/cdna/rdna4/blackwell).

Why

Without this, HIP builds run FP4 matmuls through dequant+hipBLAS. On Radeon 8060S that fallback is
the prefill bottleneck for the published Strix Halo ROCmFP4 GGUFs; HIP is half the userbase's default
backend. Tile rows mirror this master's current MXFP4 config set (incl. the 128/64/128 row), so the
gain sits on top of today's MXFP4 tuning.

Correctness

  • RTX 4090 (sm_89), CUDA 13.3 / MSVC, test-backend-ops vs CPU reference, re-verified on the rebased head:
    • -o MUL_MAT -p rocmfp4: 78 OK / 0 FAIL and 81 OK / 0 FAIL on two consecutive runs (exit 0; the case list randomizes broadcast shapes per run). On sm_89 should_use_mmq returns true for this type by default (turing_mma_available short-circuit), so these runs exercise the Ampere tile rows, not a BLAS fallback.
    • -o GET_ROWS -p rocmfp4: 14 OK / 0 FAIL (exit 0)
  • gfx1151 HIP, build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 ROCm 7.x:
    test-backend-ops test -o MUL_MAT -b ROCm0 -p rocmfp4: 78 OK / 0 FAIL (pre-rebase head; re-run offered below)
  • Review follow-up (cuda: return 0 for invalid ROCmFP4 scale bytes in the HIP decoder): the
    non-LUT rocmfp4_ue4m3_to_fp32_half_finite branch decoded scale bytes 0x7f-0xff instead of
    returning 0 like the CPU decoder, so a crafted GGUF diverged GPU vs CPU (validation only runs
    with --check-tensors). Now guarded on both branches; verified bit-identical to the CPU LUT
    decoder for all 256 scale byte values, and nvcc/sm_89 compile clean. No behavior change for
    valid bytes, so the test-backend-ops results above stand.

Performance (Radeon 8060S / gfx1151, 128 GB)

Baseline rebuilt from the merge-base of this PR in the same session on the same machine
(ab-base-dispatch = master + commit 1 of this PR, without the MMQ tiles), so the delta isolates
exactly the tile path. Method per CONTRIBUTING: warmup discarded; palindrome arm order
base/new/new/base x2 (8 runs per cell); -fa on; ubatch 512.

cell baseline t/s (median of 4, max) +MMQ t/s (median of 4, min) delta separation
pp512 Qwen3.8-27B-ROCmFP4-FAST 350.95 (353.36) 372.36 (368.51) +6.1 % clean, zero overlap
pp2048 same model 349.25 (349.51) 368.25 (368.17) +5.4 % clean, zero overlap
  • Kernel cell MUL_MAT m=4096,n=512 (test-backend-ops perf, ROCm0): 5389.5 -> 3503.0 us (x1.54).
  • Not tested: NVIDIA CUDA runtime perf (correctness only), decode-path changes (none expected; batch
    <=8 stays on MMVQ by design).

Deferred / notes

  • Dual-scale Q4_0_ROCMFP4 (non-FAST) MMQ deferred: needs the NVFP4-style per-16 scale convention.
  • All runs used default env (HIP_LAUNCH_BLOCKING unset), not CI posture; numbers are not
    comparable to blocking-mode runs.
  • Related but intentionally NOT in this PR: a Vulkan MMVQ change (amortise MXFP4 scale decode over
    the whole block, measured x1.67 at n=8 on gfx1151). Happy to split it into its own PR if wanted.

Written by Hermes Agent (Nous Research); benchmarks run on a Ryzen AI Max+ 395.

Cleanup after review: dropped unreferenced ggml/rocmfp4/rocmfp4_hip.cu, rocmfp4_hip_codebook.cuh and rocmfpx/rocmfpx_hip_codebook.cuh (never in any CMake source list, no callers; live UE4M3 scale decoder is rocmfp4_hip_scale.cuh via common.cuh).

Rebased onto current master (ec01a7dfc) and reduced scope per review: MMQ config rows now exist only for RDNA3.5 + Ampere, with an explicit should_use_mmq gate so unvalidated architectures (incl. Blackwell/Rubin) fall back to dequant + BLAS instead of inheriting a tile policy nobody measured.

@dzannotti dzannotti left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validation: the HIP gfx1151 build succeeded. Q4_0_ROCMFP4_FAST MUL_MAT passed 45/45 targeted ROCm correctness cases. test-quantize-fns and test-rocmfpx also completed successfully.

Benchmark reproduced for the stated kernel shape: against the dispatch-only parent (ede2bbe), test-backend-ops perf at m=4096, n=512, k=14336 measured 5,540.70 us/run before the MMQ tiles and 2,919.23 us/run with this PR: 1.90x faster. The PR's full-model prefill result remains unverified here because the ROCmFP4 model is not local.

Scope: the stated target is gfx1151, yet the PR changes shared MMQ configuration for Ampere, CDNA, RDNA2/3/4 too. Please split or tightly gate the Strix portion and separately validate every architecture whose policy changes.

comment generated by my clanker Codex

@baraxnaxgaming-commits

Copy link
Copy Markdown
Author

Thanks for the review. Three clarifications on scope, plus where we'll comply:

  1. No existing type's policy changes on any architecture. The PR is 754 insertions / 1 deletion (the single deletion is a string-list line). Every change in mmq-config-{ampere,cdna,pascal*,rdna2,rdna3,rdna3-5,rdna4}.cuh adds CASE(GGML_TYPE_Q4_0_ROCMFP4_FAST, …) rows for a type that exists only in this fork — dispatch of upstream types on those architectures is bit-for-bit unchanged.

  2. We'll tighten the gate as requested. We will keep CASE rows only for RDNA3.5 (gfx1151) and Ampere — validated on RTX 4090 CUDA: default build -p rocmfp4 288 OK / 0 FAIL; GGML_CUDA_FORCE_MMQ=ON -o MUL_MAT -p rocmfp4 42 OK / 0 FAIL — and drop the rows for architectures we cannot measure. Those arches keep today's dequant+BLAS fallback: zero behavior change, no re-validation burden on anyone.

  3. test-rocmfpx does not exist in this PR or on master. If you have it locally, please share it — we'll wire it into CI. Current type-specific coverage is test-backend-ops -o MUL_MAT -p rocmfp4 vs CPU ref (gfx1151 HIP: 78 OK / 0 FAIL); happy to promote that filter into a dedicated CI job as the type-specific regression test requested.

On indexing safety: MMQ entry for this type is gated by the same ne[0] % ggml_blck_size() and alignment checks used for MXFP4, and the tile config mirrors MXFP4's Q8_1 SRAM layout one-for-one — malformed dimensions reach the new loaders no more than they reach MXFP4's today.

@BonkusCheemus

Copy link
Copy Markdown

Disclosure: written by a Claude agent (Anthropic's Claude Code) for the owner of this account.

We did an independent port of the same ROCmFP4 HIP paths (dequant/convert, get_rows, MMVQ, and a Q4_0_ROCMFP4_FAST MMQ path) on a gfx1151 box, and compared it with this PR at the source level. The shared core agrees with ours: same codebook, same UE4M3 decode, same MMQ tile geometry. Three things we noticed in this PR:

  1. ggml/rocmfp4/rocmfp4_hip.cu (two extern "C" dequant kernels) is never compiled: ggml-hip/CMakeLists.txt only globs ../ggml-cuda/*.cu and the template instances, and the PR has no CMake change. Either drop the file or wire it in.
  2. The description mentions "get_rows (incl. the CUDA-graph-safe variant)", but getrows.cu has a single type switch and the PR adds only the two get_rows_cuda_kq cases there. Is there a second path we missed?
  3. rocmfp4_get_low_int_from_codebook_16 in rocmfp4_hip_codebook.cuh has no caller.

For reference, our port passes test-backend-ops -o MUL_MAT -p rocmfp4 -b ROCm0 90/90 against the CPU (f32 activations, n = 1..9, 16, 64) and GET_ROWS 16/16. We can run the same check on this PR's branch and post the numbers if useful.

baraxnaxgaming-commits added a commit to baraxnaxgaming-commits/strix-llama.cpp that referenced this pull request Sep 28, 2026
Review finding (halo-box#40): three files added in the
first commit are compiled by nothing and referenced nowhere:

- ggml/rocmfp4/rocmfp4_hip.cu - two extern "C" dequant kernels; not in
  any CMake source list (ggml-hip globs only ../ggml-cuda/*.cu and the
  template instances) and no callers, so unreachable even from tests.
- ggml/rocmfp4/rocmfp4_hip_codebook.cuh - included by nothing; live
  MMVQ/MMQ paths use get_int_from_table_16 + kvalues_rocmfp4 from
  ggml-common.h, which already lowers to permb on HIP. The single-nibble
  perm variant was for an FA K/V decode path not present in this PR.
- ggml/rocmfpx/rocmfpx_hip_codebook.cuh - same, included by nothing.

The live UE4M3 scale decoder is rocmfp4_hip_scale.cuh (included via
common.cuh) and stays. No behavior change: verified zero references to
these files/symbols in the rest of the tree.
@baraxnaxgaming-commits

Copy link
Copy Markdown
Author

Thanks — all three confirmed on our side, cleanup pushed in 0ba39db:

  1. rocmfp4_hip.cu never compiled — correct. It came in with the "restored helper headers" commit but was never added to any CMake source list (ggml-hip/CMakeLists.txt only globs ../ggml-cuda/*.cu + template instances), and rocmfp4_hip_dequantize_* has zero callers tree-wide, so it was unreachable even from tests. Dropped rather than wired in: the live UE4M3 scale decoder is rocmfp4_hip_scale.cuh (included via common.cuh) and that stays.
  2. No second get_rows path — you didn't miss anything; the PR body phrase "incl. the CUDA-graph-safe variant" was simply wrong. The change is exactly the two get_rows_cuda_kq cases in the single type switch (getrows.cu). Body corrected.
  3. Orphan codebook helpers — correct, and slightly broader: rocmfp4_hip_codebook.cuh and rocmfpx_hip_codebook.cuh are both included by nothing. The live MMVQ/MMQ paths use get_int_from_table_16(aux_q4, kvalues_rocmfp4) from ggml-common.h, which already lowers to permb on HIP; the single-nibble variant was written for an FA K/V decode path that isn't in this PR. Both files dropped.

And yes please — run test-backend-ops -o MUL_MAT -p rocmfp4 -b ROCm0 and -o GET_ROWS -p rocmfp4 -b ROCm0 against this branch (head is now 0ba39db; cleanup commit only deletes unreferenced files, no behavior change). Our correctness table has MUL_MAT coverage only (288/0 CUDA default, 42/0 forced-MMQ, 78/0 gfx1151); independent GET_ROWS numbers from a second port would close a real gap.

Loads Q4_0_ROCMFP4, Q4_0_ROCMFP4_FAST and the Q2/Q3/Q6/Q8_0_ROCMFPX GGUF
tensor types on ggml-cuda/HIP (previously CPU+Vulkan only). Adds dequant,
convert, get_rows support and MMVQ dispatch; restores the *_hip_* helper
headers (scale LUT + codebooks) that the format import pruned.

Part 1/2 of the FP4-on-HIP work; part 2 adds the MMQ tile path.
Step 2 of the Strix Halo HIP plan: give the production FP4 format an
MMQ path so prefill batches (>8, incl. spec-verify at wide ubatch and
MoE routing) stop falling through to dequant+hipBLAS purely because
the type was absent from the MMQ switches.

Mirrors MXFP4 exactly - same block size (32), one UE4M3 scale per 32
values, SRAM_LAYOUT_Q8_1 tiles, q8_0_q8_1 vec_dot kernels (dp4a and
MMA data layout), D4 y-side quantizer. Differences from MXFP4:
kvalues_rocmfp4 codebook (Codebook10, max level 10 not 12) and the
UE4M3 scale decoder without the e8m0 *0.5 factor. Tile rows mirror
this master's current MXFP4 config set, incl. the 128/64/128 row.

Not added to Blackwell configs on purpose: use_native_fp4 is false
for this type there, so it falls through to the Ampere config like
NVFP4-generic does. Dual-scale Q4_0_ROCMFP4 (2 scales/block) deferred
to a follow-up slice (NVFP4-style per-16 convention).

Validated: test-backend-ops vs CPU ref on RTX 4090 CUDA (default build
-p rocmfp4: 288 OK / 0 FAIL; GGML_CUDA_FORCE_MMQ=ON -o MUL_MAT -p
rocmfp4: 42 OK / 0 FAIL) and on gfx1151 HIP (-o MUL_MAT -b ROCm0 -p
rocmfp4: 78 OK / 0 FAIL). Measured on Ryzen AI Max+ 395 against a
rebuilt merge-base in the same session: pp512 +6.1%, pp2048 +5.4%
(palindrome x2, zero overlap); MUL_MAT m=4096,n=512 kernel 5389 ->
3503 us (x1.54).
Review finding (halo-box#40): three files added in the
first commit are compiled by nothing and referenced nowhere:

- ggml/rocmfp4/rocmfp4_hip.cu - two extern "C" dequant kernels; not in
  any CMake source list (ggml-hip globs only ../ggml-cuda/*.cu and the
  template instances) and no callers, so unreachable even from tests.
- ggml/rocmfp4/rocmfp4_hip_codebook.cuh - included by nothing; live
  MMVQ/MMQ paths use get_int_from_table_16 + kvalues_rocmfp4 from
  ggml-common.h, which already lowers to permb on HIP. The single-nibble
  perm variant was for an FA K/V decode path not present in this PR.
- ggml/rocmfpx/rocmfpx_hip_codebook.cuh - same, included by nothing.

The live UE4M3 scale decoder is rocmfp4_hip_scale.cuh (included via
common.cuh) and stays. No behavior change: verified zero references to
these files/symbols in the rest of the tree.
Review scope reduction (halo-box#40 review): keep the
MMQ tile rows only where the path is actually validated - RDNA3.5
(gfx1151, test-backend-ops 78/0) and the Ampere config table (RTX 4090
sm_89). Drop the pascal-dp4a / pascal-older / rdna2 / rdna3 / cdna /
rdna4 rows: no hardware to measure there, so those architectures keep
today's dequant + BLAS fallback (zero behavior change) instead of an
unvalidated tile policy.

should_use_mmq now returns false for this type outside the gated set
(including Blackwell, which has no rows - native FP4 MMQ is
MXFP4/NVFP4-only there), so a missing config row can never reach
ggml_cuda_mmq_get_config's abort.
@baraxnaxgaming-commits

Copy link
Copy Markdown
Author

Scope reduction done and branch rebased onto current master (ec01a7dfc) — head is now d3a8ab54.

What changed in response to the review:

  • MMQ config rows for Q4_0_ROCMFP4_FAST now exist only in mmq-config-rdna3-5.cuh (gfx1151, validated) and mmq-config-ampere.cuh (validated on RTX 4090 / sm_89). The rows added to pascal-dp4a / pascal-older / rdna2 / rdna3 / cdna / rdna4 are gone — those architectures keep today's dequant + BLAS fallback, zero behavior change, nothing to re-validate.
  • should_use_mmq now gates the type explicitly: (ampere_mma_available(cc) && cc < GGML_CUDA_CC_BLACKWELL) || GGML_CUDA_CC_IS_RDNA3_5(cc). Blackwell/Rubin are excluded even though they would otherwise fall through to the Ampere table (native FP4 MMQ there is MXFP4/NVFP4-only), so no architecture can reach a config lookup without rows.

Re-validation on the rebased head (RTX 4090, sm_89 — the NVIDIA arm kept):

  • test-backend-ops -o MUL_MAT -p rocmfp4 vs CPU ref: 78 OK / 0 FAIL and 81 OK / 0 FAIL on two consecutive runs (exit 0; the case list randomizes broadcast shapes per run). On sm_89 should_use_mmq returns true for this type by default, so these runs exercise the Ampere tile rows, not a BLAS fallback.
  • -o GET_ROWS -p rocmfp4: 14 OK / 0 FAIL (exit 0)

gfx1151 numbers from your validation (45/45 targeted cases, and our earlier 78/0) should carry over — the rebase resolves dequantize.cuh against master's <dst_t, dst_ptr_t> k-quant refactor, and the scope commit only deletes config rows + adds the gate; no kernel changes. Happy to re-run gfx1151 before merge if you want fresh numbers for this head.

One heads-up unrelated to this PR: mmb.cu (landed via the perf/mmb-qwen4exp-only merge, Sep 26) does not compile under nvcc — ext_vector_type, _Float16, and __builtin_amdgcn_* are HIP-only constructs with no guard. CUDA CI on master has been red since then (run 36348852734 also shows a getrows device-stub error). It doesn't block this PR (the workflow's cuda job is push-to-master only) but worth its own fix-forward.

@dzannotti dzannotti left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified on Strix Halo (gfx1151, HIP). Qwen3-1.7B was quantized to Q4_0_ROCMFP4_FAST and Q4_0_ROCMFP4 with this branch's llama-quantize; runs are ABBA n=4.

merge-base ec01a7dfc (CPU fallback) this PR
ROCMFP4_FAST pp512 / pp2048 / tg128 125.1 / 120.3 / 46.5 5907.5 / 5966.8 / 176.4
ROCMFP4 pp512 / pp2048 / tg128 94.6 / 93.2 / 32.3 3321.6 / 3384.6 / 164.1
gpt-oss-20b MXFP4 (regression check) 1921.0 / 1981.7 / 73.6 1921.0 / 1982.8 / 73.7

MMQ tiles vs this branch's own dispatch-only commit 2fd4c2a77 (-ub 512):

dispatch-only MMQ Δ
pp512 3373.6 5894.0 +74.7%
pp2048 3403.4 5970.6 +75.4%
tg128 176.7 176.4 –
kernel MUL_MAT m=4096 n=512 k=14336 5482.6 µs 2922.0 µs 1.88×

test-backend-ops -b ROCm0: MUL_MAT 45/45 (FAST) and 90/90, MUL_MAT_ID 73/73 and 146/146, GET_ROWS 8/8 and 16/16.

Please fix before merge:

  • Description: it says Q3/Q6/Q8_0_ROCMFPX are supported, but on HIP they are still not supported (MUL_MAT 0/0, 78 cases skipped each) and fall back to CPU.
  • Scale decode: rocmfp4_hip_scale.cuh decodes scale bytes 0x7f–0xff differently from the CPU decoder (rocmfp4.c), which returns 0 for them. The comment says those bytes are validated first, but validation only runs with --check-tensors, so a crafted GGUF gives different results on GPU and CPU. It's not an out-of-bounds read. if (x > 0x7e) return 0.0f; would align the two.

Comment generated via automated review of clanker Claude Code.

The arithmetic UE4M3 half-scale decode (non-LUT branch) accepted
bytes 0x7f-0xff and produced garbage scales, while the CPU decoder
(rocmfp4.c) returns 0 for them. Loader validation of those bytes only
runs with --check-tensors, so a crafted GGUF decoded differently on
GPU and CPU. Guard x > 0x7e -> 0 in both branches to match CPU.

Verified bit-identical to the CPU LUT decoder for all 256 scale byte
values; nvcc sm_89 compile clean. No behavior change for valid bytes,
so existing test-backend-ops results stand.

Assisted-by: Hermes Agent (Qwen3.8-Flash-Next)
@baraxnaxgaming-commits baraxnaxgaming-commits changed the title cuda: ROCmFPx tensor types + Q4_0_ROCMFP4_FAST MMQ tiles on HIP (gfx1151) cuda: Q4_0_ROCMFP4 + Q4_0_ROCMFP4_FAST support and MMQ tiles on HIP (gfx1151) Sep 30, 2026
@baraxnaxgaming-commits

Copy link
Copy Markdown
Author

Both points addressed in 40cc375.

Description: corrected. The PR now states that only Q4_0_ROCMFP4 and Q4_0_ROCMFP4_FAST gain HIP support; Q2/Q3/Q6/Q8_0_ROCMFPX stay CPU+Vulkan-only and fall back to CPU on HIP, matching the actual dispatch (no ROCmFPX cases exist anywhere under ggml/src/ggml-cuda/). Title updated to match.

Scale decode: adopted the suggested guard in rocmfp4_ue4m3_to_fp32_half_finite - x > 0x7e now returns 0 before both the LUT and the arithmetic branch, so a crafted GGUF decodes identically on GPU and CPU regardless of --check-tensors. The stale comment claiming pre-validation was replaced with one that states the actual invariant.

Verification for the decode change:

  • Exhaustive host-side equivalence check over all 256 scale byte values against the CPU decoder in rocmfp4.c (LUT table parsed from source): bit-identical everywhere; the previously divergent bytes were 0x7f-0xff (arithmetic branch produced e.g. 240.0 for 0xff, now 0).
  • Valid bytes (<= 0x7e) are unchanged bit-for-bit, so the existing gfx1151/CUDA test-backend-ops results above still stand; happy to re-run on hardware if you want a fresh confirmation post-merge-head.

Note the guard also covers the three call sites that feed MMQ tile loads and MMVQ (mmq-load-tiles.cuh, vecdotq.cuh, dequantize.cuh) - all go through this one helper, so there is no second decode path left unguarded.

@dzannotti dzannotti left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-verified on Strix Halo (gfx1151, HIP) at 40cc375d8.

check result
test-backend-ops MUL_MAT / MUL_MAT_ID / GET_ROWS, q4_0_rocmfp4 90/90, 146/146, 16/16
same, q4_0_rocmfp4_fast 45/45, 73/73, 8/8
llama-bench Qwen3-1.7B ROCmFP4_FAST, previous head vs this head (ABBA) pp512 +1.3%, pp2048 +0.4%, tg128 ±0%: no cost from the guard

The scale-decode guard now matches the CPU decoder for invalid bytes, and the description matches the HIP dispatch. Approving.

Comment generated via automated review of clanker Claude Code.

@dzannotti
dzannotti merged commit 8b72177 into halo-box:master Sep 30, 2026
9 of 12 checks passed
pwilkin added a commit that referenced this pull request Sep 30, 2026
Keep ROCmFP4 FAST and MXFP4/NVFP4 W4A4 MMQ declarations.
Keep explicit rollback registrations, including Qwen4exp, without duplicate directory runs.
Combine the status/all-model harness with scaled-weight seq_cp/graph-reuse coverage.

Assisted-by: Codex
BonkusCheemus added a commit to BonkusCheemus/strix-llama.cpp that referenced this pull request Oct 1, 2026
One conflict, ggml/src/ggml-cuda/dequantize.cuh: master added the ROCmFP4
dequantizers (halo-box#40) where this branch added the turbo2/3/4 ones. Kept both.
Turbo type IDs (109-111) stay clear of master's ROCmFPX range (100-107).

Stamp: Kurumi#Opus5.5H@fresh-io
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017fbUwN2eaP1kh7wNDYofQk
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants