Skip to content

[AIMIGRAPHX-1005] [AIMIGRAPHX-1227] [AIMIGRAPHX-1228] - #5114

Draft
bdevorem wants to merge 14 commits into
developfrom
bdevorem/flash_decode_bug
Draft

[AIMIGRAPHX-1005] [AIMIGRAPHX-1227] [AIMIGRAPHX-1228]#5114
bdevorem wants to merge 14 commits into
developfrom
bdevorem/flash_decode_bug

Conversation

@bdevorem

@bdevorem bdevorem commented Aug 5, 2026

Copy link
Copy Markdown
Member

Motivation

Technical Details

Changelog Category

Add a CHANGELOG.md entry for any option other than Not Applicable

    • Added: New functionality.
    • Changed: Changes to existing functionality.
    • Removed: Functionality or support that has been removed. (Compared to a previous release)
    • Optimized: Component performance that has been optimized or improved.
    • Resolved Issues: Known issues from a previous version that have been resolved.
    • Not Applicable: This PR is not to be included in the changelog.

Follow the LLVM AI Tool Use Policy for contributions using AI.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR strengthens the fuse_attention flash-decoding path by making flash-decoding submodule rebuilding more robust in the presence of additional submodule inputs (e.g., masks), constants/materialized literals, and non-trivial instruction ordering. It also expands the unit tests to cover these additional graph shapes and edge cases, and makes existing tests less sensitive to host environment configuration.

Changes:

  • Add test coverage for flash decoding with embedded @literal, extra @param mask inputs, unary ops affecting broadcast shape inference, outline nodes, and rebuild ordering constraints.
  • Update flash-decoding submodule rebuild to handle @literal/@outline, extra score-shaped params, and dependency-aware rebuild ordering (incl. broadcast shape inference).
  • Ensure certain tests run fuse_attention without inheriting flash-decoding settings from environment variables.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.

File Description
test/fuse_attention.cpp Adds new flash-decoding tests and introduces an “isolated” pass runner to avoid environment-dependent behavior.
src/fuse_attention.cpp Improves flash-decoding submodule reconstruction (params handling, literals/outlines, topo-ish rebuild, broadcast shape handling).

Comment thread src/fuse_attention.cpp Outdated
Comment thread test/fuse_attention.cpp Outdated
Comment thread src/fuse_attention.cpp
@codecov

codecov Bot commented Aug 5, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.37017% with 12 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/fuse_attention.cpp 93.37% 12 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##           develop    #5114      +/-   ##
===========================================
- Coverage    93.26%   93.22%   -0.04%     
===========================================
  Files          623      623              
  Lines        33097    33191      +94     
===========================================
+ Hits         30866    30942      +76     
- Misses        2231     2249      +18     
Files with missing lines Coverage Δ
src/fuse_attention.cpp 94.95% <93.37%> (-2.55%) ⬇️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@gh-app-migraphx-bot-pr-write

Copy link
Copy Markdown
Test Batch New Rate (38fc3e) Old Rate (3a503c)* Diff Status
torchvision-resnet50 64 3,325.83 3,264.92 1.87%
torchvision-resnet50_fp16 64 7,879.49 7,548.67 4.38%
torchvision-densenet121 32 2,491.18 2,483.99 0.29%
torchvision-densenet121_fp16 32 5,042.87 5,004.24 0.77%
torchvision-inceptionv3 32 2,077.09 2,058.51 0.90%
torchvision-inceptionv3_fp16 32 4,491.64 4,416.99 1.69%
cadene-inceptionv4 16 819.91 820.61 -0.08%
cadene-resnext64x4 16 780.48 782.78 -0.29%
slim-mobilenet 64 8,388.56 8,386.36 0.03%
slim-nasnetalarge 64 229.03 228.86 0.08%
slim-resnet50v2 64 3,223.87 3,180.91 1.35%
bert-mrpc-onnx 8 1,172.60 1,168.84 0.32%
bert-mrpc-tf 1 493.56 498.63 -1.02%
pytorch-examples-wlang-gru 1 474.38 473.35 0.22%
pytorch-examples-wlang-lstm 1 384.37 384.83 -0.12%
torchvision-resnet50_1 1 1,052.69 1,046.63 0.58%
cadene-dpn92_1 1 451.31 437.32 3.20%
cadene-resnext101_1 1 364.75 365.89 -0.31%
onnx-taau-downsample 1 843.63 844.09 -0.05%
dlrm-criteoterabyte 1 32.38 32.42 -0.11%
dlrm-criteoterabyte_fp16 1 51.89 51.80 0.17%
agentmodel 1 9,729.42 9,209.12 5.65% 🔆
unet_fp16 2 39.61 58.80 -32.63% 🔴
resnet50v1_fp16 1 1,400.75 1,366.11 2.54%
resnet50v1_int8 1 1,183.84 1,883.96 -37.16% 🔴
bert_base_cased_fp16 64 914.58 1,098.16 -16.72% 🔴
bert_large_uncased_fp16 32 345.49 345.59 -0.03%
bert_large_fp16 1 207.41 206.59 0.40%
distilgpt2_fp16 16 2,090.77 2,092.89 -0.10%
yolov5s 1 284.16 558.33 -49.11% 🔴
tinyllama 1 45.85 45.83 0.04%
vicuna-fastchat 1 44.12 44.20 -0.17%
whisper-tiny-encoder 1 412.63 411.87 0.18%
whisper-tiny-decoder 1 189.83 408.48 -53.53% 🔴
llama2_7b 1 8.66 20.84 -58.45% 🔴
qwen1.5-7b 1 23.49 23.58 -0.37%
phi3-3.8b 1 23.46 26.72 -12.17% 🔴
llama3-8b 1 13.82 21.80 -36.64% 🔴
whisper-large-encoder 1 10.18 10.18 0.02%
whisper-large-decoder 1 103.05 105.30 -2.13%
mistral-7b 1 5.89 23.78 -75.24% 🔴
FLUX.1-schnell 1 104.69 755.22 -86.14% 🔴

Regressions detected 🔴

* No develop baseline was found for this PR's branch point; compared against the latest available develop run instead.

@gh-app-migraphx-bot-pr-write

Copy link
Copy Markdown
Test Status Result
bert-mrpc-onnx PASSED: MIGraphX meets tolerance
bert-mrpc-tf PASSED: MIGraphX meets tolerance
pytorch-examples-wlang-gru PASSED: MIGraphX meets tolerance
pytorch-examples-wlang-lstm PASSED: MIGraphX meets tolerance
dlrm-criteoterabyte PASSED: MIGraphX meets tolerance
agentmodel PASSED: MIGraphX meets tolerance
unet PASSED: MIGraphX meets tolerance
resnet50v1 PASSED: MIGraphX meets tolerance
bert_base_cased_fp16 PASSED: MIGraphX meets tolerance
bert_large_uncased_fp16 🔴 FAILED: MIGraphX is not within tolerance - check verbose output
bert_large PASSED: MIGraphX meets tolerance
yolov5s PASSED: MIGraphX meets tolerance
tinyllama PASSED: MIGraphX meets tolerance
vicuna-fastchat PASSED: MIGraphX meets tolerance
whisper-tiny-encoder PASSED: MIGraphX meets tolerance
whisper-tiny-decoder PASSED: MIGraphX meets tolerance
distilgpt2_fp16 🔴 FAILED: MIGraphX is not within tolerance - check verbose output
llama2_7b PASSED: MIGraphX meets tolerance
qwen1.5-7b PASSED: MIGraphX meets tolerance
phi3-3.8b PASSED: MIGraphX meets tolerance
llama3-8b PASSED: MIGraphX meets tolerance
whisper-large-encoder PASSED: MIGraphX meets tolerance
whisper-large-decoder PASSED: MIGraphX meets tolerance
mistral-7b PASSED: MIGraphX meets tolerance
FLUX.1-schnell PASSED: MIGraphX meets tolerance

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants