Qualcomm AI Engine Direct - [GenAI Pipeline] Multimodal LLM & LLM Preparation and Quantization - #23050
Qualcomm AI Engine Direct - [GenAI Pipeline] Multimodal LLM & LLM Preparation and Quantization#23050DannyYuyang-quic wants to merge 1 commit into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23050
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New FailureAs of commit 7fdb147 with merge base 91d26b3 ( NEW FAILURE - The following job has failed:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@pytorchbot label "release notes: qualcomm" |
|
Hi @psiddh, TL;DR:This PR makes the GenAI Pipeline runnable through The existing Together with PR B1, Compilation Foundation (#22846), the following commands LLMpython -m backends.qualcomm.genai_pipeline.cli --model gemma3-1b --soc SM8750 --calib-samples examples/qualcomm/oss_scripts/llama/assets/samples/text.json --compile-only --max-seq-len 1024Multimodal LLMVision-Language modelpython -m backends.qualcomm.genai_pipeline.cli --model internvl3_1b --soc SM8750 --calib-samples examples/qualcomm/oss_scripts/llama/assets/samples/vision.json --compile-only --max-seq-len 1024Audio-Language modelpython -m backends.qualcomm.genai_pipeline.cli --model granite_speech_3_3-2b --soc SM8750 --calib-samples examples/qualcomm/oss_scripts/llama/assets/samples/audio --compile-only --max-seq-len 1024Please have a look. Thanks! |
4e19eb9 to
5389338
Compare
…Quantization
Add the LLM/MLLM GenAI pipeline integration for model preparation, dataset-driven calibration, and ExecuTorch quantization.
Summary:
- Add the GenAI pipeline CLI and stage context wiring.
- Add model registry lookup helpers for configs, graph builders, source
transforms, checkpoint loaders, quantization settings, and adapters.
- Add component-aware LLM/MLLM model preparation adapters for decoder,
embedding, vision, and audio modules.
- Add dataset adapters, collators, and dataset option for
calibration, training, and evaluation.
- Add ExecuTorch quantization support for export/prepare, encoding
initialization, calibration, encoding override, conversion, and QDQ EP save.
- Add source transforms for checkpoint remapping, dtype override, embedding
scaling, RoPE layout, RMSNorm offset, and linear-to-conv2d conversion.
- Expand unit coverage for model preparation, quantization, source transforms,
datasets, and pipeline stage behavior.
5389338 to
7fdb147
Compare
| return LLMCalibrationDataAdapter(**kwargs) | ||
|
|
||
|
|
||
| def get_training_dataset_adapter(dataset_options, is_multimodal=False): |
There was a problem hiding this comment.
get_training_dataset_adapter / get_eval_dataset_adapter
Both call get_dataset_adapter(...), which only reads the calib_* options — so --train-tasks, --train-hf-dataset, --train-limit, --eval-tasks, --eval-limit and --eval-num-fewshot never reach a loader, and training/eval get the calibration corpus instead. Is purpose-specific selection planned for a later PR, or should get_dataset_adapter take the purpose now?
| config = process_model_args( | ||
| control_args, | ||
| ModelArgs(**base_args), | ||
| quant_recipe(mode == Mode.CALIBRATE), |
There was a problem hiding this comment.
quant_recipe can be None — get_quant_recipe returns getattr(config, "quant_recipe", None) — and the gemma4 branch guards for exactly that (quant_recipe().get_kv_io_bit_width() if quant_recipe else 32). This call doesn't, so any registry row without a recipe raises TypeError: 'NoneType' object is not callable.
| "attention sink is not yet supported in GenAI Pipeline.", | ||
| args.max_seq_len, | ||
| ) | ||
| args.max_context_len = args.max_seq_len |
There was a problem hiding this comment.
--max-context-len is accepted, then unconditionally overwritten. Since attention sink isn't supported yet, I'd reject it in _validate_args alongside --qat and --use-attention-sink rather than silently ignoring a value the user set.
Same thought for the rest of the parser: with compilation_stage=None and inference_stage=None, --prompt, --temperature, --audio-path, --image-path, --pre-gen-pte and --eval-methods are all no-ops, and --compile-only just logs a line. Worth gating them the same way so the CLI doesn't advertise more than it does.
| _GRAPH_ROLES = { | ||
| # The quantize graph supplies encodings for the decoder's deployed variants. | ||
| ARTIFACT_TEXT_DECODER: { | ||
| GRAPH_FORWARD: _GraphRole.QUANTIZE, |
There was a problem hiding this comment.
This is correct today (GRAPH_FORWARD = "forward" isn't in DECODER_GRAPH_NAMES), but it's order-dependent: if GRAPH_FORWARD ever joins that list the ** overwrites QUANTIZE with DEPLOY, the quantize graph stops receiving calibration data, and nothing errors — you just get uncalibrated encodings. One assert GRAPH_FORWARD not in DECODER_GRAPH_NAMES would pin it.
| input_config: Graph inputs and metadata used for observer execution. | ||
| """ | ||
| for component, prepared_graphs in prepared_modules.items(): | ||
| for graph_name, graph_module in prepared_graphs.items(): |
There was a problem hiding this comment.
No break, so the last non-deploy graph wins — while _flatten_calibration_example_inputs does break and takes the first. Benign while each component has exactly one QUANTIZE graph, but the module and its example inputs could silently come from different graphs if that ever changes. Worth making the two loops agree.
| attr = getattr(attr, target_atom) | ||
| return attr | ||
|
|
||
| def _set_attr( |
There was a problem hiding this comment.
These duplicate the get_attr_from_target / set_attr_from_target helpers that the HF LLM Optimization PR adds to backends/qualcomm/builders/utils.py (and removes from llm_wrappers.py for the same reason). Worth coordinating so one copy survives.
|
do you know what these failure are ? MLTEC_elastic - GitHubPOC - executorch dev1-danny-pr-A1 — The build failed |
|
@claude can you thoroughly review this large PR from arch and direction point of view . Also review my comments and overall sugegst fi there are any blockers / concerns |
|
Claude finished @psiddh's task in 1m 26s —— View job Summary: Request changes. The pipeline structure works, but I found three correctness blockers: multimodal quantization fails in PR Review: #23050 (architecture / direction, plus your comments)SummaryThis PR adds the front half of the Qualcomm GenAI Pipeline: CLI, registry lookup, source transforms, LLM/MLLM model preparation, dataset-backed calibration, and component/graph-keyed PTQ that produces QDQ graphs. All 77 changed files are under Everything below comes from reading the code. I could not run Python or the tests in this environment. 🚫 Blockers1. Multimodal quantization always fails at the last step.
2. The "random fallback" calibration path is broken.
3. Mixed-case model names silently take the wrong code path.
4. The quantization output no longer matches what compilation reads (merge-order risk).
Architecture / direction concernsA. The dependency direction is inverted, so Phase 3 can't delete the legacy flow as planned.
B. Legacy logic is copied rather than shared, with no parity check, so the two copies will drift.
C. Model-family branching is spread across
D. The LLM/MLLM adapter split doesn't separate the logic it is meant to separate.
E. There are behavioral differences from legacy that A2's parity test will hit.
F. One module is used to export every graph, picked by dict order.
G. Phase 1 interfaces changed in breaking ways, and the default adapters stop working.
H. Hidden state is passed around, against the CLAUDE.md rule on explicit state.
Code Quality
Testing
Your review comments: verdicts
NXP Bamboo failure: that CI is NXP-hosted, and I can't see its logs. This PR changes only files under RecommendationRequest Changes
|
Summary
First PR of Phase 2 Stream A. This PR makes the GenAI Pipeline runnable for
LLM and MLLM models through the front half of the flow:
examples/and the legacyllama.pypath remain unchanged. Compilation anddevice inference remain in Stream B / PR-A2.
What's Included
CLI Entry Point:
cli.pyAdds a GenAI Pipeline CLI that resolves model, dataset, and quantization options, builds
PipelineContext, and invokes model preparation and quantization. FP16 omits quantization; QAT, embedding quantization, and attention sink are rejected for now.The rejected feature will be raised in following PR.
Registry-Backed Model lookup:
model_lookup.pyMaps a user-facing model name to the model-specific inputs consumed by the pipeline. Callers resolve one registry config, then derive from that same entry:
decoder decode/quantize graphs, optional prefill graphs, and multimodal
embedding variants.
module G2G transforms, plus the Hugging Face state-dict loader when needed.
recipe classes, and LLM or MLLM loader/quantizer adapters.
This centralizes model-family branching at the registry boundary; the model preparation and quantization stages consume only the derived component and graph keyed configuration.
LLM and Multimodal LLM model components:
model_components/Adds pipeline import surfaces for decoder, token embedding, and encoder components over existing
examplesimplementations.Module G2G Transforms:
source_transform/:Collects existing
llama.pytransformations into reusable pipeline functions:LLM and MLLM Model Loading and Preparation:
strategies/model_preparation/:ExecuTorchModelPreparationStrategyimplements the shared preparation flow:extra_options.LLMLoaderAdapterhandles text decoders;MLLMLoaderAdapterhandles multimodal models. The default adapter owns shared tokenizer and transforms logic.Dataset stack:
datasets/Adds typed dataset options, lookup, loaders, calibration adapters, and LLM/MLLM collators. It supports random fallback data, lm-eval samples, JSON messages, Hugging Face chat datasets, and multimodal messages.
generate_calibration_data()now receivesexample_inputsexplicitly because collators require the calibration-graph signature.Quantization:
strategies/quantization/:ExecuTorchQuantizationStrategyimplements native PTQ for component and graph keyed models:Decoder and token embedding use separate quantize and deploy graphs; encoders use one shared graph. Encodings are propagated before lowering.
Quantizer creation belongs to the strategy; LLM/MLLM adapters own only model-family-specific export, preparation, conversion, and calibration.
Quantization Helpers:
quant_utilities.py:Provides QNN quantizer construction, recipe application, QDQ saving, logits/KV-cache attributes, and encoding propagation. KV-cache override is enabled only when
n_cache_layersis provided.Quantization Output
QuantizationOutputConfignow returns component and graph keyedGraphBundleobjects. Each bundle carries the quantized graph module, export inputs, metadata, and optional quantized IO dtypes.Quantize-only graphs are removed after their encodings have been propagated, so downstream compilation sees only deployment graphs.
Tests
Covers CLI/config construction, source transforms, datasets and collators, loader adapters, model preparation, quantization adapters and strategy, and preparation/quantization stage integration.
PR Review Checklist
TokenizerWrapper,torch.export, PT2E, QNN quantizer construction, datasetloading, and calibration execution sit behind adapters or lookup-created
callables. - Yes.
examples/is not modified andllama.pyremains the reference path. - YesRelated PRs
Phase 2 flow. PR-A1 and PR-B1 are independent and can land in either order.
PR-B2 depends on PR-B1. PR-A2 depends on PR-A1 and PR-B2, and closes Phase 2 by wiring the runner and adding the parity test that gates Phase 3 deletion of the legacy flow.
Phase 1 (merged):
Phase 2:
dataset-backed calibration, component/graph-aware quantization, and encoding
reconciliation. [This PR].
Depends on PR-B1.
test. Depends on PR-A1 and PR-B2.
Test plan
Run only tests added in this PR:
Result:
Run all genai_pipeline tests:
Result:
Run all tests with coverage:
Result:
Confirm the legacy flow is unaffected:
Result: