Skip to content

Quantized inference support #16

Description

@AIWintermuteAI

From the README I see that it is currently possible to quantize the models:

the loader keeps the packed weights and lets ggml_mul_mat dequantize at compute, so the file just loads and runs like the bf16 one

However the inference still runs in BF16, with additional overhead from on-the-fly conversion?
Since llama.cpp supports quantized inference for language models and also for vision models, do you have adding support for quantized inference on the roadmap? It would be a great addition for CPU based inference, which is otherwise very slow, especially so with BF16.

Activity

  1. PushpakAg commented on Sep 17, 2026

    @PushpakAg

    Interesting point about the overhead from on-the-fly conversion with BF16. Curious if there's a specific reason for not having quantized inference support yet.

  2. hungho77 commented on Oct 5, 2026

    @hungho77
    Collaborator

    Hi! Quantized inference is now implemented in #33 (feat/quantized-inference).

    Quick note on the README: the stock Q8_0/Q4_0 repack isn't BF16 compute on CPU. ggml runs int8 dot products on the packed blocks, but it's uncalibrated and the action expert stays float.

    #33 adds FoldQuant W8A8 / W4A4 for GR00T N1.5/N1.6/N1.7 and π0.5: INT8/INT4 weights, per-token activation quantization, real integer GEMMs. It runs on CUDA, SYCL (Intel GPUs) and OpenVINO (CPU/GPU); other backends fall back to dequantized weights. The checkpoints come from my other project, FoldQuantVLA (paper, models on HF).

    π0.5 LIBERO numbers:

    • Arc B390 iGPU (SYCL): bf16 326 ms → W4A4 291 ms
    • Core Ultra X7 CPU (OpenVINO): bf16 6.6 s / 16.5 GB → W8A8 3.7 s / 8.6 GB

    For CPU, OpenVINO is the fast path for now. Native ggml-CPU kernels aren't there yet, so I'll leave this issue open until they are.

    Feedback on the PR is very welcome, and if FoldQuantVLA is useful to you, a ⭐ would mean a lot!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions