Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
85 changes: 85 additions & 0 deletions .github/workflows/build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,8 @@ jobs:
run: |
cp -r build/_deps/llama-src /tmp/llama-patchtest
python3 scripts/patch_ggml_cuda_ext_hook.py /tmp/llama-patchtest
python3 scripts/patch_ggml_sycl_ext_hook.py /tmp/llama-patchtest
python3 scripts/patch_ggml_sycl_ext_hook.py /tmp/llama-patchtest # idempotent
python3 scripts/patch_ggml_openvino.py /tmp/llama-patchtest
python3 scripts/patch_ggml_openvino.py /tmp/llama-patchtest # idempotent

Expand Down Expand Up @@ -88,6 +90,89 @@ jobs:
-DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined -DVLA_BUILD_TESTS=ON -DVLA_WERROR=ON
cmake --build build -j"$(nproc)"

# Intel GPUs through oneAPI SYCL. The runner has no GPU, so ggml-sycl itself
# cannot start; ctest still runs everything, and test_foldquant_sycl_op drives
# the FoldQuant kernels straight through the extension hook on the OpenCL CPU
# device that the DPC++ compiler package ships, against the CPU reference.
build-sycl:
runs-on: ubuntu-24.04
steps:
- name: free disk (oneAPI + a SYCL build need ~15 GB)
run: sudo rm -rf /usr/share/dotnet /usr/local/lib/android /opt/ghc /opt/hostedtoolcache/CodeQL
- uses: actions/checkout@v7
- name: deps
run: |
wget -qO- https://apt.repos.intel.com/intel-gpg-keys/GPG-PUB-KEY-INTEL-SW-PRODUCTS.PUB \
| sudo gpg --dearmor -o /usr/share/keyrings/oneapi-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/oneapi-archive-keyring.gpg] https://apt.repos.intel.com/oneapi all main" \
| sudo tee /etc/apt/sources.list.d/oneAPI.list
sudo apt-get update -qq
sudo apt-get install -y -qq --no-install-recommends \
build-essential cmake git ca-certificates pkg-config python3 \
libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compiler \
intel-oneapi-compiler-dpcpp-cpp-2026.1 intel-oneapi-mkl-devel-2026.1 intel-oneapi-dnnl-devel-2026.0
- name: read llama.cpp pin
id: pin
run: echo "tag=$(bash scripts/llama_tag.sh)" >> "$GITHUB_OUTPUT"
- uses: actions/cache@v6
with:
path: build/_deps
key: llama-${{ steps.pin.outputs.tag }}-${{ runner.os }}-sycl-${{ hashFiles('scripts/patch_ggml_sycl_ext_hook.py') }}
- name: build + ctest (OpenCL CPU device)
shell: bash
env:
ONEAPI_DEVICE_SELECTOR: opencl:cpu
VLA_FQ_TEST_REQUIRE_DEVICE: 1
run: |
source /opt/intel/oneapi/setvars.sh
sycl-ls
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_SYCL=ON \
-DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DVLA_BUILD_TESTS=ON -DVLA_WERROR=ON
cmake --build build -j"$(nproc)"
ctest --test-dir build --output-on-failure

# Intel CPUs, GPUs and NPUs through OpenVINO. The CPU plugin runs on the
# runner, so test_foldquant_ov_op executes the translated FoldQuant graph and
# checks it against the CPU reference.
build-openvino:
runs-on: ubuntu-24.04
env:
OV_ARCHIVE: openvino_toolkit_ubuntu24_2026.4.0.22959.99c81491cc3_x86_64
steps:
- uses: actions/checkout@v7
- name: deps
run: |
sudo apt-get update -qq
sudo apt-get install -y -qq --no-install-recommends \
build-essential cmake git ca-certificates pkg-config python3 \
libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compiler \
opencl-clhpp-headers ocl-icd-opencl-dev opencl-headers
- uses: actions/cache@v6
id: ov
with:
path: /opt/openvino
key: ${{ env.OV_ARCHIVE }}
- name: OpenVINO runtime
if: steps.ov.outputs.cache-hit != 'true'
run: |
curl -fsSL "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/linux/${OV_ARCHIVE}.tgz" \
| tar xz -C /tmp
sudo mv "/tmp/${OV_ARCHIVE}" /opt/openvino
- name: read llama.cpp pin
id: pin
run: echo "tag=$(bash scripts/llama_tag.sh)" >> "$GITHUB_OUTPUT"
- uses: actions/cache@v6
with:
path: build/_deps
key: llama-${{ steps.pin.outputs.tag }}-${{ runner.os }}-openvino-${{ hashFiles('scripts/patch_ggml_openvino.py') }}
- name: build + ctest (CPU plugin)
shell: bash
run: |
source /opt/openvino/setupvars.sh
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_OPENVINO=ON -DVLA_BUILD_TESTS=ON -DVLA_WERROR=ON
cmake --build build -j"$(nproc)"
ctest --test-dir build --output-on-failure

build-backends:
runs-on: ${{ matrix.os }}
strategy:
Expand Down
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,30 @@ Notable changes to vla.cpp. Format loosely follows [Keep a Changelog](https://ke

## [Unreleased]

### Added

- **FoldQuant W8A8 / W4A4 inference** for GR00T N1.5 / N1.6 / N1.7 and π0.5. A
FoldQuant GGUF carries the language backbone and the action
module as INT8 or INT4 codes with per-row scales in a block-Hadamard,
SmoothQuant-folded frame; activations are quantized per token and the
projections run as integer GEMMs (`src/kernels/foldquant/` on CUDA,
`src/sycl/vla_sycl_foldquant.cpp` on SYCL with oneDNN's int8 matmul, both
bit-identical to the reference in `src/foldquant_ref.cpp` that the CPU runs;
OpenVINO ops on the OpenVINO CPU and GPU plugins with INT8/INT4 weight
constants). On every other backend
(Metal, Hexagon, OpenCL, the OpenVINO NPU) the sites are read back as
float weights with FoldQuant's rounding and run the stock float path;
`VLA_FQ_DEQUANT=1` does the same on CUDA or CPU. The format and arithmetic are
in `docs/QUANTIZATION.md`.
- `scripts/convert_quantized_model_to_gguf.py` converts a FoldQuantVLA quantized
model (its quantized checkpoint or its earlier fake-quant state) to a FoldQuant
GGUF with no calibration and nothing re-rounded; `--check-onnx` byte-compares
every site against the TensorRT plugin graphs. The family converters expose
`convert(ckpt, out, writer_factory=...)` for it.
- `scripts/foldquant_fake_export.py` (uncalibrated file for bring-up),
`scripts/inspect_gguf_quant.py` (contract check) and `scripts/foldquant_ref.py`
(numpy reference).

## [0.4.0] - 2026-09-30

### Added
Expand Down
91 changes: 88 additions & 3 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,12 @@ set(LLAMA_BUILD_SERVER OFF CACHE BOOL "" FORCE)
set(LLAMA_BUILD_EXAMPLES OFF CACHE BOOL "" FORCE)
set(LLAMA_BUILD_TESTS OFF CACHE BOOL "" FORCE)
set(_vla_llama_patch "")
if(GGML_SYCL AND NOT GGML_CUDA)
# The hook FoldQuant's SYCL kernels (src/sycl/vla_sycl_foldquant.cpp) run behind.
find_package(Python3 COMPONENTS Interpreter REQUIRED)
set(_vla_llama_patch PATCH_COMMAND ${Python3_EXECUTABLE}
${CMAKE_CURRENT_SOURCE_DIR}/scripts/patch_ggml_sycl_ext_hook.py <SOURCE_DIR>)
endif()
if(GGML_CUDA)
find_package(Python3 COMPONENTS Interpreter REQUIRED)
set(_vla_llama_patch PATCH_COMMAND ${Python3_EXECUTABLE}
Expand Down Expand Up @@ -117,6 +123,17 @@ FetchContent_Declare(llama
)
FetchContent_MakeAvailable(llama)

# FoldQuant on OpenVINO: the translator for FoldQuant's two custom nodes is
# in-tree and compiled into ggml's OpenVINO backend, which reaches it through the
# hook scripts/patch_ggml_openvino.py adds (op table entry, supports_op).
if(GGML_OPENVINO)
target_sources(ggml-openvino PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/src/openvino/foldquant_ov.cpp)
target_include_directories(ggml-openvino PRIVATE
${llama_SOURCE_DIR}/ggml/src/ggml-openvino
${llama_SOURCE_DIR}/ggml/src
${CMAKE_CURRENT_SOURCE_DIR}/src)
endif()

# Drop the whole fetched tree from the default build: llama.cpp's ~20 tool
# binaries and their impl libraries are dead weight here. What we link (llama,
# ggml, mtmd, llama-common) is still built, pulled in as a dependency. Walking
Expand Down Expand Up @@ -189,6 +206,15 @@ if(VLA_SPM)
endif()
endif()

# oneAPI's icpx (the compiler of a SYCL build) defaults to -fp-model=fast, which
# reassociates sums and drops IEEE division and NaN semantics; vla.cpp's code and
# tests are written for what GCC, Clang and MSVC do by default. Set here, after
# the dependencies, so ggml keeps its own flags. The FoldQuant sources then add
# -ffp-contract=off on top, which icpx reports as an override; that is the intent.
if(CMAKE_CXX_COMPILER_ID STREQUAL "IntelLLVM")
add_compile_options(-fp-model=precise -Wno-overriding-option)
endif()

# llama.cpp turns BUILD_SHARED_LIBS on, so these follow it. That is harmless on
# ELF, which exports everything, but a Windows DLL exports nothing unannotated:
# the import library comes out empty and every consumer fails to link. Nothing
Expand All @@ -203,6 +229,8 @@ add_library(vla_core ${VLA_CORE_LIB_TYPE}
src/model.cpp
src/loader.cpp
src/options.cpp
src/foldquant.cpp
src/foldquant_ref.cpp
src/modules/action_expert.cpp
src/modules/dit_head.cpp
src/modules/encoder.cpp
Expand Down Expand Up @@ -231,6 +259,10 @@ target_include_directories(vla_core
)
# The VLA archs call no llama_* API; only vlm_core needs llama.
target_link_libraries(vla_core PUBLIC ggml)
# The FoldQuant CPU reference must match the CUDA and SYCL kernels bit for bit,
# so no FMA contraction on either side (aarch64 GCC contracts by default; under
# icpx, after the project-wide -fp-model=precise, which would allow it).
set_source_files_properties(src/foldquant.cpp src/foldquant_ref.cpp PROPERTIES COMPILE_OPTIONS "-ffp-contract=off")
if(VLA_SPM)
target_include_directories(vla_core PRIVATE ${sentencepiece_SOURCE_DIR}/src)
target_link_libraries(vla_core PRIVATE sentencepiece-static ${VLA_SPM_PROTOBUF})
Expand Down Expand Up @@ -268,8 +300,36 @@ if(GGML_CUDA)
target_compile_definitions(vla_core PUBLIC VLA_BITVLA_CUDA_KERNELS)
target_include_directories(vla_core PUBLIC ${CUDAToolkit_INCLUDE_DIRS})

# FoldQuant kernels: their own archive, device-linked on its own like bitvla
# above, because they compile with -fmad=false (bit-exact against the CPU
# reference in src/foldquant_ref.h) and no other archive does.
add_library(vla_fq_kernels STATIC
src/kernels/foldquant/fq_prologue.cu
src/kernels/foldquant/fq_gemm_i8.cu
src/kernels/foldquant/fq_gemm_mma.cu
)
set_target_properties(vla_fq_kernels PROPERTIES
CUDA_SEPARABLE_COMPILATION ON
CUDA_RESOLVE_DEVICE_SYMBOLS ON
POSITION_INDEPENDENT_CODE ON
)
target_include_directories(vla_fq_kernels PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/src)
target_compile_features(vla_fq_kernels PRIVATE cxx_std_17)
target_compile_options(vla_fq_kernels PRIVATE
$<$<COMPILE_LANGUAGE:CUDA>:-O3 -Xptxas=-O3 -fmad=false>
)
target_link_libraries(vla_fq_kernels PUBLIC CUDA::cudart)

# ggml-cuda link: the hook resolves ggml_cuda_ext_forward out of it.
add_library(vla_cuda_ops STATIC src/cuda/vla_cuda_bf16.cu)
add_library(vla_cuda_ops STATIC
src/cuda/vla_cuda_ext.cu
src/cuda/vla_cuda_bf16.cu
src/cuda/vla_cuda_foldquant.cu
)
# The staging shim in vla_cuda_foldquant.cu runs the CPU reference on the
# host; it must round exactly like vla_core's copy.
set_source_files_properties(src/cuda/vla_cuda_foldquant.cu PROPERTIES
COMPILE_OPTIONS "-Xcompiler=-ffp-contract=off")
target_include_directories(vla_cuda_ops PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/src
${llama_SOURCE_DIR}/ggml/include
Expand All @@ -283,7 +343,7 @@ if(GGML_CUDA)
target_compile_options(vla_cuda_ops PRIVATE
$<$<COMPILE_LANGUAGE:CUDA>:-O3 -Xptxas=-O3>
)
target_link_libraries(vla_cuda_ops PUBLIC CUDA::cublas CUDA::cudart ggml-cuda)
target_link_libraries(vla_cuda_ops PUBLIC CUDA::cublas CUDA::cudart ggml-cuda vla_fq_kernels)
target_link_libraries(vla_core PRIVATE vla_cuda_ops)
endif()

Expand All @@ -296,6 +356,29 @@ if(GGML_SYCL AND NOT GGML_CUDA)
"See docs/backend/sycl.md.")
endif()
target_compile_definitions(vla_core PUBLIC GGML_USE_SYCL)

# FoldQuant kernels for Intel GPUs, behind the hook the patch above adds. Built
# without FMA contraction, and with the device JIT told to round division and
# sqrt correctly (the GPU driver approximates both otherwise), so the codes
# match the CPU reference bit for bit. The JIT option is a link option: it
# rides in the executable's device image. One image per kernel, so a device
# JIT-compiles only what it launches (not the XMX kernel on a device without
# XMX: the OpenCL CPU device's compiler has crashed on the whole module).
target_sources(vla_core PRIVATE src/sycl/vla_sycl_foldquant.cpp)
set_source_files_properties(src/sycl/vla_sycl_foldquant.cpp PROPERTIES
COMPILE_OPTIONS "-fsycl;-ffp-contract=off")
target_link_options(vla_core PUBLIC -fsycl -fsycl-device-code-split=per_kernel
"SHELL:-Xsycl-target-backend=spir64 \"-cl-fp32-correctly-rounded-divide-sqrt\"")
target_compile_definitions(vla_core PRIVATE VLA_FQ_SYCL)
# Prefill-sized FoldQuant GEMMs go to oneDNN's int8 matmul when ggml-sycl
# was built with it (GGML_SYCL_DNN, the default).
if(GGML_SYCL_DNN)
find_package(DNNL QUIET)
if(DNNL_FOUND)
target_link_libraries(vla_core PRIVATE DNNL::dnnl)
target_compile_definitions(vla_core PRIVATE VLA_FQ_SYCL_DNNL)
endif()
endif()
endif()

# Precedence matches the ladder in src/backend.h.
Expand All @@ -307,6 +390,8 @@ if(GGML_OPENVINO AND NOT GGML_CUDA AND NOT GGML_SYCL AND NOT GGML_METAL)
# Only tells the archs which branch of the backend.h ladder to compile; the
# toolkit was located before the fetch above.
target_compile_definitions(vla_core PUBLIC GGML_USE_OPENVINO)
# FoldQuant nodes run natively through the translator compiled into ggml-openvino.
target_compile_definitions(vla_core PRIVATE VLA_FQ_OPENVINO)
endif()

# Hexagon and OpenCL reject some ops the archs build; the wrapper in
Expand Down Expand Up @@ -461,7 +546,7 @@ if(VLA_BUILD_SERVER)
endif()
# The CUDA targets were the only first-party code compiled without warnings.
if(GGML_CUDA)
list(APPEND VLA_FIRST_PARTY_TARGETS bitvla_cuda_kernels vla_cuda_ops)
list(APPEND VLA_FIRST_PARTY_TARGETS bitvla_cuda_kernels vla_cuda_ops vla_fq_kernels)
endif()

option(VLA_WERROR "Treat warnings in first-party code as errors" OFF)
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -225,6 +225,7 @@ Wiring, recording, training and queue sizing are in the
| [docs/USAGE.md](docs/USAGE.md) | `vla-cli`, `vla-server`, `vla-bench`: prompt tokenization, `-hf` tags, runtime flags, environment variables |
| [docs/EVAL.md](docs/EVAL.md) | Installing LIBERO and SimplerEnv, running the eval clients against `vla-server` |
| [docs/MODELS.md](docs/MODELS.md) | Converting safetensors checkpoints to GGUF, quantizing to Q8_0/Q4_0 |
| [docs/QUANTIZATION.md](docs/QUANTIZATION.md) | FoldQuant W8A8 / W4A4: the GGUF contract, the integer arithmetic, converting a FoldQuantVLA quantized model |
| [docs/DOCKER.md](docs/DOCKER.md) | Building and running the eval in containers |
| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | Engine design: layers, the prediction path, backends, adding an architecture |
| [docs/backend/](docs/backend) | Per-backend build and run notes: [SYCL](docs/backend/sycl.md), [OpenVINO](docs/backend/ov.md), [Metal](docs/backend/metal.md), [Hexagon](docs/backend/hexagon.md), [Hexagon on Windows](docs/backend/hexagon-windows.md), [WSL2](docs/backend/wsl.md) |
Expand Down
16 changes: 13 additions & 3 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,8 @@ source is the detail.
- `src/serving/` - `vla-server` (ZeroMQ + protobuf, action prediction), `vlm-server`
(chat), `vla-cli` (one-shot inference) and `vla-bench` (timing).
- `src/kernels/bitvla/` - custom 1.58-bit ternary CUDA kernels for BitVLA.
- `src/foldquant.h`, `src/kernels/foldquant/`, `src/cuda/` - FoldQuant INT8 linears
and the in-tree CUDA kernels behind the ggml extension hook.

## The prediction path

Expand Down Expand Up @@ -70,9 +72,17 @@ llama.cpp is fetched by CMake `FetchContent` and pinned by `VLA_LLAMA_TAG` in
(`--weight-dtype` overrides it), and a GGUF can be repacked to Q8_0/Q4_0 with
`scripts/quantize_gguf.py`.
The loader keeps packed weights packed. On CPU and CUDA, ggml quantizes the
activations to 8 bits and runs int8 dot products on the blocks. CPU thread count
scales to the machine core count; the GPU backends run the towers and the
transformer on the device.
activations to 8 bits and runs int8 dot products on the blocks.

A FoldQuant GGUF (see [QUANTIZATION.md](QUANTIZATION.md))
carries INT8 or INT4 codes plus sidecar scales instead. `src/foldquant.h` declares those
sites, `src/layers/fq_linear.h` turns each into two `GGML_OP_CUSTOM` nodes, the
CPU backend runs the reference in `src/foldquant_ref.cpp`, and on CUDA the
`src/kernels/foldquant/` integer kernels claim the same nodes through the ggml
extension hook (`src/cuda/`), and on OpenVINO `src/openvino/foldquant_ov.cpp`
translates them into OpenVINO ops. Every other backend reads the sites back as
float weights through `WeightLoader::as_float` and runs the arch's float path. CPU thread count scales to the machine core count;
the GPU backends run the towers and the transformer on the device.

## Adding an architecture

Expand Down
26 changes: 26 additions & 0 deletions docs/MODELS.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,3 +43,29 @@ compute, not BF16.
a bigger cut (`--type` also takes `Q4_1`, `Q5_0`, `Q5_1`).
Embeddings, the output head, norms and the action expert stay float; pass `--vision` to
pack the vision tower too (smaller, but more accuracy loss).

### FoldQuant W8A8 / W4A4

The stock repack keeps the action head float and rounds the LM weights block by
block with no calibration. A FoldQuant GGUF ships the language backbone and the
action head as INT8 or INT4 codes in a Hadamard-rotated, SmoothQuant-folded frame
with per-row scales, calibrated by FoldQuantVLA; vla.cpp quantizes the activations
per token (to 4 bits for W4A4) and runs the GEMMs on the integer tensor cores. It loads like any other checkpoint on every backend: CUDA and SYCL (Intel GPUs) run
integer kernels, CPU the exact reference and OpenVINO (CPU and GPU) the same
arithmetic as OpenVINO ops with INT8/INT4 weight constants; the others read the sites back as float
weights with FoldQuant's rounding (weight-only quantization, bf16-GGUF speed); the format
and the arithmetic are in [QUANTIZATION.md](QUANTIZATION.md).

A calibrated arm saved as a quantized model by
[FoldQuantVLA](https://github.com/VinRobotics/FoldQuantVLA) converts to that file with
no calibration and nothing re-rounded (GR00T N1.5 / N1.6 / N1.7 and π0.5):

```bash
python scripts/convert_quantized_model_to_gguf.py --quantized-model <quantized model dir> --out model-fq.gguf
./build/vla-server model-fq.gguf --bind tcp://*:5556

# uncalibrated stand-in for bring-up and benchmarks (rotation + per-row INT8, no SmoothQuant)
python scripts/foldquant_fake_export.py --in model-bf16.gguf --out model-fq.gguf
python scripts/inspect_gguf_quant.py model-fq.gguf
./build/vla-bench --ckpt model-fq.gguf --images 1 --size 256 --tokens 16
```
Loading
Loading