vla.cpp runs Vision-Language-Action (VLA) policies on the ggml/llama.cpp runtime. Every model is a self-contained GGUF that the engine loads, detects, and drives on CPU, CUDA, Metal, SYCL, OpenVINO, OpenCL or Hexagon. This page is the map; the source is the detail.
src/model.h- public API:model_load,predict,model_config,last_stats.src/arch.h- theArchenum, theModelArchBaseinterface, and one*_createfactory per architecture.src/model.cpp- loads a checkpoint, detects the architecture from its GGUF keys (or safetensors namespace), and dispatches to the matching factory.src/models/*.cpp- one translation unit per architecture. Each owns its ggml contexts, vision tower, weights, and compute graph.src/gguf_reader.h- the shared GGUF reader (metadata, tensor bytes, on-demand embedding rows).src/loader.huploads weights at the resident dtype.src/layers/- ggml building blocks: linear, attention, FFN, norms, RoPE, time embeddings.src/modules/- parts shared between archs: the SigLIP, Qwen3-VL and DINOv2+SigLIP vision towers, the Gemma action expert, the Qwen3 LM, the DiT head, and image preprocessing (preprocess.h: view checks, CHW normalization).src/scratch_ctx.h- compute contexts andgraph_cache, reused across calls.src/backend.h- backend selection;src/backend_fallback.cppruns the ops an accelerator rejects on the CPU.src/tokenizer.h- SentencePiece tokenizers stored in the GGUF, forvla-cli --text.include/vla.h,src/vla_c_api.cpp- the C ABI (libvla).src/serving/-vla-server(ZeroMQ + protobuf, action prediction),vlm-server(chat),vla-cli(one-shot inference) andvla-bench(timing).src/kernels/bitvla/- custom 1.58-bit ternary CUDA kernels for BitVLA.src/foldquant.h,src/kernels/foldquant/,src/cuda/- FoldQuant INT8 linears and the in-tree CUDA kernels behind the ggml extension hook.
A forward pass has two stages that most architectures share.
-
Prefix. Camera views go through a vision tower, language tokens through the embedding table, and proprioception through a small projection. Concatenated, they form the prefix that the language backbone attends over (bidirectionally, minus any padded language tokens).
-
Action head. A smaller expert reads the prefix and produces an action chunk of shape
[num_steps, max_action_dim], where only the firstreal_action_dimcolumns carry values and the rest are zero padding. The head comes in four flavours:- Flow-matching expert (SmolVLA, pi0, pi0.5): integrates a velocity field with
Euler steps from noise at
t=1to the action att=0. SmolVLA alternates self-attention among action tokens with cross-attention to the prefix. - DiT (GR00T N1.5/1.6/1.7, Evo-1, VLA-JEPA): a diffusion transformer head, also integrated with flow-matching Euler steps.
- Diffusion (Octo): DDPM with a small MLP denoiser over the transformer readout.
- Parallel decode (BitVLA, OpenVLA-OFT, VLA-Adapter): OpenVLA-OFT-style bidirectional decode of the action tokens in a single pass. TurboVLA is also one pass: an ACT decoder whose learned queries read the fused tokens.
- Flow-matching expert (SmolVLA, pi0, pi0.5): integrates a velocity field with
Euler steps from noise at
Actions leave predict in world units when Config::denormalized is true (the
default). GR00T N1.5/1.6/1.7 and VLA-JEPA return normalised actions and the caller
un-normalises.
Every architecture ships its vision tower inside the one GGUF. model_load and the
binaries still take an mmproj path so older command lines work, but every arch
ignores it.
llama.cpp is fetched by CMake FetchContent and pinned by VLA_LLAMA_TAG in
CMakeLists.txt. Each arch picks a resident dtype for its GEMM weights
(--weight-dtype overrides it), and a GGUF can be repacked to Q8_0/Q4_0 with
scripts/quantize_gguf.py.
The loader keeps packed weights packed. On CPU and CUDA, ggml quantizes the
activations to 8 bits and runs int8 dot products on the blocks.
A FoldQuant GGUF (see QUANTIZATION.md)
carries INT8 or INT4 codes plus sidecar scales instead. src/foldquant.h declares those
sites, src/layers/fq_linear.h turns each into two GGML_OP_CUSTOM nodes, the
CPU backend runs the reference in src/foldquant_ref.cpp, and on CUDA the
src/kernels/foldquant/ integer kernels claim the same nodes through the ggml
extension hook (src/cuda/), and on OpenVINO src/openvino/foldquant_ov.cpp
translates them into OpenVINO ops. Every other backend reads the sites back as
float weights through WeightLoader::as_float and runs the arch's float path. CPU thread count scales to the machine core count;
the GPU backends run the towers and the transformer on the device.
Extend the Arch enum, declare a *_create factory in arch.h, implement it under
src/models/, wire detection and dispatch in src/model.cpp, and add a converter in
scripts/. Reuse gguf_reader.h, loader.h, layers/ and modules/ rather than
copying them. CONTRIBUTING.md lists the exact sites.