A C++ inference engine for Vision-Language-Action (VLA) models, built on llama.cpp.
It runs the open VLA policies - SmolVLA, π0, BitVLA, Evo-1, GR00T N1.5/1.6/1.7 and more -
under one runtime, each packaged as a single self-contained GGUF that needs no Python or
PyTorch at inference time. The binaries drive robots on CPU, Apple Silicon, CUDA -
from consumer GPUs down to Jetson-class boards - Intel GPUs and NPUs via
SYCL and OpenVINO, or Qualcomm Snapdragon CPUs, Adreno GPUs and Hexagon NPUs
via OpenCL and the Hexagon backend.
Learn vla.cpp walks through the engine design and how each policy is implemented on ggml.
- CMake ≥ 3.22
- A C++17 compiler (GCC 11+ or Clang 14+)
- CUDA 12.x (optional - required only for CUDA GPU builds)
- Intel oneAPI 2025.x + GPU compute runtime (optional - only for Intel GPU builds, see docs/backend/sycl.md)
- OpenVINO 2026.x runtime (optional - only for Intel CPU/GPU/NPU builds via OpenVINO, see docs/backend/ov.md)
libzmq3-dev,cppzmq-dev,libprotobuf-dev,protobuf-compiler
sudo apt-get install -y libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compilerIdentify your machine CUDA architecture:
| GPU family | Example cards | CUDA_ARCHITECTURE |
|---|---|---|
| Ampere (Jetson) | Orin Nano, Orin NX | 87 |
| Ampere (consumer) | RTX 30-series, A40 | 86 |
| Ada Lovelace | RTX 40-series, L40 | 89 |
| Hopper | H100, H200 | 90 |
| Blackwell (consumer) | RTX 50-series | 120 |
| Blackwell (datacenter) | B100, B200, GB200 | 100 |
Then configure and build. CMake fetches and pins llama.cpp automatically (no patch, no submodule):
# CPU build:
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
# CUDA build (set CMAKE_CUDA_ARCHITECTURES for your GPU):
cmake -B build \
-DGGML_CUDA=ON \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_CUDA_ARCHITECTURES=$CUDA_ARCHITECTURE
cmake --build build -j$(nproc)If CMake cannot find CUDA, point the environment at it explicitly:
export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATHCheck docs/backend for compiling vla.cpp on other platforms.
WSL2, Apple Silicon, and Intel GPU are all tested.
To build and run in containers instead, see docs/DOCKER.md.
Once the binaries are built, run one CPU prediction without a server or simulator:
pip install -U "huggingface_hub[cli]" transformers
# -hf fetches and caches the checkpoint (under $VLA_CACHE, default ~/.cache/vla)
./build/vla-cli -hf vrfai/smolvla-libero-gguf \
--image assets/front.jpg --text "pick up the black bowl" --pretty
# or point at a file you already have
./build/vla-cli --ckpt models/smolvla/smolvla-libero.gguf \
--image assets/front.jpg --text "pick up the black bowl" --prettyvla-cli runs a single prediction without a server or simulator: give it a model,
an image, and an instruction, and it prints the action chunk. Handy for
smoke-testing a GGUF or scripting a quick inference.
There is no tokenizer in the C++ core, so --text calls
scripts/tokenize_prompt.py with the tokenizer the architecture was trained on
(VLA_PYTHON picks the interpreter, VLA_TOKENIZE_SCRIPT the script). Pass
--tokens 1,100,200,2 instead if you already have ids.
--pretty prints one action row per line;
--state sets proprioception (defaults to zeros).
For the design overview see docs/ARCHITECTURE.md, for the long-running path see Running the server, and for the other checkpoints see Roadmap.
The rest of this README refers to a few shell variables:
export VLA_GGUF=models/smolvla/smolvla-libero.gguf # the checkpoint to serve
export VLA_ARCH=smolvla # client-side arch preset, see --helpThe eval scaffold under eval/ supports two simulators end-to-end. Each setup script bootstraps an isolated Python 3.10 uv venv next to itself and clones the upstream sim repo. Both require uv on PATH.
bash eval/sim/libero/setup_libero.shClones LIBERO into eval/sim/libero/LIBERO/, creates eval/sim/libero/libero_uv/.venv/, and pins compatible versions of torch, lerobot, transformers, and gymnasium.
bash eval/sim/simpler/setup_SimplerEnv.shClones SimplerEnv (and its nested ManiSkill2_real2sim) into eval/sim/simpler/SimplerEnv/, creates eval/sim/simpler/simpler_uv/.venv/.
vla-server loads the model once at startup and answers ZeroMQ REQ/REP requests synchronously.
./build/vla-server "$VLA_GGUF"When ready, the server prints:
vla-server: bound to tcp://*:5555. ready.
Use --bind to change the address and port. Stop the server with Ctrl-C.
vla-server also takes -hf user/repo[:file.gguf] in place of a checkpoint path.
Precision flags (vla-server --help for the full list).
The fastest configuration per model, with measured latency and success rate, is
in CHANGELOG.md:
--weight-dtype f32|bf16- resident dtype for GEMM weights.--act-dtype f32|bf16- activation dtype; needs CUDA and bf16 weights.--flash-attn- faster on the larger towers, but changes numerics.--mm-prec default|f32- matmul accumulation precision.
Environment knobs that apply to every arch:
VLA_N_THREADS- CPU backend thread count, default core count capped at 16.VLA_DEVICE- GPU ordinal for CUDA and SYCL builds, default 0.VLA_CACHE- where-hfstores checkpoints, default~/.cache/vla.
eval/client/ ships an end-to-end LIBERO benchmark runner that drives vla-server directly over the protobuf protocol. Make sure the LIBERO venv from Install simulators is set up first.
With vla-server already running:
source eval/sim/libero/libero_uv/.venv/bin/activate
python eval/client/run_sim_client_direct.py \
--task libero_object --task-id 0 --n-episodes 1 \
--output-dir /tmp/libero_outputs \
--arch "$VLA_ARCH"The GR00T models need two extras:
- client side:
--stats-json /path/to/dataset_statistics.json - server side:
VLA_GR00T_EMBODIMENT(new_embodimentfor N1.5,libero_pandafor N1.6,libero_simfor N1.7).
So far only GR00T-N1.6 is wired (the gr00t-n1d6-bridge checkpoint with the oxe_widowx embodiment). Start vla-server on port 5566 with oxe_widowx embodiment:
VLA_GR00T_EMBODIMENT=oxe_widowx \
./build/vla-server "$GR00T_N1D6_GGUF"Then drive it from the SimplerEnv venv (set up via Install simulators):
source eval/sim/simpler/simpler_uv/.venv/bin/activate
python eval/client/run_simpler_client_direct.py \
--arch gr00t_n1_6 \
--task-id oxe_widowx/widowx_spoon_on_towel --n-episodes 1 \
--embodiment oxe_widowx --image-size 252 \
--stats-json "$VLA_STATS_JSON"Each model ships as a single self-contained GGUF. To convert a HuggingFace safetensors
checkpoint yourself, scripts/ has a converter per arch. Set up its venv:
python3 -m venv .venv-converter
source .venv-converter/bin/activate
pip install -e ".[convert]"Then run any of the per-arch converters (--help for the full flag list):
python scripts/convert_smolvla_to_gguf.py \
--ckpt /path/to/smolvla-libero \
--out /path/to/smolvla-libero-bf16.ggufThe shipped GGUFs are bf16. scripts/quantize_gguf.py repacks the LM-backbone weight
matrices to a smaller type and copies everything else unchanged; the loader keeps the
packed weights and lets ggml_mul_mat dequantize at compute, so the file just loads and
runs like the bf16 one.
python scripts/quantize_gguf.py --in model-bf16.gguf --out model-q8_0.gguf --type Q8_0Q8_0 is near-lossless and roughly halves the LM. Q4_0 is 4-bit for a bigger cut.
Embeddings, the output head, norms and the action expert stay float; pass --vision to
pack the vision tower too (smaller, but more accuracy loss).
vla-bench times predict() in-process on synthetic inputs: engine only, no
transport, no simulator, no claim about task success.
./build/vla-bench -hf vrfai/smolvla-libero-gguf --images 2 --size 512 --markdownRTX 5090, driver 595.84, CUDA 13.2, 24-core host, weights as shipped, 20 reps after 3 warmups, best of three sweeps, each model at its native input size and view count.
| Model | Views | Input | min ms | p50 ms | p90 ms | vision ms |
|---|---|---|---|---|---|---|
| VLA-Adapter | 1 | 224 | 18.2 | 19.8 | 21.1 | 9.4 |
| VLA-JEPA | 1 | 256 | 19.9 | 21.5 | 22.9 | 6.3 |
| BitVLA | 1 | 224 | 23.6 | 25.3 | 26.4 | 5.4 |
| GR00T N1.5 | 1 | 224 | 28.2 | 29.4 | 30.5 | 5.9 |
| GR00T N1.7 | 1 | 256 | 31.0 | 33.4 | 34.6 | 6.2 |
| GR00T N1.6 | 1 | 224 | 33.4 | 35.7 | 37.3 | 6.3 |
| OpenVLA-OFT | 1 | 224 | 47.4 | 49.2 | 50.2 | 10.3 |
| SmolVLA | 2 | 512 | 47.8 | 49.6 | 54.0 | 16.1 |
| pi0 | 2 | 224 | 48.9 | 52.1 | 55.0 | 11.6 |
| Evo-1 | 1 | 448 | 52.2 | 55.2 | 57.3 | 17.8 |
| pi0.5 | 2 | 224 | 53.4 | 56.1 | 59.3 | 11.4 |
Latency says nothing about whether a policy works. LIBERO-Object, 10 tasks and 20 episodes per model, terminated episodes counted as failures:
| Model | Chunk replay | Success rate |
|---|---|---|
| BitVLA | 8 | 100.0% |
| GR00T N1.7 | 16 | 98.0% |
| GR00T N1.5 | 16 | 96.0% |
| Evo-1 | 8 | 94.5% |
| SmolVLA | 4 | 90.5% |
| π0 | 32 | 87.5% |
| GR00T N1.6 | 16 | 86.5% |
Success rate belongs to the checkpoint, not the engine;
vla_predict_check in CONTRIBUTING.md is how a
change is shown to leave it alone.
Experimental results on other platforms can be found in eval/reports or docs/backend.
Support matrix of models (rows) against platforms (columns). Legend: Y =
supported (released and benchmarked), ~ = in progress, - = planned.
| Model | CPU (x86-64 / ARM) | CUDA | SYCL (Intel) | Metal | OpenVINO | Hexagon |
|---|---|---|---|---|---|---|
| SmolVLA | Y | Y | Y | Y | Y | Y |
| π0 | Y | Y | - | Y | Y | ~ |
| π0.5 | Y | Y | - | Y | Y | Y |
| GR00T N1.5 | Y | Y | - | Y | Y | Y |
| GR00T N1.6 | Y | Y | - | Y | Y | Y |
| GR00T N1.7 | Y | Y | - | Y | Y | Y |
| BitVLA | Y | Y | - | ~ | - | - |
| Evo-1 | Y | Y | Y | Y | Y | Y |
| VLA-Adapter | Y | Y | ~ | Y | Y | Y |
| OpenVLA-OFT | Y | Y | - | Y | Y | - |
| VLA-JEPA | Y | Y | - | Y | Y | ~ |
| Octo-Small | Y | Y | Y | Y | - | Y |
| TurboVLA | Y | Y | Y | Y | Y | Y |
See CONTRIBUTING.md for how to prove a change is numerically neutral, and the six sites you touch to add an architecture.
Licensed under the Apache License, Version 2.0.
Supported VLA models:
- SmolVLA - Hugging Face LeRobot team.
- π0,π0.5 - Physical Intelligence.
- BitVLA - Hongyu Wang et al.
- Evo-1 - Tao Lin et al.
- VLA-Adapter - Yihao Wang et al.
- OpenVLA-OFT - Moo Jin Kim et al.
- GR00T N1.x - NVIDIA Isaac.
- VLA-JEPA - Jingwen Sun et al.
- Octo - Octo Model Team, UC Berkeley RAIL.
- TurboVLA - Hengyi Xie et al.
Built on:
llama.cpp- LLM inference engine in C/C++.- LIBERO - benchmark suite for the success-rate sweeps.
- SimplerEnv - the second simulator in the eval scaffold.
