The engine works; the gap is distribution.
Done:
- C ABI.
include/vla.handlibvla.src/model.his C++ only, so without it nothing outside C++ can link the engine. - Python bindings.
bindings/python, ctypes over the ABI.pip install ./bindings/pythonbuilds a self-contained wheel with scikit-build-core. - Prebuilt binaries.
.github/workflows/release.ymlpublishes on tag: linux-x86_64 (CPU, CUDA 12.8, CUDA 13.4), linux-aarch64 (CPU, and CUDA 13.4 for Orin, Thor and DGX Spark), macos-arm64-metal, and a Docker image. The tarballs now run off the build machine: shared libraries ship next to the binaries with an$ORIGINrpath, builds useGGML_NATIVE=OFF, and CI runsvla-cli --helpwith the build tree moved away.vla-serverstill needslibzmq5from the system. - Install rules.
cmake --installgives a relocatable tree with a$ORIGIN/../librpath. - One-command model fetch.
-hf user/repo[:path/file.gguf|:tag]onvla-cli,vla-serverandvla-bench, cached under$VLA_CACHE. A repo with several GGUFs lists them instead of guessing. - Reproducible benchmarks.
vla-benchemits the README table rows. - Contributor path.
CONTRIBUTING.mdhas the six-site walkthrough for adding an architecture, plus issue and PR templates. - Instruction in, action out.
vla-cli --textbuilds each arch's real prompt. Octo carries its tokenizer in the GGUF, andscripts/add_tokenizer_to_gguf.pyadds one to pi0, pi0.5 and OpenVLA-OFT, so those need no Python at run time. The other archs callscripts/tokenize_prompt.py.
Left:
- Jetson on JetPack 6. The tarballs are built on Ubuntu 24.04 and its glibc, and the aarch64 CUDA one needs a CUDA 13 driver, so JetPack 6 Orins still build on the device.
- PyPI. The wheel builds from
bindings/python, but nothing publishes it. - Hugging Face library registration. The
vrfaimodel cards setlibrary_name: vla.cpp, but vla.cpp is not registered in huggingface.js, so the Hub shows no "Use this model" snippet. - Windows zips and a Homebrew formula. Windows builds from source only, and there is no brew tap, though the install rules now make one possible.
ci/baselines/rtx3090.jsonstill disagrees with the README table, which is now RTX 5090 numbers fromvla-bench. Re-record the baselines on one machine.- Success rates. The README table comes from a May 2026 RTX 3060 sweep and covers seven of the thirteen archs, and π0, SmolVLA and GR00T N1.7 have moved since. A fresh sweep would cover the rest.
None of these change inference behaviour.