Python bindings for vla.cpp over its
C ABI (include/vla.h). No PyTorch at inference time.
import vla_cpp
model = vla_cpp.load("smolvla-libero.gguf")
actions = model.predict(frame_hwc_uint8, tokens=[1, 100, 200, 2])predict returns [rows, max_action_dim]; only the first
model.config.real_action_dim columns carry values. model.config.denormalized
says whether they are already in world units.
From a checkout of the repo:
pip install ./bindings/pythonThis builds libvla.so with CMake and puts it inside the package, so nothing
else needs to be on the library path. You need a C++17 compiler and network
access (CMake fetches llama.cpp). The build is CPU only (Metal on macOS). It
uses GGML_NATIVE=OFF, so the wheel runs on other machines (on x86 it needs
AVX2).
pip wheel ./bindings/python -w dist gives you the wheel file. With
python -m build, pass --wheel: the sdist does not include the C++ sources.
Point VLA_LIBRARY at a libvla.so from a normal CMake build:
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j"$(nproc)" --target vla
VLA_LIBRARY=build/libvla.so python your_script.pycmake --install build --prefix <dir> gives a relocatable copy that does not
need the build tree. The libraries go to <dir>/lib, or <dir>/lib64 on
Fedora and RHEL.
load(ckpt_path, mmproj_path=None, config_path=None) |
every arch ignores mmproj_path; the runtime block of config_path sets the precision options |
Model.predict(images, tokens, state=None, noise=None, ...) |
images is one HWC array or a sequence |
Model.config |
resolved hyper-parameters |
Model.last_stats() |
per-phase timings, needs timing=TIMING_PHASE |
Model.close() |
or use as a context manager |
Output is bit-identical to vla-cli on the same inputs.