Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
61 commits
Select commit Hold shift + click to select a range
209cf8c
bump llama.cpp to b11223
anindex Sep 28, 2026
3faea0a
bump sentencepiece to v0.2.1
anindex Sep 28, 2026
eff454e
bump OpenVINO, macOS runner and python deps
anindex Sep 28, 2026
79533b2
harden vla-server and vlm-server
anindex Sep 28, 2026
66ce698
fix loader fuse, option precedence and C API races
anindex Sep 28, 2026
3a07dfc
size graphs from node counts, validate pi0/smolvla/oft input
anindex Sep 28, 2026
b7a4e7d
validate GR00T, Evo-1, VLA-JEPA and TurboVLA metadata
anindex Sep 28, 2026
d32b1ac
fix Octo crashes on bad tokens and stats
anindex Sep 28, 2026
9bf951d
fix BitVLA uploads, device choice and kernel races
anindex Sep 28, 2026
aa40537
fix BF16 CUDA hook overflow and races
anindex Sep 28, 2026
06969f6
fix quantizer, converters and eval client
anindex Sep 28, 2026
c1eaec5
harden CI, add CUDA, OpenCL and Metal build checks
anindex Sep 28, 2026
d90aee9
update changelog
anindex Sep 28, 2026
4a7a423
pair LIBERO runs with seeded noise, add octo, turbovla, vla_jepa
anindex Sep 28, 2026
e020e49
move GR00T and VLA-JEPA onto shared modules
anindex Sep 28, 2026
abf3203
share the PaliGemma stack between pi0 and pi05
anindex Sep 28, 2026
5343639
share the dual tower and RoPE helpers between OFT and VLA-Adapter
anindex Sep 28, 2026
82833df
put SmolVLA on gguf_reader and one graph builder
anindex Sep 28, 2026
a85458a
use shared layers in Octo
anindex Sep 28, 2026
aa14ea7
share image preprocessing with Evo-1 and BitVLA
anindex Sep 28, 2026
271ca9c
wire --flash-attn into GR00T N1.5/N1.6 and multi-view encoders
anindex Sep 28, 2026
61115ca
vla-cli takes runtime flags and --config
anindex Sep 28, 2026
f041e07
-hf picks files and tags, --text uses per-arch prompts
anindex Sep 28, 2026
9bb1423
add install rules and a pip wheel for the bindings
anindex Sep 28, 2026
2d289db
fix release packaging, multi-arch multi-stage Docker image
anindex Sep 28, 2026
c749729
bake pi05 adaRMS at load, GEGLU and on-device vision copy for pi0/pi05
anindex Sep 28, 2026
7e6ece2
bake DiT conditioning at load, keep one GR00T embodiment resident
anindex Sep 28, 2026
dcf9641
SmolVLA: widen weights at load, in-graph pixel shuffle, embeddings of…
anindex Sep 28, 2026
6b7039b
cache TurboVLA text encoding and Octo constants
anindex Sep 28, 2026
0e73bbf
upload Evo-1, OFT and VLA-Adapter constants once, keep vision on device
anindex Sep 28, 2026
9a36ab6
drop the BitVLA ViT pad copy
anindex Sep 28, 2026
f2f5c02
pi0: feed raw projector output to the LM
anindex Sep 28, 2026
25523b6
pi05: drop the image scale round-trip
anindex Sep 28, 2026
cf2dc60
apply DINOv3 final norm in TurboVLA
anindex Sep 28, 2026
6bd49f4
octo: match the JAX reference numerics
anindex Sep 28, 2026
db4a3ed
use erf GELU in Qwen3-VL patch mergers
anindex Sep 28, 2026
67693db
smolvla: keep cross-attn k/v projections in F32
anindex Sep 28, 2026
a8ecdf3
evo1: match the reference prompt and constant state dims
anindex Sep 28, 2026
dbf422b
bitvla: round to bf16 to nearest even, recover legacy scale
anindex Sep 28, 2026
4e14852
tap VLA-Adapter head blocks like the reference
anindex Sep 28, 2026
4f85ac4
eval client: normalize VLA-Adapter proprio, zero constant GR00T dims
anindex Sep 28, 2026
2a69d3b
changelog for the math fixes and speedups
anindex Sep 28, 2026
c753c25
tokenizer in the GGUF, Octo always built, VLA_SPM option
anindex Sep 28, 2026
55b9f01
add --num-steps as a load-time option
anindex Sep 28, 2026
e2af0b2
drop dead preprocessing and norm helpers
anindex Sep 28, 2026
335a87f
update docs for this release
anindex Sep 28, 2026
585b465
changelog for tokenizer and num-steps
anindex Sep 28, 2026
0548f1f
bitvla: convert bf16 uploads with ggml's row helper
anindex Sep 28, 2026
ad56e9b
bound Octo token count, quiet safetensors probe, refuse unused act-dtype
anindex Sep 28, 2026
33e90ba
pi05 converter QUANTILES only, vla_jepa venv override, learned noise …
anindex Sep 28, 2026
a8698f7
self-contained wheel, macOS tarball without servers, portable PTX, li…
anindex Sep 28, 2026
a427801
fix stale docs and changelog claims
anindex Sep 28, 2026
e4fed92
build sentencepiece against package abseil when protobuf needs it
anindex Sep 29, 2026
c93ca0a
drop comment in sentencepiece abseil block
anindex Sep 29, 2026
9960159
docs: drop outdated eval/reports, point to docs/benchmark
khanhnd61-vr Sep 29, 2026
1348e9e
docs(benchmark): add per-device latency and memory reports at c93ca0a
khanhnd61-vr Sep 29, 2026
3ff447c
fix(openvino): key the RoPE sin/cos cache on the position input
khanhnd61-vr Sep 30, 2026
450992c
docs(turbovla): drop the re-convert notes
khanhnd61-vr Sep 30, 2026
07196d0
release: v0.4.0 changelog, version bump, split README into docs
khanhnd61-vr Sep 30, 2026
c7e1eb2
add report libero and Intel UPX board
khanhnd61-vr Sep 30, 2026
096dc61
update CHANGELOG for v0.4.0
khanhnd61-vr Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .dockerignore
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# Keep the build context small: llama.cpp is fetched inside the image, models
# are mounted at runtime, and build/output dirs are host artifacts.
.git/
.claude/
models/
build/
build-*/
Expand Down
87 changes: 68 additions & 19 deletions .github/workflows/build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,19 +13,6 @@ concurrency:
cancel-in-progress: true

jobs:
cpp-unit:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7
# Compiled directly: pure, no llama.cpp or protobuf/zmq needed.
- name: pure unit tests
run: |
for t in test_vision_common test_rope_conventions; do
g++ -std=c++17 -Isrc -Wall -Wextra -fsanitize=address,undefined \
-fno-omit-frame-pointer "tests/$t.cpp" -o "/tmp/$t"
"/tmp/$t"
done

py-tooling:
runs-on: ubuntu-24.04
steps:
Expand Down Expand Up @@ -56,21 +43,83 @@ jobs:
- uses: actions/cache@v6
with:
path: build/_deps
key: llama-${{ steps.pin.outputs.tag }}-${{ runner.os }}
key: llama-${{ steps.pin.outputs.tag }}-${{ runner.os }}-asan
# Everything, not a target list: ctest registers tests this job must build,
# and a named list goes stale the next time one is added.
- name: build + ctest (CPU, -Wall -Wextra)
- name: build + ctest (CPU, ASan+UBSan, -Werror)
run: |
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=OFF -DVLA_BUILD_TESTS=ON
san="-fsanitize=address,undefined -fno-sanitize-recover=undefined -fno-omit-frame-pointer"
cmake -B build -DCMAKE_BUILD_TYPE=RelWithDebInfo -DGGML_CUDA=OFF -DVLA_BUILD_TESTS=ON \
-DVLA_WERROR=ON -DCMAKE_C_FLAGS="$san" -DCMAKE_CXX_FLAGS="$san"
cmake --build build -j"$(nproc)"
ctest --test-dir build --output-on-failure
# This job is the only one that fetches llama.cpp, and neither patch script
# runs on a CPU build, so their anchors would otherwise rot unnoticed until
# someone configures a CUDA or OpenVINO tree. Patch a copy: the real one is
# Neither patch script runs on a CPU build, and build-cuda patches only on a
# cache miss, so their anchors would otherwise rot unnoticed until someone
# configures a fresh CUDA or OpenVINO tree. Patch a copy: the real one is
# cached.
- name: patch anchors still apply
run: |
cp -r build/_deps/llama-src /tmp/llama-patchtest
python3 scripts/patch_ggml_cuda_ext_hook.py /tmp/llama-patchtest
python3 scripts/patch_ggml_openvino.py /tmp/llama-patchtest
python3 scripts/patch_ggml_openvino.py /tmp/llama-patchtest # idempotent

build-cuda:
runs-on: ubuntu-24.04
container: nvidia/cuda:12.9.1-devel-ubuntu24.04
steps:
- name: deps
run: |
apt-get update -qq
DEBIAN_FRONTEND=noninteractive apt-get install -y -qq --no-install-recommends \
build-essential cmake git ca-certificates pkg-config python3 \
libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compiler
- uses: actions/checkout@v7
- name: read llama.cpp pin
id: pin
run: echo "tag=$(bash scripts/llama_tag.sh)" >> "$GITHUB_OUTPUT"
- uses: actions/cache@v6
with:
path: build/_deps
key: llama-${{ steps.pin.outputs.tag }}-${{ runner.os }}-cuda-${{ hashFiles('scripts/patch_ggml_cuda_ext_hook.py') }}
- name: build (compile only, the runner has no GPU)
run: |
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89-real \
-DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined -DVLA_BUILD_TESTS=ON -DVLA_WERROR=ON
cmake --build build -j"$(nproc)"

build-backends:
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
include:
- { name: opencl, os: ubuntu-24.04, cmake: -DGGML_OPENCL=ON -DVLA_WERROR=ON }
- { name: metal, os: macos-15, cmake: -DGGML_METAL=ON, test: true }
steps:
- uses: actions/checkout@v7
- name: deps
if: runner.os == 'Linux'
run: |
sudo apt-get update -qq
sudo apt-get install -y -qq --no-install-recommends \
build-essential cmake git ca-certificates pkg-config \
libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compiler \
ocl-icd-opencl-dev opencl-headers
- name: deps
if: runner.os == 'macOS'
run: brew install cmake zeromq cppzmq protobuf
- name: read llama.cpp pin
id: pin
run: echo "tag=$(bash scripts/llama_tag.sh)" >> "$GITHUB_OUTPUT"
- uses: actions/cache@v6
with:
path: build/_deps
key: llama-${{ steps.pin.outputs.tag }}-${{ runner.os }}-${{ matrix.name }}
- name: build
run: |
cmake -B build -DCMAKE_BUILD_TYPE=Release ${{ matrix.cmake }} -DVLA_BUILD_TESTS=ON
cmake --build build -j"$(getconf _NPROCESSORS_ONLN)"
- name: ctest
if: matrix.test
run: ctest --test-dir build --output-on-failure
158 changes: 111 additions & 47 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
@@ -1,143 +1,206 @@
# Tagged binaries and a container image. build.yml already compiles all of this
# on every push; this is the same work with the artifacts kept.
# Tagged binaries and a container image. A tag push publishes them; a manual
# run builds the same artifacts on the chosen ref and publishes nothing.
name: release

on:
push:
tags: ['v*']
workflow_dispatch:
inputs:
tag:
description: Tag to build (dry run, nothing is published)
required: true

permissions:
contents: write
packages: write
contents: read

env:
BINARIES: vla-server vlm-server vla-cli vla-bench

jobs:
linux:
runs-on: ${{ matrix.runner }}
container: ${{ matrix.image }}
defaults:
run:
shell: bash
strategy:
fail-fast: false
matrix:
include:
- name: linux-x86_64-cpu
runner: ubuntu-24.04
image: ubuntu:24.04
cmake: -DGGML_CUDA=OFF
cuda: false
- name: linux-x86_64-cuda-12.8
runner: ubuntu-24.04
- name: linux-x86_64-cuda
cmake: -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES="75;86;89;120"
image: nvidia/cuda:12.8.1-devel-ubuntu24.04
cmake: -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES="75-real;80-real;86-real;89-real;90;120-real"
cuda: true
- name: linux-x86_64-cuda-13.4
runner: ubuntu-24.04
# Jetson and other aarch64 boards. Native arm64 runner, CPU only: the
# hosted images carry no CUDA for arm64, so a Jetson GPU build still
# has to happen on the device.
image: nvidia/cuda:13.4.1-devel-ubuntu24.04
cmake: -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES="75-real;80-real;86-real;89-real;90;120-real"
cuda: true
- name: linux-aarch64-cpu
cmake: -DGGML_CUDA=OFF
cuda: false
runner: ubuntu-24.04-arm
image: ubuntu:24.04
cmake: -DGGML_CUDA=OFF -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16
- name: linux-aarch64-cuda-13.4
runner: ubuntu-24.04-arm
image: nvidia/cuda:13.4.1-devel-ubuntu24.04
cmake: -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES="87-real;110-real;121-real" -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16
cuda: true
steps:
- uses: actions/checkout@v7

- name: deps
run: |
sudo apt-get update -qq
sudo apt-get install -y -qq --no-install-recommends \
build-essential cmake git ca-certificates pkg-config \
apt-get update -qq
DEBIAN_FRONTEND=noninteractive apt-get install -y -qq --no-install-recommends \
build-essential cmake git ca-certificates pkg-config python3 \
libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compiler

- name: cuda toolkit
if: matrix.cuda
run: |
wget -q https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update -qq
sudo apt-get install -y -qq --no-install-recommends cuda-toolkit-12-8
echo "/usr/local/cuda/bin" >> "$GITHUB_PATH"
- uses: actions/checkout@v7

- name: build
run: |
cmake -B build -DCMAKE_BUILD_TYPE=Release ${{ matrix.cmake }}
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=OFF \
-DCMAKE_INSTALL_RPATH='$ORIGIN' -DCMAKE_BUILD_WITH_INSTALL_RPATH=ON \
${{ matrix.cuda && '-DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined -DGGML_CUDA_NCCL=OFF' || '' }} ${{ matrix.cmake }}
cmake --build build -j"$(nproc)" --target $BINARIES vla

- name: package
run: |
out="vla.cpp-${{ github.ref_name }}-${{ matrix.name }}"
mkdir -p "$out"
out="vla.cpp-${GITHUB_REF_NAME//\//-}-${{ matrix.name }}"
mkdir -p "$out/scripts"
for b in $BINARIES; do cp "build/$b" "$out/"; done
cp build/libvla.so "$out/"
cp -P build/*.so* build/bin/*.so* "$out/"
ldd "$out/vla-server" | awk '/libgomp|libprotobuf/ {print $3}' | xargs cp -L -t "$out/"
cp /usr/share/doc/libprotobuf32t64/copyright "$out/LICENSE.protobuf"
cp /usr/share/doc/libgomp1/copyright "$out/LICENSE.libgomp"
cp build/_deps/llama-src/LICENSE "$out/LICENSE.llama.cpp"
cp build/_deps/llama-src/licenses/LICENSE-jsonhpp "$out/LICENSE.jsonhpp"
cp build/_deps/sentencepiece-src/LICENSE "$out/LICENSE.sentencepiece"
cp build/_deps/sentencepiece-src/third_party/darts_clone/LICENSE "$out/LICENSE.darts_clone"
cp include/vla.h LICENSE.md README.md "$out/"
# vla-cli --text runs this; VLA_TOKENIZE_SCRIPT points at it.
mkdir -p "$out/scripts" && cp scripts/tokenize_prompt.py "$out/scripts/"
# vla-cli --text finds this in scripts/ next to itself.
cp scripts/tokenize_prompt.py "$out/scripts/"
tar -czf "$out.tar.gz" "$out"

- name: smoke
run: |
out="vla.cpp-${GITHUB_REF_NAME//\//-}-${{ matrix.name }}"
mv build build.moved
ldd "$out"/vl*-* "$out"/*.so* > ldd.txt
if grep -v libcuda.so.1 ldd.txt | grep 'not found'; then exit 1; fi
LD_LIBRARY_PATH=/usr/local/cuda/compat "$out/vla-cli" --help

- name: cuda runtime
if: matrix.cuda
run: |
out="vla.cpp-${GITHUB_REF_NAME//\//-}-${{ matrix.name }}"
mkdir cudart
cp -L /usr/local/cuda/lib64/lib{cudart,cublas,cublasLt}.so."${CUDA_VERSION%%.*}" cudart/
tar -czf "cudart-$out.tar.gz" --transform "s,^cudart,$out," cudart

- uses: actions/upload-artifact@v7
with:
name: ${{ matrix.name }}
path: '*.tar.gz'
if-no-files-found: error

macos:
runs-on: macos-14
runs-on: macos-15
env:
BINARIES: vla-cli vla-bench
steps:
- uses: actions/checkout@v7

- name: deps
run: brew install cmake zeromq cppzmq protobuf
run: brew install cmake

- name: build
run: |
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON -DGGML_NATIVE=OFF \
-DVLA_BUILD_SERVER=OFF -DVLA_SPM=OFF \
-DCMAKE_INSTALL_RPATH='@loader_path' -DCMAKE_BUILD_WITH_INSTALL_RPATH=ON
cmake --build build -j"$(sysctl -n hw.ncpu)" --target $BINARIES vla

- name: package
run: |
out="vla.cpp-${{ github.ref_name }}-macos-arm64-metal"
mkdir -p "$out"
out="vla.cpp-${GITHUB_REF_NAME//\//-}-macos-arm64-metal"
mkdir -p "$out/scripts"
for b in $BINARIES; do cp "build/$b" "$out/"; done
cp build/libvla.dylib "$out/"
cp -a build/*.dylib build/bin/*.dylib "$out/"
cp build/_deps/llama-src/LICENSE "$out/LICENSE.llama.cpp"
cp build/_deps/llama-src/licenses/LICENSE-jsonhpp "$out/LICENSE.jsonhpp"
cp include/vla.h LICENSE.md README.md "$out/"
mkdir -p "$out/scripts" && cp scripts/tokenize_prompt.py "$out/scripts/"
cp scripts/tokenize_prompt.py "$out/scripts/"
# Metal needs the shader library next to the binary.
find build -name 'default.metallib' -exec cp {} "$out/" \;
tar -czf "$out.tar.gz" "$out"

- name: smoke
run: |
out="vla.cpp-${GITHUB_REF_NAME//\//-}-macos-arm64-metal"
mv build build.moved
otool -L "$out"/vl*-* "$out"/*.dylib
if otool -L "$out"/vl*-* "$out"/*.dylib | grep -E '/opt/homebrew|/usr/local/'; then exit 1; fi
for lib in $(otool -L "$out"/vl*-* "$out"/*.dylib | awk '/@rpath\//{print $1}' | sort -u); do
test -e "$out/${lib#@rpath/}" || { echo "missing $lib"; exit 1; }
done
"$out/vla-cli" --help

- uses: actions/upload-artifact@v7
with:
name: macos-arm64-metal
path: '*.tar.gz'
if-no-files-found: error

docker:
runs-on: ubuntu-24.04
permissions:
contents: read
packages: write
steps:
- name: free disk
run: sudo rm -rf /usr/share/dotnet /usr/local/lib/android /opt/ghc /opt/hostedtoolcache
- uses: actions/checkout@v7
- uses: docker/setup-buildx-action@v4
- uses: docker/login-action@v4
if: startsWith(github.ref, 'refs/tags/')
if: github.event_name == 'push'
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: image name
run: echo "IMAGE=ghcr.io/${GITHUB_REPOSITORY,,}" >> "$GITHUB_ENV"
run: |
echo "IMAGE=ghcr.io/${GITHUB_REPOSITORY,,}" >> "$GITHUB_ENV"
echo "REF=${GITHUB_REF_NAME//\//-}" >> "$GITHUB_ENV"
- uses: docker/build-push-action@v7
with:
context: .
push: ${{ startsWith(github.ref, 'refs/tags/') }}
push: ${{ github.event_name == 'push' }}
build-args: |
CUDA_ARCH=75-real;80-real;86-real;89-real;90;120-real
GGML_NATIVE=OFF
tags: |
${{ env.IMAGE }}:${{ github.ref_name }}
${{ env.IMAGE }}:${{ env.REF }}
${{ env.IMAGE }}:latest
cache-from: type=gha
cache-to: type=gha,mode=max

publish:
needs: [linux, macos, docker]
if: startsWith(github.ref, 'refs/tags/')
if: github.event_name == 'push'
runs-on: ubuntu-24.04
permissions:
contents: write
steps:
- uses: actions/checkout@v7
# The release body is this tag's CHANGELOG.md section; a tag with none
# fails here rather than publishing an empty release.
- name: release notes
run: |
ver="${GITHUB_REF_NAME#v}"
awk -v h="## [$ver]" 'index($0, h) == 1 {on=1; next} on && /^## \[/ {exit} on' \
CHANGELOG.md > notes.md
test -s notes.md || { echo "no CHANGELOG.md section for $ver"; exit 1; }
# Tarballs only. The docker job also leaves a .dockerbuild build record
# artifact behind, and pulling that one fails the whole download.
- uses: actions/download-artifact@v8
Expand All @@ -146,4 +209,5 @@ jobs:
with:
files: dist/*.tar.gz
fail_on_unmatched_files: true
body_path: notes.md
generate_release_notes: true
Loading
Loading