Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,19 @@ Notable changes to vla.cpp. Format loosely follows [Keep a Changelog](https://ke
- `scripts/foldquant_fake_export.py` (uncalibrated file for bring-up),
`scripts/inspect_gguf_quant.py` (contract check) and `scripts/foldquant_ref.py`
(numpy reference).
- **PicoVLA** (`fast_smolvla`): DINOv3 ConvNeXt-T with a language-gated 2x2 token
merge, a five-layer Llama backbone with 16 record tokens, and a four-layer
action expert that reads the backbone's per-layer prefix K/V; three MeanFlow
steps. `scripts/convert_picovla_to_gguf.py` converts the LeRobot checkpoint
(the latent head is dropped). The record tokens are a one-frame memory that
the model keeps between calls; the LIBERO client sets `reset_memory` at each
episode start and applies the reference server's cross-chunk blending.

### Changed

- `vla_inputs` gains `reset_memory` (C ABI version 2), as do `vla::Inputs`, the
`PredictRequest` proto and the Python bindings' `predict`. It clears the
cross-call memory of the archs that keep one (PicoVLA); the others ignore it.

## [0.4.0] - 2026-09-30

Expand Down
1 change: 1 addition & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -249,6 +249,7 @@ add_library(vla_core ${VLA_CORE_LIB_TYPE}
src/models/openvla_oft.cpp
src/models/vla_jepa.cpp
src/models/turbovla.cpp
src/models/picovla.cpp
src/models/octo.cpp
src/tokenizer.cpp
)
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,7 @@ supported (released and benchmarked), `~` = in progress, `-` = planned.
| [VLA-JEPA](https://hf.co/vrfai/vla-jepa-libero) | Y | Y | - | Y | Y | ~ |
| [Octo-Small](https://hf.co/vrfai/octo-small-libero-gguf) | Y | Y | Y | Y | - | Y |
| [TurboVLA](https://hf.co/vrfai/turbovla-libero-gguf) | Y | Y | Y | Y | Y | Y |
| [PicoVLA](https://hf.co/khanhnd61/picovla-libero-pretrained-gguf) | Y | Y | - | - | - | - |

---

Expand Down
5 changes: 4 additions & 1 deletion bindings/python/vla_cpp/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -80,11 +80,13 @@ def __del__(self):
self.close()

def predict(self, images, tokens: Sequence[int], state=None, noise=None,
pixel_format: int = PIXEL_U8, timing: int = TIMING_NONE):
pixel_format: int = PIXEL_U8, timing: int = TIMING_NONE, reset_memory: bool = False):
"""Run one forward pass.

images: one HWC array, or a sequence of them for multi-view. uint8 RGB by
default; pass pixel_format=PIXEL_F32_RGB_01 for float RGB in [0, 1].
reset_memory: set on an episode's first call; clears the cross-call
memory of the archs that keep one (PicoVLA).
"""
views = images if isinstance(images, (list, tuple)) else [images]
if not views:
Expand Down Expand Up @@ -133,6 +135,7 @@ def predict(self, images, tokens: Sequence[int], state=None, noise=None,
if noise_ptr is not None:
cin.noise = ctypes.cast(noise_ptr, POINTER(c_float))
cin.timing_detail = int(timing)
cin.reset_memory = int(bool(reset_memory))

out = POINTER(c_float)()
n = c_int64()
Expand Down
3 changes: 2 additions & 1 deletion bindings/python/vla_cpp/_ffi.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@
c_void_p,
)

ABI_VERSION = 1
ABI_VERSION = 2

OK = 0
ERR_ARG = -1
Expand Down Expand Up @@ -86,6 +86,7 @@ class Inputs(ctypes.Structure):
("attention_mask", POINTER(c_int32)),
("attention_mask_n", c_int32),
("timing_detail", c_int32),
("reset_memory", c_int32),
]


Expand Down
5 changes: 5 additions & 0 deletions docs/EVAL.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,11 @@ The GR00T models need two extras:
- server side: `VLA_GR00T_EMBODIMENT` (`new_embodiment` for N1.5, `libero_panda`
for N1.6, `libero_sim` for N1.7).

PicoVLA (`--arch picovla`) matches its reference client with `--n-action-steps 1`: it re-plans every
step, the client blends each chunk with the previous one, and `reset()` clears
the server's memory at the next request. With more steps per chunk the blending
is off.

### SimplerEnv

So far only **GR00T-N1.6** is wired (the `gr00t-n1d6-bridge` checkpoint with the
Expand Down
2 changes: 1 addition & 1 deletion docs/MODELS.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ python scripts/convert_smolvla_to_gguf.py \

## Quantization

Most shipped GGUFs are BF16. π0.5, Octo and TurboVLA ship F32, and GR00T N1.5
Most shipped GGUFs are BF16. π0.5, Octo, TurboVLA and PicoVLA ship F32, and GR00T N1.5
and N1.6 are mostly F32. `scripts/quantize_gguf.py` repacks the LM-backbone weight
matrices to a smaller type and copies everything else unchanged. The loader keeps
the packed weights, so the file loads and runs like the original.
Expand Down
5 changes: 5 additions & 0 deletions docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,11 @@ vla-server: bound to tcp://*:5555. ready.
Use `--bind` to change the address and port. Stop the server with `Ctrl-C`.
`vla-server` also takes `-hf user/repo[:file.gguf|:tag]` in place of a checkpoint path.

PicoVLA keeps a one-frame memory between requests (its record tokens), so a
server holds one episode's state: serve one client per server, and set
`reset_memory` on the first `PredictRequest` of every episode. The other
architectures are stateless and ignore the field.

Clients: the LIBERO and SimplerEnv runners in [EVAL.md](EVAL.md), and the
real-robot client in the README's
[Rollout on a real robot](../README.md#rollout-on-a-real-robot).
Expand Down
36 changes: 35 additions & 1 deletion eval/client/vla_cpp_client.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,8 @@
"max_length": 21,
},
"gr00t_n1_7": {"image_size": 256, "tokenizer": "nvidia/Cosmos-Reason2-2B", "max_state_dim": 132},
# 256x256 LIBERO frames upsampled to the ConvNeXt's 448, as its DINOv3 processor does.
"picovla": {"image_size": 448, "tokenizer": "HuggingFaceTB/SmolVLM2-500M-Instruct", "max_state_dim": 32},

"gr00t_n1_5": {"image_size": 224, "tokenizer": "lerobot/eagle2hg-processor-groot-n1p5",
"max_state_dim": 64, "trust_remote_code": True},
Expand Down Expand Up @@ -210,6 +212,15 @@ def __init__(
self.n_action_steps = n_action_steps
self._action_queue: deque = deque(maxlen=n_action_steps)

# PicoVLA keeps a one-frame memory on the server, cleared by the first
# request of each episode, and its reference server blends every chunk
# with the previous one (_picovla_blend).
self._picovla_prev = None
self._picovla_calls = 0
if arch == "picovla" and n_action_steps != 1:
print(f"vla-cpp-direct[arch=picovla]: n_action_steps={n_action_steps}, cross-chunk "
"blending off (the reference re-plans every step)", flush=True)

self._bitvla_proprio_norm = None
self._bitvla_unnorm_key = None
if arch in ("bitvla", "vla_adapter"):
Expand Down Expand Up @@ -648,6 +659,8 @@ def reset(self) -> None:
self._action_queue.clear()
self._episode += 1
self._step = 0
self._picovla_prev = None
self._picovla_calls = 0

def get_action(self, observations: dict[str, Any]) -> np.ndarray:

Expand All @@ -662,6 +675,8 @@ def get_action(self, observations: dict[str, Any]) -> np.ndarray:
chunk = self._predict_chunk_octo(observations)
else:
chunk = self._predict_chunk(observations)
if self.arch == "picovla":
chunk = self._picovla_blend(chunk)
for row in chunk[: self.n_action_steps, : self.real_action_dim]:
self._action_queue.append(np.ascontiguousarray(row, dtype=np.float32))
return self._action_queue.popleft()
Expand Down Expand Up @@ -726,6 +741,8 @@ def _predict_chunk(self, observations: dict[str, Any]) -> np.ndarray:
ip.data = img.tobytes()
req.lang_tokens.extend(lang.tolist())
req.state.extend(state_padded.tolist())
if self.arch == "picovla":
req.reset_memory = self._picovla_calls == 0

self._maybe_add_fixed_noise(req)
self.sock.send(req.SerializeToString())
Expand All @@ -740,6 +757,23 @@ def _predict_chunk(self, observations: dict[str, Any]) -> np.ndarray:
return (np.array(resp.action_chunk, dtype=np.float32)
.reshape(resp.chunk_size, resp.action_dim))

def _picovla_blend(self, chunk: np.ndarray) -> np.ndarray:
"""PicoVLA's cross-chunk blending (service/server.py, utils/policy_utils.py
make_prev_action_chunk): from an episode's third call on, the new chunk
is mixed with the previous one shifted by the one step executed since,
w_h = arccos(1 - 2(H-h-1)/H) / pi, 0.84 on the first step down to 0 on
the last. The reference blends normalized actions; the un-normalization
is affine per dimension, so world units blend the same."""
if self._picovla_calls > 1 and self.n_action_steps == 1:
horizon = chunk.shape[0]
prev = np.zeros_like(chunk)
prev[:-1] = self._picovla_prev[1:]
w = (np.arccos(1.0 - 2.0 * (horizon - np.arange(horizon) - 1) / horizon) / np.pi)[:, None]
chunk = ((1.0 - w) * chunk + w * prev).astype(np.float32)
self._picovla_prev = chunk.copy()
self._picovla_calls += 1
return chunk

def _predict_chunk_turbovla(self, observations: dict[str, Any]) -> np.ndarray:
if self._turbovla_state_norm is None or self._turbovla_action_unnorm is None:
raise RuntimeError("TurboVLA statistics were not initialized")
Expand Down Expand Up @@ -957,7 +991,7 @@ def _predict_chunk_pi05(self, observations: dict[str, Any]) -> np.ndarray:
_EVO1_MAX_TEXT_LENGTH = 1024
_FIXED_NOISE_ARCHS = {
"smolvla", "pi0", "pi05", "evo1", "gr00t_n1_5", "gr00t_n1_6", "gr00t_n1_7",
"vla_jepa", "octo",
"vla_jepa", "octo", "picovla",
}

def _maybe_add_fixed_noise(self, req) -> None:
Expand Down
15 changes: 12 additions & 3 deletions eval/run_libero.sh
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ Usage: $(basename "$0") -i <MODELS_ROOT> [-o <OUTPUT_ROOT>] [-n <N_EPISODES>] [-
-m MODEL which model to run: smol | pi0 | pi05 | bit | evo1 |
vla_adapter | openvla_oft |
gr00t_n1_5 | gr00t_n1_6 | gr00t_n1_7 |
octo | turbovla | vla_jepa | all
octo | turbovla | vla_jepa | picovla | all
(default: all)
-h show this help

Expand Down Expand Up @@ -74,9 +74,9 @@ done
shift $((OPTIND - 1))

case "${MODEL}" in
smol|pi0|pi05|bit|evo1|vla_adapter|openvla_oft|gr00t_n1_5|gr00t_n1_6|gr00t_n1_7|octo|turbovla|vla_jepa|all) ;;
smol|pi0|pi05|bit|evo1|vla_adapter|openvla_oft|gr00t_n1_5|gr00t_n1_6|gr00t_n1_7|octo|turbovla|vla_jepa|picovla|all) ;;
*)
echo "ERROR: -m must be one of: smol | pi0 | pi05 | bit | evo1 | vla_adapter | openvla_oft | gr00t_n1_5 | gr00t_n1_6 | gr00t_n1_7 | octo | turbovla | vla_jepa | all (got '${MODEL}')" >&2
echo "ERROR: -m must be one of: smol | pi0 | pi05 | bit | evo1 | vla_adapter | openvla_oft | gr00t_n1_5 | gr00t_n1_6 | gr00t_n1_7 | octo | turbovla | vla_jepa | picovla | all (got '${MODEL}')" >&2
exit 1
;;
esac
Expand Down Expand Up @@ -147,6 +147,7 @@ N_ACTION_STEPS_GR00T_N1_7="${N_ACTION_STEPS_GR00T_N1_7:-16}" # N1.7 H4 closeout
N_ACTION_STEPS_OCTO="${N_ACTION_STEPS_OCTO:-4}"
N_ACTION_STEPS_TURBOVLA="${N_ACTION_STEPS_TURBOVLA:-12}"
N_ACTION_STEPS_VLA_JEPA="${N_ACTION_STEPS_VLA_JEPA:-7}"
N_ACTION_STEPS_PICOVLA="${N_ACTION_STEPS_PICOVLA:-1}" # reference sync client re-plans every step

mkdir -p "${OUTPUT_ROOT}"
OUTPUT_ROOT="$(cd "${OUTPUT_ROOT}" && pwd)"
Expand Down Expand Up @@ -552,5 +553,13 @@ if should_run vla_jepa; then
fi
fi

if should_run picovla; then
run_model picovla \
"${MODELS_ROOT}/picovla-libero-pretrained-gguf" \
"${N_ACTION_STEPS_PICOVLA}" \
"" \
"${MODELS_ROOT}/picovla-libero-pretrained-gguf/picovla-libero-f32.gguf"
fi

echo "===================="
echo "Done. Results under ${OUTPUT_ROOT}"
6 changes: 5 additions & 1 deletion include/vla.h
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ extern "C" {
#endif

// Bumped on any incompatible change to the structs or functions below.
#define VLA_ABI_VERSION 1
#define VLA_ABI_VERSION 2

typedef struct vla_model vla_model;

Expand Down Expand Up @@ -124,6 +124,10 @@ typedef struct {
int32_t attention_mask_n;

int32_t timing_detail; ///< A vla_timing_detail value.

/// Non-zero at an episode start: clear the model's cross-call memory first.
/// Only PicoVLA keeps any; other archs ignore it.
int32_t reset_memory;
} vla_inputs;

/// Milliseconds. Phase fields are zero unless timing_detail was VLA_TIMING_PHASE.
Expand Down
Loading
Loading