The eval scaffold under eval/ runs vla-server against two
simulators end-to-end: LIBERO and SimplerEnv. To run it in containers instead,
see DOCKER.md.
Each setup script bootstraps an isolated Python 3.10 uv venv next to itself and
clones the upstream sim repo. Both require uv
on PATH.
bash eval/sim/libero/setup_libero.shClones LIBERO into eval/sim/libero/LIBERO/,
creates eval/sim/libero/libero_uv/.venv/, and pins compatible versions of
torch, lerobot, transformers, and gymnasium.
bash eval/sim/simpler/setup_SimplerEnv.shClones SimplerEnv (and its nested ManiSkill2_real2sim) into
eval/sim/simpler/SimplerEnv/, creates
eval/sim/simpler/simpler_uv/.venv/.
eval/client/ drives vla-server directly over the protobuf
protocol. Start the server first, see USAGE.md.
export VLA_GGUF=models/smolvla/smolvla-libero.gguf # the checkpoint to serve
export VLA_ARCH=smolvla # client-side arch preset, see --helpWith vla-server already running:
source eval/sim/libero/libero_uv/.venv/bin/activate
python eval/client/run_sim_client_direct.py \
--task libero_object --task-id 0 --n-episodes 1 \
--output-dir /tmp/libero_outputs \
--arch "$VLA_ARCH"The GR00T models need two extras:
- client side:
--stats-json /path/to/dataset_statistics.json - server side:
VLA_GR00T_EMBODIMENT(new_embodimentfor N1.5,libero_pandafor N1.6,libero_simfor N1.7).
So far only GR00T-N1.6 is wired (the gr00t-n1d6-bridge checkpoint with the
oxe_widowx embodiment). Start vla-server with the oxe_widowx embodiment:
VLA_GR00T_EMBODIMENT=oxe_widowx \
./build/vla-server "$GR00T_N1D6_GGUF"Then drive it from the SimplerEnv venv:
source eval/sim/simpler/simpler_uv/.venv/bin/activate
python eval/client/run_simpler_client_direct.py \
--arch gr00t_n1_6 \
--task-id oxe_widowx/widowx_spoon_on_towel --n-episodes 1 \
--embodiment oxe_widowx --image-size 252 \
--stats-json "$VLA_STATS_JSON"Success rate belongs to the checkpoint, not the engine; vla_predict_check in
CONTRIBUTING.md is how a change is shown to leave it
alone. Measured LIBERO success rates are in CHANGELOG.md.