Skip to content

Deployment pipeline for SmolVLA on a real SO-101 arm (Qualcomm Dragonwing IQ-9075, ARM CPU) #31

Description

@gigwegbe

I've fine-tuned a SmolVLA policy in LeRobot on my own SO-101 dataset and want to deploy it on a Qualcomm Dragonwing IQ-9075 board (aarch64, no CUDA) driving a physical SO-101 arm. I've built vla.cpp on the board and can convert/quantize my checkpoint to GGUF.

The eval clients I see target LIBERO/SimplerEnv simulators. Is there a recommended pipeline (or existing bridge) for driving a real robot from vla-server, connecting LeRobot's robot interface (get_observation → action chunk → send_action) to the vla-server protocol? If not, is there a preferred pattern for writing one, using the sim clients as a template?

Happy to share benchmark numbers back as part of an edge-deployment writeup.

Activity

  1. khanhnd61-vr commented on Sep 30, 2026

    @khanhnd61-vr
    Collaborator

    Thanks for the detailed write-up - you've identified the gap accurately.
    No real-robot bridge in this repo yet - eval/client/ only targets LIBERO and SimplerEnv.

    We're building one in a LeRobot fork

    • A CLI wiring LeRobot's robot interface to the vla-server protocol.
    • Covers both vla.cpp and our latest project vla.simd engine (engine to inference VLA on CPU).
    • SmolVLA on an SO-101 is the target case.

    Two caveats:

    • Experimental and not mature yet.
    • Async inference is not supported yet. vla-server is synchronous REQ/REP, so the control loop is synchronous too: a chunk is executed before the next request goes out, with no overlap of inference and execution.

    We'll comment here once it's ready- should be soon.
    Benchmark numbers from the Dragonwing IQ-9075 would be very welcome.

  2. ravediamond commented on Oct 5, 2026

    @ravediamond
    Contributor

    Here are some real-arm numbers for the new client, in case they're useful.

    Setup: Jetson Orin NX 16 GB (JetPack 6.2, MAXN_SUPER), vla.cpp f7e0f7f built on the board, and my own SmolVLA fine-tune on an SO-101 (front + wrist cameras), converted with convert_smolvla_to_gguf.py. I plugged VlaCppClient into my LeRobot rollout harness, so PyTorch and vla.cpp share the same robot I/O, recording and labelling. Task "Put the blue bowl in the pink bowl.", 25 s trials at 30 Hz, 50-action chunks, 10 trials each.

    PyTorch vla.cpp vla.cpp --flash-attn --mm-prec default
    Success 9/10 9/10 8/10
    Chunk inference 974 ms 390 ms 238 ms
    Loop blocked on inference 39% 20% 14%
    Control rate 19.3 Hz 24.6 Hz 26.6 Hz

    All failures were missed grasps, and at 10 trials the success rates are within noise.

    A few notes:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions