Run PyTorch models on Arduino microcontrollers using ExecuTorch.
This directory contains everything needed to package ExecuTorch as an Arduino library. A build script vendors the runtime sources from this repository into a self-contained library that Arduino users install through the Library Manager or by copying into their libraries folder.
Who this is for. This README is for maintainers of the packaging.
build_arduino_library.shis a release tool, not something an Arduino developer ever runs. Users install a prebuilt library from the Library Manager and never see this directory; their documentation lives in meta-pytorch/executorch-arduino.If you are here to change ExecuTorch and want to know whether you broke the Arduino library, read Keeping this working.
PyTorch Model ──► torch.export ──► .pte file ──► model.h (C array)
│
Arduino Sketch (.ino)
#include <ExecuTorch.h>
#include "model.h"
│
arduino-cli compile ──► Upload ──► Runs on board
-
The library (
arduino_lib/ExecuTorch/) — the ExecuTorch runtime, CMSIS-NN kernels, and portable ops packaged for the Arduino build system. Generated bybuild_arduino_library.sh; not checked in. -
The model (
model.h) — a.ptefile converted to a C byte array. Each user brings their own model, exported from PyTorch with the Cortex-M backend. -
The sketch (
.ino) — a standard Arduino program that loads the model, feeds it input, and reads the output. Uses the native ExecuTorch C++ API (Program::load,Method::execute, etc.).
| Board | MCU | Status |
|---|---|---|
| Arduino Uno Q | STM32U585 (Cortex-M33) | Tested |
| Arduino Nano 33 BLE | nRF52840 (Cortex-M4F) | Planned (requires mbed PAL) |
| Arduino Giga R1 WiFi | STM32H747 (Cortex-M7) | Planned (requires mbed PAL) |
| Arduino Portenta H7 | STM32H747 (Cortex-M7) | Planned (requires mbed PAL) |
The library currently requires the Zephyr board core. Non-Zephyr boards (mbed) need a platform abstraction layer port before they can compile. CMSIS-NN accelerated ops work on any ARM Cortex-M with DSP extensions. Portable ops work on any architecture.
cd examples/arduino
./build_arduino_library.shThis copies the required ExecuTorch sources from the repository into
arduino_lib/ExecuTorch/, ready for Arduino.
Copy the generated library into your Arduino libraries folder:
# macOS:
cp -r arduino_lib/ExecuTorch ~/Documents/Arduino/libraries/
# Linux:
cp -r arduino_lib/ExecuTorch ~/Arduino/libraries/Or with arduino-cli:
cd arduino_lib && zip -r ExecuTorch.zip ExecuTorch && cd ..
arduino-cli lib install --zip-path arduino_lib/ExecuTorch.zipEach sketch needs a model.h file — a .pte model converted to a C
byte array. Use pte_to_header.py to convert
any .pte file:
python examples/arduino/pte_to_header.py \
-p model.pte -d examples/arduino/examples/AddModel -o model.hAddModel — export a simple add model (no dataset needed):
python -c "
import torch
from executorch.exir import to_edge
from torch.export import export
class Add(torch.nn.Module):
def forward(self, x): return x + 1.0
et = to_edge(export(Add().eval(), (torch.tensor([1.,2.,3.]),))).to_executorch()
with open('add.pte','wb') as f: f.write(bytes(et.buffer))"
python examples/arduino/pte_to_header.py \
-p add.pte -d examples/arduino/examples/AddModel -o model.hHelloExecuTorch — uses any valid model; the AddModel .pte works:
cp examples/arduino/examples/AddModel/model.h \
examples/arduino/examples/HelloExecuTorch/model.hKeywordSpotting — requires a quantized DS-CNN model. Generate it
with export_model.py:
# Download the dataset (one time, ~2.3 GB) — run from repo root:
python -c "import torchaudio; torchaudio.datasets.SPEECHCOMMANDS(
root='outputs/speech_commands', download=True)"
# Train DS-CNN, quantize with CMSIS-NN, and export model.h:
python examples/arduino/export_model.py \
--output examples/arduino/examples/KeywordSpotting/model.hThis trains DS-CNN on Google Speech Commands v2 (100 samples/class),
quantizes to int8 via CortexMQuantizer, calibrates with real MFCC
audio data, and exports a 54 KB .pte as a C header.
To export with a pre-trained checkpoint instead of training:
python examples/arduino/export_model.py --checkpoint my_weights.pth \
--output examples/arduino/examples/KeywordSpotting/model.hNote: build_arduino_library.sh requires schema headers from a prior
cmake build. If you haven't built ExecuTorch yet, run
./install_executorch.sh first.
#include <ExecuTorch.h>
#include "model.h"
using executorch::extension::BufferDataLoader;
using executorch::runtime::Error;
using executorch::runtime::HierarchicalAllocator;
using executorch::runtime::MemoryAllocator;
using executorch::runtime::MemoryManager;
using executorch::runtime::Method;
using executorch::runtime::MethodMeta;
using executorch::runtime::Program;
using executorch::runtime::Result;
using executorch::runtime::Span;
alignas(16) uint8_t method_pool[28 * 1024];
void setup() {
Serial.begin(115200);
delay(2000);
executorch::runtime::runtime_init();
auto loader = BufferDataLoader(model_pte, sizeof(model_pte));
Result<Program> program = Program::load(&loader);
if (!program.ok()) {
Serial.println("Failed to load program");
return;
}
// ... load method, set inputs, execute, read outputs
// See examples/ for complete working sketches.
}
void loop() {
// Run inference periodically
delay(2000);
}The sketch uses the native ExecuTorch C++ API — the same API used on Linux, Android, and bare-metal targets. No wrapper layer, no Arduino-specific abstractions.
arduino-cli compile --fqbn arduino:zephyr:unoq:link_mode=static MySketch
arduino-cli upload --fqbn arduino:zephyr:unoq:link_mode=static -p /dev/cu.usbmodem* MySketch
arduino-cli monitor -p /dev/cu.usbmodem* --config baudrate=115200The library is a generated artifact. Everything under the generated
src/ is copied out of this repository, and the example models are
exported by this repository's Python. That gives one failure mode, and it
has cost multiple days:
The model and the library must come from the same ExecuTorch commit.
Cortex-M operator schemas change. scratch was added to the conv operators
on 2026-06-09 and to avg_pool2d later still. A .pte exported before a
schema change passes Program::load, resolves every operator, and then
fails inside Method::execute with InvalidProgram (0x23), because the
generated kernel wrapper expects one more argument than the model supplies.
Nothing about that error names the real cause.
This bites hardest when the Python package and the C++ sources come from
different places. pip install executorch gives a release wheel that can be
months behind this checkout; the library you build here is current. Check
which one you are exporting with:
python -c "import executorch.backends.cortex_m.ops.operators as o; print(o.__file__)"If that prints a site-packages path rather than your checkout, run
./install_executorch.sh first. Note that ExecuTorch refuses to build from a
directory not named exactly executorch (pytorch#6475), which is
a common reason people end up on a stale wheel without realising.
To check a model against a library without a board, decode the .pte and
compare each KernelCall's argument count against the stack.size() == N
in the generated src/executorch/codegen/RegisterCodegenUnboxedKernels*.cpp.
A mismatch there is the bug, found in seconds instead of hours.
link_mode=staticis mandatory, and Dynamic is the default. In the IDE that isTools > Link mode > Static, which is per-sketch and resets every time you open another example. Dynamic does a relocatable link (-r), which makes--gc-sectionsinert: nothing unused is stripped, so every vendored operator survives and the build overflows flash withSketch too big; text section exceeds available space. Same sketch, core 0.90.0: 507,876 bytes on Static, 787,508 on Dynamic. It fails at compile time, so there is nothing to see on the serial port either way.ET_LOGhas to be routed somewhere.zephyr.cpplogs throughfprintf, andplatform_stubs.cstubsfprintfout. The build script rewrites the logger to call a weaket_arduino_loghook, which the examples implement againstSerial. Without it every runtime failure is a bare hex code.- Only one platform backend may ship.
minimal.cppandzephyr.cppboth defineet_pal_*; shipping both leaves the choice to link order, andminimal's logger is empty and its allocator returnsnullptr. - Compiling proves very little. Every failure worth finding here compiled cleanly first. Flash a board.
| Symptom | Cause |
|---|---|
Sketch too big; text section exceeds available space |
Built in Dynamic link mode. Set Tools > Link mode > Static, per sketch |
| No serial output at all | Arduino_RouterBridge missing, or the monitor attached after setup() had already printed |
Program::load -> 0x23 |
Model header put the array in a section the linker discards; use pte_to_header.py from this directory, not the Ethos-U one |
load_method -> 0x14 |
Operator not in the registered set; regenerate with ROOT_OPS= |
load_method -> 0x21 |
method_pool too small; the log line gives the exact shortfall |
execute -> 0x23 |
Model and library built from different ExecuTorch commits |
The build_arduino_library.sh script assembles these components from
the ExecuTorch repository:
| Component | Source in repo | Purpose |
|---|---|---|
| ET Runtime | runtime/executor/, runtime/core/, runtime/kernel/, runtime/platform/ |
Model loading, memory management, op dispatch |
| Portable Ops | kernels/portable/ |
Software op implementations (any CPU) |
| Cortex-M Ops | backends/cortex_m/ops/ |
CMSIS-NN accelerated int8 ops |
| CMSIS-NN | fetched by cmake / Zephyr module | ARM's optimized DSP kernels |
| flatcc | third-party/flatcc/ |
.pte file parsing |
| flatbuffers | third-party/flatbuffers/ |
Schema headers |
| c10 | runtime/core/portable_type/c10/ |
Core type definitions |
The library uses no external dependencies beyond what the Arduino board core provides.
The build script applies these patches to make ExecuTorch compile under Arduino's build system:
-
#include <exception>before<variant>— Arduino's custom<new>header omits<exception>, breakingstd::bad_variant_access. -
cmake_macros.hstub — c10/torch headers expect a cmake-generated file. The build script generates a stub;C10_USING_CUSTOM_GENERATED_MACROSis defined inExecuTorch.hto skip the include. -
platform_stubs.c— provides weak stubs for_Exit(),fprintf(), and__aeabi_f2lzfor the LLEXT environment on boards that lack them. -
Compile-time defines —
ExecuTorch.hsetsET_ENABLE_DEPRECATED_CONSTANT_BUFFER=0(requires models exported with current ExecuTorch) andFLATBUFFERS_MAX_ALIGNMENT=1024.
./build_arduino_library.sh # rebuild
./build_arduino_library.sh --clean # remove generated output
ROOT_OPS="aten::add.out,..." ./build_arduino_library.sh # pick the op set
ALL_OPS=1 ./build_arduino_library.sh # every portable opThe op set is a size decision. Registering every portable kernel costs about 1.6 MB of text, twice the Uno Q's flash, because portable kernels are dtype-templated. The default registers the Cortex-M operators plus a small portable set, which lands around a quarter of flash.
Each example ships a model.pte that the build script converts to the
model.h its sketch includes. Regenerate them whenever an operator schema
changes, or the models will fail at execute against the new runtime:
# keyword spotting, from the checked-in checkpoint (no retraining)
python export_model.py --checkpoint examples/KeywordSpotting/model.pth \
--output /tmp/kws.hRELEASING.md is the checklist for cutting a version of the published library: what to verify, how to publish without leaving stale files behind, and what has gone wrong before.
The published library records the commit it was generated from in
extras/PROVENANCE.txt, and pins that commit in executorch_pin.txt
alongside it — the same one-SHA-per-file convention ExecuTorch uses in
.ci/docker/ci_commit_pins/. To move it forward:
- Update
executorch_pin.txtto the new ExecuTorch commit - Regenerate the library from a checkout at that commit
- Re-export the example models from the same checkout
- Confirm each model's
KernelCallargument counts match the regeneratedRegisterCodegenUnboxedKernels*.cpp - Compile every example at
link_mode=static, and flash at least one
Steps 2 and 3 have to happen together. Bumping the library without re-exporting the models is the mismatch described in Keeping this working.
arduino-cli compile --fqbn arduino:zephyr:unoq:link_mode=static examples/HelloExecuTorch
arduino-cli upload --fqbn arduino:zephyr:unoq:link_mode=static -p /dev/cu.usbmodem* examples/HelloExecuTorch
arduino-cli monitor -p /dev/cu.usbmodem* --config baudrate=115200The library is published by adding its repository URL to the Arduino Library Registry. After the initial registration, new git tags are picked up automatically.
Tested on Arduino Uno Q (STM32U585, Cortex-M33 @ 160 MHz):
- Portable ops: Add model (
x + 1.0) produces correct output[1,2,3] + 1 = [2.0, 3.0, 4.0] - CMSIS-NN linear: Quantized linear model (int8, 2.2 KB) runs
arm_fully_connected_s8viacortex_m::quantized_linear - CMSIS-NN keyword spotting: DS-CNN (MLPerf Tiny KWS benchmark, 54 KB, int8) correctly classifies real audio from Google Speech Commands dataset via 16 CMSIS-NN accelerated ops (conv2d, depthwise conv2d, avgpool, linear, quantize, dequantize, pad)
Verified with real audio on hardware:
"yes" → [yes]=7.82 >>> Detected: yes CORRECT!
"no" → [no]=1.60 >>> Detected: no CORRECT!
To test different keywords, change one line in the sketch:
// In KeywordSpotting.ino, change this line:
#include "mfcc_yes.h" // → detects "yes"
// #include "mfcc_no.h" // → detects "no"
// #include "mfcc_stop.h" // → detects "stop"
// Available: mfcc_yes.h, mfcc_no.h, mfcc_up.h, mfcc_down.h,
// mfcc_left.h, mfcc_right.h, mfcc_on.h, mfcc_off.h,
// mfcc_stop.h, mfcc_go.hTo test with your own audio recording:
python generate_test_input.py --input my_recording.wav --output mfcc_custom.h
# Then: #include "mfcc_custom.h" in the sketchGoogle Speech Commands "yes" audio (.wav, 16kHz, 1 second)
→ MFCC extraction (49 time frames × 10 coefficients)
→ DS-CNN model (23K params, trained on MacBook CPU)
→ CortexMQuantizer → int8 (calibrated with real MFCC data)
→ CMSIS-NN ops (conv2d, depthwise_conv2d, avgpool, linear)
→ Export to .pte (54 KB)
→ Arduino library → arduino-cli compile → upload
→ Cortex-M33 @ 160 MHz (Arduino Uno Q, STM32U585)
→ Serial output: ">>> yes" ✅
Training and test audio from Google Speech Commands v2
— 65,000 one-second recordings of 35 words spoken by thousands of
people. Standard dataset used by the MLPerf Tiny benchmark. Download
via torchaudio.datasets.SPEECHCOMMANDS (2.3 GB).
The dataset is © Google, released under
CC BY 4.0, which asks for
attribution. The keyword spotting weights checked in here
(examples/KeywordSpotting/model.pth and the .pte generated from it) are
trained on it and carry the same attribution.
Only the ten keyword classes are needed, so the full archive never has to land on disk:
mkdir -p outputs/speech_commands/SpeechCommands/speech_commands_v0.02
cd outputs/speech_commands/SpeechCommands/speech_commands_v0.02
curl -sL http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz \
| tar xz ./yes ./no ./up ./down ./left ./right ./on ./off ./stop ./goThat is 1.2 GB extracted instead of 2.3 GB downloaded plus 2.4 GB unpacked.
download.tensorflow.org serves no usable HTTPS -- its certificate does not
cover that hostname -- which is why the URL is plain HTTP and why
torchaudio.datasets.SPEECHCOMMANDS uses HTTP for it too. If transport
integrity matters, download the archive first and check it against the SHA-256
torchaudio pins for v0.02 before extracting:
af14739ee7dc311471de98f5f9d2c9191b18aedfe957f4a6ff791c709868ff58You only need this to retrain. The exported model and its checkpoint are both checked in, so nothing here is required to build the library.
The DS-CNN KWS benchmark uses 12 output classes (silence, unknown, plus 10 keywords). The Arduino export script trains the 10 keyword classes: yes, no, up, down, left, right, on, off, stop, go.
This applies to the Zephyr board core, which is the only core the library
currently supports (architectures=zephyr). Other Arduino cores do not run
Zephyr and have no link mode setting; they need a platform abstraction layer
port before they can compile at all, and their memory behaviour is untested.
On the Zephyr core the Uno Q defaults to Dynamic, and every example fails to
build that way. In the IDE, set Tools > Link mode > Static. It is a per-sketch
setting: opening another example puts it back to Dynamic.
Dynamic builds the sketch as a Zephyr loadable extension via a relocatable link
(-r). --gc-sections is passed in both modes but can only work in a final
link, where the linker has an entry point to trace reachability from. Under -r
there is nothing to trace from and the output will be linked again later, so
every section has to be kept. Nothing unused is stripped, and because a loadable
extension is loaded into RAM, the retained code is charged against RAM as well
as flash:
| AddModel, core 0.90.0 | Flash | RAM |
|---|---|---|
link_mode=static |
507,876 (64%) | 12,268 (4%) |
link_mode=dynamic |
787,508 (100%) — overflows | 249,021 (94%) |
This is not something the library can work around. Trimming the vendored operator set to only the registered ops — 15 sources instead of 172 — moved the Dynamic build by 852 bytes, still over the limit. The bulk is the runtime, flatbuffers and CMSIS-NN, all of which Static strips and Dynamic cannot.
Measured on an Arduino Uno Q, board core 0.90.0, at link_mode=static,
against 786,432 bytes of flash and 262,144 bytes of RAM:
| Build | Flash | RAM | On hardware |
|---|---|---|---|
| HelloExecuTorch | 473,368 (60%) | 3,060 (1%) | Model loaded OK!, 1 method |
| AddModel | 509,456 (64%) | 12,276 (4%) | [1,2,3] + 1 = [2.00, 3.00, 4.00] |
| KeywordSpotting (CMSIS-NN) | 564,664 (71%) | 45,044 (17%) | detects yes, logit 8.95 |
Core 0.55.2 reported a 131,072-byte RAM ceiling; 0.90.0 reports 262,144. Figures from before that change are not comparable.
All CMSIS-NN sources are compiled, but the linker's
--gc-sections discards unused functions from the final binary.
RAM is the binding constraint, not flash. Zephyr reserves a main stack and a heap
before the sketch gets any, and the arena the sketch hands to MemoryManager
comes out of what remains. KeywordSpotting's DS-CNN plans 16 KB of buffers but
needs considerably more for the method's own structures, which leaves a usable
band rather than a floor to clear.
The numbers move with the board core, so both are recorded. On 0.90.0, which is what Library Manager installs today:
| Arena | Result |
|---|---|
| 28 KB | load_method fails, MemoryAllocationFailed (0x21), 180 B short |
| 40 KB | Works. What the sketch ships |
| 48 KB | Works, with more margin. Globals reach 47 KB of the 256 KB the core reports |
On 0.55.2, where Zephyr took a 32 KB stack and a 32 KB heap out of 128 KB:
| Arena | Result |
|---|---|
| 28 KB | Worked. This is what the example shipped with |
| 64 KB | Globals reach 70,644, past Zephyr's reservation. It ran and printed the right answer, which is what makes it dangerous rather than safe |
That 28 KB row is the whole point of recording both. It was the shipped value and it was correct, until the core changed underneath it. 64 KB has not been retried on 0.90.0, so whether the same reservation cliff exists there is unknown.
The sketch ships 40 KB for that reason, and growing it is not automatically
safer: arduino-cli reports RAM against the whole region and knows nothing about
Zephyr's stack and heap, so the 64 KB build still reported a comfortable 53%
while overlapping reserved memory. Read the arena as bounded on both sides.
The upper bound has not been re-measured on core 0.90.0, which reports twice the RAM (262,144). 40 KB is confirmed working there — 47,084 bytes of globals, 17% — but whether the ceiling moved with the reported total is unverified. Treat anything above 40 KB as untested on 0.90.0.