Export nvidia/parakeet-tdt-0.6b-v3 speech recognition model to ExecuTorch.
pip install -r install_requirements.txtExport the model:
python export_parakeet_tdt.pyTest transcription on an audio file and compare eager vs lowered results:
python export_parakeet_tdt.py --audio /path/to/audio.wav| Argument | Description |
|---|---|
--output-dir |
Output directory for exports (default: ./parakeet_tdt_exports) |
--backend |
Backend for acceleration: portable, xnnpack, vulkan, metal, mlx, cuda, cuda-windows (default: xnnpack) |
--dtype |
Data type: fp32, bf16, fp16 (default: fp32). Metal backend supports fp32 and bf16 only (no fp16). |
--audio |
Path to audio file for transcription test |
Note: The preprocessor is always lowered with the portable backend regardless of the --backend setting.
The export script supports quantizing encoder and decoder linear layers using torchao.
| Argument | Description |
|---|---|
--qlinear_encoder |
Quantization config for encoder linear layers: 4w, 8w, 8da4w, 8da8w, fpa4w, nvfp4 |
--qlinear_encoder_group_size |
Group size for encoder linear quantization (default: auto) |
--qlinear_encoder_packing_format |
Packing format for encoder: tile_packed_to_4d |
--qlinear |
Quantization config for decoder linear layers: 4w, 8w, 8da4w, 8da8w, fpa4w, nvfp4 |
--qlinear_group_size |
Group size for decoder linear quantization (default: auto) |
--qlinear_packing_format |
Packing format for decoder: tile_packed_to_4d |
--qembedding |
Quantization config for decoder embedding layer: 4w, 8w, nvfp4 |
--qembedding_group_size |
Group size for embedding quantization (default: auto) |
| Config | Description | Backends |
|---|---|---|
4w |
4-bit weight only quantization | CUDA, MLX, XNNPACK (embedding only) |
8w |
8-bit weight only quantization | CUDA, MLX, XNNPACK (embedding only) |
8da4w |
8-bit dynamic activation, 4-bit weight | Vulkan, XNNPACK |
8da8w |
8-bit dynamic activation, 8-bit weight | XNNPACK |
fpa4w |
Floating point activation, 4-bit weight | Metal |
nvfp4 |
4-bit weight only quantization using NVIDIA's FP4 dtype | MLX |
python export_parakeet_tdt.py \
--backend xnnpack \
--qlinear_encoder 8da4w \
--qlinear_encoder_group_size 32 \
--qlinear 8da4w \
--qlinear_group_size 32 \
--output-dir ./parakeet_quantized_xnnpackpython export_parakeet_tdt.py \
--backend vulkan \
--qlinear_encoder 8da4w \
--qlinear_encoder_group_size 32 \
--qlinear 8da4w \
--qlinear_group_size 32 \
--vulkan_force_fp16 \
--output-dir ./parakeet_quantized_vulkanAn additional --vulkan_force_fp16 flag is available to have the Vulkan backend internally downcast FP32 tensors to FP16 within the Vulkan backend, forcing half-precision computation. Note that input/output tensors are still FP32, and the delegate will automatically convert them to/from FP16 upon entering and exiting the delegate. This will significantly improve latency but may slightly reduce transcription accuracy.
python export_parakeet_tdt.py \
--backend cuda \
--qlinear_encoder 4w \
--qlinear_encoder_group_size 32 \
--qlinear_encoder_packing_format tile_packed_to_4d \
--qlinear 4w \
--qlinear_group_size 32 \
--qlinear_packing_format tile_packed_to_4d \
--qembedding 8w \
--output-dir ./parakeet_quantizedNote: The tile_packed_to_4d packing format is optimized for CUDA.
python export_parakeet_tdt.py \
--backend metal \
--qlinear_encoder fpa4w \
--qlinear_encoder_group_size 32 \
--qlinear fpa4w \
--qlinear_group_size 32 \
--output-dir ./parakeet_metal_quantizedNote: Metal 4-bit quantization requires torchao built with experimental MPS (Metal) ops.
You can install torchao with Metal support from the ao repo:
USE_CPP=1 TORCHAO_BUILD_EXPERIMENTAL_MPS=1 pip install . --no-build-isolationAlternatively, you can build torchao with Metal support while installing ExecuTorch:
EXECUTORCH_BUILD_KERNELS_TORCHAO=1 TORCHAO_BUILD_EXPERIMENTAL_MPS=1 ./install_executorch.shpython export_parakeet_tdt.py --backend metal --output-dir ./parakeet_metalThis generates:
model.pte- The compiled Parakeet TDT model (includes Metal kernel blob)tokenizer.model- SentencePiece tokenizer
python export_parakeet_tdt.py --backend cuda --output-dir ./parakeet_cudaThis generates:
model.pte- The compiled Parakeet TDT modelaoti_cuda_blob.ptd- CUDA kernel blob required at runtimetokenizer.model- SentencePiece tokenizer
Before running cuda-windows export, make sure these requirements are set up:
x86_64-w64-mingw32-g++is installed and onPATH(mingw-w64 cross-compiler).WINDOWS_CUDA_HOMEpoints to the extracted Windows CUDA package directory.
Example setup on Ubuntu:
# 1) Install cross-compiler + extraction tools
sudo apt-get update
sudo apt-get install -y --no-install-recommends \
g++-mingw-w64-x86-64-posix mingw-w64-tools p7zip-full wget
# 2) Verify cross-compiler
x86_64-w64-mingw32-g++ --version
# 3) Download and extract Windows CUDA installer package
CUDA_VERSION=12.8.1
CUDA_DRIVER_VERSION=572.61
CUDA_INSTALLER="cuda_${CUDA_VERSION}_${CUDA_DRIVER_VERSION}_windows.exe"
CUDA_URL="https://developer.download.nvidia.com/compute/cuda/${CUDA_VERSION}/local_installers/${CUDA_INSTALLER}"
mkdir -p /opt/cuda-windows
cd /opt/cuda-windows
wget -q "${CUDA_URL}" -O "${CUDA_INSTALLER}"
7z x "${CUDA_INSTALLER}" -oextracted -y
# 4) Point WINDOWS_CUDA_HOME to extracted Windows CUDA payload
export WINDOWS_CUDA_HOME=/opt/cuda-windows/extracted/cuda_cudart/cudartpython export_parakeet_tdt.py --backend cuda-windows --output-dir ./parakeet_cuda_windowsThis generates:
model.pte- The compiled Parakeet TDT modelaoti_cuda_blob.ptd- CUDA kernel blob required at runtime
Export with MLX backend (bf16, int4 quantized, group size 128):
python export_parakeet_tdt.py \
--backend mlx \
--dtype bf16 \
--qlinear_encoder 4w \
--qlinear_encoder_group_size 128 \
--qlinear 4w \
--qlinear_group_size 128 \
--output-dir ./parakeet_mlx_4wExport with MLX backend (bf16, NVFP4 quantized):
python export_parakeet_tdt.py \
--backend mlx \
--dtype bf16 \
--qlinear_encoder nvfp4 \
--qlinear nvfp4 \
--qembedding 4w \
--output-dir ./parakeet_mlx_nvfp4Note: Although MLX supports NVFP4 embedding quantization, Parakeet's embedding layer has dimensions not divisible by 16, which is incompatible with NVFP4. Use
4wfor embeddings instead.
This generates:
model.pte- The compiled model with MLX delegate (~470 MB)tokenizer.model- SentencePiece tokenizer
From the executorch root directory:
# CPU/XNNPACK build
make parakeet-cpu
# Metal build (macOS)
make parakeet-metal
# Vulkan build (Linux / Android)
make parakeet-vulkan
# CUDA build (Linux)
make parakeet-cuda
# MLX build (macOS)
make parakeet-mlxEach Parakeet build now produces both:
parakeet_runnerfor one-shot CLI transcription from an audio fileparakeet_helperfor long-lived host integrations that keep the model warm and stream PCM requests over stdin/stdout
On Windows (PowerShell), use CMake workflow presets directly:
cmake --workflow --preset llm-release-cuda
Push-Location examples/models/parakeet
cmake --workflow --preset parakeet-cuda
Pop-LocationFrom the executorch root directory:
# CPU/XNNPACK
./cmake-out/examples/models/parakeet/parakeet_runner \
--model_path examples/models/parakeet/parakeet_tdt_exports/model.pte \
--audio_path /path/to/audio.wav \
--tokenizer_path examples/models/parakeet/parakeet_tdt_exports/tokenizer.model
# Metal
DYLD_LIBRARY_PATH=/usr/lib ./cmake-out/examples/models/parakeet/parakeet_runner \
--model_path examples/models/parakeet/parakeet_metal/model.pte \
--audio_path /path/to/audio.wav \
--tokenizer_path examples/models/parakeet/parakeet_metal/tokenizer.model
# Vulkan
./cmake-out/examples/models/parakeet/parakeet_runner \
--model_path examples/models/parakeet/parakeet_vulkan/model.pte \
--audio_path /path/to/audio.wav \
--tokenizer_path examples/models/parakeet/parakeet_vulkan/tokenizer.model
# CUDA (include .ptd data file)
./cmake-out/examples/models/parakeet/parakeet_runner \
--model_path examples/models/parakeet/parakeet_cuda/model.pte \
--data_path examples/models/parakeet/parakeet_cuda/aoti_cuda_blob.ptd \
--audio_path /path/to/audio.wav \
--tokenizer_path examples/models/parakeet/parakeet_cuda/tokenizer.model
# MLX
./cmake-out/examples/models/parakeet/parakeet_runner \
--model_path examples/models/parakeet/parakeet_mlx_4w/model.pte \
--audio_path /path/to/audio.wav \
--tokenizer_path examples/models/parakeet/parakeet_mlx_4w/tokenizer.modelWindows (PowerShell):
.\cmake-out\examples\models\parakeet\Release\parakeet_runner.exe `
--model_path C:\path\to\parakeet_cuda_windows\model.pte `
--data_path C:\path\to\parakeet_cuda_windows\aoti_cuda_blob.ptd `
--audio_path C:\path\to\audio.wav `
--tokenizer_path C:\path\to\parakeet_cuda_windows\tokenizer.modelIf your generator is single-config, the runner may be at .\cmake-out\examples\models\parakeet\parakeet_runner.exe instead.
| Argument | Description |
|---|---|
--model_path |
Path to Parakeet model (.pte) |
--audio_path |
Path to input audio file (.wav) |
--tokenizer_path |
Path to tokenizer file (default: tokenizer.json) |
--data_path |
Path to data file (.ptd) for delegate data (required for CUDA/CUDA-Windows) |
--timestamps |
Timestamp output mode: none|token|word|segment|all (default: segment) |
The helper binary uses the same Parakeet transcription stack as parakeet_runner,
but keeps the model loaded across multiple requests so host apps can avoid repeated
startup and model load overhead.
Example:
# Metal
DYLD_LIBRARY_PATH=/usr/lib ./cmake-out/examples/models/parakeet/parakeet_helper \
--model_path examples/models/parakeet/parakeet_metal/model.pte \
--tokenizer_path examples/models/parakeet/parakeet_metal/tokenizer.modelThe helper accepts framed requests over stdin, validates 16 kHz mono float32 PCM
payloads, and returns status/result messages over stdout. It is intended for app
integrations such as the macOS ExecuWhisper frontend in the separate
executorch-examples repository.
Check out a demo Android app for Parakeet in the separate executorch-examples repository.