On-device speech and audio AI in C++17 on ggml: speech-to-text, speaker diarization, end-of-utterance detection, text-to-speech, voice cloning, sound effects, speech-to-speech, enhancement, and music generation.
| Property | Value |
|---|---|
| Build | Feature-gated CMake superbuild over Whisper and three native engines |
| Runtime | No Python, PyTorch, NeMo, or ONNX Runtime at inference time |
| Model formats | Whisper/Silero GGML .bin; Parakeet, TTS, and AudioGen GGUF |
| Desktop | Linux, macOS, Windows |
| Mobile | Android arm64-v8a, iOS arm64; model and backend coverage varies by engine |
| Backends | CPU, Metal, Vulkan, OpenCL, CUDA; selected models also support explicit Hexagon placement or optional Apple Core ML sidecars |
| Shared ggml | One ggml-speech dependency from qvac-ext-ggml@speech |
The engine guides own model formats, quantization, backend validation, and input requirements. An available ggml backend does not imply every model has been validated on it.
| Engine | Model families | Capabilities and limitations |
|---|---|---|
| Whisper | tiny, base, small, medium, large v1/v2/v3, large-v3-turbo; Silero v5.1.2/v6.2.0 | ASR, multilingual-to-English translation, VAD, tinydiarize speaker turns; .en checkpoints are English-only and turbo is intended for transcription |
| Parakeet | CTC, Unified RNN-T, TDT, IndicConformer, realtime EOU, Nemotron 3.5 ASR | Offline, streaming, long-form ASR and end-of-turn detection; IndicConformer CTC requires a language selection |
| Parakeet | Sortformer v1/v2/v2.1, Nemotron 3 Diarization | Speaker diarization and attributed ASR; Sortformer supports up to four speakers, Nemotron up to eight |
| Parakeet | MOSS-Transcribe-Diarize | One-pass transcription, speakers, timestamps and hotwords; CPU/Metal validated on Spanish/Chinese; English long-form skips spans in both native and reference runs |
| TTS | Chatterbox Turbo/Multilingual, Supertonic 1/2/3, Parler mini/large/Indic, CosyVoice3, Audio8 | Synthesis with preset, described, or cloned voices, depending on the model; streaming support varies by API and CLI |
| TTS | Pocket TTS | English synthesis and incremental CPU streaming; cloning requires encoder-enabled weights |
| TTS | MOSS-TTS-v1.5 / MOSS-TTSD | Multilingual synthesis, cloning, dialogue and streaming; CPU/Metal validated, numerical reference parity pending |
| TTS | MOSS-SoundEffect-v2, MOSS-Speech | Sound-effect generation and spoken replies respectively; CPU/Metal validated, other GPU backends untested |
| TTS | LavaSR denoiser/enhancer | Speech denoising and bandwidth extension |
| AudioGen | ACE-Step v15 turbo/sft/base | Music generation and editing; base additionally supports lego stems |
| AudioGen | MiniMax-Music3 | Desktop music generation on CPU or GPU; unavailable on Android/iOS |
Optional sidecars accelerate a stage while the rest of the pipeline stays on ggml. They default to disabled and have model-specific input and fallback contracts; see the engine guides before exporting or deploying one.
| Engine | Accelerated stage | Guide |
|---|---|---|
| Whisper | Encoder | Whisper Core ML |
| Parakeet | Eligible ASR and Sortformer v2.1 encoder paths | Backend and routing contracts |
| TTS | Supertonic vocoder, Audio8 codec synthesis | Supertonic, Audio8 |
| AudioGen | ACE-Step VAE decoder | AudioGen backends |
Install a shared speech-branch ggml using docs/BUILD.md, then configure from this checkout with its install prefix:
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCMAKE_PREFIX_PATH=/path/to/ggml-install
cmake --build build --parallelFor a first transcription on a single-config build:
bash third_party/whisper.cpp/models/download-ggml-model.sh base.en
./build/bin/whisper-cli -m third_party/whisper.cpp/models/ggml-base.en.bin \
-f third_party/whisper.cpp/samples/jfk.wavThe build guide provides Windows and multi-config commands. Use docs/CLI.md to choose a tool, and the engine's model guide to acquire its required weights.
RTF = inference_time / audio_duration; lower is faster. These are dated
measurements with different workloads and timing boundaries, not universal
speed guarantees. Full methods, build pins and limitations live in
docs/PERFORMANCE.md and the linked engine reports.
| Task | Model | Recorded measurements |
|---|---|---|
| ASR | Parakeet TDT 0.6b v3 | September 2026 GPU lanes: short-clip RTF 0.0007–0.0055; method and full table |
| TTS | Supertonic 3 | September 2026 GPU lanes: process wall 0.61–0.82 s, including model load; full table |
| TTS | Audio8 | September 2026 GPU campaign: RTF 0.20–0.65, with later RTX 5090 measurements down to 0.116; full table and follow-up |
| Music | ACE-Step 1.5 | September 2026 CUDA/Vulkan/Metal lanes: generation 1,338–9,016 ms, RTF 0.139–0.957; full table |
| Music | MiniMax-Music3 f16 | September 2026 RTX 5090: two-minute song in 73.9 s CUDA / 84.2 s Vulkan, excluding model load; full table |
These engines ship inside QVAC as SDK
addons consuming the speech-cpp vcpkg port. For JavaScript/TypeScript APIs,
Bare runtime integration, and desktop/mobile applications, see QVAC.
| QVAC addon | Wraps | speech-cpp features |
|---|---|---|
@qvac/asr-ggml |
ASR, diarization, end-of-utterance | whisper, parakeet |
@qvac/tts-ggml |
Synthesis, voice cloning, enhancement | tts |
@qvac/audiogen-ggml |
Music generation | audiogen |
@qvac/bci-whispercpp |
Brain-computer interface transcription | whisper |
Code and model weights have separate terms. No model weights are shipped in
this source tree. Engine license sections and NOTICE files identify sources
and model-specific exceptions.
| Component | Code | Model weights |
|---|---|---|
| Whisper | MIT | OpenAI Whisper MIT; Silero under its model terms |
| Parakeet | Apache-2.0 | NVIDIA Parakeet/Sortformer CC-BY-4.0; EOU NVIDIA Open Model License; Nemotron OpenMDW-1.1; IndicConformer MIT; MOSS-Transcribe-Diarize Apache-2.0 |
| TTS | MIT | Chatterbox MIT; Pocket CC-BY-4.0; Supertonic OpenRAIL-M; Parler, CosyVoice3, Audio8, LavaSR and MOSS models Apache-2.0 |
| AudioGen | MIT | ACE-Step 1.5 MIT; Qwen3-Embedding Apache-2.0; MiniMax-Music3 Community License |
| Topic | Where |
|---|---|
| Architecture and repository layout | docs/ARCHITECTURE.md |
| Building and consuming packages | docs/BUILD.md |
| CLI selection and examples | docs/CLI.md |
| Performance overview and reports | docs/PERFORMANCE.md |
| Whisper models and formats | docs/WHISPER.md |
| Whisper subtree deltas and synchronization | PATCHES.md, docs/UPSTREAM-SYNC.md |
| ASR, diarization, end-of-utterance | Parakeet |
| Synthesis, voice cloning, sound effects, speech-to-speech, enhancement | TTS |
| Music generation and editing | AudioGen |
| Benchmark quality diagnostics and model preparation | Music alignment guide |
| Historical development and integration reports | Chatterbox, Supertonic, Pocket, Parakeet, Strix Halo |