Skip to content
 
 

Latest commit

 

History

4,866 Commits

Folders and files

Repository files navigation

qvac-fabric-speech.cpp

On-device speech and audio AI in C++17 on ggml: speech-to-text, speaker diarization, end-of-utterance detection, text-to-speech, voice cloning, sound effects, speech-to-speech, enhancement, and music generation.

Property Value
Build Feature-gated CMake superbuild over Whisper and three native engines
Runtime No Python, PyTorch, NeMo, or ONNX Runtime at inference time
Model formats Whisper/Silero GGML .bin; Parakeet, TTS, and AudioGen GGUF
Desktop Linux, macOS, Windows
Mobile Android arm64-v8a, iOS arm64; model and backend coverage varies by engine
Backends CPU, Metal, Vulkan, OpenCL, CUDA; selected models also support explicit Hexagon placement or optional Apple Core ML sidecars
Shared ggml One ggml-speech dependency from qvac-ext-ggml@speech

Supported models

The engine guides own model formats, quantization, backend validation, and input requirements. An available ggml backend does not imply every model has been validated on it.

Engine Model families Capabilities and limitations
Whisper tiny, base, small, medium, large v1/v2/v3, large-v3-turbo; Silero v5.1.2/v6.2.0 ASR, multilingual-to-English translation, VAD, tinydiarize speaker turns; .en checkpoints are English-only and turbo is intended for transcription
Parakeet CTC, Unified RNN-T, TDT, IndicConformer, realtime EOU, Nemotron 3.5 ASR Offline, streaming, long-form ASR and end-of-turn detection; IndicConformer CTC requires a language selection
Parakeet Sortformer v1/v2/v2.1, Nemotron 3 Diarization Speaker diarization and attributed ASR; Sortformer supports up to four speakers, Nemotron up to eight
Parakeet MOSS-Transcribe-Diarize One-pass transcription, speakers, timestamps and hotwords; CPU/Metal validated on Spanish/Chinese; English long-form skips spans in both native and reference runs
TTS Chatterbox Turbo/Multilingual, Supertonic 1/2/3, Parler mini/large/Indic, CosyVoice3, Audio8 Synthesis with preset, described, or cloned voices, depending on the model; streaming support varies by API and CLI
TTS Pocket TTS English synthesis and incremental CPU streaming; cloning requires encoder-enabled weights
TTS MOSS-TTS-v1.5 / MOSS-TTSD Multilingual synthesis, cloning, dialogue and streaming; CPU/Metal validated, numerical reference parity pending
TTS MOSS-SoundEffect-v2, MOSS-Speech Sound-effect generation and spoken replies respectively; CPU/Metal validated, other GPU backends untested
TTS LavaSR denoiser/enhancer Speech denoising and bandwidth extension
AudioGen ACE-Step v15 turbo/sft/base Music generation and editing; base additionally supports lego stems
AudioGen MiniMax-Music3 Desktop music generation on CPU or GPU; unavailable on Android/iOS

Apple Core ML sidecars

Optional sidecars accelerate a stage while the rest of the pipeline stays on ggml. They default to disabled and have model-specific input and fallback contracts; see the engine guides before exporting or deploying one.

Engine Accelerated stage Guide
Whisper Encoder Whisper Core ML
Parakeet Eligible ASR and Sortformer v2.1 encoder paths Backend and routing contracts
TTS Supertonic vocoder, Audio8 codec synthesis Supertonic, Audio8
AudioGen ACE-Step VAE decoder AudioGen backends

Getting started

Install a shared speech-branch ggml using docs/BUILD.md, then configure from this checkout with its install prefix:

cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCMAKE_PREFIX_PATH=/path/to/ggml-install
cmake --build build --parallel

For a first transcription on a single-config build:

bash third_party/whisper.cpp/models/download-ggml-model.sh base.en
./build/bin/whisper-cli -m third_party/whisper.cpp/models/ggml-base.en.bin \
  -f third_party/whisper.cpp/samples/jfk.wav

The build guide provides Windows and multi-config commands. Use docs/CLI.md to choose a tool, and the engine's model guide to acquire its required weights.

Performance

RTF = inference_time / audio_duration; lower is faster. These are dated measurements with different workloads and timing boundaries, not universal speed guarantees. Full methods, build pins and limitations live in docs/PERFORMANCE.md and the linked engine reports.

Task Model Recorded measurements
ASR Parakeet TDT 0.6b v3 September 2026 GPU lanes: short-clip RTF 0.0007–0.0055; method and full table
TTS Supertonic 3 September 2026 GPU lanes: process wall 0.61–0.82 s, including model load; full table
TTS Audio8 September 2026 GPU campaign: RTF 0.20–0.65, with later RTX 5090 measurements down to 0.116; full table and follow-up
Music ACE-Step 1.5 September 2026 CUDA/Vulkan/Metal lanes: generation 1,338–9,016 ms, RTF 0.139–0.957; full table
Music MiniMax-Music3 f16 September 2026 RTX 5090: two-minute song in 73.9 s CUDA / 84.2 s Vulkan, excluding model load; full table

Use in QVAC

These engines ship inside QVAC as SDK addons consuming the speech-cpp vcpkg port. For JavaScript/TypeScript APIs, Bare runtime integration, and desktop/mobile applications, see QVAC.

QVAC addon Wraps speech-cpp features
@qvac/asr-ggml ASR, diarization, end-of-utterance whisper, parakeet
@qvac/tts-ggml Synthesis, voice cloning, enhancement tts
@qvac/audiogen-ggml Music generation audiogen
@qvac/bci-whispercpp Brain-computer interface transcription whisper

Licenses

Code and model weights have separate terms. No model weights are shipped in this source tree. Engine license sections and NOTICE files identify sources and model-specific exceptions.

Component Code Model weights
Whisper MIT OpenAI Whisper MIT; Silero under its model terms
Parakeet Apache-2.0 NVIDIA Parakeet/Sortformer CC-BY-4.0; EOU NVIDIA Open Model License; Nemotron OpenMDW-1.1; IndicConformer MIT; MOSS-Transcribe-Diarize Apache-2.0
TTS MIT Chatterbox MIT; Pocket CC-BY-4.0; Supertonic OpenRAIL-M; Parler, CosyVoice3, Audio8, LavaSR and MOSS models Apache-2.0
AudioGen MIT ACE-Step 1.5 MIT; Qwen3-Embedding Apache-2.0; MiniMax-Music3 Community License

Documentation

Topic Where
Architecture and repository layout docs/ARCHITECTURE.md
Building and consuming packages docs/BUILD.md
CLI selection and examples docs/CLI.md
Performance overview and reports docs/PERFORMANCE.md
Whisper models and formats docs/WHISPER.md
Whisper subtree deltas and synchronization PATCHES.md, docs/UPSTREAM-SYNC.md
ASR, diarization, end-of-utterance Parakeet
Synthesis, voice cloning, sound effects, speech-to-speech, enhancement TTS
Music generation and editing AudioGen
Benchmark quality diagnostics and model preparation Music alignment guide
Historical development and integration reports Chatterbox, Supertonic, Pocket, Parakeet, Strix Halo

About

QVAC Fabric: cross-platform Speech inference engine, optimized for edge devices

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages