diff --git a/AGENTS.md b/AGENTS.md index f131286..db8baf0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -203,6 +203,8 @@ docs/ diarization.md , speaker diarization + speaker-attributed ASR: parity, C-API, speed sound.md , sound-event detection (CED) and the combined scene stream speaker.md , speaker identification: enroll, scene naming, C-API v9 and v10, measured numbers + vad.md, vad-benchmarks.md, ultra-redux.md, batching.md, performance.md, cli.md, capi.md, docker.md, licenses.md, tdt-nbest.md + , see the documentation table in README.md .github/workflows/ ci.yml , build job (per-push) + closed-loop job (pull_request + dispatch) ``` @@ -223,7 +225,7 @@ cmake -B build -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=ON && cmake --build build | `PARAKEET_GGML_CUDA` | OFF | Forward GGML_CUDA to the submodule | | `PARAKEET_GGML_METAL` | OFF | Forward GGML_METAL to the submodule | | `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule | -| `PARAKEET_GGML_HIPBLAS` | OFF | Forward GGML_HIPBLAS to the submodule | +| `PARAKEET_GGML_HIP` | OFF | Forward GGML_HIP (ROCm) to the submodule | | `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp | | `PARAKEET_WITH_VOICEDETECT` | ON | Speaker identification through voice-detect.cpp | diff --git a/README.md b/README.md index 7fb3c20..ceebfa7 100644 --- a/README.md +++ b/README.md @@ -6,554 +6,355 @@ [![License](https://img.shields.io/badge/License-MIT-green)](LICENSE) [![LocalAI](https://img.shields.io/badge/LocalAI-Run_Locally-orange)](https://github.com/mudler/LocalAI) -parakeet.cpp is a C++17 inference port of NVIDIA's [NeMo](https://github.com/NVIDIA-NeMo/NeMo) Parakeet speech-recognition models, built on [ggml](https://github.com/ggml-org/ggml). It gives you fast, dependency-light automatic speech recognition on CPU (and on GPU through ggml's backends), with no Python runtime needed at inference time. +A C++17/[ggml](https://github.com/ggml-org/ggml) port of NVIDIA's [NeMo](https://github.com/NVIDIA-NeMo/NeMo) Parakeet speech recognition models, with voice activity detection, speaker diarization and identification, and sound-event tagging. It runs on CPU and on GPU backends, reads self-contained GGUF files, and needs no Python at inference time. Transcripts match NeMo (WER 0 on every published NeMo checkpoint). -It covers all the offline Parakeet families (CTC, RNNT, TDT, and hybrid TDT-CTC, in 0.6B/1.1B/110M sizes, English plus multilingual v3), each validated at WER 0 against NeMo on every published checkpoint. It also does **cache-aware streaming with end-of-utterance (EOU) detection** for `parakeet_realtime_eou_120m-v1`, where the streaming transcript matches NeMo's cache-aware streaming byte for byte. And it supports the **multilingual, prompt-conditioned streaming model** `nvidia/nemotron-3.5-asr-streaming-0.6b` (40+ locales): pass a target language with `--lang ` (default `auto`) and both the offline and the cache-aware streaming transcripts match NeMo per language at WER 0. The full coverage matrix lives in `docs/parity.md`. - -It's faster than NeMo's PyTorch runtime on both CPU and GPU, with byte-identical transcripts. The full numbers, methodology, and all the plots are in [benchmarks/BENCHMARK.md](benchmarks/BENCHMARK.md). - -

- CPU speedup vs NeMo (RTFx ratio per dtype) - GPU speedup vs NeMo on the NVIDIA GB10 -

- -It also runs circles around whisper.cpp on the same audio: the 110M Parakeet is faster than whisper base.en and far faster than large-v3-turbo, while the larger Parakeets match or beat whisper's accuracy (see [benchmarks/BENCHMARK.md](benchmarks/BENCHMARK.md)). +![parakeet.cpp vs NeMo on GPU: identical output, parakeet.cpp finishes first](benchmarks/media/gpu_duel.gif) -

- parakeet.cpp vs whisper.cpp RTFx on CPU and GPU -

+> The same clip, side by side: identical output, parakeet.cpp finishes first (slowed down so a sub-100 ms race is watchable). More clips: [Demos](#demos). ---- +**On this page:** [What is new](#what-is-new) | [Demos](#demos) | [Quick start](#quick-start) | [Models](#models) | [CLI cheat sheet](#cli-cheat-sheet) | [Server and Docker](#server-and-docker) | [C API](#c-api) | [Build](#build) | [Benchmarks](#benchmarks) | [Documentation](#documentation) | [Limits](#limits) | [Contributing](#contributing) | [License](#license-and-credits) -## Supported models - -Every model below is validated at WER 0 against NeMo and published as GGUF (f16, q8_0, q6_k, q5_k, q4_k) in the single collection repo [mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf). Convert any of them yourself with `scripts/convert_parakeet_to_gguf.py`. The per-model parity matrix is in [docs/parity.md](docs/parity.md). - -| Model | Type | Size | Notes | Source | -| ----- | ---- | ---- | ----- | ------ | -| [parakeet-tdt_ctc-110m](https://huggingface.co/nvidia/parakeet-tdt_ctc-110m) | hybrid TDT+CTC | 110M | English, the small anchor checkpoint | NVIDIA | -| [parakeet-ctc-0.6b](https://huggingface.co/nvidia/parakeet-ctc-0.6b) | CTC | 0.6B | English | NVIDIA | -| [parakeet-rnnt-0.6b](https://huggingface.co/nvidia/parakeet-rnnt-0.6b) | RNNT | 0.6B | English | NVIDIA | -| [parakeet-tdt-0.6b-v2](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2) | TDT | 0.6B | English | NVIDIA | -| [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) | TDT | 0.6B | multilingual (25 European languages) | NVIDIA | -| [parakeet-ctc-1.1b](https://huggingface.co/nvidia/parakeet-ctc-1.1b) | CTC | 1.1B | English | NVIDIA | -| [parakeet-rnnt-1.1b](https://huggingface.co/nvidia/parakeet-rnnt-1.1b) | RNNT | 1.1B | English | NVIDIA | -| [parakeet-tdt-1.1b](https://huggingface.co/nvidia/parakeet-tdt-1.1b) | TDT | 1.1B | English | NVIDIA | -| [parakeet-tdt_ctc-1.1b](https://huggingface.co/nvidia/parakeet-tdt_ctc-1.1b) | hybrid TDT+CTC | 1.1B | English | NVIDIA | -| [parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1) | RNNT, streaming | 120M | cache-aware streaming with end-of-utterance detection (`--stream`) | NVIDIA | -| [nemotron-3.5-asr-streaming-0.6b](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b) | RNNT, streaming | 0.6B | multilingual (40+ locales), prompt-conditioned, offline and cache-aware streaming, pick a language with `--lang` (default `auto`). OpenMDW-1.1 | NVIDIA | -| [parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) | TDT | 0.6B | Moondream's post-trained derivative of parakeet-tdt-0.6b-v3, with a VAD head. CC-BY-4.0. Not NeMo-validated, see below | Moondream, from NVIDIA | -| [parakeet-redux](https://huggingface.co/moondream/parakeet-redux) | TDT | 0.6B | Moondream's ternary-encoder derivative of parakeet-tdt-0.6b-v3, with a VAD head. CPU only. CC-BY-4.0. Not NeMo-validated, see below | Moondream, from NVIDIA | - - -### Moondream Ultra and Redux - -[moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and -[moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) are Moondream's -post-trained (Ultra, F16) and ternary-encoder (Redux) derivatives of NVIDIA's -[parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). Both are released under -[CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). They are HF safetensors, converted with -`scripts/convert_hf_parakeet_to_gguf.py`. They are not part of the NeMo-validated set above: there is -no NeMo baseline for them, so parity is transcript-level against our own v3 path (see -[`docs/parity.md`](docs/parity.md)), and the GGUFs are published in -[mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf): `ultra-f16.gguf`, `ultra-q8_0.gguf`, `redux-packed.gguf` -(packed ternary), `redux-f16.gguf` and `redux-q8_0.gguf` (dequantized). Sizes and SHA-256 sums are in -[`models/MANIFEST.md`](models/MANIFEST.md). - -The models were trained by NVIDIA (the base) and Moondream (Ultra and Redux). parakeet.cpp only -converts and quantizes the weights; nothing is trained or fine-tuned here. A dequantized Redux file -(`--ternary dequant`, the converter default) holds ordinary F16 or Q8_0 weights expanded from the -ternary ones. - -- Redux packs the encoder as ternary weights: a 213 MB GGUF, 6.8x smaller than F16. It runs on CPU - only and offline only. On x86 with AVX-512 VNNI it reaches median RTF 75.6 per utterance on - LibriSpeech-100 (8 threads) against 46.1 for the same model in F16; on a single 180 s clip the - gain is about 10 percent; WER on the 100 LibriSpeech utterances is 1.96 percent. See - [`docs/ternary.md`](docs/ternary.md). SIMD kernels exist for x86-64 with AVX2 or AVX-512 VNNI and - aarch64 with dotprod; MSVC builds, Windows on ARM and aarch64 without dotprod use a slow scalar - kernel (about 1 GMAC/s), and the load logs a warning. The packed file also stays resident next to - the repacked planes, so memory use is more than the file size. -- Both carry a voice-activity head, used by `transcribe --vad` to cut long audio at pauses. Speech - is a probability of at least 0.5; pauses of at least 0.2 s are candidate cuts, segments are at most - 30 s, and segments without speech are dropped. On long-form clips it does not change WER - meaningfully. Details and measurements: [`docs/ternary.md`](docs/ternary.md). -- The same head runs on its own, without transcribing: `parakeet-cli vad`, or - `parakeet_capi_vad_pcm_json` / `parakeet_capi_vad_path_json` from the C-API. They return speech - segments (start and end in seconds) as JSON, and optionally the per-frame probabilities. A model - without the head fails with `model has no VAD head`. See [`docs/vad.md`](docs/vad.md). -- Silero VAD (MIT, 32 ms frames, 16 kHz and 8 kHz) runs from its own small GGUF ([download](https://huggingface.co/mudler/parakeet-cpp-gguf)) through the same - functions, as a stream (`parakeet_capi_vad_stream_*`), and as the cutter for any ASR model: - `parakeet-cli transcribe --vad --vad-model silero.gguf`. See [`docs/vad.md`](docs/vad.md). -- VAD-only slices of Ultra and Redux (6 to 10 MB, `redux-vad.gguf` and `ultra-vad-q8_0.gguf` in the same repo) hold just the head and its front end. They run `vad` and the `parakeet_capi_vad_*` calls and cannot transcribe. See [`docs/vad.md`](docs/vad.md). -- The head alone gives false alarms on audio without speech: on speech-free noise it calls about - 99 percent of the frames speech (Ultra 99.4, Redux 97.8), and over a 30 s noise stretch inside a - file with speech the false-alarm frame rate was 17.7 percent for Ultra and 55 percent for Redux, - against 0 percent for Silero (synthetic LibriSpeech with added noise). Prefer Silero as the - always-on gate or when the audio can have long non-speech stretches; use the head on audio known - to be mostly speech, or where its higher recall matters. An offline experiment that lets Silero - decide and the head move the edges (not implemented here) is in - [`docs/vad-benchmarks.md`](docs/vad-benchmarks.md#fusing-silero-and-the-head-offline-experiment). -- Accuracy, speed and size of both detectors, with the method and the scripts to repeat - them: [`docs/vad-benchmarks.md`](docs/vad-benchmarks.md). --- -## Performance - -parakeet.cpp is faster than NeMo's PyTorch runtime on every Parakeet model, on both CPU and GPU, and the transcripts come out byte-identical (WER 0 vs NeMo). Full methodology, all 10 models, quantization tradeoffs, and plots are in [`benchmarks/BENCHMARK.md`](benchmarks/BENCHMARK.md). +## What is new -### See it run +Latest tagged release: **v0.5.0** (2026-08-01). Entries dated after it are on master, not in a release yet: build from source (see [Build](#build)) or use the `:latest` [Docker images](docs/docker.md). -The same clip fed to parakeet.cpp and to NeMo's own PyTorch runtime on the same GPU. The output comes out byte-for-byte identical, parakeet.cpp just gets there first (slowed down so the sub-100ms race is watchable): +- ๐Ÿ” **Speaker fingerprint** (2026-10-04): the speaker registry records which encoder made each voice print and refuses a model mismatch before naming. [docs](docs/diarization.md) +- ๐Ÿ“ฆ **Bundle GGUF** (2026-10-04): ASR, VAD, diarization, sound events and speaker voice models in one file, with one licence per component; three bundles are published. [docs](docs/bundle.md) +- โœ‚๏ธ **VAD-only files** (2026-10-04): 6 to 10 MB slices of the Ultra and Redux VAD head that run `vad` and cannot transcribe. [docs](docs/vad.md) +- ๐Ÿ—ฃ๏ธ **Standalone VAD** (2026-10-04): voice activity detection from the Ultra/Redux head or Silero, with a streaming API, and `transcribe --vad` to cut long audio at pauses. [docs](docs/vad.md) +- ๐Ÿงฎ **Exact batched decode** (2026-10-03): on CPU, batched transducer decode gives the same tokens as decoding each clip alone. [docs](docs/batching.md) +- ๐Ÿงต **Concurrent requests** (2026-10-03): an opt-in pool of CPU backends lets several requests run at once on one loaded model. [docs](docs/concurrency.md) +- ๐Ÿชถ **Moondream Parakeet Ultra and Redux** (2026-10-03): Moondream's derivatives of parakeet-tdt-0.6b-v3, including a ternary Redux encoder with a packed CPU kernel in a 213 MB file. [docs](docs/ultra-redux.md) +- ๐Ÿชช **Speaker identification** (2026-09-30): enroll people from short clips and name the speakers in a diarized scene. [docs](docs/speaker.md) +- ๐Ÿ”” **Sound events and the scene stream** (2026-09-29): CED tags 527 sound classes, and one time-ordered feed carries words, speakers and sounds. [docs](docs/sound.md) +- ๐Ÿ‘ฅ **Diarization** (2026-09-28): Nemotron-3-Diarization answers who spoke when, and speaker-attributed ASR says who said what. [docs](docs/diarization.md) +- ๐ŸŒ **Nemotron 3.5 streaming** (2026-06-06, in v0.5.0): multilingual (40+ locales), prompt-conditioned, offline and cache-aware streaming. [docs](docs/parity.md) +- ๐ŸŽ **Apple Metal** (2026-06-02, in v0.5.0): the encoder runs on Apple GPUs; CUDA and Vulkan are also supported (see [Build](#build)). [docs](benchmarks/BENCHMARK.md#apple-metal-m4) -![parakeet.cpp vs NeMo on GPU: identical output, parakeet.cpp finishes first](benchmarks/media/gpu_duel.gif) - -The same race on CPU, against NeMo's own PyTorch runtime: [parakeet.cpp vs NeMo on CPU](benchmarks/media/cpu_nemo_duel.mp4) (about 1.5x faster, still byte-for-byte identical). And vs whisper.cpp turbo, same accuracy and far less compute: [on GPU](benchmarks/media/gpu_whisper_duel.mp4) (about 12x faster) and [on CPU](benchmarks/media/cpu_duel.mp4) (about 27x faster). - -CPU numbers (20-core x86, vs NeMo PyTorch-CPU, LibriSpeech test-clean, threads=8; RTFx is audio-seconds over processing-seconds, so higher is faster): - -| dtype | size vs f32 | speedup vs NeMo | accuracy | -| ----- | ----------- | --------------- | -------- | -| f32 | 100% | 1.11 to 1.69x (median 1.40x) | WER 0, byte-identical to NeMo | -| f16 | 57% | up to 1.70x | near-lossless | -| q8_0 | 37% | up to 1.86x | near-lossless | -| q4_k | 26% | n/a | small, monotonic WER cost | - -Peak RAM is also roughly 2x lower than NeMo, and lower still once quantized. - -GPU numbers (NVIDIA GB10, Grace-Blackwell, vs NeMo-GPU in the `nvcr.io/nvidia/nemo` container, since NeMo can't run on the host's torch/CUDA stack directly): parakeet.cpp wins on all 10 models, with a median of 1.25x and up to 4.3x on the large TDT/hybrid models. NeMo's TDT greedy decode isn't CUDA-graph accelerated and ours is a lean C++ loop, which is most of that gap. The log-mel front end runs on the GPU via a ggml DFT-matmul graph; the CPU path is unchanged. +The models are in [mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf), and [LocalAI](https://localai.io) runs parakeet.cpp as its `parakeet-cpp` backend. The other docs are listed in the [Documentation](#documentation) table. --- -## Pre-built binaries +## Demos -Every [release](https://github.com/mudler/parakeet.cpp/releases) ships pre-built `parakeet-cli` bundles, so there is no need to compile from source: +Each clip is a short animated preview. The full video is linked below it. The clips below the first one were made for posts on X; the post links are not collected in this repository yet. -| Platform | Variants | -| -------- | -------- | -| Linux x64 | cpu, vulkan, cuda | -| Linux arm64 | cpu | -| macOS arm64 | metal | -| macOS x64 | cpu | -| Windows x64 | cpu, vulkan, cuda | + + + + + + + + + +
-The cuda bundles target Turing (sm_75) and newer, including Blackwell. On Linux the CUDA runtime libraries are bundled in the tarball; on Windows download the `cudart-parakeet-bin-win-cuda-x64.zip` asset alongside the binary zip unless you already have the CUDA toolkit installed. The vulkan binaries need the Vulkan loader on the system (`libvulkan1` on Debian/Ubuntu; on Windows the GPU driver provides it). +**Words, speakers and sounds in one pass** -## Build +[![Scene stream on a film scene: transcript, speakers and sound tags](benchmarks/media/scene_sprite_fright.gif)](benchmarks/media/scene_sprite_fright.mp4) -Clone with submodules (ggml is vendored at `third_party/ggml`): +Excerpt of the scene demo: ASR, diarization and CED sound tags together, on CPU ([MP4 excerpt](benchmarks/media/scene_sprite_fright.mp4)). Film: Sprite Fright, CC BY 4.0, Blender Studio. -```sh -git clone --recursive https://github.com/mudler/parakeet.cpp -cd parakeet.cpp -cmake -B build -DPARAKEET_BUILD_TESTS=ON && cmake --build build -j -``` +https://github.com/user-attachments/assets/02c29d27-ce26-46f2-8677-661c59686323 -Use `-DGGML_NATIVE=OFF` for portable or CI builds (it disables host-specific ISA extensions). For the shared library (LocalAI / dlopen): + -```sh -cmake -B build-shared -DPARAKEET_SHARED=ON -DPARAKEET_BUILD_CLI=ON -cmake --build build-shared -j -# -> build-shared/libparakeet.so -``` +**Batched decode: one loop, 16 clips** -### CMake options +[![Batched against one-at-a-time decode on a GPU](benchmarks/media/batch_decode_race.gif)](benchmarks/media/batch_decode_race.mp4) -| Option | Default | Purpose | -| ------------------------ | ------- | ------------------------------------------ | -| `PARAKEET_BUILD_TESTS` | OFF | Compile and register ctest targets | -| `PARAKEET_BUILD_CLI` | ON | Build `parakeet-cli` | -| `PARAKEET_SHARED` | OFF | Build libparakeet as a shared library | -| `PARAKEET_VERSION` | 0.0.1 | Version string returned by `--version` | -| `PARAKEET_GGML_CUDA` | OFF | Forward GGML_CUDA to the submodule | -| `PARAKEET_GGML_METAL` | OFF | Forward GGML_METAL to the submodule | -| `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule | -| `PARAKEET_GGML_HIP` | OFF | Forward GGML_HIP (ROCm) to the submodule | -| `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp | -| `PARAKEET_WITH_VOICEDETECT` | ON | Speaker identification through voice-detect.cpp | +Serial against batched decode of 16 clips, with identical output ([MP4](benchmarks/media/batch_decode_race.mp4)). Details: [batching.md](docs/batching.md). -To build for a GPU backend, forward its flag, e.g. Apple Metal: -```sh -cmake -B build -DPARAKEET_GGML_METAL=ON && cmake --build build -j -``` +https://github.com/user-attachments/assets/9a541488-b03f-4c68-8d59-2a6a9e631cc1 -The CLI auto-selects the first GPU device the ggml registry reports (including integrated GPUs such as Ryzen APUs), so no runtime flag is needed. Use `PARAKEET_DEVICE` to override: set it to `cpu` to force CPU, or to a specific device name like `CUDA0` or `Vulkan1` (case-insensitive) to pick that device. Ops the chosen backend has no kernel for run on the CPU automatically, so a model always runs even when one op lacks a GPU kernel. On an Apple M4, Metal is up to about 5x faster than CPU on the larger models; see [Apple Metal](benchmarks/BENCHMARK.md#apple-metal-m4). ---- -## Docker +
-Two prebuilt images are published to GitHub Container Registry on every push to `master`, one per binary: +**Nemotron 3.5 streaming against NeMo, on CPU** -- `ghcr.io/mudler/parakeet.cpp-cli`: the command-line transcriber. -- `ghcr.io/mudler/parakeet.cpp-server`: the [OpenAI-compatible server](#openai-compatible-server). +[![parakeet.cpp q8_0 against NeMo on the same CPU](benchmarks/media/nemotron_streaming_race.gif)](benchmarks/media/nemotron_streaming_race.mp4) -Each comes in a CPU and a CUDA variant (the CUDA tag is suffixed `-cuda`), and both are multi-arch (`linux/amd64` and `linux/arm64`), so the right one is pulled for your host automatically. They contain just the binary, so mount a converted `.gguf` model (and, for the cli, your audio) at runtime: +Same model, same CPU, identical output ([MP4](benchmarks/media/nemotron_streaming_race.mp4)). -```sh -# CLI, CPU -docker run --rm \ - -v "$PWD/models:/models:ro" \ - -v "$PWD/audio:/audio:ro" \ - ghcr.io/mudler/parakeet.cpp-cli:latest \ - transcribe --model /models/parakeet-tdt_ctc-110m-q5_k.gguf --input /audio/speech.wav --decoder tdt - -# CLI, CUDA (needs the nvidia container toolkit on the host) -docker run --rm --gpus all \ - -v "$PWD/models:/models:ro" -v "$PWD/audio:/audio:ro" \ - ghcr.io/mudler/parakeet.cpp-cli:latest-cuda \ - transcribe --model /models/parakeet-tdt_ctc-110m-q5_k.gguf --input /audio/speech.wav --decoder tdt - -# Server: binds 0.0.0.0 and exposes 8080. Fetch a model by alias on first run, -# or mount a local .gguf. Add --gpus all with the :latest-cuda tag for GPU. -docker run --rm -p 8080:8080 ghcr.io/mudler/parakeet.cpp-server:latest --model tdt_ctc-110m -``` -The CUDA image is built on CUDA 13, so it covers everything from Turing up through Blackwell, including GB10 / Grace-Blackwell (DGX Spark) on arm64. +https://github.com/user-attachments/assets/3811082d-5c5e-42d5-bbd8-fd6412c8775c -To build the images yourself, see the build args at the top of the [`Dockerfile`](Dockerfile); the cli is the default target and the server is `--target runtime-server`. The CPU image is the portable `GGML_NATIVE=OFF` build, so it runs on any amd64 or arm64 host. ---- + -## Python environment setup +**More races** -You need this once, for model conversion and validation. It's not needed for inference: +- [parakeet.cpp against NeMo on CPU](benchmarks/media/cpu_nemo_duel.mp4) (about 1.5x faster, same output) +- [against whisper.cpp turbo on GPU](benchmarks/media/gpu_whisper_duel.mp4) (about 12x faster) +- [against whisper.cpp turbo on CPU](benchmarks/media/cpu_duel.mp4) (about 27x faster) -```sh -python3 -m venv .venv -.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu -.venv/bin/pip install -r scripts/requirements.txt # nemo_toolkit[asr] + gguf -``` +These are single-clip demos, not benchmarks. See [Benchmarks](#benchmarks). -NeMo 2.7.3 is the validated version. The anchor checkpoint `nvidia/parakeet-tdt_ctc-110m` (about 440 MB) is downloaded automatically by NeMo on first use. +
--- -## Converting a model +## Quick start -Convert a HuggingFace or local `.nemo` checkpoint to GGUF: +Get the CLI from a [release](https://github.com/mudler/parakeet.cpp/releases) (see [Build](#build) to compile it yourself; the VAD and bundle commands below need a build from `master`). Then download a model. Everything is in [mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf): ```sh -# Default (F32), lossless and largest -.venv/bin/python scripts/convert_parakeet_to_gguf.py \ - --model nvidia/parakeet-tdt_ctc-110m \ - --output m.gguf - -# F16, about 0.58x the size, WER 0 vs NeMo -.venv/bin/python scripts/convert_parakeet_to_gguf.py \ - --model nvidia/parakeet-tdt_ctc-110m --dtype f16 --output m.gguf - -# Q8_0, about 0.39x the size, WER 0 vs NeMo -.venv/bin/python scripts/convert_parakeet_to_gguf.py \ - --model nvidia/parakeet-tdt_ctc-110m --dtype q8_0 --output m.gguf +HF=https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main +curl -LO $HF/tdt_ctc-110m-q8_0.gguf # 178 MB, English, hybrid TDT+CTC ``` -Supported `--dtype`: `f32` (default), `f16`, `q8_0`. - ---- - -## Quantization - -The Python `gguf` writer can't produce K-quants (`q4_k`, `q5_k`, `q6_k`), so re-quantize an existing F32 GGUF with the CLI instead: +**Transcribe.** `tests/fixtures/speech.wav` is in this repository: ```sh -parakeet-cli quantize -# e.g. -parakeet-cli quantize m.gguf m_q4k.gguf q4_k -parakeet-cli quantize m.gguf m_q6k.gguf q6_k +parakeet-cli transcribe --model tdt_ctc-110m-q8_0.gguf --input tests/fixtures/speech.wav +# Well, I don't wish to see it any more, observed Phoebe, turning away her eyes. It is certainly very like the old portrait. + +parakeet-cli transcribe --model tdt_ctc-110m-q8_0.gguf --input tests/fixtures/speech.wav --timestamps +# 0.48-0.64 Well, (0.79) +# 0.80-0.88 I (1.00) +# ... ``` -Supported types: `q4_0`, `q5_0`, `q8_0`, `q4_k`, `q5_k`, `q6_k`. +**Find speech with VAD.** Silero is a 1.3 MB file: -Only the large linear `ggml_mul_mat`-consumed weights (encoder FFN, attention projections, joint enc/pred projections, subsampling output projection) get quantized. The conv, LSTM, featurizer, batch_norm, and bias tensors stay F32. See `docs/quantization.md` for the full policy, allowlist, and measured size and WER per type. +```sh +curl -LO $HF/silero-vad-f16.gguf +parakeet-cli vad --model silero-vad-f16.gguf --input tests/fixtures/two_speakers.wav +# {"mode":"speech","duration":23.605,"frame_sec":0.032,"backend":"cpu","segments":[{"start":0.514,"end":5.534}, ...]} ---- +# Cut long audio at pauses with Silero, then transcribe each piece with any ASR model +parakeet-cli transcribe --model tdt_ctc-110m-q8_0.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf +``` -## Running inference +**Use a bundle.** One 338 MB file holds the 110M ASR model, Silero, diarization, CED sound tagging and a speaker encoder: ```sh -# Default decoder (TDT for hybrid/TDT models, CTC for standalone CTC) -parakeet-cli transcribe --model m.gguf --input audio.wav - -# Force a decoder -parakeet-cli transcribe --model m.gguf --input audio.wav --decoder ctc -parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt - -# Per-word timestamps + confidence: one line per word -# - () (times in seconds) -parakeet-cli transcribe --model m.gguf --input audio.wav --timestamps - -# JSON with the flat text plus per-word and per-token timestamps + confidence: -# {"text":"...","words":[{"w":..,"start":..,"end":..,"conf":..}], -# "tokens":[{"id":..,"t":..,"conf":..}]} -parakeet-cli transcribe --model m.gguf --input audio.wav --json - -# Offline TDT N-best hypotheses as ranked JSON -parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt \ - --beam-size 4 --nbest 4 - -# Read WAV bytes from stdin (useful with ffmpeg/curl pipelines) -ffmpeg -i input.mp3 -f wav - | parakeet-cli transcribe --model m.gguf --input - - -# Long audio on Ultra/Redux: cut at VAD pauses, transcribe each piece (offline only). -# Tune with --vad-threshold F (0.5), --vad-min-pause SEC (0.2), --vad-max-seg SEC (30) -parakeet-cli transcribe --model ultra.gguf --input long.wav --vad - -# Voice activity detection only, no transcript: speech regions as JSON -# {"mode":"speech","duration":..,"frame_sec":0.08,"backend":"cpu", -# "segments":[{"start":..,"end":..}]} (seconds; models with a VAD head only) -# --mode segments gives the cuts that `transcribe --vad` uses; --probabilities adds p per 80 ms frame. -# Tune with --threshold F (0.5), --min-pause SEC (0.2), --min-speech SEC (0.1), --max-segment SEC (30) -parakeet-cli vad --model ultra.gguf --input audio.wav - -# Only the VAD head is needed? Use a 6 to 10 MB slice instead of the full model -# (it cannot transcribe). Files and checksums: docs/vad.md. -curl -LO https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/redux-vad.gguf -parakeet-cli vad --model redux-vad.gguf --input audio.wav - -# The same with a Silero VAD GGUF (frame_sec 0.032; defaults 250 ms min speech, -# 100 ms min pause, 30 ms pad). Any ASR model can then cut long audio with it. -# Download it from the collection repo (F16 is 1.3 MB, F32 is 2.2 MB). To make the -# file yourself, see scripts/convert_silero_vad_to_gguf.py, docs/vad.md and docs/conversion.md. -curl -LO https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/silero-vad-f16.gguf -parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav -parakeet-cli transcribe --model tdt-0.6b-v3.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf - -# Print model metadata (arch, dims, mel params, vocab size, TDT durations) -parakeet-cli info m.gguf - -# Cache-aware streaming (EOU model parakeet_realtime_eou_120m-v1): feeds the WAV -# in the model's chunk schedule, prints partial text incrementally and -# [EOU @ s] / [EOB @ s] event markers, then the finalized tail. Add -# --timestamps to also print per-word [start-end] (conf) lines as words finalize. -parakeet-cli transcribe --model eou.gguf --input audio.wav --stream +curl -LO $HF/parakeet-bundle-small.gguf +parakeet-cli info parakeet-bundle-small.gguf # components, licences, sizes +parakeet-cli transcribe --model parakeet-bundle-small.gguf --input tests/fixtures/speech.wav --vad +parakeet-cli scene --model parakeet-bundle-small.gguf --diar parakeet-bundle-small.gguf \ + --sound parakeet-bundle-small.gguf --input tests/fixtures/two_speakers.wav +# [00:00.4 - 00:05.4] Speaker 0: mister Quilter is the apostle of the middle classes, and we are glad to welcome his gospel. +# [00:06.8 - 00:10.8] Speaker 1: Well, I don't wish to see it any more, observed Phoebe, turning away her eyes. +# ... ``` -Timestamps and confidence match NeMo's `transcribe(timestamps=True)` with the `max_prob` confidence method exactly (word offsets to 0.0 s, per-token and per-word confidence within `5e-6`), for both the TDT and CTC heads. See `docs/parity.md`. Word start and end are in seconds (`frame x hop x subsampling / sample_rate`, which works out to 0.08 s/frame here); confidence is the rescaled softmax probability of the emitted token, aggregated per word with NeMo's `min`. +The published bundles are `parakeet-bundle-small` (110M ASR, 338 MB), `parakeet-bundle-standard` (0.6B v3 ASR, 1.1 GB) and `parakeet-bundle-moondream-redux` (packed Redux plus Silero, 215 MB). Each has a `NOTICE-*.txt` file with the credits and licences. See [bundle.md](docs/bundle.md). -The optional TDT beam decoder follows NeMo's default sequence-level beam -search and exposes raw/normalized scores plus token frame/duration metadata. -See [`docs/tdt-nbest.md`](docs/tdt-nbest.md). - -The `parakeet-cli` binary lands at `build/examples/cli/parakeet-cli`. +To use it from a program, see [C API](#c-api). To serve it over HTTP, see [Server and Docker](#server-and-docker). --- -## OpenAI-compatible server - -`parakeet-server` is a small HTTP server that speaks the OpenAI transcription -API, so any OpenAI client works by pointing its `base_url` at it. It is built by -default (`PARAKEET_BUILD_SERVER=ON`) and lands at `build/examples/server/parakeet-server`. - -```sh -# Serve a model. --model takes a local .gguf, an http(s) URL, a .gguf in -# mudler/parakeet-cpp-gguf, or an alias (downloaded and cached on first run). -parakeet-server --model tdt_ctc-110m --port 8080 +## Models -# Transcribe over HTTP -curl -F file=@audio.wav -F response_format=verbose_json \ - http://localhost:8080/v1/audio/transcriptions -``` +All models below are published as GGUF in [mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf) (f16, q8_0, q6_k, q5_k and q4_k for the ASR models). The NVIDIA models are validated at WER 0 against NeMo. Per-model parity: [parity.md](docs/parity.md). Licences per model: [licenses.md](docs/licenses.md). -```python -from openai import OpenAI -client = OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed") -with open("audio.wav", "rb") as f: - print(client.audio.transcriptions.create(model="parakeet", file=f).text) -``` +| Model | Type | Size | Notes | +| ----- | ---- | ---- | ----- | +| [parakeet-tdt_ctc-110m](https://huggingface.co/nvidia/parakeet-tdt_ctc-110m) | hybrid TDT+CTC | 110M | English, the small anchor checkpoint | +| [parakeet-ctc-0.6b](https://huggingface.co/nvidia/parakeet-ctc-0.6b), [parakeet-ctc-1.1b](https://huggingface.co/nvidia/parakeet-ctc-1.1b) | CTC | 0.6B, 1.1B | English | +| [parakeet-rnnt-0.6b](https://huggingface.co/nvidia/parakeet-rnnt-0.6b), [parakeet-rnnt-1.1b](https://huggingface.co/nvidia/parakeet-rnnt-1.1b) | RNNT | 0.6B, 1.1B | English | +| [parakeet-tdt-0.6b-v2](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2), [parakeet-tdt-1.1b](https://huggingface.co/nvidia/parakeet-tdt-1.1b) | TDT | 0.6B, 1.1B | English | +| [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) | TDT | 0.6B | 25 European languages | +| [parakeet-tdt_ctc-1.1b](https://huggingface.co/nvidia/parakeet-tdt_ctc-1.1b) | hybrid TDT+CTC | 1.1B | English | +| [parakeet_realtime_eou_120m-v1](https://huggingface.co/nvidia/parakeet_realtime_eou_120m-v1) | RNNT, streaming | 120M | Cache-aware streaming with end-of-utterance events (`--stream`) | +| [nemotron-3.5-asr-streaming-0.6b](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b) | RNNT, streaming | 0.6B | 40+ locales, language set with `--lang` (default `auto`), offline and streaming. OpenMDW-1.1 | +| [Nemotron-3-Diarization](https://huggingface.co/nvidia/Nemotron-3-Diarization) | Sortformer | n/a | Who spoke when, up to 8 speakers. See [diarization.md](docs/diarization.md) | +| [parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) | TDT + VAD head | 0.6B | Moondream, from parakeet-tdt-0.6b-v3. CC-BY-4.0. Not NeMo-validated | +| [parakeet-redux](https://huggingface.co/moondream/parakeet-redux) | TDT + VAD head | 0.6B | Moondream, ternary encoder. CPU only, offline only. CC-BY-4.0. Not NeMo-validated | +| [Silero VAD](https://github.com/snakers4/silero-vad) | VAD | 1.3 MB | MIT. 16 kHz and 8 kHz | +| [CED](https://huggingface.co/mudler/ced-gguf) | sound events | 6 to 88 MB | 527 classes, from [ced.cpp](https://github.com/localai-org/ced.cpp). See [sound.md](docs/sound.md) | -It supports `response_format` `json` / `text` / `verbose_json` and -`timestamp_granularities[]=word`. This is a single-model, one-request-at-a-time -example that accepts WAV uploads only; see [`examples/server/README.md`](examples/server/README.md) -for the full list of options and known simplifications. **For a production -deployment, use [LocalAI](https://localai.io)**, which embeds parakeet.cpp as a -backend and adds a model gallery, concurrency, multi-model serving, the full -OpenAI API surface, auth, and metrics. +Convert your own checkpoint with `scripts/convert_parakeet_to_gguf.py` (see [conversion.md](docs/conversion.md)) and quantize with `parakeet-cli quantize` (see [quantization.md](docs/quantization.md)). --- -## Batching - -Single-clip transcription is the default and needs no flags: every `transcribe` call runs one clip at a time, byte-for-byte identical to before. Batching is an opt-in path for decoding several clips together, which matters when you serve many concurrent requests on a GPU. - -The win is on the **decode** side. A transducer (TDT/RNN-T) decodes autoregressively with tiny per-step prediction-LSTM and joint GEMMs; one clip launches hundreds of these matvec-sized kernels and leaves the GPU mostly idle between launches. Decoding N clips together coalesces each step into one batched GEMM, so the device stays busy. On the NVIDIA GB10 this reaches about **10-12x** at batch size 16 (CPU about 3-5x); the encoder is already compute-bound, so batching it gives no throughput win. CTC has no autoregressive decode, so batching does not apply to standalone CTC models. On CPU the batched decode step is bit-identical to decoding each clip alone: its matmuls call the same dot-product kernel as a single column, so logits, token ids, frames and confidences are equal (`tests/test_exact_batch.cpp` checks this for F32, F16 and Q8_0 decoder weights). On GPU backends the batched matmul is the ordinary ggml kernel, so logits agree to about 1e-4 rather than exactly. The batched encoder is separate: its output is close to the single-clip encoder output, not equal, so a whole batched transcript can still differ from a single-clip one, and the tests compare with a tolerance there. Full numbers and per-model tables are in [`benchmarks/BENCHMARK.md`](benchmarks/BENCHMARK.md#batched-decode-throughput). +## CLI cheat sheet -Measure it yourself: +The binary is `build/examples/cli/parakeet-cli`. The full list of examples and options is in [cli.md](docs/cli.md). -```bash -# Decode-only: serial vs batched decode of one clip replicated B times (the win in isolation). -parakeet-cli bench-decode --model --audio [--batch-sizes 1,4,8,16] [--threads N] [--reps R] [--json ] - -# Full transcribe (encoder + decode) over a manifest at several batch sizes. -parakeet-cli bench-batch --model --manifest [--decoder ctc|tdt] [--threads N] [--batch-sizes 1,4,8] [--json ] +```sh +parakeet-cli info [--component NAME] # metadata; for a bundle, the component list +parakeet-cli transcribe --model M --input A.wav # default decoder; "--input -" reads WAV from stdin +parakeet-cli transcribe ... --decoder ctc|tdt # force a decoder +parakeet-cli transcribe ... --timestamps | --json # per-word times and confidence +parakeet-cli transcribe ... --beam-size 4 --nbest 4 # TDT N-best, see docs/tdt-nbest.md +parakeet-cli transcribe ... --stream # cache-aware streaming (EOU and Nemotron models) +parakeet-cli transcribe ... --lang # Nemotron 3.5 language, default auto +parakeet-cli transcribe ... --vad [--vad-model silero.gguf] # cut long audio at pauses (offline, greedy only) +parakeet-cli vad --model M --input A.wav [--mode segments] [--probabilities] # speech regions as JSON +parakeet-cli scene --model ASR --diar DIAR --sound CED --input A.wav # words + speakers + sounds +parakeet-cli scene ... --speakers SPK.gguf --registry R # name the speakers +parakeet-cli enroll --model SPK.gguf --name Ada --input ada.wav --registry R +parakeet-cli quantize +parakeet-cli bench --model M --manifest F [--concurrency K] # throughput +parakeet-cli bench-decode ... | bench-batch ... # batched decode, see docs/batching.md ``` -To batch from code, use the batched entry points (single-clip B=1 is just N=1): - -- C++ (`src/model.hpp`): `Model::transcribe_16k_batch(pcms16k, decoder)` and `transcribe_16k_batch_with_timestamps(...)` take N clips of 16 kHz mono float PCM and return N results. -- C-API (`include/parakeet_capi.h`): `parakeet_capi_transcribe_pcm_batch(...)` (N transcripts) and `parakeet_capi_transcribe_pcm_batch_json(...)` (one JSON array of N `{text,words,tokens}` objects). These are what LocalAI's `parakeet-cpp` backend calls to coalesce concurrent requests; it leaves batching off by default and exposes a `batch_max_size` option to opt in. +Device selection is automatic: the CLI uses the first GPU the ggml registry reports. Set `PARAKEET_DEVICE=cpu` to force CPU, or a device name such as `CUDA0` or `Vulkan1`. Ops that a backend lacks run on the CPU. --- -## Concurrent requests +## Server and Docker -One loaded model runs one request at a time by default. To serve several requests in parallel, give the model a pool of CPU backends: `parakeet_capi_set_concurrency(ctx, backends, threads_each)`, `pk::Model::set_concurrency`, or `--concurrency K` on `parakeet-server` and `parakeet-cli bench`. Results are identical to the single-backend run, aggregate throughput can go up or down depending on model size and core count (about 1.2x to 1.3x on a 110M model with 8 cores, but slower than one backend on 0.6B models with 8 threads; all measured on loaded machines, see [`docs/concurrency.md`](docs/concurrency.md)), and each request gets somewhat slower because it has fewer threads. Try it only when `backends x threads_each` fits the physical cores, measure before you enable it, and expect extra memory per backend (`Model::pool_working_set_bytes()`). The default is one backend and behaves as before. - ---- - -## Sound events - -parakeet.cpp can also tag everyday sounds (dog bark, glass breaking, applause, -alarms, music, and the rest of the 527-class AudioSet ontology) with -[CED](https://github.com/RicherMans/CED), through the -[ced.cpp](https://github.com/localai-org/ced.cpp) submodule (`PARAKEET_WITH_CED`, -on by default). `parakeet-cli scene` combines it with ASR and diarization into -one time-ordered feed: +`parakeet-server` is a small OpenAI-compatible HTTP server (`POST /v1/audio/transcriptions`). It serves one model, one request at a time, WAV uploads only, so treat it as an example. For production use [LocalAI](https://localai.io), which embeds parakeet.cpp as a backend and adds a gallery, concurrency, multi-model serving, auth and metrics. See [examples/server/README.md](examples/server/README.md). ```sh -parakeet-cli scene --model asr.gguf --diar diar.gguf --sound ced-base-q8_0.gguf \ - --latency low --input audio.wav -[00:00.4 - 00:03.2] Speaker 0: mister Quilter is the apostle of the middle classes, and -[00:24.0 - 00:30.0] (Chicken, rooster 0.86) +parakeet-server --model tdt_ctc-110m --port 8080 +curl -F file=@audio.wav -F response_format=verbose_json http://localhost:8080/v1/audio/transcriptions ``` -See [`docs/sound.md`](docs/sound.md) for the CED GGUFs, the sound and scene -stream C-API (ABI v8), and the `--sound-model` server option. - -### Naming speakers - -With a voice-detect.cpp speaker encoder (`PARAKEET_WITH_VOICEDETECT`, on by -default) the scene stream can say who is talking instead of `Speaker 0`. -Enroll each person from a short clip with `parakeet-cli enroll`, then pass -`--speakers --registry ` to `scene`. Only one two-voice -fixture has been measured so far. See [`docs/speaker.md`](docs/speaker.md) for -the models, the commands, the C-API (ABI v9 and v10) and what is still untested. +Images for the CLI and the server are published to GHCR on every push to `master` (CPU and CUDA, `linux/amd64` and `linux/arm64`): `ghcr.io/mudler/parakeet.cpp-cli` and `ghcr.io/mudler/parakeet.cpp-server`. See [docker.md](docs/docker.md). --- -## C-API (`libparakeet.so`) +## C API -`include/parakeet_capi.h` defines a flat, exception-free C-API meant for `dlopen` / FFI / LocalAI integration. Build the shared library with `-DPARAKEET_SHARED=ON`: +`include/parakeet_capi.h` is a flat, exception-free C API for `dlopen`, FFI and LocalAI. Build the shared library with `-DPARAKEET_SHARED=ON` (see [Build](#build)). ```c #include "parakeet_capi.h" -parakeet_ctx *ctx = parakeet_capi_load("model.gguf"); // load ONCE +parakeet_ctx *ctx = parakeet_capi_load("model.gguf"); // load once, reuse if (!ctx) { fprintf(stderr, "%s\n", parakeet_capi_last_error(ctx)); return 1; } -char *text = parakeet_capi_transcribe_path(ctx, "audio.wav", 0 /*default*/); +char *text = parakeet_capi_transcribe_path(ctx, "audio.wav", 0 /*default decoder*/); if (text) { printf("%s\n", text); parakeet_capi_free_string(text); } - parakeet_capi_free(ctx); ``` -In-memory PCM: -```c -char *text = parakeet_capi_transcribe_pcm(ctx, samples, n_samples, - sample_rate, 0 /*default*/); -``` - -Timestamps and confidence as JSON (matches NeMo `timestamps=True` + `max_prob`): -```c -char *json = parakeet_capi_transcribe_path_json(ctx, "audio.wav", 0 /*default*/); -// {"text":"...", -// "frame_sec":0.080000, -// "words":[{"w":"Well,","start":0.480,"end":0.640,"conf":0.7859}, ...], -// "tokens":[{"id":639,"t":0.480,"conf":0.9969}, ...]} -if (json) { printf("%s\n", json); parakeet_capi_free_string(json); } -``` -`start`/`end`/`t` are in seconds; `conf` is the rescaled softmax probability of the emitted token in `(0,1]` (a word's `conf` is the `min` over its tokens). `frame_sec` is the encoder frame stride in seconds (`hop x subsampling / sample_rate`); multiply a frame-unit segment gap threshold (NeMo's `segment_gap_threshold`) by it to get the seconds gap between words when forming segments. +- **Surface:** offline and streaming transcription (text or JSON with word and token timestamps), batched transcription, VAD (offline and streaming), diarization, speaker-attributed ASR, sound events, the scene stream, speaker identification, bundle loading and the concurrency pool. +- **ABI:** the current version is **10** (`parakeet_capi_abi_version()`). Later additions (VAD, bundle, encoder fingerprint, concurrency) are additive and keep ABI 10. LocalAI depends on the offline and streaming transcription symbols, so do not change their signatures without a coordinated bump. +- More examples and the JSON shapes: [capi.md](docs/capi.md). Exact signatures: `include/parakeet_capi.h`. -### Streaming (cache-aware EOU model) +--- -For `parakeet_realtime_eou_120m-v1`, a streaming session decodes 16 kHz mono f32 PCM as it arrives, returning newly-finalized text and signalling EOU/EOB events: +## Build -```c -parakeet_stream *s = parakeet_capi_stream_begin(ctx); -int eou = 0; -char *t = parakeet_capi_stream_feed(s, pcm, n_samples, &eou); // "" if none yet -if (t) { printf("%s", t); parakeet_capi_free_string(t); } -if (eou) printf(" [EOU]"); -// ...feed more chunks... -char *tail = parakeet_capi_stream_finalize(s); // flush the tail -if (tail) { printf("%s\n", tail); parakeet_capi_free_string(tail); } -parakeet_capi_stream_free(s); +```sh +git clone --recursive https://github.com/mudler/parakeet.cpp +cd parakeet.cpp +cmake -B build && cmake --build build -j +# CLI: build/examples/cli/parakeet-cli ``` -`` (end-of-utterance) and `` (backchannel) are stripped from the text and surfaced via `*eou_out` (the CLI `--stream` prints them as `[EOU @ s]` markers). The streaming transcript matches NeMo's cache-aware streaming exactly, and `finalize` flushes the end-of-stream tail without fabricating an `` that NeMo would not emit. - -The LocalAI backend (in the LocalAI repo) dlopens `libparakeet.so` and uses these symbols directly: the offline `parakeet_capi_transcribe_*` / `parakeet_capi_transcribe_path_json` and the streaming `parakeet_capi_stream_*`. See `include/parakeet_capi.h` for the full API. The C++ streaming session (`pk::StreamingSession`) also exposes per-word timestamps and confidence as words finalize, via `drain_words()` alongside the EOU events, which the CLI `--stream --timestamps` path prints. - ---- - -## Model coverage +Use `-DGGML_NATIVE=OFF` for a portable binary. For the shared library: `cmake -B build-shared -DPARAKEET_SHARED=ON && cmake --build build-shared -j`, which produces `libparakeet.so`. -See `docs/parity.md` for the full coverage matrix. In short: +| GPU backend | CMake flag | Status | +| ----------- | ---------- | ------ | +| CUDA | `-DPARAKEET_GGML_CUDA=ON` | Release binaries for Turing (sm_75) and newer. Benchmarked on a GB10. | +| Metal | `-DPARAKEET_GGML_METAL=ON` | Release binary for macOS arm64. Benchmarked on an M4. | +| Vulkan | `-DPARAKEET_GGML_VULKAN=ON` | Release binaries for Linux and Windows. Needs the Vulkan loader. | +| ROCm (HIP) | `-DPARAKEET_GGML_HIP=ON` | Forwarded to ggml. Not tested by us. | -| Family | Representative checkpoints | Heads | WER vs NeMo | -| --- | --- | --- | --- | -| Hybrid TDT+CTC | `parakeet-tdt_ctc-110m`, `parakeet-tdt_ctc-1.1b` | TDT + CTC | 0.0 | -| TDT (hybrid) | `parakeet-tdt-0.6b-v2`, `parakeet-tdt-0.6b-v3` (multilingual) | TDT | 0.0 | -| Pure TDT | `parakeet-tdt-1.1b` | TDT | 0.0 | -| CTC | `parakeet-ctc-0.6b`, `parakeet-ctc-1.1b` | CTC | 0.0 | -| RNNT | `parakeet-rnnt-0.6b`, `parakeet-rnnt-1.1b` | RNNT | 0.0 | +GPU backends are not exercised in CI. Other options: -All 10 published offline checkpoints are validated at WER 0 vs NeMo 2.7.3. Sizes: 110M (512/17 layers), 0.6B (1024/24), 1.1B (1024/42). +| Option | Default | Purpose | +| ------ | ------- | ------- | +| `PARAKEET_BUILD_TESTS` | OFF | ctest targets | +| `PARAKEET_BUILD_CLI` | ON | `parakeet-cli` | +| `PARAKEET_BUILD_SERVER` | ON | `parakeet-server` | +| `PARAKEET_SHARED` | OFF | `libparakeet` as a shared library | +| `PARAKEET_WITH_CED` | ON | Sound-event tagging (ced.cpp) | +| `PARAKEET_WITH_VOICEDETECT` | ON | Speaker identification (voice-detect.cpp) | -Cache-aware streaming and EOU (`parakeet_realtime_eou_120m-v1`) is implemented too: `layer_norm` plus causal conv, causal subsampling, chunked-limited attention, per-layer conv/attention caches, carried RNN-T decoder state, and ``/`` events. The streaming transcript matches NeMo's cache-aware streaming byte for byte. See `docs/parity.md` (the Streaming + EOU section). +**Pre-built binaries.** Each [release](https://github.com/mudler/parakeet.cpp/releases) ships `parakeet-cli` bundles: Linux x64 (cpu, vulkan, cuda), Linux arm64 (cpu, vulkan), macOS arm64 (metal), macOS x64 (cpu), Windows x64 (cpu, vulkan, cuda), plus AppImages and library tarballs. On Windows with CUDA, also download `cudart-parakeet-bin-win-cuda-x64.zip` unless the CUDA toolkit is installed. The newest release (v0.5.0) does not include the features marked **master** above. --- -## Running tests +## Benchmarks -Model-independent (run anywhere): +Speed is audio seconds over processing seconds (RTFx), against NeMo's PyTorch runtime on the same machine, batch size 1. Higher is faster. Numbers come from [benchmarks/BENCHMARK.md](benchmarks/BENCHMARK.md): CPU runs use 8 threads on a 20-core x86 host, GPU runs use one NVIDIA GB10. These were shared development machines, not quiet lab hosts, so read the method before you quote a number. -```sh -ctest --test-dir build --output-on-failure -LE model -``` +| Measurement | Result | Source | +| ----------- | ------ | ------ | +| CPU, f32, 10 models, LibriSpeech test-clean | 1.11x to 1.69x faster than NeMo (mean 1.40x), same transcripts | BENCHMARK.md, Headline | +| CPU, q8_0 | mean 1.56x, up to 1.89x, 37% of the f32 size | BENCHMARK.md, Quantization | +| GPU (GB10), f32 | median 1.25x, up to 4.3x (`tdt_ctc-110m`) | [performance.md](docs/performance.md) | +| Nemotron 3.5, one 7.4 s clip, CPU | 2.40x at f32, 2.52x at q8_0 | BENCHMARK.md, Nemotron | +| Batched decode, batch 16 | about 10x to 12x on the GB10 (f16), about 3x to 5x on CPU (q5_k) | BENCHMARK.md, Batched decode | +| Apple M4, Metal against CPU | 1.3x to 5.6x, most on the 0.6B and 1.1B models | BENCHMARK.md, Apple Metal | +| Redux (packed ternary) against the same model in F16, CPU | median RTF 75.6 against 46.1 per utterance (8 threads, AVX-512 VNNI); 6.8x smaller file; WER 1.96% on 100 LibriSpeech utterances | [ternary.md](docs/ternary.md) | +| Peak RAM against NeMo | about 2x lower at f32 (for example 2582 MB against 5598 MB on `tdt-0.6b-v3`) | BENCHMARK.md, Headline | -Model-dependent (need venv + checkpoint): +VAD accuracy, speed and size, Silero against the Parakeet head against whisper.cpp, with the scripts to repeat them: [vad-benchmarks.md](docs/vad-benchmarks.md). Some of those runs were not on a quiet machine, and the page says which. Transcript parity with NeMo, stage by stage: [parity.md](docs/parity.md). Concurrency throughput (it can go up or down depending on model size and cores): [concurrency.md](docs/concurrency.md). -```sh -export PARAKEET_TEST_GGUF=/tmp/pk110m.gguf -export PARAKEET_TEST_BASELINE=/tmp/baseline.gguf -export PARAKEET_TEST_BASELINE_SPEECH=/tmp/baseline_speech.gguf -ctest --test-dir build --output-on-failure -``` +

+ CPU speedup vs NeMo (RTFx ratio per dtype) + GPU speedup vs NeMo on the NVIDIA GB10 +

+ +--- -Tests labelled `model` return exit code 77 (ctest SKIP) when their required env vars are absent, so they never break a CI environment that has no model. +## Documentation + +| Page | What it covers | +| ---- | -------------- | +| [cli.md](docs/cli.md) | Full CLI examples, the server, quantize | +| [capi.md](docs/capi.md) | C API examples, JSON shapes, streaming | +| [vad.md](docs/vad.md) | The two detectors, options, which one to use, VAD-only files | +| [vad-benchmarks.md](docs/vad-benchmarks.md) | Every VAD measurement and how to repeat it | +| [ultra-redux.md](docs/ultra-redux.md) | The Moondream models: files, credits, dequantized Redux | +| [ternary.md](docs/ternary.md) | Packed ternary format, kernels, limits, speed | +| [bundle.md](docs/bundle.md) | Bundle GGUF format, selection rules, licences, build and verify | +| [diarization.md](docs/diarization.md) | Diarization, speaker-attributed ASR, encoder fingerprint, speed | +| [sound.md](docs/sound.md) | Sound events (CED) and the scene stream | +| [speaker.md](docs/speaker.md) | Enroll and name speakers, C API v9 and v10, what is not measured | +| [batching.md](docs/batching.md) | Batched decode, exactness, how to measure | +| [concurrency.md](docs/concurrency.md) | Backend pool, thread rules, measured throughput | +| [tdt-nbest.md](docs/tdt-nbest.md) | TDT beam search and N-best output | +| [parity.md](docs/parity.md) | Coverage matrix and numerical parity against NeMo | +| [performance.md](docs/performance.md) | Headline speed numbers and where they come from | +| [quantization.md](docs/quantization.md) | Which weights are quantized, size and WER per type | +| [conversion.md](docs/conversion.md) | GGUF schema, Python setup, converting a model | +| [docker.md](docs/docker.md) | Docker images | +| [licenses.md](docs/licenses.md) | Licence of every published model | +| [benchmarks/BENCHMARK.md](benchmarks/BENCHMARK.md) | Full CPU and GPU benchmark, plots, methodology | +| [AGENTS.md](AGENTS.md) | Repository layout, test list, rules for contributors and agents | --- -## Roadmap / TODO +## Limits -- **Tune the GPU encoder kernels.** On the GB10 GPU, parakeet.cpp is faster than NeMo on all 10 models (median 1.25x, up to 4.3x), but the gains are smallest on the pure-encoder CTC models (around 1.2x), because ggml's generic CUDA conv/attention kernels still trail NeMo's tuned cuDNN. Closing that gap (better conv1d and flash-attention paths for the FastConformer encoder) is the main remaining GPU headroom. The log-mel already runs on the backend (`GpuMel`); the CPU path is unaffected. +- **Redux** runs on CPU only and offline only. SIMD kernels exist for x86-64 (AVX2, AVX-512 VNNI) and aarch64 with dotprod; other targets use a slow scalar kernel. See [ternary.md](docs/ternary.md). +- **Ultra and Redux** are not NeMo-validated. Their parity is transcript-level against our own v3 path. +- **The VAD head** in Ultra and Redux gives false alarms on audio without speech. Use Silero as an always-on gate. See [vad.md](docs/vad.md). +- **`transcribe --vad`** is offline and greedy decoding only. VAD-only files cannot transcribe. +- **Bundles:** `bench` and the streaming ASR modes do not take a bundle. No bundle was run on a GPU. See [bundle.md](docs/bundle.md). +- **Speaker identification** was measured on one two-voice fixture. See [speaker.md](docs/speaker.md). +- **Batching** does not apply to standalone CTC models. On CPU the batched decode step is bit-identical to single-clip decode; on GPU it agrees to about 1e-4, and the batched encoder is close to, not equal to, the single-clip one. +- **Concurrency** uses CPU backends only, can be slower than one backend on large models, and uses more memory per backend. +- **`parakeet-server`** serves one model, one request at a time, WAV only. +- **GPU:** CUDA, Metal and Vulkan are supported; ROCm is untested. GPU backends are not exercised in CI. +- **Open work:** the GPU encoder kernels. ggml's generic CUDA conv and attention kernels still trail NeMo's tuned cuDNN, so the gain is smallest on the CTC models (about 1.2x). --- -## Why parakeet.cpp +## Contributing -NeMo is a great training framework, but running Parakeet just for inference drags in a heavy Python/PyTorch stack. parakeet.cpp is a from-scratch C++17/ggml port focused purely on inference: +Issues and pull requests are welcome. Read [AGENTS.md](AGENTS.md) first: it has the repository layout, the performance invariants that must not regress, and the policy for AI-assisted contributions (an `Assisted-by:` trailer, no `Co-Authored-By` for AI). -- **No Python at inference.** A single `libparakeet.so` (or static lib) behind a flat C API (`include/parakeet_capi.h`), easy to embed from C, C++, Go, or Rust. -- **Faster than NeMo** on CPU and GPU (see [Performance](#performance)), with byte-identical output. -- **Small and portable.** GGUF models with f16 / q8_0 / K-quant variants, running on CPU and any ggml GPU backend (CUDA, Metal, Vulkan, HIP). -- **Full family coverage.** CTC, RNNT, TDT, hybrid TDT-CTC, multilingual, and cache-aware streaming with EOU, all validated at WER 0 vs NeMo. +```sh +cmake -B build -DPARAKEET_BUILD_TESTS=ON && cmake --build build -j +ctest --test-dir build --output-on-failure -LE model # no model files needed +``` ---- +Tests labelled `model` need a converted checkpoint (set `PARAKEET_TEST_GGUF` and the other variables listed in AGENTS.md) and skip with exit code 77 when it is missing. -## Community projects +**Community projects** (not maintained by the core team): [parakeet-ios-demo](https://github.com/Kashif-E/parakeet-ios-demo), live on-device streaming speech-to-text on iOS (SwiftUI) over the streaming C API, by [@Kashif-E](https://github.com/Kashif-E). -Built on parakeet.cpp by the community (not maintained or tested by the core team): +--- -- [**parakeet-ios-demo**](https://github.com/Kashif-E/parakeet-ios-demo) โ€” live, - on-device streaming speech-to-text on iOS (SwiftUI) over the streaming C-API, - with a side-by-side compare against Apple's SpeechTranscriber and Moonshine. - By [@Kashif-E](https://github.com/Kashif-E). +## License and credits ---- +parakeet.cpp is released under the [MIT License](LICENSE). The model weights keep the licences of the original models, so check each model card. Most NVIDIA Parakeet models are CC-BY-4.0. `nemotron-3.5-asr-streaming` and Nemotron-3-Diarization are OpenMDW-1.1, `parakeet_realtime_eou_120m-v1` is under the NVIDIA Open Model License, and the Silero VAD files are MIT. Moondream's [parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [parakeet-redux](https://huggingface.co/moondream/parakeet-redux) are CC-BY-4.0: credit Moondream and NVIDIA ([parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)), link the [license](https://creativecommons.org/licenses/by/4.0/), and note that GGUF files made here are converted (and quantized, or dequantized for Redux) copies, not retrained models. The full table is in [docs/licenses.md](docs/licenses.md). A bundle keeps one licence per component (see [bundle.md](docs/bundle.md)). -## Citation +The Parakeet models are by NVIDIA NeMo ([NVIDIA-NeMo/NeMo](https://github.com/NVIDIA-NeMo/NeMo)). Parakeet Ultra and Redux are by [Moondream](https://huggingface.co/moondream), derived from NVIDIA's parakeet-tdt-0.6b-v3. Sound tagging uses [CED](https://github.com/RicherMans/CED) by Heinrich Dinkel and colleagues at Xiaomi. The demo film is Sprite Fright, CC BY 4.0, Blender Studio. If you use parakeet.cpp, please cite this repository and the original models: @@ -566,13 +367,8 @@ If you use parakeet.cpp, please cite this repository and the original models: } ``` -The Parakeet models are by NVIDIA NeMo ([NVIDIA-NeMo/NeMo](https://github.com/NVIDIA-NeMo/NeMo)). -Parakeet Ultra and Redux are by [Moondream](https://huggingface.co/moondream), derived from NVIDIA's parakeet-tdt-0.6b-v3. +Author: Ettore Di Giacinto ([@mudler](https://github.com/mudler)). -## Author - -Ettore Di Giacinto ([@mudler](https://github.com/mudler)). - -## License +--- -parakeet.cpp is released under the [MIT License](LICENSE). The model weights are governed by the licenses of the original models, so check each model card on HuggingFace. The NVIDIA Parakeet models are mostly CC-BY-4.0 (nemotron-3.5-asr-streaming and Nemotron-3-Diarization are OpenMDW-1.1, parakeet_realtime_eou_120m-v1 is the NVIDIA Open Model License, and the Silero VAD files are MIT). See [docs/licenses.md](docs/licenses.md) for the full table. Moondream's [parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [parakeet-redux](https://huggingface.co/moondream/parakeet-redux) are CC-BY-4.0 too: credit Moondream and NVIDIA ([parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)), link the [license](https://creativecommons.org/licenses/by/4.0/), and note that GGUF files made here are converted (and quantized, or dequantized for Redux) copies, not retrained models. +Built by the [LocalAI](https://github.com/mudler/LocalAI) team. If you want to run speech recognition (and LLMs, vision, voice, image, and video models) locally on any hardware with an OpenAI-compatible API, [give LocalAI a star](https://github.com/mudler/LocalAI). diff --git a/benchmarks/media/batch_decode_race.gif b/benchmarks/media/batch_decode_race.gif new file mode 100644 index 0000000..492182f Binary files /dev/null and b/benchmarks/media/batch_decode_race.gif differ diff --git a/benchmarks/media/batch_decode_race.mp4 b/benchmarks/media/batch_decode_race.mp4 new file mode 100644 index 0000000..7b35d3d Binary files /dev/null and b/benchmarks/media/batch_decode_race.mp4 differ diff --git a/benchmarks/media/nemotron_streaming_race.gif b/benchmarks/media/nemotron_streaming_race.gif new file mode 100644 index 0000000..48a6c74 Binary files /dev/null and b/benchmarks/media/nemotron_streaming_race.gif differ diff --git a/benchmarks/media/nemotron_streaming_race.mp4 b/benchmarks/media/nemotron_streaming_race.mp4 new file mode 100644 index 0000000..bd29646 Binary files /dev/null and b/benchmarks/media/nemotron_streaming_race.mp4 differ diff --git a/benchmarks/media/scene_sprite_fright.gif b/benchmarks/media/scene_sprite_fright.gif new file mode 100644 index 0000000..aee93cb Binary files /dev/null and b/benchmarks/media/scene_sprite_fright.gif differ diff --git a/benchmarks/media/scene_sprite_fright.mp4 b/benchmarks/media/scene_sprite_fright.mp4 new file mode 100644 index 0000000..bebfed2 Binary files /dev/null and b/benchmarks/media/scene_sprite_fright.mp4 differ diff --git a/docs/batching.md b/docs/batching.md new file mode 100644 index 0000000..cffac5f --- /dev/null +++ b/docs/batching.md @@ -0,0 +1,20 @@ +# Batched decode + +Single-clip transcription is the default and needs no flags: every `transcribe` call runs one clip at a time, byte-for-byte identical to before. Batching is an opt-in path for decoding several clips together, which matters when you serve many concurrent requests on a GPU. + +The win is on the **decode** side. A transducer (TDT/RNN-T) decodes autoregressively with tiny per-step prediction-LSTM and joint GEMMs; one clip launches hundreds of these matvec-sized kernels and leaves the GPU mostly idle between launches. Decoding N clips together coalesces each step into one batched GEMM, so the device stays busy. On the NVIDIA GB10 this reaches about **10-12x** at batch size 16 (CPU about 3-5x); the encoder is already compute-bound, so batching it gives no throughput win. CTC has no autoregressive decode, so batching does not apply to standalone CTC models. On CPU the batched decode step is bit-identical to decoding each clip alone: its matmuls call the same dot-product kernel as a single column, so logits, token ids, frames and confidences are equal (`tests/test_exact_batch.cpp` checks this for F32, F16 and Q8_0 decoder weights). On GPU backends the batched matmul is the ordinary ggml kernel, so logits agree to about 1e-4 rather than exactly. The batched encoder is separate: its output is close to the single-clip encoder output, not equal, so a whole batched transcript can still differ from a single-clip one, and the tests compare with a tolerance there. Full numbers and per-model tables are in [`../benchmarks/BENCHMARK.md`](../benchmarks/BENCHMARK.md#batched-decode-throughput). + +Measure it yourself: + +```bash +# Decode-only: serial vs batched decode of one clip replicated B times (the win in isolation). +parakeet-cli bench-decode --model --audio [--batch-sizes 1,4,8,16] [--threads N] [--reps R] [--json ] + +# Full transcribe (encoder + decode) over a manifest at several batch sizes. +parakeet-cli bench-batch --model --manifest [--decoder ctc|tdt] [--threads N] [--batch-sizes 1,4,8] [--json ] +``` + +To batch from code, use the batched entry points (single-clip B=1 is just N=1): + +- C++ (`src/model.hpp`): `Model::transcribe_16k_batch(pcms16k, decoder)` and `transcribe_16k_batch_with_timestamps(...)` take N clips of 16 kHz mono float PCM and return N results. +- C-API (`include/parakeet_capi.h`): `parakeet_capi_transcribe_pcm_batch(...)` (N transcripts) and `parakeet_capi_transcribe_pcm_batch_json(...)` (one JSON array of N `{text,words,tokens}` objects). These are what LocalAI's `parakeet-cpp` backend calls to coalesce concurrent requests; it leaves batching off by default and exposes a `batch_max_size` option to opt in. diff --git a/docs/capi.md b/docs/capi.md new file mode 100644 index 0000000..121a51f --- /dev/null +++ b/docs/capi.md @@ -0,0 +1,56 @@ +# C API (`libparakeet.so`) + +The README has the overview and the ABI note. This page has the longer examples. +The current ABI version is 10 (`parakeet_capi_abi_version()`); all additions since v5 are additive. +The full symbol list and the exact signatures are in `include/parakeet_capi.h`. + +`include/parakeet_capi.h` defines a flat, exception-free C-API meant for `dlopen` / FFI / LocalAI integration. Build the shared library with `-DPARAKEET_SHARED=ON`: + +```c +#include "parakeet_capi.h" + +parakeet_ctx *ctx = parakeet_capi_load("model.gguf"); // load ONCE +if (!ctx) { fprintf(stderr, "%s\n", parakeet_capi_last_error(ctx)); return 1; } + +char *text = parakeet_capi_transcribe_path(ctx, "audio.wav", 0 /*default*/); +if (text) { printf("%s\n", text); parakeet_capi_free_string(text); } + +parakeet_capi_free(ctx); +``` + +In-memory PCM: +```c +char *text = parakeet_capi_transcribe_pcm(ctx, samples, n_samples, + sample_rate, 0 /*default*/); +``` + +Timestamps and confidence as JSON (matches NeMo `timestamps=True` + `max_prob`): +```c +char *json = parakeet_capi_transcribe_path_json(ctx, "audio.wav", 0 /*default*/); +// {"text":"...", +// "frame_sec":0.080000, +// "words":[{"w":"Well,","start":0.480,"end":0.640,"conf":0.7859}, ...], +// "tokens":[{"id":639,"t":0.480,"conf":0.9969}, ...]} +if (json) { printf("%s\n", json); parakeet_capi_free_string(json); } +``` +`start`/`end`/`t` are in seconds; `conf` is the rescaled softmax probability of the emitted token in `(0,1]` (a word's `conf` is the `min` over its tokens). `frame_sec` is the encoder frame stride in seconds (`hop x subsampling / sample_rate`); multiply a frame-unit segment gap threshold (NeMo's `segment_gap_threshold`) by it to get the seconds gap between words when forming segments. + +## Streaming (cache-aware EOU model) + +For `parakeet_realtime_eou_120m-v1`, a streaming session decodes 16 kHz mono f32 PCM as it arrives, returning newly-finalized text and signalling EOU/EOB events: + +```c +parakeet_stream *s = parakeet_capi_stream_begin(ctx); +int eou = 0; +char *t = parakeet_capi_stream_feed(s, pcm, n_samples, &eou); // "" if none yet +if (t) { printf("%s", t); parakeet_capi_free_string(t); } +if (eou) printf(" [EOU]"); +// ...feed more chunks... +char *tail = parakeet_capi_stream_finalize(s); // flush the tail +if (tail) { printf("%s\n", tail); parakeet_capi_free_string(tail); } +parakeet_capi_stream_free(s); +``` + +`` (end-of-utterance) and `` (backchannel) are stripped from the text and surfaced via `*eou_out` (the CLI `--stream` prints them as `[EOU @ s]` markers). The streaming transcript matches NeMo's cache-aware streaming exactly, and `finalize` flushes the end-of-stream tail without fabricating an `` that NeMo would not emit. + +The LocalAI backend (in the LocalAI repo) dlopens `libparakeet.so` and uses these symbols directly: the offline `parakeet_capi_transcribe_*` / `parakeet_capi_transcribe_path_json` and the streaming `parakeet_capi_stream_*`. See `include/parakeet_capi.h` for the full API. The C++ streaming session (`pk::StreamingSession`) also exposes per-word timestamps and confidence as words finalize, via `drain_words()` alongside the EOU events, which the CLI `--stream --timestamps` path prints. diff --git a/docs/cli.md b/docs/cli.md new file mode 100644 index 0000000..04a6430 --- /dev/null +++ b/docs/cli.md @@ -0,0 +1,117 @@ +# Command-line reference + +`parakeet-cli` lands at `build/examples/cli/parakeet-cli`. The README has a short cheat sheet; this page has the full set of examples. + +## Transcribe, VAD, info and streaming + +```sh +# Default decoder (TDT for hybrid/TDT models, CTC for standalone CTC) +parakeet-cli transcribe --model m.gguf --input audio.wav + +# Force a decoder +parakeet-cli transcribe --model m.gguf --input audio.wav --decoder ctc +parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt + +# Per-word timestamps + confidence: one line per word +# - () (times in seconds) +parakeet-cli transcribe --model m.gguf --input audio.wav --timestamps + +# JSON with the flat text plus per-word and per-token timestamps + confidence: +# {"text":"...","words":[{"w":..,"start":..,"end":..,"conf":..}], +# "tokens":[{"id":..,"t":..,"conf":..}]} +parakeet-cli transcribe --model m.gguf --input audio.wav --json + +# Offline TDT N-best hypotheses as ranked JSON +parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt \ + --beam-size 4 --nbest 4 + +# Read WAV bytes from stdin (useful with ffmpeg/curl pipelines) +ffmpeg -i input.mp3 -f wav - | parakeet-cli transcribe --model m.gguf --input - + +# Long audio on Ultra/Redux: cut at VAD pauses, transcribe each piece (offline only). +# Tune with --vad-threshold F (0.5), --vad-min-pause SEC (0.2), --vad-max-seg SEC (30) +parakeet-cli transcribe --model ultra.gguf --input long.wav --vad + +# Voice activity detection only, no transcript: speech regions as JSON +# {"mode":"speech","duration":..,"frame_sec":0.08,"backend":"cpu", +# "segments":[{"start":..,"end":..}]} (seconds; models with a VAD head only) +# --mode segments gives the cuts that `transcribe --vad` uses; --probabilities adds p per 80 ms frame. +# Tune with --threshold F (0.5), --min-pause SEC (0.2), --min-speech SEC (0.1), --max-segment SEC (30) +parakeet-cli vad --model ultra.gguf --input audio.wav + +# Only the VAD head is needed? Use a 6 to 10 MB slice instead of the full model +# (it cannot transcribe). Files and checksums: ./vad.md. +curl -LO https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/redux-vad.gguf +parakeet-cli vad --model redux-vad.gguf --input audio.wav + +# The same with a Silero VAD GGUF (frame_sec 0.032; defaults 250 ms min speech, +# 100 ms min pause, 30 ms pad). Any ASR model can then cut long audio with it. +# Download it from the collection repo (F16 is 1.3 MB, F32 is 2.2 MB). To make the +# file yourself, see scripts/convert_silero_vad_to_gguf.py, ./vad.md and ./conversion.md. +curl -LO https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/silero-vad-f16.gguf +parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav +parakeet-cli transcribe --model tdt-0.6b-v3.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf + +# Print model metadata (arch, dims, mel params, vocab size, TDT durations) +parakeet-cli info m.gguf + +# Cache-aware streaming (EOU model parakeet_realtime_eou_120m-v1): feeds the WAV +# in the model's chunk schedule, prints partial text incrementally and +# [EOU @ s] / [EOB @ s] event markers, then the finalized tail. Add +# --timestamps to also print per-word [start-end] (conf) lines as words finalize. +parakeet-cli transcribe --model eou.gguf --input audio.wav --stream +``` + +Timestamps and confidence match NeMo's `transcribe(timestamps=True)` with the `max_prob` confidence method exactly (word offsets to 0.0 s, per-token and per-word confidence within `5e-6`), for both the TDT and CTC heads. See `./parity.md`. Word start and end are in seconds (`frame x hop x subsampling / sample_rate`, which works out to 0.08 s/frame here); confidence is the rescaled softmax probability of the emitted token, aggregated per word with NeMo's `min`. + +The optional TDT beam decoder follows NeMo's default sequence-level beam +search and exposes raw/normalized scores plus token frame/duration metadata. +See [`./tdt-nbest.md`](./tdt-nbest.md). + +The `parakeet-cli` binary lands at `build/examples/cli/parakeet-cli`. + +## OpenAI-compatible server + +`parakeet-server` is a small HTTP server that speaks the OpenAI transcription +API, so any OpenAI client works by pointing its `base_url` at it. It is built by +default (`PARAKEET_BUILD_SERVER=ON`) and lands at `build/examples/server/parakeet-server`. + +```sh +# Serve a model. --model takes a local .gguf, an http(s) URL, a .gguf in +# mudler/parakeet-cpp-gguf, or an alias (downloaded and cached on first run). +parakeet-server --model tdt_ctc-110m --port 8080 + +# Transcribe over HTTP +curl -F file=@audio.wav -F response_format=verbose_json \ + http://localhost:8080/v1/audio/transcriptions +``` + +```python +from openai import OpenAI +client = OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed") +with open("audio.wav", "rb") as f: + print(client.audio.transcriptions.create(model="parakeet", file=f).text) +``` + +It supports `response_format` `json` / `text` / `verbose_json` and +`timestamp_granularities[]=word`. This is a single-model, one-request-at-a-time +example that accepts WAV uploads only; see [`examples/server/README.md`](../examples/server/README.md) +for the full list of options and known simplifications. **For a production +deployment, use [LocalAI](https://localai.io)**, which embeds parakeet.cpp as a +backend and adds a model gallery, concurrency, multi-model serving, the full +OpenAI API surface, auth, and metrics. + +## Quantize + +The Python `gguf` writer can't produce K-quants (`q4_k`, `q5_k`, `q6_k`), so re-quantize an existing F32 GGUF with the CLI instead: + +```sh +parakeet-cli quantize +# e.g. +parakeet-cli quantize m.gguf m_q4k.gguf q4_k +parakeet-cli quantize m.gguf m_q6k.gguf q6_k +``` + +Supported types: `q4_0`, `q5_0`, `q8_0`, `q4_k`, `q5_k`, `q6_k`. + +Only the large linear `ggml_mul_mat`-consumed weights (encoder FFN, attention projections, joint enc/pred projections, subsampling output projection) get quantized. The conv, LSTM, featurizer, batch_norm, and bias tensors stay F32. See [quantization.md](quantization.md) for the full policy, allowlist, and measured size and WER per type. diff --git a/docs/conversion.md b/docs/conversion.md index bf57f97..b75d37d 100644 --- a/docs/conversion.md +++ b/docs/conversion.md @@ -332,3 +332,38 @@ only makes the file smaller; it does not change the compute path. The LSTM gate order is the PyTorch one: input, forget, cell, output. The encoder stride pattern reduces the 4 STFT frames of one chunk to one vector. + +## Setting up Python and converting a model + +You need this once, for model conversion and validation. It's not needed for inference: + +```sh +python3 -m venv .venv +.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu +.venv/bin/pip install -r scripts/requirements.txt # nemo_toolkit[asr] + gguf +``` + +NeMo 2.7.3 is the validated version. The anchor checkpoint `nvidia/parakeet-tdt_ctc-110m` (about 440 MB) is downloaded automatically by NeMo on first use. + +--- + +## Converting a model + +Convert a HuggingFace or local `.nemo` checkpoint to GGUF: + +```sh +# Default (F32), lossless and largest +.venv/bin/python scripts/convert_parakeet_to_gguf.py \ + --model nvidia/parakeet-tdt_ctc-110m \ + --output m.gguf + +# F16, about 0.58x the size, WER 0 vs NeMo +.venv/bin/python scripts/convert_parakeet_to_gguf.py \ + --model nvidia/parakeet-tdt_ctc-110m --dtype f16 --output m.gguf + +# Q8_0, about 0.39x the size, WER 0 vs NeMo +.venv/bin/python scripts/convert_parakeet_to_gguf.py \ + --model nvidia/parakeet-tdt_ctc-110m --dtype q8_0 --output m.gguf +``` + +Supported `--dtype`: `f32` (default), `f16`, `q8_0`. diff --git a/docs/docker.md b/docs/docker.md new file mode 100644 index 0000000..7abdef8 --- /dev/null +++ b/docs/docker.md @@ -0,0 +1,31 @@ +# Docker images + +Two prebuilt images are published to GitHub Container Registry on every push to `master`, one per binary: + +- `ghcr.io/mudler/parakeet.cpp-cli`: the command-line transcriber. +- `ghcr.io/mudler/parakeet.cpp-server`: the [OpenAI-compatible server](cli.md#openai-compatible-server). + +Each comes in a CPU and a CUDA variant (the CUDA tag is suffixed `-cuda`), and both are multi-arch (`linux/amd64` and `linux/arm64`), so the right one is pulled for your host automatically. They contain just the binary, so mount a converted `.gguf` model (and, for the cli, your audio) at runtime: + +```sh +# CLI, CPU +docker run --rm \ + -v "$PWD/models:/models:ro" \ + -v "$PWD/audio:/audio:ro" \ + ghcr.io/mudler/parakeet.cpp-cli:latest \ + transcribe --model /models/parakeet-tdt_ctc-110m-q5_k.gguf --input /audio/speech.wav --decoder tdt + +# CLI, CUDA (needs the nvidia container toolkit on the host) +docker run --rm --gpus all \ + -v "$PWD/models:/models:ro" -v "$PWD/audio:/audio:ro" \ + ghcr.io/mudler/parakeet.cpp-cli:latest-cuda \ + transcribe --model /models/parakeet-tdt_ctc-110m-q5_k.gguf --input /audio/speech.wav --decoder tdt + +# Server: binds 0.0.0.0 and exposes 8080. Fetch a model by alias on first run, +# or mount a local .gguf. Add --gpus all with the :latest-cuda tag for GPU. +docker run --rm -p 8080:8080 ghcr.io/mudler/parakeet.cpp-server:latest --model tdt_ctc-110m +``` + +The CUDA image is built on CUDA 13, so it covers everything from Turing up through Blackwell, including GB10 / Grace-Blackwell (DGX Spark) on arm64. + +To build the images yourself, see the build args at the top of the [`Dockerfile`](../Dockerfile); the cli is the default target and the server is `--target runtime-server`. The CPU image is the portable `GGML_NATIVE=OFF` build, so it runs on any amd64 or arm64 host. diff --git a/docs/performance.md b/docs/performance.md new file mode 100644 index 0000000..30b28cc --- /dev/null +++ b/docs/performance.md @@ -0,0 +1,59 @@ +# Performance summary + +This page keeps the headline speed numbers that the README links to. The full +methodology, all ten models, the plots and the raw results are in +[../benchmarks/BENCHMARK.md](../benchmarks/BENCHMARK.md). Read the method before you +quote a number: the CPU runs used 8 threads on a 20-core x86 host, the GPU runs used +one NVIDIA GB10, and each machine was a shared development host, not a quiet lab box. + +RTFx is audio seconds divided by processing seconds. Higher is faster. Speedup is our +RTFx divided by the NeMo (PyTorch) RTFx on the same machine, batch size 1. + +## CPU against NeMo (LibriSpeech test-clean, 100 utterances, 8 threads) + +| dtype | size vs f32 | mean speedup vs NeMo | accuracy | +| ----- | ----------- | -------------------- | -------- | +| f32 | 100% | 1.40x (range 1.11x to 1.69x over 10 models) | mean agreement WER 0.015%, byte-identical on most models | +| f16 | 57% | 1.70x | same as f32 | +| q8_0 | 37% | 1.56x (up to 1.89x) | mean agreement WER 0.16% | +| q4_k | 26% | 1.25x | agreement WER 1.08%, small and monotonic accuracy cost | + +Source: the "Quantization" and "Headline" tables in BENCHMARK.md. Peak RAM is roughly +2x lower than NeMo at f32 (for example 2582 MB against 5598 MB for `tdt-0.6b-v3`) and +lower still once quantized. + +## GPU against NeMo (NVIDIA GB10, f32, LibriSpeech) + +Median speedup over the 10 offline and streaming-EOU models: 1.25x. Best case: 4.3x on +`tdt_ctc-110m`. The smallest gains are on the pure-encoder CTC models (about 1.2x), +because ggml's generic CUDA conv and attention kernels still trail NeMo's tuned cuDNN. +The log-mel front end runs on the GPU through a ggml DFT-matmul graph; the CPU path is +unchanged. NeMo's TDT greedy decode is not CUDA-graph accelerated here and ours is a lean +C++ loop, which explains most of the gap on the TDT and hybrid models. + +These numbers come from `benchmarks/results_gpu/` (`scripts/plot_gpu.py` plots them); +the median and maximum above were recomputed from those files when this page was written. +NeMo ran in the `nvcr.io/nvidia/nemo` container for these runs. + +## Single clip, newer models + +`nemotron-3.5-asr-streaming-0.6b` on a 7.43 s clip (Ryzen 9 9950X3D, 8 threads, median of +7 passes): 2.40x NeMo at f32 and 2.52x at q8_0, byte-identical transcripts. On the GB10 GPU: +1.16x at f32 and 1.30x at q8_0. See the Nemotron section of BENCHMARK.md. + +## Against whisper.cpp + +The comparison plot [../benchmarks/plots/vs_whisper.png](../benchmarks/plots/vs_whisper.png) shows RTFx and accuracy for parakeet.cpp and whisper.cpp on CPU and GPU (the earlier README text said the 110M Parakeet is faster than whisper base.en and far faster than large-v3-turbo; read the plot for the exact values). The side-by-side clips measure about 12x (GPU) and about 27x (CPU) against whisper.cpp +turbo on one clip with WER 1.6% for both. One clip is a demo, not a benchmark. See +[../benchmarks/media/gpu_whisper_duel.mp4](../benchmarks/media/gpu_whisper_duel.mp4) and +[../benchmarks/media/cpu_duel.mp4](../benchmarks/media/cpu_duel.mp4). + +## Apple Metal + +On an Apple M4, Metal is about 1.3x to 5.6x faster than CPU, most on the larger models. See +[BENCHMARK.md, Apple Metal](../benchmarks/BENCHMARK.md#apple-metal-m4). + +## Batched decode + +Up to about 10x to 12x on the GB10 at batch 16 and about 3x to 5x on CPU. See +[batching.md](batching.md). diff --git a/docs/ultra-redux.md b/docs/ultra-redux.md new file mode 100644 index 0000000..9a8f311 --- /dev/null +++ b/docs/ultra-redux.md @@ -0,0 +1,54 @@ +# Moondream Parakeet Ultra and Redux + +This page holds the model notes for the two Moondream derivatives of Parakeet. The +runtime details are in other pages: the ternary kernels and their speed in +[ternary.md](ternary.md), the VAD head and Silero in [vad.md](vad.md), and the +measurements in [vad-benchmarks.md](vad-benchmarks.md). + +[moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and +[moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux) are Moondream's +post-trained (Ultra, F16) and ternary-encoder (Redux) derivatives of NVIDIA's +[parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). Both are released under +[CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). They are HF safetensors, converted with +`../scripts/convert_hf_parakeet_to_gguf.py`. They are not part of the NeMo-validated set above: there is +no NeMo baseline for them, so parity is transcript-level against our own v3 path (see +[`parity.md`](parity.md)), and the GGUFs are published in +[mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf): `ultra-f16.gguf`, `ultra-q8_0.gguf`, `redux-packed.gguf` +(packed ternary), `redux-f16.gguf` and `redux-q8_0.gguf` (dequantized). Sizes and SHA-256 sums are in +[`models/MANIFEST.md`](../models/MANIFEST.md). + +The models were trained by NVIDIA (the base) and Moondream (Ultra and Redux). parakeet.cpp only +converts and quantizes the weights; nothing is trained or fine-tuned here. A dequantized Redux file +(`--ternary dequant`, the converter default) holds ordinary F16 or Q8_0 weights expanded from the +ternary ones. + +- Redux packs the encoder as ternary weights: a 213 MB GGUF, 6.8x smaller than F16. It runs on CPU + only and offline only. On x86 with AVX-512 VNNI it reaches median RTF 75.6 per utterance on + LibriSpeech-100 (8 threads) against 46.1 for the same model in F16; on a single 180 s clip the + gain is about 10 percent; WER on the 100 LibriSpeech utterances is 1.96 percent. See + [`ternary.md`](ternary.md). SIMD kernels exist for x86-64 with AVX2 or AVX-512 VNNI and + aarch64 with dotprod; MSVC builds, Windows on ARM and aarch64 without dotprod use a slow scalar + kernel (about 1 GMAC/s), and the load logs a warning. The packed file also stays resident next to + the repacked planes, so memory use is more than the file size. +- Both carry a voice-activity head, used by `transcribe --vad` to cut long audio at pauses. Speech + is a probability of at least 0.5; pauses of at least 0.2 s are candidate cuts, segments are at most + 30 s, and segments without speech are dropped. On long-form clips it does not change WER + meaningfully. Details and measurements: [`ternary.md`](ternary.md). +- The same head runs on its own, without transcribing: `parakeet-cli vad`, or + `parakeet_capi_vad_pcm_json` / `parakeet_capi_vad_path_json` from the C-API. They return speech + segments (start and end in seconds) as JSON, and optionally the per-frame probabilities. A model + without the head fails with `model has no VAD head`. See [`vad.md`](vad.md). +- Silero VAD (MIT, 32 ms frames, 16 kHz and 8 kHz) runs from its own small GGUF ([download](https://huggingface.co/mudler/parakeet-cpp-gguf)) through the same + functions, as a stream (`parakeet_capi_vad_stream_*`), and as the cutter for any ASR model: + `parakeet-cli transcribe --vad --vad-model silero.gguf`. See [`vad.md`](vad.md). +- VAD-only slices of Ultra and Redux (6 to 10 MB, `redux-vad.gguf` and `ultra-vad-q8_0.gguf` in the same repo) hold just the head and its front end. They run `vad` and the `parakeet_capi_vad_*` calls and cannot transcribe. See [`vad.md`](vad.md). +- The head alone gives false alarms on audio without speech: on speech-free noise it calls about + 99 percent of the frames speech (Ultra 99.4, Redux 97.8), and over a 30 s noise stretch inside a + file with speech the false-alarm frame rate was 17.7 percent for Ultra and 55 percent for Redux, + against 0 percent for Silero (synthetic LibriSpeech with added noise). Prefer Silero as the + always-on gate or when the audio can have long non-speech stretches; use the head on audio known + to be mostly speech, or where its higher recall matters. An offline experiment that lets Silero + decide and the head move the edges (not implemented here) is in + [`vad-benchmarks.md`](vad-benchmarks.md#fusing-silero-and-the-head-offline-experiment). +- Accuracy, speed and size of both detectors, with the method and the scripts to repeat + them: [`vad-benchmarks.md`](vad-benchmarks.md).