Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -203,6 +203,8 @@ docs/
diarization.md , speaker diarization + speaker-attributed ASR: parity, C-API, speed
sound.md , sound-event detection (CED) and the combined scene stream
speaker.md , speaker identification: enroll, scene naming, C-API v9 and v10, measured numbers
vad.md, vad-benchmarks.md, ultra-redux.md, batching.md, performance.md, cli.md, capi.md, docker.md, licenses.md, tdt-nbest.md
, see the documentation table in README.md
.github/workflows/
ci.yml , build job (per-push) + closed-loop job (pull_request + dispatch)
```
Expand All @@ -223,7 +225,7 @@ cmake -B build -DPARAKEET_BUILD_TESTS=ON -DGGML_NATIVE=ON && cmake --build build
| `PARAKEET_GGML_CUDA` | OFF | Forward GGML_CUDA to the submodule |
| `PARAKEET_GGML_METAL` | OFF | Forward GGML_METAL to the submodule |
| `PARAKEET_GGML_VULKAN` | OFF | Forward GGML_VULKAN to the submodule |
| `PARAKEET_GGML_HIPBLAS` | OFF | Forward GGML_HIPBLAS to the submodule |
| `PARAKEET_GGML_HIP` | OFF | Forward GGML_HIP (ROCm) to the submodule |
| `PARAKEET_WITH_CED` | ON | Sound-event detection through ced.cpp |
| `PARAKEET_WITH_VOICEDETECT` | ON | Speaker identification through voice-detect.cpp |

Expand Down
658 changes: 227 additions & 431 deletions README.md

Large diffs are not rendered by default.

Binary file added benchmarks/media/batch_decode_race.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added benchmarks/media/batch_decode_race.mp4
Binary file not shown.
Binary file added benchmarks/media/nemotron_streaming_race.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added benchmarks/media/nemotron_streaming_race.mp4
Binary file not shown.
Binary file added benchmarks/media/scene_sprite_fright.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added benchmarks/media/scene_sprite_fright.mp4
Binary file not shown.
20 changes: 20 additions & 0 deletions docs/batching.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Batched decode

Single-clip transcription is the default and needs no flags: every `transcribe` call runs one clip at a time, byte-for-byte identical to before. Batching is an opt-in path for decoding several clips together, which matters when you serve many concurrent requests on a GPU.

The win is on the **decode** side. A transducer (TDT/RNN-T) decodes autoregressively with tiny per-step prediction-LSTM and joint GEMMs; one clip launches hundreds of these matvec-sized kernels and leaves the GPU mostly idle between launches. Decoding N clips together coalesces each step into one batched GEMM, so the device stays busy. On the NVIDIA GB10 this reaches about **10-12x** at batch size 16 (CPU about 3-5x); the encoder is already compute-bound, so batching it gives no throughput win. CTC has no autoregressive decode, so batching does not apply to standalone CTC models. On CPU the batched decode step is bit-identical to decoding each clip alone: its matmuls call the same dot-product kernel as a single column, so logits, token ids, frames and confidences are equal (`tests/test_exact_batch.cpp` checks this for F32, F16 and Q8_0 decoder weights). On GPU backends the batched matmul is the ordinary ggml kernel, so logits agree to about 1e-4 rather than exactly. The batched encoder is separate: its output is close to the single-clip encoder output, not equal, so a whole batched transcript can still differ from a single-clip one, and the tests compare with a tolerance there. Full numbers and per-model tables are in [`../benchmarks/BENCHMARK.md`](../benchmarks/BENCHMARK.md#batched-decode-throughput).

Measure it yourself:

```bash
# Decode-only: serial vs batched decode of one clip replicated B times (the win in isolation).
parakeet-cli bench-decode --model <model.gguf> --audio <wav> [--batch-sizes 1,4,8,16] [--threads N] [--reps R] [--json <out>]

# Full transcribe (encoder + decode) over a manifest at several batch sizes.
parakeet-cli bench-batch --model <model.gguf> --manifest <file> [--decoder ctc|tdt] [--threads N] [--batch-sizes 1,4,8] [--json <out>]
```

To batch from code, use the batched entry points (single-clip B=1 is just N=1):

- C++ (`src/model.hpp`): `Model::transcribe_16k_batch(pcms16k, decoder)` and `transcribe_16k_batch_with_timestamps(...)` take N clips of 16 kHz mono float PCM and return N results.
- C-API (`include/parakeet_capi.h`): `parakeet_capi_transcribe_pcm_batch(...)` (N transcripts) and `parakeet_capi_transcribe_pcm_batch_json(...)` (one JSON array of N `{text,words,tokens}` objects). These are what LocalAI's `parakeet-cpp` backend calls to coalesce concurrent requests; it leaves batching off by default and exposes a `batch_max_size` option to opt in.
56 changes: 56 additions & 0 deletions docs/capi.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# C API (`libparakeet.so`)

The README has the overview and the ABI note. This page has the longer examples.
The current ABI version is 10 (`parakeet_capi_abi_version()`); all additions since v5 are additive.
The full symbol list and the exact signatures are in `include/parakeet_capi.h`.

`include/parakeet_capi.h` defines a flat, exception-free C-API meant for `dlopen` / FFI / LocalAI integration. Build the shared library with `-DPARAKEET_SHARED=ON`:

```c
#include "parakeet_capi.h"

parakeet_ctx *ctx = parakeet_capi_load("model.gguf"); // load ONCE
if (!ctx) { fprintf(stderr, "%s\n", parakeet_capi_last_error(ctx)); return 1; }

char *text = parakeet_capi_transcribe_path(ctx, "audio.wav", 0 /*default*/);
if (text) { printf("%s\n", text); parakeet_capi_free_string(text); }

parakeet_capi_free(ctx);
```

In-memory PCM:
```c
char *text = parakeet_capi_transcribe_pcm(ctx, samples, n_samples,
sample_rate, 0 /*default*/);
```

Timestamps and confidence as JSON (matches NeMo `timestamps=True` + `max_prob`):
```c
char *json = parakeet_capi_transcribe_path_json(ctx, "audio.wav", 0 /*default*/);
// {"text":"...",
// "frame_sec":0.080000,
// "words":[{"w":"Well,","start":0.480,"end":0.640,"conf":0.7859}, ...],
// "tokens":[{"id":639,"t":0.480,"conf":0.9969}, ...]}
if (json) { printf("%s\n", json); parakeet_capi_free_string(json); }
```
`start`/`end`/`t` are in seconds; `conf` is the rescaled softmax probability of the emitted token in `(0,1]` (a word's `conf` is the `min` over its tokens). `frame_sec` is the encoder frame stride in seconds (`hop x subsampling / sample_rate`); multiply a frame-unit segment gap threshold (NeMo's `segment_gap_threshold`) by it to get the seconds gap between words when forming segments.

## Streaming (cache-aware EOU model)

For `parakeet_realtime_eou_120m-v1`, a streaming session decodes 16 kHz mono f32 PCM as it arrives, returning newly-finalized text and signalling EOU/EOB events:

```c
parakeet_stream *s = parakeet_capi_stream_begin(ctx);
int eou = 0;
char *t = parakeet_capi_stream_feed(s, pcm, n_samples, &eou); // "" if none yet
if (t) { printf("%s", t); parakeet_capi_free_string(t); }
if (eou) printf(" [EOU]");
// ...feed more chunks...
char *tail = parakeet_capi_stream_finalize(s); // flush the tail
if (tail) { printf("%s\n", tail); parakeet_capi_free_string(tail); }
parakeet_capi_stream_free(s);
```

`<EOU>` (end-of-utterance) and `<EOB>` (backchannel) are stripped from the text and surfaced via `*eou_out` (the CLI `--stream` prints them as `[EOU @ <t>s]` markers). The streaming transcript matches NeMo's cache-aware streaming exactly, and `finalize` flushes the end-of-stream tail without fabricating an `<EOU>` that NeMo would not emit.

The LocalAI backend (in the LocalAI repo) dlopens `libparakeet.so` and uses these symbols directly: the offline `parakeet_capi_transcribe_*` / `parakeet_capi_transcribe_path_json` and the streaming `parakeet_capi_stream_*`. See `include/parakeet_capi.h` for the full API. The C++ streaming session (`pk::StreamingSession`) also exposes per-word timestamps and confidence as words finalize, via `drain_words()` alongside the EOU events, which the CLI `--stream --timestamps` path prints.
117 changes: 117 additions & 0 deletions docs/cli.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
# Command-line reference

`parakeet-cli` lands at `build/examples/cli/parakeet-cli`. The README has a short cheat sheet; this page has the full set of examples.

## Transcribe, VAD, info and streaming

```sh
# Default decoder (TDT for hybrid/TDT models, CTC for standalone CTC)
parakeet-cli transcribe --model m.gguf --input audio.wav

# Force a decoder
parakeet-cli transcribe --model m.gguf --input audio.wav --decoder ctc
parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt

# Per-word timestamps + confidence: one line per word
# <start>-<end> <word> (<conf>) (times in seconds)
parakeet-cli transcribe --model m.gguf --input audio.wav --timestamps

# JSON with the flat text plus per-word and per-token timestamps + confidence:
# {"text":"...","words":[{"w":..,"start":..,"end":..,"conf":..}],
# "tokens":[{"id":..,"t":..,"conf":..}]}
parakeet-cli transcribe --model m.gguf --input audio.wav --json

# Offline TDT N-best hypotheses as ranked JSON
parakeet-cli transcribe --model m.gguf --input audio.wav --decoder tdt \
--beam-size 4 --nbest 4

# Read WAV bytes from stdin (useful with ffmpeg/curl pipelines)
ffmpeg -i input.mp3 -f wav - | parakeet-cli transcribe --model m.gguf --input -

# Long audio on Ultra/Redux: cut at VAD pauses, transcribe each piece (offline only).
# Tune with --vad-threshold F (0.5), --vad-min-pause SEC (0.2), --vad-max-seg SEC (30)
parakeet-cli transcribe --model ultra.gguf --input long.wav --vad

# Voice activity detection only, no transcript: speech regions as JSON
# {"mode":"speech","duration":..,"frame_sec":0.08,"backend":"cpu",
# "segments":[{"start":..,"end":..}]} (seconds; models with a VAD head only)
# --mode segments gives the cuts that `transcribe --vad` uses; --probabilities adds p per 80 ms frame.
# Tune with --threshold F (0.5), --min-pause SEC (0.2), --min-speech SEC (0.1), --max-segment SEC (30)
parakeet-cli vad --model ultra.gguf --input audio.wav

# Only the VAD head is needed? Use a 6 to 10 MB slice instead of the full model
# (it cannot transcribe). Files and checksums: ./vad.md.
curl -LO https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/redux-vad.gguf
parakeet-cli vad --model redux-vad.gguf --input audio.wav

# The same with a Silero VAD GGUF (frame_sec 0.032; defaults 250 ms min speech,
# 100 ms min pause, 30 ms pad). Any ASR model can then cut long audio with it.
# Download it from the collection repo (F16 is 1.3 MB, F32 is 2.2 MB). To make the
# file yourself, see scripts/convert_silero_vad_to_gguf.py, ./vad.md and ./conversion.md.
curl -LO https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/silero-vad-f16.gguf
parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav
parakeet-cli transcribe --model tdt-0.6b-v3.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf

# Print model metadata (arch, dims, mel params, vocab size, TDT durations)
parakeet-cli info m.gguf

# Cache-aware streaming (EOU model parakeet_realtime_eou_120m-v1): feeds the WAV
# in the model's chunk schedule, prints partial text incrementally and
# [EOU @ <t>s] / [EOB @ <t>s] event markers, then the finalized tail. Add
# --timestamps to also print per-word [start-end] (conf) lines as words finalize.
parakeet-cli transcribe --model eou.gguf --input audio.wav --stream
```

Timestamps and confidence match NeMo's `transcribe(timestamps=True)` with the `max_prob` confidence method exactly (word offsets to 0.0 s, per-token and per-word confidence within `5e-6`), for both the TDT and CTC heads. See `./parity.md`. Word start and end are in seconds (`frame x hop x subsampling / sample_rate`, which works out to 0.08 s/frame here); confidence is the rescaled softmax probability of the emitted token, aggregated per word with NeMo's `min`.

The optional TDT beam decoder follows NeMo's default sequence-level beam
search and exposes raw/normalized scores plus token frame/duration metadata.
See [`./tdt-nbest.md`](./tdt-nbest.md).

The `parakeet-cli` binary lands at `build/examples/cli/parakeet-cli`.

## OpenAI-compatible server

`parakeet-server` is a small HTTP server that speaks the OpenAI transcription
API, so any OpenAI client works by pointing its `base_url` at it. It is built by
default (`PARAKEET_BUILD_SERVER=ON`) and lands at `build/examples/server/parakeet-server`.

```sh
# Serve a model. --model takes a local .gguf, an http(s) URL, a <name>.gguf in
# mudler/parakeet-cpp-gguf, or an alias (downloaded and cached on first run).
parakeet-server --model tdt_ctc-110m --port 8080

# Transcribe over HTTP
curl -F file=@audio.wav -F response_format=verbose_json \
http://localhost:8080/v1/audio/transcriptions
```

```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="not-needed")
with open("audio.wav", "rb") as f:
print(client.audio.transcriptions.create(model="parakeet", file=f).text)
```

It supports `response_format` `json` / `text` / `verbose_json` and
`timestamp_granularities[]=word`. This is a single-model, one-request-at-a-time
example that accepts WAV uploads only; see [`examples/server/README.md`](../examples/server/README.md)
for the full list of options and known simplifications. **For a production
deployment, use [LocalAI](https://localai.io)**, which embeds parakeet.cpp as a
backend and adds a model gallery, concurrency, multi-model serving, the full
OpenAI API surface, auth, and metrics.

## Quantize

The Python `gguf` writer can't produce K-quants (`q4_k`, `q5_k`, `q6_k`), so re-quantize an existing F32 GGUF with the CLI instead:

```sh
parakeet-cli quantize <in.gguf> <out.gguf> <type>
# e.g.
parakeet-cli quantize m.gguf m_q4k.gguf q4_k
parakeet-cli quantize m.gguf m_q6k.gguf q6_k
```

Supported types: `q4_0`, `q5_0`, `q8_0`, `q4_k`, `q5_k`, `q6_k`.

Only the large linear `ggml_mul_mat`-consumed weights (encoder FFN, attention projections, joint enc/pred projections, subsampling output projection) get quantized. The conv, LSTM, featurizer, batch_norm, and bias tensors stay F32. See [quantization.md](quantization.md) for the full policy, allowlist, and measured size and WER per type.
35 changes: 35 additions & 0 deletions docs/conversion.md
Original file line number Diff line number Diff line change
Expand Up @@ -332,3 +332,38 @@ only makes the file smaller; it does not change the compute path.

The LSTM gate order is the PyTorch one: input, forget, cell, output. The
encoder stride pattern reduces the 4 STFT frames of one chunk to one vector.

## Setting up Python and converting a model

You need this once, for model conversion and validation. It's not needed for inference:

```sh
python3 -m venv .venv
.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu
.venv/bin/pip install -r scripts/requirements.txt # nemo_toolkit[asr] + gguf
```

NeMo 2.7.3 is the validated version. The anchor checkpoint `nvidia/parakeet-tdt_ctc-110m` (about 440 MB) is downloaded automatically by NeMo on first use.

---

## Converting a model

Convert a HuggingFace or local `.nemo` checkpoint to GGUF:

```sh
# Default (F32), lossless and largest
.venv/bin/python scripts/convert_parakeet_to_gguf.py \
--model nvidia/parakeet-tdt_ctc-110m \
--output m.gguf

# F16, about 0.58x the size, WER 0 vs NeMo
.venv/bin/python scripts/convert_parakeet_to_gguf.py \
--model nvidia/parakeet-tdt_ctc-110m --dtype f16 --output m.gguf

# Q8_0, about 0.39x the size, WER 0 vs NeMo
.venv/bin/python scripts/convert_parakeet_to_gguf.py \
--model nvidia/parakeet-tdt_ctc-110m --dtype q8_0 --output m.gguf
```

Supported `--dtype`: `f32` (default), `f16`, `q8_0`.
31 changes: 31 additions & 0 deletions docs/docker.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Docker images

Two prebuilt images are published to GitHub Container Registry on every push to `master`, one per binary:

- `ghcr.io/mudler/parakeet.cpp-cli`: the command-line transcriber.
- `ghcr.io/mudler/parakeet.cpp-server`: the [OpenAI-compatible server](cli.md#openai-compatible-server).

Each comes in a CPU and a CUDA variant (the CUDA tag is suffixed `-cuda`), and both are multi-arch (`linux/amd64` and `linux/arm64`), so the right one is pulled for your host automatically. They contain just the binary, so mount a converted `.gguf` model (and, for the cli, your audio) at runtime:

```sh
# CLI, CPU
docker run --rm \
-v "$PWD/models:/models:ro" \
-v "$PWD/audio:/audio:ro" \
ghcr.io/mudler/parakeet.cpp-cli:latest \
transcribe --model /models/parakeet-tdt_ctc-110m-q5_k.gguf --input /audio/speech.wav --decoder tdt

# CLI, CUDA (needs the nvidia container toolkit on the host)
docker run --rm --gpus all \
-v "$PWD/models:/models:ro" -v "$PWD/audio:/audio:ro" \
ghcr.io/mudler/parakeet.cpp-cli:latest-cuda \
transcribe --model /models/parakeet-tdt_ctc-110m-q5_k.gguf --input /audio/speech.wav --decoder tdt

# Server: binds 0.0.0.0 and exposes 8080. Fetch a model by alias on first run,
# or mount a local .gguf. Add --gpus all with the :latest-cuda tag for GPU.
docker run --rm -p 8080:8080 ghcr.io/mudler/parakeet.cpp-server:latest --model tdt_ctc-110m
```

The CUDA image is built on CUDA 13, so it covers everything from Turing up through Blackwell, including GB10 / Grace-Blackwell (DGX Spark) on arm64.

To build the images yourself, see the build args at the top of the [`Dockerfile`](../Dockerfile); the cli is the default target and the server is `--target runtime-server`. The CPU image is the portable `GGML_NATIVE=OFF` build, so it runs on any amd64 or arm64 host.
59 changes: 59 additions & 0 deletions docs/performance.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
# Performance summary

This page keeps the headline speed numbers that the README links to. The full
methodology, all ten models, the plots and the raw results are in
[../benchmarks/BENCHMARK.md](../benchmarks/BENCHMARK.md). Read the method before you
quote a number: the CPU runs used 8 threads on a 20-core x86 host, the GPU runs used
one NVIDIA GB10, and each machine was a shared development host, not a quiet lab box.

RTFx is audio seconds divided by processing seconds. Higher is faster. Speedup is our
RTFx divided by the NeMo (PyTorch) RTFx on the same machine, batch size 1.

## CPU against NeMo (LibriSpeech test-clean, 100 utterances, 8 threads)

| dtype | size vs f32 | mean speedup vs NeMo | accuracy |
| ----- | ----------- | -------------------- | -------- |
| f32 | 100% | 1.40x (range 1.11x to 1.69x over 10 models) | mean agreement WER 0.015%, byte-identical on most models |
| f16 | 57% | 1.70x | same as f32 |
| q8_0 | 37% | 1.56x (up to 1.89x) | mean agreement WER 0.16% |
| q4_k | 26% | 1.25x | agreement WER 1.08%, small and monotonic accuracy cost |

Source: the "Quantization" and "Headline" tables in BENCHMARK.md. Peak RAM is roughly
2x lower than NeMo at f32 (for example 2582 MB against 5598 MB for `tdt-0.6b-v3`) and
lower still once quantized.

## GPU against NeMo (NVIDIA GB10, f32, LibriSpeech)

Median speedup over the 10 offline and streaming-EOU models: 1.25x. Best case: 4.3x on
`tdt_ctc-110m`. The smallest gains are on the pure-encoder CTC models (about 1.2x),
because ggml's generic CUDA conv and attention kernels still trail NeMo's tuned cuDNN.
The log-mel front end runs on the GPU through a ggml DFT-matmul graph; the CPU path is
unchanged. NeMo's TDT greedy decode is not CUDA-graph accelerated here and ours is a lean
C++ loop, which explains most of the gap on the TDT and hybrid models.

These numbers come from `benchmarks/results_gpu/` (`scripts/plot_gpu.py` plots them);
the median and maximum above were recomputed from those files when this page was written.
NeMo ran in the `nvcr.io/nvidia/nemo` container for these runs.

## Single clip, newer models

`nemotron-3.5-asr-streaming-0.6b` on a 7.43 s clip (Ryzen 9 9950X3D, 8 threads, median of
7 passes): 2.40x NeMo at f32 and 2.52x at q8_0, byte-identical transcripts. On the GB10 GPU:
1.16x at f32 and 1.30x at q8_0. See the Nemotron section of BENCHMARK.md.

## Against whisper.cpp

The comparison plot [../benchmarks/plots/vs_whisper.png](../benchmarks/plots/vs_whisper.png) shows RTFx and accuracy for parakeet.cpp and whisper.cpp on CPU and GPU (the earlier README text said the 110M Parakeet is faster than whisper base.en and far faster than large-v3-turbo; read the plot for the exact values). The side-by-side clips measure about 12x (GPU) and about 27x (CPU) against whisper.cpp
turbo on one clip with WER 1.6% for both. One clip is a demo, not a benchmark. See
[../benchmarks/media/gpu_whisper_duel.mp4](../benchmarks/media/gpu_whisper_duel.mp4) and
[../benchmarks/media/cpu_duel.mp4](../benchmarks/media/cpu_duel.mp4).

## Apple Metal

On an Apple M4, Metal is about 1.3x to 5.6x faster than CPU, most on the larger models. See
[BENCHMARK.md, Apple Metal](../benchmarks/BENCHMARK.md#apple-metal-m4).

## Batched decode

Up to about 10x to 12x on the GB10 at batch 16 and about 3x to 5x on CPU. See
[batching.md](batching.md).
Loading
Loading