Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -187,6 +187,8 @@ parakeet-cli transcribe ... --beam-size 4 --nbest 4 # TDT N-best, se
parakeet-cli transcribe ... --stream # cache-aware streaming (EOU and Nemotron models)
parakeet-cli transcribe ... --lang <locale> # Nemotron 3.5 language, default auto
parakeet-cli transcribe ... --vad [--vad-model silero.gguf] # cut long audio at pauses (offline, greedy only)
parakeet-cli transcribe ... --vad-trim SEC # trim each piece to its speech plus SEC (default 0.3, 0 = whole cuts)
parakeet-cli transcribe ... --min-local-conf 0.5 # opt-in: drop words invented on noise (docs/vad.md)
parakeet-cli vad --model M --input A.wav [--mode segments] [--probabilities] # speech regions as JSON
parakeet-cli scene --model ASR --diar DIAR --sound CED --input A.wav # words + speakers + sounds
parakeet-cli scene ... --speakers SPK.gguf --registry R # name the speakers
Expand Down
84 changes: 84 additions & 0 deletions docs/vad-benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -925,6 +925,89 @@ compared with the model runs, but it was not timed separately.
Scripts, the small result files and how to reproduce:
[`scripts/vad_bench/fusion/`](../scripts/vad_bench/fusion/README.md).

## Trimming segments and the word filter

Two changes to what `transcribe --vad` hands to the decoder. Scripts and the full output are in
[scripts/vad_bench/decoder_guards](../scripts/vad_bench/decoder_guards) (`results/tables.txt`). The
options are described in [vad.md](vad.md).

The run was not made on a quiet machine: the load average was between 5 and 50 (see the first line
of `results/tables.txt`). Word error rates and word counts do not depend on load; no timing is
claimed.

### Setup

- Old: `parakeet-cli` before the change (the cuts of the previous segmenter). New: the default,
each segment trimmed to its speech plus 0.3 s. `--vad-trim 0` gave the old output byte for byte on
all 45 files that were run both ways (15 per detector).
- Detectors: the Ultra head, the Redux head (dequantized F16), and TDT 0.6B v3 with Silero.
The model was Ultra F16 for the head runs.
- Talks: three whole TED-LIUM long-form talks that were not used to choose any setting (5627
reference words). Words are lower-cased, punctuation removed.
- Speech in noise: 4 sets of 6 LibriSpeech test-clean utterances with gaps (about 50 s each, so the
segmenter cuts them), clean, with white noise at 5 dB SNR, or with pink noise at 0 dB SNR. 429
reference words per condition.
- Noise block: 90 s of a talk with 60 s of synthetic noise (white, pink, clicks, music-like tones) put
in at a quiet point, 30 files. Any word the decoder returns inside the block is an invented word.
"Seconds decoded" is the overlap of the cuts with the block, from `vad --mode segments`.
- Noise alone: 63 files of 30 s (seven noise types, three levels), decoded whole.

### Trimming: word error rate

| Set | Detector | Old | Trim 0.3 | Change |
| --- | --- | ---: | ---: | ---: |
| 3 talks | Ultra head | 3.68 | 3.68 | +0.00 |
| 3 talks | Redux head | 4.39 | 4.51 | +0.12 |
| 3 talks | v3 + Silero | 3.45 | 3.45 | +0.00 |
| clean | Ultra / Redux / v3 + Silero | 2.10 / 3.26 / 1.86 | 1.86 / 3.03 / 1.86 | -0.23 / -0.23 / +0.00 |
| white noise 5 dB | Ultra / Redux / v3 + Silero | 4.90 / 7.23 / 5.59 | 4.43 / 6.76 / 5.13 | -0.47 / -0.47 / -0.47 |
| pink noise 0 dB | Ultra / Redux / v3 + Silero | 6.53 / 12.35 / 7.93 | 5.59 / 13.29 / 7.69 | -0.93 / +0.93 / -0.23 |

On talks the change is within 0.12 points; the Redux head loses a little on two of the three talks
and on pink noise. The speech-in-noise sets are small (one word is 0.23 points), so read them as
"neutral", not as a gain.

### Trimming: the noise block

| Detector | Seconds of the 60 s block decoded, old | New | Invented words, old | New |
| --- | ---: | ---: | ---: | ---: |
| Ultra head | 33.8 | 12.0 | 0 | 0 |
| Redux head | 29.7 | 0.3 | 0 | 0 |
| v3 + Silero | 27.7 | 0.1 | 11 (5 files) | 0 |

The seconds are means over the 30 files. The head still passes loud white and pink noise (-5 dB
against the speech) as speech, and then the trim cannot remove it: Ultra decodes 52 to 58 s of
those blocks. The Redux head also fires on digital silence. The words the decoder invented were
rare in this run (Ultra and Redux gave none, in the block or in 63 noise-only files), so the
trim's benefit here is mostly seconds saved, and for v3 the 11 invented words.

### Word filter at 0.5

| Check | Ultra head | v3 + Silero |
| --- | ---: | ---: |
| WER on the 3 talks, trim 0.3 / plus filter | 3.68 / 3.68 | 3.45 / 3.45 |
| WER on speech in noise (12 files, 1287 words): off / 0.5 / 0.7 / 0.9 | 3.96 / 3.96 / 3.96 / 6.06 | 4.90 / 4.90 / 4.90 / 16.08 |
| Words dropped at 0.9 on those files | 31 | 164 |
| Words dropped at 0.5 on all files with real speech | 0 of 15375 | 3 of 15406 |
| Invented words in the 5 blocks that had them (old cuts, filter alone) | no events | 11 -> 1, 0 of 1484 other words lost |
| Invented words in the 63 noise-only files | 0 | 1 -> 0 (v3) |

At 0.5 the filter is free on these sets and it removes most of the few invented words there were.
At 0.9 it costs real words, most on v3. Hence 0.5 is the suggested value.

### Limits

- The noise is synthetic. The event counts are small: 11 invented words in 5 files, all from v3. A
filter that removes 10 of 11 is a weak estimate of a rate.
- The filter was not run against invented words on Ultra, Redux or an RNN-T model (none occurred), and
not at all on a CTC model of 0.6B (a unit test runs it on one fixture). The `drop_punct_only`
option for CTC rests on the 110M CTC head, not measured here.
- A single lone invented word between real speech was not tested.
- Silero followed by the head (the two-stage rule), GPU backends and the quantized files other than the
F16 ones were not tested.
- The trim changes transcripts of long audio through the VAD paths slightly; the numbers above are
from three talks and 12 clips per condition.

## When to use which

This is limited to what the numbers above support.
Expand Down Expand Up @@ -977,6 +1060,7 @@ are committed. In short:
`silero_collect.py`, `silero_analyze.py`; speed with `speed_silero.py`.
5. Long talks: `longform_b1.sh`.
6. Noise root-cause study: [`noise_dive/`](../scripts/vad_bench/noise_dive/README.md).
7. Segment trim and word filter: [`decoder_guards/`](../scripts/vad_bench/decoder_guards/README.md).

Small result files of the runs on this page (tables, per run timings and load logs)
are in `scripts/vad_bench/results/`. The raw per clip predictions are not committed;
Expand Down
63 changes: 61 additions & 2 deletions docs/vad.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,6 +98,7 @@ the same keys). Unknown keys and out of range values are errors.
| `min_speech` | seconds; shorter speech runs are dropped | 0.1 | 0.25 |
| `speech_pad` | seconds >= 0, `speech` mode: widen each region on both sides | 0 | 0.03 |
| `max_segment` | seconds; cap in `segments` mode | 30 | 30 |
| `trim` | seconds >= 0; `segments` mode and the transcribe functions: shrink each cut to its speech plus this much on each side, 0 = keep the whole cut | 0.3 | 0.3 |
| `mode` | `speech` or `segments` | `speech` | `speech` |
| `probabilities` | add the per frame probabilities | false | false |

Expand All @@ -106,7 +107,64 @@ shorter than 0.1 s are bridged, runs shorter than `min_speech` are dropped,
regions closer than `min_pause` merge, then each region is padded (and two
regions that would overlap meet in the middle of the gap). `segments` is the cut
that `transcribe --vad` decodes: pieces of at most `max_segment` seconds cut at
pauses, pieces without speech dropped, audio within the cap returned whole.
pauses, pieces without speech dropped, audio within the cap returned whole. Each
kept piece is then trimmed (see below).

### Trimming the segments

A piece cut from long audio used to carry everything between its two cut points
to the decoder, including the noise and silence the VAD had already flagged as
non-speech. That is where an ASR model invents words. Since this change each
piece shrinks to its first speech frame minus `trim` and its last speech frame
plus `trim` (default 0.3 s, never past the cut itself). Speech is the smoothed
mask, so the rules above still decide what counts as speech. Audio within the
cap is not cut and not trimmed. Word and token times are still relative to the
whole file. `trim` 0 (`--vad-trim 0`) gives the previous cuts exactly.

This is a change of default behaviour for `transcribe --vad`,
`parakeet_capi_transcribe_path_json_vad*` and the `segments` mode of the VAD
functions, for the head, for Silero and for VAD-only slices (they share the
segmenter). On talks, transcripts of long audio can shift slightly (a word WER
cost of about 0.1 point in our runs); on audio with long noisy stretches the
decoder sees much less noise. Numbers: [vad-benchmarks.md](vad-benchmarks.md#trimming-segments-and-the-word-filter).

## Word filter (opt-in)

A confidence filter can remove the words a model invents on noise. It is
post-processing of the decode: the model, the VAD and the cuts are unchanged, and
it is off by default (output is then byte for byte the same).

A word is dropped when the mean confidence of the words that start within
`local_radius` seconds of it (the word itself included, the same decode unit) is
below `min_local_conf`. A low confidence word between confident ones keeps a high
mean and stays; a word that stands alone, or among other low confidence words,
goes. A decode unit is the whole clip, or one VAD segment. `drop_punct_only` also
removes words that consist only of punctuation.

| Option | Meaning | Default |
| --- | --- | --- |
| `min_local_conf` | 0 to 1; 0 = off. 0.5 is the suggested value | 0 |
| `local_radius` | seconds, both sides | 5 |
| `drop_punct_only` | drop words with no letter or digit; recommended for CTC models | false |

```
parakeet-cli transcribe --model m.gguf --input a.wav --min-local-conf 0.5 \
[--local-radius SEC] [--drop-punct-only] [--json] [--vad ...]

char* parakeet_capi_transcribe_path_json_with(parakeet_ctx* ctx, const char* wav_path,
int decoder, const char* options_json);
```

The options JSON of `parakeet_capi_transcribe_path_json_with` takes the three
keys above. `parakeet_capi_transcribe_path_json_vad_with` takes them too, next to
the VAD keys (`trim` included). With a filter on, the JSON document gets one more
member, `"guard":{"dropped_words":N}` (N is 0 when nothing was dropped), and the
dropped words are also removed from `text`, `words` and `tokens`.

Limits: 0.5 removes hallucinated words on Ultra, Redux and RNN-T models at no
cost in WER on clean speech. On v3 and on CTC models, and at higher thresholds,
it also removes real words (see the benchmark page). It does not save time: the
decoder still runs on everything it is given.

### Why the Silero defaults differ

Expand Down Expand Up @@ -167,7 +225,8 @@ audio with Silero:

```
parakeet-cli transcribe --model tdt-0.6b-v3.gguf --input long.wav --vad --vad-model silero.gguf \
[--vad-threshold F] [--vad-min-pause SEC] [--vad-min-speech SEC] [--vad-max-seg SEC]
[--vad-threshold F] [--vad-min-pause SEC] [--vad-min-speech SEC] [--vad-max-seg SEC] \
[--vad-trim SEC]

char* parakeet_capi_transcribe_path_json_vad_with(parakeet_ctx* asr, parakeet_ctx* silero,
const char* wav_path, int decoder,
Expand Down
Loading
Loading