Skip to content

feat: trim VAD segments to speech (default 0.3 s) and an opt-in local-confidence word filter - #94

Merged
mudler merged 4 commits into
masterfrom
feat/vad-trim-and-word-filter
Oct 5, 2026
Merged

mudler merged 4 commits into
masterfrom
feat/vad-trim-and-word-filter

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

Why

transcribe --vad cuts long audio at pauses, or at the 30 s cap when there is no pause. Each piece used to carry everything between its two cut points to the decoder, including noise the VAD had already called non-speech. That is where an ASR model invents words. Two things follow from this and from a study of decoder-side guards (not part of this PR):

  • A guard inside the decoder cannot save time. The cheap fix is to stop sending the noise: trim each piece to its speech.
  • A confidence filter on the words the decoder returns removes invented words at no cost in decoding, but only if it looks at the neighbours of a word, not at the word alone (a low-confidence real word between confident ones must stay).

What changed

  1. Segment trim, on by default. SegmenterOpts::trim_sec (default 0.3). Each kept segment of long audio shrinks to its first speech frame minus 0.3 s and its last speech frame plus 0.3 s, never past the cut itself. Segments without speech are still dropped as before, and a segment with speech is never trimmed to nothing. Audio of 30 s or less is not cut and not trimmed. The segmenter is shared, so this applies to the Ultra/Redux head, to Silero and to VAD-only slices. Word and token times are still relative to the whole file. trim 0 gives the previous output byte for byte.
  2. Word filter, off by default. pk::WordFilter / apply_word_filter: a word is dropped when the mean confidence of the words that start within local_radius (default 5 s) of it, itself included, in the same decode unit (the whole clip, or one VAD segment) is below min_local_conf (0 = off, 0.5 suggested). drop_punct_only (default off, meant for CTC models) also drops words that have no letter or digit. It is post-processing of the existing per-word confidence for TDT, RNN-T and CTC. Dropped words leave text, words and tokens. When nothing is dropped, or the filter is off, the result is unchanged.
  3. API and CLI.
    • --vad-trim SEC (transcribe), --trim SEC (vad); trim key in the VAD options JSON (segments mode of parakeet_capi_vad_*, and parakeet_capi_transcribe_path_json_vad_with).
    • --min-local-conf F, --local-radius SEC, --drop-punct-only, with or without --vad.
    • New parakeet_capi_transcribe_path_json_with(ctx, wav, decoder, options_json) for a plain transcribe with the filter keys (min_local_conf, local_radius, drop_punct_only). The same keys work in ..._json_vad_with. Additive, ABI stays 10.
    • When a filter ran, the JSON gets "guard":{"dropped_words":N} (0 when nothing was dropped). Without a filter the document is the same as before.

Behaviour change: the trim default

This is a deliberate change of default for --vad, parakeet_capi_transcribe_path_json_vad* and the segments mode of the VAD functions. Transcripts of long audio through the VAD paths can shift slightly. Release-note worthy. --vad-trim 0 (or "trim":0) restores the old cuts.

Measured (details and limits in docs/vad-benchmarks.md, scripts and output in scripts/vad_bench/decoder_guards; old = master before this change, trim = new default):

WER % Detector Old Trim 0.3
3 held-out TED talks (5627 words) Ultra head 3.68 3.68
Redux head 4.39 4.51
v3 + Silero 3.45 3.45
Speech in noise, 4 sets x 6 utterances: clean / white 5 dB / pink 0 dB (429 words each) Ultra 2.10 / 4.90 / 6.53 1.86 / 4.43 / 5.59
Redux 3.26 / 7.23 / 12.35 3.03 / 6.76 / 13.29
v3 + Silero 1.86 / 5.59 / 7.93 1.86 / 5.13 / 7.69
60 s synthetic noise block inside 90 s of speech (30 files) Ultra Redux v3 + Silero
Seconds of the block decoded, old -> trim 33.8 -> 12.0 29.7 -> 0.3 27.7 -> 0.1
Invented words in the block, old -> trim 0 -> 0 0 -> 0 11 -> 0

On talks the cost is 0 to 0.12 points; on the small speech-in-noise sets the change is neutral (one word is 0.23 points). Loud white and pink noise (-5 dB) still passes the Ultra head as speech, and the trim cannot remove it (52 to 58 s of those blocks are still decoded).

The word filter at 0.5

Ultra v3 + Silero
WER on the 3 talks, trim / trim + filter 3.68 / 3.68 3.45 / 3.45
Real words removed at 0.5 over all files with speech 0 of 15375 3 of 15406
Invented words in the 5 blocks that had some (old cuts, filter alone) no events 11 -> 1 (0 of 1484 other words lost)
Invented words in 63 noise-only files of 30 s 0 1 -> 0
WER on speech in noise (1287 words), filter off / 0.5 / 0.7 / 0.9 3.96 / 3.96 / 3.96 / 6.06 4.90 / 4.90 / 4.90 / 16.08

0.5 is the suggested value because it removed invented words and no real ones in these runs; at 0.9 the filter costs many real words, most on v3 (164 words dropped). CTC models drop real words at higher thresholds too and lone punctuation is what drop_punct_only is for; see limits.

Tests

  • New test_word_filter (scripted words and confidences: off is identical, an isolated low word between confident words is kept, a run of low words is dropped, radius and boundaries, punctuation-only, token renumbering, group_words token ranges).
  • test_vad_segmenter: exact segment bounds for 0.08 s and 0.032 s frames (speech at the end of a window, at both edges, none, two runs), and a digest of the previous implementation on seeded random streams that trim 0 must equal.
  • test_vad_options: trim and the filter keys, errors, parse_filter_options; ctest cases for the CLI flag errors.
  • Model test test_vad_trim_filter (Ultra, both Redux files, v3 + Silero, a CTC 0.6B model): trimmed offsets, trim 0 against the old cuts (also in test_vad_batched), filter off identical (Model call and C-API JSON byte for byte), filter on, errors.
  • ctest -LE model: 42 tests, all passed (4 skipped for missing data). ctest -L model with Ultra, Redux keep and deq, Silero, v3 and CTC 0.6B: 86 tests; the only failures were test_relpos_attention_local_chunked and test_capi_timestamps (allowed failures) and, in the first run, test_vad_trim_filter (a test bug: the Redux head fires on digital silence; fixed, and it passes). --vad-trim 0 equals the old binary byte for byte on 45 files (15 per detector).

Limits

  • The noise is synthetic and the events are few: 11 invented words in 5 files, all from v3. Ultra and Redux gave no invented words on this noise, so the filter is not shown to remove any on them here (an earlier, uncommitted study saw it do so on Ultra, Redux and RNN-T).
  • CTC was not run on a 0.6B model beyond one fixture in a unit test; drop_punct_only rests on earlier work with a 110M CTC head.
  • A single lone invented word in real speech, the two-stage VAD (Silero then head), GPU backends and quantized files other than F16 were not tested.
  • Not measured on a quiet machine (load average 5 to 50). WER and word counts do not depend on load; no timing is claimed.

🤖 Generated with Claude Code

mudler added 4 commits October 4, 2026 23:56
segment_by_vad cuts long audio at pauses, or at the 30 s cap when there
is no pause. A cut piece carried everything between its cut points to
the decoder, including long stretches of noise and silence that the
VAD had already flagged as non-speech. That is where an ASR model
invents words.

Add SegmenterOpts::trim_sec (default 0.3). Each kept segment shrinks to
its first speech frame minus trim_sec and its last speech frame plus
trim_sec, limited to the segment itself. Speech is the smoothed mask, so
the existing rules still decide what counts: a piece without speech is
still dropped, and a piece with speech keeps at least that run. Audio of
at most max_seg_sec is not cut and not trimmed. Trim 0 gives the
previous output exactly.

The change applies to every caller of the segmenter: the Ultra and Redux
head, Silero, and the VAD-only slices. Transcripts of long audio
through the VAD paths can shift slightly.

Tests use synthetic probability streams at 0.08 s and 0.032 s frames
with exact expected bounds, and a digest of the previous implementation
over seeded random streams to check that trim 0 is unchanged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
An ASR model run on noise can emit words that were never said. Their
per-word confidence is low, but so is the confidence of some real words,
so a cut at one confidence value costs real words. The mean confidence
of the neighbouring words separates the two much better: a doubtful
word between confident ones stays, and a word that stands alone or among
other doubtful words goes.

Add pk::WordFilter and pk::apply_word_filter. A word is dropped when the
mean confidence of the words that start within local_radius_sec (default
5 s) of it, the word included, is below min_local_conf. With
drop_punct_only a word that is only punctuation is dropped too, for CTC
models that emit a lone mark on noise. The filter works on one decode
unit, removes the words from text, words and tokens, and leaves the
transcription untouched when nothing is dropped or when it is off, which
is the default. The count goes in Transcription::dropped_words (-1 when
no filter ran). Words now record the index of their first and last token
so the tokens can be removed with them.

The filter is not wired to any entry point yet.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…e CLI

Segment trim: a "trim" key (seconds >= 0, 0 = off, default 0.3) in the VAD
options JSON, so it applies to the "segments" mode of the VAD functions
and to parakeet_capi_transcribe_path_json_vad_with. CLI: --vad-trim on
transcribe and --trim on the vad command.

Word filter, off by default:
- Model::transcribe_pcm_vad and transcribe_pcm_vad_with_timestamps take
  an optional WordFilter that runs on each decode unit, the whole clip or
  one VAD segment, before the segment offsets are added.
- New parakeet_capi_transcribe_path_json_with(ctx, wav, decoder, options)
  for a plain transcribe with the filter keys min_local_conf,
  local_radius and drop_punct_only. Additive, ABI stays 10. The same
  keys work in parakeet_capi_transcribe_path_json_vad_with.
- CLI: --min-local-conf, --local-radius, --drop-punct-only, with or
  without --vad.
- When a filter ran, the JSON document has "guard":{"dropped_words":N}.
  Without a filter the document is the same as before, byte for byte.

Tests: option parsing and errors (test_vad_options, ctest cases for the
CLI flags), and model tests (test_vad_trim_filter) for the trimmed
offsets, trim 0 against the old cuts, the filter off being identical,
and the filter on, with the VAD head, Silero and a CTC model.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Describe trim and the opt-in word filter in docs/vad.md and the README.
Add a section to docs/vad-benchmarks.md with the word error rates, the
seconds of noise decoded and the invented words before and after the
trim, and the effect of the filter at 0.5, 0.7 and 0.9. The scripts that
fetch the public data, build the noise files with fixed seeds, run the
CLI and print the tables are in scripts/vad_bench/decoder_guards, with
the printed tables. No audio or model is committed.

The trim default changes the transcripts of long audio through the VAD
paths slightly, so the docs say so.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 2de154c into master Oct 5, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants