Repository navigation
feat: trim VAD segments to speech (default 0.3 s) and an opt-in local-confidence word filter - #94
Merged
Merged
Conversation
segment_by_vad cuts long audio at pauses, or at the 30 s cap when there is no pause. A cut piece carried everything between its cut points to the decoder, including long stretches of noise and silence that the VAD had already flagged as non-speech. That is where an ASR model invents words. Add SegmenterOpts::trim_sec (default 0.3). Each kept segment shrinks to its first speech frame minus trim_sec and its last speech frame plus trim_sec, limited to the segment itself. Speech is the smoothed mask, so the existing rules still decide what counts: a piece without speech is still dropped, and a piece with speech keeps at least that run. Audio of at most max_seg_sec is not cut and not trimmed. Trim 0 gives the previous output exactly. The change applies to every caller of the segmenter: the Ultra and Redux head, Silero, and the VAD-only slices. Transcripts of long audio through the VAD paths can shift slightly. Tests use synthetic probability streams at 0.08 s and 0.032 s frames with exact expected bounds, and a digest of the previous implementation over seeded random streams to check that trim 0 is unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
An ASR model run on noise can emit words that were never said. Their per-word confidence is low, but so is the confidence of some real words, so a cut at one confidence value costs real words. The mean confidence of the neighbouring words separates the two much better: a doubtful word between confident ones stays, and a word that stands alone or among other doubtful words goes. Add pk::WordFilter and pk::apply_word_filter. A word is dropped when the mean confidence of the words that start within local_radius_sec (default 5 s) of it, the word included, is below min_local_conf. With drop_punct_only a word that is only punctuation is dropped too, for CTC models that emit a lone mark on noise. The filter works on one decode unit, removes the words from text, words and tokens, and leaves the transcription untouched when nothing is dropped or when it is off, which is the default. The count goes in Transcription::dropped_words (-1 when no filter ran). Words now record the index of their first and last token so the tokens can be removed with them. The filter is not wired to any entry point yet. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…e CLI
Segment trim: a "trim" key (seconds >= 0, 0 = off, default 0.3) in the VAD
options JSON, so it applies to the "segments" mode of the VAD functions
and to parakeet_capi_transcribe_path_json_vad_with. CLI: --vad-trim on
transcribe and --trim on the vad command.
Word filter, off by default:
- Model::transcribe_pcm_vad and transcribe_pcm_vad_with_timestamps take
an optional WordFilter that runs on each decode unit, the whole clip or
one VAD segment, before the segment offsets are added.
- New parakeet_capi_transcribe_path_json_with(ctx, wav, decoder, options)
for a plain transcribe with the filter keys min_local_conf,
local_radius and drop_punct_only. Additive, ABI stays 10. The same
keys work in parakeet_capi_transcribe_path_json_vad_with.
- CLI: --min-local-conf, --local-radius, --drop-punct-only, with or
without --vad.
- When a filter ran, the JSON document has "guard":{"dropped_words":N}.
Without a filter the document is the same as before, byte for byte.
Tests: option parsing and errors (test_vad_options, ctest cases for the
CLI flags), and model tests (test_vad_trim_filter) for the trimmed
offsets, trim 0 against the old cuts, the filter off being identical,
and the filter on, with the VAD head, Silero and a CTC model.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Describe trim and the opt-in word filter in docs/vad.md and the README. Add a section to docs/vad-benchmarks.md with the word error rates, the seconds of noise decoded and the invented words before and after the trim, and the effect of the filter at 0.5, 0.7 and 0.9. The scripts that fetch the public data, build the noise files with fixed seeds, run the CLI and print the tables are in scripts/vad_bench/decoder_guards, with the printed tables. No audio or model is committed. The trim default changes the transcripts of long audio through the VAD paths slightly, so the docs say so. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
transcribe --vadcuts long audio at pauses, or at the 30 s cap when there is no pause. Each piece used to carry everything between its two cut points to the decoder, including noise the VAD had already called non-speech. That is where an ASR model invents words. Two things follow from this and from a study of decoder-side guards (not part of this PR):What changed
SegmenterOpts::trim_sec(default 0.3). Each kept segment of long audio shrinks to its first speech frame minus 0.3 s and its last speech frame plus 0.3 s, never past the cut itself. Segments without speech are still dropped as before, and a segment with speech is never trimmed to nothing. Audio of 30 s or less is not cut and not trimmed. The segmenter is shared, so this applies to the Ultra/Redux head, to Silero and to VAD-only slices. Word and token times are still relative to the whole file.trim0 gives the previous output byte for byte.pk::WordFilter/apply_word_filter: a word is dropped when the mean confidence of the words that start withinlocal_radius(default 5 s) of it, itself included, in the same decode unit (the whole clip, or one VAD segment) is belowmin_local_conf(0 = off, 0.5 suggested).drop_punct_only(default off, meant for CTC models) also drops words that have no letter or digit. It is post-processing of the existing per-word confidence for TDT, RNN-T and CTC. Dropped words leavetext,wordsandtokens. When nothing is dropped, or the filter is off, the result is unchanged.--vad-trim SEC(transcribe),--trim SEC(vad);trimkey in the VAD options JSON (segmentsmode ofparakeet_capi_vad_*, andparakeet_capi_transcribe_path_json_vad_with).--min-local-conf F,--local-radius SEC,--drop-punct-only, with or without--vad.parakeet_capi_transcribe_path_json_with(ctx, wav, decoder, options_json)for a plain transcribe with the filter keys (min_local_conf,local_radius,drop_punct_only). The same keys work in..._json_vad_with. Additive, ABI stays 10."guard":{"dropped_words":N}(0 when nothing was dropped). Without a filter the document is the same as before.Behaviour change: the trim default
This is a deliberate change of default for
--vad,parakeet_capi_transcribe_path_json_vad*and thesegmentsmode of the VAD functions. Transcripts of long audio through the VAD paths can shift slightly. Release-note worthy.--vad-trim 0(or"trim":0) restores the old cuts.Measured (details and limits in
docs/vad-benchmarks.md, scripts and output inscripts/vad_bench/decoder_guards; old = master before this change, trim = new default):On talks the cost is 0 to 0.12 points; on the small speech-in-noise sets the change is neutral (one word is 0.23 points). Loud white and pink noise (-5 dB) still passes the Ultra head as speech, and the trim cannot remove it (52 to 58 s of those blocks are still decoded).
The word filter at 0.5
0.5 is the suggested value because it removed invented words and no real ones in these runs; at 0.9 the filter costs many real words, most on v3 (164 words dropped). CTC models drop real words at higher thresholds too and lone punctuation is what
drop_punct_onlyis for; see limits.Tests
test_word_filter(scripted words and confidences: off is identical, an isolated low word between confident words is kept, a run of low words is dropped, radius and boundaries, punctuation-only, token renumbering,group_wordstoken ranges).test_vad_segmenter: exact segment bounds for 0.08 s and 0.032 s frames (speech at the end of a window, at both edges, none, two runs), and a digest of the previous implementation on seeded random streams that trim 0 must equal.test_vad_options:trimand the filter keys, errors,parse_filter_options; ctest cases for the CLI flag errors.test_vad_trim_filter(Ultra, both Redux files, v3 + Silero, a CTC 0.6B model): trimmed offsets, trim 0 against the old cuts (also intest_vad_batched), filter off identical (Model call and C-API JSON byte for byte), filter on, errors.ctest -LE model: 42 tests, all passed (4 skipped for missing data).ctest -L modelwith Ultra, Redux keep and deq, Silero, v3 and CTC 0.6B: 86 tests; the only failures weretest_relpos_attention_local_chunkedandtest_capi_timestamps(allowed failures) and, in the first run,test_vad_trim_filter(a test bug: the Redux head fires on digital silence; fixed, and it passes).--vad-trim 0equals the old binary byte for byte on 45 files (15 per detector).Limits
drop_punct_onlyrests on earlier work with a 110M CTC head.🤖 Generated with Claude Code