Skip to content

feat: opt-in VAD run gate for the Ultra/Redux heads, and the real-recording VAD evidence - #96

Merged
mudler merged 3 commits into
masterfrom
feat/vad-run-gate
Oct 6, 2026
Merged

mudler merged 3 commits into
masterfrom
feat/vad-run-gate

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

What this adds

An opt-in run gate for the VAD segmenter, and the evidence from a study of Silero, the Ultra head and the Redux head on real recordings.

The gate: a speech run (a stretch of frames with p >= threshold, found before bridging) is dropped when the median of its frame probabilities is below run_gate. A median equal to the gate keeps the run. For an even number of frames the median is the mean of the two middle values. The gate acts first, so the frames of a dropped run are silence for bridging, min_speech, the pauses, the cuts and the trim. It works for any detector and frame size and in both modes (speech and segments). It is meant for the Ultra and Redux heads, whose noise runs sit on a plateau (0.8 to 0.9) while speech sits above 0.99.

It is off by default (run_gate 0). With the gate off the output is byte for byte what it was.

How to use it

  • Options JSON key run_gate, a number in [0, 1), in parakeet_capi_vad_pcm_json, parakeet_capi_vad_path_json and parakeet_capi_transcribe_path_json_vad_with. Other values and unknown keys are errors, like the other keys.
  • CLI: parakeet-cli vad ... --run-gate P and parakeet-cli transcribe ... --vad --vad-run-gate P.
  • Suggested value: 0.92 to 0.96 for the Redux head. Ultra gains less.
  • Offline only. The streaming tracker decides frame by frame and has no run median, so parakeet_capi_vad_stream_begin refuses a non-zero run_gate with a clear error. Audio of at most max_segment seconds is returned whole without the segmenter, so the gate has no effect there (as for the trim).

The numbers

59 recordings with human labels (VoxConverse, AMI far-field, AVA-Speech film clips: 12.1 h of audio, 9.5 h of labelled speech) and 2.6 h without speech (MUSAN music and noise, ESC-50). Recording-disjoint splits, bootstrap intervals.

System VoxConverse F1 AMI F1 AVA F1 Pooled F1 False alarm on non-speech, s/h
Silero, own defaults 96.4 86.7 84.7 90.6 48
Ultra head 96.4 89.7 82.7 90.8 2174
Ultra head, gate 0.92 96.4 85.4 85.4 90.2 826
Redux head 97.3 92.2 85.4 92.8 1685
Redux head, gate 0.92 96.9 90.8 86.2 92.5 269
  • The gate at 0.98 on the Redux head: music 107, noise 1, ESC-50 15 s/h, but it costs 4.2 F1 points. On the Ultra head at 0.98 it costs 7.4 points and music is still 339 s/h.
  • Share of an hour of non-speech that transcribe --vad decodes (trim 0.3): Silero 6 percent of music, 0.03 of noise and 0.03 of ESC-50. Ultra head 95 / 84 / 94, Redux head 91 / 82 / 94. With the gate at 0.92: Ultra 53 / 14 / 27, Redux 27 / 1.1 / 5.6.
  • Word error rate on clean TED talks: every system is within 0.17 points, gate or not. On three talks with 40 s of music, 30 s of noise and 40 s of vocal music inserted at pauses: Ultra head 5.91, gate 4.24, Silero 3.87. Redux head 6.94, gate 5.22, Silero 4.76. The gate recovers about 80 percent of the head's penalty.
  • Recommendation recorded in the docs: Silero with its own defaults is the always-on gate (lower its threshold to 0.2 to 0.3 for recall). The Redux head suits long recordings that are mostly speech (best default F1 92.8, least speech lost, same WER as Silero on clean talks). Silero is better for recordings with long music or noise. The gate is an opt-in option for the heads.

What the gate does not do

  • Music still triggers the head. At 0.92 the Redux head still calls 644 s of an hour of music speech and the Ultra head 1625 s. Music runs have a high median like speech.
  • It is not a noise rejector. Some noise false alarms remain (Redux 23 s/h of noise and 131 s/h of ESC-50, Ultra 398 and 395), against about 1 s/h for Silero.
  • It costs some speech: on AMI the F1 falls by 1.4 (Redux) and 4.3 (Ultra) points, and in segments mode on AVA the Redux gate loses 7.8 percent of the labelled speech (the head alone loses 0.2). The pooled cost at 0.92 is 0.3 and 0.6 points, and neither interval excludes zero.

What was not built, and why

  • Fusion of Silero and a head (--vad-fusion). The synthetic study gained +0.75 and +1.13 F1 points. On real recordings the tuned fusion gained +0.10 (Ultra) and +0.46 (Redux) over the best tuned single detector, and gave no WER gain. It needs both models to run. Not worth a flag. The docs keep the old section as an experiment note and say so.
  • A change to the 30 s hard cut. After the trim it is not the main cost: a forced mid-speech cut costs about 0.18 to 0.2 word errors, and there are 10 to 35 of them per hour. Cutting at every pause of 2 s or more lowers the decoded non-speech by 22 to 59 percent but loses more speech, so the cut rule is unchanged.

Tests

  • New test_vad_run_gate (no model): synthetic probability streams with exact expected runs at 0.08 s and 0.032 s frames. Covers runs above and below the gate, a low median under a high peak, bridged gaps (the raw run, not the bridged one, is gated), a one-frame run, the boundary (median equal to the gate keeps), gate then trim, the early return for short audio, random streams (a gate never adds speech, a higher gate never keeps more), NaN as a degenerate option, and that the event tracker ignores the gate. Gate 0 equals the old output (the existing digest test also covers it).
  • test_vad_options: the run_gate key, ranges, errors, and that the word-filter parser rejects it. CLI flag errors for --run-gate and --vad-run-gate as ctest cases. test_capi_vad_silero: a stream refuses a non-zero gate.
  • New test_vad_run_gate_model (label model): two_speakers.wav, seeded white or pink noise, speech.wav. For the Redux slice, packed and dequantized Redux, Ultra and Silero: gate 0 gives the same JSON as no key, in both modes; for Redux at 0.92 at most 25 percent of the seconds the head called speech in the noise stretch remain (none did in the runs here) and the speech stretches are unchanged; the transcribe path honours the key.
  • Run locally on CPU: ctest -LE model passes (52 tests); the model tests pass with the Ultra, Redux (packed and dequantized), Silero and TDT 0.6B files, apart from test_relpos_attention_local_chunked and test_capi_timestamps, which are known failures that this change does not touch (test_combined_offline was skipped, its inputs were not available). parakeet-cli vad output with the gate off is byte-identical to master on 13 files, for the Redux head slice and for Silero, in both modes and with --probabilities. transcribe --vad --json is byte-identical on 4 files for Ultra, Redux and Silero + TDT. The C++ gate gives the same speech regions as the Python gate of the study on 688 combinations of recording, detector and gate value on the stored probabilities.

Limits

DIHARD, CallHome and the MUSAN speech files were not used (their speech labels were unusable). WER was measured on TED talks only. The composite recordings are synthetic. Cut policies were scored by non-speech seconds, not WER. The Silero ONNX model was not run. Only 20 recordings are held out (3 AMI meetings), so the held-out intervals are wide. In VoxConverse and AMI the frames Silero "misses" are mostly pauses inside annotated turns, so tuning pushes the pause to 0.5 s and the pad to 0.1 s, which is a labelling convention; a 0.1 s collar raises every F1 by about 0.7 to 0.9 points and keeps the ranking.

The scripts and result tables of the study are in scripts/vad_bench/real_recordings/ (no audio, probabilities or models).

🤖 Generated with Claude Code

mudler added 3 commits October 6, 2026 07:39
The Ultra and Redux heads call steady noise speech, but a noise run
sits on a plateau while a speech run sits high. Add
SegmenterOpts::run_gate (default 0 = off): a speech run, a stretch of
frames with p >= threshold found before bridging, is dropped when the
median of its frame probabilities is below the gate. A median equal to
the gate keeps the run. The gate acts first, so dropped frames are
silence for bridging, min_speech, pauses and trim. It works for any
detector and frame size, in both modes. The streaming event tracker has
no run median and ignores it.

Expose it as the "run_gate" key (a number in [0, 1)) of the VAD options
JSON, so the vad functions and the transcribe functions that take VAD
options accept it, and as --run-gate on `parakeet-cli vad` and
--vad-run-gate on `transcribe --vad`. A VAD stream refuses a non-zero
value instead of ignoring it. With the gate off the output is unchanged
byte for byte.

Tests use synthetic probability streams at 0.08 s and 0.032 s frames
with exact expected runs, the boundary, bridged gaps, a low median under
a high peak, the trim, the option parser and CLI errors. A model test
runs the gate on a clip with a seeded noise stretch for the head slice,
Ultra, Redux and Silero.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Describe the run gate in docs/vad.md: the definition of the run median,
how to use it, what it costs and what it does not do (music still
triggers a head, and it is not a noise rejector).

Add a Real recordings section to docs/vad-benchmarks.md: 59 recordings
with human labels and 2.6 h without speech. It reports frame F1 untuned
and tuned, the reference noise and a collar test, false alarms on audio
without speech, the audio the decoder receives, word error rate on clean
talks and on composites with inserted music and noise, the cost of the
30 s hard cut, and the limits.

Update the guidance. Silero stays the always-on gate, the Redux head
suits long recordings that are mostly speech, and the gate is opt-in.
The fusion gain of the synthetic experiment did not hold on real
recordings (+0.10 and +0.46 F1 points, no word error rate gain), so
fusion is not built.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Add the scripts of the study (data fetch, probability dump, a Python
replica of the segmenter, frame and decoder scoring, WER reports) and the
small result tables, with a README on how to rerun them. Paths are
environment variables. verify_gate.py checks the C++ run gate against the
Python gate on stored probabilities. No audio, probabilities or models.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 9a28a3c into master Oct 6, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants