Repository navigation
feat: opt-in VAD run gate for the Ultra/Redux heads, and the real-recording VAD evidence - #96
Merged
Merged
Conversation
The Ultra and Redux heads call steady noise speech, but a noise run sits on a plateau while a speech run sits high. Add SegmenterOpts::run_gate (default 0 = off): a speech run, a stretch of frames with p >= threshold found before bridging, is dropped when the median of its frame probabilities is below the gate. A median equal to the gate keeps the run. The gate acts first, so dropped frames are silence for bridging, min_speech, pauses and trim. It works for any detector and frame size, in both modes. The streaming event tracker has no run median and ignores it. Expose it as the "run_gate" key (a number in [0, 1)) of the VAD options JSON, so the vad functions and the transcribe functions that take VAD options accept it, and as --run-gate on `parakeet-cli vad` and --vad-run-gate on `transcribe --vad`. A VAD stream refuses a non-zero value instead of ignoring it. With the gate off the output is unchanged byte for byte. Tests use synthetic probability streams at 0.08 s and 0.032 s frames with exact expected runs, the boundary, bridged gaps, a low median under a high peak, the trim, the option parser and CLI errors. A model test runs the gate on a clip with a seeded noise stretch for the head slice, Ultra, Redux and Silero. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Describe the run gate in docs/vad.md: the definition of the run median, how to use it, what it costs and what it does not do (music still triggers a head, and it is not a noise rejector). Add a Real recordings section to docs/vad-benchmarks.md: 59 recordings with human labels and 2.6 h without speech. It reports frame F1 untuned and tuned, the reference noise and a collar test, false alarms on audio without speech, the audio the decoder receives, word error rate on clean talks and on composites with inserted music and noise, the cost of the 30 s hard cut, and the limits. Update the guidance. Silero stays the always-on gate, the Redux head suits long recordings that are mostly speech, and the gate is opt-in. The fusion gain of the synthetic experiment did not hold on real recordings (+0.10 and +0.46 F1 points, no word error rate gain), so fusion is not built. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Add the scripts of the study (data fetch, probability dump, a Python replica of the segmenter, frame and decoder scoring, WER reports) and the small result tables, with a README on how to rerun them. Paths are environment variables. verify_gate.py checks the C++ run gate against the Python gate on stored probabilities. No audio, probabilities or models. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this adds
An opt-in run gate for the VAD segmenter, and the evidence from a study of Silero, the Ultra head and the Redux head on real recordings.
The gate: a speech run (a stretch of frames with
p >= threshold, found before bridging) is dropped when the median of its frame probabilities is belowrun_gate. A median equal to the gate keeps the run. For an even number of frames the median is the mean of the two middle values. The gate acts first, so the frames of a dropped run are silence for bridging,min_speech, the pauses, the cuts and the trim. It works for any detector and frame size and in both modes (speechandsegments). It is meant for the Ultra and Redux heads, whose noise runs sit on a plateau (0.8 to 0.9) while speech sits above 0.99.It is off by default (
run_gate0). With the gate off the output is byte for byte what it was.How to use it
run_gate, a number in [0, 1), inparakeet_capi_vad_pcm_json,parakeet_capi_vad_path_jsonandparakeet_capi_transcribe_path_json_vad_with. Other values and unknown keys are errors, like the other keys.parakeet-cli vad ... --run-gate Pandparakeet-cli transcribe ... --vad --vad-run-gate P.parakeet_capi_vad_stream_beginrefuses a non-zerorun_gatewith a clear error. Audio of at mostmax_segmentseconds is returned whole without the segmenter, so the gate has no effect there (as for the trim).The numbers
59 recordings with human labels (VoxConverse, AMI far-field, AVA-Speech film clips: 12.1 h of audio, 9.5 h of labelled speech) and 2.6 h without speech (MUSAN music and noise, ESC-50). Recording-disjoint splits, bootstrap intervals.
transcribe --vaddecodes (trim 0.3): Silero 6 percent of music, 0.03 of noise and 0.03 of ESC-50. Ultra head 95 / 84 / 94, Redux head 91 / 82 / 94. With the gate at 0.92: Ultra 53 / 14 / 27, Redux 27 / 1.1 / 5.6.What the gate does not do
segmentsmode on AVA the Redux gate loses 7.8 percent of the labelled speech (the head alone loses 0.2). The pooled cost at 0.92 is 0.3 and 0.6 points, and neither interval excludes zero.What was not built, and why
--vad-fusion). The synthetic study gained +0.75 and +1.13 F1 points. On real recordings the tuned fusion gained +0.10 (Ultra) and +0.46 (Redux) over the best tuned single detector, and gave no WER gain. It needs both models to run. Not worth a flag. The docs keep the old section as an experiment note and say so.Tests
test_vad_run_gate(no model): synthetic probability streams with exact expected runs at 0.08 s and 0.032 s frames. Covers runs above and below the gate, a low median under a high peak, bridged gaps (the raw run, not the bridged one, is gated), a one-frame run, the boundary (median equal to the gate keeps), gate then trim, the early return for short audio, random streams (a gate never adds speech, a higher gate never keeps more), NaN as a degenerate option, and that the event tracker ignores the gate. Gate 0 equals the old output (the existing digest test also covers it).test_vad_options: therun_gatekey, ranges, errors, and that the word-filter parser rejects it. CLI flag errors for--run-gateand--vad-run-gateas ctest cases.test_capi_vad_silero: a stream refuses a non-zero gate.test_vad_run_gate_model(labelmodel): two_speakers.wav, seeded white or pink noise, speech.wav. For the Redux slice, packed and dequantized Redux, Ultra and Silero: gate 0 gives the same JSON as no key, in both modes; for Redux at 0.92 at most 25 percent of the seconds the head called speech in the noise stretch remain (none did in the runs here) and the speech stretches are unchanged; the transcribe path honours the key.ctest -LE modelpasses (52 tests); the model tests pass with the Ultra, Redux (packed and dequantized), Silero and TDT 0.6B files, apart fromtest_relpos_attention_local_chunkedandtest_capi_timestamps, which are known failures that this change does not touch (test_combined_offlinewas skipped, its inputs were not available).parakeet-cli vadoutput with the gate off is byte-identical to master on 13 files, for the Redux head slice and for Silero, in both modes and with--probabilities.transcribe --vad --jsonis byte-identical on 4 files for Ultra, Redux and Silero + TDT. The C++ gate gives the same speech regions as the Python gate of the study on 688 combinations of recording, detector and gate value on the stored probabilities.Limits
DIHARD, CallHome and the MUSAN speech files were not used (their speech labels were unusable). WER was measured on TED talks only. The composite recordings are synthetic. Cut policies were scored by non-speech seconds, not WER. The Silero ONNX model was not run. Only 20 recordings are held out (3 AMI meetings), so the held-out intervals are wide. In VoxConverse and AMI the frames Silero "misses" are mostly pauses inside annotated turns, so tuning pushes the pause to 0.5 s and the pad to 0.1 s, which is a labelling convention; a 0.1 s collar raises every F1 by about 0.7 to 0.9 points and keeps the ranking.
The scripts and result tables of the study are in
scripts/vad_bench/real_recordings/(no audio, probabilities or models).🤖 Generated with Claude Code