Voice activity detection, or VAD, identifies which parts of an audio stream contain speech. A voice system uses it to start capturing a turn, stop after silence, avoid sending empty audio, and decide when to respond. VAD detects acoustic speech activity, not whether a person's thought is semantically complete.
How does VAD set speech boundaries?
The detector scores short audio frames, then applies thresholds and timing rules for speech start and stop. Aggressive settings respond quickly but can clip quiet words or pauses, while conservative settings increase delay. Noise suppression can improve the signal, but it cannot remove every competing voice or sound.
How is VAD used in a voice agent?
VAD feeds endpointing and turn detection so the system knows when to process input. Silence alone does not prove that a caller has finished speaking. Semantic VAD adds meaning based completion, while a basic detector remains useful for bandwidth control, recording boundaries, and fast interruption signals in voice agents.
Frequently asked questions
Does VAD transcribe speech?
No. It marks speech and non-speech regions, while automatic speech recognition converts spoken audio into text.
Why does VAD sometimes cut people off?
A silence threshold can mistake a natural pause for the end of a turn, especially in noisy audio or slower speech.
No advertising or tracking cookies, and our visitor counts are anonymous. The Cal.com booking widget loads only if you allow it. Privacy Policy.
The page itself, anything our host sets to serve and secure it, and the anonymous visitor count. Always on, and none of it stores anything on your device.
The Cal.com booking widget. Left off, a booking link opens the booking page instead of a popup, so you can still book a call.