Voice activity detection

ProductionModels and inferencePublished By Simon Budziak

Voice activity detection, or VAD, identifies which parts of an audio stream contain speech. A voice system uses it to start capturing a turn, stop after silence, avoid sending empty audio, and decide when to respond. VAD detects acoustic speech activity, not whether a person's thought is semantically complete.

How does VAD set speech boundaries?

The detector scores short audio frames, then applies thresholds and timing rules for speech start and stop. Aggressive settings respond quickly but can clip quiet words or pauses, while conservative settings increase delay. Noise suppression can improve the signal, but it cannot remove every competing voice or sound.

How is VAD used in a voice agent?

VAD feeds endpointing and turn detection so the system knows when to process input. Silence alone does not prove that a caller has finished speaking. Semantic VAD adds meaning based completion, while a basic detector remains useful for bandwidth control, recording boundaries, and fast interruption signals in voice agents.

Frequently asked questions

Does VAD transcribe speech?

No. It marks speech and non-speech regions, while automatic speech recognition converts spoken audio into text.

Why does VAD sometimes cut people off?

A silence threshold can mistake a natural pause for the end of a turn, especially in noisy audio or slower speech.

Summarize this page with

Train your team to build this