Automatic speech recognition

LLM foundationsModels and inferencePublished By Simon Budziak

Automatic speech recognition, or ASR, converts spoken audio into machine-readable text. It is also called speech-to-text, while streaming speech recognition returns partial transcripts as a person speaks. Accuracy depends on language, accent, audio quality, domain vocabulary, overlapping speakers, and how the system decides that an utterance has ended.

What does an ASR system produce?

ASR maps an audio signal to words, often with timestamps and confidence information. Streaming ASR emits provisional text before the speaker finishes, then may revise it as more context arrives. The NIST speech recognition evaluation tradition commonly uses word error rate, but a production voice agent also needs correct names, numbers, and intent.

What affects speech recognition quality?

Microphone quality, compression, noise, accents, specialist vocabulary, and overlapping speakers all matter. Measure errors on the calls your system will actually receive, not only a public benchmark. Combine ASR with speaker diarization when several people speak, endpointing to finalize turns, and human confirmation before consequential actions based on uncertain transcripts.

Frequently asked questions

Are ASR and speech-to-text the same thing?

Yes. Speech-to-text is the common product name for automatic speech recognition.

What is streaming speech recognition?

It returns provisional transcript segments during speech, then revises or finalizes them when the utterance ends.

Summarize this page with

Train your team to build this