Speech Recognition intermediate 8 min read 6 flashcards

Semantic End-of-Turn Detection

How voice agents replaced the silence threshold with a model that predicts whether the user is finished, the three families of end-of-turn predictor, and what each one costs in latency, language coverage and failure mode.

"My account number is four, seven, two" has a 900 ms pause in the middle of it, and a voice agent with a 500 ms silence threshold will answer before the digits are done. The transcript is perfect. The turn-taking decision is wrong, and no improvement in recognition accuracy fixes it, because the question being asked is not "what did they say" but "are they finished".

End-of-turn (or end-of-utterance) detection is the component that answers it. The structure almost everybody converged on is a gate: a cheap voice activity detector notices silence, and that silence triggers a model that decides whether the silence means finished or thinking (LiveKit, Turn detection). The VAD keeps the expensive model from running continuously; the model keeps the VAD from deciding something it cannot know.

Three families, three sets of evidence

Audio-native continuous predictors. Voice Activity Projection trains a transformer on two-channel conversational audio to predict, at every frame, the voice activity of both participants over the next two seconds, which turns turn-taking into a self-supervised prediction task rather than a classification of the present moment (Ekstedt and Skantze, 2022, Voice Activity Projection: Self-supervised Learning of Turn-taking Events, arXiv:2205.09812). Because the target is the near future, the same model also predicts backchannel opportunities and shift-versus-hold, and later work made it run in real time and across languages (Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection, 2024, arXiv:2401.04868; Multilingual Turn-taking Prediction Using Voice Activity Projection, 2024, arXiv:2403.06487). It sees prosody, which is where much of the human signal lives.

Text-based completion classifiers. Run streaming recognition, hand the partial transcript to a small language model, and ask whether the utterance is syntactically and pragmatically complete. LiveKit's turn detector is a 0.5B model distilled from Qwen2.5-0.5B-Instruct for exactly this, and because it consumes text it inherits recognition latency and cannot see that the speaker's pitch is still rising.

Small audio classifiers on the silence. Pipecat's Smart Turn line puts a lightweight classification head on a Whisper-tiny encoder and runs it only when the VAD reports silence: roughly 8M parameters, an 8 MB int8 ONNX file, tens of milliseconds on CPU, no GPU in the serving path. It sees prosody like VAP but costs almost nothing, at the price of judging a single moment rather than tracking the conversation continuously.

What the predictor actually buys

It does not remove the threshold. It changes what the threshold is applied to: instead of thresholding silence duration, you threshold \(P(\text{finished} \mid \text{audio}, \text{text})\), and the two error types become separately steerable. A fixed 700 ms wait treats a trailing "so, um" and a crisp "that's all" identically. A predictor can wait 1.2 s on the first and respond in 150 ms on the second, which lowers mean latency and false-interruption rate at the same time, something no single silence threshold can do.

The second gain is that the same model supports a third action. With a probability rather than a boolean, a system can emit a backchannel or a short filler when the user is mid-turn but has paused, which buys time without claiming the floor.

When it breaks

It learned somebody else's turn conventions. These models are trained on corpora with particular languages, recording conditions and interaction styles. A predictor trained on relaxed dyadic English conversation will misjudge a hurried bilingual call-centre exchange, and the failure is systematic rather than noisy.

Domain phrasing defeats completeness. Addresses, order numbers, spelled names and dictated code are grammatically complete at many points where the speaker is not done. These are exactly the utterances that matter in production, so a general predictor usually needs a task-specific override: during a digit-collection step, require a longer silence or a digit count, not a completeness score.

It is still in the critical path. A text-based predictor cannot fire before the recogniser has emitted the words, so its decision inherits recognition delay; an audio predictor avoids that but needs its own inference slot every time the VAD trips. Budget for the component, and measure its contribution separately from the model's time to first token.

Nothing projects on behalf of a cascade. Even a perfect end-of-turn model only tells a pipeline when to start, and the generation has not begun yet. Human timing comes from planning during the other party's turn, which is a property of full-duplex architectures, not of better detection.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. LiveKit, Turn detection docs.livekit.io
  2. Ekstedt and Skantze, 2022, Voice Activity Projection: Self-supervised Learning of Turn-taking Events, arXiv:2205.09812 arxiv.org
  3. Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection, 2024, arXiv:2401.04868 arxiv.org
  4. Multilingual Turn-taking Prediction Using Voice Activity Projection, 2024, arXiv:2403.06487 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track