Turn-Taking and Barge-In in Voice Agents
Why deciding when to speak is a harder problem than transcribing what was said, how a silence threshold trades response latency against false interruptions, and what barge-in requires beyond detecting that the user started talking.
Human conversation runs on gaps of roughly 200 ms. Across ten languages on five continents, the modal gap between one speaker finishing and the next starting falls in the 0 to 200 ms band, with cultural variation measured in tens of milliseconds rather than orders of magnitude (Stivers et al., 2009, Universals and cultural variation in turn-taking in conversation, PNAS 106(26)). Producing a single word takes a speaker about 600 ms of planning. The arithmetic only works if people plan their reply while the other person is still talking (Levinson and Torreira, 2015, Timing in turn-taking and its implications for processing models of language, Frontiers in Psychology 6:731). A system that waits for silence, then starts work, is not slow because its components are slow. It is slow because of where it put the decision.
The decision, not the transcript
Conversation analysis has described the structure since the 1970s: talk is organised so that one party speaks at a time, transitions are managed locally, and participants project the end of a turn before it arrives rather than reacting to it (Sacks, Schegloff and Jefferson, 1974, A Simplest Systematics for the Organization of Turn-Taking for Conversation, Language 50(4), 696-735). Projection is the operative word. A listener hearing "could you send it to me by" already knows a prepositional phrase is coming and can begin planning.
A conventional voice pipeline has no projection stage. Voice activity detection reports energy, an endpointer declares the turn over after a silence of \(T\) milliseconds, and only then does the transcript reach the model. Everything downstream is blocked on \(T\), so \(T\) is added in full to every single response. See endpointing and voice activity detection for the detection mechanics; the point here is that endpointing is being asked to answer a question about intent using a feature, silence, that only weakly indicates it.
Why the silence threshold cannot be tuned right
Set \(T\) short and the agent interrupts anyone who pauses to think, breathe, or read a number off a card. Set it long and every reply arrives late, including the trivial ones. The expected cost of a threshold has two terms that move in opposite directions:
\(L_{\text{pipeline}}\) is the fixed cost of recognition, generation and synthesis. \(C_{\text{recovery}}\) is what a false interruption costs: the user stops, the agent stops, somebody repeats themselves, and two or three seconds disappear. Because mid-turn pauses are common in natural speech and recovery is expensive, the optimum sits well above the latency a product wants and well below the threshold that would make false interruptions rare. There is no value of \(T\) that is simply correct, which is why the industry has spent the last two years replacing the threshold with a model rather than tuning it.
Barge-in is two problems wearing one name
Detection is noticing that the user has started speaking while the agent is talking. It is hard mostly because of acoustics: the agent's own audio comes back through the microphone, so without echo cancellation the system reliably interrupts itself, and on a speakerphone in a car it does so with conviction.
State repair is the harder half. When the agent stops mid-sentence, the dialogue state now contains text the user never heard, and the synthesised audio was cut at an arbitrary point. A system that keeps the full generated response in history will later refer to information it never delivered. Correct handling truncates the history at the point where audio actually stopped playing, which means the orchestrator needs the playback cursor, not just the generation log.
Not every incoming sound is a bid for the floor. Backchannels ("mm-hm", "right", "yeah") are the listener confirming attention while explicitly declining the turn, and a system that treats them as interruptions makes an attentive user unable to listen politely. Distinguishing a backchannel from a barge-in cannot be done from energy alone; it needs the content and the prosody of the incoming audio, which is what semantic end-of-turn detection addresses.
When it breaks
Threshold tuning dressed as personalisation. Per-user or per-call-type thresholds help at the margin and hide the structural problem. The same user pauses differently when reciting a card number and when telling a story.
Overlap is normal, not an error. Measured conversation contains overlapping talk as a routine feature. A design whose only states are LISTEN and SPEAK has no representation for the thing humans do constantly, and the benchmark literature now treats pause handling, backchannelling, smooth transition and interruption as four separate capabilities precisely because systems can be good at one and hopeless at another (Lin et al., 2025, Full-Duplex-Bench, arXiv:2503.04721).
Telephony hides the problem and then exposes it. On a PSTN leg, 20 ms packets, jitter buffering and codec delay add their own tens of milliseconds, and packet loss concealment can present a gap that looks exactly like a pause. Turn-taking logic tuned on WebRTC audio in an office degrades on a phone call from a train.
The cost is asymmetric and your metric probably is not. Interrupting a user is far worse than answering 200 ms late, but a dashboard reporting mean response latency will reward the opposite trade. Measure false-interruption rate alongside latency or the optimisation will quietly make the product rude.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Stivers et al., 2009, Universals and cultural variation in turn-taking in conversation, PNAS 106(26) pnas.org
- Levinson and Torreira, 2015, Timing in turn-taking and its implications for processing models of language, Frontiers in Psychology 6:731 frontiersin.org
- Lin et al., 2025, Full-Duplex-Bench, arXiv:2503.04721 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.