Inference & Serving

When to Speak: The Turn-Taking Problem Behind Full-Duplex Voice Agents

People leave gaps of about 200 milliseconds between turns, and need about 600 milliseconds to plan a single word. The gap only exists because listeners plan while the other person is still talking. Voice agents have spent three years rebuilding their architecture around that one fact, and the benchmarks say the hard part is still unsolved.

Measure the gaps in a recorded conversation and you get a distribution with a mode between zero and 200 milliseconds, and it holds across ten languages on five continents, with cultural variation measured in tens of milliseconds rather than factors (Stivers et al., 2009, Universals and cultural variation in turn-taking in conversation, PNAS 106(26)). Now measure how long a speaker needs to plan and articulate a single word: roughly 600 milliseconds. The two numbers are incompatible unless listeners start planning their reply before the other person has finished, which is exactly what they do (Levinson and Torreira, 2015, Timing in turn-taking and its implications for processing models of language, Frontiers in Psychology 6:731).

Every voice assistant built on recognition followed by generation followed by synthesis violates that constraint by construction. It cannot plan during your turn, because it does not know what you said until you stop, and it does not know you stopped until a timer says so. The last three years of work on voice agents is, read generously, one long attempt to fix the placement of a single decision: who gets to decide when the machine starts talking, and what evidence that decision is allowed to see.

Why this matters: Latency in voice AI is usually discussed as a performance problem, something faster models and faster kernels will eventually dissolve. It is not. The dominant term is a turn-taking decision made with the wrong evidence at the wrong layer, and the architectures now shipping to production differ mainly in where they put that decision and what they give up to move it.

TL;DR

  • Human turn gaps are about 200 ms while word planning takes about 600 ms, so human timing requires planning during the other party's turn. A pipeline that starts work at end-of-turn detection can never reach it by getting faster.
  • A 2026 cascade sums to roughly 540 ms of round trip before any endpointing wait, against ITU-T G.114's 150 ms comfortable and 400 ms tolerable one-way targets.
  • Thresholding silence is a bad estimator of intent, and no threshold value is correct. In the worked example below, the cost-optimal setting still interrupts the user on about one turn in six.
  • Full-duplex models remove the threshold by generating and listening on separate streams every frame. Moshi runs 80 ms frames with 160 ms theoretical and about 200 ms practical latency (Défossez et al., 2024, arXiv:2410.00037).
  • Moving the decision inside the model does not make it correct. Full-Duplex-Bench measured Moshi taking the floor in 98.5% of mid-turn pauses, with the most patient system in the comparison also the slowest to reply (Lin et al., 2025, arXiv:2503.04721).
  • Speech-native models pay for duplex behaviour in measured intelligence: adding speech modelling to SpiritLM moved MMLU from 45.3 to 36.9, and GLM-4-Voice answers around 37% of VoxEval against a text counterpart above 60% on MMLU (Fang et al., 2025, S2SBench, arXiv:2505.14438).
  • The useful mental model is a ladder of four levels, L0 to L3, describing where the duplex decision lives; descending it buys naturalness and costs inspectability and control (A Survey of Full-Duplex Spoken Dialogue Systems, 2026, arXiv:2606.19453).

At a Glance

One decision, four possible homes. Everything else in a voice stack is downstream of which box owns it.

flowchart LR
  Audio["User audio stream"] --> L0["L0 external VAD plus EOU model"]
  Audio --> L1["L1 head on hidden states"]
  Audio --> L2["L2 interleaved token streams"]
  Audio --> L3["L3 shared latent, no control signal"]
  L0 --> Speak["Decision: speak, wait, backchannel"]
  L1 --> Speak
  L2 --> Speak
  L3 --> Speak
  Speak --> Cost["Lower levels: faster and more natural"]
  Speak --> Risk["Lower levels: less inspectable and less steerable"]
  classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
  classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
  classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
  classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
  classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
  class Audio blue
  class L0,L1,L2,L3 purple
  class Speak teal
  class Cost emerald
  class Risk rose

[IMAGE: Two stacked waveform strips labelled "user" and "agent" over a shared time axis, annotated with four marked regions: a 200 ms clean transition, a 900 ms mid-turn pause wrongly treated as end of turn, a backchannel overlapping the agent's speech, and a genuine barge-in. Caption: "Four events that a single silence threshold cannot tell apart."]

Before Duplex: Turn-Taking as an Afterthought

Conversation analysts described the machinery long before anyone had to implement it. The founding paper observed that talk is organised so one party speaks at a time, that transitions are managed locally rather than scheduled, and that participants project the end of a turn rather than react to it (Sacks, Schegloff and Jefferson, 1974, A Simplest Systematics for the Organization of Turn-Taking for Conversation, Language 50(4), 696-735). Projection is the part engineering skipped: the spoken-dialogue systems of the 2000s used energy-based voice activity detection and a silence timer, and the timer survived every later upgrade of the components around it.

Google Duplex, demonstrated in May 2018, made the gap between fluent synthesis and fluent interaction publicly obvious: the voice was convincing, and the turn structure was scripted. The stacks that followed removed any excuse that the components were the bottleneck, as weakly supervised recognition made transcripts cheap and streaming vocoders dropped time to first audio into the tens of milliseconds. What remained was the timer.

The first genuine attack on turn-taking as a modelling problem came from two directions. One was prediction: train on two-channel audio to forecast who will be speaking in the near future, which makes turn-taking self-supervised rather than a classification of the present moment (Ekstedt and Skantze, 2022, Voice Activity Projection, arXiv:2205.09812). The other was generation. dGSLM trained a dual-tower transformer with cross-attention on 2,000 hours of two-channel Fisher conversations with no text at all, and produced laughter, backchannels and overlapping talk with turn-taking statistics closer to human conversation than a cascaded text baseline (Nguyen et al., 2022, Generative Spoken Dialogue Language Modeling, arXiv:2203.16502). It also had very little to say, because nothing in it modelled language above the level of acoustic units.

timeline
  title From silence timers to duplex models
  1974 : Sacks, Schegloff and Jefferson describe locally managed turn-taking and projection
  2018 : Google Duplex shows fluent synthesis with a scripted turn structure
  2022 : Voice Activity Projection predicts near-future speech activity from two channels
       : dGSLM models both channels of a dialogue with no text, gaining overlap and losing semantics
  2024 : GPT-4o claims 232 ms minimum and 320 ms average audio response
       : Moshi ships dual audio streams plus an inner monologue at about 200 ms
       : Freeze-Omni adds duplex behaviour around a frozen language model
  2025 : Full-Duplex-Bench separates pause handling, backchannels, transitions and interruptions
       : Commercial realtime speech-to-speech APIs reach general availability
  2026 : Multi-turn duplex benchmarks, a challenge with dual-channel human data, and an L0 to L3 taxonomy

The 2024 inflection was commercial. GPT-4o was announced on 13 May 2024 with the claim of responding to audio in as little as 232 ms and 320 ms on average, numbers chosen precisely because they sit inside the human distribution. Moshi arrived in the autumn as an open, documented system with the same property and an architecture anyone could read. After that the question stopped being whether duplex interaction was possible and became what it costs.

How the Decision Moved Inside the Model

A per-frame question instead of a per-turn verdict

A half-duplex system asks "has the user finished?" once per turn and acts on the answer. A full-duplex model asks "should I be emitting audio right now?" once per frame, and the answer is a token it samples like any other. Moshi's frame rate comes from its codec: Mimi compresses 24 kHz audio to a 12.5 Hz token stream at 1.1 kbps, causally and in streaming mode, which puts one frame at 80 ms and makes a 160 ms theoretical response latency arithmetically available. The practical figure, about 200 ms, is what the human distribution looks like.

The codec detail that matters most is not the bitrate. Mimi adds a distillation loss that pushes WavLM-style self-supervised semantic information into its first codebook, so a single causal tokeniser carries both meaning and acoustics. Without that, a dialogue model needs separate semantic and acoustic token streams and the alignment problem between them. With it, the model reasons over the same tokens it will render back into sound.

The text stream nobody hears

dGSLM showed that two-channel audio modelling produces good interaction and weak content. Moshi's answer was to predict its own text tokens alongside its audio tokens, an inner monologue that scaffolds generation. This is the single most consequential design choice in the system, for three reasons. It restores the linguistic competence that textless modelling gives up. It provides a readable surface, which is where transcripts, logging and tool calls attach, because audio tokens are not something an application can branch on. And it creates a controllable interface: steering the text stream steers the speech without retraining the audio path.

[IMAGE: Diagram of three parallel token lanes over the same time axis, labelled "user audio tokens", "agent audio tokens" and "agent text tokens (inner monologue)", with the text lane slightly ahead of the agent audio lane and an annotation showing where a tool call is emitted. Caption: "The private text stream is both the semantic scaffold and the only machine-readable handle."]

The L0 to L3 ladder

A 2026 survey organises this design space by the level at which the duplex decision is taken, which is more useful than the usual cascade-versus-end-to-end dichotomy (A Survey of Full-Duplex Spoken Dialogue Systems, 2026, arXiv:2606.19453):

  • L0, external module. A voice activity detector plus an end-of-turn classifier outside the model. Every decision is visible and overridable, and every decision is made from outside the model's understanding of the conversation.
  • L1, hidden-state predictor. A head on the model's own states predicts take, wait or backchannel. Freeze-Omni is the clean example, attaching speech encoders and decoders around a language model whose parameters stay frozen, trained with text-speech pairs and 60,000 multi-round text dialogues on eight GPUs (Wang et al., 2024, Freeze-Omni, arXiv:2411.00774).
  • L2, token-level synchronisation. Both streams are interleaved in one sequence and the decision is a sampled token, as in Moshi.
  • L3, shared latent. Both parties live in one representation with no separable control signal at all.

The survey pairs the ladder with a state machine of IDLE, LISTEN, SPEAK, WAIT and DUAL. That list is worth holding onto, because most production systems implement LISTEN and SPEAK, and the two states they omit are where the naturalness lives: WAIT is "you paused, I am holding the floor open for you", and DUAL is "we are both talking and that is fine".

Routing the user's stream changes behaviour

Once the model listens continuously, a design question appears that does not exist in a cascade: how does the user's audio enter the network? Injecting it directly into the input sequence, channel fusion, grounds the speech in the same representation the model reasons over and measurably improves spoken question answering. Keeping it outside as memory reached through cross-attention adapters is semantically weaker, but it contains the damage, because fused overlapping speech can corrupt the context the model is currently generating from (How Should LLMs Listen While Speaking?, 2026, arXiv:2605.10199). Semantic grounding and context integrity are in tension, and the choice between them is architectural rather than a tuning parameter.

What the cascade did instead: a model in place of the timer

The other branch kept the pipeline and replaced the timer with a predictor. The pattern that won is a gate: a cheap detector notices silence, and that silence triggers a model that decides whether the silence means finished or thinking (LiveKit, Turn detection). Two flavours ship in production. Text classifiers read the streaming transcript and judge completeness; LiveKit's is a 0.5B model distilled from Qwen2.5-0.5B-Instruct, so it inherits recognition latency and cannot hear prosody. Audio classifiers judge the waveform directly; Pipecat's Smart Turn line puts a classification head on a Whisper-tiny encoder at roughly 8M parameters, around 8 MB quantised to int8, running in tens of milliseconds on CPU.

This does not remove the threshold. It changes what is thresholded, from silence duration to \(P(\text{finished} \mid \text{audio}, \text{text})\), and that is a real gain: the two error types become separately steerable, so the system can wait 1.2 s after a trailing "so, um" and answer in 150 ms after a crisp "that's all". A single silence threshold cannot reduce mean latency and false interruptions at the same time. A predictor can.

Seeing It in Motion

The first view is the cascade's critical path on one turn, including the barge-in that invalidates it. Note where the wait sits: before anything else can start.

sequenceDiagram
  participant U as User
  participant G as VAD plus EOU
  participant A as Recogniser
  participant M as Language model
  participant T as Synthesiser
  U->>G: Speech, then 900 ms pause
  G->>G: Is this finished or thinking
  G->>A: Endpoint declared
  A->>M: Final transcript
  M->>T: First tokens
  T->>U: First audio
  U->>G: Starts speaking again
  G->>T: Stop playback
  Note over M,T: History must be truncated at the playback cursor, not the generation log

The second view is the state machine the duplex literature converged on, with the two states most systems leave out drawn in.

stateDiagram-v2
  [*] --> Idle
  Idle --> Listen: user speech detected
  Listen --> Wait: pause, intent unresolved
  Wait --> Listen: user resumes
  Wait --> Speak: turn judged complete
  Listen --> Speak: transition point projected
  Speak --> Dual: user starts while agent speaks
  Dual --> Listen: agent yields the floor
  Dual --> Speak: incoming audio was a backchannel
  Speak --> Idle: response complete

[IMAGE: Side-by-side state diagrams, left showing the two-state LISTEN and SPEAK machine most products implement, right showing the five-state machine with WAIT and DUAL highlighted, annotated with the user-visible symptom each missing state causes. Caption: "Missing states are not abstractions; each one maps to a complaint."]

Watch It Run

Animated diagram of user audio flowing into a duplex model that emits audio and text streams while a turn-taking decision loops every frame.
Solid animated edges carry audio frames in the direction of flow; the animated self-loop on the decision node is the per-frame speak-or-wait judgement that replaces the silence timer; the amber feedback edge is the agent's own output returning through the microphone, which is why echo cancellation belongs inside the correctness argument. The static Mermaid figures above show the same structure if the animation is absent.

By the Numbers

Quantity Value Where it comes from
Modal human turn gap 0 to 200 ms, stable across 10 languages Stivers et al., 2009, PNAS
Single-word production latency about 600 ms Levinson and Torreira, 2015
ITU-T G.114 one-way target 150 ms comfortable, 400 ms tolerable ITU-T Recommendation G.114
2026 cascade round trip about 540 ms before endpointing wait Component figures from vendor documentation, aggregated
GPT-4o audio response 232 ms minimum, 320 ms average OpenAI announcement, 13 May 2024
Moshi frame and latency 80 ms frames, 160 ms theoretical, about 200 ms practical Défossez et al., 2024
Mimi codec 24 kHz in, 12.5 Hz tokens, 1.1 kbps, 80 ms streaming latency Défossez et al., 2024
Takeover rate at mid-turn pauses dGSLM 0.934, Moshi 0.985, Freeze-Omni 0.642 (lower is better) Lin et al., 2025, Full-Duplex-Bench
Response latency on the same benchmark dGSLM 0.352 s, Moshi 0.265 s, Freeze-Omni 0.953 s Lin et al., 2025
Intelligence cost of speech modelling SpiritLM MMLU 45.3 to 36.9; GLM-4-Voice about 37% VoxEval versus 60%+ text MMLU Fang et al., 2025, S2SBench
Turn-detector model sizes 0.5B text-based; about 8M audio-native, 8 MB int8 LiveKit and Pipecat documentation
HumDial dual-channel corpus over 100 hours, Chinese and English, interruption and rejection scenarios ICASSP 2026 HumDial challenge

Sources: human timing from Stivers et al., 2009 and Levinson and Torreira, 2015; model figures from Défossez et al., 2024 and Fang et al., 2025; benchmark figures from Lin et al., 2025 and the ICASSP 2026 HumDial study. Component latencies and model sizes for commercial stacks are vendor-reported best cases rather than independent measurements under load, and the GPT-4o figures are the vendor's own announcement numbers.

A Concrete Example

Take a phone support agent and cost out its turn-taking policy. The assumptions below are illustrative, chosen to be plausible rather than measured, and the point is the shape of the result, not the exact milliseconds.

Setup. Fixed pipeline cost once the turn is declared over: 60 ms network, 150 ms to finalise recognition, 150 ms to the model's first token, 90 ms to first audio. That is \(L = 450\) ms. Assume 35% of user turns contain at least one mid-turn pause, and that those pauses are exponentially distributed with a mean of 400 ms. A false interruption costs 2,500 ms of stop-start repair.

Step 1: price a silence threshold. With threshold \(T\), the agent pays \(L + T\) on every turn, and pays the recovery cost whenever a mid-turn pause exceeds \(T\):

\[\mathbb{E}[\text{cost}](T) = 450 + T + 0.35 \cdot 2500 \cdot e^{-T/400}\]
\(T\) Latency paid every turn \(P(\text{false interrupt})\) Expected penalty Total
200 ms 650 ms 0.212 531 ms 1,181 ms
300 ms 750 ms 0.165 413 ms 1,163 ms
500 ms 950 ms 0.100 251 ms 1,201 ms
700 ms 1,150 ms 0.061 152 ms 1,302 ms
1,000 ms 1,450 ms 0.029 72 ms 1,522 ms

Step 2: find the optimum. Differentiate and set to zero: \(1 = (875/400)\,e^{-T/400}\), so \(e^{-T/400} = 0.457\) and \(T^{\star} = 313\) ms, with a total of 1,163 ms. Read the third column at that point. The cost-optimal threshold interrupts the user on \(0.35 \times 0.457 = 16\%\) of turns. Not a misconfiguration; the optimum.

Step 3: add a turn predictor. Assume an audio-native end-of-turn classifier costs 30 ms of inference and cuts the probability of a wrong "finished" judgement by a factor of four at the same wait, and that a 120 ms VAD hangover is the floor. Expected cost becomes \(480 + 120 + 218.75\,e^{-0.3} = 762\) ms, with a false-interruption rate of 6.5%. Both numbers improved, which is the thing no threshold setting can do: latency fell by about a third and interruptions by more than half.

Step 4: price the duplex model. No threshold at all, so the raw response is one frame plus network: about 260 ms, less than half the guarded cascade. Then apply the measured takeover rate. At Moshi's 0.985, a mid-turn pause is taken as a turn almost every time:

\[260 + 0.35 \times 0.985 \times 2500 = 1{,}122 \text{ ms}\]

The architecture that responds more than twice as fast ends up worse on total expected cost than the cascade with a small turn predictor in front of it, because eagerness is expensive. Run the same arithmetic for Freeze-Omni, at 1,013 ms of latency and a 0.642 takeover rate, and it lands near 1,575 ms: patient, and too slow to collect the benefit.

What the example is really showing. Latency and interruption are one joint quantity, and each architecture picks a point on that curve rather than a number on the latency axis. A duplex model's 200 ms is only worth having with a turn policy that does not spend it, which is why teams end up bolting an external guard in front of the duplex model and reintroducing the L0 module the architecture existed to remove.

[IMAGE: Line chart of expected cost in milliseconds against endpointing threshold T from 0 to 1200 ms, showing the U-shaped total with its minimum marked at 313 ms, plus two horizontal reference lines for the guarded cascade at 762 ms and the duplex model at 1,122 ms. Caption: "The optimum is a point on a curve, not a latency target."]

[IMAGE: Scatter plot with takeover rate at mid-turn pauses on the x-axis and response latency on the y-axis, plotting dGSLM, Moshi and Freeze-Omni from the benchmark table, with an annotated frontier curve and the guarded-cascade point from the worked example marked separately. Caption: "Patience and speed trade along a frontier; no system in the comparison escapes it."]

Where It Breaks

Taking the floor becomes the default

The benchmark numbers are blunt. A takeover rate of 0.985 at mid-turn pauses is not a turn-taking policy; it is a bias to speak that happens to score well on response latency. The system in the same comparison with genuinely cautious behaviour achieved it with explicit state control and paid roughly three times the response latency for it (Lin et al., 2025, arXiv:2503.04721). Nothing in the training objective of a duplex model rewards patience, because the loss is next-token prediction over streams where somebody is usually talking.

Overlap corrupts the context it is supposed to inform

Fusing the user's channel into the model's input is what makes grounding work, and it is also what lets a stray cough or a third voice in the room enter the sequence the model is generating from. The routing study found exactly that trade. On a speakerphone with background conversation, "listening continuously" and "being interruptible by noise" are the same property.

The model hears itself

The agent's own audio returns through the microphone. In a cascade, failed echo cancellation causes a self-interruption that is annoying and obvious. In a duplex model, the returning audio enters the user stream and becomes evidence in the model's own turn-taking decision, which can drive it into a feedback loop of yielding to itself. Echo cancellation is part of model correctness here, not an audio-chain nicety, and a system evaluated on clean separated channels will not show the failure.

Knowledge leaks out of the architecture

Speech-native models pay in measured intelligence, and the mechanism is not mysterious: audio tokens carry less meaning per token, sequences get far longer for the same content, and prosody and speaker identity consume capacity (Fang et al., 2025, arXiv:2505.14438). The reported drops are large enough to change product viability, which is why hybrids exist: a frozen language model with speech wrapped around it, or a tandem design where a fast speech-to-speech front end converses while a text model answers the parts that need knowledge (KAME, 2025, arXiv:2510.02327). A voice benchmark that scores dialogue quality can look healthy while this collapse is underway, so run a text-equivalent knowledge suite alongside it (VoiceBench, 2024, arXiv:2410.17196).

Multi-turn failures that single-event probes miss

Scripted probes test one pause or one interruption. Drive a system through a multi-turn conversation with an examiner that pursues staged goals, varies pacing and rephrases on failure, and new failures appear: confusion under overlapping speech, poor correction handling, and losing track of which entity is under discussion (Lin et al., 2025, Full-Duplex-Bench-v2, arXiv:2510.07838). A 2026 extension adds tool use under real disfluency (Full-Duplex-Bench-v3, 2026, arXiv:2604.04847), which is where commercial deployments live.

The knobs are gone, and the transport is hostile

A requirement like "never interrupt during payment confirmation" is a one-line change at L0 and has no implementation at L2, where the policy lives in weights: the interventions left are prompting, fine-tuning, or an external guard that costs back the latency the architecture was chosen for. Decide which requirements are hard before picking a level on the ladder. The transport compounds it. A PSTN leg delivers 20 ms packets through a jitter buffer with codec delay, and packet loss concealment can manufacture a gap indistinguishable from a pause, so turn-taking calibrated on clean low-latency audio becomes both slower and more trigger-happy on the network that carries most production calls.

Alternative Designs

Design How it works Key advantage Key limitation Best when
Silence threshold only VAD plus fixed endpointing timer Trivial, inspectable, free No value of the threshold is correct; interruptions and latency trade directly Short command-and-control turns with predictable phrasing
Cascade plus semantic end-of-turn model Small text or audio classifier gated on VAD silence Improves latency and interruption rate together; keeps full text tooling Still cannot plan during the user's turn; predictor inherits its training conventions Most production voice agents, especially with tools and compliance steps
L1 duplex around a frozen language model Speech encoders and decoders plus a state head, base model frozen Keeps text-model knowledge; explicit controllable states Acoustics never influence reasoning; higher response latency in measurements Knowledge-heavy assistants that must not regress on text tasks
L2 token-interleaved duplex Two audio streams plus an inner monologue in one sequence Human-range latency; overlap and backchannels representable Floor-taking bias; context corruption under overlap; idle cost equals busy cost Open conversation where naturalness is the product
Tandem duplex plus text back end Fast speech-to-speech front end, text model behind it for knowledge Recovers knowledge without giving up timing Two systems to coordinate and to keep consistent Products that need both conversational feel and reliable facts

The ranking depends on the cost ratio in your own version of the worked example. If a false interruption means repeating two words, speed wins and a bare threshold is defensible. If it means a customer repeating a 16-digit number, patience wins and the extra 300 ms is the cheapest thing in the budget.

How It Is Used in Practice

Commercial realtime APIs are the mainstream route into duplex behaviour, and they are deliberately hybrid: a speech-to-speech model with text-side tooling, function calling and transcripts exposed alongside the audio, so an application can read and steer what the model does. OpenAI's Realtime API reached general availability with a dedicated speech-to-speech model in August 2025 after about a year in beta, and the line has since added reasoning-grade variants, as reported by developer-facing coverage rather than independent measurement. The open-source orchestration frameworks took the complementary path, shipping the L0 pieces as swappable components: an activity detector, a small end-of-turn model on top, and the rest of the pipeline replaceable.

Two operational details separate teams that ship from teams that demo. Playback state is part of dialogue state: when a barge-in truncates speech, history has to be cut at the playback cursor, because an agent that believes it delivered information it never spoke will contradict itself two turns later. And the turn policy needs per-step overrides, because digit collection, address capture and spelled names are grammatically complete at many points where the speaker is not finished. Require a longer wait or a digit count during those steps, and accept the latency.

Evaluation practice is catching up slowly. The useful discipline is to report latency percentiles separately for clean end-of-turn, mid-turn pause and barge-in cases, with false-interruption rate on the same dashboard. The ICASSP 2026 HumDial challenge contributed what the field was missing most, over 100 hours of dual-channel human-recorded conversation in Chinese and English annotated for interruption and rejection scenarios, because synthetic probes understate how messy a real pause is (Full-Duplex Interaction in Spoken Dialogue Systems, 2026, arXiv:2604.21406).

[IMAGE: Four-panel comparison of one 12-second conversation under a bare threshold, a predictor-gated cascade, an eager duplex model and a tandem design, each panel showing user and agent audio bands with interruptions in red and dead air in grey. Caption: "One conversation, four turn policies, four failure signatures."]

Insights Worth Remembering

  1. Latency in voice AI is a decision-placement problem, not a throughput problem. The human 200 ms gap requires planning during the other party's turn. Any architecture that begins work at end-of-turn detection is bounded away from human timing no matter how fast its components get.

  2. A silence threshold is an estimator of intent built from the wrong feature. Silence is weak evidence about whether a speaker is finished, which is why the optimal threshold still interrupts a sixth of turns in the worked example. Replacing it with a predictor is not tuning; it changes what is being estimated.

  3. Moving a decision into a model relocates it rather than solving it. Moshi's 0.985 takeover rate at mid-turn pauses is the clearest evidence in the literature that duplex architecture buys timing, not judgement, and the judgement is harder to inspect afterwards.

  4. Speed and patience are one quantity measured twice. Every system that improves response latency without a turn-taking model gets it by being more willing to interrupt. Any report that gives one number without the other is unreadable.

  5. The inner monologue is an interface, not just a trick. The text stream in a duplex model is what restores semantics, and it is also the only part an application can log, branch on, or attach tools to. A duplex design without a readable stream is a product with no seams.

  6. Duplex models bill for silence. Listening is a forward pass, so idle cost equals busy cost and capacity planning shifts from talk time to call duration. This often changes the economics more than it changes any quality metric.

Open Questions

Can a duplex model be trained to be patient without being slow? Current objectives optimise next-token prediction over conversations where someone is usually speaking, and the measured result is a floor-taking bias. Whether an auxiliary objective on turn-taking events, or preference training over interaction traces, can move patience and latency independently is untested at the scale that matters.

What is the right target for turn-taking, given that humans disagree? Benchmarks score latency, takeover and backchannel behaviour against statistics from recorded human conversation, which is a reasonable proxy and not a preference model. A dictation assistant and a sales call want opposite settings, and the field has no calibrated human preference data over these axes.

Does continuous listening have an acceptable security story? A model that treats incoming audio as context it reasons over is taking instructions from whatever is audible in the room, and prompt injection through a third voice or a television is a known attack shape in text systems that has barely been studied here.

Will evaluation reach deployment conditions? The benchmarks improved quickly from single-event probes to multi-turn examiners and dual-channel human corpora, but almost all of it is clean-channel audio in two languages. How the current rankings survive on a mobile call from a noisy street, with code-switching and a jitter buffer in the path, is not yet known.

Sources and Further Reading

  1. Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., & Levinson, S. C. (2009). "Universals and cultural variation in turn-taking in conversation." PNAS, 106(26), 10587-10592. doi:10.1073/pnas.0903616106
  2. Levinson, S. C., & Torreira, F. (2015). "Timing in turn-taking and its implications for processing models of language." Frontiers in Psychology, 6:731. Full text
  3. Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). "A Simplest Systematics for the Organization of Turn-Taking for Conversation." Language, 50(4), 696-735.
  4. Nguyen, T. A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W.-N., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., & Dupoux, E. (2022). "Generative Spoken Dialogue Language Modeling." TACL (2023). arXiv:2203.16502
  5. Ekstedt, E., & Skantze, G. (2022). "Voice Activity Projection: Self-supervised Learning of Turn-taking Events." arXiv:2205.09812
  6. Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., & Zeghidour, N. (2024). "Moshi: a speech-text foundation model for real-time dialogue." arXiv:2410.00037
  7. Wang, X., Li, Y., et al. (2024). "Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM." arXiv:2411.00774
  8. Zeng, A., Du, Z., Liu, M., Wang, K., Jiang, S., Zhao, L., Dong, Y., & Tang, J. (2024). "GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot." arXiv:2412.02612
  9. Lin, G.-T., et al. (2025). "Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities." ASRU 2025. arXiv:2503.04721
  10. Lin, G.-T., Kuan, S.-Y. S., Shi, J., Chang, K.-W., Arora, S., Watanabe, S., & Lee, H.-y. (2025). "Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner." ACL 2026. arXiv:2510.07838
  11. "Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency" (2026). arXiv:2604.04847
  12. Fang, Y., Sun, H., Liu, J., Zhang, T., Zhou, Z., Chen, W., Xing, X., & Xu, X. (2025). "S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models." arXiv:2505.14438
  13. "VoiceBench: Benchmarking LLM-Based Voice Assistants" (2024). arXiv:2410.17196
  14. "KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI" (2025). arXiv:2510.02327
  15. "A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine" (2026). arXiv:2606.19453
  16. "How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue" (2026). arXiv:2605.10199
  17. "Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge" (2026). arXiv:2604.21406
  18. LiveKit. "Turn detection." LiveKit Agents documentation. docs.livekit.io
  19. ITU-T Recommendation G.114, "One-way transmission time," International Telecommunication Union.

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.