Speech Synthesis advanced 9 min read 6 flashcards

Full-Duplex Spoken Dialogue Models

How models that listen and speak on separate simultaneous streams move the turn-taking decision inside the network, the levels at which that decision can live, and what always-on listening costs in compute, context integrity and controllability.

A half-duplex voice system has two states and one question: has the user finished? A full-duplex model never asks it. It consumes the user's audio and emits its own on separate streams at every timestep, so "should I be speaking right now" is answered once per frame by the same forward pass that generates the audio. Moshi, the first real-time full-duplex speech-text model, runs on 80 ms frames with a theoretical latency of 160 ms and about 200 ms in practice, which is inside the human gap distribution rather than an order of magnitude above it (Défossez et al., 2024, Moshi: a speech-text foundation model for real-time dialogue, arXiv:2410.00037).

Two streams, and a text stream nobody hears

The textless ancestor of this design is dGSLM, a dual-tower transformer with cross-attention trained on 2,000 hours of two-channel conversational audio from the Fisher corpus, with no text or labels anywhere in the pipeline. Each tower models one participant's channel, cross-attention lets each see the other, and the result generates laughter, backchannels and overlapping talk, with turn-taking statistics closer to human conversation than a cascaded text baseline (Nguyen et al., 2022, Generative Spoken Dialogue Language Modeling, arXiv:2203.16502, TACL 2023). Textless also meant semantically weak: fluent conversational behaviour with little to say.

Moshi keeps the two audio streams and adds a third, private one. Its Mimi codec compresses 24 kHz audio to a 12.5 Hz token rate at 1.1 kbps in a streaming causal encoder, with a distillation loss that pushes WavLM-style semantic information into the first codebook so one tokeniser carries both semantics and acoustics. The model then predicts its own text tokens alongside its audio tokens, an "inner monologue" that acts as scaffolding for generation. The text stream is what recovers the linguistic competence dGSLM lacked, and it is also a convenient place to hang tool calls and transcripts, because it is the only part of the model's output that is readable.

Where the duplex decision lives

A 2026 survey organises the design space by the level at which the decision is taken, from an external module down to a shared latent representation, which is a more useful axis than the usual cascade-versus-end-to-end split (A Survey of Full-Duplex Spoken Dialogue Systems, 2026, arXiv:2606.19453):

  • L0, external module. A VAD and an end-of-turn classifier outside the model. Fully inspectable, fully threshold-bound.
  • L1, hidden-state predictor. A head on the model's own states predicts take-turn, wait or backchannel. Freeze-Omni takes this route with the language model's parameters frozen, trained with text-speech pairs and 60,000 multi-round text question-answer dialogues on eight GPUs (Wang et al., 2024, Freeze-Omni, arXiv:2411.00774).
  • L2, token-level synchronisation. The user's stream and the model's stream are interleaved in one token sequence, as in Moshi. The decision is a sampled token.
  • L3, shared latent. Both parties' audio live in one representation with no separate control signal at all.

The survey pairs this with a state machine of IDLE, LISTEN, SPEAK, WAIT and DUAL, which is worth internalising because most production systems implement only LISTEN and SPEAK and then discover that WAIT (user paused, hold the floor open) and DUAL (both talking on purpose) are where the naturalness lives.

Routing the user's stream is a real design choice

How the user's audio enters the model changes its behaviour measurably. Injecting it directly into the input sequence, "channel fusion", gives stronger spoken question answering because the speech is grounded in the same representation the model reasons over. Keeping it outside as memory reached through cross-attention adapters is weaker semantically but contains the damage, because fused overlapping speech can corrupt the context the model is currently generating from (How Should LLMs Listen While Speaking?, 2026, arXiv:2605.10199). Semantic grounding and context integrity pull against each other, and the choice between them is not a hyperparameter.

When it breaks

It listens to itself. The model's own output reaches the microphone. Echo cancellation is now part of the model's correctness, not an audio nicety, and a duplex system evaluated on clean separated channels will behave differently on a single-channel phone call.

Always-on inference costs always-on compute. A frame must be processed whether anyone is talking or not, so the idle cost is the busy cost. A cascade can be silent for free; a duplex model cannot, which changes the shape of capacity planning for a fleet of concurrent calls.

Taking the floor becomes the default failure. Full-Duplex-Bench measured takeover rate during mid-turn pauses and found Moshi taking the turn in about 98 percent of them, which reads less like a turn-taking policy than a bias to speak (Lin et al., 2025, arXiv:2503.04721). Moving the decision into the model does not make it a good decision; it makes it harder to inspect.

The knobs are gone. A product that needs "never interrupt during the confirmation step" has no threshold to raise. The policy now lives in weights, and the available interventions are prompting, fine-tuning, or an external guard that re-introduces the L0 module the architecture was meant to remove.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Défossez et al., 2024, Moshi: a speech-text foundation model for real-time dialogue, arXiv:2410.00037 arxiv.org
  2. Nguyen et al., 2022, Generative Spoken Dialogue Language Modeling, arXiv:2203.16502 arxiv.org
  3. A Survey of Full-Duplex Spoken Dialogue Systems, 2026, arXiv:2606.19453 arxiv.org
  4. Wang et al., 2024, Freeze-Omni, arXiv:2411.00774 arxiv.org
  5. How Should LLMs Listen While Speaking?, 2026, arXiv:2605.10199 arxiv.org
  6. Lin et al., 2025, arXiv:2503.04721 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track