Evaluating Full-Duplex Dialogue
Why word error rate and mean opinion score say nothing about whether a voice agent takes turns well, the four interaction behaviours current benchmarks isolate, and how to read metrics whose best value is neither high nor low.
A voice agent can transcribe at 4 percent word error rate, synthesise at a mean opinion score above 4, answer every question correctly, and still be unusable because it talks over people. None of the three standard numbers measures interaction. Turn-taking quality is a property of the timing relationship between two audio streams, and it is invisible to any metric computed on one stream at a time (word error rate and MOS are both single-stream measures).
Four behaviours, measured separately
Full-Duplex-Bench established the decomposition most later work follows, scoring a model on pause handling, backchannelling, smooth turn taking and user interruption, using controlled audio scenarios that provoke each behaviour (Lin et al., 2025, Full-Duplex-Bench, arXiv:2503.04721, ASRU 2025). The reported numbers are the useful part, because they show the capabilities are not correlated:
| Model | Takeover rate at mid-turn pauses | Backchannel frequency | Response latency |
|---|---|---|---|
| dGSLM | 0.934 | 0.015 | 0.352 s |
| Moshi | 0.985 | 0.001 | 0.265 s |
| Freeze-Omni | 0.642 | 0.001 | 0.953 s |
Takeover rate here is the fraction of mid-turn pauses in which the model grabbed the floor, so lower is better; latency is how long it waits once the user genuinely finishes, so lower is better too. The two columns move together in the wrong direction. Freeze-Omni's explicit state control is three times more patient and roughly three times slower to respond. Moshi is fast because it is eager, not because it is well calibrated, and reading its 0.265 s latency without its 0.985 takeover rate would invert the conclusion.
Backchannel frequency carries the same trap. dGSLM produces fifteen times more backchannels than the others and its timing distribution is closer to human, while a model that never backchannels can score well on "did not interrupt" by being mute. Any single column of this table can be gamed by a system that does nothing.
From scenarios to conversations
Single-event probes miss the failures that accumulate. The second version of the benchmark makes an automated examiner drive a multi-turn conversation under two pacing regimes, pushing the system through daily tasks, corrections, entity tracking and safety-relevant requests, and rephrasing or repeating when the system fails to comply, which turns recovery into something measurable rather than anecdotal. Systems are rated on turn-taking fluency, instruction following and task competence, and the headline result is that duplex systems become confused under overlapping speech, handle corrections poorly, and lose track of entities across turns (Lin et al., 2025, Full-Duplex-Bench-v2, arXiv:2510.07838, ACL 2026). A third version adds tool use under real-world disfluency (Full-Duplex-Bench-v3, 2026, arXiv:2604.04847), and the ICASSP 2026 HumDial challenge contributed over 100 hours of dual-channel human-recorded Chinese and English conversation annotated for interruption and rejection scenarios, which is the data type this evaluation needs and the data type that barely existed (Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge, 2026, arXiv:2604.21406).
Reading these numbers without fooling yourself
Report distributions, not means. Response latency has a long right tail and the tail is the user experience. A median of 300 ms with a p95 of 2.4 s is a different product from a flat 500 ms.
Pair every speed metric with its error metric. Latency without false-interruption rate, or takeover rate without latency, admits a trivially optimal degenerate system. The pair is the measurement.
Synthetic probes are a lower bound on difficulty. A scripted pause is cleaner than a real one. Benchmarks built from synthesised or elicited audio will overstate performance relative to a call from a kitchen with a child in it.
Judges inherit the problem. Scoring turn-taking fluency with a language model means scoring timing through a transcript that has had the timing removed, unless the judge sees the actual two-channel audio with offsets. Check what the judge was given before trusting the score.
When it breaks
The deeper issue is that these benchmarks measure what a system does, not whether a person liked it. Preference for a conversational partner is not a monotone function of patience or speed: a model that always waits is annoying in a different way from one that always interrupts, and the optimum is task-dependent (a dictation assistant and a sales call want opposite settings). Until the field has calibrated human preference data over these axes, benchmark scores are best read as diagnostics that tell you which behaviour is broken, not as a ranking of which agent is better.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Lin et al., 2025, Full-Duplex-Bench, arXiv:2503.04721 arxiv.org
- Lin et al., 2025, Full-Duplex-Bench-v2, arXiv:2510.07838 arxiv.org
- Full-Duplex-Bench-v3, 2026, arXiv:2604.04847 arxiv.org
- Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge, 2026, arXiv:2604.21406 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.