Cascaded Versus Speech-Native Voice Pipelines
What a recognition-generation-synthesis cascade actually costs in latency and lost acoustic information, what a speech-native model gives up in measured knowledge and controllability, and the hybrid designs that try to keep both.
Add up a 2026 voice stack and the number is uncomfortable. Streaming recognition finalising in roughly 250 ms, a fast model's first token at around 150 ms, synthesis producing first audio at about 90 ms, and 50 ms of orchestration, before any turn-taking wait, lands near 540 ms of round trip, against ITU-T G.114's recommendation of 150 ms one-way mouth-to-ear as the comfortable target and 400 ms as the tolerable limit for interactive speech. Vendors publish component figures in the same range, with first-byte synthesis latencies advertised between roughly 40 and 90 ms, and those are vendor-reported best cases rather than measurements under load.
The cascade is not slow because any stage is badly built. It is slow because the stages are sequential and because the first stage cannot begin until the turn is declared over.
What each stage throws away
Recognition converts a waveform into a string. Everything not in the string is gone: pitch, speaking rate, hesitation, laughter, emphasis on the third word, the shake in the voice of a customer who has called three times. Synthesis then invents all of it again from scratch, guided by punctuation. Two lossy conversions sit in the middle of a medium whose point is that it carries more than words.
A speech-native model skips both conversions. It also inherits a different problem, now measured rather than suspected. S2SBench quantifies the intelligence degradation of speech-to-speech models against their text counterparts and attributes it to three causes: audio tokens carry less meaning per token, sequences get much longer, and prosody and speaker identity add variance the model must absorb (Fang et al., 2025, S2SBench, arXiv:2505.14438). The reported gaps are large. Adding speech modelling to SpiritLM moved MMLU from 45.3 to 36.9; GLM-4-Voice answers around 37 percent of VoxEval against a text counterpart above 60 percent on MMLU; some end-to-end systems show gaps above 35 points between text and speech input on the same questions (Zeng et al., 2024, GLM-4-Voice, arXiv:2412.02612). Fixed capacity spent on acoustics is capacity not spent on knowing things.
The comparison that matters per product
| Property | Cascade | Speech-native |
|---|---|---|
| Floor latency | Sum of stages plus endpointing wait | One frame of the model's rate |
| Paralinguistics | Lost at recognition, re-invented at synthesis | Preserved end to end |
| Measured knowledge | The text model's, unchanged | Degraded relative to the same-size text model |
| Overlapping speech | Needs an external policy | Representable natively |
| Tool calls and grounding | Ordinary text tooling | Needs a text stream or a tandem back end |
| Debuggability | Transcript at every boundary | Weights, plus whatever text stream exists |
| Idle cost | Near zero | One forward pass per frame |
Hybrids are where most production systems actually sit
Freezing the language model and attaching speech encoders and decoders around it keeps text-model competence and adds speech behaviour, at the cost of the acoustic signal never influencing the reasoning itself (Wang et al., 2024, Freeze-Omni, arXiv:2411.00774). A tandem arrangement goes further: a fast speech-to-speech front end holds the conversation while a text model answers the parts that need knowledge, which is the explicit design of KAME (KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI, 2025, arXiv:2510.02327). Both are admissions that the two architectures fail in complementary places.
The commercial realtime APIs are the same bet in a product wrapper, exposing a speech-to-speech model with text-side tooling, function calling and transcripts alongside the audio, so that an application can read and steer what the model is doing. OpenAI's Realtime API reached general availability with a dedicated speech-to-speech model in August 2025 after nearly a year in beta.
When it breaks
The cascade's latency is not uniformly distributed. The mean looks fine and the tail does not. A long recognition finalisation on one utterance, or a slow first token under load, produces the two-second silence that users remember from an otherwise fast call.
Speech-native models fail on the parts you did not benchmark. Voice benchmarks that score instruction following and dialogue quality can still miss a knowledge collapse that a text benchmark would catch immediately (VoiceBench: Benchmarking LLM-Based Voice Assistants, 2024, arXiv:2410.17196). Run both.
Paralinguistics preserved is not paralinguistics used. A model that hears frustration and responds in the same cheerful tone has kept the signal and ignored it. The capability needs its own evaluation, not an architectural claim.
Cost structure differs more than quality does. Audio tokens are numerous and priced differently from text, and a duplex model bills for silence. For a given call volume, the architecture choice can change the bill by a larger factor than it changes any quality metric, which is worth knowing before the quality argument starts.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Fang et al., 2025, S2SBench, arXiv:2505.14438 arxiv.org
- Zeng et al., 2024, GLM-4-Voice, arXiv:2412.02612 arxiv.org
- Wang et al., 2024, Freeze-Omni, arXiv:2411.00774 arxiv.org
- KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI, 2025, arXiv:2510.02327 arxiv.org
- VoiceBench: Benchmarking LLM-Based Voice Assistants, 2024, arXiv:2410.17196 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.