Definition
Speech-to-speech AI models process audio input and generate audio output natively, bypassing cascaded ASR→LLM→TTS pipelines. This reduces latency and preserves conversational prosody for real-time voice agents.
Key Points
- 2026-09: gemini-38-live Extended Thinking ranks #1 on Artificial Analysis Speech-to-Speech index (82.6)
- Supports background reasoning, async tool calls, visual grounding during live sessions
- Competes with OpenAI GPT-Live and enterprise voice agent stacks
- 97-language mid-conversation switching without session restart