Definition

Speech-to-speech AI models process audio input and generate audio output natively, bypassing cascaded ASR→LLM→TTS pipelines. This reduces latency and preserves conversational prosody for real-time voice agents.

Key Points

  • 2026-09: gemini-38-live Extended Thinking ranks #1 on Artificial Analysis Speech-to-Speech index (82.6)
  • Supports background reasoning, async tool calls, visual grounding during live sessions
  • Competes with OpenAI GPT-Live and enterprise voice agent stacks
  • 97-language mid-conversation switching without session restart

Sources