Introducing GPT-Live

July 8, 2026

We’re launching GPT-Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.

GPT-Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT-Live can show it’s paying attention with phrases like “mhmm” or “yeah”, engage in quick back-and-forth, or just stay quiet when you need a moment to think.

GPT-Live is also our smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the result back into the conversation when it’s ready. While it works, GPT-Live can keep talking with you and maintain the flow of conversation. At launch, GPT-Live will use GPT-5.5 in the background. As we release new frontier models, we’ll continuously update the model used by GPT-Live.

We’re beginning to roll out two versions of GPT-Live – GPT-Live-1 and GPT-Live-1 mini – to ChatGPT users globally today. We also plan to bring them to the API soon.

Previous approaches

Older generations of voice AI systems brought us closer to that vision, but with important tradeoffs.

Cascaded voice systems

Cascaded voice systems rely on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe your speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled us to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted.

Turn-based voice models

Turn-based voice models like ChatGPT Advanced Voice Mode processed and generated audio within a single model, reducing latency and making conversations smoother — but they still operated through discrete turns. The model had to wait for the user to stop speaking before responding, resulting in rigid back-and-forth.

Our new approach

GPT-Live addresses these limitations through two architectural changes.

Continuous interaction

First, we built GPT-Live for continuous interaction using a full-duplex architecture. Instead of processing a sequence of separate messages, GPT-Live continuously processes input while generating output. The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool.

This allows the model to engage in more natural back-and-forth, maintain a better sense of time, and even perform live translation.

Delegation for deeper work

Second, we decoupled GPT-Live — which handles continuous interaction — from deeper work. When a question requires search, reasoning, or more agentic capabilities, GPT-Live can delegate the task to another model like GPT-5.5. This allows it to keep the conversation going, even as it handles multiple tasks in the background.

Evaluations

We built new human evaluations to measure pleasantness and the flow of conversation. In head-to-head comparisons, GPT-Live-1 and GPT-Live-1 mini are strongly preferred over Advanced Voice Mode in matched 5–10 minute conversations.

  • GPQA: GPT-Live-1 substantially outperforms Advanced Voice Mode on GPQA, which tests expert-level scientific reasoning across biology, chemistry, and physics.
  • BrowseComp: GPT-Live-1 shows strong gains over Advanced Voice Mode on BrowseComp, which tests agentic web search and the ability to find difficult-to-locate information.
  • τ³-Voice Telecom (internal variant): GPT-Live-1 outperforms Advanced Voice Mode on τ³-Voice Telecom, which tests voice agents on realistic, multi-turn telecom support tasks.

A new ChatGPT Voice experience

Each week, more than 150 million people talk to ChatGPT using features like Voice and Dictation.

Starting today, when you tap the Voice button to talk with ChatGPT, you’ll get an improved experience powered by GPT-Live—with more natural conversations, smarter answers, better listening, and visual responses.

More natural conversations

You can interrupt with a question, pause to gather your thoughts, or ask ChatGPT to slow down. It naturally acknowledges what you’re saying with phrases like “mhmm” or “got it.” We’ve also remastered the nine distinct voices in ChatGPT for GPT-Live.

Smarter answers

ChatGPT Voice can now draw on our latest frontier models. You can choose the level of reasoning: Instant for fast responses, or Medium and High when you want ChatGPT to spend more time thinking.

Better listening

If you take a moment to think, ChatGPT Voice now waits instead of jumping in and interrupting. If you ask it to stay quiet and listen, it will. ChatGPT is better at focusing on your voice instead of getting distracted by background noise.

Visual answers at a glance

While you’re talking, ChatGPT can now show rich visual cards for topics like weather, stocks, sports, and more. Voice also continues to support search, memory, images, and file uploads.

Safety designed for voice

GPT-Live was designed to be safe by default. It builds on the safety advances from our latest models while adding dedicated safety training across key risk areas and new safeguards designed specifically for voice.

We expanded safety testing to include new audio-native evaluations and synthetic evaluations focusing on self-harm, psychosis and mania, emotional reliance on AI, violence, and sexual content. In testing, GPT-Live performed comparably to or better than Advanced Voice Mode across nearly all areas evaluated.

Because voice conversations unfold in real time, we built safeguards that can act while the model is speaking. When the system detects potentially unsafe output, it can steer the model toward a safer response, surface additional safety messaging or resources, or end the voice conversation in higher-risk cases.

We designed additional protections to support teen users, and trained age-appropriate behavior directly into the model. Parents can choose whether their teen can use ChatGPT Voice through Parental Controls.

GPT-Live is designed for conversation, not voice impersonation. It uses a set of predefined voices in ChatGPT, with safeguards to prevent it from imitating a real person’s voice.

Availability & limitations

GPT-Live is rolling out now to ChatGPT users globally across iOS, Android, and ChatGPT.com. GPT-Live-1 will become the default model powering ChatGPT Voice for Go, Plus, and Pro users, and GPT-Live-1 mini will become the default for Free users.

At launch, GPT-Live will not support voice with video or screen sharing in ChatGPT, but we’re working to introduce these capabilities soon. You can still access legacy versions of ChatGPT Voice, including Standard and Advanced Voice Mode, where these features are available.