ZAX ZAX
AI Models 5 min read

OpenAI Launches GPT-Live: Voice Models That Listen and Speak in Real-Time

Eric Leroy
OpenAI Launches GPT-Live: Voice Models That Listen and Speak in Real-Time

OpenAI has launched GPT-Live, a new family of voice models capable of listening and speaking simultaneously. With GPT-Live-1 and its lightweight version GPT-Live-1 mini, AI conversations finally lose that robotic feel of pauses and latencies. Natural interruption becomes possible: you can cut off the AI mid-sentence, just as you would with a human interlocutor.

The End of Turn-Based Conversations

Until now, voice assistants operated in "walkie-talkie" mode: the user speaks, the AI waits for the end, processes, then responds. This delay, even reduced to a few hundred milliseconds, is enough to make the conversation feel artificial. With GPT-Live, the model processes audio continuously, enabling real-time exchanges where the AI can be interrupted, relaunched, or respond before the user has finished their sentence.

In practice, this means a voice assistant can now handle complex conversations: negotiate an appointment, take an order with modifications, answer nested questions. Use cases in customer support, phone ordering, or internal assistance become significantly smoother.

Two Models for Two Uses

GPT-Live-1 is the full model, optimized for sophisticated conversations requiring advanced reasoning. It's suited for high-end assistants: virtual concierges, level 2 technical support, executive assistants. The cost per minute of conversation is higher, but the response quality justifies the investment in these premium use cases.

GPT-Live-1 mini targets high-volume uses where speed takes precedence over reasoning depth: appointment booking, order confirmation, voice FAQs. The reduced cost allows deploying voice assistants at significant volumes without blowing the budget.

Impact for Businesses

For companies considering automating their phone reception or deploying internal voice assistants, GPT-Live changes the game. Experience quality finally approaches what users expect from a phone conversation. Abandonment cases linked to frustration with slow, rigid voice systems should decrease significantly.

The API is available now, with pricing based on conversation duration rather than token count. For companies that already have OpenAI integrations, adding real-time voice capabilities naturally fits into the existing architecture.

This announcement is part of a broader trend: conversational AI is leaving text behind to become multimodal. Tomorrow's business chatbots will speak as much as they write. To explore how these new capabilities can integrate into your processes, our AI audit identifies relevant use cases for your context.

Related Articles