Introducing GPT-Live-1 in the API

GPT-Live-1 brings ChatGPT’s natural, full-duplex voice conversations to the API. It listens and speaks simultaneously, handles pauses, interruptions, and backchannels, and adapts when a conversation changes direction.

The model manages the live conversation while delegating deeper reasoning and actions to your choice of backend model, tools, or agent framework.

Performance on real tasks

Paired with GPT-6 Astra at medium reasoning effort, GPT-Live-1 completed 83.6% of Tau3 tasks on the first attempt, versus 45.7% for GPT-Realtime-2.1. Tau3 covers airline, retail, and telecom support.

The same pairing scored 38.1% on TauBanking, which tests document retrieval and account-tool use.

Additional benchmark results
  • 97.3% on Artificial Analysis’s Conversational Dynamics benchmark.

  • 80.1% on Full Duplex Bench v1.5 interactivity.

  • 0.798-second turn-taking latency, compared with 1.41 seconds for GPT-Realtime-2.1.

  • 87% tool-calling success and 90% response quality on separate Full Duplex Bench v3 evaluations.

Voices and production features

GPT-Live-1 adds 12 real-time voices:

Quartz · Ripple · Vesper · Willow · Stone · Gleam · Meridian · Bossa · Tempo · Beacon · Delta · Cinder

The expanded selection covers more accents, dialects, and languages. System prompts can shape tone, pace, speaking style, and conversational behavior.

Listen to the new voices

For production applications, the model provides native ASR transcripts and response text, keyword biasing, explicit turn detection, and improved handling of silence and background noise.

Connect using WebRTC for browsers, WebSockets for server-side audio, or Telephony and SIP for phone agents.

Delegation and Codex

Use managed Responses delegation or connect an existing model, agent, or service through client delegation. Your application controls permissions and durable task state; interrupting speech does not automatically cancel backend work.

You can also pair GPT-Live-1 with the Codex SDK. Your app passes conversation context to a Codex thread and returns the result to the active voice session, allowing Codex to investigate a repository or complete delegated work while the conversation continues.

See the GPT-Live-1 and Codex integration

Pricing

Voice sessions cost $0.05 per minute, billed per second. Backend model and tool usage is billed separately.

Read the launch post · Get started with GPT-Live

Building with GPT-Live-1? Share what you learn about interruptions, delegation, and production tool use.

OpenAI has released full duplex GPT-Live-1 on the API with pricing at 5c/minute. As someone who loves building voice assistants, this is very exciting (full duplex models are amazing).

There are hints in the announcement: “GPT Live backend: Astra (medium).”

Multimodal finally powered by a newer underlying model than training of gpt-4o?

@platypus out of interest, are you using a telco?

No just self hosting on my home server and connecting via WebRTC. I’ve only been using local models so far (tried ElevenLabs and some other services in the past but I can get very similar quality with Fish Audio hosted at home, with unlimited speaking time). BUT…it aint full duplex :wink:

O que está faltando é esse Canal de Voz conseguir executar as Ferramentas e Pluings como o chat normal.

Dear @sps

Hope you well. Question for you.

I am currently using the gpt realtime api 2.1. And now we finished our test application on gpt live 1. Our tests proved successful in every possible way so now I have a question:

Should we Migrate completely to Gpt Live 1 or should I Add GPT Live 1 and offer both realtime 2.1 and gptlive 1?

To give you some context, yes, we see that gpt live 1 is better at what we need than realtime because of the following factors:

Gpt Live 1 is better at tool calling, charged per second and not rounded, faster, and supports interruptions. All these we need.

So what would openai advise? and what is going to happen with the realtime 2.1? Why would openai Keep realtime 2.1 when gptlive 1 is so much superior?

Thanks so much and great work to the team

God bless

Unfortunately, the system didn’t perform well for us across all our criteria, for the following reason:
Unlike Realtime, we can’t retrieve the conversation transcript because GPT-Live doesn’t transcribe the user’s voice. The recorded transcript only contains the agent’s messages, not the user’s words. It seems that OpenAI doesn’t currently offer this functionality via its API, which is very inconvenient. We can only retrieve the LLM’s response, which isn’t exactly what GPT Live says. I think this is related to the dual architecture (Live + LLM), but this limitation prevents us from considering a production deployment, even though the conversations are truly amazing.

Shipped a real build on this in the first week: a voice-driven system design mock interviewer that watches an Excalidraw whiteboard while the candidate draws.

A few protocol notes that might save other builders time, all verified against a live session:

  • The voice layer can’t take images, so the diagram reaches the model two ways: a compact text summary via session.thinking.append on each drawing pause (kept under the 500-token cap), and a JPEG pushed to the delegated Responses backend via response.item.create when the backend calls a view_whiteboard tool, followed by response.create. Verified with a visual-only sentinel: the backend read a two-character code that existed only in the image.
  • WebRTC data channel: the negotiated pc.sctp.maxMessageSize came back 262144. A 1024px JPEG of the board was ~9 KB as a message. dc.send throws on oversize; if you swallow the exception the image path fails silently.
  • Send session.close and wait for session.closed before tearing down the peer connection. Ack took ~1.3 s in my test. With store: true you get the stereo WAV back for replay.
  • Put anything the voice layer needs to answer instantly (in my case a hidden fact sheet for clarifying questions) in the live instructions rather than only in the backend. Otherwise every quick question round-trips through delegation.

Happy to compare notes with anyone else building on the delegation model.

GitHub: jkhoffman/system-design-coach

demo

We get the user’s side in real time over the data channel: session.input_transcript.delta carries their speech with start_ms/end_ms, alongside session.output_transcript.delta for the agent. We persist both ourselves (grouped by speaker and a ~1.5 s gap) and pair them with the stereo WAV from store: true (user left, agent right) so the transcript can be checked against the audio. If you’re pulling the transcript after the fact from the stored session, the live deltas are the workaround.


Thanks for the insights on transcript handling! We’re doing something similar — persisting both session.input_transcript.delta and session.output_transcript.delta client-side. The ~1.5s gap grouping with start_ms/end_ms is a smart approach, we’ll adopt that for better segmentation.

Since we’re mid-migration from Realtime to GPT-Live, here are the main gotchas we’ve hit in case it helps others:

Delegation split — The biggest paradigm shift. The voice model talks to the user but has no tools; the backend model has the tools but no voice. The voice model can respond directly without delegating (e.g., to a “goodbye”), which means tool-dependent flows like end_conversation simply never fire unless you build workarounds on the client side.

response.done output is always empty — In Realtime, function calls come through response.done.output. In GPT-Live delegation mode, output is consistently []. Function calls arrive individually via response.output_item.done wrapped inside response.event envelopes. Any logic that relied on inspecting response.done.output needs to be rewritten.

Double response.completed events — Each delegation turn fires both an inner (delegation-level) and an outer (session-level) response.completed. If you have state machines counting these events, they’ll advance twice. We had to tag inner events with a _fromDelegation flag to distinguish them.

No session.update / response.cancel / input_audio_buffer.clear — In Realtime you can override instructions mid-session, cancel responses, and clear the audio buffer. None of these seem available in GPT-Live via the data channel, so goodbye flows and interruption handling need different patterns.

Voice model commentary — The voice model can speak (fillers, brief responses) before session.delegation.created fires. This breaks any turn-detection logic that assumes “voice spoke = delegation happened.”

Would love to hear if others have found cleaner workarounds for any of these, especially the end_conversation / graceful shutdown flow.


@fred5 we’re also having a lot of trouble ending the conversation. With gpt-realtime we have a finish_session tool that the assistant calls when the conversation is done, and we then hang up the SIP call when the assistant finishes its goodbye speech.

With gpt-live, the frontend often just says “goodbye” or “if you need anything else, I’m here” and never delegates to the backend… so the hangup never happens. Same thing you’re experiencing. We’ve tried adjusting the prompt and it improved a bit, but it’s nowhere near as reliable as with gpt-realtime.

In general, after a week of testing, we feel like the frontend/backend split makes the assistant more capable, but the conversation is more unpredictable and less suitable for semi-rigid flows. The frontend sometimes doesn’t delegate when it should, or the preambles get ahead of the backend when it does. It’s like the frontend lacks context. This can be improved with prompting up to a point, but the more we tweak it to get the desired behavior, the more the voice degrades.

Would be great to hear if anyone successfully transitioned a gpt-realtime assistant that relied on a “# Conversation Flow” into gpt-live, and especially if anyone has found a clean way to end the conversation gracefully.

I’ve tested this quite extensively over the past few days, and it’s easily one of the most impressive features I’ve tried so far — even more so than Astra, although I haven’t been able to test Astra nearly as much because of the token usage.

The latency and overall quality are remarkable. I only noticed very occasional dialect-related issues, but what impressed me most is how artfully the model seems to have been designed. The interaction feels remarkably humanlike — not just fast or technically polished, but natural in its timing, reactions, tone, and conversational flow.

At the same time, you can still notice the limits in intelligence. In some situations the reasoning feels weaker, especially !!! compared with the earlier setup where you could explicitly choose a higher reasoning level.

Another issue (as info) I noticed is that, when it starts looking for information or when a conversation becomes longer, it can sometimes reach a point where it no longer seems able to progress properly. It feels as if the conversation gets stuck and the model cannot really continue from there. I’m not sure whether this is an intentional limitation, a context-management issue, or simply a bug, but it happened often enough for me to notice it.

You are meant to pair it with a reasoning model for which you can configure the reasoning level - can you elaborate what you mean?

After several days of testing migrating a RealTime 2 application to GPT-Live-1, here are my comments. I hope they help:
Positive points:

  • The human-sounding voice quality is a real improvement; the progress compared to RealTime 2 is quite impressive.

  • The background noise filtering is excellent, as this was a real problem with the RealTime application.

Negative point:

  • The voices: We conduct most of our tests in French, and only two voices are free of an English or Canadian accent: Marin and Cedar, which is not many. But the real problem is that it’s inconsistent; the quality of conversations is uneven: sometimes the French is accent-free, sometimes it has an accent, or it deteriorates during the conversation.

Regarding user interruptions during conversations, I haven’t noticed any significant improvement compared to Realtime-2.

The main difficulty encountered is the increased complexity of managing conversations with GPT-Voice on the client side and LLM on the backend. Migrating from Realtime-2 required us to completely rewrite our module that manages and directs the conversation according to predefined rules, particularly at the beginning and end of the conversation, as well as the transcription of the conversation for later use.

Advice on how to better manage the conversation workflow would be appreciated:
The agent introduces themselves and starts the conversation → the conversation proceeds → the agent detects the end of the conversation, concludes, and ends the session.

The problem is that LLMs aren’t designed to end a conversation, but rather to always restart it. That’s not the case in real life :slight_smile:

Congratulations on the work done; we eagerly await the next version, which I hope will better handle accents based on the languages ​​spoken.

Cheers

Yeah, the voices are a bit odd - it’s a bit “generative random” sometimes - I swear the voice is sometimes quite different between calls and even in British English or Irish English you never seem to get a perfectly “indigenous” accent, there’s often a tinge of North American, lol.