# New Realtime Voice Models in the API

**URL:** <https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471>\
**Category:** Announcements\
**Tags:** gpt-realtime\
**Created:** [May 7, 2026, 7:49pm UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471 "2026-05-07T19:49:27Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![VeitB](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/veitb/32/712987_2.png) [@VeitB](https://community.openai.com/u/VeitB)\
**Post date:** [May 7, 2026, 7:49pm UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/1 "2026-05-07T19:49:27Z")

</div>

OpenAI announced a new generation of realtime voice models for the API:

[Advancing voice intelligence with new models in the API](https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/)

The release includes three new audio models for building more capable voice experiences:

[![realtime2](https://us1.discourse-cdn.com/openai1/original/4X/4/a/f/4af93a85e803a3dc890ec952045383f209283680.jpeg)](https://www.youtube.com/watch?v=JOu8v6CBjkE)

- **GPT-Realtime-2** , a voice model with GPT-5-class reasoning for harder requests, better context handling, and more natural conversations.
- **GPT-Realtime-Translate** , a live translation model that supports speech from 70+ input languages into 13 output languages.
- **GPT-Realtime-Whisper** , a new streaming speech-to-text model for live transcription while someone is speaking.

The larger shift here is that realtime voice is moving beyond simple call-and-response. These models are aimed at voice agents that can listen, reason, translate, transcribe, use tools, and take action while the conversation is still unfolding.

Some of the use cases OpenAI highlights include voice-to-action workflows, live spoken guidance from software, and voice-to-voice conversations across languages.

 ![Audio MultiChallenge Instruction Following](https://us1.discourse-cdn.com/openai1/original/4X/d/7/3/d7322097a64dbfecb75d4d47f9411d4535235a57.png)  
 ![Big Bench AudioIntelligence](https://us1.discourse-cdn.com/openai1/original/4X/9/1/7/917dc59b06d893fc70aab11ceb75ae1c3f5454b6.png)

[Definitely check out the official announcement. It includes more context and examples of what these new models can do.](https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/)

---

<div class="post-metadata">

**Author:** ![VeitB](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/veitb/32/712987_2.png) [@VeitB](https://community.openai.com/u/VeitB)\
**Post date:** [May 7, 2026, 8:42pm UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/3 "2026-05-07T20:42:13Z")

</div>

Additional Documentation and Guides related to this release:

- [Realtime and audio](https://developers.openai.com/api/docs/guides/realtime)  
Updated overview for choosing between voice agents, realtime translation, realtime transcription, and request-based audio APIs. It explicitly routes low-latency voice agents to `gpt-realtime-2`.

- [Using realtime models](https://developers.openai.com/api/docs/guides/realtime-models-prompting)  
New/updated prompting guide for `gpt-realtime-2`, including reasoning effort, preambles, tool policies, unclear audio handling, exact entity capture, and long-session behavior.

- [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents)  
Updated guide for building speech-to-speech agents with `RealtimeAgent` / `RealtimeSession`, WebRTC, tools, handoffs, and guardrails.

- [Realtime translation](https://developers.openai.com/api/docs/guides/realtime-translation)  
Dedicated guide for `gpt-realtime-translate`, including `/v1/realtime/translations`, WebRTC/WebSocket patterns, listen-along translation, conversational translation, and production checklist.

- [Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)  
Dedicated/refreshed guide for `gpt-realtime-whisper`, streaming transcript deltas, latency/accuracy tuning, vocabulary guidance, and production checklist.

- [Realtime with tools](https://developers.openai.com/api/docs/guides/realtime-mcp)  
Guide for function tools, remote MCP servers, and built-in connectors in Realtime sessions with `gpt-realtime-2`.

- [gpt-realtime-2 model page](https://developers.openai.com/api/docs/models/gpt-realtime-2)

Pricing info:

- `GPT-Realtime-2`: `$32 / 1M` audio input tokens, `$0.40 / 1M` cached input tokens, `$64 / 1M` audio output tokens
- `GPT-Realtime-Translate`: `$0.034 / minute`
- `GPT-Realtime-Whisper`: `$0.017 / minute`

---

<div class="post-metadata">

**Author:** ![LarisaHaster](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/larisahaster/32/761614_2.png) [@LarisaHaster](https://community.openai.com/u/LarisaHaster)\
**Post date:** [May 7, 2026, 8:52pm UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/4 "2026-05-07T20:52:38Z")

</div>

I already did a quick test with gpt-realtime-2 using my Cheeky-Razor persona.

The audio output felt more natural than with previous models: clearer, better paced and easier to follow.

I tested it in English and as a non-native English speaker, I found it easy to understand without having to concentrate too hard, which is actually a very useful improvement​🙌

---

<div class="post-metadata">

**Author:** ![sam.saffron](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/sam.saffron/32/31387_2.png) [@sam.saffron](https://community.openai.com/u/sam.saffron)\
**Post date:** [May 8, 2026, 3:39am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/5 "2026-05-08T03:39:55Z")

</div>

I just tried this out, it feels great

Some good news is that this is available on codex plans, I can mint a webrtc token and use it! (At least the 200 dollar one for me)

Hebrew though is no good sadly, heavy annoying accent, English is flawless

---

<div class="post-metadata">

**Author:** ![sam.saffron](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/sam.saffron/32/31387_2.png) [@sam.saffron](https://community.openai.com/u/sam.saffron)\
**Post date:** [June 6, 2026, 6:39am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/6 "2026-06-06T06:39:11Z")

</div>

Update, gpt-realtime-2 no longer works via codex account, you can mind keys but when using them you will get a 500 error. Not sure if this is on purpose or not, but as it stands today to use it you need to use API direct and not codex route.

---

<div class="post-metadata">

**Author:** ![LarisaHaster](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/larisahaster/32/761614_2.png) [@LarisaHaster](https://community.openai.com/u/LarisaHaster)\
**Post date:** [June 6, 2026, 6:46am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/7 "2026-06-06T06:46:46Z")

</div>

According to VB, yesterday when it was OpenAI’s developers office hour, he did mention gpt-realtime being available in Codex. So it should work normally.

I can test that a little later and see if I can reproduce the 500 error.

---

<div class="post-metadata">

**Author:** ![LarisaHaster](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/larisahaster/32/761614_2.png) [@LarisaHaster](https://community.openai.com/u/LarisaHaster)\
**Post date:** [June 6, 2026, 9:40am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/8 "2026-06-06T09:40:36Z")

</div>

I may have misunderstood what you mean by “Codex route,” so please correct me if I’m testing the wrong thing here.

I tested this from Codex with a project API key created through the OpenAI Platform / Developers flow, calling the GA Realtime endpoint directly: `wss://api.openai.com/v1/realtime?model=gpt-realtime-2`

I couldn’t reproduce the 500 error. `gpt-realtime-2` returned successfully for me:

```auto
json

{

  "ok": true,

  "model": "gpt-realtime-2",

  "transcript": "realtime 2 route ok",

  "sawSessionCreated": true,

  "sawResponseDone": true,

  "error": null

}

```

One thing I noticed: the old beta header `OpenAI-Beta: realtime=v1` now fails and `response.modalities` also seems outdated for the GA schema.

My test only confirms that the direct Platform API route works for me.

* * *

I also hope I understood yesterday’s OAI Developer Office Hours correctly. VB mentioned this quickly and it was a little hard to follow as a non-native English speaker when that part went fast. Hopefully someone else can correct or confirm me here.

---

<div class="post-metadata">

**Author:** ![sam.saffron](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/sam.saffron/32/31387_2.png) [@sam.saffron](https://community.openai.com/u/sam.saffron)\
**Post date:** [June 6, 2026, 10:19am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/9 "2026-06-06T10:19:12Z")

</div>

Sorry what I meant is that if you use the oauth route via the 100/200 dollar monthly codex plans the ephemeral key can be cut fine but the ephemeral key I get is not usable, it was usable 24 hours ago so this is a new change

I tried the metered pay as you go key and it works 100% fine

---

<div class="post-metadata">

**Author:** ![LarisaHaster](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/larisahaster/32/761614_2.png) [@LarisaHaster](https://community.openai.com/u/LarisaHaster)\
**Post date:** [June 6, 2026, 10:20am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/10 "2026-06-06T10:20:20Z")

</div>

Ah thanks for clarifying!

---

<div class="post-metadata">

**Author:** ![sam.saffron](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/sam.saffron/32/31387_2.png) [@sam.saffron](https://community.openai.com/u/sam.saffron)\
**Post date:** [July 25, 2026, 4:52am UTC](https://community.openai.com/t/new-realtime-voice-models-in-the-api/1380471/11 "2026-07-25T04:52:31Z")

</div>

Some good news, codex plan now includes gpt-realtime-2.1 … I can mint tokens and use them, a bit unclear what the limit is, but this is super handy on my iPhone where I can quickly ask for a song, it can search apple music and play something.
