ENDPOINT:
wss://api.openai.com/v1/realtime?model=gpt-4o-mini-realtime-preview-2024-12-17
When using the OpenAI Realtime API with output_modalities explicitly set to ["audio"], the API intermittently returns text-only responses instead of audio.
To reproduce:
-
Provide a conversation history with
inputmessages of typeoutput_text(see sample below). -
Call the API with:
{
"type": "response.create",
"response": {
"instructions": prompt,
"output_modalities": ["audio"],
"input": inputs
}
}
- Observe that the API sometimes responds with only
output_text, despiteoutput_modalitiescontainingaudio.
Example Input
[
{
"type": "message",
"content": [{ "type": "output_text", "text": "Hello" }],
"role": "assistant"
},
{
"type": "message",
"content": [{ "type": "output_text", "text": "Hi" }],
"role": "assistant"
},
{
"type": "message",
"content": [{ "type": "output_text", "text": "Hello" }],
"role": "assistant"
},
{
"type": "message",
"content": [{ "type": "output_text", "text": "Hi" }],
"role": "assistant"
},
{
"type": "message",
"content": [{ "type": "output_text", "text": "What's your name?" }],
"role": "assistant"
}
]
Example Response (unexpected)
{
"type": "response.done",
"response": {
"output_modalities": ["audio"],
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "How old are you?"
}
]
}
]
}
}
Expected Behavior
The API should always return an audio response (output_audio content) when output_modalities is set to ["audio"]. Returning text-only responses violates the expected modality setting.