Realtime API sometimes returns text-only responses even when output_modalities is set to audio

ENDPOINT:

wss://api.openai.com/v1/realtime?model=gpt-4o-mini-realtime-preview-2024-12-17

When using the OpenAI Realtime API with output_modalities explicitly set to ["audio"], the API intermittently returns text-only responses instead of audio.

To reproduce:

  1. Provide a conversation history with input messages of type output_text (see sample below).

  2. Call the API with:

{
  "type": "response.create",
  "response": {
    "instructions": prompt,
    "output_modalities": ["audio"],
    "input": inputs
  }
}

  1. Observe that the API sometimes responds with only output_text, despite output_modalities containing audio.

Example Input

[
  {
    "type": "message",
    "content": [{ "type": "output_text", "text": "Hello" }],
    "role": "assistant"
  },
  {
    "type": "message",
    "content": [{ "type": "output_text", "text": "Hi" }],
    "role": "assistant"
  },
  {
    "type": "message",
    "content": [{ "type": "output_text", "text": "Hello" }],
    "role": "assistant"
  },
  {
    "type": "message",
    "content": [{ "type": "output_text", "text": "Hi" }],
    "role": "assistant"
  },
  {
    "type": "message",
    "content": [{ "type": "output_text", "text": "What's your name?" }],
    "role": "assistant"
  }
]

Example Response (unexpected)

{
  "type": "response.done",
  "response": {
    "output_modalities": ["audio"],
    "output": [
      {
        "type": "message",
        "role": "assistant",
        "content": [
          {
            "type": "output_text",
            "text": "How old are you?"
          }
        ]
      }
    ]
  }
}

Expected Behavior
The API should always return an audio response (output_audio content) when output_modalities is set to ["audio"]. Returning text-only responses violates the expected modality setting.