Important Feedback/improvements for gpt-live-1 API

We’re building production voice agents on gpt-live-1. These are the issues that matter most to us, starting with the most important one.

  1. session.instructions.append is unreliable (top priority)

This is the most important primitive for building real voice agents. We use it to steer the model mid-call with tool results, next steps, and required disclosures. Right now it is not reliable at all, and it’s worst when the instruction asks the model to say something. The model often ignores the instruction or only partly follows it.

session.commentary.append works sometimes, but it’s inconsistent too and we can’t use it for important instructions.

Request:

Guarantee that appended instructions are followed, especially “speak this / say this” instructions.
Document the expected behavior and timing, meaning when an appended instruction takes effect relative to the current turn.

  1. There’s no way to send instructions without interrupting the model

session.instructions.append currently cuts the model off mid-speech. Most mid-call guidance isn’t urgent enough to justify that. It sounds jarring to the caller and breaks the flow of the conversation.

Request: Add a non-interrupting delivery mode, for example interrupt: false or mode: “after_current_utterance” | “immediate”. With it, the model finishes its current thought and then moves smoothly into the new instruction. This is critical for us. Reliable delivery plus a no-interrupt option would unlock most of our production use cases.

  1. The initial greeting is inconsistent

The first greeting doesn’t always follow the configured instructions, and in some sessions the voice sounds distorted or wrong on the first utterance.

Request: Make the first response follow the session config deterministically, with a stable voice from the very first audio frame.

  1. The voice sometimes switches mid-sentence (rare but severe)

In rare cases the output voice switches mid-sentence. For example, we configured a female voice and it changed to a male voice partway through a sentence.

Request: Lock the voice to the one configured for the session. Even rare switches break caller trust in production calls.

I hope this will be solved particularly the session.instruction.append.

Thank You! Open to share more info if needed.

Hi akarsh, until append is reliable, here’s how I’d make sure a missed instruction can’t go unnoticed on a live call:

  1. Verify instead of assuming. After each appended “say X” instruction, check the output transcript of the next turn or two for a few key words of the disclosure. If they’re missing, you get a logged, countable failure instead of a silent one, and you can re-send or escalate.
  2. Take must-say content away from the model. For the greeting and legal disclosures, play fixed audio from your telephony layer before you connect the caller to the live session. It sounds the same every time and can’t be paraphrased or skipped.
  3. Measure the follow rate. Script 30–50 synthetic calls per scenario (greeting, mid-call instruction, disclosure after a tool result) and track how often the instruction was actually spoken. OpenAI’s voice agents guide recommends the same kind of fixed synthetic-speech runs. That number turns “not reliable” into something you can show OpenAI and re-test after each model update.

Adam

Welcome to the OpenAI Developer Forum, and thanks for sharing such detailed feedback!

What you’re describing is especially important for production voice agents, particularly the behavior of session.instructions.append.

According to the current GPT-Live documentation, the session.instructions.appended event confirms that an instruction has been accepted into the session timeline, but it does not necessarily guarantee that the model has acted on that instruction. The documentation also notes that appending instructions while the model is speaking can interrupt the current response.

Here are a few relevant resources:

GPT-Live conversation and session management

GPT-Live prompting guide

GPT-Live delegation and tools

Live API reference

Your suggestion for something like:

interrupt: false

or

mode: “after_current_utterance”

would address an important production use case: allowing developers to provide new instructions during a call without abruptly cutting off the current response.

For the initial greeting and especially the rare mid-sentence voice-switching issue, it would also be helpful to provide session/request IDs, timestamps, transport type (WebRTC or WebSocket), SDK/version, and a minimal reproduction when possible. Please make sure any caller or other sensitive information is removed before posting.

Thanks again for documenting these issues so clearly I added feature-request to your tag string. The additional examples you offered could be particularly useful for reproducing the session.instructions.append behavior.