Gpt-live-1 responses delegation: voice model sometimes never delegates (acknowledges verbally, sends nothing)

We’re running gpt-live-1 over WebSocket (/v1/live/sessions, audio/pcmu @ 8kHz, telephony) with responses-mode delegation (gpt-5.6 backend, tool_choice: auto, tools for transfer/hangup/knowledge lookups).

Intermittently the voice model doesn’t delegate at all. It verbally commits — caller: “can you hang up the call”, model: “Sure, I’ll end this call now. Goodbye!” — but nothing arrives on the wire for that request: no session.delegation.created, no response.event, no error. We log every server event at debug, so this is absence, not a handling gap.

What we’ve established:

  • Happens on clean, new sessions — including the very first request after the greeting.
  • Request-level, not session-level: one session delegated a transfer correctly, then ignored a hangup request two minutes later.
  • Delegation config is echoed back in session.started on every failing session.
  • Not our changes: with a frozen client config we saw ~50% failures one morning, 18/18 success from ~09:00 UTC that same day, then regression overnight — coinciding with the session ID format changing (longer suffix) and the model starting to paraphrase a greeting our instructions mark “do not vary”, which it had honoured all the previous day.

Prompting follows the delegation-policy template from the prompting guide, plus the migration guide’s voice/backend prompt split.

Anyone else seeing delegation skip requests like this? And has anyone actually received session.delegation.created in responses mode? We have paired session IDs (fail vs success, same config, seconds apart) if anyone from OpenAI wants traces.

Hi Jiayu from OpenAI here.

high levelly you should put a delegation policy in your frontend model first, something like:

“when you hear end the call or bye bye, you should delegate”

you don’t need to ask the frontend model (gpt-live-1) to understand which tool to call. it only needs to understand when to delegate.

then your backend model, which is way smarter can understand how to call the tools to end the call.

Hi Jiayu,

Have faced some issues with the session.instructions.append and session.commentary.append.

for instructions.append, sometimes it doesn’t follow the instructions particularly in cases where the model is expected to initiate, like if I send 'Ask the user…" it doesn’t follow it and remains silent instead. Also another annoying thing is that sending instructions interrupts the flow, ideally there must be a way to send instructions it can follow without interrupting speech. I tried commentary.append and here’s the issues with that:

Although it handles flow without interrupting but issue is it doesn’t work reliably. Sometimes it says, sometimes it doesn’t. It’s also not a great way to send instructions as obv it’s not meant for that.

These events are not reliable enough to be used in production where deterministic flow is needed which tbh is almost all use-cases. I was wondering if you guys are aware about it and maybe already working on improving.

Please add a way to reliably instruct the model without interrupting the overall flow. This alone would make it incredibly useful in production.

We’re seeing intermittent missed delegation with gpt-live-1 → gpt-6-luna, using the official Python SDK 3.20.0, Responses delegation, low reasoning, automatic tool selection, and PCMU 8 kHz audio. In a standalone synthetic test, the caller requests a transfer and then says “Cancel that transfer.” Our application pauses the pending action on new caller captions; a backend tool must explicitly cancel or resume it. The voice model sometimes acknowledges cancellation—or correctly describes the action as paused—but produces no new delegation, backend response, or resolution tool call, and no API error. We followed the documented frontend/backend prompt split and concrete delegation conditions. Across 16 comparisons, we tested structured versus concise factual thinking.append updates and shortened frontend capability descriptions from 370 to 256 words while preserving full backend procedures. Immediate cancellation passed only 1/8 cases; cancellation after a settled pause acknowledgment passed 2/2. All sessions closed normally, and no Twilio actions occurred. We have exact prompts, recordings, event timelines, and paired successful/failing session IDs. We recognize that context acknowledgments don’t guarantee the model used the update. What documented pattern should reliably delegate corrections to an already-pending action, and would these traces help investigate?

Nice write-up. Two thoughts, one on the design and one on the numbers.

Design: I wouldn’t let cancelling an already-pending action depend on the voice model choosing to delegate. You already pause the action on new caller captions, so the backend could own the decision: a small classifier (or your backend model) reads the caption that triggered the pause and decides cancel, resume or unclear. The voice model then only has to say the result, which fits Jiayu’s advice above about keeping tool decisions in the backend. Add a watchdog too: if a paused action has no resolution within N seconds, keep it paused and have the agent ask “just to confirm, cancel the transfer?”. Then “acknowledged but nothing happened” can’t sit there silently.

Numbers: 1/8 against 2/2 fits the immediate-vs-settled difference, but with n=2 the settled arm can’t be told apart from noise. I’d run 30+ per arm with the same scripted audio before choosing a prompt shape, and keep that set as a regression suite for each model update. OpenAI’s voice agents guide lists corrections while backend work runs as something to measure on its own, so this case deserves a fixed test.

Adam