Introducing GPT-Live-1 in the API

@fred5 we’re also having a lot of trouble ending the conversation. With gpt-realtime we have a finish_session tool that the assistant calls when the conversation is done, and we then hang up the SIP call when the assistant finishes its goodbye speech.

With gpt-live, the frontend often just says “goodbye” or “if you need anything else, I’m here” and never delegates to the backend… so the hangup never happens. Same thing you’re experiencing. We’ve tried adjusting the prompt and it improved a bit, but it’s nowhere near as reliable as with gpt-realtime.

In general, after a week of testing, we feel like the frontend/backend split makes the assistant more capable, but the conversation is more unpredictable and less suitable for semi-rigid flows. The frontend sometimes doesn’t delegate when it should, or the preambles get ahead of the backend when it does. It’s like the frontend lacks context. This can be improved with prompting up to a point, but the more we tweak it to get the desired behavior, the more the voice degrades.

Would be great to hear if anyone successfully transitioned a gpt-realtime assistant that relied on a “# Conversation Flow” into gpt-live, and especially if anyone has found a clean way to end the conversation gracefully.