How can I switch from text generation to audio generation?

I had the same problem, and the only mitigation I found is to give all the history of the conversation as if it came with the user but with something like “[Assistant]” and “[User]” in front of the messages. I think combining that with the session.update tricks you mention made it more reliable as well, but that might just be placebo.

As a side note, the other scenario where I bumped into this is trying to reduce API cost by deleting previous audio conversation items and replacing them by the text transcript (since the cost per input text token is so much lower than the cost per audio token). That results in the same problem where the model, but if I keep the last 2 or ideally 3 assistant messages as audio, it seems to nearly always keep responding in audio, even though the older history is text.

Emphasis on nearly always, none of this seems 100% reliable, which would be a big problem in production (well, if the costs wouldn’t immediately bankrupt anyone using this in production anyway) - we really need a way to force the model to reply with audio…