Tts-1-hd returns silence instead of audio

We use tts-1-hd heavily (and pay a few 1000 USD / month for it). So far it worked acceptable, sometimes it timed-out or returned an empty response (audio length == 0). We detected that and with a few retries we got the correct audio.

But a few weeks ago it started to return valid audio files containing nothing but exact silence (all wave data == 0), and even worse it now returns valid audio file containing mostly silence and then a few spoken words. For example, I asked it to speak the following German paragraph:

“Sehr, sehr wichtig ist die Qualität und die Tiefe der internationalen Delegationen, die nach Iran gereist sind, um dem verstorbenen Führer Ayatollah Khamenei ihre Ehrerbietung zu erweisen. Über hundert Länder aus dem Globalen Süden waren vertreten – praktisch niemand aus dem sogenannten NATO-Raum, also aus dem Westen, aus dem kriegstreibenden Westen. \n\nWir sehen hier die Folgen dieser Krise, ein deutliches Zeichen des Respekts der iranischen Bevölkerung, das sich im gesamten Globalen Süden widerspiegelt. Man sieht das in Afrika, in Südostasien, in Russland, in China und in Lateinamerika. \n\nUnd eine dieser Delegationen war besonders bedeutend: die pakistanische Delegation. Dabei waren der Premierminister von Pakistan, Shehbaz Sharif, und der Armeechef, Feldmarschall Asim Munir. \n\nWarum war das so bemerkenswert? Weil Pakistan Feldmarschall Asim Munir entsandt hat, um direkt sowohl mit der politischen als auch mit der militärischen Führung Irans zu sprechen.”

The spoken audio returned by tts-1-hd is this:

It contains 54 seconds of silence and then 6 seconds of the last words of the paragraph: “um direkt sowohl mit der politischen als auch mit der militärischen Führung Irans zu sprechen.”

This is very hard to detect in a retry-loop. PLEASE fix that.

Thank You!

That sounds like a very annoying issue.

I ran a few tests but could not reproduce it right away. Can you share a bit more detail about how often this happens? Maybe you can share the code you are using to call the API?

We let it speak chunks of ~ 1 Minute, sometimes shorter if the speaker says something like “Yes” or similar. Currently we only detect complete silence and retry. The job I referred to had 570 audio snippets (in 7 languages) and 5 times tts returned complete silence, and after a retry the audio was correct

The problem I described here is that tts also replies with mixed audio and silence. This is not yet detected by our system so I don’t know the frequency of that. We were pointed to this particular case by human listeners.

We use ‘com.openai:openai-java:4.38.0’ to call the API.

Maybe I should add that we send requests in parallel up to 5,000 tts requests / minute (tier 5).