# Streaming from Text-to-Speech api

**URL:** <https://community.openai.com/t/streaming-from-text-to-speech-api/493784>\
**Category:** API\
**Tags:** api, python, tts\
**Created:** [November 10, 2023, 9:04pm UTC](https://community.openai.com/t/streaming-from-text-to-speech-api/493784 "2023-11-10T21:04:37Z")\
**Posts on this page:** 1\
**Showing post:** 25

<div class="post-metadata">

**Author:** ![nimobeeren](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/nimobeeren/32/315017_2.png) [@nimobeeren](https://community.openai.com/u/nimobeeren)\
**Post date:** [February 22, 2024, 9:57pm UTC](https://community.openai.com/t/streaming-from-text-to-speech-api/493784/25 "2024-02-22T21:57:36Z")

</div>

I was able to get some preliminary results with streaming + playing back audio in real-time using [pyaudio](https://people.csail.mit.edu/hubert/pyaudio/):

```python
import os
import requests
from time import time
import pyaudio

url = "https://api.openai.com/v1/audio/speech"
headers = {
    "Authorization": f'Bearer {os.getenv("OPENAI_API_KEY")}',
}

data = {
    "model": "tts-1",
    "input": "This is a test",
    "voice": "shimmer",
    "response_format": "wav",
}

start_time = time()
response = requests.post(url, headers=headers, json=data, stream=True)
if response.status_code == 200:
    print(f"Time to first byte: {int((time() - start_time) * 1000)} ms")
    p = pyaudio.PyAudio()
    stream = p.open(format=8, channels=1, rate=24000, output=True)
    for chunk in response.iter_content(chunk_size=1024):
        stream.write(chunk)
    print(f"Time to complete: {int((time() - start_time) * 1000)} ms")

```

**HEADPHONE WARNING:** this can cause very harsh noise, especially on longer inputs. Keep volume low.

The best part is that this drastically reduces latency to about 200-500 ms (time to first byte). I found latency was lowest with `"response_format": "wav"`, though the trade-off is larger file size. But on any decent connection, the bottleneck will still be generation speed, not network.

I got the values for `format`, `channels` and `rate` by writing the stream to a `.wav` file and analyzing it as per [pyaudio docs](https://people.csail.mit.edu/hubert/pyaudio/docs/):

```python
import os
import requests
import io
import wave
import pyaudio

# url = ...
# headers = ...
# data = ...

response = requests.post(url, headers=headers, json=data, stream=True)
if response.status_code == 200:
    buffer = io.BytesIO()
    for chunk in response.iter_content(chunk_size=1024):
        buffer.write(chunk)

with open("speech.wav", "wb") as f:
    f.write(buffer.getvalue())

with wave.open('speech.wav', 'rb') as wf:
    p = pyaudio.PyAudio()

    print('format', p.get_format_from_width(wf.getsampwidth()))
    print('channels', wf.getnchannels())
    print('rate', wf.getframerate())

```

But my guess is that this is wrong, and is the cause of the intermittent noise.

---

_[View the full topic](https://community.openai.com/t/streaming-from-text-to-speech-api/493784)._
