GPT-Live-Transcribe and GPT-Transcribe: Two New Transcription Models in the API

We’re introducing two new transcription models in the API:

• GPT-Live-Transcribe: built for low-latency live transcription.
• GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads.

Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for short phrases, numbers, specialized terminology, and speech with loud background noise.

Builders can improve live audio transcription by providing:

• Free-form context about the recording
• Keywords for names and domain-specific terms
• Expected input languages
• Earlier transcribed turns as context

On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it.

Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1.

Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

Been a long time arriving!

So many meeting note taker apps that actually work are on the horizon now.

I have so many GB to re-transcribe and re-embed now. :sweat_smile: