I frequently use Voice while walking to have substantive conversations about software architecture, agentic engineering, product design, and future directions. My priority in these conversations is reasoning quality, not instantaneous response.
The current low-latency Voice experience is excellent at conversational responsiveness, but there is a substantial intelligence/reasoning gap compared with the stronger text models ( eg. 5.6 Sol High ). That makes it unsuitable for some of the conversations for which Voice would otherwise be most valuable to me.
I would willingly accept several seconds—or considerably more—of latency between turns in exchange for having my speech transcribed, sent to a high-capability reasoning model, and the resulting response spoken back to me.
In other words, please don’t assume that Voice = real-time latency is paramount. There is another important use case:
speech in → powerful reasoning model → wait → speech out.
Ideally Pro users could choose something analogous to:
- Live — optimize for conversational latency.
- Thoughtful Voice — optimize for reasoning quality, accepting latency.
Ideally the user could choose the model + intensity that they are chatting with, but I will take what I can get. I don’t need interruption at arbitrary points while the model is speaking, sub-second response times, or continuous full-duplex audio nearly as much as I need the model on the other end of the conversation to be capable of genuinely contributing to a difficult intellectual discussion.
As an aside, I find the newer voice mode output speech rate frustratingly slower than the previous generation’s. It feels to me around ~6% slower, but I understand that is a subjective perception.
Thank you for your time!