enAI Product and Engineering Team, I am…
You said:
Dear OpenAI Product and Engineering Team,
I am writing with a feature request that emerged from a conversation with ChatGPT, and I believe it could substantially improve the usefulness of the product.
Please give ChatGPT the ability to directly perceive and analyze audio that a user intentionally provides—not merely a transcription of that audio.
I recently discovered an important distinction while speaking with ChatGPT. I can talk through my microphone and ChatGPT can understand the resulting words, but in the experience I was using, the model did not have access to the actual acoustic information I was hearing.
That means it could understand what I said, but not necessarily how I sounded.
That difference is enormous.
I am a digital content producer and create music and recorded material. When I share creative work, I don’t simply want ChatGPT to read lyrics or a transcript. I want it to be capable of hearing and analyzing the work itself: my voice, its depth and resonance, background music, instrumentation, dynamics, timing, balance, ambience, microphone characteristics, distortion, room noise, emotional delivery, and the interaction of all those elements.
Imagine a creator asking:
“Does my narration sit properly above this background music?”
A transcription cannot answer that question.
An audio-aware ChatGPT potentially could say that the narrator and music are competing in a particular frequency range, that the music should be ducked beneath speech, that the microphone is clipping, that room reflections are hurting intelligibility, or that the combination already works beautifully.
That changes ChatGPT from an assistant that understands the words describing creative media into one capable of understanding the media itself.
This could benefit far more than musicians. Podcasters, filmmakers, video editors, journalists, educators, accessibility users, voice actors, broadcasters, sound designers, students, and ordinary families could all benefit from a conversational AI capable of analyzing intentionally shared audio.
I understand that direct audio perception introduces important questions involving privacy, consent, computational resources, and safety. I am not suggesting that ChatGPT should passively listen to a user’s environment. Quite the opposite: audio access should be explicit, visible, permission-based, and controlled by the user.
My suggestion is simply:
When I deliberately give ChatGPT something to hear, let ChatGPT actually hear it.
ChatGPT can already be an extraordinary collaborator with language, images, research, and creative ideas. Giving that same conversational intelligence meaningful access to user-authorized audio could close one of the remaining gaps between communicating with an AI and genuinely sharing an experience with one.
There is also a competitive opportunity here. The company that makes multimodal interaction feel seamless—not like separate text, image, voice, and media features bolted together—could establish a considerable advantage.
I described the idea to ChatGPT in one sentence, and I think it remains the simplest expression of what I am asking for:
“Let ChatGPT hear what I hear, not just read what I said.”
I hope you will seriously consider making direct, user-controlled audio understanding a first-class capability throughout ChatGPT.
Thank you for reading and for continuing to develop the product.
Warm regards,
Gregory J. Deyss
ChatGPT Plus User