# SORA: wait but does it come with sound?!

**URL:** <https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185>\
**Category:** Community\
**Tags:** sora\
**Created:** [February 17, 2024, 4:40am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185 "2024-02-17T04:40:33Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bestbubbldev](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/bestbubbldev/32/13232_2.png) [@Bestbubbldev](https://community.openai.com/u/Bestbubbldev)\
**Post date:** [February 17, 2024, 4:40am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/1 "2024-02-17T04:40:33Z")

</div>

The primary concern is how to incorporate sound into the videos we create. I’m not referring to a soundtrack or background music, but rather scenarios like a dialogue scene or a scene where a character is singing.

Manually syncing voice and audio with lip movements would be a huge headache. Are there any existing applications that can synchronize audio and video?

How do you plan to tackle this problem?

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [February 17, 2024, 5:12am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/2 "2024-02-17T05:12:00Z")

</div>

The way this kind of algorithm works is the AI uses both prompt and understanding of natural motions of labeled and learned objects to continue to step through likely progress of a video scene, powered by machine learning.

It appears superb, but is not unique:

[![](https://us1.discourse-cdn.com/openai1/original/4X/1/3/a/13ab77f559c15464863d4bb94879fc5a95be505a.jpeg "Lumiere") ](https://www.youtube.com/watch?v=wxLr02Dz2Sc)

Temporal specificity doesn’t appear to be a target.

If you have the AI elucidate on a singer in a jazz club, you’ll have to write your own song to come up with whatever is output.

Don’t expect this to be an AI newscaster…

Don’t expect to see this in the wild any time soon.

---

<div class="post-metadata">

**Author:** ![anon22939549](https://avatars.discourse-cdn.com/v4/letter/a/f14d63/32.png) [@anon22939549](https://community.openai.com/u/anon22939549)\
**Post date:** [February 17, 2024, 6:07am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/3 "2024-02-17T06:07:18Z")

</div>

> [@Bestbubbldev](#):
>
> The primary concern is how to incorporate sound into the videos we create.

 ![1000034423](https://us1.discourse-cdn.com/openai1/original/4X/3/3/b/33b51ba28afa4ec1d8f54969301560956a365693.webp)

---

<div class="post-metadata">

**Author:** ![anon22939549](https://avatars.discourse-cdn.com/v4/letter/a/f14d63/32.png) [@anon22939549](https://community.openai.com/u/anon22939549)\
**Post date:** [February 17, 2024, 6:27am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/5 "2024-02-17T06:27:19Z")

</div>

> [@\_j](#):
>
> If you have the AI elucidate on a singer in a jazz club, you’ll have to write your own song to come up with whatever is output.
> 
> Don’t expect this to be an AI newscaster…

_Maybe…_ AI music has come a long way in the last two years. OpenAI smartly (at least not publicly) hasn’t delved into music generation. But, I’m sure you could get at least a rendition of a song to lay over a video. It would be absolutely trash right now—humans are incredibly sensitive to really minor audio/video sync issues^{[https://www.itu.int/dms\_pubrec/itu-r/rec/bt/R-REC-BT.1359-0-199802-S!!PDF-E.pdf](https://www.itu.int/dms_pubrec/itu-r/rec/bt/R-REC-BT.1359-0-199802-S!!PDF-E.pdf)].

I absolutely forsee models coming in the future that target video-to-audio and audio-to-video. That just a natural progression and we already have huge training sets available.

Then I imagine it’s only a very short amount of time before sometime inevitably pits them against each other with the goal of convergence.

> [@\_j](#):
>
> Don’t expect to see this in the wild any time soon.

I think you’re right, but I think we need to qualify what _soon_ means in this context.

I think we’ll\[1\] be able to generate text-to-AV that’s _mostly_ "good enough\* within 5-years but definitely within 10.

But, I also think that, if you don’t need a monolithic model to do it in one go, we can do these things today—with significant post-processing and editing.

It’s an exciting and scary time…

* * *

1. Meaning humanity.

---

<div class="post-metadata">

**Author:** ![VeitB](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/veitb/32/712987_2.png) [@VeitB](https://community.openai.com/u/VeitB)\
**Post date:** [February 17, 2024, 8:53am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/6 "2024-02-17T08:53:36Z")

</div>

> [@anon22939549](#):
>
> OpenAI smartly (at least not publicly) hasn’t delved into music generation.

Take a look. It’s ‘old’ but it exists.

[https://openai.com/research/jukebox](https://openai.com/research/jukebox)

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [February 17, 2024, 9:01am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/7 "2024-02-17T09:01:54Z")

</div>

> [@anon22939549](#):
>
> ![33b51ba28afa4ec1d8f54969301560956a365693_2_690x394](https://us1.discourse-cdn.com/openai1/original/4X/4/a/e/4ae1e7bc9747485631d3665d75bd0137aff59b81.webp)

![ezgif-6-76af43d06b](https://us1.discourse-cdn.com/openai1/original/4X/f/6/b/f6bdba1467cf25c05018fe8970f27f72794b4919.webp)

---

<div class="post-metadata">

**Author:** ![h.alesso](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/h.alesso/32/546696_2.png) [@h.alesso](https://community.openai.com/u/h.alesso)\
**Post date:** [February 17, 2024, 10:11pm UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/8 "2024-02-17T22:11:53Z")

</div>

I think [D-ID.com](http://D-ID.com) does a pretty good job of audio sync with avatars. True the avatars are usually just small motion heads, but the lips sync is good. This same tech could probably be incorporated in Sora. Don’t you think?

---

<div class="post-metadata">

**Author:** ![curt.kennedy](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/curt.kennedy/32/709249_2.png) [@curt.kennedy](https://community.openai.com/u/curt.kennedy)\
**Post date:** [February 17, 2024, 10:22pm UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/9 "2024-02-17T22:22:17Z")

</div>

Yeah exactly … the OpenAI Jukebox.

Wondering if they could generate another model, that after the video is rendered, you feed the video and a prompt describing the desired soundtrack to the video, and this other model would generate the sound and sync it to the video, creating another video with this prompted AI soundtrack in it.

---

<div class="post-metadata">

**Author:** ![Agha.khan](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/agha.khan/32/309097_2.png) [@Agha.khan](https://community.openai.com/u/Agha.khan)\
**Post date:** [February 18, 2024, 5:34am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/11 "2024-02-18T05:34:16Z")

</div>

I think sora is still in its infancy and think of the possibilities in coming days. I think it will be a complete package with lip sync and dialogue prompting in a single prompt.

---

<div class="post-metadata">

**Author:** ![Agha.khan](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/agha.khan/32/309097_2.png) [@Agha.khan](https://community.openai.com/u/Agha.khan)\
**Post date:** [February 18, 2024, 5:37am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/12 "2024-02-18T05:37:53Z")

</div>

Thanks for the share, I don’t know how I missed it 🙂

---

<div class="post-metadata">

**Author:** ![curt.kennedy](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/curt.kennedy/32/709249_2.png) [@curt.kennedy](https://community.openai.com/u/curt.kennedy)\
**Post date:** [February 18, 2024, 5:54am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/14 "2024-02-18T05:54:29Z")

</div>

Yes, all early days … but think of these scenarios as “prompts”:

1. Input a video, _without_ sound. Get back the same video _with_ sound.

- Option 1 - Infer the soundtrack from the video, no text prompt
- Option 2 - Infer the soundtrack from the joint prompt/video information

1. Input a sound, no video. Get back a video that corresponds to the sound.

- Option 1 - Infer the video from the sound, no text prompt
- Option 2 - Infer the video from the joint prompt/sound

But think of the different permutations. You could input sounds, get back videos. Input videos, get back sounds. You could use the joint information to sync video/sound together. And control both with a text prompt as well.

Lot’s of permutations.

Right now, Sora is input text, get a video without sound. Adding sound and syncing it back to the video is a next logical step.

Having this a separate rendering steps would give creators more control. For example, you like the video, but need to iterate on the sound a bunch to dial it in. Or you love the sound, but need to iterate on the video side.

If you are limited to doing both in one pass, you risk _losing both_ since you are changing two variables at the same time.

---

<div class="post-metadata">

**Author:** ![videoaierc](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/videoaierc/32/310265_2.png) [@videoaierc](https://community.openai.com/u/videoaierc)\
**Post date:** [February 18, 2024, 8:37am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/15 "2024-02-18T08:37:52Z")

</div>

This would be incredible, plus it will change the world we know. We’ll see what happens from now on. I’m waiting to do something amazing with Sora.

---

<div class="post-metadata">

**Author:** ![dignity\_for\_all](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/dignity_for_all/32/300058_2.png) [@dignity\_for\_all](https://community.openai.com/u/dignity_for_all)\
**Post date:** [February 18, 2024, 9:42am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/16 "2024-02-18T09:42:24Z")

</div>

I understand the excitement, and it seems like a fantastic technology, but I think it’s best not to consider it as something urgent, and I don’t know when it will be generally available.  
Currently, access is exclusive to Red Teamers, and there is no waiting list.  
We’ll have to be patient for a while.

* * *

It may remind you of a leader from a European country who introduced one popular policy after another. Despite being ridiculed as a butterfly or a soap bubble, they only ended up disappointing the people.

---

<div class="post-metadata">

**Author:** ![tradingfhk](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/tradingfhk/32/239245_2.png) [@tradingfhk](https://community.openai.com/u/tradingfhk)\
**Post date:** [February 18, 2024, 9:00pm UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/17 "2024-02-18T21:00:12Z")

</div>

I started seeing fake SORA scams on the internet.

---

<div class="post-metadata">

**Author:** ![N2U](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/n2u/32/54438_2.png) [@N2U](https://community.openai.com/u/N2U)\
**Post date:** [February 18, 2024, 9:04pm UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/18 "2024-02-18T21:04:39Z")

</div>

Yeah there will probably be more of those 😅

If you see any scams out there, you can send a link to me or one of the other moderators, we do report these on a regular basis. ❤

---

<div class="post-metadata">

**Author:** ![JRMazarri](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jrmazarri/32/148564_2.png) [@JRMazarri](https://community.openai.com/u/JRMazarri)\
**Post date:** [February 19, 2024, 12:35am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/19 "2024-02-19T00:35:07Z")

</div>

Everyone share your thoughts on Sora! And how it can help you!?  
I can’t wait to use it to generate marketing videos for my clients, automating the whole process for video content creation for business brands! I’m sure a lot of people already have so many ideas on how it can help them! If you are open to sharing, I would love to build and teach you how to use this ai automation in a way that helps you, imagine running it while you sleep… You can find me on all socials with this same name.

I am such an AI ethuist, I started building ai automations with open ai over 2+ years ago. And I tell you the different applications and use case scenarios are so beneficial for businesses of all kinds even small businesses! A lot of businesses today are so outdated, and not automating their whole process… which is very unfortunate… Utilizing AI to automate your business, can cut so much cost saving your business $1000’s! You can then turn around and use that money to grow your company or run paid ads! Business Owners do not understand that yet… One day they will. And I am excited to be a part of it!

HAPPY AI! everyone ❤

---

<div class="post-metadata">

**Author:** ![JRMazarri](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jrmazarri/32/148564_2.png) [@JRMazarri](https://community.openai.com/u/JRMazarri)\
**Post date:** [February 19, 2024, 1:05am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/20 "2024-02-19T01:05:01Z")

</div>

Wow I am just so amazed with Open ai, on Sora video generator! I can’t wait to make marketing content and movies! Who with me?!

---

<div class="post-metadata">

**Author:** ![timeislight](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/timeislight/32/310390_2.png) [@timeislight](https://community.openai.com/u/timeislight)\
**Post date:** [February 19, 2024, 1:22am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/21 "2024-02-19T01:22:25Z")

</div>

Of course @JRMazarri I am with you! ))

---

<div class="post-metadata">

**Author:** ![timeislight](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/timeislight/32/310390_2.png) [@timeislight](https://community.openai.com/u/timeislight)\
**Post date:** [February 19, 2024, 1:37am UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/22 "2024-02-19T01:37:03Z")

</div>

## Hello, do you know this other incredible news? Yet another AI revolution! :

> <https://x.com/LinusEkenstam/status/1759298104593952867?s=20>

---

<div class="post-metadata">

**Author:** ![minjaben](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/minjaben/32/314666_2.png) [@minjaben](https://community.openai.com/u/minjaben)\
**Post date:** [February 20, 2024, 4:09pm UTC](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185/23 "2024-02-20T16:09:37Z")

</div>

Great, no need for a sound designer! Guess I’ll just go look for another career that isn’t being encroached by AI progress?  
Or find a way to use it? I don’t even know anymore.

[Next page](https://community.openai.com/t/sora-wait-but-does-it-come-with-sound/632185.md?page=2)
