# Using embeddings for semantic search on transcripts

**URL:** <https://community.openai.com/t/using-embeddings-for-semantic-search-on-transcripts/108449>\
**Category:** API\
**Created:** [March 20, 2023, 4:00pm UTC](https://community.openai.com/t/using-embeddings-for-semantic-search-on-transcripts/108449 "2023-03-20T16:00:38Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![enknamel](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/enknamel/32/37911_2.png) [@enknamel](https://community.openai.com/u/enknamel)\
**Post date:** [March 20, 2023, 4:00pm UTC](https://community.openai.com/t/using-embeddings-for-semantic-search-on-transcripts/108449/1 "2023-03-20T16:00:38Z")

</div>

Hello,

I would like to do semantic search on audio transcripts. I’ve made a few proof of concepts using embeddings and have some questions.

1. If I embed the whole transcript (which often doesn’t fit) it seems wasteful. If I create embeddings with simple heuristics like sliding windows, it’s hard to optimize for relevancy given how freeform transcripts can be. Is there a method for identifying the optimal text to create an embedding for?

2. The transcript document quality is a bit low but there is separate metadata I can add to it. Such as who is speaking, or a summary of the conversation to that point, topic labels, etc. Are there any best practices for cleaning text to create an embedding?

---

<div class="post-metadata">

**Author:** ![wfhbrian](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/wfhbrian/32/19388_2.png) [@wfhbrian](https://community.openai.com/u/wfhbrian)\
**Post date:** [March 20, 2023, 4:12pm UTC](https://community.openai.com/t/using-embeddings-for-semantic-search-on-transcripts/108449/2 "2023-03-20T16:12:13Z")

</div>

I embed based on who is talking. Each embedding also include metadata like the subject of the meeting. I also skip embedding if what the person said is less than a present number of words.

Something I’ve been thinking about is including both the last thing someone else said, and the following thing, to give more context. But I haven’t tested that yet.
