# Are vectors generated by text-embedding-3-small always the same for the same text input?

**URL:** <https://community.openai.com/t/are-vectors-generated-by-text-embedding-3-small-always-the-same-for-the-same-text-input/739757>\
**Category:** API\
**Tags:** embeddings\
**Created:** [May 8, 2024, 12:19pm UTC](https://community.openai.com/t/are-vectors-generated-by-text-embedding-3-small-always-the-same-for-the-same-text-input/739757 "2024-05-08T12:19:44Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![suhas.chatekar](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/suhas.chatekar/32/335613_2.png) [@suhas.chatekar](https://community.openai.com/u/suhas.chatekar)\
**Post date:** [May 8, 2024, 12:19pm UTC](https://community.openai.com/t/are-vectors-generated-by-text-embedding-3-small-always-the-same-for-the-same-text-input/739757/1 "2024-05-08T12:19:44Z")

</div>

I am in the process of building a prototype for continuous ingestion of content into a vector database. I am using text-embedding-3-small model to generate vectors before storing a Azure Search index.

The internal API I am using to get the newly created content is not perfect which means I might be fetching content that is already stored in the index. I am thinking this should not be a problem if the model produces the same vector representation then when I send that vector to Azure Search index, it would simply be replaced in place of the current vector in the index. So my question is - will the model generate same vector for the same input text on second and subsequent calls to the model API?

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [May 8, 2024, 12:53pm UTC](https://community.openai.com/t/are-vectors-generated-by-text-embedding-3-small-always-the-same-for-the-same-text-input/739757/2 "2024-05-08T12:53:46Z")

</div>

The embeddings models should not be used to verify identical contents or avoid duplication - you can use a hash algorithm for that for free.

> [@Splitting text into chunks versus reducing the text](https://community.openai.com/t/splitting-text-into-chunks-versus-reducing-the-text/696028/3):
>
> Determinism means you always get the same output from the same input. The way you expect computer code to work, basically. The current embeddings models are not deterministic. They do not produce the same output for the same input. I performed 10 embeddings runs, on 3-small model, of the same 600 tokens of text, and got three unique results with tensor comparisons: np.unique(em\_ndarray, axis=0) array([[ 0.01526077, 0.01427238, 0.06942467, ..., 0.00670789, -0.01064828, -0.0110238…

Results are close enough between successive runs that it would be effective to almost always return the same top results, and embeddings quality is kind of subjective anyway..

---

<div class="post-metadata">

**Author:** ![suhas.chatekar](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/suhas.chatekar/32/335613_2.png) [@suhas.chatekar](https://community.openai.com/u/suhas.chatekar)\
**Post date:** [May 8, 2024, 1:52pm UTC](https://community.openai.com/t/are-vectors-generated-by-text-embedding-3-small-always-the-same-for-the-same-text-input/739757/3 "2024-05-08T13:52:28Z")

</div>

Thanks for that response. Very useful.

I am aware of hashing techniques and aware that I can employ them to make sure I am not vectoring the same content multiple times. I was wondering if there is a way to avoid having to hash the content separately.

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [May 8, 2024, 3:53pm UTC](https://community.openai.com/t/are-vectors-generated-by-text-embedding-3-small-always-the-same-for-the-same-text-input/739757/4 "2024-05-08T15:53:12Z")

</div>

The only way I see would be more expensive and slower. Which is to not add what you paid embeddings for if there is an embeddings result \>.999 from an exhaustive search and the text returned from the database matches.
