# Non-deterministic embedding models?

**URL:** <https://community.openai.com/t/non-deterministic-embedding-models/634880>\
**Category:** API\
**Created:** [February 18, 2024, 12:03am UTC](https://community.openai.com/t/non-deterministic-embedding-models/634880 "2024-02-18T00:03:31Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![leaf](https://avatars.discourse-cdn.com/v4/letter/l/ec9cab/32.png) [@leaf](https://community.openai.com/u/leaf)\
**Post date:** [February 18, 2024, 12:03am UTC](https://community.openai.com/t/non-deterministic-embedding-models/634880/1 "2024-02-18T00:03:31Z")

</div>

My understanding of embedding models is that they are a deterministic thing, mapping text to a numerical vector.

I repeatedly regenerated an embedding for two words about 10-15 times. They were the same most of the time, but, in two cases we got a different embedding.

Am I misunderstanding how embedding models work, or, is there something going on under the hood of the API - bug or otherwise?

The model was `text-embedding-3-small`

---

<div class="post-metadata">

**Author:** ![Diet](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/diet/32/562749_2.png) [@Diet](https://community.openai.com/u/Diet)\
**Post date:** [February 18, 2024, 12:10am UTC](https://community.openai.com/t/non-deterministic-embedding-models/634880/2 "2024-02-18T00:10:49Z")

</div>

Embedding models _are_ large language models, and determinism isn’t something that OpenAI really guarantees at the moment.

Were they significantly different? Sometimes you get rounding errors if you don’t specify the encoding format [https://platform.openai.com/docs/api-reference/embeddings/create#embeddings-create-encoding\_format](https://platform.openai.com/docs/api-reference/embeddings/create#embeddings-create-encoding_format)

How different were the embeddings? If you get a cosine similarity exceeding 0.99999, they might as well be the same.
