# Non-deterministic embedding results using text-embedding-ada-002

**URL:** <https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733>\
**Category:** API\
**Created:** [February 25, 2023, 8:08pm UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733 "2023-02-25T20:08:38Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![thiboeri](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/thiboeri/32/34696_2.png) [@thiboeri](https://community.openai.com/u/thiboeri)\
**Post date:** [February 25, 2023, 8:08pm UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/1 "2023-02-25T20:08:38Z")

</div>

Hi all, I am getting different embeddings for the same texts.

Using the following script:

```auto
import time

import numpy as np
import openai

openai.api_key = ...

model = 'text-embedding-ada-002'

def test():
    def get_openai_embeddings(texts, model):
        result = openai.Embedding.create(
            model=model,
            input=texts,
        )
        return result

    texts = [
        "The Lake Street Transfer station was a rapid transit station on the Chicago \"L\" that linked its Lake Street Elevated with the Logan Square branch of its Metropolitan West Side Elevated Railroad from 1913 to 1951.",
        "The Lake Street and Metropolitan were both constructed in the 1890s by different companies. The two companies owning the lines, along with two others, unified their operations in the early 1910s; as part of the merger, the Lake Street's owner had to close its nearby station on Wood Street and build a new one to form a transfer with the Metropolitan.",
        "This transfer station had a double-decked construction (depicted), with the Metropolitan's infrastructure crossing over the Lake Street. This arrangement continued until the Dearborn Street subway opened on February 25, 1951, replacing the Logan Square branch in the area and leading to the station's closure. The site would eventually serve as the junction of the modern Pink Line to the Green Line.",
    ]

    for text in texts:
        print(text)
        a = get_openai_embeddings(texts=text, model=model)
        b = get_openai_embeddings(texts=text, model=model)
        a_e = np.array(a['data'][0]['embedding'])
        b_e = np.array(b['data'][0]['embedding'])
        print('Rounded vectors to 5 decimals equal', a_e.round(5) == b_e.round(5))
        print('Max elementwise ratio', (a_e / b_e).max())
        print('Min elementwise ratio', (a_e / b_e).min())
        print('normalized norm', ((a_e - b_e)**2).sum()**0.5 / (a_e**2).sum()**0.5)
        print()
        time.sleep(1)

if __name__ == ' __main__':
    test()

```

If I run it a few times I get:

```auto
The Lake Street Transfer station was a rapid transit station on the Chicago "L" that linked its Lake Street Elevated with the Logan Square branch of its Metropolitan West Side Elevated Railroad from 1913 to 1951.
Rounded vectors to 5 decimals equal [False False True ... False False False]
Max elementwise ratio 2.9404772399566323
Min elementwise ratio -7.386667323598624
normalized norm 0.0020847767758131125

The Lake Street and Metropolitan were both constructed in the 1890s by different companies. The two companies owning the lines, along with two others, unified their operations in the early 1910s; as part of the merger, the Lake Street's owner had to close its nearby station on Wood Street and build a new one to form a transfer with the Metropolitan.
Rounded vectors to 5 decimals equal [True True True ... True True True]
Max elementwise ratio 1.0
Min elementwise ratio 1.0
normalized norm 0.0

This transfer station had a double-decked construction (depicted), with the Metropolitan's infrastructure crossing over the Lake Street. This arrangement continued until the Dearborn Street subway opened on February 25, 1951, replacing the Logan Square branch in the area and leading to the station's closure. The site would eventually serve as the junction of the modern Pink Line to the Green Line.
Rounded vectors to 5 decimals equal [False False False ... False False False]
Max elementwise ratio 5.401572035423635
Min elementwise ratio -13.42049661181725
normalized norm 0.004099269126226605

```

It seems the vectors returned can sometimes have very different results! I understand that a certain amount of stochasticity is possible, but it seems that some elements are very different.

---

<div class="post-metadata">

**Author:** ![ruby\_coder](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ruby_coder/32/19875_2.png) [@ruby\_coder](https://community.openai.com/u/ruby_coder)\
**Post date:** [February 26, 2023, 1:52am UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/2 "2023-02-26T01:52:02Z")

</div>

> [@thiboeri](#):
>
> ` "The Lake Street Transfer station was a rapid transit station on the Chicago \"L\" that linked its Lake Street Elevated with the Logan Square branch of its Metropolitan West Side Elevated Railroad from 1913 to 1951.",`

Hi @thiboeri

I ran your text (the first example) 10 times and got the same embedded vector 10 times, using a Ruby API wrapper (not Python).

HTH

🙂

---

<div class="post-metadata">

**Author:** ![anon10827405](https://avatars.discourse-cdn.com/v4/letter/a/ed8c4c/32.png) [@anon10827405](https://community.openai.com/u/anon10827405)\
**Post date:** [February 26, 2023, 2:12am UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/3 "2023-02-26T02:12:56Z")

</div>

Yes. The OpenAI python library will return many more decimal places; essentially noise. I believe it’s something to do with the conversion to base64? Can’t remember exactly.

Regardless, the similarity scores will be the same.

---

<div class="post-metadata">

**Author:** ![danielbichuetti](https://avatars.discourse-cdn.com/v4/letter/d/fbc32d/32.png) [@danielbichuetti](https://community.openai.com/u/danielbichuetti)\
**Post date:** [March 6, 2023, 8:33pm UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/4 "2023-03-06T20:33:00Z")

</div>

We just faced the same issues for the first time here when using the _openai-python_ package.

We did some tests and around 11% of them were considerably different, even being near in the vector space.

**UPDATE:** For anyone facing this issue, the embeddings’ endpoint is deterministic. The reason to this difference is caused by the OpenAI Python package, as it uses base64 as the default encoding format, while others don’t.

---

<div class="post-metadata">

**Author:** ![curt.kennedy](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/curt.kennedy/32/709249_2.png) [@curt.kennedy](https://community.openai.com/u/curt.kennedy)\
**Post date:** [March 6, 2023, 8:44pm UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/5 "2023-03-06T20:44:02Z")

</div>

See this post for more discussion on embedding decimal places [Discrepancy in embeddings precision - #8 by curt.kennedy](https://community.openai.com/t/discrepancy-in-embeddings-precision/78666/8)

---

<div class="post-metadata">

**Author:** ![tb.songipark](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/tb.songipark/32/39092_2.png) [@tb.songipark](https://community.openai.com/u/tb.songipark)\
**Post date:** [March 15, 2023, 8:05am UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/6 "2023-03-15T08:05:25Z")

</div>

hi, faced same issue with python here. 🥲 In my dataset, Im getting only 2 decimal precision.

did you fix the prob?  
I’ve read all the instructions and guide but still don’t know what to do.

Is it the only way not to use python? 🥲

---

<div class="post-metadata">

**Author:** ![rav](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/rav/32/75732_2.png) [@rav](https://community.openai.com/u/rav)\
**Post date:** [April 18, 2023, 6:37am UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/7 "2023-04-18T06:37:40Z")

</div>

I think it’s worth to separate the python precision issue discussed in [Discrepancy in embeddings precision - #7 by RonaldGRuckus](https://community.openai.com/t/discrepancy-in-embeddings-precision/78666/7) from the issue of the OpenAI API returning slightly different embeddings for the same exact input. Instead of using python, let’s use `curl` directly:

```sh
curl https://api.openai.com/v1/embeddings \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{
    "input": "<TEXT>",
    "model": "text-embedding-ada-002"
  }' | jq '.data[0].embedding[0]'

```

Make sure to set your `OPENAI_API_KEY`, the `jq` command will return the 1st number from the embeddings, if you run this a couple of times, you will see that the numbers can be slight different, in my case for example: -0.026714837, -0.026664866 (difference of about 5e-05).

---

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [December 24, 2023, 1:36pm UTC](https://community.openai.com/t/non-deterministic-embedding-results-using-text-embedding-ada-002/74733/8 "2023-12-24T13:36:19Z")

</div>


