# Embedding very sensitive to punctuation

**URL:** <https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205>\
**Category:** API\
**Tags:** embeddings, ada, ada002\
**Created:** [December 6, 2023, 1:13pm UTC](https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205 "2023-12-06T13:13:25Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![leo.bachelier.ext](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/leo.bachelier.ext/32/147246_2.png) [@leo.bachelier.ext](https://community.openai.com/u/leo.bachelier.ext)\
**Post date:** [December 6, 2023, 1:13pm UTC](https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205/1 "2023-12-06T13:13:25Z")

</div>

I have been using OpenAI Embeddings specifically text-embedding-ada-002 and noticed it was very sensitive to punctuation even. I have around 1000 chunks and need to extract each time the 15 most similar chunks to my query. I have been testing my query without punctuation and when I add a dot ‘.’ at the end of my query it changes the initial set I got from the retriever with the query without punctuation (some chunks are the same but new ones may appear or the initial order is different).

- Have you noticed anything similar ?
- Is it the basic behaviour of this embedding to be that sensitive to punctuation ?
- Is there a way to make it more robust to minor changes in the query ?

FYI: I am using PGvector to store my chunks vectors

---

<div class="post-metadata">

**Author:** ![lmccallum](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/lmccallum/32/363_2.png) [@lmccallum](https://community.openai.com/u/lmccallum)\
**Post date:** [December 6, 2023, 8:29pm UTC](https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205/2 "2023-12-06T20:29:32Z")

</div>

I haven’t noticed punctuation, but I have noticed a significant downgrade in performance using ada-002 compared to the davinci-001 embeddings model. I am really frustrated because I re-embedded all my texts, and now the results aren’t fit for my use case. Before, the most relevant text always showed up in the first 1-5 search results, now it’s the 50th search result!

---

<div class="post-metadata">

**Author:** ![anon10827405](https://avatars.discourse-cdn.com/v4/letter/a/ed8c4c/32.png) [@anon10827405](https://community.openai.com/u/anon10827405)\
**Post date:** [December 6, 2023, 8:32pm UTC](https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205/3 "2023-12-06T20:32:14Z")

</div>

> [@leo.bachelier.ext](#):
>
> Have you noticed anything similar ?  
> Is it the basic behaviour of this embedding to be that sensitive to punctuation ?

Yes. Grammar do be like that.

> [@leo.bachelier.ext](#):
>
> Is there a way to make it more robust to minor changes in the query ?

Have you tried normalizing/correcting the text using GPT and then embedding it?

---

<div class="post-metadata">

**Author:** ![leo.bachelier.ext](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/leo.bachelier.ext/32/147246_2.png) [@leo.bachelier.ext](https://community.openai.com/u/leo.bachelier.ext)\
**Post date:** [December 8, 2023, 10:12am UTC](https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205/4 "2023-12-08T10:12:58Z")

</div>

Actually, it’s more in the query I am sending to the retriever, if I add a dot at the end it changes the set returned. Depending on the query sometimes the version with the dot returns what I am looking for and sometimes it’s the version without the dot.  
So not sure how to normalize it correclty.

---

<div class="post-metadata">

**Author:** ![anon10827405](https://avatars.discourse-cdn.com/v4/letter/a/ed8c4c/32.png) [@anon10827405](https://community.openai.com/u/anon10827405)\
**Post date:** [December 8, 2023, 12:43pm UTC](https://community.openai.com/t/embedding-very-sensitive-to-punctuation/546205/5 "2023-12-08T12:43:14Z")

</div>

What was your chunking strategy? Could you provide some examples of the data you chunked?
