Some questions about text-embedding-ada-002’s embedding

@mbabayev

I am just following the “Algorithm 1” in the “All But The Top” paper mentioned above.

They pick out the top D dimensions of the PCA basis vectors to represent the embedding, which increases the isotropy, or spread, of the embedding vectors.

They don’t reduce the dimensions in the correlation, at least that is what “Algorithm 1” looks like to me, but if you insist on reducing dimensions, it looks like you want the v_jj scalers in the code. But I didn’t use those, since it looked like the paper didn’t use them.

Also, picking the top PCA components and reframing the embedding vector with these components and keeping the high 1536 dimensions has the advantage of: if you shift how many dimensions you want over time, you still have the same length vectors, and don’t need to “re-reduce” all the vectors again.

I was thinking of using this operationally on large databases, so I didn’t want to re-compute everything if I made a minor change. Minor changes just lead to minor shifts in the vectors, not dimensional changes that make them incompatible.

So think of D as a tunable parameter, where if D = 1536, you get exactly the original Ada-002 vector back, and if D < 1536, you are getting back a vector with less correlation and anisotropy and more meaningful spread and variation, which gives a much higher range to your cosine similarity, which was the original goal.

All vectors produced have 1536 dimensions, which makes them compatible with each other as you vary the tunable parameter D.