Some questions about text-embedding-ada-002’s embedding

No idea if internal embeddings are changed with a fine-tune. I just assumed the neural weights changed. The main reasoning is that the semantics, once trained, shouldn’t change, and the fine-tune just reshapes the output from the input (unchanged) semantics.

Yes the codex is primarily trained on code. But rolling everything into one new model, and deprecating the rest, suggest they went with a Mixture of Experts (MoE) thing similar to GPT-4 (rumors, I know), and this consolidation would basically get the new models all on the same architecture, possible saving some money in the process.

Go for it!