I’ve figured out that issue with not reproducing same output with defined seed is related with calculating question embeddings which tend to produce varying results. So for testing purposes to achieve comparable results you need to retrieve once and then reuse (mock) embeddings for defined set of test questions. More about this issue you can find in thread below:
egils
15
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Deterministic Results Impossible for GPT-4o | 7 | 2708 | August 26, 2025 | |
| ChatCompletions are not deterministic even with seed set, temperature=0, top_p=0, n=1 | 9 | 2403 | October 7, 2024 | |
| Is the seed parameter getting deprecated? | 1 | 1324 | October 19, 2025 | |
| AI model fingerprints are not unique, making them fairly useless for tracking model updates | 15 | 3871 | May 22, 2024 | |
| Question about the Use of Seed Parameter and Deterministic Outputs | 3 | 5251 | May 29, 2024 |