Did you see:
which leads to
which notes
Evaluation, tuning, and shipping safely
- Evals API for eval-driven development.
- Reinforcement fine-tuning (RFT) using programmable graders.
- Supervised fine-tuning / distillation for pushing quality down into smaller, cheaper models once you’ve validated a task with a larger one.
- Graders and the Prompt optimizer helped teams run a tighter “eval → improve → re-eval” loop.
Since the question did not note if this was just for ChatGPT and/or API, including all of the info.
Also check out
the noted tools AFAIK are not public but the ideas are valid.