Hi, have a bunch of multi-turn agent traces (approx 12-16 steps long) generated from o4-mini, want to try the new RFT fine-tuning API. In the examples posted online, the setup is single-turn convo, but since I am using for an agentic usecase, curious if submitting an entire agent trace of multiple Q/A’s along with a grader for the entire agent sequence would work well? Or am I better off isolating and picking out specific agentic turns?
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Autoregressive Fine-Tuning for Chat Models | 0 | 199 | July 10, 2024 | |
| Reinforcement Fine Tuning using gpt-4.1-mini | 2 | 141 | May 15, 2026 | |
| Fine tuning with function calling / tools help! | 2 | 290 | November 27, 2024 | |
| Instrruction tuning for GPT API | 7 | 669 | March 11, 2024 | |
| How does gpt-3.5-turbo fine-tuning work? | 10 | 2051 | September 11, 2023 |