Set Astra - Low/Medium as a replacement for 5.6 - Sol High

OpenAI’s advice is to treat GPT-6-Astra with reasoning set to low as the replacement for GPT-5.6-Sol with reasoning set to high.

While this can provide measurable improvements in cost and latency, it still allows you to use higher reasoning levels by having Astra with reasoning set to low delegate tasks to Astra with reasoning set to high.

Essentially, Astra-low is the new recommendation for agent orchestration. GPT-5.6-Luna and Terra are still available if you want to realize additional token, cost, and latency savings.

Of course, all of this depends on the specific use case, but it should provide a useful baseline for your evaluations.

What are you seeing and measuring in your applications and use cases? Is there anything else you would like to know to get more out of the Astra release?

You should add that to the Codex tips topic.

Even if you do not, this should not appear as a link under the first topic.

Interesting comparison. I’d be curious to see how Astra Low performs on longer tasks where consistency matters more than raw speed. The cost difference looks promising, though.

Good point: reasoning also would influence the persistence at long term goals and multi-turn use, and pregenerating the plans persisted internally as context.

Examine the graph another way. The lowest Astra reasoning setting is costing you as much as Sol “xhigh”, with its starting point of costs being vertically aligned with the second-to-last Sol pip.

Terminal Bench shows that the “right answer” might be judged with more successes at an overlapping cost point, but when you factor the cost amplification that translates Astra tokens to the right at the same reasoning points, we see a “easy” 20% accuracy IQ but multi-faceted task that simply takes extended tool use token production to complete, and doesn’t result in AI model failures and retries and churn, might be instead delivering you merely high model costs without an “Astra lite” available.

So “replacement” would depend on how much agency you give your agent beyond the size of one Posix sequential benchmark - and how much production besides the right bash command line is needed. “Translate each page of this book” - a high token output task that is not easy to benchmark or measure a word error rate, but is easy to observe a linear relation in token costs.

Positive: Astra can actually understand better who the job is targeting and write for that audience through several layers of system, developer, user, and different documents and voices.