OpenAI tasked GPT-5.6 with optimizing its own runtime efficiency.
The results:
- 20% lower serving costs through production GPU kernel improvements.
- More than 15% better token-generation efficiency through improved speculative decoding.
These improvements are being passed on to everyone using the API, Codex, and ChatGPT.
Starting today, GPT-5.6 Luna, the fastest and most affordable model, will cost 80% less. GPT-5.6 Terra, the balanced model for everyday work, will cost 20% less.
Along the price reductions for GPT-5.6 Luna and Terra, Fast mode for GPT-5.6 Sol in the API delivers up to 2.5x the speed of Standard processing at twice the Standard price. It gives API customers faster access to GPT-5.6 Sol without changing the model’s intelligence.
OpenAI is also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna. Combined with Luna’s new pricing, Auto-review is expected to cost about 10x less. This affects auto-approve mode (“review for me” in the app), which now uses Luna to prevent many high risk actions from the main agent
