I’ve been using GPT-5.6 Sol on Plus with High reasoning, and over the last couple of weeks the quality has noticeably deteriorated.
The main issue is not just that responses are shorter. The model often starts answering almost immediately on prompts that clearly require deeper analysis, even though High is selected.
OpenAI’s own documentation says High = extended thinking with GPT-5.6 Sol.
What I’m seeing instead:
- complex prompts sometimes get an almost instant response
- reasoning is much shallower than it used to be
- instructions are followed less consistently
- the model jumps to conclusions instead of checking alternatives
- answers often feel like a quick summary rather than an extended-reasoning response
- retrying the same task can produce dramatically different quality
I initially thought this might simply be a UI issue or that ChatGPT was silently using Instant instead of High, so I checked the requests in DevTools.
In affected turns, the client was explicitly sending:
model: gpt-5-6-thinking
thinking_effort: extended
and the response metadata still showed:
model_slug: gpt-5-6-thinking
requested_model_experience: thinking
did_auto_switch_to_reasoning: false
I also tested this in a brand-new chat with Fast answers disabled. The same general behavior still occurred.
So at least in my case, this does not look like simply selecting the wrong model in the UI.
I contacted OpenAI Support and provided request IDs, timestamps, and sanitized EventStream metadata. Support could confirm the documented behavior of High/extended thinking, but could not confirm any non-public routing, serving, or reasoning-budget changes.
I’m also clearly not the only person seeing this. There are multiple recent reports describing essentially the same symptom: GPT-5.6 Sol High responding extremely quickly and producing much shallower results despite Thinking/High being selected. Recent reports were still appearing on September 27-28.
One user even captured a clean GPT-5.6 Thinking/extended run where the server still reported the correct model, but time-to-first-token was only about 2 seconds and the expected reasoning behavior appeared absent.
Another Plus user reported that the initial response had become extremely fast and shallow, while using Try Again GPT-5.6 Sol sometimes produced a substantially deeper response.
I understand that response time alone does not prove how much internal reasoning occurred, and model behavior naturally varies. But the combination of:
- High explicitly selected
- extended thinking explicitly requested
- correct GPT-5.6 Thinking model reported
- no obvious model switch
- much faster responses on complex tasks
- noticeably worse reasoning quality
- multiple users independently reporting the same behavior
makes this feel like a real regression rather than normal output variance.
GPT-5.6 Sol High used to be much more reliable for complex technical, mathematical, and analytical work. Right now it often feels closer to a fast-response mode, despite High being selected.
Is anyone else still seeing this as of September 28?
It would be useful if OpenAI could clarify whether there have been any recent changes to the serving/execution path for GPT-5.6 Sol High, or at least acknowledge whether this is being investigated.