Why is GPT-5.6 failing long-context tasks?

I’m still seeing this as well, and one detail in your report may be more significant than it first appears.

Your failed long GPT-5.6 Sol run lasted approximately 25m54s, which is 1,554 seconds.

I have been independently investigating a long-running ordinary-Chat execution regression with native ChatGPT metadata and HAR captures. My preserved genuine GPT-5.6 Thinking / Extended foreground specimens currently terminate at:

1541, 1549, 1557, 1560, 1564 and 1564 seconds.

So your independent 1554-second run falls directly inside that very narrow cluster.

My strongest 1564s specimen is particularly useful because the network request itself did not fail at 26:04. The same HTTP/2 SSE connection remained healthy for another ~60.9 seconds and completed normally with HTTP 200, while the required workload was objectively unfinished. That substantially localizes the boundary to active execution/reasoning rather than a simple browser/network timeout.

I cannot establish from your public evidence that your 1554s run used the same backend execution class, so I do not want to overclaim that.

But the timing is close enough that I think Engineering should compare them.

There is another important development from today.

Earlier I had a separate browser/session problem where GPT-5.6 Thinking requests were resolving to GPT-5.5-mini. That environment now appears to have recovered sustained GPT-5.6-like behavior.

However, the long-running execution-routing problem remains separate.

Today I captured ordinary Chat and Work on the same Plus account, exact same frontend client build/version and exact same server conversation build.

Ordinary Chat remained on the foreground path with:

  • no stream_handoff

  • no wfr_ worker identity

  • no SAServer source

  • no per-turn worker topic

At essentially the same moment, GPT-5.6 Sol (max) in Work received:

temporal_conversation_turn=true
→ stream_handoff
→ exact conversation-turn-* topic
→ wfr_ worker execution
→ SAServer

The second Work worker was actually admitted while the ordinary-Chat foreground request was still running, with approximately 1.54 seconds of overlap.

That means the worker infrastructure was demonstrably available to the same account at the same instant.

So I now think there are at least two separable issues:

  1. the reported GPT-5.6 → mini routing problem, which can change with browser/session treatment;

  2. a broader ordinary-Chat execution/admission regression where genuine GPT-5.6 remains foreground/non-Temporal while Work still receives worker routing.

I documented the native execution-class evidence separately in my Community topic 1391943.

Your 25m54s / 1554s result is one of the strongest independent timing observations I’ve seen because it lands directly inside the native foreground cluster above.

If OpenAI Engineering investigates these reports, I think the useful comparison is not merely “does GPT-5.6 feel worse?” but whether affected long-running turns are being assigned a different execution/admission class than they were before the Aug 19–20 change.