GPT-5.6 High execution regression: ordinary Chat went 102m26s -> repeated ~25–26m stops on same Plus account; Work still gets long-run handoff

5.6 is back to serious regression

Its refusing to do tasks,it is repeating itself and is seriously regression why do I feel like this is due to gpt 6 sol and Luna releasing Everything was fine but as soon as gpt 6 sol and Luna released in work chat , project chats and normal chats for 5.6 is seriously regressing again so we are back to the same issues we had weeks ago great

Follow-up / instrumentation correction to #163.

I found and fixed a flaw in my forensic recorder that is important enough to put on the public record.

The September 24 control is still a foreground/no-worker turn, but one piece of live instrumentation was wrong.

During that turn the V3.3.1 toolbar showed SPAINCENTRAL, which initially looked like a current-turn SAServer worker region.

It was not.

The ChatGPT user WebSocket is account-scoped. While the foreground turn was running, it delivered a background conversation-update for a completely different chat on the same account. That unrelated chat really did have a SpainCentral SAServer async_source.

Recorder V3.3.1 was scanning every received account-WebSocket frame for worker metadata without first proving the frame belonged to the open conversation/turn.

So the region badge was contaminated by another chat.

That bug could also contaminate the derived WORK/LIVE state, so I am withdrawing V3.3.1’s live region/vitality pill as standalone proof of current-turn worker identity.

The raw HAR evidence is much cleaner and does not depend on that pill.

For the September 24 GPT-5.6 Thinking / Extra High control, the actual current-turn route was:

  • plain /backend-api/conversation HTTP/SSE;
  • one resume_conversation_token(kind=topic);
  • zero stream_handoff;
  • zero current-turn SAServer async_source;
  • zero current-turn worker heartbeat;
  • zero current-turn worker-topic WebSocket stream;
  • zero typed current-turn safety_review_update.

The turn then completed normally after 5m51s with:

finished_successfully -> end_turn:true -> message_stream_complete

So this is another useful counterexample:

foreground does not automatically mean failure, but this turn still did not receive the historical regular-worker admission path.

There is another correction that matters here: the resume_conversation_token by itself is not worker admission.

I compared it against preserved healthy GPT-5.6 worker captures.

The healthy sequence is:

resume_conversation_token
→ stream_handoff
→ subscribe to the matching conversation-turn-* topic
→ current-turn worker stream / heartbeats
→ SAServer async_source

The Sep-24 foreground turn stopped after the first step.

So from now on I am treating stream_handoff + scoped matching current-turn worker traffic as the worker-admission evidence, not the presence of a topic token or a region label alone.

I rebuilt the recorder as V3.3.2 with two corrections:

  1. account-level WebSocket worker evidence is now scoped to the active conversation/turn before it can affect region or vitality;
  2. the newer direct-SSE terminal format seen on Sep 24 is recognized, so a completed foreground turn can correctly move ADMIT HTTP -> DONE instead of remaining stuck.

I replayed the exact Sep-24 capture through the fixed logic:

  • 88 account-WebSocket frames observed;
  • 1 genuine SAServer async_source frame existed, but it belonged to the other conversation;
  • 0 frames were accepted as current-worker traffic;
  • region = unknown;
  • heartbeat = none;
  • final state = DONE.

A known-good September 18 worker capture still passes the same filter, keeps its matching conversation-turn-* worker traffic, heartbeat evidence and CENTRALUS canonical worker region.

So this correction does not weaken the main conclusion in #163.

It makes the boundary cleaner:

Sep 21 and earlier positive controls: real regular worker handoff exists.

Sep 22/24 affected controls: foreground SSE, no stream_handoff, no matching current-turn worker stream.

What I am correcting is the instrumentation rule around region/vitality, not retroactively converting these foreground turns into workers.

No one needs to rerun a paid workload for this. If you already have a natural capture, the useful discriminator is still the current turn’s actual handoff/topic/heartbeat evidence, not a model label, region badge, or topic token by itself.

Same here. To get a task completed, I now have to provide baby-level hand-holding. Without it, GPT5.6 keeps repeating pointless things or fails to accurately grasp the actual point of a GPT-5.6 task.
At this point, continuing to pay for Pro for any reason other than Codex feels like a waste of money.

Yeah its really frustrating :confused:

Also, GPT-5.6 Sol Extra High in ChatGPT Chat suddenly gets disconnected even on simple tasks, almost like hitting a Codex capacity warning, so I can’t even complete tasks in ChatGPT Chat
GPT-6 Sol in Codex is excellent, but I feel incredibly stupid for having paid ¥16,800 for Pro

Its really frustrating how broken normal chats and projects are seriously broken but yet work chat works just fine :unamused_face:

What does that do and how will that help

Reports of chat getting combined limited usage with work and codex next devday, that’s less than a week away, enjoy while it lasts :+1:

It seems that even after setting DNS to 1.1.1.1, the random disconnections still occur, even when I simply leave ChatGPT running without touching anything
Things have worsened since September 14. Considering that I have evidence showing that Extra High was able to run for an hour at that time, it is now difficult to say that even Pro’s 5x capacity can completely avoid the 25-minute limit on High
In fact, back when ChatGPT could run for 100 minutes, I was able to implement production-level code for Secure MCP Tunnel in just eight days, with everything fully debugged. By contrast, considering how unproductive my life has been since August 20, 2026, it is clear just how much a ChatGPT session capable of running for 100 minutes improves my productivity
The way Extra High disconnects is also random. Sometimes it simply stops silently, while at other times requests to the tunnel are suddenly interrupted. The agent does not respond, and the turn simply ends
Requests that finish within one minute, however, complete normally in almost every case

Hi, can you show me the source to that claim? Thank you

If that happens then many users will not be happy and where are you getting this information

Hi everyone,

I completely understand why there is a lot of anxiety and speculation about this right now! With OpenAI DevDay coming up on September 29th, everyone is on high alert for major platform changes, and it’s easy to see how the recent discussions could point to that conclusion.

To help clarify what is currently known:

• The Rumor Source: The speculation largely stems from community interpretations of recent tweets by Tibo (OpenAI’s engineering lead for Chat and Codex) regarding upcoming updates to workspaces and interface management.

• Current Structure: Right now, standard web Chat remains entirely separate from the heavily metered ChatGPT Work and Codex pools. Work and Codex modes handle resource-heavy tasks (like code execution and the newer GPT-6 Astra features), which is why they carry those strict 5-hour and weekly caps.

• What’s Actually Expected: While a merge hasn’t been announced, platform trackers (like TestingCatalog) suggest OpenAI is actually preparing to introduce a brand-new, higher-tier plan (potentially ChatGPT Pro Max) aimed at power-users who need higher limits for managed AI agents.

While OpenAI hasn’t officially announced plans to restrict standard chat with a Work-style limit, it’s definitely worth keeping a close eye on the official keynotes on September 29th. Hopefully, we get a clear answer on the future of workspace limits then!

Extra High times out after 13 minutes and 3 seconds with 66 tool calls through local MCP, while High times out after 14 minutes and 12 seconds with 78 tool calls. This is only about one-quarter of the previous 60-minute limit
These timeout limits are completely unreasonable, and it has become impossible to make any meaningful progress on tasks without using Codex
Codex is weaker with very large contexts and quickly loses track of instructions, so I preferred ChatGPT. However, ChatGPT now has an even shorter limit than the previous High mode, and I can no longer entrust it with bug fixing even with Pro 5x
Today, absolutely no progress was made on the bug fix because ChatGPT kept modifying pointless code instead of addressing the actual issue. This is the first time I have ever experienced this
Also, the timeout behavior does not appear to have changed even in the new UI that is currently being A/B tested
As a result, it is now impossible to move the project forward without Codex, so I have no choice but to use Codex

It is ridiculous but yet work chat works fine but normal chats and project chats remain broken

Given this disastrous situation, the time might come when we have to use GPT-6 Pro to fix the bugs.
I am considering not using the Codex harness because it is inefficient.

This is exactly what worries me about the direction things seem to be moving in.
OpenAI’s own consumer research says roughly 70% of ChatGPT usage is non-work related, and only around 30% is work related. So most users are clearly not here because they want to hand off huge professional workflows to agents all day.
And yet more and more attention seems to be going toward Work, Codex, long-running agents, enterprise use cases, etc.

I have nothing against those tools. If someone wants an agent to work through a huge coding task for an hour, great. But that is not the same thing as normal Chat.
Creative writing, art projects, worldbuilding, research, long discussions, learning, continuity-heavy conversations these can all be complex and demanding without being “agentic work.” Sometimes I do not want Codex. I do not want an autonomous worker. I want a strong conversational model that can reason properly, remember the context, follow instructions, and stay coherent over a long session.

If normal Chat keeps getting shorter limits, weaker long-context behavior, worse reasoning continuity, or more aggressive cutoffs while Work/Codex gets the stronger execution path, then people are effectively being pushed into Work whether they want that workflow or not.
And that seems risky considering how large the consumer side of ChatGPT actually is.

I also cannot help wondering how much of this direction is being driven by the race with Anthropic and by the fact that enterprise/agentic AI is probably much more attractive to large corporate customers and investors. I am not claiming that is definitely the reason, but strategically it would make sense.

The problem is that this is still a much smaller and very different audience from the broader consumer base. If OpenAI optimizes too hard for enterprise agents and heavy professional workflows while normal Chat becomes the weaker tier underneath them, I think they risk alienating a lot of the users who made ChatGPT popular in the first place.

Chat should not become a crippled waiting room for Work and Codex.