GPT-5.6 High execution regression: ordinary Chat went 102m26s -> repeated ~25–26m stops on same Plus account; Work still gets long-run handoff

Aug 25 update: the question is no longer “is there a ~26-minute pattern?”

@OpenAI_Support

I am updating this thread because the evidence has moved well beyond one suspicious timer or a subjective “Thinking feels shorter” report.

I now have:

  • a fresh Aug 25 native reproduction;
  • a historical/current native comparison on the same Plus account;
  • a near-simultaneous Chat-vs-Work control on the same current builds;
  • and independent users reporting the same hour-scale → ~25-minute collapse.

The question is now much narrower:

Why did affected ordinary Chat lose the longer-running execution path it demonstrably received before?


1. Fresh Aug 25 native reproduction: 25m50s

Today one of my genuine GPT-5.6 Thinking / Extended ordinary-Chat turns showed:

Worked for 25m 50s

I did not treat the UI timer as proof.

Earlier in this investigation I found a misleading 46m18s visible turn that actually contained two backend request/exchange segments. Since then, I have bound long-looking turns to their native request objects before using their duration as evidence.

This Aug 25 specimen is clean:

  • finished_duration_sec = 1550
  • actual reasoning elapsed: 1550.434s
  • model: gpt-5-6-thinking
  • thinking effort: extended
  • one working turn
  • one request ID
  • one turn_exchange_id
  • no async_source
  • no per-turn stream topic
  • no Temporal marker
  • no wfr_ worker request

So this is not a multi-exchange UI-timer artifact.

My current native genuine GPT-5.6 / Extended foreground-duration set is now:

1541, 1549, 1550, 1557, 1560, 1564, 1564 seconds

That is seven native turns across three conversations in a remarkably narrow range. Five of the seven are repeated turns in one conversation, so I am not presenting them as seven independent reproductions.


2. The same account previously had a 102m26s ordinary-Chat turn

The historical control is also native evidence, not a screenshot.

Historical ordinary Chat

  • 6146s / 102m26s
  • GPT-5.6 Thinking / Extended
  • one request/exchange
  • wfr_ request
  • SAServer async_source
  • longer-running worker execution path

Fresh Aug 25 ordinary Chat

  • 1550s / 25m50s
  • GPT-5.6 Thinking / Extended
  • one request/exchange
  • no wfr_
  • no SAServer async_source
  • no equivalent worker/handoff markers

So the comparison is:

Historical ordinary Chat: 102m26s, worker-routed

Current ordinary Chat: 25m50s, foreground/non-worker

Same GPT-5.6 Thinking family.

Same Extended reasoning effort.

Single request/exchange on both sides.

The duration ratio is approximately 3.97x.

I am not claiming that the route difference alone is proven to cause the entire runtime difference.

I am saying it is now the strongest observable discriminator between preserved hour-scale ordinary Chat and the current ~25–26m execution class.


3. The Aug 24 Chat-vs-Work A/B makes an account-wide explanation difficult

On Aug 24 I captured ordinary Chat and Work on the same Plus account, using the same current client/server build identifiers, within the same few-second window.

The ordinary-Chat request remained on its original foreground SSE execution path.

A Work turn began 1.540 seconds before that Chat SSE had even closed.

The Work turn received:

  • stream_handoff
  • a conversation-turn-* topic
  • a wfr_ request ID
  • SAServer execution
  • WebSocket continuation
  • temporal_conversation_turn = true

So, during essentially the same moment on the same paid account:

ordinary Chat → foreground/non-worker execution

Work → per-turn handoff/worker execution

This makes several simple explanations fit badly.

The account was not globally unable to receive worker execution. Work received it while the ordinary-Chat turn was still active.

The client/server build was not the differentiator in that A/B.

And this is not dependent on the separate GPT-5.5-mini routing problem: today’s 1550s specimen genuinely resolved to GPT-5.6 Thinking / Extended.


4. Independent users are reporting the same before/after runtime collapse

I am treating the following as timing corroboration only.

These reports do not prove that the users have my exact backend route.

But the pattern is no longer confined to one account, one browser or one workload.

Independent report #1: >90–120m before, ~25m now

In the Reddit thread “Limited ChatGPT execution time”, one user reports that their ChatGPT tasks previously ran:

well beyond 90 minutes, often 120 minutes

but recently began cutting out at around:

25 minutes

That is a direct same-user before/after comparison: hour-scale execution before, ~25-minute execution now.

Independent report #2: 107m last week, now 25–26m repeatedly

In the separate Reddit thread “ChatGPT Plus Thinking limited to 25 minutes”, another High user reports:

  • a 107-minute ordinary-Chat High run last week;
  • current stops consistently between 25 and 26 minutes;
  • roughly 10–15 occurrences across four different projects;
  • and failure wording indicating that the tool window ended before the requested work was complete.

This is particularly useful because it is not one anomalous turn.

The same user reports the new ~25–26m behavior repeatedly across several projects after previously receiving a 107-minute High turn.

Independent report #3: Pro High at 25m56s, Extra High later lasts longer

The same 25-minute discussion contains a ChatGPT Pro user reporting ordinary Chat:

  • High: 25m56s, unfinished
  • Extra High: approximately 35m, also unfinished

I am not claiming that ~35 minutes is a universal Extra High limit.

What makes this useful is that the observed execution boundary changes when the reasoning effort changes.

That is difficult to reconcile with a generic browser timeout or a single fixed network lifetime.

Independent report #4: two Plus accounts

The original poster in that same GPT-5.6 High timing discussion reports testing two separate Plus accounts, with both newly reaching roughly the same ~25-minute maximum on substantial GPT-5.6 Sol High work.

That does not prove a universal account policy.

It does weaken an explanation based on one damaged account or one isolated conversation.

Independent Work control

The same discussion also contains a separate user reporting current Work runs of:

  • 33m57s
  • 41m06s
  • approximately 43m49s

That independently points in the same direction as my native Chat-vs-Work control:

ordinary Chat clustering near ~25–26m while Work can continue materially beyond that range

Again: these third-party reports are not backend-route proof.

They are independent timing evidence.

The important repeated pattern is:

hour-scale ordinary Chat before → ~25–26m ordinary Chat now

For reference, the two public Reddit discussions are:

Limited ChatGPT execution time
r/codex, thread ID 1vwwxo0

ChatGPT Plus Thinking limited to 25 minutes
r/ChatGPT, thread ID 1vuns4u


5. Ordinary browser Chat historically used a per-turn handoff mechanism

There is also independent technical evidence from March 2026, completely separate from my captures.

A browser reverse-engineering report documented ordinary Chat with high/extended reasoning using a two-stage continuation flow:

/backend-api/f/conversation

then:

resume_conversation_token

then:

stream_handoff

then:

conversation-turn-*

then WebSocket continuation on that turn topic.

The technical report is:

xtekky/gpt4free issue #3404
“[Bug] [OpenaiChat] extended thinking_effort failed using API, however good on chrome browser within vnc”

I am not claiming that this external capture independently proves SAServer, wfr_, or Temporal execution.

It does independently establish that ordinary browser Chat high/extended historically used a server-issued per-turn handoff/continuation architecture.

That matters because current affected ordinary-Chat turns are not receiving that treatment.


6. I tried to make the route/admission hypothesis fail

I do not want to find evidence only for the explanation I already prefer.

So I tried to make it fail.

Browser/network timeout?

Poor fit.

In my strongest 1564s specimen, active reasoning stopped at 1564 seconds while the same HTTP/2 SSE connection remained healthy for another ~60.9 seconds.

The connection later closed normally with HTTP 200.

The required workload was still objectively unfinished.

So the network connection itself did not die at the ~26-minute boundary.

GPT-5.5-mini fallback?

Not required.

The fresh Aug 25 1550s specimen genuinely resolved to GPT-5.6 Thinking / Extended.

The mini-routing problem is separate.

Plus accounts simply cannot receive long-running execution anymore?

Poor fit.

Work received worker execution on the same Plus account while the ordinary-Chat request was still active.

Client/server build difference?

Poor fit for the Aug 24 A/B.

Chat and Work were captured using the same current build identifiers within the same few-second window, yet were assigned observably different execution architectures.

Did Aug 20 create a brand-new 26-minute timer?

I no longer think that is the best explanation.

Public ~25–26m examples existed before Aug 20.

That means the strongest current model is not:

“OpenAI created a new 26-minute timer on Aug 20.”

A better fit is:

A shorter foreground/tool-execution class already existed.

Ordinary Chat could also receive a longer-running handoff/worker execution class.

Affected ordinary-Chat High traffic that previously received the longer-running class now appears to remain in the shorter class instead.

The exact admission predicate is still unknown.

Product surface, model variant, reasoning effort, conversation mode and other fields co-vary between Chat and Work.

I am therefore not claiming to know the exact server-side decision rule.


7. The ~25–26m class itself appears to predate this regression

There are older public reports of ChatGPT/tool workloads repeatedly landing around the same ~25–26-minute range.

That suggests the short execution envelope itself is not necessarily new.

What appears new for the affected cohort is the loss of the hour-scale path.

This interpretation fits both sides of the evidence better:

Before

Some ordinary-Chat High/Extended turns received a server-issued handoff/continuation path and could run for well over an hour.

My preserved native example:

102m26s

Independent users report historical examples around:

90m
107m
120m

Now

Affected genuine GPT-5.6 High/Extended ordinary-Chat turns repeatedly remain in the shorter execution class.

My native set:

1541, 1549, 1550, 1557, 1560, 1564, 1564 seconds

Independent current reports:

~25m
25m56s
25–26m repeatedly

This is why I am describing the issue as an execution-routing/admission regression, not simply “a new timer.”


8. I looked for counterexamples too

I specifically searched for fresh post-regression ordinary-Chat High single turns clearly exceeding 60 minutes.

I did not find a clean one.

That is absence of found evidence, not proof that none exist.

I would actually welcome a valid current >60m ordinary-Chat High specimen.

If some users can still receive hour-scale ordinary-Chat High execution while others repeatedly land at ~25–26m, that would weaken a universal-cap theory but make a conditional execution/admission regression even more interesting.

Useful details would be:

  • plan
  • model
  • reasoning effort
  • date
  • Worked for duration
  • whether the requested work completed
  • whether it was ordinary Chat rather than Work/Codex
  • ideally native metadata showing whether the visible duration spans one request/exchange or several

A counterexample would be useful evidence, not something I am trying to argue away.


9. I cannot find a documented ~25–26m High per-turn policy

OpenAI’s current GPT-5.6 Help documentation still describes High as:

Extended reasoning

It documents reasoning usage allowances and fallback behavior.

I cannot find a disclosed ~25–26 minute ordinary-Chat High/Extended per-turn execution ceiling.

The Aug 20 Thinking-mode incident is also marked resolved.

I am not claiming that incident caused this regression.

The clean Aug 25 reproduction simply establishes that this behavior remains observable after that incident was declared recovered.

If a ~25–26m High execution envelope is intentional product behavior, documenting that would resolve a major part of this report immediately.


10. What is actually proven, and what remains unknown

To keep this report falsifiable:

Directly demonstrated on my account

  • ordinary Chat historically produced a 102m26s GPT-5.6 Thinking / Extended single turn;
  • that historical turn carried worker/SAServer execution markers;
  • current genuine GPT-5.6 Thinking / Extended ordinary-Chat turns repeatedly cluster around 1541–1564s;
  • today’s fresh reproduction is 1550s / 25m50s;
  • the 1550s specimen is one request/exchange, not a combined UI timer;
  • the strongest 1564s specimen did not coincide with loss of its HTTP/SSE connection;
  • Work on the same Plus account still receives the per-turn handoff/worker execution architecture;
  • in the Aug 24 A/B, Work began while the ordinary-Chat SSE was still active.

Strongly supported by independent public reports

  • other users previously received ordinary-Chat High execution in the 90–120m range;
  • affected users now independently report ~25–26m unfinished execution;
  • one user reports 107m before → 25–26m now, repeated 10–15 times across four projects;
  • the behavior is reported on more than one Plus account and also by a Pro user;
  • current Work can still exceed the ~25–26m range.

Not established

  • that every ChatGPT account is affected;
  • that ~25–26m is a universal hard cap;
  • that Aug 20 caused the regression;
  • that GPT-5.5-mini fallback causes this regression;
  • that one particular client field controls worker admission;
  • that Work and historical ordinary Chat use every identical backend component;
  • why the server is making the current execution-class decision;
  • whether the change is intentional.

Those last two are exactly what I am asking OpenAI to clarify.


The engineering question is now very narrow

@OpenAI_Support

I am not asking for speculative explanations.

I am not asking you to disclose proprietary infrastructure.

I am asking for a classification.

Q1: Is this expected behavior?

Is approximately 25–26 minutes now an intentional per-turn execution envelope for ordinary Chat High/Extended on affected traffic?

If yes:

Please confirm that and point to the documentation or effective date for the change.

Q2: If it is not expected, is affected ordinary Chat failing admission to the longer-running execution path?

Why do current affected High/Extended turns no longer show the longer-running handoff/execution treatment that preserved historical ordinary Chat demonstrably received, while Work on the same account still does?

Q3: Can Engineering correlate the preserved historical/current request IDs?

I can provide privately:

  • exact historical/current request IDs
  • turn IDs
  • timestamps
  • validated sanitized Aug-23 HAR transfer copies
  • native conversation backups
  • client/server build identifiers
  • SHA-256 hashes of the preserved evidence

Those artifacts should allow the relevant team to determine what execution treatment the historical and current turns actually received.

I do not need internal implementation details posted publicly.

I need one of two answers:

If ~25–26m is now expected ordinary-Chat High behavior, please say so and document the change.

If it is not expected, please log/escalate the historical/current request IDs as a conversation-execution/admission regression.

That is the issue I am trying to get resolved.