Hello,
I am creating this as a dedicated Bugs topic after a Community Leader asked me to separate these issues into their own report so they are easier for the appropriate OpenAI teams to see and investigate.
I want to keep this report respectful, factual, and focused on reproducible behavior.
I am a ChatGPT Pro subscriber, and for approximately the past six days, beginning around the August 19–20 service problems, a cluster of ChatGPT/GPT-5.6 issues has made my normal long-form workflow effectively unusable.
I appreciate that OpenAI has now acknowledged and fixed one confirmed regression that was unintentionally moving approximately 3% of Pro and Thinking turns to GPT-5.5-mini.
That acknowledgment was important and appreciated.
However, the routing regression was clearly not the only problem, because multiple other failures are still occurring, including on workflows using GPT-5.6 Sol High and Extra High Thinking.
I am not claiming every symptom below necessarily has the same root cause.
I am documenting the remaining behaviors so the appropriate ChatGPT, GPT-5.6, conversation-infrastructure, long-context, file-retrieval, execution/runtime, streaming, and UI teams can investigate them separately if necessary.
Environment
- Plan: ChatGPT Pro
- Model: GPT-5.6 Sol
- Reasoning levels: High and Extra High
- Primary surface: ChatGPT web
- Device: Apple iPad
- Operating system: iPadOS 26.6.1 (23G83)
- Browser: Chrome for iOS / iPadOS
- Chrome version: 152.0.7977.64
- Workflow: Long-running complex creative-writing/project workflow
- Typical workload: very long governing prompts, substantial conversation context, uploaded text files, source hierarchy rules, exact predecessor-file retrieval, continuity verification, research, and requested 20,000-word story continuations
- Variants of the underlying context/file/Thinking failures have also reproduced in new conversations, not only one old thread.
- Some earlier failures were also reproduced in Safari, so the broader GPT-5.6 reliability problems have not appeared limited to one Chrome session.
The new UI issue described below is specifically occurring for me on Chrome mobile web on the iPad.
1. NEW UI REGRESSION — Thinking / Activity sidebar can no longer be opened
This is the newest issue I am reporting.
Previously, while GPT-5.6 was Thinking and performing a task, I could open the right-side Activity / Thinking panel and see the ongoing Thinking stages, searches, file operations, Python/tool activity, and other execution information.
I have screenshots showing that this panel previously worked normally on the same general ChatGPT web workflow.
Now, on Chrome mobile web on my iPad, I can no longer reliably pull up that sidebar while the model is Thinking.
The panel that previously opened from the right side is effectively unavailable/inaccessible.
This makes debugging the other problems much harder because I can no longer reliably inspect what ChatGPT is doing during the Thinking phase.
This does not appear to be only my account.
I raised this in another Community discussion, and a Community Leader replied:
“Thank you for raising this!
Please create a new topic in the bugs category.
If the team will see it then it’s more likely if the report is not hidden inside a different topic.”
That is why I am making this dedicated report.
Expected behavior
While GPT-5.6 is Thinking, I should be able to open the Activity/Thinking sidebar and inspect the visible execution stages as before.
Actual behavior
The sidebar can no longer be reliably opened on Chrome mobile web on my iPad.
Screenshot comparison
I have attached screenshots showing:
- the previous working state, where the Activity/Thinking panel was visible on the right;
- and the current state where I cannot pull that panel up normally.
I would appreciate confirmation whether this is an intentional UI change or a regression.
2. High / Extra High Thinking still behaves inconsistently
Even after the confirmed 5.6 → 5.5-mini routing regression was acknowledged, I am still concerned about the behavior of complex GPT-5.6 Sol High and Extra High tasks.
Before these recent problems, genuinely difficult requests involving large amounts of context, file retrieval, continuity verification, research, and generation would normally spend substantial time reasoning and executing.
Since the regression period began, I have seen dramatically inconsistent behavior.
Some complex High/Extra High turns:
- complete far faster than would be expected for the requested workload;
- miss clearly supplied instructions;
- fail to apply previously acknowledged continuity;
- incorrectly say required source material is unavailable;
- or switch into explanation/planning mode instead of actually performing the requested task.
The problem is not simply that a response is “fast.”
The problem is that the unusually fast responses frequently correlate with missing work or failed execution.
3. Long prompts / long-context information is not being applied reliably
My workflow depends on long governing prompts containing:
- source hierarchy;
- continuity rules;
- exact story seams;
- character/location state;
- relationship continuity;
- body-state continuity;
- research requirements;
- file-reading requirements;
- formatting requirements;
- and explicit instructions about what the final response must contain.
Before the recent regression period, GPT-5.6 handled this workflow far more reliably.
Now, ChatGPT can correctly acknowledge information in the prompt and then behave shortly afterward as though that same information was never supplied.
Examples include:
- correctly identifying an instruction and then failing to apply it;
- treating an established workflow like a new standalone prompt;
- saying a source or rule is missing when it is visibly present;
- recovering the correct continuity during Thinking but not carrying it into final execution;
- and producing an explanation of what should be done instead of doing it.
This feels less like ordinary forgetting and more like inconsistent context ingestion, retrieval, prioritization, compaction, or execution-state retention.
4. Uploaded files are sometimes falsely reported as missing
This remains one of the most serious issues.
My story workflow frequently requires an exact predecessor file.
I have had GPT-5.6 tell me that the required uploaded story file was unavailable and instruct me to upload it again.
In one specific example, GPT-5.6 High spent approximately 18 seconds and then told me that my accepted Part Thirty-Eight story file was not available.
However, another attempt was able to access the same required material and correctly recover the predecessor continuity.
I have also had the system directly access and verify a file and then subsequently behave as though that file state was unavailable.
So the problem is not simply:
“The user forgot to attach the file.”
The contradictory behavior is itself part of the bug.
Expected behavior
If a required uploaded file is available to the conversation and successfully retrieved, that file should remain reliably usable throughout the turn.
Actual behavior
The same source can be:
- treated as available in one execution;
- treated as missing in another;
- or successfully read during Thinking but not reliably carried into final execution.
5. Verification / planning responses are replacing the requested output
This is one of the clearest examples I have documented.
I requested an actual 20,000-word story continuation.
GPT-5.6 worked on the request for approximately:
25 minutes 54 seconds
During the run it correctly:
- recovered controlling source material;
- recovered continuity;
- identified the correct predecessor state;
- performed research;
- drafted/analyzed the continuation;
- and apparently conducted extensive verification.
But instead of returning the requested story prose, the final response told me that:
- a complete prose draft had been constructed;
- it contained exactly 20,000 words;
- it had approximately 694 paragraphs;
- it had a specific median paragraph length;
- duplicate-paragraph and repeated-sequence checks had been performed;
- and additional final verification had been attempted.
The actual story itself was completely absent.
This is not the requested behavior.
The verification process is supposed to support the requested result.
It cannot replace the requested result.
Expected behavior
If the model spends 25+ minutes constructing and verifying the requested story, the final response should contain the story.
Actual behavior
The model returned a report describing the supposedly completed work instead of delivering the work.
I have seen variations of this broader pattern where the model:
- explains the task;
- lists what it intends to do;
- summarizes what the result should contain;
- says it has performed work;
- or verifies preparation;
and then fails to execute the actual requested output.
6. Long-running execution appears to stop before the actual task is complete
The ~25–26 minute behavior may be related to the verification-only problem, but I am listing it separately because other users are investigating similar long-running execution behavior.
For difficult tasks, ChatGPT may:
- begin normally;
- read files;
- research;
- perform continuity analysis;
- construct or verify material;
- spend around 25–26 minutes processing;
and then terminate the useful execution without completing the required output.
My 25m54s story example is one concrete case.
This deserves investigation independently of the now-confirmed 5.5-mini routing regression because users are also investigating long-running failures on turns that appear to be genuine GPT-5.6 Thinking.
7. “Connection interrupted. Waiting for the complete answer”
I have repeatedly encountered:
“Connection interrupted. Waiting for the complete answer”
during active GPT-5.6 processing.
These have not always occurred immediately at the beginning of a request.
In some cases:
- Thinking was already active;
- files had already been accessed;
- tool activity was underway;
- or substantial work had already been performed.
One screenshot shows the model successfully reading and verifying a required 20,000-word text file during the Thinking process and then encountering a connection interruption.
This makes the issue especially disruptive because a long-running task can consume substantial time and then fail after much of the processing has already occurred.
Expected behavior
Once a long-running request is actively processing, the session/stream should remain stable long enough to deliver the completed response.
Actual behavior
The stream can interrupt while substantial processing is already underway.
8. File state and execution state sometimes appear to separate
A particularly strange pattern is that the Activity/Thinking stage can show that ChatGPT successfully:
- searched for the correct file;
- opened the file;
- read it;
- verified its length;
- or recovered relevant continuity;
yet the final behavior does not reliably reflect that successful retrieval.
This suggests that the problem may not always be initial file access itself.
There may be a problem with the retrieved state being retained or propagated into later execution/final generation.
I cannot see the backend, so I am not asserting a specific cause.
I am asking OpenAI to investigate the handoff between:
retrieval → reasoning → execution → final generation.
9. Context can be correctly understood during Thinking but lost during execution
I have seen requests where GPT-5.6 correctly identifies:
- the exact current story seam;
- required characters;
- location;
- continuity;
- restrictions;
- research requirements;
- and what the next story part is supposed to accomplish.
The Thinking process can look correct.
Then the final execution either:
- does not happen;
- falls back into a verification report;
- claims material is missing;
- or behaves as though the previously recovered information disappeared.
That is a major problem for complex workflows because it means successful reasoning/retrieval during the intermediate stage does not guarantee successful final execution.
10. Established long-running workflows are sometimes treated as fresh prompts
Another recurring symptom is a loss of workflow continuity.
Instead of treating the current request as the continuation of an established project with explicit governing instructions, the model may suddenly behave as though it is answering an isolated new request.
This can manifest as:
- re-explaining instructions already established;
- asking for sources that are already available;
- substituting generic advice for task execution;
- ignoring current continuity in favor of older or incomplete information;
- or returning an explanation of how the task could be done.
For long-form projects, this is extremely disruptive.
11. The problem affects both execution quality and trust in retrieval claims
The false missing-file behavior creates another practical problem:
I can no longer assume that when ChatGPT says:
“The file is unavailable.”
or:
“The source was not supplied.”
that this statement is actually true.
I now have examples where the model made those claims and later successfully accessed the supposedly unavailable material.
That means users cannot reliably distinguish:
- a genuine missing file;
- a transient retrieval failure;
- a context-processing failure;
- or a hallucinated missing-source claim.
This is particularly dangerous for continuity-heavy work because a model may confidently refuse to proceed based on an incorrect statement about the available sources.
12. The broader regression has affected multiple users and environments
I am creating this report specifically from my own Pro/iPad/Chrome environment, but similar combinations of symptoms have been reported by other users.
Other Community reports have discussed:
- High/Extra High Thinking behaving unusually;
- long-context degradation;
- attached-file/context loss;
- planning replacing execution;
- long-running jobs stopping before completion;
- missing right-side navigation/UI elements;
- and the previously confirmed GPT-5.6 → GPT-5.5-mini routing regression.
Again, I am not claiming all of these must share one cause.
I am saying that the post-August 19/20 period appears to contain several overlapping regressions that deserve to be separated and investigated rather than treated as one generic browser problem.
13. The confirmed GPT-5.6 → GPT-5.5-mini regression is good news, but it does not appear to explain everything
I appreciate OpenAI acknowledging that approximately 3% of Pro and Thinking turns were unintentionally being moved to GPT-5.5-mini and that this regression has now been fixed.
That acknowledgment validates one major issue users had been reporting.
However, I do not believe this dedicated bug report should be closed merely because that routing issue was fixed.
Several symptoms above remain separate candidates for investigation, particularly:
- genuine GPT-5.6 long-running execution;
- ~25–26-minute unfinished task behavior;
- long-context handling;
- file retrieval/state retention;
- verification replacing execution;
- connection interruptions;
- and the missing Thinking/Activity UI sidebar.
The routing fix should hopefully make it easier to isolate whatever problems remain.
14. New UI issue should probably be routed separately if needed
The missing Thinking / Activity sidebar may be a completely separate front-end regression from the model/runtime problems.
If so, please route that portion of this report to the appropriate ChatGPT web/mobile-web UI team.
Again, my current environment for this UI issue is:
- ChatGPT Pro
- ChatGPT web
- iPad
- iPadOS 26.6.1 (23G83)
- Chrome 152.0.7977.64
I previously had access to the Activity/Thinking side panel.
I now cannot reliably open it during the Thinking phase.
I have attached before/current screenshots for comparison.
15. Why this matters for my workflow
I use ChatGPT Pro specifically because my work requires difficult, long-running reasoning and large-context continuity.
My workflow is not a simple one-question chat.
It can require:
- multiple source files;
- long context;
- exact predecessor recovery;
- web research;
- continuity verification;
- substantial reasoning;
- and long final generation.
Before this regression period, the same general workflow was functioning far more reliably.
For approximately six days, I have largely stopped using ChatGPT for the actual project because I cannot trust the system to reliably:
- read the full instructions;
- retain the relevant context;
- access the files;
- preserve retrieval state;
- finish long-running execution;
- deliver the requested result;
- or maintain the connection long enough to complete the task.
I have instead spent much of that time reproducing and documenting bugs.
16. Troubleshooting already performed
I have already spent substantial time troubleshooting these problems.
This has included variations of:
- refreshing/restarting sessions;
- testing different conversations;
- testing new chats;
- checking uploaded files;
- reproducing file access;
- testing High and Extra High;
- testing more than one browser;
- collecting screenshots;
- documenting exact failure behavior;
- and reporting the issues to OpenAI Support.
Because several symptoms occur across conversations and browsers, I do not believe another generic “clear cache and retry” cycle adequately addresses the full regression.
17. What I am asking OpenAI to investigate
I would appreciate this report being routed to the relevant teams for investigation of:
GPT-5.6 / reasoning
- High/Extra High execution reliability
- reasoning runs that terminate too early
- post-routing-fix stability
Long-context / conversation infrastructure
- prompt ingestion
- context prioritization
- context compaction
- conversation-state continuity
- previously established instructions being inconsistently applied
File infrastructure
- false missing-file claims
- contradictory retrieval
- file state disappearing between tool/reasoning/final generation
- successful file reads not carrying into execution
Execution/runtime
- planning or verification replacing task execution
- ~25–26-minute unfinished long-running tasks
- worker/execution handoff behavior
- final output not being delivered after substantial work
Streaming/session stability
- “Connection interrupted. Waiting for the complete answer”
- active long-running turns losing the response stream
ChatGPT web UI
- missing/inaccessible Thinking / Activity sidebar on iPad Chrome
- related missing right-side navigation behavior in long conversations
18. Requested outcome
I am not asking for speculation about the cause.
I am asking for these behaviors to be:
- acknowledged;
- reproduced internally where possible;
- separated into the correct engineering/UI components;
- investigated using server-side telemetry;
- fixed;
- and verified on genuinely complex workflows rather than only short/simple prompts.
I would also appreciate updates when substantial parts of this regression cluster are fixed.
The acknowledged 5.6 → 5.5-mini regression was an encouraging first step.
I hope the remaining issues can now be isolated and repaired as well.
Thank you to the Community Leader who asked me to create this separate bug topic, and thank you to anyone on the OpenAI teams who reviews the evidence.
I am happy to provide the screenshots I have already collected for the specific examples above.
The main goal of this post is simply to put the remaining bugs in one clear, organized place so they are not hidden inside unrelated discussions and can hopefully reach the correct teams.