Apologies for the length, but these issues are interconnected. I’ve tried to keep this concise without stripping out the evidence needed to show the pattern.
I’ve been using ChatGPT heavily for long-running professional and personal workflows for several years. For me, the deterioration became noticeable after o3 was retired.
This is not nostalgia for o3, and it is not about preferring the personality or writing style of an older model.
GPT-5.6 can be extremely intelligent turn by turn. When I first started using it, I was genuinely impressed.
The problem is what happens over time.
Local intelligence is high. Longitudinal intelligence is unreliable.
In long-running threads, the model increasingly forgets what the thread is actually for. It loses decisions that were already settled, reintroduces things that were explicitly rejected, forgets corrections that have already been made, conflates separate workflows, and starts treating the newest prompt as though it were the entire task.
Worst of all, it can confidently answer from an incorrect reconstruction of the project state.
And the most exhausting part is that corrections do not reliably stick.
You correct it. It acknowledges the correction. It works properly for a while. Then the same error returns and you have to establish the state all over again.
That is not simply forgetting some obscure detail from hundreds of messages ago.
It is losing the identity and accumulated logic of the task.
At times it genuinely feels like goldfish-memory GPT.
The transcript exists, but the working state apparently does not
The strange part is that I can still scroll upward and see the information.
The conversation is visibly there.
Yet ChatGPT can behave as though major parts of that history are no longer functionally influencing the response.
This looks like a long-thread context-compaction failure.
The full transcript can remain visible in the UI while the model appears to be operating from a compressed representation of older turns. If important instructions, corrections or decisions are dropped during that compression, you get context drift: the model reverts to earlier behaviour, forgets established constraints, resurrects rejected ideas, or starts treating the latest prompt as though it were the whole task.
Obviously I cannot see the internal implementation of ChatGPT Web, so I am not claiming to know precisely what mechanism is being used internally.
But the user-facing failure is clear:
something that was previously part of the accumulated state of the conversation stops influencing the model reliably.
And once that happens, a response can be intelligent locally while being completely wrong globally.
This is not one broken thread
I have experienced the same class of failure across completely unrelated work: long-form writing, project development, technical workflows, financial analysis psychological work and long-running personal conversations.
That matters.
If one thread became confused, fine. Long conversations are complicated.
But when the same pattern repeatedly appears across unrelated projects, the issue stops looking like one badly managed conversation and starts looking systemic.
The recurring pattern is:
the model remains intelligent about what is immediately in front of it, but becomes unreliable about everything that gives that message its historical meaning.
That is what I mean by longitudinal intelligence.
Corrections that do not persist are worse than ordinary mistakes
Models make mistakes. That is expected.
The problem here is different.
A correction should become part of the state of the conversation.
Instead, the cycle increasingly becomes:
the model makes an error, I correct it, it acknowledges the correction, it behaves correctly temporarily, and then several turns later it repeats essentially the same error as though the correction never happened.
That forces the user to do two jobs.
First, catch the mistake.
Then repeatedly rebuild the context that was already established.
Over months of work, that becomes exhausting.
And in serious workflows it can become dangerous.
The psychological and health implications are not trivial
Many colleagues and psychologists absolutely use AI to help them organise complicated histories, navigate diagnoses, prepare for appointments, identify patterns, keep track of previous conclusions, and move through years of information more efficiently.
Imagine someone who has been using a long-running conversation to help navigate a complicated psychological history.
Over time that conversation may contain symptoms, diagnoses, clinician conclusions, major life events, things that were investigated and ruled out, previous interpretations that were corrected, medications, changes in symptoms, and years of context.
If the model suddenly loses part of that accumulated history, mixes chronology, resurrects an interpretation that had already been rejected, or starts treating the last few messages as though they represent the entire history, the response may still sound perfectly intelligent.
But it is reasoning from an inaccurate reconstruction of that person’s history.
That makes it unreliable for the very thing a longitudinal assistant should be good at.
And the most worrying part is that it does not necessarily say:
“I have lost part of your earlier context.”
It simply answers!!
So the user may have no idea that the foundation underneath the answer has shifted.
Financially, the consequences can be very real
This has also affected financial workflows for me.
Failures around established state, previous decisions, strategy and the actual purpose of long-running analytical threads have had an effective financial impact on me of more than $22,000.
I made the decisions. I am responsible for them.
That is not what I am disputing.
The point is that if you have spent months building an analytical workflow with specific objectives, constraints, indicators, previous decisions and corrections, and the model silently forgets what that workflow is actually trying to achieve, the resulting analysis can be sophisticated and completely inappropriate at the same time.
That is not simply “AI made a mistake.”
It is a state failure.
For high-consequence longitudinal work, that distinction matters enormously.
The previous memory system helped mitigate this
I also deliberately used the previous memory system.
Whenever ChatGPT said that memory had been updated, I could inspect what had actually been stored.
And I did.
Important objectives, project rules, portfolio information, relationships and operating instructions were explicitly preserved there.
So this was never simply a case of expecting ChatGPT to magically remember everything forever.
I used the persistence mechanisms the product gave me and checked what they contained.
The newer memory experience gives me much less visibility into that granular operational state, while long-running conversations themselves appear less dependable at preserving it.
That combination is a serious regression for people who use ChatGPT longitudinally.
Exporting everything and uploading it again does not fix it
I even tried the obvious recovery route.
I exported my ChatGPT data and conversations into the JSON files provided by the export system and uploaded them again.
That does not recreate the original working state.
The text may be there, but the model does not suddenly regain the accumulated relationships between decisions, corrections, rejected approaches, chronology, project rules, objectives, current state and the actual purpose of the thread.
A transcript archive is not the same thing as restoring a functioning longitudinal conversation.
This matters because “export everything and start again” sounds like a solution until you actually try it.
It is not continuity.
It simply transfers the reconstruction work onto the user.
After months or years of accumulated work, that is an enormous amount of labour.
And now custom GPTs are being replaced by plugins
This is another change that makes very little sense to me from the user side.
Many people built custom GPTs precisely because they provided stable, specialised working environments with their own instructions, files, knowledge and purpose.
I have several myself.
Some are for writing.
Some are for psychological work.
Some are for particular creative projects.
Some are finacial.
They work.
I am happy with them.
Now those environments are being migrated toward plugins.
Even if plugins eventually provide greater flexibility, replacing something that already works with something that may behave differently and has to be migrated, retested and potentially rebuilt is not automatically an upgrade simply because the underlying architecture may be cleaner.
From the customer side the question is much simpler:
Does the thing I built and rely on continue working the way it currently works?
If the answer is no, then the user has been handed another migration project.
And the timing makes this especially frustrating.
Long-running ordinary conversations are already having continuity problems.
Memory is less transparent.
Exporting and re-uploading the data does not recreate the working state.
And now another mechanism people used to create specialised, persistent environments is being replaced.
So users are being asked to absorb more migration, more retesting and more responsibility for reconstructing continuity at exactly the moment continuity itself appears to be getting worse.
From the user side, that is not simplification.
It is additional complexity.
This is clearly affecting many users
Initially I wondered whether my experience was unusual because I use extremely long-running threads.
After looking through OpenAI’s own community forum, I no longer think that is remotely plausible.
The forum is riddled with complaints describing overlapping symptoms: long-context failures, established instructions being ignored, corrections not sticking, models reverting to previously corrected behaviour, project state being lost, file and tool workflows breaking, responses becoming shallow or generic, and users having to repeat information that already exists in the conversation.
The exact technical cause may not be identical in every case.
That is important to acknowledge.
But the user-facing pattern is clearly affecting many people.
And forum complaints will inevitably only represent the visible fraction.
For every person who takes the time to document the problem, create a thread and argue their case, how many others experience the same deterioration and simply do not post?
How many assume they are doing something wrong?
How many stop using long threads?
How many quietly downgrade their usage?
How many move to another platform?
We cannot quantify that from outside.
OpenAI can.
OpenAI has the telemetry.
Users should not have to spend hours collecting anecdotes to prove there is a longitudinal reliability problem.
Please investigate it at cohort level.
Look specifically at users with genuinely long-running conversations.
Measure correction recurrence.
Measure how often users have to restate something already supplied.
Measure contradictions with established state.
Measure task-identity drift.
Measure reintroduction of rejected decisions.
Measure how often the model changes an established objective.
And measure whether these failures become more frequent as conversations grow.
Those are longitudinal quality metrics.
Traditional reasoning benchmarks do not measure this experience.
Benchmarks are not enough
This is why statements about a model being “better” because it scores higher on benchmarks are increasingly meaningless to me and countless others here.
A model can outperform another model on coding, maths, reasoning or knowledge benchmarks while simultaneously becoming substantially worse for someone whose work depends on months of accumulated context.
If I have to spend more time supervising the model, correcting it, checking whether it forgot something, rebuilding context and catching regressions, then the model has not become more capable for my use case.
It has become more expensive in time and cognitive load.
One of my threads even suggested moving to a different AI platform that “might” perform better.
Intelligence without dependable state preservation is not enough.
Please give us a Full Context / Maximum Continuity mode
If aggressive context compaction is necessary because carrying enormous histories costs more compute, fine.
Give users a choice.
Make full context slower.
Make it consume more resources.
Make it part of plus option.
Make it a plugin.
I genuinely do not care how it is implemented.
Just allow users whose work depends on long-running conversations to choose continuity over compaction.
Something as simple as:
Full Context / Maximum Continuity: slower or more expensive, maximum preservation
versus:
Compact Context: faster or cheaper, older conversation may be summarised
would at least make the trade-off visible and voluntary.
Right now the user can see an entire conversation in the interface while having no reliable way of knowing how much of that conversation is actually influencing the current response.
That is a fundamental transparency problem for long-running work.
At minimum, let us pin context that can never be compacted away
If keeping an entire conversation verbatim in active context is technically unrealistic, there is an obvious middle ground.
Allow users to select messages, instructions, decisions or sections of a conversation and mark them:
Preserve verbatim
or
Never compact this
The core objective of a project, current state, locked decisions, rejected approaches, important corrections, critical relationships and project rules should not disappear because an automated summarisation process decided they were less important.
Give the user some control over what absolutely must survive.
That alone would make a huge difference.
Please do not solve this by making users continuously restate their projects
That defeats the entire point.
If every serious conversation eventually requires me to paste a giant summary explaining what the thread is for, what has happened previously, what was rejected, what was decided, what my current state is and what the model must not forget, then maintaining a long-running thread has lost much of its purpose.
The product should carry accumulated state forward.
The user should not have to continuously reconstruct it.
And “create a new thread” is not a satisfactory answer either.
Starting a new conversation every time the existing one becomes unreliable destroys the longitudinal value that made ChatGPT useful in the first place.
This is not simply “bring o3 back because many of us liked it”
I want to make this distinction very clear.
I am not asking OpenAI to restore o3 because I preferred its personality.
I am saying that my long-running workflows were substantially more reliable before its retirement, and that the deterioration became particularly noticeable around that transition.
If GPT-5.6 can restore that reliability, excellent.
If it requires a different context architecture, better memory, pinned context, a maximum-continuity mode, another model, or some other technical solution, also excellent.
The model label is secondary.
What matters is restoring the ability to work cumulatively.
Because right now the fundamental problem is simple:
GPT-5.6 can be extraordinarily intelligent about the message directly in front of it while forgetting the accumulated structure that gives that message meaning.
Local intelligence is high.
Longitudinal intelligence is unreliable.
For people doing serious long-term work, particularly where accumulated history affects professional, financial, creative, medical or psychological decisions, that is not an upgrade.
Please treat this as a product-level regression rather than a collection of isolated conversation failures.
And please give users a way to choose continuity over compaction.
Thank you,
K