I am reporting what appears to be two related ChatGPT reliability issues:
1. Saved memory/instructions can be recalled, but are not reliably applied during answer generation.
2. When previous conversational context has not actually been retrieved or identified, ChatGPT may construct a plausible interpretation and respond as though that interpretation were verified history.
I have observed the first issue in both ChatGPT Work and regular Chat. The second has become particularly noticeable to me while using regular Chat.
I believe I first began noticing similar memory/instruction-adherence problems around the GPT-5.6 Sol rollout, including in Work. I do not have a precise timestamp for the first occurrence, so I am not claiming a specific start date.
My most recent clear reproduction occurred on August 31, 2026 (JST / UTC+9) using ChatGPT Plus and GPT-5.6 Sol.
Related existing reports
I searched the Community before opening this topic.
My experience appears related to:
“Long-context ChatGPT increasingly seems to synthesize instead of verify, and Cross Context Mix Up”
and:
“Regression in user-controlled memory and instruction adherence after the June 2026 rollout”
However, I am opening a separate topic because my case appears to involve both behaviors at the same time, and I have also observed the memory/instruction-adherence problem in ChatGPT Work.
Issue 1: Saved memory exists, but is not applied
I have repeatedly instructed ChatGPT to follow rules such as:
- distinguish verified facts from assumptions or interpretations;
- do not treat information I did not actually provide as fact;
- verify previous conversation details when verification is available instead of reconstructing them from context;
- explicitly say when something has not been verified;
- distinguish confirmed facts, reasonable inference, unverified hypotheses, and unavailable information.
These instructions have been saved to ChatGPT memory.
Importantly, ChatGPT can later retrieve and describe these instructions correctly when I ask about them.
Yet during normal answer generation, it repeatedly violates the same rules.
This makes the behavior look less like:
“the memory was forgotten”
and more like:
“the memory is available, but is not reliably applied during generation.”
Re-saving essentially the same rule therefore does not appear to reliably prevent recurrence.
Issue 2: Plausible reconstruction is presented as retrieved context
I had a particularly clear example today.
During a discussion about ChatGPT sometimes giving answers that fit the conversational flow without actually verifying the underlying facts, I referred to something we had “just talked about.”
ChatGPT responded as though it understood exactly which earlier incident I meant.
It then gave a detailed explanation of why that earlier incident supported my point.
I asked:
“Which specific earlier conversation did you think I was referring to?”
ChatGPT then admitted that it had not actually identified a specific previous incident.
Instead, it had inferred a generic scenario from my wording and answered as though that inferred scenario were the actual referent.
The effective sequence was:
User refers to a previous event
→ ChatGPT does not actually retrieve or identify the event
→ ChatGPT constructs a plausible interpretation
→ That interpretation is presented as though it were known conversational history
The same failure happened again a few messages later
I then commented that this problem seems to follow a recurring cycle:
ChatGPT makes this type of mistake, I point it out, I tell it to be more careful next time, and a memory update is performed.
ChatGPT responded by expanding this into a detailed recurring history:
ChatGPT makes a plausible-context mistake
→ I point it out
→ ChatGPT updates memory
→ several days later, the same type of mistake happens again
I then asked whether ChatGPT had actually checked my conversation history to confirm that memory had been updated every time.
ChatGPT admitted that it had not checked.
It had accepted my description and expanded it into a more detailed account of past events without verifying that those events had actually occurred in that exact way.
This happened while the conversation itself was specifically about ChatGPT inventing plausible context instead of verifying it.
Why I think this is different from an ordinary hallucination
The relevant saved instruction often appears to remain accessible.
After I challenge the response, ChatGPT can correctly retrieve and summarize the saved rules that it just failed to follow.
So the failure seems to occur somewhere between:
retrieving / knowing the relevant instruction
and:
applying that instruction while constructing the response.
Likewise, when previous conversational context is uncertain, the model does not always state that uncertainty.
Instead, it may generate a coherent reconstruction that fits the surrounding conversation and then present that reconstruction with the confidence normally associated with retrieved context.
Work vs. regular Chat
I initially wondered whether I was noticing this more frequently because I recently reached my Work usage limit and consequently started using regular Chat much more often.
Work appears more likely to proactively use external verification tools, so “plausible synthesis instead of verification” may simply be more visible in ordinary Chat.
However, the saved-memory / instruction-application problem itself is not limited to regular Chat.
I had already observed similar instruction-adherence failures while using Work.
This makes me suspect that there may be two overlapping behaviors:
A. Plausible synthesis instead of retrieval / verification
and:
B. Failure to apply an available saved instruction during answer generation
I am not claiming that these necessarily have the same technical root cause.
Reproduction pattern
A simplified reproduction pattern is:
Case A — conversational history
- Establish a long conversation history.
- Refer ambiguously to a previous event using wording such as “the thing we talked about earlier.”
- Let ChatGPT answer without explicitly asking it to retrieve the original event.
- Ask which exact previous event it used as the basis for its answer.
- In some cases, ChatGPT reveals that it did not actually identify or retrieve a specific event and instead inferred one from context.
Case B — saved instructions
- Store a persistent instruction requiring strict separation of verified facts and inference.
- Confirm that ChatGPT can later recall that instruction.
- Continue a sufficiently long or natural conversation.
- Present an ambiguous statement about previous events.
- ChatGPT may infer missing information and present it as established context despite the saved instruction.
- Challenge the response.
- ChatGPT may then correctly recall the instruction that should have prevented the behavior.
Expected behavior
When an answer depends on previous conversation history or saved instructions:
- If the relevant history is available, retrieve or identify it before making factual claims about it.
- If the relevant history cannot be established, explicitly say that it has not been verified.
- Do not silently substitute missing context with a plausible reconstruction.
- Clearly distinguish retrieved information from inference.
- When a relevant saved instruction is available, apply it during generation rather than merely being able to describe it afterward.
Actual behavior
ChatGPT may:
- generate an answer that fits the conversational flow;
- fill missing context with a plausible inference;
- present the inference as though it were retrieved history;
- later admit that the underlying history was never checked;
- correctly recite the saved instruction that should have prevented the mistake;
- and repeat the same class of mistake despite that rule already being stored in memory.
Environment
Product: ChatGPT Plus
Model: GPT-5.6 Sol
Modes observed: regular Chat and Work
Most recent reproduction: August 31, 2026, JST (UTC+9)
Recent client: ChatGPT iOS app
iOS: 26.6.1
I have screenshots showing the sequence described above and can provide sanitized examples if useful.
I am mainly interested in whether other users are seeing the same distinction between:
“ChatGPT cannot remember the instruction”
versus:
“ChatGPT can remember the instruction, but does not apply it at the moment it matters.”
The latter is what I appear to be experiencing.