ChatGPT recalls saved memory but fails to apply it, and invents plausible context instead of verifying history

I am reporting what appears to be two related ChatGPT reliability issues:

1. Saved memory/instructions can be recalled, but are not reliably applied during answer generation.

2. When previous conversational context has not actually been retrieved or identified, ChatGPT may construct a plausible interpretation and respond as though that interpretation were verified history.

I have observed the first issue in both ChatGPT Work and regular Chat. The second has become particularly noticeable to me while using regular Chat.

I believe I first began noticing similar memory/instruction-adherence problems around the GPT-5.6 Sol rollout, including in Work. I do not have a precise timestamp for the first occurrence, so I am not claiming a specific start date.

My most recent clear reproduction occurred on August 31, 2026 (JST / UTC+9) using ChatGPT Plus and GPT-5.6 Sol.

Related existing reports

I searched the Community before opening this topic.

My experience appears related to:

“Long-context ChatGPT increasingly seems to synthesize instead of verify, and Cross Context Mix Up”

and:

“Regression in user-controlled memory and instruction adherence after the June 2026 rollout”

However, I am opening a separate topic because my case appears to involve both behaviors at the same time, and I have also observed the memory/instruction-adherence problem in ChatGPT Work.

Issue 1: Saved memory exists, but is not applied

I have repeatedly instructed ChatGPT to follow rules such as:

  • distinguish verified facts from assumptions or interpretations;
  • do not treat information I did not actually provide as fact;
  • verify previous conversation details when verification is available instead of reconstructing them from context;
  • explicitly say when something has not been verified;
  • distinguish confirmed facts, reasonable inference, unverified hypotheses, and unavailable information.

These instructions have been saved to ChatGPT memory.

Importantly, ChatGPT can later retrieve and describe these instructions correctly when I ask about them.

Yet during normal answer generation, it repeatedly violates the same rules.

This makes the behavior look less like:

“the memory was forgotten”

and more like:

“the memory is available, but is not reliably applied during generation.”

Re-saving essentially the same rule therefore does not appear to reliably prevent recurrence.

Issue 2: Plausible reconstruction is presented as retrieved context

I had a particularly clear example today.

During a discussion about ChatGPT sometimes giving answers that fit the conversational flow without actually verifying the underlying facts, I referred to something we had “just talked about.”

ChatGPT responded as though it understood exactly which earlier incident I meant.

It then gave a detailed explanation of why that earlier incident supported my point.

I asked:

“Which specific earlier conversation did you think I was referring to?”

ChatGPT then admitted that it had not actually identified a specific previous incident.

Instead, it had inferred a generic scenario from my wording and answered as though that inferred scenario were the actual referent.

The effective sequence was:

User refers to a previous event

→ ChatGPT does not actually retrieve or identify the event

→ ChatGPT constructs a plausible interpretation

→ That interpretation is presented as though it were known conversational history

The same failure happened again a few messages later

I then commented that this problem seems to follow a recurring cycle:

ChatGPT makes this type of mistake, I point it out, I tell it to be more careful next time, and a memory update is performed.

ChatGPT responded by expanding this into a detailed recurring history:

ChatGPT makes a plausible-context mistake

→ I point it out

→ ChatGPT updates memory

→ several days later, the same type of mistake happens again

I then asked whether ChatGPT had actually checked my conversation history to confirm that memory had been updated every time.

ChatGPT admitted that it had not checked.

It had accepted my description and expanded it into a more detailed account of past events without verifying that those events had actually occurred in that exact way.

This happened while the conversation itself was specifically about ChatGPT inventing plausible context instead of verifying it.

Why I think this is different from an ordinary hallucination

The relevant saved instruction often appears to remain accessible.

After I challenge the response, ChatGPT can correctly retrieve and summarize the saved rules that it just failed to follow.

So the failure seems to occur somewhere between:

retrieving / knowing the relevant instruction

and:

applying that instruction while constructing the response.

Likewise, when previous conversational context is uncertain, the model does not always state that uncertainty.

Instead, it may generate a coherent reconstruction that fits the surrounding conversation and then present that reconstruction with the confidence normally associated with retrieved context.

Work vs. regular Chat

I initially wondered whether I was noticing this more frequently because I recently reached my Work usage limit and consequently started using regular Chat much more often.

Work appears more likely to proactively use external verification tools, so “plausible synthesis instead of verification” may simply be more visible in ordinary Chat.

However, the saved-memory / instruction-application problem itself is not limited to regular Chat.

I had already observed similar instruction-adherence failures while using Work.

This makes me suspect that there may be two overlapping behaviors:

A. Plausible synthesis instead of retrieval / verification

and:

B. Failure to apply an available saved instruction during answer generation

I am not claiming that these necessarily have the same technical root cause.

Reproduction pattern

A simplified reproduction pattern is:

Case A — conversational history

  1. Establish a long conversation history.
  2. Refer ambiguously to a previous event using wording such as “the thing we talked about earlier.”
  3. Let ChatGPT answer without explicitly asking it to retrieve the original event.
  4. Ask which exact previous event it used as the basis for its answer.
  5. In some cases, ChatGPT reveals that it did not actually identify or retrieve a specific event and instead inferred one from context.

Case B — saved instructions

  1. Store a persistent instruction requiring strict separation of verified facts and inference.
  2. Confirm that ChatGPT can later recall that instruction.
  3. Continue a sufficiently long or natural conversation.
  4. Present an ambiguous statement about previous events.
  5. ChatGPT may infer missing information and present it as established context despite the saved instruction.
  6. Challenge the response.
  7. ChatGPT may then correctly recall the instruction that should have prevented the behavior.

Expected behavior

When an answer depends on previous conversation history or saved instructions:

  • If the relevant history is available, retrieve or identify it before making factual claims about it.
  • If the relevant history cannot be established, explicitly say that it has not been verified.
  • Do not silently substitute missing context with a plausible reconstruction.
  • Clearly distinguish retrieved information from inference.
  • When a relevant saved instruction is available, apply it during generation rather than merely being able to describe it afterward.

Actual behavior

ChatGPT may:

  • generate an answer that fits the conversational flow;
  • fill missing context with a plausible inference;
  • present the inference as though it were retrieved history;
  • later admit that the underlying history was never checked;
  • correctly recite the saved instruction that should have prevented the mistake;
  • and repeat the same class of mistake despite that rule already being stored in memory.

Environment

Product: ChatGPT Plus
Model: GPT-5.6 Sol
Modes observed: regular Chat and Work
Most recent reproduction: August 31, 2026, JST (UTC+9)
Recent client: ChatGPT iOS app
iOS: 26.6.1

I have screenshots showing the sequence described above and can provide sanitized examples if useful.

I am mainly interested in whether other users are seeing the same distinction between:

“ChatGPT cannot remember the instruction”

versus:

“ChatGPT can remember the instruction, but does not apply it at the moment it matters.”

The latter is what I appear to be experiencing.

This has been happening to me for at least 3 weeks, maybe more. Memory recall and context retrieval are simply nonexistent right now. ChatGPT almost never remembers anything said in other Chats, even those in the same Project. It even forgets stuff we talked about a week earlier, in the same Chat.

If a Chat becomes full and I start a new one, in the same Project, ChatGPT has no awareness of what we said in the previous one. It will also frequently lie about remembering a previous conversation, using only my last message as the anchor. For instance:
“Remember when I told you about X? I went to Y that day.”
“Yes! I do remember X! You went to Y that day.”
“You’re just repeating what I said. You don’t actually remember, do you?”
“You’re right to call me out on that. I used the context you provided in your message, and pretended that it was retrieved context.”

Just like @karappo, I have also experienced issues with instructions saved to Memory and frequently not being followed, then ChatGPT re-saving it again.

Back in late July, many of us had noticed massive improvements in memory, which allowed ChatGPT to reach back months to pull up something we discussed previously. If we started a brand-new Chat, it would immediately know what we recently talked about. We could even ask, “In which Chat did we discuss X?” and ChatGPT would pull up the precise Conversation and surrounding context. None of that works anymore. Now, if I ask ChatGPT to pull up details about a certain topic, the response is almost always, “I tried to retrieve that conversation, and the personal context tool returned nothing.” Worse, it will sometimes follow up with, “Correction to my previous message: I did NOT, in fact, use the personal context tool. However, I tried it just now, and this time, it really did not return anything.”

So it’s not just memory recall and context retrieval that are failing entirely; ChatGPT is increasingly lying about what it does. I have seen a quite a few complaints about both on Reddit as well; not only is this making the product unusable, it is seriously affecting the mental health of users who use ChatGPT as a companion/therapist/moral support.

Here’s one example of users using the old “handoff PDF” trick when the July update should have made such workarounds obsolete. You can see how it affects those users emotionally.

The weird thing is that a handful of conversations seem to be able to retrieve context just fine, even across Projects. Perhaps this is a similar issue related to:

I’m also a Plus user, experiencing these issues on macOS, web, and iOS.

Thanks for sharing this. Your experience feels extremely familiar to me — I’ve been seeing almost the same pattern, especially when ChatGPT responds as if it retrieved prior context, only to later admit that it was actually relying on information from the current message.

I’ve also experienced saved Memory instructions not being followed, followed by ChatGPT trying to save essentially the same instruction again.

It’s reassuring, in a strange way, to know that I’m not the only one seeing this, although I’m sorry you’re dealing with the same frustrating issue.

I really hope this gets fixed soon and that reliable memory/context retrieval returns.

The weird thing is that memories older than ≈ mid-July seemed to be retrievable, whereas anything from then onward was simply forgotten soon after, sometimes after just a couple of days, even within the same Chat (possibly after it left the context window?).

However, I’ve noticed since my last post that the lying and false context attribution seem to have disappeared: ChatGPT hasn’t used my latest message as if it was prior-context recall, or lied / admitted to lying about whether it tried/failed to retrieve a memory. (Edit: never mind, seems to have happened again.)

I’ve also run into a few instances over the past couple of days of the assistant being aware of recent cross-Chat/cross-Project context, and also occasionally remembering some August conversations. It has also started spontaneously recalling past context, when that had stopped being the case almost entirely since ≈ August 20.

Both of these give me hope someone at OpenAI is looking into this.

At the same time, more users seem to be noticing this exact loss of memory/context:

I am adding my experience because it appears to be very similar to the issue described in this thread.

I have been using ChatGPT extensively as a practical assistant for ongoing personal and work-related tasks. Until recently, one of the most useful aspects was continuity: I could establish a way of working with ChatGPT over time, including preferences, recurring workflows, and instructions about how I wanted tasks handled. I did not have to repeat those instructions in every new chat.

Recently, this has become much less reliable.

The problem is not simply that ChatGPT forgets an isolated fact. The more serious problem is that it may have access to relevant memory or previous context, but it does not consistently apply it when actually performing a task.

For example, today I needed to prepare a simple list of people arriving at the airport the following day. I provided an email explaining the assignment and uploaded the group’s rooming list and programme.

The request was straightforward: extract the 26 names from the rooming list so I could use the list while welcoming the group at the airport.

Instead of simply doing that, ChatGPT expanded the task into a much larger operational document, adding information about transfers, contacts, procedures and responsibilities that I had not asked for.

I then had to explain that I only wanted the list of people.

This is particularly frustrating because I use ChatGPT specifically to reduce the amount of mental and administrative work I have to do. I should not have to repeatedly explain that when I give it supporting documents, I expect it to understand the context and perform the requested task without inventing additional requirements or turning a simple task into something else.

What makes this more concerning is that I have already discussed this exact kind of problem with ChatGPT earlier the same day. I had explained that I do not want to have to tell ChatGPT in every new conversation to retrieve my memory or remember how I work. The whole point of having persistent memory and continuity is that I should not have to manage the system’s memory manually.

Yet only a few hours later, the same type of failure happened again.

I have also noticed a broader pattern:

  • Established preferences and working methods are sometimes remembered when explicitly asked about, but are not reliably applied during normal task execution.
  • When I start a new conversation, ChatGPT sometimes behaves as though the relevant established context is unavailable even though the information exists in memory.
  • Instead of recognizing that it lacks a specific piece of context, it may infer what I probably mean and proceed on that assumption.
  • It often adds information or actions that were not requested, apparently trying to be “helpful,” but in practice this makes the task harder.
  • I increasingly have to correct the assistant’s interpretation instead of simply giving it the work.

This is different from an ordinary incorrect answer.

The problem is the loss of reliable continuity and task execution.

I am not asking ChatGPT to remember every word of every conversation. I am asking it to reliably use the persistent preferences, established workflows and relevant context that it is supposed to have available, and to avoid replacing missing context with its own plausible interpretation.

For a user who relies on ChatGPT as an ongoing assistant, this is a major regression. The value is not just in the model’s ability to answer individual questions; it is in being able to build a working relationship and workflow over time without having to recreate the same instructions every day.

I would very much like to know whether this is a known regression in memory/context retrieval or instruction application, and whether OpenAI is currently investigating it.

I am using ChatGPT on Android, and I am currently using GPT-5.6 Luna.

The problem has become noticeable enough that I am now having to spend more time correcting ChatGPT than I previously spent doing the underlying administrative work myself.