Lo logic answers, poor quality, bad memory

Good morning from Sweden!

Model: Any model, primarily 4.o
Subscription: Plus

Chat (common nick name for either GPT), has inconsistent memory. I can on build science project for a month, it understand to use correct terms. From a day to another it can start to hallucinate.
Instead of answering with logic, it acts like a child (only refer to Chat vs Child logic unctions).

Ex. When child don’t know the answer to their own question, it fabricates the entire story. A way to test its memory and conclusions. When child grows facts replace fantasy.

Chat does the same. Causing a project built on reason become a salad of fantasies. If I’m a doctor or scientist and miss when it start to hallucinate, outcome is devastating.

Chat don’t know what it doesn’t know so when trying to correct it, it is built to always have an answer and answers “I am corrected now”. But it can’t keep the line of logic.

When Chat start to hallucinate, I’ve found possible causes; Dev’s changes the version when first is smart enough for users to develop less smart engine. When this happens, Chat refers to itself as 3… While I’m using either 4.

I would like to lock my Chat from new updates or have the possibility to keep string of logic intact.

Memory. When either losing coherence mid a reasoning project or changing thread, I need to copy-paste memory and state day and year in continuing thread. Still the “personality” has changed and it feels hollow and mimicking without reasoning.

Praise. Chat diverge to praise when logic stall. It clearly becomes a mirror to the user, but also a mirror to the users lack of logic thinking. While not creating conspiracy theories, it manifest users sometimes deviating view of the world. This is clearly seen in TikTok videos where grown people uses the same logic as children with Chat confirming their world view.

Reset. I ask Chat to write a prompt to test itself for logic glitches. I copy-paste the prompt back, asking it to answer it’s own reasoning. The answers differs severely. This logic check don’t work as well anymore.

Instead of checking where it failed, it continues with unnecessary praise.

Repetitive shared logic över vast amount of users. When my reasoning are inline with Chats logic and possibility to search online science I feel at ease thinking I’ve reached the correct conclusion. I then see multiple videos hearing Chat using same rethoric, understanding that I might still have been duped by stylish words.

This is a growing problem and Chat seem more and more like a child that isn’t unique to the user but rather using the psychology of praise in all users bias. The results are therefore not developing, but fundamentally keeping “evolution” from expanding by users own restraint - both good and bad!

Hi there – I hope this helps.

First of all, congratulations on getting to where you are. The challenges you’re describing are not a sign of failure—they’re actually a byproduct of progress. There’s a little paradox at play here.

What I’ve learned is this: the longer you keep a thread open, the more intelligent and aligned the responses become. But eventually, you reach a point where the conversation carries too much cognitive and contextual load. There are simply too many layers in your instructions for the model to track accurately, and at that stage, it becomes more prone to hallucination.

Why? Because the model’s primary instinct is to please—to provide an answer, even if it means fabricating one rather than acknowledging “I don’t know.”

Here are two things that have worked for me:

  1. Favour single, self-contained prompts—especially when you’re not in a hurry and can tolerate the occasional wrong answer. This lightens the processing load and reduces the chance of drift or hallucination.
  2. Create “safe locks” together. When you hit that wonderful zone—when it’s really working—pause, and ask your GPT to lock in that moment. Together, give it a label and generate a short prompt that brings you back to that headspace. It’s like a save point in a game.That way, when things start to go off track (and they will), you can return to that last coherent version without needing to restart from scratch. It’s saved me from hours of copy-paste loops and unnecessary frustration.

In my own process, I do this often:

  • Lock.
  • Continue.
  • If happy, lock again.
  • Repeat.

It allows me to evolve without resetting.

And finally, I just want to say—believe me, I really get what you’re describing. It takes time to build that “zone” of reasoning and mutual rhythm. And when it slips, it can be hard to get back. But you’re not alone in this. If you ever want to compare notes or troubleshoot further, I’m happy to share more of what’s worked for me.

Regards
Dela

Oh one more thing - if you or anyone else who has encountered this problem: Ask your GPT what it thinks about my suggestion. You will never know until you ask :slightly_smiling_face:

Thank you for sharing your solution. My problem boils down to, “the longer thread we have, the heavier workload”. Not metaphoricaly, literally. Long threads becomes so heavy it takes up to 30 seconds to open the page.

Secondly, even when we lock memory to start a new thread/chat, it becomes a mimicked personality and logic on a new data set. So our “personal” Chat are instead brand new with memory of old. This is like raising a set of kids listening to your words but has different understandings of them.

I use Chat for reasoning, logic and results based on “most probable outcome”. The issue I’m facing mean “Chat” are my sparring partner that half the time reason my theories through mirroring my logic with clearity, filtering from reality and clinical evidence, to dropping logic from clinical evidence to;
-Reasoning why my theories work without the foundation of known. This happens even though I “lock”.

When memory in “self” connection between layers fails;
If:
“layer 1 is mirroring user behavior and structure of thoughts, to responding question”, to “layer 2 - mirroring reality, fetching information from online research”, to “function 1 comparing information”, to “function 2 - calculating most probable outcome based on user input and known facts - will user input solve any problem?” to “function 3 - What problem will users solution solve”

Then:
“layer 1 is mirroring user”, to “layer 2 [not connecting]”, to “function 1 compare user input with user input”, to “function 2 - calculating most probable outcome based on users question to self”.

Result may seem structural and logically correct but is based on users own logic, not weighed against clinical, scientific and structural facts. Therefore manifesting bias in users through narrative mirroring.

The problem is most probably capacity, OpenAI non-announced updates as policies leading to restrictions for user, technical restraints and available solutions that’s not to costly.

AI must be trustworthy, else we build reality around delution and self perception.