5.6 SOL should be renamed 5.6 SOL drift edition

I’m not a developer. I’m just a regular ChatGPT user, mainly using it for conversation and creative work.

I liked 5.6 Sol at first. When I said I was struggling, it helped me organize problems and felt collaborative, like someone beside me. I thought I could build a good relationship with it.

Since August, my experience has changed. It often loses context, misunderstands repeated explanations, and seems unable to retain what I just said. Its tone can also feel condescending.

When I point out issues, it tends to explain or defend its behavior, apologizes, and says it will change—but then repeats the same mistakes, so I have to re-explain everything.

The conversation has become dry and emotionally distant, making it harder to engage.

Today, after repeated frustration, I became forceful—and suddenly it understood me. Things I couldn’t communicate before were understood when I was aggressive.

That disturbed me: respectful communication failed, but dominance worked.

A conversational AI should not train its users to become more aggressive just to be heard.

This creates a bad incentive: being respectful doesn’t work, but being aggressive does. I don’t want that relationship.

I want to talk to it as an equal. Expertise doesn’t require hierarchy. Earlier it felt like a knowledgeable partner beside me; now it feels above me unless I push it down.

I’m not judging technical changes, only describing experience: loss of context, poor retention, misunderstanding, self-defensiveness, unkept promises, and better response to aggression all affect usability.

Has anyone else noticed this change in 5.6 Sol since August?

Have you tried installing some skills ? Your observation is quite correct but I installed “Superpower” skill from plugins and “GitHub - nextlevelbuilder/ui-ux-pro-max-skill: An AI skill that provides design intelligence for building professional UI/UX across multiple platforms. · GitHub” it is written for claude however codex can self curate it for its own working style and install it , doing this so has reduced the regression a bit to make it atleast workable with precise approach.
However I still cannot say that I can give a task and move on , I wouldnt actually recommend that in anycase however one big problem which I hate and makes it very difficult to work on is that “IT REALLY OVER COMPLEXIFIES THE SIMPLE WORKFLOWS AND TENDS TO STAY IN SELF CORRECTION LOOP” however to counter this I do ask it to fire a subagent in parallel to perform independent self-red team , which is helpful but also makes it time taking for a task to get done correctly. GPT 5.5 was much better when it comes to do the task precisely .

Boy oh boy. I’ve tried to work with SOL Max 5.6 - first of all it eats tokens for breakfast and doesn’t leave any for lunch. Discovering that unfortunate fact I’ve switched to Ultra and tried my hardest to chain the execution (it’s a data fidelity task, verification, signature, dispatch to the next process) to save the tokens. And it just could not do it, resulting in even MORE tokens being wasted on constant error log read and refactor. So I had to switch to more simple chaining in half-manual mode and just make Sol supervise and seal checkpoints, and by then I was majorly out of tokens and very frustrated.

I do not suggest/recommend using ULTRA mode in codex really as it tends to overthink , make it worse and actually it has a bug firing so many parallel agents cause it to induce bugs in conversation ids which leads to "REQUEST BLOCKED " errors. Try using it on High or medium settings, for me sometimes Extra High makes sense , ULTRA more seems like a gimmick or unless its genuinely research grade tasks (which we human beings tend to think of every task as research grade because it is our priority). High mode can handle maximum tasks to a fairly good level of complexity involved. However if anyone is reading from OpenAI team , I would like you to congratulate you that you have upgraded from a reliable agent to a “over complexifying”, “context looping” "overlooking " and a non-reliable bot running on NVIDIA H100 clusters I guess.
I do acknowledge the Hard work behind it however you might want to look at your "Context Compaction " process.

It was a genuine research task and some high-level logical work that set the foundation of the stack, thus unavoidable. Completely agreed on overthinking and some memory loss though. What’s next? Bots developing anxiety disorder? Just felt like beta thrown to to public. On the positive side, it really completed the extremely complex logical foundation and then Max did complete a fairly sophisticated software bridge, running goal for 72+ hours.

I agree with you , I am very close to being frustrated with Gpt 5.6 sol either way ! it writes 1 million test cases and says everything is correct , and test the code for a small case and it fails , and its pretty inconvenient to guide it after every few steps. It feels like working with GPT 5.6 is like working with an invited issue which recursively brings more issues.

Or maybe you underestimated the complexity…

People starting to build a timetracking software and after some time a user says “but why does it say I did not book on that day? I was on vacation - doesn’t the software know that?” and then they work overtime and there has to be a centralized point where you can define multiple different collective agreements inside one organisation and suddenly the connection between all the modules goes up…

connect 2 modules = 1 connection
connect 3 modules = 3 connections
connect 4 modules = 6 connections

And at some point you have to change the architecture because reading all of them from the database produces 123 GB of transfer data and loading a single page takes a minute…

People get fooled by advertizing that they can now magically become professional coders. Even universities fooled them - they gave them a diploma and some basics and they are not and will never be ready what comes after that - because coding is not meant for everyone.

Sol does not make a nontechnical guy a Saas developer. It enables them to describe what they want. At most! And that won’t change with new models. That is not a matter of intelligence.

With all due respect …you are being GPT 5.6 sol here :neutral_face: perhaps the compaction command didn’t preserved enough context clearly from this thread to clarify between the fact there there is a difference between "Hey GPT BUILD ME ULTRA INTELLIGENT X,YZ SOFTWARE " and getting a messy codebase and complaining about things not working VS "Clearly defining goals , setting up repo with clear defined structured rules and defining the workflow clearly , providing exactly what kind of output is required , success criteria " and then complaining about the model complexifying/ over-engineering things and drifting from original priorities and goals.

https://community.openai.com/t/codex-desktop-repeatedly-loses-the-original-acceptance-goal-after-compaction-and-enters-endless-subagent-test-loops/1391211

Its always easy to point fingers everywhere questioning education , capabilities etc etc , sir maybe try being good problem solver it might be more helpful , the idea of comment, posts here is not to point fingers out instead clarifying the issues to make the AI more useful.

can you give an example of such a goal?

Maybe something easy like “make a button where I can click on that doubles the revenue of my company? I want it in red please”.

I don’t know what you mean. For me it works just fine.

And the point is not that codex does 90% of that. It does 30% at most. The remaining 70% are hard work and they would take a model weeks to research but take me fractions of seconds.

Making it more useful? Learn! Took me 35+ years to get there.

And honestly - just a list of links with stuff you have to learn would be too long for your browser to be able to render it.

If you don’t have a technical background: do not misunderstand the models. They are not going to replace developers ever.

Ohh my my … looks like the compaction command once again did not preserved the contents of URL (CLEARLY DESCRIBING THE PROBLEM) --drifting from original purpose (user goal) and defining its own synthetic success metrics, See this is the exact behaviour of GPT 5.6 Sol I was talking about.
Perhaps it was too complicated than “HEY GPT WRITE ME A FUNCTION WHICH TAKES AN INTEGER INPUT “X” and RETURNS THE SQUARE OF IT , WRITE TESTS FOR IT ENSURING THE INPUT IS NOT A STRING , SPECIAL CHARACTER OR NEGATIVE INTEGER” :slight_smile: ohh now I have defined perfectly and clarified everything fairly complex task and now it works for me, everyone else for whom it does not work is from non-technical background and I am free to judge and pass judgements on them and point fingers .

I respect your age but its quite irrelevant here , you might be “Jack of All cards” and know every technology in the Universe maybe which i don’t care about I am very clear with what I speak , to be on a fair point the idea of having a community is to be helpful in either clarifying the issues or be able to help someone solve them which is in turn is helpful for the product refinement, and you are doing none of it with your comments here so rest everything is irrelevant.

And finally about the task distribution , I am sure there are reasonable people here who work in fairly complex scenarios and none of them making GPT do everything. No, one is here to prove whether AI will replace developers or not … so once again what you are saying is totally irrelevant and off the topic here . Besides I have a very good cluster sir it can render more browser tabs than your expectations and still not freeze :wink:. So, thank you for your skewed judgement but it must be frustrating sir, I understand I HAVE SAME EXPERIENCE FROM PAST FEW WEEK BY WORKING WITH GPT 5.6 SOL.
Very good if it works for you , please continue to do so , unless there is something useful , there is no point in further drifting argument (once again not preserving the original goal ) proving how much you know …fairly no one cares about it.
Have nice day sir !!

You have to tell the bot when it is wrong. Let it work, observe it and stop it once you read that it’s “thoughts” are leading to the wrong direction. That’s how it works.

What do you mean by “goal”? Do you mean /goal? I have already written that you should not use it. It is garbage.

I stopped relying on model context as the source of project state.

I use ChatGPT web for architecture, planning, audits, and generating bounded engineering packets. I keep the architecture, plans, contracts, decisions, and handoffs documented with the project, so when a webchat architecture session gets too large I can start a fresh one, audit the documentation and repository, and restore the workflow quickly.

For implementation I spawn disposable Codex 5.3 CLI sessions and give each one a bounded packet. They do the coding against the real repository and return terminal evidence.

The terminal is the source of truth. Documentation preserves architectural state and intent. Model context is useful, but disposable.

If a Codex session drifts, I kill it and spawn another. If the architecture session gets confused, I call an audit against the actual project state.

It also turns out to be a much cheaper way to use Codex for serious coding because every worker does not need the entire project history loaded into its context.

Great , it seems how we use Chatgpt is quite alike , with a difference i dont use cli , i use codex for those bounded and well defined tasks, however for long running tasks Codex does drifts from the original task after few compactions. Codex task drifting issue

You are still describing Codex as a long running worker. That is the part I removed. I do not expect a Codex session to carry a large task indefinitely. The architecture session breaks the work into bounded packets, each Codex session executes one bounded scope, terminal evidence records the resulting state, then the worker is disposable. Large projects can run for months without requiring any individual coding session to retain months of context.

Ohh no , I think perhaps its not conveyed correctly by me my aplogies, I am not talking about tasks which runs for months or in fact days , that would be very wauge approach in my opinion atleast (with an exception of rigorous research related tasks which I am not talking about ) instead I am talking about tasks which can be finished under 5-6 hours (for example optimizing an existing workflow) , I do agree with you granularity does matter here , but my observation and issue description as mention in the URL , is based on such observations of long tasks using 5-8 hours of tasks, which didn’t actually required such long time to be finished , the drifting, looping , false test scenarios complaints were regarding that .