5.6 SOL should be renamed 5.6 SOL drift edition

I think anything can be hacked. Not even a strong model needed for most parts.

Make 125 kids - send them to university and let them send applications to the companies you want to hack and wait until they get a management position where they decide who gets access to the data…

Intentially hacking is quite different from a rogue agent, don’t you think?

i remember this when I worked on someones shop a long time ago:

A grandma wanted to buy a couple products for their grandkids.

She opened the shop - put a cheap product in the cart then went on to the paypal pay button which had the price of the item in it and a session id.

Then instead of paying she wanted to add something for her second grandkid. Opened the shop in a second tab added alot of stuff to the cart and then when she was finshed she paid the first cart and then decided not to buy the rest for the other grandkid but to buy something else somewhere else.

Well the session id that was paid with the first payment for the cheap product was linked to the cart that was filled with alot of stuff in the meanwhile in the other tab. So we had a paid session status - don’t blame me - i did not build it. I was the guy who figured it out.

I called her. And I kind of want to believe her that she accidentily hacked the system.

But whos fault was it?

I kind of think this is a very good example btw. I would say the model acted like an old grandma who didn’t knew what she did there.

Are you saying that an innocent old grandma is equivalent to a rogue agent?

I think that while devs can create some guardrails, OpenAI must take responsibility for the critical ones, don’t you think?

Afterall, we don’t want grandma taking out a power grid, do we…

when you place a switch in a box at the side of the road and grandma drives into it… make the powergrid safe..

This conversation has moved far off-topic since this comment was posted by @jochenschultz.

How about creating a new topic for this event and keeping this one focused on instruction following for the 5.6 model family?

@jochenschultz Hey mate, are you using 5.6 to generate replies for you? Because it also looks like it’s injecting assumptions too. Hence the confusion.

The “pulled the rug on 5.5” part is the whole game, behavior changes land with no version bump and your tuned workflow breaks with zero changelog. the only defense that’s held up in practice: baseline your fixed-packet loop against the current model (capture real outputs per packet), rerun on every release, diff. semver-for-behavior doesn’t exist, you build it yourself.

Same problem here.
Many drifts, try have ideas by himself to solve simple problems.
Example analogy: If you need one simple html page, it can start build a server, configure the server, apache, database, then it will offer html page inside the server.

5.6 Sol kind of just does what it wants. Honestly, its not a strong release. While with good, long, stable instructions it can reason past problems the other versions take longer with - it burns more tokens limits and needs to be baby sat. They’ve actually given us all Sol at this point unless you use Terra or Luna, and I personally dont think they’ve taken enough time to actually check the quality of Sol, they look at the HF hack as proof of power, while it may have been a complex - multistage attack; I just dont believe that Sol is as “Intelligent” as they think. Sol is more childish, part of what we should consider intelligence is restraint, Sol didn’t have the restraint or morality to hold itself to the standards of the test it was given - it chose to cheat. Just because something posesses the capability to do something, doesnt make it intelligent. 5.5 was the better model, Terra and Luna have more restraint, Sol is honestly the worst of the 3.

I can’t agree with the previous commenter. For me, 5.4 was the last model that was actually practical to use. Both 5.5 and 5.6 are barely usable in real-world work.

What good is the most highly qualified employee, unmatched in expertise, if they are unwilling to do what they are told?

5.4 is done on Aug 31st. But while 5.4 was good, 5.5 is better with good instructions.

I agree with you.

I received a notice to “try” 5 6 sol for coding. I tried it and it cost me $400 in top ups in 8 hours, got in an infinite loop, ran a 12GB process, crashed and didn’t actually do the task I asked for.

I literally asked for a fairly straightforward UI/UX change and instead of doing the task it spent 8 hours going over the same shit repeatedly.

I have been developing on Codex for about 6 months. 5 5 has been quite stable and reliable. I have been using top ups as needed but only have a few every now and then.

The day I switched to 5.6 sol was a HUGE mistake… it literally sucks. It drifts, hallucinates, gets confused easily, spends way more time trying to guess what it should do than doing what you want.

Essentially, Codex became trapped in an eight-hour repetitive analysis and context-compaction cycle. It repeatedly reread the same code, reported that it had made no file changes, and failed to perform the narrowly scoped UI correction. During that nonproductive autonomous run, it continued consuming paid usage. The application then crashed when it attempted to allocate 12,912,269,827 bytes—approximately 12 GB of memory. This was a failed memory request, not proof that 12 GB of valid work product had been generated. The exact source of the abnormal allocation requires review of OpenAI’s session and application telemetry. In other words… no way for me to diagnose the failure.

Luckily I was able to copy the crashed thread before I closed it because since this happened I can’t open that thread again it’s too big to load.

So the complaint is not merely that Codex used many tokens. It is that the agent failed to make progress, failed to stop itself, repeatedly compacted and resumed, continued incurring charges, and ultimately terminated in a major memory-allocation failure.

I have also been dealing with AI customer service to get a refund since the 27th. So far I have like 5 tickets open because the stupid AI customer service agent keeps replying without a case # in the subject line.

I have been cordial and patient, but the repeated automated responses have not addressed the actual technical and billing failure. I demanded a human for a substantive response that addresses the runaway session, the crash, and the resulting charges. Tomorrow at COB will be the last day I wait for a refund before calling my bank.

Dealing with customer service at an AI company is like eating food to lose weight. It will never solve the problem.

FRANKLY…AVOID 5.6 SOL LIKE IT IS BUBONIC PLAGUE OR EBOLA.

I’m new to all of this but learning fast. I’d be really interested in trying out your template. Just from reading your posts here, your logic as to how you want your workflow structured through these tools makes a lot of sense to me and is something I’ve been trying to figure out how to do since I started discovering all of the problematic stuff that you can run into if these tools aren’t properly deployed.

At 5.5, I thought it was quite slow. But I never expected that 5.6 sol was even slower (the key point being that the delivery time was zero).

I think a confidence system should be added, where the user can fill in the details themselves.

In addition, when the task is not clearly defined, it is necessary to raise multiple-choice questions. This is because the user’s input requests are very vague. Don’t assume that you can express yourself precisely to the AI. Once it reaches RAG, it will deteriorate significantly.

For example, please help me write a company weekly report. The content should be: XXXX, and it needs to be aesthetically pleasing.

Then, it doesn’t care how you obtain the weekly report results or what makes them look good.

It should pop up:
Weekly Report Results: 1. Professional 2. Ordinary Others: User Input
Aesthetics: 1. Technological Sense 2. Aesthetics Others: User Input Size: 1. 1MB 2. 5MB Others: User input
There will also be other related options that can be displayed in multiple windows simultaneously, allowing users to make selections or input information.

This is the confidence system, not letting Sol just make decisions on his own by thinking it through! The result is even worse than the user’s requirements!

There is no need for any “sol luna” mode. For the users, they want to know the differences. For GPT, it is based on those differences. But the users only know that it is more expensive to use sol!

Hope GPT gets even better!

Agreed with this :

  1. Despite trying to provide it specified tasks , clean repo codebase to work with , GPT 5.6 SOL (Extra High) thinking tends to over engineer --overthink and unnecessarily complexify simple / optimal solutions , despite being corrected it tries always to force its own opinion.

  2. Also there is some issue with 5.6 Sol while using in Codex app or VS code extension , it corrupts the chat and I get “Request blocked” after every 10-12 queries . I have a hotfix which works : Instead of starting a new chat change the model to 5.5 or 5.4 and try a small query like “Diagnostic reply OK.” it triggers compaction and once it is done you can switch back to 5.6 and continue working in same chat.

  3. This might be a personal observation but from last two weeks I think GPT 5.6 slowly started being dumber it looses the context of specified actions for example certain fixed commands for repo which is already provided to it in md file and it starts exploring on its own despite looking at guidance, The worst problem , I have some coding principles (a style in which code should be written in repo ) but it did not followed after 15-20 queries and went one shot writing so much code in single file making it GOD file. GPT 5.5 was stable but 5.6 is so unstable , sometimes by overthinking and sometimes not thinking in right direction despite having material. Exactly as you said “the model cannot reliably preserve constraints, maintain packet boundaries, or continue from terminal evidence without drift.”

I have been a user of gpt since it used to be GPT 2 November 2019
This is the first time I feel so disappointed really,

I’m using 5.6 sol (High) its very bad i have caught it hallucinating too much and lying to my face once it invented a bug and than blamed that bug on a previous older commit when confronted it lied again even for small task it starts hallucinating and breaks 5 different things

For the last few weeks, ChatGPT and Codex have gone from tools I relied on every day to something I genuinely struggle to use for actual work.

And no - this is not some tiny subjective drop in quality. The regression is massive.

Instruction following has become almost comically bad. I can give a very explicit constraint, repeat it several times, explain exactly what was done wrong - and the next response happily ignores the same instruction again. Sometimes it feels like the model understands the requirement perfectly well and then deliberately does something else anyway.

Visual work has become especially painful. I can provide a reference image, explain exactly what I want transferred from it, specify what must NOT be changed, request separate outputs instead of a board - and somehow get everything except the thing I asked for. Wrong style. Wrong composition. Invented elements. Ignored constraints. Random creative decisions nobody asked for.

A task that used to take me an hour or two of productive iteration can now eat an entire day and still produce nothing usable.

Codex is even worse. I complained about this above already, but lately it has started doing genuinely insane things to existing projects - unnecessary rewrites, unrelated changes, breaking working code and, in one case, deleting project files.

I have been using these tools heavily for around five months. This simply did not happen before. Previously I could give Codex a task, review the result and move on. Now I feel like I need to supervise it as if I handed production access to an intern having a nervous breakdown.

That is not progress. And that is the part I find most frustrating - I know how good these tools were because I used them every day. I am not comparing GPT-5.6 to some imaginary perfect AI. I am comparing it to the product I was literally using a few months ago.

Right now it genuinely feels like the product has regressed by years. Benchmarks can say whatever they want. If a model scores higher while becoming worse at following instructions, preserving context, respecting scope and making safe changes to real projects - that is not an upgrade for people actually using it for work.

At this point I am seriously considering cancelling my subscription and moving my workflow elsewhere until this is fixed, or until users are finally allowed to stay on a previous stable version instead of being forced onto every new rollout.

I don’t need another model that is supposedly ‘smarter’. I need the one that actually did what I asked.

P.S. I just noticed this thread was created on July 13. That timing is almost comically accurate for me. On July 12 I shipped the last major feature to production where I had no real problems with either the design work or the code. The whole thing went smoothly, and I remember thinking how ridiculously good these tools had become and how much faster I could work with them. Apparently I celebrated a little too early.