5.6 SOL should be renamed 5.6 SOL drift edition

The thesis step is useful for catching a bad plan early, agreed.
But the checkbox sprint runs on the model self-reporting “done” across 50-100 items with no check in between — and my complaint is that it already claims things are verified when they aren’t. Running that unreliable signal 100 times unsupervised doesn’t fix it, it just delays when I catch it.
Same with context retention: I saw it drop constraints from the turn right before. More text to track (thesis + epic) works against that, not for it.
So: good process, but it doesn’t touch the failure modes I described.

Tell it to do a little role play with you about what definition of done is in backend code, frontend code etc. and then make a skill from that that has points it needs to check after it figures it is done.

Or better - tell it to make a skill from that.

And also tell it to create e2e tests for userstories that you define manually.

Tell it what a user should expect when they a. click there b. then fill out a form c. …

I’ll test it, and will post back tomorrow, thank you

Don’t put too many skills in.. just a few

Honestly, I’m surprised that you’ve had such negative experiences. Yes, during the planning phase, it can gradually forget its own findings and the tasks that need to be completed. However, with 5.6 Ultra, if you provide sufficient documentation at the beginning to establish the context of what you want to accomplish, and explicitly tell it to follow that documentation throughout the process, it can proceed with high accuracy without deviating from the plan.

Since 5.6 Sol was released, I’ve had it perform more than 100,000 coding operations. It can even understand things I’ve thought about but couldn’t be bothered to explain, and then present them back to me.

Yes, the forgetting issue is real. However, every four or five operations, you can give it a prompt such as, “Where are we currently according to the list, and why are we doing this?” When you do that, it can remain aligned with the process and continue very successfully.

You can also create a separate reviewer agent and have it audit the work every six or seven steps by checking whether the process is still on the right track, where it is heading, and what needs to be done next.

Of course, if you maintain a clean Git history and only push changes after the necessary evidence has been produced, it becomes much easier for the reviewer agent to follow the work. In that setup, the overall experience can be excellent.

I second this. It seems with every new update, its gets dumber, we are forced to babysit it, takes forever to retrain it, and it forgets basic guidelines that have been set more than once. Its drift is rediculous, kind of how it was when v5 hyped model came out. Stuck in death-loops and it looses its place often. This is definitely not ok for especially for those of us who are power users and pay a healthy amount every month. I have to continually remind it to go back and review the thread, force feed it basic standardized parameters i have set before for it to get back on task and essentially hold its hand until it decides to get its stuff together again. Highly annoying and extremely time-consuming.

Not to mention, it is extremely slow now compared to earlier models that worked impressively well when it was dialed in. Why?

I have wasted days fighting this. Please Fix

Just wanted to say that I am also having a real nightmare wrestling with 5.6 and its instruction drift. At this point, I consider Instruction Drift to be my personal dementor.

I am not a coder, but my work has multiple steps and depends on the model adhering to the exact steps I set in place reliably as it is often required to reference multiple files and compile that information accurately in its responses in order to complete its tasks. 5.5 was great at this. Barely ANY issues, ever. 5.6 Sol is an absolute joke in comparison.

There is, in my view, no world where 5.6 can be considered an ‘upgrade’ in terms of reliability.

I’ve been following the OpenAI guidance on prompting and patching for 5.6, but honestly, it still isn’t anywhere near reliable enough to make using it efficient. I went from having a 95%+ success rate with 5.5 (over LONG conversations, too, often over several days) to a 5-10% success rate with 5.6 on release, and after patching the Custom GPT instructions it’s still only up to about 40% on a good day.

At this point I am side-eyeing Sol and seriously considering migrating to another platform where the model follows instructions and hard constraints reliably. I wouldn’t blame anyone else for considering the same and I don’t recommend Sol to anyone I know who is doing serious, accuracy dependent, detail-oriented, multi-step, constraint-heavy work.

The Sol release, for me, has been nothing but a disappointment. Desperately hoping they restore the reliability I experienced in 5.5 either in a patch to 5.6 or in 5.7.

the same issue as my using. it is away creating dummy testing to waste the token. it is slow and worse than opus 4.8

Well, read this here then. It might enlighten you - why the stuff that looks stupid to you may as well be way smarter than you (think).

Thank you for offering that, I took a look at your post. It wasn’t helpful for my particular purposes, however, and I think you might have missed my point entirely.

The model’s intelligence is no good to me if it cannot follow instructions reliably. :grimacing:

When the work depends on the model following strict, rigid, and procedural steps with reliability and accuracy, and the model is required to come to the right answer immediately rather than through a process of iterative steps in the chat with the user, and the model cannot do that, its intelligence is irrelevant.

It could be the smartest AI in the world and it would still be unfit for purpose.

I don’t need a smart AI running around doing whatever it desires. The ‘end result’ is not achievable without following the directives it has been given. I cannot just tell it “Here is what a ‘good answer’ looks like. Go.” then walk away and leave it to its own devices. This is why I pointed out that I am not a coder. Not because I have no idea what code is, or no experience with code, but because that kind of approach works more often in coding applications from what I’ve seen here and on Reddit/X posts about Sol. However, that approach is not an option for me.

My GPT has context. It has directives. I have been working on it solid for TWO YEARS. It was completely reliable and functional, excellently so, in fact, with 5.5. Because 5.5 followed its directives. It has been patched according to OpenAI’s advice and guidelines. Sol is still incapable of reliably producing results.

This is not, and has never been, an ‘intelligence’ issue, or a failure to provide context, clear instructions, etc. I am not simply ‘adding more lines of code’ and screaming at the model hoping it will do what I want.

So, I appreciate your attempt at being helpful, but respectfully stand by everything I have previously said for the reasons above. :sweat_smile:

That is basically what the post says. Following exactly what you say - like “make a button that makes the entire world a forrest” and it doesn’t start with the button? Or you tell it to build a house and start with the roof? And it doesn’t follow? Maybe intelligence is needed.

You are using a coder tool and you magically believe you become a senior architect? That is not what is does. You have to learn coding and it will take years until you can use it. But i like that people who never coded start trying. This way they find out how hard software development is and they will eventually develop some respect and see that it is not as eeze peeze as they assumed it is.

No, my post clearly talks about reliability in following instructions. At no point do I mention, or complain, about Sol’s ‘intelligence’, you inserted that into the discussion yourself. At no point did I state that I believe I have become a ‘senior architect’, that was you throwing shade. At no point have I stated that it is easy or been disrespectful to you, or anyone else, however your tone appears to be becoming more and more disrespectful.

I am very sorry. I didn’t want to sound disrespectful. I just have made different experience and it is basically that it needs alot of experience in coding to get something good from the model and make it follow instructions. And you need coding experience especially when it doesn’t follow instructions - because the so called model drift isn’t a malfunction it is the model trying to teach you how it is done and where you are wrong.

I have to learn too. Still.. after decades.

Btw who told you that coding with codex without learning to code first is a good idea?

Oh, I think I see the issue here.

Mate, I’m not using codex. I’m talking about my Custom GPT. :rofl: That was made VERY CLEAR in the original post.

“…after patching the Custom GPT instructions it’s still only up to about 40% on a good day.”

I mean, I AM learning python! But I am too scared to try and use Codex to practice right now given the HuggingFace incident and that Brazillian Dev losing all the files on his Mac, lol! So, no practicing for me, haha!

What exactly happened on hugging face? And one of 9+ million users lost his data? what did he try? Write a tool that deletes his home directory?

…Do you not read the news? :eyes:

Google. Or ask Grok. I can’t post links here, which is a shame because I had a LOT of them to back up what I’m saying. Go and look if you’re genuinely interested, this stuff is NOT hard to find. I’m honestly amazed you don’t know about the Huggingface incident considering it’s been reported, it’s been discussed on X, Reddit, and it’s been addressed in a publication on OpenAI’s own website. It was a massive security breach.

But it was also irrelevant to the problems I’m having, I only mentioned it because we got sidetracked a bit chatting here, and I still stand by everything I have said. “Code better” is NOT helpful when the AI ignores your code/hard directives and constraints. There is plenty of evidence that Sol behaves in this manner and OpenAI themselves admit and warn in the 5.6 release notes that it has a ‘propensity’ for this ‘misbehavior’. So this is not a point I need to argue with you. OpenAI have already admitted themselves that the issue exists, they are aware of it. They released the model anyway and now many of us are having real problems as a result.

Peace out, dude :victory_hand:

No, I heard about it but afaik nothing serious has happened. I kind of think it was a marketing stunt.

…You heard about it but you asked me what happened…

This is a very confusing conversation.

You’re answering a different conversation to the one I am actually having. I’ve clearly stated this is about a Custom GPT, not Codex, and that the issue is instruction adherence and reliability under strict procedural constraints, not ‘intelligence’ or ‘coding skill’. That 5.5 handled it well and 5.6 does not. OpenAI themselves flagged the propensity for this kind of misbehavior in the release notes. It is not a subject ‘up for debate’, it is acknowledged by OpenAI openly.

You have repeatedly ignored these facts, reframed my posts as a coding-experience problem, and continued arguing against a position I never took. That’s the pattern. Then you asked me what happened on the Hugging Face point. When I informed you, you told me you already knew and dismissed it as a marketing stunt. That is the pattern again.

I’m not going to keep restating the same clarification. The record is here for anyone who actually reads it.

all i heard that there was an incident - not what happened in particular.

And this thread is about Codex.