Need to have serious discussion with the $200/month plan

Hi everyone,

I’ve been a long time user and supporter of codex since September last year with the $200/month plan.

I’ve seen ups and downs but lately I pointed out the major degraded perfromance from GPT 5.5.

To tell you how serious this is, I’ve had to push back production releases due to numerous issues that was performed by codex and its failure to adhere to instructions.

I’ve never had codex cause this much problems to the point of delaying releases but is happening and we are still unable to get it to rectify and fix the regressions it has caused throughout our codebase.

I’m currently relying on Opus. 4.8 to fix all of its mistakes and the reality is frightening. Not only are there serious security lapses caused by its degraded ability to follow instructions, I find that a lot of the sloppiness coincides around the time where the usage sync issue was declared fixed. This means that any work it has done since then now is under scrutiny and Opus 4.8 is finding a ton of critical issues and worse xhigh appears to have refactored the code base in a strange way that does not address the prompt instructions but hallucinated big time writing .

I’m unable to any other feedbacks but it is quite worrying since we rely heavily on codex and it has been excellent as of date but we are still seeing the same degradation issues I addressed a week ago and others have pointed out.

Also there is the larger issue with GPT Pro. This does not seem like the old Pro model where it has been able to think through a problem for extended period of time. It often responds way too quickly and it makes me think that there is not enough inference being performed.

Lastly, the usage issue. Overtime I’ve noticed gradual inflation and reduction of usages. Now that the promotion period is coming to an end as of this month, all of these collective issues is making me rethink my workflows purely centered around codex.

I hate to say this but this seems to happen repeatedly with new model releases in the past. There’s a honeymoon phase where we seemingly have maximum and optimal inference then it slowly degrades but i’ve never seen it this bad.

As a result I am now using Claude , in particular Opus 4.8 and leaning on it to correct the errors caused by gpt 5.5 xhigh. A month ago, this was unthinkable. 5.5 xhigh got a ton of work done and could be trusted but here we are, we seem to be back in the era of 5.3~5.4 suddenly.

Once again, I implore the OpenAI team to take this issue seriously and investigate not just my reports but others raised on here and other social media platforms. I want Codex to succeed and I have benefited from it and want others to do the same continuously.

Thank you.

Yes, I was also one of the people who felt that there had been a regression in that regard. However, at this point, I don’t think the decline is nearly as severe as some people suggest. There are clear signs of recovery.

Also, when GPT-5.6 is released soon, I think switching to Opus 4.8 might be something you regret, because OpenAI has consistently been very strong at pushing the frontier forward and releasing models that stay one step ahead.

To me, Opus 4.8 feels like a balancing model somewhere between Gemini and GPT, taking certain strengths from both. Gemini tends to be highly creative and energetic, but it also hallucinates more and can be overly enthusiastic. Its tool usage and context tracking are not as reliable. ChatGPT, on the other hand, tends to be less imaginative and more conservative, almost the opposite of Gemini in that sense.

Claude models feel like they sit somewhere in the middle, trying to balance those two extremes.

However, when it comes to heavy reasoning, serious coding, and complex architectural work, I currently don’t think anything beats GPT models. If you are building a large, complex architecture, moving away from GPT would likely be a serious strategic mistake.

There hasn’t been degraded performance from GPT 5.5.

These issues are caused by how you use GPT, you are likely feeding it WAY too much irrelevant context, which causes the issues you are seeing now.

The reason it works fine with Opus 4.8 is because you don’t have these context issues, and you are essentially starting fresh with Opus 4.8.

How you can confirm this: go back to one of the issues that you claim could “only be fixed by Opus 4.8” and instead use GPT 5.5 via the API the same way you used Opus 4.8

Humans love to place blame on anyone besides themselves, but 99% of the time it is indeed the users fault.

Please keep discussions professional this isn’t r/codex

Update: I am still not seeing any noticeable improvement after I started a few threads on GPT 5.5 not performing consistently. Starting a new chat doesn’t seem to help. I am stumped why this is so difficult , it is not a hard frontend fix but all the 5.5 models seem to be on their 50th attempt and still not able to fix it.

In contrast Claude and opus 4.8 one shots it. How is there this much of a gap between GPT 5.5 and Opus 4.8 ?

It’s not a context issue because starting a new chat doesn’t change the outcome. Mean while the same problem when asked in Claude, it fixes in far fewer steps. This points to a real gap between the two models. Let’s try to keep the discussion professional and relevant please.

I’m switching to Opus 4.8 when GPT 5.5 can’t manage it and not sure what you mean by regret when it unblocks GPT 5.5 on problems that it can’t solve. I really don’t understand this constant downplaying of Claude, both have strengths but when it comes to frontend bugs, there’s just no comparison, Claude wins everytime and I have the results in front of me.

The model’s conscientiousness has clearly declined; it no longer engages in effective thinking but instead executes mechanically without consideration, even appearing to be outwardly compliant but inwardly defiant.