Codex Rate Limits Discussion Thread

i’d say the opposite?

Do you understand there’s costs for OpenAI and they’re giving access to billions around the world? They’re trying to keep access available to all while keeping the servers up until they can acquire more compute/electricity.

You can choose to use the API which still has rate limits, but should give you a bit more stability for professional jobs. ChatGPT is more of a consumer product, imho.

Welcome to the community! Be sure to take a few minutes to check out our history here!

1 Like

I wonder if part of the recent sudden increase in Codex usage could be related to larger context usage, particularly in long-running sessions.

There are a few interesting facts that seem worth connecting.

According to OpenAI’s official GPT-5.6 Sol model documentation, the model supports a 1,050,000-token context window and up to 128K output tokens.

The current Codex model metadata for GPT-5.6 Sol shows:

context_window: 272000
max_context_window: 872000

So 272K remains the default, while Codex now supports opting into a substantially larger context window.

More importantly, OpenAI’s GPT-5.6 Sol documentation states that prompts with more than 272K input tokens are priced at 2x input and 1.5x output for the full request.

OpenAI’s Codex usage documentation also explains that larger codebases, long-running tasks, and extended sessions that require Codex to hold more context can use significantly more usage per message.

Tibo Sottiaux from the Codex team also recently explained how to enable the larger context window for GPT-5.6 Sol, and in a separate post noted that there had been a recent change to Codex’s context handling, while also emphasizing that the normal Codex context limit was tuned with performance and cost in mind.

That makes me wonder whether some of the unexpectedly high Codex quota consumption users are seeing could be related to substantially more context being retained or processed during long-running sessions — particularly for users using the extended context window, or if anything recently changed in how much context Codex retains between turns.

I am not claiming this is definitely the cause. The >272K multiplier refers to API pricing, and we do not know exactly how ChatGPT/Codex subscription quota accounting maps to API token pricing.

But it would be useful if the Codex team could clarify:

Has anything recently changed in the amount of context Codex retains or processes during long-running sessions, and does using the extended context window — particularly beyond ~272K tokens — materially increase the rate at which ChatGPT/Codex subscription usage limits are consumed?

If so, that could explain at least part of the sudden usage increase users are reporting.

Can you give us the link to where Tibo Sottiaux explains how to set larger context window in 5.6?

Because I’m unable to use more than default for months.
Also, the sudden drain of usage limit occurs on the default context window of about 250k tokens for me. So it’s not because of larger context windows, yet I’m still curious how to change it now so that it’s visible it changed.

ChatGPT 5.6 on very high reasoning says you can’t do that in Codex 5.6 at all. That it’s been disabled.

this one?

this might also be relevant…

i’ve been thinking of starting a thread for Codex tips and tricks… any interest from anyone?

It kind of sounds like you’re thinking we should be thankful it’s not worse.

i mean, honestly, it could be better or worse depending on how you look at it? :sweat_smile:

take a step back and think about what they’re trying to do and how they’re trying to be fair to as many people as possible. not an easy task with so many bad actors out and about.

the normal devs like us end up paying for some taking advantage however they can, but i do see things improving relatively soon. we’re due another “breakthrough” soon, imho! :wink:

but yea, is the glass half-empty or half-full currently when it comes to OpenAI?

growth pains have gotta be crazy for them to deal with internally and externally.

i’ve been thinking of starting a thread for Codex tips and tricks… any interest from anyone?

I spent $500 of my personal cash in the last 4 days just trying to keep my work project alive because I enjoy working with ChatGPT and doing things I only dreamed about before. That’s at the regular usage that in the weeks before was within the limits of my OpenAI subscription.

In one case, last week when Sol was affordable, my system did in 15 minutes what a top-tier sql developer said would take so long that it was impractical. True story. Probably saved the company $5k minimum had they funded that development, but the truth is that they wouldn’t have funded it at $5k.

For a business, it’s not a matter of brand loyalty or appreciation of “fairness”. If users can’t afford it, they will seek alternatives. At $1000 a week, hardware for an on-prem llm that will run GLM is looking like a reasonable investment, as insane as that sounds. haha…

Wow… Locked out of my account now…

{“error”:{“code”:502,“message”:“Bad gateway.”,“param”:null,“type”:“cf_bad_gateway”}}

Anyone else having this issue?

There is an Analytics page in Codex Cloud that provides detailed usage data by model. You can access it by going to Codex Cloud, opening Settings, and selecting Analytics.

A lot of discussion over the past few weeks over codex usage limits, a lot of complaints and a lot of people saying others aren’t using Codex right so I’ll leave here this example

Today, brand new chat with Terra Medium, requested Terra Medium to run through the project my project’s Roadmap md file and check which phases are marked complete and plan the development for the day. Simple request. Worked for 5 minutes, consumed 2%! Weeks ago something like this would have barely scraped 1%.

I’m going to risk saying that’s not ok. 3 to 4 weeks ago I was running full steam ahead with Sol High and Extra high, several hours a day on complex tasks and I’d get to the end of the week with over 30% left. There’s either changes in how token usage is being calculated, or limits have been reduced. I never thought I’d say this, but I’m actually getting more usage out of Claude Fable at the moment than I am getting out of Terra or Sol…imagine that! Also, my 5 Hour usage limit seems to have disappeared…unsure why, no longer visible on my usage display, there’s only the weekly limit.

And NO I haven’t used the RESETS, never needed them before

codex is celebrating 20m user today, yet no recognizing there is a problem here is not a good brand move.
loyalty is everything which they will lose if they don’t address something, don’t be an anthropic.

I’m experiencing strange drops in my usage limits. Last week, I already noticed that my limits were decreasing much faster than usual.

With Sol, the limit was dropping extremely fast, which is why I switched to GPT-5.3-Codex-Spark. But now Spark has also been exhausted despite very little usage. Previously, I could use Codex for hours every day without reaching the limit. Over the last 1–2 weeks, however, my limits have been decreasing unusually and noticeably fast.

Yesterday, my GPT-5.3-Codex-Spark usage limit was reset to 100%. Today it is already completely exhausted again, even though I have barely used Codex at all.

Something clearly seems wrong. I suspect the rate/usage is being calculated incorrectly somewhere.

I’m on the $100/month plan.

you are not the only, its rank out of my 100% reset earlier today in few hours

Hi, everyone.

Evidence is increasingly pointing to the fact that the Codex Pro x5 (and possibly x20) weekly allowance has dropped by ~30–33% in recent weeks — measured from local rollout logs, not vibes. I encourage you to check out the analysis below.

I initially thought I was imagining this. However, after finding more than a few similar opinions on this topic (including within this Community):

For previous months, I could use Codex noticeably more on a Prolite (x5) plan account, and the weekly allowance behaved reasonably and consistently. Recently, however, (roughly since the beginning of August), my weekly limit started disappearing dramatically faster - even though I deliberately moved most of the workload to GPT-5.6 Terra, which is substantially cheaper than Sol - as much as 60% less per token! So instead of guessing, I audited my local Codex rollout history, and the result is concerning…

I reconstructed several historical Codex weekly windows from local rollout-*.jsonl files and compared the amount of locally recorded model workload required for the backend to report 100% weekly usage. The comparison uses the exact same normalization function and the same published Codex rate card for every historical window. I am NOT comparing raw token counts between models.

I normalize:

  • uncached input
  • cached input
  • output

using the same model-specific rates for every period. I also use cumulative ‘total_token_usage’ deltas instead of summing ‘last_token_usage’, so duplicated ‘token_count’ snapshots are not counted repeatedly.

Historical weekly windows may appear under “secondary”, while newer weekly windows appear under “primary”; both are identified by the 10080-minute weekly window.

For 100% windows, accounting stops at the first observed used_percent = 100.

Here are the clean results:

Period Account Backend usage Normalized workload at limit
Jun 11–16 Pro x5 account A 100% 14,381
Jun 30 – Jul 5 Pro x5 account B 99% ~15,024 implied
Jul 12–14 Pro x5 account B 100% 14,202
Aug 13–17 Pro x5 account B 100% 13,514
Aug 20–22 Pro x5 account B 100% 9,613

All of these sessions report the same Codex plan type: Prolite the older clean Pro x5 windows cluster around roughly 14–15k normalized credit-equivalent workload. The current weekly window reached 100% at only: 9,613 (!). Compared with the historical clean median of approximately: 14,381, that is about a 33% reduction in effective observed weekly capacity. Even compared only with the immediately previous weekly window on the SAME account: 13,514 → 9,613

that is approximately 28.9% less normalized workload before reaching 100%.

The model mix does not explain this. The current weekly workload was:

  • GPT-5.6 Terra: 87.4% of tokens
  • GPT-5.6 Sol: 12.5%
  • GPT-5.6 Luna: 0.08%

After normalization:

  • Terra accounted for ~69.8% of the weighted workload
  • Sol accounted for ~30.2%

So yes, Sol requests were relatively expensive, but that cost is already included in the comparison. The entire point of the normalization is that an August week with a high Terra volume can be compared against an older Sol-heavy week using the same cost function. For example:

  • Jul 12–14 was essentially 100% GPT-5.6 Sol and reached 100% around 14,202 normalized units.
  • Aug 20–22 was overwhelmingly GPT-5.6 Terra and reached 100% around 9,613 normalized units.

This is not simply “Terra used more tokens”. The current week actually processed approximately 1.19 billion aggregate token traffic before reaching 100%, because most input was cached and Terra is cheaper.

So the question is: Why did the backend weekly percentage reach 100% after dramatically less normalized workload?

I am NOT claiming that this proves OpenAI secretly reduced Pro x5 limits by exactly 33%, although it is tempting…

There are several possible explanations:

  • the underlying weekly allowance was reduced
  • model-specific weekly metering multipliers changed
  • the backend is charging additional activity that is not represented in local ‘total_token_usage’
  • there is a quota/accounting bug
  • some other Codex metering rule changed

However, something measurable changed. The strongest evidence is that multiple earlier weekly windows, across two separate Pro x5 accounts, cluster around the same ~14–15k normalized threshold, while the newest window stops at ~9.6k. This is not based on how the progress bar “felt”. It is reconstructed from thousands of local Codex ‘token_count’ records and backend-reported weekly ‘used_percent’

My current weekly window:

  • started: Aug 20, ~08:30 CEST
  • reached 100%: Aug 22, ~17:12 CEST
  • normalized workload before 100%: ~9,613

I would really like OpenAI to clarify:

  • Did the effective Codex weekly allowance for Pro x5 / prolite change in August?
  • If not, what changed in the metering formula that causes ‘used_percent’ to reach 100% after ~30–33% less normalized workload than historical windows?

If other Pro users have local Codex rollout history (and you definitely do), especially from June/July and August, I strongly encourage you to compare your own 10080-minute weekly windows. This should be reproducible. If you don’t know how to do this, please let me know in this thread. I’ll explain how it’s done. Best regards.

Obviously something is wrong, the capacity is running out much faster than normal.

but the audacity of not calling this out as an issue is poor judgment on openai’s end
this isn’t the time to be way over there head to be thinking customer wont leave them. if they show now reponsibility, we also wont be as loyal if another play comes up with a better option

This describes my observations - very unfortunate. Thanks for writing it up.

Things have started to slowly, go back to normal

You are in a house of foxes indicating that chickens are missing

I’m a longtime subscriber, currently on a Pro Plan, and I find it impossible to plan work when Limits are reset randomly while also resetting the date applied to the limit.

A typical situation:
I have 70% usage left for 2 days and plan heavy work, i wake up in the morning ready to go to find: 100% usage for 7 days, meaning once i finished the work i am left with 30% usage for 5 days, where i would otherwise have had 100% for 7 days, effectively eliminating a lot of tokens and value.

These resets can act in favor of developers and can act against developers - the core message is that they are completely and utterly random, making planning impossible. More often than not, they punish reasonable use while rewarding excessive use and a gambling mentality.