Codex Rate Limits Discussion Thread

I think the thing that most people don’t seem to realize, whether they want to or not, is that this is all intentional.

When a new flagship model releases, the collective intelligence of all other models decreases (in some cases heavily), but the consumption stays the same. This allows them to hype new models while shifting the blame of ridiculous usage to “You were using a new model - of course, it drains fast.”

Then they “Solve” the issue, but everytime, it only gets worse not better. The resets they issue are just meant to make people feel like there is progress, but even then everyone is still angry :sweat_smile:.

The whole thing is just to mitigate the loss of capital from compute they are currently facing. The amazing codex limits we used to have were just to build a user base that cannot function without cheap agents, since the code most of you wrote got too over your head to maintain quite quickly. This intelligence level won’t be cheap for years. It WILL get much worse…

Something changed overnight: the same ChatGPT Work workflow suddenly started consuming 3–4x more of my Plus allowance.

For context: I primarily use ChatGPT Work, not the separate Codex workspace. My Usage dashboard shows that on September 18, about 98% of my agentic usage was Work and 2% was Codex.

I have been using Work heavily for software development for quite some time.

My workflow is very consistent:

  • local repository work
  • large implementation/debugging prompts
  • code inspection
  • code changes
  • tests
  • analysis
  • Standard speed
  • mostly GPT-5.6 Sol High

The important part is this:

My workflow did not materially change. The usage did.

Until September 17, a substantial Sol High task lasting roughly 15–25 minutes would typically consume around 8–12 percentage points of my 5-hour allowance.

I could normally work for several hours and run multiple large tasks.

Then on September 18, the behaviour changed dramatically.

I started taking screenshots before and after individual runs.

Example 1 — GPT-6 Astra Low

  • Standard speed
  • ~12 minutes
  • 5-hour allowance: 100% → 64%
  • weekly allowance: 77% → 71%

That already surprised me.

Example 2 — GPT-5.6 Sol Medium immediately afterwards

Same Work conversation.

  • Standard speed
  • ~22 minutes
  • 5-hour allowance: 64% → 20%
  • weekly allowance: 71% → 64%

So one Sol Medium task consumed approximately 44 percentage points of the entire 5-hour allowance.

At first I wondered whether switching from Astra to Sol in the same Work conversation had caused unusual context/compression costs.

So I later measured several additional Sol High Work runs as well.

Additional Sol High runs later the same day

  • ~6m 22s: 80% → 70%
  • ~10m 25s: 95% → 80%
  • ~24m 27s: 70% → 30%

The last one consumed 40 percentage points of the 5-hour allowance.

These were normal development tasks of the same general kind I had been doing on previous days.

That is why the generic explanation that “usage varies depending on model, tools, reasoning, context, task size, etc.” does not really explain what I am seeing.

Of course usage varies.

But my development workflow did not suddenly become 3–4 times heavier from one day to the next.

The practical amount of Work I can perform within the same Plus allowance appears to have dropped dramatically.

Support case

I have opened a support case and I am currently going through this with OpenAI Support.

I have provided:

  • timestamps
  • model/reasoning settings
  • Standard speed
  • before/after allowance screenshots
  • Usage dashboard screenshots

Support asked me for per-run token/credit details, but my account currently does not expose them.

My Usage UI shows:

  • per-chat usage: “Chat usage could not be loaded”
  • 5-hour historical usage: unavailable
  • weekly historical usage: unavailable
  • no Token History breakdown for cached/uncached input/output
  • no per-chat usage indicator in either the Work or Codex view

So I have asked Support to inspect the backend accounting for the actual runs.

Why I’m posting this

I want to know whether other Plus users saw the same abrupt change around September 17–18.

If you did, please include actual measurements if possible:

  • plan
  • Work or Codex
  • model + reasoning
  • Standard/Fast speed
  • task duration
  • 5-hour % before/after
  • weekly % before/after

For me, the biggest problem is not simply that there is a limit.

I can work within a known limit.

The problem is that the same kind of work suddenly appears to consume several times more allowance than it did the day before, with no clear explanation of what changed.

I am getting fed up with this garbage, many youtubers speaking about how they can leave astra running long tasks and here i am after 8mins of usage already almost used up all my tokens. this is upsetting beyond any reason.

bot forgot to bring subject over: 60% usage in 8mins astra medium

get a second subscription in that case.. and also stop using ultra the whole time. It is like using the afterburner on a jet plane.. and when you do stuff with that for paying customers you can easily charge them 50-100k per week.. easily!

Just give the market fair prices!

I’m on a ChatGPT Business Premium seat and I’m trying to understand how the Codex/Work usage limits interact.

Right now my usage shows:

  • Weekly allowance: 0% remaining, resets in 2 days
  • Monthly allowance: 0% remaining, resets in 11 days
  • 2 banked resets available

The weekly allowance has been resetting normally over the past few weeks, while the monthly allowance has continued counting down and has now also reached 0%.

What I’m trying to understand is whether the monthly allowance is a hard cap on Codex/Work usage.

When my weekly allowance resets in 2 days, what happens?

  1. Does my weekly allowance return to 100% and Codex/Work become usable again even though the monthly allowance is still at 0%?

  2. Or does the monthly 0% override the weekly reset, meaning Codex/Work remains unavailable until the monthly allowance resets in 11 days?

I’m also trying to understand how banked resets interact with the monthly limit.

The UI says a banked reset can restore the 5-hour window, weekly window, or both. If I use one while my monthly allowance is at 0%, will it actually restore usable Codex capacity?

Or will it simply reset the shorter windows while the 0% monthly cap continues to block usage?

I’m specifically asking about ChatGPT Business Premium, not Plus, Pro, Enterprise, or Edu.

Has anyone on Business Premium actually reached 0% on both weekly and monthly usage and then observed what happens when either:

  • the weekly allowance resets, or
  • a banked reset is redeemed?

I’m mainly looking for confirmation from someone who has actually encountered this state, or clarification from OpenAI on whether the monthly allowance functions as a hard cap.

I have no clue why anyone expects a $200/mo subscription to build, run, and/or maintain anything going into widespread production, let alone a $20/mo subscription. What we already get under current subscription limits is insane for the costs, and this entire thread sounds like a bunch of complainers to me. If you are on a Plus subscription and complaining, welcome to capitalism. Nut up and upgrade, or become an actual developer and use the API… Oh, wait, you can’t upgrade to Pro anymore?… YEAH!… That is because all subscription tiers are already structured in ways that produce economic loss for OpenAI when they are maximized, with Pro being the most extreme of them. Exhibit A:

… I’m a Pro subscription user, utilizing no API or credits, and I’m able to rack up billions of tokens per week. 2 billion in one day, in fact. That is not sustainable, and this clearly feeds into why OpenAI isn’t anywhere close to posting net profits as a company… It’s for the “public good” and all though, so gotta keep that in mind :rofl::rofl::rofl:…

No… API is the more easily measurable and directly billable usage method, and thus, anything that is being called “professional usage”, “professional coding”, “production deployment”, “advanced usage”, etc., etc. here is something OpenAI most likely wants routed through API. That is the dynamic I see at least, so if anyone is expecting the usage dynamics under subscriptions to get better, just look at the suspension of new Pro subscribers, and I think you’ll see the writing on the wall. If you’re serious about “production”, you better brush up on your API skills and start building your war chest.

The issue people have is that there is no consistency or transparency regarding what they are paying for. Maybe you have no issue paying and getting an arbitrary amount of usage that fluctuates, but most people care about what their $200, or even $20 is getting them.

I think an important point is getting lost in parts of this discussion: this is not about expecting unlimited Codex usage from a Pro subscription.

I fully understand that Codex is computationally expensive, that different models and reasoning levels consume different amounts of resources, and that a subscription needs to have limits.

The problem is transparency and predictability.

It starts with the way the plans themselves are described. Pro is advertised as providing 5x or 20x the Codex usage of Plus. But 5x or 20x of what, exactly?

If the underlying Plus allowance is not expressed in a stable, understandable unit, then a relative multiplier does not tell customers how much usable capacity they are actually buying. And even the published model-dependent usage ranges are so broad that they are of limited help when planning real work.

More importantly, once we start using Codex, we still cannot meaningfully see why our allowance is being consumed at the rate it is.

Many people in this thread are describing essentially the same experience: workflows that we were able to run for weeks without coming close to our weekly limit are now consuming a very substantial percentage of the allowance, sometimes even for relatively simple tasks. As a Pro user, I can now literally watch percentage points disappear during tasks that previously would barely have been noticeable.

That does not necessarily mean OpenAI intentionally reduced the limits. There may be changes in model behavior, context handling, reasoning, tool calls, compaction, metering, or something else entirely.

But from the users perspective, without transparent accounting, all of those possibilities look exactly the same: the effective amount of work we can do with the subscription has suddenly decreased.

That is why I think what several people here have already requested would make such a big difference. At minimum, users should be able to see:

  • the actual consumption of each completed task
  • how much was deducted from the 5-hour and weekly allowances
  • which model/reasoning configuration caused that consumption
  • a usable history of consumption;

It does not have to be perfectly predictable down to the last token. Even approximate accounting would be a huge improvement. What matters is giving customers enough information to understand their usage and make informed decisions about how to spend their allowance.

The new option to purchase weekly resets or additional usage is useful as an emergency valve, but I don’t think it solves the underlying problem. If I don’t know what consumed my existing allowance, buying another reset simply gives me another opaque bucket that may disappear just as unpredictably. For a professional workflow, that is very difficult to budget around.

There is another reason why this currently feels uncomfortable. Tibo wrote publicly on August 21:

“We’ve investigated a few messages about codex usage limits being different. That’s not something we change without engaging the community and being transparent.”

I appreciated that statement, and I still take it at face value. But the current experience described by many users in this thread feels indistinguishable from a substantial reduction in effective usage.

That is exactly why more visibility would help both users and OpenAI. If the limits themselves have not changed, transparent per-task accounting would make that much easier to demonstrate. It would also allow us to identify whether a specific model, setting, tool, repository, or workflow is responsible for unexpectedly high consumption instead of everyone having to speculate.

I like Codex a lot. That is precisely why this matters. People are trying to integrate it into serious development workflows and are willing to pay for substantial usage. We are not asking for unlimited compute.

We are asking to know what we are buying, what we are consuming, and how to plan around it.

This morning I used up my 5 hr window after only 5 codex prompts in vs code ( 1 hour of usage) using gpt 5.5 light mainly. this was after waiting 5 days to resume usage on my plan. Not even heavy lifting or anything. This is not the same as before. i have also had to go through several reset usage prompts which barely lasted long at all. OpenAi needs to come clean here. The limits have very much changed this past month when they rolled out Astra. even if im not using that model. I also do not a clear path to getting personal support anymore via email or a message. So what is going on? Is it time to start requesting chargebacks for those who used credit card? _) look at reddit posts about this this past month. big problem. But i wasnt even using sol or astra much

Same is happening here. I’ve started noticing serious degradation on tokens consumption after the release of Astra even when nothing in my process has changed: I keep using the same models with the same capabilities.

It’s really frustrating. I’m burning my 5h session limit in only a few minutes, and the weekly limits in less then a day when I use it more frequently.

If nothing changes, unfortunately seriously considering ditching OpenAI for some other like Cursor :frowning:

is there a better service out there we should be using? Claude doesnt give much usage either. My guess is they are trying to steer all the plus user to the higher tier then give us the same usage we used to have with plus.

Yup. Because they are pressured to make more profit. But what I can’t stand is the gaslighting by OpenAI. That somehow all those people who have been reporting big differences in token consumption from previous models are all somehow “wrong”. That they looked into it and it all seemed “fine”.

As I look at other alternatives I am reminded that this was bound to happen.

anyone using claude $20 plan is that working out well ?

When are the expected long-running 5h overflow limit behaviour (as per Using Codex with your ChatGPT plan and Pricing | ChatGPT Learn) coming back? Hard stopping long-running tasks at the 5h limit is hindering many of my projects already…

Edit.: It puts y’all under a very bad light - the honest users’ are penalised because a bunch of bad apples decided to trick the system.

I tried this approach for a couple of hours and it was an absolute disaster. I ran 5.6 SOL as the main agent, and then specifically instructred it to select Luna or Terra based on the task. The work done by Luna took me more time to fix then if i wrote it myself.

I also tried some basic non coding stuff with Luna on high, and the model is simply nonsensical.

I get more out of my $20 claude plan than my $100 codex plan since Sept 3rd.

This will be my last month of codex.

interesting. could you share how you’re getting more in Claude for $20 than codex at $100?

i’ve tried claude a few times over last 2 years.
each time the macos app fails me; glitchy, buggy, unusable. connectors and mcp’s unreliable, auth brittle and much more.
that, and ALL usage counts against your account.

I’m paying $200 per month for Pro, and I used 44% of my weekly Codex allowance in a single day. Very little work was actually completed for that usage. I barely got to accomplish anything before nearly half of my weekly allowance was gone.

I’ve tried switching to lower model/settings than GPT-6 Extra High, but I’m still not getting enough useful work out of the allowance.

The problem is the amount of completed work I’m getting for $200 a month. Using 44% of a seven-day allowance in one day would be frustrating enough; having so little finished work to show for it makes the value much harder to justify.

I want OpenAI to review this and make it clearer what is consuming the allowance. I’m also asking support for substantial free usage refills or complimentary credits. If OpenAI cannot provide enough usable access to make this subscription worthwhile, I will have to cancel. It’s too expensive for too little use.

Are other Pro subscribers seeing their weekly Codex allowance disappear this quickly, even after lowering their model/settings? How much finished work are you getting for that usage?

Last week, my entire allowance ran out in about three days. I then waited roughly four or five days for the weekly refill. Now, about one day after that refill, nearly half is already gone again—with very little work completed. For $200 a month, that isn’t enough usable access.

I’m just getting tired of having to complain to Openai and them giving me free refills. How about we give our Pro customers enough usage that they can use this the entire week especially if they are a power user? Makes sense right?

Does anyone know which other company gives more generous weekly usage? I’m thinking about trying Grok. I like Claude also, I switched to Codex because the response time is faster but I will switch back. This is absolutely ridiculous. Build the data centers in the oceans like China and lets drive the cost of this down! Lets go!

One must naturally conclude that IF your usage prior to the release of GPT-6 is the same as after, and you’re not using GPT-6 and your tokens are being burned faster, for the same work, THEN openai is at fault.

that is easy to test: check out an old git hash, push the same prompt and see what happens. I call out bullshit