Codex Rate Limits Discussion Thread

I want to add one concrete data point from someone who uses BOTH Codex and ChatGPT Work extensively for substantial professional tasks.The 5-hour limit is now actively breaking my agentic workflows.

This literally happened to me today:I gave Codex a substantial job. It was working on it when the job was stopped in the middle because my 5-hour limit was exhausted — even though I still had weekly capacity available.

I now have to wait for the 5-hour limit to reset, and it is not even clear to me whether I will be able to continue the existing job afterwards or whether some of the work already performed by the agent has effectively been lost. For an agentic development tool, this is a serious reliability problem.

And I am seeing the same fundamental problem with ChatGPT Work. I use Work for larger autonomous tasks: working across repositories, refactoring websites, extracting and processing larger datasets, research, and other multi-step deliverables.

These are precisely the tasks for which Work and Codex are most valuable: I give an agent a substantial job and let it work autonomously. A short-term limit that can interrupt such a job halfway through — while weekly capacity is still available — directly undermines that value proposition.

To be clear: I am NOT asking for unlimited usage. I understand why OpenAI needs usage limits, and I have no fundamental problem with a weekly allowance.

My problem is that OpenAI gives me a weekly agentic allowance and then prevents me from deciding when to use that allowance. If I have 30% of my weekly capacity remaining but cannot use it for a substantial task because the 5-hour limit has been reached, then that remaining weekly capacity is not actually available to me when I need it.

For me, this is not primarily a question of whether the limits are “generous enough”. It is a product-design and reliability problem.

Agentic tools need to be able to finish substantial jobs reliably. If I cannot know whether a Work or Codex job will actually finish before an unrelated short-term limit stops it, I cannot confidently use these tools for exactly the larger tasks they are designed for.

I already submitted this as product feedback through OpenAI Support. The AI-assisted support response captured the issue very accurately: “I’ve captured your feedback about the 5-hour rolling usage limit for ChatGPT Work/Codex and its impact on long, agentic workflows, and I’ll pass it along to the team that owns Work/Codex limits (including your point that it can block usage even when weekly capacity remains).”

That last sentence is exactly the problem. Please reconsider the 5-hour limit, or redesign it so that substantial agentic jobs are not interrupted or prevented while weekly capacity is still available.

Otherwise, Work and Codex risk becoming unreliable for some of the very use cases where agentic tools provide the greatest value.

I think the problem goes beyond the confusing “5-hour limit” label.

Today my weekly Codex allowance reset. After roughly 30–40 minutes, around 15% of my entire weekly allowance was already gone, and I had also exhausted the so-called 5-hour limit.

That would already be surprising. But what makes it much worse is the amount of useful work I actually received for that consumption.

I was testing Astra on a real development project. Yesterday it consumed a very significant amount of my allowance working on a set of corrections. Today it resumed that work, spent around 10 minutes on it, reported that the corrections had been completed and tested — and when I opened the application, several of the obvious problems were still there. In one case, the result was essentially the same broken behaviour I had reported before.

So I had to ask Codex to review its own previous work again. It started doing that, found some of the problems… and before it could even finish the correction, I hit the usage limit again.

That is the part I find unacceptable as a paying user.

Astra is presented as a more capable model that can solve difficult tasks more effectively. I can accept that a more capable model may consume more resources. What is very difficult to accept is consuming approximately 15% of a weekly allowance in around half an hour while still ending up with an unfinished task and having to spend even more allowance correcting work that was already reported as completed.

And on top of that, the UI calls another restriction a “5-hour usage limit”, even though those “5 hours” can apparently be exhausted in a fraction of that time depending on compute, context, tools and model usage.

This makes it almost impossible to plan serious development work.

I am not asking for unlimited usage. I am asking for three things:

  1. Clear information about what the “5-hour” limit actually measures.
  2. Transparency about the relative consumption of models such as Astra versus Sol.
  3. Usage limits that make sense in relation to the amount of effective work delivered.

If a model consumes substantially more allowance, that can be reasonable if it also produces substantially better results or completes the work faster. In my experience so far, I have seen the higher consumption, but I have not seen a corresponding improvement in the final result.

For a professional tool, predictability matters. I need to know whether I can start a development task with a reasonable expectation of finishing it — not discover 30 minutes later that I am locked out for several hours with the task still unfinished.

Exactly my experience (frustration) with Astra so far. On top of it when “5 hours” is up it stops whatever it’s doing and there is no way knowing if it did managed to finish what it supposed to do or just went on brake for 5 hours. So i had to ask:


Looks like it got “smart” with me about previous vs latest use and that three lines quip string used 11% of another “5 hours”. And it just stopped. I had to poke it.

17 minutes later and just one more prompt asking to correct a few very obvious errors and the remaining 89% gone leaving me with another usage limit and not finished task.

I was born and raised in Soviet and remember how you could stand in a line for hours just to get a “Closed for lunch” sign in your face making you to wait for another hour until they open again. Long forgotten feeling is back.

I’m on ChatGPT Plus. My weekly Codex allowance reset today, so I started at 100%. I began a small, self hosted media project for my home network. Nothing exotic: a Synology NAS with my own movies, a Linux container as backend, and an Android TV box on the TV. The goal is a lightweight library and player without the overhead of Plex or Jellyfin.

Codex did solid groundwork. It inspected the Android box, including the Android version, CPU architecture, hardware decoders and ADB. It helped mount the NAS share read only and used ffprobe to inventory the collection, around 654 files and 775 GB. To be clear, it did not transcode 775 GB. It read metadata only, and the NAS stayed read only the whole time. The result was useful. Almost everything should Direct Play, a handful of old AVI, Xvid and DivX files need practical testing, and two files need special handling. Then it prepared the build directory. In other words, discovery was done and we were ready to actually start building.

Then I checked the meter: 23% remaining.

Roughly two hours, around 77% of my weekly allowance, and the actual app did not exist yet. No APK, no backend, no library, no metadata, no UI and no deployment. Just technical discovery and preparation.

I understand that agentic development costs compute, and I am not asking for unlimited usage on Plus. But a “weekly” allowance that loses three quarters of its capacity during two hours of preparation does not feel like a weekly development budget. It feels like a few hours of development spread across an entire week. What am I supposed to do with the remaining 23%? Start building the APK and hope the meter survives? Stop debugging halfway through and wait for the next reset?

The worst part is the unpredictability. I had no way of knowing beforehand that this work would consume 77% of the weekly allowance.

So, concretely, what would actually help:

  1. Show per task consumption and break it down. Repository context, reasoning, tool calls or whatever is actually consuming the allowance.

  2. Warn me before an expensive task is likely to consume a large share of the remaining weekly allowance, so I can decide whether I want to run it.

  3. Include enough predictable capacity to finish at least one meaningful development milestone in a single working session.

The capability of Codex is not the problem. It is genuinely capable and careful with my systems, which is exactly why this is so frustrating. I want to use it. But the limits are unpredictable enough that I cannot trust Codex to carry a project from start to finish.

And this is what happens in practice when the meter runs out: I hand the unfinished project over to Claude and continue working there. I have worked through long development sessions with Claude and personally have never run into a comparable situation where I suddenly had to stop working because the available usage was gone. With Codex, this keeps happening to me after relatively short periods of serious work.

That is the part I do not understand from a product perspective. I am paying for ChatGPT Plus. I want to use Codex. Codex is capable of doing the work. But the usage limits repeatedly push me toward a competing product simply because I can continue working there.

I’m not asking for infinite compute. I’m asking for predictable limits and enough usable capacity to finish a meaningful piece of work.

Adding one more data point to my earlier post: GPT-5.6 Sol Ultra used up my entire 5-hour Codex/Work usage window in about 20 minutes.

That is exactly the kind of burn rate I was talking about. I understand that Ultra uses more compute, but consuming a full 5-hour allowance that quickly makes it extremely hard to use on Plus for sustained coding or research.

This is why I think OpenAI needs either a higher Plus allowance, a lower Sol Ultra usage cost, or much clearer per-task cost estimates before a run starts.

After reading through the more recent reports here, I think this is clearly bigger than just my Sol Ultra example.

People with very different usage patterns are reporting essentially the same thing: the 5-hour allowance can suddenly disappear after only a few minutes or a small number of tasks, even when a large amount of weekly capacity is still available.

That creates several problems at once:

  • Light users are hitting limits they previously almost never reached.

  • Long-running agent tasks can get cut off in the middle instead of being allowed to finish.

  • Users can still have plenty of weekly allowance left but be unable to use it because the 5-hour bucket is empty.

  • Similar workflows seem to consume dramatically different amounts of quota from one day or week to another.

  • Resets don’t necessarily solve the problem if the new allowance can also disappear within minutes.

  • We still cannot see exactly how much each individual task deducted from the 5-hour and weekly limits.

My own example was Sol Ultra consuming my entire 5-hour allowance in roughly 20 minutes, but other reports in this thread make it clear that the underlying problem is not limited to one person or one model.

At minimum, I think Codex/Work needs:

  1. Exact per-task usage deductions.

  2. An estimated cost before starting unusually expensive tasks.

  3. A warning when a task is likely to consume most of the remaining allowance.

  4. Active tasks being allowed to finish instead of abruptly dying at the 5-hour cap.

  5. More flexibility to spend the weekly allowance instead of being blocked by such a restrictive short-term bucket.

  6. Clear notices whenever model weighting, default settings, or effective limits change.

A five-hour window should be a usage-control mechanism, not something that unpredictably turns into 5, 15, or 20 minutes of actual productive work.

The biggest issue for me is predictability. I can manage a limited allowance if I know what things cost. I cannot realistically plan development or research work when the same subscription can provide dramatically different amounts of usable work without explaining why.

Your 5-hour example is bad enough, but I’m currently seeing the same basic problem with the weekly limit.

My weekly Codex allowance reset to 100% today, September 7, at 09:46. I then worked on one relatively small Android TV media project. Codex analyzed the existing media collection, built a small test app and HTTP range server, installed the APK on the Android TV box, and ran some playback tests.

After roughly two hours of actual project work, I was down to about 8% of my entire weekly allowance.

I stopped all development and testing at that point because I didn’t want to burn the rest. We only had a short conversation afterward about the rate limits. When I checked the usage page again, I was at 6% remaining.

So in my case this isn’t even primarily about the 5-hour window. I have consumed 94% of a freshly reset weekly allowance on the same day it reset, on one project, before the prototype is even finished. My next weekly reset is September 14 at 09:46.

I don’t have a problem with Codex taking five minutes longer or using more compute for a difficult task. The quality of the work is not my complaint either. The problem is that I cannot use it as a serious development tool if almost an entire weekly allowance can disappear in a couple of hours.

That’s why I completely agree with your point about predictability. Show us what individual tasks cost, warn us before a task is going to consume a significant part of the allowance, and give us enough information to decide whether a task is worth running. Right now I simply cannot plan around it.

This is exactly why I think the problem is bigger than just one model or one 5-hour window.

If one user can burn through nearly an entire weekly allowance in a couple of hours, while another can lose a full 5-hour allowance in around 20 minutes, then the real problem is predictability and transparency.

Right now users are being asked to treat Codex like a serious development tool without being given the information needed to budget usage like one.

At minimum, OpenAI should show:

  • the estimated cost of a task before it starts

  • the actual cost after it finishes

  • how much was deducted from the 5-hour and weekly pools

  • a warning before a task is likely to consume a large percentage of either limit

  • a usage history so people can identify what actually caused a spike

The issue is not that difficult tasks use more compute. That is understandable.

The issue is that a user can go from a fresh allowance to nearly exhausted without any meaningful way to predict it beforehand.

If Codex is meant to be used for real projects, people need to be able to plan around its limits instead of discovering the cost only after the quota is already gone.

With all due respect, guys! I love you, I really do, but for gods sake, DON’T interrupt a running task and risk consistent continuation with the new 5h-limit workflow interruption, PLEASE! Interrupt on weekly limit used: OK! Let Astra eat up half of the weekly limit for 1 medium-sized task: COMPLETELY FINE WITH THAT! I shall buy tokens if needed. But DON’T, PLEASE DON’T kill my task mid execution for a 5h-limit hit only.

Thank you! I love you!! Love working with you and paying for it. Do me and 50 million other users that favor. Return to weekly limits only (best case) or don’t let 5h limits interrupt workflows (second best) :wink: THANKS!!

Something changed since Sept 4th, I have not been able to get codex to complete any prompt without it telling me half way through that I’m out of tokens.

This is exactly why I think the problem is bigger than the 5-hour rate limit.

In my case, the 5-hour limit is not even the problem. I don’t get anywhere near it.

My weekly allowance had just reset, and after roughly two hours of work on a real development project, 93% of the entire weekly allowance was already gone.

That means I have only 7% left for the rest of the week after a single two-hour development session.

The problem is not that difficult tasks use more compute. That’s understandable. The problem is that there is no meaningful way to know beforehand whether a task will consume 2%, 10%, 30%, or most of the entire weekly allowance.

At minimum, OpenAI should show:

  • an estimated usage cost before a task starts

  • the actual usage cost after it finishes

  • exactly how much was deducted from the weekly allowance

  • a warning before a task is likely to consume a significant percentage of the weekly limit

  • a usage history showing which tasks caused the biggest consumption

If Codex is meant to be used for serious development work, weekly capacity needs to be predictable.

A weekly allowance that can be 93% exhausted within roughly two hours, without the user knowing beforehand what individual tasks will cost, is impossible to plan around.

This is exactly the opposite of what I am experiencing.

The 5-hour limit is basically irrelevant in my case because I don’t even get close to it.

My weekly allowance had just reset, and after roughly two hours of working on a real development project, 93% of my entire weekly allowance was already gone.

So I am now sitting at 7% remaining for the rest of the week after about two hours of work.

I completely understand that complex tasks can consume significantly more compute. That’s not my complaint.

What I cannot understand is how users are supposed to plan serious development work when there is no indication beforehand whether a task is going to consume 2%, 10%, 30%, or a huge part of the entire weekly allowance.

And seeing other users reporting that something seems to have changed since September 4 makes me wonder whether there has actually been a change in how Codex usage is being calculated.

I’m stuck developing with Astra because of the Plus plan’s 5-hour limit. Even with my tasks already broken down, just one 15-minute command wipes out the whole 5-hour window. I think we need a totally different workflow to make this work.

Astra’s Token Burn Is Unsustainable

I’ve been testing GPT-6 Astra in Codex against the type of long-running repository work I normally perform with GPT-5.6 Sol, and the difference in usage consumption is extreme enough that Astra is currently impractical for my workflow.

Test Configuration / Observed Results

• Model: GPT-6 Astra
• Reasoning effort: Medium
• Workload: Standard software-development task in an existing project repository
• Concurrent Codex activity: None
• Runtime: Approximately 40 minutes
• Files affected: 12 Python files
• Code created or modified: 758 lines
• Code deleted: 46 lines
• Weekly allowance before run: Approximately 87% remaining
• Weekly allowance after run: Approximately 47% remaining
• Weekly allowance consumed: Approximately 40 percentage points during a single run

I ran Astra against one of my standard project repositories. Before starting the task, my Codex weekly allowance showed approximately 87% remaining. Nothing else was running against the allowance.

The Astra task ran for approximately 40 minutes.

During that run, Astra modified or created 758 lines of code across 12 Python files and deleted 46 lines of code.

When the task completed, my weekly Codex allowance had fallen from approximately 87% to 47%.

That means a single approximately 40-minute Astra run consumed roughly 40 percentage points of my entire weekly Codex allowance.

The task produced useful work. The issue isn’t whether Astra can perform the work. The issue is the amount of weekly capacity required to produce it.

For comparison, I routinely perform this same general class of repository and software-engineering work using GPT-5.6 Sol. Sol can work continuously for extended periods and typically moves my weekly usage meter by roughly 2 to 3 percentage point per hour. It may take somewhat longer to complete a difficult task, but it gets the work done without making sustained use of Codex impractical.

I don’t question that Astra is a more capable model according to OpenAI’s evaluations. The relevant metric for my workflow, however, is useful completed work per unit of available capacity.

Astra produced approximately 758 new or modified lines across 12 Python files, with 46 lines deleted, in roughly 40 minutes. That’s certainly productive. But consuming approximately 40% of a weekly allowance to accomplish it makes that productivity difficult to use in practice.

At that rate, I could potentially perform only a few hours of serious Astra work before exhausting a resource that takes a week to replenish.

I initially wondered whether my experience was unusual, so I searched for other reports. It isn’t isolated.

Other Codex users are reporting very rapid Astra allowance depletion, including users exhausting available capacity within approximately 20 minutes, consuming substantial portions of Pro allowances in a single session, and switching back to GPT-5.6 Sol because Astra was consuming their allowance too quickly. There are now multiple reports in the OpenAI Codex GitHub issue tracker concerning Astra usage limits, quota depletion, and discrepancies between apparent token activity and allowance consumption.

OpenAI’s current documentation also explicitly states that Astra can consume included Work and Codex allowance faster than GPT-5.6 Sol.

What remains unclear is why the difference can be this large.

Is Astra intentionally weighted much more heavily against the weekly Codex allowance? Is additional reasoning or server-side agent activity being counted that isn’t readily visible to users? Is cached context accounted for differently? Or is there currently an issue with Astra usage metering during the rollout?

Whatever the mechanism, users need much better visibility into it.

Ideally, Codex should expose per-task and per-model allowance attribution so that we can see exactly how much weekly capacity a particular run consumed and why. Right now, the practical way to determine this is to record the usage percentage before a run, execute the task, and compare the meter afterward.

For the moment, I’ve switched back to GPT-5.6 Sol for production work.

Astra may be substantially more capable, but at the consumption rate I observed, I simply cannot justify using it. A model that consumes approximately 40% of my weekly allowance during a single 40-minute repository task prevents me from doing the sustained development work I use Codex for in the first place.

I’d be particularly interested in hearing from other developers who have compared Astra directly with GPT-5.6 Sol on sustained repository work. Are you seeing a similar difference in weekly allowance consumption?

References:

Related GitHub Issues in the OpenAI Codex Repository

GitHub Issue #43029 — GPT-6 Astra Ultra consumes approximately 30% of a Pro 20x weekly quota during a roughly one-hour task.

GitHub Issue #43201 — GPT-6 Astra usage limits consuming rapidly / short session length. User reports exhausting available capacity in approximately 20 minutes.

GitHub Issue #43222 — GPT-6 Astra weekly quota depletion appears disproportionate to local token telemetry in a same-day Astra-to-Sol comparison.

GitHub Issue #43141 — Astra usage limit reached again in under an hour after applying a usage reset.

GitHub Issue #43230 — Astra token burn drastically increased; Pro 20x user reports approximately 20% depletion shortly after a reset and a substantial difference compared with Sol.

No Mac, no fast mode.

The point is not how to decrease usage, the point is - increased consumption made on purpose and without benefit to users.

I agree, was just trying to offer a possible solution. I apologize that my kindness offended you.

Hey guys, im having the same issue, even using 5.6 SOL or 5.5 is consuming to fast, this is getting frustrating, i mostly use GPT for excel, i got a temprary workaround using copilot in microsoft 365 suscription, at least in formula reasoning its a temporary workaround.

I’m seeing a severe regression in Codex 5-hour usage efficiency, with some tightly defined examples.

Historically, across ~100 five-hour sessions doing similar repo work, I only exhausted the allowance once or twice. I used GPT-5.6 Sol/Terra Medium for many of those sessions and typical tasks were often ~5–13 minutes.

The first abnormal run I noticed yesterday was on Astra and exhausted the 5-hour allowance. I then deliberately reverted to GPT-5.6 Sol Medium as a temporary workaround/test, because that was the configuration I had previously used without issue.

That did not solve it:

  • one local WSL GPT-5.6 Sol Medium run took ~55 minutes and used ~74–76% of a fresh 5-hour window;
  • a later very small compatibility fix used another ~12% of the full window, taking me from 18% remaining to 6%.

I’ve ruled out the obvious local changes: local WSL not cloud, no relevant AGENTS.md, no external MCP servers, no explicit Fast service-tier config.

The large recovery run was itself only necessary because an earlier task had already hit the limit mid-run, so it was a consequence of the problem, not the cause.

Practical impact: Codex has gone from excellent for continuous development to effectively unusable for the same kind of work, because even ordinary tasks now consume such a large fraction of the 5-hour allowance that I repeatedly have to stop and wait for reset.

I’m not claiming something changed exactly on 7 Sept — that is simply when I first observed the abnormal behaviour. The previous time I used Codex it was behaving normally.

Same experience here. I eventually gave up and moved my development work over to Claude.

And to be honest, working there has been much more relaxed. I can actually focus on the project instead of constantly watching a usage meter and wondering whether the next task will kill my session. So far, despite doing a lot of real development work, I haven’t even come close to exhausting a 5-hour limit there.

That’s what makes the current Codex situation so frustrating. I really like Codex. The quality and speed are not my problem at all. But if roughly two hours of normal project work can consume almost an entire weekly allowance, I simply can’t use it for continuous development.

I’m genuinely disappointed with OpenAI here. I would much rather keep using Codex alongside my other tools, but right now the limits are pushing me away from it.

Hopefully this gets fixed, because telling users to reduce usage isn’t really a solution when the same kind of work was perfectly possible before.

Just posting to hopefully get OpenAI’s attention to this issue. I really like ChatGPT web and Codex and have been using both extensively.

Since Astra rolled out, however, I’ve been exhausting my 5-hour Codex limit in mere minutes, despite continuing to use Sol as before. This never used to happen.

I don’t want to switch to Claude, but if this is the new normal, I’ll definitely have to look at alternatives.