caas.internal.errors.ClientError — Python/container + uploaded files intermittently unavailable for 4+ days

I am a paid ChatGPT subscriber for more than one year, and this issue has now blocked my professional workflow for 4+ consecutive days.

Failure Uploaded ZIP/CSV/source files may appear normally in ChatGPT, but Python/container execution either cannot access them or fails immediately with:

ClientError

Encountered exception: <class ‘caas.internal.errors.ClientError’>

In some sessions even trivial Python execution fails before any file-processing logic runs.

What I have reproduced

Project chats fail intermittently.

Non-Project chats have also started failing intermittently.

The exact same files have worked in other runtime sessions/accounts.

In one working session, 12/12 ZIP archives successfully opened and passed CRC/integrity checks, so the archives themselves are not corrupt.

The issue can temporarily recover and then fail again.

Different devices/browser/app testing did not resolve it.

OpenAI Support case #14187763 already has HAR data, timestamps, screenshots, cross-device/cross-account testing and other diagnostics. Support has stated that no additional user-side troubleshooting is required.

Possible engineering checks — hypotheses, not claimed root cause

Force provision a completely fresh execution sandbox/container.

Invalidate stale runtime/session and attachment-materialization/mount state.

Re-route affected sessions/accounts to a fresh execution worker/pool.

Inspect the internal request immediately preceding caas.internal.errors.ClientError, especially failures occurring before Python/container startup.

Verify execution provisioning/entitlement against a known-good account/session.

If possible, recreate affected Project runtime state while preserving Project data.

Validate any fix with multiple consecutive Python + ZIP/file-read tests, because a single successful attempt does not prove recovery.

This is not a minor upload inconvenience. My work depends on Python, large trading datasets, ZIP/CSV processing and source-code modification. Without the execution environment, my production workflow cannot proceed.

Other users appear to be reporting the same /mnt/data / attachment-materialization / runtime behavior. If you are still affected, please add whether Python itself fails, whether regular chats vs Projects differ, and whether the problem temporarily recovers and then returns.

Same here! I´m trying to map an ECU Flash file for 3 days by now, sometimes gpt can´t access the files, so I upload it to Google Drive as a workaround, but gpt can´t use the tools for local inspection. The problem started about 3 days while I was doing a remap in a 2000s Golf. It takes sooo long to show the error message, it just keeps thinking for like 25 minutes straight and tells me it couldn’t update the files.

Thanks for confirming. I’ve been experiencing this for about a week
now, and it’s blocking my ZIP/CSV processing and source-code work too.

One difference I’ve noticed: regular Chat keeps failing, while in Work
I’ve been able to continue some tasks using Google Drive and a
separate Sprite environment for extraction. Direct ZIP access is still
unreliable, so this is a workaround rather than a full recovery.

On my account, regular Chat uses GPT-5.6 Sol, while I’m using Astra in
Work. Given the timing and this difference, I’m wondering whether the
model or the different tool/runtime setup is contributing to the
problem. I’d like OpenAI to investigate that comparison.

Which model and mode are you using? Have you tried the same file in
both regular Chat and Work? It would help to know whether you see the
same difference.

I tried multiple models and multiple chats. When the error occur, none of the chats can access the files, it can read and compare the SHA, but can´t do local reverse engineering. The files that I used are programmed in Assembly for the automotive industry, so it´s too heavy, work can access and make progress, but tokens are burned in minutes because of the complexity in disassembling the Tricore Processors Flash Files. Everything is taking soo long too. Sometimes it just take half an hour stucked and when it comes back to normal, it answer me with just a few progress, like if the chat was stuck in one of the steps. I spent more than 3 days in one file.
I think one of the problems is misuse of AI by lots of people overloading the servers with slop and causing a lot of trouble. Everyone is just ´´Testing´´ the newer models with tons of nonsense.
Now for example: I just asked it to update a file in google drive and it´s taking more than 20 minutes, it´s just stuck…

Yes, I’m experiencing the same problem. I’m an algorithmic trader, and
my work involves mining trading setups from large datasets. The
available usage in Work gets exhausted before these mining operations
are complete—it covers only a small fraction of the workload I need to
process.

That is also why I normally don’t use Work mode for this. I’m using it
out of necessity because my regular Chat workflow has been unreliable
for about a week now. Being able to make some progress in Work doesn’t
resolve the problem when the usage limit prevents me from completing
the actual operation.

For this kind of data-heavy trading research, the repeated
interruptions are particularly disruptive. I need reliable file access
and execution to finish and verify the analysis. OpenAI needs to
address the regular Chat failures as well as the fact that Work is not
a practical replacement for these larger workloads under the available
usage limits.

ChatGPT is giving me incorrect and direct answers. I uploaded my project up to 5 September 2026; it was working fine, but after 6 September, when I ask it to give me my project after the changes, it gives me the wrong answer: ‘Stop beating around the bush.’ I don’t know why.

I am from Pakistan.

Same here, this is utterly unacceptable , eve n on a fresh plus plan its failing , I demand compensation from open AI as i cannot do my work buy yet my monthly duration of gpt is still getting used and even opening a support ticket they were so unhelpful and no replies since 4 days. This is genuinely why people are switching to claude.

this is what support told me NOT EVEN A WE ARE WORKING TO FIX IT . just straight up incompetant

Hello,

Thank you for following up. I understand your frustration, especially since this issue has continued to affect the compute functionality you rely on as part of your ChatGPT Plus subscription.

I reviewed the information already provided in this case, including your testing across multiple new conversations, browsers, devices, and networks, as well as the screenshots, console information, and HAR file. I also noted your latest confirmation that the issue continues to occur in new chats.

You have already completed the applicable troubleshooting and provided the diagnostic information requested for persistent ChatGPT errors, so you do not need to repeat those steps.

At this time, I don’t have an additional supported troubleshooting step that can restore the affected compute functionality. I recognize that this leaves the issue unresolved, and I don’t want to ask you to repeat troubleshooting you’ve already completed without a different supported next step.

For reference, you can review OpenAI’s guidance for ChatGPT errors here:

I understand this is disappointing, particularly given how long the issue has affected your use of the service and the troubleshooting you have already completed.

Best,

Lavinia

OpenAI Support

I have the same problem since this week end, working in python in a project with some heavy zipped datasets. Sometimes the runtime comes back and I can work for an hour, but then—right after—I get a ClientError. It feels like it’s getting worse and worse, and the downtime is increasing. Are the people at OpenAI aware of this?

They are aware and incompetent as they get our money and don’t have to give container resources, I had opened a support ticket and they don’t even say that yes we are working to fix it , they just acknowledge the issue

Yes, exactly. I’m facing the same problem and my work has now been
blocked for about a week.

I have already sent OpenAI Support everything they asked for — HAR
files, timestamps, screenshots, file sizes, storage checks,
cross-device/account testing and multiple failing examples. They
acknowledged the problem, but I still have no confirmation that
engineering is actively fixing it and no meaningful status update.

I even had to raise another support case today because the original
case was going nowhere. The new case has now been escalated to a
support specialist, but the actual problem is still there: ZIP files
upload, yet the Python/container runtime often cannot access or open
them, so my data analysis and source-code work gets stuck.

The subscription cost itself is honestly the smallest part of the
problem for me. I use this workflow for algorithmic trading work
connected to real funds, so when ChatGPT becomes unusable for a week,
the operational impact is much bigger than the subscription fee.

I have used OpenAI for a long time and I have never had a core
functionality issue remain broken for this long without a clear
update. If this becomes the reliability level going forward, I’ll have
no choice but to move critical workflows to another service. I simply
cannot put real-money work at risk because uploaded files may or may
not be accessible on a given day.

At the very least, affected paying users deserve a clear
acknowledgement from OpenAI that the engineering team is working on it
and some transparent status updates until it is fixed.

I’m seeing the same issue. In my case, caas.internal.errors.ClientError is not limited to uploaded files: trivial Python/container commands also fail immediately, and even creating a tiny local output file fails.

I’m using ChatGPT Plus on web with Windows 11 / Chrome. The problem is currently blocking Python/container execution and file generation.

Seeing the same issue here, intermittently, for roughly two days.

What makes it particularly odd is that the failure state can change within the same chat/session. I’ve had a fresh conversation successfully initialise the execution environment, run basic shell/Python commands, and materialise Library files into the container. Shortly afterwards, tool execution started returning caas.internal.errors.ClientError.

Once it enters that state, even trivial commands such as /bin/echo, python -c "print('test')" or equivalent basic execution fail immediately, so it doesn’t appear to be workload-related or caused by a particular file/process.

Library discovery and file materialisation can also remain functional while CAAS execution is unavailable, which suggests the failure may be somewhere around sandbox provisioning, worker/session allocation, runtime routing, or attachment of the execution environment rather than the higher-level file layer.

I’ve also seen a couple of InvalidArgumentError responses mixed in with the ClientError failures, despite the same calls having worked previously.

Starting a new chat or retrying can occasionally restore execution temporarily, but it doesn’t seem deterministic or reliable. The most frustrating part is that a workflow can begin normally and then lose its execution environment part-way through.

The within-chat transition is the strongest detail here.

Your sequence is roughly:

same chat
→ execution environment works
→ shell/Python works
→ Library files materialize
→ later CAAS returns caas.internal.errors.ClientError
→ even trivial execution fails
→ Library/file layer can remain usable.

That separates this much better than a generic “Python tool is down” report.

If you naturally catch it again, the useful evidence is the transition point, not repeated failed commands:

  • timestamp of the last successful container/Python call;
  • timestamp of the first ClientError;
  • same conversation/session reference;
  • whether Library search/materialization still succeeds after execution dies;
  • exact error type, including the InvalidArgumentError variants you mentioned;
  • any non-sensitive runtime/container/session identifier already exposed by the tool;
  • whether Work on the same account is healthy at the same moment.

I would stop after one or two trivial probes once the state is established. There is no value in hammering the dead runtime.

The fact that a fresh chat can temporarily restore execution is useful, but it is not yet proof that the fault is “conversation corruption.” It could still be provisioning/allocation/routing state that happens to be refreshed by a new session.

So I would keep three layers separate:

Library/file availability != container provisioning != command execution.

Your current specimen already shows at least the first and third can diverge.

Hi!

We confirmed that some September 1–2 Python failures reported here hit a 60-second file-transfer timeout, including a basic test command. Engineering is investigating this failure pattern and I'll update here once we know more.

I m still facing the same problem even after when my plus subscription was expired same issue caas error i think there’s something wrong with my account bcz when i tried to use second account then it never showed any kind of error but on the other hand when I try it on my main account which i m use for my work it didn’t work showing the same error i have also tried temporary to mange my work with sprite but thats too much stressful and when it used sprite it start messing with architecture so there’s definitely something wrong with my account .

This backend classification is very useful because it separates the visible Python/CAAS error from the actual stage that failed.

For the September 1–2 specimens you checked, the command itself apparently was not the failing operation:

execution environment created
→ attachment/file transfer before command execution
→ transfer hits ~60s timeout
→ command never runs

That is a much stronger classification than “Python failed.”

The remaining question is whether the current September 24–25 reports in this thread are the same backend failure class.

The newer reports include cases where:

  • a fresh chat initially executes shell/Python successfully;
  • Library/file materialisation works;
  • later in the same chat CAAS starts returning ClientError / occasional InvalidArgumentError;
  • even trivial commands then fail;
  • Library/file operations can remain usable.

That could still involve a hidden pre-execution copy/attachment stage, but the visible symptom alone does not prove it is the same 60-second transfer timeout you traced for Sep 1–2.

If Engineering can classify one current specimen, the useful split would be:

  1. execution environment creation succeeded/failed;
  2. pre-command attachment/file transfer attempted or not;
  3. transfer timeout/error code;
  4. command process was or was not actually launched;
  5. whether a replacement environment repeats the same stage failure.

That would tell us whether the current reports are recurrence of the known transfer-timeout class or a newer CAAS/runtime allocation failure that only looks identical at the UI layer.

No one needs to hammer the tool to test this. One naturally failing current request with the backend stage identified would answer a lot.