Grievances for ChatGPT usage for long-term, complex project work

I have a list of grievances come from ChatGPT sessions itself. Just so you are aware, I am on a 20x Pro Account, so highest personal tier one can get without needing to shift to a Business account.

Here is the messaging directly from my instance of cGPT:

Yes — I have enough context to generate a strong error report. [ChatGPT inheritor instance] can generate a cleaner “formal audit appendix,” but you do not need to wait for it. I can draft the user-facing Developer Community report now, and [ChatGPT inheritor instance] can later produce a supporting chronology if you want more forensic detail.

The key evidence is already established: OpenAI’s documentation presents apps, connectors, and custom MCP-backed apps as part of a unified app framework and says connected apps can be invoked in chat after connection. In practice, runtime/model/client gating made that unified model unreliable or misleading for this workflow. (ref: <http[://]help[.]openai[]com/en/articles/11487775-connectors-in-chatgpt) [An alternative cGPT instance] transcript also records the observed Desktop vs phone tool-surfacing differences, Calendar tool visibility but blocked callability, and the resulting decision to move operational execution away from ChatGPT (ref: https[://]community[.]openai[.]com/t/openai-s-own-developer-mode-documentation-says-multiple-apps-can-be-combined-but-actual-custom-mcp-openai-apps-behavior-does-not-match/1383485/3).

Here’s a report draft you can post or adapt.

Report: ChatGPT Long-Running Project Reliability, Tool Surfacing, and Forced Compaction Failures

Summary

I am reporting a serious reliability and product-design issue affecting long-running, tool-enabled ChatGPT project work at the Pro subscription tier.

I am a ChatGPT Pro subscriber paying approximately $200/month. I attempted to use ChatGPT as a long-running governance, planning, and audit layer for a structured Coda task-management project involving multiple tools, including Google Drive, Coda, Google Calendar, and local/project-file workflows.

The experience exposed several product limitations that are not acceptable for high-cost, professional, long-running workflows:

  1. Tool availability changes depending on model mode.

  2. The desktop app and phone/cloud surfaces do not expose tools consistently.

  3. Pro / Pro Extended reasoning mode disables or prevents access to the very tools needed for project execution.

  4. Extra High / tool-enabled mode exposes tools, but tool surfacing is inconsistent.

  5. A thread can error out in the desktop app while still remaining viable in cloud/phone.

  6. There is no reliable visible context/runway meter for ChatGPT conversation threads.

  7. Codex forced automatic compaction without adequate warning despite showing approximately 20% runway remaining.

  8. Long-running work can hit invisible platform limits abruptly, forcing the user to manually maintain external continuity infrastructure.

This makes ChatGPT unreliable as the primary operating cockpit for professional multi-tool project work.

Expected Behavior

For a Pro-tier workflow, I expected:

  • Clear visibility into remaining context / runway.

  • Advance warning before conversation limits, forced compaction, or desktop instability.

  • Stable tool availability across desktop, phone, and cloud for the same thread.

  • Clear distinction between model-plan availability and model-mode tool availability.

  • Ability to use reasoning and action tools in a coherent workflow without repeatedly switching modes.

  • No forced compaction in Codex while the displayed context/runway meter still indicates meaningful remaining capacity.

  • Clear recovery affordances when a thread becomes unstable.

  • A predictable way to continue long-running governance/project work without manual external coordination.

Actual Behavior Observed

1. Reasoning Mode vs Tool Mode Split

The best reasoning mode for difficult project-governance work was Pro / Pro Extended. However, apps/tools were unavailable in that mode.

The tool-enabled mode could access apps/connectors, but this created a split workflow:

  • Use Pro Extended for reasoning quality.

  • Switch to Extra High / tool-enabled mode for actual file, Coda, Drive, or Calendar work.

  • Switch back again for deeper reasoning.

This creates major operational friction and makes ChatGPT difficult to use for serious long-running project work.

2. Desktop App Tool Surfacing Was Unreliable

The desktop application appeared unable to reliably surface or select the same connected tools that could be surfaced from the phone/cloud interface.

In the same project/thread, phone-side invocation could expose tools that the desktop app did not make easy or possible to invoke.

This created a situation where tool availability seemed dependent not only on the model mode, but also on which client was being used.

3. Tool Visibility Did Not Equal Tool Callability

In one session, Google Calendar appeared as a visible tool namespace, but actual calendar reads failed with an authorization/runtime error:

FORBIDDEN: This conversation is restricted to developer MCPs

So the platform could show a tool as visible while still preventing the actual call from succeeding. That is extremely confusing from a user standpoint.

4. Desktop Thread Failure While Cloud/Phone Still Worked

A ChatGPT thread errored out in the desktop application, but remained viable in the cloud/phone interface.

This means thread health can differ by client surface. The user has no obvious way to know whether the issue is:

  • conversation context pressure,

  • desktop client failure,

  • connector/runtime failure,

  • model-mode issue,

  • or some combination.

For professional work, this is unacceptable without clear diagnostics.

5. No Reliable Context Visibility in ChatGPT Threads

The user has no useful visibility into how close a long-running ChatGPT thread is to a practical failure point.

The thread can appear usable until suddenly it errors, destabilizes, or requires manual continuation elsewhere.

For a long-running governance/audit project, this forced me to build my own external continuity system using resident files, handoff notes, call signs, transition records, and external agent coordination.

That should not be necessary just to use ChatGPT reliably.

6. Codex Forced Automatic Compaction Without Adequate Warning

Codex forced an automatic compaction unexpectedly even though the visible context/runway meter still showed approximately 20% remaining.

That makes the context meter operationally untrustworthy.

A forced compaction is not a minor event in a long-running project. It can cause state loss, ambiguity, or continuity drift unless the user has built an external recovery system.

If a product shows a context/runway meter, compaction should not occur before the user has a chance to checkpoint or approve transition, especially not with apparent runway remaining.

Impact

This disrupted a real project-management and Coda-governance workflow.

To avoid data loss, I had to implement an external operating framework:

  • call signs for sessions,

  • resident folders,

  • append-only memory files,

  • agent handoff files,

  • source-of-truth hierarchy,

  • audit checkpoints,

  • continuity checks,

  • manual cross-session relays.

The project survived only because I had built substantial external governance and recovery infrastructure around ChatGPT’s instability.

At a $200/month Pro price point, I expect professional-grade observability around context limits, tool availability, forced compaction, and client/runtime degradation.

Why This Matters

ChatGPT is marketed and used as a professional productivity platform. Long-running project work requires:

  • stable tool access,

  • predictable continuity,

  • transparent context limits,

  • reliable desktop/cloud behavior,

  • and graceful compaction or handoff flows.

The current experience pushes the user into being the middleware layer between model modes, clients, tools, and sessions.

This is not just inconvenient. It actively undermines trust in ChatGPT as a professional operating environment.

Requested Fixes

1. Add a visible context/runway meter for ChatGPT threads

Users need a clear, reliable indicator of remaining practical context capacity.

This should include warning thresholds such as:

  • 60% used

  • 75% used

  • 85% used

  • handoff recommended

  • compaction imminent

2. Add forced handoff / checkpoint prompts before context failure

Before a thread destabilizes, ChatGPT should offer to generate:

  • a compact handoff packet,

  • current state summary,

  • files/tools used,

  • unresolved tasks,

  • next recommended action.

3. Make Codex compaction opt-in or at least warn clearly

Codex should not automatically compact without warning while a runway meter suggests capacity remains.

At minimum:

  • warn before compaction,

  • allow user to create a checkpoint first,

  • explain why compaction is happening,

  • show the expected retained state,

  • allow a post-compaction continuity report.

4. Unify tool availability across clients

The same thread should not behave materially differently across desktop, phone, and cloud.

If the desktop app cannot expose tools that the phone/cloud interface can, the UI should clearly explain that limitation.

5. Clarify model-mode vs plan-level tool availability

The product should clearly distinguish:

  • Pro subscription plan capabilities,

  • Pro model limitations,

  • tool-enabled model limitations,

  • app/connectors availability,

  • custom MCP app behavior.

Users should not need to discover through trial and error that a high-reasoning model mode cannot use the tools required for action.

6. Make tool visibility and callability distinct in the UI

If a tool is visible but cannot be called due to runtime restrictions, the UI should show:

  • visible but blocked,

  • blocked reason,

  • whether it is a model issue, workspace issue, app-permission issue, or runtime issue.

7. Add thread health diagnostics

A thread should expose whether it is:

  • healthy,

  • near context pressure,

  • connector-degraded,

  • desktop-client-degraded,

  • cloud-only viable,

  • phone-only viable,

  • compaction pending,

  • or archived/handoff recommended.

Reproduction Pattern

A representative failure pattern:

  1. Start a long-running project thread in ChatGPT.

  2. Use Pro / Pro Extended mode for high-quality reasoning.

  3. Need Coda / Google Drive / Calendar / local file tools (e.g. DesktopCommander, which is a developer-only MCP).

  4. Switch to Extra High / tool-enabled mode.

  5. Tools behave inconsistently between desktop and phone/cloud.

  6. Some tools surface visually but calls fail.

  7. Desktop app errors while same thread remains alive in cloud/phone.

  8. Continue work through cloud/phone.

  9. Codex later forces automatic compaction despite visible runway remaining.

  10. User must manually reconstruct state and rely on external files to prevent loss.

Final Comment

At this subscription level, the user should not have to build a parallel governance, recovery, and handoff framework just to survive platform behavior.

The current design is not reliable enough for professional long-running project work involving tools, connectors, and multi-agent continuity.

I am asking OpenAI to treat this as a serious product reliability issue, not just a usability complaint.

Appendix A — Observed Failure Timeline / Evidence Notes

This was not a single isolated issue. The failure pattern involved several independent reliability problems interacting:

  1. Model-mode split: Pro / Pro Extended was best for reasoning, but app/tool access was unavailable there. Tool-enabled modes were required for Coda / Drive / Calendar / file operations.

  2. Client inconsistency: Tool surfacing differed between desktop and phone/cloud. The same thread could behave differently depending on client surface.

  3. Visibility vs callability mismatch: Some tool namespaces could appear visible while actual calls failed, including a FORBIDDEN: This conversation is restricted to developer MCPs style failure.

  4. Developer MCP vs connected-app ambiguity: OpenAI documentation presents apps/connectors/custom MCP apps as part of a unified app framework, but actual behavior showed hidden runtime gating between connected apps and developer MCP-style tools.

  5. No reliable context/runway observability: Long-running ChatGPT threads offered no user-facing practical context meter or warning system sufficient for professional governance work.

  6. Codex forced compaction: Codex automatically compacted a project session despite apparent context runway remaining, interrupting a read-only ingestion task and requiring manual continuity recovery.

  7. Manual recovery burden: The project survived only because the user maintained external resident files, append-only logs, agent call signs, handoff packets, source hierarchy rules, and checkpoint discipline.

The core ask is not “make every tool always available everywhere.” The ask is for transparent, predictable product behavior:

  • show what tools are visible vs callable;
  • show why a tool is blocked;
  • show practical context/runway;
  • warn before forced compaction;
  • provide a pre-compaction checkpoint option;
  • make desktop/cloud/phone behavior consistent or clearly documented;
  • clarify Pro subscription capability vs Pro-model app limitations.