Reproducible context loss after crashes and transport failures in recent Codex desktop releases

I am reporting a reproducible context-loss and recovery problem that I have encountered across several recent Codex/ChatGPT desktop releases.

After a long-running task encounters HTTP 429/rate limiting, a transient network or transport failure, or a desktop application crash, reopening the conversation in a new window can show a recent sidebar timestamp while restoring an older or incomplete context. Continuing from that state may repeat steps that had already completed.

Environment

  • Platform: Apple Silicon (ARM64)
  • OS: macOS 26.2 (build 25C56)
  • Application builds observed: 7119, 7345, and 7377
  • Observation window: 2026-08-27 through 2026-09-01 (UTC)

Evidence

  • Raw append-only session logs remained present and parseable, including their latest tail records.
  • For multiple sessions, the local thread-history/UI projection was materially behind the raw log.
  • Several turns ended as interrupted without a useful error being surfaced in the UI.
  • Six desktop crash reports were collected: five EXC_BREAKPOINT/SIGTRAP incidents and one EXC_BAD_ACCESS/SIGSEGV incident.
  • Repeated stack signatures included ares_llist_replace_destructor, reading_mode$cxxbridge1$194$parse_distilled_html, and V8/cppgc/Node frames.
  • A separate white-screen incident (corresponding to one of these six reports, not an additional crash) contained a V8 brk #0 fatal path with nearby OOM error in V8: / OOM detail: literals.
  • Two independent read-only recovery audits also ended with stream disconnected / transport decoding errors.
  • In the white-screen incident, one turn ended as turn_aborted at 2026-09-01T03:16:38Z, and a continuation completed at 2026-09-01T03:42:17Z. The raw log and projection eventually caught up in that case, showing that apparent context loss can occur even when raw data survives.
  • A late command-completion event from the old turn arrived at 2026-09-01T03:18:06Z, after the continuation had already started. This may indicate an ordering or projection race across turn boundaries.

Current diagnosis

The evidence strongly suggests that the append-only session log and the local thread-history/UI projection can become desynchronized after a desktop renderer or transport failure. The raw data may still be present while the UI restores stale or incomplete state. I would like OpenAI to confirm whether this is a known product issue and provide an official recovery procedure.

Questions

  1. Is there a supported way to rebuild or replay the local thread-history projection from an intact raw session log?
  2. How are interrupted turns and goal checkpoints persisted across desktop crashes?
  3. Can the application expose a durable recovery checkpoint for long-running tasks?
  4. Can completed steps be protected by an idempotency mechanism so recovery cannot repeat them?
  5. What is the safest supported procedure when raw session data exists but the UI projection is stale?

Unless there is a reliable recovery path and a concrete mitigation or fix before the next billing cycle, I may not renew my current 20x subscription.

I can provide redacted UTC timestamps, crash signatures, and small log excerpts privately. I will not publish raw session files, identifiers, or conversation content.

Environment:

  • macOS

  • ChatGPT Desktop 26.831.21537

  • Embedded codex-cli 0.152.1

  • Standalone codex-cli 0.152.1

  • Both use the same $HOME/.codex

Reproduction:

1. Open thread 01a0381b-9216-7273-8b5a-87bf967db864 in ChatGPT Desktop.

2. Switch to another conversation.

3. Reload Desktop or leave the original thread visually inactive.

4. Run:

 codex resume 01a0381b-9216-7273-8b5a-87bf967db864

Actual:

thread/resume failed during TUI bootstrap:

thread 01a0381b-9216-7273-8b5a-87bf967db864 already has an active writer

(code -32600)

Local evidence:

  • Desktop embedded app-server PID 34092 holds the thread lock.

  • The same process holds the rollout JSONL file.

  • Independent nonblocking flock probe returns EAGAIN / Resource temporarily unavailable.

  • Reload starts a new embedded app-server that automatically resumes the same thread and reacquires the lock.

  • Desktop logs also repeatedly report “Conversation state not found” for avatarOverlay after successful resume.

Expected:

  • Do not allow concurrent writers, but provide a deterministic handoff.

  • Show the owning client/process or a clear “release ownership” action.

  • On route switch, either release the writer immediately or clearly indicate that the thread remains loaded for the grace period.

  • TUI should distinguish “active writer elsewhere” from generic unavailable state and offer retry/backoff or handoff.

  • Frontend overlay state should not drop valid thread updates after successful resume.