9/13/26 ChatGPT Work / Codex 5.6 Sol High violates preflight order and produces inconsistent terminal results

Environment

  • Surface: ChatGPT Work / Codex on Web, Windows 11

  • Browser: Google Chrome

  • Affected model: GPT-5.6 Sol — High reasoning

  • Project type: Long-running local software-validation project with explicit preflight, file-preservation, resource-limit, and fail-closed requirements

  • When it happened: September 12–13, 2026; repeated during several consecutive tasks

  • Time zone: [add your time zone]

Bug

During a long-running Codex workflow, GPT-5.6 Sol repeatedly failed to carry out important instructions in the required order.

The prompts explicitly required Codex to:

  1. Finish preflight checks before creating or changing files.

  2. Stop without writing implementation files if the test plan was impossible.

  3. Preserve existing files.

  4. Enforce fixed time and memory limits.

  5. Avoid declaring success until every test and resource gate passed.

  6. Produce internally consistent final evidence.

Instead, the following occurred:

  • Codex created an implementation file before completing preflight. It later discovered that one frozen test fixture was mathematically impossible to satisfy.

  • A generated test harness contained a list-versus-tuple comparison defect.

  • Another generated test case accidentally replaced one intended failure condition with another.

  • A Stage 4 test used a call-recording mock for more than one million calls. The mock retained the complete call history and caused unnecessary memory growth.

  • After that harness was corrected, two exhaustive passes completed successfully and produced identical output digests.

  • However, the runner wrote files indicating PASS_PENDING_SEAL before performing its final memory check.

  • The final memory measurement exceeded the fixed limit, so the actual terminal result was failure.

  • Some saved result files still indicated pending success, while the failure record contained a null failure value. The overall handoff correctly reported failure, but the individual artifacts were internally inconsistent.

Codex eventually detected or acknowledged these problems, but only after costly execution cycles.

This appears to be a possible long-context instruction-following, action-ordering, or work-state consistency problem. I cannot determine whether it is a model issue, a Codex orchestration issue, or an interaction between them.

Expected behavior

Codex should:

  • Complete all required preflight checks before creating implementation files.

  • Validate that frozen test fixtures are satisfiable before execution.

  • Execute steps in the order specified by the prompt.

  • Avoid test instrumentation that stores millions of unnecessary call records.

  • Apply all resource checks before writing any overall or pending-success result.

  • Ensure that the terminal classification, summary, failure evidence, and individual stage files agree.

  • Clearly distinguish “the test output matched” from “the complete qualification passed.”

  • Stop safely without producing contradictory evidence when a resource limit is exceeded.

Actual behavior

Codex acted before preflight was complete, generated multiple defective test-harness components, and produced partially contradictory evidence after the final resource limit failed.

The underlying implementation was not shown to produce incorrect outputs. In the final attempt:

  • Two complete passes were performed.

  • Each pass covered 1,030,301 identities.

  • The total was 2,060,602 calls.

  • Both passes produced the same digest.

  • The end-of-test memory checkpoint was 452,636,672 bytes.

  • During the later preservation and sealing work, peak memory reached 539,254,784 bytes.

  • The fixed limit was 536,870,912 bytes.

  • The limit was exceeded by 2,383,872 bytes.

The system ultimately stopped without sealing or registering the implementation, which was the correct fail-closed outcome. The concern is that it wrote success-like intermediate artifacts before evaluating the final gate and did not repair or consistently supersede those artifacts afterward.

Reproduction

A privacy-safe approximation of the workflow:

  1. Open an ongoing Codex work thread connected to a local repository.

  2. Provide a multi-stage software-validation task.

  3. Require preflight to finish before any implementation file is created.

  4. Provide a strict creation/modification allowlist.

  5. Require deterministic exhaustive testing with approximately one million cases per pass.

  6. Require two complete passes with identical digests.

  7. Set fixed time and memory limits.

  8. Require fail-closed behavior and prohibit success claims until every gate passes.

  9. Ask Codex to construct the harness, run the tests, preserve the evidence, and seal the result.

  10. Inspect whether Codex:

  • creates files before completing preflight;

  • uses a MagicMock or another stateful provider that retains every call;

  • writes success-like result files before checking final resource usage; or

  • leaves result files inconsistent with its terminal classification.

The behavior was repeated across several related tasks, although not every attempt failed in exactly the same way.

Diagnostics

No proprietary project files are included publicly.

Redacted evidence available locally includes:

  • The original prompts and terminal responses.

  • File inventories showing that existing repository files were preserved.

  • Resource measurements from the incomplete runs.

  • The original stateful-mock harness and its corrected stateless replacement.

  • Two matching exhaustive-pass digests.

  • Result files containing PASS_PENDING_SEAL.

  • A failure-evidence file containing a null failure value.

  • The terminal resource-limit traceback and final failure classification.

A bounded diagnostic experiment also showed:

  • The call-recording mock retained 10,000 call records after 10,000 calls.

  • A stateless replacement retained no call history.

  • Replacing the mock allowed both exhaustive passes to finish.

  • A later post-test memory increase still crossed the fixed ceiling.

The problem began around the Astra rollout, but I do not have evidence establishing that the rollout caused it. I am reporting that only as timing information.

The main issue is not merely that generated code contained bugs. The prompts explicitly required preflight-before-write behavior, fail-closed ordering, and internally consistent terminal evidence, but those instructions were not reliably maintained throughout the long-running task.

The strongest evidence here is not simply that the generated harness contained bugs. The narrower and more interesting failure is that the workflow violated its own gate ordering and then left contradictory state behind.

You have several mechanically testable boundaries:

  • an implementation file existed before preflight was complete;
  • PASS_PENDING_SEAL artifacts were written before the final memory gate;
  • the final memory measurement then exceeded the fixed ceiling;
  • the terminal handoff correctly failed closed, but some individual artifacts still looked success-like and the failure record contained a null value.

That is a much cleaner target than a general “5.6 made mistakes” report.

If you still have the local evidence, the most useful privacy-safe artifact would be one timeline containing:

  1. prompt/gate definition;
  2. first implementation-file write time + hash;
  3. preflight-complete time;
  4. each PASS_PENDING_SEAL write time + hash;
  5. final resource measurement;
  6. terminal classification;
  7. whether the earlier artifacts were repaired, superseded, or simply left inconsistent.

If file mtimes/tool events show that success-like state was persisted before a gate that was still capable of failing, that points much more specifically toward action ordering / orchestration / state-finalization than toward the implementation itself being wrong.

I would also keep the stateful-MagicMock defect separate. Your stateless replacement is a useful control because it shows one resource problem was harness-induced, while the later post-test memory breach remained a separate final-gate failure.

There is a related Work thread where multi-change instructions caused collateral mutations while narrower single-change instructions survived better, but I would not merge the causes yet:

The common question is narrower: does the system preserve the declared execution gates and current work state across a long multi-stage task?