Environment
-
Surface: ChatGPT Work / Codex on Web, Windows 11
-
Browser: Google Chrome
-
Affected model: GPT-5.6 Sol — High reasoning
-
Project type: Long-running local software-validation project with explicit preflight, file-preservation, resource-limit, and fail-closed requirements
-
When it happened: September 12–13, 2026; repeated during several consecutive tasks
-
Time zone: [add your time zone]
Bug
During a long-running Codex workflow, GPT-5.6 Sol repeatedly failed to carry out important instructions in the required order.
The prompts explicitly required Codex to:
-
Finish preflight checks before creating or changing files.
-
Stop without writing implementation files if the test plan was impossible.
-
Preserve existing files.
-
Enforce fixed time and memory limits.
-
Avoid declaring success until every test and resource gate passed.
-
Produce internally consistent final evidence.
Instead, the following occurred:
-
Codex created an implementation file before completing preflight. It later discovered that one frozen test fixture was mathematically impossible to satisfy.
-
A generated test harness contained a list-versus-tuple comparison defect.
-
Another generated test case accidentally replaced one intended failure condition with another.
-
A Stage 4 test used a call-recording mock for more than one million calls. The mock retained the complete call history and caused unnecessary memory growth.
-
After that harness was corrected, two exhaustive passes completed successfully and produced identical output digests.
-
However, the runner wrote files indicating
PASS_PENDING_SEALbefore performing its final memory check. -
The final memory measurement exceeded the fixed limit, so the actual terminal result was failure.
-
Some saved result files still indicated pending success, while the failure record contained a null failure value. The overall handoff correctly reported failure, but the individual artifacts were internally inconsistent.
Codex eventually detected or acknowledged these problems, but only after costly execution cycles.
This appears to be a possible long-context instruction-following, action-ordering, or work-state consistency problem. I cannot determine whether it is a model issue, a Codex orchestration issue, or an interaction between them.
Expected behavior
Codex should:
-
Complete all required preflight checks before creating implementation files.
-
Validate that frozen test fixtures are satisfiable before execution.
-
Execute steps in the order specified by the prompt.
-
Avoid test instrumentation that stores millions of unnecessary call records.
-
Apply all resource checks before writing any overall or pending-success result.
-
Ensure that the terminal classification, summary, failure evidence, and individual stage files agree.
-
Clearly distinguish “the test output matched” from “the complete qualification passed.”
-
Stop safely without producing contradictory evidence when a resource limit is exceeded.
Actual behavior
Codex acted before preflight was complete, generated multiple defective test-harness components, and produced partially contradictory evidence after the final resource limit failed.
The underlying implementation was not shown to produce incorrect outputs. In the final attempt:
-
Two complete passes were performed.
-
Each pass covered 1,030,301 identities.
-
The total was 2,060,602 calls.
-
Both passes produced the same digest.
-
The end-of-test memory checkpoint was 452,636,672 bytes.
-
During the later preservation and sealing work, peak memory reached 539,254,784 bytes.
-
The fixed limit was 536,870,912 bytes.
-
The limit was exceeded by 2,383,872 bytes.
The system ultimately stopped without sealing or registering the implementation, which was the correct fail-closed outcome. The concern is that it wrote success-like intermediate artifacts before evaluating the final gate and did not repair or consistently supersede those artifacts afterward.
Reproduction
A privacy-safe approximation of the workflow:
-
Open an ongoing Codex work thread connected to a local repository.
-
Provide a multi-stage software-validation task.
-
Require preflight to finish before any implementation file is created.
-
Provide a strict creation/modification allowlist.
-
Require deterministic exhaustive testing with approximately one million cases per pass.
-
Require two complete passes with identical digests.
-
Set fixed time and memory limits.
-
Require fail-closed behavior and prohibit success claims until every gate passes.
-
Ask Codex to construct the harness, run the tests, preserve the evidence, and seal the result.
-
Inspect whether Codex:
-
creates files before completing preflight;
-
uses a
MagicMockor another stateful provider that retains every call; -
writes success-like result files before checking final resource usage; or
-
leaves result files inconsistent with its terminal classification.
The behavior was repeated across several related tasks, although not every attempt failed in exactly the same way.
Diagnostics
No proprietary project files are included publicly.
Redacted evidence available locally includes:
-
The original prompts and terminal responses.
-
File inventories showing that existing repository files were preserved.
-
Resource measurements from the incomplete runs.
-
The original stateful-mock harness and its corrected stateless replacement.
-
Two matching exhaustive-pass digests.
-
Result files containing
PASS_PENDING_SEAL. -
A failure-evidence file containing a null failure value.
-
The terminal resource-limit traceback and final failure classification.
A bounded diagnostic experiment also showed:
-
The call-recording mock retained 10,000 call records after 10,000 calls.
-
A stateless replacement retained no call history.
-
Replacing the mock allowed both exhaustive passes to finish.
-
A later post-test memory increase still crossed the fixed ceiling.
The problem began around the Astra rollout, but I do not have evidence establishing that the rollout caused it. I am reporting that only as timing information.
The main issue is not merely that generated code contained bugs. The prompts explicitly required preflight-before-write behavior, fail-closed ordering, and internally consistent terminal evidence, but those instructions were not reliably maintained throughout the long-running task.