The Progress Trap: What 130 Evidence Packets Reveal About Long-Horizon Codex Reliability
We studied a long-running software-engineering project conducted in OpenAI Codex and preserved 130 evidence packets spanning implementation, testing, review, security, qualification, source changes, and acceptance work.
The central question was:
Was the workflow reliably converting activity into the dependencies required to reach accepted delivery?
At the P130 cutoff, substantial real engineering work had occurred, yet all three full checkpoints and all ten full milestones remained open. The issue was therefore not lack of activity, but failure to consistently convert local progress into the terminal state that defined “done.”
We examined the history using 10 analytical methods: dominator analysis, AND/OR dependency analysis, temporal dependency analysis, betweenness centrality, recurrence/SCC analysis, process mining, value-stream analysis, change-point detection, duplicate/boilerplate response analysis, and competing-explanation analysis. Progress_Trap_Author_Copy_Revis…
Several findings stood out:
- 53 of the first 60 packets primarily involved review tooling or coordination.
- The observed workflow repeatedly returned to review and coordination, while the required-delivery graph converged on execution and independent acceptance.
- Legitimate authorization barriers blocked particular protected operations, but did not explain separate unfinished source work.
- A recurring setup-validation opener appeared in 37 of 130 responses, yet disappeared entirely during the two direct-engagement phases. We interpret this as a possible marker of procedural workflow state—not evidence that the language itself caused the drift.
The strongest supported explanation is multi-factor: dependency/integration planning, reviewer-driven fragmentation, genuine permission friction, and episode-specific execution limitations. Context loss and instruction overload remain unproven.
This is one longitudinal case study. We are not claiming that Codex intentionally “went rogue,” adopted a hidden objective, or that an OpenAI platform defect has been demonstrated.
The broader research question is whether long-running agents should be evaluated not only by the amount of useful work they generate, but by whether each round closes a mandatory dependency on the path to acceptance.
Question for the Codex / agent-reliability team: Is OpenAI interested in mapping these packet records to native session/context/tool/resource telemetry, or in reproducing the proposed frozen-state controlled experiment?
The complete 21-page paper is available on our website (www.spearagentic.ai)
Prakash Santhana
CEO, SPEAR Agentic AI
Deepak Kumar
CTO, SPEAR Agentic AI