I asked codex to launch a codex review in max reasoning prior to landing a PR. Over the past 16 hours it has attempted a few reviews but so far I have gotten billed >$1000 without a successful completion.
Codex with GPT 5.6 Sol is arguably the most powerful model, at least for the types of work I am doing. However the harness shows serious signs of immaturity, as codex is unable to successfully spin up sub agents or delegate work as would be commonplace with other harnesses.
This is the kind of information codex is giving me:
• The second reviewer was not returning a stable verdict; it had stalled at
the account’s long-window limit after streaming an unrecoverable oversized
trace, so I stopped it cleanly. I’m now doing the bounded final audit myself
at Max reasoning, focused on the two repaired failure classes and the
receipt boundary’s completeness; I won’t spend another multi-hour cycle
repeating already-green coverage.
and
What failed here was the wrapper I used: it accumulated a very large event
stream, the orchestration layer truncated the result, and the reviewer then
sat alive without yielding another usable checkpoint. So I could see that
the process was alive, but not recover a reliable final verdict from that
invocation. That is not an acceptable operating mode for paid review.
If we ever use an external Codex reviewer again, I’ll run it with:
- event output captured incrementally to a bounded log;
- periodic progress/tail checks;
- a hard time/token ceiling;
- explicit artifact paths for the final verdict;
- termination when it produces no meaningful progress.
I still cannot inspect its hidden chain of thought, but I can make the work,
tool calls, elapsed time, output, and final deliverable observable. For now
I am not launching another reviewer.
its final proposal is that it should not be allowed to launch Codex CLI:
• I agree: the safe standing rule should be no nested Codex
processes unless you explicitly authorize a specific, bounded run.
I do know a better invocation now—stream JSONL with --json, persist the
final answer separately with --output-last-message, tail progress, and
enforce a wall-clock timeout. But that fixes observability, not cost
control. The installed CLI exposes no hard token or dollar ceiling, so a
process can still consume significant credits before a timeout. That is not
a sufficient safeguard.
My recommendation:
- I do not launch Codex CLI, Claude CLI, or another paid model process on my
own.
- I use this active Max session, local source inspection, mutation tests,
Buck verification, and governance checks.
- If genuinely independent model review would materially change a high-risk
decision, I stop and tell you exactly why.
- You may then authorize one bounded run, initiate it yourself, or decline
it.
- “Reviewer stalled” is an immediate termination condition; no multi-hour
grace period.