Issues with how codex calls itself

I asked codex to launch a codex review in max reasoning prior to landing a PR. Over the past 16 hours it has attempted a few reviews but so far I have gotten billed >$1000 without a successful completion.

Codex with GPT 5.6 Sol is arguably the most powerful model, at least for the types of work I am doing. However the harness shows serious signs of immaturity, as codex is unable to successfully spin up sub agents or delegate work as would be commonplace with other harnesses.

This is the kind of information codex is giving me:

• The second reviewer was not returning a stable verdict; it had stalled at

  the account’s long-window limit after streaming an unrecoverable oversized

  trace, so I stopped it cleanly. I’m now doing the bounded final audit myself

  at Max reasoning, focused on the two repaired failure classes and the

  receipt boundary’s completeness; I won’t spend another multi-hour cycle

  repeating already-green coverage.

and

  What failed here was the wrapper I used: it accumulated a very large event

  stream, the orchestration layer truncated the result, and the reviewer then

  sat alive without yielding another usable checkpoint. So I could see that

  the process was alive, but not recover a reliable final verdict from that

  invocation. That is not an acceptable operating mode for paid review.



  If we ever use an external Codex reviewer again, I’ll run it with:



  - event output captured incrementally to a bounded log;

  - periodic progress/tail checks;

  - a hard time/token ceiling;

  - explicit artifact paths for the final verdict;

  - termination when it produces no meaningful progress.



  I still cannot inspect its hidden chain of thought, but I can make the work,

  tool calls, elapsed time, output, and final deliverable observable. For now

  I am not launching another reviewer.

its final proposal is that it should not be allowed to launch Codex CLI:

• I agree: the safe standing rule should be no nested Codex

  processes unless you explicitly authorize a specific, bounded run.



  I do know a better invocation now—stream JSONL with --json, persist the

  final answer separately with --output-last-message, tail progress, and

  enforce a wall-clock timeout. But that fixes observability, not cost

  control. The installed CLI exposes no hard token or dollar ceiling, so a

  process can still consume significant credits before a timeout. That is not

  a sufficient safeguard.



  My recommendation:



  - I do not launch Codex CLI, Claude CLI, or another paid model process on my

    own.



  - I use this active Max session, local source inspection, mutation tests,

    Buck verification, and governance checks.



  - If genuinely independent model review would materially change a high-risk

    decision, I stop and tell you exactly why.



  - You may then authorize one bounded run, initiate it yourself, or decline

    it.



  - “Reviewer stalled” is an immediate termination condition; no multi-hour

    grace period.

Agreed that a $1,000 review with no usable result is an unacceptable failure mode. But the failure described here isn’t really that Codex couldn’t spawn a reviewer. It did, then the verdict had to travel back through a conversational event stream that grew too large and was truncated. Max reasoning can’t fix that boundary. The structural fix is for each bounded review unit to write its verdict to a known path with a fixed schema, and for the orchestrator to read that artifact instead of treating the stream as the result, so a kill loses one small unit rather than the entire 16-hour run.

Sure. Totally agree. This has nothing to do with reasoning mode. My point is codex should be able to automatically do what you describe.

Codex has a native subagent workflow now: https://learn.chatgpt.com/docs/agent-configuration/subagents. Sorry the review didn’t produce a usable result. Please share the session IDs, client version and usage details privately with Support so it can be investigated. We’ll pass on the request for stronger bounds and reliable review results.