I’m trying to work out whether anyone has already solved a workflow problem that I suspect is better described as agent orchestration rather than simply automation.
At this point in the development workflow I already have the architecture, development strategy, roadmap, experiment methodology, decision boundaries, and human-approval points reasonably well defined.
The remaining bottleneck is mostly execution/transport.
The rough model is:
ChatGPT / reasoning layer
→ determines the next governed task or experiment
→ Codex executes it against the repository
→ files/state are modified and experiments/tests are run
→ results/artifacts come back to the reasoning layer
→ the reasoning layer reviews the evidence and decides the next step
→ repeat, or escalate to the human when an actual judgement call is required.
Right now the manual part remains unnecessarily involved in parts of that loop: transferring instructions, making sure the right files are available, starting execution, bringing experimental results/state back, etc.
I have explicitly governed state, explicit decision points, experiment records, and places where the system must stop for human judgement, and I’m trying to automate the mechanical layer around that process.
Something roughly like:
planner/reviewer → Codex → repo/files → experiment → results → planner/reviewer → next task
Ideally with:
-
the repository/files acting as durable state rather than relying on chat memory;
-
automatic artifact/result handoff;
-
traceability of what happened;
-
retries/failure handling;
-
explicit human approval gates where required;
-
the ability for the planner to generate the next task based on the resulting state.
Has anyone built something like this successfully?
In particular, I’m curious whether the best current approach is:
-
OpenAI Agents API;
-
Codex App Server;
-
Codex SDK / Codex exec;
-
MCP;
-
GitHub Actions or another external orchestrator;
-
ChatGPT Work;
-
some combination of the above;
-
or something I’m completely overlooking.
The part I’m especially unsure about is the ChatGPT side. Can the existing ChatGPT experience realistically remain the reasoning/orchestration layer, or does a reliable implementation effectively require recreating that layer programmatically using an API agent?
I’d be particularly interested in hearing from anyone who has actually implemented a loop like this, including what broke once you tried to make it unattended.



