I made one small change to the skill and predefined the models I want to use for different types of tasks, basically adapting the routing to my own setup. So I guess I ended up making my own little v3.0 version of it.
Thanks a lot for the great advice. I’ve already started using it in practice, and it genuinely works really well.
The deeper solution is to treat Codex as a hierarchical agent runtime, not simply a system that can spawn more models: the orchestrator should first decompose the objective into independent, measurable subtasks, then perform capability-based routing—assigning each task according to reasoning difficulty, context requirements, latency, cost, and risk rather than model popularity. A cheap Luna subagent might implement a small function or generate tests, while a stronger model handles architecture or ambiguous debugging; importantly, each subagent should receive a minimal task-specific context fork, not the orchestrator’s entire reasoning history, because inherited context increases cost and can cause anchoring or propagate an incorrect assumption. Subagents should also be treated as workers, not autonomous peers: they execute a bounded assignment, produce structured evidence and artifacts, and return control to the orchestrator. The orchestrator then performs dependency resolution, conflict handling, integration, and final verification. For high-impact changes, independent verification should be performed by a different agent or deterministic test suite rather than trusting the worker that produced the result. The resulting loop is Plan → Decompose → Route → Fork minimal context → Execute → Test → Collect evidence → Review → Integrate → Verify, with explicit budgets and stop conditions. This makes multi-agent delegation economically useful: cheap models absorb high-volume predictable work, powerful models are reserved for high-uncertainty decisions, and the orchestrator spends its reasoning budget only where coordination or judgment is genuinely required. The real engineering challenge is therefore not “how many agents can we spawn?” but how to minimize total system cost and failure propagation while maximizing verified task completion—and that requires capability routing, context isolation, structured outputs, observability, permissions, rollback, and a hard verification gate before an agent’s work becomes part of the final system.
Hello there,
For what i can tell, these model/Reasoning efforts, for spawn agents is being done by skills. However today i configured config.toml, created 4 agents under /agents (explored, worker, reviewer, architect) maybe one day i add more agents. Also added instructions under Agents.md . I did not test this yet because, well… i’m currently at 0% usage . I wonder if this works better/worse than using a skill.
Something i noticed was that i was using a plugin (superpowers) and i believe this plugin, with all its skills, was burning a lot of my usage on my projects.
Thanks for sharing the tips and experiences! The current guide covers model choices and custom agents: https://learn.chatgpt.com/docs/agent-configuration/subagents. Usage depends on the whole workflow, so adding agents won't always reduce it.