Executive summary
I am currently running a large software milestone through Claude’s UltraCode orchestration.
The current configuration is intentionally heavy:
-
Opus 5.5
-
high reasoning effort
-
specialist reviewers across security, payments, concurrency, UX, reliability, etc.
-
independent finding verification
-
fixers
-
fix verification
-
regression gates
The workflow is finding real P0 defects, so the additional reasoning clearly has value.
But I have hit a different scaling problem:
The size of the engineering task and the size of the orchestration graph are beginning to multiply each other.
My conclusion so far is not “use weaker models.”
It is:
Keep the brain. Reduce the bureaucracy.
What I observed
At one checkpoint, only 8 agents had started, but their conversations already contained more than 1,200 responses.
Some individual specialist reviews had produced:
-
security / instance boundary: 208 responses
-
concurrency and failure: 186
-
payments and audit: 163
-
UX: 232
-
backup/restore: 177 and still running
And this was before most finding checks, fixers, fix validation and possible second-fix rounds.
This matters because every new agent has to reconstruct some portion of:
requirements → architecture → implementation → invariants → tests → evidence
As the repository grows, that working set grows.
The approximate scaling model I now have in mind is:
Working-set size × agent fan-out × reasoning depth × verification depth
None of these variables is individually bad. Their combination can become inefficient.
Why I am not reducing reasoning quality yet
The current workflow has caught serious defects, including identity/security-boundary issues.
Those are exactly the classes of problems where I want strong independent reasoning.
I also have another useful data point: a separate narrow inventory task running on a cheaper Sonnet model still consumed roughly 360K tokens.
That suggests model choice is only part of the problem.
Working-set management and orchestration topology may matter more.
The pivot
The current run will finish unchanged. I want a clean maximum-rigour baseline.
For the next milestones, I plan to move toward risk-adaptive orchestration:
Direct execution
The parent agent handles bounded, well-specified work without mandatory delegation.
Focused delegation
Specialists receive narrow questions such as:
“Attempt to violate invariant X across this authorization boundary.”
rather than:
“Review this entire subsystem.”
Adversarial assurance
Multiple independent high-reasoning agents remain for security, identity, money, concurrency, destructive operations, data isolation and final release assurance.
I also expect to simplify verification topology.
A P0 may justify:
reviewer → independent reproduction → fixer → independent validation
A normal P1 may only need:
finding → fixer reproduces → fix → targeted verification
Similarly, agents should normally run targeted/subsystem tests. Full regression belongs primarily at the milestone gate unless blast radius is genuinely uncertain.
What I will measure
At the end of this run I want evidence, not intuition.
We are capturing:
-
tokens and runtime per agent/stage
-
unique defects versus duplicate discoveries
-
tokens per confirmed defect
-
verifier reversal rate
-
fix rejection rate by severity
-
value of second-fix rounds
-
targeted-test versus full-suite cost
-
marginal defect yield of each additional assurance layer
The most interesting metric is conceptually:
Marginal Assurance Yield = additional useful engineering information / additional inference
The objective is not simply lower token usage.
It is identifying the point where another independent reasoning process stops materially improving engineering confidence.
Current plan
M3: maximum-rigour baseline
M4–M5: risk-adaptive orchestration
M6: deliberately heavy adversarial review against a frozen candidate
The question I started with was:
“Is the orchestration mode too heavy?”
The question I now find more interesting is:
How should an agentic coding system determine that another reviewer, verifier or reasoning pass is unlikely to materially change the outcome?
I suspect solving that well will matter as much as making the underlying models smarter.