Scaling Multi-Agent Coding: Keep the intelligence, Reduce the Orchestration

Executive summary

I am currently running a large software milestone through Claude’s UltraCode orchestration.

The current configuration is intentionally heavy:

  • Opus 5.5

  • high reasoning effort

  • specialist reviewers across security, payments, concurrency, UX, reliability, etc.

  • independent finding verification

  • fixers

  • fix verification

  • regression gates

The workflow is finding real P0 defects, so the additional reasoning clearly has value.

But I have hit a different scaling problem:

The size of the engineering task and the size of the orchestration graph are beginning to multiply each other.

My conclusion so far is not “use weaker models.”

It is:

Keep the brain. Reduce the bureaucracy.

What I observed

At one checkpoint, only 8 agents had started, but their conversations already contained more than 1,200 responses.

Some individual specialist reviews had produced:

  • security / instance boundary: 208 responses

  • concurrency and failure: 186

  • payments and audit: 163

  • UX: 232

  • backup/restore: 177 and still running

And this was before most finding checks, fixers, fix validation and possible second-fix rounds.

This matters because every new agent has to reconstruct some portion of:

requirements → architecture → implementation → invariants → tests → evidence

As the repository grows, that working set grows.

The approximate scaling model I now have in mind is:

Working-set size × agent fan-out × reasoning depth × verification depth

None of these variables is individually bad. Their combination can become inefficient.

Why I am not reducing reasoning quality yet

The current workflow has caught serious defects, including identity/security-boundary issues.

Those are exactly the classes of problems where I want strong independent reasoning.

I also have another useful data point: a separate narrow inventory task running on a cheaper Sonnet model still consumed roughly 360K tokens.

That suggests model choice is only part of the problem.

Working-set management and orchestration topology may matter more.

The pivot

The current run will finish unchanged. I want a clean maximum-rigour baseline.

For the next milestones, I plan to move toward risk-adaptive orchestration:

Direct execution
The parent agent handles bounded, well-specified work without mandatory delegation.

Focused delegation
Specialists receive narrow questions such as:

“Attempt to violate invariant X across this authorization boundary.”

rather than:

“Review this entire subsystem.”

Adversarial assurance
Multiple independent high-reasoning agents remain for security, identity, money, concurrency, destructive operations, data isolation and final release assurance.

I also expect to simplify verification topology.

A P0 may justify:

reviewer → independent reproduction → fixer → independent validation

A normal P1 may only need:

finding → fixer reproduces → fix → targeted verification

Similarly, agents should normally run targeted/subsystem tests. Full regression belongs primarily at the milestone gate unless blast radius is genuinely uncertain.

What I will measure

At the end of this run I want evidence, not intuition.

We are capturing:

  • tokens and runtime per agent/stage

  • unique defects versus duplicate discoveries

  • tokens per confirmed defect

  • verifier reversal rate

  • fix rejection rate by severity

  • value of second-fix rounds

  • targeted-test versus full-suite cost

  • marginal defect yield of each additional assurance layer

The most interesting metric is conceptually:

Marginal Assurance Yield = additional useful engineering information / additional inference

The objective is not simply lower token usage.

It is identifying the point where another independent reasoning process stops materially improving engineering confidence.

Current plan

M3: maximum-rigour baseline
M4–M5: risk-adaptive orchestration
M6: deliberately heavy adversarial review against a frozen candidate

The question I started with was:

“Is the orchestration mode too heavy?”

The question I now find more interesting is:

How should an agentic coding system determine that another reviewer, verifier or reasoning pass is unlikely to materially change the outcome?

I suspect solving that well will matter as much as making the underlying models smarter.

fancy words for my whole code base and the tools available was being communicated to AI everytime on every prompt and i stopped that.

Fair point, and context duplication is certainly one component. But TBHm What surprised me is that I don’t think it explains the whole behaviour here.

The larger amplification appears to be orchestration topology: UC creates multiple task owners, reviewers, finding verifiers and fixers, many of which independently reconstruct overlapping working sets. Their outputs then flow back into the parent, which itself eventually has to compact.

So I think there are two different optimizations: reduce context passed to each worker, and reduce workers/verification layers whose marginal assurance value is low.

I still have the UC task running on my machine and testing my hopes and patience. I would love to capture enough telemetry data so that I can quantify both rather than guess.
Cheers!!