For about eight months, I’ve been building a fairly large, multi-stage AI-assisted system.
In another thread, I’ve been documenting a different class of problems around long-running conversations, context continuity, branching, and tasks that eventually become too large or too tightly coupled to execute reliably as a single unit:
[link to previous thread]
During the current large remediation effort, however, I ran into a different problem that seems worth discussing separately.
This is not a typical implementation bug.
It is not:
the application no longer starts after a change,
or:
we fixed A and accidentally broke B.
It is also not quite the usual form of technical debt where the code still works but gradually becomes difficult to maintain.
The problem is deeper.
A dependency error can remain invisible to tests
The project has a large number of tests and deterministic verification scripts. They are executed throughout subsequent implementation stages.
Despite that, the current audit process has started exposing errors that originated much earlier in the project.
The important part is that these errors did not necessarily cause the earlier tests to fail.
The issue appears to exist one level above:
“Does the implementation correctly satisfy the contract?”
The more dangerous question is:
“Is the contract itself correct, and are the assumptions it depends on valid?”
That distinction matters.
A test can correctly prove:
the implementation does exactly what the specification requires.
But if the specification contains a wrong assumption, then a green test may only prove that:
the wrong assumption was implemented correctly.
The most dangerous failure may be the one nobody notices
I do not yet know whether the final system would actually have failed visibly because of the specific issues we are finding. The complete product does not exist yet, so there is no final end-to-end validation beyond the current tests and verifiers.
But there is a realistic possibility that a system could:
- start normally,
- pass its tests,
- perform its expected functions,
- and still have part of its logic built on an invalid foundation.
That worries me more than a crash.
A crash tells you:
something is wrong.
A foundational error may tell you:
everything is working.
And then the next implementation inherits the same assumption.
After enough time, you no longer have one bug.
You have something closer to:
incorrect contract → dependent module → later specification → later implementation → later tests based on the same assumption
Every individual element in that chain may be locally correct.
The chain itself may still be wrong.
In my case, some of the problem comes from an older generation of the workflow
The earliest parts of this project were built using a much simpler process:
human → main design chat → Codex
At that point I did not yet have the independent audit layer, formal stage gates, dependency checks, and strict artifact/version validation that exist in the project today.
Those controls evolved later.
The interesting part is that the current remediation is now discovering problems originating from that earlier period.
So the stricter process did not create the problem.
It exposed a problem that had remained hidden.
If those assumptions had survived for several more months and more layers had been built on top of them, fixing the foundation at the end might have required revisiting a significant amount of roughly eight months of work.
What I am trying to separate now
One lesson from this remediation is that I no longer treat a passing test suite as sufficient evidence that the entire system is correct.
Tests remain essential, but I am increasingly separating several different kinds of verification:
- implementation verification — did we correctly implement what was defined?
- contract verification — is the contract internally coherent?
- dependency verification — are its assumptions compatible with the things it depends on?
- independent review — can another role/context challenge the decision rather than inherit the same assumptions?
- provenance and versioning — which exact state and version produced this decision?
- stage gates — can uncertain or inconsistent state become input to the next stage?
There is another lesson from the remediation:
Understand the problem globally, but execute the repair locally.
The dependency chain has to be understood as a whole. Fixing one node without understanding what depends on it can simply move the inconsistency somewhere else.
At the same time, the remediation itself became too large and too interconnected to execute safely as one unit.
We have already had to decompose it repeatedly into smaller, independently verifiable execution units.
That connects back to the problem I mentioned in my previous thread, but it is not quite the same problem.
Two different kinds of correctness
The distinction I keep coming back to is:
Did the system correctly do what we told it to do?
versus:
Did we tell the system to do the right thing in the first place?
The first question can often be covered extremely well by tests.
The second one is much harder.
And I think this second class of failure becomes especially dangerous in large AI-assisted projects, because a wrong assumption can be inherited by many later stages without producing an obvious failure signal.
Why I am writing this
If I were starting the same project again, I would rather spend an additional month on architecture, contracts, dependencies, documentation, and validation rules before serious implementation began.
For a large system, I increasingly think the documentation and contracts should be far ahead of the implementation rather than being generated afterward as a description of what the model just built.
The very natural workflow today is:
prompt → implementation → tests → works → next prompt → next implementation
For a small project, that can be incredibly effective.
For a large, multi-stage system, I now see a different risk:
each locally correct implementation may simply make an earlier incorrect assumption more deeply embedded.
Adding more agents does not automatically solve that problem either.
Changing the process to something like:
agent → plan → agent → implementation → tests
is better, but it can still fail if every participant inherits the same underlying assumptions.
If the first design decision defines the dependency incorrectly, several agents can simply implement that mistake much more efficiently.
That is why I am becoming increasingly convinced that complex AI-assisted development needs independent roles with deliberately conflicting objectives.
The designer’s job is to create a solution.
The auditor’s job is not to help the designer finish it. The auditor’s job is to try to prove that it should not pass.
A separate critical reviewer should challenge contracts, dependencies and assumptions rather than only reviewing code.
The executor should not be able to expand its own scope.
The human remains the final decision boundary.
And, most importantly, the workflow needs hard blockers capable of saying:
STOP. There is not enough evidence to allow the next stage to begin.
Even when the implementation looks good.
Even when the local tests are green.
That costs time.
Sometimes a lot of time.
But after what I am seeing now, one month spent establishing good contracts and independent validation may be much cheaper than discovering eight months later that dozens of later components inherited an error from the foundation.
So if someone is starting a larger AI-assisted software project today, my main warning would be:
Don’t only ask whether AI can build the next thing. First build a system that can stop AI when the next thing should not be built yet.
I’m curious whether others have encountered something similar: a system that passed its tests and looked correct, only to reveal later that the implementation was fine but an earlier specification, contract, or dependency assumption was wrong.