Observed behavior: branching near the context limit appears to extend usable conversation continuity

I’ve been testing the behavior of very long ChatGPT conversations used for actual technical work.

This is not a chat with a few dozen messages. One of my conversations acts as a long-running audit/workflow chat, processing documents, decisions, files, execution results, and project state over many days.

I noticed an interesting behavior with Branch in new chat.

My first audit chat lasted roughly a month. Near the end, I created a state/memory snapshot and used it to bootstrap a successor chat.

That second chat ran from August 10 to August 20 before reaching the point where its usable context was effectively exhausted.

Instead of starting another empty chat, I went back one message before the end and created a Branch.

I deliberately did not provide the previously generated memory snapshot to that branch. I wanted to see what the branch itself would inherit.

The branch continued working normally from August 20 to August 24, completing additional technical tasks beyond the point where the parent conversation could no longer continue usefully.

When that branch also reached its limit, I repeated the experiment:

I created another Branch from the branch, again one message before the end.

The important part is that the first new task in Branch #2 was to generate a fresh project-state/memory snapshot itself. It was not given that snapshot as input.

It successfully reconstructed the state and is still completing further technical tasks.

Current sequence:

Chat #1: ~1 month
→ Chat #2: Aug 10–20
→ Branch #1: Aug 20–24
→ Branch #2: Aug 24–present, still working

I am not claiming that branching creates an infinite context window, and I do not know what ChatGPT is doing internally with inherited/compacted context.

But I can currently separate two interesting behaviors.

1. Continuity

The new branch can continue producing multiple useful responses beyond the point where its parent conversation stopped being usable. In my test, this is not just one final response after branching.

2. Context fidelity

Continuity is not perfect fidelity.

Branch #1 eventually forgot one previously established file-output structure. A single reminder restored the expected behavior, so the broader project context was still present, but at least one detailed rule had lost salience.

I am therefore tracking these as separate variables:

  • continuity — can the conversation keep doing the work?
  • context fidelity — how much of the detailed prior state/rules remain correctly active?

One additional observation: the first responses after creating a new branch are noticeably slower for me, while later responses become faster again. I do not know whether this is related to history reconstruction, caching, compaction, frontend rendering, or something else, so I am treating it only as an observation.

The test is still running. If Branch #2 reaches the same point, I plan to branch it again and compare the lifetime and fidelity of each generation.

Has anyone tested Branch in new chat specifically near the end of a very large conversation context and seen similar behavior?

Update 1 — another branch, and the first data on actual conversation load

Since my first post, the experiment has already gone through two more full stages, and the next one has now started.

At the moment, it looks like this:

  • Base conversation: about 10 days of intensive technical work.
  • Branch 1: Aug 20, 19:36 → Aug 24, 18:38 — 3 days, 23 hours, 2 minutes.
  • Branch 2: Aug 24, 18:38 → Aug 29, 13:11 — 4 days, 18 hours, 33 minutes.
  • Branch 3: opened on August 29 and is currently continuing the work. Interestingly, continuity was preserved after branching, and the new branch has already answered two more follow-up questions correctly.

The number of calendar days turns out to be a fairly poor measure of how heavily these conversations are actually being used.

These are not chats where a few short messages are exchanged over four or ten days. They are very dense technical sessions involving the design, auditing, and development of a large system.

A single task may involve many consecutive analyses, comparisons between document versions, tracking dependencies between earlier decisions, tool use, verification of physical artifacts, hash checks, corrections, repeated audits, and comparison against the current Git state.

A useful reference point is an earlier experiment I did with long-form text.

For that test, I used one of my own manuscripts, roughly 440,000 characters long. I ran the entire manuscript through a single chat three times, chapter by chapter, performing full rounds of revision and rewriting.

That amounts to roughly 1.3 million characters of manuscript text alone, not counting the surrounding discussion, instructions, individual corrections, or alternative versions.

And a single chat handled the entire process. More importantly, that conversation still exists and is still usable — it never reached its practical limit.

The technical conversations behave very differently.

For Branch 2, which lasted about 4 days and 18.5 hours, I reviewed the work performed during that period in more detail.

Within that span, I can identify 5 major multi-stage work cycles. I am not counting every correction of the same problem as a separate task.

Together, those cycles contained roughly 18 larger analysis or revision stages.

The heaviest single cycle went through 5 major audit rounds, with each round triggered by a new concrete issue discovered in the previous version.

Other tasks involved physical packages containing roughly 30 and 93 controlled objects, including their versions, dependencies, and hashes.

One of the earlier branches had a similar example: a single task involved around 11 consecutive analyses, and at the end of that task the chat produced a package of 11 files, calculated their hashes, and compared the result against the relevant Git branch.

That was one task, not an entire day of work.

So the difference between the manuscript experiment and the technical conversations may not come down only to token count or the amount of visible text.

The structure of the work may matter as well: the number of documents, previous decisions, dependencies between artifacts, tool outputs, revision rounds, and pieces of project state that must remain relevant to later responses.

My current working hypothesis is therefore:

two conversations with a similar number of tokens may have very different “context density” and behave very differently in practice.

From now on, for each new branch, I plan to track not only how long it remains usable, but also the approximate number of major work cycles, larger analysis rounds, and the scale of the artifacts it has to manage.

So far, the sequence is:

base conversation → branch 1 → branch 2 → branch 3

More data will be needed before any real pattern can be claimed.

But there is already a second, very practical observation.

Opening a new branch may be a real way to extend the useful life of a very long conversation without abruptly losing continuity.

Instead of abandoning a conversation that has been running for weeks or months when it approaches its practical context limit, it may be possible to branch it and continue the same line of work.

In this experiment, the branches did not give me only a handful of extra messages. The first branch remained useful for almost four days of intensive work. The second lasted about 4 days and 18.5 hours. The third picked up the conversation immediately and has already continued answering follow-up questions correctly.

I would not treat this as a guarantee that every new branch will always provide a similar amount of additional life.

But after two complete consecutive cases, I think it is reasonable to say that opening a new branch can provide several more days of continuous work on the same very long conversation, rather than forcing you to start over from scratch.

If that pattern continues, this could become a very practical way to maintain long-running technical conversations over many months — not as one infinite session, but as a sequence of branches that preserve continuity.

Appreciate the extended details but a short summary would be helpful

If I understand correctly say you chatting long convo get a hard limit so you go to the reply before branch it out and can continue chatting for much longer than the original convo?

That’s good only concern is the earlier old chat contents as well as any attachments being transferred?

Yes, that’s basically it.

One important detail: I’m not branching arbitrarily.

I create the new branch from the message immediately before the conversation hits the practical context limit / becomes unusable. In other words, I branch as late as possible, one message back from the point where the original thread effectively stops being useful.

So far, that new branch has preserved enough context to continue the same technical work for several more days.

The conversation context seems to carry over surprisingly well between branches, but attachments are less consistent.

I’ve tested this already. In some cases the new branch remembers that a specific file or ZIP package existed, what it was used for, and even that it had previously audited it — but it may then say that the actual file is no longer available in its current accessible sources.

So there seems to be a difference between:

  • remembering information about a previous attachment,
  • and still having direct access to the attachment itself.

I’ve seen this with both files I uploaded and files generated during the conversation. Some remain accessible, some don’t, and I don’t yet know what determines that.

So the short version is: conversation continuity transfers much better than attachment availability does.

For anything important, I currently wouldn’t assume that an old attachment will remain directly accessible after branching. I’d keep the original files separately and reattach them if the new branch actually needs to inspect the file again.

Personally this does not affect me I just like to transfer well in advance when a certain point is reached as I break down a lot so only dealing with small issues at a time

You know what would be cool to test this before reaching an actual limit but still having like a realistically long conversation to test drifting compression hallucination losing any attachments etc from both you and ai

So like start a legitimate practical convo about something real with attachments thinking and then when it gets long branch out and then like ask the original to do a proper cross chat continuity hand off transfer

So you end up with a branch and then you start a new chat with a hand off transfer

And like a same prompt to both explaining what you are testing so they can check the difference

Just my 2c

Yes — and that’s exactly why I don’t rely on branching alone.

When a conversation reaches its practical limit, I branch from the message immediately before that point. The first thing I do in the new branch is generate a structured snapshot using a fixed template: current project state, important decisions, open issues, completed work, and the exact point from which the next session should continue.

Then I keep working normally. When that branch reaches its limit, I repeat the same process: new branch → new snapshot → continue.

So I’m not relying only on the conversation’s memory.

I also assume that at some point branching may stop being enough or I may hit a limit I haven’t seen yet. For that reason I keep separate recovery documents that can be used to start a completely new conversation and rebuild the project context.

In practice my continuity chain is:

current conversation → next branch → snapshots → recovery documents

Attachments are the weaker part. A branch may remember that a file existed and what was done with it, but it may no longer have direct access to the actual file. So I keep important source files and generated packages separately and reattach them when necessary.

For long-running projects, this feels much safer than relying on the memory of a single chat.

Update 2 — I’m ending the branching experiment

I think I now have enough data to end this part of the experiment.

After publishing Update 1, I opened two more branches in exactly the same way as before — from the message immediately before the point where the previous conversation reached its practical limit.

The full sequence now looks like this:

  • Base conversation: about 10 days of intensive technical work.

  • Branch 1: 3 days, 23 hours, 2 minutes.

  • Branch 2: 4 days, 18 hours, 33 minutes.

  • Branch 3: about 4 hours, 30 minutes.

  • Branch 4: about 2 hours, 41 minutes.

And this is where I’m stopping the creation of further branches.

The most interesting part is how sharply their behavior changed.

The first two branches still gave me full days of intensive, continuous work on the same project. Then the next branch lasted only about 4.5 hours, and the following one less than 3 hours.

At this point, that no longer looks like a random difference between two conversations.

Of course, I do not know exactly how ChatGPT implements branching internally, so I do not want to claim that I know the mechanism behind this behavior.

But from a user perspective, the result is fairly clear:

branching can significantly extend the useful life of a very long conversation, but it does not appear to behave like an unlimited context reset.

In my case, the first branches provided several additional days of work, while later branches showed a very noticeable drop in useful lifetime.

That suggests, at least in practice, the possibility of diminishing returns from successive branches.

Importantly, even the final short branches still preserved project continuity. The problem was not that the conversation suddenly “forgot everything.” It simply reached its practical limit again much faster.

This changes my original conclusion slightly in favor of a more conservative approach.

I still think branching is a very useful recovery mechanism for long-running conversations. Instead of losing a thread that has been active for weeks or months all at once, branching can buy several more days of work and give you time to prepare a safe transition.

But I would no longer treat repeated branching as a way to keep the same conversation alive indefinitely.

For my use case, the better model now looks like this:

long conversation → branch → another branch → checkpoints/snapshots → full recovery documents → new chat

And that is the final stage I am starting to test now.

Instead of opening a fifth branch, I am creating a completely new conversation and rebuilding the project state from previously prepared snapshots, continuity documents, and current artifacts.

That will be a separate test — not of how long one chain of branches can be stretched, but of how well a completely new conversation can take over a complex, multi-month project without relying on the history of the previous chat.

For me, this is the most important practical result of the experiment so far:

branching can buy additional time, but the real safety mechanism for a long-running project should still be an external, reconstructable project state — not the memory of a single conversation.

Really interesting — I went through a somewhat similar progression in my own research.

At first I was mostly working in long chats. Once continuity started becoming a problem, I moved more of the work into Projects and had the chat generate a kind of master ledger for the research.

That helped, because the ledger could sit in the Project files and different chats could use it as a shared reference. But the annoying part was that the research kept changing, so I had to keep generating a new version of the ledger and replacing the Project file manually.

Later I moved that durable state into Git, which made the whole thing much easier to maintain. The AI could work against the same external state and update it as the project evolved, instead of me constantly shuttling updated files back into the Project.

So I’m really curious about your snapshot process. How are you generating and maintaining them now? Do you use a fixed template, and how do you decide what is important enough to carry forward versus what can just stay in the conversation?

That’s really interesting, because I think we arrived at almost the same problem, just from slightly different directions.

In my case, the snapshot is not really a master project log or a copy of all project documents. I treat it more as a snapshot of the current conversation and process state.

I generate it using a fixed template. It records things such as:

  • where the project currently is,
  • what has been completed,
  • which decisions were made,
  • what is still blocking progress,
  • what the next chat should not repeat,
  • and exactly where the next session should continue from.

One important rule is that the snapshot is not the source of truth. If it ever conflicts with an actual project document, artifact, audit result, or repository state, the canonical source wins.

I keep the important files separately as well.

I started doing that because I noticed there is a real difference between:

“the chat remembers that it worked with a file”

and

“the chat still has direct access to that physical file.”

Across branches, information about attachments can carry over surprisingly well. A new branch may remember the name of a package, what it was used for, and even that it previously audited it, while at the same time being unable to locate the actual file in its currently available sources.

More recently, I have run into another problem: some ZIP packages are becoming difficult or impossible for me to attach again to new chats, even though they are relatively small and mostly contain text documents and JSON files.

I still don’t know exactly what causes this, so for now I’m treating it as a separate issue to investigate.

And this is why I really like the direction you took with Git.

For projects like these, my ideal solution would be to connect ordinary ChatGPT conversations to a persistent, version-controlled project workspace, potentially Git-based.

And ideally not read-only.

I would like to be able to tell a normal ChatGPT conversation:

“Save this generated artifact into the project repository.”

Then, after explicit user approval, ChatGPT could store it at a defined location and version.

Later, in a completely fresh conversation, I could simply say:

“Read the package stored here.”

and that new chat could retrieve the exact files from the repository itself.

That would remove the current workflow where I have to download a ZIP from one conversation, upload it into another one, and then repeat the same process again when a third conversation needs to work with it.

For a small project, this is probably just an inconvenience.

For a project that runs for months and contains a large number of documents, versions, dependencies, audits, and interacting workstreams, it becomes a real state-management problem.

So my current separation is roughly:

Chat = active reasoning and work
Snapshot = continuity between conversations
Versioned artifacts / repository = durable project state

It looks like we independently arrived at a very similar separation of responsibilities.

I’d also be very interested to hear from anyone at openai working on ChatGPT Projects, long-running workflows, or GitHub integration who happens to read this thread.

Is there currently a recommended approach for maintaining this kind of long-running, multi-chat project around one durable, version-controlled state?

And if not, is a bidirectional project workspace something OpenAI has considered — one where ordinary ChatGPT conversations could both read project artifacts and, with explicit user approval, write new or updated artifacts back into that shared workspace?

I think that could solve an entire class of problems that only really becomes visible once projects get sufficiently long-lived and dependency-heavy.

P.S. There is another reason this shared workspace problem has suddenly become much more important in my project.

I’ve just run into what appears to be a different kind of practical limit: one remediation task became so large and internally dependent that a completely fresh execution session could not complete it as a single run.

We also tried decomposing it across agents, and that still wasn’t enough.

I’m going to write that up separately as Update 3, because it seems to be a different problem from conversation longevity itself — more about the practical size of a single executable task.

Interestingly, it leads straight back to the same issue you raised: once the work has to be split across multiple fresh chats or agents, having one persistent, version-controlled project state becomes much more important.

My current workflow is to use gpt project as a workspace, then use codex Git connector to access my private git repo. The connector works fine for read/write git. I use git repo to store durable study states, for other raw materials and documents or other large size files, I use google drive. I’ve talked with gpt and decided what human do and what AI do and decided how the AI works and when it should stop to wait for human decision. We’ve made a set of rules. Now it works fine for me.

That’s very close to the direction I’m moving toward as well, and it sounds much cleaner than manually passing files between conversations.

I especially like the separation:

Git = durable/versioned project state
Google Drive = larger raw materials and files
GPT Project = active workspace
Human/AI rules = who is allowed to do what, and when the AI must stop and wait for a human decision

I’ve actually just started using Google Drive for ZIP packages and larger handoff files for the same reason — moving those manually between chats was starting to become a problem.

I have one question about the handoff between GPT Project and Codex, though.

Does the normal chat inside the GPT Project have direct access to the current repository state after Codex has changed it, or do you still have to tell the chat what changed or point it to a specific commit/file?

And in the other direction: if the normal Project chat creates a new document or changes some project state, can you simply tell Codex to persist that result into Git, or is there still a manual step between the two?

That handoff is the part I’m most interested in right now.

What I would ideally like is a closed loop like:

GPT Project → Codex → Git → GPT Project

with all of them working against the same durable state, instead of me repeatedly moving ZIP packages and updated documents between separate conversations.

If your current setup already closes that loop reliably, I’d be very interested to hear more about exactly how you structured it.

Actually, it’s include the rules we’ve created, the gpt chats have the authorities to update project states. And codex git connetor is a plug-in which you can find in ChatGPT. Human’s attention is considered an important resource and the whole system values it. So we’ve reached a conclusion:

"Human decision boundary

Human attention should concentrate on:
Goal / Why / material Judgment / Priority / Permission

Human confirmation is required for material changes to research direction/meaning, authority/release/validation promotion, public disclosure/IP disposition, major architecture/implementation commitment, destructive archive/delete batches when governance requires it, and material cross-project adoption.

Routine reversible research/tool/file maintenance may proceed within existing permissions and boundaries."

You may have a try if you’re interested, I can share my workflow with you.

Thanks — I just checked this on my side and learned something new.

It turns out my regular ChatGPT can not only read from my private GitHub repo, but also write to it. Until now, I had treated writing as a separate step handled by Codex, so this changes how I’m thinking about the workflow.

My structure is actually quite similar to yours:

main chat = project direction and decisions
independent second chat = audit and error detection
additional chats = specialized tasks
Codex = code changes

Until now, I’ve been manually moving packages, documents, and snapshots between them. Now I can see that the repository itself could probably act as the shared durable state. I’ve also started using Google Drive for larger ZIP packages and raw materials.

I’m especially interested in your Human Decision Boundary, because we arrived at a similar idea, although my governance became much heavier.

If you’re willing to share your workflow, I’d be very interested in:

  • how several chats work against the same repo state,
  • how you avoid conflicts,
  • when AI is allowed to update state on its own,
  • when it must stop and wait for a human,
  • and how your GPT Project → Codex → Git → GPT Project flow works.

One more question: do you trigger your audit/review chat manually, or have you automated part of that process, for example after a state change, commit, or Codex execution?

Thanks for pointing me in this direction — it looks like part of the problem I was trying to solve manually may already be much easier with the tools that are available.

I’m glad you’ve found my idea helpful.

For audit and review, I currently trigger them manually in chat. I think that’s still a human choice — if you’re not going to look into the result, there’s probably no need to automatically generate all that overhead.

Right now it’s mostly triggered by human judgment. If I feel a thread is getting too long and starting to lag or lose context, I’ll ask the chat to prepare a handover message. If I feel the progress in one thread could affect other parts of the project, I’ll kick off a project-level review or audit.

Normally I do a broader audit either at the end of a day’s research or at the beginning of the next session. So it’s not automated after every commit or Codex run. I prefer to trigger it when I feel there’s actually something worth reviewing.

I’ll send you a message about the workflow I’ve been experimenting with — I think it’ll be easier to explain the rest there.

Thanks, that makes a lot more sense now. In your setup, audit/review is something you trigger when you feel it’s actually worth doing, rather than after every stage.

Mine is a little different because I have an independent audit chat as part of the workflow itself, and its result often determines who is allowed to continue next. So with larger chains, a lot of my time is spent not on the analysis itself, but on manually handing work from one chat to another.

I’d definitely like to see the workflow you offered to share. This is genuinely interesting to me, especially because it looks like we reached similar problems but are solving them in different ways.

While thinking about your setup, we also came up with a possible way to automate a large part of my handoff process without moving the whole system to the API.

Each of my chats already ends its work with a formal status block that includes, among other things, who should act next. All project documents are versioned, hashed, and stored in Git.

So the idea is:

chat finishes → writes its result to Git → sets NEXT_ACTOR → a small local router detects the new commit → opens the appropriate existing ChatGPT chat and tells it to start the next stage.

The router itself would not analyze anything and would not need an LLM. The next chat would read the exact commit and file versions from the repo, perform its role, write the result, and indicate the next actor.

That would let me keep using regular ChatGPT chats under the subscription instead of moving the whole orchestration layer to the API.

The same idea could simplify ZIP packages too. Right now they are often used as transport between stages. In the new model, their contents could remain unpacked in the repo during the whole process, and the final ZIP would only be created once, at the very end, before execution.

It’s still just an experiment, but it looks like it could remove a lot of manual overhead. So I’d really like to see how you’ve structured your workflow — there may be parts of your approach that fit very well with this.

That’s very interesting — especially the separation between the router and the actual reasoning.

Your NEXT_ACTOR idea makes a lot of sense to me. The router doesn’t need to understand the research itself; it only needs to move control to the right chat, while Git remains the durable source of project state.

I ended up moving in a somewhat similar direction, although I haven’t automated chat-to-chat routing. In my current setup, Git holds the shared research state, and individual chats recover the state they need from there. Audit, review, handoff, and other roles operate on that shared state when needed.

So the chats are increasingly becoming execution surfaces rather than the place where the authoritative project state has to live.

I’ll share the experimental workflow I mentioned here. I think the comparison is particularly interesting because we seem to have reached a similar underlying problem from different directions: your approach is more explicit about routing between roles, while mine has focused more on continuity, recovery, and keeping the research state governable.

Update 3 — I may have hit a different kind of limit: one task became too large to execute

This is a slightly different problem from the one I described earlier with long conversations and branching.

This time, the issue is not a conversation gradually exhausting its practical context over many days.

The problem appeared with one specific execution task.

Over the last few weeks, I have been working on repairing one stage of a large project. Successive execution attempts kept revealing new problems only after execution had already started: incorrect dependencies, outdated artifact versions, inconsistencies between files, implementation errors, or cases that in principle could have been detected before the actual execution began.

In response, we gradually built a stricter and stricter validation process.

Every new failure added another case that should be detected earlier.

After many iterations, we reached a process that appears to be much more robust. Before the real execution begins, it now checks a large portion of the dependencies, versions, inputs, and previously discovered failure conditions.

That means an invalid state can now be stopped before consuming the actual execution attempt, instead of starting another run only to discover an error that could have been caught earlier.

That part worked.

The current remediation package now accounts for all currently known problems discovered by the previous attempts and audits. It also includes safeguards intended to prevent a fix in one area from damaging something that is already correct.

Its individual components have been tested repeatedly, and the earlier validation stages are passing.

That obviously does not mean I can guarantee that the whole process is flawless.

But at this point, the problem stopped looking like:

“we still don’t know what needs to be fixed”

and started looking more like:

“we know what needs to be done, but the complete task has become too large and too internally dependent to execute as a single run.”

And that produced a fairly ironic result.

For weeks, we were building a process designed to prevent execution from failing in increasingly subtle ways.

Eventually, we ended up with a remediation package strict enough to control practically all of the problems we currently know about — but the package itself became so large and interconnected that a single execution session could no longer complete the whole process.

At first, I suspected the problem might simply be an old, heavily loaded session.

So we started a completely fresh execution session, with no previous conversation history, and supplied it with the required sources and precise instructions.

The problem remained.

We then tried decomposing the task across agents and smaller scopes, so that each agent would have a limited responsibility and would not need to understand the entire project.

That still was not enough to complete the whole process.

For that reason, I do not want to call this simply a “context window limit,” because I do not yet know what the actual limiting factor is.

It could be:

  • active context size,
  • the number of dependencies that must be maintained simultaneously,
  • the practical length of a single execution run,
  • the number of operations and tool calls,
  • the amount of source material,
  • coordination overhead between subtasks,
  • or some combination of these.

What I can say from the user side is simpler:

a completely fresh session was still unable to complete this task as one execution unit.

And I find that interesting for another reason.

Current models can generate very large things in a single run: multi-file applications, frontends, backends, databases, configuration, dependencies, and entire project structures.

In this case, the task should theoretically be easier because the model does not have to invent the architecture from scratch.

It already has defined sources, scope, requirements, validation rules, STOP conditions, and a precise description of what must be changed.

And yet it appears to be harder to execute as a single unit.

My current hypothesis is that task size is not determined only by how much code or text must be generated.

The density of dependencies may matter just as much.

Generating 50 new files may be easier than modifying 10 existing files if every one of those changes must remain consistent with dozens of earlier contracts, tests, versions, and validation conditions.

I will now probably have to do exactly what I was trying to avoid: decompose one coherent remediation package into a sequence of much smaller, independently executable and independently validated stages.

Each stage will have to:

repair its own scope → pass validation → preserve the already-correct state → leave an unambiguous input state for the next stage

And this is where the experiment connects back to the earlier discussion about persistent project state.

If one task has to be distributed across multiple fresh conversations, sessions, or agents, then manually moving the same evolving packages and documents between them starts becoming part of the problem itself.

A shared, persistent, version-controlled workspace stops being merely convenient and starts becoming part of the safe orchestration of the work.

So after experimenting with very long conversations, I may now have reached a second interesting boundary case:

not “how long can one conversation remain useful?” but “how large can one strongly interdependent task become before it has to be decomposed into separate execution units?”

I do not know the answer yet.

But it looks like I have just started that second experiment.

I’ll document this case separately in more detail — how we got there, what actually failed, and what approaches might make this class of task more reliable. I’m moving it to a separate thread because it appears to be a different problem from long-chat continuity: the issue is increasingly about dependency density and how much tightly coupled work can realistically be handled as a single execution unit.

For about eight months, I’ve been building a fairly large, multi-stage AI-assisted system.

In another thread, I’ve been documenting a different class of problems around long-running conversations, context continuity, branching, and tasks that eventually become too large or too tightly coupled to execute reliably as a single unit:

[link to previous thread]

During the current large remediation effort, however, I ran into a different problem that seems worth discussing separately.

This is not a typical implementation bug.

It is not:

the application no longer starts after a change,

or:

we fixed A and accidentally broke B.

It is also not quite the usual form of technical debt where the code still works but gradually becomes difficult to maintain.

The problem is deeper.

A dependency error can remain invisible to tests

The project has a large number of tests and deterministic verification scripts. They are executed throughout subsequent implementation stages.

Despite that, the current audit process has started exposing errors that originated much earlier in the project.

The important part is that these errors did not necessarily cause the earlier tests to fail.

The issue appears to exist one level above:

“Does the implementation correctly satisfy the contract?”

The more dangerous question is:

“Is the contract itself correct, and are the assumptions it depends on valid?”

That distinction matters.

A test can correctly prove:

the implementation does exactly what the specification requires.

But if the specification contains a wrong assumption, then a green test may only prove that:

the wrong assumption was implemented correctly.

The most dangerous failure may be the one nobody notices

I do not yet know whether the final system would actually have failed visibly because of the specific issues we are finding. The complete product does not exist yet, so there is no final end-to-end validation beyond the current tests and verifiers.

But there is a realistic possibility that a system could:

  • start normally,
  • pass its tests,
  • perform its expected functions,
  • and still have part of its logic built on an invalid foundation.

That worries me more than a crash.

A crash tells you:

something is wrong.

A foundational error may tell you:

everything is working.

And then the next implementation inherits the same assumption.

After enough time, you no longer have one bug.

You have something closer to:

incorrect contract → dependent module → later specification → later implementation → later tests based on the same assumption

Every individual element in that chain may be locally correct.

The chain itself may still be wrong.

In my case, some of the problem comes from an older generation of the workflow

The earliest parts of this project were built using a much simpler process:

human → main design chat → Codex

At that point I did not yet have the independent audit layer, formal stage gates, dependency checks, and strict artifact/version validation that exist in the project today.

Those controls evolved later.

The interesting part is that the current remediation is now discovering problems originating from that earlier period.

So the stricter process did not create the problem.

It exposed a problem that had remained hidden.

If those assumptions had survived for several more months and more layers had been built on top of them, fixing the foundation at the end might have required revisiting a significant amount of roughly eight months of work.

What I am trying to separate now

One lesson from this remediation is that I no longer treat a passing test suite as sufficient evidence that the entire system is correct.

Tests remain essential, but I am increasingly separating several different kinds of verification:

  • implementation verification — did we correctly implement what was defined?
  • contract verification — is the contract internally coherent?
  • dependency verification — are its assumptions compatible with the things it depends on?
  • independent review — can another role/context challenge the decision rather than inherit the same assumptions?
  • provenance and versioning — which exact state and version produced this decision?
  • stage gates — can uncertain or inconsistent state become input to the next stage?

There is another lesson from the remediation:

Understand the problem globally, but execute the repair locally.

The dependency chain has to be understood as a whole. Fixing one node without understanding what depends on it can simply move the inconsistency somewhere else.

At the same time, the remediation itself became too large and too interconnected to execute safely as one unit.

We have already had to decompose it repeatedly into smaller, independently verifiable execution units.

That connects back to the problem I mentioned in my previous thread, but it is not quite the same problem.

Two different kinds of correctness

The distinction I keep coming back to is:

Did the system correctly do what we told it to do?

versus:

Did we tell the system to do the right thing in the first place?

The first question can often be covered extremely well by tests.

The second one is much harder.

And I think this second class of failure becomes especially dangerous in large AI-assisted projects, because a wrong assumption can be inherited by many later stages without producing an obvious failure signal.

Why I am writing this

If I were starting the same project again, I would rather spend an additional month on architecture, contracts, dependencies, documentation, and validation rules before serious implementation began.

For a large system, I increasingly think the documentation and contracts should be far ahead of the implementation rather than being generated afterward as a description of what the model just built.

The very natural workflow today is:

prompt → implementation → tests → works → next prompt → next implementation

For a small project, that can be incredibly effective.

For a large, multi-stage system, I now see a different risk:

each locally correct implementation may simply make an earlier incorrect assumption more deeply embedded.

Adding more agents does not automatically solve that problem either.

Changing the process to something like:

agent → plan → agent → implementation → tests

is better, but it can still fail if every participant inherits the same underlying assumptions.

If the first design decision defines the dependency incorrectly, several agents can simply implement that mistake much more efficiently.

That is why I am becoming increasingly convinced that complex AI-assisted development needs independent roles with deliberately conflicting objectives.

The designer’s job is to create a solution.

The auditor’s job is not to help the designer finish it. The auditor’s job is to try to prove that it should not pass.

A separate critical reviewer should challenge contracts, dependencies and assumptions rather than only reviewing code.

The executor should not be able to expand its own scope.

The human remains the final decision boundary.

And, most importantly, the workflow needs hard blockers capable of saying:

STOP. There is not enough evidence to allow the next stage to begin.

Even when the implementation looks good.

Even when the local tests are green.

That costs time.

Sometimes a lot of time.

But after what I am seeing now, one month spent establishing good contracts and independent validation may be much cheaper than discovering eight months later that dozens of later components inherited an error from the foundation.

So if someone is starting a larger AI-assisted software project today, my main warning would be:

Don’t only ask whether AI can build the next thing. First build a system that can stop AI when the next thing should not be built yet.

I’m curious whether others have encountered something similar: a system that passed its tests and looked correct, only to reveal later that the implementation was fine but an earlier specification, contract, or dependency assumption was wrong.

Sure, this happens. In fact this always happened - if your tests are based on unsafe assumptions, they would often pass but this would bite you at some point in Production.

Systems evolve partly by exercising them in Production and discovering things that weren’t immediately obvious - just as it happened before 2025.

Absolutely — and I agree that this is not a new software engineering problem. Production has always exposed assumptions that tests and pre-production validation failed to capture.

What surprised me in this case was seeing this failure mode before the complete system even exists as an end-to-end product. We are still building and validating the underlying structure, contracts and execution layers, so there is not yet a finished system that can simply be run in production-like conditions to expose the problem.

The implementation may correctly satisfy its contract, and its tests may correctly verify that implementation, while the contract itself already contains an incorrect upstream assumption. Later specifications, implementations and tests can then inherit the same assumption without any individual stage appearing obviously broken.

So the part I am focusing on is not “tests can miss things” — that is well understood. What caught my attention was how efficiently a wrong assumption can propagate through a large AI-assisted workflow when each later stage treats the previous one as valid input.

That is really why I wanted to share the observation. At this stage of the project, finding something originating much earlier in the workflow was a useful reminder that local correctness and system-level correctness are not the same thing.

I am therefore trying to move part of the validation one level earlier: not only “does the code satisfy the contract?”, but also “is this contract sufficiently justified to become a dependency for the next stage?”