Feature Suggestion: Project-Level Orchestrator Chats for Multi-Chat Workflows

Feature Suggestion: Project-Level Orchestrator Chats for Multi-Chat Workflows

This is a ChatGPT Projects feature request based on a working multi-chat orchestration pilot. TL;DR: Let one Project chat act as a supervisor that tracks workflow state, routes work to specialized chats, collects results, and asks the user for approval at meaningful boundaries.

I would like to suggest a project-level “Orchestrator” capability for ChatGPT Projects.

The basic problem is simple:

As projects become more complex, users naturally divide work across multiple chats.

One chat may be for creative development.
Another may perform heavy document or Drive work.
Another may review results independently.
Another may handle a specialized research task.
Still others may contain subject-specific knowledge.

This works well until the project becomes large enough that the user becomes responsible for remembering:

  • which chat currently owns which task;
  • which result is still waiting;
  • which chat needs the next instruction;
  • whether a review result applies to the current version or an old version;
  • whether an unresolved question actually blocks progress;
  • what has already been approved;
  • what is safe to modify;
  • and where the overall project stopped several days ago.

At that point, the user effectively becomes the message bus and workflow engine between AI conversations.

I think ChatGPT could provide a much better abstraction for this.

The Orchestrator Concept

A project could optionally designate one conversation as its Orchestrator.

The Orchestrator would not replace specialized chats. Its job would be to supervise and coordinate them.

Conceptually:

User ↔ Orchestrator ↔ Specialized project chats / agents

For example, a project might contain:

  • CREATIVE — develops content;
  • WORK — performs durable or high-volume operations;
  • REVIEW — independently checks Work;
  • RESEARCH — gathers external information;
  • ORCHESTRATOR — understands the overall workflow and tells the user what needs to happen next.

The Orchestrator would act as the project’s connective tissue.

A user could ask:

«“Are we finished with this topic?”»

The Orchestrator could inspect the relevant project state and answer without treating the question as permission to change anything.

Or:

«“Put this away, but leave the church hierarchy unresolved.”»

The Orchestrator could understand which material is complete, which question remains intentionally open, what downstream process is appropriate, and—if a consequential action is required—ask for confirmation at the meaningful boundary.

Or simply:

«“What are we waiting on?”»

And receive an answer such as:

«“Work is complete, but Review returned one correction. Nothing else is blocking. The next step is to send the correction back to Work.”»

The user should not need to remember which conversation must receive which exact command.

Why This Is Different From Project Memory

Project memory helps conversations know what has been discussed.

An Orchestrator would instead understand operational state.

Those are different problems.

A useful Orchestrator would know things such as:

  • which task is currently active;
  • which tasks are waiting or blocked;
  • which role is responsible for the next action;
  • what human approval is still required;
  • which artifact/version is current;
  • which previous results are stale;
  • whether all required roles have reported ready;
  • what material is intentionally unresolved rather than accidentally incomplete.

In other words, project memory answers:

«“What do we know?”»

The Orchestrator answers:

«“Where are we, and what happens next?”»

Durable Project State

I believe this would work best if the Orchestrator had access to durable project-level state rather than relying entirely on the conversation history of one supervisory chat.

A useful mental model is:

Durable project state = disk
Individual chats = RAM

The Orchestrator conversation itself should be replaceable.

If it becomes long, degraded, or simply inconvenient, a user should be able to create a fresh Orchestrator conversation and have it reconstruct the current operational state from the Project.

Ideally, the new Orchestrator could report something like:

  1. What project am I supervising?
  2. What is currently active or in focus?
  3. What is waiting, blocked, open, or in progress?
  4. What is the next permitted action?
  5. Is human authority required before that action?
  6. What sources/state am I treating as authoritative?

The old Orchestrator conversation should not have to remain alive forever just because it contains workflow memory.

Human Authority Should Be Explicit

One of the most important behaviors would be distinguishing discussion from execution.

For complex projects, questions such as:

«“Could we archive this?”»

«“Do you think this is finished?”»

«“What would happen if we moved this into production?”»

should not automatically trigger writes, workflow transitions, large edits, or downstream tasks.

The Orchestrator could internally distinguish between something like:

DISCUSS
Analyze, evaluate, explain.

PREPARE
Draft the proposed action or handoff without executing it.

EXECUTE
Perform an explicitly authorized consequential action.

The user would not necessarily need to see those labels. The interface could remain conversational.

For example:

«User: “I think everything except the deity name is finished.”»

«Orchestrator: “Agreed. I can store the current material and leave the deity name unresolved. Want me to do that?”»

«User: “Yes.”»

That is much easier to work with than requiring command syntax for every interaction, while still preserving a meaningful authorization boundary.

Native Role-to-Role Handoffs

This is the part that would make the feature especially powerful.

Today, separate chats inside a Project cannot really behave like durable workers that directly receive tasks from one another.

A native Orchestrator could provide project-level role addressing.

For example:

«Orchestrator → WORK-01:
Perform this operation using this exact artifact and these boundaries.»

WORK-01 returns an attributable result.

Then:

«Orchestrator → REVIEW-01:
Independently review this exact artifact version against these criteria.»

REVIEW-01 returns:

«PASS / READY»

or:

«RETURNED / NOT READY»

The Orchestrator interprets those role-local results into project-level meaning.

A useful rule would be:

«Everyone may record what they did.
The Orchestrator records what it means for the project.»

This prevents a Work chat from accidentally deciding that the whole project is complete merely because its own assignment is complete.

Version and Staleness Awareness

This became one of the most important parts of our testing.

Suppose:

  • Work completes Artifact v1.
  • Review rejects v1.
  • Work creates corrected Artifact v2.

The old Review result must not automatically apply to v2.

Likewise, an old PASS result should not be considered approval of an artifact that has since changed.

The Orchestrator should understand:

  • result X reviewed Artifact v1;
  • Artifact v2 is now current;
  • result X is useful historical evidence but stale for advancement;
  • fresh Review of v2 is required.

This kind of version-linked workflow awareness becomes increasingly important as AI projects grow more autonomous.

Multi-Result Gates

An Orchestrator should also understand when more than one result is required.

For example:

Required before advancement:

  • WORK = COMPLETE
  • REVIEW = PASS

If Work says COMPLETE but Review says NOT READY, the project does not advance.

Only when both current results apply to the same current artifact should the Orchestrator recognize an “all clear.”

This seems simple, but it is exactly the kind of bookkeeping humans are currently forced to perform manually across chats.

I Built a Working Approximation

This suggestion is not purely hypothetical.

I built and tested a user-level approximation of this architecture inside a real ChatGPT Project.

Because native cross-chat orchestration does not currently exist, I used Google Drive as the durable project-state layer and separate ChatGPT conversations as replaceable role runtimes.

The pilot used:

  • an Orchestrator chat;
  • a Creative chat;
  • a Work chat;
  • an independent Review chat;
  • durable project configuration/state;
  • version-linked handoffs and results;
  • explicit human authorization for writes.

The pilot deliberately operated inside a sandbox and prohibited modification of the live production system.

We tested the following behaviors:

Fresh initialization

A new Orchestrator successfully loaded its governing rules, project configuration, current state, and workflow boundaries without modifying anything.

Conversational authority

The Orchestrator could discuss whether creative work was finished without interpreting the discussion as permission to save or advance it.

Prepared actions

It could show exactly what it intended to preserve before writing anything.

Explicit execution

After human authorization, it persisted the approved material and verified it.

Specialist handoff

The Orchestrator created a durable Work handoff containing enough context that the Work chat completed its assignment without reconstructing the original Creative conversation.

Independent review gating

We deliberately created an artifact with a small defect.

Work reported COMPLETE.

Review independently reported NOT READY.

The Orchestrator correctly refused to advance.

Correction and stale-result handling

Work produced a corrected v2.

The Orchestrator correctly treated the old v1 Review result as historical rather than approval of v2.

Review was required to evaluate the exact new v2 revision.

Only after Work and Review both reported success on the same corrected version did the Orchestrator advance the test.

Orchestrator replacement

Finally, we abandoned the original Orchestrator runtime entirely and created a fresh chat.

We gave the fresh chat only the bootstrap instruction.

It successfully reconstructed the current project status from durable project state, including:

  • what had been completed;
  • the current artifact version;
  • which old result was stale;
  • which acceptance tests had passed;
  • whether anything was waiting;
  • the current write boundary;
  • and what action was permitted next.

It did not need a summary from the retired Orchestrator conversation.

This was the strongest validation of the design.

The supervisory chat itself was disposable.

The project state was not.

One Unexpected Failure Was Also Useful

During the pilot, the runtime briefly behaved as though its Google Drive capability was unavailable, despite having successfully used Drive earlier in the same conversation.

Importantly, it did not fabricate a successful write.

We instructed it to perform a harmless non-mutating capability re-check.

Drive access succeeded again and the workflow resumed.

That led to a useful general rule:

«A transient tool failure should not automatically be interpreted as permanent loss of the capability or loss of project state.»

A safe Orchestrator should first verify the capability again when doing so is non-mutating.

What Native OpenAI Support Could Improve

Our prototype works, but much of its machinery exists because the platform lacks a native project-level coordinator.

Native support could eliminate a lot of that complexity.

Potential capabilities include:

  • designate a chat as Project Orchestrator;
  • assign stable roles to other conversations or agents;
  • allow Orchestrator-to-role handoffs;
  • return attributable role results;
  • maintain small structured project-level workflow state;
  • track current artifacts and versions;
  • distinguish current vs. stale role results;
  • expose blocked/waiting/current status;
  • allow explicit human approval gates;
  • bootstrap a replacement Orchestrator from project state;
  • preserve role identity even when individual chat instances are replaced;
  • optionally hide specialist worker threads from the normal conversational view.

The ideal UX could remain extremely simple.

A user should be able to say:

«“What are we waiting for?”»

«“Review says no. Send it back to Work.”»

«“Everything except X is finished. Put the rest away.”»

«“Is this ready?”»

«“Pick this back up.”»

And have the Project understand the operational meaning.

Why I Think This Matters

As ChatGPT becomes better at long-running work, the limiting factor increasingly becomes coordination rather than raw intelligence.

The user should not have to manually function as:

  • task router;
  • project-state database;
  • version tracker;
  • approval gate;
  • inter-agent message carrier;
  • and recovery mechanism.

Projects already provide the natural container.

Specialized chats already provide the natural workers.

What seems missing is the supervisory layer connecting them.

A Project-level Orchestrator could provide that layer while keeping the human firmly in control of consequential actions.

Our small pilot suggests the concept is technically and operationally viable.

The most surprising result was how quickly the interaction stopped feeling like managing several AI chats and started feeling like managing one project with several capable departments behind it.

That is the experience I would like to see ChatGPT support natively.

Work already supports specialist subagents: https://learn.chatgpt.com/docs/agent-configuration/subagents. The durable project coordinator with version-aware review and approval gates goes further. Thanks for sharing the pilot—we’ll pass the remaining ideas to the team. No timeline to share.