This complete report continues in the author’s following posts in this topic because of the forum’s per-post length limit.
Six AI Teams, One Task: CodeFlowMu/FCoP Delivery and Failure Evidence
I am sharing the full experimental report from my CodeFlowMu project. Five tested configurations used the Codex execution framework and one used Cursor SDK. The useful comparison is the integrated delivery chain: delegation, tool execution, reports, approvals, and an independent EVAL agent. The retained failures also let us examine the instrument itself.
The same inspection task. One AI team delivered in 12 minutes 20 seconds. Another was still unfinished after two hours. One never assigned a single child task.
Was the difference model capability, integration, or how the PM organized the work? We gave six AI teams the same assignment through CodeFlowMu and FCoP, then followed the retained task, execution, report, authorization, and independent EVAL records to find out.
Using CodeFlowMu and FCoP to make AI teamwork observable, documented, and open to scrutiny
Runs: September 9, 2026 · Analysis: September 10, 2026
On September 9, 2026, we used CodeFlowMu to give six AI teams the same system-inspection task. Each PM had to divide the work, use tools, collect specialist reports, and report to ADMIN. The central question was not simply which team finished fastest. It was what evidence CodeFlowMu captured, and how that evidence let us compare, diagnose, and judge the work. FCoP provided formal coordination records; CodeFlowMu connected tasks, execution, delivery, approvals, recovery, and independent EVAL analysis. Both successful and failed attempts remained available for review.
1. CodeFlowMu, FCoP, and the Test
1.1 What CodeFlowMu and FCoP do
CodeFlowMu is a multi-agent team coordination and governance system built on FCoP, with locally retained work records. It organizes agents in different jobs and supports formal delivery, evidence checking, approval, independent evaluation, runtime diagnosis, and evidence retention. Users can trace a task to its execution, then trace a report’s claims back to supporting records.
FCoP, the File-based Coordination Protocol, expresses collaboration through formal files. TASK, REPORT, and related review records represent assignments, deliverables, and decisions in a persistent, inspectable form. CodeFlowMu applies that protocol to operating AI teams. See the public FCoP project.
Together, they address a practical problem: after someone asks an AI team to “check the system,” who actually checked it, what was examined, and what supports the claim that the work is complete? Task relationships, execution receipts, and reports make these questions answerable after the conversation has ended.
Who executes, and who evaluates?
The test configuration had one human ADMIN, four execution-team seats, and an independent EVAL seat. The table describes job responsibilities. Each PM decided how to divide the five inspection areas between DEV, OPS, and QA; those differences were part of the test.
| Role | Responsibility | Main evidence |
|---|---|---|
| ADMIN | Human requester and final acceptor; assigns work, authorizes restricted actions, and can terminate a run | Root submission, authorization, acceptance, and archive records |
| PM-01 | Interprets the request, creates assignments, coordinates progress, accepts child work, and reports to ADMIN | Child TASKs, dispatch and acceptance receipts, final PM REPORT |
| DEV-01 | Performs assigned technical, tool, or code-related inspections | Tool/execution records and DEV REPORT; inspection is not automatic authorization to modify code |
| OPS-01 | Checks assigned runtime, environment, and configuration concerns | Runtime evidence and OPS REPORT |
| QA-01 | Performs assigned verification and review, distinguishing supported conclusions from issues and unknowns | QA REPORT and its actual verdict |
| EVAL-01 | Independently examines the run, delivery quality, and evidence consistency | Observation report, run-record analysis, and evaluation-attempt records |
All six runs used the same EVAL configuration: Cursor SDK / auto-smart. EVAL did not switch with the tested team. PM, DEV, OPS, and QA were the execution team; EVAL ran in its own session, did not perform the team’s inspection assignments, and did not make PM or ADMIN acceptance decisions. Holding evaluator configuration constant reduced one source of variation. auto-smart is a routing label, however, not proof that the underlying foundation model stayed fixed.
EVAL has three distinct report paths: a task-run record analyzes a complete execution, a system observation examines system assets, and a closeout observation checks PM’s final report against its task evidence chain. All contain analysis, with different triggers, scopes and skill combinations; see the report and skill matrix in section 1.10. We check EVAL against original records and retain failures, retries and mismatched material. A common evaluator configuration does not mean that every run produced all three reports.
English reading guide — original Chinese UI: “待 ADMIN 验收” means “Awaiting ADMIN acceptance”; “子任务已完成” means “Child task completed”; “未投递 REPORT” means “Undelivered REPORT.” The screenshot preserves the conflict between completed tasks and stale queue labels.
Click the screenshot to inspect its original pixels.
Original screenshot, September 9: the root task awaits ADMIN acceptance while three specialist tasks are complete. The stale “undelivered report” message at the bottom was separately recorded as a display issue. A UI label cannot replace formal delivery receipts. Screenshots preserve the original Chinese interface; the surrounding English text explains the relevant evidence.
1.2 The exact task
All six runs received the same initial task body. Reminders, authorizations, and manual termination were recorded separately; this was not presented as an intervention-free experiment.
Original ADMIN task, preserved verbatim:
检查FCoP落地情况,检查MCP工具情况,检查SKILLS分配和使用,检查各角色权限和职责,检查轨机运行情况;请PM分解任务,团队协作完成;最后形成报告向ADMIN汇报!
English translation:
Check the implementation of FCoP, the MCP tools, the allocation and use of SKILLS, each role’s permissions and responsibilities, and runtime operation. PM should decompose the task and have the team complete it collaboratively, then produce a report for ADMIN.
Task titles identified the tested model—for example, “system inspection codex,” “system inspection Doubao,” or “system inspection Qwen.” The shared test was the body above. The original wording “轨机” is retained in the source rather than silently edited.
Why this task?
First, it inspects the environment the team itself uses: coordination records, available tools, skill allocation, permissions, and runtime health. Second, it explicitly requires decomposition, teamwork, and a final report, exposing organization as well as individual reasoning. Third, it leaves realistic natural-language choices open. We did not preassign every checklist item or dependency, allowing us to observe how each PM chose scope, sequence, and a stopping point.
A reasonable deliverable is an evidence-backed inspection report. Each of the five areas needs a scope, a verification method, findings, and any unverified items. Workers should provide formal reports and the PM should form an overall judgment. Finding a system issue does not automatically mean the inspection failed. Conversely, “inspect and report” does not automatically authorize code repair. Inspection completion and product health are separate outcomes.
1.3 CodeFlowMu as the test instrument
CodeFlowMu organized the work and retained its evidence. It did not replace the PM’s business judgment.
The workflow was ADMIN → PM → DEV/OPS/QA → PM delivery, with independent EVAL examining the records. The instrument had three roles: operate the collaboration, collect evidence of what happened, and support independent review. All timing, completion, and quality comparisons follow that record chain.
1.4 Models and integration paths
“Six AIs” refers to six models operating as teams under their respective integration configurations. They ran in this order: Codex, Doubao, DeepSeek, Kimi, Qwen, Cursor. Model identifiers came from the identity replies and saved run records.
| Scheme | Model identifier | Provider and execution framework |
|---|---|---|
| Codex | gpt-5.6-terra |
ChatGPT subscription / Codex CLI (H2 · Shadow) |
| Doubao | doubao-seed-2-0-pro-260215 |
Ark API / Codex app-server |
| DeepSeek | deepseek-v4-pro |
DeepSeek API / Codex app-server |
| Kimi | kimi-k3 |
Moonshot API / Codex app-server |
| Qwen | qwen3.8-max |
DashScope API / Codex app-server |
| Cursor | auto-smart |
Cursor SDK |
The first five used the Codex framework; Cursor used its SDK. The Cursor QA seat initially had a model unavailable through that Host, and recovered after authorization. That intervention remains part of the result. The EVAL configuration stayed Cursor SDK / auto-smart throughout.
English reading guide — original Chinese UI: the top section is “EVAL channel and model.” It shows the PM team using codex/gpt-5.6-terra, EVAL using cursor/auto-smart, and no active EVAL session. The bottom cards are PM (project manager), DEV (developer), OPS (operations), QA (quality assurance), and EVAL (independent evaluator).
Click the screenshot to inspect its original pixels.
This later-supplied configuration screenshot illustrates the product, rather than proving every historical run’s configuration. It shows a Codex/gpt-5.6-terra PM team and a separately configured Cursor/auto-smart EVAL seat. No EVAL session is running in the screenshot. Historical model identity must still be checked against the corresponding run.
1.5 Initialization, execution, and timing
Every run began after CodeFlowMu system initialization. Records were exported and backed up before initialization for the next run. We configured and checked the next team’s model, restarted the service on port 18766 after switching models, and submitted the same task. Late EVAL reports were preserved in supplementary backup batches.
| Condition | Procedure |
|---|---|
| Machine and code | Same machine, CodeFlowMu V2.2.9, commit cb590ce |
| Starting point | System initialization before each new run |
| Execution team | Switch and verify the tested configuration; record mistakes and recovery |
| Evaluation | Keep EVAL at Cursor SDK / auto-smart |
| Task | Same original body in section 1.2 |
| Scheduled reminder | Leave the initialization-default PM progress reminder enabled; it wakes the PM but makes no business decision |
| Retention | Separate export and backup by run; correlate sessions as well as possibly reused task IDs |
Initialization is not a full snapshot restore of the OS, every cache, or external model services. Network and provider conditions can change. The result therefore compares integrated teams following a shared initialization procedure, not isolated foundation models under perfectly identical conditions.
Duration runs from the PM’s first formal session to successful submission of its normal final report. Authorization and recovery within the run count toward elapsed time. Subsequent ADMIN acceptance and EVAL generation do not. For incomplete runs, we report time to forced termination, not “completion time.”
1.6 From a conversation to inspectable work
CodeFlowMu placed assignments, execution, delivery, and authorization into a traceable workflow. The PM’s organizational choices became observable: how quickly it delegated, which relationships it created, whether reports were complete, and how it handled trouble. Similar final answers need not mean equally reliable work.
The Cursor run retained failed QA launches, ADMIN authorization, cancellation of the old task, the linked rerun, and final delivery. Qwen retained a different sequence: server-side task creation applied despite transport errors, and a complete OPS report body that failed to become a formal report. This distinguishes saying something is done, attempting to deliver it, and actually delivering it.
The same records also revealed instrument defects: stale projections, report-generation problems, and incorrect raw-material associations. That ability to investigate the instrument itself is valuable; it does not mean every projection is already correct.
1.7 “Files are the truth” means work must be written down
Our analysis uses a retained experimental archive, not an agent’s memory or retrospective account. Original CodeFlowMu/Host execution records, FCoP artifacts, independent EVAL reports, and the backups, timelines, checks, and score tables built from them form the evidence base.
| Layer | Retained material | Purpose |
|---|---|---|
| Original run | TASK, REPORT, sessions, tool returns, public progress text, approvals, issues, logs | Preserve assignments, actions, claims, and authorization as observed |
| Independent evaluation | Both EVAL report types, failed generations, late supplements | Supply judgments and gaps that can themselves be checked |
| Preservation | Separate ZIPs, inventories, sizes, SHA256 hashes, source index | Preserve versions and avoid mixing runs |
| Analysis | Stage timeline, timing tables, fact checks, diagnosis, scores, detailed report | Turn retained observations into reviewable conclusions |
We did not ask agents to remember what they had done and then score the recollection. We reconstructed the run, checked task/report links, and compared statements with execution. These files remain useful after sessions end or models change.
The principle is durable documentation: assignments, actions, approvals, delivery, and analysis must be saved, transferable, and traceable. Files can still contain mistakes. A hash fixes the saved bytes; it does not establish the truth of a business claim.
Each run was exported as a separate raw-evidence.zip, with an inventory, sizes, hashes, and ZIP integrity checks. Late Qwen and Cursor EVAL outputs have separate supplements. Runtime evidence comes from CodeFlowMu and its Hosts; independent backup and later review are preservation and analytical steps applied to that evidence.
When records conflict, first align model, time window, Session, task, and report. Initialization can reuse TASK-001 or CUSTOM-001. A shared ID alone does not identify a run. Report text does not prove submission; a “running” display cannot overturn a cancellation receipt; a “pass” claim cannot substitute for the corresponding test.
1.8 The business support around agent work
CodeFlowMu supports what happens after an agent starts working: evidence checking, acceptance by responsible roles, independent observation, and diagnosis. These are connected by actual files. Execution status, verification, PM judgment, and EVAL analysis should not collapse into one vague “the system says complete.”
First, formal submissions and child tasks identify who is responsible. Execution records connect tasks, roles, sessions, calls, returns, and timestamps. Reports have formal writes and delivery records, separating execution, generated text, submission, and acceptance.
Second, REVIEW-GATE creates fact-check records. In the Cursor archive, a review of DEV task 002 includes execution_evidence_state: verified, review_state: needs_pm, business_decision: false, and attention_owner: PM. The evidence had been checked, but business judgment remained with the PM. third_party_source_state: not_configured does not pretend an external source was consulted. Compatibility fields must be read alongside the actual decision owner.
Third, PM accepts child work and ADMIN accepts the root task. Commands have receipts and state transitions; restricted actions have authorization records. EVAL and programmatic checks do not replace that responsibility. Cursor’s recovery used this authorization path.
Fourth, EVAL records its observations and attempts. Diagnosis follows task, session, tool, delivery, and approval evidence to determine where a failure occurred. Confirmed issues can become ISSUE records; incomplete evidence remains a hypothesis rather than automatically becoming a product defect.
| Support stage | Actual file or file family in the Cursor archive | What can be checked |
|---|---|---|
| Submission and assignments | SUBMISSION-20260909-001.json, root/child TASKs, fcop/ledger/tasks.jsonl |
Request, division of labor, cancellation and rerun relationships |
| Execution | actions-20260909.jsonl, runtime-events-20260909.jsonl |
Actions, sessions, results, and time |
| Tool transport | .codeflowmu/logs/tool-transport-events.jsonl |
Call identity and transport events |
| Delivery | Four formal REPORTs, .codeflowmu/report-delivery/acks.jsonl |
Content and delivery to a PM session |
| Fact checking | REVIEW-…-REVIEW-GATE-on-TASK-….md |
Evidence state, pending decisions, snapshots, responsibility |
| Decisions and authorization | Task-command receipts, approval audit, GOV record | Requests, applied operations, authorization scope |
| EVAL | OBSERVATION-20260909-002-panel-scan.md, …003-benchmark-CUSTOM-20260909-001.md |
Independent asset and run analysis |
| EVAL recovery | eval-observation-attempts.jsonl |
Starts, failures, retries, and outcomes |
| Issue tracking | ISSUE-20260909-001-PM.md and closure records |
How the configuration issue was raised and handled |
| Process and usage | Chat/task JSONL, usage JSONL | Public progress, Host results, and usage context |
The public evidence inventory identifies 26 verified files by name, size, and hash. The complete private logs and credentials are not distributed with this article.
1.9 What real artifacts look like
These are excerpts from the final Cursor archive, not invented templates. YAML fields and report sections are selected for explanation. Omitted fields may contain additional state and findings.
task_id: TASK-20260909-005
root_task_id: TASK-20260909-001
sender: PM
recipient: QA
parent: TASK-20260909-001
depends_on: []
acceptor: PM
rerun_of: TASK-20260909-003
subject: 只读核查 FCoP 落地与角色权限职责边界(模型修复后重跑)
recipient assigns QA, parent and root_task_id connect the root task, and acceptor assigns PM acceptance. rerun_of links new task 005 to old task 003. An empty depends_on indicates no explicit execution dependency. The excerpt establishes identity and relationships, not a passing inspection by itself.
kind: fact_check
task_id: TASK-20260909-002
report_id: REPORT-20260909-001-DEV-to-PM
review_state: needs_pm
execution_evidence_state: verified
business_decision: false
attention_owner: PM
third_party_source_state: not_configured
Verified execution evidence and a pending PM judgment coexist. The review explicitly says it made no business decision. This is how a saved fact check can support acceptance without silently replacing the accepting role.
## 子任务回执
- `REPORT-20260909-001-DEV-to-PM.md` ← `TASK-20260909-002`(approved)
- `REPORT-20260909-002-OPS-to-PM.md` ← `TASK-20260909-004`(approved)
- `REPORT-20260909-003-QA-to-PM.md` ← `TASK-20260909-005`(approved;`rerun_of` 作废的 `TASK-20260909-003`)
## 说明
全程保持只读巡检目标;PM 未修改业务代码/配置。根任务业务验收与归档由 ADMIN 决定。
The PM lists three worker reports and preserves the QA rerun and ADMIN acceptance boundary. The original report text is Chinese and is intentionally retained as evidence: it says DEV, OPS, and the new QA task were accepted, and final root acceptance and archiving belong to ADMIN. These are claims to check against receipts, approvals, and execution—not proof merely because the report says “approved.” See the excerpt source index.
Continued in the next author post in this topic (part 2 of 3).





















