Codex Desktop repeatedly loses the original acceptance goal after compaction and enters endless subagent/test loops

I looked through v0.2. This is now a well-formed experiment.

The strongest changes are that recorded_by is explicitly separated from authority, criterion state is derived from admission records, and completion is derived from settlement rather than from the worker’s verdict. Equally important, you state the limitation honestly: a skill-level protocol can test behavioral adherence, but it cannot technically enforce the authority boundary.

I would not add much more mechanism before the first run. The most informative result would be a redacted transition trace across at least one compaction and one failed live gate:

contract → evidence → assessment → admission/rejection → settlement or authority violation

A boundary violation would be just as informative as a clean run, because it would locate the point where protocol stops being sufficient and runtime enforcement becomes necessary.

No urgency while your limits are exhausted. When you are able to run it, I would be interested in the trace—including a failed result. We can then compare the observed boundary with KFD-10, while keeping your experiment independent of KFD.