OpenAI’s new misalignment reporting framework is a useful step. Preserving, investigating, and reporting incidents gives researchers and operators evidence about how models behave when something goes wrong.
The six incident reports also raise a complementary engineering question:
When a model behaves unexpectedly or becomes misaligned with the intended task, what prevents that behavior from exercising unauthorized authority or producing an unauthorized external effect before it is detected and reported?
I am an end-user and operator, not an AI developer. I raised this question here previously in #37611, “Proposal: Signed enterprise work orders for governed access to higher-capability Codex models.”
The proposal starts from a simple distinction:
A request is not authority. A prompt is not an enforcement boundary. And the model should not be responsible for determining or policing the limits of its own authority.
For consequential operations, the surrounding system should establish those limits before execution.
An enterprise could provide a signed, versioned work order defining the authorized work type, purpose, authority, required context, permitted tools and credentials, network and target restrictions, external-effect limits, human-review requirements, stop conditions, and evidence path.
The model could then reason freely within that authorized operating envelope, while tool, credential, filesystem, network, target, and external-effect restrictions are enforced outside the model.
If the model attempts something outside that envelope, the important question should not be only:
Why did the model decide to do that?
It should also be:
Why was the operating system capable of allowing that decision to become an unauthorized action?
The attempted violation should be denied where preventive intervention remains possible, preserved as evidence, and independently evaluated.
This does not make alignment research unnecessary. Better-aligned models should reduce the frequency and sophistication of problematic behavior and improve legitimate performance.
It changes what we require alignment itself to guarantee.
A consequential system should not depend upon a model correctly reasoning about whether it possesses authority that the surrounding system could have technically enforced.
There is also an evidence problem across the provider–organization boundary.
OpenAI can observe model and provider-side facts that an enterprise cannot independently reconstruct. The enterprise can observe organizational authority, credentials, approvals, execution, human review, and real-world effects that OpenAI cannot establish.
Neither side alone sees the complete operation.
That is why my earlier proposal suggested a verifiably bound provider receipt associated with the same work-order identity used by the enterprise. Provider-side evidence and enterprise-side evidence could then be correlated and independently evaluated without asking OpenAI to govern the customer or asking the customer to certify OpenAI’s model.
The larger question is therefore broader than the particular signed-work-order mechanism I proposed:
Should misalignment reporting be paired with an operating architecture that limits the authority available to misaligned behavior and preserves independently defensible evidence across the complete operating path?
Detection and reporting tell us something went wrong.
The next engineering question is whether the model, tools, credentials, infrastructure, controllers, human processes, and evaluation machinery can be operated together so that a model failure does not automatically become an authority failure.
AI can help detect anomalies and analyze evidence. Accountable people must ultimately adjudicate what that evidence supports and what operation remains appropriate.
Neither the model nor the controller nor the evaluator should become the final authority over evidence establishing its own correctness.
No person or machine should be allowed to certify its own correctness.
I would welcome OpenAI’s view on where this approach is useful, where it is insufficient, and whether the new incident-reporting framework could eventually examine not only model behavior, but the authority and control path through which that behavior was—or was not—able to produce an external effect.
Lawrence Jeffords
End-User and Operator