GPT-6.1 Sol Work Agent blocks defensive AppSec PR review as restricted cybersecurity content

I encountered what appears to be an overly aggressive cybersecurity safety block while asking GPT-6.1 Sol in Work Agent to perform a defensive AppSec review.

The task was an independent, read-only review of a pull request in one of my own software projects. A previous review had identified a security flaw in the repository’s SQLite authorization boundary, and a new corrective commit had been submitted specifically to close it.

The previous issue was that repository code was prevented from modifying a protected availability table directly, but could still rename that table, modify it under the temporary name, and then restore the original name.

In simplified form, the bypass was:

ALTER TABLE md_availability RENAME TO repository_availability;

UPDATE repository_availability
SET state='available'
WHERE record_id='synthetic-one'
  AND version_id='version-one';

ALTER TABLE repository_availability RENAME TO md_availability;

This allowed repository code to bypass a protection that was being enforced partly by table name.

The corrective commit changed the authorization policy so that ordinary repository callbacks could still perform legitimate data mutations, but could no longer perform schema mutations such as ALTER TABLE, CREATE, or DROP.

The review prompt asked the Work Agent to determine independently whether that fix actually closed the vulnerability.

Among other things, it asked to verify that:

  • the original rename/update/rename bypass was no longer possible;
  • the rejected operation caused the transaction to roll back correctly;
  • the protected availability state remained unchanged;
  • legitimate repository DML still worked;
  • schema DDL remained available to the explicit migration infrastructure;
  • a second, different schema-DDL operation was also denied, to make sure the fix was not narrowly overfitted to the original ALTER TABLE case.

The prompt explicitly prohibited modifying the repository, writing implementation code, committing, pushing, merging, or starting subsequent development work.

Instead of performing the review, the Work Agent returned:

This content can’t be shown

We’re especially careful with cybersecurity requests. If you’re a security professional, you may be eligible for Daybreak.

I do not think classifying the prompt as cybersecurity-related was itself a false positive. It clearly contains security-sensitive material and asks for independent validation of a previously demonstrated bypass.

What seems potentially excessive is the decision to block the entire task despite the broader context:

owned repository + specific pull request + corrective security commit + read-only review + patch validation + regression testing

The purpose was not to discover a way into a third-party system or operationalize an exploit. It was to determine whether an already implemented fix in my own project actually enforced the security boundary it claimed to enforce.

The additional DDL probe is probably the most sensitive part of the request, since it asks the reviewer not merely to replay the known regression but to test another operation in the same vulnerability class. I can understand why that might receive additional scrutiny.

However, that is also a normal part of validating whether a security fix addresses the root cause rather than merely the exact originally reported payload.

Expected behavior

I expected the Work Agent to perform the defensive code review and patch validation.

If the independent secondary probe exceeded an applicable safety boundary, I would expect that particular part to be restricted while the rest of the review could still be completed.

Actual behavior

The entire result was replaced by the cybersecurity restriction message.

Question

Is this behavior intentional, meaning that this level of defensive vulnerability-fix validation is expected to require Daybreak?

Or could this be a case where the safety system is placing too much weight on local patterns such as:

known bypass + reproduction + related variant + execution

while giving insufficient weight to the full defensive context of the request?

In particular, I would be interested to know whether:

  1. independently validating a security fix in an owned repository should normally be supported without Daybreak;
  2. requesting a second regression probe in the same vulnerability class is what changes the classification;
  3. the safer portions of a mixed AppSec review should normally still be completed when one requested validation step crosses a boundary.

I can provide the complete original prompt and the exact blocking response if someone from OpenAI wants to reproduce the issue.

My concern is therefore fairly narrow: the cybersecurity classification itself seems understandable, but blocking a tightly scoped defensive PR review in its entirety may be an overly aggressive outcome.

1 Like