AGENTS.md and project Markdown instructions are not reliably followed across tasks

Even though I’m submitting this bug report, I am DONE with OpenAI and the ChatGPT app. It is so broken in so many areas that I’m spending half my time fixing issues instead of getting work done. Business Plan customer with 4 standard and 1 premium seat.

App: ChatGPT macOS / Codex project tasks, including scheduled and delegated tasks. Mac Studio and MacBook. Last verified version: 26.930.31730, build 12947, on October 3, 2026; current build needs reconfirming.

Problem

We maintain AGENTS.md, BCP-WRITING-PROCESS.md and WRITER-RUNTIME.md so we don’t have to repeat our requirements for every task. Tasks still miss documented checks, use outdated copies, or continue following older instructions after corrections.

The assistant acknowledges the instructions and reports fixes, but related tasks repeat the same mistakes. There is no clear way to see which instruction files and versions each task actually used.

Examples from October 5–7

  • Our rules require APIs and scrapers first, with browsers as a last resort. Browser use continued without first checking the documented alternatives.
  • Articles omitted links to existing reviews despite internal-link checks being required. I had to catch this manually.
  • The main WRITER-RUNTIME.md was revision 10 while a handoff copy remained revision 9. This was discovered and repaired later.
  • I instructed the assistant to stop routine Slack updates. After it reported updating the rules and notification controls, an independent scheduled task posted again the next morning.
  • A project-folder rename left an existing task pointing to the old directory, preventing it from starting. This may be a separate path-handling bug, but there was no useful warning before the scheduled failure.

Reproduction pattern

  1. Set up a project with AGENTS.md and explicitly referenced workflow files.
  2. Create related scheduled or delegated tasks.
  3. Update the instructions or correct a recurring mistake.
  4. Run existing tasks and compare their actions with the current project instructions.

These are observed symptoms; I don’t yet have a minimal reproduction proving they all share one cause.

Expected behaviour and requested fixes

  • Show which instruction files each task actually loaded, their versions/hashes and when they were read.
  • Clearly define how updated instructions reach existing scheduled and delegated tasks.
  • Warn about missing files, conflicting copies and stale project paths before execution.
  • Make required checks verifiable, rather than relying on the assistant saying it followed them.

Important distinction

Our required Jev review also encountered a security denial. In that case, recorded hashes matched the correct workflow files, so missing instructions were not established as the cause. The app should explain when an action is blocked by security rather than presenting it as a new requirement in our workflow. This is not a request to let Markdown files override security controls.

Impact

Repeated failed runs, manual checking and repeated “fixes” that don’t carry through to related tasks. The instruction files are supposed to make the workflow dependable; currently I still have to supervise each step.