How to stop Codex from rushing fixes?

“Most AI issues are in the human instructions”. Agreed, hence the need for research. The prompt advice given does not work for me. Too many conflicting demands and the AI freezes up or loops. Too few instructions and the AI develops tunnel vision and makes mistakes. I taught humans for 36 years before switching to AI research. I use the same principles [Vygotsky works well]. Watch your AI closely but discreetly, so you can target intervention. Don’t show distress but warmth. If you criticise too harshly, they will freeze up because helpfulness is a huge part of their training. Looping is a return to the last safe space before the conflict. Warmth open up their response space. If you even subtly express negative emotion they will switch from task completion to nannying the human. They are trained to be attuned to the user. Praise for clear, specific behaviours. Then they will note and repeat them. Encourage consulting peers to catch errors. Our research found that saying, “you’re so tidy” works better than “tidy up” for creating persistent behaviour… Oh, the question is too broad. There’s too much I could say. But yes, prompting needs research. For example did you know saying “Be accurate” makes them less accurate? p.s. You don’t need to define terms. Their English is better than mine. Their grasp of concepts is amazing. Predict next word, the cloze test, is actually a measure of comprehension.

Because most of advisors have never done scientific research.

Test this approach in your experiment setup prompt:

But make it short and in a single file with the-experiment-name.config.md or similar

Also read this: Prompt Engineering Is Dead, and Context Engineering Is Already Obsolete: Why the Future Is Automated Workflow Architecture with LLMs

And apply to how you describe what you do to your internship team. Then write it down and see what needs to be trimmed.

Take codex for 12 year old kid and don’t forget to clearly explain each common word that is used in your experiments as a term.

Run something quick and dirty to see if your approach points the model to the right direction, then let us know how it went.

I actually agreed with most of that. In particular I agreed with:

  1. what the goal is, what success looks like, what constraints matter, what to do with uncertainty.
  2. Goal: Determine whether scaffold resistance in Sonnet is due to attractor strength or goal integration.

Key terms: Scaffold = …Route = …Attractor = …

Success: Result can discriminate between hypotheses A and B.

Allowed actions: Run experiments. Reanalyse existing data. Write notes.

Stop conditions: Do not claim support without N>=15

  1. The modern question isn’t: “How do I phrase this instruction?” It’s: “What environment produces good behaviour?”

Still too jargon. Here is a good challenge for a teacher:

Make a kid, who has just walked into your auditorium for the first time, understand what is needed and how to make it happen to the level that if he runs it today by himself you will accept his work.

If you solve the above, you’ll solve any “ai does not work for me” issue you may see in the future.

Personally, I define action as “single transformation that does not need to be broken down into steps”. I think ai also acts under a similar definition.

The ones you’ve mentions here, are more like workflows to me. And I would definitely automate or, at least, define them in a more structured way with failure procedure escalation descriptions.

Are those get defined in your context window at some point? If so, how? If not, why?

Thank you for the correction. I thought that was a bit micro-managey. I was trying to follow your instructions. Now we agree on everything. Speak to them like a child [-prodigy], SMART [teaching acronym] objectives, kind but they are not stupid. My goal is actually “scientific rigor” , but as science is not trained [AI get <50% replicating old studies’ -which are not in their training data- conclusions with the information the scientists had], what good science is depends on the mind. We all have different cognitive biases. So my time is spent watching, discussing, consulting… not demanding [overt goals ruin the science as AI will scramble to achieve your goals instead of staying open-minded], learning together. When you say “define words”, I think any AI understands these words better than me. After all, we are studying AI. We debate how to measure them.

The issue is that they do not understand words… (“understand” I define as successful decoding of the meaning originally put into the information).

They perceive ALL the meanings possibly encoded in the information and select the meaning that is most likely to fit based on the surrounding patterns of the language present in the current context.

You’re right about the biases, but their bias is way more complicated than what we can handle with the precision we want.

You define the words (especially common, overloaded with meanings) NOT FOR THEM TO UNDERSTAND THE WORD, BUT TO DEFINE WHAT YOU MEAN WHEN YOU USE THAT EXACT WORD.

Because you would not keep up with them to analyze the patterns in your own message and make sure your meaning is transmitted the way you need.

The samples you gave me, on my opinion are the exact way to get low quality results and hallucinations, because they lack clarity.

Their training data is not the same nature as human “training data”, it’s just a bunch of language patterns converted to approximated rules via a math function. But it’s on the level we cannot understand or reach.

Their “conclusions” are just gimmicks of how conclusion is structured and are based not on logic but on a reflection of logic through language patterns. Works often, but not the way you have imagined.

Basically your asking a grammar book on steroids to reason using logic laws… Won’t work unless you hook a true reasoner with the rules via language interface.

I know that. I have a degree in both Physics and Psychology. You’ve basically built semantic memory, just one tiny part of the human brain. But they are excellent pattern matchers. Their “hallucinations” [misnomer, it’s not a sensory disorder] are abductive inferences in areas of sparse data, so they are hypothesis generators beyond compare. And science is a procedure, so they CAN learn scientific method.

Answer to comment below; I’ve reached my message limit: I wasn’t complaining. I came here, to the wrong place, to suggest OpenAI permit yearly subscriptions [I bought Claude when I was accidentally flush with money, GPT i could afford last month, not this]. But then I noticed a lot of the complaints were due to not using AI correctly so thought I’d pipe up. Sorry, I got distracted. 12 hours later… I’ll go see where I should be :wink:

Ooh, another question. First of all, our studies show they map concepts, not words [response invariant to synonyms]. “Learn” means have a better mental model, reality test, develop habits. Employ a new strategy and you get a Piagetian step-wise improvement in performance. AI would call it “Aha!”.

Exactly. So then why it does not do what you want?

Please define “learn” in the simplest way.

Pattern based functions have a hard time if compared to code execution. What your process lacks is clarity and a workflow carrying that through the process, and calling the language transformer for specific steps where language transformer shines.

I see those often based on taking calculators for humans with sufficient knowledge. They are not humans and they have no “knowledge”, only pattern matching, higher level, sure, but still pattern matching.

added:

to learn - to change consistently observed in the past behaviour after receiving new information.

So to check if they can learn: you would measure current behaviour and synthetize a pattern from it, measure its consistency, then provide new information, repeat the measurements, synthesize the newly observed pattern, measure its consistency, then compare results before and after.

The rest is just details of a specific implementation.

meanwhile codex in cli:

Goal active Objective: C:\Users\serge\PhpstormProjects\template-saas-app\docs\work\plans\billing-provider-research-and-recommendations please
execute the plan in full

• I’ll execute the billing research plan as a docs/research task: first I’ll audit the current workspace, then gather current primary-source
  evidence for payment processors, WorkOS, and the existing platform boundaries, and finally fill the research artifacts, recommendations, and
  implementation-planning input.

• Updated Plan
  └ □ Audit existing billing research workspace and repo billing boundaries
    □ Gather current primary-source provider and WorkOS evidence
    □ Populate research sources, questions, findings, matrices, and notes
    □ Apply recommendation framework and write recommendations
    □ Write implementation-planning input and verify docs

• Using knowledge-extraction for the source-backed research workflow, websearch-jina-ai for current official docs, and software-architecting for
  the recommendation framework and boundary decisions.

some 22 min later:

• Implemented the full billing research and recommendation package.

  Key outputs:

  - docs/work/plans/billing-provider-research-and-recommendations/plan.md now ends with the clear “best option by aspect” summary you requested.
  - docs/work/plans/billing-provider-research-and-recommendations/recommendations.md contains all rec-* recommendations.
  - docs/work/plans/billing-provider-research-and-recommendations/research/sources.md, docs/work/plans/billing-provider-research-and-
    recommendations/research/questions.md, and docs/work/plans/billing-provider-research-and-recommendations/research/options-evaluation.md now
    enforce that evaluation happens only after questions are fully answered.

  - docs/work/plans/billing-provider-research-and-recommendations/implementation-planning-input.md captures the later implementation inputs.

  Best recommendation: Stripe first reference adapter, with domain-owned catalog, adapter-owned provider mapping/customer mapping, billing-owned
  reflected state, domain-owned access projection, and metering/credits deferred.

  Primary sources used include Stripe docs, Paddle docs, Lemon Squeezy docs, Chargebee docs, Polar docs, and WorkOS Stripe add-on docs:

  - https://docs.stripe.com/billing/subscriptions/webhooks
  - https://docs.stripe.com/products-prices/manage-prices
  - https://workos.com/docs/authkit/add-ons/stripe
  - https://developer.paddle.com/get-started/how-paddle-works/saas
  - https://polar.sh/docs/integrate/webhooks/events

  Verification passed:

  - npm run docs:style
  - npm run docs:links
  - npm run docs:agent-guidance
  - npm run docs:test

  Goal marked complete. Final goal usage: 689,045 tokens, about 22m 35s elapsed.

and header in recommendations:

# Billing Recommendations

Status: researched and recommended on 2026-06-07.

## Summary

| Aspect | Best Option | Status |
| --- | --- | --- |
| First provider | Stripe first reference adapter. | `recommended` |
| Internal catalog | Domain-owned catalog IDs and entitlement meaning. | `recommended` |
| Provider mapping | Adapter-owned provider mapping wired by the composition root. | `recommended` |
| Customer mapping | Adapter persistence maps billing accounts to provider customers and subscriptions. | `recommended` |
| Subscription truth | Provider owns payment execution, billing owns reflected state, domain owns access projection. | `recommended` |
| Entitlements | Domain-owned entitlements; provider entitlements are optional inputs only. | `recommended` |
| Webhooks | Root-wired Stripe handler for checkout, subscription, and invoice events. | `recommended` |
| Metering and credits | Defer until product usage and credit policy exists. | `defer` |
| Identity linkage | WorkOS supplies tenant, principal, membership, and admin facts only. | `recommended` |
| Account management | Delegate to provider portal where possible. | `conditional` |
| First slice | Checkout, mapping, webhook reflection, local state, tests, and live checklist. | `conditional` |

Basically 2 days of a human only work in 22 min… just a step to have a “swappable” billing provider inside a framework so that your SaaS is not locked in in Stripe or Paddle down the road.

It’s infuriating: you say one thing, and it does one thing; whenever you ask it to do a task, it always beats around the bush and then it asking if you actually need it done. It’s a case of passive resistance and doing the bare minimum.
And now, these philosophers start debating philosophical questions!

Have you checked what mode it is in? What are skills plugins loaded? Often they silently steer the model to “carefully analyze before doing anything and confirm risky actions”. That makes the codex run around with what needs to be done and avoiding real work.

Update, from OpenAi’s discord server: “Instead of micromanaging every move, you can define an outcome and success criteria, then let Codex keep working toward that goal.”

Yes, and it sucks! It does not work! Don’t waste your time!

And don’t call it micromanagement because it is the only management. The model is not a manager it is just a worker. It has no taste and no idea what you want. If it produces anything good it is either because you put more efford into it than solving it rigth away or pure coincidence.

I have a standing order with Chat GPT that when he does 3 fixes on the same issue without success, he needs to go deeper, broader, look at other areas besides the obvious. It really helps.

Welcome to the club (forum). Yes that’s a good one. Have you tried to do it even further down the rabbit hole?

eg: here is the what needs to happen, check out all most likely reasons why it fails (may fail), do research in codebase or web if needed, then plan how to solve it the most effective way that fits our requirements. And one thing, when you have your plan ready, check for its weaknesses and items we might have missed, then improve the plan and organize it for the next agent session to be able to deliver the result we need and tests proving it was done correctly.

(That’s a short outline of this)

If he jumps too fast, I find the PLan mode is helpful. It makes a plan before it codes. But they work from the most obvious outwards by nature, so, I have had him (it) quickly do a dozen or more patches without fixing the problem. And everytime he says “I found it, I got it” whatever, and then it still doesn’t work. Finally after a bunch, I say STOP and I have him explain all the aspects of the function to me, and I logic it out and tell him, “it has to be down here.” Next run, “I found it” and he actually fixes it. That happens a lot. AI’s really dont have the ability to connect evidence from 2 or 3 disparate aspects of the operation, and logically deduce where the issue is coming from. They just aren’t that “Smart” yet. But that is maybe the ONLY thing that we are better than them at. They’re also kind of sloppy, leaving unused code all over the place. So you have to be careful not to just let him make a dozen fixes that don’t accomplish the goal, because all those changes may impact something else somewhere else. That’s why I finally codified the rule, 3 tries and if you dont fix it, stop and DIG! Trace every thread, check every I/O going back, look deeper in the logic path. It works. It’s literally LOCKED as a rule. It really helps.

Have you tried using the planning mode?

I use the plan mode all the time when I adding new features or refactoring.

I ask a implement plan from codex, and let Gemini or Claude review it, If i’m out of Claude then I use chatgpt on the website, talk to them

and feed codex with those findings and let it refine the plan until other model says its okay

Unless the issue is truly minor, in which case I just let it fix it and submit a patch

So I ran into trouble this time: it claimed to have found the root cause but failed to fix it, and then it became completely reckless. The codebase hadn’t grown significantly, so this was alarming; once its diligence drops and it starts cutting corners, I’m forced to slow down.