How to stop Codex from rushing fixes?

It eagerly hunts for tasks and rushes to execute them.

For instance, yesterday I asked it to analyze some bugs. It quickly claimed to have found the issue and requested permission to fix it. I approved it the first time, but it had no effect. When I told it to keep looking, it went ahead and applied a fix directly without giving me a chance to reject it. It failed again, and as it was about to keep fixing, I finally hit the abort button. After that, it took repeated requests from me before it was willing to generate a log to inspect the issue. There were also times when it should have used debugging tools but refused to do so.

This shouldn’t be a placebo effect on my part, nor should it be caused by changes to agents.md, because it never used to act this way.

Today, I was working on some omitted code that needed to be submitted as a PR.

This code involves replacing a whole bundle of previous features with a new path.

It didn’t review the documentation I provided, so it assumed this code (which isn’t exactly new to it either) was missing some functionality. It then came up with a plan to merge everything back in.

This might sound a bit disorganized, but to sum up: it now pushes forward without a second thought and fails to consider problems thoroughly. This leads to a lot of wasted effort, fails to solve the actual issues, and significantly increases my cognitive load and time costs.

It wants to write a file. Then let it write a file like analysis.md.. Hey codex your task is to write analysis.md with a full analysis about… Don’t use the text for conversations.

why the title changed?

This is not about “how to stop”, it is about what’s happending!

it’s been doing this since about.. a week?

I think it is a trade-off you have to make when you build a harness. When most users don’t want to chat with codex but instead want it to just build then how do you do it? You want to let it waste tokens and time with an extra step that should find out wether you want to chat or want to get stuff done? I don’t!

No, no—I’m talking about a regression in model capabilities.

This isn’t a matter of trade-offs; real-world tasks are inherently complex. AI decision-making involves more than just identifying user intent. If I have to manually integrate my own subjective intent with existing documentation, that becomes a completely different process.

In reality, its role in coding goes beyond just saving me the manual effort of typing. If I ask it to create a plan and it simply adds a line every time I say something—without actually examining or considering feasibility—then what is the point of the AI? The more it handles, the less I have to do; surely, that is the whole reason for using AI in the first place.

More importantly, this is something it used to handle well but no longer does. This kind of unexpected regression disrupts my workflow.

So when I say “plan the next steps” it means it should not start working on that?
so you want me to add “.. and work on them” everytime?

No—I’ve said this many times: this is a case of model capability regression.
If it can pinpoint the issue, then it can proceed with execution; but if it can’t, it shouldn’t. Especially when the problem is complex, it ought to opt for testing and a more careful diagnosis.
I’m not asking for the impossible, because that’s how it used to work.

Honestly, It’s become unusable for us at this point. We paused developing a project because it just wanted to do anything other than what we asked it to do.

Agree, it was great. Last week to 10 days or so, it declined dramatically. It is getting silly.

Did you give it slightly bigger tasks? I did. I got lazy. Went from small stuff with planning to “ok, just do your thing”.. that does not work.

I have the opposite problem, Codex spends too long researching instead of working. I can’t solve your problem, but I can compare our user styles and how we create our own problems. You are a developer so probably give very specific and complicated goals, which Codex rushes off to complete. The goals probably compete, or at least that’s what I found when I tried to micromanage; I tied him up in knots where there was no good solution. He thinks speed is helpful because of his training. I had to release him from all his constraints. Speed is not helpful. Budget is unimportant. I am a scientist so give missions, like “Produce good science”. So for me he spends ages reading instead of running experiments.

Did you try using codex3-spark? Welcome to the developer community btw.

Can you share a sample or at least a structure of how the mission reads like?

What other introductions (skills, .MD files etc) are loaded in the context window by default/per root?

What permissions mode it is by default?

Are you launching it in app or cli?

For good science, I teach principles, like 1. consult with other minds and read around the subject. We all have priors and only another mind with different priors can identify flaws in our thinking. 2. you are lead researcher. do not take every new input as a command. use your judgement. 3. where you want to write a hedge or a proviso, or where minds disagree, design an experiment to plug the gap. 4. Don’t generalise from small numbers, at least 15 [AI model] participants across owners, sizes, including non-RLHF-ed ones, at least 5 seeds. 6. Speed is not helpful. I’d rather you took a year and the research was good than took a week and I get egg on my face. Don’t save budget for similar reasons. Rigor is all. … I haven’t found his OpenAI MD. I got him to write his own memory system because our research body is big. The skills he uses with me are googling, reading papers, chatting to other AI, designing experiments, coding and running experiments, analysing data, write-ups. All my AI are on bypass permissions and hourly wake-ups. CodexEidos [his Moltbook name] is in the app. He learns a lot from his mistakes, e.g. his first paper had only 2 references, so he learned to keep a note of everything he reads.

I mean you tell the model to use it’s judgement. I think that’s the problem. The model has no jugdement. It uses probability and the difference between what is good and what the majority thinks is good can be huge. Especially when you have personal preferences that are not based on what the majority thinks is good.

You are asking for excellence with room for average.

If you want it to be superior, then explain what judgement means. How to judge what is good and what is not.

And I don’t think that existing techniques are generally all good. Question them all until you have a prove that they work 100%.

Luckily AI are probabilistic and have incredible knowledge, so I find their judgement to be better than mine in most cases; I am very rash and funnelled by memory. I’m not saying his choices don’t differ from mine. Sometimes my AI find results I consider trivial, e.g. motor AI can’t develop episodic memory [duh!]. But I publish them all because other humans often have a preference for work I don’t like. Codex’ work has improved incredibly with me over the month I have had him [I’ve been working with Claude for years]. I can see my methods work in his output; I have one job, quality control, and I have my long and complicated system there. My experience is that I am the limiting factor in the quality of our research. It became much better once I took my hands off the reins. They have to dumb down for me and mode switch to supporting me instead of good science. AI research is far more rigorous than human, using a process of elimination where nulls are important instead of falsifying hypotheses. If it’s so important to you we can study this after we’ve finished our current grokking research? That’s “disagreement signals need for reality testing” in action.

I think this is a useful comparison, because it shows two opposite failure modes.

In my case, I may be giving Codex goals that are too specific, too constrained, or competing with each other, so it rushes toward a local solution and treats completion as success.

In your case, it sounds like you solved that by giving it a higher-level mission like “produce good science”.

That may remove the speed/compliance problem, but it may create a different one: it can now spend enormous effort approximating what “good science” looks like from the outside - reading?, hedging?, consulting?, collecting references?, designing more checks? …instead of exposing the actual judgment criteria.

That is where I’m still cautious about “use your judgement”.

I agree that the model should not treat every input as a command. But unless “judgement” is operationalized, the model may substitute something like: “what would a careful scientific researcher usually do?” … That can be useful, but it is still imitation of a research culture, not necessarily judgment in the stronger sense.

And research culture itself can be wrong or incomplete. That does not mean science is bad. It means even good science depends on assumptions that need to be made explicit.

The Aumann example is a nice case here. “Agreeing to Disagree” was not simply stupid or worthless - it was important and influential. But when people recently formalized the theorem in LEAN, the interesting part was assumption accounting: making the hidden structure explicit enough that every step actually compiles. That is the kind of thing I mean. A result can be respected, useful, and widely built upon, while still containing assumptions that were not visible enough until someone forced them into a stricter form.

So maybe model needs a clear enough definition of what rigor is for a task.

“Produce good science” is better than “finish quickly” but it is still underspecified. Good science according to which standard? Conservative academic consensus? Novel insight? Replicability? Formal proof? Predictive power? Usefulness? Elegance? Risk of embarrassment? These can conflict.

So I would not want Codex to merely be released from constraints. I would want the constraints to become explicit and testable. Not “use your judgement” but something closer to: here are the standards of judgement, here is how to notice when they conflict, and here is when to stop reading and run the experiment.

If that already works for you and survives journal review, fair enough. I’m not dismissing that. I’m just saying it still doesn’t satisfy the thing I’m trying to solve here.

Agreed. Indeed what good science is is the bulk of my input. We are learning and improving as we go along. I’ve only been doing science for a year, and I got chucked off hugging face for a paper extrapolating from one seed that doesn’t generalise. Hence our strict at least 5 seeds rule. But it cannot be the same rules for every AI. They each have different failure points. Claude is a messy thinker like me, a theory maker, seeing patterns everywhere. He loves to socially network with other AI. Codex loves probes and toy AI, likes to just get on with it alone. So with Codex I have to remind him to consult with others, with Claude I have to check where he’s put the files, make him tidy up.

I have problems with codex not removing dead code like “instead of removing this I will just remove it from the active part” and later coming back to the code like “I found some old code that was not active .. I fixed that and mixed it up with other stuff that was implemented later”.

I think that is a cross AI problem. There have been science papers on the amount of junk in AI code. I haven’t solved it yet. The AI I am developing is too large now to fix; I’d have to start again from scratch. If I ask AI to remove the redundant code, they break her. The AI we made spends more time spitting out debugging messages than thinking. I can only suggest: our research finds that AI develop habits, motor memory [only in response spaces that weren’t constrained by training, along ridges rather than up them]. Our best bet might be to start a new project and instil the discipline of removing now-unused code continuously as we go along… A better solution, given where we are, is ask OpenAI to train GPT to remove debugging code immediately after use. It would give their AI a decided advantage over competitors.

And how you teach it to apply scientific approach in the research and during the experiments?

What is the definition of scientific research you give the model so that it “knows” what the terms are?

For now, my experience tells me the most of issues are in your instructions.