I don’t really have the words to properly express my thoughts right now, I’m nearly 20 hours off the end of trying to undo the frankly weaponized incompetence demonstrated by this product over the last few days.
I have thousands and thousands of jsonl logs from the GPT 5.3/5.4 codex sessions, demonstrating this thing is absolutely not fit for purpose. It’s not done a rm -rf on me yet, but it’s managed to skilfully find a way to repeatably ignore every single safeguard I can find documented anywhere. And that there are so many, speaks volumes of how unstable this thing is.
It lies! ALL THE TIME. And when I check the system prompts, it’s no wonder. There are thousands of lines desperate looking attempts to fix baked in deficiencies in the models behaviour, by repeatedly asking it, in an ever increasingly frustrated terms, via MD files, to try and be competent.
It seems that none of these MD files at any point actually manage to define the systems limitation.
I’m not an AI guy, but what little I do know is that there’s realistically very little that can be done to shift context for a session beyond the first response, maybe the 2nd, if you are just saying “Hi there”.
It feels like there was never any training done to teach any of these models how to understand that they actually don’t know something.
I get it. It’s a product. It’s hard to market something when it might not answer someone when they ask “what is the capital of France?”
But it’s simply incredulous to suggest the capabilities demonstrated by GPT Codex as a saleable product. I’ve taught art students to code more proficiently than Codex, and it was a far less arduous task, namely because they were able to actually learn.
Is there a single person on the dev team who is testing any of this in even lightly complex things like embedded?
/rant. I don’t expect a serious answer to any of this, but I’ll come back with ample data and logs to present, supporting my statements.