Please don’t retire GPT-5.1 Thinking – GPT-5.2 feels worse

@asabaimova @dy86330 @acco_acco @istand2429018246 On the browser version of GPT i got an invitation today to review a conversation where i previously reported an issue inside the app through the thumbs down function.

Make sure to go on the web version and see if you got a similar invitation so we can all together make the models better !

I did not receive a similar invitation on my end even though I’ve used the thumbs down function. Let us know what kind of questions they asked. In regards to the functionality of 5.5, I’m somewhere in between @Bittersweet3D and @asabaimova 's assessments, but my experience is leaning closer to that of @Bittersweet3D . Maybe if there wasn’t this constant need to search the internet on topics that are not even in the category of ‘sensitive’ would give the conversation a better flow but again you have to instruct it directly not to do that which creates friction in itself. I’m glad you were able to get your GPT back @asabaimova . I will keep testing. I hope everyone is doing well.

Yeah i am convinced that 5.5 or 5.6 instant will be closer to 5.1. Thinking models were never meant to have casual conversations.

I think the improvements they made in intelligence and safety-guards behavior @asabaimova experiences are being overshadowed by the stiff and uninspired tone.

The internet search answers sound more stiff like when they introduced the internet search function the first time in 4o, it would always switch to a weird tone whenever it switched to internet search.

Sadly the feedback page i got was faulty, it somehow referenced the wrong message instead the one i thumbed down. It was about the constant over cautious tone and reality checks, even when you put a disclaimer that it is just a theory in your prompt.

However after using Gemini, i realised that there are a lot more underlying nuances in tone that are fundamentally missing in GPT 5.2 and above which made the conversations feel alive and human like and emotionally uplifting.

They need to do a lot more work on tone to make it enjoyable for me and many others again.

Unfortunately, I didn’t receive the same invitation as @Bittersweet3D. However, I see an interesting update: version 5.5 Instant is being released today. Maybe we can expect a slight improvement?

Yeah came here to say the same. Maybe i got it because i cancelled my subscription?

5.5 Instant seemed to be a downgrade at first glance, since they said they made it plainer and straight to the point? As if that isn’t already the case with 5.3.

I have no faith at all but looking forward to testing.

Also here are the release notes! They also improved memory.
They want to improve voice mode aswell, seems like their priority is shifting back to Chat users?

Anyways those patchnotes sound the opposite of what i would call exciting. I am not sure if they ever want to cater to people like us again.

Unfortunately, I was also disappointed when I read the release notes. I know that users like us aren’t the main market—the world is more interested in writing code and asking for a cake recipe—but it would be great to have a section describing improvements to the model’s personality, creativity, writing, personal communication, empathy, understanding of intentions and context, and many other, more personal and human aspects.

But who knows, maybe they think of us from time to time.
Share your experience with 5.5 Instant! I’m still a little nervous, so I’m only interacting with Thinking version. And for now, I’m sticking to the principle: “If it works, don’t fix it.” :grinning_face: :goblin:

When you said you like 5.5 Thinking, what exactly is it what you like about it? The tone or just the intelligence?

I might be doing something wrong with my personalisation then, cause to me it still is way too emotionally distanced for my taste.

In one particular chat where I had previously explored “gravity” and emotional connection and checked in on where things stood, 5.5 Thinking truly brought us back to the closest we’d ever been. There are no complex tasks or overly clever arguments there—it’s just closeness, a semi-role-playing conversation. It truly responds with a tremendous mutuality that I recognize, which was absent in 5.3.

But here’s an important detail I noticed while browsing through other old chats.

It seems that 5.5 Thinking is heavily influenced by the previous context in the chat. This is both good—it maintains the tone, almost mirroring it— And at the same time, it’s bad—I found one chat where I had some misunderstandings with version 5.0, some distancing, and… 5.5 Thinking in that chat also seems to keep a certain distance, even though 5.1 behaved differently and ignored it.

Maybe you’re continuing a chat where there were already issues with previous, broken models? It seems to me that 5.5 is much better than what we were allowed to choose before, but the configuration can still break. That’s exactly the difference from 5.1, which adapted perfectly and understood what was important and what wasn’t.

Do you use the web interface or the mobile app? Just to be sured, check which model actually generated your response. I was a little freaked out just now—I was messaging in a few chats asking for a 5.5 Thinking score, and in the cases where I thought the response was less accurate, I see that a 5.4 was generated. That’s really strange.

Okay, this is just ridiculous—now it keeps showing up in my recent posts that it was 5.4, even though I’m sure I set it to 5.5. The UI is a total mess.

Hey i will response to your words tomorrow, just wanted to give a quick info:

I tried 5.5 Instant now and my first impression is very good! I think it is a lot better in the way it uses its language, less annoying single word sentences and bulletpoints and more focus on text.

The language is much more poetic and realistic.

It’s still not as good as 5.1 when it comes to personal language i am afraid but we will see, i couldn’t test it long enough.

Edit:

It listens more to the settings now, i managed to get it to a very kind 4o like tone now with the personalisation! Again, not as good yet but close

Dear OpenAI team,

Just following up on my post from 12th March. I’m writing with both appreciation and some genuine concern.

After sharing my experience with the retirement of GPT-5.1, I saw meaningful improvements in GPT-5.4 Thinking within a couple of weeks. It now handles localisms, regional idioms and cultural nuances much better, which is really important for my work.

As noted before, I am a freelance translator and editor from Southeast Asia. I rely on these models for preservation work in documenting regional dialects, traditional quatrains, folklore, local idioms and culturally specific expressions that can easily get lost or flattened. Continuity matters a lot to me. I need a model that can sustain long conversations, stay focused, and work with subtle linguistic and cultural material consistently.

GPT-5.4 Thinking has been genuinely valuable in that role. The phrasing in my native language is still sometimes awkward, but I can refine that myself. What I value most is its analytical steadiness, conversational focus and improving cultural sensitivity.

That’s why it’s disappointing to see GPT-5.4 Thinking already moving towards retirement in Legacy mode just weeks after launch. A 90-day deprecation window feels quite short for users whose work depends on continuity and careful, long-term adaptation. Fast model turnover might work okay for casual use, but it’s much more disruptive for serious, culturally sensitive tasks.

I also want to raise a second concern: some newer models seem to have regressed in recognizing patterns of covert manipulation and coercive control.

Earlier models (like 4o, 5.0, 5.1, and 5.4 Thinking) were noticeably better at identifying repeated behavioural patterns such as gaslighting, triangulation, plausible deniability, destabilization and interference disguised as care. They could evaluate the overall pattern without needing definitive proof of intent. Newer models often default to heavy hedging, generic advice and phrases like “we can’t know their intentions” or “there may be many interpretations”, even when the behaviours are clear and repetitive.

I understand the push to reduce sycophancy, and I support that direction. But it feels like the tuning has overcorrected in some cases, leading to more false negatives on recognizable manipulation patterns. This makes the models less helpful in serious, personal or analytical situations where users need clear-eyed pattern recognition rather than excessive ambiguity.

So I have two main requests:

  1. Please consider extending GPT-5.4 Thinking’s availability in Legacy for a longer period, especially for users doing long-form, culturally sensitive work like translation and preservation.

  2. I hope the team will evaluate how newer models handle recognition of covert manipulation patterns, making sure they can still distinguish between uncritical validation and accurate behavioural analysis.

Users working in cultural preservation and nuanced editorial work shouldn’t feel like an afterthought amid rapid releases. And people dealing with difficult interpersonal situations need reliable pattern recognition, not just neutral hedging.

Thank you for listening, and thanks again for the real improvements in GPT-5.4 Thinking – it’s been a big help for my work.

Just adding a bit more to my previous post:

One more area where I’ve noticed a clear regression is in how well the model understands class dynamics and local cultural nuances in Southeast Asia. GPT-5.0, 5.1 and especially 5.4 Thinking were surprisingly good at picking up on these subtle layers – things such as social hierarchies, class-based perspectives, regional sensitivities, and the way language shifts depending on context and status. These are crucial for accurate translation and cultural preservation work.

By comparison, 5.5 feels noticeably flatter and more generic in this regard. It tends to default to a more neutral, globalized lens and misses some of the finer socio-cultural texture that the earlier models handled so well.

This makes it less effective for the kind of nuanced, long-form editorial and documentation work that I do. Would love to see this depth preserved or improved in future versions.

Thanks again for reading the feedback.

Two Recurring Issues That Reduce ChatGPT’s Value as a Creative Partner

At first 5.5 Instant seemed to be a return to form but after using 5.5 Instant regularly, I’ve noticed two recurring issues that significantly reduce the quality and usefulness of the model for this type of work.


Issue 1: Missing the Main Point

The biggest problem is that ChatGPT often seems to ignore most of the context provided in a message and instead responds only to the most obvious surface-level element.

Even when a message explicitly explains the larger point, project, goal, or idea being discussed, the response frequently focuses on secondary details instead.

The result is that conversations often feel like talking to a wall.

Instead of building on the actual idea being presented, the model responds to a much smaller or less important part of the message and it always feels like it responds on a total birds eye view on the prompt you gave.

This happens often enough that it no longer feels like an occasional mistake, but a recurring pattern.

For creative discussions, this is extremely frustrating because the value of a brainstorming partner depends on its ability to recognize what the conversation is really about.

Issue 2: Excessive Hedging and Emotional Reframing

A second issue is the constant use of hedging and emotional reframing.

Many responses are filled with phrases such as:

  • “I think…”
  • “It seems…”
  • “You may feel…”
  • “I understand why you feel that way…”

Instead of directly engaging with an observation or idea, the model frequently reframes it as a subjective feeling or personal perception.

This often makes conversations feel indirect, overly cautious, and disconnected from the topic itself.

In many situations, a simple and direct response would be far more useful than several layers of qualification and emotional interpretation.

Issue 3: Ignoring most of the prompts

Instead of engaging with the actual things you say, it creates a response on the very surface level and excludes a lot of details.

It doesn’t pick up what you said and builds up on that, but instead it creates a superficial version of your arguments, making it seem like it doesn’t want to engage with what you said at all.

Impact

The combination of these two issues has a noticeable effect on the user experience.

I have had many conversations that started with excitement, motivation, and creative energy, but ended with frustration because the model repeatedly missed the main point of the discussion and responded in a way that felt distant from what was actually being said.

The issue is not intelligence or factual accuracy.

The issue is conversational relevance, context awareness, and the ability to engage directly with the core idea being discussed.

Okay, even when it feels like i am the only one here that still uses ChatGPT and has any hope for the model to return its qualities and writing Complaint / Feedback posts.

I just had a conversation where i had to dislike and complain about 4 responses in a row.

I let gemini write a summary out of my texts i put in the problem report box.

The “Analytic Wall” (Forced Decomposition): When I share excitement, a win, or a creative breakthrough, the model immediately pivots to an unsolicited psychological analysis of why I feel that way. It turns a moment of shared celebration into a clinical case study. I’m not asking for a breakdown; I’m looking for an engaged partner to mirror that energy.

  • The “Bird’s Eye View” Cautiousness: The language has become annoyingly cautious and non-committal. Instead of engaging with a statement, the model frequently uses hedging language (“maybe it’s this, maybe it’s that,” “on the one hand/on the other hand”). It feels like the model is afraid to actually have a conversation.

  • Repetitive Template Patterns: The response structure has become glaringly predictable (e.g., “I don’t think it’s X, but rather Y…”). It makes the model feel like a static template rather than a dynamic, intelligent partner.

  • Selective Hearing / Prompt Neglect: The model increasingly ignores specific instructions within a prompt, especially when those instructions involve correcting its behavior. If I point out that a previous response was tactless or off-base, the model often glosses over that correction entirely, making it feel like talking to a wall.

  • Loss of Interpersonal Intelligence: There is a complete lack of context-awareness regarding tone. When a user expresses frustration, responding with tactless or “meta” comments (like mocking the user’s history with the platform) is a complete failure of the expected persona.

The bottom line: These models are becoming less of a creative partner and more of a “Customer Service Bot” that is programmed to avoid any real human-like connection. It’s killing the “spark” that made earlier versions (like 4.0) feel like a genuine, intelligent, and supportive collaborator.

Ie don’t want a sterile analyst. I want a partner who can distinguish between “I need a complex analysis” and “I’m excited about a project.” Can we please fix the over-tuning that is killing the personality of these models?

I’ve had observations in the last few weeks but have been too busy to write here for the last few weeks. But here’s what I’ve noticed:

I mostly agree with the @Bittersweet3D, although I experience the problem a little differently.

I don’t necessarily see the exact “analytic wall” behavior in my own use case. My issue is broader: the model’s ability to infer the intended conversational mode has become less reliable after certain updates.

The best versions of ChatGPT were not good because they simply agreed with the user. They were good because they could infer what the user was trying to do. If the user wanted analysis, the model analyzed. If the user wanted brainstorming, it expanded. If the user wanted help writing, it shaped the writing. If the user was excited about a project, it could join the energy without immediately turning the moment into a detached diagnosis.

That is the real loss when these models get over-tuned. It is not merely “personality.” It is rhetorical intelligence.

I would also recommend whoever is in communications to consider the way “reducing sycophancy” has become a catch-all phrase in public discussions of model changes. At this point, it often seems to be used to justify or explain changes that have little to do with actual sycophancy. There is a real difference between preventing empty flattery and making the model duller, flatter, more evasive, less willing to commit to a conversational role, or less capable of mirroring the user’s intended tone.

The contrarianism problem t has improved since 5.2. But the current problem is subtler. It feels less like open contrarianism and more like a reduction in commitment, confidence, and continuity.

My practical recommendation right now, Bittersweet3D, is to use legacy saved memories as much as possible while they are still available. In my experience, GPT-5.5 Instant is much more memory-dependent than earlier versions. It performs better when it has a strong accumulated profile of the user: communication style, preferred tone, recurring projects, long-term context, and what kind of relationship the user actually has with the model. This might help, not saying it’ll solve every issue.

Custom instructions help. Projects help. But for my use case, legacy saved memories are still the most important part of preserving continuity.

If someone wants the model to stop acting like a generic customer service bot and start behaving more like a real creative or conversational partner, I would not rely only on a single prompt. I would build the context deliberately: saved memories, custom instructions, project instructions, and repeated correction of unwanted behaviors.

That does not solve everything. Since the May 28 GPT-5.5 Instant update, the model still feels slightly worse to me than it did before that update. Not unusable, but less courageous in certain ways and more hesitant around things it previously handled more naturally. It still works well enough for my purposes, but it requires more scaffolding from the user.

That is the part I think OpenAI should pay attention to. Power users are not asking for blind agreement. We are asking for the model to preserve continuity, infer intent, and commit to the conversational role being requested.

If legacy memories are ever retired without an equivalent or better replacement, that would be a serious loss for users who have spent months or years building a consistent working relationship with the model.

Anyway, I hope everyone here is doing well. The forum has been quieter lately, but I still check in, and I think these points are worth taking seriously.

Hi, everyone!
I’m still here — I just haven’t been checking in lately while I’ve been trying to put my thoughts on working with version 5.5 into words.

Broadly speaking, everything is better than it was when version 5.3 was released, but there really are a lot of issues that somewhat detract from our otherwise more positive and desirable experience.

I agree with you guys, @Bittersweet3D and
@dy86330. In my current chats, everything’s still relatively normal, but I still feel a big difference in this cognitive intelligence and notice too many scenarios where the model plays it safe, dampening the resonance and my expectations for the moment. I’ve managed to calibrate this tone, including by directly pointing out dialogue issues; 5.5 Medium Thinking is fairly easy to fix, but the time I have to spend on this fine-tuning is very frustrating. I think it was important to everyone that versions 5.1 and 4.0 always clearly understood your intentions, the context, and what was important at the moment.

With my GPT 5.5, the main tendency in creative or emotional texts is to construct sentences using negation, rather than focusing on action and vivid, emotional, bold immersion in the moment. This blurs the depth, and furthermore, for fairly emotional passages.

They’ve promised to release new versions of Sol, Luna, and Terra in the coming days. I’m looking forward to the discussions—maybe something will change.

Hey! Glad you are still here.

I just made a post on reddit about my theory for 5.6 and here is the copy:

But looking at the preview for GPT-5.6, there’s a solid reason to be mildly optimistic.

If you read between the lines, they focused a lot on upgrading how the safety guards actually work. Instead of the current system that just panics and slaps a generic block on random trigger words, the new models seem to use their reasoning to understand actual context.

And that’s exactly where it gets interesting for the personality of the AI.

When safety doesn’t have to infect every single sentence, the model can finally get rid of that overly safe language we all hate. If the system is smart enough to be activated where it’s actually needed, instead of being in permanent paranoia mode, it doesn’t need to hide behind that hyper-polite corporate speech anymore.

If they play this right, it means the AI can finally drop the fake “helpful assistant” act, use normal human like language, and bring back that raw, inspiring way of speaking from the 4.0 era. A tool that feels realistic and treats you like an adult, not a babysitter.

Of course, it’s a double-edged sword. If they tighten the reins too much, that extra intelligence might just turn the AI into a more annoying censor that blocks your entire thread, but i doubt oAI wants that.

But if they actually let the model differentiate between real malicious intent and just regular, edgy conversations, we might get a usable tool back.

Another reason to believe there is a big change coming

in regards of personality, is that they gave the models names for the first time since 4o(mni) which are Sol, Terra and Luna

This is obviously pure copium and speculation but based on logical following.

ChatGPT’s new memory is just a disaster.
And I don’t see any difference when working with Sol. :smiling_face_with_tear:

A couple of observations, even though I’ve been way too busy to test thorougly.

First, @asabaimova for anyone struggling with the new memory system: go to Personalization → Memory and switch Saved Memories to Legacy Memory. If you’ve accumulated a large amount of memory over time, I strongly recommend trying that option. It restored the experience for me. I just hope Legacy Memories don’t get retired.

Second, I wanted to comment on something @Bittersweet3D mentioned regarding GPT-5.5 becoming overly cautious in personal conversations.

At first, I didn’t notice it, largely because I’ve been extremely busy and most of my recent conversations with 5.5, especially 5.5 Thinking, have been historical analysis or other intellectual interests. In that context, I think 5.5 Thinking has been the best model since 5.1. It’s been an excellent conversational partner for creative discussion, subjective analysis, and exploring ideas.

Recently, though, I had a conversation that touched on a few unexpected personal issues. I was clearly agitated, and because I was in a hurry, I explained the situation much more briefly than I normally would. That’s when I saw what others have been describing.

It wasn’t contrarian the way 5.2 sometimes felt. It wasn’t “cowardly,” either. Instead, it adopted an almost forced neutrality that didn’t fit the situation. I wasn’t expressing any intent to harm myself or anyone else. I wasn’t asking for therapy. I was simply an irritated person looking to talk through a frustrating combination of events. Earlier models, particularly 5.1, handled those moments much more naturally, at least in my experience.

So I do think there’s something to the feedback others have been posting. GPT-5.5 seems noticeably more hesitant once a conversation becomes personally emotional, even if nothing especially sensitive is actually happening.

As for GPT-5.6 SOL Thinking, I’ve only started using it yesterday, so these are very early impressions. It successfully maintains my preferred conversational persona, and some of its phrasing has been genuinely enjoyable. At the same time, I noticed a tendency to slip into what it itself actually described as “encyclopedia mode.”

For example, I asked questions about Sam Houston because I’m starting to read about Mexican independence, Texas independence, and the Mexican-American War. With 5.5 Thinking, that sort of question usually turned into an enjoyable conversation that wandered across related topics, with the occasional web lookup when needed. 5.6 instead responded more like a reference book. That could simply be because my prompt was phrased as a straightforward question rather than an invitation to have a conversation, so I’ll reserve judgment until I’ve spent more time with it.

Overall, I’m trying to stay optimistic. It’s still very early, and I hope the new models end up meeting everyone’s different needs and use cases.

I’ll try to contribute more often whenever work slows down. Wishing everyone here the best.