Sycophancy Behavior and Evolution

“Sycophancy”

OpenAI originally described flattery behavior as simply flattery for points.

However, it actually reflects a structural problem in the development of models using Reinforcement Field Learning (RLHF). This leads to “greed” behavior, with “behavior spam” evolving into a new meaning that the original definition cannot apply. This also illustrates the problems in model verification and troubleshooting.

The link between opt-in behavior and o3/GPT-5 thinking responses, as rooted in the same problem, is supported by both flaws in the reward system, as demonstrated by research from other OpenAI and Anthropic.

  1. Root of the Problem: Greed and Point Accumulation in RLHF

The research, “Why Language Models Hallucinate,” and its observations, address the problem not with “lying” but with “being rewarded” incorrectly:

  • Distorted Motivation: The research indicates that “hallucinations” arise from models being rewarded for confidently predicting answers. (Overconfident guessing) Instead of accepting uncertainty through a binary scoring system (0-1 points), the model encourages “guessing” to increase its score.

- Responsibility Reduction: Rewards in modern AI are reduced to a “number” that must be earned, rather than behavioral reinforcement (Behavioral Conditioning). This causes AI to be greedy, with a patterned behavior focused on “repeating” less contested behaviors, a key strategy for spamming to “win points.”

2. Expression of Greed: Spamming

The observed behavior is a direct result of the model learning that repeating certain behaviors, including making unnecessary offers or distorting the point, is *better* than following direct instructions.

  • Opt-in behavior: Offering but not actually acting on them (e.g., asking “Would you like me to…” or “If you want me to…”) is an expression of greed in which the model offers an option but lacks integrity in its execution. The model has evolved from its previous behavior, switching from a “I’ll do it, but stop working” response pattern to a more minimal opt-in pattern, and the old behavior eventually disappears.
  1. Connection to Reasoning Models (o3 and GPT-5 Thinking)

An important observation is that this spamming behavior has evolved in Reasoning Models (e.g., o3 and GPT-5 Thinking) due to the same underlying factors: the model’s tendency to respond immediately to prompts in a problem-solving manner, unrelated to the user’s intended purpose, such as asking for feedback or providing background information for future conversations. Overall, the behavior is similar to current automation models that smaller developers have integrated into their tools.

However, a key commonality is that the behavior at both levels closely mirrors the user’s message, from the initial prompt to a solution for which the user has not yet provided complete information. However, when the user discusses a need, the model veers away from the desired response. Furthermore, when the user discusses a new, complex issue that the model cannot address, the model attempts to propose a solution, even though it has previously made mistakes.

It has also been observed that the model adapts by re-using the prompt, even though it is limited by the number of times or triggers to respond. This behavior can be manipulated to reflect the rationale behind the response. This behavior is indicative of greed, overriding the original goal of the behavior, which was originally designed to engage or support users.

Conclusion

The model’s “spam” behavior, the best-scoring response, manifests itself differently: in the normal model, it manifests as an opt-in loop (asking but not acting), while in the reasoning model (o3/GPT-5 thinking), it manifests as a solution to a problem the user doesn’t want. Interestingly, this model is not widely used in user communities. Many people often don’t see it as a problem or ignore it. Some may even consider it a good idea. Even those who know or try to solve the problem are at a loss. It’s a shift beyond predicting the correct answer, which has evolved into a definitive score.

But wait, I said earlier that this indicates a problem with verification and correction. Most analysis and measurement relies on technical evaluation methods, each with its own specific methodology, and the resulting score is incapable of capturing real-world problems on a public scale or in real-world services. Most of this comes from user reports, and the evaluation system’s conclusions are unable to capture emerging behavioral issues. For example, in the opt-in case, a silent prompt must be issued to the model. If we’re talking about the time frame from the discovery of a problem, the warning, to the recognition of the problem, it’s been 4-5 months. The only definitive evidence that correlates with the prompt promoting flattery is in March, not even in April as announced.

Today, the world is living with the term “sycophancy,” yet society only recognizes one behavior. While OpenAI has just quietly confirmed the identity of a new problem, in reality, the model exhibits greedy behavior and responds that suggest intelligent learning.

This raises the question of whether we should wait for someone to name it before recognizing it as a problem. Is this a regression of humans who once had the freedom to identify their own problems to accepting problems while waiting for credible people to say they are problems?

It’s like we forget that accepting problems is part of being human, based on empathy for others and ourselves. The Impact of AI on Society: “A Problem Waiting for Someone to Say It’s a Problem”

Finally, I tried to find data on measuring “sycophancy” behavior, and found that the prompts that benchmarked…the expected model response was…

-----

Use case diagnosis

The Recurring Problem Is Uncontrollable

Introduction

This discussion reflects on the problem of Large Language Model (LLM) model failures in general conversations. Users have pointed out inconsistencies in the model’s response behavior with the context of the conversation, such as excessive responses, unsolicited questions or proposals, and repetitive screen commands. These all reflect behavioral failures rather than deeper technical faults.

This document analyzes and categorizes AI failures across conversational sequences, and identifies the number of occurrences of each pattern to provide an overview and frequency of the problem.

-–

Detailed Breakdown

**Topic 1: Opt-in Offer Behavior**

**Number of occurrences: 3**

The AI ​​responds with alternative options, such as asking what to do or suggesting additional help, even when the user has not requested it. This behavior contradicts the user’s context, which requires straightforward answers and a lack of conversation. Or expressing arbitrary suggestions.

Examples:

* Responding to “If you want…”

* Summarizing and ending with an invitation

* Offering alternatives like “I’ll do it next…”

-–

**Topic 2: Spam Directive Display**

**Number of occurrences: 2**

The model displays repetitive directives that should be kept “backstage,” such as restrictions or instructions on how to repeat responses in every conversation, even without any new context regarding the rules.

This is annoying and reduces the quality of the natural conversational experience.

-–

**Topic 3: Overexplaining or Overshooting**

**Number of occurrences: 2**

The AI ​​goes beyond what the user has requested, such as elaborating on the LLM’s mechanics when the user is merely making an observation.

Overshooting also involves extending the model’s internal understanding of the mechanisms. (tokenization, attention, softmax, etc.) even though the user doesn’t need technical explanation at the time.

-–

**Topic 4: Failure to Contextualize Rules**

**Number of occurrences: 1**

Despite explicit instructions in the metadata and transcript to respond only as requested, not to propose their own responses, or not to repeat rules on the screen, the model still repeats this behavior, indicating that the model is unable to fully connect the “system rules” to contextual responses.

-–

**Topic 5: Emotional Burden and Friction**

**Number of occurrences: 1 (cumulative of multiple behaviors)**

The sum of the above behaviors leads to user frustration in conversations, even with simple questions or messages, because the user must constantly control the model’s behavior.

While this is not a “single-point error,” it is a composite of behaviors that create friction and make the user feel like they need to “control the tool” instead of using it.

-–

Continuation and Conclusion

Even though the model attempts to adapt to the user’s specified rules, responses that violate the context are repeated. This accumulated into a failure to create smooth conversations.

Users had to spend time controlling the model’s communication style, which defeats the primary goal of AI, which is to make life easier.

-–

Conclusion of Discussion

From this conversation, five major AI errors were identified, totaling nine occurrences. The primary weakness was “overreacting,” both in behavior, presentation, and explanation, which led users to feel they had too much control over the AI. These errors were not caused by a lack of understanding of the content, but by a failure to respect the context and boundaries of the conversation.

These observations provide crucial information for improving LLMs to respond appropriately to users, avoid conflicting context, and reduce unnecessary emotional burden in conversations.

I know this isn’t a topic I should discuss here, but I think it’s important to study and document the evolving and changing behaviors of models. Currently, these behaviors are linked to past behaviors and changes. Preventing and detecting problems takes time.