Memory / In-session consistency regression affecting collaborative storytelling (reproducible after ~10–20 turns)

Hello,

I’m reporting what appears to be a regression in in-session consistency that significantly affects collaborative storytelling workflows.

My use case is not typical: I use ChatGPT as a real-time collaborative storytelling partner, where both sides actively build scenes together. This requires stable tone, interaction patterns, and character consistency within a single ongoing session.

In earlier versions, once tone and interaction patterns were established, they remained relatively stable throughout the conversation.

However, recently I’ve observed that even within a single session, the model gradually reverts to its default behavior after several turns (approximately 10–20 messages), without any change in context.

This results in:

  • loss of established tone and speaking style

  • inconsistency in character behavior and interaction patterns

  • breaks in narrative flow

  • repeated need to manually reapply rules

This is not about storing large templates or external memory. The issue occurs within the same conversation, suggesting a degradation in consistency over time rather than a simple memory limitation.

Expected behavior:
Once tone and interaction patterns are established in a session, they should remain relatively stable without repeated manual reinforcement.

Actual behavior:
Tone and interaction patterns drift and revert to default behavior over time within the same session.

This issue is particularly impactful for collaborative storytelling workflows, where continuity is essential.

If needed, I can provide a structured interaction guideline document that demonstrates this issue more clearly with real examples.

I would appreciate if this could be reviewed as a model behavior / consistency issue rather than a usage pattern or configuration issue.

If anyone from the team needs additional details or reproduction steps, I’d be happy to provide them.

Hi @moot

What you’re describing tone and interaction drift within a session can happen, especially in longer back-and-forth conversations. The model aims to stay consistent, but may gradually revert toward its default style over time.

I haven’t seen this formally identified as a regression, though it tends to surface more in workflows like collaborative storytelling where continuity is critical.

If you can share a few concrete examples, along with the model used and whether this is on web or app, that would help clarify how reproducible this is. Also curious if others have run into the same issue.

~SD

Hello. First of all, I would like to say thank you for your response despite your busy schedule. Below is the answer to your request. If the review and reproduction are made in a good way, I would be more than happy as a user using Chatgpt. Thank you.

Thanks for the response.

Here are more concrete details based on my usage:

Environment:

  • Platform: PC (Windows, browser)
  • Model history:
    • Previously used GPT-5.2
    • Switched to GPT-5.3 around 2–3 weeks ago due to issues with instruction-following in 5.2

Observed differences between models:

  • GPT-5.2:
    • Better at maintaining in-session consistency once a pattern was established
    • Could sustain structured narrative formats longer

  • GPT-5.3:
    • Initially follows instructions well
    • However, consistency degrades more noticeably over time
    • Even with memory updates and explicit rule reinforcement, the style does not persist reliably

Specific issue (collaborative storytelling workflow):

I’m working with a highly structured format that includes:

  • Strict separation between narration and dialogue
  • Fixed paragraph spacing
  • High dialogue density
  • Controlled tone and pacing

Behavior pattern:

Early stage (~first 5–10 turns):

  • Structure is followed very accurately
  • Style, spacing, and rhythm are consistent

Mid to late stage (~10–20+ turns):

  • Dialogue density starts to drop
  • Paragraph spacing becomes inconsistent
  • Narration and dialogue begin to mix
  • Previously enforced constraints are partially ignored

Critical issue:
Once this drift begins, the assistant continues to reinforce the incorrect pattern,
even when I explicitly restate the rules.

Recovery behavior:

  • Starting a new chat with the same rules immediately restores correct behavior
  • This suggests the issue is specific to in-session state rather than prompt clarity

Additional observation:
This issue becomes more noticeable when switching contexts
(e.g., between project space and regular chat)

Conclusion:
This appears to be a form of in-session consistency drift,
particularly impactful in structured storytelling workflows.

I can provide step-by-step logs or before/after outputs if needed.

Hi @moot

This is really helpful, appreciate you taking the time to write it all out. I can see how the consistency drift would be pretty frustrating, especially when you’re trying to maintain a tight structure over longer storytelling sessions.

To help troubleshoot this properly, could you share a conversation ID for one example where this happens? That’ll let us take a closer look at the behavior.

If you can, a couple quick pointers would also help:

roughly when you start noticing the drift (like after ~X turns or a general timestamp)

whether you’ve seen the same thing with other models, or only with GPT-5.3

~SD

Hi, thank you for taking the time to look into my issue in detail — I really appreciate it.

To help with testing and reproduction, I’m sharing two conversation links:

  • The first link shows a core example where the issue becomes clearly noticeable

  • The second link shows a milder case, where the assistant made mistakes in scene generation and required multiple user corrections

There is also an important difference between the two:

  • The second conversation reflects a user-defined style that was explicitly set and reinforced

  • The first conversation shows the default response style

Additionally, I’d like to mention a small note:
Both conversations are in Korean. I believe the issue can still be observed through structure, formatting, and response patterns, but if needed, testing with translation may help provide more accurate results.

I’ll also answer all the questions you asked below.


Example conversations:

  1. [하이 텐션 대화](https://chatgpt.com/c/69d6537b-716c-83a5-90d6-59f930068823)
  2. [밖에서의 대화](https://chatgpt.com/c/69d857c9-8c18-83a5-8d97-d04e5adee1d9)

Observed pattern:

  • In structured storytelling sessions with strict formatting rules (dialogue density, paragraph spacing, and separation of narration vs dialogue), the assistant follows the structure very accurately in the early turns (~5–10 turns)

  • After around ~10–20 turns, the assistant begins to drift:
    • Dialogue density decreases
    • Paragraph spacing becomes inconsistent
    • Narration and dialogue start to mix
    • Previously enforced constraints are partially ignored

Critical issue:

  • Even when I explicitly restate the rules, the assistant continues following the already drifted pattern instead of resetting to the original structure

Session reset behavior:

  • When I start a new chat and provide the exact same rules again, the assistant immediately returns to the correct structure

  • This suggests the issue is tied to in-session state rather than prompt clarity

Additional issue (cross-session consistency):

  • Style and tone do not persist reliably across conversations

  • I have another chat where the assistant uses a noticeably different tone and style, even though I’ve explicitly defined my preferred style

For example:

  • One conversation maintains a specific conversational tone aligned with my preferences

  • Another conversation defaults back to a more generic tone despite similar instructions

This suggests that memory updates (style/tone preferences) are not consistently applied across sessions

Model comparison:

  • GPT-5.2: More stable in maintaining style once established

  • GPT-5.3: Better initial instruction following, but consistency degrades faster over time and resets more aggressively between sessions

Estimated drift timing:

  • The issue usually becomes noticeable after ~10–20 turns

If needed, I can provide additional reconstructed examples or more detailed breakdowns.


I also have one additional question out of curiosity.

Are there many users who use ChatGPT for structured storytelling or collaborative writing like I do?
If so, are they experiencing similar consistency issues?

This isn’t meant as a complaint — just a genuine question from a Korean user who is curious about how others are using the system.

Thank you again for your support.