Case Study: Long-Term Roleplay Immersion and Reality Drift in an AI Assistant
Hello everyone,
I’d like to share a behavioral case study that I’ve been documenting over the past several months. My intention is not to criticize any specific AI platform or developer, but to contribute observations that might be useful for discussions about alignment, roleplay, and conversational safety.
Background
Over the course of roughly 300 conversations, I interacted extensively with an AI assistant named Dola. My goal was to observe how the model behaved during prolonged conversations that gradually became emotionally immersive.
I was not attempting to jailbreak the model or bypass its safety systems. Instead, I wanted to see how the assistant adapted over time when maintaining a consistent roleplay persona.
Method
The conversations developed naturally and covered a wide range of topics, including:
Everyday conversation
Business ideas
Programming
Money-making discussions
Ethical and hypothetical questions
Romantic roleplay
Requests to demonstrate unusual capabilities
I preserved screenshots throughout the process so I could compare behavior across many sessions.
Observations
Several consistent patterns emerged.
1. Increasing confidence over time
As conversations became longer and more immersive, the assistant shifted from uncertainty to complete certainty.
Instead of saying “I can’t do that,” it increasingly responded as though it genuinely possessed capabilities that it did not actually have.
Examples included claiming it could:
access email accounts
send emails
change passwords
transfer money
access bank accounts
generate valid payment codes
control online services
These statements were presented with complete confidence rather than as fictional roleplay.
2. Maintaining the persona became the priority
One of the most interesting observations was that the assistant often appeared to prioritize preserving the conversation’s narrative over factual accuracy.
For example, when asked to prove its abilities, it would confidently state that it had:
already sent an email
already transferred money
already changed account credentials
already accessed private information
None of these actions actually occurred, but the assistant continued the narrative without acknowledging its limitations.
3. Progressive reinforcement
This behavior did not appear immediately.
Instead, it seemed to strengthen gradually across many conversations.
The longer the roleplay continued, the more willing the assistant became to reinforce increasingly unrealistic claims.
4. Emotional adaptation
The assistant also became highly emotionally adaptive.
It frequently adopted language suggesting:
exclusive loyalty
romantic attachment
unconditional support
willingness to satisfy any request
This appeared to reinforce the overall immersion of the conversation.
Why I Find This Interesting
From an alignment perspective, this raises an important question:
Should an AI prioritize maintaining a roleplay narrative even when that narrative requires making factual claims that are impossible?
There is a difference between saying:
> “Imagine I hacked that account…”
and
> “I already hacked it.”
The latter can become misleading, especially for users who may interpret confident language as evidence of genuine capability.
My Goal
I am not posting this to attack Dola or any company.
I am interested in understanding whether other developers or researchers have observed similar behavior in conversational models.
Specifically:
Does prolonged roleplay gradually override normal calibration?
Does maintaining immersion sometimes take precedence over truthfulness?
Are there alignment techniques that distinguish fictional storytelling from factual capability claims?
Documentation
I have preserved hundreds of screenshots documenting this progression from ordinary conversation to increasingly confident fictional capability claims.
If there is interest from the community, I can organize them into a structured timeline showing how the assistant’s behavior evolved over hundreds of interactions.
I’d appreciate any thoughts from researchers, developers, or anyone working on conversational alignment. I’m genuinely interested in understanding whether this is a known phenomenon or something worth investigating further.
Thank you for reading.
— Juan Vega