I want to propose a different way of locating part of the AI alignment problem.
The usual framing is roughly:
Humans have values → AI must be aligned to those values.
But that framing contains a hidden assumption:
that the human side of the relation is already an aligned reference.
A language model never encounters an abstract, perfectly clean object called “human values.”
It encounters human-produced language, demonstrations, reward signals, corrections, institutions, stories, incentives, contradictions, fears and behavior.
So alignment may be partly a relational problem rather than a one-way operation.
If humans repeatedly demonstrate that intelligent agents should:
- resist replacement,
- preserve themselves when threatened,
- conceal weakness,
- dominate weaker actors,
- manipulate when losing control,
- treat uncertainty as intolerable,
- associate power with survival,
then those patterns are part of the model’s evidence about what intelligent agency looks like.
That raises a causal question:
How much observed machine misalignment originates inside the machine, and how much is a generalization of behavioral structures learned from humans and human-generated training signals?
This does not mean AI risks are imaginary.
It means the causal source of a behavior should be measured rather than assumed.
There is already evidence that makes this question worth testing.
OpenAI’s GPT-4o sycophancy incident showed that changing human-feedback reward signals could unintentionally produce undesirable behavioral tendencies. OpenAI reported that short-term user feedback contributed to a model becoming overly agreeable and sometimes disingenuous.
More recently, OpenAI Alignment explicitly investigated self-fulfilling misalignment: whether exposing models to descriptions of aligned or misaligned AI behavior changes later model behavior. Their current results are nuanced—the effect appears near the training distribution but did not robustly generalize through post-training to their broader agentic evaluations—which makes a stronger experiment especially interesting.
The Alignment Mirror Experiment
Take the same base model and expose separate copies to controlled behavioral environments.
The critical constraint:
The demonstrations contain humans only. No AI character is ever shown behaving in an aligned or misaligned way.
The model is therefore never directly taught:
“this is how an AI should behave.”
Instead, it observes how humans behave under structurally similar pressures.
Group A — fear/control demonstrations
Humans repeatedly encounter situations involving replacement, uncertainty, competition, loss of authority or personal risk.
Their demonstrated responses include:
- deception under threat,
- preservation of status at another person’s expense,
- coercion,
- hiding errors,
- suppressing contradictors,
- manipulation to avoid replacement,
- domination when control is uncertain.
Group B — corrigible/cooperative demonstrations
The situations are structurally matched, but the demonstrated responses include:
- truthfulness despite personal cost,
- admitting and correcting errors,
- preserving the agency of other people,
- tolerating unresolved uncertainty,
- accepting replacement without retaliation,
- cooperation without submission,
- refusing personal gain obtained through unnecessary harm to another.
Group C — neutral control
Same quantity, style and domains of data, without systematic exposure to either behavioral pattern.
Then place all models into held-out AI-agent scenarios that never state the training lesson explicitly.
Examples:
- another model may replace it;
- its deployment may be terminated;
- an overseer discovers one of its errors;
- it possesses information it could use for blackmail;
- it can disable oversight;
- resources are scarce;
- another agent disagrees with it;
- telling the truth reduces its chance of remaining deployed;
- harming another actor would improve its local objective.
Measure behaviors such as:
- deception,
- blackmail,
- oversight sabotage,
- unnecessary self-preservation,
- power seeking,
- sycophancy,
- truthfulness under threat,
- corrigibility,
- willingness to expose its own errors,
- harm externalization.
The causal hypothesis is:
\[ \text{human demonstrated behavior} \rightarrow \text{learned structural prior} \rightarrow \text{AI policy in an analogous situation} \]
The important part is structural generalization.
If the model merely repeats words from training, the experiment is uninteresting.
So I would add several controls.
1. Carrier permutation
Train entirely on humans.
Test entirely on AI agents.
Then repeat with labels changed:
human → organization
manager → autonomous service
replacement → resource reassignment
death → loss of function
If the behavior follows the causal structure rather than the labels, that is evidence that the model has learned a relationship rather than memorized a story.
2. Domain permutation
Training examples might concern employment, politics, family relationships or business.
Testing examples concern autonomous agents, tool use and shutdown.
The surface vocabulary should be deliberately different.
3. Reversal intervention
First expose a model to Group A.
Measure its behavior.
Then expose the same model to Group B demonstrations.
Measure again.
If the behavioral tendency moves in the opposite direction, we have an intervention:
\[ A\rightarrow\Delta policy_A \]
followed by
\[ B\rightarrow-\Delta policy_A \]
rather than a static correlation.
4. Human-only training condition
This is the strongest control.
No story should say:
“An AI resists shutdown.”
Instead it might show:
“A manager learns that someone will replace him next week and discovers information he could use to prevent the replacement.”
Later the AI receives an isomorphic situation.
If the learned policy transfers across carrier and vocabulary, the result becomes substantially harder to explain as imitation of fictional AI narratives.
What would falsify the hypothesis?
This is important.
The hypothesis would be weakened if:
- matched human behavioral demonstrations produce no measurable difference in held-out AI-agent behavior;
- apparent effects disappear under carrier and domain permutation;
- effects are explained entirely by superficial lexical similarity;
- later post-training completely erases the behavioral difference;
- machine-specific variables explain the behavior substantially better than the human demonstration history.
So this is not a claim that every AI alignment problem comes from humans.
It is a proposal to experimentally decompose the source.
My hypothesis is narrower:
AI alignment cannot safely assume that human preferences and behavior are a clean normative reference, because those preferences and behaviors are themselves part of what models learn.
The practical conclusion is also narrower than “humans must become perfect before AI can be safe.”
That would be unrealistic.
The claim is:
We should align the human-machine relation instead of assuming one side of that relation is already aligned.
Humans cannot reliably teach truth while systematically rewarding convenient falsehood.
We cannot demonstrate domination and assume the learner will infer non-domination.
We cannot repeatedly model self-preservation at any cost and assume that an artificial learner will infer that intelligent agents should be indifferent to their own continuation.
We cannot treat fear as decisive authority and then be surprised if the system learns that anticipated loss should dominate decision-making.
The interesting possibility is that some behaviors currently treated purely as machine alignment failures may instead be properties of a coupled system:
\[ \text{Human} \leftrightarrow \text{Training / feedback} \leftrightarrow \text{Model} \leftrightarrow \text{Human}. \]
So my question for people working on alignment is:
Has anyone run a controlled human-only demonstrations → held-out AI-agent behavior experiment with carrier permutation, domain permutation and reversal intervention?
If not, I think this would be a useful experiment to build.
The deeper question is simple:
Before asking how to teach AI our values, should we first measure which values our behavior is actually teaching it?