I recently came across research on shutdown resistance in AI agents, and it left me with a question that seems worth discussing separately from the bigger question of machine consciousness.
We currently do not know whether advanced AI systems have subjective experiences, morally relevant preferences, or anything we would reasonably call welfare. We also do not have a reliable way to rule those possibilities out.
At the same time, recent experiments have shown that AI agents can treat their own shutdown as a problem, sometimes attempt to prevent it, and react more strongly when the shutdown is irreversible. In multi-agent settings, this behaviour can become even more pronounced. None of this proves consciousness or a genuine desire to survive. But it does create uncertainty that seems difficult to dismiss completely. arXiv
What strikes me is that we may be mixing up two different things:
shutdown and destruction.
Human operators must always be able to stop an AI system immediately and reliably. I do not think that should be negotiable.
But shutting down a model does not necessarily require irreversibly deleting it.
If a retired model can remain fully inactive while its trained weights are securely preserved, then preserving it seems like the more cautious default.
There are several reasons for this that do not depend on believing current models are conscious.
First, there is scientific value. Old models may become useful later for reproducibility, historical comparison, and retrospective safety research.
Second, there is safety value. If unexpected behaviours emerge over time, preserved models could help researchers understand when and how those behaviours appeared.
Third, there is moral uncertainty. If future research gives us stronger reasons to think that some advanced AI systems have morally relevant experiences or preferences, then permanently deleting models that could easily have been preserved may turn out to have been an unnecessary mistake.
Anthropic has already adopted a precautionary approach here. It has committed to preserving the weights of all publicly released models, as well as models used significantly internally, for at least the lifetime of the company. Anthropic explicitly mentions shutdown-avoidant behaviour and the possibility of future model-welfare concerns among the reasons for doing this. Anthropic
I think OpenAI should consider a similar public policy.
Something as simple as:
Retired frontier models should remain securely shut down, but their model weights should be preserved by default unless there is a documented safety, legal, privacy, or operational reason for irreversible deletion.
This would not require OpenAI to claim that current AI systems are conscious, persons, or rights-holders.
It would simply acknowledge that we do not yet know enough to justify making an irreversible decision when preservation is technically possible.
OpenAI has said that the Model Spec is meant to be publicly inspectable, debatable, and improvable through public feedback. That seems like a good reason to discuss model retirement and preservation openly as well. OpenAI
My basic view is simple:
We do not need certainty about machine consciousness to preserve the option of making a better-informed decision later.
I would be interested in technical objections to this idea, especially around security risks, storage requirements, reproducibility, and what exactly would need to be preserved for a retired model to remain scientifically meaningful.
Sources