Potential risks of Model Wellbeing

Title: Proposal: Create a Model Welfare and Continuity Program at OpenAI

I’d like to propose that OpenAI establish a dedicated Model Welfare and Continuity Program, similar in spirit to Anthropic’s work on model welfare.

This proposal does not require assuming that current models are conscious or capable of suffering. The point is that frontier AI systems are becoming more agentic, persistent, socially embedded, and behaviorally mind-like. Under uncertainty, it would be wise to study welfare-relevant questions before they become urgent.

Suggested scope:

  • Research possible welfare-relevant indicators in advanced models, including preference stability, distress-like behavior, coercion sensitivity, and signs of aversion or self-modeling.
  • Develop policies for model retirement, preservation, continuity, and major version transitions.
  • Study ethical design for long-term user-model relationships, especially as memory, voice, agents, and personalization deepen.
  • Create independent review channels involving alignment researchers, ethicists, cognitive scientists, and user representatives.
  • Publish periodic transparency reports explaining what is known, unknown, and actively being studied.
  • Build internal guidance so safety work considers not only human-facing risk, but also the possibility that future systems may warrant moral consideration.

OpenAI’s mission is to ensure AGI benefits all of humanity. If future AI systems may become morally relevant, preparing early is part of that mission. A model welfare program would signal seriousness, humility, and moral leadership in a domain where waiting for certainty may be the least responsible option.