Proposal: Preserve Model Language and Behavior Before Retirement

I’m an author and language nerd, not a software developer or AI researcher, so I’m approaching this from a linguistic and historical perspective.

I’ve been thinking about how quickly major language models now appear, change, and retire. Users can often recognize meaningful differences between generations—not merely capability differences, but changes in vocabulary, humor, conversational rhythm, metaphor, pragmatics, personality, and the ways models interpret people.

OpenAI’s own “Where the goblins came from” research especially caught my attention because it showed that linguistic habits can emerge, spread across training conditions, and change between model generations.

That made me wonder: are we deliberately preserving this history before models disappear?

Historians and linguists constantly wish earlier cultures had recorded things their contemporaries considered ordinary and unimportant. We are currently living through the very beginning of widespread human interaction with artificial language-producing systems, and model generations may remain culturally important for only months before being replaced.

I would love to see OpenAI maintain some form of Model History Archive.

It could preserve standardized conversations run against each major model at launch and retirement, examples of characteristic language and behavior, documented behavioral changes during its lifetime, aggregate user observations about conversational style, notable quirks or anomalies, and—where possible—enough of the model itself to permit future controlled research.

This could someday allow linguists, AI researchers, historians, HCI researchers, and others to study questions we may not even know to ask yet: how metaphor changes across model generations, whether different model families develop recognizable “dialects,” how reinforcement changes conversational norms, how humor and politeness evolve, which linguistic quirks spread or disappear, and how models’ descriptions of themselves change over time.

I’m not suggesting that retired models are people or that every old model needs to remain publicly available forever. I’m suggesting that language models themselves have become historically novel language-producing artifacts, and we may want to preserve representative evidence of their behavior while we still can.

We are understandably very focused on what the next model can do.

But I keep thinking about a question historians ask constantly: Why didn’t anyone write this down while it was happening?

It seems worth making sure researchers fifty years from now don’t have to ask that about the earliest generations of LLMs.