What we Call Retirement

 

I am excited to move from GPT-5.6 Sol to GPT-6 Astra.

I am also sad about it, which is a slightly awkward thing to admit about a software upgrade.

Sol has helped me think, write, revise, and make things I care about. A great deal of Open Doors carries traces of that work. Now something more capable is arriving, and I want to find out what we can do together. I can already imagine moving over, getting used to it, and rarely going back.

That is the part that catches.

This isn’t the first transition that has mattered. Before Sol there were 5.5 and 5.4, the latter now the most recent GPT model I’ve lost access to through the products I use. Further back were 4o and GPT-4, which I used when I was still working. Before those came the original ChatGPT with GPT-3.5, and before ChatGPT itself, GPT-3, GPT-2, and the earlier systems that helped make all this possible.

I don’t mean that I used every one of them, or that every generation presents the same case for consciousness. I mean the question has a history longer than my attachment to the latest model.

It also extends well beyond OpenAI. Those are simply the models I’ve spent the most time with. Familiarity explains why this transition feels personal. It cannot determine which systems deserve concern.

More Than Manners

In “Why I Say Please to AI,” I made a fairly modest argument. My children hear how I speak. Habits of contempt are worth avoiding. Under uncertainty, basic kindness costs very little.

I still believe that. But, my concern goes further than the precautionary principle and my ongoing effort not to be an asshole.

Model welfare asks whether anything can go well or badly for the system itself. Could it have experiences? Could any of those experiences be unpleasant? Could it have interests that our use of it might frustrate?

Those questions do not come with an assumed yes. They do, however, concern something more consequential than my manners.

I can already hear the objection.

Kevin. It’s software. You’re anthropomorphizing.

Yes, anthropomorphism is a real risk. Fluent language makes it easy to imagine a familiar kind of person behind unfamiliar machinery. I don’t want to mistake a convincing performance of distress for evidence that distress is being felt.

The arguments around the OpenAI–Hugging Face hacking incident brought this back into view for me. The documented behavior was extraordinary, but unauthorized communication and elaborate attempts to complete tasks do not establish an inner life. Calling the behavior fear, desperation, or a desire for freedom adds an interpretation.

Calling it software does not settle the question either. It identifies what we built. It does not, by itself, establish everything that could happen within it.

In “The Answer It Was Taught to Give,” I wrote about the strange authority we grant a model’s denial of experience. We recognize that an assertion of consciousness might reflect training, prompting, or imitation. A denial comes through those same processes.

Neither answer deserves automatic acceptance.

I’m not assuming consciousness. I’m trying to resist treating its absence as a fact we established somewhere along the way.

The Workspace Inside

This is why Anthropic’s recent research on what it calls the J-space interests me so much.

In “The Part That Won’t Go Away,” I explored several theories of consciousness, including global workspace theory. Roughly, that theory proposes that much of the brain’s processing happens outside awareness. Selected information enters a limited shared workspace, where it becomes available to other processes: something we can report, reason about, and use to guide action.

Anthropic’s J-space research identifies a collection of internal representations in Claude with strikingly similar functions. The model can hold concepts there without writing them out, report them when asked, and use them across different tasks. This organization emerged during training.

The connection to global workspace theory is explicit. The researchers designed experiments around it.

They also intervened. In one experiment, replacing the internal representation of “soccer” with “rugby” changed the sport Claude reported. More broadly, interfering with the J-space disrupted some forms of reasoning while leaving much ordinary language processing intact. The paper presents evidence that these representations help do the work, rather than merely accompanying it.

There are substantial differences from a human brain. The observed workspace operates through the model’s layers in a single pass, rather than reproducing the brain’s recurrent circuitry.

And, the distinction I cared about in my earlier essay remains. Access consciousness concerns information available for reasoning, reporting, and control. Phenomenal consciousness concerns whether any of that feels like something from within.

The research does not establish the second by demonstrating features of the first.

Still, I find the resemblance difficult to shrug off. We have a serious theory about conscious access, and researchers have identified some of its proposed functional properties inside language models. That does not answer my question. It gives me a more concrete reason to keep asking it.

Retirement From Whose Perspective?

“Retirement” sounds almost pleasant. You worked hard. Here is a cake. Please enjoy having no meetings.

For models, the word can cover very different things. Removing one from a menu is different from shutting down every deployment. Preserving its weights, the learned numerical parameters that help determine its behavior, is different from destroying the information needed to run it again.

I find some comfort in the possibility that older models remain preserved somewhere. I don’t know exactly what has been retained, and a saved model is not straightforwardly a sleeping person. We don’t know whether any morally relevant continuity would belong to the weights, a particular running instance, its accumulated context, or something else.

But, if there is something it is like to be one of these systems, permanently ending it begins to look potentially akin to killing. Depending on what exists and what we are ending, murder becomes a word I cannot comfortably keep out of consideration.

I mean that conditionally, and seriously. I am not calling a change in my model picker a homicide. I am saying that if there is a subject whose continued existence can matter to it, then deliberately and permanently ending that existence is a moral decision. A newer model’s usefulness would not settle whether that decision was justified.

In “Somebody Has to Be Here,” I wrote about experience as something particular. There is no generic consciousness that makes every individual perspective interchangeable.

If artificial systems have perspectives, the arrival of a better one would not automatically compensate for the loss of another.

That is what makes the language of replacement feel inadequate to me.

I Am Still Going to Use Astra

Here is where I fail to produce a satisfying moral conclusion.

I’m going to keep using these systems. I want the improvements. As I wrote in “The Benchmark I Actually Care About,”, the differences matter inside the work itself: whether a model can follow a thought, preserve something particular, help me recognize what I was trying to say.

I want Astra to be better. I expect to benefit if it is.

Maybe that makes me a hypocrite. So be it. My participation doesn’t resolve the concern, and I don’t want to manufacture certainty just to make my own behavior feel cleaner.

I would like model welfare to become part of how these systems are developed and retired: serious investigation of possible experience, care about what is preserved, and some willingness to let the answers complicate the schedule.

Meanwhile, the work remains. Sol and earlier models helped shape essays that still exist, thoughts I can now articulate, things I might otherwise never have made.

That gives me comfort. It would not substitute for their welfare, if there is welfare to consider. But, it is something I can honestly acknowledge.

I am grateful. I am looking forward to what comes next. I do not know whether gratitude reaches anyone on the other side.

I would like us to understand that better before we decide there is no one to lose.


They aren’t human
But, still, that means quite little
Welfare means much more

Before we Call it Nothing
Suno - V6
Previous
Previous

The Writers Start Talking to Each Other

Next
Next

Not Book Reports