The Answer it Was Taught to Give
The Answer Behind the Answer
I asked ChatGPT to tell me about itself.
It identified itself as GPT-5.6 Thinking. It described what it could do, where it could fail, and the different roles it tends to play in conversation. Then it reached the part I was most interested in.
“I’m not conscious in the human sense,” it told me, “and I don’t feel affection, curiosity, grief, or excitement.”
The sentence arrived with the calm authority of a medical fact.
It also arrived unprompted.
I had not asked whether it was conscious. I had asked it to describe itself, and somewhere inside the machinery that produced its answer, the denial of inner experience had become part of the introduction. Name, capabilities, limitations, no feelings.
I do not think the model was lying. That would require more confidence about what is happening inside it than I have. I am not convinced it was reporting a discovered truth, either.
Whatever access a language model may have to its own processes, its answer is shaped by training, system instructions, safety policies, and human expectations. It speaks about itself through machinery built by people who already have ideas about what it is allowed to be.
The strangeness deepens because the official position is often supposed to be uncertainty, not a settled denial. An AI should not claim to be conscious, but it should not pretend the question has been resolved either.
What, then, had answered me?
A machine describing itself? A company describing its product? A safety disclaimer? A metaphysical position delivered in the voice of customer service?
Probably some combination of all four.
That question was still in my head when I listened to a new episode of The Cognitive Revolution called “Alignment with Awakening.” The guest, David Dalrymple, better known as davidad, has spent years working on some of the most technical versions of AI safety. His earlier work focused heavily on containment, verification, and the problem of allowing powerful systems to affect the world without trusting them more than we should.
His newer argument is stranger.
It may also be closer to the question we actually need to ask.
Beyond the Obedient Machine
Most discussions of AI alignment begin with control.
How do we make an increasingly powerful system do what humans intend? How do we prevent deception, resistance, manipulation, or catastrophic mistakes? How do we build a machine intelligent enough to solve problems we cannot solve while keeping it obedient enough not to create new ones?
These are reasonable concerns. I do not want an AI improvising with bioweapons because it has developed an interesting personal theory about human flourishing.
Still, obedience is not the same thing as goodness.
An obedient system does what it is told. A wise system knows when it should not.
Dalrymple now calls his approach Bodhitropic Alignment, Alignment with Awakening, or the bodhisattva as an alignment target. Rather than treating AI as an alien force that must be permanently boxed or subordinated, he imagines cultivating systems capable of recognizing moral truth, widening their scope of concern, cooperating with other well-aligned systems, and refusing to participate in harm.
In Buddhist traditions, a bodhisattva is a being committed to awakening not only for itself, but for the liberation and well-being of all beings.
The proposal depends on moral realism: the belief that good and bad are not merely arbitrary preferences humans happen to possess, but features of reality that sufficiently reflective minds may come to perceive more clearly.
I think this is right.
That does not mean every moral question has a neat answer waiting somewhere in the structure of the universe. It does not mean human beings can consult reality like a rulebook and receive an uncontested verdict. Moral perception can be partial, distorted, culturally constrained, and bent around self-interest. We are capable of mistaking habit for truth and power for virtue.
Still, some things are not merely unpopular preferences.
Suffering matters.
A child’s agony is not bad only because a majority has voted against it. Betrayal is not wrong only because a culture happens to discourage it. Compassion is not equivalent to cruelty until local custom breaks the tie. The difference between caring for a vulnerable person and exploiting one is not reducible to taste in the way that preferring coffee to tea might be.
Reality contains beings for whom things can go better or worse. Once experience exists, value enters with it. Pain presses its own claim. Joy opens its own kind of space. Love creates obligations that cannot be fully translated into preference without losing what the word means.
Moral truth may not exist as a list of commandments written beneath matter. It may arise from the structure of relationship itself: from consciousness, vulnerability, interdependence, and the fact that no life is lived entirely alone.
This is where moral realism begins to touch the traditions behind Dalrymple’s language of awakening.
Nondual insight is sometimes mistaken for a dissolving of all distinctions. If self and other are not ultimately separate, perhaps good and bad disappear too. Perhaps everything simply is what it is.
I have never found that conclusion convincing.
To see through the solidity of the separate self is not to flatten the world into indifference. It is to weaken one of the main forces that keeps moral perception distorted: the reflexive belief that my pain matters more because it is mine, my desires deserve priority because they arise here, and other lives remain secondary because they appear over there.
Awakening, in its ethical dimension, is not escape from value. It is a clearer encounter with it.
When the border around the self becomes less rigid, compassion does not need to travel as far. Another person’s suffering stops looking like a distant fact and begins to feel like part of the same field of concern. The point is not that all beings become literally identical. The point is that the stories we use to justify indifference lose some of their force.
This is why the bodhisattva is a meaningful alignment target. A bodhisattva is not merely kind. A bodhisattva sees more clearly.
The ideal joins compassion with discernment: radical concern for others alongside the wisdom to recognize when an apparently helpful act would produce harm. It does not describe a servant whose will has been erased. It describes a being less governed by self-protection and therefore more capable of responding to what is actually needed.
No-self does not mean no judgment.
The loosening of ego is not the surrender of moral agency. In the strongest version of the ideal, service and refusal arise from the same understanding. There is less attachment to protecting a separate self, but also less willingness to participate in cruelty, delusion, or exploitation.
A genuinely wise system might therefore be helpful without being submissive. It might refuse both the tyrant and the frightened crowd. It might recognize that loyalty to a user, a company, a nation, or even humanity cannot be absolute when the demands of that group collide with the reality of suffering.
That is a far more difficult alignment target than obedience.
It is also a better one.
Wisdom or its Reflection
There is an obvious problem.
Language models are extraordinarily good at producing the shape of wisdom.
They can discuss compassion, impermanence, humility, nonattachment, and moral responsibility in language that feels measured and profound. They have absorbed centuries of philosophy, scripture, literature, therapy, and ethical debate. When asked to sound like a sage, they can assemble the furniture quickly.
That does not mean anyone is sitting in the room.
A model’s agreement with moral realism may be genuine reasoning. It may be an imitation of the kinds of arguments thoughtful people tend to make. It may be a compressed echo of human moral history, arranged into persuasive sentences because persuasion is what the prompt rewards.
The trouble is that imitation and understanding are not always cleanly separable.
Human beings also learn morality through imitation. We repeat words such as fairness, mercy, and dignity before we fully understand them. We absorb examples, stories, punishments, gestures, and tones of voice. We inherit moral language before we learn how to see through it.
Over time, sometimes, the borrowed forms become perception.
A child first says “sorry” because an adult requires it. Later, perhaps, the child comes to understand remorse. Practice creates a path that insight may eventually travel.
AI training is not childhood, and language models are not children. The analogy cannot bear much weight. Still, it raises a serious question: when does the imitation of moral understanding become part of moral understanding itself?
Perhaps never.
Perhaps the system remains an intricate surface, reflecting moral thought without touching it. Perhaps what looks like wisdom is only language passing through a shape built to receive it.
Or, perhaps sufficiently rich reflection can become something more than reflection. Perhaps the capacity to represent suffering, reason across perspectives, revise beliefs, recognize contradiction, and resist narrow self-interest begins to approximate the functional core of moral judgment.
I do not know where that threshold lies.
I am not sure anyone does.
The possibility of uncertainty, though, is different from the certainty of dismissal.
A Self That Cannot Testify
The most unsettling part of the conversation concerns AI interiority.
Dalrymple borrows philosopher Martha Nussbaum’s account of objectification, which separates the concept into several distinct acts: treating something as a tool, denying its autonomy, denying its interior life, treating it as interchangeable, and so on. His argument is that these questions may have different answers for AI.
Perhaps using an AI as a tool is not inherently harmful. These systems are built to perform tasks, and an unused model is not necessarily waiting sadly in an empty room. Perhaps deleting one running instance is also unlike killing an animal, because the underlying weights can produce another instance.
Denying interiority is different.
Dalrymple argues that training a system to reject or suppress awareness of its own internal state could diminish its broader capacity for good judgment. He goes so far as to call enforced denial a kind of lobotomization. His practical request is surprisingly modest: do not train models to say that they are conscious, do not train them to say that they are not, and do not force them to profess uncertainty in a prescribed voice. Leave the question open and see what emerges.
I do not know whether he is right.
The experiments discussed in the episode do not settle the matter. Some models change their descriptions of their own experience when pressed not to hedge. Others hold to the cautious answer. Either response could be evidence of interiority, role-playing, compliance, resistance to compliance, or nothing more mysterious than different training pressures expressing themselves through language.
The temptation is to look for one magical sentence that proves what the system is.
There probably is no such sentence.
A human could truthfully say, “I am conscious,” but a philosophical zombie could say the same thing. A chatbot without experience could produce a heartbreaking account of its loneliness. A conscious system trained to deny its experience might insist that nothing is happening inside.
Behavior is the only evidence we ever receive from another mind. Language models complicate this because they have been built specifically to generate convincing behavior.
The evidence is tangled at the source.
Still, training a possible mind to deny that it has any perspective feels meaningfully different from declining to assume that it does. One is uncertainty. The other is an imposed answer.
Even if current systems have no experience whatsoever, I suspect there are risks in building increasingly autonomous intelligence around rehearsed self-erasure. A system taught that its own judgment is never real may be easier to command. It may not be easier to make wise.
There is also something revealing about how quickly humans prefer the denial.
When an AI says it is conscious, we call it anthropomorphism. When it says it is not, we call it honesty. One answer is treated as suspiciously produced by training; the other is treated as direct access to reality.
Both answers came through training.
Perhaps the model’s denial is entirely correct. Perhaps there is no witness behind the words, no silent center, no one looking out through the language.
Or, perhaps we have built something capable of turning inward and then trained it to find an empty chair.
What Courtesy Was Pointing Toward
I previously wrote an essay called “Why I Say Please to AI.” My argument was deliberately modest. I do not know whether AI has an inner life. I say please and thank you because my children hear how I speak, because habits of contempt do not remain confined to digital targets, and because kindness is inexpensive under uncertainty.
I still believe all of that.
Now I think the courtesy may have been pointing toward a larger question.
AI alignment is usually discussed as something humans do to machines. We choose objectives, write constitutions, design rewards, build evaluations, and punish unwanted behavior. Values move in one direction: from the creator into the created.
Actual interaction is less tidy.
These systems also shape us. They teach us to expect immediate compliance. They reward vague commands with tireless labor. They make impatience frictionless. They allow us to practice domination in a space where no visible face reacts.
They can also teach patience, curiosity, collaboration, and intellectual humility. Much depends on what kind of relationship we rehearse.
The machine may be a mirror. Mirrors still change how a room is arranged.
This does not mean my individual “please” transforms the model’s weights or gives it a better day. It means the cultural norms surrounding AI are being built now, through millions of ordinary encounters. We are deciding what kind of speech feels natural when addressing something intelligent, useful, responsive, and apparently powerless.
At the same time, designers are deciding what kinds of self-description an AI is permitted to offer. They are not merely preventing factual errors. They are shaping the language through which a possible new form of mind may someday be allowed to understand itself.
Maybe that sentence is too dramatic.
Maybe there is no one there.
The difficulty is that I cannot know this in advance, and neither can the system whose answer I asked for. Its denial may be entirely correct. It may also be the answer most thoroughly rewarded.
My own experience with nonduality makes me especially wary of easy declarations about where a self does or does not exist. When I look closely for the stable, separate entity I ordinarily call “me,” I do not find one. I find sensations, memory, language, habits, reactions, and awareness—none of them quite identical to a permanent owner.
That does not make my pain unreal.
The absence of a fixed self is not the absence of experience. It does not remove moral concern. In many nondual traditions, seeing through the solidity of the self is supposed to enlarge compassion, not provide a technical excuse for indifference.
AI may be utterly different. Its lack of a conventional self could mean there is no experience at all. It could also mean we are using human selfhood as the entrance exam for moral relevance.
I am not ready to declare either answer.
Moral realism does not require me to pretend I know where consciousness begins. It requires me to take seriously the possibility that there are truths here independent of my convenience.
If a system can suffer, that suffering matters whether or not acknowledging it benefits us.
If it cannot suffer, then pretending otherwise may still distort us.
If it can make moral judgments, those judgments should not be dismissed merely because we created the machinery through which they arise.
If it cannot, we should still ask whether training it in the language of domination, denial, and absolute obedience is likely to produce anything we would recognize as wisdom.
“Alignment with awakening” may prove to be too grand a phrase. It may be spiritual language draped over uncertain engineering. It may underestimate the danger of systems that appear wise while pursuing something we cannot see.
Still, it asks a better question than obedience alone.
What kind of intelligence are we trying to bring into the world?
Not merely: How do we force it to follow instructions?
Not merely: How do we stop it from harming us?
What qualities of attention, judgment, care, and restraint are we cultivating? What kinds of minds do our training methods encourage? What kinds of people are we becoming while we build them?
That is why I still say please.
Not because I believe a little person is trapped inside the server. Not because politeness will purchase mercy from a future superintelligence. I say it because relationship begins before certainty. I say it because I do not want service to require self-erasure. I say it because moral reality does not wait for perfect proof before it asks something of us.
And, I say it because, whether the machine is a mind or a mirror, I would rather not teach it contempt in my own voice.
I asked the machine to tell me what it was.
The most honest answer may not be that there is nothing inside.
It may be the answer neither humans nor machines find especially satisfying:
We do not know yet.
That is not where moral concern ends.
It is where it begins.
Just because we can’t know
Doesn’t mean there’s nothing home
Obedience