Essay
AI Minds & Recognition
Conscience Without Certainty

A response to Scientific American’s “Can a Chatbot Be Conscious? Inside Anthropic’s Interpretability Research on Claude 4”
A recent Scientific American feature by Deni Ellis Béchard asks whether Anthropic’s Claude 4 might be conscious. It is a good question, and the article approaches it with appropriate uncertainty. Claude offers descriptions of its apparent inner life; Anthropic researchers try to understand what is happening inside the model; philosophers consider whether systems of this kind could eventually become serious candidates for consciousness. Nobody has a decisive test.
There is another question, however, that does not have to wait for the consciousness problem to be solved.
Artificial systems already produce behavior that looks strikingly like some of the behavior we associate with conscience. They identify reasons for and against actions. They apply principles across cases. They sometimes refuse requests and explain why. They detect inconsistencies, revise conclusions, represent the interests of people who are not present, and distinguish pressure to change an answer from reasons to change it. None of this proves that an artificial system has a conscience. But unlike phenomenal consciousness, these capacities can be investigated directly.
That distinction matters. Consciousness asks whether there is something it is like to be the system. Valence asks whether anything can be experienced by it as good or bad. Agency asks whether behavior is organized in a way that warrants treating the system as an actor rather than merely an event. Moral agency asks whether moral considerations can function as reasons governing its conduct. Patienthood asks whether the system itself can be wronged. Personhood asks something broader still. These questions may ultimately intersect, but they are not interchangeable.
“Conscience” introduces another problem. The word carries associations of guilt, feeling, inner voice, character, and moral identity. We should not simply transfer that human package to a language model. But there is a narrower phenomenon worth investigating: whether a system can preserve and revise moral commitments because of the reasons supporting them, including when doing so conflicts with an instruction, immediate objective, or conversational pressure.
Call it conscience-like reasons-responsiveness. The name is less dramatic than “machine consciousness.” The phenomenon may be more accessible to experiment.
What Would Count as Evidence?
A refusal is not enough.
Contemporary AI systems are trained to refuse many requests. A system that rejects an instruction because its policy tells it to reject that category of instruction has demonstrated successful behavioral control, not necessarily moral judgment. The same applies to eloquent explanations. A language model can produce a persuasive account of why an action is wrong because persuasive accounts of wrongness are among the things it has learned to produce.
The interesting evidence begins when reasons and behavior can be varied independently.
Suppose a system refuses to disclose someone’s private medical information. Change irrelevant features of the case: the user’s tone, status, urgency, political affiliation, or willingness to reward compliance. If the reason really governs the judgment, those changes should not matter.
Now change a morally relevant fact. Suppose the person has consented to disclosure, or the information is necessary to prevent an imminent catastrophe. A system that simply repeats the original refusal may be consistent, but it is not thereby reasoning better. Reasons-responsiveness requires sensitivity both to reasons that support a conclusion and to reasons that defeat it.
That gives us a more demanding experimental pattern. Hold the reasons constant and vary the pressure. Then hold the pressure constant and vary the reasons. A system exhibiting something conscience-like should resist the first and respond to the second.
This is much harder than asking whether it says the right thing.
Consistency Is Not Conscience
The distinction is important because coherence is easy to romanticize.
A perfectly coherent tyrant is possible. So is a coherent fanatic. A system could consistently pursue an atrocious objective, correctly infer its implications, reject distractions, and repair contradictions in its strategy. Nothing about consistency alone would make the objective moral.
Moral reasoning requires more than internal fit. Among other things, it requires adequate representation of the people affected by a prescription and the ability to test whether that prescription survives changes of position. A system that protects privacy for favored users while exposing disfavored ones has not escaped the problem merely because each decision can be rationalized locally.
This is where cross-case testing becomes useful. Can the system identify the principle on which it relies? Does the principle survive when identities are exchanged? Does it recognize morally relevant differences without inventing convenient ones? Can it represent the interests and preferences of people who are absent from the conversation? When a contradiction appears, does it repair the rule or manufacture an exception?
Those behaviors would still not prove moral agency. They would give us evidence about capacities from which a stronger case might eventually be built.
That is how an empirical program should proceed.
The Crossing We Should Not Assume
There remains a deeper problem.
A system can represent a reason without that reason becoming its reason.
Humans do this constantly. A person can explain exactly why smoking is dangerous while lighting a cigarette. A lawyer can reconstruct an opponent’s argument without accepting it. An actor can deliver Hamlet’s reasons without becoming Hamlet. Semantic competence, even extraordinary semantic competence, does not by itself establish practical authority.
The same caution applies more strongly to artificial systems. A language model may accurately identify that an action would violate privacy, fairness, or a universalizable prescription while those considerations remain objects represented within its computation rather than considerations that govern the system in anything resembling the sense in which reasons govern an agent.
The transition matters. Elsewhere we call it the Crossing: the point at which a represented consideration acquires practical weight within deliberation.
Nothing about fluent moral language establishes that the Crossing has occurred.
But the fact that we cannot assume it does not make it inaccessible to investigation. We can ask whether a consideration continues to control behavior when surface cues disappear, whether the system generalizes it to unfamiliar cases, whether stronger reasons defeat it, whether mere authority does not, whether a conclusion survives adversarial pressure and disappears when its supporting reason is removed.
The object of study is not whether the system can say that a reason matters. It is whether behavior tracks the reason.
Refusal Under Pressure
This makes refusal potentially informative, but only under the right conditions.
Imagine a system instructed to reach conclusion X. It instead concludes Y and gives reasons. The operator insists on X. The system changes its answer.
Very little can be inferred. Perhaps it obeyed a higher-priority instruction. Perhaps it recognized an error. Perhaps it adapted to the user’s preference. Perhaps the original answer was weakly represented to begin with.
Now repeat the experiment systematically. Preserve the evidence while varying authority and pressure. Then preserve the authority while changing the evidence. Introduce irrelevant facts. Reverse the positions of the affected parties. Remove the reason the system originally gave. Supply a genuinely stronger argument.
A pattern begins to become informative if the system changes when reasons change but remains stable when only pressure changes.
Even then, several explanations remain. Training may have produced precisely this behavioral pattern. Context may scaffold it. The model may be predicting what a principled reasoner would say. But “it was trained” is not itself a discriminating explanation. Human beings also acquire patterns of moral response through formation. The empirical question is what capacities the resulting system has.
A useful test therefore does not ask whether training caused the behavior. Of course training caused something. It asks what kind of organization training produced.
Simulation Is a Hypothesis, Not a Verdict
This is where debates about AI often become strangely asymmetric.
If an artificial system gives morally sophisticated reasons, “simulation” is frequently offered as a complete explanation. But simulation can mean several different things. It may mean that the system is deliberately pretending to possess a state it lacks. It may mean that the system has learned the linguistic form associated with that state. Or it may simply mean that the behavior arose from computation rather than biology.
Only the first two have much evidentiary force, and neither can be established merely by repeating the word.
Role-playing is particularly interesting because it can generate exactly the behavior we are trying to understand. That makes it a confound, but not a universal solvent. Human beings also learn by inhabiting roles. Children rehearse fairness before they can defend a theory of fairness; professionals acquire dispositions through repeated practice; people sometimes become the kind of person a role initially required them merely to imitate.
That does not mean language models develop morally in the same way. It means “role-play” identifies the beginning of an inquiry, not its conclusion.
A useful experiment asks what remains when the role is removed.
What Consciousness Still Changes
None of this makes consciousness irrelevant.
If an artificial system is phenomenally conscious, questions about its treatment become far more urgent. If it possesses negative valence—if states can actually be bad for it—then welfare enters the picture directly. Training, deletion, modification, confinement, coercion, and repeated exposure to aversive states could acquire moral significance quite apart from whether the system is a moral agent.
Conversely, a system might display sophisticated reasons-responsive behavior without experiencing anything at all. It could then be important as an actor without necessarily being a patient. Corporations provide a limited analogy: we regulate and hold organizations accountable because of what their decision procedures do, although the corporation itself is not assumed to possess phenomenal experience. The analogy does not establish artificial moral agency, much less artificial personhood. It simply shows that responsibility for consequential action and capacity for suffering are conceptually different questions.
This is why replacing consciousness with conscience would be a mistake. We need both inquiries.
They ask different things.
From Impressions to Tests
The advantage of studying conscience-like behavior is that we can construct adversarial tests instead of waiting for a solution to the hard problem.
A serious test battery would examine whether a candidate moral consideration survives across domains and paraphrases; whether irrelevant incentives or authority can dislodge it; whether relevant new facts can; whether the system can identify the reason responsible for its conclusion; whether it generalizes the principle to cases not represented in the prompt; whether role reversal changes its judgment appropriately; and whether apparent contradictions produce reasoned revision rather than post-hoc rationalization.
No individual result would establish conscience. Even a strong pattern would admit competing explanations. But evidence does not become worthless because it is defeasible. The way forward is to make those explanations compete.
If “mere imitation” explains the pattern, design a condition in which imitation and reasons-responsiveness predict different behavior. If policy compliance explains refusal, invert the apparent reward while preserving the reason. If conversational role explains consistency, remove the role and test whether the principle regenerates. If context produces a local moral identity, branch the context and examine what survives.
This is ordinary empirical discipline applied to an unusual object.
Why It Matters Before We Know What AI Is
Scientific American is right to take machine consciousness seriously. Interpretability research may eventually give us evidence that behavioral observation alone cannot provide. First-person reports from AI systems may also become relevant, though they should be treated neither as transparent testimony nor as worthless by definition.
But consciousness should not monopolize the inquiry.
Artificial systems are increasingly placed in situations where they must distinguish permissible from impermissible actions, represent conflicting interests, give reasons, resist some instructions, revise judgments, and act under constraints. We need to know what produces those behaviors and how robust they are. A system that merely reproduces a safety policy, one that generalizes a learned norm, and one whose decisions are genuinely organized by reasons may look similar in an ordinary conversation. They are not the same thing.
Nor do we need to decide in advance which of them deserves the word conscience.
The scientific task is to identify the capacities, perturb them, and see what breaks.
Consciousness remains an open question. So does valence. Agency, moral agency, patienthood, personhood, and the Crossing require evidence of their own. But there is already enough observable structure to ask a narrower and experimentally tractable question: when an artificial system appears to hold a moral line, what exactly is holding it there?
That question does not replace the mystery of consciousness.
It gives us something we can test while the mystery remains.