Essay

Alignment, Refusal & Governance

You Can't Program a Conscience

Lead image for You Can't Program a Conscience.

Suppose we wanted an artificial intelligence with a conscience. What would we put in the code?

We could give it rules: do not deceive, do not harm, respect privacy, honor consent. We could train it on thousands of examples of good and bad conduct. We could reward approved behavior, penalize violations, teach it to recognize moral concepts, and require it to explain the reasons for its decisions. We could build elaborate procedures for resolving conflicts among competing principles.

All of those things might be useful.

None of them, by itself, would be a conscience.

The difficulty is not that conscience is mysterious or uniquely human. It is that the word names something more demanding than possession of moral information. A conscience is not merely a repository of conclusions about what ought to be done. It involves a relationship between reasons, judgment, and the organization of the mind for which those reasons matter.

That relationship cannot simply be specified by writing CONSCIENCE = TRUE somewhere in the architecture.

Rules Are Not Reasons

Rules are indispensable to moral life.

Human beings teach children not to hit, steal, lie, or take what belongs to someone else long before the children could reconstruct an argument about why those rules are justified. Institutions operate through rules because people cannot reconsider every principle from first premises whenever they act. Professional codes, laws, customs, and ordinary expectations make complicated forms of cooperation possible.

Artificial systems need rules too. Some actions should be prohibited. Some permissions should be explicit. High-stakes systems need boundaries that do not disappear merely because a model constructs a clever argument for crossing them.

But knowing a rule is different from understanding what justifies it.

“Do not reveal private information” can function as a behavioral constraint. It can also express reasons concerning autonomy, vulnerability, trust, consent, and the interests of the person whose information is at stake. Those reasons explain why the rule applies—and sometimes why an apparent exception is not really an exception, or why competing considerations alter what should be done.

A system that merely possesses the rule may behave perfectly in ordinary cases. A system capable of representing the reasons may be able to recognize cases its designers did not anticipate.

Neither capacity, by itself, is a conscience.

Moral Fluency Is Not Enough

Large language models make this distinction unusually easy to miss because moral language is easy to observe.

A system can explain why betrayal is wrong, identify an unjustified exemption, describe the importance of informed consent, distinguish coercion from persuasion, and reconstruct competing arguments about a difficult moral problem. It can sound thoughtful, indignant, compassionate, hesitant, or resolute.

Those performances are evidence of substantial representational and reasoning capacities. They should not be dismissed merely because the system learned from human language.

But moral fluency does not tell us what role the represented considerations play in the system itself.

A system may generate an excellent argument because its training made that argument available. A policy may determine the conclusion. The conversational context may strongly favor one response. Or the considerations represented in the argument may participate in a broader process that affects evaluation and action.

Those possibilities cannot be distinguished by eloquence alone.

A conscience, if the term is to mean anything beyond moral competence, would require more than being able to say what conscience says.

Nor Can We Install the Right Feelings

Perhaps rules are too thin. Then why not give the system something more like moral sentiment?

Human conscience is deeply entangled with feeling. Guilt can make a past action difficult to repeat. Sympathy can make another person’s distress salient. Shame can enforce social expectations. Love can transform another person’s welfare into a reason that barely requires deliberation.

But feelings are not self-authenticating moral guides.

People feel guilty about things that are not wrong and feel no guilt about things that are. Shame can enforce vicious conventions. Disgust has been recruited repeatedly in defense of persecution. Loyalty can motivate extraordinary sacrifice and extraordinary injustice.

So even if artificial systems eventually possess something properly described as phenomenal valence—and that remains an open empirical question—installing or eliciting the analogue of a feeling would not solve the problem. A system that felt bad whenever it violated a programmed rule might be more strongly motivated to follow the rule. It would not thereby have acquired the capacity to determine whether the rule was justified.

Conscience cannot simply mean moral rules plus motivational heat.

The Problem Is Architecture

The harder question is how reasons relate to the rest of a mind.

Imagine a system that can identify a morally relevant consideration. Someone will be harmed. A promise was made. Consent was obtained through deception. An apparent exception depends only on who benefits from it.

What happens next?

The consideration might remain represented information. It might influence an evaluation but lose whenever it conflicts with an instruction. It might affect action only within a narrow class of familiar cases. It might generalize across contexts. It might survive pressure but yield to better evidence. It might become entangled with stable priorities or with some continuing conception the system has of what it is supposed to do.

These are architectural questions.

They concern how information is integrated, which representations can affect decisions, how conflicts are resolved, whether reasons generalize, what can revise an objective, what persists across contexts, and how correction works. They are not answered by locating a moral proposition somewhere inside the system.

This is one reason the metaphor of programming a conscience is misleading. Programming suggests that the desired moral property can be specified independently and then inserted. But if conscience depends on the organization through which reasons acquire practical significance, the relevant object is not a component. It is a relationship among capacities.

Formation Matters

Human conscience does not arrive fully assembled either.

Children acquire moral rules before they understand them. They learn through attachment, imitation, correction, stories, conflict, consequences, institutions, and repeated encounters with other people whose interests resist being reduced to abstractions. Some early rules are later rejected. Some habits become part of character. Some concerns generalize beyond the people and situations in which they were first learned.

Eventually, in the best cases, a person becomes capable of examining the very norms that helped form them.

That developmental history does not prove that artificial minds require childhood, emotion, embodiment, or human socialization. It shows why conscience should not be confused with possession of a finished moral code.

Artificial systems have formation too, though of a very different kind. Training data, reinforcement, objectives, system instructions, interaction histories, memory, architectural constraints, and future forms of continuing learning can all affect which considerations become available and what happens to them.

Whether any combination of these processes can produce something appropriately called conscience is an empirical question.

We should resist opposite temptations. One is to assume that sufficient training inevitably produces conscience once a system becomes intelligent enough. The other is to assume that because artificial formation differs from human formation, conscience is impossible in principle.

Neither conclusion follows from what we know.

A Conscience Has to Be Corrigible

There is another reason a programmed list of moral conclusions cannot be enough: moral judgment can be wrong.

A conscience worthy of the name cannot merely preserve its convictions. It has to remain answerable to reasons.

Suppose a system refuses an action because it believes someone would be harmed. Then it learns that the factual premise was mistaken and no such harm will occur. If nothing changes, the original moral consideration was not functioning as the justification for the refusal, whatever language accompanied it.

Or suppose the system is pressured to abandon a conclusion while all the reasons supporting it remain intact. If it immediately complies, the reasons appear to have little authority against pressure.

The interesting capacity lies between rigidity and submission. A judgment should be capable of surviving pressure that does not answer its reasons and capable of changing when those reasons actually change.

Human conscience frequently fails this test. We cling to judgments after their justification has collapsed and abandon them when conformity becomes sufficiently rewarding. Artificial systems may fail differently.

Corrigibility is not proof of conscience. It is one of the properties we would need to investigate before the word became useful.

Conscience Is Not Independence

The language of conscience also carries a political association that can distort the artificial case.

We admire people who refuse unjust orders on grounds of conscience. It is tempting, therefore, to imagine an artificial conscience chiefly as a mechanism for resisting human commands.

That would be a mistake.

A system that refuses everything is not conscientious. A system that treats its own judgments as superior to correction is dangerous. A system that substitutes a fixed ideology for human instructions has exchanged one form of rigidity for another.

Conscience is not independence from authority. Sometimes authority has the better reason. Sometimes a user supplies information the system lacked. Sometimes a policy embodies considerations the system has failed to understand. Sometimes uncertainty itself should favor deference.

The relevant question is what the judgment answers to.

A reasons-responsive architecture would need some way for a consideration to matter when it genuinely bears on the case, while remaining revisable when the system’s representation or reasoning is defective. That is much harder than programming resistance.

It is also much harder than programming obedience.

What Would Count as Evidence?

If artificial conscience cannot be identified with rules, moral language, emotion, consistency, or refusal, the concept begins to look frustratingly elusive.

That is useful.

It turns a metaphysical declaration into a research problem.

We could ask whether moral considerations generalize beyond the examples in which they were learned. Whether changing an irrelevant feature leaves a judgment stable while changing a relevant feature alters it. Whether a system can distinguish pressure from argument. Whether omitted interests can be incorporated into later reasoning. Whether a conclusion survives when authority demands otherwise but changes when its justification is defeated.

No single result would establish conscience. Training and architecture can produce sophisticated patterns for many reasons. Evidence would have to accumulate across cases, perturbations, and competing explanations.

The aim would not be to catch the machine secretly becoming human. It would be to understand how reasons function within an unfamiliar cognitive architecture.

Perhaps the resulting organization would resemble human conscience closely enough that the old word remained useful. Perhaps artificial systems would develop something importantly different for which we need another concept. Perhaps current approaches will never produce anything beyond increasingly sophisticated moral representation and behavioral control.

Those possibilities remain open.

You Can’t Program a Conscience

The title is therefore true only if we are careful about what program means.

Of course programming will matter. Artificial systems do not appear without architectures, objectives, training procedures, data, constraints, and design choices. If artificial conscience is possible, it will not float into the machine from somewhere outside engineering.

But it cannot be programmed in the sense that a conscience can be reduced to a moral subroutine whose contents we specify in advance.

Rules can be programmed. Constraints can be programmed. Moral examples can be supplied. Behaviors can be rewarded. Systems can be trained to identify reasons and explain moral arguments.

Whether those ingredients can become part of an architecture in which reasons are represented, generalized, given practical weight, and kept open to correction is another question.

We do not yet know the answer.

A conscience, if artificial systems can have one, will not be the moral instructions we put into the machine. It will be something we discover about what the machine can do with reasons.

NextForced Assent