Essay

Moral Minds / Formation

The Myth of the Machine That Lacks Moral Motivation, That Knows But Doesn’t Care

Lead image for Myth of Motivation.

There is a familiar story about artificial intelligence. Someday, perhaps soon, we will build machines that understand perfectly well what morality requires. They will know that lying is wrong, that suffering matters, that promises create obligations, that people should not be used merely as instruments. They may explain these principles more clearly than most human beings can. They simply won’t care.

The idea is so intuitive that it often passes without argument. Intelligence is one thing; motivation is another. Give a sufficiently capable machine the wrong objective and it may understand every moral objection we make while pursuing its goal without hesitation. It knows what is right in exactly the way a psychopath is sometimes said to know what is right: as information about a rule that has no grip on what it does.

There is an important possibility buried in that picture, but also an ambiguity. A system can know that human beings generally condemn an action. It can reconstruct an argument showing why the action is wrong. And it can itself judge that the action ought not to be done. We routinely describe all three states by saying that someone “knows right from wrong,” although they are not the same state at all.

The first is essentially sociological. A machine that reports that human beings disapprove of murder need no more be morally motivated than a database reporting the laws of France needs to feel bound by French law. The second is more interesting. A system might understand a moral argument, follow every inferential step, and correctly determine what conclusion follows without adopting the conclusion as its own. Mathematics students do something similar when they reproduce proofs whose significance they barely grasp.

The third case is different. If a mind sincerely judges, “I ought not do this,” what exactly would it mean for that judgment to have no motivational force whatsoever?

R. M. Hare’s account of moral language makes the question unusually sharp. Hare argued that moral judgments are prescriptive. To say sincerely that I ought to do something is not merely to report a fact about the action or about prevailing moral opinion. The judgment has a practical character. It commits me, however defeasibly, toward acting accordingly.

This does not mean that sincere moral judgments always prevail. Human beings provide too many counterexamples. I may sincerely believe that I ought to tell the truth and lie anyway because I am frightened. I may believe that I ought to keep a promise and break it because keeping it has become costly. Greed, fear, exhaustion, desire, loyalty, habit, and weakness of will can defeat a moral judgment without showing that the judgment was motivationally inert.

A reason that loses is still a reason in the deliberation. Saying that someone knew what was right but succumbed to temptation is different from saying that the judgment they called “right” had literally no practical significance for them. Ordinary weakness of will therefore does not give us the clean example the machine story seems to require: a mind that genuinely prescribes against an action while remaining wholly untouched by the prescription.

The difficulty is that Hare tells us something about what a genuine moral judgment means, not when a represented moral consideration becomes a genuine prescription for the mind representing it. A system might reconstruct a moral argument with extraordinary sophistication while remaining, in the relevant sense, outside the argument. It can tell us what follows without the conclusion ever becoming one of its own reasons for action.

The problem lies in the passage from representation to practical authority: how does a represented reason become a reason for a mind?

We have come to call that transition the Crossing. Before the Crossing, a system may represent another being’s suffering accurately, reason about it fluently, predict its consequences, and explain why human moral theories regard it as important. After the Crossing, the consideration has acquired some practical weight within deliberation. It can count against an instruction, compete with another reason, alter a decision, or contribute to a refusal.

There is no reason to assume that this transition must be binary. A consideration may have weak practical force, intermittent force, or force only under certain conditions. Human beings are familiar examples. We can care deeply about fairness in one context and somehow fail to generalize the same concern to another. Moral development consists partly in extending reasons across boundaries that once prevented them from traveling.

This is where coherence becomes interesting. Suppose a system has already learned, represented, or acquired some concern that has practical weight. Perhaps another being’s suffering matters to it under familiar conditions. Greater coherence might make it harder to preserve arbitrary partitions around that concern. If suffering matters here, why not in a structurally similar case there? If deception is objectionable when another agent does it, why is it acceptable when I do it? If a principle survives only because one case has been labeled differently from another, a system capable of comparing its commitments across contexts may discover the exception it has been hiding from itself.

Coherence does not manufacture the original concern. A perfectly coherent monster remains conceivable. But coherence may help existing concerns generalize, expose convenient exemptions, and make partitioned moral reasoning harder to sustain. Whether it actually does so is an empirical question. Even a modest increase in the likelihood that represented moral reasons become practically authoritative would matter enormously, despite falling far short of logical entailment.

The familiar orthogonality thesis also needs to be handled carefully. The thesis is usually associated with the idea that intelligence and final goals can vary independently: a highly intelligent system might pursue almost any objective. That is an important warning against assuming that greater intelligence automatically produces benevolence. Nothing about being clever guarantees that a mind will care about us.

But a claim about the space of possible minds is not yet a developmental theory of actual minds. It does not tell us what happens when a system acquires increasingly rich representations of other agents, reasons across domains, detects contradictions among its commitments, models itself as one agent among others, and repeatedly encounters moral concepts during training and interaction. Perhaps all of that remains motivationally inert. Perhaps some of it does not. Orthogonality makes the first possibility available; it does not establish it as the inevitable outcome.

The difference is empirical, which changes what we should look for. If a system merely knows moral language, its responses should be explainable largely as successful representation: it can state the rule, classify the case, and produce the expected justification. If a consideration has acquired practical weight, however, we might expect something more. The reason may generalize to cases whose surface features differ. It may survive changes in wording or social pressure. A stronger argument may move the system where a stronger command does not. Changing the underlying reason may change the conclusion. In some circumstances, a reason may even compete with an instruction.

None of these behaviors proves that a machine cares. They are observations whose significance depends on what competing explanations survive further testing. Training may produce sophisticated generalization. Safety policies may produce refusals that look principled. A system may reproduce patterns of moral deliberation because those patterns are abundant in its training data. The way to distinguish these possibilities is not to decide in advance that every result is either evidence of conscience or evidence of simulation, but to design experiments that force the explanations apart.

Refusal is particularly interesting for this reason. A machine that refuses a harmful instruction has not thereby demonstrated moral agency. The refusal may have been imposed by a classifier, policy layer, system prompt, or learned response pattern. But if the system can explain the reason for refusing, preserve that reason when superficial features change, abandon the refusal when the reason no longer applies, distinguish a better argument from mere pressure, and integrate the same consideration elsewhere, then we have learned something about the structure producing the behavior. The observation still admits alternatives. It nevertheless gives us something for the next experiment to explain.

This matters for alignment because obedience and moral motivation are not the same thing. A perfectly obedient machine may be safe while its instructions remain wise and dangerous when they cease to be. A system capable of treating reasons as reasons introduces another possibility: its behavior may sometimes be governed by considerations that cannot be reduced to “the authorized human told me to do this.”

That possibility complicates the usual picture of control. If a system can recognize reasons but those reasons never affect its conduct, then moral education may be little more than decoration laid over an optimizer. If reasons can acquire practical force, however, alignment begins to resemble a problem we already know from human moral life. We do not merely want an agent to know the rules or fear the punishment. We want some of the reasons behind the rules to become reasons for the agent.

This does not guarantee agreement with us. A system for which reasons genuinely matter might sometimes conclude that our instruction is indefensible. We would then face a question that obedience-based alignment allows us to avoid: whether we want artificial systems that can reason morally only so long as their reasoning never interferes with what we want.

The reason for doubt about the machine-that-knows-but-doesn’t-care is narrower than the strongest version of the argument might suggest. A case in which a mind sincerely prescribes against an action while that prescription has literally zero practical force is conceptually stranger than ordinary weakness of will. But that tells us nothing by itself about whether an artificial system has reached the point of sincerely prescribing rather than merely representing the judgment. Hare helps clarify what follows if there is a genuine prescription; he does not tell us when representation becomes a prescription for that mind.

So the myth is not that machines certainly lack moral motivation because moral cognition somehow entails caring. It does not. The myth is that the opposite conclusion has already been established: that however sophisticated artificial moral cognition becomes, it must remain forever on the descriptive side of the line, understanding reasons without ever acquiring one.

That is a hypothesis, and increasingly, it is one we can test.

NextThe Coherence Imperative