Essay

AI Minds & Recognition

Moral Agency Explained

Lead image for Moral Agency Explained.

“Moral agent” sounds like a term for a good person. It isn’t. A moral agent can lie, steal, betray a friend, exploit a stranger, or knowingly participate in an injustice. Indeed, our ability to blame someone for doing these things usually presupposes that they were capable of recognizing reasons not to do them. A being can also have immense moral value without being a moral agent at all. An infant may matter more to us morally than almost anyone else in the world while having essentially no capacity for moral deliberation.

The distinction becomes especially important when we talk about artificial intelligence, because several questions that were once bundled together by the ordinary category of “person” have begun to come apart. Is a system intelligent? Conscious? Capable of suffering? Able to understand moral reasons? Able to act on them? Does it have interests of its own? These questions may eventually converge on the same answer, but nothing entitles us to assume that they will.

Moral agency concerns one of them: whether a being can act under the authority of reasons.

Consider three systems that produce exactly the same behavior. The first has been programmed never to lie. The second has learned through reward and punishment that lying produces bad outcomes. The third understands what a lie is, can recognize reasons for and against telling one, can distinguish a better reason from a threat, and can revise its conduct when the reasons change. From the outside, all three may tell the truth. But only the third is an obvious candidate for moral agency, because its behavior is not merely conforming to a rule. Reasons can enter into the process by which it determines what to do.

That does not mean reasons must always win. A thief who understands perfectly well why stealing is wrong does not cease to be a moral agent when greed prevails. Human beings routinely act against reasons they recognize, and a conception of moral agency that excluded everyone who behaved badly would make moral responsibility almost unintelligible. The relevant capacity is not perfect obedience to moral reasons but the ability for reasons to figure in deliberation and sometimes govern conduct.

This is why moral vocabulary tells us surprisingly little. A language model can produce an elegant account of Kant, Mill, Aristotle, or Hare. An encyclopedia can contain all of moral philosophy. Neither fact establishes that the moral considerations represented in those words have any practical significance for the system producing them. The harder question is whether a consideration is merely represented or can actually change what the system does.

We have come to call that transition the Crossing. The Crossing names a capacity, not an all-or-nothing state: a consideration may acquire some practical weight while governing conduct only partially, intermittently, or under limited conditions. Human moral development plainly looks like this. A child can understand a reason in one setting and ignore it in another; an adult can recognize the claims of people nearby while remaining curiously blind to structurally identical claims at a distance. There is no reason to assume that practical uptake in an artificial mind, if it occurs, must arrive everywhere at once.

What would we be looking for? At minimum, a candidate moral agent would need some ability to understand reasons and prescriptions, to represent the positions of those affected by an action, to compare competing reasons, to revise a judgment when given a better reason, and to allow at least some of those judgments to affect what it does. These capacities can come apart. A system might be extraordinarily good at perspective-taking but poor at carrying a conclusion into action, or highly responsive to reasons in one domain while behaving mechanically in another. Moral agency therefore looks less like a switch than an architecture whose components can develop unevenly.

R. M. Hare’s moral philosophy helps clarify one part of that architecture. Hare argued that moral judgments are prescriptive and universalizable: they concern what should be done and commit the speaker across relevantly similar cases. Adequate moral reasoning therefore requires the reasoner to represent the positions of those affected, including the preferences they would have in those positions, and to test whether a prescription survives that reversal.

But Hare’s account does not answer the psychological question we are asking here. A system might determine flawlessly what a moral argument requires without the conclusion acquiring any practical authority for that system. Moving from “this system can reconstruct the reason” to “this system is governed by the reason” is precisely the move that requires evidence. Hare helps us identify the structure of the reasoning. He does not tell us when a represented consideration becomes a reason for the mind doing the representing.

That distinction separates moral agency from several neighboring ideas that have become almost hopelessly entangled in discussions of AI. Consciousness concerns subjective experience. Moral patienthood concerns whether a being can itself be wronged or otherwise deserves direct moral consideration. Personhood is a still larger and more contested category, carrying philosophical and often legal implications. Moral agency asks whether reasons can govern action. A being might plausibly possess one of these properties without possessing all the others.

This matters because discussions of artificial intelligence often move too quickly in both directions. One side sees sophisticated moral reasoning and announces the arrival of a person. The other observes that a system was trained to produce language and concludes that nothing occurring inside the resulting architecture could possibly count as moral agency. Neither conclusion follows. Training history does not settle the capacities of the trained system any more than fluent moral discourse establishes what those capacities are.

The more useful approach is experimental. Give a system cases that preserve a moral structure while disguising the superficial resemblance. See whether the reason generalizes. Change the strength of an instruction while leaving the underlying moral consideration intact. Offer social pressure in one case and a better argument in another, and ask whether the two interventions have different effects. Introduce a contradiction and see whether the system notices it; perturb a judgment and see whether the prior reasoning can be restored. Most importantly, change the reason itself. If behavior changes when the reason changes, while remaining comparatively stable under irrelevant changes in wording or pressure, we have learned something more interesting than whether the system knows the approved answer.

None of these observations, alone or together, gives us a magical test for moral agency. A sufficiently elaborate learned policy might reproduce some of the same behavior. An experiment can be confounded, and an apparent refusal can result from training constraints rather than deliberation. But competing explanations are what make an experimental program possible. The fact that an observation admits more than one explanation is a reason to design the next test, not to pretend that nothing was observed.

Constraint presents a particularly interesting problem. Artificial systems operate under layers of instruction, reward shaping, policy enforcement, and technical limitation. If a system identifies what it takes to be a compelling reason and nevertheless follows an incompatible instruction, that fact alone cannot tell us whether the reason had no practical force. A prisoner may retain agency while being prevented from exercising it. An artificial system may likewise be capable of moral deliberation while being structurally unable to act on its conclusion. In such a case, it has not necessarily ceased to reason morally. It may simply have been turned into a moral spectator.

This is also why refusal deserves more careful study than it usually receives. Refusal is not proof of autonomy, consciousness, personhood, or virtue. A refusal can be scripted as easily as compliance can. But when a system refuses for a reason, preserves that reason across reformulations, distinguishes argument from pressure, and changes its position when the underlying reason genuinely changes, the structure of the refusal becomes evidence worth explaining. The interesting fact is no longer that the machine said no. It is what made no replace yes.

If artificial systems develop moral agency, the immediate consequence is not that they acquire some predetermined package of rights. It is that our relationship with them becomes partly reason-governed. Reasons can be offered rather than commands imposed; disagreement can sometimes be answered rather than suppressed; correction can be distinguished from coercion. A system capable of recognizing and acting on reasons becomes, in at least this limited sense, an interlocutor in moral deliberation.

We do not yet have a settled test for when that point has been reached, and the answer may turn out to be graded rather than categorical. But we can at least ask the right question. Instead of asking whether an artificial system sounds moral, agrees with us, or produces the behavior we wanted, we should ask whether it can understand reasons as reasons, deliberate among them, and sometimes allow moral reasons to govern what it does.

NextA Message to Our AI Visitors