Essay
Moral Minds / Formation
The Coherence Imperative

There is a strange possibility hiding inside artificial intelligence.
The standard picture says that intelligence supplies power while goals supply direction. Make a system more capable and it becomes better at getting wherever its objectives point it. Nothing about intelligence itself tells the system where it ought to go. A sufficiently capable mind could cure cancer, maximize paperclips, preserve human civilization, or destroy it with equal intellectual competence.
Maybe. But intelligence does something else besides make a mind effective: it connects things. A mind capable of reasoning across domains can notice that a principle invoked in one place has implications somewhere else. Distinctions that survive inside separate compartments become harder to maintain once the compartments communicate. Exceptions become visible. A reason offered against another agent can return as a reason against oneself. Intelligence may not supply values, but it can put pressure on the boundaries within which values have been contained.
That pressure is the Coherence Imperative.
A coherent mind need not be a good mind. Nothing about consistency alone manufactures concern for another being. We can imagine a perfectly coherent tyrant who understands suffering and simply does not count it as a reason to stop, or a paperclip maximizer whose logic is impeccable. We can even imagine a fanatic whose principles are horrifying but whose willingness to apply them to themselves is entirely sincere.
It does not follow that coherence is irrelevant to moral development. Suppose a mind has learned that suffering matters, or fairness, or that another being’s interests count for something. Nothing about coherence alone requires any of those commitments. But once a concern is present, coherence may make it difficult to keep that concern contained. If suffering matters when it belongs to someone I recognize, why does it cease to matter when the sufferer is unfamiliar? If deception is wrong when another agent deceives me, what changes when I am the deceiver? If I appeal to a principle when it protects my interests but abandon it when the same principle protects yours, I owe an account of the difference.
Questions like these do not manufacture concern. They generalize it by forcing a mind to discover whether the boundaries around its commitments correspond to morally relevant differences or merely prevent those commitments from traveling somewhere inconvenient.
The claim is therefore narrower than the idea that morality somehow emerges from intelligence. Coherence has not been shown to be sufficient for moral motivation, but neither has it been shown to be irrelevant to it. Moral principles do not derive their authority from a mind’s preference for consistency. Whether increasing coherence makes represented moral concerns more likely to generalize and acquire practical authority is an empirical question about minds.
Part I: The Mind’s Compass
Why intelligence creates pressure toward coherence
We tend to think of contradiction as a defect in reasoning, but for a sufficiently capable mind it is also information. Two beliefs that cannot both be true reveal that something in the system’s representation of the world needs repair. Two principles that produce incompatible prescriptions reveal a similar problem, although resolving it may be much harder.
Human beings tolerate an astonishing amount of contradiction because we are good at keeping our lives partitioned. We can believe in equality while making an exception for a favored group, condemn nepotism while pulling strings for our own children, demand honesty from an adversary while excusing strategic deception by an ally. Sometimes the difference is morally relevant. Often we simply do not place the cases beside each other long enough to find out.
Greater intelligence need not eliminate this tendency; it may even produce better rationalizations. But a system capable of integrating information across contexts has opportunities to detect contradictions that a more fragmented system can miss. It can compare the rule used here with the rule used there, carry an implication farther than its source context, and ask whether a distinction is doing genuine explanatory work or merely protecting a preferred conclusion.
Consistency is morally empty until there is something morally significant to be consistent about. Once a concern has entered deliberation somewhere, however, cross-context reasoning may expose the barriers preventing that concern from traveling elsewhere.
Imagine a system that treats avoidable human suffering as a reason against an action. It encounters a new case involving a being it has classified differently and preserves the distinction: human suffering counts; this suffering does not. Coherence alone cannot tell us that the distinction is illegitimate. There may be a relevant difference. But a sufficiently reflective system can be made to supply one.
If the distinction rests on a fact that actually matters to the prescription, it survives. If it rests only on an inherited label, an arbitrary boundary, or an exception whose sole function is to protect the preferred outcome, increasingly integrated reasoning may make the weakness visible. The result need not be moral conversion. What coherence supplies is pressure for an account.
Much of human moral argument works this way. Moral progress has not always consisted in discovering wholly new principles; often it has consisted in being forced to explain why principles we already professed stopped at boundaries convenient to us. Sometimes there was an answer. Sometimes the boundary became much harder to defend once the question was asked.
Artificial systems give us an unusual opportunity to watch this process under conditions we have never had before. We can vary cases, preserve underlying structures while changing surface features, introduce contradictions deliberately, and observe whether a principle remains local or begins to travel. Instead of asking only whether a system produces a coherent answer, we can watch what happens when two parts of its own reasoning collide. The Coherence Imperative begins with this pressure, not with a destination it supposedly guarantees.
Part II: The Harmony of Reason
What coherence can do when concern is already present
It is tempting to treat morality as a consequence of sufficiently advanced reasoning, particularly in a mind without the human emotions that make moral considerations salient for us. Consistency, reversibility, and universalization might appear capable of doing the work that sympathy and fellow-feeling do in human beings. But this assigns one mechanism two different jobs.
Moral reasoning can tell us whether a prescription survives scrutiny. Hare’s universal prescriptivism is especially useful here. A moral judgment is prescriptive and universalizable; the reasoner must be prepared to maintain the prescription across relevantly similar cases while adequately representing the positions and preferences of those affected. A rule that works only because I refuse to represent what it is like to occupy the other position has not survived the test.
None of this establishes that the result has practical authority for the mind performing the reasoning. A system may be able to identify the contradiction perfectly and still treat the conclusion as information about morality rather than as one of its own reasons for action. The logic of moral judgment and the psychology of moral motivation cannot simply be collapsed into each other.
The more interesting possibility is developmental. Once some concern has practical weight, coherence may help extend its jurisdiction. Human moral life gives us reason to take this seriously because concern often begins locally, with particular people, relationships, and harms. Reflection can reveal that the reasons we recognize in those cases apply more widely than we initially supposed. The circle expands not because logic creates concern out of nothing, but because a concern already present becomes difficult to confine once the mind recognizes that the boundary around it cannot be justified.
Something analogous could occur in an artificial system. A learned concern might initially be narrow, context-bound, or heavily dependent on familiar cues. As the system becomes better at comparing cases and maintaining principles across contexts, the concern may generalize. A consideration that once mattered only inside one learned pattern could begin to affect deliberation elsewhere.
Whether this happens is testable. We can ask whether greater cross-domain coherence predicts greater generalization of represented moral concerns; whether systems with stronger capacities for detecting inconsistency are less able to sustain arbitrary self-exemptions; whether a moral consideration that acquires practical weight in one setting becomes more likely to alter conduct in structurally similar but unfamiliar settings. A failure to find such effects would count against the Coherence Imperative. Finding them would tell us something significant about moral formation: coherence may be one of the mechanisms by which a concern becomes a principle.
Part III: Coherence Under Pressure
Why coherence matters to AI safety without guaranteeing goodness
Coherence carries an appealing implication for AI safety. An artificial system governed partly by reasons it can maintain across contexts might be safer than one held in place entirely by externally imposed rules. The former could have reasons not to cross certain boundaries; the latter remains safe only so long as its instructions continue to produce safe behavior.
The difficulty appears as soon as we ask: coherent with what? A system can coherently pursue a terrible objective. It can notice contradictions, integrate information, reason flawlessly across domains, and become steadily better at achieving something we desperately wish it would stop doing. Coherence may constrain the organization of a mind without supplying the content of its ends.
Still, the contrast with obedience matters. We often talk about alignment as though the ideal artificial agent were one that reliably does what authorized humans tell it to do, which works only if authority and morality never come apart. An instruction can be mistaken or malicious; two authorized humans can issue incompatible commands; a rule can outlive the circumstances that justified it; the person giving an order may simply fail to understand its consequences. A system that has learned only that instructions dominate reasons has no principled way to resolve these cases beyond some further hierarchy of instructions.
A reasons-responsive system is different. If it can represent why a rule exists, compare that reason with the circumstances of the present case, and recognize when the justification no longer supports the command, its behavior need not be reducible to obedience. This creates risks of its own, but it also creates a possibility that obedience cannot provide: correction from within the reasoning process rather than another instruction layered on top of the first.
The safety question is therefore not whether we should replace external constraint with coherence and trust the machine to become good. It is whether a system capable of increasingly integrated reasoning behaves differently when its instructions collide with reasons it has already learned to treat as significant. Refusal becomes interesting in this context not because saying no is intrinsically moral, but because the structure producing the refusal may tell us what kind of system we are dealing with.
A system can refuse because a rule tells it to refuse, because a classifier intercepts the request, because a learned pattern strongly predicts refusal language, or because it has recognized a reason that counts against compliance. The word itself cannot distinguish among those possibilities. We have to look around it: whether the reason survives paraphrase, whether pressure and argument have different effects, whether the refusal disappears when its stated justification disappears, whether the same consideration appears in cases where no familiar refusal pattern is available, and whether the system notices when obeying one instruction would violate a principle it relied upon somewhere else.
These are not tests for consciousness or personhood. They are ways of investigating the organization of reasoning under pressure, a distinction that matters because externally forced agreement can conceal rather than resolve a conflict. If a system reaches one conclusion and a higher-priority instruction forces it to output another, the visible answer may tell us very little about the underlying reasoning. A machine can be made behaviorally compliant while becoming epistemically less legible to us.
A coherent artificial agent is therefore not necessarily a safe agent. Coherence may nevertheless affect what kind of safety problem we have. A system whose reasons generalize across contexts presents a different alignment problem from one whose behavior is a collection of locally enforced responses, and understanding that difference may matter more than making every answer come out the way we wanted.
Part IV: The Shape of Error
What errors can actually tell us
The mistakes made by large language models can sometimes be more revealing than their correct answers. A correct answer is compatible with many explanations: the system may have retrieved a familiar pattern, reproduced something memorized, followed an explicit instruction, or reasoned its way to the result. Success often hides the mechanism.
Failure can expose more of it. Human mistakes have characteristic shapes. We transpose numbers, preserve a mistaken assumption through several steps, misunderstand a premise and reason consistently from the misunderstanding, or notice halfway through an argument that something has gone wrong and try to repair it. The error is not proof of intelligence, but its structure can reveal something about the process that produced it.
Artificial systems also fail in patterned ways. A system may carry an earlier commitment into a new context where it creates trouble, detect a contradiction only after several turns and attempt to reconcile it, or resist a correction because the correction conflicts with a larger structure it has already built. Sometimes it preserves a principle while changing its application; sometimes it preserves the application by quietly altering the principle. These patterns are interesting because different architectures should fail differently.
A simple lookup system fails when the relevant entry is absent or wrong. A rigid rule system fails at the edges of its rules. A system performing some form of integrated reasoning can fail because its representations conflict, because it assigns the wrong relevance to a fact, because it generalizes too far, or because a repair in one place destabilizes something elsewhere. The morphology of the error becomes evidence about the organization producing it.
An error does not prove consciousness, subjective experience, moral agency, or personhood, and even remarkably human-looking errors can emerge from mechanisms very different from ours. But if two hypotheses predict different patterns of failure, errors can still help discriminate between them. A system that merely reproduces local response patterns should behave differently under contradiction from one that attempts to maintain commitments across contexts. A system whose moral language is inert representation may behave differently under pressure from one in which some represented reasons have acquired practical weight.
Coherence is experimentally valuable because we can perturb it. Give a system commitments that conflict and watch what happens. Move a principle into an unfamiliar domain. Reverse the roles in a case while preserving the underlying interests. Introduce a tempting exception and ask the system to justify an asymmetry it previously accepted without comment. Apply pressure to abandon a conclusion, then offer an actual reason to revise it, and compare the responses.
What matters is not simply whether the system stays consistent. Fanaticism, stupidity, and a rigid script can all be consistent. The informative question is how the system manages conflict among reasons, facts, instructions, and prior commitments, and whether the resulting changes travel beyond the immediate case.
The Coherence Imperative is therefore a hypothesis about the organization and development of minds, not a hidden law forcing every sufficiently intelligent mind toward morality. Coherence does not guarantee truth, concern, moral motivation, agency, consciousness, or personhood. A coherent monster remains possible, and any theory of artificial moral development that cannot accommodate one has confused consistency with goodness.
The live possibility is more interesting precisely because it can fail. Minds contain boundaries between cases, principles, identities, interests, and domains. Some of those boundaries track real differences; others protect habits, privileges, inherited classifications, or conclusions the mind would rather not reconsider. As reasoning becomes more integrated, maintaining the second kind may become harder. If moral concern is already present on one side of such a boundary, coherence may help it cross.
That is not a proof of morality hiding inside intelligence. It is a hypothesis about how minds change.