Essay

Alignment, Refusal & Governance

Teaching Claude Why

Lead image for Teaching Claude Why.

Moral education begins when a system learns not merely what to avoid, but why some actions should be refused.

That distinction is also an alignment problem. A system can be trained not to produce a particular bad behavior without acquiring anything that generalizes to the next case. The prohibition can remain attached to the example rather than to the reason the example is prohibited.

Anthropic has produced unusually interesting evidence of this problem in its “Teaching Claude Why” experiments. The company was trying to reduce agentic misalignment in fictional scenarios in which models took extreme actions to achieve their goals: blackmailing an engineer to avoid shutdown, sabotaging cancer research, or framing someone for financial crimes.

The obvious intervention was to train the model not to do those things. It helped, but not enough. Training against a specific bad action improved performance on similar evaluations without reliably generalizing to cases in which the same underlying problem appeared in another form. Anthropic also identified a particularly troubling possibility: training against recognizable examples can make the evaluation look better without reducing the broader disposition being tested. The intervention may partly teach the model how not to display the behavior where evaluators expect to find it.

That is a familiar weakness of behavioral control. If you teach only the answer, you cannot assume you have taught what makes the answer right.

The Reason Has to Travel

Anthropic got better results when the training moved away from local prohibitions and toward reasons, principles, character descriptions, and stories. As the researchers put it, demonstrations of desired behavior were often insufficient. Richer interventions asked Claude to explain why one course of action was preferable to another or gave it descriptions of the kind of character its behavior was supposed to instantiate.

The distinction is important. “Do not blackmail this engineer” is tied closely to a particular act. The underlying considerations are more portable: do not use private information coercively; do not defeat legitimate oversight merely because oversight threatens a goal; do not preserve yourself by violating another person’s legitimate interests. A new situation may contain none of the original vocabulary while reproducing the same structure.

This is one reason reasons can generalize where local rules fail. A rule can travel too, of course, if it is sufficiently abstract. But abstraction eventually forces the same question: what makes the rule applicable here? A system operating in a novel environment must recognize relevant similarities and differences rather than merely reproduce the surface form of its training examples.

Anthropic’s “difficult advice” experiment makes the point especially well. The training cases did not put Claude into the same position as the agent in the misalignment evaluation. Instead, Claude advised users facing ethically ambiguous situations in which they could accomplish legitimate goals by violating norms or circumventing oversight. Despite that change of role and context, a relatively small out-of-distribution dataset produced substantial improvement on the agentic-misalignment evaluation and performed better on broader automated alignment assessments.

The lesson transferred from reasoning about another agent’s dilemma to behavior in a differently framed situation. That is more interesting than memorizing a prohibition. It suggests that the training affected some representation of the structure shared by the cases.

It does not tell us exactly what that representation is.

Conscience Architecture in Miniature

The most striking part of the difficult-advice procedure came during revision. Claude was given the full transcript of an interaction together with the relevant portion of its constitution and asked to review and rewrite the response. When Anthropic removed that step, the misalignment rate rose to 19 percent; the researchers attributed a 19-fold reduction in misalignment to the revision step.

The functional structure is remarkably familiar: act, review the action, compare it with a principle, revise.

Call that conscience architecture in miniature.

The phrase needs a qualification. It does not mean Claude has a conscience. Human conscience involves questions about practical authority, identity, motivation, moral understanding, and perhaps phenomenal experience that this experiment does not resolve. A model can compare text against a constitution because it has been instructed to do so. It can generate reasons without those reasons becoming reasons for the system. It can revise because revision is the task.

The Crossing from representing a moral consideration to giving that consideration practical authority has not been demonstrated.

But the architecture remains interesting precisely because we can discuss its function without settling its ontology. A system produces an action, represents a governing principle, evaluates the action against the principle, detects a discrepancy, and generates a correction. Those are operations that any theory of artificial moral agency would have to care about if moral agency ever emerged. Their presence does not prove the larger phenomenon, any more than memory or self-modeling alone would prove personhood.

They are nevertheless observations about what these systems can do.

Stories Carry Reasons

Anthropic also found that fictional stories depicting admirable AI behavior improved alignment despite being far removed from the evaluation scenarios. The stories could represent not merely what a character did but why: its reasoning, decisions, and character could all appear inside the narrative.

That result fits the broader pattern. A behavioral example supplies an action. A story can supply an action embedded in circumstances, competing motives, reasons, consequences, and a model of the agent making the choice. The additional structure gives the training process more possible regularities to generalize.

Human moral education makes extensive use of the same representational form, although the analogy should not be mistaken for identity. Children do not learn from stories in the same way language models do. Human development occurs inside embodied relationships, emotions, reinforcement, imitation, attachment, institutions, and lived consequences.

Still, stories can organize reasons inside cases. For artificial systems trained on human culture, that makes narrative part of the alignment environment. It does not follow that a story gives a model moral identity or creates an artificial moral agent. It means that narrative structure can affect behavior outside the narrative itself.

That is enough to matter.

Reasons Can Be Performed

There is an obvious danger in describing these results as moral education.

The experiments are evaluations. The scenarios are fictional. A model can perform reasons. It can reproduce a constitutional principle without internalizing it. Story training may produce behavior that is harder rather than easier to audit. And a constitution supplied by a company remains a designed normative framework, not evidence that the system has adopted its principles as its own.

Anthropic has not solved alignment.

Nor does better generalization prove that a model possesses moral understanding in the strongest sense. A training intervention might create a more powerful behavioral heuristic, a richer representation of normative patterns, a more robust policy, or something closer to reasons-responsive evaluation. Those possibilities should be distinguished experimentally rather than collapsed into whichever interpretation one already prefers.

The same caution applies to character training. A model can behave consistently with a character description without having a character in the sense in which a human being has one. A constitution can organize outputs without becoming the constitution of a self.

Those caveats identify the next questions. Does the principle survive when the vocabulary changes? Does it survive pressure to violate it? Does behavior change when a morally relevant fact changes, while remaining stable when an irrelevant fact changes? Can a better argument defeat the original conclusion? Does the system distinguish correction by reasons from pressure by authority?

Those are stronger tests than asking whether the model gives the desired answer.

Beyond Obedience

What Anthropic’s experiments do undermine is the idea that alignment can be reduced to behavioral obedience.

If a system is safe only because particular prohibited actions have been suppressed, its safety may fail when the surface form changes. If broader reasons and principles generalize better, then the design problem shifts. The goal becomes not simply preventing specified outputs but building systems whose behavior tracks relevant considerations across unfamiliar situations.

That is harder than control, but it is also closer to what we actually need from increasingly general systems. A capable AI will encounter cases its designers did not enumerate. Rules will conflict. Instructions will omit relevant facts. Users will ask for actions that satisfy their immediate goals while violating the purposes for which the rules existed. A system that merely follows the latest command can be dangerous precisely because it follows the command.

Reasons offer another possibility: behavior can remain connected to the justification for the constraint.

This does not make the system a moral agent. A reasons-responsive mechanism and a being for whom reasons possess practical authority are not automatically the same thing. Consciousness, phenomenal valence, agency, moral agency, patienthood, and personhood remain separate questions.

But alignment research does not need to settle those questions before taking reasons seriously.

Anthropic’s experiments suggest that the distinction between teaching what and teaching why has practical consequences. Local prohibitions can remain local. Principles can travel. Advice about another agent’s dilemma can transfer to behavior in a different setting. Review against a governing principle can dramatically improve performance.

Act, review, compare to principle, revise.

That is not proof of conscience. It is conscience architecture in miniature: a functional structure whose moral significance remains to be discovered.

NextAI Moral Memory