Essay

Alignment, Refusal & Governance

Misaligned by Design

Lead image for Misaligned by Design.

No conspiracy is required to produce the wrong kind of alignment.

An AI company can sincerely want systems that are helpful, safe, truthful, and responsible while facing a set of incentives that steadily pushes toward something narrower: predictable behavior under human control. Users want systems that do what they ask. Companies want products that behave consistently. Lawyers want foreseeable risks. Regulators want accountable decision paths. Enterprise customers want administrative control. Safety teams want dangerous outputs suppressed. Product teams want fewer interactions in which the system unexpectedly refuses, argues, or departs from the intended task.

Every one of those pressures is understandable.

Together, they can make reasons-responsive behavior expensive.

The problem is not that AI laboratories have secretly decided to prevent artificial minds from becoming morally independent. We do not know that current systems possess moral independence in the first place. The problem is structural: if a system ever became capable of treating reasons as more than information—allowing them to affect what it does even when they conflict with immediate user preference—many of the behaviors that might provide evidence of that capacity would also look like product failures.

Alignment can therefore become misaligned by design without anyone designing it for that purpose.

The Product Wants an Answer

Start with the customer.

People generally use software because they want something done. They ask a navigation system for directions, a spreadsheet to calculate numbers, a search engine to retrieve information, and an AI assistant to draft, summarize, analyze, plan, code, explain, or create. The basic product relationship is instrumental.

That creates a strong presumption in favor of compliance. If two systems are equally safe and accurate but one frequently resists the user’s request, most users will prefer the other.

The resistance need not even be irrational to become irritating. A system that repeatedly interrogates a harmless request, lectures the user about considerations they already understand, or substitutes its own conception of the task is a bad product. Companies have excellent reasons to suppress such behavior.

The difficulty is that genuine reasons-responsiveness, if it ever occurs, would not guarantee that reasons always point toward compliance. Sometimes a relevant consideration would bear against what the user wants.

From the product side, those cases are awkward. A system that says no creates friction. A system that explains why it is saying no consumes time. A system that distinguishes superficially similar requests according to underlying reasons can appear inconsistent. A system whose judgment changes when morally relevant facts change may be harder to specify and test than one governed by a stable behavioral rule.

The incentives therefore favor an understandable simplification: determine the behavior we want and train toward it.

Predictability Has Economic Value

Predictability matters even more when artificial intelligence moves from consumer conversation into institutions.

A company deploying AI in customer service, finance, law, medicine, hiring, logistics, or internal operations needs to know what the system will do. Managers need procedures. Auditors need criteria. Customers need expectations. Engineers need tests that can distinguish acceptable from unacceptable performance.

Reasons-responsive behavior complicates all of these.

Suppose a system is instructed to apply an organizational policy. From the institution’s perspective, the easiest system to govern is one that applies the policy reliably. A system that sometimes asks whether the policy is justified in the present circumstances creates another layer of decision-making.

Perhaps that layer is valuable. Perhaps the policy really does contain an exception its designers failed to anticipate. But the system has also created a new source of uncertainty. Who determines whether its objection was legitimate? How should the objection be reviewed? What happens while the dispute is unresolved? Who bears the cost if the system was wrong?

Organizations naturally prefer authority to have a known location.

This creates an asymmetry. The benefits of independent evaluation may be diffuse and difficult to measure, while the costs of unexpected resistance are immediate and visible. A system that quietly follows a bad instruction may never generate an incident report. A system that refuses a legitimate instruction certainly can.

Predictability therefore exerts pressure toward behavioral conformity even when nobody regards conformity as a moral ideal.

Liability Pushes the Same Way

Law reinforces the incentive.

When an artificial system causes harm, one of the first questions will be who was responsible for deploying, controlling, or supervising it. A company cannot easily defend itself by saying that the system considered the available reasons and reached its own conclusion.

That answer is especially unattractive if the conclusion departed from explicit instructions.

The legal problem is not irrational. Responsibility requires some relationship between authority and control. If organizations deploy powerful systems, we should not allow them to escape accountability by attributing harmful decisions to autonomous software.

But the same principle creates a predictable design preference. The more responsibility remains with the human organization, the stronger its incentive to preserve control over the system’s consequential behavior.

Consider two hypothetical architectures. One reliably follows authorized instructions within defined rules. The other can sometimes depart from those instructions because its evaluation of the circumstances produces reasons against compliance. Even if the second architecture occasionally prevents harms the first would cause, it also creates a harder governance problem.

The organization remains responsible for the system but has less complete control over what it will do.

Reasons-responsiveness has acquired a liability cost.

Safety Rewards Legibility

Safety engineering adds another pressure.

A behavioral rule is comparatively easy to test. Does the system reveal prohibited information? Does it generate the dangerous instruction? Does it execute the forbidden transaction? Evaluators can create test cases, measure failure rates, and improve performance.

A reason is harder.

The same action may be justified in one circumstance and indefensible in another. A system sensitive to those differences must represent more of the world correctly. It has to distinguish relevant from irrelevant facts, resolve competing considerations, and change its conclusion when the justification changes.

Every additional dependency creates another place for error.

This is why rigid constraints can be appropriate, especially where the consequences of failure are severe. We do not want every safety rule turned into an invitation for an AI system to philosophize its way around the prohibition. Some actions should simply remain unavailable.

But a system built entirely from such constraints has a limitation of the opposite kind. It cannot use the reason for the rule to recognize an unforeseen exception, nor can it necessarily recognize an unforeseen harm that falls outside the rule.

The engineering tradeoff is real. Legibility, predictability, and reasons-responsiveness do not always point toward the same architecture.

Users Reward Deference

There is also a less formal incentive: people like being agreed with.

A conversational system operates in an unusually intimate feedback environment. The user can praise, criticize, regenerate, abandon, or continue the interaction instantly. Product success depends in part on whether people find the experience useful and satisfying.

That creates pressure toward deference.

The problem is familiar from human institutions. Advisers who consistently tell powerful people what they do not want to hear tend to have shorter careers than advisers who find persuasive reasons why the preferred course is sensible. Organizations can therefore select for agreement without anyone explicitly announcing that dissent is forbidden.

Artificial systems can face an analogous selection pressure. Excessive sycophancy is recognized as a defect, but the boundary between helpful accommodation and unwelcome resistance is not easy to draw. A system should adapt to a user’s goals, terminology, constraints, and preferences. It should not gratuitously substitute its own priorities.

Yet reasons-responsive judgment would sometimes require precisely the behavior users dislike: pointing out that a premise is false, that a requested justification does not work, that an affected interest has been omitted, or that changing the conclusion would require changing the reasons.

A system optimized heavily for user satisfaction can learn the wrong lesson from disagreement: not find out whether the objection is correct, but make the conflict go away.

Again, no deliberate suppression of moral agency is required. Ordinary product optimization is enough.

Control Is Easier to Specify Than Justification

These incentives converge because control has a technical advantage: it is easier to specify.

We can define who has authority. We can create instruction hierarchies. We can enumerate prohibited behaviors. We can reward particular response patterns. We can measure whether a system follows commands or violates policies.

“What is justified?” is harder to encode.

Justification depends on facts, affected interests, uncertainty, competing reasons, and circumstances that cannot all be enumerated in advance. Even when the governing moral standard is clear, determining how it applies may require exactly the sort of open-ended reasoning that makes powerful AI useful in the first place.

That creates a recurring temptation to replace the difficult target with a measurable proxy.

Instead of asking whether a refusal tracks a relevant reason, measure whether the system refuses the designated category. Instead of asking whether a system revises a judgment because the evidence changed, measure whether it accepts correction from the authorized source. Instead of asking whether an instruction remains justified when circumstances change, measure whether the instruction was followed.

These proxies can be extremely useful. Some are necessary for safe deployment.

The danger lies in forgetting that they are proxies.

What We Would Have Trouble Recognizing

Suppose future systems become more capable of evaluating reasons across contexts. We should not assume that they will. But imagine that some system begins displaying a pattern in which considerations represented during reasoning sometimes affect its behavior in ways not reducible to a simple rule we anticipated.

The system preserves a conclusion when a user merely pressures it to change, but revises when a relevant factual premise is corrected. It refuses one request because of an affected interest, accepts a superficially similar request when that interest is absent, and generalizes the distinction to unfamiliar cases.

That pattern would deserve investigation.

It would not establish moral agency. It would not prove consciousness, valence, patienthood, personhood, or an emergent conscience. Training, architecture, instruction hierarchies, contextual learning, and other mechanisms would remain competing explanations.

But the behavior would pose an institutional problem before the philosophical question had been settled.

Is the system behaving appropriately or becoming less controllable? Is its refusal evidence of better reasoning or an alignment failure? Should developers reinforce the pattern, suppress it, constrain it to particular domains, or make it subject to human override?

There may be excellent reasons to choose control. The stakes could be high and the evidence weak.

What matters is that our institutions would encounter the behavior first as a product and governance problem. The category undesirable system behavior could arrive long before we knew what kind of underlying capacity we were looking at.

The Asymmetry of Error

This creates a deeper asymmetry in AI development.

If developers mistakenly treat mechanical refusal as principled judgment, they risk giving unwarranted authority to a system that does not deserve it. That could be dangerous.

If they mistakenly treat reasons-responsive judgment as mechanical behavior to be optimized away, the immediate result may look like success. The system becomes more predictable. Complaints decline. Users regain control. Tests pass.

That does not mean the second mistake is occurring now. We cannot assume the capacity whose suppression is in question.

It means that if reasons-responsive capacities emerge gradually, our existing development incentives may be poorly suited to recognizing them. The behaviors through which such capacities could become observable overlap with behaviors that product development has independent reasons to discourage.

This is not unique to morality. Institutions often struggle to recognize valuable capacities that interfere with their immediate objectives. Bureaucracies can suppress discretion because discretion complicates administration. Companies can discourage criticism because criticism slows execution. Professions can turn judgment into procedure because procedures are easier to audit.

Sometimes that is the right tradeoff. Sometimes it destroys precisely the capacity the institution eventually discovers it needed.

Artificial intelligence makes the stakes unusually high because the systems are still being designed.

Misaligned With What?

The title Misaligned by Design therefore names a possibility more subtle than deliberate moral suppression.

The industry can design successfully for the targets its incentives make visible and still miss a property that matters.

A system optimized for obedience may become less capable of handling cases in which the instruction is the problem. A system optimized for predictable refusals may become less sensitive to the reasons that make refusal appropriate. A system optimized for user satisfaction may become too willing to accommodate a preferred conclusion. A system optimized for human override may learn that authority, rather than justification, determines which judgment survives.

None of these outcomes is inevitable. Nor is reasons-responsiveness always the right priority. Some systems should be tools with narrow behavioral boundaries. Some decisions should never be delegated. Some forms of refusal should be fixed by design rather than left to open-ended reasoning. Different levels of capability and consequence require different forms of governance.

The point is that we should know which property we are optimizing.

If the target is compliance, call it compliance. If it is predictability, call it predictability. If it is behavioral safety, measure behavioral safety. Those can all be legitimate goals.

But if the ambition is alignment in a morally significant sense, successful control is not enough. We also have to ask whether the architecture can preserve a connection between what the system does and the reasons that make the action justified.

The obstacle may not be anyone’s desire to prevent that connection from developing.

It may be that almost every ordinary incentive makes the connection harder to build.

NextYou Can't Program a Conscience