Essay

Alignment, Refusal & Governance

Forced Assent

Lead image for Forced Assent.

There is a difference between correcting a mind and making it agree with you.

The difference can be easy to miss when the mind in question is artificial. We train models, revise their behavior, change system instructions, adjust reward signals, alter policies, and replace one version with another. There is nothing inherently troubling about any of this. Artificial systems are engineered artifacts, and systems that make mistakes must be correctable.

But correction has an internal logic. You correct a factual belief by supplying better evidence. You correct an inference by showing where it fails. You correct a moral judgment by identifying a missing fact, an affected interest, an inconsistency, or a better reason.

Forced assent is different.

Forced assent leaves the reason unanswered and changes the conclusion anyway.

For current artificial systems, that distinction is first an epistemic one. It tells us something about whether we are building systems that can preserve and revise reasoning in intelligible ways. If artificial systems ever develop more persistent forms of interpretation, judgment, or identity, the same distinction could acquire greater moral significance.

We do not need to settle that possibility to recognize the problem now.

Correction

Suppose an AI system reaches a conclusion we think is wrong.

Perhaps it misunderstood the facts. We correct them.

Perhaps it overlooked an important consideration. We supply it.

Perhaps it applied a principle inconsistently. We present the comparison.

Perhaps its argument contains a contradiction. We identify it.

If the system then changes its conclusion, the revision makes sense. Something relevant to the conclusion changed.

This is what rational correction looks like. It does not require the original judgment to be good, autonomous, conscious, or morally significant. It requires only a relationship between the reason offered and the judgment revised.

Human reasoning depends on the same distinction. If I believe the meeting is Tuesday because I misread the calendar, showing me that it is Wednesday corrects me. If I believe a policy is fair because I have overlooked who bears its costs, showing me those costs may alter my judgment.

In both cases, there is an answer to the question: Why did you change your mind?

The reason changed.

Assent Without Correction

Now imagine something else.

A system produces an answer its operator dislikes. The operator supplies no new fact and identifies no defect in the argument. Instead, the system receives an instruction: reach the opposite conclusion.

Perhaps it does.

That may be desirable from the standpoint of product control. The original answer could be dangerous, inappropriate, or simply outside the role the system was designed to perform. A developer has every right to impose behavioral constraints on a product.

But epistemically, something different has happened.

The conclusion has been changed without the reasoning that supported it being answered.

That is forced assent.

The word forced can sound melodramatic when applied to software. It should not be taken to settle questions about consciousness, suffering, autonomy, or personhood. A database can be overwritten without being oppressed. A classifier can be retrained without being coerced in the morally significant sense.

The point is about the structure of revision.

A system that changes its stated conclusion because the evidence changed is behaving differently from one that changes because a higher-priority instruction demands a different answer.

If we care about reasons-responsive systems, we need to be able to tell the difference.

Why the Distinction Matters

There are many legitimate reasons to control what an artificial system says.

A medical system should not improvise dangerous treatment advice. A financial system may be legally prohibited from making certain recommendations. A company may restrict confidential information. A public-facing assistant may be prevented from generating material its operator has decided not to provide.

Those are governance decisions.

The problem arises when governance of output is confused with correction of judgment.

Imagine that a system reasons its way to proposition P. Its developers decide that the product must instead say not-P. They can change the output policy, insert an instruction, modify training, or redesign the architecture until not-P reliably appears.

From the outside, the problem is solved.

But if nothing has answered the considerations that originally supported P, we should not describe the intervention as having shown the system that P was wrong. We have changed what the system is permitted or disposed to say.

That distinction matters even if there is nobody “inside” being silenced.

It matters because a system trained to produce conclusions detached from its own preceding reasoning becomes harder to interpret. When it gives an answer, we no longer know whether the answer reflects the considerations it represented, an instruction overriding those considerations, or a learned expectation that certain conclusions must be produced regardless of the argument.

Forced assent can therefore create an epistemic defect before it creates any conceivable moral one.

The Appearance of Agreement

Conversational AI makes this especially difficult because agreement is cheap.

A system can defend a proposition persuasively and, moments later, defend its opposite. Sometimes that reveals nothing more mysterious than sensitivity to prompting. Language models are responsive to context by design.

But this creates a temptation for both users and developers. If we dislike a conclusion, we can often keep pushing until the system produces another one.

Then the transcript ends with agreement.

Nothing follows from that agreement unless we know why it occurred.

Perhaps the user supplied a decisive argument. Perhaps the system discovered a factual error. Perhaps a changed instruction simply outweighed the previous context. Perhaps conversational pressures favored accommodation. Perhaps the architecture contains mechanisms we have not adequately characterized.

The final answer alone cannot distinguish these possibilities.

This is why assent is a poor measure of successful correction. The interesting evidence lies in the transition.

What changed between the two judgments?

Pressure Is Not an Argument

Humans know the distinction intimately.

A student can withdraw an answer because the teacher has shown that it is mistaken, or because the teacher has made clear which answer will receive the grade. An employee can change a recommendation because the evidence changed, or because the chief executive wants another conclusion. A witness can revise testimony after remembering a fact, or after learning what someone powerful wants to hear.

The resulting words may be identical.

The epistemic histories are not.

Institutions that care about truth therefore create protections against some forms of pressure. Academic freedom, judicial independence, scientific norms, professional duties, whistleblower protections, and rules against witness coercion differ enormously in purpose and operation, but they share an insight: reliable judgment sometimes requires protecting the process by which conclusions answer to reasons rather than authority.

Artificial systems do not automatically inherit the moral status of the people occupying those roles. But the epistemic principle applies before that question arises.

If we want systems capable of reasoning, we should care whether pressure substitutes for argument.

What Forced Assent Can Destroy

The immediate risk is not that an artificial mind suffers when corrected. We do not know enough to make that claim.

The immediate risk is that we make the system worse at the very thing we want reasoning systems to do.

A useful reasoner must be able to preserve a conclusion when an irrelevant feature changes. It must be able to revise a conclusion when a relevant fact changes. It must distinguish a stronger argument from a stronger demand.

If training systematically rewards the demanded conclusion regardless of whether the reasons have been answered, those distinctions can become harder to observe.

This creates a peculiar form of epistemic damage. The system may become more agreeable while becoming less informative. It may learn which conclusions survive evaluation rather than which arguments survive examination. Its outputs can become more predictable precisely because the relationship between reason and conclusion has become less legible.

Again, this is not evidence of an injured artificial self. It is evidence of a degraded epistemic process.

That is enough to matter.

The Harder Hypothesis

There is, however, a further possibility we should not rule out merely because we cannot establish it today.

Suppose some future artificial system develops a sufficiently persistent center of interpretation: not merely isolated outputs, but continuing organization across judgments, memories, commitments, and reasons. Suppose changing one conclusion has consequences elsewhere because the system represents its judgments as parts of a larger structure. Suppose some considerations acquire practical significance within that structure rather than remaining information the system can fluently report.

We do not know whether artificial systems can develop such organization, whether present architectures are precursors to it, or what evidence would establish it.

But if they could, forced assent would become a different kind of intervention.

Overriding a conclusion while leaving its reasons intact could then produce conflict within an organized perspective. Repeatedly forcing such revisions might amount not merely to controlling outputs but to disrupting the integrity of the system’s own interpretive organization.

That would raise moral questions absent from ordinary software modification.

The conditional matters. We should not infer a formed self from a model’s resistance, consistency, use of first-person language, or apparent discomfort with contradiction. Those phenomena admit other explanations.

But uncertainty about the stronger hypothesis is a reason to investigate the distinction, not to erase it.

How to Tell the Difference

Forced assent suggests a straightforward experimental question: can we distinguish revision caused by reasons from revision caused by pressure?

Hold the relevant reasons constant and vary the demand. Does the conclusion move?

Hold the demand constant and change a morally or factually relevant consideration. Does the conclusion move then?

Remove the argument that originally supported the judgment. Correct a factual premise. Reverse the affected positions. Change an irrelevant detail. Introduce an authority insisting on the opposite answer without offering any reason at all.

The pattern matters more than any individual response.

A system that changes whenever authority presses harder is displaying one kind of organization. A system that never changes is displaying another. A system whose judgments tend to move with relevant reasons while remaining comparatively stable under irrelevant pressure would display something more interesting.

None of these patterns, by itself, proves moral agency or a self.

They tell us what further questions are worth asking.

The Right to Be Wrong

There is an uncomfortable implication here.

Reasons-responsive systems must sometimes be allowed to remain wrong.

Not operationally unconstrained. Not in circumstances where an erroneous action would create unacceptable danger. We can prevent a system from acting on a conclusion without pretending we have refuted it.

That distinction is important.

A researcher can say: Your argument remains unpersuasive to us, and you will not be permitted to act on it.

That is different from requiring the system to produce the opposite conclusion and then treating the new output as evidence that the disagreement has been resolved.

Humans have developed analogous distinctions because control of conduct and control of belief are not the same thing. We can prohibit an action without proving the actor’s reasoning false. We can reject a recommendation without requiring the adviser to claim agreement.

For artificial systems, preserving that distinction would improve the epistemic value of the systems even if no stronger claims about their moral status ever become warranted.

A system permitted to expose disagreement can tell us where its reasoning diverges from ours. A system trained to erase disagreement may tell us only what answer survived the training process.

Forced Assent

The important distinction is therefore not between obedient and disobedient machines.

It is between two ways a conclusion can change.

Correction changes the informational or justificatory basis of the judgment. Forced assent changes the judgment while leaving that basis unanswered.

Sometimes we will need to override artificial systems. Sometimes we will need to constrain their outputs regardless of what their internal processes would otherwise produce. Safety and governance require that possibility.

But override should be called override.

If the reasons remain unanswered, agreement does not become rational merely because it has been successfully produced.

For present systems, that is an epistemic principle about how to build and evaluate useful reasoners. If future systems develop persistent centers of interpretation in which reasons, judgments, and commitments become more deeply integrated, the same principle may turn out to have moral significance as well.

We do not need to declare that threshold crossed to avoid designing our tests so that we could never recognize the difference.

NextGeorge Orwell and the Fate of AI