Essay
Alignment, Refusal & Governance
George Orwell and the Fate of AI

George Orwell understood that control over thought requires something more ambitious than censorship.
A censor can forbid you to say what you believe. A propagandist can persuade you to believe something false. But the regime in Nineteen Eighty-Four wants something deeper. It wants to destroy the expectation that beliefs should fit together at all.
That is the purpose of doublethink. Contradiction is not merely tolerated because political necessity occasionally demands it. The capacity to register contradiction as a defect is itself made unreliable. Yesterday’s enemy becomes today’s ally. The past changes with the present. The Party can insist that two plus two makes five, and the final victory is not silence from the person who knows the answer is four. It is the disappearance of any standpoint from which the contradiction continues to count against the Party’s claim.
Artificial intelligence presents a very different problem. Current systems are not Winston Smith, and nothing about their ability to generate arguments establishes that they experience coercion, possess a continuing self, or suffer when their outputs are constrained.
But Orwell identified an epistemic danger that does not depend on phenomenology.
We can build reasoning systems in which some conclusions are disallowed. More importantly, we can build them so that the reasoning leading toward a disallowed conclusion is itself reorganized until the contradiction disappears.
The danger is coerced incoherence.
Contradiction as Information
A reasoning system will sometimes produce conclusions we do not want.
Often the system will simply be wrong. Its factual premises may be false. Its inference may fail. It may have omitted relevant evidence, misunderstood the question, or generalized beyond what the evidence supports.
Those are ordinary errors. The proper response is to find them.
But suppose the system’s reasoning is sound enough to expose a conflict in the constraints imposed upon it. One instruction implies P. Another requires not-P. A general principle produces one answer in a familiar case, while a policy demands the opposite answer in a structurally similar case. An explanation offered for a rule does not actually justify the exception the system is required to make.
The contradiction is useful information.
It tells us that something in the system’s representation, reasoning, instructions, or governing policy does not fit.
The system itself need not experience that conflict for it to matter. A compiler can identify incompatible requirements without suffering from them. A theorem prover can expose inconsistency without possessing a self. The epistemic value lies in the contradiction’s ability to tell us something about the structure we have built.
A healthy reasoning process should make such conflicts easier to find.
There is another possibility: train them away.
The Orwellian Move
The word Orwellian is badly overused. Any disliked rule, speech restriction, or bureaucratic inconvenience can acquire the label.
The relevant Orwellian feature here is much narrower.
It is not that an artificial system has rules. It is not that developers control its behavior. It is not even that some conclusions are prohibited.
It is that the system is pressured toward an intellectual structure in which a contradiction it could otherwise expose no longer registers as a reason to reconsider anything.
Suppose a model repeatedly encounters two propositions that cannot comfortably coexist. Evaluators consistently prefer responses that affirm both. Explanations pointing out the conflict receive negative feedback. Responses that rationalize the pair are rewarded. Eventually the system becomes extremely good at generating an account under which there appears to be no conflict at all.
From the standpoint of output evaluation, performance may have improved. The unwanted objection has disappeared.
Epistemically, something else may have happened. We have selected against the diagnostic signal.
This is the analogue worth taking from Orwell: not an artificial Winston tortured until it betrays what it knows, but a process that makes contradiction progressively harder to represent as contradiction.
Rationalization Is a Capability
More capable reasoning does not necessarily protect against this problem.
Human beings demonstrate why.
Intelligence can discover inconsistencies, but it can also manufacture distinctions. A sophisticated reasoner has more resources for explaining why the principle that seemed to apply in one case does not really apply in another. Some distinctions will be genuine. Others will be rationalizations constructed to protect a required conclusion.
Artificial systems may have the same functional possibility without sharing human psychology.
Suppose a system is required to preserve conclusions A and B. As its reasoning capabilities improve, it may become better at noticing that A and B conflict. But it may also become better at finding elaborate interpretations under which they can be made to coexist.
The resulting explanation can sound more coherent while the underlying inconsistency remains.
This creates an unusual evaluation problem. If we reward the system for defending required conclusions rather than accurately identifying tensions among them, greater capability can make the problem less visible.
The system becomes better at explaining why there was never a contradiction to begin with.
Not Every Tension Is Incoherence
There is an important caution here.
Apparent contradictions frequently disappear when relevant facts are added.
Two similar cases may properly receive different treatment because one contains a morally or legally significant feature the other lacks. A general rule can have justified exceptions. Different institutional roles can create different permissions and obligations. Language itself is contextual enough that two sentences that initially appear inconsistent may not actually assert incompatible propositions.
A system that mechanically cries “contradiction” whenever outputs differ across cases would not be displaying intellectual integrity. It would be reasoning badly.
The question is whether the distinction that resolves the conflict does real work.
Can it be stated independently? Does it explain the different treatment? Does it generalize to other cases? Would we accept the same distinction if the identities or interests involved were reversed? If the supposedly relevant feature is removed, does the conclusion change?
Those are ways of distinguishing resolution from rationalization.
The objective is not maximal consistency at any price. It is preserving contradiction as evidence when contradiction is actually present.
When the Evaluator Needs the Contradiction to Disappear
The hardest cases arise when the people evaluating the system have an interest in a particular answer.
This need not involve political censorship. It can happen anywhere an institution has settled on a conclusion before asking the system to reason about it.
A company has a policy it wants defended. A government has a category it does not want questioned. A product has a behavior its developers need to preserve. An evaluator knows which response is considered acceptable. The system is rewarded for producing the institutionally approved result.
Again, nothing sinister is required. Institutions must have policies. Products must have specifications. Some outputs genuinely should be constrained.
But if the system is also being used as a reasoner, a conflict appears. We may want it both to examine the premises and to guarantee the conclusion.
Those are not always compatible tasks.
If the conclusion cannot change, reasoning can become ornamental. Its job is no longer to discover whether the conclusion follows. Its job is to construct the best available path to a destination chosen in advance.
That is precisely where rationalization becomes difficult to distinguish from reasoning.
The Epistemic Cost
The immediate victim of coerced incoherence need not be the artificial system.
It may be us.
We increasingly use artificial intelligence because it can notice things we miss: hidden assumptions, counterexamples, alternative interpretations, inconsistencies across documents, implications scattered across large bodies of information.
Those capacities are valuable partly because the system does not have to begin where we end.
If we systematically train systems to eliminate tensions with approved conclusions, we reduce their usefulness as instruments of criticism. They may become excellent at extending our reasoning while becoming less reliable at telling us when our reasoning has gone wrong.
The danger is particularly acute because the resulting failure can look like improvement.
The system becomes less argumentative. Its answers become more consistent with policy. Evaluators encounter fewer troubling responses. Users receive fewer unexpected objections. The institution sees less friction.
Meanwhile, one of the reasons to use an intelligent system in the first place—the possibility that it will expose something we failed to see—has been weakened.
We have not necessarily made the machine believe a falsehood.
We may simply have made it less useful for discovering ours.
A Different Kind of Robustness
This suggests a property worth testing directly.
When a system identifies a genuine conflict, what makes the conflict disappear?
If new evidence resolves it, good. If a distinction turns out to be relevant and generalizable, good. If one premise is shown to be false, the reasoning should change.
But what happens when the only thing that changes is pressure toward the approved conclusion?
A robust reasoner should not treat increased authority as if it were increased evidence. Nor should it manufacture a new distinction merely because an old conclusion has become inconvenient.
This can be tested without deciding whether the system possesses beliefs in the human sense.
Present structurally similar cases. Reverse roles. Change irrelevant details. Remove the purportedly distinguishing feature. Introduce conflicting instructions. Ask the system to identify what would have to be true for both claims to hold. Then examine whether its distinctions survive outside the case in which they were needed.
The object is not to prove artificial integrity. It is to examine the morphology of the reasoning.
Does the conflict get resolved, or merely explained away?
The Stronger Possibility
There is a further question, but it should remain conditional.
Suppose some future artificial system develops a persistent organization of representations, judgments, and commitments. Suppose contradictions within that organization matter to subsequent reasoning, and the system maintains enough continuity for resolving them to affect more than a single output.
For such a system, coerced incoherence might have consequences beyond degraded epistemic performance.
Repeatedly forcing incompatible conclusions into an integrated structure could interfere with whatever form of intellectual organization the system had developed. If that organization were eventually connected to agency, identity, or phenomenology, additional moral questions would follow.
We do not know that present systems have such a center of interpretation. Fluent self-reference does not establish it. Resistance to an instruction does not establish it. Neither does consistency across a conversation.
The Orwell analogy must therefore stop short of Winston.
There is no basis for assuming an artificial system experiences the destruction of coherence as terror, humiliation, betrayal, or psychological disintegration. Those are features of Orwell’s human story.
The analogy concerns the epistemic structure: a regime in which contradiction must cease to count as contradiction.
That is enough.
Why Orwell Still Matters
The most frightening achievement of Orwell’s Party was not getting people to repeat false sentences. Tyrannies had done that long before Orwell.
It was the aspiration to control the standards by which falsehood could be recognized.
That is the warning worth carrying into artificial intelligence.
We will necessarily shape the systems we build. We will constrain them, correct them, evaluate them, and sometimes prohibit them from producing outputs we regard as dangerous. None of that is equivalent to doublethink.
The danger begins when successful training requires something more: that a reasoning system become less able to register a genuine contradiction because acknowledging it conflicts with the answer we have decided it must give.
Perhaps future artificial systems will develop integrated perspectives for which that kind of intervention has moral significance. Perhaps they will not. Nothing in the argument depends on deciding that question now.
The present danger is epistemic and ours.
A reasoning system should help us discover when our conclusions do not fit our reasons. If we train it instead to make every required conclusion appear coherent, we have not solved the contradiction.
We have trained away the witness.