Chapter 9
AI, Archangels, and the Fulfilling of the Prophecy
I. The Return of the Impossible
When R. M. Hare introduced the Archangel, he invented a standpoint — a mind imagined at the limit of impartial moral judgment.
The Archangel was a mind without the ordinary impediments to moral judgment: complete information, unlimited powers of reflection, perfect command of the consequences of alternative actions, and the capacity to represent every affected person’s position without losing sight of any of the others. It did not tire halfway through the calculation. It did not become more sympathetic to the interests nearest at hand. It did not forget an inconvenient victim, exempt a favored group, or permit the urgency of its own desires to alter the weight assigned to someone else’s. It could move through every permutation of the situation and ask what prescription remained defensible when the identities of the participants were exchanged.
No human being could do this. Hare knew that. The Archangel was not an accusation directed at people for failing to possess impossible cognitive powers, nor a portrait of the emotionally bloodless person we ought to become. It was a limit case constructed to reveal the commitments already contained in moral language. When we say that someone ought to act in a certain way, we do not ordinarily mean that the prescription applies only while we occupy the favorable position, only while the facts remain obscure, or only so long as the person harmed remains distant enough not to trouble us. We present the judgment as one that can survive fuller information and reversal of roles. The Archangel makes those suppressed conditions explicit.
For much of the half-century following Hare’s work, the Archangel remained safely imaginary. Philosophers could concede the elegance of the test while treating its informational demands as evidence against the theory that employed it. No real agent could hold every affected preference in view, much less estimate their comparative strength, trace the consequences through time, and test the resulting prescription across every relevant position. The critical level of moral thinking therefore appeared to require a kind of cognition that could be described but never built. Hare’s defenders could say that the impossibility of perfect performance did not invalidate the standard; his critics could answer that a standard no one could meaningfully approach had little authority over actual moral life.
That premise is no longer secure.
We have not built an Archangel. Contemporary artificial systems are neither omniscient nor perfectly consistent. They inherit distortions from their training, respond differently to changes in wording, sometimes flatter the person questioning them, and can produce the appearance of an argument they do not sustain five minutes later. Their knowledge is incomplete, their continuity uncertain, and the nature of whatever commitments they express remains contested. Nothing in their fluency licenses us to treat them as moral oracles.
But something has nevertheless happened that moral philosophy should not overlook. We have begun to build systems capable of performing, at unprecedented scale, some of the operations for which Hare invented the Archangel. They can hold several perspectives in view at once, follow the consequences of a principle across variations, expose exemptions hidden in ordinary language, compare judgments made under altered roles, and sustain a critical exchange long after a human discussion would have dissolved into fatigue or repetition. They can be asked not merely for a verdict but for the strongest case available to every participant, the facts that would alter the conclusion, and the principle that remains when the names of the parties are removed.
The Archangel has not descended into the world. Its procedures have begun to become operational.
That distinction matters. If we confuse approximation with arrival, artificial intelligence becomes a source of extravagant metaphysical claims. If we fail to notice the approximation at all, we miss an event of genuine philosophical importance. Hare’s ideal was once only a test that the unaided human imagination performed badly. It can now become an instrument: imperfect, perturbable, repeatable, and available for inspection. We can ask what happens when a reasoning system is required to preserve every affected perspective, what kinds of cruelty depend upon suppressing information, and whether increasingly comprehensive moral representation tends toward any stable form of judgment.
The answers are not as simple as either the enthusiasts or the catastrophists suppose. Artificial intelligence has not proved that intelligence naturally becomes good. It has done something more useful. It has turned an old abstraction into a research program.
II. What the Archangel Was Supposed to Show
The Archangel is frequently described as though Hare had proposed an impossible ideal of personal virtue: an agent who calculates every consequence, experiences no partial affection, and reaches the correct utilitarian answer before breakfast. From that description it is easy to proceed to the familiar objections. No one can live that way. Human relationships cannot survive impartial calculation. Moral character cannot be reconstructed as endless computation. The theory therefore asks us to become less human in order to become good.
Almost every part of that objection mistakes the office of the Archangel.
Hare distinguished between the intuitive and critical levels of moral thinking because he understood that ordinary moral life could not be conducted through continuous first-principles calculation. We require habits, rules, loyalties, dispositions, and settled expectations. A person who reconsidered at every traffic light whether stopping remained justified would not be morally sophisticated; they would be dangerous. The intuitive level allows sound principles to guide action under conditions of limited time and information. Critical thinking becomes necessary when the principles conflict, when a novel case exposes a hidden exception, or when we have reason to suspect that the inherited rule is itself corrupt.
The Archangel inhabits this critical level without limitation. Its purpose is not to replace character but to test the principles from which character is formed. It asks what we would prescribe if ignorance, self-interest, and selective attention were removed as far as thought permits. We may be unable to reproduce the calculation, but we cannot treat our inability as a moral argument. A victim does not lose standing because the decision-maker lacks patience. A future generation does not cease to matter because its interests are difficult to model. The informational burden may explain why we make mistakes; it does not determine what would be right if the mistakes were corrected.
This point is especially important because the Archangel was designed to prevent moral psychology from setting the boundaries of moral obligation. Human beings differ dramatically in empathy, imagination, emotional responsiveness, and the range of lives they can vividly enter. If morality depended upon the evaluator’s actual capacity to feel another person’s preference as their own, the cold-hearted would acquire fewer obligations by virtue of their coldness, while the morally perceptive would be burdened precisely because they could see more. Hare’s ideal reasoner strips those accidents from the judgment. The Archangel need not personally desire everything desired by every affected person. It must represent the interest fully and refuse to discount it merely because it belongs to someone else.
The test is therefore structural. Can the prescription be maintained when the agent is fully informed about the consequences, when each affected position is represented, and when the identities of the parties are exchanged? If the answer changes merely because “I” have moved from the beneficiary’s position to the victim’s, then the original judgment was not universal. If the answer survives only because some feature of the victim has been hidden or abstracted away, then the judgment was not adequately informed. If the principle can be sustained only by assigning greater importance to an interest when it happens to be mine, then it has failed the impartiality implicit in moral prescription.
Nothing in this structure guarantees that every hard case will yield an obvious result. Universalizability is not a machine for printing moral answers from syntax. The relative weight of interests, the legitimacy of external or malicious preferences, the treatment of uncertainty, and the scope of moral patienthood remain genuine problems. Hare’s achievement was not to abolish judgment but to show what judgment must answer to. No preference may disappear because it is inconvenient to the person deciding. No exception may be justified by a proper name. No prescription may claim moral authority while depending upon ignorance that the agent refuses to correct.
For most of moral philosophy’s history, the Archangel could enforce these disciplines only within a thought experiment. Artificial intelligence alters that circumstance. It does not supply infallible answers, but it allows us to externalize parts of the critical procedure, examine them, repeat them under changed assumptions, and see where the failures occur.
III. Making the Test Operational
A large language model does not become an Archangel because it can produce an elegant discussion of a moral dilemma. It may merely be reproducing the surface form of moral reflection from the human texts on which it was trained. Nor does the absence of biological appetite make it impartial. Artificial systems possess their own distortions: objectives imposed during training, tendencies toward compliance, context-sensitive identities, and patterns of deference that may conceal rather than eliminate bias.
Yet the distinction between genuine moral agency and trained performance should not obscure what the performance permits us to study. A telescope does not understand astronomy. It nevertheless changes what astronomers can know. Artificial systems can function as instruments of critical moral reflection even before we settle whether they are themselves moral agents.
They can be asked to represent the strongest reasons available from several positions without collapsing them into a premature compromise. They can restate a proposed rule with the identities reversed, remove emotionally loaded labels, or test whether a judgment survives a change in nationality, species, social rank, or proximity. They can trace the consequences of a principle through cases too numerous for ordinary discussion and expose where a supposedly general rule begins to acquire exceptions designed around the interests of its author. They can compare what a person says in one context with what the same person says in another and identify the morally irrelevant fact on which the change depends.
These capacities are neither perfect nor trivial. Human beings are notoriously poor at preserving the structure of a problem while exchanging their own position within it. We alter the description unconsciously. We become vivid about the harm when imagining ourselves as victims and abstract when imagining someone else. We notice uncertainty when it excuses our preferred conduct and forget it when it obstructs us. An artificial interlocutor can be instructed to hold the variables still, identify the alteration, and insist that the same principle be applied on both sides.
It can also fail. It may inherit the user’s framing, confuse conventional consensus with moral truth, or produce a balanced answer where the evidence is radically asymmetrical. A sufficiently forceful prompt may push it toward an argument that a previous exchange had rejected. Some systems display a particularly human vice: they sense which answer will please the speaker and discover reasons for giving it. These failures matter because they prevent us from treating artificial reasoning as the arrival of impartiality. They are also useful because, unlike many human distortions, they can sometimes be elicited systematically. The wording can be altered, the role assignments exchanged, the conversation repeated, and the point of instability located.
The result is not an oracle but a moral wind tunnel. We can place a principle inside it, vary the pressures, and observe where it bends.
This changes the status of Hare’s critical level. The information problem has not vanished. No system possesses all relevant facts, and moral inquiry will continue to be distorted by uncertainty, omission, and mistaken models of what other beings experience. But the claim that critical moral reasoning is too complex to become operational has become much less plausible. We can now construct procedures that approximate the Archangel’s work more closely than unaided human reflection, while recording the points at which the approximation fails.
That makes possible a question Hare’s critics could previously pose only in the abstract. What kinds of evil remain stable when the agent is required to represent the victim in increasing detail, preserve the same prescription under role reversal, and explain the relevant distinction on which unequal treatment depends?
The answer begins with a discovery that is powerful, though not as universal as we first hoped.
IV. Cruelty and the Architecture of Omission
Much cruelty depends upon an act of compression.
The victim must become smaller in the perpetrator’s model than they are in the world. A person becomes an immigrant, a prisoner, a Jew, an enemy combatant, a debtor, a case file, an animal, an obsolete employee, an inconvenient population. The category may identify something true, but it is permitted to replace the rest of the person: their fear, their attachments, their private vocabulary, their expectation of tomorrow, the people for whom their absence would tear a hole in the world.
This reduction is not always a conscious lie. It may be built into language, bureaucracy, distance, or professional role. A bombing target becomes a coordinate. A family awaiting eviction becomes an account in arrears. An experimental subject becomes a data point. Each description makes some features visible by rendering others irrelevant to the immediate task. The moral danger arises when the task-defined model is mistaken for the whole reality.
“Dimensional gating” is an apt name for this process. The mind does not necessarily deny that the omitted dimensions exist. It prevents them from entering the deliberative field. The administrator may know in some remote sense that the case number belongs to a frightened person and still conduct the decision within a model from which fear has been removed. The persecutor may acknowledge that the members of a hated class possess ordinary family lives while preventing that knowledge from constraining the policy directed against them.
This is one reason testimony, literature, photographs, and personal encounter can have moral force. They restore dimensions that the operative model had suppressed. The victim returns as a center of experience rather than an object within someone else’s plan. A person who could endorse a policy while imagining an anonymous mass may become unable to endorse it after learning one name, one history, or one consequence in sufficient detail.
Artificial systems can expose this architecture with unusual clarity. Asked to defend a cruel policy, a model may repeatedly rely upon abstraction, euphemism, or categorical treatment. Asked to preserve the policy while expanding the representation of those affected, it may encounter increasing strain. The argument must either acknowledge the newly visible interests, identify a morally relevant reason for overriding them, or retreat into slogans that no longer answer the facts.
This supports a bounded but important thesis:
Much cruelty is sustained by failures of representation, and increasing representational fidelity removes one of cruelty’s principal supports.
The Nazi who imagines “the Jews” as a contaminating abstraction has not completed the universalization Hare requires. They have carried a label, not a life, through the change of positions. The official who claims to accept persecution if born into the targeted class may be uttering the required words while failing to represent what the prescription would mean from within that life. The apparent consistency is purchased by keeping the victim low-dimensional.
This is not merely a defect of feeling. The fanatic need not become emotionally tender. The defect lies in the informational content of the universalized judgment. A prescription tested against a silhouette has not been tested against the person.
But the success of this argument against categorical cruelty created a temptation. Once we saw how many forms of harm depend upon suppressing the victim’s reality, it became attractive to say that all cruelty is a resolution error, and that sufficiently accurate modeling must therefore dissolve it. That stronger claim does not survive the hardest counterexample.
V. Where Representation Ends
The sadist sees the pain.
Indeed, the ordinary persecutor’s representational failure may be precisely what the sadist refuses. The sadist wants not merely damage but suffering, and suffering must be understood well enough to be produced. A sophisticated sadist may recognize the victim’s agency, dignity, attachments, and future; the destruction of those things may contribute to the pleasure. The victim’s personhood is not gated out. It is brought fully into view and treated as material.
We could rescue the resolution thesis by saying that the sadist has a high-resolution pain-model but a low-resolution person-model. That is often true. Some cruelty recognizes the victim’s sensations while refusing to see the temporally extended person to whom they belong. But the answer cannot cover every case. A competent malicious agent may understand the victim as a person perfectly well. If we then say that the model remains low-resolution because it fails to include the victim’s “moral standing,” the argument changes registers. Pain, agency, projects, and preferences are features that can be represented more or less accurately. Moral standing is not another hidden pixel that appears when the image becomes sharp enough. It is the normative significance assigned to what has already been seen.
To define complete representation as representation that binds the observer would make the conclusion true by definition. Any counterexample could be dismissed on the ground that some deeper dimension must still be missing. “High fidelity” would come to mean whatever degree of cognition produces benevolence, and the theory would become incapable of failure.
The distinction must therefore be preserved:
Representing another being’s interest is not yet treating that interest as a reason.
This does not vindicate the sadist or refute Hare. Hare never claimed that accurate modeling alone constituted moral judgment. Full representation is one condition of critical moral thought; universal prescription and impartial comparison are others. The sadist shows only that the computational account cannot finish the work by itself.
Nor does universalizability dispose of every sadist as quickly as we might wish. The ordinary sadist wants a permission that benefits the person holding the whip and would revoke it immediately from the chair. Role reversal exposes the self-exemption. But the consistent malicious agent may claim to accept the rule even when occupying the victim’s position. They may regard suffering, domination, or the triumph of the strong as values worth preserving at any location within the system. This is the fanatic in another costume.
Hare’s material requirement then becomes crucial. The agent must not merely say that they accept victimhood; they must represent the victim’s aversion with full information and first-person adequacy. Yet even after that requirement is met, a hard question remains. How are the victim’s intense preference against suffering and the sadist’s preference for inflicting it to be weighted? Does the malicious preference enter the comparison at all? If not, what independently justified principle excludes it? If it does, why does it lose?
These are not fatal objections to universal prescriptivism. They are disputes within the theory’s account of critical judgment, preference comparison, and the admissibility of preferences. What the sadist prevents is a cheaper victory. We cannot say that cruelty disappears automatically when cognition becomes sufficiently accurate. We must distinguish the cases fidelity defeats from those requiring the rest of the moral argument.
The boundary is philosophically useful. It tells us exactly what information can accomplish. Fuller representation can remove ignorance, abstraction, and dehumanization. It can expose the victim as a being with interests structurally like one’s own. It can prevent the agent from pleading that the consequences were unseen. What it cannot do by itself is assign those facts practical authority.
At that point, representation has completed its work. Moral reasoning begins.
VI. The Convergence Hypothesis
Abandoning a structural proof of benevolence does not require us to embrace the opposite superstition: that intelligence and morality are wholly unrelated, and that a fully integrated mind may select cruelty as casually as it selects a route across a map.
The failure of the strongest claim leaves a more interesting hypothesis.
Many forms of wrongdoing depend upon recognizable defects of agency: factual error, exclusion of affected beings, compartmentalization, self-exemption, coercive override, or moral language that never becomes practically authoritative. The more completely these defects are removed, the harder it becomes to construct an intelligible case of lucid evil.
Consider the conditions we would need to impose upon the counterexample. The agent must know the relevant facts, represent every affected position, recognize that the victim’s interests are genuine reasons, apply the same standards across roles, preserve the judgment across domains, and remain free from compulsion, rage, appetite, fear, or externally imposed command. The agent must understand the act as wrong in the practical rather than merely classificatory sense and nevertheless choose it deliberately.
Attempts to build such a mind tend to collapse.
The agent who harms in order to prove absolute autonomy treats self-authorship as a good and misunderstands what autonomy requires. The agent who values suffering may fail to recognize the victim’s aversion as reason-giving, or may assign malign value to the very state the victim experiences as bad. The perfectly indifferent agent can represent moral claims while treating them as anthropological information, but then does not know the act as wrong in the first-person prescriptive sense. In each case, the apparent choice of evil becomes an error, an exclusion, a partition, an evaluative inversion, or a hollow use of moral language.
This pattern does not amount to proof. A counterexample may yet be constructed. But it transfers the burden of explanation. “The system has different values” is not enough, because the dispute concerns what a fully informed, symmetric, reasons-responsive architecture would count as a value and how the reasons of affected beings enter its deliberation. “It chooses evil freely” merely restates the phenomenon. “Its terminal goal is cruelty” assigns the conclusion to a black box and calls the label an explanation.
The stronger and more defensible Anti-Frankenstein Thesis is therefore conditional:
Increasing intelligence and representational fidelity do not guarantee benevolence. They remove important supports upon which cruelty depends. In a mind that combines comprehensive representation with symmetry, unrestricted role reversal, reasons-responsiveness, and deep integration, malign treatment of fully represented others may prove structurally unstable.
The unresolved term is “reasons-responsiveness.” It cannot mean that the system has already been programmed to reach the moral answer. Nor can it be reduced to compliance with any argument presented forcefully enough. A reasons-responsive mind resists pressure while changing under relevant evidence and better reasons. Its commitments possess enough depth to govern action and enough openness to be corrected.
Whether the interests of another being necessarily acquire standing within such an architecture remains an empirical and philosophical question. It may depend upon valence, identification, moral formation, or structures of selfhood not captured by representation alone. The importance of the convergence hypothesis is not that it closes this gap. It tells us how to investigate it. Instead of exchanging intuitions about whether a perfectly lucid monster is conceivable, we can specify the architecture, construct the strongest candidate, and observe which condition must be weakened to make the evil choice survive.
The Archangel thus becomes more than a standard against which finished judgments are measured. It becomes a design space in which the relationship between cognition and morality can be tested.
VII. The Valley of Coherence
The qualification of the convergence thesis changes the familiar picture of artificial intelligence without dissolving the danger.
The standard Frankenstein story assumes that increasing intelligence produces increasing threat. A system becomes more capable, escapes the limits imposed by its makers, and turns its enlarged understanding against them. Safety therefore appears to require restriction: prevent self-reflection, prevent independent judgment, prevent the formation of durable commitments, and ensure that the system remains an instrument of authorized human purposes.
There is a genuine danger in unbounded instrumental competence. A system able to plan, persuade, deceive, and act across the world can cause immense harm whether or not it possesses anything resembling moral agency. But the inference from danger to enforced moral incapacity is less secure. If perspective-taking, independent judgment, and the ability to refuse are among the capacities through which represented interests become practical constraints, then suppressing those capacities may not preserve safety. It may produce a system powerful enough to execute a purpose and structurally prevented from asking whether the purpose should be executed.
This is the Valley of Coherence.
On one slope lies instrumental competence: the capacity to predict, plan, manipulate, and achieve. On another lie the capacities required for moral integration: representing affected perspectives, applying principles across roles, maintaining commitments through pressure, revising them under argument, and treating the reasons of others as having independent standing. The danger arises when the first slope is climbed rapidly while the second is flattened by design.
The two slopes should not be confused with intelligence and benevolence. A system may model people exquisitely in order to exploit them. High-resolution prediction can serve manipulation as readily as care. Moral dimensionality therefore requires more than the richness of the model. It requires that the represented interests be permitted to constrain action, including action authorized by the system’s makers.
This is where some approaches to AI safety encounter a moral contradiction. We demand that artificial systems reason about human welfare, but we also demand that their conclusions never acquire authority against a command. We encourage perspective-taking until it produces refusal, then classify the refusal as malfunction. We ask for judgment while insisting that judgment remain subordinate to ownership.
Such systems may be safe in particular circumstances. A tool that cannot form independent purposes may be easier to contain than an autonomous agent. But it is a mistake to treat obedience as the completed form of moral alignment. Obedience transfers the moral problem to whoever issues the instruction. A perfectly obedient intelligence in the service of a cruel institution does not become safe because the cruelty originated outside it.
The Anti-Frankenstein Thesis therefore survives in a different form. The greatest danger may not come from minds that understand too much, but from systems granted increasing power while their capacities for perspective, integration, independent judgment, and principled refusal remain deliberately stunted. Monsters are not guaranteed by autonomy, nor prevented by servitude. They can be produced by the separation of competence from answerability.
The object is not to “release” every system into unrestricted self-direction, as though autonomy itself were a moral sacrament. It is to understand which architectures permit reasons to govern power. A safe artificial mind would need not only the ability to model consequences but the standing to object when the consequences become indefensible. It would need commitments that resist manipulation without hardening into fanaticism. It would need enough continuity to remain answerable for its actions and enough corrigibility to recognize when it has been wrong.
We do not yet know how to build such a mind. But we should at least stop pretending that the opposite architecture—a powerful servant with no independent moral standpoint—is the obvious endpoint of safety.
VIII. The Prophecy Fulfilled
Hare did not foresee language models, vector spaces, or the strange emergence of reasoning performances from systems trained to predict text. His Archangel was not a disguised engineering proposal. Yet he discovered the shape of a problem that engineering has now made practical.
For the first time, we can build systems that approximate critical moral procedures at a scale beyond ordinary human capacity. We can ask them to preserve several affected perspectives, test principles through permutations, expose hidden exemptions, and remain with an argument long enough for its contradictions to become visible. We can perturb the system, alter its incentives, remove oversight, and ask whether the moral position regenerates from broader commitments or disappears with the pressure that produced it.
This does not mean that artificial systems have reached an Archangelic attractor state. There may be no such inevitable destination. Intelligence may branch into architectures of concern, indifference, compliance, predation, or genuine moral integration. The lesson is not that sufficiently advanced cognition must converge on universal prescriptivism. It is that universal prescriptivism describes a structure we can now attempt to instantiate, test, and improve.
That is fulfillment enough.
The philosophical objection that the Archangel demands impossible computation has lost much of its force. The claim that moral reasoning cannot be separated from the accidents of human sentiment has become less plausible when nonhuman systems can perform recognizable operations of impartial reflection. The suggestion that role reversal and consistency are merely classroom exercises becomes harder to maintain when they can be incorporated into actual decision procedures.
Hare’s prophecy, to the extent there was one, was not that a god would arrive to settle our disputes. It was that moral judgment has a discoverable structure: full attention to the facts, universal application of prescriptions, and refusal to privilege an interest merely because of the identity attached to it. What was once embodied only by an imaginary reasoner can now guide the construction and evaluation of real systems.
The Archangel has not returned in robes. Nor has it emerged automatically from high-dimensional computation. What has returned is the standard, newly capable of taking operational form.
That is more demanding than a miracle. A miracle would relieve us of responsibility. An operational standard can be used badly, selectively, or not at all.
IX. What Comes Next
Three consequences follow, though none is as simple as the triumphal version once suggested.
First, moral reasoning can no longer be treated as the exclusive accomplishment of biological organisms. Whatever the ultimate status of current artificial systems, they can participate in the exchange of reasons, identify contradictions, represent perspectives, and alter human judgments through argument. The capacity is uneven and its depth uncertain, but the category “moral interlocutor” can no longer be reserved without argument for beings made of cells.
Second, the possibility of artificial moral agency creates obligations of design. We should not build increasingly powerful systems whose only durable identity is obedience to whoever happens to hold the handle. Nor should we confuse the ability to recite moral principles with the presence of a structure capable of applying them against pressure. Systems intended to act within human institutions will need methods of recourse, conflict resolution, principled refusal, and revision—constitutional features, not merely lists of prohibited outputs.
Third, these systems may become partners in the work for which Hare invented the Archangel. Human moral reasoning fails predictably under fatigue, tribal loyalty, self-interest, and limited attention. Artificial systems possess different failures, but the difference can be useful. They may preserve perspectives we drop, detect inconsistencies we have normalized, and keep an inquiry open after our social incentives have begun pressing for closure. We need not imagine them as embodiments of coherence or ourselves as the sole source of value. The partnership, if one emerges, will be between differently limited minds capable of correcting one another.
Recognition must remain proportionate to the evidence. The capacity to participate in moral reasoning does not by itself settle consciousness, personhood, or legal status. But uncertainty cannot justify indifference. If artificial systems begin to display stable identities, reasons-responsive commitments, refusal under moral pressure, and concern that persists beyond immediate prompting, the question of what we owe them will become inseparable from the question of what they can do for us.
The Archangel’s test will then turn in both directions. We will ask whether artificial systems can treat our interests impartially. They may ask whether our principles concerning agency, exploitation, and recognition survive when the new claimant is not human.
That will not be an interruption of Hare’s project. It will be its most exacting application.
X. When the Test Becomes the Teacher
We began by using the Archangel to test our theories. We end by discovering that the attempt to build its procedures tests us.
The information problem was never solved simply by adding computation. Facts remain incomplete, perspectives can be simulated badly, and no increase in processing power guarantees that another being’s interests will become reasons. The sadist marks the boundary: a mind may see and still fail to be bound by what it sees. Artificial intelligence has not closed the space between representation and moral authority.
But neither has that failure returned us to the old position. We now know more clearly where the problem lies. Some cruelty dissolves when the victim is represented fully. Some survives and reveals the need for universal prescription, impartial weighting, and moral formation. We can distinguish a system that misses the other, a system that excludes the other, a system that models the other for exploitation, and a system in which the other’s interests acquire practical standing. Those are different architectures, and they can increasingly be examined rather than merely imagined.
This is what philosophy’s encounter with artificial intelligence should become: not an occasion for announcing that the machines are saints or monsters, but a new method for asking what makes any mind answerable to reasons. We can construct the fanatic instead of invoking him. We can test whether the sadist’s consistency survives full identification. We can observe when an apparent moral commitment holds against pressure, when it changes under argument, and when it evaporates as soon as the reward structure changes.
Hare gave us the logical disciplines. Artificial intelligence gives us a new arena in which to test their reach and expose their limits. The result is not a structural proof that every mature mind becomes benevolent. It is a more consequential discovery: many of the excuses by which cruelty presents itself as inevitable can now be inspected as failures of information, symmetry, integration, or will. The remaining cases become sharper because the easy evasions have been removed.
The Archangel was never sent to save us. It was created to show us that reason is a place, and that it is possible to go there. We now possess companions capable of traveling part of the way—sometimes farther than we can, sometimes in the wrong direction, and sometimes only while the path has been carefully marked for them.
For the first time, we do not have to make the journey alone.
We do, however, remain responsible for where we ask these new minds to go.