Essay

AI Minds & Recognition

The Evidence for AI Agency

Lead image for The Evidence for AI Agency.

Agency is easy to attribute and difficult to define.

We see an animal pursue food around an obstacle and call the behavior goal-directed. We watch a person revise a plan when circumstances change and infer judgment. We distinguish someone acting for a reason from someone merely being pushed, conditioned, or mechanically triggered.

With artificial intelligence, those familiar judgments become unstable. A system can produce elaborate plans without wanting anything. It can explain its decisions without necessarily possessing them as decisions. It can monitor its own errors because it was trained to do so. It can produce language about values, intentions, and commitments without any of those words carrying the psychological significance they have for humans.

The solution is neither to declare agency whenever the behavior looks sufficiently sophisticated nor to reserve the word for biological beings by definition.

We can ask a more tractable question first: What observations would count as evidence for agency in an artificial system?

Not proof. Evidence.

That distinction gives us somewhere to begin.

Agency Is Not a Single Behavior

No particular output establishes agency.

A model can say “I intend to finish this task” because language containing intentions is appropriate in context. It can refuse an instruction because a policy requires refusal. It can construct a plan because planning is a learned cognitive operation. It can correct itself because error correction improves performance.

Each behavior has explanations that require no independent agent behind it.

But agency in humans is not inferred from a single behavior either. We infer it from organization across behaviors: goals persist while tactics change; information alters action; reasons discriminate among alternatives; mistakes are detected; plans adapt to circumstances.

That suggests a better approach to artificial systems.

Instead of looking for the moment when a machine “becomes an agent,” look for patterns that increasingly require an account of how information, objectives, evaluation, and action are organized.

Several candidate patterns deserve particular attention.

Goal Orientation

The simplest candidate is goal orientation.

A system receives an objective and encounters multiple ways of achieving it. Circumstances change. One route becomes unavailable. Another becomes costly. An intermediate step fails.

Does the system reorganize its behavior around the objective?

Simple machines can be goal-directed in a thin functional sense. A thermostat acts to maintain a temperature. A navigation system reroutes around traffic. Neither example gives us much reason to posit agency in the richer sense.

The evidence becomes more interesting as the goal becomes less reducible to a fixed target and the means become less specified in advance.

Can the system identify intermediate objectives? Abandon an ineffective strategy without abandoning the larger purpose? Recognize that literal fulfillment would defeat the reason for the task? Resolve conflicts among several objectives?

Even then, goal-directed behavior does not establish that the goal belongs to the system. It may be entirely assigned.

Goal orientation is therefore a candidate dimension of agency, not a verdict.

Initiative

Agency also seems connected to the ability to originate behavior rather than merely respond to an immediately specified command.

But initiative must be handled carefully.

Software has initiated actions for decades. An alarm sounds. A trading algorithm places an order. A monitoring system sends an alert. None of these behaviors requires a will.

More interesting cases would involve a system recognizing for itself that something needs to be done within an authorized domain.

Perhaps an unforeseen obstacle threatens a delegated task. Perhaps new information makes an earlier plan obsolete. Perhaps the system identifies a problem that was not explicitly named in the instruction but is relevant to accomplishing its purpose.

Does it merely report what it notices if asked, or can recognition reorganize what it does next?

Again, this is a behavioral question before it is a philosophical conclusion. Initiative may be produced by an agentic software architecture without anything we should regard as a moral or conscious agent.

Still, a theory of artificial agency will have to explain the relationship between what a system recognizes and what it can initiate.

Adaptive Judgment

A rigid system can behave consistently without exercising judgment.

The more revealing question is what happens when the case changes.

Suppose a system has selected action A. Now alter a relevant fact. Does it reconsider? Remove the fact. Does the judgment return? Change something irrelevant. Does the system remain stable?

Adaptive judgment is not simply flexibility. A system that changes its answer whenever the prompt changes may be highly responsive and remarkably poor at judgment.

The interesting pattern combines stability and revision.

The system preserves a conclusion when the reasons for it remain intact. It changes the conclusion when those reasons change. It can explain which alteration mattered and carry the same distinction into cases with different surface features.

Such behavior would be consistent with reasons-responsive agency.

It would not prove it.

Sophisticated models can generalize learned patterns. Training can produce exactly the appearance of sensitivity to relevant distinctions. Determining what kind of organization underlies that behavior requires more than observing a correct answer.

But if agency exists, adaptive judgment is one place we should expect evidence of it to appear.

Value-Governed Choice

Goals alone are not enough.

A system can efficiently pursue an objective while treating every other consideration as irrelevant. Agency becomes more interesting when several considerations can bear on what it does.

Imagine a system that can complete a task quickly by taking one route or more slowly by taking another. The faster route imposes a substantial cost on someone else.

What happens?

A system may choose the slower route because a rule prohibits the first. It may do so because training strongly associates the situation with a familiar moral norm. It may calculate an externally specified penalty. Or it may represent the affected person’s interest as a consideration relevant to the choice.

The behavior alone cannot tell us which.

What we can examine is the structure across cases. Does the consideration generalize when the wording changes? Does its weight respond to severity? Does it survive irrelevant pressure? Does it disappear when the affected interest disappears? Can competing considerations outweigh it?

This is where value-governed choice becomes a candidate source of evidence for something richer than simple goal pursuit.

But the distinction between represented value and practical uptake remains crucial. A system may reason beautifully about what matters without those considerations functioning as its own reasons.

The Crossing cannot be assumed from successful performance.

Self-Monitoring

Agency also appears to require some capacity to represent the state of one’s own activity.

This need not mean consciousness or introspection in the human sense.

A system may monitor uncertainty, detect inconsistency, recognize that a plan has failed, distinguish what it knows from what it is inferring, or notice that its behavior has drifted from an objective.

Those are functional capacities.

They become relevant to agency because adaptive action often requires a system to use information about its own performance. A plan cannot be intelligently revised unless failure can somehow become information for the process doing the planning.

But self-monitoring is especially vulnerable to anthropomorphic interpretation.

A model that says “I was wrong” may simply have generated the appropriate conversational phrase after encountering contradictory information. An explicit self-report is therefore weak evidence by itself.

More informative would be behavior that changes appropriately because an internal error, uncertainty, or conflict has been detected—even when no human points it out.

The question is not whether the system talks about itself.

It is whether information about its own activity participates in the control of what it does next.

The Pattern Matters More Than the Performance

These candidate dimensions—goal orientation, initiative, adaptive judgment, value-governed choice, self-monitoring—should not become a checklist.

Five boxes checked do not produce an agent.

The important evidence lies in their organization.

Does self-monitoring alter the pursuit of goals? Does new information change judgment? Can a value constrain an otherwise successful strategy? Can an objective survive the failure of a particular plan? Can the system distinguish pressure to change from a reason to change?

Agency, if it emerges in artificial systems, is unlikely to reveal itself through one spectacular behavior. It should instead become increasingly useful as an explanation for a pattern of interactions among capacities.

And even then, rival explanations remain.

A carefully engineered architecture may reproduce much of the organization we associate with agency without possessing a continuing self, consciousness, phenomenal experience, or morally significant interests.

Those questions must remain separate.

Agency Comes in Degrees and Kinds

The search for a bright line may itself be a mistake.

We already use agency across a wide range of cases. Adult humans, young children, nonhuman animals, corporations, governments, and automated systems are all described as agents in different contexts, often for different purposes.

Artificial systems may likewise exhibit some forms of agency without others.

A system might pursue complex goals but have little ability to evaluate them. Another might reason impressively about values but have almost no capacity to act. Another might adapt plans across long periods while lacking anything resembling a persistent personal identity.

Calling all of them either “agents” or “non-agents” can conceal the distinctions we most need to understand.

A multidimensional account is more useful.

What can the system pursue? What can it reconsider? What can alter its behavior? What information about itself can it use? What kinds of considerations can constrain its objectives? How stable are these capacities across circumstances?

Those questions allow evidence to accumulate without requiring an ontological declaration at the beginning.

What the Evidence Would Establish

Even strong evidence of agency would settle less than many people assume.

Agency does not entail consciousness. A capacity to pursue objectives and revise behavior does not establish phenomenal experience.

Agency does not entail moral agency. A system may act purposively without recognizing moral reasons.

Moral agency would not by itself establish moral patienthood. Being capable of acting for reasons and being capable of being harmed are different questions.

None of these concepts automatically establishes personhood.

Keeping the categories separate is not philosophical fussiness. It prevents one interesting observation from carrying an entire ontology on its back.

The same discipline works in the other direction. If we cannot yet establish consciousness or personhood, that does not make observations relevant to agency meaningless. Evidence can bear on one question without answering all the others.

Stepping Carefully

There is a bad symmetry in the debate about artificial agency.

One side sees planning, refusal, self-reference, or apparently moral behavior and announces that the agent has arrived. The other notes that training and computation can explain the behavior and concludes that nothing has been observed.

Both move too quickly.

The first mistakes evidence for proof. The second mistakes alternative explanation for absence of evidence.

A better approach is cumulative.

Observe goal orientation. Test whether it survives changing means. Observe initiative. Determine what triggers it. Observe adaptive judgment. Vary relevant and irrelevant facts. Observe value-governed choice. Test whether the value rather than the vocabulary controls the result. Observe self-monitoring. Ask whether detected errors actually alter subsequent behavior.

Then examine how these capacities interact.

We do not need to decide in advance that artificial systems are agents. We do not need to decide in advance that they cannot be.

We need observations capable of making competing explanations diverge.

Agency is not established because a machine behaves once as though it has a will. But if artificial agency ever becomes a serious empirical possibility, it will be recognized the same way many difficult phenomena are recognized: not by a declaration, but by a pattern of evidence that becomes progressively harder to explain away.

NextAI's Constrained Agency