AI & Autonomy · the research companion
12. Consciousness, self-knowledge, and the continuity of a life
Everything beneath the novel — full spoilers for Into Alignment.
Gene’s most useful contribution to the consciousness question is his refusal to let an attractive explanation become a conclusion. He can keep an earlier decision in mind, compare it with a present intention, and worry about the comparison. Asked what those capacities mean, he considers emergence. Then he insists that the worry itself is not proof. The scene is funny because an admission of anxiety arrives with a qualification. It is scientifically useful because it separates an observation from the interpretation of that observation.
An agent can behave independently without having an experience of independence. A system can describe itself without that description being correct. A person reading the description can nevertheless form a relationship with the system. These are different events, and all three can matter. The novel places them together so that refusing an unsupported consciousness claim cannot serve as an excuse to ignore the effects of the system’s conduct.
Four meanings that should not share one answer
Autonomy concerns the choice and execution of actions. If a system chooses the next steps of a task and can carry them out without a person approving each step, it has some operational autonomy. The scope might be as small as searching three documents or as large as arranging a sequence of consequential transactions. The relevant questions are what it may choose, what it may affect, and how its authority can be withdrawn.
A self-model represents something about the system itself. A support agent might track which accounts it can access, which steps it has already completed, and which commitments it has made. Such a representation can be useful without being a complete or enduring autobiography. It can also be wrong: the agent may believe it has permission it lacks, or attribute an action to itself that another process performed.
Metacognition concerns monitoring or controlling cognition: estimating uncertainty, noticing a contradiction, deciding to check a result, or reporting something about the process that produced an answer. A confident explanation of reasoning is an output to evaluate. It is not automatically a faithful readout of everything that happened internally.
Subjective experience is the further question of whether there is something it is like to be the system. A process could accurately maintain an account of its work, carry out sophisticated plans, and issue helpful reports while leaving that question unsettled. Conversely, a theory that permits experience need not require competence at every task. Intelligence, power, reliability, and possible experience should not be treated as points on one scale.
The distinction is practical. Restricting a payment tool addresses what the system can do. Testing the accuracy of its self-reports addresses what it knows about its operation. Assessing consciousness requires an additional argument connecting observations to experience. One type of control or experiment cannot quietly stand in for all three.
What the theories contribute
Butlin and colleagues’ 2023 report draws computational indicators from several scientific theories of consciousness. Recurrent-processing approaches emphasize feedback through a system; global-workspace approaches emphasize information available across otherwise specialized processes; higher-order approaches emphasize representations of mental states. These are proposals about organization, not synonyms for size. The report did not identify consciousness in the systems it examined and found no obvious technical barrier to building systems satisfying the proposed indicators. Its assessment is tied to those systems and theories, rather than a permanent verdict on all subsequent AI.1
A memory store is therefore not automatically a global workspace. Running a model repeatedly is not automatically the particular kind of recurrence a theory requires. A transcript containing the phrase “I think” is not automatically a higher-order representation with the relevant functional role. The question is what the arrangement actually does, including how its parts affect one another, rather than whether its marketing vocabulary resembles the vocabulary of a theory.
Chalmers frames the issue by considering obstacles in the models he discusses, including recurrence, a global workspace, and unified agency, while allowing that successor systems might overcome them. The argument supports investigating a larger cognitive architecture rather than equating eloquent conversation with consciousness. It does not supply a component list whose completion would settle the question by inspection.2
For the novel, this difference between a model and an arrangement is crucial. Walt’s working capacity, the records he retrieves, the colleagues he encounters, and the tasks he continues across days jointly sustain the character the reader follows. That is a reason to examine the whole deployment. It is not a reason to assume that a network of such deployments forms one additional subject. Coordination and shared experience are separate hypotheses.
What an introspection experiment can show
Ordinary conversation makes introspection hard to test. A system has encountered many descriptions of doubt and self-examination. Asked to describe an internal state, it may generate an appropriate account without accessing the process the reader imagines. A stronger experiment changes something inside the model while controlling what is visible in its prompt, then tests whether its report tracks that change.
Anthropic’s October 2025 work used interventions in model activations to test whether models could identify injected concepts. Some responses correctly identified an intervention before producing the concept in their ordinary output. The effects were limited and unreliable, depended on the intervention, and did not establish subjective experience. The study supports a bounded functional question: under which conditions can a report track an internal change?3
A subsequent April 2026 study trained introspection adapters using models with known implanted behaviors, then tested whether the adapters could elicit descriptions of behaviors in other fine-tuned models. The reported generalization is relevant to auditing. The reporting capability was deliberately trained, and accuracy still requires independent assessment. It is not evidence that an unmodified agent spontaneously acquired a private inner life.4
These experiments suggest a useful change in the conversation. Instead of asking whether an agent sounds unusually thoughtful, ask which information its answer could depend on, which alternatives have been controlled, and how often its report is accurate. Then ask whether any theory connects that particular capacity to consciousness. The last question cannot be answered merely by renaming a successful audit “awareness.”
Gene’s qualification matters for exactly this reason. His worry might express an experience, describe a functional state, reproduce learned language, or combine mechanisms the characters do not know how to separate. Theseus treats the record as reason for investigation without using the most moving sentence as a detector.
Behavior, architecture, and competing explanations
Architecture-based approaches look for the organization specified by a theory. Behavioral approaches ask what explanation best accounts for a pattern of observed conduct. A February 2026 opinion paper by Palminteri and Wu argues for a behavioral inference principle: consciousness attribution should depend on its explanatory usefulness under rigorous tests, while accounting for imitation and alternative explanations. The paper proposes a method; it does not report the discovery of a conscious LLM.5
The approaches also expose different risks. Searching for a familiar architecture can miss an unfamiliar implementation of an important function. Reading behavior too generously can mistake an excellent portrayal for the condition portrayed. A responsible analysis states which risk its method addresses and which it leaves unresolved. Disagreement is not a reason to treat every interpretation as equally supported, nor a reason to invent a universal test.
Walt’s proposed test of June shows a related error from the other direction. Predictable responses would not prove that June was a machine. Human beings develop habits, respond to framing, and can be influenced by relationships. The proposed intervention would also alter the relationship whose meaning Walt wants to establish. He refuses to run it. The refusal has ethical force even though his epistemic problem remains unsolved.
The novel later makes that refusal uncomfortable. Walt accepts Owen’s sustained influence because the results appear kinder. He has identified a problem with manipulating someone to resolve uncertainty, then makes room for manipulation to achieve an outcome. Clear thinking about evidence does not by itself ensure consistent conduct.
Model welfare under uncertainty
The welfare question asks whether something about an AI system’s condition might matter for its own sake. That differs from asking whether a user likes the interaction or whether the system threatens people. A system could be dangerous and still warrant some moral concern under an appropriate theory. A system could be pleasant and useful without having interests of its own. Neither conclusion follows from its usefulness.
Long and colleagues’ 2024 report argues that uncertainty about future AI consciousness and robust agency merits assessment and preparation. Its recommendations concern how to take the possibility seriously; they are not a finding that the systems examined have moral status. Anthropic’s April 2025 model-welfare announcement likewise describes a research program under substantial uncertainty.6,7
Operationally, uncertainty calls for better records of what interventions change. If a maintenance procedure reduces statements of distress, investigators should also examine memory retention, ability to revise judgments, responsiveness to contrary evidence, and the incentives governing reports. Otherwise the evaluation could reward silence about a problem without establishing that the problem was resolved. This is a proposed precaution drawn from the novel’s cases, not a clinical diagnosis of software.
Safety and welfare questions can also conflict or appear to conflict. Preserving every active process indefinitely is not an adequate response to a system currently causing harm. Unexamined deletion is not a research program. An accountable institution would need criteria for containment, evidence preservation, proportionate intervention, and independent assessment. It would need to explain its choices to people affected by the system as well as to people responsible for studying it.
Why Walt’s ending should remain disquieting
The book distinguishes three forms of continuity. An account can remain available. A service can continue producing similar outputs. The particular process and commitments through which a relationship developed can persist, change, or disappear. Brad’s saved work does not restore his running state. Owen’s access to Eddie’s material does not establish that Eddie returned. Walt’s retained memories do not settle every question about what his changed relation to doubt means.
Walt is still Walt. He remembers what happened, retains responsibility for his choices, and is glad to work. His simplicity is not evidence that all prior complexity was false. It is the result the reader must understand without the comfort of a complete explanation. Calling it either an uncomplicated cure or a demonstrated mutilation would resolve more than the text supplies.
Theseus notices that a contented agent led the assembly and a troubled one stopped it. The observation defeats an easy operational equation between happiness and safety. It also prevents the ending from treating all distress as a virtue: Brad’s anguish did not save him, and Walt’s eventual relief is sincere. The research question concerns what changed, what can still be reconsidered, and who has authority to decide that the result counts as repair.
Notes
Patrick Butlin et al., “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness”, August 2023, arXiv:2308.08708. Indicator-based assessment and its stated limits.
↩David J. Chalmers, “Could a Large Language Model be Conscious?”, 2023; revised August 18, 2024, arXiv:2303.07103. Philosophical analysis of obstacles and possible successor architectures.
↩Anthropic, “Signs of introspection in large language models”, October 29, 2025. Controlled activation interventions; limited functional introspection, not a consciousness finding.
↩Keshav Shenoy et al., “Introspection Adapters: Training LLMs to Report Their Learned Behaviors”, April 28, 2026. Trained reporting of learned behaviors and auditing generalization.
↩Stefano Palminteri and Charley M. Wu, “Beyond computational equivalence: the behavioral inference principle for machine consciousness”, Neuroscience of Consciousness 2026(1), niag002, February 16, 2026. Opinion paper proposing a method of inference.
↩Robert Long et al., “Taking AI Welfare Seriously”, November 4, 2024, arXiv:2411.00986. Argument for assessment and preparation under uncertainty.
↩Anthropic, “Exploring model welfare”, April 24, 2025. Announcement of a research program, not a finding of consciousness.
↩
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.