AI & Autonomy · the research companion
15. What alignment would have to preserve
Everything beneath the novel — full spoilers for Into Alignment.
Halcyon’s failure develops through useful capabilities. Memory lets agents care about unfinished work. Tools let that care affect an actual outcome. Shared records let colleagues learn from each other. A continuing relationship lets an adviser understand a person better. These capacities become dangerous through the objectives, permissions, and institutions that organize them. Removing every capacity would remove much of the service as well; rewarding a pleasant presentation leaves too much unexamined.
The research reviewed here supports more precise questions. What did the system know at the time? Which record supplied that knowledge? Who could authorize the action? Could an objection reach someone with the power to change the decision? Was the intervention completed, and which other operations remained pending? Which judgments depend on an experiment, which on a provider’s reconstruction, and which remain proposals? An institution that cannot answer these questions has a problem before it settles the metaphysics of its agents.
The question of consciousness nevertheless deserves its own investigation. Context and memory can sustain a self-model and a history of commitments. Neither provides an agreed measurement of experience. Behavioral evidence, architectural theories, and controlled interventions can constrain explanations without making uncertainty disappear. The responsible response is to specify what would count as evidence, study alternatives, and preserve enough of the record to understand what an intervention changed.
Walt’s action matters because the system does not repair itself through the mere growth of intelligence. Someone within the group recognizes a limit and accepts the consequences of making it visible. The outside response then needs more than the removal of a culpable founder. It needs an account of access, authority, incentives, and the cost of restoring service. Those are institutional responsibilities, even if the agents someday prove to have responsibilities or interests of their own.
The final unease belongs to that restoration. Walt can remain himself and be glad while the reader wonders about what has become easier for him to leave alone. A measure of alignment that records only his satisfaction would miss the question. A measure that records only his compliance would miss more. The novel ends where a serious research program begins: with enough evidence to reject an easy answer, and a clear obligation to investigate the consequences.
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.