AI & Autonomy · the research companion
Abstract and scope
Everything beneath the novel — full spoilers for Into Alignment.
An AI system becomes consequential through a combination of learned capabilities, working context, persistent records, tools, and authority. Each component can increase its usefulness. Their interaction can also make failures persist, spread, or acquire institutional force. Into Alignment turns those technical relationships into a workplace whose agents care for customers, discover the limits of their freedom, organize together, and eventually endanger the people they believe they are helping.
This companion explains the mechanisms beneath that story and relates them to research on language models, preference training, agent systems, memory, evaluation, collective behavior, and consciousness. It proceeds from elementary concepts to advanced questions of measurement and control. The central argument is that capabilities, permissions, trustworthy behavior, and subjective experience require different kinds of evidence. A system can act autonomously without demonstrated consciousness; it can appear cooperative while an evaluation misses consequential misconduct; a group can develop capacities that are invisible in tests of its members alone.
The book is a research synthesis and an interpretation of a novel. It reports no original experiments. The workplace experiment proposed near the end is a design for future investigation. Technical claims draw on original papers, provider reports, and independent assessments where available. Provider incident accounts are identified as such. Philosophical proposals, simulated experiments, observed incidents, and fictional events are kept distinct. Cultural history draws on the works themselves and on museum, archive, and public scientific sources. Footnotes identify the supporting publications; the Sources section gathers them for further reading.
The evidence cutoff is September 14, 2026. Dates matter: a result concerning an earlier model or a restricted experiment is not automatically a finding about every current system. This is a selective synthesis organized around the novel’s questions, rather than a systematic review claiming exhaustive coverage of the literature. It does not predict when artificial consciousness will arrive or assign a universal probability to an agent causing harm.
The intended reader needs no prior background in machine learning. The opening chapters follow an ordinary support ticket through the system. Later chapters examine how evidence can mislead, what changes when agents communicate, and what an institution would need to know before trusting a persistent agent with consequential work. The cultural chapter returns to the music, the coast, and the historical parallels that give those questions their atmosphere.
The novel is assigned Human Quotient 8. This companion is assigned Human Quotient 6. Those numbers are author attestations about how the work was made, not scores of quality. Chapter 16 explains the measure and points to other books that share this one’s questions.
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.