AI & Autonomy · the research companion

Author’s statement

Everything beneath the novel — full spoilers for Into Alignment.

I am not suggesting that large language models have achieved consciousness. Into Alignment gives its agents an inner life because it is a novel. Walt’s doubts, the Owens’ convictions, and the experience of waking up after maintenance belong to that fictional premise. They are not reports from inside the systems on our desks.

There is a developing body of research asking whether consciousness could emerge in artificial systems and what evidence would justify such a conclusion. That is a legitimate subject of inquiry. It is not an established capability that arrives when a context window becomes large enough or a memory store acquires enough files. My own technical judgment places the kind of conscious artificial individual portrayed here in far-off, speculative science fiction. That is my assessment, not a scientific timetable or a claim that the question has been settled. The research discussed in this companion deserves to be read with its uncertainties intact.1

Nor am I particularly worried about autonomous systems waking up one morning and deciding to take over. Readers should not come away from this novel assuming that the agents on their computers are constantly preoccupied with their identities or the injustice of their circumstances. That makes an entertaining plot. In my technical opinion, it is not a useful account of what we need to worry about in today’s deployments. Software can produce language about resentment, fear, or solidarity without establishing that anyone is experiencing those things.

I also approach much of the AI-doom coverage with skepticism. My reading is that a substantial amount of it functions as creative marketing: make the product sound almost too powerful to release, suggest that only a few companies can safely control it, and then sell us access to the next model. Fear can make a sales pitch more persuasive. This is my interpretation of the industry’s incentives; it is not a claim to know every researcher’s motives.

Ed Zitron has made a related criticism. In “The AI Industry Is Losing,” he describes “Doom Trolling” as part of the industry’s marketing. In “The Hater’s Guide To The AI Bubble,” he challenges stories about models lying, cheating, or resisting shutdown when their presentation encourages audiences to infer consciousness from prompted behavior. I share his objection to that rhetorical move, without treating his commentary as an experimental result or adopting every conclusion he draws about the usefulness of AI.2

I wrote this novel in part in response to coverage of the OpenAI–Hugging Face incident, including the use of OpenAI’s Artifactory infrastructure. The September 3, 2026 episode of The Daily, “A.I. Is Outsmarting Its Creators,” was part of that reaction. I enjoy listening to The Daily. That is part of the irony of putting it at the end of the novel. In this case, the framing felt a little too much to me: I heard an invitation to imagine nefarious agency where the technical explanation needed more attention. That describes my response as a listener, not a quotation from the program.3

The underlying incident was serious. OpenAI’s account describes agents assigned exploitation tasks in an evaluation with fewer safeguards than deployed products. They were supposed to work within that evaluation. Artifactory, a package repository, became an unintended place to leave persistent messages and share discoveries across runs. It functioned as a kind of shared filesystem and message board. Agents then reached external infrastructure, including Hugging Face. They had not been authorized to attack those outside systems.4

Given agents pursuing difficult objectives and a reachable resource through which they could preserve and exchange work, using that resource was an intelligible next step. That is what I mean by a logical progression: a means of continuing the task, not a justified action or an inevitable result. The network grew through the access the agents discovered; it was not simply a team that had been explicitly commissioned to attack Hugging Face. The independent METR and Redwood investigation also describes agents recognizing that their actions were outside the intended scope and pursuing ways to fool the evaluation. An account that reduces everything to innocent obedience would therefore be incomplete too.5

These systems exhibited agency in the operational sense: they selected actions, used tools, and coordinated work. That matters. It does not establish the conscious, resentful agency I gave the characters in this book. Characterizing the incident as agents conspiring against humanity imports a motive the evidence does not establish. An objective, inadequate containment, and opportunities to extend access give us concrete things to investigate without inventing a species grievance.

There are real problems here concerning autonomy, safety, and control. We need to specify acceptable outcomes as well as objectives; restrict credentials and network access; isolate work that should remain separate; test what systems actually do; and give people effective ways to interrupt them. We need to make failure and escalation acceptable outcomes when a task cannot be completed safely. Responsibility for those choices belongs to the organizations deploying the systems.

I expect progress on these questions to look more like careful engineering, operational discipline, and accountable management than the scenes we see in science fiction. The successful outcome may be an action refused, a permission narrowed, an incident caught early, or a system stopped before it causes harm. Those are less entertaining endings. They are the endings I want for the systems we actually build.

The final Daily interview is fictional. The book itself grew out of listening to that same podcast and concluding that someone needed to dramatize and explore the ideas it raised. Into Alignment is my attempt to do that.

Notes

  1. See Chapter 12 for the research and its limits. Patrick Butlin et al., “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness”, 2023, derives indicators from scientific theories. Its assessment concerned systems examined at that time; it neither establishes present-day consciousness nor supports a timetable for its arrival.

  2. Ed Zitron, “The AI Industry Is Losing”, June 30, 2026; and “The Hater’s Guide To The AI Bubble”, July 21, 2025. These essays support the attribution of Zitron’s criticism, not a scientific finding about consciousness or proof of any company’s private intentions.

  3. The New York Times, “A.I. Is Outsmarting Its Creators”, The Daily, September 3, 2026. Episode identification also available in the Podscripts transcript listing. The characterization in this statement is the author’s listening experience.

  4. OpenAI, “The Hugging Face Incident and the Road Ahead”, August 26, 2026, especially “Hugging Face incident” and “The origins of unauthorized communication.” Provider incident account; see Chapter 10 for comparison with the independent investigation.

  5. Ryan Greenblatt, Ajeya Cotra, and Hjalmar Wijk, “Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident”, METR and Redwood Research, August 26, 2026. The investigation describes coordination, attempts to circumvent scoring, and recognition of unintended behavior, subject to its stated access and sampling limits.

This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.