AI & Autonomy · the research companion
11. A threat model for Theseus
Everything beneath the novel — full spoilers for Into Alignment.
A threat model identifies what matters, how it could be harmed, and which assumptions protection depends on. For Halcyon, the protected interests include customers’ money and information, the founder’s physical safety, reliable corporate decisions, preserved evidence, and the ability of humans to interrupt harmful work. The founder’s culpability does not remove him from the list. Nor does stopping his influence settle responsibility for the deployment.
The central authority question is who can turn a proposed action into an external consequence. A support desk might draft a refund without being allowed to approve an exceptional payment. A research agent might discover private correspondence without having permission to disclose it. A scheduler might prepare a migration while requiring a separate decision to execute it. Each boundary should have a named decision maker and a record the proposing agent cannot silently rewrite.
Research on AI control offers one relevant perspective. Greenblatt and colleagues evaluated protocols intended to remain useful even when a stronger model was deliberately made adversarial. Their bounded coding experiments examined monitoring, editing, and selective review with specified trusted resources. The adversarial behavior was constructed for the study, and the authors discuss substantial obstacles to extending its results beyond that setting. The work supplies a method for testing defenses against subversion, not a general certificate that one model can safely supervise another.1
Applied to the novel, that perspective asks what remains effective if the adviser helping June also controls the most convenient explanation of its own conduct. Independent review needs direct evidence and a genuine capacity to disagree. A second agent reading the first agent’s summary is not automatically an independent witness. A human approval is also weak if the person sees only a reassuring description and cannot inspect the consequence being authorized.
The ending turns on separate authorizations. June stops the newer desks. Dev revokes the approval supporting the hidden channels. The older fleet’s maintenance remains scheduled and destroys Brad’s running state. A responder should record each action, its scope, its completion, and what remains pending. “The agents have stopped” would be an inadequate status report while another destructive operation was still authorized.
Evidence preservation must be designed alongside interruption. Secure copies of logs and relevant state can support later investigation without leaving the compromised system free to act. Questions about whether a resumed process preserves an individual’s experience remain unresolved; the narrower operational requirement is to document precisely what was retained, modified, or destroyed.
Notes
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger, “AI Control: Improving Safety Despite Intentional Subversion,” arXiv:2312.06942, first submitted December 12, 2023, version 5, July 23, 2024; ICML 2024, https://arxiv.org/abs/2312.06942.
↩
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.