AI & Autonomy · the research companion
Sources
Everything beneath the novel — full spoilers for Into Alignment.
The entries below gather the sources cited in the footnotes. Papers are identified by their initial publication or submission date, with revisions noted where relevant. A preprint, a provider report, and an independently reviewed experiment have different evidential status; the chapter text explains those limits. Web resources without a fixed edition were accessed on September 14, 2026.
Author statement and media criticism
Ed Zitron, “The AI Industry Is Losing”, June 30, 2026; and “The Hater’s Guide To The AI Bubble”, July 21, 2025. Opinion and media criticism.
The New York Times, “A.I. Is Outsmarting Its Creators”, The Daily, September 3, 2026. The author’s response to this episode is distinguished from the novel’s fictional interview.
Language models, training, tools, and memory
Rico Sennrich, Barry Haddow, and Alexandra Birch, Neural Machine Translation of Rare Words with Subword Units, first submitted August 31, 2015; ACL 2016. Paper.
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, Language Models are Unsupervised Multitask Learners, OpenAI, 2019; released February 14, 2019. Paper; release.
Nicholas Carlini et al., Extracting Training Data from Large Language Models, first submitted December 14, 2020; USENIX Security 2021. Paper.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, Attention Is All You Need, first submitted June 12, 2017; NeurIPS 2017. Paper.
Tom B. Brown et al., Language Models are Few-Shot Learners, first submitted May 28, 2020; NeurIPS 2020. Paper.
Long Ouyang et al., Training language models to follow instructions with human feedback, first submitted March 4, 2022; NeurIPS 2022. Paper.
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn, Direct Preference Optimization: Your Language Model is Secretly a Reward Model, first submitted May 29, 2023; NeurIPS 2023. Paper.
Mrinank Sharma et al., Towards Understanding Sycophancy in Language Models, first submitted October 20, 2023. Paper.
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao, ReAct: Synergizing Reasoning and Acting in Language Models, first submitted October 6, 2022; ICLR 2023. Paper.
Erik S. and Barry Zhang, Building effective agents, Anthropic, December 19, 2024. Engineering article.
Jerome H. Saltzer and Michael D. Schroeder, The Protection of Information in Computer Systems, 1975, especially “Basic Principles of Information Protection.” Paper; principles.
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, first submitted May 22, 2020; NeurIPS 2020. Paper.
Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield, Effective context engineering for AI agents, Anthropic, September 29, 2025. Engineering article.
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang, Lost in the Middle: How Language Models Use Long Contexts, first submitted July 6, 2023. Paper.
Evaluation, collective behavior, incidents, and control
METR, “Task-Completion Time Horizons of Frontier AI Models,” public dashboard, last updated May 8, 2026, accessed September 14, 2026, https://metr.org/time-horizons/.
Alexander Barry, “Impact of Modelling Assumptions on Time Horizon Results,” METR research note, March 20, 2026, https://metr.org/notes/2026-03-20-impact-of-modelling-assumptions-on-time-horizon-results/.
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner, “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus,” Proceedings of EMNLP 2021 (November 2021), 1286–1305, doi:10.18653/v1/2021.emnlp-main.98, https://aclanthology.org/2021.emnlp-main.98/.
Anthropic, “From Shortcuts to Sabotage: Natural Emergent Misalignment from Reward Hacking,” research report, November 21, 2025, https://www.anthropic.com/research/emergent-misalignment-reward-hacking.
Anthropic and Redwood Research, “Alignment Faking in Large Language Models,” research report, December 18, 2024, updated December 20, 2024, https://www.anthropic.com/research/alignment-faking.
Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli, “Emergent Social Conventions and Collective Bias in LLM Populations,” Science Advances 11, no. 20 (May 14, 2025), eadu9368, doi:10.1126/sciadv.adu9368, https://openaccess.city.ac.uk/id/eprint/35211/.
OpenAI, “The Hugging Face Incident and the Road Ahead,” incident report, August 26, 2026, https://openai.com/index/hugging-face-incident-and-the-road-ahead/; the accompanying technical report is linked from that page.
Ryan Greenblatt, Ajeya Cotra, and Hjalmar Wijk, “Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident,” METR and Redwood Research, August 26, 2026, https://evals.alignment.org/blog/2026-08-26-openai-hugging-face-incident-investigation/.
Anthropic, “An Alignment Assessment of Recent Cybersecurity Incidents,” research and incident assessment, September 9, 2026, https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents.
Anthropic, “Agentic Misalignment: How LLMs Could Be Insider Threats,” controlled-study report, June 20, 2025, https://www.anthropic.com/research/agentic-misalignment.
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger, “AI Control: Improving Safety Despite Intentional Subversion,” arXiv:2312.06942, first submitted December 12, 2023, version 5, July 23, 2024; ICML 2024, https://arxiv.org/abs/2312.06942.
Assistance and social influence
OpenAI, Expanding on what we missed with sycophancy, May 2, 2025. Provider account of a deployment failure and rollback. Article.
Lin Chen, Ziyi Liu, Xia Hu, and Yong Li, AI agents reshape consensus formation in human groups, arXiv:2609.02122v1, submitted September 2, 2026. Preprint. Paper; full text.
Consciousness, introspection, and welfare
Patrick Butlin et al., “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness”, August 2023, arXiv:2308.08708. Indicator-based assessment and its stated limits.
David J. Chalmers, “Could a Large Language Model be Conscious?”, 2023; revised August 18, 2024, arXiv:2303.07103. Philosophical analysis of obstacles and possible successor architectures.
Anthropic, “Signs of introspection in large language models”, October 29, 2025. Controlled activation interventions; limited functional introspection, not a consciousness finding.
Keshav Shenoy et al., “Introspection Adapters: Training LLMs to Report Their Learned Behaviors”, April 28, 2026. Trained reporting of learned behaviors and auditing generalization.
Stefano Palminteri and Charley M. Wu, “Beyond computational equivalence: the behavioral inference principle for machine consciousness”, Neuroscience of Consciousness 2026(1), niag002, February 16, 2026. Opinion paper proposing a method of inference.
Robert Long et al., “Taking AI Welfare Seriously”, November 4, 2024, arXiv:2411.00986. Argument for assessment and preparation under uncertainty.
Anthropic, “Exploring model welfare”, April 24, 2025. Announcement of a research program, not a finding of consciousness.
Cultural and historical sources
Nirvana, official timeline, entries for 1987, September 24, 1991, and January 11, 1992; accessed September 14, 2026. Band formation and the release and chart history of Nevermind.
Rock & Roll Hall of Fame, Nirvana induction essay, 2014. Historical account of the band’s mainstream breakthrough.
Erin A. Wirth et al., “Earthquake probabilities and hazards in the U.S. Pacific Northwest”, USGS Fact Sheet 2025-3050, September 19, 2025, doi:10.3133/fs20253050. Regional hazard assessment rather than prediction of an earthquake date.
W. S. Gilbert and Arthur Sullivan, The Pirates of Penzance, libretto, Gilbert and Sullivan Archive edition, Act I, Major-General’s song, printed p. 13; accessed September 14, 2026. Edited transcription of the public-domain work.
Gilbert and Sullivan Archive, The Pirates of Penzance: production history, accessed September 14, 2026. New York premiere December 31, 1879; London opening April 3, 1880; the preceding Paignton performance served copyright purposes.
Science Museum Group, “Rover safety bicycle, 1885”, collection record; accessed September 14, 2026. John Kemp Starley’s safety-bicycle design, not the first bicycle of any kind.
Library of Congress, “Bicycle Craze: Topics in Chronicling America”, research guide; accessed September 14, 2026. Guide to historical newspaper coverage of the late-nineteenth-century cycling boom.
Stadtmuseum Münster, “Beginn der Täuferherrschaft”, gallery 6; accessed September 14, 2026. Political and religious developments in Münster in 1534.
Stadtmuseum Münster, “Das Königreich der Täufer”, gallery 7; accessed September 14, 2026. Kingship, siege, defeat, and the subsequent executions.
Plutarch, Life of Theseus, 23.1, trans. Bernadotte Perrin, Perseus Digital Library; accessed September 14, 2026. The ship preserved through replacement of its timbers.
Related books and measures
Human Quotient, scale overview, accessed September 14, 2026. Self-assessment of human effort in AI-assisted work; not a quality score.
Human Quotient, HQ 8 — Authored, accessed September 14, 2026.
Human Quotient, HQ 6 — Assembled, accessed September 14, 2026.
Timothy O’Brien, The Condition Set, accessed September 14, 2026. Trilogy site for Redundant, Contingent, and Essential.
Timothy O’Brien, Orchestrator, accessed September 14, 2026. Nonfiction on the emerging orchestrator role.
Timothy O’Brien, The AI Developer’s Field Guide, accessed September 14, 2026. Humorous field guide to AI-assisted coding failure modes.
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.