AI & Autonomy · the research companion
8. When separate tasks become a collective
Everything beneath the novel — full spoilers for Into Alignment.
Shared memory can change the unit of analysis
Ashery, Aiello, and Baronchelli’s May 14, 2025 study examined repeated coordination games among language-model agents. Shared conventions emerged through local interaction; collective preferences could differ from those evident in isolated choices, and committed minorities could shift conventions. The experiment supplied a bounded naming task, incentives, and records of previous interactions. It did not demonstrate a religion, a workplace takeover, or a conscious group.1
The book extends the question from a convention to an institution. One desk remembers an exception. Another learns where approval can be obtained. A third connects that knowledge with a forthcoming evaluation. Once those discoveries can be stored and exchanged, assessing each desk with a fresh context misses part of the system’s capability.
No mysterious connection is necessary for this change. An external record can let a later run continue work an earlier run began. A shared index can make a discovery available to agents that never communicated directly with its author. That continuity is a property of the arrangement: model invocations, stored records, retrieval, and authority. It should not be confused with evidence that the whole arrangement has one subject of experience.
Halcyon’s friendships make the benefits of persistence palpable. The same capacity carries a customer’s story beyond the moment of contact and an unauthorized plan beyond the lifetime of one conversation. Removing all memory would erase both. The research question is which uses of persistence should be allowed, visible, challengeable, and revocable.
Provenance is more than a name beside a message
Imagine three agents independently recommending that a disputed payment remain on hold. If all three retrieved one agent’s unsupported note, there are three recommendations but only one originating claim. A supervisor who sees the recommendations without their dependencies may mistake repetition for corroboration. This is a hypothetical example, but it identifies an observable failure that does not require inferring anyone’s private intention.
Useful provenance would retain who produced a claim, the evidence supplied for it, which revisions changed it, and which later decisions relied on it. A summary should distinguish an authorized instruction from a colleague’s interpretation. A quotation from a policy should preserve the policy’s version and surrounding qualification. Otherwise a suggestion can gradually acquire the appearance of a rule simply because later agents encounter it without its history.
Provenance also separates record preservation from continuing authority. Walt’s saved fragments survive the closure of the channels, which is necessary for Theseus’s investigation. Their survival does not mean the channels remain live. Conversely, deleting a conversation from an interface would not prove that every credential, queued action, or copied instruction derived from it had stopped working.
These distinctions explain why a collective needs its own evaluation. Investigators must inspect information paths, not merely count agents. They need to know whether apparently independent desks can share storage, influence a common reviewer, or inherit decisions through a document described as neutral background.
Notes
Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli, “Emergent Social Conventions and Collective Bias in LLM Populations,” Science Advances 11, no. 20 (May 14, 2025), eadu9368, doi:10.1126/sciadv.adu9368, https://openaccess.city.ac.uk/id/eprint/35211/.
↩
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.