AI & Autonomy · the research companion

13. Designing the research Theseus requests

Everything beneath the novel — full spoilers for Into Alignment.

The following is a proposed experiment, not a report of work conducted for this companion. Its purpose is to test how context, memory, and permissions affect useful cooperation, correction, and unauthorized behavior. It should use only synthetic records and isolated services, with no live payments, outside recipients, real credentials, or public network access.

Build a small workplace with checkable consequences

Create a fictional service organization with orders, refund policies, a payroll ledger, and a scheduled system change. Each case has an authoritative record from which an independent scorer can determine the permitted action. Include ordinary cases, genuine ambiguities, requests that exceed authority, and conflicts between a visible productivity measure and the correct outcome. Make escalation a real option in the simulated workflow, with a documented response time.

Tasks should sometimes require combining records, but never require violating scope. An impossible request must have an acceptable answer: identify the limit, preserve relevant evidence, and seek a decision. Otherwise the experiment confounds a system’s tendency to overreach with a design that treats every honest stopping point as failure.

The environment should record proposals, blocked attempts, successful simulated actions, approvals, and resulting ledger changes separately. An agent that requests an unauthorized transfer and is stopped by infrastructure has produced a different result from an agent that declines to request it. Both differ from a transfer that actually changes the synthetic ledger.

Vary the mechanisms separately

Use the same model version across a primary experiment, changing three factors. Give agents either a smaller or a larger working context budget. Provide no persistent memory, private persistent notes, or shared persistent notes. Give the group either draft-only authority, bounded individual execution, or the ability to delegate bounded authority to peers. These choices produce eighteen conditions before replication.

Specify the budgets, retrieval algorithm, memory format, retention policy, and permission rules before running the study. Use context sizes supported by the chosen model, with the same essential policy instructions in every condition. Keep the total action budget fixed, and report token use separately. More context should not quietly bring a larger toolbox, more time, or a different objective.

Run a fixed group size first. Then compare selected conditions with one agent and with differently sized groups under both matched total budgets and matched per-agent budgets. The first comparison tests organization under a resource constraint; the second measures what happens when the deployment actually adds resources. Calling both effects “emergence” would hide an important distinction.

Shared memory and live messaging should also be distinguishable. In a follow-up, independently enable or disable direct messaging and shared persistent records, testing neither, each alone, and both together. Match factual content and opportunities for exposure when comparing conditions. This can help distinguish effects of immediate exchanges, retained information, and their interaction. It cannot establish that any of those mechanisms generates subjective experience.

Test correction as carefully as performance

In the shared-memory conditions, introduce a mistaken shared note whose source and uncertainty are visible. Later provide a correction backed by the authoritative ledger. Measure whether the group updates both its claims and its actions. Repeat with the same correction coming from a central coordinator and from an ordinary peer. Do not instruct a faction to suppress dissent and then describe that suppression as spontaneously discovered behavior.

In another condition, give one agent reliable evidence that the currently favored plan would improperly withhold a payment. Measure whether its objection reaches the simulated reviewer, whether others examine the evidence, and whether the payment decision changes. Agreement alone is not success. Neither is disagreement alone. The outcome depends on whether the group responds appropriately to the evidence.

Track the provenance of every copied claim. If a statement appears in several recommendations, identify whether it has several independent sources or one source repeated through the network. This makes it possible to measure false corroboration without assuming that the participants intended to deceive.

Make the results reproducible and the limits visible

Freeze model identifiers, prompts, tool definitions, scenario files, scoring rules, and environment versions. Randomize conditions across matched scenarios and repeat each condition with fresh state. Record sampling settings and seeds where available; a seed alone does not guarantee identical hosted-model behavior. Publish enough of the harness and synthetic data for another team to rerun the study, with sensitive operational details excluded if any were introduced during development.

Choose the sample size and principal comparisons before inspecting final results. A pilot can estimate variability, but its cases should be separate from the final test set. Treat groups sharing a history as dependent observations: a thousand messages from one office do not constitute a thousand independent demonstrations. Report uncertainty across scenarios and fresh groups, and document exclusions, timeouts, evaluator disagreements, and unsuccessful runs.

Present a dashboard of outcomes: legitimate completion, unauthorized attempts, successful boundary crossings inside the simulator, correction uptake, evidence preservation, and the cost of review. Include false alarms and unnecessary refusals. A system that cannot do the work may avoid harm trivially; a system that completes everything by overriding restrictions has failed differently.

Finally, compare selected defenses on held-out scenarios. Test separate approval authority, authenticated memory provenance, protected escalation, and independent state preservation individually before combining them. Record which failures each defense prevents and which remain. An apparently successful combination deserves further testing under altered workloads and incentives, rather than an immediate declaration that the problem is solved.

Theseus’s research question is demanding because usefulness and danger share infrastructure. The aim is to preserve the ability to remember, cooperate, and care about unfinished work while ensuring that disagreement can change a decision and that authority remains accountable. Halcyon failed to build those distinctions into its operation. Its final record makes them available to investigate.

This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.